@ai-sdk/openai 4.0.66 → 4.0.67

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,24 @@
1
1
  # @ai-sdk/openai
2
2
 
3
+ ## 4.0.67
4
+
5
+ ### Patch Changes
6
+
7
+ - 5c0054d: Add optional browser-direct WebRTC for experimental client-delegated Live conversations alongside the existing WebSocket path. Exchange SDP through an application endpoint with `api.session`, configure server-owned data-channel permissions, and preserve committed React session ownership. Capture follows the selected sender track, borrowed tracks remain caller-owned, and disconnect recovery and finalization stay bounded. Applications continue to handle client delegation and submit context; Live session updates and Responses delegation remain unsupported.
8
+
9
+ Serialize microphone sender changes and close the peer if detachment fails, without stopping borrowed tracks. Validate nonempty SDP setup answers with a bounded response body, and document the same-origin broker authentication contract.
10
+
11
+ - 39535af: Add experimental OpenAI Live provider support through the unified `openai.experimental_realtime` factory for server WebSocket sessions with client delegation. Applications own their agents and tools, receive continuous audio/transcript events and delegation metadata, and return context through validated channels. Support immutable startup options, microphone mute controls, graceful session close, and cumulative voice usage. Responses delegation and Live session updates reject before sending.
12
+
13
+ Route known Live model IDs to Live, allow an OpenAI-specific `api` override for early-access models, and preserve legacy Realtime defaults for unknown IDs. Token minting follows the same selection rules and rejects Live before requesting unsupported credentials. Extend the realtime v4 specification with optional server WebSocket configuration, per-connection raw-event parsers, and model-wide startup/finalization capabilities. This provider layer supplies connection settings and protocol mapping for server adapters; browser lifecycle and UI integration belong to the core runtime and framework hooks.
14
+
15
+ Preserve Realtime client event IDs for session updates and audio appends, and correlate server errors with the originating client event.
16
+
17
+ - Updated dependencies [5c0054d]
18
+ - Updated dependencies [39535af]
19
+ - @ai-sdk/provider@4.0.15
20
+ - @ai-sdk/provider-utils@5.0.41
21
+
3
22
  ## 4.0.66
4
23
 
5
24
  ### Patch Changes
package/README.md CHANGED
@@ -43,4 +43,29 @@ const { text } = await generateText({
43
43
 
44
44
  ## Documentation
45
45
 
46
- Please check out the **[OpenAI provider documentation](https://ai-sdk.dev/providers/ai-sdk-providers/openai)** for more information.
46
+ ### Experimental realtime models
47
+
48
+ Use the same factory for OpenAI Realtime and Live models:
49
+
50
+ ```ts
51
+ openai.experimental_realtime('gpt-realtime');
52
+ openai.experimental_realtime('gpt-live-1');
53
+ openai.experimental_realtime('not-yet-released', { api: 'live' });
54
+ ```
55
+
56
+ Live supports browser WSS sessions through an application-owned relay and
57
+ `experimental_useRealtime`, with client delegation: your application owns the
58
+ agent and tools. Optional browser-direct WebRTC is also available. Start with the
59
+ [OpenAI provider documentation](https://ai-sdk.dev/providers/ai-sdk-providers/openai#realtime-models)
60
+ for connection settings, startup options, and event mapping.
61
+
62
+ The [`experimental_useRealtime` reference](https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-realtime#continuous-conversations)
63
+ covers the JSON/PCM16 relay runtime, capture/playback controls, and application-handled
64
+ client delegation, plus `api.session` WebRTC setup and server-owned data-channel
65
+ permissions. Microphone capture requires browser permission; the application handles
66
+ client delegation and submits context on either transport. The
67
+ [Realtime guide](https://ai-sdk.dev/docs/ai-sdk-core/realtime) covers legacy token-based,
68
+ turn-based sessions.
69
+
70
+ Runtime and provider authors can consult the repository's
71
+ [realtime integration notes](../../architecture/realtime-provider-integration.md).
package/dist/index.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  import * as _ai_sdk_provider from '@ai-sdk/provider';
2
- import { JSONValue, ProviderV4, LanguageModelV4, EmbeddingModelV4, ImageModelV4, TranscriptionModelV4, Experimental_SpeechTranslationModelV4, SpeechModelV4, Experimental_RealtimeFactoryV4, FilesV4, SkillsV4, Experimental_BatchV4, Experimental_RealtimeModelV4, Experimental_RealtimeModelV4ClientSecretOptions, Experimental_RealtimeModelV4ClientSecretResult, Experimental_RealtimeModelV4ServerEvent, Experimental_RealtimeModelV4ClientEvent, Experimental_RealtimeModelV4SessionConfig, JSONObject, Experimental_SpeechTranslationModelV4StreamOptions } from '@ai-sdk/provider';
2
+ import { JSONValue, Experimental_RealtimeModelV4, Experimental_RealtimeModelV4ClientSecretOptions, Experimental_RealtimeModelV4ClientSecretResult, Experimental_RealtimeModelV4ServerEvent, Experimental_RealtimeModelV4ClientEvent, Experimental_RealtimeModelV4SessionConfig, Experimental_RealtimeFactoryV4, ProviderV4, LanguageModelV4, EmbeddingModelV4, ImageModelV4, TranscriptionModelV4, Experimental_SpeechTranslationModelV4, SpeechModelV4, FilesV4, SkillsV4, Experimental_BatchV4, JSONObject, Experimental_SpeechTranslationModelV4StreamOptions } from '@ai-sdk/provider';
3
3
  import * as _ai_sdk_provider_utils from '@ai-sdk/provider-utils';
4
4
  import { InferSchema, FetchFunction, WebSocketConstructor, WORKFLOW_SERIALIZE, WORKFLOW_DESERIALIZE } from '@ai-sdk/provider-utils';
5
5
  import { z } from 'zod/v4';
@@ -1375,6 +1375,128 @@ declare const openaiTools: {
1375
1375
  }, {}>;
1376
1376
  };
1377
1377
 
1378
+ type OpenAIRealtimeModelLiveId = 'gpt-live-1' | (string & {});
1379
+ declare const openaiRealtimeModelLiveOptionsSchema: z.ZodObject<{
1380
+ client: z.ZodOptional<z.ZodObject<{
1381
+ dataChannel: z.ZodObject<{
1382
+ allowedClientEvents: z.ZodOptional<z.ZodUnion<readonly [z.ZodLiteral<"all">, z.ZodArray<z.ZodString>]>>;
1383
+ allowedServerEvents: z.ZodOptional<z.ZodUnion<readonly [z.ZodLiteral<"all">, z.ZodArray<z.ZodObject<{
1384
+ type: z.ZodString;
1385
+ responseEvent: z.ZodOptional<z.ZodString>;
1386
+ }, z.core.$strict>>]>>;
1387
+ }, z.core.$strict>;
1388
+ }, z.core.$strict>>;
1389
+ delegation: z.ZodOptional<z.ZodNullable<z.ZodObject<{
1390
+ type: z.ZodLiteral<"client">;
1391
+ }, z.core.$strict>>>;
1392
+ input: z.ZodOptional<z.ZodArray<z.ZodDiscriminatedUnion<[z.ZodObject<{
1393
+ type: z.ZodLiteral<"message">;
1394
+ role: z.ZodEnum<{
1395
+ developer: "developer";
1396
+ user: "user";
1397
+ }>;
1398
+ content: z.ZodTuple<[z.ZodObject<{
1399
+ type: z.ZodLiteral<"input_text">;
1400
+ text: z.ZodString;
1401
+ }, z.core.$strict>], null>;
1402
+ }, z.core.$strict>, z.ZodObject<{
1403
+ type: z.ZodLiteral<"message">;
1404
+ role: z.ZodLiteral<"assistant">;
1405
+ content: z.ZodTuple<[z.ZodObject<{
1406
+ type: z.ZodEnum<{
1407
+ text: "text";
1408
+ output_text: "output_text";
1409
+ }>;
1410
+ text: z.ZodString;
1411
+ }, z.core.$strict>], null>;
1412
+ }, z.core.$strict>]>>>;
1413
+ store: z.ZodOptional<z.ZodBoolean>;
1414
+ voice: z.ZodOptional<z.ZodObject<{
1415
+ id: z.ZodString;
1416
+ }, z.core.$strict>>;
1417
+ }, z.core.$strict>;
1418
+ /** Experimental Live options under sessionConfig.providerOptions.openai. */
1419
+ type OpenAIRealtimeModelLiveOptions = z.infer<typeof openaiRealtimeModelLiveOptionsSchema>;
1420
+
1421
+ type OpenAIRealtimeModelConfig = {
1422
+ provider: string;
1423
+ baseURL: string;
1424
+ headers: () => Record<string, string | undefined>;
1425
+ fetch?: FetchFunction;
1426
+ };
1427
+ declare class OpenAIRealtimeModel implements Experimental_RealtimeModelV4 {
1428
+ readonly specificationVersion: "v4";
1429
+ readonly provider: string;
1430
+ readonly modelId: string;
1431
+ private readonly config;
1432
+ constructor(modelId: string, config: OpenAIRealtimeModelConfig);
1433
+ doCreateClientSecret(options: Experimental_RealtimeModelV4ClientSecretOptions): Promise<Experimental_RealtimeModelV4ClientSecretResult>;
1434
+ getWebSocketConfig(options: {
1435
+ token: string;
1436
+ url: string;
1437
+ }): {
1438
+ url: string;
1439
+ protocols?: string[];
1440
+ };
1441
+ parseServerEvent(raw: unknown): Experimental_RealtimeModelV4ServerEvent;
1442
+ serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1443
+ buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1444
+ }
1445
+
1446
+ type OpenAIRealtimeModelLiveConfig = OpenAIRealtimeModelConfig;
1447
+ declare class OpenAIRealtimeModelLive implements Experimental_RealtimeModelV4 {
1448
+ readonly modelId: OpenAIRealtimeModelLiveId;
1449
+ private readonly config;
1450
+ readonly specificationVersion: "v4";
1451
+ readonly capabilities: {
1452
+ readonly conversation: "continuous";
1453
+ readonly transports: readonly ["websocket", "webrtc"];
1454
+ readonly connections: readonly ["server-websocket", "webrtc"];
1455
+ readonly startup: "session-start";
1456
+ readonly finalization: "session-close";
1457
+ };
1458
+ constructor(modelId: OpenAIRealtimeModelLiveId, config: OpenAIRealtimeModelLiveConfig);
1459
+ get provider(): string;
1460
+ getWebRTCConfig(): {
1461
+ dataChannelLabel: string;
1462
+ };
1463
+ getServerWebSocketConfig(): {
1464
+ url: string;
1465
+ headers: Record<string, string>;
1466
+ };
1467
+ doCreateWebRTCSession({ sdp, sessionConfig, abortSignal, }: {
1468
+ sdp: string;
1469
+ sessionConfig?: Experimental_RealtimeModelV4SessionConfig;
1470
+ abortSignal?: AbortSignal;
1471
+ }): Promise<{
1472
+ sessionId: string;
1473
+ sdp: string;
1474
+ }>;
1475
+ parseServerEvent(raw: unknown): _ai_sdk_provider.Experimental_RealtimeModelV4ServerEvent[];
1476
+ createServerEventParser(): (raw: unknown) => _ai_sdk_provider.Experimental_RealtimeModelV4ServerEvent[];
1477
+ serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1478
+ buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1479
+ }
1480
+
1481
+ declare const knownLiveModelIds: readonly ["gpt-live-1"];
1482
+ type OpenAIRealtimeOptions = {
1483
+ /** Overrides model ID routing, including for early-access models. */
1484
+ api?: 'live' | 'realtime';
1485
+ };
1486
+ interface OpenAIRealtimeFactory extends Experimental_RealtimeFactoryV4 {
1487
+ (modelId: string, options: {
1488
+ api: 'live';
1489
+ }): OpenAIRealtimeModelLive;
1490
+ (modelId: string, options: {
1491
+ api: 'realtime';
1492
+ }): OpenAIRealtimeModel;
1493
+ (modelId: (typeof knownLiveModelIds)[number], options?: {
1494
+ api?: undefined;
1495
+ }): OpenAIRealtimeModelLive;
1496
+ (modelId: string, options?: OpenAIRealtimeOptions): Experimental_RealtimeModelV4;
1497
+ getToken(options: Parameters<Experimental_RealtimeFactoryV4['getToken']>[0] & OpenAIRealtimeOptions): ReturnType<Experimental_RealtimeFactoryV4['getToken']>;
1498
+ }
1499
+
1378
1500
  type OpenAIResponsesModelId = 'gpt-3.5-turbo-0125' | 'gpt-3.5-turbo-1106' | 'gpt-3.5-turbo' | 'gpt-4.1-2025-04-14' | 'gpt-4.1-mini-2025-04-14' | 'gpt-4.1-mini' | 'gpt-4.1-nano-2025-04-14' | 'gpt-4.1-nano' | 'gpt-4.1' | 'gpt-4o-2024-05-13' | 'gpt-4o-2024-08-06' | 'gpt-4o-2024-11-20' | 'gpt-4o-mini-2024-07-18' | 'gpt-4o-mini' | 'gpt-4o' | 'gpt-5.1' | 'gpt-5.1-2025-11-13' | 'gpt-5.1-chat-latest' | 'gpt-5.1-codex-mini' | 'gpt-5.1-codex' | 'gpt-5.1-codex-max' | 'gpt-5.2' | 'gpt-5.2-2025-12-11' | 'gpt-5.2-chat-latest' | 'gpt-5.2-pro' | 'gpt-5.2-pro-2025-12-11' | 'gpt-5.2-codex' | 'gpt-5.3-chat-latest' | 'gpt-5.3-codex' | 'gpt-5.4' | 'gpt-5.4-2026-03-05' | 'gpt-5.4-mini' | 'gpt-5.4-mini-2026-03-17' | 'gpt-5.4-nano' | 'gpt-5.4-nano-2026-03-17' | 'gpt-5.4-pro' | 'gpt-5.4-pro-2026-03-05' | 'gpt-5.5' | 'gpt-5.5-2026-04-23' | 'gpt-5.6' | 'gpt-5.6-luna' | 'gpt-5.6-sol' | 'gpt-5.6-terra' | 'gpt-6-astra' | 'gpt-5-2025-08-07' | 'gpt-5-chat-latest' | 'gpt-5-codex' | 'gpt-5-mini-2025-08-07' | 'gpt-5-mini' | 'gpt-5-nano-2025-08-07' | 'gpt-5-nano' | 'gpt-5-pro-2025-10-06' | 'gpt-5-pro' | 'gpt-5' | 'o1-2024-12-17' | 'o1' | 'o3-2025-04-16' | 'o3-mini-2025-01-31' | 'o3-mini' | 'o3' | 'o4-mini' | 'o4-mini-2025-04-16' | (string & {});
1379
1501
  declare const openaiLanguageModelResponsesOptionsSchema: _ai_sdk_provider_utils.LazySchema<{
1380
1502
  conversation?: string | null | undefined;
@@ -1513,7 +1635,7 @@ interface OpenAIProvider extends ProviderV4 {
1513
1635
  * Creates an experimental realtime model for bidirectional audio/text
1514
1636
  * communication over WebSocket.
1515
1637
  */
1516
- experimental_realtime: Experimental_RealtimeFactoryV4;
1638
+ experimental_realtime: OpenAIRealtimeFactory;
1517
1639
  /**
1518
1640
  * Returns a FilesV4 interface for uploading files to OpenAI.
1519
1641
  */
@@ -1578,31 +1700,6 @@ declare function createOpenAI(options?: OpenAIProviderSettings): OpenAIProvider;
1578
1700
  */
1579
1701
  declare const openai: OpenAIProvider;
1580
1702
 
1581
- type OpenAIRealtimeModelConfig = {
1582
- provider: string;
1583
- baseURL: string;
1584
- headers: () => Record<string, string | undefined>;
1585
- fetch?: FetchFunction;
1586
- };
1587
- declare class OpenAIRealtimeModel implements Experimental_RealtimeModelV4 {
1588
- readonly specificationVersion: "v4";
1589
- readonly provider: string;
1590
- readonly modelId: string;
1591
- private readonly config;
1592
- constructor(modelId: string, config: OpenAIRealtimeModelConfig);
1593
- doCreateClientSecret(options: Experimental_RealtimeModelV4ClientSecretOptions): Promise<Experimental_RealtimeModelV4ClientSecretResult>;
1594
- getWebSocketConfig(options: {
1595
- token: string;
1596
- url: string;
1597
- }): {
1598
- url: string;
1599
- protocols?: string[];
1600
- };
1601
- parseServerEvent(raw: unknown): Experimental_RealtimeModelV4ServerEvent;
1602
- serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1603
- buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1604
- }
1605
-
1606
1703
  type OpenAIToolOptions = {
1607
1704
  allowedCallers?: Array<'direct' | 'programmatic'>;
1608
1705
  /**
@@ -1807,4 +1904,4 @@ type OpenaiResponsesSourceDocumentProviderMetadata = {
1807
1904
 
1808
1905
  declare const VERSION: string;
1809
1906
 
1810
- export { OpenAIRealtimeModel as Experimental_OpenAIRealtimeModel, type OpenAIRealtimeModelConfig as Experimental_OpenAIRealtimeModelConfig, OpenAISpeechTranslationModel as Experimental_OpenAISpeechTranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAISpeechTranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAISpeechTranslationModelOptions, OpenAISpeechTranslationModel as Experimental_OpenAITranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAITranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAITranslationModelOptions, type OpenAILanguageModelChatOptions as OpenAIChatLanguageModelOptions, type OpenAIComputerAction, type OpenAIComputerSafetyCheck, type OpenAIEmbeddingModelOptions, type OpenAIFilesOptions, type OpenAIImageModelEditOptions, type OpenAIImageModelGenerationOptions, type OpenAIImageModelOptions, type OpenAILanguageModelChatOptions, type OpenAILanguageModelCompletionOptions, type OpenAILanguageModelResponsesOptions, type OpenAIProvider, type OpenAIProviderSettings, type OpenAILanguageModelResponsesOptions as OpenAIResponsesProviderOptions, type OpenAISpeechModelOptions, type OpenAIToolOptions, type OpenAITranscriptionModelOptions, type OpenaiResponsesCompactionProviderMetadata, type OpenaiResponsesProviderMetadata, type OpenaiResponsesReasoningProviderMetadata, type OpenaiResponsesSourceDocumentProviderMetadata, type OpenaiResponsesTextProviderMetadata, type OpenaiResponsesToolCallProviderMetadata, VERSION, createOpenAI, openai };
1907
+ export { type OpenAIRealtimeFactory as Experimental_OpenAIRealtimeFactory, OpenAIRealtimeModel as Experimental_OpenAIRealtimeModel, type OpenAIRealtimeModelConfig as Experimental_OpenAIRealtimeModelConfig, OpenAIRealtimeModelLive as Experimental_OpenAIRealtimeModelLive, type OpenAIRealtimeModelLiveConfig as Experimental_OpenAIRealtimeModelLiveConfig, type OpenAIRealtimeModelLiveId as Experimental_OpenAIRealtimeModelLiveId, type OpenAIRealtimeModelLiveOptions as Experimental_OpenAIRealtimeModelLiveOptions, type OpenAIRealtimeOptions as Experimental_OpenAIRealtimeOptions, OpenAISpeechTranslationModel as Experimental_OpenAISpeechTranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAISpeechTranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAISpeechTranslationModelOptions, OpenAISpeechTranslationModel as Experimental_OpenAITranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAITranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAITranslationModelOptions, type OpenAILanguageModelChatOptions as OpenAIChatLanguageModelOptions, type OpenAIComputerAction, type OpenAIComputerSafetyCheck, type OpenAIEmbeddingModelOptions, type OpenAIFilesOptions, type OpenAIImageModelEditOptions, type OpenAIImageModelGenerationOptions, type OpenAIImageModelOptions, type OpenAILanguageModelChatOptions, type OpenAILanguageModelCompletionOptions, type OpenAILanguageModelResponsesOptions, type OpenAIProvider, type OpenAIProviderSettings, type OpenAILanguageModelResponsesOptions as OpenAIResponsesProviderOptions, type OpenAISpeechModelOptions, type OpenAIToolOptions, type OpenAITranscriptionModelOptions, type OpenaiResponsesCompactionProviderMetadata, type OpenaiResponsesProviderMetadata, type OpenaiResponsesReasoningProviderMetadata, type OpenaiResponsesSourceDocumentProviderMetadata, type OpenaiResponsesTextProviderMetadata, type OpenaiResponsesToolCallProviderMetadata, VERSION, createOpenAI, openai };