@ai-sdk/openai 4.0.65 → 4.0.67

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,39 @@
1
1
  # @ai-sdk/openai
2
2
 
3
+ ## 4.0.67
4
+
5
+ ### Patch Changes
6
+
7
+ - 5c0054d: Add optional browser-direct WebRTC for experimental client-delegated Live conversations alongside the existing WebSocket path. Exchange SDP through an application endpoint with `api.session`, configure server-owned data-channel permissions, and preserve committed React session ownership. Capture follows the selected sender track, borrowed tracks remain caller-owned, and disconnect recovery and finalization stay bounded. Applications continue to handle client delegation and submit context; Live session updates and Responses delegation remain unsupported.
8
+
9
+ Serialize microphone sender changes and close the peer if detachment fails, without stopping borrowed tracks. Validate nonempty SDP setup answers with a bounded response body, and document the same-origin broker authentication contract.
10
+
11
+ - 39535af: Add experimental OpenAI Live provider support through the unified `openai.experimental_realtime` factory for server WebSocket sessions with client delegation. Applications own their agents and tools, receive continuous audio/transcript events and delegation metadata, and return context through validated channels. Support immutable startup options, microphone mute controls, graceful session close, and cumulative voice usage. Responses delegation and Live session updates reject before sending.
12
+
13
+ Route known Live model IDs to Live, allow an OpenAI-specific `api` override for early-access models, and preserve legacy Realtime defaults for unknown IDs. Token minting follows the same selection rules and rejects Live before requesting unsupported credentials. Extend the realtime v4 specification with optional server WebSocket configuration, per-connection raw-event parsers, and model-wide startup/finalization capabilities. This provider layer supplies connection settings and protocol mapping for server adapters; browser lifecycle and UI integration belong to the core runtime and framework hooks.
14
+
15
+ Preserve Realtime client event IDs for session updates and audio appends, and correlate server errors with the originating client event.
16
+
17
+ - Updated dependencies [5c0054d]
18
+ - Updated dependencies [39535af]
19
+ - @ai-sdk/provider@4.0.15
20
+ - @ai-sdk/provider-utils@5.0.41
21
+
22
+ ## 4.0.66
23
+
24
+ ### Patch Changes
25
+
26
+ - 5ec21a6: fix: reject unsupported batch request types
27
+ - 7469a3b: feat: support image generation requests in batches
28
+ - 0de8886: fix(openai): allow providers to disable web search source includes
29
+ - d5e3024: fix(openai): remove propertyNames from JSON Schema
30
+ - Updated dependencies [5ec21a6]
31
+ - Updated dependencies [7469a3b]
32
+ - Updated dependencies [813bb36]
33
+ - Updated dependencies [c43e4b7]
34
+ - @ai-sdk/provider@4.0.14
35
+ - @ai-sdk/provider-utils@5.0.40
36
+
3
37
  ## 4.0.65
4
38
 
5
39
  ### Patch Changes
package/README.md CHANGED
@@ -43,4 +43,29 @@ const { text } = await generateText({
43
43
 
44
44
  ## Documentation
45
45
 
46
- Please check out the **[OpenAI provider documentation](https://ai-sdk.dev/providers/ai-sdk-providers/openai)** for more information.
46
+ ### Experimental realtime models
47
+
48
+ Use the same factory for OpenAI Realtime and Live models:
49
+
50
+ ```ts
51
+ openai.experimental_realtime('gpt-realtime');
52
+ openai.experimental_realtime('gpt-live-1');
53
+ openai.experimental_realtime('not-yet-released', { api: 'live' });
54
+ ```
55
+
56
+ Live supports browser WSS sessions through an application-owned relay and
57
+ `experimental_useRealtime`, with client delegation: your application owns the
58
+ agent and tools. Optional browser-direct WebRTC is also available. Start with the
59
+ [OpenAI provider documentation](https://ai-sdk.dev/providers/ai-sdk-providers/openai#realtime-models)
60
+ for connection settings, startup options, and event mapping.
61
+
62
+ The [`experimental_useRealtime` reference](https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-realtime#continuous-conversations)
63
+ covers the JSON/PCM16 relay runtime, capture/playback controls, and application-handled
64
+ client delegation, plus `api.session` WebRTC setup and server-owned data-channel
65
+ permissions. Microphone capture requires browser permission; the application handles
66
+ client delegation and submits context on either transport. The
67
+ [Realtime guide](https://ai-sdk.dev/docs/ai-sdk-core/realtime) covers legacy token-based,
68
+ turn-based sessions.
69
+
70
+ Runtime and provider authors can consult the repository's
71
+ [realtime integration notes](../../architecture/realtime-provider-integration.md).
package/dist/index.d.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  import * as _ai_sdk_provider from '@ai-sdk/provider';
2
- import { JSONValue, ProviderV4, LanguageModelV4, EmbeddingModelV4, ImageModelV4, TranscriptionModelV4, Experimental_SpeechTranslationModelV4, SpeechModelV4, Experimental_RealtimeFactoryV4, FilesV4, SkillsV4, Experimental_BatchV4, Experimental_RealtimeModelV4, Experimental_RealtimeModelV4ClientSecretOptions, Experimental_RealtimeModelV4ClientSecretResult, Experimental_RealtimeModelV4ServerEvent, Experimental_RealtimeModelV4ClientEvent, Experimental_RealtimeModelV4SessionConfig, JSONObject, Experimental_SpeechTranslationModelV4StreamOptions } from '@ai-sdk/provider';
2
+ import { JSONValue, Experimental_RealtimeModelV4, Experimental_RealtimeModelV4ClientSecretOptions, Experimental_RealtimeModelV4ClientSecretResult, Experimental_RealtimeModelV4ServerEvent, Experimental_RealtimeModelV4ClientEvent, Experimental_RealtimeModelV4SessionConfig, Experimental_RealtimeFactoryV4, ProviderV4, LanguageModelV4, EmbeddingModelV4, ImageModelV4, TranscriptionModelV4, Experimental_SpeechTranslationModelV4, SpeechModelV4, FilesV4, SkillsV4, Experimental_BatchV4, JSONObject, Experimental_SpeechTranslationModelV4StreamOptions } from '@ai-sdk/provider';
3
3
  import * as _ai_sdk_provider_utils from '@ai-sdk/provider-utils';
4
4
  import { InferSchema, FetchFunction, WebSocketConstructor, WORKFLOW_SERIALIZE, WORKFLOW_DESERIALIZE } from '@ai-sdk/provider-utils';
5
5
  import { z } from 'zod/v4';
@@ -1375,10 +1375,133 @@ declare const openaiTools: {
1375
1375
  }, {}>;
1376
1376
  };
1377
1377
 
1378
+ type OpenAIRealtimeModelLiveId = 'gpt-live-1' | (string & {});
1379
+ declare const openaiRealtimeModelLiveOptionsSchema: z.ZodObject<{
1380
+ client: z.ZodOptional<z.ZodObject<{
1381
+ dataChannel: z.ZodObject<{
1382
+ allowedClientEvents: z.ZodOptional<z.ZodUnion<readonly [z.ZodLiteral<"all">, z.ZodArray<z.ZodString>]>>;
1383
+ allowedServerEvents: z.ZodOptional<z.ZodUnion<readonly [z.ZodLiteral<"all">, z.ZodArray<z.ZodObject<{
1384
+ type: z.ZodString;
1385
+ responseEvent: z.ZodOptional<z.ZodString>;
1386
+ }, z.core.$strict>>]>>;
1387
+ }, z.core.$strict>;
1388
+ }, z.core.$strict>>;
1389
+ delegation: z.ZodOptional<z.ZodNullable<z.ZodObject<{
1390
+ type: z.ZodLiteral<"client">;
1391
+ }, z.core.$strict>>>;
1392
+ input: z.ZodOptional<z.ZodArray<z.ZodDiscriminatedUnion<[z.ZodObject<{
1393
+ type: z.ZodLiteral<"message">;
1394
+ role: z.ZodEnum<{
1395
+ developer: "developer";
1396
+ user: "user";
1397
+ }>;
1398
+ content: z.ZodTuple<[z.ZodObject<{
1399
+ type: z.ZodLiteral<"input_text">;
1400
+ text: z.ZodString;
1401
+ }, z.core.$strict>], null>;
1402
+ }, z.core.$strict>, z.ZodObject<{
1403
+ type: z.ZodLiteral<"message">;
1404
+ role: z.ZodLiteral<"assistant">;
1405
+ content: z.ZodTuple<[z.ZodObject<{
1406
+ type: z.ZodEnum<{
1407
+ text: "text";
1408
+ output_text: "output_text";
1409
+ }>;
1410
+ text: z.ZodString;
1411
+ }, z.core.$strict>], null>;
1412
+ }, z.core.$strict>]>>>;
1413
+ store: z.ZodOptional<z.ZodBoolean>;
1414
+ voice: z.ZodOptional<z.ZodObject<{
1415
+ id: z.ZodString;
1416
+ }, z.core.$strict>>;
1417
+ }, z.core.$strict>;
1418
+ /** Experimental Live options under sessionConfig.providerOptions.openai. */
1419
+ type OpenAIRealtimeModelLiveOptions = z.infer<typeof openaiRealtimeModelLiveOptionsSchema>;
1420
+
1421
+ type OpenAIRealtimeModelConfig = {
1422
+ provider: string;
1423
+ baseURL: string;
1424
+ headers: () => Record<string, string | undefined>;
1425
+ fetch?: FetchFunction;
1426
+ };
1427
+ declare class OpenAIRealtimeModel implements Experimental_RealtimeModelV4 {
1428
+ readonly specificationVersion: "v4";
1429
+ readonly provider: string;
1430
+ readonly modelId: string;
1431
+ private readonly config;
1432
+ constructor(modelId: string, config: OpenAIRealtimeModelConfig);
1433
+ doCreateClientSecret(options: Experimental_RealtimeModelV4ClientSecretOptions): Promise<Experimental_RealtimeModelV4ClientSecretResult>;
1434
+ getWebSocketConfig(options: {
1435
+ token: string;
1436
+ url: string;
1437
+ }): {
1438
+ url: string;
1439
+ protocols?: string[];
1440
+ };
1441
+ parseServerEvent(raw: unknown): Experimental_RealtimeModelV4ServerEvent;
1442
+ serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1443
+ buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1444
+ }
1445
+
1446
+ type OpenAIRealtimeModelLiveConfig = OpenAIRealtimeModelConfig;
1447
+ declare class OpenAIRealtimeModelLive implements Experimental_RealtimeModelV4 {
1448
+ readonly modelId: OpenAIRealtimeModelLiveId;
1449
+ private readonly config;
1450
+ readonly specificationVersion: "v4";
1451
+ readonly capabilities: {
1452
+ readonly conversation: "continuous";
1453
+ readonly transports: readonly ["websocket", "webrtc"];
1454
+ readonly connections: readonly ["server-websocket", "webrtc"];
1455
+ readonly startup: "session-start";
1456
+ readonly finalization: "session-close";
1457
+ };
1458
+ constructor(modelId: OpenAIRealtimeModelLiveId, config: OpenAIRealtimeModelLiveConfig);
1459
+ get provider(): string;
1460
+ getWebRTCConfig(): {
1461
+ dataChannelLabel: string;
1462
+ };
1463
+ getServerWebSocketConfig(): {
1464
+ url: string;
1465
+ headers: Record<string, string>;
1466
+ };
1467
+ doCreateWebRTCSession({ sdp, sessionConfig, abortSignal, }: {
1468
+ sdp: string;
1469
+ sessionConfig?: Experimental_RealtimeModelV4SessionConfig;
1470
+ abortSignal?: AbortSignal;
1471
+ }): Promise<{
1472
+ sessionId: string;
1473
+ sdp: string;
1474
+ }>;
1475
+ parseServerEvent(raw: unknown): _ai_sdk_provider.Experimental_RealtimeModelV4ServerEvent[];
1476
+ createServerEventParser(): (raw: unknown) => _ai_sdk_provider.Experimental_RealtimeModelV4ServerEvent[];
1477
+ serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1478
+ buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1479
+ }
1480
+
1481
+ declare const knownLiveModelIds: readonly ["gpt-live-1"];
1482
+ type OpenAIRealtimeOptions = {
1483
+ /** Overrides model ID routing, including for early-access models. */
1484
+ api?: 'live' | 'realtime';
1485
+ };
1486
+ interface OpenAIRealtimeFactory extends Experimental_RealtimeFactoryV4 {
1487
+ (modelId: string, options: {
1488
+ api: 'live';
1489
+ }): OpenAIRealtimeModelLive;
1490
+ (modelId: string, options: {
1491
+ api: 'realtime';
1492
+ }): OpenAIRealtimeModel;
1493
+ (modelId: (typeof knownLiveModelIds)[number], options?: {
1494
+ api?: undefined;
1495
+ }): OpenAIRealtimeModelLive;
1496
+ (modelId: string, options?: OpenAIRealtimeOptions): Experimental_RealtimeModelV4;
1497
+ getToken(options: Parameters<Experimental_RealtimeFactoryV4['getToken']>[0] & OpenAIRealtimeOptions): ReturnType<Experimental_RealtimeFactoryV4['getToken']>;
1498
+ }
1499
+
1378
1500
  type OpenAIResponsesModelId = 'gpt-3.5-turbo-0125' | 'gpt-3.5-turbo-1106' | 'gpt-3.5-turbo' | 'gpt-4.1-2025-04-14' | 'gpt-4.1-mini-2025-04-14' | 'gpt-4.1-mini' | 'gpt-4.1-nano-2025-04-14' | 'gpt-4.1-nano' | 'gpt-4.1' | 'gpt-4o-2024-05-13' | 'gpt-4o-2024-08-06' | 'gpt-4o-2024-11-20' | 'gpt-4o-mini-2024-07-18' | 'gpt-4o-mini' | 'gpt-4o' | 'gpt-5.1' | 'gpt-5.1-2025-11-13' | 'gpt-5.1-chat-latest' | 'gpt-5.1-codex-mini' | 'gpt-5.1-codex' | 'gpt-5.1-codex-max' | 'gpt-5.2' | 'gpt-5.2-2025-12-11' | 'gpt-5.2-chat-latest' | 'gpt-5.2-pro' | 'gpt-5.2-pro-2025-12-11' | 'gpt-5.2-codex' | 'gpt-5.3-chat-latest' | 'gpt-5.3-codex' | 'gpt-5.4' | 'gpt-5.4-2026-03-05' | 'gpt-5.4-mini' | 'gpt-5.4-mini-2026-03-17' | 'gpt-5.4-nano' | 'gpt-5.4-nano-2026-03-17' | 'gpt-5.4-pro' | 'gpt-5.4-pro-2026-03-05' | 'gpt-5.5' | 'gpt-5.5-2026-04-23' | 'gpt-5.6' | 'gpt-5.6-luna' | 'gpt-5.6-sol' | 'gpt-5.6-terra' | 'gpt-6-astra' | 'gpt-5-2025-08-07' | 'gpt-5-chat-latest' | 'gpt-5-codex' | 'gpt-5-mini-2025-08-07' | 'gpt-5-mini' | 'gpt-5-nano-2025-08-07' | 'gpt-5-nano' | 'gpt-5-pro-2025-10-06' | 'gpt-5-pro' | 'gpt-5' | 'o1-2024-12-17' | 'o1' | 'o3-2025-04-16' | 'o3-mini-2025-01-31' | 'o3-mini' | 'o3' | 'o4-mini' | 'o4-mini-2025-04-16' | (string & {});
1379
1501
  declare const openaiLanguageModelResponsesOptionsSchema: _ai_sdk_provider_utils.LazySchema<{
1380
1502
  conversation?: string | null | undefined;
1381
1503
  include?: ("web_search_call.results" | "file_search_call.results" | "message.output_text.logprobs" | "reasoning.encrypted_content")[] | null | undefined;
1504
+ includeWebSearchSources?: boolean | undefined;
1382
1505
  instructions?: string | null | undefined;
1383
1506
  logprobs?: number | boolean | undefined;
1384
1507
  maxToolCalls?: number | null | undefined;
@@ -1512,7 +1635,7 @@ interface OpenAIProvider extends ProviderV4 {
1512
1635
  * Creates an experimental realtime model for bidirectional audio/text
1513
1636
  * communication over WebSocket.
1514
1637
  */
1515
- experimental_realtime: Experimental_RealtimeFactoryV4;
1638
+ experimental_realtime: OpenAIRealtimeFactory;
1516
1639
  /**
1517
1640
  * Returns a FilesV4 interface for uploading files to OpenAI.
1518
1641
  */
@@ -1577,31 +1700,6 @@ declare function createOpenAI(options?: OpenAIProviderSettings): OpenAIProvider;
1577
1700
  */
1578
1701
  declare const openai: OpenAIProvider;
1579
1702
 
1580
- type OpenAIRealtimeModelConfig = {
1581
- provider: string;
1582
- baseURL: string;
1583
- headers: () => Record<string, string | undefined>;
1584
- fetch?: FetchFunction;
1585
- };
1586
- declare class OpenAIRealtimeModel implements Experimental_RealtimeModelV4 {
1587
- readonly specificationVersion: "v4";
1588
- readonly provider: string;
1589
- readonly modelId: string;
1590
- private readonly config;
1591
- constructor(modelId: string, config: OpenAIRealtimeModelConfig);
1592
- doCreateClientSecret(options: Experimental_RealtimeModelV4ClientSecretOptions): Promise<Experimental_RealtimeModelV4ClientSecretResult>;
1593
- getWebSocketConfig(options: {
1594
- token: string;
1595
- url: string;
1596
- }): {
1597
- url: string;
1598
- protocols?: string[];
1599
- };
1600
- parseServerEvent(raw: unknown): Experimental_RealtimeModelV4ServerEvent;
1601
- serializeClientEvent(event: Experimental_RealtimeModelV4ClientEvent): unknown;
1602
- buildSessionConfig(config: Experimental_RealtimeModelV4SessionConfig): Record<string, unknown>;
1603
- }
1604
-
1605
1703
  type OpenAIToolOptions = {
1606
1704
  allowedCallers?: Array<'direct' | 'programmatic'>;
1607
1705
  /**
@@ -1636,6 +1734,13 @@ type OpenAIConfig = {
1636
1734
  * @see https://github.com/vercel/ai/issues/20180
1637
1735
  */
1638
1736
  explicitMessageItemType?: boolean;
1737
+ /**
1738
+ * Whether the provider supports the
1739
+ * `web_search_call.action.sources` Responses API include value.
1740
+ *
1741
+ * Defaults to `true`.
1742
+ */
1743
+ supportsWebSearchSourcesInclude?: boolean;
1639
1744
  /**
1640
1745
  * This is soft-deprecated. Use provider references (e.g. `{ openai: 'file-abc123' }`)
1641
1746
  * in file part data instead. File ID prefixes used to identify file IDs
@@ -1799,4 +1904,4 @@ type OpenaiResponsesSourceDocumentProviderMetadata = {
1799
1904
 
1800
1905
  declare const VERSION: string;
1801
1906
 
1802
- export { OpenAIRealtimeModel as Experimental_OpenAIRealtimeModel, type OpenAIRealtimeModelConfig as Experimental_OpenAIRealtimeModelConfig, OpenAISpeechTranslationModel as Experimental_OpenAISpeechTranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAISpeechTranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAISpeechTranslationModelOptions, OpenAISpeechTranslationModel as Experimental_OpenAITranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAITranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAITranslationModelOptions, type OpenAILanguageModelChatOptions as OpenAIChatLanguageModelOptions, type OpenAIComputerAction, type OpenAIComputerSafetyCheck, type OpenAIEmbeddingModelOptions, type OpenAIFilesOptions, type OpenAIImageModelEditOptions, type OpenAIImageModelGenerationOptions, type OpenAIImageModelOptions, type OpenAILanguageModelChatOptions, type OpenAILanguageModelCompletionOptions, type OpenAILanguageModelResponsesOptions, type OpenAIProvider, type OpenAIProviderSettings, type OpenAILanguageModelResponsesOptions as OpenAIResponsesProviderOptions, type OpenAISpeechModelOptions, type OpenAIToolOptions, type OpenAITranscriptionModelOptions, type OpenaiResponsesCompactionProviderMetadata, type OpenaiResponsesProviderMetadata, type OpenaiResponsesReasoningProviderMetadata, type OpenaiResponsesSourceDocumentProviderMetadata, type OpenaiResponsesTextProviderMetadata, type OpenaiResponsesToolCallProviderMetadata, VERSION, createOpenAI, openai };
1907
+ export { type OpenAIRealtimeFactory as Experimental_OpenAIRealtimeFactory, OpenAIRealtimeModel as Experimental_OpenAIRealtimeModel, type OpenAIRealtimeModelConfig as Experimental_OpenAIRealtimeModelConfig, OpenAIRealtimeModelLive as Experimental_OpenAIRealtimeModelLive, type OpenAIRealtimeModelLiveConfig as Experimental_OpenAIRealtimeModelLiveConfig, type OpenAIRealtimeModelLiveId as Experimental_OpenAIRealtimeModelLiveId, type OpenAIRealtimeModelLiveOptions as Experimental_OpenAIRealtimeModelLiveOptions, type OpenAIRealtimeOptions as Experimental_OpenAIRealtimeOptions, OpenAISpeechTranslationModel as Experimental_OpenAISpeechTranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAISpeechTranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAISpeechTranslationModelOptions, OpenAISpeechTranslationModel as Experimental_OpenAITranslationModel, type OpenAISpeechTranslationModelId as Experimental_OpenAITranslationModelId, type OpenAISpeechTranslationModelOptions as Experimental_OpenAITranslationModelOptions, type OpenAILanguageModelChatOptions as OpenAIChatLanguageModelOptions, type OpenAIComputerAction, type OpenAIComputerSafetyCheck, type OpenAIEmbeddingModelOptions, type OpenAIFilesOptions, type OpenAIImageModelEditOptions, type OpenAIImageModelGenerationOptions, type OpenAIImageModelOptions, type OpenAILanguageModelChatOptions, type OpenAILanguageModelCompletionOptions, type OpenAILanguageModelResponsesOptions, type OpenAIProvider, type OpenAIProviderSettings, type OpenAILanguageModelResponsesOptions as OpenAIResponsesProviderOptions, type OpenAISpeechModelOptions, type OpenAIToolOptions, type OpenAITranscriptionModelOptions, type OpenaiResponsesCompactionProviderMetadata, type OpenaiResponsesProviderMetadata, type OpenaiResponsesReasoningProviderMetadata, type OpenaiResponsesSourceDocumentProviderMetadata, type OpenaiResponsesTextProviderMetadata, type OpenaiResponsesToolCallProviderMetadata, VERSION, createOpenAI, openai };