@tanstack/ai 0.21.2 → 0.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -181,15 +181,27 @@ The terminal event is a `CUSTOM` chunk: `{ type: 'CUSTOM', name: 'structured-out
181
181
 
182
182
  **Adapter coverage for streaming:**
183
183
 
184
- | Adapter | `outputSchema` + `stream: true` |
185
- | ------------------------------------------------- | --------------------------------------------------------------------------------------------- |
186
- | `@tanstack/ai-openai` | Native single-request stream (Responses API) |
187
- | `@tanstack/ai-openrouter` | Native single-request stream |
188
- | `@tanstack/ai-grok` | Native single-request stream (Chat Completions) |
189
- | `@tanstack/ai-groq` | Native single-request stream (Chat Completions) |
190
- | All other adapters (anthropic, gemini, ollama, …) | Fallback: runs non-streaming `structuredOutput`, emits one `structured-output.complete` event |
191
-
192
- Consumer code is identical across providers — always read the final object off `structured-output.complete`. You only see incremental `TEXT_MESSAGE_CONTENT` deltas when the adapter implements `structuredOutputStream` natively.
184
+ | Adapter | `outputSchema` + `stream: true` |
185
+ | --------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
186
+ | `@tanstack/ai-openai` (Responses + Chat Completions) | **Native combined mode (#605)** — schema wired into the regular `chatStream` call alongside `tools`; engine harvests JSON, no finalization round-trip |
187
+ | `@tanstack/ai-anthropic` (Claude 4.5+ only) | **Native combined mode (#605)** — `output_config.format` + `tools` in one beta Messages call. Older Claude models fall back |
188
+ | `@tanstack/ai-gemini` (Gemini 3.x only) | **Native combined mode (#605)** — `responseSchema` + `tools` in one `generateContentStream`. Gemini 2.x falls back |
189
+ | `@tanstack/ai-grok` (Grok 4 family only) | **Native combined mode (#605)** — `response_format: json_schema` + `tools`. Grok 2 / 3 fall back |
190
+ | `@tanstack/ai-openrouter` | Native single-request stream (legacy `structuredOutputStream` path; per-call combined-mode lookup is a follow-up) |
191
+ | `@tanstack/ai-groq` | Legacy `structuredOutputStream` only (no tools — Groq's API rejects schema + tools + stream) |
192
+ | All other adapters (ollama, older Claude, Gemini 2.x, Grok 2/3) | Fallback: runs non-streaming `structuredOutput`, emits one `structured-output.complete` event |
193
+
194
+ **Native combined mode vs fallback** is signaled by the adapter's
195
+ optional `supportsCombinedToolsAndSchema(modelOptions)` method. When
196
+ it returns `true`, the engine wires the JSON Schema into the regular
197
+ `chatStream` call and harvests the final-turn text — middleware sees
198
+ the run through `beforeModel` / `modelStream` as usual, and the
199
+ `'structuredOutput'` middleware phase does **not** fire. When it
200
+ returns `false` (or is omitted), the engine takes the legacy
201
+ finalization path: agent loop, then a separate `structuredOutput` /
202
+ `structuredOutputStream` call with `'structuredOutput'` phase tagging.
203
+
204
+ Consumer code is identical across providers — always read the final object off `structured-output.complete`.
193
205
 
194
206
  ### Pattern 4: useChat with outputSchema (progressive UI)
195
207
 
@@ -123,6 +123,29 @@ export interface TextAdapter<
123
123
  structuredOutputStream?: (
124
124
  options: StructuredOutputOptions<TProviderOptions>,
125
125
  ) => AsyncIterable<StreamChunk>
126
+
127
+ /**
128
+ * Declares whether the adapter supports combining `tools` and a
129
+ * schema-constrained final answer in a single streaming request.
130
+ *
131
+ * When `true`, the engine wires `outputSchema` into the regular
132
+ * `chatStream()` call and skips the separate `runStructuredFinalization`
133
+ * round-trip. The model's natural final turn carries the
134
+ * schema-constrained JSON text and the engine harvests it from the agent
135
+ * loop's accumulated content.
136
+ *
137
+ * When `false`, `undefined`, or the method is omitted, the engine runs
138
+ * the agent loop without `outputSchema` and then issues a separate
139
+ * `structuredOutput` / `structuredOutputStream` call against the JSON
140
+ * schema for finalization (the legacy path).
141
+ *
142
+ * The method receives the per-call `modelOptions` so providers whose
143
+ * support depends on the resolved upstream model (e.g. OpenRouter) can
144
+ * answer per-request. Most adapters can return a constant.
145
+ */
146
+ supportsCombinedToolsAndSchema?: (
147
+ modelOptions?: TProviderOptions | undefined,
148
+ ) => boolean
126
149
  }
127
150
 
128
151
  /**
@@ -312,11 +312,20 @@ interface TextEngineConfig<
312
312
  * as the validated result and retrievable via
313
313
  * `getValidatedStructuredOutput()`. Used by `runAgenticStructuredOutput`
314
314
  * to perform Standard Schema validation inside the engine.
315
+ * - nativeCombined: when true, the adapter declared
316
+ * `supportsCombinedToolsAndSchema()` and the engine wires `jsonSchema`
317
+ * into the regular `chatStream` call instead of running a separate
318
+ * finalization round-trip. The agent loop's final-turn text is the
319
+ * schema-constrained JSON; the engine parses it from accumulated
320
+ * content. The `'structuredOutput'` middleware phase does NOT fire on
321
+ * this path — middleware sees the run through `beforeModel` /
322
+ * `modelStream` as usual.
315
323
  */
316
324
  finalStructuredOutput?: {
317
325
  jsonSchema: JSONSchema
318
326
  yieldChunks: boolean
319
327
  validate?: (data: unknown) => unknown
328
+ nativeCombined?: boolean
320
329
  }
321
330
  }
322
331
 
@@ -379,6 +388,16 @@ class TextEngine<
379
388
  // Structured-output finalization state (populated by runStructuredFinalization)
380
389
  private structuredOutputResult: { data: unknown; rawText: string } | null =
381
390
  null
391
+ // Native combined mode: tracks whether we've already emitted the synthetic
392
+ // `structured-output.start` event before the schema-constrained final-turn
393
+ // text begins streaming. The event must precede the first
394
+ // TEXT_MESSAGE_START so the client-side StreamProcessor routes the JSON
395
+ // deltas into a StructuredOutputPart instead of a plain TextPart.
396
+ private combinedStartEmitted = false
397
+ // Native combined mode: messageId we want the synthetic
398
+ // `structured-output.start` (and any error emitted before deltas arrive)
399
+ // to carry, so the client matches it to the streaming text deltas.
400
+ private combinedStructuredMessageId: string | null = null
382
401
  // Holds the validated value when `finalStructuredOutput.validate` is provided
383
402
  // and succeeds. Distinct from `structuredOutputResult.data` (the raw,
384
403
  // unvalidated payload from the structured-output.complete chunk).
@@ -393,6 +412,7 @@ class TextEngine<
393
412
  jsonSchema: JSONSchema
394
413
  yieldChunks: boolean
395
414
  validate?: (data: unknown) => unknown
415
+ nativeCombined?: boolean
396
416
  }
397
417
 
398
418
  constructor(
@@ -560,12 +580,19 @@ class TextEngine<
560
580
  return
561
581
  }
562
582
 
563
- // Skip the agent loop entirely when there are no tools AND a structured-
564
- // output finalization will run. Without tools the model has nothing to
565
- // do in the loop, so executing one iteration would burn an extra
566
- // provider call before the finalization request.
583
+ // Skip the agent loop entirely when there are no tools AND a separate
584
+ // structured-output finalization will run. Without tools the model has
585
+ // nothing to do in the loop, so executing one iteration would burn an
586
+ // extra provider call before the finalization request.
587
+ //
588
+ // Native combined mode does NOT skip — the agent loop itself produces
589
+ // the schema-constrained final answer in one pass (model emits the
590
+ // schema-constrained text on its natural final turn). Even with zero
591
+ // tools, the single chatStream call IS the structured-output call.
567
592
  const skipAgentLoop =
568
- !!this.finalStructuredOutput && this.tools.length === 0
593
+ !!this.finalStructuredOutput &&
594
+ this.tools.length === 0 &&
595
+ this.finalStructuredOutput.nativeCombined !== true
569
596
 
570
597
  if (!skipAgentLoop) {
571
598
  do {
@@ -584,11 +611,12 @@ class TextEngine<
584
611
  this.middlewareCtx.phase = 'beforeModel'
585
612
  this.middlewareCtx.iteration = this.iterationCount
586
613
  const iterConfig = this.buildMiddlewareConfig()
587
- const transformedConfig = await this.middlewareRunner.runOnConfig(
588
- this.middlewareCtx,
589
- iterConfig,
590
- )
591
- this.applyMiddlewareConfig(transformedConfig)
614
+ const iterTransformedConfig =
615
+ await this.middlewareRunner.runOnConfig(
616
+ this.middlewareCtx,
617
+ iterConfig,
618
+ )
619
+ this.applyMiddlewareConfig(iterTransformedConfig)
592
620
 
593
621
  yield* this.streamModelResponse()
594
622
  } else {
@@ -607,12 +635,20 @@ class TextEngine<
607
635
  // requested AND the run hasn't already errored/aborted, run it through
608
636
  // the middleware pipeline. The terminal hook fires once at the very
609
637
  // end (after finalization), not after the agent loop.
638
+ //
639
+ // Native combined mode takes a different path: the agent loop's final-
640
+ // turn text IS the schema-constrained JSON, so we harvest it from
641
+ // `accumulatedContent` instead of issuing a second provider call.
610
642
  if (
611
643
  this.finalStructuredOutput &&
612
644
  !this.isCancelled() &&
613
645
  !this.finalizationError
614
646
  ) {
615
- yield* this.runStructuredFinalization()
647
+ if (this.finalStructuredOutput.nativeCombined === true) {
648
+ yield* this.harvestCombinedStructuredOutput()
649
+ } else {
650
+ yield* this.runStructuredFinalization()
651
+ }
616
652
  }
617
653
 
618
654
  // Call terminal hook (skip when waiting for client — stream is paused, not finished).
@@ -777,6 +813,18 @@ class TextEngine<
777
813
  },
778
814
  )
779
815
 
816
+ // When the adapter declared `supportsCombinedToolsAndSchema()`, the
817
+ // activity layer set `nativeCombined: true` and we forward the
818
+ // pre-converted JSON Schema into the regular chatStream call. The
819
+ // adapter wires it into the upstream request (e.g. `response_format`,
820
+ // `text.format`, `output_format`) so the model's final-turn text is
821
+ // schema-constrained and the engine can harvest it from the agent loop
822
+ // without a separate finalization round-trip.
823
+ const combinedSchema =
824
+ this.finalStructuredOutput?.nativeCombined === true
825
+ ? this.finalStructuredOutput.jsonSchema
826
+ : undefined
827
+
780
828
  for await (const chunk of this.adapter.chatStream({
781
829
  model: this.params.model,
782
830
  messages: this.messages,
@@ -792,6 +840,7 @@ class TextEngine<
792
840
  threadId: this.threadId,
793
841
  runId: this.runIdOverride,
794
842
  parentRunId: this.parentRunIdOverride,
843
+ ...(combinedSchema ? { outputSchema: combinedSchema } : {}),
795
844
  })) {
796
845
  if (this.isCancelled()) {
797
846
  break
@@ -803,6 +852,44 @@ class TextEngine<
803
852
  // BEFORE middleware, so fields like finishReason, delta, etc. are available
804
853
  this.handleStreamChunk(chunk)
805
854
 
855
+ // Native combined mode: synthesize `structured-output.start` BEFORE
856
+ // the first TEXT_MESSAGE_START so the client-side StreamProcessor
857
+ // routes the schema-constrained JSON deltas into a
858
+ // StructuredOutputPart. We delay synthesis until we actually see
859
+ // text starting — intermediate tool-call iterations don't need it,
860
+ // and emitting at run-start would wrap tool-call commentary into a
861
+ // structured-output part too.
862
+ if (
863
+ this.finalStructuredOutput?.nativeCombined === true &&
864
+ this.finalStructuredOutput.yieldChunks &&
865
+ !this.combinedStartEmitted &&
866
+ chunk.type === EventType.TEXT_MESSAGE_START
867
+ ) {
868
+ this.combinedStartEmitted = true
869
+ const messageId =
870
+ typeof chunk.messageId === 'string' && chunk.messageId !== ''
871
+ ? chunk.messageId
872
+ : generateMessageId()
873
+ this.combinedStructuredMessageId = messageId
874
+ const synthStart: StreamChunk = {
875
+ type: EventType.CUSTOM,
876
+ name: 'structured-output.start',
877
+ value: { messageId },
878
+ model: this.params.model,
879
+ timestamp: Date.now(),
880
+ threadId: this.threadId,
881
+ ...(this.runIdOverride ? { runId: this.runIdOverride } : {}),
882
+ }
883
+ const synthOutputs = await this.middlewareRunner.runOnChunk(
884
+ this.middlewareCtx,
885
+ synthStart,
886
+ )
887
+ for (const outputChunk of synthOutputs) {
888
+ yield outputChunk
889
+ this.middlewareCtx.chunkIndex++
890
+ }
891
+ }
892
+
806
893
  // Pipe chunk through middleware (devtools middleware observes; strip-to-spec cleans)
807
894
  const outputChunks = await this.middlewareRunner.runOnChunk(
808
895
  this.middlewareCtx,
@@ -812,8 +899,13 @@ class TextEngine<
812
899
  // the agent loop, suppress the agent-loop's RUN_STARTED/RUN_FINISHED
813
900
  // here — the finalization step emits the single outer lifecycle pair
814
901
  // that reaches the consumer.
902
+ //
903
+ // Native combined mode does NOT issue a second adapter stream — the
904
+ // agent loop's lifecycle IS the outer pair the consumer sees.
815
905
  const suppressAgentLifecycle =
816
- !!this.finalStructuredOutput && this.finalStructuredOutput.yieldChunks
906
+ !!this.finalStructuredOutput &&
907
+ this.finalStructuredOutput.yieldChunks &&
908
+ this.finalStructuredOutput.nativeCombined !== true
817
909
  for (const outputChunk of outputChunks) {
818
910
  if (
819
911
  suppressAgentLifecycle &&
@@ -1948,6 +2040,179 @@ class TextEngine<
1948
2040
  }
1949
2041
  }
1950
2042
 
2043
+ /**
2044
+ * Native combined mode: harvest the structured output from the agent
2045
+ * loop's accumulated final-turn text (no separate provider call).
2046
+ *
2047
+ * The adapter wired `outputSchema` into the regular `chatStream` request,
2048
+ * so the model's final-turn text is the schema-constrained JSON. We parse
2049
+ * `this.accumulatedContent`, populate `this.structuredOutputResult`, emit
2050
+ * a synthetic `structured-output.complete` (and a `structured-output.start`
2051
+ * if one wasn't emitted earlier — only happens on the streaming path when
2052
+ * the model returned no text at all), and run the validate callback when
2053
+ * present. Failures populate `this.finalizationError` so the engine's
2054
+ * terminal-hook chooser routes to `onError` (per spec §7.3).
2055
+ *
2056
+ * The `'structuredOutput'` middleware phase intentionally does NOT fire on
2057
+ * this path — middleware sees the run through `beforeModel` / `modelStream`
2058
+ * as usual. See PR #605 / issue #605 for the design rationale.
2059
+ */
2060
+ private async *harvestCombinedStructuredOutput(): AsyncGenerator<StreamChunk> {
2061
+ if (!this.finalStructuredOutput) {
2062
+ throw new Error(
2063
+ 'harvestCombinedStructuredOutput called without finalStructuredOutput config',
2064
+ )
2065
+ }
2066
+
2067
+ const yieldChunks = this.finalStructuredOutput.yieldChunks
2068
+ const rawText = this.accumulatedContent
2069
+
2070
+ // Empty final-turn text means the agent loop terminated without the
2071
+ // model emitting any assistant content (e.g. early termination after
2072
+ // tool calls). Mirror the fallback path's "missing structured result"
2073
+ // error rather than silently returning undefined.
2074
+ if (rawText.length === 0) {
2075
+ this.finalizationError = {
2076
+ message: 'missing structured result',
2077
+ code: 'structured-output-missing-result',
2078
+ }
2079
+ } else {
2080
+ try {
2081
+ const parsed: unknown = JSON.parse(rawText)
2082
+ this.structuredOutputResult = { data: parsed, rawText }
2083
+ } catch (err: unknown) {
2084
+ const detail =
2085
+ rawText.slice(0, 200) + (rawText.length > 200 ? '...' : '')
2086
+ this.finalizationError = {
2087
+ message: `Failed to parse structured output as JSON. Content: ${detail}`,
2088
+ code: 'structured-output-parse-failed',
2089
+ cause: err,
2090
+ }
2091
+ }
2092
+ }
2093
+
2094
+ // Validate against the Standard Schema (when supplied). Validation
2095
+ // failures route through onError just like the fallback path.
2096
+ if (
2097
+ this.structuredOutputResult &&
2098
+ !this.finalizationError &&
2099
+ this.finalStructuredOutput.validate
2100
+ ) {
2101
+ try {
2102
+ const validated = this.finalStructuredOutput.validate(
2103
+ this.structuredOutputResult.data,
2104
+ )
2105
+ this.validatedStructuredOutput = validated
2106
+ this.hasValidatedStructuredOutput = true
2107
+ } catch (err: unknown) {
2108
+ const message = err instanceof Error ? err.message : String(err)
2109
+ this.finalizationError = {
2110
+ message,
2111
+ code: 'structured-output-validation-failed',
2112
+ cause: err,
2113
+ }
2114
+ }
2115
+ }
2116
+
2117
+ if (!yieldChunks) {
2118
+ // Promise<T> path: state is populated, nothing to yield. The
2119
+ // activity-layer caller pulls `structuredOutputResult` /
2120
+ // `validatedStructuredOutput` directly.
2121
+ return
2122
+ }
2123
+
2124
+ // Streaming path: emit a synthetic `structured-output.start` if the
2125
+ // model produced no text at all (so the client snaps an errored
2126
+ // StructuredOutputPart rather than nothing). The normal path already
2127
+ // emitted start before the first TEXT_MESSAGE_START in
2128
+ // `streamModelResponse`.
2129
+ if (!this.combinedStartEmitted) {
2130
+ this.combinedStartEmitted = true
2131
+ const messageId = this.combinedStructuredMessageId ?? generateMessageId()
2132
+ this.combinedStructuredMessageId = messageId
2133
+ const synthStart: StreamChunk = {
2134
+ type: EventType.CUSTOM,
2135
+ name: 'structured-output.start',
2136
+ value: { messageId },
2137
+ model: this.params.model,
2138
+ timestamp: Date.now(),
2139
+ threadId: this.threadId,
2140
+ ...(this.runIdOverride ? { runId: this.runIdOverride } : {}),
2141
+ }
2142
+ const startOutputs = await this.middlewareRunner.runOnChunk(
2143
+ this.middlewareCtx,
2144
+ synthStart,
2145
+ )
2146
+ for (const outputChunk of startOutputs) {
2147
+ yield outputChunk
2148
+ this.middlewareCtx.chunkIndex++
2149
+ }
2150
+ }
2151
+
2152
+ // On success, emit the synthetic `structured-output.complete` carrying
2153
+ // the parsed object + raw text. Pin the messageId so the client-side
2154
+ // handler can target the right UIMessage even when the agent loop's
2155
+ // terminal RUN_FINISHED has already cleared `activeMessageIds` (the
2156
+ // complete event yields AFTER the loop ends, by which point
2157
+ // `getActiveAssistantMessageId()` returns null and would otherwise drop
2158
+ // the event silently).
2159
+ if (this.structuredOutputResult && !this.finalizationError) {
2160
+ const completeChunk: StreamChunk = {
2161
+ type: EventType.CUSTOM,
2162
+ name: 'structured-output.complete',
2163
+ value: {
2164
+ object: this.structuredOutputResult.data,
2165
+ raw: this.structuredOutputResult.rawText,
2166
+ ...(this.combinedStructuredMessageId
2167
+ ? { messageId: this.combinedStructuredMessageId }
2168
+ : {}),
2169
+ },
2170
+ model: this.params.model,
2171
+ timestamp: Date.now(),
2172
+ threadId: this.threadId,
2173
+ ...(this.runIdOverride ? { runId: this.runIdOverride } : {}),
2174
+ }
2175
+ const completeOutputs = await this.middlewareRunner.runOnChunk(
2176
+ this.middlewareCtx,
2177
+ completeChunk,
2178
+ )
2179
+ for (const outputChunk of completeOutputs) {
2180
+ yield outputChunk
2181
+ this.middlewareCtx.chunkIndex++
2182
+ }
2183
+ }
2184
+
2185
+ // On failure, emit a synthetic RUN_ERROR so the streaming consumer's
2186
+ // `for await` doesn't end silently. Mirrors the fallback path.
2187
+ if (this.finalizationError) {
2188
+ const errChunk: StreamChunk = {
2189
+ type: EventType.RUN_ERROR,
2190
+ runId: this.runIdOverride ?? this.requestId,
2191
+ model: this.params.model,
2192
+ timestamp: Date.now(),
2193
+ threadId: this.threadId,
2194
+ message: this.finalizationError.message,
2195
+ ...(this.finalizationError.code
2196
+ ? { code: this.finalizationError.code }
2197
+ : {}),
2198
+ error: {
2199
+ message: this.finalizationError.message,
2200
+ ...(this.finalizationError.code
2201
+ ? { code: this.finalizationError.code }
2202
+ : {}),
2203
+ },
2204
+ }
2205
+ const errOutputs = await this.middlewareRunner.runOnChunk(
2206
+ this.middlewareCtx,
2207
+ errChunk,
2208
+ )
2209
+ for (const outputChunk of errOutputs) {
2210
+ yield outputChunk
2211
+ this.middlewareCtx.chunkIndex++
2212
+ }
2213
+ }
2214
+ }
2215
+
1951
2216
  private buildMiddlewareConfig(): ChatMiddlewareConfig {
1952
2217
  return {
1953
2218
  messages: this.messages,
@@ -2243,6 +2508,13 @@ async function runAgenticStructuredOutput<TSchema extends SchemaInput>(
2243
2508
  parseWithStandardSchema<InferSchemaType<TSchema>>(outputSchema, data)
2244
2509
  : undefined
2245
2510
 
2511
+ // Per issue #605: same capability check as the streaming path. When the
2512
+ // adapter handles tools + schema natively, the engine skips the separate
2513
+ // structured-output finalization call and harvests the JSON from the
2514
+ // agent loop's accumulated final-turn text.
2515
+ const nativeCombined =
2516
+ adapter.supportsCombinedToolsAndSchema?.(options.modelOptions) === true
2517
+
2246
2518
  const engine = new TextEngine(
2247
2519
  {
2248
2520
  adapter,
@@ -2256,6 +2528,7 @@ async function runAgenticStructuredOutput<TSchema extends SchemaInput>(
2256
2528
  jsonSchema,
2257
2529
  yieldChunks: false,
2258
2530
  ...(validate ? { validate } : {}),
2531
+ ...(nativeCombined ? { nativeCombined: true } : {}),
2259
2532
  },
2260
2533
  },
2261
2534
  logger,
@@ -2493,6 +2766,16 @@ async function* runStreamingStructuredOutputImpl<TSchema extends SchemaInput>(
2493
2766
  const model = adapter.model
2494
2767
  const logger = resolveDebugOption(debug)
2495
2768
 
2769
+ // Per issue #605: adapters that natively combine tools + schema-constrained
2770
+ // output in one streaming call (modern OpenAI, Anthropic 4.5+, Gemini 3+,
2771
+ // Grok 4+) opt in via `supportsCombinedToolsAndSchema()`. The engine then
2772
+ // forwards the schema into the regular `chatStream` call and harvests the
2773
+ // structured result from the agent loop's accumulated text — no separate
2774
+ // finalization round-trip, and the `'structuredOutput'` middleware phase
2775
+ // does not fire.
2776
+ const nativeCombined =
2777
+ adapter.supportsCombinedToolsAndSchema?.(options.modelOptions) === true
2778
+
2496
2779
  // Inputs may be UIMessages (from useChat) or ModelMessages (from server-side
2497
2780
  // callers). TextEngine handles the conversion uniformly.
2498
2781
  const engine = new TextEngine(
@@ -2504,7 +2787,11 @@ async function* runStreamingStructuredOutputImpl<TSchema extends SchemaInput>(
2504
2787
  >,
2505
2788
  middleware,
2506
2789
  context,
2507
- finalStructuredOutput: { jsonSchema, yieldChunks: true },
2790
+ finalStructuredOutput: {
2791
+ jsonSchema,
2792
+ yieldChunks: true,
2793
+ ...(nativeCombined ? { nativeCombined: true } : {}),
2794
+ },
2508
2795
  },
2509
2796
  logger,
2510
2797
  )
@@ -32,6 +32,14 @@ function safeJsonStringify(value: unknown): string {
32
32
  }
33
33
  }
34
34
 
35
+ function parseToolResultContent(content: string): unknown {
36
+ try {
37
+ return JSON.parse(content)
38
+ } catch {
39
+ return content
40
+ }
41
+ }
42
+
35
43
  /**
36
44
  * Collapse an array of ContentParts into the most compact ModelMessage content:
37
45
  * - Empty array → null
@@ -200,6 +208,7 @@ function createSegment(): AssistantSegment {
200
208
  function isToolCallIncluded(part: ToolCallPart): boolean {
201
209
  return (
202
210
  part.state === 'input-complete' ||
211
+ part.state === 'complete' ||
203
212
  part.state === 'approval-responded' ||
204
213
  part.output !== undefined
205
214
  )
@@ -467,10 +476,21 @@ export function modelMessagesToUIMessages(
467
476
  currentAssistantMessage &&
468
477
  currentAssistantMessage.role === 'assistant'
469
478
  ) {
479
+ const content = getTextContent(msg.content)
480
+ const toolCallPart = currentAssistantMessage.parts.find(
481
+ (part): part is ToolCallPart =>
482
+ part.type === 'tool-call' && part.id === msg.toolCallId,
483
+ )
484
+
485
+ if (toolCallPart) {
486
+ toolCallPart.output = parseToolResultContent(content)
487
+ toolCallPart.state = 'complete'
488
+ }
489
+
470
490
  currentAssistantMessage.parts.push({
471
491
  type: 'tool-result',
472
492
  toolCallId: msg.toolCallId,
473
- content: getTextContent(msg.content),
493
+ content,
474
494
  state: 'complete',
475
495
  })
476
496
  } else {
@@ -217,11 +217,7 @@ export function updateToolCallWithOutput(
217
217
 
218
218
  if (toolCallPart) {
219
219
  toolCallPart.output = errorText ? { error: errorText } : output
220
- if (state) {
221
- toolCallPart.state = state
222
- } else {
223
- toolCallPart.state = 'input-complete'
224
- }
220
+ toolCallPart.state = state ?? (errorText ? 'input-complete' : 'complete')
225
221
  }
226
222
 
227
223
  return { ...msg, parts }
@@ -378,6 +378,7 @@ export class StreamProcessor {
378
378
  // 3. It has a corresponding tool-result part (server tool completed)
379
379
  return toolParts.every(
380
380
  (part) =>
381
+ part.state === 'complete' ||
381
382
  part.state === 'approval-responded' ||
382
383
  (part.output !== undefined && !part.approval) ||
383
384
  toolResultIds.has(part.id),
@@ -103,10 +103,10 @@ function createId(prefix: string): string {
103
103
  * @example Generate speech from text
104
104
  * ```ts
105
105
  * import { generateSpeech } from '@tanstack/ai'
106
- * import { openaiTTS } from '@tanstack/ai-openai'
106
+ * import { openaiSpeech } from '@tanstack/ai-openai'
107
107
  *
108
108
  * const result = await generateSpeech({
109
- * adapter: openaiTTS('tts-1-hd'),
109
+ * adapter: openaiSpeech('tts-1-hd'),
110
110
  * text: 'Hello, welcome to TanStack AI!',
111
111
  * voice: 'nova'
112
112
  * })
@@ -117,7 +117,7 @@ function createId(prefix: string): string {
117
117
  * @example With format and speed options
118
118
  * ```ts
119
119
  * const result = await generateSpeech({
120
- * adapter: openaiTTS('tts-1'),
120
+ * adapter: openaiSpeech('tts-1'),
121
121
  * text: 'This is slower speech.',
122
122
  * voice: 'alloy',
123
123
  * format: 'wav',
package/src/types.ts CHANGED
@@ -40,6 +40,7 @@ export type ToolCallState =
40
40
  | 'input-complete' // All arguments received
41
41
  | 'approval-requested' // Waiting for user approval
42
42
  | 'approval-responded' // User has approved/denied
43
+ | 'complete' // Result is complete
43
44
 
44
45
  /**
45
46
  * Tool result states - track the lifecycle of a tool result
@@ -799,10 +800,26 @@ export interface TextOptions<
799
800
 
800
801
  /**
801
802
  * Schema for structured output.
802
- * When provided, the adapter should use the provider's native structured output API
803
- * to ensure the response conforms to this schema.
804
- * The schema will be converted to JSON Schema format before being sent to the provider.
805
- * Supports any Standard JSON Schema compliant library (Zod, ArkType, Valibot, etc.).
803
+ *
804
+ * **Two distinct use sites:**
805
+ *
806
+ * 1. **User-facing (activity layer):** accepts any
807
+ * {@link SchemaInput} — Zod, ArkType, Valibot, or a raw JSON Schema.
808
+ * The activity layer converts to JSON Schema before handing off.
809
+ *
810
+ * 2. **Adapter-facing (`chatStream` call):** the engine populates this with
811
+ * a pre-converted JSON Schema **only** when the adapter declared
812
+ * `supportsCombinedToolsAndSchema(modelOptions) === true`. The adapter
813
+ * should then wire the schema into the upstream request (e.g.
814
+ * `response_format: { type: 'json_schema', ... }`, `text.format`,
815
+ * `output_format`) alongside any `tools`. The model's natural final
816
+ * turn carries the schema-constrained JSON text and the engine
817
+ * harvests it from the agent loop without a separate finalization
818
+ * round-trip.
819
+ *
820
+ * Adapters that did NOT declare the capability never see this field
821
+ * populated — the engine instead invokes `structuredOutput` /
822
+ * `structuredOutputStream` after the agent loop.
806
823
  */
807
824
  outputSchema?: SchemaInput
808
825
  /**