@bastani/pi-ai 0.9.20-alpha.1 → 0.9.20-alpha.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,9 +1,17 @@
1
1
  # Changelog
2
2
 
3
- This package is a Bastani fork of `@earendil-works/pi-ai`. Upstream history at the audited Pi `main` sync point (`92d8e2d17d4f357788381c49ce2cdb3f4ed1f21c`) lives in [earendil-works/pi](https://github.com/earendil-works/pi/blob/92d8e2d17d4f357788381c49ce2cdb3f4ed1f21c/packages/ai/CHANGELOG.md).
3
+ This package is a Bastani fork of `@earendil-works/pi-ai`. Upstream history at the audited Pi `main` sync point (`6671c604766b3670ed95f405aa7856835d0ca702`) lives in [earendil-works/pi](https://github.com/earendil-works/pi/blob/6671c604766b3670ed95f405aa7856835d0ca702/packages/ai/CHANGELOG.md).
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.9.20-alpha.3] - 2026-09-16
8
+
9
+ ### Fixed
10
+
11
+ - Fixed OpenAI-compatible Responses errors to identify the actual provider instead of always labeling them as OpenAI errors ([#9298](https://github.com/earendil-works/pi/issues/9298)).
12
+ - Fixed Amazon Bedrock one-hour cache writes being priced at the five-minute rate ([#9457](https://github.com/earendil-works/pi/issues/9457)).
13
+ - Fixed Baseten requests to send session-affinity headers from `sessionId` for automatic prompt-cache routing ([#9629](https://github.com/earendil-works/pi/issues/9629)).
14
+
7
15
  ## [0.9.19] - 2026-09-13
8
16
 
9
17
  Cumulative release of the `0.9.19-alpha.2` through `0.9.19-alpha.4` prereleases. Per-change details remain in the unchanged prerelease sections below.
package/NOTICE.md CHANGED
@@ -6,7 +6,7 @@ monorepo at `packages/ai` and publishes at the same version as `@bastani/atomic`
6
6
 
7
7
  - Upstream package: [`@earendil-works/pi-ai`](https://www.npmjs.com/package/@earendil-works/pi-ai)
8
8
  - Original fork point: `v0.84.2` (`914cf1472e715297caa30db4b9535d534a9eb718`)
9
- - Pi AI fixes and generated image catalog synced through audited upstream `main`: `earendil-works/pi@92d8e2d17d4f357788381c49ce2cdb3f4ed1f21c`
9
+ - Pi AI fixes and generated image catalog synced through audited upstream `main`: `earendil-works/pi@6671c604766b3670ed95f405aa7856835d0ca702`
10
10
  - Catalog JSON under `src/providers/data/` is generated at build time from models.dev, matching upstream. It is not committed.
11
11
 
12
12
  Original work is Copyright (c) 2025 Mario Zechner and is licensed under the MIT License.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # @bastani/pi-ai
2
2
 
3
- Bastani-branded fork of [`@earendil-works/pi-ai`](https://www.npmjs.com/package/@earendil-works/pi-ai) from [earendil-works/pi](https://github.com/earendil-works/pi). Originally forked at **v0.84.2** (`914cf1472e715297caa30db4b9535d534a9eb718`); upstream Pi AI fixes and the generated image catalog are synced through [`92d8e2d17d4f357788381c49ce2cdb3f4ed1f21c`](https://github.com/earendil-works/pi/commit/92d8e2d17d4f357788381c49ce2cdb3f4ed1f21c), the audited Pi `main` sync point. `@bastani/pi-ai` publishes at the same version as Atomic. `npm run build` refreshes the models.dev catalog, same as upstream.
3
+ Bastani-branded fork of [`@earendil-works/pi-ai`](https://www.npmjs.com/package/@earendil-works/pi-ai) from [earendil-works/pi](https://github.com/earendil-works/pi). Originally forked at **v0.84.2** (`914cf1472e715297caa30db4b9535d534a9eb718`); upstream Pi AI fixes and the generated image catalog are synced through [`6671c604766b3670ed95f405aa7856835d0ca702`](https://github.com/earendil-works/pi/commit/6671c604766b3670ed95f405aa7856835d0ca702), the audited Pi `main` sync point. `@bastani/pi-ai` publishes at the same version as Atomic. `npm run build` refreshes the models.dev catalog, same as upstream.
4
4
 
5
5
  The public API is a drop-in replacement: install `@bastani/pi-ai` and import from `@bastani/pi-ai` instead of `@earendil-works/pi-ai`. See [NOTICE.md](NOTICE.md). This package lives in the Atomic monorepo and publishes from `.github/workflows/publish.yml`. The first npm version must be published by hand so trusted publishing can be attached.
6
6
 
@@ -1577,6 +1577,8 @@ Provider notes:
1577
1577
 
1578
1578
  **OpenAI Codex**: Requires a ChatGPT Plus or Pro subscription. Provides access to GPT-5.x Codex models with extended context windows and reasoning capabilities. The library automatically handles session-based prompt caching when `sessionId` is provided in stream options unless `cacheRetention` is `"none"`. You can set `transport` in stream options to `"sse"`, `"websocket"`, or `"auto"` for Codex Responses transport selection. When using WebSocket with a `sessionId` and cache retention enabled, connections are reused per session and expire after 5 minutes of inactivity.
1579
1579
 
1580
+ Call `cleanupSessionResources(sessionId)` when finished with a Codex session so its pooled WebSocket connection does not keep the process alive. Import it from `@bastani/pi-ai`.
1581
+
1580
1582
  **Azure OpenAI (Responses)**: Uses the Responses API only. Set `AZURE_OPENAI_API_KEY` and either `AZURE_OPENAI_BASE_URL` or `AZURE_OPENAI_RESOURCE_NAME`. `AZURE_OPENAI_BASE_URL` supports both `https://<resource>.openai.azure.com` and `https://<resource>.cognitiveservices.azure.com`; root endpoints are normalized to `.../openai/v1` automatically. Use `AZURE_OPENAI_API_VERSION` (defaults to `v1`) to override the API version if needed. Deployment names are treated as model IDs by default, override with `azureDeploymentName` or `AZURE_OPENAI_DEPLOYMENT_NAME_MAP` using comma-separated `model-id=deployment` pairs (for example `gpt-4o-mini=my-deployment,gpt-4o=prod`). Legacy deployment-based URLs are intentionally unsupported.
1581
1583
 
1582
1584
  **GitHub Copilot**: If you get "The requested model is not supported" error, enable the model manually in VS Code: open Copilot Chat, click the model selector, select the model (warning icon), and click "Enable".
@@ -1 +1 @@
1
- {"version":3,"file":"bedrock-converse-stream.d.ts","sourceRoot":"","sources":["../../src/api/bedrock-converse-stream.ts"],"names":[],"mappings":"AA8BA,OAAO,KAAK,EAUX,mBAAmB,EAEnB,cAAc,EACd,aAAa,EAEb,eAAe,EAEf,aAAa,EAIb,MAAM,aAAa,CAAC;AAmBrB,MAAM,MAAM,sBAAsB,GAAG,YAAY,GAAG,SAAS,CAAC;AAE9D,MAAM,WAAW,cAAe,SAAQ,aAAa;IACpD,MAAM,CAAC,EAAE,MAAM,CAAC;IAChB,OAAO,CAAC,EAAE,MAAM,CAAC;IACjB,UAAU,CAAC,EAAE,MAAM,GAAG,KAAK,GAAG,MAAM,GAAG;QAAE,IAAI,EAAE,MAAM,CAAC;QAAC,IAAI,EAAE,MAAM,CAAA;KAAE,CAAC;IAEtE,SAAS,CAAC,EAAE,aAAa,CAAC;IAE1B,eAAe,CAAC,EAAE,eAAe,CAAC;IAElC,mBAAmB,CAAC,EAAE,OAAO,CAAC;IAC9B;;;;;;;;;OASG;IACH,eAAe,CAAC,EAAE,sBAAsB,CAAC;IACzC;;;sGAGkG;IAClG,eAAe,CAAC,EAAE,MAAM,CAAC,MAAM,EAAE,MAAM,CAAC,CAAC;IACzC;;;;yGAIqG;IACrG,WAAW,CAAC,EAAE,MAAM,CAAC;CACrB;AAcD,eAAO,MAAM,MAAM,EAAE,cAAc,CAAC,yBAAyB,EAAE,cAAc,CAwO5E,CAAC;AAwKF,eAAO,MAAM,YAAY,EAAE,cAAc,CAAC,yBAAyB,EAAE,mBAAmB,CAiDvF,CAAC","sourcesContent":["import type { Agent as HttpsAgent } from \"node:https\";\nimport {\n\tBedrockRuntimeClient,\n\ttype BedrockRuntimeClientConfig,\n\tBedrockRuntimeServiceException,\n\tStopReason as BedrockStopReason,\n\ttype Tool as BedrockTool,\n\tCachePointType,\n\tCacheTTL,\n\ttype ContentBlock,\n\ttype ContentBlockDeltaEvent,\n\ttype ContentBlockStartEvent,\n\ttype ContentBlockStopEvent,\n\tConversationRole,\n\tConverseStreamCommand,\n\ttype ConverseStreamMetadataEvent,\n\tDocumentFormat,\n\tImageFormat,\n\ttype Message,\n\ttype SystemContentBlock,\n\ttype ToolChoice,\n\ttype ToolConfiguration,\n\ttype ToolResultContentBlock,\n\tToolResultStatus,\n} from \"@aws-sdk/client-bedrock-runtime\";\nimport { NodeHttpHandler } from \"@smithy/node-http-handler\";\nimport type { BuildMiddleware, DeserializeMiddleware, DocumentType, HttpResponse, MetadataBearer } from \"@smithy/types\";\nimport { HttpProxyAgent } from \"http-proxy-agent\";\nimport { HttpsProxyAgent } from \"https-proxy-agent\";\nimport { calculateCost } from \"../models.ts\";\nimport type {\n\tApi,\n\tAssistantMessage,\n\tCacheRetention,\n\tContext,\n\tDocumentContent,\n\tImageContent,\n\tModel,\n\tProviderEnv,\n\tProviderResponse,\n\tSimpleStreamOptions,\n\tStopReason,\n\tStreamFunction,\n\tStreamOptions,\n\tTextContent,\n\tThinkingBudgets,\n\tThinkingContent,\n\tThinkingLevel,\n\tTool,\n\tToolCall,\n\tToolResultMessage,\n} from \"../types.ts\";\nimport { appendAssistantMessageDiagnostic } from \"../utils/diagnostics.ts\";\nimport { assertSupportedDocumentMimeType } from \"../utils/document-input.ts\";\nimport { normalizeProviderError } from \"../utils/error-body.ts\";\nimport { AssistantMessageEventStream } from \"../utils/event-stream.ts\";\nimport { providerHeadersToRecord } from \"../utils/headers.ts\";\nimport { parseStreamingJson } from \"../utils/json-parse.ts\";\nimport { resolveHttpProxyUrlForTarget } from \"../utils/node-http-proxy.ts\";\nimport { getProviderEnvValue } from \"../utils/provider-env.ts\";\nimport { sanitizeSurrogates } from \"../utils/sanitize-unicode.ts\";\nimport { getJsonSchemaToolParameters, resolveJsonSchemaStrictSampling } from \"./constrained-sampling.ts\";\nimport {\n\tadjustMaxTokensForThinking,\n\tbuildBaseOptions,\n\tclampMaxTokensToContext,\n\tclampReasoning,\n} from \"./simple-options.ts\";\nimport { transformMessages } from \"./transform-messages.ts\";\n\nexport type BedrockThinkingDisplay = \"summarized\" | \"omitted\";\n\nexport interface BedrockOptions extends StreamOptions {\n\tregion?: string;\n\tprofile?: string;\n\ttoolChoice?: \"auto\" | \"any\" | \"none\" | { type: \"tool\"; name: string };\n\t/* See https://docs.aws.amazon.com/bedrock/latest/userguide/inference-reasoning.html for supported models. */\n\treasoning?: ThinkingLevel;\n\t/* Custom token budgets per thinking level. Overrides default budgets. */\n\tthinkingBudgets?: ThinkingBudgets;\n\t/* Only supported by Claude 4.x models, see https://docs.aws.amazon.com/bedrock/latest/userguide/claude-messages-extended-thinking.html#claude-messages-extended-thinking-tool-use-interleaved */\n\tinterleavedThinking?: boolean;\n\t/**\n\t * Controls how Claude's thinking content is returned in responses.\n\t * - \"summarized\": Thinking blocks contain summarized thinking text (default here).\n\t * - \"omitted\": Thinking content is redacted but the signature still travels back\n\t * for multi-turn continuity, reducing time-to-first-text-token.\n\t *\n\t * Note: Anthropic's API default for Claude Opus 4.8 and Mythos Preview is\n\t * \"omitted\". We default to \"summarized\" here to keep behavior consistent with\n\t * older Claude 4 models. Only applies to Claude models on Bedrock.\n\t */\n\tthinkingDisplay?: BedrockThinkingDisplay;\n\t/** Key-value pairs attached to the inference request for cost allocation tagging.\n\t * Keys: max 64 chars, no `aws:` prefix. Values: max 256 chars. Max 50 pairs.\n\t * Tags appear in AWS Cost Explorer split cost allocation data.\n\t * @see https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ConverseStream.html */\n\trequestMetadata?: Record<string, string>;\n\t/** Bearer token for Bedrock API key authentication.\n\t * When set, bypasses SigV4 signing and sends Authorization: Bearer <token> instead.\n\t * Requires `bedrock:CallWithBearerToken` IAM permission on the token's identity.\n\t * Set via AWS_BEARER_TOKEN_BEDROCK env var or pass directly.\n\t * @see https://docs.aws.amazon.com/service-authorization/latest/reference/list_amazonbedrock.html */\n\tbearerToken?: string;\n}\n\ntype Block = (TextContent | ThinkingContent | ToolCall) & {\n\tindex?: number;\n\tpartialJson?: string;\n\t/** Scratch buffer for encrypted reasoning deltas, joined into `thinkingSignature`. */\n\tredactedChunks?: Uint8Array[];\n};\n\nconst EMPTY_TEXT_PLACEHOLDER = \"<empty>\";\n\n/** Matches the placeholder the Anthropic API path uses for redacted thinking. */\nconst REDACTED_THINKING_PLACEHOLDER = \"[Reasoning redacted]\";\n\nexport const stream: StreamFunction<\"bedrock-converse-stream\", BedrockOptions> = (\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcontext: Context,\n\toptions: BedrockOptions = {},\n): AssistantMessageEventStream => {\n\tconst stream = new AssistantMessageEventStream();\n\n\t(async () => {\n\t\tconst output: AssistantMessage = {\n\t\t\trole: \"assistant\",\n\t\t\tcontent: [],\n\t\t\tapi: \"bedrock-converse-stream\" as Api,\n\t\t\tprovider: model.provider,\n\t\t\tmodel: model.id,\n\t\t\tusage: {\n\t\t\t\tinput: 0,\n\t\t\t\toutput: 0,\n\t\t\t\tcacheRead: 0,\n\t\t\t\tcacheWrite: 0,\n\t\t\t\ttotalTokens: 0,\n\t\t\t\tcost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },\n\t\t\t},\n\t\t\tstopReason: \"pending\",\n\t\t\ttimestamp: Date.now(),\n\t\t};\n\n\t\tconst blocks = output.content as Block[];\n\n\t\t// A profile explicitly configured through pi's auth flow (the `profile`\n\t\t// option or scoped `AWS_PROFILE` on the stored credential's env) must win\n\t\t// over ambient AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY. The SDK default\n\t\t// chain already prefers a configured profile over env keys, but only when\n\t\t// `credentials` is not set on the client config. See #6957.\n\t\tconst optionsProfile = options.profile || options.env?.AWS_PROFILE;\n\t\tconst config: BedrockRuntimeClientConfig = {\n\t\t\tprofile: optionsProfile || getProviderEnvValue(\"AWS_PROFILE\", options.env),\n\t\t};\n\t\tconst configuredRegion = getConfiguredBedrockRegion(options);\n\t\tconst hasAmbientConfiguredProfile = Boolean(getProviderEnvValue(\"AWS_PROFILE\"));\n\t\tconst endpointRegion = getStandardBedrockEndpointRegion(model.baseUrl);\n\t\tconst useExplicitEndpoint = shouldUseExplicitBedrockEndpoint(\n\t\t\tmodel.baseUrl,\n\t\t\tconfiguredRegion,\n\t\t\thasAmbientConfiguredProfile,\n\t\t);\n\n\t\t// Only pin standard AWS Bedrock runtime endpoints when no region or ambient AWS_PROFILE is configured.\n\t\t// This preserves custom endpoints (VPC/proxy) from #3402 without forcing built-in\n\t\t// catalog defaults such as us-east-1 to override AWS_REGION/AWS_PROFILE.\n\t\tif (useExplicitEndpoint) {\n\t\t\tconfig.endpoint = model.baseUrl;\n\t\t}\n\n\t\t// Resolve bearer token for Bedrock API key auth.\n\t\tconst skipAuth = getProviderEnvValue(\"AWS_BEDROCK_SKIP_AUTH\", options.env) === \"1\";\n\t\tconst bearerToken =\n\t\t\toptions.bearerToken ||\n\t\t\toptions.apiKey ||\n\t\t\tgetProviderEnvValue(\"AWS_BEARER_TOKEN_BEDROCK\", options.env) ||\n\t\t\tundefined;\n\t\tconst useBearerToken = bearerToken !== undefined && !skipAuth;\n\n\t\t// in Node.js/Bun environment only\n\t\tif (typeof process !== \"undefined\" && (process.versions?.node || process.versions?.bun)) {\n\t\t\t// Region resolution: ARN-embedded > explicit option > env vars > SDK default chain.\n\t\t\t// When the model ID is an inference profile ARN, extract the region from it.\n\t\t\t// This avoids conflicts with AWS_REGION set for other services.\n\t\t\tconst arnRegionMatch = model.id.match(/^arn:aws(?:-[a-z0-9-]+)?:bedrock:([a-z0-9-]+):/);\n\t\t\tif (arnRegionMatch) {\n\t\t\t\tconfig.region = arnRegionMatch[1];\n\t\t\t} else if (configuredRegion) {\n\t\t\t\tconfig.region = configuredRegion;\n\t\t\t} else if (endpointRegion && useExplicitEndpoint) {\n\t\t\t\tconfig.region = endpointRegion;\n\t\t\t} else if (!hasAmbientConfiguredProfile) {\n\t\t\t\tconfig.region = \"us-east-1\";\n\t\t\t}\n\n\t\t\t// Support proxies that don't need authentication\n\t\t\tif (skipAuth) {\n\t\t\t\tconfig.credentials = {\n\t\t\t\t\taccessKeyId: \"dummy-access-key\",\n\t\t\t\t\tsecretAccessKey: \"dummy-secret-key\",\n\t\t\t\t};\n\t\t\t}\n\n\t\t\tconst credentials = getConfiguredBedrockCredentials(options.env);\n\t\t\tif (!skipAuth && credentials && !optionsProfile) {\n\t\t\t\tconfig.credentials = credentials;\n\t\t\t}\n\n\t\t\tconst proxyUrl = resolveHttpProxyUrlForTarget(model.baseUrl, options.env);\n\t\t\tif (proxyUrl) {\n\t\t\t\t// Bedrock runtime uses NodeHttp2Handler by default since v3.798.0, which is based\n\t\t\t\t// on `http2` module and has no support for http agent.\n\t\t\t\t// Use NodeHttpHandler to support HTTP(S) proxy agents.\n\t\t\t\tconfig.requestHandler = new NodeHttpHandler({\n\t\t\t\t\thttpAgent: new HttpProxyAgent(proxyUrl),\n\t\t\t\t\thttpsAgent: new HttpsProxyAgent(proxyUrl) as unknown as HttpsAgent,\n\t\t\t\t});\n\t\t\t} else if (getProviderEnvValue(\"AWS_BEDROCK_FORCE_HTTP1\", options.env) === \"1\") {\n\t\t\t\t// Some custom endpoints require HTTP/1.1 instead of HTTP/2\n\t\t\t\tconfig.requestHandler = new NodeHttpHandler();\n\t\t\t}\n\t\t} else {\n\t\t\t// Non-Node environment (browser): fall back to us-east-1 since\n\t\t\t// there's no config file resolution available.\n\t\t\tconfig.region =\n\t\t\t\tconfiguredRegion || (endpointRegion && useExplicitEndpoint ? endpointRegion : undefined) || \"us-east-1\";\n\t\t}\n\n\t\tif (useBearerToken) {\n\t\t\tconfig.token = { token: bearerToken };\n\t\t\tconfig.authSchemePreference = [\"httpBearerAuth\"];\n\t\t}\n\n\t\t// Kept outside the try so the catch can still correlate a mid-stream failure:\n\t\t// exceptions delivered as stream events carry no HTTP metadata of their own.\n\t\tlet responseRequestId: string | undefined;\n\n\t\ttry {\n\t\t\tconst supportsStrictMode = model.compat?.supportsStrictMode ?? false;\n\t\t\tconst client = new BedrockRuntimeClient(config);\n\t\t\tlet observedRawResponse = false;\n\t\t\tif (options.onResponse) {\n\t\t\t\taddResponseHeadersMiddleware(client, options.onResponse, model, () => {\n\t\t\t\t\tobservedRawResponse = true;\n\t\t\t\t});\n\t\t\t}\n\t\t\tconst customHeaders = providerHeadersToRecord({ ...model.headers, ...options.headers });\n\t\t\tif (customHeaders) {\n\t\t\t\taddCustomHeadersMiddleware(client, customHeaders);\n\t\t\t}\n\t\t\tconst cacheRetention = resolveCacheRetention(options.cacheRetention, options.env);\n\t\t\tconst inferenceMaxTokens = options.maxTokens ?? (isAnthropicClaudeModel(model) ? model.maxTokens : undefined);\n\t\t\tlet commandInput = {\n\t\t\t\tmodelId: model.id,\n\t\t\t\tmessages: convertMessages(context, model, cacheRetention, options.env),\n\t\t\t\tsystem: buildSystemPrompt(context.systemPrompt, model, cacheRetention, options.env),\n\t\t\t\tinferenceConfig: {\n\t\t\t\t\t...(inferenceMaxTokens !== undefined && { maxTokens: inferenceMaxTokens }),\n\t\t\t\t\t// Claude Fable 5.1 rejects non-default `temperature`, `top_p`, and `top_k` on every\n\t\t\t\t\t// request. Generated metadata marks such models `supportsTemperature: false`; every\n\t\t\t\t\t// other model keeps sending the field exactly as before.\n\t\t\t\t\t// https://platform.claude.com/docs/en/build-with-claude/thinking\n\t\t\t\t\t...(options.temperature !== undefined &&\n\t\t\t\t\t\tmodel.compat?.supportsTemperature !== false && { temperature: options.temperature }),\n\t\t\t\t},\n\t\t\t\ttoolConfig: convertToolConfig(context.tools, options.toolChoice, supportsStrictMode, model),\n\t\t\t\tadditionalModelRequestFields: buildAdditionalModelRequestFields(model, options),\n\t\t\t\t...(options.requestMetadata !== undefined && { requestMetadata: options.requestMetadata }),\n\t\t\t};\n\t\t\tconst nextCommandInput = await options?.onPayload?.(commandInput, model);\n\t\t\tif (nextCommandInput !== undefined) {\n\t\t\t\tcommandInput = nextCommandInput as typeof commandInput;\n\t\t\t}\n\t\t\tconst command = new ConverseStreamCommand(commandInput);\n\n\t\t\tconst response = await client.send(command, { abortSignal: options.signal });\n\t\t\tresponseRequestId = normalizeDiagnosticValue(response.$metadata.requestId);\n\t\t\tif (!observedRawResponse && response.$metadata.httpStatusCode !== undefined) {\n\t\t\t\tconst responseHeaders: Record<string, string> = {};\n\t\t\t\tif (response.$metadata.requestId) {\n\t\t\t\t\tresponseHeaders[\"x-amzn-requestid\"] = response.$metadata.requestId;\n\t\t\t\t}\n\t\t\t\tawait options?.onResponse?.({ status: response.$metadata.httpStatusCode, headers: responseHeaders }, model);\n\t\t\t}\n\n\t\t\tfor await (const item of response.stream!) {\n\t\t\t\tif (item.messageStart) {\n\t\t\t\t\tif (item.messageStart.role !== ConversationRole.ASSISTANT) {\n\t\t\t\t\t\tthrow new Error(\"Unexpected assistant message start but got user message start instead\");\n\t\t\t\t\t}\n\t\t\t\t\tstream.push({ type: \"start\", partial: output });\n\t\t\t\t} else if (item.contentBlockStart) {\n\t\t\t\t\thandleContentBlockStart(item.contentBlockStart, blocks, output, stream);\n\t\t\t\t} else if (item.contentBlockDelta) {\n\t\t\t\t\thandleContentBlockDelta(item.contentBlockDelta, blocks, output, stream);\n\t\t\t\t} else if (item.contentBlockStop) {\n\t\t\t\t\thandleContentBlockStop(item.contentBlockStop, blocks, output, stream);\n\t\t\t\t} else if (item.messageStop) {\n\t\t\t\t\toutput.rawStopReason = item.messageStop.stopReason;\n\t\t\t\t\tconst { stopReason, errorMessage } = mapStopReason(item.messageStop.stopReason);\n\t\t\t\t\toutput.stopReason = stopReason;\n\t\t\t\t\tif (errorMessage) {\n\t\t\t\t\t\toutput.errorMessage = errorMessage;\n\t\t\t\t\t}\n\t\t\t\t} else if (item.metadata) {\n\t\t\t\t\thandleMetadata(item.metadata, model, output);\n\t\t\t\t} else if (item.internalServerException) {\n\t\t\t\t\tthrow item.internalServerException;\n\t\t\t\t} else if (item.modelStreamErrorException) {\n\t\t\t\t\tthrow item.modelStreamErrorException;\n\t\t\t\t} else if (item.validationException) {\n\t\t\t\t\tthrow item.validationException;\n\t\t\t\t} else if (item.throttlingException) {\n\t\t\t\t\tthrow item.throttlingException;\n\t\t\t\t} else if (item.serviceUnavailableException) {\n\t\t\t\t\tthrow item.serviceUnavailableException;\n\t\t\t\t}\n\t\t\t}\n\n\t\t\tif (options.signal?.aborted) {\n\t\t\t\tthrow new Error(\"Request was aborted\");\n\t\t\t}\n\n\t\t\tif (output.stopReason === \"pending\") {\n\t\t\t\tthrow new Error(\"Bedrock stream ended without a stop reason\");\n\t\t\t}\n\t\t\tif (output.stopReason === \"error\" || output.stopReason === \"aborted\") {\n\t\t\t\tthrow new Error(output.errorMessage || \"An unknown error occurred\");\n\t\t\t}\n\n\t\t\t// A stream can settle without stopping every block, so finalize here too.\n\t\t\tfor (const block of output.content) finalizeStreamingBlock(block as Block);\n\t\t\tstream.push({ type: \"done\", reason: output.stopReason, message: output });\n\t\t\tstream.end();\n\t\t} catch (error) {\n\t\t\tfor (const block of output.content) {\n\t\t\t\tfinalizeStreamingBlock(block as Block);\n\t\t\t}\n\t\t\toutput.stopReason = options.signal?.aborted ? \"aborted\" : \"error\";\n\t\t\toutput.errorMessage = formatBedrockError(error);\n\t\t\tif (output.stopReason === \"error\") {\n\t\t\t\tappendBedrockFailureDiagnostic(output, error, responseRequestId);\n\t\t\t}\n\t\t\tstream.push({ type: \"error\", reason: output.stopReason, error: output });\n\t\t\tstream.end();\n\t\t}\n\t})();\n\n\treturn stream;\n};\n\n/**\n * Human-readable prefixes for Bedrock SDK exception names.\n * The downstream retry logic in agent-session matches patterns like\n * `server.?error` and `service.?unavailable`, so we preserve the legacy\n * prefix format rather than using the raw SDK exception name.\n */\nconst BEDROCK_ERROR_PREFIXES: Record<string, string> = {\n\tInternalServerException: \"Internal server error\",\n\tModelStreamErrorException: \"Model stream error\",\n\tValidationException: \"Validation error\",\n\tThrottlingException: \"Throttling error\",\n\tServiceUnavailableException: \"Service unavailable\",\n};\n\n/**\n * Some models reject the account/profile's configured Bedrock data retention mode\n * (e.g. \"data retention mode 'default' is not available for this model\"). Point\n * users at the AWS docs explaining how to configure a supported mode.\n */\nconst BEDROCK_DATA_RETENTION_DOCS_URL = \"https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html\";\n\n/**\n * Format a Bedrock error with a human-readable prefix.\n * AWS SDK exceptions (both from `client.send()` and from stream event items)\n * extend BedrockRuntimeServiceException. We map the `.name` to a stable\n * human-readable prefix so downstream consumers (retry logic, context-overflow\n * detection) can distinguish error categories via simple string matching.\n */\nfunction formatBedrockError(error: unknown): string {\n\tconst norm = normalizeProviderError(error);\n\t// Surface the raw HTTP body (with status) when the SDK did not fold it into\n\t// the message; otherwise fall back to the message. This is what stops a\n\t// gateway 403 from collapsing to `Unknown: UnknownError`.\n\tconst core =\n\t\t!norm.messageCarriesBody && norm.status !== undefined && norm.body !== undefined\n\t\t\t? `${norm.status}: ${norm.body}`\n\t\t\t: norm.message;\n\tconst dataRetentionHint = /data retention mode/i.test(core)\n\t\t? ` See ${BEDROCK_DATA_RETENTION_DOCS_URL} for supported data retention modes.`\n\t\t: \"\";\n\tif (error instanceof BedrockRuntimeServiceException) {\n\t\tconst prefix = BEDROCK_ERROR_PREFIXES[error.name] ?? error.name;\n\t\treturn `${prefix}: ${core}${dataRetentionHint}`;\n\t}\n\treturn `${core}${dataRetentionHint}`;\n}\n\ntype SdkErrorMetadata = { $metadata?: { httpStatusCode?: unknown; requestId?: unknown } };\n\n/** Over-long header values are dropped rather than truncated: a truncated request id is not a request id. */\nconst MAX_BEDROCK_DIAGNOSTIC_VALUE_CHARS = 200;\n\nfunction normalizeDiagnosticValue(value: unknown): string | undefined {\n\tif (typeof value !== \"string\") return undefined;\n\tconst trimmed = value.trim();\n\tif (trimmed.length === 0 || trimmed.length > MAX_BEDROCK_DIAGNOSTIC_VALUE_CHARS) return undefined;\n\treturn trimmed;\n}\n\n/**\n * The SDK puts the modeled code on `error.name` for service exceptions and unmodeled stream errors alike, so\n * do not narrow to `BedrockRuntimeServiceException`. Modeled Bedrock errors all end in `Exception`, unlike\n * transport names such as `TimeoutError`.\n */\nfunction extractBedrockErrorCode(error: unknown): string | undefined {\n\tif (!(error instanceof Error) || !error.name.endsWith(\"Exception\")) return undefined;\n\treturn normalizeDiagnosticValue(error.name);\n}\n\n/**\n * Structured metadata alongside `errorMessage`, which stays byte-identical because `isRetryableAssistantError`\n * matches against it. Unknown fields are omitted, never guessed: a modeled mid-stream exception reaches us as\n * a bare object literal, leaving only `fallbackRequestId`. `details` only, as the throw is not always `Error`.\n */\nfunction appendBedrockFailureDiagnostic(\n\toutput: AssistantMessage,\n\terror: unknown,\n\tfallbackRequestId: string | undefined,\n): void {\n\tconst metadata = (error as SdkErrorMetadata)?.$metadata;\n\tconst details: Record<string, unknown> = {};\n\n\tif (typeof metadata?.httpStatusCode === \"number\") details.status = metadata.httpStatusCode;\n\n\tconst errorCode = extractBedrockErrorCode(error);\n\tif (errorCode !== undefined) details.errorCode = errorCode;\n\n\tconst requestId = normalizeDiagnosticValue(metadata?.requestId) ?? fallbackRequestId;\n\tif (requestId !== undefined) details.requestId = requestId;\n\n\tif (Object.keys(details).length === 0) return;\n\n\tappendAssistantMessageDiagnostic(output, { type: \"bedrock_response_failure\", timestamp: Date.now(), details });\n}\n\n/**\n * Header keys that must never be overwritten by caller-supplied headers.\n * `host` and `x-amz-*` participate in the SigV4 canonical request; `authorization`\n * is owned by SigV4 or the bearer-token path (config.token + authSchemePreference).\n * Compared case-insensitively (caller key is lower-cased before lookup).\n */\nconst RESERVED_HEADER_EXACT = new Set([\"authorization\", \"host\"]);\n\nfunction isReservedHeader(key: string): boolean {\n\tconst lower = key.toLowerCase();\n\treturn lower.startsWith(\"x-amz-\") || RESERVED_HEADER_EXACT.has(lower);\n}\n\n/**\n * Attach caller-supplied headers to the outgoing Bedrock request via a Smithy\n * `build`-step middleware. The `build` step runs after request serialisation but\n * before SigV4 signing, so injected headers are covered by the signature. Reserved\n * SigV4 / auth headers (`x-amz-*`, `authorization`, `host`) are silently skipped;\n * all other caller headers override any existing same-named header on the request.\n */\nfunction addCustomHeadersMiddleware(client: BedrockRuntimeClient, headers: Record<string, string>): void {\n\tconst middleware: BuildMiddleware<object, MetadataBearer> = (next) => async (args) => {\n\t\tconst request = args.request;\n\t\tif (request && typeof request === \"object\" && \"headers\" in request) {\n\t\t\tconst requestHeaders = (request as { headers: Record<string, string> }).headers;\n\t\t\tfor (const [key, value] of Object.entries(headers)) {\n\t\t\t\tif (!isReservedHeader(key)) {\n\t\t\t\t\trequestHeaders[key] = value;\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t\treturn next(args);\n\t};\n\tclient.middlewareStack.add(middleware, { step: \"build\", name: \"pi-ai-custom-headers\", priority: \"low\" });\n}\n\nfunction isSmithyHttpResponse(response: unknown): response is HttpResponse {\n\tif (!response || typeof response !== \"object\") return false;\n\tconst candidate = response as Partial<HttpResponse>;\n\treturn typeof candidate.statusCode === \"number\" && !!candidate.headers && typeof candidate.headers === \"object\";\n}\n\nfunction toProviderResponse(response: unknown): ProviderResponse | undefined {\n\tif (!isSmithyHttpResponse(response)) return undefined;\n\treturn { status: response.statusCode, headers: { ...response.headers } };\n}\n\n/**\n * Bedrock's modeled `$metadata` only preserves selected HTTP metadata (for example\n * requestId), so custom gateway headers are otherwise lost before callers see\n * `onResponse`. Capture the raw Smithy HTTP response at the deserialize step,\n * after the SDK receives the response but before the event stream is consumed.\n */\nfunction addResponseHeadersMiddleware(\n\tclient: BedrockRuntimeClient,\n\tonResponse: NonNullable<BedrockOptions[\"onResponse\"]>,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tonObserved: () => void,\n): void {\n\tconst middleware: DeserializeMiddleware<object, MetadataBearer> = (next) => async (args) => {\n\t\tconst result = await next(args);\n\t\tconst providerResponse = toProviderResponse(result.response);\n\t\tif (providerResponse) {\n\t\t\tonObserved();\n\t\t\tawait onResponse(providerResponse, model);\n\t\t}\n\t\treturn result;\n\t};\n\tclient.middlewareStack.add(middleware, { step: \"deserialize\", name: \"pi-ai-response-headers\" });\n}\n\nexport const streamSimple: StreamFunction<\"bedrock-converse-stream\", SimpleStreamOptions> = (\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcontext: Context,\n\toptions?: SimpleStreamOptions,\n): AssistantMessageEventStream => {\n\tconst base = {\n\t\t...buildBaseOptions(model, context, options, undefined),\n\t\ttoolChoice: options?.toolChoice,\n\t} satisfies BedrockOptions;\n\tif (!options?.reasoning) {\n\t\treturn stream(model, context, { ...base, reasoning: undefined } satisfies BedrockOptions);\n\t}\n\n\tif (isAnthropicClaudeModel(model)) {\n\t\tif (supportsAdaptiveThinking(model.id, model.name)) {\n\t\t\treturn stream(model, context, {\n\t\t\t\t...base,\n\t\t\t\treasoning: options.reasoning,\n\t\t\t\tthinkingBudgets: options.thinkingBudgets,\n\t\t\t} satisfies BedrockOptions);\n\t\t}\n\n\t\t// Undefined means the caller did not request an output cap; let the helper use the model cap.\n\t\t// Do not coerce to 0 here, or the thinking budget would become the entire maxTokens value.\n\t\tconst adjusted = adjustMaxTokensForThinking(\n\t\t\tbase.maxTokens,\n\t\t\tmodel.maxTokens,\n\t\t\toptions.reasoning,\n\t\t\toptions.thinkingBudgets,\n\t\t);\n\n\t\tconst maxTokens = clampMaxTokensToContext(model, context, adjusted.maxTokens);\n\n\t\treturn stream(model, context, {\n\t\t\t...base,\n\t\t\tmaxTokens,\n\t\t\treasoning: options.reasoning,\n\t\t\tthinkingBudgets: {\n\t\t\t\t...(options.thinkingBudgets || {}),\n\t\t\t\t[clampReasoning(options.reasoning)!]: Math.min(adjusted.thinkingBudget, Math.max(0, maxTokens - 1024)),\n\t\t\t},\n\t\t} satisfies BedrockOptions);\n\t}\n\n\treturn stream(model, context, {\n\t\t...base,\n\t\treasoning: options.reasoning,\n\t\tthinkingBudgets: options.thinkingBudgets,\n\t} satisfies BedrockOptions);\n};\n\nfunction handleContentBlockStart(\n\tevent: ContentBlockStartEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst index = event.contentBlockIndex!;\n\tconst start = event.start;\n\n\tif (start?.toolUse) {\n\t\tconst block: Block = {\n\t\t\ttype: \"toolCall\",\n\t\t\tid: start.toolUse.toolUseId || \"\",\n\t\t\tname: start.toolUse.name || \"\",\n\t\t\targuments: {},\n\t\t\tpartialJson: \"\",\n\t\t\tindex,\n\t\t};\n\t\toutput.content.push(block);\n\t\tstream.push({ type: \"toolcall_start\", contentIndex: blocks.length - 1, partial: output });\n\t}\n}\n\nfunction handleContentBlockDelta(\n\tevent: ContentBlockDeltaEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst contentBlockIndex = event.contentBlockIndex!;\n\tconst delta = event.delta;\n\tlet index = blocks.findIndex((b) => b.index === contentBlockIndex);\n\tlet block = blocks[index];\n\n\tif (delta?.text !== undefined) {\n\t\t// If no text block exists yet, create one, as `handleContentBlockStart` is not sent for text blocks\n\t\tif (!block) {\n\t\t\tconst newBlock: Block = { type: \"text\", text: \"\", index: contentBlockIndex };\n\t\t\toutput.content.push(newBlock);\n\t\t\tindex = blocks.length - 1;\n\t\t\tblock = blocks[index];\n\t\t\tstream.push({ type: \"text_start\", contentIndex: index, partial: output });\n\t\t}\n\t\tif (block.type === \"text\") {\n\t\t\tblock.text += delta.text;\n\t\t\tstream.push({ type: \"text_delta\", contentIndex: index, delta: delta.text, partial: output });\n\t\t}\n\t} else if (delta?.toolUse && block?.type === \"toolCall\") {\n\t\tblock.partialJson = (block.partialJson || \"\") + (delta.toolUse.input || \"\");\n\t\tblock.arguments = parseStreamingJson(block.partialJson);\n\t\tstream.push({ type: \"toolcall_delta\", contentIndex: index, delta: delta.toolUse.input || \"\", partial: output });\n\t} else if (delta?.reasoningContent) {\n\t\tlet thinkingBlock = block;\n\t\tlet thinkingIndex = index;\n\n\t\tif (!thinkingBlock) {\n\t\t\tconst newBlock: Block = { type: \"thinking\", thinking: \"\", thinkingSignature: \"\", index: contentBlockIndex };\n\t\t\toutput.content.push(newBlock);\n\t\t\tthinkingIndex = blocks.length - 1;\n\t\t\tthinkingBlock = blocks[thinkingIndex];\n\t\t\tstream.push({ type: \"thinking_start\", contentIndex: thinkingIndex, partial: output });\n\t\t}\n\n\t\tif (thinkingBlock?.type === \"thinking\") {\n\t\t\tif (delta.reasoningContent.text) {\n\t\t\t\tthinkingBlock.thinking += delta.reasoningContent.text;\n\t\t\t\tstream.push({\n\t\t\t\t\ttype: \"thinking_delta\",\n\t\t\t\t\tcontentIndex: thinkingIndex,\n\t\t\t\t\tdelta: delta.reasoningContent.text,\n\t\t\t\t\tpartial: output,\n\t\t\t\t});\n\t\t\t}\n\t\t\t// `thinkingSignature` holds either an Anthropic signature or an opaque redacted\n\t\t\t// payload, never both: mixing them would corrupt whichever arrived first.\n\t\t\tif (delta.reasoningContent.signature && !thinkingBlock.redacted) {\n\t\t\t\tthinkingBlock.thinkingSignature =\n\t\t\t\t\t(thinkingBlock.thinkingSignature || \"\") + delta.reasoningContent.signature;\n\t\t\t}\n\t\t\tif (delta.reasoningContent.redactedContent?.length) {\n\t\t\t\t// Encrypted reasoning from non-Anthropic models on Bedrock (e.g. OpenAI GPT-5.6).\n\t\t\t\t// The payload is opaque, so keep it verbatim in `thinkingSignature` the way the\n\t\t\t\t// Anthropic path stores redacted thinking, and replay it on the next turn.\n\t\t\t\tif (!thinkingBlock.redacted) {\n\t\t\t\t\tthinkingBlock.redacted = true;\n\t\t\t\t\tthinkingBlock.thinkingSignature = \"\";\n\t\t\t\t\tthinkingBlock.thinking += REDACTED_THINKING_PLACEHOLDER;\n\t\t\t\t\tstream.push({\n\t\t\t\t\t\ttype: \"thinking_delta\",\n\t\t\t\t\t\tcontentIndex: thinkingIndex,\n\t\t\t\t\t\tdelta: REDACTED_THINKING_PLACEHOLDER,\n\t\t\t\t\t\tpartial: output,\n\t\t\t\t\t});\n\t\t\t\t}\n\t\t\t\tthinkingBlock.redactedChunks ??= [];\n\t\t\t\tthinkingBlock.redactedChunks.push(delta.reasoningContent.redactedContent);\n\t\t\t}\n\t\t}\n\t}\n}\n\n/**\n * Encodes buffered encrypted reasoning into `thinkingSignature` and drops the scratch\n * buffer, which must never reach a persisted message: `Uint8Array` serializes to an\n * index-keyed object roughly ten times the size of the base64 payload.\n */\nfunction flushRedactedContent(block: Block): void {\n\tif (block.type !== \"thinking\" || !block.redactedChunks) return;\n\tblock.thinkingSignature = bytesToBase64(block.redactedChunks);\n\tdelete block.redactedChunks;\n}\n\n/**\n * Strips every streaming scratch field. Runs from the terminal paths as well as\n * `contentBlockStop`, because a stream can settle without stopping each block.\n */\nfunction finalizeStreamingBlock(block: Block): void {\n\tdelete block.index;\n\t// partialJson is only a streaming scratch buffer; never persist it.\n\tdelete block.partialJson;\n\tflushRedactedContent(block);\n}\n\nfunction handleMetadata(\n\tevent: ConverseStreamMetadataEvent,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\toutput: AssistantMessage,\n): void {\n\tif (event.usage) {\n\t\toutput.usage.input = event.usage.inputTokens || 0;\n\t\toutput.usage.output = event.usage.outputTokens || 0;\n\t\toutput.usage.cacheRead = event.usage.cacheReadInputTokens || 0;\n\t\toutput.usage.cacheWrite = event.usage.cacheWriteInputTokens || 0;\n\t\toutput.usage.totalTokens = event.usage.totalTokens || output.usage.input + output.usage.output;\n\t\tcalculateCost(model, output.usage);\n\t}\n}\n\nfunction handleContentBlockStop(\n\tevent: ContentBlockStopEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst index = blocks.findIndex((b) => b.index === event.contentBlockIndex);\n\tconst block = blocks[index];\n\tif (!block) return;\n\tdelete (block as Block).index;\n\n\tswitch (block.type) {\n\t\tcase \"text\":\n\t\t\tstream.push({ type: \"text_end\", contentIndex: index, content: block.text, partial: output });\n\t\t\tbreak;\n\t\tcase \"thinking\":\n\t\t\tflushRedactedContent(block);\n\t\t\tstream.push({ type: \"thinking_end\", contentIndex: index, content: block.thinking, partial: output });\n\t\t\tbreak;\n\t\tcase \"toolCall\":\n\t\t\tblock.arguments = parseStreamingJson(block.partialJson);\n\t\t\t// Finalize in-place and strip the scratch buffer so replay only\n\t\t\t// carries parsed arguments.\n\t\t\tdelete (block as Block).partialJson;\n\t\t\tstream.push({ type: \"toolcall_end\", contentIndex: index, toolCall: block, partial: output });\n\t\t\tbreak;\n\t}\n}\n\n/**\n * Check if the model supports adaptive thinking (Opus 4.6+, Sonnet 4.6).\n * Checks both model ID and model name to support application inference profiles\n * whose ARNs don't contain the model name.\n */\nfunction getModelMatchCandidates(modelId: string, modelName?: string): string[] {\n\tconst values = modelName ? [modelId, modelName] : [modelId];\n\treturn values.flatMap((value) => {\n\t\tconst lower = value.toLowerCase();\n\t\treturn [lower, lower.replace(/[\\s_.:]+/g, \"-\")];\n\t});\n}\n\nfunction supportsAdaptiveThinking(modelId: string, modelName?: string): boolean {\n\tconst candidates = getModelMatchCandidates(modelId, modelName);\n\treturn candidates.some(\n\t\t(s) =>\n\t\t\ts.includes(\"opus-4-6\") ||\n\t\t\ts.includes(\"opus-4-7\") ||\n\t\t\ts.includes(\"opus-4-8\") ||\n\t\t\ts.includes(\"opus-5\") ||\n\t\t\ts.includes(\"sonnet-4-6\") ||\n\t\t\ts.includes(\"sonnet-5\") ||\n\t\t\ts.includes(\"fable-5\"),\n\t);\n}\n\nfunction supportsNativeXhighEffort(model: Model<\"bedrock-converse-stream\">): boolean {\n\tconst candidates = getModelMatchCandidates(model.id, model.name);\n\treturn candidates.some(\n\t\t(s) =>\n\t\t\ts.includes(\"opus-4-7\") ||\n\t\t\ts.includes(\"opus-4-8\") ||\n\t\t\ts.includes(\"opus-5\") ||\n\t\t\ts.includes(\"sonnet-5\") ||\n\t\t\ts.includes(\"fable-5\"),\n\t);\n}\n\nfunction mapThinkingLevelToEffort(\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tlevel: SimpleStreamOptions[\"reasoning\"],\n): \"low\" | \"medium\" | \"high\" | \"xhigh\" | \"max\" {\n\tif (level === \"xhigh\" && supportsNativeXhighEffort(model)) return \"xhigh\";\n\n\tconst mapped = level ? model.thinkingLevelMap?.[level] : undefined;\n\tif (typeof mapped === \"string\") return mapped as \"low\" | \"medium\" | \"high\" | \"xhigh\" | \"max\";\n\n\tswitch (level) {\n\t\tcase \"minimal\":\n\t\tcase \"low\":\n\t\t\treturn \"low\";\n\t\tcase \"medium\":\n\t\t\treturn \"medium\";\n\t\tcase \"high\":\n\t\t\treturn \"high\";\n\t\tdefault:\n\t\t\treturn \"high\";\n\t}\n}\n\n/**\n * Resolve cache retention preference.\n * Defaults to \"short\" and uses PI_CACHE_RETENTION for backward compatibility.\n */\nfunction resolveCacheRetention(cacheRetention?: CacheRetention, env?: ProviderEnv): CacheRetention {\n\tif (cacheRetention) {\n\t\treturn cacheRetention;\n\t}\n\tif (getProviderEnvValue(\"PI_CACHE_RETENTION\", env) === \"long\") {\n\t\treturn \"long\";\n\t}\n\treturn \"short\";\n}\n\n/**\n * Check if the model is an Anthropic Claude model on Bedrock.\n * Checks both model ID and model name to support application inference profiles\n * whose ARNs don't contain the model name.\n */\nfunction isAnthropicClaudeModel(model: Model<\"bedrock-converse-stream\">): boolean {\n\tconst id = model.id.toLowerCase();\n\tconst name = model.name?.toLowerCase() ?? \"\";\n\treturn (\n\t\tid.includes(\"anthropic.claude\") ||\n\t\tid.includes(\"anthropic/claude\") ||\n\t\tname.includes(\"anthropic.claude\") ||\n\t\tname.includes(\"anthropic/claude\") ||\n\t\tname.includes(\"claude\")\n\t);\n}\n\n/**\n * Check if the model supports prompt caching.\n * Supported: Claude 3.5 Haiku, Claude 3.7 Sonnet, Claude 4.x models, Claude 5 models\n *\n * For base models and system-defined inference profiles the model ID / ARN\n * contains the model name, so we can decide locally.\n *\n * For application inference profiles (whose ARNs don't contain the model name),\n * also checks model.name which is user-controlled via models.json or registerProvider.\n * As a last resort, set AWS_BEDROCK_FORCE_CACHE=1 to enable cache points.\n * Amazon Nova models have automatic caching and don't need explicit cache points.\n */\nfunction supportsPromptCaching(model: Model<\"bedrock-converse-stream\">, env?: ProviderEnv): boolean {\n\tconst candidates = getModelMatchCandidates(model.id, model.name);\n\n\tconst hasClaudeRef = candidates.some((s) => s.includes(\"claude\"));\n\tif (!hasClaudeRef) {\n\t\t// Application inference profiles don't contain the model name in the ARN.\n\t\t// Allow users to force cache points via environment variable.\n\t\tif (getProviderEnvValue(\"AWS_BEDROCK_FORCE_CACHE\", env) === \"1\") return true;\n\t\treturn false;\n\t}\n\t// Claude 5 models (fable-5, opus-5, sonnet-5)\n\tif (candidates.some((s) => s.includes(\"fable-5\") || s.includes(\"opus-5\") || s.includes(\"sonnet-5\"))) return true;\n\t// Claude 4.x models (opus-4, sonnet-4, haiku-4)\n\tif (candidates.some((s) => s.includes(\"-4-\"))) return true;\n\t// Claude 3.7 Sonnet\n\tif (candidates.some((s) => s.includes(\"claude-3-7-sonnet\"))) return true;\n\t// Claude 3.5 Haiku\n\tif (candidates.some((s) => s.includes(\"claude-3-5-haiku\"))) return true;\n\treturn false;\n}\n\n/**\n * Check if the model supports thinking signatures in reasoningContent.\n * Only Anthropic Claude models support the signature field.\n * Other models (OpenAI, Qwen, Minimax, Moonshot, etc.) reject it with:\n * \"This model doesn't support the reasoningContent.reasoningText.signature field\"\n *\n * Checks both model ID and model name to support application inference profiles.\n */\nfunction supportsThinkingSignature(model: Model<\"bedrock-converse-stream\">): boolean {\n\treturn isAnthropicClaudeModel(model);\n}\n\nfunction buildSystemPrompt(\n\tsystemPrompt: string | undefined,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcacheRetention: CacheRetention,\n\tenv?: ProviderEnv,\n): SystemContentBlock[] | undefined {\n\tif (!systemPrompt) return undefined;\n\n\tconst blocks: SystemContentBlock[] = [{ text: sanitizeSurrogates(systemPrompt) }];\n\n\t// Add cache point for supported Claude models when caching is enabled\n\tif (cacheRetention !== \"none\" && supportsPromptCaching(model, env)) {\n\t\tblocks.push({\n\t\t\tcachePoint: { type: CachePointType.DEFAULT, ...(cacheRetention === \"long\" ? { ttl: CacheTTL.ONE_HOUR } : {}) },\n\t\t});\n\t}\n\n\treturn blocks;\n}\n\nfunction normalizeToolCallId(id: string): string {\n\tconst sanitized = id.replace(/[^a-zA-Z0-9_-]/g, \"_\");\n\treturn sanitized.length > 64 ? sanitized.slice(0, 64) : sanitized;\n}\n\nfunction createNonBlankTextBlock(text: string): ContentBlock.TextMember | undefined {\n\tconst sanitized = sanitizeSurrogates(text);\n\treturn sanitized.trim().length === 0 ? undefined : { text: sanitized };\n}\n\nfunction createRequiredTextBlock(text: string): ContentBlock.TextMember {\n\treturn createNonBlankTextBlock(text) ?? { text: EMPTY_TEXT_PLACEHOLDER };\n}\n\nfunction sanitizeBedrockDocument(value: DocumentType): DocumentType {\n\tif (Array.isArray(value)) {\n\t\treturn value.map(sanitizeBedrockDocument);\n\t}\n\tif (value !== null && typeof value === \"object\") {\n\t\treturn Object.fromEntries(\n\t\t\tObject.entries(value)\n\t\t\t\t.filter(([key]) => key.length > 0)\n\t\t\t\t.map(([key, nestedValue]) => [key, sanitizeBedrockDocument(nestedValue)]),\n\t\t);\n\t}\n\treturn value;\n}\n\nfunction convertToolResultContent(content: (TextContent | ImageContent)[]): ToolResultContentBlock[] {\n\tconst result: ToolResultContentBlock[] = [];\n\tfor (const c of content) {\n\t\tif (c.type === \"image\") {\n\t\t\tresult.push({ image: createImageBlock(c.mimeType, c.data) });\n\t\t} else {\n\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\tif (textBlock) result.push(textBlock);\n\t\t}\n\t}\n\tif (result.length === 0) result.push({ text: EMPTY_TEXT_PLACEHOLDER });\n\treturn result;\n}\n\nfunction convertMessages(\n\tcontext: Context,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcacheRetention: CacheRetention,\n\tenv?: ProviderEnv,\n): Message[] {\n\tconst result: Message[] = [];\n\tconst transformedMessages = transformMessages(context.messages, model, normalizeToolCallId);\n\n\tfor (let i = 0; i < transformedMessages.length; i++) {\n\t\tconst m = transformedMessages[i];\n\n\t\tswitch (m.role) {\n\t\t\tcase \"user\": {\n\t\t\t\tconst content: ContentBlock[] = [];\n\t\t\t\tif (typeof m.content === \"string\") {\n\t\t\t\t\tcontent.push(createRequiredTextBlock(m.content));\n\t\t\t\t} else {\n\t\t\t\t\tfor (const c of m.content) {\n\t\t\t\t\t\tswitch (c.type) {\n\t\t\t\t\t\t\tcase \"text\": {\n\t\t\t\t\t\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\t\t\t\t\t\tif (textBlock) content.push(textBlock);\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\tcase \"image\":\n\t\t\t\t\t\t\t\tcontent.push({ image: createImageBlock(c.mimeType, c.data) });\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\tcase \"document\":\n\t\t\t\t\t\t\t\tcontent.push({ document: createDocumentBlock(c, content.length) });\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\tdefault:\n\t\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t\t}\n\t\t\t\t\t}\n\t\t\t\t\tif (content.length === 0) content.push({ text: EMPTY_TEXT_PLACEHOLDER });\n\t\t\t\t}\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.USER,\n\t\t\t\t\tcontent,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tcase \"assistant\": {\n\t\t\t\t// Skip assistant messages with empty content (e.g., from aborted requests)\n\t\t\t\t// Bedrock rejects messages with empty content arrays\n\t\t\t\tif (m.content.length === 0) {\n\t\t\t\t\tcontinue;\n\t\t\t\t}\n\t\t\t\tconst contentBlocks: ContentBlock[] = [];\n\t\t\t\tfor (const c of m.content) {\n\t\t\t\t\tswitch (c.type) {\n\t\t\t\t\t\tcase \"text\": {\n\t\t\t\t\t\t\t// Skip empty text blocks\n\t\t\t\t\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\t\t\t\t\tif (!textBlock) continue;\n\t\t\t\t\t\t\tcontentBlocks.push(textBlock);\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t}\n\t\t\t\t\t\tcase \"toolCall\":\n\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\ttoolUse: { toolUseId: c.id, name: c.name, input: sanitizeBedrockDocument(c.arguments) },\n\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\tcase \"thinking\": {\n\t\t\t\t\t\t\t// Encrypted reasoning is opaque: replay the stored payload as the\n\t\t\t\t\t\t\t// `redactedContent` member instead of lowering it to reasoning text.\n\t\t\t\t\t\t\tif (c.redacted) {\n\t\t\t\t\t\t\t\tconst redactedContent = decodeRedactedContent(c.thinkingSignature);\n\t\t\t\t\t\t\t\tif (redactedContent?.length) {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({ reasoningContent: { redactedContent } });\n\t\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t// Skip empty thinking blocks\n\t\t\t\t\t\t\tconst thinking = sanitizeSurrogates(c.thinking);\n\t\t\t\t\t\t\tif (thinking.trim().length === 0) continue;\n\t\t\t\t\t\t\t// Only Anthropic models support the signature field in reasoningText.\n\t\t\t\t\t\t\t// For other models, we omit the signature to avoid errors like:\n\t\t\t\t\t\t\t// \"This model doesn't support the reasoningContent.reasoningText.signature field\"\n\t\t\t\t\t\t\tif (supportsThinkingSignature(model)) {\n\t\t\t\t\t\t\t\t// Signatures arrive after thinking deltas. If a partial or externally\n\t\t\t\t\t\t\t\t// persisted message lacks a signature, Bedrock rejects the replayed\n\t\t\t\t\t\t\t\t// reasoning block. Fall back to plain text, matching Anthropic.\n\t\t\t\t\t\t\t\tif (!c.thinkingSignature || c.thinkingSignature.trim().length === 0) {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({ text: thinking });\n\t\t\t\t\t\t\t\t} else {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\t\t\treasoningContent: {\n\t\t\t\t\t\t\t\t\t\t\treasoningText: {\n\t\t\t\t\t\t\t\t\t\t\t\ttext: thinking,\n\t\t\t\t\t\t\t\t\t\t\t\tsignature: c.thinkingSignature,\n\t\t\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t} else {\n\t\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\t\treasoningContent: {\n\t\t\t\t\t\t\t\t\t\treasoningText: { text: thinking },\n\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t}\n\t\t\t\t\t\tdefault:\n\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t}\n\t\t\t\t}\n\t\t\t\t// Skip if all content blocks were filtered out\n\t\t\t\tif (contentBlocks.length === 0) {\n\t\t\t\t\tcontinue;\n\t\t\t\t}\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.ASSISTANT,\n\t\t\t\t\tcontent: contentBlocks,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tcase \"toolResult\": {\n\t\t\t\t// Collect all consecutive toolResult messages into a single user message\n\t\t\t\t// Bedrock requires all tool results to be in one message\n\t\t\t\tconst toolResults: ContentBlock.ToolResultMember[] = [];\n\n\t\t\t\t// Add current tool result with all content blocks combined\n\t\t\t\ttoolResults.push({\n\t\t\t\t\ttoolResult: {\n\t\t\t\t\t\ttoolUseId: m.toolCallId,\n\t\t\t\t\t\tcontent: convertToolResultContent(m.content),\n\t\t\t\t\t\tstatus: m.isError ? ToolResultStatus.ERROR : ToolResultStatus.SUCCESS,\n\t\t\t\t\t},\n\t\t\t\t});\n\n\t\t\t\t// Look ahead for consecutive toolResult messages\n\t\t\t\tlet j = i + 1;\n\t\t\t\twhile (j < transformedMessages.length && transformedMessages[j].role === \"toolResult\") {\n\t\t\t\t\tconst nextMsg = transformedMessages[j] as ToolResultMessage;\n\t\t\t\t\ttoolResults.push({\n\t\t\t\t\t\ttoolResult: {\n\t\t\t\t\t\t\ttoolUseId: nextMsg.toolCallId,\n\t\t\t\t\t\t\tcontent: convertToolResultContent(nextMsg.content),\n\t\t\t\t\t\t\tstatus: nextMsg.isError ? ToolResultStatus.ERROR : ToolResultStatus.SUCCESS,\n\t\t\t\t\t\t},\n\t\t\t\t\t});\n\t\t\t\t\tj++;\n\t\t\t\t}\n\n\t\t\t\t// Skip the messages we've already processed\n\t\t\t\ti = j - 1;\n\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.USER,\n\t\t\t\t\tcontent: toolResults,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tdefault:\n\t\t\t\tcontinue;\n\t\t}\n\t}\n\n\t// Add cache point to the last user message for supported Claude models when caching is enabled\n\tif (cacheRetention !== \"none\" && supportsPromptCaching(model, env) && result.length > 0) {\n\t\tconst lastMessage = result[result.length - 1];\n\t\tif (lastMessage.role === ConversationRole.USER && lastMessage.content) {\n\t\t\t(lastMessage.content as ContentBlock[]).push({\n\t\t\t\tcachePoint: {\n\t\t\t\t\ttype: CachePointType.DEFAULT,\n\t\t\t\t\t...(cacheRetention === \"long\" ? { ttl: CacheTTL.ONE_HOUR } : {}),\n\t\t\t\t},\n\t\t\t});\n\t\t}\n\t}\n\n\treturn result;\n}\n\nfunction convertToolConfig(\n\ttools: Tool[] | undefined,\n\ttoolChoice: BedrockOptions[\"toolChoice\"],\n\tsupportsStrictMode: boolean,\n\tmodel: Model<\"bedrock-converse-stream\">,\n): ToolConfiguration | undefined {\n\t// Validate the requested choice before any early return. A forced choice on a model that\n\t// rejects it is an error regardless of whether tools were supplied: returning `undefined` for\n\t// an empty or absent tool list would discard the caller's instruction silently, which is the\n\t// behavior this guard exists to prevent.\n\t//\n\t// Claude Fable 5.1 rejects forced tool use on every request with a 400, whichever platform\n\t// serves it. Reject rather than rewriting to `auto`, which would discard an explicit caller\n\t// instruction. Mirrors the guards in `anthropic-messages.ts` and `openai-completions.ts`.\n\t// https://platform.claude.com/docs/en/build-with-claude/thinking\n\tif (model.compat?.supportsForcedToolChoice === false) {\n\t\tconst isForced = toolChoice === \"any\" || (typeof toolChoice === \"object\" && toolChoice.type === \"tool\");\n\t\tif (isForced) {\n\t\t\tconst requestedLabel = typeof toolChoice === \"string\" ? toolChoice : `tool \"${toolChoice.name}\"`;\n\t\t\tthrow new Error(\n\t\t\t\t`Model ${model.id} does not support forced tool choice (requested: ${requestedLabel}). ` +\n\t\t\t\t\t`Use toolChoice \"auto\" with strict tool use or structured outputs instead.`,\n\t\t\t);\n\t\t}\n\t}\n\n\tif (!tools?.length) return undefined;\n\t// `none` is never a forced choice, so this return can never skip a rejection above.\n\tif (toolChoice === \"none\") return undefined;\n\n\tconst bedrockTools: BedrockTool[] = tools.map((tool) => {\n\t\tconst strict = resolveJsonSchemaStrictSampling(tool, supportsStrictMode);\n\t\treturn {\n\t\t\ttoolSpec: {\n\t\t\t\tname: tool.name,\n\t\t\t\tdescription: tool.description,\n\t\t\t\tinputSchema: { json: getJsonSchemaToolParameters(tool, strict) as unknown as DocumentType },\n\t\t\t\t...(strict === true ? { strict: true } : {}),\n\t\t\t},\n\t\t};\n\t});\n\n\tlet bedrockToolChoice: ToolChoice | undefined;\n\tswitch (toolChoice) {\n\t\tcase \"auto\":\n\t\t\tbedrockToolChoice = { auto: {} };\n\t\t\tbreak;\n\t\tcase \"any\":\n\t\t\tbedrockToolChoice = { any: {} };\n\t\t\tbreak;\n\t\tdefault:\n\t\t\tif (toolChoice?.type === \"tool\") {\n\t\t\t\tbedrockToolChoice = { tool: { name: toolChoice.name } };\n\t\t\t}\n\t}\n\n\treturn { tools: bedrockTools, toolChoice: bedrockToolChoice };\n}\n\nfunction mapStopReason(reason: string | undefined): { stopReason: StopReason; errorMessage?: string } {\n\tswitch (reason) {\n\t\tcase BedrockStopReason.END_TURN:\n\t\tcase BedrockStopReason.STOP_SEQUENCE:\n\t\t\treturn { stopReason: \"stop\" };\n\t\tcase BedrockStopReason.MAX_TOKENS:\n\t\tcase BedrockStopReason.MODEL_CONTEXT_WINDOW_EXCEEDED:\n\t\t\treturn { stopReason: \"length\" };\n\t\tcase BedrockStopReason.TOOL_USE:\n\t\t\treturn { stopReason: \"toolUse\" };\n\t\tdefault:\n\t\t\treturn reason\n\t\t\t\t? { stopReason: \"error\", errorMessage: `Provider stopped with: ${reason}` }\n\t\t\t\t: { stopReason: \"error\" };\n\t}\n}\n\nfunction getConfiguredBedrockRegion(options: BedrockOptions): string | undefined {\n\treturn (\n\t\toptions.region ||\n\t\tgetProviderEnvValue(\"AWS_REGION\", options.env) ||\n\t\tgetProviderEnvValue(\"AWS_DEFAULT_REGION\", options.env) ||\n\t\tundefined\n\t);\n}\n\nfunction getConfiguredBedrockCredentials(env?: ProviderEnv): BedrockRuntimeClientConfig[\"credentials\"] | undefined {\n\tconst accessKeyId = getProviderEnvValue(\"AWS_ACCESS_KEY_ID\", env);\n\tconst secretAccessKey = getProviderEnvValue(\"AWS_SECRET_ACCESS_KEY\", env);\n\tif (!accessKeyId || !secretAccessKey) {\n\t\treturn undefined;\n\t}\n\tconst sessionToken = getProviderEnvValue(\"AWS_SESSION_TOKEN\", env);\n\treturn {\n\t\taccessKeyId,\n\t\tsecretAccessKey,\n\t\t...(sessionToken ? { sessionToken } : {}),\n\t};\n}\n\nfunction getStandardBedrockEndpointRegion(baseUrl: string | undefined): string | undefined {\n\tif (!baseUrl) {\n\t\treturn undefined;\n\t}\n\n\ttry {\n\t\tconst { hostname } = new URL(baseUrl);\n\t\tconst match = hostname.toLowerCase().match(/^bedrock-runtime(?:-fips)?\\.([a-z0-9-]+)\\.amazonaws\\.com(?:\\.cn)?$/);\n\t\treturn match?.[1];\n\t} catch {\n\t\treturn undefined;\n\t}\n}\n\nfunction shouldUseExplicitBedrockEndpoint(\n\tbaseUrl: string,\n\tconfiguredRegion: string | undefined,\n\thasAmbientConfiguredProfile: boolean,\n): boolean {\n\tconst endpointRegion = getStandardBedrockEndpointRegion(baseUrl);\n\tif (!endpointRegion) {\n\t\treturn true;\n\t}\n\n\treturn !configuredRegion && !hasAmbientConfiguredProfile;\n}\n\nfunction isGovCloudBedrockTarget(model: Model<\"bedrock-converse-stream\">, options: BedrockOptions): boolean {\n\tconst region = getConfiguredBedrockRegion(options);\n\tif (region?.toLowerCase().startsWith(\"us-gov-\")) {\n\t\treturn true;\n\t}\n\n\tconst modelId = model.id.toLowerCase();\n\treturn modelId.startsWith(\"us-gov.\") || modelId.startsWith(\"arn:aws-us-gov:\");\n}\n\nfunction isOpenAiGpt6AstraBedrockModel(model: Model<\"bedrock-converse-stream\">): boolean {\n\treturn /^(?:global\\.|us\\.)?openai\\.gpt-6-astra$/.test(model.id);\n}\n\nfunction buildAdditionalModelRequestFields(\n\tmodel: Model<\"bedrock-converse-stream\">,\n\toptions: BedrockOptions,\n): Record<string, any> | undefined {\n\tif (!options.reasoning || !model.reasoning) {\n\t\treturn undefined;\n\t}\n\n\tif (isAnthropicClaudeModel(model)) {\n\t\t// GovCloud Bedrock currently rejects the Claude thinking.display field.\n\t\t// Omit it there until the GovCloud Converse schema catches up.\n\t\tconst display = isGovCloudBedrockTarget(model, options) ? undefined : (options.thinkingDisplay ?? \"summarized\");\n\t\tconst result: Record<string, any> = supportsAdaptiveThinking(model.id, model.name)\n\t\t\t? {\n\t\t\t\t\tthinking: { type: \"adaptive\", ...(display !== undefined ? { display } : {}) },\n\t\t\t\t\toutput_config: { effort: mapThinkingLevelToEffort(model, options.reasoning) },\n\t\t\t\t}\n\t\t\t: (() => {\n\t\t\t\t\tconst defaultBudgets: Record<ThinkingLevel, number> = {\n\t\t\t\t\t\tminimal: 1024,\n\t\t\t\t\t\tlow: 2048,\n\t\t\t\t\t\tmedium: 8192,\n\t\t\t\t\t\thigh: 16384,\n\t\t\t\t\t\txhigh: 16384, // Budget-based Claude clamps extended levels to high\n\t\t\t\t\t\tmax: 16384,\n\t\t\t\t\t};\n\n\t\t\t\t\t// Custom budgets only cover token-based levels through high.\n\t\t\t\t\tconst level = options.reasoning === \"xhigh\" || options.reasoning === \"max\" ? \"high\" : options.reasoning;\n\t\t\t\t\tconst budget = options.thinkingBudgets?.[level] ?? defaultBudgets[options.reasoning];\n\n\t\t\t\t\treturn {\n\t\t\t\t\t\tthinking: {\n\t\t\t\t\t\t\ttype: \"enabled\",\n\t\t\t\t\t\t\tbudget_tokens: budget,\n\t\t\t\t\t\t\t...(display !== undefined ? { display } : {}),\n\t\t\t\t\t\t},\n\t\t\t\t\t};\n\t\t\t\t})();\n\n\t\tif (!supportsAdaptiveThinking(model.id, model.name) && (options.interleavedThinking ?? true)) {\n\t\t\tresult.anthropic_beta = [\"interleaved-thinking-2025-05-14\"];\n\t\t}\n\n\t\treturn result;\n\t}\n\n\tif (isOpenAiGpt6AstraBedrockModel(model)) {\n\t\t// Bedrock maps OpenAI Chat Completions fields that are not part of Converse's\n\t\t// inferenceConfig into additionalModelRequestFields. GPT-6-Astra uses the Chat\n\t\t// Completions spelling rather than the Responses API's nested reasoning object.\n\t\t// https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-openai.html\n\t\treturn { reasoning_effort: mapThinkingLevelToEffort(model, options.reasoning) };\n\t}\n\n\treturn undefined;\n}\n\n/**\n * Build a Bedrock `DocumentBlock`. Three things differ from the Anthropic path:\n *\n * - `name` is **required**, and AWS restricts it to alphanumerics, single whitespace runs,\n * hyphens, parentheses, and square brackets — then warns the field \"is vulnerable to prompt\n * injections… we recommend that you specify a neutral name.\" Any caller-supplied name is\n * sanitized to that character set, and anything left empty falls back to a generated label.\n * - `source` takes **raw bytes**, not base64: \"If you use an Amazon Web Services SDK, you don't\n * need to encode the bytes in base64.\" So the base64 payload is decoded here, the opposite of\n * what the Anthropic converter does.\n * - Without citations enabled, Converse degrades to plain text extraction rather than full visual\n * PDF understanding. That limitation is documented rather than worked around here.\n * https://platform.claude.com/docs/en/build-with-claude/pdf-support\n *\n * `DocumentFormat` also has eight non-PDF members, but `format` is hardcoded to PDF because that\n * is the only media type `DocumentContent` declares — and the seven office/markup formats have no\n * Anthropic source variant to map onto. The block's own media type is therefore verified rather\n * than read.\n */\nfunction createDocumentBlock(block: DocumentContent, index: number) {\n\tassertSupportedDocumentMimeType(block);\n\tconst sanitized = (block.name ?? \"\")\n\t\t.replace(/[^a-zA-Z0-9\\s\\-()[\\]]/g, \" \")\n\t\t.replace(/\\s+/g, \" \")\n\t\t.trim();\n\treturn {\n\t\tformat: DocumentFormat.PDF,\n\t\tname: sanitized.length > 0 ? sanitized : `document ${index + 1}`,\n\t\tsource: { bytes: base64ToBytes(block.data) },\n\t};\n}\n\nfunction createImageBlock(mimeType: string, data: string) {\n\tlet format: ImageFormat;\n\tswitch (mimeType) {\n\t\tcase \"image/jpeg\":\n\t\tcase \"image/jpg\":\n\t\t\tformat = ImageFormat.JPEG;\n\t\t\tbreak;\n\t\tcase \"image/png\":\n\t\t\tformat = ImageFormat.PNG;\n\t\t\tbreak;\n\t\tcase \"image/gif\":\n\t\t\tformat = ImageFormat.GIF;\n\t\t\tbreak;\n\t\tcase \"image/webp\":\n\t\t\tformat = ImageFormat.WEBP;\n\t\t\tbreak;\n\t\tdefault:\n\t\t\tthrow new Error(`Unknown image type: ${mimeType}`);\n\t}\n\n\treturn { source: { bytes: base64ToBytes(data) }, format };\n}\n\nfunction base64ToBytes(data: string): Uint8Array {\n\tconst binaryString = atob(data);\n\tconst bytes = new Uint8Array(binaryString.length);\n\tfor (let i = 0; i < binaryString.length; i++) {\n\t\tbytes[i] = binaryString.charCodeAt(i);\n\t}\n\treturn bytes;\n}\n\n/**\n * Decodes a stored redacted payload. The AWS SDK hands the blob over as bytes, but a\n * persisted session carries it as base64. A hand-edited or externally produced session\n * can hold a signature that is not base64; drop that block instead of failing the\n * whole request.\n */\nfunction decodeRedactedContent(signature: string | undefined): Uint8Array | undefined {\n\tif (!signature) return undefined;\n\ttry {\n\t\treturn base64ToBytes(signature);\n\t} catch {\n\t\treturn undefined;\n\t}\n}\n\nfunction bytesToBase64(chunks: Uint8Array[]): string {\n\t// Encrypted reasoning runs to tens of KB, so build the binary string in slices\n\t// rather than one concatenation per byte. The window stays under the engine's\n\t// argument-count limit for spread calls.\n\tconst WINDOW = 0x8000;\n\tlet binary = \"\";\n\tfor (const chunk of chunks) {\n\t\tfor (let i = 0; i < chunk.length; i += WINDOW) {\n\t\t\tbinary += String.fromCharCode(...chunk.subarray(i, i + WINDOW));\n\t\t}\n\t}\n\treturn btoa(binary);\n}\n"]}
1
+ {"version":3,"file":"bedrock-converse-stream.d.ts","sourceRoot":"","sources":["../../src/api/bedrock-converse-stream.ts"],"names":[],"mappings":"AA8BA,OAAO,KAAK,EAUX,mBAAmB,EAEnB,cAAc,EACd,aAAa,EAEb,eAAe,EAEf,aAAa,EAIb,MAAM,aAAa,CAAC;AAmBrB,MAAM,MAAM,sBAAsB,GAAG,YAAY,GAAG,SAAS,CAAC;AAE9D,MAAM,WAAW,cAAe,SAAQ,aAAa;IACpD,MAAM,CAAC,EAAE,MAAM,CAAC;IAChB,OAAO,CAAC,EAAE,MAAM,CAAC;IACjB,UAAU,CAAC,EAAE,MAAM,GAAG,KAAK,GAAG,MAAM,GAAG;QAAE,IAAI,EAAE,MAAM,CAAC;QAAC,IAAI,EAAE,MAAM,CAAA;KAAE,CAAC;IAEtE,SAAS,CAAC,EAAE,aAAa,CAAC;IAE1B,eAAe,CAAC,EAAE,eAAe,CAAC;IAElC,mBAAmB,CAAC,EAAE,OAAO,CAAC;IAC9B;;;;;;;;;OASG;IACH,eAAe,CAAC,EAAE,sBAAsB,CAAC;IACzC;;;sGAGkG;IAClG,eAAe,CAAC,EAAE,MAAM,CAAC,MAAM,EAAE,MAAM,CAAC,CAAC;IACzC;;;;yGAIqG;IACrG,WAAW,CAAC,EAAE,MAAM,CAAC;CACrB;AAcD,eAAO,MAAM,MAAM,EAAE,cAAc,CAAC,yBAAyB,EAAE,cAAc,CAwO5E,CAAC;AAwKF,eAAO,MAAM,YAAY,EAAE,cAAc,CAAC,yBAAyB,EAAE,mBAAmB,CAiDvF,CAAC","sourcesContent":["import type { Agent as HttpsAgent } from \"node:https\";\nimport {\n\tBedrockRuntimeClient,\n\ttype BedrockRuntimeClientConfig,\n\tBedrockRuntimeServiceException,\n\tStopReason as BedrockStopReason,\n\ttype Tool as BedrockTool,\n\tCachePointType,\n\tCacheTTL,\n\ttype ContentBlock,\n\ttype ContentBlockDeltaEvent,\n\ttype ContentBlockStartEvent,\n\ttype ContentBlockStopEvent,\n\tConversationRole,\n\tConverseStreamCommand,\n\ttype ConverseStreamMetadataEvent,\n\tDocumentFormat,\n\tImageFormat,\n\ttype Message,\n\ttype SystemContentBlock,\n\ttype ToolChoice,\n\ttype ToolConfiguration,\n\ttype ToolResultContentBlock,\n\tToolResultStatus,\n} from \"@aws-sdk/client-bedrock-runtime\";\nimport { NodeHttpHandler } from \"@smithy/node-http-handler\";\nimport type { BuildMiddleware, DeserializeMiddleware, DocumentType, HttpResponse, MetadataBearer } from \"@smithy/types\";\nimport { HttpProxyAgent } from \"http-proxy-agent\";\nimport { HttpsProxyAgent } from \"https-proxy-agent\";\nimport { calculateCost } from \"../models.ts\";\nimport type {\n\tApi,\n\tAssistantMessage,\n\tCacheRetention,\n\tContext,\n\tDocumentContent,\n\tImageContent,\n\tModel,\n\tProviderEnv,\n\tProviderResponse,\n\tSimpleStreamOptions,\n\tStopReason,\n\tStreamFunction,\n\tStreamOptions,\n\tTextContent,\n\tThinkingBudgets,\n\tThinkingContent,\n\tThinkingLevel,\n\tTool,\n\tToolCall,\n\tToolResultMessage,\n} from \"../types.ts\";\nimport { appendAssistantMessageDiagnostic } from \"../utils/diagnostics.ts\";\nimport { assertSupportedDocumentMimeType } from \"../utils/document-input.ts\";\nimport { normalizeProviderError } from \"../utils/error-body.ts\";\nimport { AssistantMessageEventStream } from \"../utils/event-stream.ts\";\nimport { providerHeadersToRecord } from \"../utils/headers.ts\";\nimport { parseStreamingJson } from \"../utils/json-parse.ts\";\nimport { resolveHttpProxyUrlForTarget } from \"../utils/node-http-proxy.ts\";\nimport { getProviderEnvValue } from \"../utils/provider-env.ts\";\nimport { sanitizeSurrogates } from \"../utils/sanitize-unicode.ts\";\nimport { getJsonSchemaToolParameters, resolveJsonSchemaStrictSampling } from \"./constrained-sampling.ts\";\nimport {\n\tadjustMaxTokensForThinking,\n\tbuildBaseOptions,\n\tclampMaxTokensToContext,\n\tclampReasoning,\n} from \"./simple-options.ts\";\nimport { transformMessages } from \"./transform-messages.ts\";\n\nexport type BedrockThinkingDisplay = \"summarized\" | \"omitted\";\n\nexport interface BedrockOptions extends StreamOptions {\n\tregion?: string;\n\tprofile?: string;\n\ttoolChoice?: \"auto\" | \"any\" | \"none\" | { type: \"tool\"; name: string };\n\t/* See https://docs.aws.amazon.com/bedrock/latest/userguide/inference-reasoning.html for supported models. */\n\treasoning?: ThinkingLevel;\n\t/* Custom token budgets per thinking level. Overrides default budgets. */\n\tthinkingBudgets?: ThinkingBudgets;\n\t/* Only supported by Claude 4.x models, see https://docs.aws.amazon.com/bedrock/latest/userguide/claude-messages-extended-thinking.html#claude-messages-extended-thinking-tool-use-interleaved */\n\tinterleavedThinking?: boolean;\n\t/**\n\t * Controls how Claude's thinking content is returned in responses.\n\t * - \"summarized\": Thinking blocks contain summarized thinking text (default here).\n\t * - \"omitted\": Thinking content is redacted but the signature still travels back\n\t * for multi-turn continuity, reducing time-to-first-text-token.\n\t *\n\t * Note: Anthropic's API default for Claude Opus 4.8 and Mythos Preview is\n\t * \"omitted\". We default to \"summarized\" here to keep behavior consistent with\n\t * older Claude 4 models. Only applies to Claude models on Bedrock.\n\t */\n\tthinkingDisplay?: BedrockThinkingDisplay;\n\t/** Key-value pairs attached to the inference request for cost allocation tagging.\n\t * Keys: max 64 chars, no `aws:` prefix. Values: max 256 chars. Max 50 pairs.\n\t * Tags appear in AWS Cost Explorer split cost allocation data.\n\t * @see https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ConverseStream.html */\n\trequestMetadata?: Record<string, string>;\n\t/** Bearer token for Bedrock API key authentication.\n\t * When set, bypasses SigV4 signing and sends Authorization: Bearer <token> instead.\n\t * Requires `bedrock:CallWithBearerToken` IAM permission on the token's identity.\n\t * Set via AWS_BEARER_TOKEN_BEDROCK env var or pass directly.\n\t * @see https://docs.aws.amazon.com/service-authorization/latest/reference/list_amazonbedrock.html */\n\tbearerToken?: string;\n}\n\ntype Block = (TextContent | ThinkingContent | ToolCall) & {\n\tindex?: number;\n\tpartialJson?: string;\n\t/** Scratch buffer for encrypted reasoning deltas, joined into `thinkingSignature`. */\n\tredactedChunks?: Uint8Array[];\n};\n\nconst EMPTY_TEXT_PLACEHOLDER = \"<empty>\";\n\n/** Matches the placeholder the Anthropic API path uses for redacted thinking. */\nconst REDACTED_THINKING_PLACEHOLDER = \"[Reasoning redacted]\";\n\nexport const stream: StreamFunction<\"bedrock-converse-stream\", BedrockOptions> = (\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcontext: Context,\n\toptions: BedrockOptions = {},\n): AssistantMessageEventStream => {\n\tconst stream = new AssistantMessageEventStream();\n\n\t(async () => {\n\t\tconst output: AssistantMessage = {\n\t\t\trole: \"assistant\",\n\t\t\tcontent: [],\n\t\t\tapi: \"bedrock-converse-stream\" as Api,\n\t\t\tprovider: model.provider,\n\t\t\tmodel: model.id,\n\t\t\tusage: {\n\t\t\t\tinput: 0,\n\t\t\t\toutput: 0,\n\t\t\t\tcacheRead: 0,\n\t\t\t\tcacheWrite: 0,\n\t\t\t\ttotalTokens: 0,\n\t\t\t\tcost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },\n\t\t\t},\n\t\t\tstopReason: \"pending\",\n\t\t\ttimestamp: Date.now(),\n\t\t};\n\n\t\tconst blocks = output.content as Block[];\n\n\t\t// A profile explicitly configured through pi's auth flow (the `profile`\n\t\t// option or scoped `AWS_PROFILE` on the stored credential's env) must win\n\t\t// over ambient AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY. The SDK default\n\t\t// chain already prefers a configured profile over env keys, but only when\n\t\t// `credentials` is not set on the client config. See #6957.\n\t\tconst optionsProfile = options.profile || options.env?.AWS_PROFILE;\n\t\tconst config: BedrockRuntimeClientConfig = {\n\t\t\tprofile: optionsProfile || getProviderEnvValue(\"AWS_PROFILE\", options.env),\n\t\t};\n\t\tconst configuredRegion = getConfiguredBedrockRegion(options);\n\t\tconst hasAmbientConfiguredProfile = Boolean(getProviderEnvValue(\"AWS_PROFILE\"));\n\t\tconst endpointRegion = getStandardBedrockEndpointRegion(model.baseUrl);\n\t\tconst useExplicitEndpoint = shouldUseExplicitBedrockEndpoint(\n\t\t\tmodel.baseUrl,\n\t\t\tconfiguredRegion,\n\t\t\thasAmbientConfiguredProfile,\n\t\t);\n\n\t\t// Only pin standard AWS Bedrock runtime endpoints when no region or ambient AWS_PROFILE is configured.\n\t\t// This preserves custom endpoints (VPC/proxy) from #3402 without forcing built-in\n\t\t// catalog defaults such as us-east-1 to override AWS_REGION/AWS_PROFILE.\n\t\tif (useExplicitEndpoint) {\n\t\t\tconfig.endpoint = model.baseUrl;\n\t\t}\n\n\t\t// Resolve bearer token for Bedrock API key auth.\n\t\tconst skipAuth = getProviderEnvValue(\"AWS_BEDROCK_SKIP_AUTH\", options.env) === \"1\";\n\t\tconst bearerToken =\n\t\t\toptions.bearerToken ||\n\t\t\toptions.apiKey ||\n\t\t\tgetProviderEnvValue(\"AWS_BEARER_TOKEN_BEDROCK\", options.env) ||\n\t\t\tundefined;\n\t\tconst useBearerToken = bearerToken !== undefined && !skipAuth;\n\n\t\t// in Node.js/Bun environment only\n\t\tif (typeof process !== \"undefined\" && (process.versions?.node || process.versions?.bun)) {\n\t\t\t// Region resolution: ARN-embedded > explicit option > env vars > SDK default chain.\n\t\t\t// When the model ID is an inference profile ARN, extract the region from it.\n\t\t\t// This avoids conflicts with AWS_REGION set for other services.\n\t\t\tconst arnRegionMatch = model.id.match(/^arn:aws(?:-[a-z0-9-]+)?:bedrock:([a-z0-9-]+):/);\n\t\t\tif (arnRegionMatch) {\n\t\t\t\tconfig.region = arnRegionMatch[1];\n\t\t\t} else if (configuredRegion) {\n\t\t\t\tconfig.region = configuredRegion;\n\t\t\t} else if (endpointRegion && useExplicitEndpoint) {\n\t\t\t\tconfig.region = endpointRegion;\n\t\t\t} else if (!hasAmbientConfiguredProfile) {\n\t\t\t\tconfig.region = \"us-east-1\";\n\t\t\t}\n\n\t\t\t// Support proxies that don't need authentication\n\t\t\tif (skipAuth) {\n\t\t\t\tconfig.credentials = {\n\t\t\t\t\taccessKeyId: \"dummy-access-key\",\n\t\t\t\t\tsecretAccessKey: \"dummy-secret-key\",\n\t\t\t\t};\n\t\t\t}\n\n\t\t\tconst credentials = getConfiguredBedrockCredentials(options.env);\n\t\t\tif (!skipAuth && credentials && !optionsProfile) {\n\t\t\t\tconfig.credentials = credentials;\n\t\t\t}\n\n\t\t\tconst proxyUrl = resolveHttpProxyUrlForTarget(model.baseUrl, options.env);\n\t\t\tif (proxyUrl) {\n\t\t\t\t// Bedrock runtime uses NodeHttp2Handler by default since v3.798.0, which is based\n\t\t\t\t// on `http2` module and has no support for http agent.\n\t\t\t\t// Use NodeHttpHandler to support HTTP(S) proxy agents.\n\t\t\t\tconfig.requestHandler = new NodeHttpHandler({\n\t\t\t\t\thttpAgent: new HttpProxyAgent(proxyUrl),\n\t\t\t\t\thttpsAgent: new HttpsProxyAgent(proxyUrl) as unknown as HttpsAgent,\n\t\t\t\t});\n\t\t\t} else if (getProviderEnvValue(\"AWS_BEDROCK_FORCE_HTTP1\", options.env) === \"1\") {\n\t\t\t\t// Some custom endpoints require HTTP/1.1 instead of HTTP/2\n\t\t\t\tconfig.requestHandler = new NodeHttpHandler();\n\t\t\t}\n\t\t} else {\n\t\t\t// Non-Node environment (browser): fall back to us-east-1 since\n\t\t\t// there's no config file resolution available.\n\t\t\tconfig.region =\n\t\t\t\tconfiguredRegion || (endpointRegion && useExplicitEndpoint ? endpointRegion : undefined) || \"us-east-1\";\n\t\t}\n\n\t\tif (useBearerToken) {\n\t\t\tconfig.token = { token: bearerToken };\n\t\t\tconfig.authSchemePreference = [\"httpBearerAuth\"];\n\t\t}\n\n\t\t// Kept outside the try so the catch can still correlate a mid-stream failure:\n\t\t// exceptions delivered as stream events carry no HTTP metadata of their own.\n\t\tlet responseRequestId: string | undefined;\n\n\t\ttry {\n\t\t\tconst supportsStrictMode = model.compat?.supportsStrictMode ?? false;\n\t\t\tconst client = new BedrockRuntimeClient(config);\n\t\t\tlet observedRawResponse = false;\n\t\t\tif (options.onResponse) {\n\t\t\t\taddResponseHeadersMiddleware(client, options.onResponse, model, () => {\n\t\t\t\t\tobservedRawResponse = true;\n\t\t\t\t});\n\t\t\t}\n\t\t\tconst customHeaders = providerHeadersToRecord({ ...model.headers, ...options.headers });\n\t\t\tif (customHeaders) {\n\t\t\t\taddCustomHeadersMiddleware(client, customHeaders);\n\t\t\t}\n\t\t\tconst cacheRetention = resolveCacheRetention(options.cacheRetention, options.env);\n\t\t\tconst inferenceMaxTokens = options.maxTokens ?? (isAnthropicClaudeModel(model) ? model.maxTokens : undefined);\n\t\t\tlet commandInput = {\n\t\t\t\tmodelId: model.id,\n\t\t\t\tmessages: convertMessages(context, model, cacheRetention, options.env),\n\t\t\t\tsystem: buildSystemPrompt(context.systemPrompt, model, cacheRetention, options.env),\n\t\t\t\tinferenceConfig: {\n\t\t\t\t\t...(inferenceMaxTokens !== undefined && { maxTokens: inferenceMaxTokens }),\n\t\t\t\t\t// Claude Fable 5.1 rejects non-default `temperature`, `top_p`, and `top_k` on every\n\t\t\t\t\t// request. Generated metadata marks such models `supportsTemperature: false`; every\n\t\t\t\t\t// other model keeps sending the field exactly as before.\n\t\t\t\t\t// https://platform.claude.com/docs/en/build-with-claude/thinking\n\t\t\t\t\t...(options.temperature !== undefined &&\n\t\t\t\t\t\tmodel.compat?.supportsTemperature !== false && { temperature: options.temperature }),\n\t\t\t\t},\n\t\t\t\ttoolConfig: convertToolConfig(context.tools, options.toolChoice, supportsStrictMode, model),\n\t\t\t\tadditionalModelRequestFields: buildAdditionalModelRequestFields(model, options),\n\t\t\t\t...(options.requestMetadata !== undefined && { requestMetadata: options.requestMetadata }),\n\t\t\t};\n\t\t\tconst nextCommandInput = await options?.onPayload?.(commandInput, model);\n\t\t\tif (nextCommandInput !== undefined) {\n\t\t\t\tcommandInput = nextCommandInput as typeof commandInput;\n\t\t\t}\n\t\t\tconst command = new ConverseStreamCommand(commandInput);\n\n\t\t\tconst response = await client.send(command, { abortSignal: options.signal });\n\t\t\tresponseRequestId = normalizeDiagnosticValue(response.$metadata.requestId);\n\t\t\tif (!observedRawResponse && response.$metadata.httpStatusCode !== undefined) {\n\t\t\t\tconst responseHeaders: Record<string, string> = {};\n\t\t\t\tif (response.$metadata.requestId) {\n\t\t\t\t\tresponseHeaders[\"x-amzn-requestid\"] = response.$metadata.requestId;\n\t\t\t\t}\n\t\t\t\tawait options?.onResponse?.({ status: response.$metadata.httpStatusCode, headers: responseHeaders }, model);\n\t\t\t}\n\n\t\t\tfor await (const item of response.stream!) {\n\t\t\t\tif (item.messageStart) {\n\t\t\t\t\tif (item.messageStart.role !== ConversationRole.ASSISTANT) {\n\t\t\t\t\t\tthrow new Error(\"Unexpected assistant message start but got user message start instead\");\n\t\t\t\t\t}\n\t\t\t\t\tstream.push({ type: \"start\", partial: output });\n\t\t\t\t} else if (item.contentBlockStart) {\n\t\t\t\t\thandleContentBlockStart(item.contentBlockStart, blocks, output, stream);\n\t\t\t\t} else if (item.contentBlockDelta) {\n\t\t\t\t\thandleContentBlockDelta(item.contentBlockDelta, blocks, output, stream);\n\t\t\t\t} else if (item.contentBlockStop) {\n\t\t\t\t\thandleContentBlockStop(item.contentBlockStop, blocks, output, stream);\n\t\t\t\t} else if (item.messageStop) {\n\t\t\t\t\toutput.rawStopReason = item.messageStop.stopReason;\n\t\t\t\t\tconst { stopReason, errorMessage } = mapStopReason(item.messageStop.stopReason);\n\t\t\t\t\toutput.stopReason = stopReason;\n\t\t\t\t\tif (errorMessage) {\n\t\t\t\t\t\toutput.errorMessage = errorMessage;\n\t\t\t\t\t}\n\t\t\t\t} else if (item.metadata) {\n\t\t\t\t\thandleMetadata(item.metadata, model, output);\n\t\t\t\t} else if (item.internalServerException) {\n\t\t\t\t\tthrow item.internalServerException;\n\t\t\t\t} else if (item.modelStreamErrorException) {\n\t\t\t\t\tthrow item.modelStreamErrorException;\n\t\t\t\t} else if (item.validationException) {\n\t\t\t\t\tthrow item.validationException;\n\t\t\t\t} else if (item.throttlingException) {\n\t\t\t\t\tthrow item.throttlingException;\n\t\t\t\t} else if (item.serviceUnavailableException) {\n\t\t\t\t\tthrow item.serviceUnavailableException;\n\t\t\t\t}\n\t\t\t}\n\n\t\t\tif (options.signal?.aborted) {\n\t\t\t\tthrow new Error(\"Request was aborted\");\n\t\t\t}\n\n\t\t\tif (output.stopReason === \"pending\") {\n\t\t\t\tthrow new Error(\"Bedrock stream ended without a stop reason\");\n\t\t\t}\n\t\t\tif (output.stopReason === \"error\" || output.stopReason === \"aborted\") {\n\t\t\t\tthrow new Error(output.errorMessage || \"An unknown error occurred\");\n\t\t\t}\n\n\t\t\t// A stream can settle without stopping every block, so finalize here too.\n\t\t\tfor (const block of output.content) finalizeStreamingBlock(block as Block);\n\t\t\tstream.push({ type: \"done\", reason: output.stopReason, message: output });\n\t\t\tstream.end();\n\t\t} catch (error) {\n\t\t\tfor (const block of output.content) {\n\t\t\t\tfinalizeStreamingBlock(block as Block);\n\t\t\t}\n\t\t\toutput.stopReason = options.signal?.aborted ? \"aborted\" : \"error\";\n\t\t\toutput.errorMessage = formatBedrockError(error);\n\t\t\tif (output.stopReason === \"error\") {\n\t\t\t\tappendBedrockFailureDiagnostic(output, error, responseRequestId);\n\t\t\t}\n\t\t\tstream.push({ type: \"error\", reason: output.stopReason, error: output });\n\t\t\tstream.end();\n\t\t}\n\t})();\n\n\treturn stream;\n};\n\n/**\n * Human-readable prefixes for Bedrock SDK exception names.\n * The downstream retry logic in agent-session matches patterns like\n * `server.?error` and `service.?unavailable`, so we preserve the legacy\n * prefix format rather than using the raw SDK exception name.\n */\nconst BEDROCK_ERROR_PREFIXES: Record<string, string> = {\n\tInternalServerException: \"Internal server error\",\n\tModelStreamErrorException: \"Model stream error\",\n\tValidationException: \"Validation error\",\n\tThrottlingException: \"Throttling error\",\n\tServiceUnavailableException: \"Service unavailable\",\n};\n\n/**\n * Some models reject the account/profile's configured Bedrock data retention mode\n * (e.g. \"data retention mode 'default' is not available for this model\"). Point\n * users at the AWS docs explaining how to configure a supported mode.\n */\nconst BEDROCK_DATA_RETENTION_DOCS_URL = \"https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html\";\n\n/**\n * Format a Bedrock error with a human-readable prefix.\n * AWS SDK exceptions (both from `client.send()` and from stream event items)\n * extend BedrockRuntimeServiceException. We map the `.name` to a stable\n * human-readable prefix so downstream consumers (retry logic, context-overflow\n * detection) can distinguish error categories via simple string matching.\n */\nfunction formatBedrockError(error: unknown): string {\n\tconst norm = normalizeProviderError(error);\n\t// Surface the raw HTTP body (with status) when the SDK did not fold it into\n\t// the message; otherwise fall back to the message. This is what stops a\n\t// gateway 403 from collapsing to `Unknown: UnknownError`.\n\tconst core =\n\t\t!norm.messageCarriesBody && norm.status !== undefined && norm.body !== undefined\n\t\t\t? `${norm.status}: ${norm.body}`\n\t\t\t: norm.message;\n\tconst dataRetentionHint = /data retention mode/i.test(core)\n\t\t? ` See ${BEDROCK_DATA_RETENTION_DOCS_URL} for supported data retention modes.`\n\t\t: \"\";\n\tif (error instanceof BedrockRuntimeServiceException) {\n\t\tconst prefix = BEDROCK_ERROR_PREFIXES[error.name] ?? error.name;\n\t\treturn `${prefix}: ${core}${dataRetentionHint}`;\n\t}\n\treturn `${core}${dataRetentionHint}`;\n}\n\ntype SdkErrorMetadata = { $metadata?: { httpStatusCode?: unknown; requestId?: unknown } };\n\n/** Over-long header values are dropped rather than truncated: a truncated request id is not a request id. */\nconst MAX_BEDROCK_DIAGNOSTIC_VALUE_CHARS = 200;\n\nfunction normalizeDiagnosticValue(value: unknown): string | undefined {\n\tif (typeof value !== \"string\") return undefined;\n\tconst trimmed = value.trim();\n\tif (trimmed.length === 0 || trimmed.length > MAX_BEDROCK_DIAGNOSTIC_VALUE_CHARS) return undefined;\n\treturn trimmed;\n}\n\n/**\n * The SDK puts the modeled code on `error.name` for service exceptions and unmodeled stream errors alike, so\n * do not narrow to `BedrockRuntimeServiceException`. Modeled Bedrock errors all end in `Exception`, unlike\n * transport names such as `TimeoutError`.\n */\nfunction extractBedrockErrorCode(error: unknown): string | undefined {\n\tif (!(error instanceof Error) || !error.name.endsWith(\"Exception\")) return undefined;\n\treturn normalizeDiagnosticValue(error.name);\n}\n\n/**\n * Structured metadata alongside `errorMessage`, which stays byte-identical because `isRetryableAssistantError`\n * matches against it. Unknown fields are omitted, never guessed: a modeled mid-stream exception reaches us as\n * a bare object literal, leaving only `fallbackRequestId`. `details` only, as the throw is not always `Error`.\n */\nfunction appendBedrockFailureDiagnostic(\n\toutput: AssistantMessage,\n\terror: unknown,\n\tfallbackRequestId: string | undefined,\n): void {\n\tconst metadata = (error as SdkErrorMetadata)?.$metadata;\n\tconst details: Record<string, unknown> = {};\n\n\tif (typeof metadata?.httpStatusCode === \"number\") details.status = metadata.httpStatusCode;\n\n\tconst errorCode = extractBedrockErrorCode(error);\n\tif (errorCode !== undefined) details.errorCode = errorCode;\n\n\tconst requestId = normalizeDiagnosticValue(metadata?.requestId) ?? fallbackRequestId;\n\tif (requestId !== undefined) details.requestId = requestId;\n\n\tif (Object.keys(details).length === 0) return;\n\n\tappendAssistantMessageDiagnostic(output, { type: \"bedrock_response_failure\", timestamp: Date.now(), details });\n}\n\n/**\n * Header keys that must never be overwritten by caller-supplied headers.\n * `host` and `x-amz-*` participate in the SigV4 canonical request; `authorization`\n * is owned by SigV4 or the bearer-token path (config.token + authSchemePreference).\n * Compared case-insensitively (caller key is lower-cased before lookup).\n */\nconst RESERVED_HEADER_EXACT = new Set([\"authorization\", \"host\"]);\n\nfunction isReservedHeader(key: string): boolean {\n\tconst lower = key.toLowerCase();\n\treturn lower.startsWith(\"x-amz-\") || RESERVED_HEADER_EXACT.has(lower);\n}\n\n/**\n * Attach caller-supplied headers to the outgoing Bedrock request via a Smithy\n * `build`-step middleware. The `build` step runs after request serialisation but\n * before SigV4 signing, so injected headers are covered by the signature. Reserved\n * SigV4 / auth headers (`x-amz-*`, `authorization`, `host`) are silently skipped;\n * all other caller headers override any existing same-named header on the request.\n */\nfunction addCustomHeadersMiddleware(client: BedrockRuntimeClient, headers: Record<string, string>): void {\n\tconst middleware: BuildMiddleware<object, MetadataBearer> = (next) => async (args) => {\n\t\tconst request = args.request;\n\t\tif (request && typeof request === \"object\" && \"headers\" in request) {\n\t\t\tconst requestHeaders = (request as { headers: Record<string, string> }).headers;\n\t\t\tfor (const [key, value] of Object.entries(headers)) {\n\t\t\t\tif (!isReservedHeader(key)) {\n\t\t\t\t\trequestHeaders[key] = value;\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t\treturn next(args);\n\t};\n\tclient.middlewareStack.add(middleware, { step: \"build\", name: \"pi-ai-custom-headers\", priority: \"low\" });\n}\n\nfunction isSmithyHttpResponse(response: unknown): response is HttpResponse {\n\tif (!response || typeof response !== \"object\") return false;\n\tconst candidate = response as Partial<HttpResponse>;\n\treturn typeof candidate.statusCode === \"number\" && !!candidate.headers && typeof candidate.headers === \"object\";\n}\n\nfunction toProviderResponse(response: unknown): ProviderResponse | undefined {\n\tif (!isSmithyHttpResponse(response)) return undefined;\n\treturn { status: response.statusCode, headers: { ...response.headers } };\n}\n\n/**\n * Bedrock's modeled `$metadata` only preserves selected HTTP metadata (for example\n * requestId), so custom gateway headers are otherwise lost before callers see\n * `onResponse`. Capture the raw Smithy HTTP response at the deserialize step,\n * after the SDK receives the response but before the event stream is consumed.\n */\nfunction addResponseHeadersMiddleware(\n\tclient: BedrockRuntimeClient,\n\tonResponse: NonNullable<BedrockOptions[\"onResponse\"]>,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tonObserved: () => void,\n): void {\n\tconst middleware: DeserializeMiddleware<object, MetadataBearer> = (next) => async (args) => {\n\t\tconst result = await next(args);\n\t\tconst providerResponse = toProviderResponse(result.response);\n\t\tif (providerResponse) {\n\t\t\tonObserved();\n\t\t\tawait onResponse(providerResponse, model);\n\t\t}\n\t\treturn result;\n\t};\n\tclient.middlewareStack.add(middleware, { step: \"deserialize\", name: \"pi-ai-response-headers\" });\n}\n\nexport const streamSimple: StreamFunction<\"bedrock-converse-stream\", SimpleStreamOptions> = (\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcontext: Context,\n\toptions?: SimpleStreamOptions,\n): AssistantMessageEventStream => {\n\tconst base = {\n\t\t...buildBaseOptions(model, context, options, undefined),\n\t\ttoolChoice: options?.toolChoice,\n\t} satisfies BedrockOptions;\n\tif (!options?.reasoning) {\n\t\treturn stream(model, context, { ...base, reasoning: undefined } satisfies BedrockOptions);\n\t}\n\n\tif (isAnthropicClaudeModel(model)) {\n\t\tif (supportsAdaptiveThinking(model.id, model.name)) {\n\t\t\treturn stream(model, context, {\n\t\t\t\t...base,\n\t\t\t\treasoning: options.reasoning,\n\t\t\t\tthinkingBudgets: options.thinkingBudgets,\n\t\t\t} satisfies BedrockOptions);\n\t\t}\n\n\t\t// Undefined means the caller did not request an output cap; let the helper use the model cap.\n\t\t// Do not coerce to 0 here, or the thinking budget would become the entire maxTokens value.\n\t\tconst adjusted = adjustMaxTokensForThinking(\n\t\t\tbase.maxTokens,\n\t\t\tmodel.maxTokens,\n\t\t\toptions.reasoning,\n\t\t\toptions.thinkingBudgets,\n\t\t);\n\n\t\tconst maxTokens = clampMaxTokensToContext(model, context, adjusted.maxTokens);\n\n\t\treturn stream(model, context, {\n\t\t\t...base,\n\t\t\tmaxTokens,\n\t\t\treasoning: options.reasoning,\n\t\t\tthinkingBudgets: {\n\t\t\t\t...(options.thinkingBudgets || {}),\n\t\t\t\t[clampReasoning(options.reasoning)!]: Math.min(adjusted.thinkingBudget, Math.max(0, maxTokens - 1024)),\n\t\t\t},\n\t\t} satisfies BedrockOptions);\n\t}\n\n\treturn stream(model, context, {\n\t\t...base,\n\t\treasoning: options.reasoning,\n\t\tthinkingBudgets: options.thinkingBudgets,\n\t} satisfies BedrockOptions);\n};\n\nfunction handleContentBlockStart(\n\tevent: ContentBlockStartEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst index = event.contentBlockIndex!;\n\tconst start = event.start;\n\n\tif (start?.toolUse) {\n\t\tconst block: Block = {\n\t\t\ttype: \"toolCall\",\n\t\t\tid: start.toolUse.toolUseId || \"\",\n\t\t\tname: start.toolUse.name || \"\",\n\t\t\targuments: {},\n\t\t\tpartialJson: \"\",\n\t\t\tindex,\n\t\t};\n\t\toutput.content.push(block);\n\t\tstream.push({ type: \"toolcall_start\", contentIndex: blocks.length - 1, partial: output });\n\t}\n}\n\nfunction handleContentBlockDelta(\n\tevent: ContentBlockDeltaEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst contentBlockIndex = event.contentBlockIndex!;\n\tconst delta = event.delta;\n\tlet index = blocks.findIndex((b) => b.index === contentBlockIndex);\n\tlet block = blocks[index];\n\n\tif (delta?.text !== undefined) {\n\t\t// If no text block exists yet, create one, as `handleContentBlockStart` is not sent for text blocks\n\t\tif (!block) {\n\t\t\tconst newBlock: Block = { type: \"text\", text: \"\", index: contentBlockIndex };\n\t\t\toutput.content.push(newBlock);\n\t\t\tindex = blocks.length - 1;\n\t\t\tblock = blocks[index];\n\t\t\tstream.push({ type: \"text_start\", contentIndex: index, partial: output });\n\t\t}\n\t\tif (block.type === \"text\") {\n\t\t\tblock.text += delta.text;\n\t\t\tstream.push({ type: \"text_delta\", contentIndex: index, delta: delta.text, partial: output });\n\t\t}\n\t} else if (delta?.toolUse && block?.type === \"toolCall\") {\n\t\tblock.partialJson = (block.partialJson || \"\") + (delta.toolUse.input || \"\");\n\t\tblock.arguments = parseStreamingJson(block.partialJson);\n\t\tstream.push({ type: \"toolcall_delta\", contentIndex: index, delta: delta.toolUse.input || \"\", partial: output });\n\t} else if (delta?.reasoningContent) {\n\t\tlet thinkingBlock = block;\n\t\tlet thinkingIndex = index;\n\n\t\tif (!thinkingBlock) {\n\t\t\tconst newBlock: Block = { type: \"thinking\", thinking: \"\", thinkingSignature: \"\", index: contentBlockIndex };\n\t\t\toutput.content.push(newBlock);\n\t\t\tthinkingIndex = blocks.length - 1;\n\t\t\tthinkingBlock = blocks[thinkingIndex];\n\t\t\tstream.push({ type: \"thinking_start\", contentIndex: thinkingIndex, partial: output });\n\t\t}\n\n\t\tif (thinkingBlock?.type === \"thinking\") {\n\t\t\tif (delta.reasoningContent.text) {\n\t\t\t\tthinkingBlock.thinking += delta.reasoningContent.text;\n\t\t\t\tstream.push({\n\t\t\t\t\ttype: \"thinking_delta\",\n\t\t\t\t\tcontentIndex: thinkingIndex,\n\t\t\t\t\tdelta: delta.reasoningContent.text,\n\t\t\t\t\tpartial: output,\n\t\t\t\t});\n\t\t\t}\n\t\t\t// `thinkingSignature` holds either an Anthropic signature or an opaque redacted\n\t\t\t// payload, never both: mixing them would corrupt whichever arrived first.\n\t\t\tif (delta.reasoningContent.signature && !thinkingBlock.redacted) {\n\t\t\t\tthinkingBlock.thinkingSignature =\n\t\t\t\t\t(thinkingBlock.thinkingSignature || \"\") + delta.reasoningContent.signature;\n\t\t\t}\n\t\t\tif (delta.reasoningContent.redactedContent?.length) {\n\t\t\t\t// Encrypted reasoning from non-Anthropic models on Bedrock (e.g. OpenAI GPT-5.6).\n\t\t\t\t// The payload is opaque, so keep it verbatim in `thinkingSignature` the way the\n\t\t\t\t// Anthropic path stores redacted thinking, and replay it on the next turn.\n\t\t\t\tif (!thinkingBlock.redacted) {\n\t\t\t\t\tthinkingBlock.redacted = true;\n\t\t\t\t\tthinkingBlock.thinkingSignature = \"\";\n\t\t\t\t\tthinkingBlock.thinking += REDACTED_THINKING_PLACEHOLDER;\n\t\t\t\t\tstream.push({\n\t\t\t\t\t\ttype: \"thinking_delta\",\n\t\t\t\t\t\tcontentIndex: thinkingIndex,\n\t\t\t\t\t\tdelta: REDACTED_THINKING_PLACEHOLDER,\n\t\t\t\t\t\tpartial: output,\n\t\t\t\t\t});\n\t\t\t\t}\n\t\t\t\tthinkingBlock.redactedChunks ??= [];\n\t\t\t\tthinkingBlock.redactedChunks.push(delta.reasoningContent.redactedContent);\n\t\t\t}\n\t\t}\n\t}\n}\n\n/**\n * Encodes buffered encrypted reasoning into `thinkingSignature` and drops the scratch\n * buffer, which must never reach a persisted message: `Uint8Array` serializes to an\n * index-keyed object roughly ten times the size of the base64 payload.\n */\nfunction flushRedactedContent(block: Block): void {\n\tif (block.type !== \"thinking\" || !block.redactedChunks) return;\n\tblock.thinkingSignature = bytesToBase64(block.redactedChunks);\n\tdelete block.redactedChunks;\n}\n\n/**\n * Strips every streaming scratch field. Runs from the terminal paths as well as\n * `contentBlockStop`, because a stream can settle without stopping each block.\n */\nfunction finalizeStreamingBlock(block: Block): void {\n\tdelete block.index;\n\t// partialJson is only a streaming scratch buffer; never persist it.\n\tdelete block.partialJson;\n\tflushRedactedContent(block);\n}\n\nfunction handleMetadata(\n\tevent: ConverseStreamMetadataEvent,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\toutput: AssistantMessage,\n): void {\n\tif (event.usage) {\n\t\toutput.usage.input = event.usage.inputTokens || 0;\n\t\toutput.usage.output = event.usage.outputTokens || 0;\n\t\toutput.usage.cacheRead = event.usage.cacheReadInputTokens || 0;\n\t\toutput.usage.cacheWrite = event.usage.cacheWriteInputTokens || 0;\n\t\toutput.usage.cacheWrite1h = event.usage.cacheDetails?.reduce(\n\t\t\t(total, detail) => total + (detail.ttl === CacheTTL.ONE_HOUR ? (detail.inputTokens ?? 0) : 0),\n\t\t\t0,\n\t\t);\n\t\toutput.usage.totalTokens = event.usage.totalTokens || output.usage.input + output.usage.output;\n\t\tcalculateCost(model, output.usage);\n\t}\n}\n\nfunction handleContentBlockStop(\n\tevent: ContentBlockStopEvent,\n\tblocks: Block[],\n\toutput: AssistantMessage,\n\tstream: AssistantMessageEventStream,\n): void {\n\tconst index = blocks.findIndex((b) => b.index === event.contentBlockIndex);\n\tconst block = blocks[index];\n\tif (!block) return;\n\tdelete (block as Block).index;\n\n\tswitch (block.type) {\n\t\tcase \"text\":\n\t\t\tstream.push({ type: \"text_end\", contentIndex: index, content: block.text, partial: output });\n\t\t\tbreak;\n\t\tcase \"thinking\":\n\t\t\tflushRedactedContent(block);\n\t\t\tstream.push({ type: \"thinking_end\", contentIndex: index, content: block.thinking, partial: output });\n\t\t\tbreak;\n\t\tcase \"toolCall\":\n\t\t\tblock.arguments = parseStreamingJson(block.partialJson);\n\t\t\t// Finalize in-place and strip the scratch buffer so replay only\n\t\t\t// carries parsed arguments.\n\t\t\tdelete (block as Block).partialJson;\n\t\t\tstream.push({ type: \"toolcall_end\", contentIndex: index, toolCall: block, partial: output });\n\t\t\tbreak;\n\t}\n}\n\n/**\n * Check if the model supports adaptive thinking (Opus 4.6+, Sonnet 4.6).\n * Checks both model ID and model name to support application inference profiles\n * whose ARNs don't contain the model name.\n */\nfunction getModelMatchCandidates(modelId: string, modelName?: string): string[] {\n\tconst values = modelName ? [modelId, modelName] : [modelId];\n\treturn values.flatMap((value) => {\n\t\tconst lower = value.toLowerCase();\n\t\treturn [lower, lower.replace(/[\\s_.:]+/g, \"-\")];\n\t});\n}\n\nfunction supportsAdaptiveThinking(modelId: string, modelName?: string): boolean {\n\tconst candidates = getModelMatchCandidates(modelId, modelName);\n\treturn candidates.some(\n\t\t(s) =>\n\t\t\ts.includes(\"opus-4-6\") ||\n\t\t\ts.includes(\"opus-4-7\") ||\n\t\t\ts.includes(\"opus-4-8\") ||\n\t\t\ts.includes(\"opus-5\") ||\n\t\t\ts.includes(\"sonnet-4-6\") ||\n\t\t\ts.includes(\"sonnet-5\") ||\n\t\t\ts.includes(\"fable-5\"),\n\t);\n}\n\nfunction supportsNativeXhighEffort(model: Model<\"bedrock-converse-stream\">): boolean {\n\tconst candidates = getModelMatchCandidates(model.id, model.name);\n\treturn candidates.some(\n\t\t(s) =>\n\t\t\ts.includes(\"opus-4-7\") ||\n\t\t\ts.includes(\"opus-4-8\") ||\n\t\t\ts.includes(\"opus-5\") ||\n\t\t\ts.includes(\"sonnet-5\") ||\n\t\t\ts.includes(\"fable-5\"),\n\t);\n}\n\nfunction mapThinkingLevelToEffort(\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tlevel: SimpleStreamOptions[\"reasoning\"],\n): \"low\" | \"medium\" | \"high\" | \"xhigh\" | \"max\" {\n\tif (level === \"xhigh\" && supportsNativeXhighEffort(model)) return \"xhigh\";\n\n\tconst mapped = level ? model.thinkingLevelMap?.[level] : undefined;\n\tif (typeof mapped === \"string\") return mapped as \"low\" | \"medium\" | \"high\" | \"xhigh\" | \"max\";\n\n\tswitch (level) {\n\t\tcase \"minimal\":\n\t\tcase \"low\":\n\t\t\treturn \"low\";\n\t\tcase \"medium\":\n\t\t\treturn \"medium\";\n\t\tcase \"high\":\n\t\t\treturn \"high\";\n\t\tdefault:\n\t\t\treturn \"high\";\n\t}\n}\n\n/**\n * Resolve cache retention preference.\n * Defaults to \"short\" and uses PI_CACHE_RETENTION for backward compatibility.\n */\nfunction resolveCacheRetention(cacheRetention?: CacheRetention, env?: ProviderEnv): CacheRetention {\n\tif (cacheRetention) {\n\t\treturn cacheRetention;\n\t}\n\tif (getProviderEnvValue(\"PI_CACHE_RETENTION\", env) === \"long\") {\n\t\treturn \"long\";\n\t}\n\treturn \"short\";\n}\n\n/**\n * Check if the model is an Anthropic Claude model on Bedrock.\n * Checks both model ID and model name to support application inference profiles\n * whose ARNs don't contain the model name.\n */\nfunction isAnthropicClaudeModel(model: Model<\"bedrock-converse-stream\">): boolean {\n\tconst id = model.id.toLowerCase();\n\tconst name = model.name?.toLowerCase() ?? \"\";\n\treturn (\n\t\tid.includes(\"anthropic.claude\") ||\n\t\tid.includes(\"anthropic/claude\") ||\n\t\tname.includes(\"anthropic.claude\") ||\n\t\tname.includes(\"anthropic/claude\") ||\n\t\tname.includes(\"claude\")\n\t);\n}\n\n/**\n * Check if the model supports prompt caching.\n * Supported: Claude 3.5 Haiku, Claude 3.7 Sonnet, Claude 4.x models, Claude 5 models\n *\n * For base models and system-defined inference profiles the model ID / ARN\n * contains the model name, so we can decide locally.\n *\n * For application inference profiles (whose ARNs don't contain the model name),\n * also checks model.name which is user-controlled via models.json or registerProvider.\n * As a last resort, set AWS_BEDROCK_FORCE_CACHE=1 to enable cache points.\n * Amazon Nova models have automatic caching and don't need explicit cache points.\n */\nfunction supportsPromptCaching(model: Model<\"bedrock-converse-stream\">, env?: ProviderEnv): boolean {\n\tconst candidates = getModelMatchCandidates(model.id, model.name);\n\n\tconst hasClaudeRef = candidates.some((s) => s.includes(\"claude\"));\n\tif (!hasClaudeRef) {\n\t\t// Application inference profiles don't contain the model name in the ARN.\n\t\t// Allow users to force cache points via environment variable.\n\t\tif (getProviderEnvValue(\"AWS_BEDROCK_FORCE_CACHE\", env) === \"1\") return true;\n\t\treturn false;\n\t}\n\t// Claude 5 models (fable-5, opus-5, sonnet-5)\n\tif (candidates.some((s) => s.includes(\"fable-5\") || s.includes(\"opus-5\") || s.includes(\"sonnet-5\"))) return true;\n\t// Claude 4.x models (opus-4, sonnet-4, haiku-4)\n\tif (candidates.some((s) => s.includes(\"-4-\"))) return true;\n\t// Claude 3.7 Sonnet\n\tif (candidates.some((s) => s.includes(\"claude-3-7-sonnet\"))) return true;\n\t// Claude 3.5 Haiku\n\tif (candidates.some((s) => s.includes(\"claude-3-5-haiku\"))) return true;\n\treturn false;\n}\n\n/**\n * Check if the model supports thinking signatures in reasoningContent.\n * Only Anthropic Claude models support the signature field.\n * Other models (OpenAI, Qwen, Minimax, Moonshot, etc.) reject it with:\n * \"This model doesn't support the reasoningContent.reasoningText.signature field\"\n *\n * Checks both model ID and model name to support application inference profiles.\n */\nfunction supportsThinkingSignature(model: Model<\"bedrock-converse-stream\">): boolean {\n\treturn isAnthropicClaudeModel(model);\n}\n\nfunction buildSystemPrompt(\n\tsystemPrompt: string | undefined,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcacheRetention: CacheRetention,\n\tenv?: ProviderEnv,\n): SystemContentBlock[] | undefined {\n\tif (!systemPrompt) return undefined;\n\n\tconst blocks: SystemContentBlock[] = [{ text: sanitizeSurrogates(systemPrompt) }];\n\n\t// Add cache point for supported Claude models when caching is enabled\n\tif (cacheRetention !== \"none\" && supportsPromptCaching(model, env)) {\n\t\tblocks.push({\n\t\t\tcachePoint: { type: CachePointType.DEFAULT, ...(cacheRetention === \"long\" ? { ttl: CacheTTL.ONE_HOUR } : {}) },\n\t\t});\n\t}\n\n\treturn blocks;\n}\n\nfunction normalizeToolCallId(id: string): string {\n\tconst sanitized = id.replace(/[^a-zA-Z0-9_-]/g, \"_\");\n\treturn sanitized.length > 64 ? sanitized.slice(0, 64) : sanitized;\n}\n\nfunction createNonBlankTextBlock(text: string): ContentBlock.TextMember | undefined {\n\tconst sanitized = sanitizeSurrogates(text);\n\treturn sanitized.trim().length === 0 ? undefined : { text: sanitized };\n}\n\nfunction createRequiredTextBlock(text: string): ContentBlock.TextMember {\n\treturn createNonBlankTextBlock(text) ?? { text: EMPTY_TEXT_PLACEHOLDER };\n}\n\nfunction sanitizeBedrockDocument(value: DocumentType): DocumentType {\n\tif (Array.isArray(value)) {\n\t\treturn value.map(sanitizeBedrockDocument);\n\t}\n\tif (value !== null && typeof value === \"object\") {\n\t\treturn Object.fromEntries(\n\t\t\tObject.entries(value)\n\t\t\t\t.filter(([key]) => key.length > 0)\n\t\t\t\t.map(([key, nestedValue]) => [key, sanitizeBedrockDocument(nestedValue)]),\n\t\t);\n\t}\n\treturn value;\n}\n\nfunction convertToolResultContent(content: (TextContent | ImageContent)[]): ToolResultContentBlock[] {\n\tconst result: ToolResultContentBlock[] = [];\n\tfor (const c of content) {\n\t\tif (c.type === \"image\") {\n\t\t\tresult.push({ image: createImageBlock(c.mimeType, c.data) });\n\t\t} else {\n\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\tif (textBlock) result.push(textBlock);\n\t\t}\n\t}\n\tif (result.length === 0) result.push({ text: EMPTY_TEXT_PLACEHOLDER });\n\treturn result;\n}\n\nfunction convertMessages(\n\tcontext: Context,\n\tmodel: Model<\"bedrock-converse-stream\">,\n\tcacheRetention: CacheRetention,\n\tenv?: ProviderEnv,\n): Message[] {\n\tconst result: Message[] = [];\n\tconst transformedMessages = transformMessages(context.messages, model, normalizeToolCallId);\n\n\tfor (let i = 0; i < transformedMessages.length; i++) {\n\t\tconst m = transformedMessages[i];\n\n\t\tswitch (m.role) {\n\t\t\tcase \"user\": {\n\t\t\t\tconst content: ContentBlock[] = [];\n\t\t\t\tif (typeof m.content === \"string\") {\n\t\t\t\t\tcontent.push(createRequiredTextBlock(m.content));\n\t\t\t\t} else {\n\t\t\t\t\tfor (const c of m.content) {\n\t\t\t\t\t\tswitch (c.type) {\n\t\t\t\t\t\t\tcase \"text\": {\n\t\t\t\t\t\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\t\t\t\t\t\tif (textBlock) content.push(textBlock);\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\tcase \"image\":\n\t\t\t\t\t\t\t\tcontent.push({ image: createImageBlock(c.mimeType, c.data) });\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\tcase \"document\":\n\t\t\t\t\t\t\t\tcontent.push({ document: createDocumentBlock(c, content.length) });\n\t\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t\tdefault:\n\t\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t\t}\n\t\t\t\t\t}\n\t\t\t\t\tif (content.length === 0) content.push({ text: EMPTY_TEXT_PLACEHOLDER });\n\t\t\t\t}\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.USER,\n\t\t\t\t\tcontent,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tcase \"assistant\": {\n\t\t\t\t// Skip assistant messages with empty content (e.g., from aborted requests)\n\t\t\t\t// Bedrock rejects messages with empty content arrays\n\t\t\t\tif (m.content.length === 0) {\n\t\t\t\t\tcontinue;\n\t\t\t\t}\n\t\t\t\tconst contentBlocks: ContentBlock[] = [];\n\t\t\t\tfor (const c of m.content) {\n\t\t\t\t\tswitch (c.type) {\n\t\t\t\t\t\tcase \"text\": {\n\t\t\t\t\t\t\t// Skip empty text blocks\n\t\t\t\t\t\t\tconst textBlock = createNonBlankTextBlock(c.text);\n\t\t\t\t\t\t\tif (!textBlock) continue;\n\t\t\t\t\t\t\tcontentBlocks.push(textBlock);\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t}\n\t\t\t\t\t\tcase \"toolCall\":\n\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\ttoolUse: { toolUseId: c.id, name: c.name, input: sanitizeBedrockDocument(c.arguments) },\n\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\tcase \"thinking\": {\n\t\t\t\t\t\t\t// Encrypted reasoning is opaque: replay the stored payload as the\n\t\t\t\t\t\t\t// `redactedContent` member instead of lowering it to reasoning text.\n\t\t\t\t\t\t\tif (c.redacted) {\n\t\t\t\t\t\t\t\tconst redactedContent = decodeRedactedContent(c.thinkingSignature);\n\t\t\t\t\t\t\t\tif (redactedContent?.length) {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({ reasoningContent: { redactedContent } });\n\t\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t// Skip empty thinking blocks\n\t\t\t\t\t\t\tconst thinking = sanitizeSurrogates(c.thinking);\n\t\t\t\t\t\t\tif (thinking.trim().length === 0) continue;\n\t\t\t\t\t\t\t// Only Anthropic models support the signature field in reasoningText.\n\t\t\t\t\t\t\t// For other models, we omit the signature to avoid errors like:\n\t\t\t\t\t\t\t// \"This model doesn't support the reasoningContent.reasoningText.signature field\"\n\t\t\t\t\t\t\tif (supportsThinkingSignature(model)) {\n\t\t\t\t\t\t\t\t// Signatures arrive after thinking deltas. If a partial or externally\n\t\t\t\t\t\t\t\t// persisted message lacks a signature, Bedrock rejects the replayed\n\t\t\t\t\t\t\t\t// reasoning block. Fall back to plain text, matching Anthropic.\n\t\t\t\t\t\t\t\tif (!c.thinkingSignature || c.thinkingSignature.trim().length === 0) {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({ text: thinking });\n\t\t\t\t\t\t\t\t} else {\n\t\t\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\t\t\treasoningContent: {\n\t\t\t\t\t\t\t\t\t\t\treasoningText: {\n\t\t\t\t\t\t\t\t\t\t\t\ttext: thinking,\n\t\t\t\t\t\t\t\t\t\t\t\tsignature: c.thinkingSignature,\n\t\t\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\t} else {\n\t\t\t\t\t\t\t\tcontentBlocks.push({\n\t\t\t\t\t\t\t\t\treasoningContent: {\n\t\t\t\t\t\t\t\t\t\treasoningText: { text: thinking },\n\t\t\t\t\t\t\t\t\t},\n\t\t\t\t\t\t\t\t});\n\t\t\t\t\t\t\t}\n\t\t\t\t\t\t\tbreak;\n\t\t\t\t\t\t}\n\t\t\t\t\t\tdefault:\n\t\t\t\t\t\t\tcontinue;\n\t\t\t\t\t}\n\t\t\t\t}\n\t\t\t\t// Skip if all content blocks were filtered out\n\t\t\t\tif (contentBlocks.length === 0) {\n\t\t\t\t\tcontinue;\n\t\t\t\t}\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.ASSISTANT,\n\t\t\t\t\tcontent: contentBlocks,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tcase \"toolResult\": {\n\t\t\t\t// Collect all consecutive toolResult messages into a single user message\n\t\t\t\t// Bedrock requires all tool results to be in one message\n\t\t\t\tconst toolResults: ContentBlock.ToolResultMember[] = [];\n\n\t\t\t\t// Add current tool result with all content blocks combined\n\t\t\t\ttoolResults.push({\n\t\t\t\t\ttoolResult: {\n\t\t\t\t\t\ttoolUseId: m.toolCallId,\n\t\t\t\t\t\tcontent: convertToolResultContent(m.content),\n\t\t\t\t\t\tstatus: m.isError ? ToolResultStatus.ERROR : ToolResultStatus.SUCCESS,\n\t\t\t\t\t},\n\t\t\t\t});\n\n\t\t\t\t// Look ahead for consecutive toolResult messages\n\t\t\t\tlet j = i + 1;\n\t\t\t\twhile (j < transformedMessages.length && transformedMessages[j].role === \"toolResult\") {\n\t\t\t\t\tconst nextMsg = transformedMessages[j] as ToolResultMessage;\n\t\t\t\t\ttoolResults.push({\n\t\t\t\t\t\ttoolResult: {\n\t\t\t\t\t\t\ttoolUseId: nextMsg.toolCallId,\n\t\t\t\t\t\t\tcontent: convertToolResultContent(nextMsg.content),\n\t\t\t\t\t\t\tstatus: nextMsg.isError ? ToolResultStatus.ERROR : ToolResultStatus.SUCCESS,\n\t\t\t\t\t\t},\n\t\t\t\t\t});\n\t\t\t\t\tj++;\n\t\t\t\t}\n\n\t\t\t\t// Skip the messages we've already processed\n\t\t\t\ti = j - 1;\n\n\t\t\t\tresult.push({\n\t\t\t\t\trole: ConversationRole.USER,\n\t\t\t\t\tcontent: toolResults,\n\t\t\t\t});\n\t\t\t\tbreak;\n\t\t\t}\n\t\t\tdefault:\n\t\t\t\tcontinue;\n\t\t}\n\t}\n\n\t// Add cache point to the last user message for supported Claude models when caching is enabled\n\tif (cacheRetention !== \"none\" && supportsPromptCaching(model, env) && result.length > 0) {\n\t\tconst lastMessage = result[result.length - 1];\n\t\tif (lastMessage.role === ConversationRole.USER && lastMessage.content) {\n\t\t\t(lastMessage.content as ContentBlock[]).push({\n\t\t\t\tcachePoint: {\n\t\t\t\t\ttype: CachePointType.DEFAULT,\n\t\t\t\t\t...(cacheRetention === \"long\" ? { ttl: CacheTTL.ONE_HOUR } : {}),\n\t\t\t\t},\n\t\t\t});\n\t\t}\n\t}\n\n\treturn result;\n}\n\nfunction convertToolConfig(\n\ttools: Tool[] | undefined,\n\ttoolChoice: BedrockOptions[\"toolChoice\"],\n\tsupportsStrictMode: boolean,\n\tmodel: Model<\"bedrock-converse-stream\">,\n): ToolConfiguration | undefined {\n\t// Validate the requested choice before any early return. A forced choice on a model that\n\t// rejects it is an error regardless of whether tools were supplied: returning `undefined` for\n\t// an empty or absent tool list would discard the caller's instruction silently, which is the\n\t// behavior this guard exists to prevent.\n\t//\n\t// Claude Fable 5.1 rejects forced tool use on every request with a 400, whichever platform\n\t// serves it. Reject rather than rewriting to `auto`, which would discard an explicit caller\n\t// instruction. Mirrors the guards in `anthropic-messages.ts` and `openai-completions.ts`.\n\t// https://platform.claude.com/docs/en/build-with-claude/thinking\n\tif (model.compat?.supportsForcedToolChoice === false) {\n\t\tconst isForced = toolChoice === \"any\" || (typeof toolChoice === \"object\" && toolChoice.type === \"tool\");\n\t\tif (isForced) {\n\t\t\tconst requestedLabel = typeof toolChoice === \"string\" ? toolChoice : `tool \"${toolChoice.name}\"`;\n\t\t\tthrow new Error(\n\t\t\t\t`Model ${model.id} does not support forced tool choice (requested: ${requestedLabel}). ` +\n\t\t\t\t\t`Use toolChoice \"auto\" with strict tool use or structured outputs instead.`,\n\t\t\t);\n\t\t}\n\t}\n\n\tif (!tools?.length) return undefined;\n\t// `none` is never a forced choice, so this return can never skip a rejection above.\n\tif (toolChoice === \"none\") return undefined;\n\n\tconst bedrockTools: BedrockTool[] = tools.map((tool) => {\n\t\tconst strict = resolveJsonSchemaStrictSampling(tool, supportsStrictMode);\n\t\treturn {\n\t\t\ttoolSpec: {\n\t\t\t\tname: tool.name,\n\t\t\t\tdescription: tool.description,\n\t\t\t\tinputSchema: { json: getJsonSchemaToolParameters(tool, strict) as unknown as DocumentType },\n\t\t\t\t...(strict === true ? { strict: true } : {}),\n\t\t\t},\n\t\t};\n\t});\n\n\tlet bedrockToolChoice: ToolChoice | undefined;\n\tswitch (toolChoice) {\n\t\tcase \"auto\":\n\t\t\tbedrockToolChoice = { auto: {} };\n\t\t\tbreak;\n\t\tcase \"any\":\n\t\t\tbedrockToolChoice = { any: {} };\n\t\t\tbreak;\n\t\tdefault:\n\t\t\tif (toolChoice?.type === \"tool\") {\n\t\t\t\tbedrockToolChoice = { tool: { name: toolChoice.name } };\n\t\t\t}\n\t}\n\n\treturn { tools: bedrockTools, toolChoice: bedrockToolChoice };\n}\n\nfunction mapStopReason(reason: string | undefined): { stopReason: StopReason; errorMessage?: string } {\n\tswitch (reason) {\n\t\tcase BedrockStopReason.END_TURN:\n\t\tcase BedrockStopReason.STOP_SEQUENCE:\n\t\t\treturn { stopReason: \"stop\" };\n\t\tcase BedrockStopReason.MAX_TOKENS:\n\t\tcase BedrockStopReason.MODEL_CONTEXT_WINDOW_EXCEEDED:\n\t\t\treturn { stopReason: \"length\" };\n\t\tcase BedrockStopReason.TOOL_USE:\n\t\t\treturn { stopReason: \"toolUse\" };\n\t\tdefault:\n\t\t\treturn reason\n\t\t\t\t? { stopReason: \"error\", errorMessage: `Provider stopped with: ${reason}` }\n\t\t\t\t: { stopReason: \"error\" };\n\t}\n}\n\nfunction getConfiguredBedrockRegion(options: BedrockOptions): string | undefined {\n\treturn (\n\t\toptions.region ||\n\t\tgetProviderEnvValue(\"AWS_REGION\", options.env) ||\n\t\tgetProviderEnvValue(\"AWS_DEFAULT_REGION\", options.env) ||\n\t\tundefined\n\t);\n}\n\nfunction getConfiguredBedrockCredentials(env?: ProviderEnv): BedrockRuntimeClientConfig[\"credentials\"] | undefined {\n\tconst accessKeyId = getProviderEnvValue(\"AWS_ACCESS_KEY_ID\", env);\n\tconst secretAccessKey = getProviderEnvValue(\"AWS_SECRET_ACCESS_KEY\", env);\n\tif (!accessKeyId || !secretAccessKey) {\n\t\treturn undefined;\n\t}\n\tconst sessionToken = getProviderEnvValue(\"AWS_SESSION_TOKEN\", env);\n\treturn {\n\t\taccessKeyId,\n\t\tsecretAccessKey,\n\t\t...(sessionToken ? { sessionToken } : {}),\n\t};\n}\n\nfunction getStandardBedrockEndpointRegion(baseUrl: string | undefined): string | undefined {\n\tif (!baseUrl) {\n\t\treturn undefined;\n\t}\n\n\ttry {\n\t\tconst { hostname } = new URL(baseUrl);\n\t\tconst match = hostname.toLowerCase().match(/^bedrock-runtime(?:-fips)?\\.([a-z0-9-]+)\\.amazonaws\\.com(?:\\.cn)?$/);\n\t\treturn match?.[1];\n\t} catch {\n\t\treturn undefined;\n\t}\n}\n\nfunction shouldUseExplicitBedrockEndpoint(\n\tbaseUrl: string,\n\tconfiguredRegion: string | undefined,\n\thasAmbientConfiguredProfile: boolean,\n): boolean {\n\tconst endpointRegion = getStandardBedrockEndpointRegion(baseUrl);\n\tif (!endpointRegion) {\n\t\treturn true;\n\t}\n\n\treturn !configuredRegion && !hasAmbientConfiguredProfile;\n}\n\nfunction isGovCloudBedrockTarget(model: Model<\"bedrock-converse-stream\">, options: BedrockOptions): boolean {\n\tconst region = getConfiguredBedrockRegion(options);\n\tif (region?.toLowerCase().startsWith(\"us-gov-\")) {\n\t\treturn true;\n\t}\n\n\tconst modelId = model.id.toLowerCase();\n\treturn modelId.startsWith(\"us-gov.\") || modelId.startsWith(\"arn:aws-us-gov:\");\n}\n\nfunction isOpenAiGpt6AstraBedrockModel(model: Model<\"bedrock-converse-stream\">): boolean {\n\treturn /^(?:global\\.|us\\.)?openai\\.gpt-6-astra$/.test(model.id);\n}\n\nfunction buildAdditionalModelRequestFields(\n\tmodel: Model<\"bedrock-converse-stream\">,\n\toptions: BedrockOptions,\n): Record<string, any> | undefined {\n\tif (!options.reasoning || !model.reasoning) {\n\t\treturn undefined;\n\t}\n\n\tif (isAnthropicClaudeModel(model)) {\n\t\t// GovCloud Bedrock currently rejects the Claude thinking.display field.\n\t\t// Omit it there until the GovCloud Converse schema catches up.\n\t\tconst display = isGovCloudBedrockTarget(model, options) ? undefined : (options.thinkingDisplay ?? \"summarized\");\n\t\tconst result: Record<string, any> = supportsAdaptiveThinking(model.id, model.name)\n\t\t\t? {\n\t\t\t\t\tthinking: { type: \"adaptive\", ...(display !== undefined ? { display } : {}) },\n\t\t\t\t\toutput_config: { effort: mapThinkingLevelToEffort(model, options.reasoning) },\n\t\t\t\t}\n\t\t\t: (() => {\n\t\t\t\t\tconst defaultBudgets: Record<ThinkingLevel, number> = {\n\t\t\t\t\t\tminimal: 1024,\n\t\t\t\t\t\tlow: 2048,\n\t\t\t\t\t\tmedium: 8192,\n\t\t\t\t\t\thigh: 16384,\n\t\t\t\t\t\txhigh: 16384, // Budget-based Claude clamps extended levels to high\n\t\t\t\t\t\tmax: 16384,\n\t\t\t\t\t};\n\n\t\t\t\t\t// Custom budgets only cover token-based levels through high.\n\t\t\t\t\tconst level = options.reasoning === \"xhigh\" || options.reasoning === \"max\" ? \"high\" : options.reasoning;\n\t\t\t\t\tconst budget = options.thinkingBudgets?.[level] ?? defaultBudgets[options.reasoning];\n\n\t\t\t\t\treturn {\n\t\t\t\t\t\tthinking: {\n\t\t\t\t\t\t\ttype: \"enabled\",\n\t\t\t\t\t\t\tbudget_tokens: budget,\n\t\t\t\t\t\t\t...(display !== undefined ? { display } : {}),\n\t\t\t\t\t\t},\n\t\t\t\t\t};\n\t\t\t\t})();\n\n\t\tif (!supportsAdaptiveThinking(model.id, model.name) && (options.interleavedThinking ?? true)) {\n\t\t\tresult.anthropic_beta = [\"interleaved-thinking-2025-05-14\"];\n\t\t}\n\n\t\treturn result;\n\t}\n\n\tif (isOpenAiGpt6AstraBedrockModel(model)) {\n\t\t// Bedrock maps OpenAI Chat Completions fields that are not part of Converse's\n\t\t// inferenceConfig into additionalModelRequestFields. GPT-6-Astra uses the Chat\n\t\t// Completions spelling rather than the Responses API's nested reasoning object.\n\t\t// https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-openai.html\n\t\treturn { reasoning_effort: mapThinkingLevelToEffort(model, options.reasoning) };\n\t}\n\n\treturn undefined;\n}\n\n/**\n * Build a Bedrock `DocumentBlock`. Three things differ from the Anthropic path:\n *\n * - `name` is **required**, and AWS restricts it to alphanumerics, single whitespace runs,\n * hyphens, parentheses, and square brackets — then warns the field \"is vulnerable to prompt\n * injections… we recommend that you specify a neutral name.\" Any caller-supplied name is\n * sanitized to that character set, and anything left empty falls back to a generated label.\n * - `source` takes **raw bytes**, not base64: \"If you use an Amazon Web Services SDK, you don't\n * need to encode the bytes in base64.\" So the base64 payload is decoded here, the opposite of\n * what the Anthropic converter does.\n * - Without citations enabled, Converse degrades to plain text extraction rather than full visual\n * PDF understanding. That limitation is documented rather than worked around here.\n * https://platform.claude.com/docs/en/build-with-claude/pdf-support\n *\n * `DocumentFormat` also has eight non-PDF members, but `format` is hardcoded to PDF because that\n * is the only media type `DocumentContent` declares — and the seven office/markup formats have no\n * Anthropic source variant to map onto. The block's own media type is therefore verified rather\n * than read.\n */\nfunction createDocumentBlock(block: DocumentContent, index: number) {\n\tassertSupportedDocumentMimeType(block);\n\tconst sanitized = (block.name ?? \"\")\n\t\t.replace(/[^a-zA-Z0-9\\s\\-()[\\]]/g, \" \")\n\t\t.replace(/\\s+/g, \" \")\n\t\t.trim();\n\treturn {\n\t\tformat: DocumentFormat.PDF,\n\t\tname: sanitized.length > 0 ? sanitized : `document ${index + 1}`,\n\t\tsource: { bytes: base64ToBytes(block.data) },\n\t};\n}\n\nfunction createImageBlock(mimeType: string, data: string) {\n\tlet format: ImageFormat;\n\tswitch (mimeType) {\n\t\tcase \"image/jpeg\":\n\t\tcase \"image/jpg\":\n\t\t\tformat = ImageFormat.JPEG;\n\t\t\tbreak;\n\t\tcase \"image/png\":\n\t\t\tformat = ImageFormat.PNG;\n\t\t\tbreak;\n\t\tcase \"image/gif\":\n\t\t\tformat = ImageFormat.GIF;\n\t\t\tbreak;\n\t\tcase \"image/webp\":\n\t\t\tformat = ImageFormat.WEBP;\n\t\t\tbreak;\n\t\tdefault:\n\t\t\tthrow new Error(`Unknown image type: ${mimeType}`);\n\t}\n\n\treturn { source: { bytes: base64ToBytes(data) }, format };\n}\n\nfunction base64ToBytes(data: string): Uint8Array {\n\tconst binaryString = atob(data);\n\tconst bytes = new Uint8Array(binaryString.length);\n\tfor (let i = 0; i < binaryString.length; i++) {\n\t\tbytes[i] = binaryString.charCodeAt(i);\n\t}\n\treturn bytes;\n}\n\n/**\n * Decodes a stored redacted payload. The AWS SDK hands the blob over as bytes, but a\n * persisted session carries it as base64. A hand-edited or externally produced session\n * can hold a signature that is not base64; drop that block instead of failing the\n * whole request.\n */\nfunction decodeRedactedContent(signature: string | undefined): Uint8Array | undefined {\n\tif (!signature) return undefined;\n\ttry {\n\t\treturn base64ToBytes(signature);\n\t} catch {\n\t\treturn undefined;\n\t}\n}\n\nfunction bytesToBase64(chunks: Uint8Array[]): string {\n\t// Encrypted reasoning runs to tens of KB, so build the binary string in slices\n\t// rather than one concatenation per byte. The window stays under the engine's\n\t// argument-count limit for spread calls.\n\tconst WINDOW = 0x8000;\n\tlet binary = \"\";\n\tfor (const chunk of chunks) {\n\t\tfor (let i = 0; i < chunk.length; i += WINDOW) {\n\t\t\tbinary += String.fromCharCode(...chunk.subarray(i, i + WINDOW));\n\t\t}\n\t}\n\treturn btoa(binary);\n}\n"]}
@@ -536,6 +536,7 @@ function handleMetadata(event, model, output) {
536
536
  output.usage.output = event.usage.outputTokens || 0;
537
537
  output.usage.cacheRead = event.usage.cacheReadInputTokens || 0;
538
538
  output.usage.cacheWrite = event.usage.cacheWriteInputTokens || 0;
539
+ output.usage.cacheWrite1h = event.usage.cacheDetails?.reduce((total, detail) => total + (detail.ttl === CacheTTL.ONE_HOUR ? (detail.inputTokens ?? 0) : 0), 0);
539
540
  output.usage.totalTokens = event.usage.totalTokens || output.usage.input + output.usage.output;
540
541
  calculateCost(model, output.usage);
541
542
  }