@kindgi/adapter-model-openai-compat 0.1.4-rc.5 → 0.1.5-rc.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +27 -6
- package/dist/embedding.d.ts +37 -0
- package/dist/embedding.d.ts.map +1 -0
- package/dist/embedding.js +61 -0
- package/dist/embedding.js.map +1 -0
- package/dist/index.d.ts +7 -2
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +4 -1
- package/dist/index.js.map +1 -1
- package/dist/provider.d.ts +46 -8
- package/dist/provider.d.ts.map +1 -1
- package/dist/provider.js +159 -43
- package/dist/provider.js.map +1 -1
- package/dist/responses.d.ts +52 -0
- package/dist/responses.d.ts.map +1 -0
- package/dist/responses.js +358 -0
- package/dist/responses.js.map +1 -0
- package/dist/wire.d.ts +73 -0
- package/dist/wire.d.ts.map +1 -0
- package/dist/wire.js +71 -0
- package/dist/wire.js.map +1 -0
- package/package.json +4 -3
- package/src/embedding.ts +108 -0
- package/src/index.ts +15 -1
- package/src/provider.ts +215 -56
- package/src/responses.ts +489 -0
- package/src/wire.ts +130 -0
package/README.md
CHANGED
|
@@ -1,16 +1,35 @@
|
|
|
1
1
|
# `@kindgi/adapter-model-openai-compat`
|
|
2
2
|
|
|
3
|
-
OpenAI-compatible `ModelProvider` for [`@kindgi/capabilities`](../../capabilities/). Wraps the official `openai` SDK client behind the framework's provider-neutral interface and works against any endpoint that speaks the Chat Completions wire format —
|
|
3
|
+
OpenAI-compatible `ModelProvider` for [`@kindgi/capabilities`](../../capabilities/). Wraps the official `openai` SDK client behind the framework's provider-neutral interface and works against OpenAI's own API and any endpoint that speaks the Chat Completions wire format — Ollama, vLLM, llama-server, LM Studio, OpenRouter, Groq, Together, Fireworks, DeepInfra, or a LiteLLM proxy. One adapter, many backends: `baseURL` selects the endpoint, and `api` the OpenAI API it speaks.
|
|
4
4
|
|
|
5
5
|
## Purpose
|
|
6
6
|
|
|
7
|
-
Translate between the framework's `ModelCallInput` / `ModelCallResult` and
|
|
7
|
+
Translate between the framework's `ModelCallInput` / `ModelCallResult` and one non-streaming call, on one of two OpenAI APIs (`api`):
|
|
8
|
+
|
|
9
|
+
- **`'responses'`** — OpenAI's Responses API (`responses.create`), the default for a `baseURL` on `api.openai.com` (and its data-residency hosts, such as `eu.api.openai.com`). OpenAI's GPT-6 models call tools only through it.
|
|
10
|
+
- **`'chat-completions'`** — Chat Completions (`chat.completions.create`), the default for every other `baseURL`: the format the OpenAI-compatible servers speak.
|
|
11
|
+
|
|
12
|
+
### Chat Completions
|
|
8
13
|
|
|
9
14
|
- Messages map role for role. Assistant `toolCalls` become `tool_calls` with JSON-encoded arguments; `tool` messages carry `tool_call_id`.
|
|
10
15
|
- `tools` become `function` tools; `structuredOutput` becomes `response_format: { type: 'json_schema', json_schema: { name, schema, strict: true } }`; `temperature` and `maxOutputTokens` (as `max_tokens`) are sent only when set.
|
|
11
16
|
- Chat Completions endpoints reject `.` in function names, so tool names are encoded `.` → `__` on send and decoded on receive (`acme.orders.lookup` ↔ `acme__orders__lookup`). Tool ids must not contain a literal `__`. Tool-call arguments that are not valid JSON decode to `{}`.
|
|
12
17
|
- `finish_reason` maps `tool_calls` / `function_call` → `tool-use`, `content_filter` → `content-filter`, `length` → `length`, anything else → `stop`.
|
|
13
|
-
- `costUsd` is computed from the matching `ModelInfo.cost
|
|
18
|
+
- `costUsd` is computed from the matching `ModelInfo.cost` (`computeCost`), on both APIs. The wire format carries no pricing, so rates come from the caller (`OpenAICompatCostRates`, the names the Gemini and Anthropic adapters price with):
|
|
19
|
+
- prompt and completion tokens per 1K tokens (`promptUsdPer1kTokens`, `completionUsdPer1kTokens`);
|
|
20
|
+
- cached prompt tokens at `cachedPromptMultiplier` of the prompt rate, and cache-write tokens at `promptCacheCreationMultiplier` (each absent: the prompt rate);
|
|
21
|
+
- `longContext: { thresholdTokens, promptUsdPer1kTokens, completionUsdPer1kTokens }`: past the threshold (cached and cache-write prompt tokens included), the whole call bills at those rates, the cache multipliers applying to the long prompt rate;
|
|
22
|
+
- `dataResidencyMultiplier`: an uplift on the whole call, only when `baseURL` is a data-residency host (`eu.api.openai.com`; `isDataResidencyHost`).
|
|
23
|
+
A rate that isn't a finite, non-negative number is ignored. The `openai` preset carries OpenAI's GPT-6 rates. Its `longContext` applies per call, counting all of the call's input tokens (cached and cache-write ones included) past 272,000: that's how we read OpenAI's pricing page, which doesn't spell it out. These costs are estimates from published prices; the provider's invoice is authoritative.
|
|
24
|
+
|
|
25
|
+
### Responses
|
|
26
|
+
|
|
27
|
+
- **Stateless:** every call sends the whole conversation with `store: false`, so OpenAI keeps no conversation state for it (no `previous_response_id`).
|
|
28
|
+
- **Input items:** `system` messages go as `developer` messages, `user` and `assistant` as messages, assistant `toolCalls` as `function_call` items, and `tool` messages as `function_call_output` items.
|
|
29
|
+
- **Tools and typed output:** `tools` become flat `function` tools (`strict: false`, the same `.` → `__` name encoding); `structuredOutput` becomes `text.format: { type: 'json_schema', name, schema, strict: true }`; `temperature` and `maxOutputTokens` (as `max_output_tokens`) are sent only when set.
|
|
30
|
+
- **The answer:** its `output_text` (its refusal when it has no text) and its `function_call` items. `incomplete` maps `max_output_tokens` → `length` and `content_filter` → `content-filter`; a `failed` response throws, naming OpenAI's error code and message.
|
|
31
|
+
- **Reasoning between tool calls:** a reasoning model's reasoning items (encrypted by OpenAI), with the item ids and `phase` that tie them to the message and calls after them, are carried in the first tool call's `ModelToolCall.signature` (opaque, versioned `oair1.`). When the conversation continues with the same model, they go back in their original order; for any other model, or a signature that isn't this adapter's, the turn goes back without them. A final answer's `phase` isn't carried (`ModelMessage` has no signature).
|
|
32
|
+
- **Usage:** `input_tokens` / `output_tokens`, with the cached, cache-write and reasoning parts reported apart when given.
|
|
14
33
|
|
|
15
34
|
## Exports
|
|
16
35
|
|
|
@@ -19,9 +38,11 @@ Translate between the framework's `ModelCallInput` / `ModelCallResult` and a non
|
|
|
19
38
|
- `baseURL: string` — endpoint root, e.g. `BASE_URLS.OLLAMA_LOCAL`.
|
|
20
39
|
- `apiKey: string | (() => string | Promise<string>)` — a static key, or a resolver called on every `invoke()`: a rotated key takes effect on the next call. Local runners that ignore keys still need a non-empty string such as `'unused'`.
|
|
21
40
|
- `metadata: ProviderMetadata` — surfaced to the router. `models[]` must list every model invoked through this connection, with its per-1K-token cost.
|
|
41
|
+
- `api?: 'responses' | 'chat-completions'` (`OPENAI_COMPAT_APIS`) — the OpenAI API to speak. Absent: `defaultOpenAICompatApi(baseURL)`, `'responses'` for `api.openai.com` and its subdomains, `'chat-completions'` for any other. Another value is refused at construction.
|
|
22
42
|
- `clientOptions?` — other `openai` client options (`timeout`, `maxRetries`, `defaultHeaders`, `fetch`, …); `apiKey` and `baseURL` are excluded.
|
|
23
|
-
- `extraBody?` — fields merged into every request body: settings an endpoint takes that the OpenAI format has no field for. A Qwen thinking model served by vLLM, SGLang or llama-server needs `{ chat_template_kwargs: { enable_thinking: false } }`, or its answer starts with its thinking and a typed (JSON) answer fails. The fields the adapter sets
|
|
24
|
-
- **`openAICompatAdapterFactory`** (`OPENAI_COMPAT_ADAPTER_ID`) — the `AdapterFactory` a runtime registers. A provider registration names its endpoint in `adapter_config.baseURL` (an http(s) URL; `openAICompatBaseUrl` reads and checks it), any extra request fields as `extraBody.*` keys (`adapter_config` is flat, so one key per field and dots nest: `"extraBody.chat_template_kwargs.enable_thinking": false`; `openAICompatExtraBody` expands and checks them), and, for an endpoint that needs a key, its secret in `secret_ref`; without one the adapter sends `'unused'`, as local runners expect.
|
|
43
|
+
- `extraBody?` — fields merged into every request body: settings an endpoint takes that the OpenAI format has no field for. A Qwen thinking model served by vLLM, SGLang or llama-server needs `{ chat_template_kwargs: { enable_thinking: false } }`, or its answer starts with its thinking and a typed (JSON) answer fails. The fields the adapter sets are refused: on Chat Completions `EXTRA_BODY_RESERVED` (`model`, `messages`, `tools`, `response_format`, `temperature`, `max_tokens`, `stream`), on Responses `EXTRA_BODY_RESERVED_RESPONSES` (`model`, `input`, `tools`, `text`, `temperature`, `max_output_tokens`, `stream`, `store`). On Responses, `{ reasoning: { effort: 'low' } }` sets a reasoning model's effort.
|
|
44
|
+
- **`openAICompatAdapterFactory`** (`OPENAI_COMPAT_ADAPTER_ID`) — the `AdapterFactory` a runtime registers. A provider registration names its endpoint in `adapter_config.baseURL` (an http(s) URL; `openAICompatBaseUrl` reads and checks it), the OpenAI API in `adapter_config.api` (`openAICompatApi` reads and checks it, with the base URL's default when absent), any extra request fields as `extraBody.*` keys (`adapter_config` is flat, so one key per field and dots nest: `"extraBody.chat_template_kwargs.enable_thinking": false`; `openAICompatExtraBody` expands and checks them), and, for an endpoint that needs a key, its secret in `secret_ref`; without one the adapter sends `'unused'`, as local runners expect.
|
|
45
|
+
- **`openAICompatAdapterEntry`** — the `AdapterFactoryEntry` a runtime registers: `openAICompatAdapterFactory` plus **`openAICompatCheckConfig(input)`**, which reports every problem the factory would throw on (`adapter_config.baseURL`, `adapter_config.api`, each `adapter_config.extraBody.*` key), with the factory's own messages, without building anything.
|
|
25
46
|
- **`BASE_URLS`** — well-known endpoints: `OPENAI`, `OLLAMA_LOCAL`, `VLLM_LOCAL`, `LLAMA_SERVER_LOCAL`, `LM_STUDIO_LOCAL`, `OPENROUTER`, `GROQ`, `TOGETHER`, `FIREWORKS`, `DEEPINFRA`, `LITELLM_LOCAL`. Any other URL works too.
|
|
26
47
|
- **`ModelProvider`** — type re-export from `@kindgi/capabilities`.
|
|
27
48
|
|
|
@@ -78,8 +99,8 @@ A local runner uses the same factory: `baseURL: BASE_URLS.OLLAMA_LOCAL`, `apiKey
|
|
|
78
99
|
## Non-goals
|
|
79
100
|
|
|
80
101
|
- **Streaming.** Requests are sent with `stream: false`; `invoke()` resolves with the complete response.
|
|
102
|
+
- **Server-side conversation state** (Responses `previous_response_id`, `store: true`). Every call carries its whole conversation.
|
|
81
103
|
- **Pricing discovery.** Cost rates are caller-supplied per model.
|
|
82
|
-
- **Cached-token accounting.** `usage` reports prompt and completion tokens only; all prompt tokens are billed at `promptUsdPer1kTokens`.
|
|
83
104
|
|
|
84
105
|
## Related
|
|
85
106
|
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
import OpenAI from 'openai';
|
|
2
|
+
import type { EmbeddingProvider } from '@kindgi/embedding';
|
|
3
|
+
/**
|
|
4
|
+
* An embeddings endpoint that speaks OpenAI's `POST /embeddings`: OpenAI
|
|
5
|
+
* itself, Ollama (`nomic-embed-text`), vLLM, Hugging Face TEI, LM Studio,
|
|
6
|
+
* llama-server, a LiteLLM proxy. Same `baseURL` as for chat
|
|
7
|
+
* (`BASE_URLS`), with the `/v1` suffix.
|
|
8
|
+
*/
|
|
9
|
+
export interface OpenAICompatEmbeddingOptions {
|
|
10
|
+
readonly baseURL: string;
|
|
11
|
+
/** The embedding model the endpoint serves (`text-embedding-3-small`, `nomic-embed-text`, …). */
|
|
12
|
+
readonly model: string;
|
|
13
|
+
/**
|
|
14
|
+
* The key, or a resolver called before each request (a rotated key
|
|
15
|
+
* takes effect on the next one). Absent: the endpoint takes none
|
|
16
|
+
* (Ollama, a local TEI); the SDK sends a placeholder.
|
|
17
|
+
*/
|
|
18
|
+
readonly apiKey?: string | (() => string | Promise<string>);
|
|
19
|
+
/** The vector length, when known; otherwise learned by `probe()` or the first `embed()`. */
|
|
20
|
+
readonly dimensions?: number;
|
|
21
|
+
/** HTTP overrides passed through to the OpenAI SDK (`fetch`, `timeout`, `maxRetries`, …). */
|
|
22
|
+
readonly clientOptions?: Omit<NonNullable<ConstructorParameters<typeof OpenAI>[0]>, 'apiKey' | 'baseURL'>;
|
|
23
|
+
}
|
|
24
|
+
export interface OpenAICompatEmbeddingProvider extends EmbeddingProvider {
|
|
25
|
+
/**
|
|
26
|
+
* Embed a short text once: checks the endpoint answers with this model,
|
|
27
|
+
* and learns the dimensions. Call it at boot, before `dimensions()`.
|
|
28
|
+
*/
|
|
29
|
+
probe(): Promise<number>;
|
|
30
|
+
}
|
|
31
|
+
/**
|
|
32
|
+
* Create an `EmbeddingProvider` over an OpenAI-compatible embeddings
|
|
33
|
+
* endpoint. The SDK client is built on first use (and rebuilt when the
|
|
34
|
+
* key rotates), so construction has no side effects.
|
|
35
|
+
*/
|
|
36
|
+
export declare function createOpenAICompatEmbeddingProvider(options: OpenAICompatEmbeddingOptions): OpenAICompatEmbeddingProvider;
|
|
37
|
+
//# sourceMappingURL=embedding.d.ts.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"embedding.d.ts","sourceRoot":"","sources":["../src/embedding.ts"],"names":[],"mappings":"AAGA,OAAO,MAAM,MAAM,QAAQ,CAAC;AAE5B,OAAO,KAAK,EAAE,iBAAiB,EAAE,MAAM,mBAAmB,CAAC;AAI3D;;;;;GAKG;AACH,MAAM,WAAW,4BAA4B;IAC3C,QAAQ,CAAC,OAAO,EAAE,MAAM,CAAC;IACzB,iGAAiG;IACjG,QAAQ,CAAC,KAAK,EAAE,MAAM,CAAC;IACvB;;;;OAIG;IACH,QAAQ,CAAC,MAAM,CAAC,EAAE,MAAM,GAAG,CAAC,MAAM,MAAM,GAAG,OAAO,CAAC,MAAM,CAAC,CAAC,CAAC;IAC5D,4FAA4F;IAC5F,QAAQ,CAAC,UAAU,CAAC,EAAE,MAAM,CAAC;IAC7B,6FAA6F;IAC7F,QAAQ,CAAC,aAAa,CAAC,EAAE,IAAI,CAC3B,WAAW,CAAC,qBAAqB,CAAC,OAAO,MAAM,CAAC,CAAC,CAAC,CAAC,CAAC,EACpD,QAAQ,GAAG,SAAS,CACrB,CAAC;CACH;AAED,MAAM,WAAW,6BAA8B,SAAQ,iBAAiB;IACtE;;;OAGG;IACH,KAAK,IAAI,OAAO,CAAC,MAAM,CAAC,CAAC;CAC1B;AAKD;;;;GAIG;AACH,wBAAgB,mCAAmC,CACjD,OAAO,EAAE,4BAA4B,GACpC,6BAA6B,CAuD/B"}
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
// Copyright (C) 2026 Kindgi Inc.
|
|
3
|
+
import OpenAI from 'openai';
|
|
4
|
+
import { OPENAI_COMPAT_ADAPTER_ID } from './provider.js';
|
|
5
|
+
/** The SDK needs a non-empty key; endpoints that take none ignore it. */
|
|
6
|
+
const NO_KEY = 'unused';
|
|
7
|
+
/**
|
|
8
|
+
* Create an `EmbeddingProvider` over an OpenAI-compatible embeddings
|
|
9
|
+
* endpoint. The SDK client is built on first use (and rebuilt when the
|
|
10
|
+
* key rotates), so construction has no side effects.
|
|
11
|
+
*/
|
|
12
|
+
export function createOpenAICompatEmbeddingProvider(options) {
|
|
13
|
+
let dims = options.dimensions;
|
|
14
|
+
let cached;
|
|
15
|
+
async function client() {
|
|
16
|
+
const resolved = typeof options.apiKey === 'function' ? await options.apiKey() : options.apiKey;
|
|
17
|
+
const key = resolved === undefined || resolved === '' ? NO_KEY : resolved;
|
|
18
|
+
if (cached !== undefined && cached.key === key)
|
|
19
|
+
return cached.client;
|
|
20
|
+
const fresh = new OpenAI({
|
|
21
|
+
apiKey: key,
|
|
22
|
+
baseURL: options.baseURL,
|
|
23
|
+
...(options.clientOptions ?? {}),
|
|
24
|
+
});
|
|
25
|
+
cached = { key, client: fresh };
|
|
26
|
+
return fresh;
|
|
27
|
+
}
|
|
28
|
+
async function embed(text) {
|
|
29
|
+
const response = await (await client()).embeddings.create({
|
|
30
|
+
model: options.model,
|
|
31
|
+
input: text,
|
|
32
|
+
encoding_format: 'float',
|
|
33
|
+
});
|
|
34
|
+
const values = response.data[0]?.embedding;
|
|
35
|
+
if (!Array.isArray(values) || values.length === 0) {
|
|
36
|
+
throw new Error(`${OPENAI_COMPAT_ADAPTER_ID}: the embeddings endpoint ${options.baseURL} returned no vector for model "${options.model}".`);
|
|
37
|
+
}
|
|
38
|
+
if (dims === undefined)
|
|
39
|
+
dims = values.length;
|
|
40
|
+
if (values.length !== dims) {
|
|
41
|
+
throw new Error(`${OPENAI_COMPAT_ADAPTER_ID}: model "${options.model}" returned a ${values.length}-dimension vector, expected ${dims}.`);
|
|
42
|
+
}
|
|
43
|
+
return Float32Array.from(values);
|
|
44
|
+
}
|
|
45
|
+
return {
|
|
46
|
+
embed,
|
|
47
|
+
async probe() {
|
|
48
|
+
return (await embed('dimension probe')).length;
|
|
49
|
+
},
|
|
50
|
+
dimensions() {
|
|
51
|
+
if (dims === undefined) {
|
|
52
|
+
throw new Error(`${OPENAI_COMPAT_ADAPTER_ID}: the dimensions of "${options.model}" aren't known yet; call probe() first.`);
|
|
53
|
+
}
|
|
54
|
+
return dims;
|
|
55
|
+
},
|
|
56
|
+
describe() {
|
|
57
|
+
return { name: OPENAI_COMPAT_ADAPTER_ID, version: '1', model: options.model };
|
|
58
|
+
},
|
|
59
|
+
};
|
|
60
|
+
}
|
|
61
|
+
//# sourceMappingURL=embedding.js.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"embedding.js","sourceRoot":"","sources":["../src/embedding.ts"],"names":[],"mappings":"AAAA,sCAAsC;AACtC,iCAAiC;AAEjC,OAAO,MAAM,MAAM,QAAQ,CAAC;AAI5B,OAAO,EAAE,wBAAwB,EAAE,MAAM,eAAe,CAAC;AAmCzD,yEAAyE;AACzE,MAAM,MAAM,GAAG,QAAQ,CAAC;AAExB;;;;GAIG;AACH,MAAM,UAAU,mCAAmC,CACjD,OAAqC;IAErC,IAAI,IAAI,GAAG,OAAO,CAAC,UAAU,CAAC;IAC9B,IAAI,MAAqE,CAAC;IAE1E,KAAK,UAAU,MAAM;QACnB,MAAM,QAAQ,GAAG,OAAO,OAAO,CAAC,MAAM,KAAK,UAAU,CAAC,CAAC,CAAC,MAAM,OAAO,CAAC,MAAM,EAAE,CAAC,CAAC,CAAC,OAAO,CAAC,MAAM,CAAC;QAChG,MAAM,GAAG,GAAG,QAAQ,KAAK,SAAS,IAAI,QAAQ,KAAK,EAAE,CAAC,CAAC,CAAC,MAAM,CAAC,CAAC,CAAC,QAAQ,CAAC;QAC1E,IAAI,MAAM,KAAK,SAAS,IAAI,MAAM,CAAC,GAAG,KAAK,GAAG;YAAE,OAAO,MAAM,CAAC,MAAM,CAAC;QACrE,MAAM,KAAK,GAAG,IAAI,MAAM,CAAC;YACvB,MAAM,EAAE,GAAG;YACX,OAAO,EAAE,OAAO,CAAC,OAAO;YACxB,GAAG,CAAC,OAAO,CAAC,aAAa,IAAI,EAAE,CAAC;SACjC,CAAC,CAAC;QACH,MAAM,GAAG,EAAE,GAAG,EAAE,MAAM,EAAE,KAAK,EAAE,CAAC;QAChC,OAAO,KAAK,CAAC;IACf,CAAC;IAED,KAAK,UAAU,KAAK,CAAC,IAAY;QAC/B,MAAM,QAAQ,GAAG,MAAM,CAAC,MAAM,MAAM,EAAE,CAAC,CAAC,UAAU,CAAC,MAAM,CAAC;YACxD,KAAK,EAAE,OAAO,CAAC,KAAK;YACpB,KAAK,EAAE,IAAI;YACX,eAAe,EAAE,OAAO;SACzB,CAAC,CAAC;QACH,MAAM,MAAM,GAAG,QAAQ,CAAC,IAAI,CAAC,CAAC,CAAC,EAAE,SAAS,CAAC;QAC3C,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,MAAM,CAAC,IAAI,MAAM,CAAC,MAAM,KAAK,CAAC,EAAE,CAAC;YAClD,MAAM,IAAI,KAAK,CACb,GAAG,wBAAwB,6BAA6B,OAAO,CAAC,OAAO,kCAAkC,OAAO,CAAC,KAAK,IAAI,CAC3H,CAAC;QACJ,CAAC;QACD,IAAI,IAAI,KAAK,SAAS;YAAE,IAAI,GAAG,MAAM,CAAC,MAAM,CAAC;QAC7C,IAAI,MAAM,CAAC,MAAM,KAAK,IAAI,EAAE,CAAC;YAC3B,MAAM,IAAI,KAAK,CACb,GAAG,wBAAwB,YAAY,OAAO,CAAC,KAAK,gBAAgB,MAAM,CAAC,MAAM,+BAA+B,IAAI,GAAG,CACxH,CAAC;QACJ,CAAC;QACD,OAAO,YAAY,CAAC,IAAI,CAAC,MAAM,CAAC,CAAC;IACnC,CAAC;IAED,OAAO;QACL,KAAK;QACL,KAAK,CAAC,KAAK;YACT,OAAO,CAAC,MAAM,KAAK,CAAC,iBAAiB,CAAC,CAAC,CAAC,MAAM,CAAC;QACjD,CAAC;QACD,UAAU;YACR,IAAI,IAAI,KAAK,SAAS,EAAE,CAAC;gBACvB,MAAM,IAAI,KAAK,CACb,GAAG,wBAAwB,wBAAwB,OAAO,CAAC,KAAK,yCAAyC,CAC1G,CAAC;YACJ,CAAC;YACD,OAAO,IAAI,CAAC;QACd,CAAC;QACD,QAAQ;YACN,OAAO,EAAE,IAAI,EAAE,wBAAwB,EAAE,OAAO,EAAE,GAAG,EAAE,KAAK,EAAE,OAAO,CAAC,KAAK,EAAE,CAAC;QAChF,CAAC;KACF,CAAC;AACJ,CAAC"}
|
package/dist/index.d.ts
CHANGED
|
@@ -1,4 +1,9 @@
|
|
|
1
|
-
export { BASE_URLS, EXTRA_BODY_PREFIX, EXTRA_BODY_RESERVED, OPENAI_COMPAT_ADAPTER_ID, createOpenAICompatModelProvider, openAICompatAdapterFactory, openAICompatBaseUrl, openAICompatExtraBody, } from './provider.js';
|
|
2
|
-
export
|
|
1
|
+
export { BASE_URLS, EXTRA_BODY_PREFIX, EXTRA_BODY_RESERVED, OPENAI_COMPAT_ADAPTER_ID, OPENAI_COMPAT_APIS, createOpenAICompatModelProvider, openAICompatAdapterEntry, openAICompatCheckConfig, defaultOpenAICompatApi, isDataResidencyHost, openAICompatAdapterFactory, openAICompatApi, openAICompatBaseUrl, openAICompatExtraBody, } from './provider.js';
|
|
2
|
+
export { EXTRA_BODY_RESERVED_RESPONSES } from './responses.js';
|
|
3
|
+
export { computeCost } from './wire.js';
|
|
4
|
+
export type { OpenAICompatCostRates, OpenAICompatModelInfo } from './wire.js';
|
|
5
|
+
export type { OpenAICompatApi, OpenAICompatProviderOptions } from './provider.js';
|
|
6
|
+
export { createOpenAICompatEmbeddingProvider } from './embedding.js';
|
|
7
|
+
export type { OpenAICompatEmbeddingOptions, OpenAICompatEmbeddingProvider, } from './embedding.js';
|
|
3
8
|
export type { ModelProvider } from '@kindgi/capabilities';
|
|
4
9
|
//# sourceMappingURL=index.d.ts.map
|
package/dist/index.d.ts.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../src/index.ts"],"names":[],"mappings":"AAGA,OAAO,EACL,SAAS,EACT,iBAAiB,EACjB,mBAAmB,EACnB,wBAAwB,EACxB,+BAA+B,EAC/B,0BAA0B,EAC1B,mBAAmB,EACnB,qBAAqB,GACtB,MAAM,eAAe,CAAC;AACvB,YAAY,EAAE,2BAA2B,EAAE,MAAM,eAAe,CAAC;
|
|
1
|
+
{"version":3,"file":"index.d.ts","sourceRoot":"","sources":["../src/index.ts"],"names":[],"mappings":"AAGA,OAAO,EACL,SAAS,EACT,iBAAiB,EACjB,mBAAmB,EACnB,wBAAwB,EACxB,kBAAkB,EAClB,+BAA+B,EAC/B,wBAAwB,EACxB,uBAAuB,EACvB,sBAAsB,EACtB,mBAAmB,EACnB,0BAA0B,EAC1B,eAAe,EACf,mBAAmB,EACnB,qBAAqB,GACtB,MAAM,eAAe,CAAC;AACvB,OAAO,EAAE,6BAA6B,EAAE,MAAM,gBAAgB,CAAC;AAC/D,OAAO,EAAE,WAAW,EAAE,MAAM,WAAW,CAAC;AACxC,YAAY,EAAE,qBAAqB,EAAE,qBAAqB,EAAE,MAAM,WAAW,CAAC;AAC9E,YAAY,EAAE,eAAe,EAAE,2BAA2B,EAAE,MAAM,eAAe,CAAC;AAClF,OAAO,EAAE,mCAAmC,EAAE,MAAM,gBAAgB,CAAC;AACrE,YAAY,EACV,4BAA4B,EAC5B,6BAA6B,GAC9B,MAAM,gBAAgB,CAAC;AACxB,YAAY,EAAE,aAAa,EAAE,MAAM,sBAAsB,CAAC"}
|
package/dist/index.js
CHANGED
|
@@ -1,4 +1,7 @@
|
|
|
1
1
|
// SPDX-License-Identifier: Apache-2.0
|
|
2
2
|
// Copyright (C) 2026 Kindgi Inc.
|
|
3
|
-
export { BASE_URLS, EXTRA_BODY_PREFIX, EXTRA_BODY_RESERVED, OPENAI_COMPAT_ADAPTER_ID, createOpenAICompatModelProvider, openAICompatAdapterFactory, openAICompatBaseUrl, openAICompatExtraBody, } from './provider.js';
|
|
3
|
+
export { BASE_URLS, EXTRA_BODY_PREFIX, EXTRA_BODY_RESERVED, OPENAI_COMPAT_ADAPTER_ID, OPENAI_COMPAT_APIS, createOpenAICompatModelProvider, openAICompatAdapterEntry, openAICompatCheckConfig, defaultOpenAICompatApi, isDataResidencyHost, openAICompatAdapterFactory, openAICompatApi, openAICompatBaseUrl, openAICompatExtraBody, } from './provider.js';
|
|
4
|
+
export { EXTRA_BODY_RESERVED_RESPONSES } from './responses.js';
|
|
5
|
+
export { computeCost } from './wire.js';
|
|
6
|
+
export { createOpenAICompatEmbeddingProvider } from './embedding.js';
|
|
4
7
|
//# sourceMappingURL=index.js.map
|
package/dist/index.js.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"index.js","sourceRoot":"","sources":["../src/index.ts"],"names":[],"mappings":"AAAA,sCAAsC;AACtC,iCAAiC;AAEjC,OAAO,EACL,SAAS,EACT,iBAAiB,EACjB,mBAAmB,EACnB,wBAAwB,EACxB,+BAA+B,EAC/B,0BAA0B,EAC1B,mBAAmB,EACnB,qBAAqB,GACtB,MAAM,eAAe,CAAC"}
|
|
1
|
+
{"version":3,"file":"index.js","sourceRoot":"","sources":["../src/index.ts"],"names":[],"mappings":"AAAA,sCAAsC;AACtC,iCAAiC;AAEjC,OAAO,EACL,SAAS,EACT,iBAAiB,EACjB,mBAAmB,EACnB,wBAAwB,EACxB,kBAAkB,EAClB,+BAA+B,EAC/B,wBAAwB,EACxB,uBAAuB,EACvB,sBAAsB,EACtB,mBAAmB,EACnB,0BAA0B,EAC1B,eAAe,EACf,mBAAmB,EACnB,qBAAqB,GACtB,MAAM,eAAe,CAAC;AACvB,OAAO,EAAE,6BAA6B,EAAE,MAAM,gBAAgB,CAAC;AAC/D,OAAO,EAAE,WAAW,EAAE,MAAM,WAAW,CAAC;AAGxC,OAAO,EAAE,mCAAmC,EAAE,MAAM,gBAAgB,CAAC"}
|
package/dist/provider.d.ts
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import OpenAI from 'openai';
|
|
2
|
-
import type { AdapterFactory, AdapterFactoryInput, Feature, ModelProvider, ProviderMetadata } from '@kindgi/capabilities';
|
|
2
|
+
import type { AdapterConfigCheckInput, AdapterConfigProblem, AdapterFactory, AdapterFactoryEntry, AdapterFactoryInput, Feature, ModelProvider, ProviderMetadata } from '@kindgi/capabilities';
|
|
3
3
|
/**
|
|
4
4
|
* Configuration for an OpenAI-compatible ModelProvider. `baseURL` is the
|
|
5
5
|
* key — same adapter, different endpoints (the well-known ones are
|
|
@@ -37,16 +37,38 @@ export interface OpenAICompatProviderOptions {
|
|
|
37
37
|
/** Optional HTTP overrides passed through to the OpenAI SDK. */
|
|
38
38
|
readonly clientOptions?: Omit<NonNullable<ConstructorParameters<typeof OpenAI>[0]>, 'apiKey' | 'baseURL'>;
|
|
39
39
|
/**
|
|
40
|
-
*
|
|
41
|
-
*
|
|
42
|
-
*
|
|
43
|
-
*
|
|
44
|
-
* `
|
|
45
|
-
*
|
|
40
|
+
* Which OpenAI API the adapter speaks (`OPENAI_COMPAT_APIS`):
|
|
41
|
+
* - `'responses'`: OpenAI's Responses API, the one OpenAI's GPT-6
|
|
42
|
+
* models call tools through. Stateless: every call sends
|
|
43
|
+
* `store: false`, so OpenAI keeps no conversation state.
|
|
44
|
+
* - `'chat-completions'`: Chat Completions, the format the
|
|
45
|
+
* OpenAI-compatible servers (Ollama, vLLM, Groq, OpenRouter, …) speak.
|
|
46
|
+
* Absent: `'responses'` for a `baseURL` on `api.openai.com` (or one of
|
|
47
|
+
* its data-residency hosts, `eu.api.openai.com`), `'chat-completions'`
|
|
48
|
+
* for any other (`defaultOpenAICompatApi`).
|
|
49
|
+
*/
|
|
50
|
+
readonly api?: OpenAICompatApi;
|
|
51
|
+
/**
|
|
52
|
+
* Fields merged into every request body: settings an endpoint takes
|
|
53
|
+
* that the OpenAI format has no field for. A Qwen thinking model served
|
|
54
|
+
* by vLLM, SGLang or llama-server answers with its thinking first unless
|
|
55
|
+
* asked not to: `{ chat_template_kwargs: { enable_thinking: false } }`.
|
|
56
|
+
* The fields the adapter sets itself are refused: `EXTRA_BODY_RESERVED`
|
|
57
|
+
* on Chat Completions, `EXTRA_BODY_RESERVED_RESPONSES` on Responses.
|
|
46
58
|
*/
|
|
47
59
|
readonly extraBody?: Readonly<Record<string, unknown>>;
|
|
48
60
|
}
|
|
49
|
-
/**
|
|
61
|
+
/** The OpenAI APIs the adapter speaks (`OpenAICompatProviderOptions.api`). */
|
|
62
|
+
export declare const OPENAI_COMPAT_APIS: readonly ["responses", "chat-completions"];
|
|
63
|
+
export type OpenAICompatApi = (typeof OPENAI_COMPAT_APIS)[number];
|
|
64
|
+
/** The API a provider speaks when it doesn't say (`OpenAICompatProviderOptions.api`). */
|
|
65
|
+
export declare function defaultOpenAICompatApi(baseURL: string): OpenAICompatApi;
|
|
66
|
+
/**
|
|
67
|
+
* Whether a base URL is one of OpenAI's data-residency hosts
|
|
68
|
+
* (`eu.api.openai.com`), where a model's `dataResidencyMultiplier` applies.
|
|
69
|
+
*/
|
|
70
|
+
export declare function isDataResidencyHost(baseURL: string): boolean;
|
|
71
|
+
/** Request fields the adapter sets itself on Chat Completions; `extraBody` can't override them. */
|
|
50
72
|
export declare const EXTRA_BODY_RESERVED: readonly ["model", "messages", "tools", "response_format", "temperature", "max_tokens", "stream"];
|
|
51
73
|
/**
|
|
52
74
|
* Create an OpenAI-compatible `ModelProvider`. The OpenAI SDK is
|
|
@@ -83,10 +105,26 @@ export declare const OPENAI_COMPAT_ADAPTER_ID = "@kindgi/adapter-model-openai-co
|
|
|
83
105
|
* a placeholder key, as local runners (Ollama, vLLM, llama-server)
|
|
84
106
|
* expect. Its `extraBody.*` keys (`EXTRA_BODY_PREFIX`) are extra request
|
|
85
107
|
* fields, merged into every request (`OpenAICompatProviderOptions.extraBody`).
|
|
108
|
+
* Its `adapter_config.api` picks the OpenAI API (`openAICompatApi`).
|
|
86
109
|
*/
|
|
87
110
|
export declare const openAICompatAdapterFactory: AdapterFactory;
|
|
111
|
+
/**
|
|
112
|
+
* What's wrong with a registration for this adapter, read without building
|
|
113
|
+
* it (`AdapterFactoryEntry.checkConfig`): its base URL, its API and its
|
|
114
|
+
* extra request fields. Each problem's message is the error the factory
|
|
115
|
+
* throws for it: both read the registration through the same functions.
|
|
116
|
+
*/
|
|
117
|
+
export declare function openAICompatCheckConfig(input: AdapterConfigCheckInput): readonly AdapterConfigProblem[];
|
|
118
|
+
/** The entry a runtime registers: the factory, and its static check. */
|
|
119
|
+
export declare const openAICompatAdapterEntry: AdapterFactoryEntry;
|
|
88
120
|
/** The endpoint a registration names (see `openAICompatAdapterFactory`). */
|
|
89
121
|
export declare function openAICompatBaseUrl(input: AdapterFactoryInput): string;
|
|
122
|
+
/**
|
|
123
|
+
* The API a registration speaks: its `adapter_config.api` (`OPENAI_COMPAT_APIS`),
|
|
124
|
+
* or without one, the default for its base URL (`defaultOpenAICompatApi`).
|
|
125
|
+
* Throws, naming the key, on an API the adapter doesn't speak.
|
|
126
|
+
*/
|
|
127
|
+
export declare function openAICompatApi(input: AdapterFactoryInput): OpenAICompatApi;
|
|
90
128
|
/**
|
|
91
129
|
* The `adapter_config` key prefix for extra request fields. `adapter_config`
|
|
92
130
|
* is flat, so each field is its own key and dots nest:
|
package/dist/provider.d.ts.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"provider.d.ts","sourceRoot":"","sources":["../src/provider.ts"],"names":[],"mappings":"AAGA,OAAO,MAAM,MAAM,QAAQ,CAAC;AAE5B,OAAO,KAAK,EACV,cAAc,EACd,mBAAmB,EACnB,OAAO,EAKP,aAAa,EAGb,gBAAgB,EAEjB,MAAM,sBAAsB,CAAC;
|
|
1
|
+
{"version":3,"file":"provider.d.ts","sourceRoot":"","sources":["../src/provider.ts"],"names":[],"mappings":"AAGA,OAAO,MAAM,MAAM,QAAQ,CAAC;AAE5B,OAAO,KAAK,EACV,uBAAuB,EACvB,oBAAoB,EACpB,cAAc,EACd,mBAAmB,EACnB,mBAAmB,EACnB,OAAO,EAKP,aAAa,EAGb,gBAAgB,EAEjB,MAAM,sBAAsB,CAAC;AAQ9B;;;;;;;;;;;;;;;;;;;;;;;;;;;;GA4BG;AACH,MAAM,WAAW,2BAA2B;IAC1C,QAAQ,CAAC,OAAO,EAAE,MAAM,CAAC;IACzB,QAAQ,CAAC,MAAM,EAAE,MAAM,GAAG,CAAC,MAAM,MAAM,GAAG,OAAO,CAAC,MAAM,CAAC,CAAC,CAAC;IAC3D,kDAAkD;IAClD,QAAQ,CAAC,QAAQ,EAAE,gBAAgB,CAAC;IACpC,gEAAgE;IAChE,QAAQ,CAAC,aAAa,CAAC,EAAE,IAAI,CAC3B,WAAW,CAAC,qBAAqB,CAAC,OAAO,MAAM,CAAC,CAAC,CAAC,CAAC,CAAC,EACpD,QAAQ,GAAG,SAAS,CACrB,CAAC;IACF;;;;;;;;;;OAUG;IACH,QAAQ,CAAC,GAAG,CAAC,EAAE,eAAe,CAAC;IAC/B;;;;;;;OAOG;IACH,QAAQ,CAAC,SAAS,CAAC,EAAE,QAAQ,CAAC,MAAM,CAAC,MAAM,EAAE,OAAO,CAAC,CAAC,CAAC;CACxD;AAED,8EAA8E;AAC9E,eAAO,MAAM,kBAAkB,4CAA6C,CAAC;AAC7E,MAAM,MAAM,eAAe,GAAG,CAAC,OAAO,kBAAkB,CAAC,CAAC,MAAM,CAAC,CAAC;AAElE,yFAAyF;AACzF,wBAAgB,sBAAsB,CAAC,OAAO,EAAE,MAAM,GAAG,eAAe,CAGvE;AAED;;;GAGG;AACH,wBAAgB,mBAAmB,CAAC,OAAO,EAAE,MAAM,GAAG,OAAO,CAE5D;AAaD,mGAAmG;AACnG,eAAO,MAAM,mBAAmB,mGAQtB,CAAC;AAEX;;;;GAIG;AACH,wBAAgB,+BAA+B,CAC7C,OAAO,EAAE,2BAA2B,GACnC,aAAa,CAkJf;AA8ED;;;;GAIG;AACH,eAAO,MAAM,SAAS;;;;;;;;;;;;CAYZ,CAAC;AAEX,MAAM,MAAM,QAAQ,GAAG,OAAO,CAAC;AAE/B,uDAAuD;AACvD,eAAO,MAAM,wBAAwB,wCAAwC,CAAC;AAK9E;;;;;;;;;GASG;AACH,eAAO,MAAM,0BAA0B,EAAE,cAYxC,CAAC;AAEF;;;;;GAKG;AACH,wBAAgB,uBAAuB,CACrC,KAAK,EAAE,uBAAuB,GAC7B,SAAS,oBAAoB,EAAE,CAIjC;AAED,wEAAwE;AACxE,eAAO,MAAM,wBAAwB,EAAE,mBAKtC,CAAC;AAWF,4EAA4E;AAC5E,wBAAgB,mBAAmB,CAAC,KAAK,EAAE,mBAAmB,GAAG,MAAM,CAGtE;AAWD;;;;GAIG;AACH,wBAAgB,eAAe,CAAC,KAAK,EAAE,mBAAmB,GAAG,eAAe,CAG3E;AAmBD;;;;;GAKG;AACH,eAAO,MAAM,iBAAiB,eAAe,CAAC;AAK9C;;;;;GAKG;AACH,wBAAgB,qBAAqB,CACnC,KAAK,EAAE,mBAAmB,GACzB,QAAQ,CAAC,MAAM,CAAC,MAAM,EAAE,OAAO,CAAC,CAAC,GAAG,SAAS,CAI/C"}
|
package/dist/provider.js
CHANGED
|
@@ -1,26 +1,38 @@
|
|
|
1
1
|
// SPDX-License-Identifier: Apache-2.0
|
|
2
2
|
// Copyright (C) 2026 Kindgi Inc.
|
|
3
3
|
import OpenAI from 'openai';
|
|
4
|
+
import { adapterConfigError, samplingFor } from '@kindgi/capabilities';
|
|
4
5
|
import { createAttemptCounter } from '@kindgi/capabilities/attempts';
|
|
6
|
+
import { nameToolsAsSent } from '@kindgi/capabilities/tool-names';
|
|
7
|
+
import { EXTRA_BODY_RESERVED_RESPONSES, invokeResponses, requestOptions } from './responses.js';
|
|
8
|
+
import { computeCost, decodeToolName, encodeToolName } from './wire.js';
|
|
9
|
+
/** The OpenAI APIs the adapter speaks (`OpenAICompatProviderOptions.api`). */
|
|
10
|
+
export const OPENAI_COMPAT_APIS = ['responses', 'chat-completions'];
|
|
11
|
+
/** The API a provider speaks when it doesn't say (`OpenAICompatProviderOptions.api`). */
|
|
12
|
+
export function defaultOpenAICompatApi(baseURL) {
|
|
13
|
+
// OpenAI's own API, its data-residency hosts (`eu.api.openai.com`) included.
|
|
14
|
+
return openAIHostOf(baseURL) !== undefined ? 'responses' : 'chat-completions';
|
|
15
|
+
}
|
|
5
16
|
/**
|
|
6
|
-
*
|
|
7
|
-
*
|
|
8
|
-
* `function.name` to `^[a-zA-Z0-9_-]{1,128}$` — dots are rejected.
|
|
9
|
-
* The framework's tool id convention is `<pack>.<tool>`, so this
|
|
10
|
-
* adapter transparently encodes on send and decodes on receive.
|
|
11
|
-
*
|
|
12
|
-
* See the sibling comment in
|
|
13
|
-
* `packages/adapters/model-anthropic/src/translate.ts` — same
|
|
14
|
-
* substitution (`.` → `__`), same reversibility caveat (authors
|
|
15
|
-
* should not put literal `__` in tool ids).
|
|
17
|
+
* Whether a base URL is one of OpenAI's data-residency hosts
|
|
18
|
+
* (`eu.api.openai.com`), where a model's `dataResidencyMultiplier` applies.
|
|
16
19
|
*/
|
|
17
|
-
function
|
|
18
|
-
return
|
|
20
|
+
export function isDataResidencyHost(baseURL) {
|
|
21
|
+
return openAIHostOf(baseURL) === 'data-residency';
|
|
19
22
|
}
|
|
20
|
-
function
|
|
21
|
-
|
|
23
|
+
function openAIHostOf(baseURL) {
|
|
24
|
+
let host;
|
|
25
|
+
try {
|
|
26
|
+
host = new URL(baseURL).hostname;
|
|
27
|
+
}
|
|
28
|
+
catch {
|
|
29
|
+
return undefined;
|
|
30
|
+
}
|
|
31
|
+
if (host === 'api.openai.com')
|
|
32
|
+
return 'global';
|
|
33
|
+
return host.endsWith('.api.openai.com') ? 'data-residency' : undefined;
|
|
22
34
|
}
|
|
23
|
-
/** Request fields the adapter sets itself; `extraBody` can't override them. */
|
|
35
|
+
/** Request fields the adapter sets itself on Chat Completions; `extraBody` can't override them. */
|
|
24
36
|
export const EXTRA_BODY_RESERVED = [
|
|
25
37
|
'model',
|
|
26
38
|
'messages',
|
|
@@ -37,7 +49,13 @@ export const EXTRA_BODY_RESERVED = [
|
|
|
37
49
|
*/
|
|
38
50
|
export function createOpenAICompatModelProvider(options) {
|
|
39
51
|
const metadata = options.metadata;
|
|
40
|
-
|
|
52
|
+
if (options.api !== undefined && !OPENAI_COMPAT_APIS.includes(options.api)) {
|
|
53
|
+
// A caller outside TypeScript can pass anything.
|
|
54
|
+
throw new Error(`${OPENAI_COMPAT_ADAPTER_ID}: provider "${metadata.id}": api must be one of ${OPENAI_COMPAT_APIS.join(', ')}`);
|
|
55
|
+
}
|
|
56
|
+
const api = options.api ?? defaultOpenAICompatApi(options.baseURL);
|
|
57
|
+
const dataResidency = isDataResidencyHost(options.baseURL);
|
|
58
|
+
const problem = extraBodyProblem(options.extraBody, api);
|
|
41
59
|
if (problem !== undefined) {
|
|
42
60
|
throw new Error(`${OPENAI_COMPAT_ADAPTER_ID}: provider "${metadata.id}": ${problem}`);
|
|
43
61
|
}
|
|
@@ -72,7 +90,33 @@ export function createOpenAICompatModelProvider(options) {
|
|
|
72
90
|
}
|
|
73
91
|
const startedAt = Date.now();
|
|
74
92
|
const openai = await clientForCall();
|
|
75
|
-
|
|
93
|
+
// The system prompt names the call's tools as they're sent (`acme__lookup_order`):
|
|
94
|
+
// a model told to call `acme.lookup_order` calls a name it wasn't given.
|
|
95
|
+
const toolIds = input.tools?.map((t) => t.name) ?? [];
|
|
96
|
+
const sent = input.messages.map((m) => m.role === 'system'
|
|
97
|
+
? { ...m, content: nameToolsAsSent(m.content, toolIds, encodeToolName) }
|
|
98
|
+
: m);
|
|
99
|
+
const sampling = samplingFor(modelInfo, input);
|
|
100
|
+
// A caller that wants as little thinking as the model allows (a judge).
|
|
101
|
+
const lowestThinking = input.thinking === 'lowest' && modelInfo.thinking !== undefined
|
|
102
|
+
? modelInfo.thinking.lowest
|
|
103
|
+
: undefined;
|
|
104
|
+
if (api === 'responses') {
|
|
105
|
+
return invokeResponses({
|
|
106
|
+
client: openai,
|
|
107
|
+
attempts,
|
|
108
|
+
input: { ...input, messages: sent },
|
|
109
|
+
modelInfo,
|
|
110
|
+
providerId: metadata.id,
|
|
111
|
+
extraBody,
|
|
112
|
+
dataResidency,
|
|
113
|
+
...(sampling.temperature !== undefined && { temperature: sampling.temperature }),
|
|
114
|
+
...(lowestThinking !== undefined && { reasoningEffort: lowestThinking }),
|
|
115
|
+
warnings: sampling.warnings,
|
|
116
|
+
startedAt,
|
|
117
|
+
});
|
|
118
|
+
}
|
|
119
|
+
const messages = sent.map(toOpenAiMessage);
|
|
76
120
|
const tools = input.tools?.map(toOpenAiTool);
|
|
77
121
|
const responseFormat = input.structuredOutput
|
|
78
122
|
? {
|
|
@@ -90,10 +134,13 @@ export function createOpenAICompatModelProvider(options) {
|
|
|
90
134
|
messages,
|
|
91
135
|
...(tools !== undefined && tools.length > 0 && { tools }),
|
|
92
136
|
...(responseFormat !== undefined && { response_format: responseFormat }),
|
|
93
|
-
...(
|
|
137
|
+
...(sampling.temperature !== undefined && { temperature: sampling.temperature }),
|
|
138
|
+
...(lowestThinking !== undefined && {
|
|
139
|
+
reasoning_effort: lowestThinking,
|
|
140
|
+
}),
|
|
94
141
|
...(input.maxOutputTokens !== undefined && { max_tokens: input.maxOutputTokens }),
|
|
95
142
|
stream: false,
|
|
96
|
-
}, input
|
|
143
|
+
}, requestOptions(input)));
|
|
97
144
|
const completion = counted.value;
|
|
98
145
|
const durationMs = Date.now() - startedAt;
|
|
99
146
|
const choice = completion.choices[0];
|
|
@@ -109,7 +156,7 @@ export function createOpenAICompatModelProvider(options) {
|
|
|
109
156
|
message: responseMessage,
|
|
110
157
|
finishReason: mapFinishReason(choice?.finish_reason),
|
|
111
158
|
usage,
|
|
112
|
-
costUsd: computeCost(modelInfo, usage
|
|
159
|
+
costUsd: computeCost(modelInfo, usage, { dataResidency }),
|
|
113
160
|
durationMs,
|
|
114
161
|
provider: { id: metadata.id, model: input.model },
|
|
115
162
|
...(completion.model !== undefined &&
|
|
@@ -122,6 +169,7 @@ export function createOpenAICompatModelProvider(options) {
|
|
|
122
169
|
// An injected client sends with its own fetch: nothing was counted.
|
|
123
170
|
...(counted.attempts > 0 && { attempts: counted.attempts }),
|
|
124
171
|
...(completion.usage !== undefined && { rawUsage: { ...completion.usage } }),
|
|
172
|
+
...(sampling.warnings.length > 0 && { warnings: sampling.warnings }),
|
|
125
173
|
};
|
|
126
174
|
},
|
|
127
175
|
};
|
|
@@ -201,11 +249,6 @@ function mapFinishReason(reason) {
|
|
|
201
249
|
return 'stop';
|
|
202
250
|
}
|
|
203
251
|
}
|
|
204
|
-
function computeCost(modelInfo, promptTokens, completionTokens) {
|
|
205
|
-
const promptCost = (promptTokens / 1000) * modelInfo.cost.promptUsdPer1kTokens;
|
|
206
|
-
const completionCost = (completionTokens / 1000) * modelInfo.cost.completionUsdPer1kTokens;
|
|
207
|
-
return promptCost + completionCost;
|
|
208
|
-
}
|
|
209
252
|
/**
|
|
210
253
|
* Well-known base URLs — surfaced as constants so callers can import
|
|
211
254
|
* without typos. Adding a new alias here doesn't lock anyone in; the
|
|
@@ -236,12 +279,14 @@ const NO_KEY = 'unused';
|
|
|
236
279
|
* a placeholder key, as local runners (Ollama, vLLM, llama-server)
|
|
237
280
|
* expect. Its `extraBody.*` keys (`EXTRA_BODY_PREFIX`) are extra request
|
|
238
281
|
* fields, merged into every request (`OpenAICompatProviderOptions.extraBody`).
|
|
282
|
+
* Its `adapter_config.api` picks the OpenAI API (`openAICompatApi`).
|
|
239
283
|
*/
|
|
240
284
|
export const openAICompatAdapterFactory = (input) => {
|
|
241
285
|
const extraBody = openAICompatExtraBody(input);
|
|
242
286
|
return createOpenAICompatModelProvider({
|
|
243
287
|
metadata: input.metadata,
|
|
244
288
|
baseURL: openAICompatBaseUrl(input),
|
|
289
|
+
api: openAICompatApi(input),
|
|
245
290
|
apiKey: input.resolveApiKey ?? NO_KEY,
|
|
246
291
|
...(extraBody !== undefined && { extraBody }),
|
|
247
292
|
// The registration chose the endpoint: the runtime's fetch decides
|
|
@@ -249,13 +294,66 @@ export const openAICompatAdapterFactory = (input) => {
|
|
|
249
294
|
...(input.fetch !== undefined && { clientOptions: { fetch: input.fetch } }),
|
|
250
295
|
});
|
|
251
296
|
};
|
|
297
|
+
/**
|
|
298
|
+
* What's wrong with a registration for this adapter, read without building
|
|
299
|
+
* it (`AdapterFactoryEntry.checkConfig`): its base URL, its API and its
|
|
300
|
+
* extra request fields. Each problem's message is the error the factory
|
|
301
|
+
* throws for it: both read the registration through the same functions.
|
|
302
|
+
*/
|
|
303
|
+
export function openAICompatCheckConfig(input) {
|
|
304
|
+
return [baseUrlProblem(input), apiProblem(input), ...readExtraBody(input).problems].filter((problem) => problem !== undefined);
|
|
305
|
+
}
|
|
306
|
+
/** The entry a runtime registers: the factory, and its static check. */
|
|
307
|
+
export const openAICompatAdapterEntry = {
|
|
308
|
+
adapterId: OPENAI_COMPAT_ADAPTER_ID,
|
|
309
|
+
capabilityKind: 'llm-inference',
|
|
310
|
+
factory: openAICompatAdapterFactory,
|
|
311
|
+
checkConfig: openAICompatCheckConfig,
|
|
312
|
+
};
|
|
313
|
+
function throwIf(input, problem) {
|
|
314
|
+
if (problem !== undefined) {
|
|
315
|
+
throw adapterConfigError(OPENAI_COMPAT_ADAPTER_ID, input.metadata.id, problem);
|
|
316
|
+
}
|
|
317
|
+
}
|
|
252
318
|
/** The endpoint a registration names (see `openAICompatAdapterFactory`). */
|
|
253
319
|
export function openAICompatBaseUrl(input) {
|
|
320
|
+
throwIf(input, baseUrlProblem(input));
|
|
321
|
+
return input.config?.baseURL;
|
|
322
|
+
}
|
|
323
|
+
function baseUrlProblem(input) {
|
|
254
324
|
const value = input.config?.baseURL;
|
|
255
|
-
if (typeof value
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
325
|
+
if (typeof value === 'string' && /^https?:\/\/[^\s/]+/.test(value))
|
|
326
|
+
return undefined;
|
|
327
|
+
return {
|
|
328
|
+
path: '/adapter_config/baseURL',
|
|
329
|
+
message: `needs adapter_config.baseURL, an http(s) URL (e.g. ${BASE_URLS.OPENAI}).`,
|
|
330
|
+
};
|
|
331
|
+
}
|
|
332
|
+
/**
|
|
333
|
+
* The API a registration speaks: its `adapter_config.api` (`OPENAI_COMPAT_APIS`),
|
|
334
|
+
* or without one, the default for its base URL (`defaultOpenAICompatApi`).
|
|
335
|
+
* Throws, naming the key, on an API the adapter doesn't speak.
|
|
336
|
+
*/
|
|
337
|
+
export function openAICompatApi(input) {
|
|
338
|
+
throwIf(input, apiProblem(input));
|
|
339
|
+
return apiOf(input);
|
|
340
|
+
}
|
|
341
|
+
function apiProblem(input) {
|
|
342
|
+
const value = input.config?.api;
|
|
343
|
+
if (value === undefined || OPENAI_COMPAT_APIS.some((known) => known === value))
|
|
344
|
+
return undefined;
|
|
345
|
+
return {
|
|
346
|
+
path: '/adapter_config/api',
|
|
347
|
+
message: `adapter_config.api must be one of ${OPENAI_COMPAT_APIS.join(', ')}.`,
|
|
348
|
+
};
|
|
349
|
+
}
|
|
350
|
+
/** The registration's API; its base URL's default when `api` is absent or not one the adapter speaks. */
|
|
351
|
+
function apiOf(input) {
|
|
352
|
+
const api = OPENAI_COMPAT_APIS.find((known) => known === input.config?.api);
|
|
353
|
+
if (api !== undefined)
|
|
354
|
+
return api;
|
|
355
|
+
const baseURL = input.config?.baseURL;
|
|
356
|
+
return defaultOpenAICompatApi(typeof baseURL === 'string' ? baseURL : '');
|
|
259
357
|
}
|
|
260
358
|
/**
|
|
261
359
|
* The `adapter_config` key prefix for extra request fields. `adapter_config`
|
|
@@ -273,12 +371,17 @@ const UNSAFE_FIELDS = new Set(['__proto__', 'constructor', 'prototype']);
|
|
|
273
371
|
* keys that collide, or a field the adapter sets.
|
|
274
372
|
*/
|
|
275
373
|
export function openAICompatExtraBody(input) {
|
|
276
|
-
const
|
|
277
|
-
|
|
278
|
-
|
|
374
|
+
const { body, problems } = readExtraBody(input);
|
|
375
|
+
throwIf(input, problems[0]);
|
|
376
|
+
return body;
|
|
377
|
+
}
|
|
378
|
+
/** A registration's extra request fields, and what's wrong with them (`openAICompatExtraBody`). */
|
|
379
|
+
function readExtraBody(input) {
|
|
380
|
+
const problems = [];
|
|
381
|
+
const at = (path, problem) => problems.push({ path, message: problem });
|
|
279
382
|
const config = input.config ?? {};
|
|
280
383
|
if (Object.hasOwn(config, 'extraBody')) {
|
|
281
|
-
|
|
384
|
+
at('/adapter_config/extraBody', `adapter_config takes extra request fields one per key, "${EXTRA_BODY_PREFIX}<field>" (dots nest), e.g. "${EXTRA_BODY_PREFIX}chat_template_kwargs.enable_thinking": false`);
|
|
282
385
|
}
|
|
283
386
|
let body;
|
|
284
387
|
for (const [key, value] of Object.entries(config)) {
|
|
@@ -287,12 +390,17 @@ export function openAICompatExtraBody(input) {
|
|
|
287
390
|
body ??= {};
|
|
288
391
|
const problem = setField(body, key.slice(EXTRA_BODY_PREFIX.length).split('.'), value);
|
|
289
392
|
if (problem !== undefined)
|
|
290
|
-
|
|
393
|
+
at(pointerTo(key), `adapter_config "${key}" ${problem}`);
|
|
291
394
|
}
|
|
292
|
-
const
|
|
293
|
-
if (
|
|
294
|
-
|
|
295
|
-
|
|
395
|
+
const reserved = reservedIn(body, apiOf(input));
|
|
396
|
+
if (reserved.length > 0) {
|
|
397
|
+
at(pointerTo(`${EXTRA_BODY_PREFIX}${reserved[0]}`), `adapter_config ${reservedProblem(reserved)}`);
|
|
398
|
+
}
|
|
399
|
+
return { ...(body !== undefined && { body }), problems };
|
|
400
|
+
}
|
|
401
|
+
/** The JSON pointer to a flat `adapter_config` key (its dots stay in one token). */
|
|
402
|
+
function pointerTo(key) {
|
|
403
|
+
return `/adapter_config/${key.replace(/~/g, '~0').replace(/\//g, '~1')}`;
|
|
296
404
|
}
|
|
297
405
|
/** Set `path` in `body` to `value`; what's wrong, if anything. */
|
|
298
406
|
function setField(body, path, value) {
|
|
@@ -320,16 +428,24 @@ function setField(body, path, value) {
|
|
|
320
428
|
node[leaf] = value;
|
|
321
429
|
return undefined;
|
|
322
430
|
}
|
|
323
|
-
function extraBodyProblem(value) {
|
|
431
|
+
function extraBodyProblem(value, api) {
|
|
324
432
|
if (value === undefined)
|
|
325
433
|
return undefined;
|
|
326
434
|
if (value === null || typeof value !== 'object' || Array.isArray(value)) {
|
|
327
435
|
return 'extraBody must be an object of request fields, e.g. { chat_template_kwargs: { enable_thinking: false } }';
|
|
328
436
|
}
|
|
329
|
-
const reserved =
|
|
330
|
-
return reserved.length > 0
|
|
331
|
-
|
|
332
|
-
|
|
437
|
+
const reserved = reservedIn(value, api);
|
|
438
|
+
return reserved.length > 0 ? reservedProblem(reserved) : undefined;
|
|
439
|
+
}
|
|
440
|
+
/** The fields of `body` the adapter sets itself on `api`. */
|
|
441
|
+
function reservedIn(body, api) {
|
|
442
|
+
if (body === undefined)
|
|
443
|
+
return [];
|
|
444
|
+
const fields = api === 'responses' ? EXTRA_BODY_RESERVED_RESPONSES : EXTRA_BODY_RESERVED;
|
|
445
|
+
return fields.filter((key) => Object.hasOwn(body, key));
|
|
446
|
+
}
|
|
447
|
+
function reservedProblem(reserved) {
|
|
448
|
+
return `extraBody can't set ${reserved.join(', ')}: the adapter sets ${reserved.length === 1 ? 'it' : 'them'}`;
|
|
333
449
|
}
|
|
334
450
|
/**
|
|
335
451
|
* The completion's usage as the framework's counters. The endpoint's
|