@morlay/dsh-llm-openai-compatible 0.0.8 → 0.0.9-alpha.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -71
- package/dist/index.mjs +1 -1
- package/dist/{translate-BzOJ1xx-.mjs → translate-CwniWJOm.mjs} +11 -6
- package/dist/wire.mjs +1 -1
- package/package.json +11 -11
- package/src/adapter.ts +5 -9
- package/src/index.ts +1 -1
- package/src/serialize.ts +22 -7
- package/src/translate.ts +1 -4
- package/src/wire.ts +0 -2
package/README.md
CHANGED
|
@@ -3,14 +3,14 @@
|
|
|
3
3
|
DeepSeek Harness 的 **OpenAI 兼容 LLM 适配器**插件。与内置 `llm-pi-ai` /
|
|
4
4
|
`llm-deepseek` 不同,本插件的配置 schema 支持 **profile 级默认采样参数**
|
|
5
5
|
(`temperature` / `topP` / `topK` / `presencePenalty` / `frequencyPenalty` /
|
|
6
|
-
`seed`),请求级 `GenerateOptions.temperature` 优先于 profile
|
|
7
|
-
|
|
8
|
-
`llm-pi-ai` 配置近乎无缝迁移。
|
|
6
|
+
`seed`),请求级 `GenerateOptions.temperature` 优先于 profile 默认值;`providers`
|
|
7
|
+
用 dict 多路由结构(与 `llm-pi-ai` 一致)。
|
|
9
8
|
|
|
10
9
|
传输层复用 **[@ai-sdk/openai-compatible](https://www.npmjs.com/package/@ai-sdk/openai-compatible)**
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
10
|
+
(wire 序列化与 SSE 解析由 SDK 负责);本插件负责 harness 消息 → AI SDK prompt
|
|
11
|
+
转换、采样默认合并、stream part → `StreamChunk` 翻译、错误归一化与凭据策略。
|
|
12
|
+
|
|
13
|
+
本文件只给配置面与用法;每条规则与取舍的 home 是各自 ADR(见下)。
|
|
14
14
|
|
|
15
15
|
## 配置
|
|
16
16
|
|
|
@@ -55,73 +55,31 @@ llm-openai-compatible:
|
|
|
55
55
|
maxRetries: 5
|
|
56
56
|
```
|
|
57
57
|
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
-
|
|
74
|
-
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
- 模型不声明 `reasoningEfforts`(或 `false`)→ 不公开 reasoning 能力。
|
|
78
|
-
|
|
79
|
-
### 模型目录
|
|
80
|
-
|
|
81
|
-
- `models` 缺省 = 服务**空目录**:`listModels` 返回空,未列出的 id 原样透传
|
|
82
|
-
(`resolveModel` 返回基础信息 + `defaultContextWindow` / `defaultMaxTokens`)。
|
|
83
|
-
- 模型 `maxTokens` 配置后成为该模型的 per-request 默认输出上限。
|
|
84
|
-
- `inputModalities` 缺省 `[text]`;声明含 `image` 的模型接受图片输入
|
|
85
|
-
(attachments seam,base64 data-URL parts)。
|
|
86
|
-
|
|
87
|
-
### 凭据
|
|
88
|
-
|
|
89
|
-
- profile 设置 `apiKeyEnv` 后:存在 `ctx.credentials` 服务时经它解析
|
|
90
|
-
(credentials-local 覆盖进程环境 / `.env`);服务缺失时回退
|
|
91
|
-
`launchEnvironmentOf(ctx)`。解析不到 → `LlmError('MISSING_CREDENTIAL')`。
|
|
92
|
-
- profile 不设置 `apiKeyEnv` → 请求不带 `authorization` 头(无认证端点,
|
|
93
|
-
如本地 Ollama)。
|
|
94
|
-
|
|
95
|
-
## 传输
|
|
96
|
-
|
|
97
|
-
- 端点 = `baseURL` + `/chat/completions`(streaming,`stream_options.include_usage`)。
|
|
98
|
-
- 每个请求携带 `attributionHeaders()` + `x-…-harness-user-id`(+ session-id /
|
|
99
|
-
compaction 标头),并带 SDK 的 `ai-sdk/openai-compatible` user-agent 后缀。
|
|
100
|
-
- `streamIdleTimeoutMs` 控制流空闲超时(`TIMEOUT`);`timeoutMs` 控制整体请求
|
|
101
|
-
超时(缺省不设)。
|
|
102
|
-
- 错误映射:401/403 → `AUTH`、429 → `RATE_LIMIT`、400+上下文 →
|
|
103
|
-
`CONTEXT_WINDOW_EXCEEDED`、5xx → `SERVER`、配额 → `QUOTA_EXCEEDED`。
|
|
104
|
-
- 用量:`prompt_tokens_details.cached_tokens` 与 DeepSeek 方言的
|
|
105
|
-
`prompt_cache_hit_tokens` 都被拆出为 `cacheReadTokens`(disjoint 计数)。
|
|
58
|
+
## 规则与取舍
|
|
59
|
+
|
|
60
|
+
规则细节(合并表、wire 字段落点、端点与错误映射)在各自 ADR 里维护,这里只列结论:
|
|
61
|
+
|
|
62
|
+
- **采样默认值与省略语义**:请求级优先于 profile 默认,缺省一律**不发送**(= 提供方
|
|
63
|
+
默认)——见 [ADR-采样默认值合并规则与省略语义](./.agents/adrs/20260917-采样默认值合并规则与省略语义.md)。
|
|
64
|
+
- **模型目录缺省为空、描述模型绝不抛错**:未列出的 id 原样透传,不支持的能力配置推迟到
|
|
65
|
+
请求执行处失败——见 [ADR-模型目录缺省为空且描述模型绝不抛错](./.agents/adrs/20260917-模型目录缺省为空且描述模型绝不抛错.md)。
|
|
66
|
+
- **传输层复用 SDK**:非标准字段(`top_k`)与用量方言在本层显式处理——见
|
|
67
|
+
[ADR-传输层复用ai-sdk-openai-compatible而非自研wire序列化](./.agents/adrs/20260917-传输层复用ai-sdk-openai-compatible而非自研wire序列化.md)。
|
|
68
|
+
- **凭据经 `ctx.credentials` 解析**(服务缺失回退 launch environment):见
|
|
69
|
+
[ADR-凭据经credentials服务解析而非直接读环境变量](./.agents/adrs/20260917-凭据经credentials服务解析而非直接读环境变量.md)。
|
|
70
|
+
- **图片超预算报错交由 durable offload 重试**(本适配器不自行裁剪请求图片):见
|
|
71
|
+
[ADR-图片超预算报错交由durable-offload重试](./.agents/adrs/20260917-图片超预算报错交由durable-offload重试.md)。
|
|
72
|
+
- **多路由结构对齐 `llm-pi-ai`**(数组式 profiles 被拒绝):见
|
|
73
|
+
[ADR-providers采用dict多路由结构对齐llm-pi-ai](./.agents/adrs/20260917-providers采用dict多路由结构对齐llm-pi-ai.md)
|
|
74
|
+
与 [ADR-起因llm-pi-ai的请求参数配置不完整](./.agents/adrs/20260917-起因llm-pi-ai的请求参数配置不完整.md)。
|
|
75
|
+
- **未做(YAGNI)**:`modelOverrides`、模型 discovery(`GET /models`)、OAuth / 非
|
|
76
|
+
bearer 认证——理由见模型目录 ADR。
|
|
106
77
|
|
|
107
78
|
## 从 `llm-pi-ai` 迁移
|
|
108
79
|
|
|
109
|
-
把 `llm-pi-ai.providers.<route>` 的 `baseURL
|
|
80
|
+
把 `llm-pi-ai.providers.<route>` 的 `baseURL` / `models` / 采样字段平移到
|
|
110
81
|
`llm-openai-compatible.providers.<route>`(无 `api` 字段——协议固定
|
|
111
|
-
chat-completions),`apiKeyEnv` 与 `retryPolicy`
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
## 暂缓能力(YAGNI)
|
|
82
|
+
chat-completions),`apiKeyEnv` 与 `retryPolicy` 原样保留;`reasoningEfforts`
|
|
83
|
+
的 `off` 空值语义一致。
|
|
115
84
|
|
|
116
|
-
|
|
117
|
-
- 不做模型 discovery(端点询问 `GET /models`);需要时手写 `models`。
|
|
118
|
-
- 不做 OAuth / 非 bearer 认证。
|
|
119
|
-
|
|
120
|
-
## 构建与验证
|
|
121
|
-
|
|
122
|
-
```bash
|
|
123
|
-
pnpm install
|
|
124
|
-
pnpm --filter @morlay/dsh-llm-openai-compatible run build # → dist/*.mjs + *.d.mts
|
|
125
|
-
pnpm exec tsc --noEmit
|
|
126
|
-
pnpm exec vitest run
|
|
127
|
-
```
|
|
85
|
+
构建、测试与类型检查的命令见根 `justfile` 与 `mise.toml`。
|
package/dist/index.mjs
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { a as serializeCallOptions, o as serializeCallOptionsWithImages, r as translate } from "./translate-
|
|
1
|
+
import { a as serializeCallOptions, o as serializeCallOptionsWithImages, r as translate } from "./translate-CwniWJOm.mjs";
|
|
2
2
|
import z from "@deepseek-ai/schemastery";
|
|
3
3
|
import { CONTEXT_WINDOW_EXCEEDED_CODE, LlmAdapter, LlmError, ProviderRequestId, QUOTA_EXCEEDED_CODE, ReasoningEffortId, RetryPolicySchema, assertUsableApiKey, attributionHeaders, contentHasImage, isContextWindowExceededError, isQuotaExceededError, resolveRetryPolicy } from "@deepseek-ai/dsh-llm";
|
|
4
4
|
import { credentialRef } from "@deepseek-ai/dsh-credentials";
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { EMPTY_RESPONSE_CODE, LlmError, ToolCallId, contentHasImage,
|
|
1
|
+
import { EMPTY_RESPONSE_CODE, IMAGE_OFFLOAD_REQUIRED_CODE, LlmError, ToolCallId, contentHasImage, offloadedImageText, projectOffloadedImages, requiredImageOffload } from "@deepseek-ai/dsh-llm";
|
|
2
2
|
import { AttachmentError } from "@deepseek-ai/dsh-attachment";
|
|
3
3
|
import { Buffer } from "node:buffer";
|
|
4
4
|
//#region src/serialize.ts
|
|
@@ -24,6 +24,14 @@ function assertTextOnly(blocks) {
|
|
|
24
24
|
function assertSupportedImageRoles(messages) {
|
|
25
25
|
for (const message of messages) if (message.role !== "user" && contentHasImage(message.content)) throw new LlmError(`The OpenAI-compatible chat-completions adapter cannot represent image content in a ${message.role} message.`, "UNSUPPORTED_CONTENT");
|
|
26
26
|
}
|
|
27
|
+
function assertRetainedImagesFit(messages, maxRequestImageBytes) {
|
|
28
|
+
const offloadImages = requiredImageOffload(messages, {
|
|
29
|
+
representation: "base64",
|
|
30
|
+
maxBytes: maxRequestImageBytes
|
|
31
|
+
}, (block) => block.attachment.bytes);
|
|
32
|
+
if (offloadImages === 0) return;
|
|
33
|
+
throw new LlmError(`OpenAI-compatible request images exceed the provider budget; ${offloadImages} more oldest occurrence(s) must be offloaded.`, IMAGE_OFFLOAD_REQUIRED_CODE, { offloadImages });
|
|
34
|
+
}
|
|
27
35
|
async function imagePart(block, attachments, signal) {
|
|
28
36
|
try {
|
|
29
37
|
const stored = await attachments.readImage(block.attachment, signal);
|
|
@@ -199,11 +207,8 @@ async function serializeCallOptions(options, profile, model) {
|
|
|
199
207
|
return callOptionsWithPrompt(options, profile, model, [...system, ...prompt]);
|
|
200
208
|
}
|
|
201
209
|
async function serializeCallOptionsWithImages(options, profile, model, images) {
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
maxBytes: images.maxRequestImageBytes,
|
|
205
|
-
placeholder: (ref) => textOnlyImageText(ref)
|
|
206
|
-
});
|
|
210
|
+
assertRetainedImagesFit(options.messages, images.maxRequestImageBytes);
|
|
211
|
+
const requestMessages = projectOffloadedImages(options.messages, (ref) => offloadedImageText(ref));
|
|
207
212
|
const resolveImage = (block, signal) => imagePart(block, images.attachments, signal);
|
|
208
213
|
const system = options.system === void 0 ? [] : [{
|
|
209
214
|
role: "system",
|
package/dist/wire.mjs
CHANGED
|
@@ -1,2 +1,2 @@
|
|
|
1
|
-
import { a as serializeCallOptions, i as resolveReasoningWire, n as mapUsage, o as serializeCallOptionsWithImages, r as translate, t as mapFinishReason } from "./translate-
|
|
1
|
+
import { a as serializeCallOptions, i as resolveReasoningWire, n as mapUsage, o as serializeCallOptionsWithImages, r as translate, t as mapFinishReason } from "./translate-CwniWJOm.mjs";
|
|
2
2
|
export { mapFinishReason, mapUsage, resolveReasoningWire, serializeCallOptions, serializeCallOptionsWithImages, translate };
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@morlay/dsh-llm-openai-compatible",
|
|
3
|
-
"version": "0.0.
|
|
3
|
+
"version": "0.0.9-alpha.1",
|
|
4
4
|
"description": "OpenAI-compatible LLM adapter plugin for DeepSeek Harness with configurable default sampling parameters (temperature / topP / topK / penalties / seed) over a providers dict.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"dsh",
|
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
"license": "MIT",
|
|
14
14
|
"repository": {
|
|
15
15
|
"type": "git",
|
|
16
|
-
"url": "https://github.com/morlay/
|
|
16
|
+
"url": "https://github.com/morlay/dsh-plugin.git"
|
|
17
17
|
},
|
|
18
18
|
"files": [
|
|
19
19
|
"dist",
|
|
@@ -32,18 +32,18 @@
|
|
|
32
32
|
"dependencies": {
|
|
33
33
|
"@ai-sdk/openai-compatible": "^3.0.32",
|
|
34
34
|
"@ai-sdk/provider": "^4.0.7",
|
|
35
|
-
"@deepseek-ai/dsh-util-values": "^0.1.
|
|
35
|
+
"@deepseek-ai/dsh-util-values": "^0.1.6-alpha.2"
|
|
36
36
|
},
|
|
37
37
|
"peerDependencies": {
|
|
38
38
|
"@deepseek-ai/cordis": "^4.0.2",
|
|
39
|
-
"@deepseek-ai/dsh-anonymous-user-id": "^0.1.
|
|
40
|
-
"@deepseek-ai/dsh-attachment": "^0.1.
|
|
41
|
-
"@deepseek-ai/dsh-credentials": "^0.1.
|
|
42
|
-
"@deepseek-ai/dsh-invariants": "^0.1.
|
|
43
|
-
"@deepseek-ai/dsh-launch-environment": "^0.1.
|
|
44
|
-
"@deepseek-ai/dsh-llm": "^0.1.
|
|
45
|
-
"@deepseek-ai/dsh-settings": "^0.1.
|
|
46
|
-
"@deepseek-ai/dsh-timeout": "^0.1.
|
|
39
|
+
"@deepseek-ai/dsh-anonymous-user-id": "^0.1.6-alpha.2",
|
|
40
|
+
"@deepseek-ai/dsh-attachment": "^0.1.6-alpha.2",
|
|
41
|
+
"@deepseek-ai/dsh-credentials": "^0.1.6-alpha.2",
|
|
42
|
+
"@deepseek-ai/dsh-invariants": "^0.1.6-alpha.2",
|
|
43
|
+
"@deepseek-ai/dsh-launch-environment": "^0.1.6-alpha.2",
|
|
44
|
+
"@deepseek-ai/dsh-llm": "^0.1.6-alpha.2",
|
|
45
|
+
"@deepseek-ai/dsh-settings": "^0.1.6-alpha.2",
|
|
46
|
+
"@deepseek-ai/dsh-timeout": "^0.1.6-alpha.2",
|
|
47
47
|
"@deepseek-ai/schemastery": "^3.18.2"
|
|
48
48
|
},
|
|
49
49
|
"dsh": {
|
package/src/adapter.ts
CHANGED
|
@@ -63,7 +63,7 @@ export interface ResolvedProviderProfile {
|
|
|
63
63
|
|
|
64
64
|
baseURL: string;
|
|
65
65
|
headers?: Readonly<Record<string, string>>;
|
|
66
|
-
|
|
66
|
+
|
|
67
67
|
temperature?: number;
|
|
68
68
|
topP?: number;
|
|
69
69
|
topK?: number;
|
|
@@ -72,11 +72,11 @@ export interface ResolvedProviderProfile {
|
|
|
72
72
|
seed?: number;
|
|
73
73
|
|
|
74
74
|
reasoning?: ReasoningEffort;
|
|
75
|
-
|
|
75
|
+
|
|
76
76
|
models: readonly ResolvedModelProfile[];
|
|
77
77
|
defaultContextWindow: number;
|
|
78
78
|
defaultMaxTokens: number;
|
|
79
|
-
|
|
79
|
+
|
|
80
80
|
maxRequestImageBytes: number;
|
|
81
81
|
streamIdleTimeoutMs: number;
|
|
82
82
|
|
|
@@ -157,9 +157,7 @@ function reasoningInfo(
|
|
|
157
157
|
return {
|
|
158
158
|
reasoning: {
|
|
159
159
|
efforts,
|
|
160
|
-
|
|
161
|
-
// here (describing a model must never throw); the request path still
|
|
162
|
-
// refuses it, which is where a bad deployment default belongs.
|
|
160
|
+
|
|
163
161
|
...(defaultEffort !== void 0 && declaration[defaultEffort] !== void 0
|
|
164
162
|
? { defaultEffort: ReasoningEffortId(defaultEffort) }
|
|
165
163
|
: {}),
|
|
@@ -390,9 +388,7 @@ export class OpenAICompatibleAdapter extends LlmAdapter {
|
|
|
390
388
|
if (!exhausted) {
|
|
391
389
|
try {
|
|
392
390
|
await iterator.return(void 0);
|
|
393
|
-
} catch {
|
|
394
|
-
// The transport already aborted; teardown is best-effort.
|
|
395
|
-
}
|
|
391
|
+
} catch {}
|
|
396
392
|
}
|
|
397
393
|
}
|
|
398
394
|
} finally {
|
package/src/index.ts
CHANGED
package/src/serialize.ts
CHANGED
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
import {
|
|
2
|
+
IMAGE_OFFLOAD_REQUIRED_CODE,
|
|
2
3
|
LlmError,
|
|
3
4
|
contentHasImage,
|
|
4
|
-
|
|
5
|
-
|
|
5
|
+
offloadedImageText,
|
|
6
|
+
projectOffloadedImages,
|
|
7
|
+
requiredImageOffload,
|
|
6
8
|
} from "@deepseek-ai/dsh-llm";
|
|
7
9
|
import type { ContentBlock, GenerateOptions, Message } from "@deepseek-ai/dsh-llm";
|
|
8
10
|
import { AttachmentError } from "@deepseek-ai/dsh-attachment";
|
|
@@ -92,6 +94,20 @@ function assertSupportedImageRoles(messages: readonly Message[]): void {
|
|
|
92
94
|
}
|
|
93
95
|
}
|
|
94
96
|
|
|
97
|
+
function assertRetainedImagesFit(messages: readonly Message[], maxRequestImageBytes: number): void {
|
|
98
|
+
const offloadImages = requiredImageOffload(
|
|
99
|
+
messages,
|
|
100
|
+
{ representation: "base64", maxBytes: maxRequestImageBytes },
|
|
101
|
+
(block) => block.attachment.bytes,
|
|
102
|
+
);
|
|
103
|
+
if (offloadImages === 0) return;
|
|
104
|
+
throw new LlmError(
|
|
105
|
+
`OpenAI-compatible request images exceed the provider budget; ${offloadImages} more oldest occurrence(s) must be offloaded.`,
|
|
106
|
+
IMAGE_OFFLOAD_REQUIRED_CODE,
|
|
107
|
+
{ offloadImages },
|
|
108
|
+
);
|
|
109
|
+
}
|
|
110
|
+
|
|
95
111
|
async function imagePart(
|
|
96
112
|
block: Extract<ContentBlock, { type: "image" }>,
|
|
97
113
|
attachments: AttachmentStore,
|
|
@@ -322,11 +338,10 @@ export async function serializeCallOptionsWithImages(
|
|
|
322
338
|
model: ResolvedModelProfile | undefined,
|
|
323
339
|
images: { attachments: AttachmentStore; maxRequestImageBytes: number; signal?: AbortSignal },
|
|
324
340
|
): Promise<OpenAICompatibleCallOptions> {
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
});
|
|
341
|
+
assertRetainedImagesFit(options.messages, images.maxRequestImageBytes);
|
|
342
|
+
const requestMessages = projectOffloadedImages(options.messages, (ref) =>
|
|
343
|
+
offloadedImageText(ref),
|
|
344
|
+
);
|
|
330
345
|
const resolveImage = (block: Extract<ContentBlock, { type: "image" }>, signal?: AbortSignal) =>
|
|
331
346
|
imagePart(block, images.attachments, signal);
|
|
332
347
|
const system =
|
package/src/translate.ts
CHANGED
|
@@ -148,8 +148,6 @@ export async function* translate(
|
|
|
148
148
|
case "file":
|
|
149
149
|
case "reasoning-file":
|
|
150
150
|
case "source":
|
|
151
|
-
// Provider-executed tools and generated files are not part of this
|
|
152
|
-
// adapter's client-executed tool loop; nothing to emit.
|
|
153
151
|
break;
|
|
154
152
|
case "finish":
|
|
155
153
|
pendingUsage = mapUsage(part.usage);
|
|
@@ -189,7 +187,6 @@ function applyToolCall(part: LanguageModelV4ToolCall, toolQueue: OpenBlock[]): v
|
|
|
189
187
|
if (block === void 0) return;
|
|
190
188
|
block.callId = part.toolCallId;
|
|
191
189
|
block.name = part.toolName;
|
|
192
|
-
|
|
193
|
-
// deltas were a partial view.
|
|
190
|
+
|
|
194
191
|
block.text = part.input;
|
|
195
192
|
}
|
package/src/wire.ts
CHANGED