@morlay/dsh-llm-openai-compatible 0.0.8 → 0.0.9-alpha.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,14 +3,14 @@
3
3
  DeepSeek Harness 的 **OpenAI 兼容 LLM 适配器**插件。与内置 `llm-pi-ai` /
4
4
  `llm-deepseek` 不同,本插件的配置 schema 支持 **profile 级默认采样参数**
5
5
  (`temperature` / `topP` / `topK` / `presencePenalty` / `frequencyPenalty` /
6
- `seed`),请求级 `GenerateOptions.temperature` 优先于 profile 默认值。用
7
- `providers` dict 多路由结构(与 `llm-pi-ai` 一致),用户可将现有
8
- `llm-pi-ai` 配置近乎无缝迁移。
6
+ `seed`),请求级 `GenerateOptions.temperature` 优先于 profile 默认值;`providers`
7
+ 用 dict 多路由结构(与 `llm-pi-ai` 一致)。
9
8
 
10
9
  传输层复用 **[@ai-sdk/openai-compatible](https://www.npmjs.com/package/@ai-sdk/openai-compatible)**
11
- (`LanguageModelV4.doStream`:wire 序列化与 SSE 解析由 SDK 负责);本插件负责
12
- harness 消息 → AI SDK prompt 转换、采样默认合并、stream part → `StreamChunk`
13
- 翻译、错误归一化与凭据策略。
10
+ (wire 序列化与 SSE 解析由 SDK 负责);本插件负责 harness 消息 → AI SDK prompt
11
+ 转换、采样默认合并、stream part → `StreamChunk` 翻译、错误归一化与凭据策略。
12
+
13
+ 本文件只给配置面与用法;每条规则与取舍的 home 是各自 ADR(见下)。
14
14
 
15
15
  ## 配置
16
16
 
@@ -55,73 +55,31 @@ llm-openai-compatible:
55
55
  maxRetries: 5
56
56
  ```
57
57
 
58
- ### 采样默认值合并规则
59
-
60
- | wire 字段 | 取值 | 省略语义 |
61
- | ----------------------------------------------------------- | ------------------------------------------------------------------ | --------------------------- |
62
- | `temperature` | `options.temperature ?? profile.temperature` | 都不给 → 不发送,提供方默认 |
63
- | `max_tokens` | `options.maxTokens ?? model.maxTokens ?? profile.defaultMaxTokens` | 都不给 → 不发送 |
64
- | `top_p` / `presence_penalty` / `frequency_penalty` / `seed` | `profile` 值 | undefined → 不发送 |
65
- | `top_k` | `profile.topK`,经 `providerOptions` 透传进请求体 | undefined → 不发送 |
66
- | `reasoning_effort` | 见下 | 解析不出 → 不发送 |
67
-
68
- ### reasoning 映射(OpenAI 风格)
69
-
70
- - 模型声明 `reasoningEfforts`(对象)后,该模型的选器公开 `efforts`(按声明
71
- 顺序)+ `defaultEffort`(= `profile.reasoning`,须在模型能力内,否则视为无
72
- 默认值——**描述模型时绝不抛错**)。
73
- - wire:非 `off` 档位发送 `reasoning_effort: <声明值>`;`off` → 不发送。
74
- - 请求级 `options.reasoningEffort` 不在模型能力内 → 网络 I/O 前抛
75
- `LlmError('UNSUPPORTED_REASONING_EFFORT')`。`profile.reasoning` 配了模型不
76
- 支持的档位 → 请求执行处失败(同一错误码),配置页面仍可编辑。
77
- - 模型不声明 `reasoningEfforts`(或 `false`)→ 不公开 reasoning 能力。
78
-
79
- ### 模型目录
80
-
81
- - `models` 缺省 = 服务**空目录**:`listModels` 返回空,未列出的 id 原样透传
82
- (`resolveModel` 返回基础信息 + `defaultContextWindow` / `defaultMaxTokens`)。
83
- - 模型 `maxTokens` 配置后成为该模型的 per-request 默认输出上限。
84
- - `inputModalities` 缺省 `[text]`;声明含 `image` 的模型接受图片输入
85
- (attachments seam,base64 data-URL parts)。
86
-
87
- ### 凭据
88
-
89
- - profile 设置 `apiKeyEnv` 后:存在 `ctx.credentials` 服务时经它解析
90
- (credentials-local 覆盖进程环境 / `.env`);服务缺失时回退
91
- `launchEnvironmentOf(ctx)`。解析不到 → `LlmError('MISSING_CREDENTIAL')`。
92
- - profile 不设置 `apiKeyEnv` → 请求不带 `authorization` 头(无认证端点,
93
- 如本地 Ollama)。
94
-
95
- ## 传输
96
-
97
- - 端点 = `baseURL` + `/chat/completions`(streaming,`stream_options.include_usage`)。
98
- - 每个请求携带 `attributionHeaders()` + `x-…-harness-user-id`(+ session-id /
99
- compaction 标头),并带 SDK 的 `ai-sdk/openai-compatible` user-agent 后缀。
100
- - `streamIdleTimeoutMs` 控制流空闲超时(`TIMEOUT`);`timeoutMs` 控制整体请求
101
- 超时(缺省不设)。
102
- - 错误映射:401/403 → `AUTH`、429 → `RATE_LIMIT`、400+上下文 →
103
- `CONTEXT_WINDOW_EXCEEDED`、5xx → `SERVER`、配额 → `QUOTA_EXCEEDED`。
104
- - 用量:`prompt_tokens_details.cached_tokens` 与 DeepSeek 方言的
105
- `prompt_cache_hit_tokens` 都被拆出为 `cacheReadTokens`(disjoint 计数)。
58
+ ## 规则与取舍
59
+
60
+ 规则细节(合并表、wire 字段落点、端点与错误映射)在各自 ADR 里维护,这里只列结论:
61
+
62
+ - **采样默认值与省略语义**:请求级优先于 profile 默认,缺省一律**不发送**(= 提供方
63
+ 默认)——见 [ADR-采样默认值合并规则与省略语义](./.agents/adrs/20260917-采样默认值合并规则与省略语义.md)。
64
+ - **模型目录缺省为空、描述模型绝不抛错**:未列出的 id 原样透传,不支持的能力配置推迟到
65
+ 请求执行处失败——见 [ADR-模型目录缺省为空且描述模型绝不抛错](./.agents/adrs/20260917-模型目录缺省为空且描述模型绝不抛错.md)。
66
+ - **传输层复用 SDK**:非标准字段(`top_k`)与用量方言在本层显式处理——见
67
+ [ADR-传输层复用ai-sdk-openai-compatible而非自研wire序列化](./.agents/adrs/20260917-传输层复用ai-sdk-openai-compatible而非自研wire序列化.md)。
68
+ - **凭据经 `ctx.credentials` 解析**(服务缺失回退 launch environment):见
69
+ [ADR-凭据经credentials服务解析而非直接读环境变量](./.agents/adrs/20260917-凭据经credentials服务解析而非直接读环境变量.md)。
70
+ - **图片超预算报错交由 durable offload 重试**(本适配器不自行裁剪请求图片):见
71
+ [ADR-图片超预算报错交由durable-offload重试](./.agents/adrs/20260917-图片超预算报错交由durable-offload重试.md)。
72
+ - **多路由结构对齐 `llm-pi-ai`**(数组式 profiles 被拒绝):见
73
+ [ADR-providers采用dict多路由结构对齐llm-pi-ai](./.agents/adrs/20260917-providers采用dict多路由结构对齐llm-pi-ai.md)
74
+ 与 [ADR-起因llm-pi-ai的请求参数配置不完整](./.agents/adrs/20260917-起因llm-pi-ai的请求参数配置不完整.md)。
75
+ - **未做(YAGNI)**:`modelOverrides`、模型 discovery(`GET /models`)、OAuth / 非
76
+ bearer 认证——理由见模型目录 ADR。
106
77
 
107
78
  ## 从 `llm-pi-ai` 迁移
108
79
 
109
- 把 `llm-pi-ai.providers.<route>` 的 `baseURL`/`models`/采样字段平移到
80
+ 把 `llm-pi-ai.providers.<route>` 的 `baseURL` / `models` / 采样字段平移到
110
81
  `llm-openai-compatible.providers.<route>`(无 `api` 字段——协议固定
111
- chat-completions),`apiKeyEnv` 与 `retryPolicy` 原样保留;
112
- `reasoningEfforts` 的 `off` 空值语义一致。
113
-
114
- ## 暂缓能力(YAGNI)
82
+ chat-completions),`apiKeyEnv` 与 `retryPolicy` 原样保留;`reasoningEfforts`
83
+ 的 `off` 空值语义一致。
115
84
 
116
- - 不做 `modelOverrides`(providers 里每个路由自己写 `models` 即可)。
117
- - 不做模型 discovery(端点询问 `GET /models`);需要时手写 `models`。
118
- - 不做 OAuth / 非 bearer 认证。
119
-
120
- ## 构建与验证
121
-
122
- ```bash
123
- pnpm install
124
- pnpm --filter @morlay/dsh-llm-openai-compatible run build # → dist/*.mjs + *.d.mts
125
- pnpm exec tsc --noEmit
126
- pnpm exec vitest run
127
- ```
85
+ 构建、测试与类型检查的命令见根 `justfile` 与 `mise.toml`。
package/dist/index.mjs CHANGED
@@ -1,4 +1,4 @@
1
- import { a as serializeCallOptions, o as serializeCallOptionsWithImages, r as translate } from "./translate-BzOJ1xx-.mjs";
1
+ import { a as serializeCallOptions, o as serializeCallOptionsWithImages, r as translate } from "./translate-CwniWJOm.mjs";
2
2
  import z from "@deepseek-ai/schemastery";
3
3
  import { CONTEXT_WINDOW_EXCEEDED_CODE, LlmAdapter, LlmError, ProviderRequestId, QUOTA_EXCEEDED_CODE, ReasoningEffortId, RetryPolicySchema, assertUsableApiKey, attributionHeaders, contentHasImage, isContextWindowExceededError, isQuotaExceededError, resolveRetryPolicy } from "@deepseek-ai/dsh-llm";
4
4
  import { credentialRef } from "@deepseek-ai/dsh-credentials";
@@ -1,4 +1,4 @@
1
- import { EMPTY_RESPONSE_CODE, LlmError, ToolCallId, contentHasImage, offloadRequestImagesWithPolicy, textOnlyImageText } from "@deepseek-ai/dsh-llm";
1
+ import { EMPTY_RESPONSE_CODE, IMAGE_OFFLOAD_REQUIRED_CODE, LlmError, ToolCallId, contentHasImage, offloadedImageText, projectOffloadedImages, requiredImageOffload } from "@deepseek-ai/dsh-llm";
2
2
  import { AttachmentError } from "@deepseek-ai/dsh-attachment";
3
3
  import { Buffer } from "node:buffer";
4
4
  //#region src/serialize.ts
@@ -24,6 +24,14 @@ function assertTextOnly(blocks) {
24
24
  function assertSupportedImageRoles(messages) {
25
25
  for (const message of messages) if (message.role !== "user" && contentHasImage(message.content)) throw new LlmError(`The OpenAI-compatible chat-completions adapter cannot represent image content in a ${message.role} message.`, "UNSUPPORTED_CONTENT");
26
26
  }
27
+ function assertRetainedImagesFit(messages, maxRequestImageBytes) {
28
+ const offloadImages = requiredImageOffload(messages, {
29
+ representation: "base64",
30
+ maxBytes: maxRequestImageBytes
31
+ }, (block) => block.attachment.bytes);
32
+ if (offloadImages === 0) return;
33
+ throw new LlmError(`OpenAI-compatible request images exceed the provider budget; ${offloadImages} more oldest occurrence(s) must be offloaded.`, IMAGE_OFFLOAD_REQUIRED_CODE, { offloadImages });
34
+ }
27
35
  async function imagePart(block, attachments, signal) {
28
36
  try {
29
37
  const stored = await attachments.readImage(block.attachment, signal);
@@ -199,11 +207,8 @@ async function serializeCallOptions(options, profile, model) {
199
207
  return callOptionsWithPrompt(options, profile, model, [...system, ...prompt]);
200
208
  }
201
209
  async function serializeCallOptionsWithImages(options, profile, model, images) {
202
- const requestMessages = offloadRequestImagesWithPolicy(options.messages, {
203
- representation: "raw",
204
- maxBytes: images.maxRequestImageBytes,
205
- placeholder: (ref) => textOnlyImageText(ref)
206
- });
210
+ assertRetainedImagesFit(options.messages, images.maxRequestImageBytes);
211
+ const requestMessages = projectOffloadedImages(options.messages, (ref) => offloadedImageText(ref));
207
212
  const resolveImage = (block, signal) => imagePart(block, images.attachments, signal);
208
213
  const system = options.system === void 0 ? [] : [{
209
214
  role: "system",
package/dist/wire.mjs CHANGED
@@ -1,2 +1,2 @@
1
- import { a as serializeCallOptions, i as resolveReasoningWire, n as mapUsage, o as serializeCallOptionsWithImages, r as translate, t as mapFinishReason } from "./translate-BzOJ1xx-.mjs";
1
+ import { a as serializeCallOptions, i as resolveReasoningWire, n as mapUsage, o as serializeCallOptionsWithImages, r as translate, t as mapFinishReason } from "./translate-CwniWJOm.mjs";
2
2
  export { mapFinishReason, mapUsage, resolveReasoningWire, serializeCallOptions, serializeCallOptionsWithImages, translate };
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@morlay/dsh-llm-openai-compatible",
3
- "version": "0.0.8",
3
+ "version": "0.0.9-alpha.0",
4
4
  "description": "OpenAI-compatible LLM adapter plugin for DeepSeek Harness with configurable default sampling parameters (temperature / topP / topK / penalties / seed) over a providers dict.",
5
5
  "keywords": [
6
6
  "dsh",
@@ -13,7 +13,7 @@
13
13
  "license": "MIT",
14
14
  "repository": {
15
15
  "type": "git",
16
- "url": "https://github.com/morlay/better-session.git"
16
+ "url": "https://github.com/morlay/dsh-plugin.git"
17
17
  },
18
18
  "files": [
19
19
  "dist",
@@ -32,18 +32,18 @@
32
32
  "dependencies": {
33
33
  "@ai-sdk/openai-compatible": "^3.0.32",
34
34
  "@ai-sdk/provider": "^4.0.7",
35
- "@deepseek-ai/dsh-util-values": "^0.1.5-rc.2"
35
+ "@deepseek-ai/dsh-util-values": "^0.1.6-alpha.2"
36
36
  },
37
37
  "peerDependencies": {
38
38
  "@deepseek-ai/cordis": "^4.0.2",
39
- "@deepseek-ai/dsh-anonymous-user-id": "^0.1.5-rc.2",
40
- "@deepseek-ai/dsh-attachment": "^0.1.5-rc.2",
41
- "@deepseek-ai/dsh-credentials": "^0.1.5-rc.2",
42
- "@deepseek-ai/dsh-invariants": "^0.1.5-rc.2",
43
- "@deepseek-ai/dsh-launch-environment": "^0.1.5-rc.2",
44
- "@deepseek-ai/dsh-llm": "^0.1.5-rc.2",
45
- "@deepseek-ai/dsh-settings": "^0.1.5-rc.2",
46
- "@deepseek-ai/dsh-timeout": "^0.1.5-rc.2",
39
+ "@deepseek-ai/dsh-anonymous-user-id": "^0.1.6-alpha.2",
40
+ "@deepseek-ai/dsh-attachment": "^0.1.6-alpha.2",
41
+ "@deepseek-ai/dsh-credentials": "^0.1.6-alpha.2",
42
+ "@deepseek-ai/dsh-invariants": "^0.1.6-alpha.2",
43
+ "@deepseek-ai/dsh-launch-environment": "^0.1.6-alpha.2",
44
+ "@deepseek-ai/dsh-llm": "^0.1.6-alpha.2",
45
+ "@deepseek-ai/dsh-settings": "^0.1.6-alpha.2",
46
+ "@deepseek-ai/dsh-timeout": "^0.1.6-alpha.2",
47
47
  "@deepseek-ai/schemastery": "^3.18.2"
48
48
  },
49
49
  "dsh": {
package/src/adapter.ts CHANGED
@@ -63,7 +63,7 @@ export interface ResolvedProviderProfile {
63
63
 
64
64
  baseURL: string;
65
65
  headers?: Readonly<Record<string, string>>;
66
- // === sampling defaults (request-level values win) ===
66
+
67
67
  temperature?: number;
68
68
  topP?: number;
69
69
  topK?: number;
@@ -72,11 +72,11 @@ export interface ResolvedProviderProfile {
72
72
  seed?: number;
73
73
 
74
74
  reasoning?: ReasoningEffort;
75
- // === model catalog ===
75
+
76
76
  models: readonly ResolvedModelProfile[];
77
77
  defaultContextWindow: number;
78
78
  defaultMaxTokens: number;
79
- // === transport ===
79
+
80
80
  maxRequestImageBytes: number;
81
81
  streamIdleTimeoutMs: number;
82
82
 
@@ -157,9 +157,7 @@ function reasoningInfo(
157
157
  return {
158
158
  reasoning: {
159
159
  efforts,
160
- // A configured default the model does not declare is silently dropped
161
- // here (describing a model must never throw); the request path still
162
- // refuses it, which is where a bad deployment default belongs.
160
+
163
161
  ...(defaultEffort !== void 0 && declaration[defaultEffort] !== void 0
164
162
  ? { defaultEffort: ReasoningEffortId(defaultEffort) }
165
163
  : {}),
@@ -390,9 +388,7 @@ export class OpenAICompatibleAdapter extends LlmAdapter {
390
388
  if (!exhausted) {
391
389
  try {
392
390
  await iterator.return(void 0);
393
- } catch {
394
- // The transport already aborted; teardown is best-effort.
395
- }
391
+ } catch {}
396
392
  }
397
393
  }
398
394
  } finally {
package/src/index.ts CHANGED
@@ -61,7 +61,7 @@ export interface ProviderProfileSource {
61
61
  baseURL: string;
62
62
 
63
63
  headers?: Record<string, string>;
64
- // === sampling defaults (request-level values win) ===
64
+
65
65
  temperature?: number;
66
66
  topP?: number;
67
67
  topK?: number;
package/src/serialize.ts CHANGED
@@ -1,8 +1,10 @@
1
1
  import {
2
+ IMAGE_OFFLOAD_REQUIRED_CODE,
2
3
  LlmError,
3
4
  contentHasImage,
4
- offloadRequestImagesWithPolicy,
5
- textOnlyImageText,
5
+ offloadedImageText,
6
+ projectOffloadedImages,
7
+ requiredImageOffload,
6
8
  } from "@deepseek-ai/dsh-llm";
7
9
  import type { ContentBlock, GenerateOptions, Message } from "@deepseek-ai/dsh-llm";
8
10
  import { AttachmentError } from "@deepseek-ai/dsh-attachment";
@@ -92,6 +94,20 @@ function assertSupportedImageRoles(messages: readonly Message[]): void {
92
94
  }
93
95
  }
94
96
 
97
+ function assertRetainedImagesFit(messages: readonly Message[], maxRequestImageBytes: number): void {
98
+ const offloadImages = requiredImageOffload(
99
+ messages,
100
+ { representation: "base64", maxBytes: maxRequestImageBytes },
101
+ (block) => block.attachment.bytes,
102
+ );
103
+ if (offloadImages === 0) return;
104
+ throw new LlmError(
105
+ `OpenAI-compatible request images exceed the provider budget; ${offloadImages} more oldest occurrence(s) must be offloaded.`,
106
+ IMAGE_OFFLOAD_REQUIRED_CODE,
107
+ { offloadImages },
108
+ );
109
+ }
110
+
95
111
  async function imagePart(
96
112
  block: Extract<ContentBlock, { type: "image" }>,
97
113
  attachments: AttachmentStore,
@@ -322,11 +338,10 @@ export async function serializeCallOptionsWithImages(
322
338
  model: ResolvedModelProfile | undefined,
323
339
  images: { attachments: AttachmentStore; maxRequestImageBytes: number; signal?: AbortSignal },
324
340
  ): Promise<OpenAICompatibleCallOptions> {
325
- const requestMessages = offloadRequestImagesWithPolicy(options.messages, {
326
- representation: "raw",
327
- maxBytes: images.maxRequestImageBytes,
328
- placeholder: (ref) => textOnlyImageText(ref),
329
- });
341
+ assertRetainedImagesFit(options.messages, images.maxRequestImageBytes);
342
+ const requestMessages = projectOffloadedImages(options.messages, (ref) =>
343
+ offloadedImageText(ref),
344
+ );
330
345
  const resolveImage = (block: Extract<ContentBlock, { type: "image" }>, signal?: AbortSignal) =>
331
346
  imagePart(block, images.attachments, signal);
332
347
  const system =
package/src/translate.ts CHANGED
@@ -148,8 +148,6 @@ export async function* translate(
148
148
  case "file":
149
149
  case "reasoning-file":
150
150
  case "source":
151
- // Provider-executed tools and generated files are not part of this
152
- // adapter's client-executed tool loop; nothing to emit.
153
151
  break;
154
152
  case "finish":
155
153
  pendingUsage = mapUsage(part.usage);
@@ -189,7 +187,6 @@ function applyToolCall(part: LanguageModelV4ToolCall, toolQueue: OpenBlock[]): v
189
187
  if (block === void 0) return;
190
188
  block.callId = part.toolCallId;
191
189
  block.name = part.toolName;
192
- // The provider emits the complete arguments on this part; the buffered
193
- // deltas were a partial view.
190
+
194
191
  block.text = part.input;
195
192
  }
package/src/wire.ts CHANGED
@@ -1,4 +1,2 @@
1
- // wire 编解码层:请求序列化(serialize)与响应翻译(translate)的 OpenAI 兼容
2
- // wire 格式处理。文件物理分开,经本入口聚合导出。
3
1
  export * from "./serialize.ts";
4
2
  export * from "./translate.ts";