@arcaneorion/dsh-model-channel-manager 0.1.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +55 -5
- package/package.json +11 -4
- package/src/client.js +159 -22
- package/src/index.js +243 -78
package/README.md
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
DSH 原生模型渠道管理。两半结构:
|
|
4
4
|
|
|
5
5
|
- **host 半** `src/index.js`:轮询故障转移引擎(`llm.registerAdapter` 虚拟路由 `roundrobin/<组id>`)+ 7 天健康流水 + 测速排序 + 单模型真实请求测试通道。
|
|
6
|
-
- **client 半** `src/client.js`:`conversation.view` 顶级页签「模型配置」,内含三个子页:**模型配置**(llm-pi-ai providers 全字段编辑、拉取上游、单模型 ⚡ 测试、供应商搜索过滤)、**轮询渠道**(groups 编辑 + ⚡测速 +
|
|
6
|
+
- **client 半** `src/client.js`:`conversation.view` 顶级页签「模型配置」,内含三个子页:**模型配置**(llm-pi-ai providers 全字段编辑、拉取上游、单模型 ⚡ 测试、供应商搜索过滤)、**轮询渠道**(groups 编辑 + ⚡测速 + 输入模态编辑;**命名单一身份:显示名 = 组唯一 ID**,保存时归一化 `virtualModel.name = id`,改 ID 即改名,永不漂移)、**健康统计**(7 天聚合)。
|
|
7
7
|
- **会话模型选择器已拆出**为独立 cordis client 插件 [`@arcaneorion/dsh-model-selector-search`](../model-selector-search/)(一个占座者一个插件单元,可独立启停/替换;座位遮蔽 + 搜索 + 近 7 天置顶 + 菜单向上展开都在该仓)。
|
|
8
8
|
|
|
9
9
|
语义参考 pi 的 `pi-provider-manager`,但完全走 DSH 原生 seam(无独立 HTTP 服务/端口/token):
|
|
@@ -19,7 +19,14 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
19
19
|
|
|
20
20
|
## 挂载
|
|
21
21
|
|
|
22
|
-
|
|
22
|
+
从 npm 装(发布版):
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
dsh plugin --profile web add @arcaneorion/dsh-model-channel-manager
|
|
26
|
+
# 然后重启 dsh --profile web 并刷新页面
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
本地开发用 `link:`(改源码即时生效)——`profiles/web/package.json`:
|
|
23
30
|
- `dependencies` 加 `"@arcaneorion/dsh-model-channel-manager": "link:/home/arcaneorion/AI/AI-DSH/plugin/model-channel-manager"`
|
|
24
31
|
- `dsh.profile.bundles` 加 `"@arcaneorion/dsh-model-channel-manager"`
|
|
25
32
|
- `pnpm install` 后重启 `dsh --profile web`
|
|
@@ -29,6 +36,26 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
29
36
|
- `llm.providers` 出现 `roundrobin/<组id>`
|
|
30
37
|
- settings describe 含 `model-channels` / `model-channel-health` 命名空间
|
|
31
38
|
|
|
39
|
+
## 兼容性(DSH 版本)
|
|
40
|
+
|
|
41
|
+
> **当前工作树已移植到 DSH `0.2.0-rc.1`**(peer 按 `0.2.0-rc.1` 声明;settings 寻址从「自建命名空间」改为「本行实例配置」——
|
|
42
|
+
> 原先的 `model-channels` + `model-channel-health` 两个命名空间合并为本行 Config 的
|
|
43
|
+
> `groups` / `providerOrder` / `effortMemory` / `health`)。下面这段 `0.1.1-rc.2` 的记录仅作历史基线参考。
|
|
44
|
+
|
|
45
|
+
本包原在 **DSH `0.1.1-rc.2`**(`dsh --version`)上开发与实测,宿主侧依赖按该版本**精确钉住**:
|
|
46
|
+
|
|
47
|
+
| 宿主包 | 声明 | 用途 |
|
|
48
|
+
|---|---|---|
|
|
49
|
+
| `@deepseek-ai/dsh-llm` | `0.1.1-rc.2` | `llm.registerAdapter` / `llm.stream`(轮询引擎与健康采集) |
|
|
50
|
+
| `@deepseek-ai/dsh-settings` | `0.1.1-rc.2` | `model-channels` / `model-channel-health` 命名空间读写 |
|
|
51
|
+
| `@deepseek-ai/dsh-client-connection` | `0.1.1-rc.2` | client 半的 `connection.api` 调用 |
|
|
52
|
+
| `@deepseek-ai/dsh-client-ui-conversation` | `0.1.1-rc.2` | `conversation.view` 页签座位 |
|
|
53
|
+
| `@deepseek-ai/cordis` | `^4.0.2` | 插件生命周期 |
|
|
54
|
+
| `@deepseek-ai/schemastery` | `>=3.18.2` | 配置 schema |
|
|
55
|
+
|
|
56
|
+
**换 DSH 版本(例如 `0.1.2-rc.1`)必须先重新验证、再放宽 peer**:宿主服务与座位契约跨版本会变,
|
|
57
|
+
精确钉住的 peer 会在安装时报冲突——这正是它存在的意义,好过装上去静默失效。
|
|
58
|
+
|
|
32
59
|
## 数据通道(全走公共 seam,无私有 RPC)
|
|
33
60
|
|
|
34
61
|
- 读配置/健康/运行态 = `api.settings.describe()` 过滤命名空间
|
|
@@ -36,13 +63,17 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
36
63
|
- 保存轮询组 = `api.settings.update({ns:'model-channels', patch:{groups}})`
|
|
37
64
|
- ⚡测速 = `api.settings.update({ns:'model-channel-health', patch:{speedRequest:{group,nonce}}})`(host watcher 消费)
|
|
38
65
|
- 单模型测试 = `settings.update({ns:'model-channel-health', patch:{testRequest:{nonce,provider,model,prompt,maxTokens}}})`;host 执行真实 `llm.stream` 后把结果写回 `testResults[nonce]`;client 轮询 describe 直到 ok/error
|
|
39
|
-
- `apiRef`
|
|
66
|
+
- `apiRef` 获取:**0.2 为 `ctx.remote`**(插件级 `inject` 声明 `remote` / `remote.settings` / `remote.credentials` / `remote.llm`,apply 时经 `ctx.inject(['remote'])` 捕获);0.1 的 `ctx.get('connection').api` 已不存在。
|
|
67
|
+
- 上述调用形状仍保留 0.1 的样子:client 半内建门面 `makeLegacyApi` 把 0.2 的**位置参数 + RemoteResult** 适配回旧的**对象入参 + `{result:{ok,value}}`**,并把 `model-channels` / `model-channel-health` 合成回旧命名空间视图(真实承载是本插件行 id `model-channel-manager` 的实例配置)。
|
|
40
68
|
|
|
41
|
-
>
|
|
69
|
+
> **宿主边界(0.1 历史,0.2 已不适用)**:0.1 的 settings RPC 走 apiproxy 暴露白名单(`exposedNamespaces()` = LLM provider ns + `WEB_/PRODUCT_SETTINGS_NAMESPACES`),当时含该边界的宿主必须放行 `model-channels` / `model-channel-health`(本仓曾在 harness `dsh-host-apiproxy` 打 `PLUGIN_SETTINGS_NAMESPACES` 补丁)。
|
|
70
|
+
> **0.2 的 settings 命名空间就是 profile 行 id**,由 `@deepseek-ai/dsh-api-settings-controller` 的 `describe` 直接投影本行实例配置,没有该白名单环节;对应地,本插件的数据落在 `~/.dsh/profiles/web/cordis.patch.yml` 的 `model-channel-manager` 行 `config` 下。
|
|
42
71
|
|
|
43
72
|
## 响应信封(重要)
|
|
44
73
|
|
|
45
|
-
|
|
74
|
+
0.1 的 `connection.api.*` 返回 `{result: {ok, value}}` 包裹(`dsh-client-connection` 的 `callUnary` + zod 校验)。
|
|
75
|
+
**0.2 的原生远程调用直接返回 `RemoteResult`(`{ok, value} | {ok:false, error}`)**,不再有 `result` 外层;
|
|
76
|
+
本插件 client 的门面把它重新包回旧形状,因此下面这层解包逻辑在 0.2 上依然成立:
|
|
46
77
|
- 成功:`resp.result.value.{...}`
|
|
47
78
|
- 失败:`resp.result.ok === false`,错误在 `resp.result.error.message`
|
|
48
79
|
- `settings.describe` 的 value = `{writable, hasDocument, namespaces:[{ns, value, base, user, revision, ...}]}`
|
|
@@ -50,6 +81,14 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
50
81
|
|
|
51
82
|
**不要把 `result.value` 当 `result` 读**——曾因少解一层导致整个面板静默空数据(describe 返回 namespaces 但全面板 0 provider,无任何错误提示)。
|
|
52
83
|
|
|
84
|
+
## Token 用量统计(健康面板)
|
|
85
|
+
|
|
86
|
+
- 健康记录条目在既有字段(ts/provider/model/ok/ttftMs/latencyMs/code)上**增量附带**上游真实 token 用量:`inputTokens` / `outputTokens` / `cacheReadTokens` / `cacheWriteTokens` / `reasoningTokens`——来自适配器在 `finish` 前发出的 `usage` StreamChunk(rc.2 运行时 `StreamChunk` 契约,pi-ai `done`/`error` 事件都带);无 usage 则这些字段不写。
|
|
87
|
+
- **计费口径与 DSH `tokenMeter` 一致**:input + cacheRead + cacheWrite + output(互斥计数,`inputTokens` 不含缓存命中)。面板「总 Token 用量」卡按统计窗口求和,副行显示 输入/输出/缓存 拆分;每模型卡底部显示该模型窗口 Tokens。
|
|
88
|
+
- 采集点与健康记录**同址**(保证 token 与请求数同记录同窗口):全局 `llm/stream` 拦截器 + 轮询引擎 `streamAttempt`(成功/失败/超时路径都尽量携带;pi-ai 的 error 事件同样上报部分 usage)。
|
|
89
|
+
- 测速(⚡测速)/ 单模型测试(⚡测试)消耗的 token **不计入**——与「测速结果不入健康流水」的既有口径一致。
|
|
90
|
+
- 记录 schema 无需改动(`records` 为 `z.array(z.any())`),无新 RPC / settings 字段;client 5s 轮询自动刷新。
|
|
91
|
+
|
|
53
92
|
## 测试通道(模型可用性)
|
|
54
93
|
|
|
55
94
|
- 模型行「⚡测试」→ 弹窗输入自定义问题 + maxTokens → 发送
|
|
@@ -79,6 +118,12 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
79
118
|
- 历史健康流水保留在原 ID 名下(历史存档不受影响)
|
|
80
119
|
- 改名后仍需点击右上「保存全部变更」落盘
|
|
81
120
|
|
|
121
|
+
## 密钥写入(凭据引用虚拟化)
|
|
122
|
+
|
|
123
|
+
- 面板主视图只出现「API Key」输入框:**粘贴或输入后失焦即自动写入** DSH 凭据存储(`~/.dsh/.credentials.yaml`,0600,write-only 读不回),无手动按钮;清空输入框不会删除已存 key。上游 llm-pi-ai 的供应商 profile 只有 `apiKeyEnv` 一个密钥字段(凭据引用名),不存在内联 key 的选项——secrets 不进 settings.yaml、不随 `settings.describe` 下发,是有意的安全设计。
|
|
124
|
+
- 「API Key 环境变量名」已收进供应商高级选项、更名「凭据引用名 (apiKeyEnv)」:新增供应商时自动生成(`normalizeCredentialRef`),并对 **ID + 引用双重去重**——改名供应商会保留旧引用(write-only 无法搬移),只按 ID 去重会复活 `provider-1` 并继承已被占用的 `PROVIDER_1_API_KEY`(两个供应商同引用 = 共用同一把 key,写入互相覆盖)。此坑已由双重去重修复,存量撞引用靠 ⚠ 警示提示手动处理(改其中一个引用 → 重新写入)。
|
|
125
|
+
- 环境优先级:启动 shell 同名变量(只读、优先)> 存储的 key > 项目 `.env` > 用户 `.env`(credentials-local 分层);想用环境注入直接在启动环境 export 即可。
|
|
126
|
+
|
|
82
127
|
**曾踩坑**:`selected` 曾初始化为 `missing`(只含"可加"),而 configured 项 checkbox 显示 `checked:true` 却不在 selected 里——应用时 `kept = models.filter(m => cs.has(m.id))` 把已配置模型全部丢弃 → **已有模型消失**。修复 = selected 初始化为 `configured ∩ 端点`。
|
|
83
128
|
|
|
84
129
|
## host 半内部接口
|
|
@@ -86,6 +131,7 @@ DSH 原生模型渠道管理。两半结构:
|
|
|
86
131
|
- settings 接入用 **`ctx.inject(['settings'], (sctx) => {...})`**(settings 服务异步初始化,apply 时 `ctx.get('settings')` 为 undefined——曾经整个引擎静默失效,命名空间从未注册)
|
|
87
132
|
- 配置 schema(schemastery):`model-channels` 的虚模型/candidates/strategy/timeoutMs/cooldownMs/maxRetriesPerCandidate/speedTest;`model-channel-health` 的 records 7 天切片(单组 ≤2000 条)/speedResults/runtime/speedRequest+lastHandledNonce/testRequest+testResults+lastTestHandledNonce
|
|
88
133
|
- 引擎:sticky/round-robin/primary 三策略;首响应超时 + 流中空闲超时(动态 = max(timeoutMs, min(120s, ttft×2)));单候选原地重试(指数退避)耗尽才换;全炸清冷却重试一轮;测速 ttft/latency/hybrid/smart 四键(smart = 0.5×ttft_norm + 0.3×(1−reliability) + 0.2×latency_norm,reliability 贝叶斯平滑 `(success+2.5)/(total+5)`);测速失败进冷却;请求隔离按组
|
|
134
|
+
- 虚拟模型元数据:`reasoning.efforts` 七档(off…max)、**defaultEffort=max**——原生 `/model` 弹窗对新模型的自动填档与展示跟随该声明;会话内显式档位的跨会话恢复由 selector 插件的档位记忆层负责(`modelDirectories` 拦截,存 `model-channels.effortMemory`)
|
|
89
135
|
- 迁移:startup 时从工作区 `.channel-manager/config.json` 一次性迁入 `model-channels`(无遗留则忽略);完成后写 `legacyMigrated` 哨兵防止「清空组后重启复活」;fs 未就绪时 5s×6 重试
|
|
90
136
|
- 遗留 `.channel-manager/` 目录不再使用
|
|
91
137
|
|
|
@@ -100,6 +146,7 @@ react 经 `require('react')`;样式用 `ctx.effect` 自管理;`dsh.client: {
|
|
|
100
146
|
|
|
101
147
|
- 动态超时实现了首响应 + 流中空闲;全炸后「清冷却重试一轮」回溯,未实现「等待最早冷却」的睡眠分支
|
|
102
148
|
- 测速结果不入健康流水(pi 记);smart 键只统计真实请求
|
|
149
|
+
- **Token 字段只在新记录上出现**:host 升级重启前的存量健康记录无 token 字段,7 天视图对重启前的调用会低估 token(请求数/可用率不受影响);数据自重启后开始累积
|
|
103
150
|
- 配置里 provider 必须非虚拟路由(防自引用)
|
|
104
151
|
- 轮询渠道/健康统计面板需要 host 新代码(重启后生效);健康流水的数据在**实际请求过轮询组**后才出现
|
|
105
152
|
- `reasoningEfforts` 缺失(undefined)的 model 正确渲染(`|| {}` 兜底)
|
|
@@ -128,3 +175,6 @@ react 经 `require('react')`;样式用 `ctx.effect` 自管理;`dsh.client: {
|
|
|
128
175
|
19. **single slot 换占必须传负 priority**:`conversation.input.model` 是单占位 seat,cell = slot 本身;原生无 priority(= 0),插件同名注册同不传 → **exact-priority 撞格直接抛错**(「already has a registration at priority 0」→ apply 失败 → 整个插件含模型配置页签加载失败,面板全白)。规则:同 cell 多 entry 按 priority **升序、数值最小者渲染**,遮蔽原生传 `priority: -1`。注意 slot-catalog 的「Do NOT pass priority」只适用于**动态包**(guard 自动分配);静态 bundle 必须自己传。另:mock 验证 slots.register 不会暴露 occupancy 检查(mock 不抛)——验证座位替换必须复刻真实 SlotCore 撞格语义。选择器拆出后,回归测试随代码迁至 `../model-selector-search/tests/slot-priority.test.cjs`。
|
|
129
176
|
20. **诊断临时实例必须独立 home(`DSH_HOME=/tmp/dsh-diag dsh ...`)**:临时实例与主实例共用 `~/.dsh` 会并发写同一会话日志与 `session_projcache.json`——两进程各自的 seq 计数器交错追加,日志出现重复 seq → `corrupt session log: seq gap in committed region` → 会话 resume 直接拒绝,表现为该会话内模型目录加载失败(选择器「暂无可用模型」)。修复:解压 jsonl 删掉多余事件即可(后续 seq 连续则天然对齐),用 `session-persistence-jsonl` 的 `scanLog` 校验后压缩回写;杀进程前务必备份。
|
|
130
177
|
21. **适配器契约以安装运行时的 d.ts 为准,不能照抄源码仓快照**:源码仓较新、rc.2 运行时的 `LlmAdapter` 多一个必需的 `prepareCall(provider, model, signal) → Promise<{model, stream}>`(主分发路径 llm.stream/llm.prepareCall 都先走它再 `adapterCall.stream(options)`;`adapter.stream` 在 rc.2 服务层从不直调)。缺它的症状极具迷惑性:注册/目录/菜单全正常,**真实发对话**才报 `registration.adapter.prepareCall is not a function`。实现对齐 llm-pi-ai 的快照模式:prepare 时捕获一份配置快照,元数据与 dispatch 都出自同一代。回归:`tests/adapter-contract.test.cjs`(T3 直接解析安装版 d.ts 的 LlmAdapter 方法集做契约同步)。
|
|
178
|
+
22. **0.2 配置写入是 HMR 独占事务**:`settings.update` → `configEditor.edit()` → `hmr.runExclusive()`;在 `loader/volatile-update` 回调里回写会抛 `HMR transactions cannot be nested`(实测一段会话内 15 次,面板“测试”结果永远落不了盘)。事务内创建的**任何**异步资源(`AsyncResource` / `setTimeout` / `setInterval`)都继承事务上下文,**只有 `AsyncLocalStorage.exit()` 能切出**:`ctx.get('hmr').executing.exit(fn)`(仅当 `getStore()` 为真时切)。写入会被 `runExclusive` 排进队列、在本次事务结束后执行;**监听器保持同步、不要在外层事务里 await 它**(队列串行,互等即死锁)。
|
|
179
|
+
23. **整字段落盘 + 内存态被配置快照覆盖**:`settings.update` 是整字段替换;`reloadFromConfig()` 每次 volatile-update 都用配置快照整体覆盖 `state.records`,而健康 flush 有 2s 防抖 → **刚记下的一笔在落盘前就被内存覆盖**(症状:面板“测试”成功不入账,失败反被全局拦截器的 catch 记上)。修法:`pendingRecords` 缓冲,快照覆盖后把未落盘的补回,flush 成功后再清账。
|
|
180
|
+
24. **nonce 落盘时机与启动竞态**:`lastTestHandledNonce` / `lastHandledNonce` 必须在**得出结果之后**写(提前写会让“未就绪”的重试被自己的持久值挡掉);启动瞬间凭据服务尚未就绪时测试/测速会以 `MISSING_CREDENTIAL` 失败(凭据其实已在 `.credentials.yaml` 里),应识别为「还没就绪」→ 释放认领 + 5s 延时重试(上限 24 次),**不要**写成渠道故障;测速还必须在整组候选都因未就绪失败时**不落盘、不冷却**,否则一次启动重放就把所有渠道误判成故障。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@arcaneorion/dsh-model-channel-manager",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"main": "src/index.js",
|
|
6
6
|
"exports": {
|
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
"author": "arcanexis",
|
|
14
14
|
"repository": {
|
|
15
15
|
"type": "git",
|
|
16
|
-
"url": "git+https://github.com/
|
|
16
|
+
"url": "git+https://github.com/ArcaneOrion/dsh-model-channel-manager.git"
|
|
17
17
|
},
|
|
18
18
|
"keywords": [
|
|
19
19
|
"deepseek",
|
|
@@ -34,14 +34,21 @@
|
|
|
34
34
|
"access": "public"
|
|
35
35
|
},
|
|
36
36
|
"peerDependencies": {
|
|
37
|
-
"@deepseek-ai/
|
|
37
|
+
"@deepseek-ai/cordis": "^4.0.4",
|
|
38
|
+
"@deepseek-ai/dsh-llm": "0.2.0-rc.1",
|
|
39
|
+
"@deepseek-ai/dsh-settings": "0.2.0-rc.1",
|
|
40
|
+
"@deepseek-ai/dsh-client-connection": "0.2.0-rc.1",
|
|
41
|
+
"@deepseek-ai/dsh-client-ui-conversation": "0.2.0-rc.1",
|
|
42
|
+
"@deepseek-ai/dsh-api-remotes": "0.2.0-rc.1",
|
|
43
|
+
"@deepseek-ai/schemastery": ">=3.18.4",
|
|
44
|
+
"react": "^18.3.1"
|
|
38
45
|
},
|
|
39
46
|
"dsh": {
|
|
40
47
|
"bundle": {
|
|
41
48
|
"patch": "./cordis.patch.yml"
|
|
42
49
|
},
|
|
43
50
|
"client": {
|
|
44
|
-
"inject": ["slots"
|
|
51
|
+
"inject": ["slots"],
|
|
45
52
|
"platform": "web"
|
|
46
53
|
}
|
|
47
54
|
}
|
package/src/client.js
CHANGED
|
@@ -12,6 +12,103 @@ window.__ModuleLoader__.load({
|
|
|
12
12
|
let apiRef = null
|
|
13
13
|
const savedKeys = {}
|
|
14
14
|
|
|
15
|
+
// dsh 0.2:客户端远程面由 connection.api 改为 ctx.remote;参数改为位置参数,结果统一为
|
|
16
|
+
// RemoteResult。本适配层把两者还原成 0.1 的调用形状(对象入参 + {result:{ok,value}})。
|
|
17
|
+
// 另:0.2 的 settings 命名空间恒等于「插件行 id」,本插件的两套旧命名空间
|
|
18
|
+
// (model-channels / model-channel-health)现由其实例配置的字段承载,在此合成回旧视图。
|
|
19
|
+
const NS_MCM = 'model-channel-manager'
|
|
20
|
+
const makeLegacyApi = (remote) => {
|
|
21
|
+
const wrap = (res) => ({ result: res || { ok: false } });
|
|
22
|
+
const readMcm = async () => {
|
|
23
|
+
const res = await remote.settings.describe()
|
|
24
|
+
if (!res || res.ok !== true) return null
|
|
25
|
+
return ((res.value && res.value.namespaces) || []).find((n) => n && n.ns === NS_MCM) || null
|
|
26
|
+
}
|
|
27
|
+
return {
|
|
28
|
+
settings: {
|
|
29
|
+
describe: async () => {
|
|
30
|
+
const res = await remote.settings.describe()
|
|
31
|
+
if (!res || res.ok !== true) return wrap(res)
|
|
32
|
+
const namespaces = (res.value && res.value.namespaces) || []
|
|
33
|
+
const mcm = namespaces.find((n) => n && n.ns === NS_MCM)
|
|
34
|
+
const value = (mcm && mcm.value) || {}
|
|
35
|
+
const user = (mcm && mcm.user) || {}
|
|
36
|
+
return { result: { ok: true, value: { namespaces: [...namespaces,
|
|
37
|
+
{ ns: 'model-channels', value: { groups: value.groups || [], providerOrder: value.providerOrder || [], effortMemory: value.effortMemory || {} }, user: { groups: user.groups, providerOrder: user.providerOrder }, revision: (mcm && mcm.revision) || 0 },
|
|
38
|
+
{ ns: 'model-channel-health', value: value.health || {}, user: (user && user.health) || {}, revision: (mcm && mcm.revision) || 0 },
|
|
39
|
+
] } } }
|
|
40
|
+
},
|
|
41
|
+
// 旧签名:update({ ns, patch })
|
|
42
|
+
update: async (args) => {
|
|
43
|
+
const ns = (args && args.ns) || NS_MCM
|
|
44
|
+
const patch = (args && args.patch) || {}
|
|
45
|
+
if (ns === 'model-channels') return wrap(await remote.settings.update(NS_MCM, patch, undefined))
|
|
46
|
+
const row = await readMcm()
|
|
47
|
+
const health = (row && row.value && row.value.health) || {}
|
|
48
|
+
return wrap(await remote.settings.update(NS_MCM, { health: Object.assign({}, health, patch) }, undefined))
|
|
49
|
+
},
|
|
50
|
+
// 旧签名:mutate({ ns, ops })——0.2 只对真实行命名空间提供 path-ops;
|
|
51
|
+
// 本插件自身的健康子树用合并写实现等价语义。
|
|
52
|
+
mutate: async (args) => {
|
|
53
|
+
const ns = (args && args.ns) || NS_MCM
|
|
54
|
+
const ops = (args && args.ops) || []
|
|
55
|
+
if (ns !== 'model-channels' && ns !== 'model-channel-health' && ns !== NS_MCM) {
|
|
56
|
+
return wrap(await remote.settings.mutate(ns, ops, undefined))
|
|
57
|
+
}
|
|
58
|
+
const row = await readMcm()
|
|
59
|
+
const health = Object.assign({}, (row && row.value && row.value.health) || {})
|
|
60
|
+
const top = {
|
|
61
|
+
groups: (row && row.value && row.value.groups) || [],
|
|
62
|
+
providerOrder: (row && row.value && row.value.providerOrder) || [],
|
|
63
|
+
}
|
|
64
|
+
for (const op of ops) {
|
|
65
|
+
const path = (op && op.path) || []
|
|
66
|
+
const root = (ns === 'model-channels') ? (path[0] === 'groups' ? top : (path[0] === 'providerOrder' ? top : health)) : health
|
|
67
|
+
if (path.length === 0) continue
|
|
68
|
+
const key = path[0]
|
|
69
|
+
const rest = path.slice(1)
|
|
70
|
+
if (rest.length === 0) {
|
|
71
|
+
if (op.op === 'set') root[key] = op.value
|
|
72
|
+
else delete root[key]
|
|
73
|
+
} else if (rest.length === 1) {
|
|
74
|
+
const child = Object.assign({}, root[key] || {})
|
|
75
|
+
if (op.op === 'set') child[rest[0]] = op.value
|
|
76
|
+
else delete child[rest[0]]
|
|
77
|
+
root[key] = child
|
|
78
|
+
}
|
|
79
|
+
}
|
|
80
|
+
if (ns === 'model-channels') {
|
|
81
|
+
return wrap(await remote.settings.update(NS_MCM, { groups: top.groups, providerOrder: top.providerOrder }, undefined))
|
|
82
|
+
}
|
|
83
|
+
return wrap(await remote.settings.update(NS_MCM, { health }, undefined))
|
|
84
|
+
},
|
|
85
|
+
},
|
|
86
|
+
credentials: {
|
|
87
|
+
// 旧签名:set({ ref, value })
|
|
88
|
+
set: async (args) => wrap(await remote.credentials.set((args && args.ref) || '', (args && args.value) || '')),
|
|
89
|
+
// 旧签名:describe({ refs: [...] }) → value.credentials[ref]
|
|
90
|
+
describe: async (args) => {
|
|
91
|
+
const refs = (args && args.refs) || []
|
|
92
|
+
const res = await remote.credentials.describe(refs)
|
|
93
|
+
if (!res || res.ok !== true) return wrap(res)
|
|
94
|
+
return { result: { ok: true, value: { credentials: res.value || {} } } }
|
|
95
|
+
},
|
|
96
|
+
},
|
|
97
|
+
llm: {
|
|
98
|
+
// 旧签名:discoverModels({ settingsNs, ...request }) → value.models
|
|
99
|
+
discoverModels: async (args) => {
|
|
100
|
+
const settingsNs = (args && args.settingsNs) || 'llm-pi-ai'
|
|
101
|
+
const request = Object.assign({}, args)
|
|
102
|
+
delete request.settingsNs
|
|
103
|
+
const res = await remote.llm.discoverModels(settingsNs, request)
|
|
104
|
+
if (!res || res.ok !== true) return wrap(res)
|
|
105
|
+
const value = res.value
|
|
106
|
+
return { result: { ok: true, value: Array.isArray(value) ? { models: value } : value } }
|
|
107
|
+
},
|
|
108
|
+
},
|
|
109
|
+
}
|
|
110
|
+
}
|
|
111
|
+
|
|
15
112
|
const CSS = `
|
|
16
113
|
.mcm-root { display:flex; flex-direction:column; height:100%; overflow:hidden; font-size:13px; color:var(--dsw-alias-label-primary); background:var(--dsw-alias-bg-base); }
|
|
17
114
|
.mcm-header { display:flex; align-items:center; justify-content:space-between; padding:12px 24px; border-bottom:1px solid var(--dsw-alias-border-l1); background:var(--dsw-alias-bg-layer-1); flex-shrink:0; }
|
|
@@ -549,7 +646,8 @@ window.__ModuleLoader__.load({
|
|
|
549
646
|
const saveKey = (name) => {
|
|
550
647
|
const ref = (providers[name] || {}).apiKeyEnv || ''
|
|
551
648
|
const val = (keyInput[name] || '').trim()
|
|
552
|
-
if (!
|
|
649
|
+
if (!val || !apiRef) return
|
|
650
|
+
if (!ref) { setLive((l) => Object.assign({}, l, { [name + '_key']: 'err: 缺少凭据引用名(见高级选项)' })); return }
|
|
553
651
|
savedKeys[name] = val
|
|
554
652
|
apiRef.credentials.set({ ref, value: val }).then((resp) => {
|
|
555
653
|
const r = resp && resp.result ? resp.result : resp
|
|
@@ -570,6 +668,13 @@ window.__ModuleLoader__.load({
|
|
|
570
668
|
if (p.displayName && String(p.displayName).toLowerCase().includes(pqLower)) return true
|
|
571
669
|
return (p.models || []).some((m) => m && ((m.id || '').toLowerCase().includes(pqLower) || (m.name || '').toLowerCase().includes(pqLower)))
|
|
572
670
|
})
|
|
671
|
+
// 凭据引用占用计数:同一 apiKeyEnv 被多个供应商引用 = 共用同一把 key(任一处
|
|
672
|
+
// 「写入存储」覆盖共享引用,全部共引供应商随之同 key)——用于卡片 ⚠ 警示
|
|
673
|
+
const refUsage = new Map()
|
|
674
|
+
for (const p of Object.values(providers || {})) {
|
|
675
|
+
const ref = p && typeof p.apiKeyEnv === 'string' ? p.apiKeyEnv.trim() : ''
|
|
676
|
+
if (ref) refUsage.set(ref, (refUsage.get(ref) || 0) + 1)
|
|
677
|
+
}
|
|
573
678
|
const cards = visibleEntries.map(({ name, p, realIdx }) => {
|
|
574
679
|
const isOpen = pqLower ? true : !!exp[name]
|
|
575
680
|
const models = p.models || []
|
|
@@ -619,21 +724,23 @@ window.__ModuleLoader__.load({
|
|
|
619
724
|
sel('协议架构 (API)', '当前支持的请求格式', p.api || 'openai-completions', APIS, (v) => updateP(name, { api: v }))
|
|
620
725
|
),
|
|
621
726
|
el('div', { className: 'mcm-row' },
|
|
622
|
-
tf('Base URL', 'API 端点基础地址', p.baseURL || '', (v) => updateP(name, { baseURL: v }), true)
|
|
623
|
-
tf('API Key 环境变量名', '凭据存储引用名', p.apiKeyEnv || '', (v) => updateP(name, { apiKeyEnv: v }), true)
|
|
727
|
+
tf('Base URL', 'API 端点基础地址', p.baseURL || '', (v) => updateP(name, { baseURL: v }), true)
|
|
624
728
|
),
|
|
625
|
-
field('
|
|
729
|
+
field('API Key', '粘贴或输入后失焦即自动写入 DSH 凭据存储(~/.dsh/.credentials.yaml,0600,write-only 读不回);清空输入框不会删除已存 key;启动环境若有同名变量则其只读优先', false,
|
|
626
730
|
el('div', { style: { display: 'flex', gap: 8, alignItems: 'center' } },
|
|
627
|
-
el('input', { type: 'password', className: 'mcm-in mono', style: { flex: 1 }, placeholder: '
|
|
628
|
-
|
|
629
|
-
live[name + '_key'] === 'saved' ? el('span', { style: { color: 'var(--dsw-alias-state-success-primary)', fontSize: 12 } }, '✓ 已保存') : null
|
|
731
|
+
el('input', { type: 'password', className: 'mcm-in mono', style: { flex: 1 }, placeholder: '粘贴 API Key (sk-...),失焦自动保存', value: keyInput[name] || '', onChange: (e) => setKeyInput((k) => Object.assign({}, k, { [name]: e.target.value })), onBlur: () => { const v = (keyInput[name] || '').trim(); if (v && v !== savedKeys[name]) saveKey(name) } }),
|
|
732
|
+
live[name + '_key'] === 'saved' ? el('span', { style: { color: 'var(--dsw-alias-state-success-primary)', fontSize: 12 } }, '✓ 已保存') : (live[name + '_key'] && String(live[name + '_key']).indexOf('err') === 0 ? el('span', { style: { color: 'var(--dsw-alias-state-error-primary)', fontSize: 12 } }, live[name + '_key']) : null)
|
|
630
733
|
)
|
|
631
734
|
),
|
|
735
|
+
(p.apiKeyEnv && (refUsage.get(p.apiKeyEnv) || 0) > 1) ? el('div', { style: { color: 'var(--dsw-alias-state-warn-primary)', fontSize: 11, lineHeight: 1.6 } }, '⚠ 该引用被 ' + refUsage.get(p.apiKeyEnv) + ' 个供应商共用 = 共用同一把 Key,任一处写入会同时覆盖所有供应商;若非有意,请在高级选项中修改其中一个「凭据引用名」并重新写入') : null,
|
|
632
736
|
el('div', { style: { borderTop: '1px dashed var(--dsw-alias-border-l2)', paddingTop: 10 } },
|
|
633
737
|
el('div', { style: { cursor: 'pointer', fontSize: 12, fontWeight: 600, color: 'var(--dsw-alias-label-secondary)', display: 'flex', alignItems: 'center', gap: 6 }, onClick: () => setSa((s) => Object.assign({}, s, { [name]: !s[name] })) },
|
|
634
738
|
el('span', null, sa[name] ? '▼' : '▶'), '供应商高级选项 (Headers, 传输, 超时, 图片预算, 重试策略, Compat)'
|
|
635
739
|
),
|
|
636
740
|
sa[name] ? el('div', { style: { display: 'flex', flexDirection: 'column', gap: 12, marginTop: 10 } },
|
|
741
|
+
el('div', { className: 'mcm-row' },
|
|
742
|
+
tf('凭据引用名 (apiKeyEnv)', 'key 在凭据存储(~/.dsh/.credentials.yaml)中的引用名,write-only 读不回已存值;启动环境同名变量只读优先;修改后需重新写入 key', p.apiKeyEnv || '', (v) => updateP(name, { apiKeyEnv: v }), true)
|
|
743
|
+
),
|
|
637
744
|
el('div', { className: 'mcm-row' },
|
|
638
745
|
nf('默认 Context Window', '未单独配置模型时的默认窗口', p.defaultContextWindow || 0, (v) => updateP(name, { defaultContextWindow: v })),
|
|
639
746
|
nf('默认 Max Tokens', '未单独配置模型时的最大输出', p.defaultMaxTokens || 0, (v) => updateP(name, { defaultMaxTokens: v }))
|
|
@@ -768,7 +875,14 @@ window.__ModuleLoader__.load({
|
|
|
768
875
|
...cards,
|
|
769
876
|
visibleEntries.length === 0 && pqLower ? el('div', { className: 'mcm-empty' }, '无匹配「' + pq.trim() + '」的供应商或模型') : null,
|
|
770
877
|
btn('+新增提供商 (Provider)', () => {
|
|
771
|
-
|
|
878
|
+
// 双重去重:provider ID 与凭据引用都不得与现有草稿冲突。改名供应商会保留旧
|
|
879
|
+
// 引用(write-only 无法搬移),只按 ID 去重会复活 provider-1 并继承已被占用
|
|
880
|
+
// 的 PROVIDER_1_API_KEY——两个供应商同引用 = 共用同一把 key(写入互相覆盖)
|
|
881
|
+
const idsTaken = new Set(Object.keys(providers || {}))
|
|
882
|
+
const refsTaken = new Set(Object.values(providers || {}).map((p) => p && p.apiKeyEnv).filter(Boolean))
|
|
883
|
+
let i = 1
|
|
884
|
+
while (idsTaken.has('provider-' + i) || refsTaken.has(normalizeCredentialRef('provider-' + i + '_api_key'))) i++
|
|
885
|
+
const nm = 'provider-' + i
|
|
772
886
|
const keyRef = normalizeCredentialRef(nm + '_api_key')
|
|
773
887
|
setDraft((d) => Object.assign({}, d || {}, { [nm]: { api: 'openai-completions', baseURL: '', apiKeyEnv: keyRef, displayName: nm, models: [], headers: Object.assign({}, DEFAULT_HEADERS) } }))
|
|
774
888
|
setExp((e) => Object.assign({}, e, { [nm]: true }))
|
|
@@ -823,10 +937,9 @@ window.__ModuleLoader__.load({
|
|
|
823
937
|
const id = uniqueSuffixName('group-', (n) => base.some((g) => g && g.id === n))
|
|
824
938
|
return base.concat([{
|
|
825
939
|
id,
|
|
826
|
-
//
|
|
827
|
-
//
|
|
828
|
-
//
|
|
829
|
-
// 附加图片直接被拒(MODEL_DOES_NOT_SUPPORT_IMAGES),按直觉默认放开
|
|
940
|
+
// 单一身份:显示名 = 组唯一 ID(save 时统一归一化 virtualModel.name = id,
|
|
941
|
+
// 不会漂移)。默认模态含 image:host 对当前会话模型做 resolveModelInfo,
|
|
942
|
+
// 缺 image 时附加图片直接被拒(MODEL_DOES_NOT_SUPPORT_IMAGES),按直觉默认放开
|
|
830
943
|
virtualModel: { name: id, reasoning: true, input: ['text', 'image'], contextWindow: 1048576, maxTokens: 131072 },
|
|
831
944
|
candidates: [],
|
|
832
945
|
strategy: 'sticky',
|
|
@@ -871,7 +984,7 @@ window.__ModuleLoader__.load({
|
|
|
871
984
|
return el('div', { className: 'mcm-card', key: 'rr_card_' + i },
|
|
872
985
|
el('div', { className: 'mcm-card-h', onClick: () => setExpanded((e) => Object.assign({}, e, { [i]: !e[i] })) },
|
|
873
986
|
el('span', { style: { fontSize: 11, color: 'var(--dsw-alias-label-tertiary)' } }, isOpen ? '▼' : '▶'),
|
|
874
|
-
el('div', { className: 'mcm-card-title' }, el('span', null,
|
|
987
|
+
el('div', { className: 'mcm-card-title' }, el('span', null, g.id), el('span', { className: 'mcm-badge brand' }, 'roundrobin/' + g.id)),
|
|
875
988
|
el('span', { className: 'mcm-badge' }, (g.candidates || []).length + ' 候选'),
|
|
876
989
|
el('span', { className: 'mcm-badge' }, g.strategy || 'sticky'),
|
|
877
990
|
el('div', { style: { marginLeft: 'auto', display: 'flex', gap: 6 }, onClick: (e) => e.stopPropagation() },
|
|
@@ -882,8 +995,7 @@ window.__ModuleLoader__.load({
|
|
|
882
995
|
),
|
|
883
996
|
isOpen ? el('div', { className: 'mcm-editor' },
|
|
884
997
|
el('div', { className: 'mcm-row' },
|
|
885
|
-
tf('组唯一 ID', '对应虚拟路由 roundrobin/<id
|
|
886
|
-
tf('虚拟模型呈现名', '对话侧栏显示的名字', vm.name || '', (v) => patchGroupPath(i, ['virtualModel', 'name'], v))
|
|
998
|
+
tf('组唯一 ID', '对应虚拟路由 roundrobin/<id>,也是模型选择器里显示的名字(改这里即改名)', g.id || '', (v) => patchGroup(i, { id: v }), true)
|
|
887
999
|
),
|
|
888
1000
|
field('输入模态', '声明虚拟模型接受的输入;不含图片时附加图片会被 host 直接拒绝(MODEL_DOES_NOT_SUPPORT_IMAGES)', false,
|
|
889
1001
|
el('div', { style: { display: 'flex', gap: 16, alignItems: 'center', height: 32 } },
|
|
@@ -979,12 +1091,18 @@ window.__ModuleLoader__.load({
|
|
|
979
1091
|
lastTs: 0,
|
|
980
1092
|
lastOk: null,
|
|
981
1093
|
lastCode: null,
|
|
982
|
-
recentErrors: 0
|
|
1094
|
+
recentErrors: 0,
|
|
1095
|
+
tokSum: 0
|
|
983
1096
|
})
|
|
984
1097
|
}
|
|
985
1098
|
}
|
|
986
1099
|
}
|
|
987
1100
|
|
|
1101
|
+
// 窗口级 token 聚合(计费口径 = input + cacheRead + cacheWrite + output,与 dsh-token-meter 一致)
|
|
1102
|
+
let totalTokIn = 0
|
|
1103
|
+
let totalTokOut = 0
|
|
1104
|
+
let totalTokCache = 0
|
|
1105
|
+
|
|
988
1106
|
// 累加时间窗口内的真实流水记录
|
|
989
1107
|
for (const e of allEvents) {
|
|
990
1108
|
if (!e || !e.provider || !e.model) continue
|
|
@@ -1004,11 +1122,19 @@ window.__ModuleLoader__.load({
|
|
|
1004
1122
|
lastTs: 0,
|
|
1005
1123
|
lastOk: null,
|
|
1006
1124
|
lastCode: null,
|
|
1007
|
-
recentErrors: 0
|
|
1125
|
+
recentErrors: 0,
|
|
1126
|
+
tokSum: 0
|
|
1008
1127
|
}
|
|
1009
1128
|
pMap.set(e.model, a)
|
|
1010
1129
|
}
|
|
1011
1130
|
a.total++
|
|
1131
|
+
const tokIn = (e && e.inputTokens) || 0
|
|
1132
|
+
const tokOut = (e && e.outputTokens) || 0
|
|
1133
|
+
const tokCache = ((e && e.cacheReadTokens) || 0) + ((e && e.cacheWriteTokens) || 0)
|
|
1134
|
+
totalTokIn += tokIn
|
|
1135
|
+
totalTokOut += tokOut
|
|
1136
|
+
totalTokCache += tokCache
|
|
1137
|
+
a.tokSum += tokIn + tokOut + tokCache
|
|
1012
1138
|
if (e.ok) {
|
|
1013
1139
|
a.success++
|
|
1014
1140
|
if (e.ttftMs != null && e.ttftMs >= 0) a.ttftSum += e.ttftMs
|
|
@@ -1088,6 +1214,14 @@ window.__ModuleLoader__.load({
|
|
|
1088
1214
|
el('span', { className: 'mcm-metric-value' }, totalRequests),
|
|
1089
1215
|
el('span', { style: { fontSize: 11, color: 'var(--dsw-alias-label-tertiary)' } }, '真实上游交互捕获')
|
|
1090
1216
|
),
|
|
1217
|
+
el('div', { className: 'mcm-metric-card' },
|
|
1218
|
+
el('span', { className: 'mcm-metric-label' }, '总 Token 用量'),
|
|
1219
|
+
el('span', { className: 'mcm-metric-value' }, (totalTokIn + totalTokOut + totalTokCache).toLocaleString('zh-CN')),
|
|
1220
|
+
el('span', { style: { fontSize: 11, color: 'var(--dsw-alias-label-tertiary)' } },
|
|
1221
|
+
'输入 ' + totalTokIn.toLocaleString('zh-CN') + ' · 输出 ' + totalTokOut.toLocaleString('zh-CN') +
|
|
1222
|
+
(totalTokCache > 0 ? ' · 缓存 ' + totalTokCache.toLocaleString('zh-CN') : '')
|
|
1223
|
+
)
|
|
1224
|
+
),
|
|
1091
1225
|
el('div', { className: 'mcm-metric-card' },
|
|
1092
1226
|
el('span', { className: 'mcm-metric-label' }, '窗口可用率'),
|
|
1093
1227
|
el('span', { className: 'mcm-metric-value', style: { color: totalSuccess === totalRequests ? 'var(--dsw-alias-state-success-primary)' : 'var(--dsw-alias-label-primary)' } }, globalRate),
|
|
@@ -1135,7 +1269,7 @@ window.__ModuleLoader__.load({
|
|
|
1135
1269
|
)
|
|
1136
1270
|
),
|
|
1137
1271
|
el('div', { style: { display: 'flex', justifyContent: 'space-between', fontSize: 11, color: 'var(--dsw-alias-label-tertiary)', borderTop: '1px solid var(--dsw-alias-border-l1)', paddingTop: 8 } },
|
|
1138
|
-
el('span', null, 'TTFT: ', el('strong', { style: { color: 'var(--dsw-alias-label-primary)' } }, m.avgTtft)),
|
|
1272
|
+
el('span', null, 'TTFT: ', el('strong', { style: { color: 'var(--dsw-alias-label-primary)' } }, m.avgTtft), ' · Tokens: ', el('strong', { style: { color: 'var(--dsw-alias-label-primary)' } }, (m.tokSum || 0).toLocaleString('zh-CN'))),
|
|
1139
1273
|
el('span', null, '最近: ', el('span', { style: { color: 'var(--dsw-alias-label-secondary)' } }, m.lastTimeStr))
|
|
1140
1274
|
)
|
|
1141
1275
|
))
|
|
@@ -1350,7 +1484,9 @@ window.__ModuleLoader__.load({
|
|
|
1350
1484
|
const orderPatch = draft ? { providerOrder: Object.keys(cleanProviders) } : {}
|
|
1351
1485
|
const p2 = apiRef.settings.update({
|
|
1352
1486
|
ns: 'model-channels',
|
|
1353
|
-
|
|
1487
|
+
// 命名单一身份:写入时统一 virtualModel.name = 组 id——旧的独立呈现名
|
|
1488
|
+
// (如遗留的 group-1)在下一次保存时自动归一,无需迁移
|
|
1489
|
+
patch: Object.assign({ groups: (channelsDraft || []).map((g) => Object.assign({}, g, { virtualModel: Object.assign({}, g.virtualModel, { name: g.id }) })) }, orderPatch)
|
|
1354
1490
|
}).then((resp) => {
|
|
1355
1491
|
const r = resp && resp.result ? resp.result : resp
|
|
1356
1492
|
if (r && r.ok === false) throw new Error((r.error && (r.error.message || r.error)) || 'model-channels save failed')
|
|
@@ -1389,8 +1525,9 @@ window.__ModuleLoader__.load({
|
|
|
1389
1525
|
|
|
1390
1526
|
|
|
1391
1527
|
function apply(ctx) {
|
|
1392
|
-
|
|
1393
|
-
|
|
1528
|
+
// dsh 0.2:远程调用面为 ctx.remote(connection.api 已移除)。保留 0.1 的调用形状,
|
|
1529
|
+
// 由下面的适配层把位置参数/RemoteResult 还原成旧的 {result:{ok,value}} 与对象入参。
|
|
1530
|
+
ctx.inject(['remote'], (scope) => { apiRef = makeLegacyApi(scope.remote) })
|
|
1394
1531
|
ctx.effect(() => {
|
|
1395
1532
|
const tag = document.createElement('style')
|
|
1396
1533
|
tag.dataset.mcmStyle = ''
|
|
@@ -1409,6 +1546,6 @@ window.__ModuleLoader__.load({
|
|
|
1409
1546
|
))
|
|
1410
1547
|
}
|
|
1411
1548
|
|
|
1412
|
-
return { name: 'model-channel-manager', inject: ['slots', '
|
|
1549
|
+
return { name: 'model-channel-manager', inject: ['slots', 'remote', 'remote.settings', 'remote.credentials', 'remote.llm'], apply }
|
|
1413
1550
|
},
|
|
1414
1551
|
})
|
package/src/index.js
CHANGED
|
@@ -5,15 +5,9 @@ const ROUTE_PREFIX = 'roundrobin/';
|
|
|
5
5
|
const GROUP_ID_RE = /^[a-z0-9][a-z0-9-]*$/;
|
|
6
6
|
const NS_CONFIG = 'model-channels';
|
|
7
7
|
const NS_HEALTH = 'model-channel-health';
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
}).loose(true)).default([]),
|
|
12
|
-
// 面板供应商顺序的持久化载体:settings-file 的 patchNode 对 map 键序是盲的
|
|
13
|
-
// (纯重排 = 零 diff),但数组是 deepEqual 不等即整值 setIn(原位、保序)。
|
|
14
|
-
// 顺序存成数组数据而不是对象键序元数据,重启后从文件真实还原。引擎不消费它。
|
|
15
|
-
providerOrder: z.array(z.string()).default([]),
|
|
16
|
-
}).loose(true);
|
|
8
|
+
// dsh 0.2:settings 不再是可动态注册的命名空间总线,而是「插件行实例配置」的表单投影。
|
|
9
|
+
// 因此原先的 model-channels / model-channel-health 两个命名空间合并为本插件的 Config
|
|
10
|
+
// (字段 groups / providerOrder / effortMemory / health),命名空间 id 即本行 id。
|
|
17
11
|
const HEALTH_SCHEMA = z.object({
|
|
18
12
|
records: z.dict(z.array(z.any())).default({}),
|
|
19
13
|
speedResults: z.dict(z.array(z.any())).default({}),
|
|
@@ -25,7 +19,34 @@ const HEALTH_SCHEMA = z.object({
|
|
|
25
19
|
lastTestHandledNonce: z.number().default(0),
|
|
26
20
|
legacyMigrated: z.any(),
|
|
27
21
|
}).loose(true);
|
|
28
|
-
export
|
|
22
|
+
export const Config = z.object({
|
|
23
|
+
groups: z.array(z.any()).default([]).volatile(),
|
|
24
|
+
providerOrder: z.array(z.string()).default([]).volatile(),
|
|
25
|
+
effortMemory: z.dict(z.string()).default({}).volatile(),
|
|
26
|
+
health: HEALTH_SCHEMA.volatile(),
|
|
27
|
+
}).loose(true);
|
|
28
|
+
export function apply(ctx, config) {
|
|
29
|
+
// 本行在 profile 中的 entry id 就是 settings 命名空间;缺失时回落到包名。
|
|
30
|
+
const SELF_NS = (ctx.fiber && ctx.fiber.entry && ctx.fiber.entry.options && ctx.fiber.entry.options.id) || 'model-channel-manager';
|
|
31
|
+
const refOf = (key, fallback) => {
|
|
32
|
+
const field = config && config[key];
|
|
33
|
+
if (field && typeof field.get === 'function') {
|
|
34
|
+
try {
|
|
35
|
+
return field.get();
|
|
36
|
+
}
|
|
37
|
+
catch (_e) {
|
|
38
|
+
return fallback;
|
|
39
|
+
}
|
|
40
|
+
}
|
|
41
|
+
return fallback;
|
|
42
|
+
};
|
|
43
|
+
const healthOf = () => refOf('health', null) || {};
|
|
44
|
+
const cfgOf = () => ({
|
|
45
|
+
groups: refOf('groups', []) || [],
|
|
46
|
+
providerOrder: refOf('providerOrder', []) || [],
|
|
47
|
+
effortMemory: refOf('effortMemory', {}) || {},
|
|
48
|
+
});
|
|
49
|
+
|
|
29
50
|
const isContentChunk = (c) => {
|
|
30
51
|
if (!c) return false;
|
|
31
52
|
if (c.type === 'text-delta' || c.type === 'reasoning-delta') return c.text !== '';
|
|
@@ -266,21 +287,41 @@ export function apply(ctx) {
|
|
|
266
287
|
const now = Date.now();
|
|
267
288
|
const key = gid || 'global';
|
|
268
289
|
const list = state.records[key] || (state.records[key] = []);
|
|
269
|
-
|
|
290
|
+
const rec = { ts: now, provider: cand.provider, model: cand.model, ok: entry.ok, ttftMs: entry.ttftMs, latencyMs: entry.latencyMs, code: entry.code || null };
|
|
291
|
+
// 上游真实 token 用量(来自 usage StreamChunk,计费口径与 dsh-token-meter 一致):
|
|
292
|
+
// 展平写入记录条目;缺失/非法字段不写,旧记录按 0 处理(向后兼容)。
|
|
293
|
+
const u = entry.usage;
|
|
294
|
+
if (u && typeof u === 'object') {
|
|
295
|
+
if (Number.isFinite(u.inputTokens)) rec.inputTokens = u.inputTokens;
|
|
296
|
+
if (Number.isFinite(u.outputTokens)) rec.outputTokens = u.outputTokens;
|
|
297
|
+
if (Number.isFinite(u.cacheReadTokens)) rec.cacheReadTokens = u.cacheReadTokens;
|
|
298
|
+
if (Number.isFinite(u.cacheWriteTokens)) rec.cacheWriteTokens = u.cacheWriteTokens;
|
|
299
|
+
if (Number.isFinite(u.reasoningTokens)) rec.reasoningTokens = u.reasoningTokens;
|
|
300
|
+
}
|
|
301
|
+
list.push(rec);
|
|
270
302
|
const cutoff = now - 7 * 24 * 3600 * 1000;
|
|
271
|
-
|
|
303
|
+
// 每键保留 300 条:健康页签 30m/24h 视图与选择器置顶绰绰有余,同时把
|
|
304
|
+
// settings.yaml 的体量压到可读(2000 条/键时该文件一度膨胀到 2 万行)
|
|
305
|
+
state.records[key] = list.filter((e) => e.ts >= cutoff).slice(-300);
|
|
306
|
+
pendingRecords.push({ key, rec });
|
|
272
307
|
persistHealth();
|
|
273
308
|
};
|
|
274
309
|
const persistSpeedResults = async (r) => {
|
|
275
310
|
if (bus === null) return;
|
|
276
|
-
await bus.
|
|
311
|
+
await bus.writeHealth({ speedResults: clone(r), runtime: await persistRuntime() });
|
|
277
312
|
};
|
|
278
313
|
// ---------- 健康持久化节流:合并 2s 窗口内的写入,避免每个请求全量重写 settings ----------
|
|
279
314
|
let healthFlushHandle = null;
|
|
280
315
|
let healthDirty = false;
|
|
316
|
+
// 尚未落盘的流水:reloadFromConfig 每次 volatile-update 都用配置快照整体覆盖 state.records,
|
|
317
|
+
// 而 flush 有 2s 防抖——不缓冲的话,刚记下的一笔会在落盘前被内存覆盖(实测「测试成功不入账」就是这个)。
|
|
318
|
+
let pendingRecords = [];
|
|
281
319
|
const flushHealthNow = async () => {
|
|
282
320
|
if (bus === null) return;
|
|
283
|
-
|
|
321
|
+
const flushed = pendingRecords.slice();
|
|
322
|
+
await bus.writeHealth({ records: clone(pullRecords()), speedResults: clone(pullSpeedResults()), runtime: await persistRuntime() });
|
|
323
|
+
// 写入成功后才清账;失败就留着,等下一次 flush 重试
|
|
324
|
+
pendingRecords = pendingRecords.filter((p) => !flushed.includes(p));
|
|
284
325
|
};
|
|
285
326
|
const persistHealth = () => {
|
|
286
327
|
healthDirty = true;
|
|
@@ -315,6 +356,7 @@ export function apply(ctx) {
|
|
|
315
356
|
let started = false;
|
|
316
357
|
let ttftMs = null;
|
|
317
358
|
let lastAt = Date.now();
|
|
359
|
+
let lastUsage = null;
|
|
318
360
|
const closeInner = async () => { try {
|
|
319
361
|
const c = inner.return ? inner.return() : null;
|
|
320
362
|
if (c && c.then)
|
|
@@ -341,6 +383,10 @@ export function apply(ctx) {
|
|
|
341
383
|
}
|
|
342
384
|
lastAt = Date.now();
|
|
343
385
|
const chunk = next.value;
|
|
386
|
+
// 适配器在 finish 前发 usage 块(pi-ai done/error 都带);捕获后随记录写入,
|
|
387
|
+
// 与请求/健康计数保持同一记录、同一窗口口径
|
|
388
|
+
if (chunk && chunk.type === 'usage' && chunk.usage && typeof chunk.usage === 'object')
|
|
389
|
+
lastUsage = chunk.usage;
|
|
344
390
|
if (chunk === undefined)
|
|
345
391
|
throw { code: 'STREAM_CLOSED', message: 'channel stream ended without a terminal chunk' };
|
|
346
392
|
if (!started) {
|
|
@@ -355,11 +401,11 @@ export function apply(ctx) {
|
|
|
355
401
|
const reason = chunk.reason || {};
|
|
356
402
|
// 引擎提前终止的内层流不会被全局拦截器完整排水记账,这里自行记录
|
|
357
403
|
if (isSuccessReason(reason)) {
|
|
358
|
-
recordHealth(cand.provider, cand, { ok: false, ttftMs: null, latencyMs: null, code: 'EMPTY_RESPONSE' });
|
|
404
|
+
recordHealth(cand.provider, cand, { ok: false, ttftMs: null, latencyMs: null, code: 'EMPTY_RESPONSE', usage: lastUsage });
|
|
359
405
|
throw { code: 'EMPTY_RESPONSE', message: 'model returned a completed response with no content' };
|
|
360
406
|
}
|
|
361
407
|
const preCode = (reason.failure && reason.failure.code) || 'STREAM_ERROR';
|
|
362
|
-
recordHealth(cand.provider, cand, { ok: false, ttftMs: null, latencyMs: null, code: preCode });
|
|
408
|
+
recordHealth(cand.provider, cand, { ok: false, ttftMs: null, latencyMs: null, code: preCode, usage: lastUsage });
|
|
363
409
|
throw { code: preCode, message: (reason.failure && reason.failure.message) || 'candidate stream failed' };
|
|
364
410
|
}
|
|
365
411
|
else {
|
|
@@ -370,14 +416,14 @@ export function apply(ctx) {
|
|
|
370
416
|
if (isTerminalChunk(chunk)) {
|
|
371
417
|
const reason = chunk.reason || {};
|
|
372
418
|
if (isSuccessReason(reason)) {
|
|
373
|
-
recordHealth(cand.provider, cand, { ok: true, ttftMs, latencyMs: Date.now() - attemptStart, code: null });
|
|
419
|
+
recordHealth(cand.provider, cand, { ok: true, ttftMs, latencyMs: Date.now() - attemptStart, code: null, usage: lastUsage });
|
|
374
420
|
yield chunk;
|
|
375
421
|
return;
|
|
376
422
|
}
|
|
377
423
|
// 内容已向下游输出:绝不能再下发该失败终止块(否则出现 finish 后继续输出/双 finish),
|
|
378
424
|
// 改为抛错并标记 emitted,由组层直接终结本次请求
|
|
379
425
|
const midCode = (reason.failure && reason.failure.code) || 'STREAM_ERROR';
|
|
380
|
-
recordHealth(cand.provider, cand, { ok: false, ttftMs, latencyMs: null, code: midCode });
|
|
426
|
+
recordHealth(cand.provider, cand, { ok: false, ttftMs, latencyMs: null, code: midCode, usage: lastUsage });
|
|
381
427
|
throw { code: midCode, message: (reason.failure && reason.failure.message) || 'candidate stream failed mid-stream' };
|
|
382
428
|
}
|
|
383
429
|
yield chunk;
|
|
@@ -389,7 +435,7 @@ export function apply(ctx) {
|
|
|
389
435
|
if (started && err && typeof err === 'object')
|
|
390
436
|
err.emitted = true;
|
|
391
437
|
if (err && err.code === 'TIMEOUT')
|
|
392
|
-
recordHealth(cand.provider, cand, { ok: false, ttftMs, latencyMs: null, code: 'TIMEOUT' });
|
|
438
|
+
recordHealth(cand.provider, cand, { ok: false, ttftMs, latencyMs: null, code: 'TIMEOUT', usage: lastUsage });
|
|
393
439
|
throw err;
|
|
394
440
|
}
|
|
395
441
|
}
|
|
@@ -503,7 +549,7 @@ export function apply(ctx) {
|
|
|
503
549
|
// ---------- 测速 ----------
|
|
504
550
|
async function measureCandidate(cfg, cand, st) {
|
|
505
551
|
const llm = ctx.llm;
|
|
506
|
-
const inner = llm.stream({ provider: cand.provider, model: cand.model, messages: [{ role: 'user', content: [{ type: 'text', text: st.prompt }] }], maxTokens: st.maxTokens });
|
|
552
|
+
const inner = llm.stream({ provider: cand.provider, model: cand.model, messages: [{ role: 'user', content: [{ type: 'text', text: st.prompt }] }], maxTokens: st.maxTokens, sessionId: 'mcm-speedtest-' + (st.nonce ?? Date.now()) });
|
|
507
553
|
const start = Date.now();
|
|
508
554
|
let ttft = null;
|
|
509
555
|
const closeInner = async () => { try {
|
|
@@ -567,11 +613,15 @@ export function apply(ctx) {
|
|
|
567
613
|
await ctx.timeout(1500);
|
|
568
614
|
}
|
|
569
615
|
}
|
|
570
|
-
return { provider: cand.provider, model: cand.model, ok: false, ttft: null, latency: null, failure: (lastErr && lastErr.message) || 'speedtest failed' };
|
|
616
|
+
return { provider: cand.provider, model: cand.model, ok: false, ttft: null, latency: null, failure: (lastErr && lastErr.message) || 'speedtest failed', code: (lastErr && lastErr.code) || null };
|
|
571
617
|
}));
|
|
572
618
|
results.push(...measured);
|
|
573
619
|
}
|
|
574
620
|
const rows = results.map((r) => ({ provider: r.provider, model: r.model, ok: r.ok, ttft: r.ttft, latency: r.latency, at: Date.now(), failure: r.failure || null }));
|
|
621
|
+
// 整组候选都因「环境未就绪」(凭据服务尚未加载等)失败时,不落盘 speedResults、不冷却候选,
|
|
622
|
+
// 交给调用方延后重试——否则启动瞬间的一次重放会把所有渠道误判成故障并冷却。
|
|
623
|
+
if (rows.length > 0 && rows.every((r, i) => !r.ok && notReady({ code: results[i] && results[i].code, message: r.failure })))
|
|
624
|
+
return { ok: false, deferred: true };
|
|
575
625
|
state.speedResults[cfg.id] = rows;
|
|
576
626
|
await persistSpeedResults(state.speedResults);
|
|
577
627
|
const now = Date.now();
|
|
@@ -601,7 +651,10 @@ export function apply(ctx) {
|
|
|
601
651
|
context: { contextWindow: cfg.virtualModel.contextWindow },
|
|
602
652
|
defaultMaxTokens: cfg.virtualModel.maxTokens,
|
|
603
653
|
inputModalities: cfg.virtualModel.input.slice(),
|
|
604
|
-
|
|
654
|
+
// defaultEffort=max:原生 /model 弹窗对新模型自动填与列表展示都跟随该声明
|
|
655
|
+
// (此前 medium 导致新会话选组即 medium);会话内显式档位由 selector 插件的
|
|
656
|
+
// 档位记忆层(modelDirectories 拦截)恢复
|
|
657
|
+
reasoning: levels.length > 0 ? { efforts: levels.map((l) => ({ id: l, name: l })), defaultEffort: 'max' } : undefined,
|
|
605
658
|
};
|
|
606
659
|
}
|
|
607
660
|
const adapter = {
|
|
@@ -649,11 +702,12 @@ export function apply(ctx) {
|
|
|
649
702
|
// ---------- 单模型真实请求测试(client 经 settings 总线下发,走 DSH 真实 llm.stream 链路) ----------
|
|
650
703
|
async function runModelTest(provider, model, prompt, maxTokens) {
|
|
651
704
|
const llm = ctx.llm;
|
|
652
|
-
const inner = llm.stream({ provider, model, messages: [{ role: 'user', content: [{ type: 'text', text: prompt || '你好' }] }], maxTokens: maxTokens || 512 });
|
|
705
|
+
const inner = llm.stream({ provider, model, messages: [{ role: 'user', content: [{ type: 'text', text: prompt || '你好' }] }], maxTokens: maxTokens || 512, sessionId: 'mcm-modeltest-' + Date.now() });
|
|
653
706
|
const start = Date.now();
|
|
654
707
|
let ttft = null;
|
|
655
708
|
let text = '';
|
|
656
709
|
let reasoning = '';
|
|
710
|
+
let lastUsage = null;
|
|
657
711
|
const closeInner = async () => { try {
|
|
658
712
|
const c = inner.return ? inner.return() : null;
|
|
659
713
|
if (c && c.then)
|
|
@@ -676,6 +730,8 @@ export function apply(ctx) {
|
|
|
676
730
|
if (ttft === null && isContentChunk(chunk)) {
|
|
677
731
|
ttft = Date.now() - start;
|
|
678
732
|
}
|
|
733
|
+
if (chunk && chunk.type === 'usage' && chunk.usage && typeof chunk.usage === 'object')
|
|
734
|
+
lastUsage = chunk.usage;
|
|
679
735
|
if (chunk && chunk.type === 'text-delta' && typeof chunk.text === 'string') {
|
|
680
736
|
text += chunk.text;
|
|
681
737
|
}
|
|
@@ -684,8 +740,12 @@ export function apply(ctx) {
|
|
|
684
740
|
}
|
|
685
741
|
if (isTerminalChunk(chunk)) {
|
|
686
742
|
const reason = chunk.reason || {};
|
|
687
|
-
if (isSuccessReason(reason))
|
|
743
|
+
if (isSuccessReason(reason)) {
|
|
744
|
+
// 命中 finish 就 return,内层流不会被全局拦截器完整排干,成功这一笔要自己补记
|
|
745
|
+
// (失败那笔由拦截器的 catch 分支记,所以此前只有失败会进健康流水)
|
|
746
|
+
recordHealth(provider, { provider, model }, { ok: true, ttftMs: ttft, latencyMs: Date.now() - start, code: null, usage: lastUsage });
|
|
688
747
|
return { ok: true, ttftMs: ttft, latencyMs: Date.now() - start, text: text.slice(0, 2000), reasoning: reasoning.slice(0, 2000) };
|
|
748
|
+
}
|
|
689
749
|
throw { code: (reason.failure && reason.failure.code) || 'STREAM_ERROR', message: (reason.failure && reason.failure.message) || 'model test failed' };
|
|
690
750
|
}
|
|
691
751
|
}
|
|
@@ -695,38 +755,56 @@ export function apply(ctx) {
|
|
|
695
755
|
throw err;
|
|
696
756
|
}
|
|
697
757
|
}
|
|
698
|
-
function handleTestRequest(
|
|
758
|
+
function handleTestRequest(next) {
|
|
699
759
|
const req = next.testRequest;
|
|
700
|
-
const last = next.lastTestHandledNonce || 0;
|
|
760
|
+
const last = Math.max(next.lastTestHandledNonce || 0, claimedTestNonce);
|
|
701
761
|
if (!req || typeof req.provider !== 'string' || typeof req.model !== 'string' || typeof req.nonce !== 'number' || req.nonce === last)
|
|
702
762
|
return;
|
|
703
|
-
|
|
704
|
-
//
|
|
705
|
-
//
|
|
706
|
-
//
|
|
707
|
-
const
|
|
708
|
-
|
|
763
|
+
claimedTestNonce = req.nonce;
|
|
764
|
+
// nonce 的落盘推迟到本次测试得出结论之后:只有「真跑过」才算 handled。
|
|
765
|
+
// 启动早于凭据服务就绪时会以 MISSING_CREDENTIAL 失败,那要释放认领重试;
|
|
766
|
+
// 若提前写了 nonce,重试会被自己的持久值挡掉。
|
|
767
|
+
const settle = (entry) => {
|
|
768
|
+
if (bus !== null)
|
|
769
|
+
bus.writeHealth({ lastTestHandledNonce: req.nonce }).catch((e) => console.error('[model-channel-manager] lastTestHandledNonce 写入失败:', e));
|
|
770
|
+
return setResult(entry);
|
|
771
|
+
};
|
|
772
|
+
// 0.2 的配置写入是「整字段落盘」,不再有 path-ops;结果集按当前值合并后整段写回。
|
|
773
|
+
const setResult = async (entry) => {
|
|
774
|
+
const key = String(req.nonce);
|
|
775
|
+
const value = Object.assign({}, entry, { nonce: req.nonce, provider: req.provider, model: req.model });
|
|
776
|
+
const cur = healthOf().testResults || {};
|
|
777
|
+
if (bus !== null)
|
|
778
|
+
await bus.writeHealth({ testResults: Object.assign({}, cur, { [key]: value }) }).catch((e) => console.error('[model-channel-manager] testResults 写入失败:', e));
|
|
709
779
|
};
|
|
710
780
|
// 修剪低频执行:只在条目数超限时砍到 50(不在每次写入时整包重写)
|
|
711
781
|
const maybePrune = () => {
|
|
712
|
-
const all = (
|
|
782
|
+
const all = healthOf().testResults || {};
|
|
713
783
|
const entries = Object.entries(all);
|
|
714
784
|
if (entries.length <= 50) return;
|
|
715
|
-
const
|
|
716
|
-
|
|
717
|
-
|
|
785
|
+
const kept = {};
|
|
786
|
+
for (const [k, v] of entries.sort((a, b) => ((b[1] && b[1].finishedAt) || 0) - ((a[1] && a[1].finishedAt) || 0)).slice(0, 50)) kept[k] = v;
|
|
787
|
+
if (bus !== null)
|
|
788
|
+
bus.writeHealth({ testResults: kept }).catch((e) => console.error('[model-channel-manager] testResults 修剪失败:', e));
|
|
718
789
|
};
|
|
719
790
|
const done = async (r) => {
|
|
720
791
|
try {
|
|
721
|
-
|
|
792
|
+
await settle(Object.assign({}, r, { finishedAt: Date.now() }));
|
|
722
793
|
maybePrune();
|
|
723
794
|
}
|
|
724
795
|
catch (_e) { }
|
|
725
796
|
};
|
|
726
|
-
setResult({ status: 'running', startedAt: Date.now() });
|
|
797
|
+
void setResult({ status: 'running', startedAt: Date.now() });
|
|
727
798
|
runModelTest(req.provider, req.model, req.prompt, req.maxTokens).then(async (r) => {
|
|
799
|
+
replayAttempts.delete('test');
|
|
728
800
|
await done(Object.assign({ status: 'ok' }, r));
|
|
729
801
|
}).catch(async (e) => {
|
|
802
|
+
if (notReady(e)) {
|
|
803
|
+
// 凭据服务尚未就绪(典型是启动瞬间的重放):不写成假 error,释放认领并延后重试
|
|
804
|
+
claimedTestNonce = 0;
|
|
805
|
+
scheduleReplay('test', () => reloadFromConfig(false));
|
|
806
|
+
return;
|
|
807
|
+
}
|
|
730
808
|
await done({ status: 'error', code: (e && e.code) || 'STREAM_ERROR', error: (e && e.message) || String(e) });
|
|
731
809
|
});
|
|
732
810
|
}
|
|
@@ -747,44 +825,126 @@ export function apply(ctx) {
|
|
|
747
825
|
}
|
|
748
826
|
// ---------- settings 总线接入(响应式:settings 服务异步初始化,apply 时查询太早) ----------
|
|
749
827
|
let bus = null;
|
|
828
|
+
let runtimeRestored = false;
|
|
829
|
+
// 非持久去重:回写要等当前 HMR 事务结束才落盘,这中间 volatile-update 可能带着同一个 nonce 再来,
|
|
830
|
+
// 内存里先认领,避免同一请求被重复执行(并顺带消掉重复的回写风暴)。
|
|
831
|
+
let claimedTestNonce = 0;
|
|
832
|
+
let claimedSpeedNonce = 0;
|
|
833
|
+
// 启动早于凭据服务就绪:此时重放测试/测速会以 MISSING_CREDENTIAL 失败。
|
|
834
|
+
// 这类错误是「环境还没准备好」而不是「渠道故障」,必须释放认领、延后重试,不能写成假 error。
|
|
835
|
+
const notReady = (e) => {
|
|
836
|
+
const code = (e && e.code) || '';
|
|
837
|
+
const msg = String((e && (e.message || e.error)) || '');
|
|
838
|
+
return code === 'MISSING_CREDENTIAL' || code === 'INVALID_CREDENTIAL' || /no credential|credential .*not set|is not set — store/i.test(msg);
|
|
839
|
+
};
|
|
840
|
+
const replayAttempts = new Map();
|
|
841
|
+
const scheduleReplay = (key, fn, ms = 5000) => {
|
|
842
|
+
const n = (replayAttempts.get(key) || 0) + 1;
|
|
843
|
+
if (n > 24)
|
|
844
|
+
return;
|
|
845
|
+
replayAttempts.set(key, n);
|
|
846
|
+
ctx.timeout(() => { fn(); }, ms);
|
|
847
|
+
};
|
|
848
|
+
// 从实例配置读取全部状态;volatile 提交后由 loader/volatile-update 触发重载。
|
|
849
|
+
function reloadFromConfig(initial) {
|
|
850
|
+
const cfg = cfgOf();
|
|
851
|
+
state.config = clone(cfg) || { groups: [] };
|
|
852
|
+
const health = healthOf();
|
|
853
|
+
state.records = clone(health.records || {});
|
|
854
|
+
// 配置快照覆盖后,把尚未落盘的流水补回来(ts+provider+model 相同视为同一笔,不重复)
|
|
855
|
+
for (const p of pendingRecords) {
|
|
856
|
+
const list = state.records[p.key] || (state.records[p.key] = []);
|
|
857
|
+
if (!list.some((e) => e.ts === p.rec.ts && e.provider === p.rec.provider && e.model === p.rec.model))
|
|
858
|
+
list.push(p.rec);
|
|
859
|
+
}
|
|
860
|
+
state.speedResults = clone(health.speedResults || {});
|
|
861
|
+
state.testResults = clone(health.testResults || {});
|
|
862
|
+
if (initial && !runtimeRestored) {
|
|
863
|
+
restoreRuntime(health.runtime || {});
|
|
864
|
+
runtimeRestored = true;
|
|
865
|
+
}
|
|
866
|
+
const live = pullConfig().groups;
|
|
867
|
+
for (const g of live)
|
|
868
|
+
groupRuntime(g.id);
|
|
869
|
+
// 清理已删除组的运行时残留,避免 runtime 持久化无限累积
|
|
870
|
+
for (const id of Array.from(runtime.keys()))
|
|
871
|
+
if (!live.some((g) => g.id === id))
|
|
872
|
+
runtime.delete(id);
|
|
873
|
+
rewireRoutes();
|
|
874
|
+
const req = health.speedRequest;
|
|
875
|
+
const last = Math.max(health.lastHandledNonce || 0, claimedSpeedNonce);
|
|
876
|
+
if (req && typeof req.group === 'string' && typeof req.nonce === 'number' && req.nonce !== last) {
|
|
877
|
+
claimedSpeedNonce = req.nonce;
|
|
878
|
+
const cfgRow = pullConfig().groups.find((g) => g.id === req.group);
|
|
879
|
+
// nonce 落盘推迟到本次测速得出结论之后(同 handleTestRequest 的理由):
|
|
880
|
+
// 提前写会把「没就绪」的重试用自己的持久值挡掉。
|
|
881
|
+
const settleSpeed = () => {
|
|
882
|
+
if (bus !== null)
|
|
883
|
+
bus.writeHealth({ lastHandledNonce: req.nonce }).catch(() => { });
|
|
884
|
+
};
|
|
885
|
+
if (cfgRow)
|
|
886
|
+
runSpeedTest(cfgRow).then((r) => {
|
|
887
|
+
if (r && r.deferred) {
|
|
888
|
+
claimedSpeedNonce = 0;
|
|
889
|
+
scheduleReplay('speed', () => reloadFromConfig(false));
|
|
890
|
+
return;
|
|
891
|
+
}
|
|
892
|
+
replayAttempts.delete('speed');
|
|
893
|
+
settleSpeed();
|
|
894
|
+
}).catch((e) => {
|
|
895
|
+
if (notReady(e)) {
|
|
896
|
+
claimedSpeedNonce = 0;
|
|
897
|
+
scheduleReplay('speed', () => reloadFromConfig(false));
|
|
898
|
+
return;
|
|
899
|
+
}
|
|
900
|
+
settleSpeed();
|
|
901
|
+
console.error('[model-channel-manager] speedtest error:', e);
|
|
902
|
+
});
|
|
903
|
+
else
|
|
904
|
+
settleSpeed();
|
|
905
|
+
}
|
|
906
|
+
handleTestRequest(health);
|
|
907
|
+
}
|
|
750
908
|
ctx.inject(['settings'], (sctx) => {
|
|
751
909
|
const settings = sctx.settings;
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
755
|
-
|
|
756
|
-
|
|
757
|
-
|
|
758
|
-
|
|
759
|
-
|
|
760
|
-
|
|
761
|
-
|
|
762
|
-
|
|
763
|
-
|
|
764
|
-
|
|
765
|
-
|
|
766
|
-
|
|
767
|
-
|
|
768
|
-
|
|
769
|
-
|
|
770
|
-
|
|
771
|
-
|
|
910
|
+
// 0.2 只暴露「按行 id 定位的实例配置」;健康字段是同一 Config 里的 health 子树,
|
|
911
|
+
// 写入为整字段合并写(旧的 path-ops mutate 在 0.2 不存在)。
|
|
912
|
+
// 0.2 的配置写入走 configEditor.edit() → hmr.runExclusive(),而 runExclusive 一旦发现已在
|
|
913
|
+
// 事务内就直接拒绝("HMR transactions cannot be nested")。loader/volatile-update 回调本身
|
|
914
|
+
// 就运行在该事务里,所以回写必须先切出事务上下文:事务内创建的异步资源(AsyncResource、
|
|
915
|
+
// setTimeout)都会继承事务上下文,只有 AsyncLocalStorage.exit() 能干净地切出去,
|
|
916
|
+
// 写入再由 runExclusive 排进队列、在本次事务结束后执行。
|
|
917
|
+
const outsideTransaction = (fn) => {
|
|
918
|
+
const hmr = ctx.get('hmr');
|
|
919
|
+
const als = hmr && hmr.executing;
|
|
920
|
+
if (als && typeof als.exit === 'function' && als.getStore()) return als.exit(fn);
|
|
921
|
+
return fn();
|
|
922
|
+
};
|
|
923
|
+
// 0.2 是「整字段落盘」,并发调用会各自基于同一份旧快照做 read-modify-write,
|
|
924
|
+
// 后写的直接覆盖先写的(实测 lastTestHandledNonce 与 records 就这样丢过)。
|
|
925
|
+
// 所以写入串行化:排队后逐个重新读当前值再合并,不让两次写入互相覆写。
|
|
926
|
+
let healthWrites = Promise.resolve();
|
|
927
|
+
const writeHealth = (patch) => {
|
|
928
|
+
const run = async () => {
|
|
929
|
+
const cur = healthOf();
|
|
930
|
+
const next = Object.assign({}, cur, patch);
|
|
931
|
+
return outsideTransaction(() => settings.update(SELF_NS, { health: next }));
|
|
932
|
+
};
|
|
933
|
+
const queued = healthWrites.then(run, run);
|
|
934
|
+
healthWrites = queued.catch(() => { });
|
|
935
|
+
return queued;
|
|
936
|
+
};
|
|
937
|
+
bus = {
|
|
938
|
+
settings,
|
|
939
|
+
writeHealth,
|
|
940
|
+
cfgScope: { get: () => cfgOf() },
|
|
941
|
+
healthScope: { get: () => healthOf() },
|
|
942
|
+
};
|
|
943
|
+
reloadFromConfig(true);
|
|
944
|
+
ctx.on('loader/volatile-update', () => {
|
|
945
|
+
reloadFromConfig(false);
|
|
772
946
|
console.log('[model-channel-manager] config hot-reloaded, routes:', pullConfig().groups.map((g) => g.id).join(', ') || '(none)');
|
|
773
947
|
});
|
|
774
|
-
healthScope.watch((next) => {
|
|
775
|
-
state.records = clone(next.records || {});
|
|
776
|
-
state.speedResults = clone(next.speedResults || {});
|
|
777
|
-
state.testResults = clone(next.testResults || {});
|
|
778
|
-
const req = next.speedRequest;
|
|
779
|
-
const last = next.lastHandledNonce || 0;
|
|
780
|
-
if (req && typeof req.group === 'string' && typeof req.nonce === 'number' && req.nonce !== last) {
|
|
781
|
-
settings.update(NS_HEALTH, { lastHandledNonce: req.nonce }).catch(() => { });
|
|
782
|
-
const cfg = pullConfig().groups.find((g) => g.id === req.group);
|
|
783
|
-
if (cfg)
|
|
784
|
-
runSpeedTest(cfg).catch((e) => console.error('[model-channel-manager] speedtest error:', e));
|
|
785
|
-
}
|
|
786
|
-
handleTestRequest(settings, next);
|
|
787
|
-
});
|
|
788
948
|
boot();
|
|
789
949
|
});
|
|
790
950
|
// ---------- 全局 LLM 请求健康拦截 (涵盖所有非虚拟路由的真实渠道模型调用) ----------
|
|
@@ -800,6 +960,7 @@ export function apply(ctx) {
|
|
|
800
960
|
let ttft = null;
|
|
801
961
|
let lastError = null;
|
|
802
962
|
let isSuccess = false;
|
|
963
|
+
let usage = null;
|
|
803
964
|
// 用户主动中止不是渠道故障:AbortError / code ABORTED / signal 已 aborted
|
|
804
965
|
// 三种形态都不进健康流水,否则污染成功率与 smart 键的 reliability 权重
|
|
805
966
|
const isAbortLike = (err) => {
|
|
@@ -813,6 +974,9 @@ export function apply(ctx) {
|
|
|
813
974
|
};
|
|
814
975
|
try {
|
|
815
976
|
for await (const chunk of next()) {
|
|
977
|
+
// 适配器在 finish 前发 usage 块(done/error 都带);捕获后随记录写入
|
|
978
|
+
if (chunk && chunk.type === 'usage' && chunk.usage && typeof chunk.usage === 'object')
|
|
979
|
+
usage = chunk.usage;
|
|
816
980
|
if (ttft === null && isContentChunk(chunk)) {
|
|
817
981
|
ttft = Date.now() - startTs;
|
|
818
982
|
}
|
|
@@ -835,6 +999,7 @@ export function apply(ctx) {
|
|
|
835
999
|
ttftMs: isSuccess ? ttft : null,
|
|
836
1000
|
latencyMs: isSuccess ? latency : null,
|
|
837
1001
|
code: isSuccess ? null : (lastError || 'UNKNOWN_TERMINAL'),
|
|
1002
|
+
usage,
|
|
838
1003
|
});
|
|
839
1004
|
}
|
|
840
1005
|
} catch (err) {
|
|
@@ -843,6 +1008,7 @@ export function apply(ctx) {
|
|
|
843
1008
|
recordHealth(p, { provider: p, model: options.model }, {
|
|
844
1009
|
ok: false,
|
|
845
1010
|
code: (err && err.code) || (err && err.message) || 'EXCEPTION',
|
|
1011
|
+
usage,
|
|
846
1012
|
});
|
|
847
1013
|
}
|
|
848
1014
|
throw err;
|
|
@@ -863,14 +1029,13 @@ export function apply(ctx) {
|
|
|
863
1029
|
}
|
|
864
1030
|
console.log('[model-channel-manager] booted, groups:', pullConfig().groups.map((g) => g.id).join(', ') || '(none)');
|
|
865
1031
|
}
|
|
866
|
-
// 工作区 .channel-manager/config.json ->
|
|
1032
|
+
// 工作区 .channel-manager/config.json -> 本行实例配置(一次性)。
|
|
867
1033
|
// 完成后写 legacyMigrated 哨兵,防止「用户清空全部组 → 重启 → 旧配置复活」;
|
|
868
1034
|
// fs/sandboxPolicy 未就绪时短暂重试,而不是静默放弃直到下次重启。
|
|
869
1035
|
async function migrateLegacyConfig() {
|
|
870
|
-
|
|
871
|
-
if (settings === undefined || pullConfig().groups.length !== 0)
|
|
1036
|
+
if (bus === null || pullConfig().groups.length !== 0)
|
|
872
1037
|
return;
|
|
873
|
-
if ((
|
|
1038
|
+
if (healthOf().legacyMigrated === true)
|
|
874
1039
|
return;
|
|
875
1040
|
for (let attempt = 0; attempt < 6; attempt++) {
|
|
876
1041
|
const fsSvc = ctx.get('fs');
|
|
@@ -883,13 +1048,13 @@ export function apply(ctx) {
|
|
|
883
1048
|
const legacy = JSON.parse(text);
|
|
884
1049
|
const migrated = normalizeConfig(legacy);
|
|
885
1050
|
if (migrated.groups.length > 0) {
|
|
886
|
-
await settings.
|
|
1051
|
+
await bus.settings.update(SELF_NS, { groups: migrated.groups });
|
|
887
1052
|
state.config = { groups: migrated.groups };
|
|
888
1053
|
console.log('[model-channel-manager] migrated legacy config from workspace .channel-manager/config.json');
|
|
889
1054
|
}
|
|
890
1055
|
}
|
|
891
1056
|
catch (_e) { /* 无遗留配置,忽略 */ }
|
|
892
|
-
|
|
1057
|
+
bus.writeHealth({ legacyMigrated: true }).catch(() => { });
|
|
893
1058
|
return;
|
|
894
1059
|
}
|
|
895
1060
|
try {
|