@arcaneorion/dsh-model-channel-manager 0.3.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -58,7 +58,7 @@ dsh plugin --profile web add @arcaneorion/dsh-model-channel-manager
58
58
 
59
59
  ## 数据通道(全走公共 seam,无私有 RPC)
60
60
 
61
- - 读配置/健康/运行态 = `api.settings.describe()` 过滤命名空间
61
+ - 读配置/运行态 = `api.settings.describe()` 过滤命名空间
62
62
  - 保存 provider = `api.settings.update({ns:'llm-pi-ai', patch:{providers}})`
63
63
  - 保存轮询组 = `api.settings.update({ns:'model-channels', patch:{groups}})`
64
64
  - ⚡测速 = `api.settings.update({ns:'model-channel-health', patch:{speedRequest:{group,nonce}}})`(host watcher 消费)
@@ -67,7 +67,28 @@ dsh plugin --profile web add @arcaneorion/dsh-model-channel-manager
67
67
  - 上述调用形状仍保留 0.1 的样子:client 半内建门面 `makeLegacyApi` 把 0.2 的**位置参数 + RemoteResult** 适配回旧的**对象入参 + `{result:{ok,value}}`**,并把 `model-channels` / `model-channel-health` 合成回旧命名空间视图(真实承载是本插件行 id `model-channel-manager` 的实例配置)。
68
68
 
69
69
  > **宿主边界(0.1 历史,0.2 已不适用)**:0.1 的 settings RPC 走 apiproxy 暴露白名单(`exposedNamespaces()` = LLM provider ns + `WEB_/PRODUCT_SETTINGS_NAMESPACES`),当时含该边界的宿主必须放行 `model-channels` / `model-channel-health`(本仓曾在 harness `dsh-host-apiproxy` 打 `PLUGIN_SETTINGS_NAMESPACES` 补丁)。
70
- > **0.2 的 settings 命名空间就是 profile 行 id**,由 `@deepseek-ai/dsh-api-settings-controller` 的 `describe` 直接投影本行实例配置,没有该白名单环节;对应地,本插件的数据落在 `~/.dsh/profiles/web/cordis.patch.yml` 的 `model-channel-manager` 行 `config` 下。
70
+ > **0.2 的 settings 命名空间就是 profile 行 id**,由 `@deepseek-ai/dsh-api-settings-controller` 的 `describe` 直接投影本行实例配置,没有该白名单环节;对应地,本插件的配置落在 `~/.dsh/profiles/web/cordis.patch.yml` 的 `model-channel-manager` 行 `config` 下。
71
+
72
+ ## 健康数据存储(0.3.1 重构:事实数据归位 storageDomain)
73
+
74
+ > 背景:0.3.0 及之前,健康流水整字段存在 settings health 子树里,每 2s 防抖整段重写
75
+ > profile patch(实测 6350 行中 health 约占 3000 行),且与 volatile 快照覆盖互相踩——
76
+ > 刚记的账在落盘前被旧快照抹掉(审计 F11/F23,「测试成功不入账」的根因)。
77
+
78
+ - **权威存储**:`storageDomain` 的 `model_channel_health` 单元(dsh-base 已组合 json 后端,
79
+ root=`~/.dsh/storages/`),`per-record` 布局——一条渠道一个桶文档,`backup-and-skip`
80
+ 容错。host 侧 `src/health-store.js` 封装:追加走原子写链(并发 `recordHealth` 不丢更新)、
81
+ 7 天窗口过期、每桶 300 条截断。
82
+ - **settings 只存小投影**:`health.digest`(host 聚合好的摘要数组:total/success/ttft/latency/
83
+ token 三分项/lastTs)+ `digestAt`,client 健康页渲染用;不再下发原始流水。
84
+ - **聚合上移 host**:`buildDigest` 在 host 折叠(口径同旧 client:计费 token = input +
85
+ cacheRead + cacheWrite + output),client 不再拉全量 describe 做原始事件折叠。
86
+ - **存量迁移**:首次启动自动把 settings 里的 `records`/`speedResults` 搬入 domain(桶已存在
87
+ 即跳过,幂等),同一次合并写里清空 settings 旧存量并落 `healthMigrated` 标记。
88
+ - **降级**:storageDomain 缺席的 profile 退化为纯内存(不持久化流水),不拒绝启动。
89
+ - **client 兼容**:旧 host(无 digest 字段)自动回落原始 records 路径,升级窗口不断供。
90
+ - **已知近似**:30m/24h 视图按「最近活跃渠道」过滤,数值仍是 7 天累计(UI 已标注);
91
+ 精确分窗口需 host 出多份 digest,后续增强。
71
92
 
72
93
  ## 响应信封(重要)
73
94
 
@@ -129,7 +150,7 @@ dsh plugin --profile web add @arcaneorion/dsh-model-channel-manager
129
150
  ## host 半内部接口
130
151
 
131
152
  - settings 接入用 **`ctx.inject(['settings'], (sctx) => {...})`**(settings 服务异步初始化,apply 时 `ctx.get('settings')` 为 undefined——曾经整个引擎静默失效,命名空间从未注册)
132
- - 配置 schema(schemastery):`model-channels` 的虚模型/candidates/strategy/timeoutMs/cooldownMs/maxRetriesPerCandidate/speedTest;`model-channel-health` 的 records 7 天切片(单组 ≤2000 条)/speedResults/runtime/speedRequest+lastHandledNonce/testRequest+testResults+lastTestHandledNonce
153
+ - 配置 schema(schemastery):`model-channels` 的虚模型/candidates/strategy/timeoutMs/cooldownMs/maxRetriesPerCandidate/speedTest;`model-channel-health` 的 runtime/speedRequest+lastHandledNonce/testRequest+testResults+lastTestHandledNonce/digest 小投影(records/speedResults 仅作 0.3.1 迁移的读取源,权威在 storageDomain)
133
154
  - 引擎:sticky/round-robin/primary 三策略;首响应超时 + 流中空闲超时(动态 = max(timeoutMs, min(120s, ttft×2)));单候选原地重试(指数退避)耗尽才换;全炸清冷却重试一轮;测速 ttft/latency/hybrid/smart 四键(smart = 0.5×ttft_norm + 0.3×(1−reliability) + 0.2×latency_norm,reliability 贝叶斯平滑 `(success+2.5)/(total+5)`);测速失败进冷却;请求隔离按组
134
155
  - 虚拟模型元数据:`reasoning.efforts` 七档(off…max)、**defaultEffort=max**——原生 `/model` 弹窗对新模型的自动填档与展示跟随该声明;会话内显式档位的跨会话恢复由 selector 插件的档位记忆层负责(`modelDirectories` 拦截,存 `model-channels.effortMemory`)
135
156
  - 迁移:startup 时从工作区 `.channel-manager/config.json` 一次性迁入 `model-channels`(无遗留则忽略);完成后写 `legacyMigrated` 哨兵防止「清空组后重启复活」;fs 未就绪时 5s×6 重试
@@ -176,5 +197,6 @@ react 经 `require('react')`;样式用 `ctx.effect` 自管理;`dsh.client: {
176
197
  20. **诊断临时实例必须独立 home(`DSH_HOME=/tmp/dsh-diag dsh ...`)**:临时实例与主实例共用 `~/.dsh` 会并发写同一会话日志与 `session_projcache.json`——两进程各自的 seq 计数器交错追加,日志出现重复 seq → `corrupt session log: seq gap in committed region` → 会话 resume 直接拒绝,表现为该会话内模型目录加载失败(选择器「暂无可用模型」)。修复:解压 jsonl 删掉多余事件即可(后续 seq 连续则天然对齐),用 `session-persistence-jsonl` 的 `scanLog` 校验后压缩回写;杀进程前务必备份。
177
198
  21. **适配器契约以安装运行时的 d.ts 为准,不能照抄源码仓快照**:源码仓较新、rc.2 运行时的 `LlmAdapter` 多一个必需的 `prepareCall(provider, model, signal) → Promise<{model, stream}>`(主分发路径 llm.stream/llm.prepareCall 都先走它再 `adapterCall.stream(options)`;`adapter.stream` 在 rc.2 服务层从不直调)。缺它的症状极具迷惑性:注册/目录/菜单全正常,**真实发对话**才报 `registration.adapter.prepareCall is not a function`。实现对齐 llm-pi-ai 的快照模式:prepare 时捕获一份配置快照,元数据与 dispatch 都出自同一代。回归:`tests/adapter-contract.test.cjs`(T3 直接解析安装版 d.ts 的 LlmAdapter 方法集做契约同步)。
178
199
  22. **0.2 配置写入是 HMR 独占事务**:`settings.update` → `configEditor.edit()` → `hmr.runExclusive()`;在 `loader/volatile-update` 回调里回写会抛 `HMR transactions cannot be nested`(实测一段会话内 15 次,面板“测试”结果永远落不了盘)。事务内创建的**任何**异步资源(`AsyncResource` / `setTimeout` / `setInterval`)都继承事务上下文,**只有 `AsyncLocalStorage.exit()` 能切出**:`ctx.get('hmr').executing.exit(fn)`(仅当 `getStore()` 为真时切)。写入会被 `runExclusive` 排进队列、在本次事务结束后执行;**监听器保持同步、不要在外层事务里 await 它**(队列串行,互等即死锁)。
179
- 23. **整字段落盘 + 内存态被配置快照覆盖**:`settings.update` 是整字段替换;`reloadFromConfig()` 每次 volatile-update 都用配置快照整体覆盖 `state.records`,而健康 flush 有 2s 防抖 → **刚记下的一笔在落盘前就被内存覆盖**(症状:面板“测试”成功不入账,失败反被全局拦截器的 catch 记上)。修法:`pendingRecords` 缓冲,快照覆盖后把未落盘的补回,flush 成功后再清账。
200
+ 23. **整字段落盘 + 内存态被配置快照覆盖**(0.3.1 已根治):`settings.update` 是整字段替换;旧实现 `reloadFromConfig()` 每次 volatile-update 都用配置快照整体覆盖 `state.records`,而健康 flush 有 2s 防抖 → **刚记下的一笔在落盘前就被内存覆盖**(症状:面板“测试”成功不入账,失败反被全局拦截器的 catch 记上)。当时的修法是 `pendingRecords` 缓冲补账;0.3.1 起权威数据搬入 storageDomain(见「健康数据存储」),settings 只存 digest 小投影,此竞态从数据模型层消除。
180
201
  24. **nonce 落盘时机与启动竞态**:`lastTestHandledNonce` / `lastHandledNonce` 必须在**得出结果之后**写(提前写会让“未就绪”的重试被自己的持久值挡掉);启动瞬间凭据服务尚未就绪时测试/测速会以 `MISSING_CREDENTIAL` 失败(凭据其实已在 `.credentials.yaml` 里),应识别为「还没就绪」→ 释放认领 + 5s 延时重试(上限 24 次),**不要**写成渠道故障;测速还必须在整组候选都因未就绪失败时**不落盘、不冷却**,否则一次启动重放就把所有渠道误判成故障。
202
+ 25. **domain 写入的并发丢失(0.3.1 review 挽救)**:`KvTable.put` 是整 record 覆盖,`get→filter→put` 的读-改-写在 put 的 IO 延迟窗口内并发调用会互相覆盖(后写盖先写,先记的账丢失)——恰好复刻了要消灭的丢账问题。**并发追加必须走原子链**:`update(key, fn)` 的 fn 在写链队列槽位看到当前值;桶不存在时 update 报 `missing-key`,先 put 初始化。另外在 promise 链里 `ctx.effect` 注册 disposer 前必须先验 fiber 活性——对 inactive fiber 注册会抛 `INACTIVE_EFFECT`,若被外层 catch 吞掉则 domain 永不 close,facility 名字被占 → HMR 重载后 `already-open` 静默降级。回归:`tests/health-domain-sync.test.cjs`。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@arcaneorion/dsh-model-channel-manager",
3
- "version": "0.3.0",
3
+ "version": "0.3.1",
4
4
  "type": "module",
5
5
  "main": "src/index.js",
6
6
  "exports": {
@@ -40,8 +40,10 @@
40
40
  "@deepseek-ai/dsh-client-connection": "0.2.0-rc.1",
41
41
  "@deepseek-ai/dsh-client-ui-conversation": "0.2.0-rc.1",
42
42
  "@deepseek-ai/dsh-api-remotes": "0.2.0-rc.1",
43
+ "@deepseek-ai/dsh-storage-domain": "0.2.0-rc.1",
43
44
  "@deepseek-ai/schemastery": ">=3.18.4",
44
- "react": "^18.3.1"
45
+ "react": "^18.3.1",
46
+ "zod": "^4.4.3"
45
47
  },
46
48
  "dsh": {
47
49
  "bundle": {
package/src/client.js CHANGED
@@ -1065,11 +1065,13 @@ window.__ModuleLoader__.load({
1065
1065
 
1066
1066
  const now = Date.now()
1067
1067
  const windowCutoff = windowMode === '30m' ? now - 30 * 60 * 1000 : windowMode === '24h' ? now - 24 * 3600 * 1000 : 0
1068
- const recsMap = health.records || {}
1069
- const rawEvents = Object.values(recsMap).flat()
1070
- // 时间窗口过滤
1071
- const allEvents = rawEvents.filter((e) => (e.ts || 0) >= windowCutoff)
1072
1068
 
1069
+ // ---------- 聚合数据源 ----------
1070
+ // 0.3.1 起优先用 host 下发的 digest 小投影(聚合已上移 host,client 不再拉原始流水)。
1071
+ // digest 是 7 天全窗口聚合;30m/24h 视图按 lastTs 近似(lastTs 在窗口内才计入),
1072
+ // 精细窗口统计属后续增强(需要 host 按窗口出多份 digest)。
1073
+ // 旧 host 兼容:无 digest 字段时走原始 records 路径(升级窗口不断供)。
1074
+ const digestRows = Array.isArray(health.digest) ? health.digest : null
1073
1075
  // 按 provider 分组
1074
1076
  const byProvider = new Map()
1075
1077
 
@@ -1092,7 +1094,10 @@ window.__ModuleLoader__.load({
1092
1094
  lastOk: null,
1093
1095
  lastCode: null,
1094
1096
  recentErrors: 0,
1095
- tokSum: 0
1097
+ tokSum: 0,
1098
+ tokIn: 0,
1099
+ tokOut: 0,
1100
+ tokCache: 0
1096
1101
  })
1097
1102
  }
1098
1103
  }
@@ -1103,57 +1108,134 @@ window.__ModuleLoader__.load({
1103
1108
  let totalTokOut = 0
1104
1109
  let totalTokCache = 0
1105
1110
 
1106
- // 累加时间窗口内的真实流水记录
1107
- for (const e of allEvents) {
1108
- if (!e || !e.provider || !e.model) continue
1109
- let pMap = byProvider.get(e.provider)
1110
- if (!pMap) { pMap = new Map(); byProvider.set(e.provider, pMap) }
1111
- let a = pMap.get(e.model)
1112
- if (!a) {
1113
- a = {
1114
- provider: e.provider,
1115
- model: e.model,
1116
- name: e.model,
1117
- total: 0,
1118
- success: 0,
1119
- fail: 0,
1120
- ttftSum: 0,
1121
- latSum: 0,
1122
- lastTs: 0,
1123
- lastOk: null,
1124
- lastCode: null,
1125
- recentErrors: 0,
1126
- tokSum: 0
1111
+ if (digestRows !== null) {
1112
+ // ---------- digest 投影路径 ----------
1113
+ for (const row of digestRows) {
1114
+ if (!row || !row.provider || !row.model) continue
1115
+ // 窗口过滤:lastTs 不在窗口内的渠道不计入该窗口(近似口径,见上注)
1116
+ if (windowCutoff > 0 && (row.lastTs || 0) < windowCutoff) continue
1117
+ let pMap = byProvider.get(row.provider)
1118
+ if (!pMap) { pMap = new Map(); byProvider.set(row.provider, pMap) }
1119
+ let a = pMap.get(row.model)
1120
+ if (!a) {
1121
+ a = {
1122
+ provider: row.provider,
1123
+ model: row.model,
1124
+ name: row.model,
1125
+ total: 0,
1126
+ success: 0,
1127
+ fail: 0,
1128
+ ttftSum: 0,
1129
+ latSum: 0,
1130
+ lastTs: 0,
1131
+ lastOk: null,
1132
+ lastCode: null,
1133
+ recentErrors: 0,
1134
+ tokSum: 0,
1135
+ tokIn: 0,
1136
+ tokOut: 0,
1137
+ tokCache: 0
1138
+ }
1139
+ pMap.set(row.model, a)
1127
1140
  }
1128
- pMap.set(e.model, a)
1129
- }
1130
- a.total++
1131
- const tokIn = (e && e.inputTokens) || 0
1132
- const tokOut = (e && e.outputTokens) || 0
1133
- const tokCache = ((e && e.cacheReadTokens) || 0) + ((e && e.cacheWriteTokens) || 0)
1134
- totalTokIn += tokIn
1135
- totalTokOut += tokOut
1136
- totalTokCache += tokCache
1137
- a.tokSum += tokIn + tokOut + tokCache
1138
- if (e.ok) {
1139
- a.success++
1140
- if (e.ttftMs != null && e.ttftMs >= 0) a.ttftSum += e.ttftMs
1141
- if (e.latencyMs != null && e.latencyMs >= 0) a.latSum += e.latencyMs
1142
- } else {
1143
- a.fail++
1144
- if (e.code) a.lastCode = e.code
1141
+ a.total = row.total || 0
1142
+ a.success = row.success || 0
1143
+ a.fail = a.total - a.success
1144
+ a.ttftSum = (row.ttftAvg || 0) * a.success
1145
+ a.latSum = (row.latAvg || 0) * a.success
1146
+ a.lastTs = row.lastTs || 0
1147
+ a.lastOk = row.lastOk
1148
+ a.lastCode = row.lastCode || null
1149
+ a.tokIn = row.tokIn || 0
1150
+ a.tokOut = row.tokOut || 0
1151
+ a.tokCache = row.tokCache || 0
1152
+ a.tokSum = a.tokIn + a.tokOut + a.tokCache
1153
+ totalTokIn += a.tokIn
1154
+ totalTokOut += a.tokOut
1155
+ totalTokCache += a.tokCache
1145
1156
  }
1146
- if ((e.ts || 0) > a.lastTs) {
1147
- a.lastTs = e.ts || 0
1148
- a.lastOk = e.ok
1157
+ } else {
1158
+ // ---------- 旧 host 原始流水路径(兼容) ----------
1159
+ const recsMap = health.records || {}
1160
+ const rawEvents = Object.values(recsMap).flat()
1161
+ // 时间窗口过滤(赋给函数级 allEvents,供顶部指标卡使用)
1162
+ allEvents = rawEvents.filter((e) => (e.ts || 0) >= windowCutoff)
1163
+ // 累加时间窗口内的真实流水记录
1164
+ for (const e of allEvents) {
1165
+ if (!e || !e.provider || !e.model) continue
1166
+ let pMap = byProvider.get(e.provider)
1167
+ if (!pMap) { pMap = new Map(); byProvider.set(e.provider, pMap) }
1168
+ let a = pMap.get(e.model)
1169
+ if (!a) {
1170
+ a = {
1171
+ provider: e.provider,
1172
+ model: e.model,
1173
+ name: e.model,
1174
+ total: 0,
1175
+ success: 0,
1176
+ fail: 0,
1177
+ ttftSum: 0,
1178
+ latSum: 0,
1179
+ lastTs: 0,
1180
+ lastOk: null,
1181
+ lastCode: null,
1182
+ recentErrors: 0,
1183
+ tokSum: 0,
1184
+ tokIn: 0,
1185
+ tokOut: 0,
1186
+ tokCache: 0
1187
+ }
1188
+ pMap.set(e.model, a)
1189
+ }
1190
+ a.total++
1191
+ const tokIn = (e && e.inputTokens) || 0
1192
+ const tokOut = (e && e.outputTokens) || 0
1193
+ const tokCache = ((e && e.cacheReadTokens) || 0) + ((e && e.cacheWriteTokens) || 0)
1194
+ totalTokIn += tokIn
1195
+ totalTokOut += tokOut
1196
+ totalTokCache += tokCache
1197
+ a.tokIn += tokIn
1198
+ a.tokOut += tokOut
1199
+ a.tokCache += tokCache
1200
+ a.tokSum += tokIn + tokOut + tokCache
1201
+ if (e.ok) {
1202
+ a.success++
1203
+ if (e.ttftMs != null && e.ttftMs >= 0) a.ttftSum += e.ttftMs
1204
+ if (e.latencyMs != null && e.latencyMs >= 0) a.latSum += e.latencyMs
1205
+ } else {
1206
+ a.fail++
1207
+ if (e.code) a.lastCode = e.code
1208
+ }
1209
+ if ((e.ts || 0) > a.lastTs) {
1210
+ a.lastTs = e.ts || 0
1211
+ a.lastOk = e.ok
1212
+ }
1149
1213
  }
1150
1214
  }
1151
1215
 
1152
- const totalRequests = allEvents.length
1153
- const totalSuccess = allEvents.filter((x) => x.ok).length
1216
+ // 顶部指标卡:digest 路径用聚合值求和,旧路径用原始事件计数
1217
+ // (allEvents 提升到函数级:旧 host 回落路径的 else 块在此作用域外定义它)
1218
+ let allEvents = []
1219
+ let totalRequests = 0
1220
+ let totalSuccess = 0
1221
+ let ttftWeightedSum = 0
1222
+ if (digestRows !== null) {
1223
+ for (const pMap of byProvider.values()) {
1224
+ for (const a of pMap.values()) {
1225
+ totalRequests += a.total
1226
+ totalSuccess += a.success
1227
+ ttftWeightedSum += a.ttftSum
1228
+ }
1229
+ }
1230
+ } else {
1231
+ totalRequests = allEvents.length
1232
+ totalSuccess = allEvents.filter((x) => x.ok).length
1233
+ for (const pMap of byProvider.values()) {
1234
+ for (const a of pMap.values()) ttftWeightedSum += a.ttftSum
1235
+ }
1236
+ }
1154
1237
  const globalRate = totalRequests > 0 ? ((totalSuccess / totalRequests) * 100).toFixed(1) + '%' : '100%'
1155
- const validTtfts = allEvents.filter((x) => x.ok && x.ttftMs != null).map((x) => x.ttftMs)
1156
- const avgGlobalTtft = validTtfts.length > 0 ? (validTtfts.reduce((a, b) => a + b, 0) / validTtfts.length / 1000).toFixed(2) + 's' : '—'
1238
+ const avgGlobalTtft = totalSuccess > 0 ? (ttftWeightedSum / totalSuccess / 1000).toFixed(2) + 's' : '—'
1157
1239
 
1158
1240
  const providerGroups = [...byProvider.entries()].map(([provName, modelMap]) => {
1159
1241
  const models = [...modelMap.values()].map((m) => {
@@ -1212,7 +1294,9 @@ window.__ModuleLoader__.load({
1212
1294
  el('div', { className: 'mcm-metric-card' },
1213
1295
  el('span', { className: 'mcm-metric-label' }, windowMode === '30m' ? '近 30 分钟请求' : windowMode === '24h' ? '近 24 小时请求' : '7 天全周期请求'),
1214
1296
  el('span', { className: 'mcm-metric-value' }, totalRequests),
1215
- el('span', { style: { fontSize: 11, color: 'var(--dsw-alias-label-tertiary)' } }, '真实上游交互捕获')
1297
+ // digest 聚合是 7 天窗口口径;30m/24h 只按「最近活跃」过滤渠道,数值为该渠道 7 天累计
1298
+ el('span', { style: { fontSize: 11, color: 'var(--dsw-alias-label-tertiary)' } },
1299
+ digestRows !== null && windowMode !== '7d' ? '窗口内活跃渠道的 7 天累计(聚合口径)' : '真实上游交互捕获')
1216
1300
  ),
1217
1301
  el('div', { className: 'mcm-metric-card' },
1218
1302
  el('span', { className: 'mcm-metric-label' }, '总 Token 用量'),
@@ -1334,7 +1418,12 @@ window.__ModuleLoader__.load({
1334
1418
  setChannels(ch)
1335
1419
  setChannelsDraft((prev) => prev || clone(ch))
1336
1420
  const hn = findNs('model-channel-health')
1337
- setHealth(hn ? { records: (hn.value && hn.value.records) || {}, speedResults: (hn.value && hn.value.speedResults) || {}, runtime: (hn.value && hn.value.runtime) || {} } : null)
1421
+ // 0.3.1:健康权威数据在 host 的 storageDomain;client 只接收小投影 digest。
1422
+ // 兼容旧 host(无 digest 字段)时回落到 records 原始视图,避免升级窗口白屏。
1423
+ const hv = (hn && hn.value) || {}
1424
+ setHealth(hv.digest || (hv.digestAt !== undefined)
1425
+ ? { digest: hv.digest || [], digestAt: hv.digestAt || 0, speedResults: hv.speedResults || {}, runtime: hv.runtime || {} }
1426
+ : (hn ? { records: hv.records || {}, speedResults: hv.speedResults || {}, runtime: hv.runtime || {} } : null))
1338
1427
  }).catch((e) => setNotice('加载失败: ' + String(e)))
1339
1428
  }
1340
1429
 
@@ -1350,11 +1439,10 @@ window.__ModuleLoader__.load({
1350
1439
  const d = unwrap(resp)
1351
1440
  const hn = ((d && d.namespaces) || []).find((n) => n && n.ns === 'model-channel-health')
1352
1441
  if (hn && hn.value) {
1353
- setHealth({
1354
- records: hn.value.records || {},
1355
- speedResults: hn.value.speedResults || {},
1356
- runtime: hn.value.runtime || {}
1357
- })
1442
+ const hv = hn.value
1443
+ setHealth(hv.digest || (hv.digestAt !== undefined)
1444
+ ? { digest: hv.digest || [], digestAt: hv.digestAt || 0, speedResults: hv.speedResults || {}, runtime: hv.runtime || {} }
1445
+ : { records: hv.records || {}, speedResults: hv.speedResults || {}, runtime: hv.runtime || {} })
1358
1446
  }
1359
1447
  }).catch(() => {})
1360
1448
  }
@@ -0,0 +1,169 @@
1
+ /**
2
+ * 健康数据存储:原始流水(事件/测速)搬入 storageDomain 的 per-record KV 表。
3
+ *
4
+ * 动机(审计 F11/F12/F23 的根因):0.1 时代健康流水整字段存在 settings 命名空间里,
5
+ * 与配置共用同一条「HMR 独占事务 + 整字段落盘」通道,导致:
6
+ * - 2s 防抖 flush 与 volatile-update 快照覆盖互相踩(刚记的账被旧快照抹掉);
7
+ * - profile patch 文件膨胀(实测 6350 行中 health 约占 3000 行);
8
+ * - 任何无关 settings 写入都会整段重写健康数据。
9
+ *
10
+ * 现在的归属:事实数据(append-only 事件)进 domain;settings 里只保留
11
+ * 一份给 client 渲染用的小投影(聚合摘要 + 运行态指针),由 host 定期刷新。
12
+ *
13
+ * 组合方式照抄 dsh-session-projection-cache 范本:
14
+ * - 依赖方 inject ['storageDomain'](dsh-base 已装 json 后端,root=storages/);
15
+ * - Service.init 里 open(spec),effect 持有 close;
16
+ * - per-record 布局:一条事件一个文档,坏记录 backup-and-skip 不阻塞整体。
17
+ */
18
+ import { z } from 'zod';
19
+ import { defineDomain, domainTable } from '@deepseek-ai/dsh-storage-domain';
20
+
21
+ const WINDOW_MS = 7 * 24 * 3600 * 1000;
22
+ // 每渠道保留上限:30m/24h 视图 + selector 置顶绰绰有余,同时控制 storages/ 目录体量
23
+ const PER_KEY_LIMIT = 300;
24
+
25
+ /** 单条健康事件记录(zod,domain 落盘边界校验)。向后兼容:token 字段可选。 */
26
+ const eventRecord = z.object({
27
+ ts: z.number().int().nonnegative(),
28
+ provider: z.string(),
29
+ model: z.string(),
30
+ ok: z.boolean(),
31
+ ttftMs: z.number().nullable().optional(),
32
+ latencyMs: z.number().nullable().optional(),
33
+ code: z.string().nullable().optional(),
34
+ inputTokens: z.number().optional(),
35
+ outputTokens: z.number().optional(),
36
+ cacheReadTokens: z.number().optional(),
37
+ cacheWriteTokens: z.number().optional(),
38
+ reasoningTokens: z.number().optional(),
39
+ });
40
+
41
+ /** 单渠道的事件桶记录:key = provider(与旧 state.records 的键一致,迁移零转换)。 */
42
+ const bucketRecord = z.object({
43
+ events: z.array(eventRecord),
44
+ });
45
+
46
+ /** 测速结果桶:key = 组 id(与旧 state.speedResults 的键一致)。 */
47
+ const speedRecord = z.object({
48
+ rows: z.array(z.object({
49
+ provider: z.string(),
50
+ model: z.string(),
51
+ ok: z.boolean(),
52
+ ttft: z.number().nullable(),
53
+ latency: z.number().nullable(),
54
+ at: z.number(),
55
+ failure: z.string().nullable().optional(),
56
+ })),
57
+ });
58
+
59
+ const healthDomainSpec = defineDomain({
60
+ name: 'model_channel_health',
61
+ version: 1,
62
+ layout: 'per-record',
63
+ invalidRecords: 'backup-and-skip',
64
+ tables: {
65
+ events: domainTable(bucketRecord),
66
+ speed: domainTable(speedRecord),
67
+ },
68
+ });
69
+
70
+ /**
71
+ * HealthStore:封装 domain 读写 + 保留策略 + 聚合。
72
+ * 非 Service(宿主插件自有生命周期),由 apply() 创建并 effect 持有 close。
73
+ */
74
+ export class HealthStore {
75
+ constructor(ctx) {
76
+ this.ctx = ctx;
77
+ this.ready = null; // Promise<Domain>
78
+ this.events = null; // KvTable
79
+ this.speed = null; // KvTable
80
+ this.eventChain = Promise.resolve(); // 类内写链:appendEvent 串行化(并发不丢更新)
81
+ }
82
+
83
+ /** 打开 domain(幂等)。dsh-base 未装 storage 栈的 profile 会 reject——调用方降级。 */
84
+ open() {
85
+ if (this.ready === null) {
86
+ this.ready = this.ctx.storageDomain.open(healthDomainSpec).then((domain) => {
87
+ this.events = domain.table('events');
88
+ this.speed = domain.table('speed');
89
+ return domain;
90
+ });
91
+ }
92
+ return this.ready;
93
+ }
94
+
95
+ /** effect 用:domain 关闭(写入队列排干后释放)。 */
96
+ async close() {
97
+ if (this.ready === null) return;
98
+ const domain = await this.ready.catch(() => null);
99
+ this.ready = null;
100
+ this.events = null;
101
+ this.speed = null;
102
+ if (domain) await domain.close().catch(() => {});
103
+ }
104
+
105
+ /** 追加一条事件到渠道桶,同时执行窗口过期 + 条数截断。
106
+ * 并发安全:桶初始化(put)与追加(update)都排队到类内写链,逐个重读当前值,
107
+ * 并发 appendEvent 不交错、不互相覆盖(等价 KvTable.update 的原子语义;
108
+ * 桶不存在时 KvTable.update 报 missing-key,故初始化先行)。
109
+ */
110
+ appendEvent(provider, rec) {
111
+ if (this.events === null) return Promise.resolve();
112
+ this.eventChain = this.eventChain.then(async () => {
113
+ const cutoff = Date.now() - WINDOW_MS;
114
+ const cur = this.events.get(provider);
115
+ if (cur === undefined) {
116
+ await this.events.put(provider, { events: [rec] });
117
+ return;
118
+ }
119
+ await this.events.update(provider, (c) => ({
120
+ events: [...c.events.filter((e) => e.ts >= cutoff), rec].slice(-PER_KEY_LIMIT),
121
+ }));
122
+ }).catch(() => { }); // 链不断:失败吞掉(调用方已有 warn),后续追加继续
123
+ return this.eventChain;
124
+ }
125
+
126
+ /** 整组测速结果替换写入。 */
127
+ async putSpeedRows(groupId, rows) {
128
+ if (this.speed === null) return;
129
+ await this.speed.put(groupId, { rows });
130
+ }
131
+
132
+ /** 读全部事件桶(迁移/聚合用)。 */
133
+ allEventBuckets() {
134
+ const out = {};
135
+ if (this.events === null) return out;
136
+ for (const [provider, bucket] of this.events.entries()) out[provider] = bucket.events;
137
+ return out;
138
+ }
139
+
140
+ /** 读全部测速桶。 */
141
+ allSpeedBuckets() {
142
+ const out = {};
143
+ if (this.speed === null) return out;
144
+ for (const [groupId, bucket] of this.speed.entries()) out[groupId] = bucket.rows;
145
+ return out;
146
+ }
147
+
148
+ /**
149
+ * 存量迁移:settings health 子树的 records/speedResults 一次性搬入 domain。
150
+ * 迁移以「桶里有数据即跳过」防重放;完成后由调用方写迁移标记。
151
+ * 并发安全:逐桶判断改走原子链(迁移进行中 recordHealth 并发创建的新桶不会被旧数据覆盖)。
152
+ */
153
+ async migrateFrom(records, speedResults) {
154
+ if (this.events === null) return false;
155
+ const cutoff = Date.now() - WINDOW_MS;
156
+ for (const [provider, list] of Object.entries(records || {})) {
157
+ if (!Array.isArray(list) || list.length === 0) continue;
158
+ if (this.events.get(provider) !== undefined) continue; // 已有桶:domain 是权威,不回搬
159
+ const events = list.filter((e) => e && e.ts >= cutoff).slice(-PER_KEY_LIMIT);
160
+ if (events.length > 0) await this.events.put(provider, { events });
161
+ }
162
+ for (const [groupId, rows] of Object.entries(speedResults || {})) {
163
+ if (!Array.isArray(rows) || rows.length === 0) continue;
164
+ if (this.speed.get(groupId) !== undefined) continue;
165
+ await this.speed.put(groupId, { rows });
166
+ }
167
+ return true;
168
+ }
169
+ }
package/src/index.js CHANGED
@@ -1,6 +1,9 @@
1
1
  import z from "@deepseek-ai/schemastery";
2
+ import { HealthStore } from './health-store.js';
2
3
  export const name = 'model-channel-manager';
3
4
  export const inject = ['llm', 'timer'];
5
+ // storageDomain 由 dsh-base 组合(json 后端,root=storages/);不进 inject 硬依赖——
6
+ // 缺栈的 profile 降级为「健康流水不持久化」而不是拒绝启动。
4
7
  const ROUTE_PREFIX = 'roundrobin/';
5
8
  const GROUP_ID_RE = /^[a-z0-9][a-z0-9-]*$/;
6
9
  const NS_CONFIG = 'model-channels';
@@ -9,6 +12,8 @@ const NS_HEALTH = 'model-channel-health';
9
12
  // 因此原先的 model-channels / model-channel-health 两个命名空间合并为本插件的 Config
10
13
  // (字段 groups / providerOrder / effortMemory / health),命名空间 id 即本行 id。
11
14
  const HEALTH_SCHEMA = z.object({
15
+ // records/speedResults 仅作旧版存量迁移的读取源(迁入 domain 后不再写入);
16
+ // 权威数据在 storageDomain 的 model_channel_health 单元。
12
17
  records: z.dict(z.array(z.any())).default({}),
13
18
  speedResults: z.dict(z.array(z.any())).default({}),
14
19
  runtime: z.dict(z.any()).default({}),
@@ -18,6 +23,10 @@ const HEALTH_SCHEMA = z.object({
18
23
  testResults: z.dict(z.any()).default({}),
19
24
  lastTestHandledNonce: z.number().default(0),
20
25
  legacyMigrated: z.any(),
26
+ healthMigrated: z.any(),
27
+ // 小投影:host 聚合好的健康摘要(client 渲染用),digestAt 为生成时刻
28
+ digest: z.array(z.any()).default([]),
29
+ digestAt: z.number().default(0),
21
30
  }).loose(true);
22
31
  export const Config = z.object({
23
32
  groups: z.array(z.any()).default([]).volatile(),
@@ -286,7 +295,6 @@ export function apply(ctx, config) {
286
295
  const recordHealth = (gid, cand, entry) => {
287
296
  const now = Date.now();
288
297
  const key = gid || 'global';
289
- const list = state.records[key] || (state.records[key] = []);
290
298
  const rec = { ts: now, provider: cand.provider, model: cand.model, ok: entry.ok, ttftMs: entry.ttftMs, latencyMs: entry.latencyMs, code: entry.code || null };
291
299
  // 上游真实 token 用量(来自 usage StreamChunk,计费口径与 dsh-token-meter 一致):
292
300
  // 展平写入记录条目;缺失/非法字段不写,旧记录按 0 处理(向后兼容)。
@@ -298,53 +306,30 @@ export function apply(ctx, config) {
298
306
  if (Number.isFinite(u.cacheWriteTokens)) rec.cacheWriteTokens = u.cacheWriteTokens;
299
307
  if (Number.isFinite(u.reasoningTokens)) rec.reasoningTokens = u.reasoningTokens;
300
308
  }
309
+ // 内存镜像(聚合与 smart 键读内存即可)+ domain 追加(事实持久化)。
310
+ // domain 写入走 per-record put(写一条只动一个渠道桶文档),
311
+ // 不再有 2s 防抖与 volatile 快照覆盖的竞态窗口。
312
+ const list = state.records[key] || (state.records[key] = []);
301
313
  list.push(rec);
302
314
  const cutoff = now - 7 * 24 * 3600 * 1000;
303
- // 每键保留 300 条:健康页签 30m/24h 视图与选择器置顶绰绰有余,同时把
304
- // settings.yaml 的体量压到可读(2000 条/键时该文件一度膨胀到 2 万行)
305
315
  state.records[key] = list.filter((e) => e.ts >= cutoff).slice(-300);
306
- pendingRecords.push({ key, rec });
307
- persistHealth();
316
+ if (healthStore !== null) {
317
+ healthStore.appendEvent(key, rec).catch((e) => console.warn('[model-channel-manager] health event append failed:', e && e.message));
318
+ }
319
+ scheduleDigest();
308
320
  };
309
321
  const persistSpeedResults = async (r) => {
322
+ if (healthStore !== null) {
323
+ await healthStoreSpeed(r).catch(() => { });
324
+ }
310
325
  if (bus === null) return;
311
- await bus.writeHealth({ speedResults: clone(r), runtime: await persistRuntime() });
312
- };
313
- // ---------- 健康持久化节流:合并 2s 窗口内的写入,避免每个请求全量重写 settings ----------
314
- let healthFlushHandle = null;
315
- let healthDirty = false;
316
- // 尚未落盘的流水:reloadFromConfig 每次 volatile-update 都用配置快照整体覆盖 state.records,
317
- // 而 flush 有 2s 防抖——不缓冲的话,刚记下的一笔会在落盘前被内存覆盖(实测「测试成功不入账」就是这个)。
318
- let pendingRecords = [];
319
- const flushHealthNow = async () => {
320
- if (bus === null) return;
321
- const flushed = pendingRecords.slice();
322
- await bus.writeHealth({ records: clone(pullRecords()), speedResults: clone(pullSpeedResults()), runtime: await persistRuntime() });
323
- // 写入成功后才清账;失败就留着,等下一次 flush 重试
324
- pendingRecords = pendingRecords.filter((p) => !flushed.includes(p));
326
+ await bus.writeHealth({ runtime: await persistRuntime() });
325
327
  };
326
- const persistHealth = () => {
327
- healthDirty = true;
328
- if (healthFlushHandle === null) {
329
- healthFlushHandle = ctx.timeout(() => {
330
- healthFlushHandle = null;
331
- if (!healthDirty) return;
332
- healthDirty = false;
333
- flushHealthNow().catch(() => { });
334
- }, 2000);
328
+ const healthStoreSpeed = async (r) => {
329
+ for (const [groupId, rows] of Object.entries(r || {})) {
330
+ await healthStore.putSpeedRows(groupId, rows);
335
331
  }
336
332
  };
337
- ctx.effect(() => () => {
338
- // 插件卸载时尽力冲刷最后一批健康数据(settings 服务仍在,失败静默)
339
- if (healthFlushHandle !== null) {
340
- healthFlushHandle();
341
- healthFlushHandle = null;
342
- }
343
- if (healthDirty) {
344
- healthDirty = false;
345
- flushHealthNow().catch(() => { });
346
- }
347
- }, 'model-channel health flush');
348
333
  // ---------- 引擎:单候选尝试 ----------
349
334
  async function* streamAttempt(cfg, cand, options, attemptStart) {
350
335
  if (cand.provider.startsWith(ROUTE_PREFIX))
@@ -823,6 +808,89 @@ export function apply(ctx, config) {
823
808
  else if (JSON.stringify(current) !== JSON.stringify(next))
824
809
  adapterHandle.replace(next);
825
810
  }
811
+ // ---------- 健康数据 domain 存储 + 存量迁移 + 小投影 ----------
812
+ // healthStore:storageDomain 存在时启用(dsh-base 组合了 json 后端)。
813
+ // 启动序列:先开 domain → 迁移 settings 存量 → reloadFromConfig 从 domain 读权威数据。
814
+ let healthStore = null;
815
+ // digest:写进 settings health 子树的小投影(client 健康页渲染用)。
816
+ // 聚合上移 host,client 不再拉原始流水全量 describe——settings 里不再出现 records。
817
+ let digestTimer = null;
818
+ const DIGEST_INTERVAL = 5000;
819
+ const buildDigest = () => {
820
+ // 与旧 client HealthPanel 聚合口径一致:total/success/ttft/latency/token(计费口径
821
+ // = input + cacheRead + cacheWrite + output),外加 lastTs/lastOk/lastCode
822
+ const by = new Map();
823
+ for (const list of Object.values(state.records)) {
824
+ for (const e of list) {
825
+ const key = e.provider + '::' + e.model;
826
+ let a = by.get(key);
827
+ if (a === undefined) {
828
+ a = { provider: e.provider, model: e.model, total: 0, success: 0, ttftSum: 0, latSum: 0, tokIn: 0, tokOut: 0, tokCache: 0, lastTs: 0, lastOk: null, lastCode: null };
829
+ by.set(key, a);
830
+ }
831
+ a.total++;
832
+ const inTok = e.inputTokens || 0;
833
+ const outTok = e.outputTokens || 0;
834
+ const cacheTok = (e.cacheReadTokens || 0) + (e.cacheWriteTokens || 0);
835
+ a.tokIn += inTok;
836
+ a.tokOut += outTok;
837
+ a.tokCache += cacheTok;
838
+ if (e.ok) {
839
+ a.success++;
840
+ if (e.ttftMs != null && e.ttftMs >= 0) a.ttftSum += e.ttftMs;
841
+ if (e.latencyMs != null && e.latencyMs >= 0) a.latSum += e.latencyMs;
842
+ }
843
+ if ((e.ts || 0) > a.lastTs) {
844
+ a.lastTs = e.ts || 0;
845
+ a.lastOk = e.ok;
846
+ a.lastCode = e.ok ? null : (e.code || null); // lastCode 随 lastTs 对齐(不再按遍历序覆盖)
847
+ }
848
+ }
849
+ }
850
+ return [...by.values()].map((a) => ({
851
+ provider: a.provider, model: a.model, total: a.total, success: a.success,
852
+ ttftAvg: a.success > 0 ? Math.round(a.ttftSum / a.success) : null,
853
+ latAvg: a.success > 0 ? Math.round(a.latSum / a.success) : null,
854
+ tokIn: a.tokIn, tokOut: a.tokOut, tokCache: a.tokCache,
855
+ lastTs: a.lastTs, lastOk: a.lastOk, lastCode: a.lastCode,
856
+ }));
857
+ };
858
+ const flushDigest = () => {
859
+ digestTimer = null;
860
+ if (bus !== null) {
861
+ bus.writeHealth({ digest: buildDigest(), digestAt: Date.now() }).catch(() => { });
862
+ }
863
+ };
864
+ const scheduleDigest = () => {
865
+ if (digestTimer === null) {
866
+ digestTimer = ctx.timeout(flushDigest, DIGEST_INTERVAL);
867
+ }
868
+ };
869
+ ctx.effect(() => () => {
870
+ if (digestTimer !== null) {
871
+ digestTimer();
872
+ digestTimer = null;
873
+ }
874
+ flushDigest();
875
+ }, 'model-channel health digest');
876
+ // settings 存量 records/speedResults → domain 一次性迁移(桶已存在即跳过,domain 权威)。
877
+ const migrateHealthToDomain = async () => {
878
+ if (healthStore === null || bus === null) return;
879
+ const health = healthOf();
880
+ if (health.healthMigrated === true) return;
881
+ let ok = false;
882
+ try {
883
+ ok = await healthStore.migrateFrom(health.records || {}, health.speedResults || {});
884
+ }
885
+ catch (_e) {
886
+ ok = false; // 下次启动再试
887
+ }
888
+ if (!ok) return; // store 已关闭(卸载竞态):不写标记,避免标记与数据永久不一致
889
+ // 同一次合并写:迁移标记 + 清空 settings 里的旧存量(否则 writeHealth 的
890
+ // Object.assign 会把 records/speedResults 永远带下去,profile 膨胀问题未真正解决)
891
+ bus.writeHealth({ healthMigrated: true, records: {}, speedResults: {} }).catch(() => { });
892
+ };
893
+
826
894
  // ---------- settings 总线接入(响应式:settings 服务异步初始化,apply 时查询太早) ----------
827
895
  let bus = null;
828
896
  let runtimeRestored = false;
@@ -846,18 +914,19 @@ export function apply(ctx, config) {
846
914
  ctx.timeout(() => { fn(); }, ms);
847
915
  };
848
916
  // 从实例配置读取全部状态;volatile 提交后由 loader/volatile-update 触发重载。
917
+ // 0.3.1 起 records/speedResults 的权威在 storageDomain(healthStore);
918
+ // settings health 子树只承载小投影(digest 给 client 渲染)+ 运行态 + 任务哨,
919
+ // volatile 快照覆盖不再能抹掉任何流水(F11/F23 的根因消除)。
849
920
  function reloadFromConfig(initial) {
850
921
  const cfg = cfgOf();
851
922
  state.config = clone(cfg) || { groups: [] };
852
923
  const health = healthOf();
853
- state.records = clone(health.records || {});
854
- // 配置快照覆盖后,把尚未落盘的流水补回来(ts+provider+model 相同视为同一笔,不重复)
855
- for (const p of pendingRecords) {
856
- const list = state.records[p.key] || (state.records[p.key] = []);
857
- if (!list.some((e) => e.ts === p.rec.ts && e.provider === p.rec.provider && e.model === p.rec.model))
858
- list.push(p.rec);
859
- }
860
- state.speedResults = clone(health.speedResults || {});
924
+ state.records = healthStore !== null
925
+ ? clone(healthStore.allEventBuckets())
926
+ : clone(health.records || {});
927
+ state.speedResults = healthStore !== null
928
+ ? clone(healthStore.allSpeedBuckets())
929
+ : clone(health.speedResults || {});
861
930
  state.testResults = clone(health.testResults || {});
862
931
  if (initial && !runtimeRestored) {
863
932
  restoreRuntime(health.runtime || {});
@@ -940,12 +1009,54 @@ export function apply(ctx, config) {
940
1009
  cfgScope: { get: () => cfgOf() },
941
1010
  healthScope: { get: () => healthOf() },
942
1011
  };
943
- reloadFromConfig(true);
1012
+ // healthStore 可用(dsh-base 的 storage 栈在场)时:开 domain → 迁移存量 →
1013
+ // 用 domain 权威数据重建内存镜像 → boot;否则同步走旧路径。
1014
+ // ctx.inject 回调不是 async 函数,domain 初始化用 promise 链表达。
1015
+ // 卸载竞态防护:then 回调先验 fiber 活性,inactive 时回滚 healthStore 并跳过
1016
+ // effect 注册(对 inactive fiber 注册 effect 会抛 INACTIVE_EFFECT,若被吞掉则
1017
+ // domain 永不 close,facility 名字被占 → HMR 重载后 already-open 静默降级)。
1018
+ // boot 也挪进链尾:domain 路径下 reloadFromConfig(true) 异步排队,若 boot 先跑,
1019
+ // state.speedResults 尚为空 → onFirstUse 组每次启动都重测;且 state.config 为空
1020
+ // 时 rewireRoutes 推迟,启动早期虚拟路由短暂不存在。
1021
+ const domainFacility = ctx.get('storageDomain');
1022
+ const fiberUid = ctx.fiber && ctx.fiber.uid;
1023
+ const fiberAlive = () => ctx.fiber !== undefined && ctx.fiber.uid === fiberUid && ctx.fiber.state !== 4 /* disposed */;
1024
+ if (domainFacility !== undefined) {
1025
+ const store = new HealthStore(ctx);
1026
+ Promise.resolve()
1027
+ .then(() => store.open())
1028
+ .then(() => {
1029
+ if (!fiberAlive()) {
1030
+ // 已卸载:回滚赋值,由 closeAll 兜底回收 domain
1031
+ void store.close().catch(() => { });
1032
+ return;
1033
+ }
1034
+ healthStore = store;
1035
+ ctx.effect(() => () => { void store.close(); }, 'model-channel health domain close');
1036
+ })
1037
+ .catch((e) => {
1038
+ console.warn('[model-channel-manager] storageDomain open failed, health records stay in-memory only:', e && e.message);
1039
+ })
1040
+ .then(() => (healthStore !== null ? migrateHealthToDomain().catch(() => { }) : null))
1041
+ .then(() => {
1042
+ if (fiberAlive())
1043
+ reloadFromConfig(true);
1044
+ })
1045
+ .then(() => {
1046
+ if (fiberAlive())
1047
+ boot();
1048
+ })
1049
+ .catch(() => { }); // 卸载竞态下 reload/boot 内部可能 throw,链尾兜底防 unhandled rejection
1050
+ }
1051
+ else {
1052
+ reloadFromConfig(true);
1053
+ boot();
1054
+ }
944
1055
  ctx.on('loader/volatile-update', () => {
945
1056
  reloadFromConfig(false);
946
1057
  console.log('[model-channel-manager] config hot-reloaded, routes:', pullConfig().groups.map((g) => g.id).join(', ') || '(none)');
947
1058
  });
948
- boot();
1059
+ // boot 已在 domain ready 链尾(或降级分支)调用,此处不再调
949
1060
  });
950
1061
  // ---------- 全局 LLM 请求健康拦截 (涵盖所有非虚拟路由的真实渠道模型调用) ----------
951
1062
  ctx.on('llm/stream', async function* (options, next) {