dsh-tap 0.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (121) hide show
  1. package/AGENTS.md +99 -0
  2. package/CHANGELOG.md +751 -0
  3. package/LICENSE +21 -0
  4. package/README.en.md +59 -0
  5. package/README.md +256 -0
  6. package/cordis.patch.yml +385 -0
  7. package/core/bridge.js +698 -0
  8. package/core/json-store.js +91 -0
  9. package/core/rotation.js +108 -0
  10. package/core/usage-meter.js +176 -0
  11. package/docs/diagnosis-cache-quota.md +181 -0
  12. package/docs/diagnosis-qoder-flash.md +67 -0
  13. package/docs/diagnosis-trae-3003.md +302 -0
  14. package/docs/goals/bridge-port-host-split.md +91 -0
  15. package/docs/goals/desktop-adaptation.md +106 -0
  16. package/docs/goals/qoder-cn-provider-design.md +308 -0
  17. package/docs/goals/settings-card-ux-redesign-plan.md +1680 -0
  18. package/docs/goals/settings-card-ux-redesign.md +162 -0
  19. package/docs/goals/trae-agent-v3.md +41 -0
  20. package/docs/goals/trae-work-cn-repair.md +44 -0
  21. package/docs/goals/v0.8-/351/242/235/345/272/246/345/217/257/350/247/201-/346/250/241/345/236/213/345/212/250/346/200/201/345/214/226-/345/244/232/346/234/215/345/212/241/345/225/206.md +85 -0
  22. package/docs/pitfalls.md +105 -0
  23. package/docs/reverse/trae-cloud-api.md +218 -0
  24. package/docs/reverse/trae-model-catalog.md +129 -0
  25. package/docs/reverse/traework-cn.md +520 -0
  26. package/docs/rules/STATE.md +198 -0
  27. package/docs/rules/content-moderation.md +54 -0
  28. package/docs/rules/dev-role-boundary.md +94 -0
  29. package/docs/rules/extra-providers.md +42 -0
  30. package/docs/rules/gateway-facts.md +91 -0
  31. package/docs/rules/oauth-handshake.md +76 -0
  32. package/docs/rules/prompt-cache.md +93 -0
  33. package/docs/rules/quota-signals.md +125 -0
  34. package/docs/rules/routing.md +84 -0
  35. package/docs/rules/templates/oauth-reverse-checklist.md +42 -0
  36. package/docs/rules/trae-surface.md +242 -0
  37. package/docs/rules/ua-validation.md +81 -0
  38. package/host-config.js +282 -0
  39. package/index.js +2306 -0
  40. package/lib/client.js +2893 -0
  41. package/local-scan.js +104 -0
  42. package/package.json +82 -0
  43. package/providers/ark/index.js +11 -0
  44. package/providers/bailian/index.js +10 -0
  45. package/providers/bigmodel/index.js +11 -0
  46. package/providers/codebuddy/agenttool.js +122 -0
  47. package/providers/codebuddy/catalog.js +227 -0
  48. package/providers/codebuddy/errors.js +44 -0
  49. package/providers/codebuddy/headers.js +36 -0
  50. package/providers/codebuddy/images.js +125 -0
  51. package/providers/codebuddy/index.js +123 -0
  52. package/providers/codebuddy/oauth.js +279 -0
  53. package/providers/deepseek/index.js +11 -0
  54. package/providers/moonshot/index.js +11 -0
  55. package/providers/openai-compat.js +177 -0
  56. package/providers/openrouter/index.js +27 -0
  57. package/providers/qoder/catalog.js +145 -0
  58. package/providers/qoder/cosy.js +419 -0
  59. package/providers/qoder/gateway.js +563 -0
  60. package/providers/qoder/index.js +116 -0
  61. package/providers/qoder/oauth.js +364 -0
  62. package/providers/qoder/qoder_auth.wasm +0 -0
  63. package/providers/qoder/quota.js +56 -0
  64. package/providers/qwen/index.js +15 -0
  65. package/providers/tool-pairing.js +129 -0
  66. package/providers/trae/catalog.js +103 -0
  67. package/providers/trae/errors.js +85 -0
  68. package/providers/trae/gateway.js +853 -0
  69. package/providers/trae/index.js +126 -0
  70. package/providers/trae/oauth.js +443 -0
  71. package/providers/trae/quota.js +75 -0
  72. package/providers/trae/remote.js +365 -0
  73. package/scripts/capture-cache.mjs +65 -0
  74. package/scripts/capture-traffic.mjs +83 -0
  75. package/scripts/hermes-probe-dev-role.mjs +263 -0
  76. package/scripts/measure-latency.mjs +253 -0
  77. package/scripts/probe-ark-thinking.mjs +298 -0
  78. package/scripts/probe-cache-decline.mjs +292 -0
  79. package/scripts/probe-cache-ttl.mjs +221 -0
  80. package/scripts/probe-cache.mjs +156 -0
  81. package/scripts/probe-codebuddy-efforts.mjs +355 -0
  82. package/scripts/probe-codebuddy-tier-wiring.mjs +244 -0
  83. package/scripts/probe-effort-gaps.mjs +133 -0
  84. package/scripts/probe-media.mjs +115 -0
  85. package/scripts/probe-moderation.mjs +159 -0
  86. package/scripts/probe-oauth.mjs +617 -0
  87. package/scripts/probe-qoder-attribution-arm8.mjs +92 -0
  88. package/scripts/probe-qoder-attribution-arm9.mjs +102 -0
  89. package/scripts/probe-qoder-attribution.mjs +266 -0
  90. package/scripts/probe-qoder-flash-confirm.mjs +94 -0
  91. package/scripts/probe-qoder-live.mjs +344 -0
  92. package/scripts/probe-qoder-matrix.mjs +274 -0
  93. package/scripts/probe-qoder-null-content.mjs +153 -0
  94. package/scripts/probe-qoder-pairing.mjs +308 -0
  95. package/scripts/probe-qoder-quota.mjs +162 -0
  96. package/scripts/probe-qoder-thinking-config.mjs +52 -0
  97. package/scripts/probe-qoder-thinking-efforts.mjs +269 -0
  98. package/scripts/probe-quota.mjs +136 -0
  99. package/scripts/probe-quota2.mjs +144 -0
  100. package/scripts/probe-routing.mjs +226 -0
  101. package/scripts/probe-trae-3003-diagnosis.mjs +139 -0
  102. package/scripts/probe-trae-agent-v3.mjs +399 -0
  103. package/scripts/probe-trae-efforts.mjs +149 -0
  104. package/scripts/probe-trae-live.mjs +147 -0
  105. package/scripts/probe-trae-max-effort.mjs +179 -0
  106. package/scripts/probe-trae-model-routing.mjs +381 -0
  107. package/scripts/probe-trae-thinking-scene.mjs +238 -0
  108. package/scripts/probe-trae-transport-outage.mjs +176 -0
  109. package/scripts/probe-ua.mjs +508 -0
  110. package/scripts/trae-model-catalog.mjs +632 -0
  111. package/scripts/verify-agents-md.mjs +45 -0
  112. package/scripts/verify-bridge.mjs +1041 -0
  113. package/scripts/verify-core-generic.mjs +302 -0
  114. package/scripts/verify-desktop-acceptance.mjs +184 -0
  115. package/scripts/verify-host-config.mjs +265 -0
  116. package/scripts/verify-models.mjs +390 -0
  117. package/scripts/verify-providers.mjs +255 -0
  118. package/scripts/verify-qoder-provider.mjs +1020 -0
  119. package/scripts/verify-rotation.mjs +308 -0
  120. package/scripts/verify-trae-model-catalog.mjs +284 -0
  121. package/scripts/verify-trae-provider.mjs +1451 -0
@@ -0,0 +1,91 @@
1
+ /**
2
+ * core/json-store.js — provider-agnostic JSON file persistence and
3
+ * environment/credentials-file key resolution.
4
+ *
5
+ * Extracted from index.js during the core/providers split. Nothing in this
6
+ * file knows anything about any specific upstream gateway.
7
+ */
8
+
9
+ import { readFileSync, writeFileSync, renameSync, unlinkSync, mkdirSync, existsSync } from 'node:fs'
10
+ import { dirname } from 'node:path'
11
+
12
+ /** Read a JSON object file; missing/corrupt/non-object all yield {}. */
13
+ export function readJson(path) {
14
+ try {
15
+ const parsed = JSON.parse(readFileSync(path, 'utf8'))
16
+ return parsed && typeof parsed === 'object' && !Array.isArray(parsed) ? parsed : {}
17
+ } catch {
18
+ return {}
19
+ }
20
+ }
21
+
22
+ /**
23
+ * Atomic file write: tmp sibling + rename, so a crash mid-write never leaves
24
+ * a truncated file behind. Truncation is worse than a clean failure here —
25
+ * a truncated settings.yaml often stays VALID yaml and silently drops config,
26
+ * and a truncated JSON token store reads back as {} (tokens/keys gone).
27
+ * Every write lands a fresh inode, so `mode` reliably applies (credentials
28
+ * and token stores pass 0600 — same discipline as the dsh credentials file).
29
+ */
30
+ export function writeTextAtomic(path, text, mode) {
31
+ mkdirSync(dirname(path), { recursive: true })
32
+ const tmp = `${path}.${process.pid}.tmp`
33
+ try {
34
+ writeFileSync(tmp, text, mode ? { mode } : undefined)
35
+ renameSync(tmp, path)
36
+ } catch (err) {
37
+ try {
38
+ unlinkSync(tmp)
39
+ } catch {
40
+ // tmp never created (e.g. mkdir failed) — nothing to sweep
41
+ }
42
+ throw err
43
+ }
44
+ }
45
+
46
+ /**
47
+ * Write a JSON object file (pretty, trailing newline) atomically.
48
+ * These files carry credentials/tokens, so they are written 0600.
49
+ */
50
+ export function writeJson(path, value) {
51
+ writeTextAtomic(path, JSON.stringify(value, null, 2) + '\n', 0o600)
52
+ }
53
+
54
+ /**
55
+ * Resolve a named secret: process environment first, then a flat YAML-ish
56
+ * credentials file (`NAME: value` per line — the ~/.dsh/.credentials.yaml
57
+ * shape dsh uses). Returns null when neither source has it.
58
+ *
59
+ * The file is scanned with string operations, NOT a RegExp built from
60
+ * `envName`: the name is settings-configurable free text, and interpolating
61
+ * it into a pattern would let a crafted name (regex metachar payloads)
62
+ * rewrite what the match accepts. Line-wise parsing keeps the exact
63
+ * previous semantics: a line whose pre-colon part equals envName, with an
64
+ * optional single quote stripped from either end of the value.
65
+ *
66
+ * @param {string|undefined} envName settings-configured reference name
67
+ * @param {string} credFilePath absolute path of the credentials file
68
+ */
69
+ export function resolveEnvKey(envName, credFilePath) {
70
+ if (envName && process.env[envName]) return process.env[envName]
71
+ if (envName && existsSync(credFilePath)) {
72
+ for (const line of readFileSync(credFilePath, 'utf8').split(/\r?\n/)) {
73
+ const trimmed = line.trim()
74
+ const colon = trimmed.indexOf(':')
75
+ if (colon <= 0) continue
76
+ // No trim on the name part: the old pattern required envName
77
+ // immediately followed by ':'.
78
+ if (trimmed.slice(0, colon) !== envName) continue
79
+ let value = trimmed.slice(colon + 1).trim()
80
+ // Same optional single-quote stripping the old pattern accepted
81
+ // (asymmetric quotes included), then the same character class:
82
+ // no whitespace, no quotes inside the value.
83
+ if (value.length >= 2 && (value[0] === '"' || value[0] === "'")) value = value.slice(1)
84
+ if (value.length >= 1 && (value.endsWith('"') || value.endsWith("'"))) {
85
+ value = value.slice(0, -1)
86
+ }
87
+ if (value && !/[\s"']/.test(value)) return value
88
+ }
89
+ }
90
+ return null
91
+ }
@@ -0,0 +1,108 @@
1
+ /**
2
+ * core/rotation.js — provider-agnostic multi-credential rotation engine.
3
+ *
4
+ * Policy: requests round-robin across the key list; a key that answers a
5
+ * failover status (401/403/429/5xx) or drops the connection is
6
+ * cooled for cooldownMs, then rejoins on its own. Cooling keys are appended
7
+ * as last resort when every fresh key is unavailable.
8
+ *
9
+ * INSTANCE state (cooldown map + round-robin cursor) lives on the KeyRotator
10
+ * object, NOT on the module — the composition root creates one instance per
11
+ * plugin module instance, so test harnesses that import the plugin twice
12
+ * (scripts/verify-rotation.mjs ?case=provider / ?case=bridge) keep the same
13
+ * isolation the old module-global state provided.
14
+ *
15
+ * Extracted from index.js during the core/providers split. This file knows
16
+ * nothing about any specific upstream: the "no credential" error for the
17
+ * empty-candidate case is injected by the caller (`emptyError`).
18
+ */
19
+
20
+ /** Statuses that fail a request over to the next key. */
21
+ function isFailoverStatus(status) {
22
+ return status === 401 || status === 403 || status === 429 || status >= 500
23
+ }
24
+
25
+ export class KeyRotator {
26
+ constructor() {
27
+ /** keyName → cooldown-until epoch ms. */
28
+ this.cooldowns = new Map()
29
+ /** Round-robin cursor (advanced once per ordered() call). */
30
+ this.cursor = 0
31
+ }
32
+
33
+ markCooling(name, cooldownMs) {
34
+ this.cooldowns.set(name, Date.now() + cooldownMs)
35
+ }
36
+
37
+ /**
38
+ * Order key entries for one outbound call: every key exactly once,
39
+ * non-cooling first in round-robin order, cooling keys appended as last
40
+ * resort. 0/1-key cases are handled by the caller (no rotation state is
41
+ * touched for them — the cursor must not advance).
42
+ *
43
+ * @param {Array<{name: string, key: string}>} keys
44
+ * @returns {Array<{authorization: string, headers: {}, keyName: string}>}
45
+ */
46
+ ordered(keys) {
47
+ const now = Date.now()
48
+ const n = keys.length
49
+ const start = this.cursor % n
50
+ this.cursor = (this.cursor + 1) % n
51
+ const rotated = Array.from({ length: n }, (_, i) => keys[(start + i) % n])
52
+ const fresh = rotated.filter((k) => (this.cooldowns.get(k.name) ?? 0) <= now)
53
+ const cooling = rotated.filter((k) => (this.cooldowns.get(k.name) ?? 0) > now)
54
+ return [...fresh, ...cooling].map((k) => ({
55
+ authorization: `Bearer ${k.key}`,
56
+ headers: {},
57
+ keyName: k.name,
58
+ }))
59
+ }
60
+
61
+ /**
62
+ * Run `attempt(cred)` over the candidates with failover: a candidate that
63
+ * throws a network error or answers a failover status is cooled and the
64
+ * next candidate takes over. The LAST candidate's response/error is
65
+ * returned as-is (no retry). `attempt(cred)` must return the fetch
66
+ * Response (body unconsumed on error statuses — it is cancelled before
67
+ * failing over) or throw.
68
+ *
69
+ * Resolves { cred, res, err }: exactly one of res/err is set.
70
+ *
71
+ * @param {Array<{authorization: string, headers: Record<string,string>, keyName: string|null}>} candidates
72
+ * @param {(cred: object) => Promise<Response>} attempt
73
+ * @param {{ cooldownMs: number, emptyError: Error }} options
74
+ */
75
+ async run(candidates, attempt, { cooldownMs, emptyError }) {
76
+ if (candidates.length === 0) return { cred: null, res: null, err: emptyError }
77
+ let lastErr = null
78
+ for (let i = 0; i < candidates.length; i++) {
79
+ const cred = candidates[i]
80
+ const isLast = i === candidates.length - 1
81
+ let res
82
+ try {
83
+ res = await attempt(cred)
84
+ } catch (err) {
85
+ // Caller-side aborts are not the key's fault: no cooldown, no failover.
86
+ // The bridge marks a client hangup with err.clientDisconnected; an
87
+ // upstream first-byte timeout (peer connects but never sends headers)
88
+ // is also an AbortError but is NOT caller-side — it must cool the key
89
+ // and fail over, or every key burns its timeout budget in sequence
90
+ // against a wedged edge. Only clientDisconnected aborts are exempt.
91
+ if (err?.name === 'AbortError' && err?.clientDisconnected === true) return { cred, res: null, err }
92
+ // Network-layer failure: cool the key and fail over (unless last).
93
+ if (cred.keyName) this.markCooling(cred.keyName, cooldownMs)
94
+ lastErr = err
95
+ if (isLast) return { cred, res: null, err }
96
+ continue
97
+ }
98
+ if (!isLast && isFailoverStatus(res.status)) {
99
+ if (cred.keyName) this.markCooling(cred.keyName, cooldownMs)
100
+ await res.body?.cancel().catch(() => {})
101
+ lastErr = null
102
+ continue
103
+ }
104
+ return { cred, res, err: null }
105
+ }
106
+ return { cred: null, res: null, err: lastErr }
107
+ }
108
+ }
@@ -0,0 +1,176 @@
1
+ /**
2
+ * core/usage-meter.js — provider-agnostic usage metering store.
3
+ *
4
+ * The upstream reports per-request billing on every chat SSE stream; the
5
+ * bridge scans for it ALWAYS and accumulates into a JSON file so a settings
6
+ * card can show live consumption. Tokens/credits only — never message
7
+ * content. Metering must never break the data path: every failure is
8
+ * swallowed, writes debounced.
9
+ *
10
+ * The recorded usage contract is the OpenAI-style SSE usage object plus an
11
+ * optional provider `credit` cost field:
12
+ * { prompt_tokens, prompt_cache_hit_tokens, prompt_cache_miss_tokens,
13
+ * completion_tokens, credit }
14
+ * Providers whose usage objects lack a field simply record 0 for it.
15
+ *
16
+ * Extracted from index.js during the core/providers split. Instance state
17
+ * lives on the meter object (created once per plugin module instance by the
18
+ * composition root) — not on this module.
19
+ */
20
+
21
+ import { readJson, writeJson } from './json-store.js'
22
+
23
+ const USAGE_RECENT_CAP = 100
24
+ const USAGE_DAYS_CAP = 31
25
+ const USAGE_FLUSH_MS = 5000
26
+ /** Two requests further apart than this belong to different (approximate) turns. */
27
+ const TURN_GAP_MS = 45_000
28
+
29
+ /** Local-day key (YYYY-MM-DD) — the display groups by the user's day. */
30
+ function dayKey(ts) {
31
+ const d = new Date(ts)
32
+ const mm = String(d.getMonth() + 1).padStart(2, '0')
33
+ const dd = String(d.getDate()).padStart(2, '0')
34
+ return `${d.getFullYear()}-${mm}-${dd}`
35
+ }
36
+
37
+ function normalizeUsageStore(raw) {
38
+ const store = raw && typeof raw === 'object' && !Array.isArray(raw) ? raw : {}
39
+ store.since = typeof store.since === 'number' ? store.since : Date.now()
40
+ store.totalCredit = typeof store.totalCredit === 'number' ? store.totalCredit : 0
41
+ store.totalRequests = typeof store.totalRequests === 'number' ? store.totalRequests : 0
42
+ store.days = store.days && typeof store.days === 'object' && !Array.isArray(store.days) ? store.days : {}
43
+ store.recent = Array.isArray(store.recent) ? store.recent.slice(-USAGE_RECENT_CAP) : []
44
+ return store
45
+ }
46
+
47
+ const roundCredit = (n) => Math.round(n * 10000) / 10000
48
+
49
+ /**
50
+ * Group recent requests into approximate turns: entries closer than
51
+ * TURN_GAP_MS merge into one row (a title call lands in the turn that
52
+ * triggered it; tool-loop steps are seconds apart by construction). The
53
+ * bridge cannot see the host's turn boundaries (no session ids on the
54
+ * wire), so this is a disclosed approximation.
55
+ */
56
+ function groupTurns(recent) {
57
+ const turns = []
58
+ for (const r of recent) {
59
+ const last = turns[turns.length - 1]
60
+ if (last && r.ts - last.end <= TURN_GAP_MS) {
61
+ last.end = r.ts
62
+ last.requests += 1
63
+ last.credit = roundCredit(last.credit + r.credit)
64
+ last.prompt += r.prompt
65
+ last.hit += r.hit
66
+ last.miss += r.miss
67
+ last.completion += r.completion ?? 0
68
+ if (r.model && !last.models.includes(r.model)) last.models.push(r.model)
69
+ if (!last.kinds.includes(r.kind)) last.kinds.push(r.kind)
70
+ } else {
71
+ turns.push({
72
+ start: r.ts,
73
+ end: r.ts,
74
+ requests: 1,
75
+ credit: r.credit,
76
+ prompt: r.prompt,
77
+ hit: r.hit,
78
+ miss: r.miss,
79
+ completion: r.completion ?? 0,
80
+ models: r.model ? [r.model] : [],
81
+ kinds: [r.kind],
82
+ })
83
+ }
84
+ }
85
+ return turns
86
+ }
87
+
88
+ /**
89
+ * @param {{ path: string }} options persistence target
90
+ * @returns {{ record: Function, view: Function, dispose: Function }}
91
+ */
92
+ export function createUsageMeter({ path }) {
93
+ const store = normalizeUsageStore(readJson(path))
94
+ let flushTimer = null
95
+
96
+ function flush() {
97
+ try {
98
+ writeJson(path, store)
99
+ } catch {
100
+ // persistence is best-effort
101
+ }
102
+ }
103
+
104
+ /** Debounced: a tool-loop burst produces one write, not one per request. */
105
+ function scheduleFlush() {
106
+ if (flushTimer) return
107
+ flushTimer = setTimeout(() => {
108
+ flushTimer = null
109
+ flush()
110
+ }, USAGE_FLUSH_MS)
111
+ flushTimer.unref?.()
112
+ }
113
+
114
+ /**
115
+ * Record one billed request. `usage` is the upstream's SSE usage object;
116
+ * only requests that actually produced one are recorded (failed/error
117
+ * responses carry no billing signal).
118
+ * kind: chat | title | compaction | image | search | fetch.
119
+ */
120
+ function record({ ts, kind, model, usage }) {
121
+ if (!usage || typeof usage !== 'object') return
122
+ try {
123
+ const credit = typeof usage.credit === 'number' && Number.isFinite(usage.credit) ? usage.credit : 0
124
+ const entry = {
125
+ ts,
126
+ kind,
127
+ model: typeof model === 'string' ? model : null,
128
+ prompt: usage.prompt_tokens ?? 0,
129
+ hit: usage.prompt_cache_hit_tokens ?? 0,
130
+ miss: usage.prompt_cache_miss_tokens ?? 0,
131
+ completion: usage.completion_tokens ?? 0,
132
+ credit,
133
+ }
134
+ store.totalCredit = roundCredit(store.totalCredit + credit)
135
+ store.totalRequests += 1
136
+ const day = dayKey(ts)
137
+ const bucket = store.days[day] ?? { credit: 0, requests: 0 }
138
+ bucket.credit = roundCredit(bucket.credit + credit)
139
+ bucket.requests += 1
140
+ store.days[day] = bucket
141
+ const dayKeys = Object.keys(store.days).sort()
142
+ while (dayKeys.length > USAGE_DAYS_CAP) delete store.days[dayKeys.shift()]
143
+ store.recent.push(entry)
144
+ if (store.recent.length > USAGE_RECENT_CAP) {
145
+ store.recent = store.recent.slice(-USAGE_RECENT_CAP)
146
+ }
147
+ scheduleFlush()
148
+ } catch {
149
+ // metering is best-effort
150
+ }
151
+ }
152
+
153
+ /** The usage-view payload: totals + approximate turns + exact recent rows. */
154
+ function view() {
155
+ const today = store.days[dayKey(Date.now())] ?? { credit: 0, requests: 0 }
156
+ return {
157
+ since: store.since,
158
+ totalCredit: store.totalCredit,
159
+ totalRequests: store.totalRequests,
160
+ today: { day: dayKey(Date.now()), credit: today.credit, requests: today.requests },
161
+ recent: store.recent.slice(-20).reverse(),
162
+ turns: groupTurns(store.recent).slice(-10).reverse(),
163
+ turnGapMs: TURN_GAP_MS,
164
+ }
165
+ }
166
+
167
+ function dispose() {
168
+ if (flushTimer) {
169
+ clearTimeout(flushTimer)
170
+ flushTimer = null
171
+ }
172
+ flush()
173
+ }
174
+
175
+ return { record, view, dispose }
176
+ }
@@ -0,0 +1,181 @@
1
+ # 诊断报告:后端重复提问与缓存命中/额度消耗
2
+
3
+ 日期:2026-08-17 | 环境:dsh 0.1.0-rc.6(web profile)+ dsh-codebuddy-plugin 0.6.0 | 网关:copilot.tencent.com
4
+ 方法:桥内新增取证日志(`CODEBUDDY_BRIDGE_LOG`,本文 §1.2)+ 受控流量(`scripts/capture-traffic.mjs`)+ 网关对照探测(`scripts/probe-cache.mjs`)+ 83 个历史会话日志挖掘 + dsh rc.3↔rc.6 代码对比。
5
+ 预算遵守:真实网关调用共 48 次(上限 50),探测 `max_tokens=16`,主聊天回答按 prompt 约束为 ≤20 token。
6
+
7
+ **原始数据**:`/tmp/diag/bridge.jsonl`(桥日志)、`/tmp/diag/probe-cache*.jsonl`(探测)、`~/.dsh/sessions/**/session.jsonl.zstd`(会话事件)。
8
+
9
+ ## TL;DR
10
+
11
+ 1. **"同一提问多次出现在后端"的主因不是 bug,是 agent 架构**:每个工具循环 step 都重发全量历史(历史实测 3275 steps / 540 turns ≈ **6.1 次/turn**);其上是标题生成(每会话 1 次、内嵌首问全文)与子代理派生(独立会话、全量历史)。dsh-llm-retry 重试机制存在且请求体逐字节相同,但 codebuddy 路径 83 会话 0 实例。
12
+ 2. **网关缓存是按内容寻址的自动前缀缓存,与会话头/prompt_cache_key 无关**——对照实验全等价。**但按模型分策略:deepseek-v4-pro 缓存工作,deepseek-v3 在 ≤16.3k tokens 全部 0 命中**。插件默认模型是 deepseek-v3——"缓存命中率下降"的首要解释。
13
+ 3. **额度(credit)实测**:v4-pro 未命中 ≈0.26 credit/1k tokens,命中 ≈0.011(约 1/24);v3 ≈0.03/1k 且无缓存折扣。一轮 15.8k token 的 turn:全 miss 4.08 credit vs 全命中 0.17。
14
+ 4. **v0.7"请求指纹→稳定会话映射"对缓存恢复无收益**(缓存本就不靠会话亲和),不建议以此为目标立项;杠杆在模型选择。
15
+
16
+ ---
17
+
18
+ ## 1. 重复提问甄别
19
+
20
+ ### 1.1 全局事实:dsh 对 LLM 的出站调用点只有 3 处
21
+
22
+ 全量 grep `llm.stream(` 证实(/usr/local/lib/node_modules/@deepseek-ai/dsh/node_modules/ 下):
23
+
24
+ - `dsh-agent-loop/lib/index.js:616` — 主循环 step(含重试重发)
25
+ - `dsh-session-title-llm/lib/index.js:228` — 会话标题
26
+ - `dsh-compaction-basic/lib/index.js:301` — 压缩摘要
27
+
28
+ 主循环每个 step 调 `session.deriveMessages()` 重发**全量历史**(`dsh-agent-loop/lib/index.js:613`;投影规则 `dsh-session/lib/index.js:278-287`)。历史 83 会话实测:3275 step / 540 turn ≈ **6.1 次请求/用户 turn**——一次提问触发 N 个工具轮次,就有 N+1 次含该提问全文的请求。这是"后端多次相同提问"的最大来源,属架构行为而非异常。
29
+
30
+ ### 1.2 取证手段:桥内请求日志(本次新增)
31
+
32
+ `index.js` 的桥在 `CODEBUDDY_BRIDGE_LOG=<path>` 时每请求落一条 JSONL:入站(头特征、body/消息/系统提示哈希、末条用户消息预览 60 字符、marker 分类)+ 出站(状态、排队/首字节/总耗时、上游 usage 含缓存计数)。不落消息明文;Authorization 只记 `sentinel`/`caller-set` 分类。回归锁在 `scripts/verify-bridge.mjs` 第 7 节。
33
+
34
+ marker 分类规则(对照 dsh 源码验证):系统提示以 `Create a concise title for an AI coding-assistant session` 开头 → `session-title`(dsh-session-title-llm/lib/index.js:147-154);末条 user 以 `You are now acting as a compaction engine` 开头 → `compaction`(dsh-compaction-basic/lib/index.js:218-253)。
35
+
36
+ ### 1.3 受控流量捕获(/tmp/diag/bridge.jsonl)
37
+
38
+ `scripts/capture-traffic.mjs` 经 dsh web RPC 驱动:会话 A 三轮短问答、会话 B 逐字重发 A 的首问、会话 C 要求派生子代理。10 个 chat 请求全记录(模型 deepseek-v4-pro,用户此前在设置页选定的模型):
39
+
40
+ | seq | 间隔 | msgs | bytes | bodySha | marker | prompt_tokens | hit | miss | credit |
41
+ |-----|------|------|-------|---------|--------|-------|------|-------|--------|
42
+ | 1 | — | 2 | 665 | 7bb6f7b9 | session-title | 113 | 0 | 113 | 0.03 |
43
+ | 2 | +26ms | 5 | 64356 | 2da804c8 | — | 15807 | 0 | 15807 | 4.07 |
44
+ | 3 | +6.3s | 7 | 64432 | 1b02fc16 | — | 15820 | 4864 | 10956 | 2.88 |
45
+ | 4 | +3.8s | 9 | 64589 | 1bab4e0f | — | 15831 | 15744 | 87 | 0.16 |
46
+ | 5 | +3.0s | 5 | 64356 | **2da804c8** | — | 15807 | 15744 | 63 | 0.18 |
47
+ | 6 | +2ms | 2 | 665 | **7bb6f7b9** | session-title | 113 | 0 | 113 | 0.03 |
48
+ | 7 | +5.7s | 2 | 711 | e73a0bcb | session-title | 135 | 0 | 135 | 0.05 |
49
+ | 8 | +1ms | 5 | 64398 | 0e25a15d | — | 15829 | 4352 | 11477 | 3.19 |
50
+ | 9 | +6.1s | 5 | 64763 | 95b38492 | — | 15818 | **0** | 15818 | 4.08 |
51
+ | 10 | +3.0s | 7 | 65347 | 5dbfd366 | — | 15972 | 15744 | 228 | 0.27 |
52
+
53
+ 注:seq1 比主请求 seq2 早 26ms 到达——标题调用在主请求路由确定后立即发出(dsh-session-title/lib/index.js:316-352)。
54
+
55
+ ### 1.4 六类候选来源逐项结论
56
+
57
+ | 候选 | 结论 | 证据 |
58
+ |------|------|------|
59
+ | **dsh-llm-retry 重试** | 机制确认存在;**codebuddy 路径排除为观察期内来源** | 机制:默认 maxRetries=2、可重试码 EMPTY_RESPONSE/RATE_LIMIT/SERVER/TIMEOUT/TRANSPORT、退避 500ms 起(dsh-llm/lib/index.js:356-366);重试重建请求体**逐字节相同、线上无标记**(agent-loop 回到 buildRequest,失败 chunk 不投影进消息,dsh-session/lib/index.js:278-287;SDK 内重试 maxRetries=0,dsh-llm-pi-ai/lib/index.js:666)。实例:历史 83 会话仅 1 起 2 次,**在 qianwenai provider**(TRANSPORT "stream ended before terminal event",session-5f800970 seq52-57);codebuddy 路径 0 起;本次捕获 10 请求无任何相同 bodySha 的亚秒级重复 |
60
+ | **会话标题生成** | **确认** | 桥日志 seq1/6/7 marker=session-title:system+user 形状与 dsh-session-title-llm/lib/index.js:147-158 完全一致,maxTokens=64,模型同主聊天;每会话首条用户消息后 1 次(first-prompt 提供者)。**用户首问全文内嵌标题 prompt**(≤4096B),即在网关出现第二次。历史 75/83 会话有 title-llm-request。seq6 与 seq1 逐字节相同(跨会话同首问 → 同标题请求)。单次 0.03-0.05 credit |
61
+ | **压缩摘要** | 机制确认存在;低频,非当前重复主因 | 触发:≥80% contextWindow 自动(thresholdRatio,dsh-compaction-basic/lib/index.js:13)、CONTEXT_WINDOW_EXCEEDED 溢出、手动 /compact。请求=全量历史+压缩指令,每次压缩 1 次 LLM 调用(pre-step 路径最多 2 次,:71)。历史 2 起:手动 /compact 成功(10.4s);自动触发一次因 glm-4.6 返回 400 失败(session-21f564ae)。本次捕获未触发(上下文远未及阈值) |
62
+ | **手动重发** | **确认** | 会话 B 首问与 A 首问逐字节相同(seq5 bodySha==seq2 的 2da804c8,间隔 3s);第二次 99.6% 缓存命中,credit 4.07→0.18。同一文本的跨会话重发在网关侧就是一次完整新请求 |
63
+ | **子代理调用** | **确认** | 会话 C 经 subagent 工具派生子会话 83dd6d24(session 事件:`origin:"subagent"`、`parentSession:session-4c2e6533…`、delegationDepth:1)。桥日志 seq9=子会话首请求(5 msgs、与父**同系统提示** sysSha 879617fd、但 0 命中——15.8k tokens 全量重新计费 4.08 credit);seq10=父会话 step2。子代理继承父 provider/model 但用独立会话(dsh-subagent/lib/index.js:779-792),其每个 step 同样全量重发 |
64
+ | **其他** | 见下 | ① **工具循环步**(最大放大器,见 §1.1,6.1 次/turn);② goal 轮驱动/定时任务注入新 turn(dsh-goal-round-driver/lib/index.js:11-18、dsh-schedule/lib/index.js:825-835):机制确认,历史 27 起 goal/change,本次未触发;③ dsh-time-context 每 step 追加时间戳消息:**默认 profile 未加载**(dsh-base/dsh-web-app/dsh-headless patch 均无),排除;④ dsh-repeat-tool-reminder 只改后续请求内容、不产生额外请求,排除 |
65
+
66
+ **附带风险提示**:默认可重试码含 `EMPTY_RESPONSE`——网关若返回 200 但空内容(如 content_filter 收尾),dsh 会静默重发同一请求体最多 2 次,桥上表现为"一模一样又问了一遍"。本次未观察到实例,列为隐患。
67
+
68
+ ## 2. 缓存机制实测(scripts/probe-cache.mjs)
69
+
70
+ ### 2.1 缓存信号可观测性
71
+
72
+ 网关 SSE 每 chunk 带 usage,末chunk含完整字段:`prompt_cache_hit_tokens` / `prompt_cache_miss_tokens` / `prompt_tokens_details.cached_tokens` / `cache_read_input_tokens` / `cache_creation_input_tokens` / `prompt_cache_write_tokens` / `credit`(单次计费额度)。实测有效的命中字段是 `prompt_cache_hit_tokens`(= `prompt_tokens_details.cached_tokens`;Anthropic 方言字段恒 0)。**缓存行为可观测,无需依赖延迟推断。**
73
+
74
+ ### 2.2 对照实验(同 payload 连发,各 arm 独立 nonce 防串扰)
75
+
76
+ | 实验 | 模型 | prompt tokens | 结果 |
77
+ |------|------|------|------|
78
+ | run1:anon/session/cachekey/session+cachekey ×3 | deepseek-v3 | ~883-958 | **12 次全 0 命中** |
79
+ | run2:同上 | deepseek-v3 | ~2611-2830 | **12 次全 0 命中** |
80
+ | run4:anon ×2 | deepseek-v3 | 16300 | **0 命中**(#2 仍 miss 16300) |
81
+ | run3:anon ×3、session ×3 | deepseek-v4-pro | 2684 | 两臂 #1 miss、#2/#3 均 hit 2560,**完全相同** |
82
+ | run5:anon ×2、session ×2 | deepseek-v4-pro | 15859 | 两臂 #2 均 hit 15744、credit 均 4.08→0.17,**完全相同** |
83
+
84
+ 结论:
85
+
86
+ - **C1 缓存按内容寻址、自动生效**,不需要会话亲和头(openai 三头)也不需要 `prompt_cache_key`;两者对命中**零影响**(2.6k 与 15.8k 两档、anon/session 完全等价)。
87
+ - **C2 缓存策略按模型分**:deepseek-v4-pro 缓存工作(阈值 ≤2684 tokens,粒度 128 tokens——命中值皆为 128 倍数);**deepseek-v3 在 883/2611/16300 tokens 三档全部 0 命中**——v3 路径上缓存对该账号不可用。插件默认模型 `agent-default-model` 恰为 deepseek-v3(cordis.patch.yml:266-269)。
88
+ - **C3 命中可用性在 15.8k 规模存在网关内部波动**:真实流量样本命中 0/4864/15744/15744/4352/0/15744(seq2-10),非单调;而受控连发(run5)第二次即稳定全命中——波动与并发/后端调度有关,**亲和头无法消除**(run5 对照)。
89
+ - 延迟信号与之一致但噪声大(v4-pro 15.8k:miss ttfb 2653ms vs hit 2453ms;2.6k 档 2313→1903ms),仅作旁证,不作结论依据。
90
+
91
+ ### 2.3 dsh 主聊天出站画像(桥日志实证,seq2)
92
+
93
+ - 头:`user-agent: deepseek-harness/0.1.0-rc.6 (+…)`(dsh attribution 强制覆盖静态 UA)、`x-ide-*/x-product-version/x-requested-with/x-private-data`(patch 静态头)、`authorization: 哨兵`(桥替换为真实凭据)。**无任何会话亲和头(sessionIn=null)、无 prompt_cache_key**——dsh 内部虽每步设 `sessionId: session.id`(dsh-agent-loop/lib/index.js:734),但 pi-ai 的亲和开关被 dsh-llm-pi-ai 的 compat 白名单剥除(白名单仅 thinkingFormat/supportsReasoningEffort,dsh-llm-pi-ai/lib/index.js:1334-1337),永不上线。
94
+ - 体:`stream:true` + `stream_options.include_usage`、全量 messages、40 个工具、reasoning_effort=max、max_tokens=50000。系统提示 17589B(≈4.3k tokens)+ 工具为主体;**系统提示哈希跨会话/子代理恒定**(879617fd),前缀结构稳定。
95
+ - 放大系数:每 turn ≈ 6.1 次全量请求(§1.1)。
96
+
97
+ ## 3. 额度消耗归因
98
+
99
+ ### 3.1 credit 单价(实测拟合)
100
+
101
+ | 模型 | 未命中 | 命中 | 证据 |
102
+ |------|--------|------|------|
103
+ | deepseek-v4-pro | ≈0.26 credit/1k tok | ≈0.01 credit/1k tok(两组拟合 1/24–1/29) | 15859 tok 全 miss=4.08;15744 hit+115 miss=0.17(run5);2684 档 0.70→0.06(run3);15807 全 miss=4.07(seq2) |
104
+ | deepseek-v3 | ≈0.03 credit/1k tok | —(无缓存可用) | 883→0.03、2611→0.08、16300→0.47(run1/2/4) |
105
+
106
+ v4-pro 单价约为 v3 的 9 倍,但缓存命中部分再打 ~96% 折扣;v3 便宜但每轮全价。
107
+
108
+ ### 3.2 全量上下文重发与额度页的关系
109
+
110
+ 额度页消耗 = Σ credit(逐请求)。一轮 turn 的输入 ≈ 系统提示+工具(本机 ≈15.7k tokens)+ 累积历史;**每个 step 全量重发**(§1.1)。因此:
111
+
112
+ - 无缓存(v3,或命中失效时):10-turn 会话输入成本随历史近似线性→累计近似二次增长。本机规模:每 turn ≈ 15.8k×6.1 steps×0.26/1k ≈ **25 credit/turn**(v4-pro 全 miss)。
113
+ - 稳命中(v4-pro):同规模 hit 部分按 ~1/24 计价,seq4 实测整轮仅 0.16 credit(vs 全 miss 4.07)。
114
+ - dsh 侧记账与网关一致:session.list 投影 `tokenUsage` 累计值与桥日志逐数吻合(26763=15807+10956;20608=4864+15744),**dsh 无额外放大**。
115
+ - 标题/压缩/子代理为附加项:标题每会话 1 次约 0.03-0.05 credit;子代理首请求全量 miss(本次 4.08)。
116
+
117
+ ## 4. 旧版 dsh 对比(rc.3 ↔ rc.6)
118
+
119
+ npm 拉取 @deepseek-ai/*@0.1.0-rc.3 逐包对比:请求路径上 6 个关键包中 5 个逐字节相同(dsh-agent-loop、dsh-llm-retry、dsh-time-context、dsh-session-title-llm、dsh-compaction-basic);唯一差异在 dsh-llm-pi-ai:rc.6 对带思考开关的模型强制 `supportsDeveloperRole:false`(rc.6 lib/index.js:1080)——思考模型的系统提示角色由 rc.3 自动检测的 `developer` 改回 `system`。该差异一次性、版本内恒定,**不构成持续的缓存 miss 机制**。亲和头/prompt_cache_key/重试参数/压缩阈值/标题形状两版全同。**结论:排除"dsh 升级导致缓存命中率下降"**(逐包 diff,/tmp/dsh-old/)。
120
+
121
+ ## 5. v0.7 可行性结论:桥内"请求指纹→稳定会话映射"
122
+
123
+ - **对缓存命中:收益 ≈ 0,不建议以此立项。** 缓存按内容寻址(C1),会话亲和头对命中零影响(run3/run5 对照,含 15.8k dsh 真实规模)。指纹映射无法恢复"本就不缺的东西";观察到的命中波动(C3)是网关内部行为,亲和头不能消除。
124
+ - **输入可缓存占比**(若模型有缓存):本机实测连续 turn 命中 15744/15807 = **99.6%**;下限为系统提示+工具 ≈4.3k+ tokens(命中波动时观察到的部分命中 4352/4864 即该量级)。可缓存性瓶颈不在 dsh 前缀稳定性(sysSha 恒定),而在**模型是否享受缓存策略**。
125
+ - 真正的杠杆(按收益排序):① 默认/推荐模型切换到缓存可用的模型(v4-pro 实测命中折扣 ~24x;v3 无任何折扣);② 命中波动属网关侧,插件不可控,但桥日志现已可长期观测命中率(`CODEBUDDY_BRIDGE_LOG`),可据此向网关方反馈;③ 指纹映射仅对 SessionLimiter 并发公平/网关会话分析有残余价值,不值得单独版本。
126
+
127
+ ## 6. 不可观测边界(诚实声明)
128
+
129
+ - 网关**为什么** v3 不缓存、命中波动的后端原因(分片/复制/淘汰策略)不可从客户端观测;本文只断言可观测行为(usage 字段 + 对照实验)。
130
+ - 额度页本身的口径(credit 与页面数值的映射、日结/周期)未验证;本文 credit 全部取自响应 usage 字段实测。
131
+ - 标题/压缩请求是否享受缓存对总额度影响可忽略(≤135 tokens/次),未单独探测。
132
+ - 子代理首请求 0 命中观察到 1 次(seq9);系统提示与父相同,miss 原因(新会话路由到冷后端 vs 其他)不可进一步区分,但其计费后果已实测。
133
+
134
+ ## 7. 补遗(2026-08-18):大上下文真实会话命中率劣化归因("36%" 事件)
135
+
136
+ 用户报告:同模型 v4-flash、high、极简模式,插件路径 25 步聚合命中 36%,DeepSeek 原生路径 90%+,且"没有做会话管理之前命中合理"。定位与结论如下。
137
+
138
+ ### 7.1 用户会话复核(证据:docs/probes/cache-session-f52c817f-2026-08-18.json)
139
+
140
+ - 会话 f52c817f:codebuddy / deepseek-v4-flash / high,web 研究任务(web_fetch 大页面),25 步,**平均上下文 ≈55.5k tokens**,dsh 记录口径聚合命中 **34.7%**——即用户所见 36%。
141
+ - UI 口径(dsh-client-ui-conversation/lib/client.js:2777-2788):`cacheHit% = ΣcacheRead / (Σuncached + ΣcacheRead + ΣcacheWrite)`,projection 来自 dsh-token-meter(uncachedInputTokens=usage.inputTokens)。两条路径记录口径一致,无插件侧记账差异。
142
+ - 逐步形态:cacheRead **吸附于历史快照尺寸**(7424 出现 3 次、32768=恰好 32k、34304、46336),coverage(cacheRead/上一步 prompt)在 21%→100%→20% 间非单调跳动;t1.13 coverage 100% 证明 46k 大前缀可写可读,下一步又回落到 12.4k 化石快照——**网关在 40k+ 规模对真实增长内容的缓存条目保留/驱逐不稳定**。
143
+
144
+ ### 7.2 对照实验(v4-flash high、**直连网关不经插件**,docs/probes/cache-flash-scale-2026-08-18.jsonl)
145
+
146
+ | 实验 | 条件 | 结果 |
147
+ |------|------|------|
148
+ | scale 连发 | 8k/16k/24k 同 prompt 各 4 发 | 复发 99.1–99.7%(阈值扫描另见 cache-flash-thresh:412 tok 起即缓存) |
149
+ | gap | 16k,间隔 20s×3 + 60s×1 | 全部 99.7%——TTL ≥60s 无影响 |
150
+ | grow | 20k→108k,+10k/轮,15s 间隔 | **coverage≈100%**,观测 hit% 与增长稀释上限逐点吻合(如 24011/35965=66.7% 实测 66.6%);末轮原样重发 100% |
151
+
152
+ → 网关**有能力**在 108k 规模全量缓存增长前缀;用户会话的劣化不由尺寸本身决定,取决于网关内部状态/内容/负载。同一现象此前已在 15.8k 档以直连探测记录(C3),本次把"波动"上限推进到 40k+ 真实内容。
153
+
154
+ ### 7.3 插件链路无责(逐项排除)
155
+
156
+ - 桥逐字节透明(2.6k 档直连==经桥,命中一致);usage 数字是网关 SSE 自报,链路无法改写。
157
+ - 会话管理(v0.5.5)对主聊天**不注入任何头**(dsh 出站无 session id,亲和开关被适配器剥除,§2.3);SessionLimiter 只排队不动内容。"没有会话管理时命中合理"的印象与版本无关——v0.5.5 前后请求内容零变化。
158
+ - 受控 A/B(极简 preset、同任务、小上下文 2–4.5k):插件 15 步 agg **92.3%** vs 原生 16 步 **97.3%**(dsh 记录口径)——插件路径在小中上下文健康。
159
+ - 用户的原生对照实为 **api.deepseek.com(另一套后端)且 avg ctx 仅 22k**,非同后端同规模;原生亦有 0 命中记录(85969bee,16k,[0,0])。
160
+
161
+ ### 7.4 实务建议
162
+
163
+ 命中率的主导变量是**后端 × 上下文规模**:≤5k 各缓存模型稳定 95%+;codebuddy 网关 40k+ 真实增长内容当前不可靠(v4-flash 实测 agg 35%);**v4-pro 在 99k 规模 215 步真实编码会话 agg 93.1%**(cbfaa68e)仍可用;glm-5.1/5.2 秒级失效、v3 恒 0。超大上下文且缓存敏感的场景优先 v4-pro 或 DeepSeek 官方 API。插件侧无可修项。
164
+
165
+ ## 8. 补遗更正(2026-09-03):§7 与 C3 的"网关侧不稳定"归因已被推翻
166
+
167
+ > 全量重判见 **docs/diagnosis-cache-decline.md**(证据链完整)。本节只记录对本报告结论的影响。
168
+
169
+ - §7.1/§7.4 与 §2.2 C3 的"网关内部波动 / 40k+ 真实内容条目保留不稳 / 吸附历史快照"**全部重判为桥缺陷**:`core/bridge.js` 用 `rawBody += c` 逐 TCP 分片隐式 utf8 解码,跨分片中文字符→3×U+FFFD 且损坏位置逐请求随机——出站前缀逐请求漂移,内容寻址缓存忠实地把损坏点当命中边界(hit 值与 dump 分叉字节数定量对应)。上述观察全部取自经桥流量,因此被污染;直连网关受控探测(24 发全 99.3%、16k/32k TTL ≥600s、增长前缀+空闲全命中)证明网关 v4-flash 缓存稳定。
170
+ - §7.2 的直连对照实验(合成内容 grow 至 108k ~100%)本身不经桥,结论不变;§7.3"桥逐字节透明(2.6k 档)"成立但**有边界**——单分片体透明,大体量中文体在分片切断多字节字符时不透明。
171
+ - §7.4"插件侧无可修项"**作废**:主修 = 桥 body 组装改 `Buffer.concat` 一次解码(见 diagnosis-cache-decline.md §6,含同型缺陷 3 处与回归用例建议)。
172
+ - §5"指纹映射无收益"结论不变(缓存按内容寻址的前提未被推翻,反而加强)。
173
+
174
+ ## 附:复现
175
+
176
+ ```sh
177
+ CODEBUDDY_BRIDGE_LOG=/tmp/diag/bridge.jsonl dsh web # 开取证日志
178
+ node scripts/capture-traffic.mjs # 受控流量(约 10 次小调用)
179
+ node scripts/probe-cache.mjs [--model M] [--arms anon,session,cachekey,session+cachekey] [--repeat N] [--calls K]
180
+ node scripts/verify-bridge.mjs # 离线回归(含日志断言)
181
+ ```
@@ -0,0 +1,67 @@
1
+ # 诊断:Qoder CN 的 Qwen3.8-Flash(`qfmodel`)为什么在 Qoder 里能用、在 dsh 里不能用
2
+
3
+ > 结论日期 2026-09-22(Windows 侧 dsh-tap)。**一句话**:不是 dsh/插件的请求形态问题——`qfmodel` 的上游后端节点 `oa_qwen-plus-main` 在 **04:21:38–04:21:51 之间进入持久失败**,此后**任何客户端**(包括 Qoder 官方客户端自己的签名器/元数据/请求体/版本号)都拿不到正文;"在 Qoder 里能用"是**04:21:38 之前的旧印象**。
4
+ >
5
+ > 复测(幂等,节点恢复即翻绿):`node scripts/probe-qoder-flash-confirm.mjs`
6
+
7
+ ## 1. 两条报错分别是什么
8
+
9
+ | 报错 | 归属 | 机制 |
10
+ |---|---|---|
11
+ | `{"code":"400","message":"[FAIL]node:oa_qwen-plus-main msg:Execution failed: null"}` | **上游侧**(Qoder 后端节点内部异常) | qfmodel 被路由到推理节点 `oa_qwen-plus-main`,该节点对任何请求抛 `Execution failed: null`(形似 Java 侧 NPE);HTTP 仍是 200,错误藏在 SSE 信封的 body 里("带内失败帧") |
12
+ | `{"code":"provider_error","message":"Error in upstream response","request_id":"…","details":"{…\"Messages with role 'tool' must be a response to a preceding message with 'tool_calls'…}"}` | **宿主侧真 bug,插件已修**(踩坑 #41 为主因、#39 为次因) | 严格家族(dmodel/kmodel/mmodel)的配对校验器把 **`content` 为 `null`/缺键的消息当"不存在"**:宿主 pi-ai 对**每个纯工具轮**都发 `{role:'assistant',content:null,tool_calls:[…]}`(openai-completions.js:961),声明因此蒸发、其后的 `role:"tool"` 被判孤儿 → 400。次因:pi-ai 还会丢弃 `stopReason=error/aborted` 的 assistant 却保留其 tool 结果(#39),产出真孤儿。首版修复补的桩自己也是 `content:null`,所以"修完仍报同一条错";0.9.9 起网关出站做可见性归一(null/缺键 → `''`)+ developer→system,实测四形态全部 400 → 200(证据 docs/probes/qoder-null-content-*.json) |
13
+
14
+ ## 2. qfmodel:为什么能断定是上游侧(逐变量排除)
15
+
16
+ **基准**:最小合法体 = `{model:"qfmodel", stream:true, stream_options:{include_usage:true}, messages:[{role:"user",content:"只回复两个字:收到"}]}`,打 `POST gateway.qoder.com.cn/algo/api/v2/service/pro/sse/agent_chat_generation`(COSY 签名 + 加密 body)。**同一账号、同一线缆路径**,同批请求里 `qmodel_38max`(Qwen3.8-Max)与 `q37fmodel`(Qwen3.7-Flash)稳定出正文。
17
+
18
+ | # | 假设 | 判定 | 判别证据 |
19
+ |---|---|---|---|
20
+ | 1 | 模型 key 猜错 | **否** | Qoder 客户端自己的日志打印 `model_config={"key":"qfmodel","display_name":"Qwen3.8-Flash",…,"source":"system"}`——客户端发的就是 `qfmodel`;两处 bundle 内**零硬编码模型 key**(key 全部来自服务端目录)。11 个臆造 key 全部被静默改派 auto(响应 `model:"auto"` + `billable:false`,踩坑 #37)→ 只有 `qfmodel` 会打到真实节点 |
21
+ | 2 | 请求体字段差异 | **否** | 复刻客户端 agent 形态(system 提示 + 3~26 个工具 schema + `tool_choice` + `max_tokens` + `temperature`,899B vs 对照 904B)仍失败;另扫 `enable_thinking`/`reasoning_effort`/`scene`+`agent_id`/`stream:false`/`context_window`(1M·200K·1048576·驼峰)/`thinking_config`/空 messages——全部同样失败 |
22
+ | 3 | 头/客户端元数据差异 | **否** | 用**官方 `QoderContext` 类**(wasm 原实现)签同一 body,扫 9 组 clientMetadata(含与客户端 fallback 逐字节一致的组合)——全部同样失败;`Cosy-ClientType` 无论传 `'qoder'`/`5`/缺省,wasm 恒出 **`5`**(已在插件侧对齐,见 §4) |
23
+ | 4 | 版本闸门 / 静默降级 | **否** | 客户端 bundle 里 `minimal_version`/`minimalVersion`/`999.999.999` **零命中**(客户端从不读该字段,目录里的 `minimal_version:{vsc:"999.999.999"}` 是死数据);把 `cosyVersion` 从 1.1.40 扫到 999.999.999,错误一字不变 → 服务端也没有版本闸门 |
24
+ | 5 | 节点本身坏了 | **是** | `oa_qwen-plus-main` **只被 qfmodel 指名**;同样的坏体(孤儿 tool)打到 `qmodel`/`qmodel_38max`/`q37fmodel` 分别回 70/52/162 个 chunk 的正文——Qwen 节点族活着,只有 Flash 的节点指派坏了 |
25
+ | + | 客户端有重试/降级 | **否** | 客户端 `contractMaxAttempts=3` 只覆盖传输层重试;三次都是 200,环内不换模型 |
26
+
27
+ **路由是头驱动的**(对将来排错有用):`X-Model-Key: qfmodel` + body `model:"auto"` → 仍失败;`X-Model-Key: auto` + body `model:"qfmodel"` → **成功**。即服务端按 `X-Model-Key`(= `prepareChat` 传的 `modelKey`)选节点,body 里的 `model` 字段不决定路由。插件正是用 `prepareChat(..., {modelKey: model})` 生成该头,行为与官方一致。
28
+
29
+ ## 3. 时间线("在 Qoder 里能用"是被推翻的前提)
30
+
31
+ | 时刻(+08) | 事件 | 证据 |
32
+ |---|---|---|
33
+ | 09-18 / 09-19 | Flash 在 Qoder 客户端**真实可用**:客户端 transcript 里累计 **754 条** `qfmodel` assistant 消息(完整 agent 环:thinking + tool_use + 正文) | `docs/probes/qoder-qfmodel-client-timeline.txt` |
34
+ | **04:21:38** | 客户端最后一次成功:assistant 正文 `"你好!有什么需要帮忙的尽管说。…当前分支是 v0.8.3…"`(`runtime-config model=qfmodel`) | `~/.qoder-cn/projects/…/f3ec00b4-…jsonl` |
35
+ | **04:21:51** | **同一客户端**的下一次运行:`model_config{"key":"qfmodel"}` → POST 同一端点 → `status=200 duration=3322ms` → `model.response.completed … output_tokens=0`、**不写 assistant 消息**(即同一条带内失败) | 原始日志:`~/.qoder-cn/logs/runs/2026-09-22T04-21-51-583+08-00-q2evm7-p26840/qodercli.log`(本机,含会话数据不入库);本次复核落盘 `docs/probes/qoder-qfmodel-final-confirm.json` |
36
+ | ~04:31 | dsh 侧首次复现同样报错(用户报障) | — |
37
+ | 04:33–05:12 | 本轮 13 次复测(差分矩阵 + 确认探针)**全部失败**,对照模型全绿 | `docs/probes/qoder-matrix-*.json`、`docs/probes/qoder-flash-confirm-*.json` |
38
+
39
+ → 节点在 04:21:38–04:21:51 之间坏掉并持续。**用户"Qoder 里能用"的印象来自坏掉之前的体验**;坏掉之后官方客户端同样不能用(客户端表现为"回合结束但没有回复",因为 HTTP 200,它不会弹错误)。
40
+
41
+ ## 4. 插件侧本轮实际修了什么(都不改变"Flash 属上游故障"的结论)
42
+
43
+ 1. **出站 tool 配对 + 可见性体检**(`providers/tool-pairing.js`,踩坑 #39 → #41):第一版只修"孤儿 tool",且在**容错家族 qmodel** 上验证 → 上线后同一条 `provider_error` 照旧。第二版(0.9.9)用单变量差分定出真因:严格家族的校验器把 `content:null`/缺键的消息**当不存在**,而宿主对每个纯工具轮都发 `content:null`,首版补的桩也是 `content:null`(修复自身即坏体)。修法 = 可见性归一(assistant/tool 的 null/缺键 content → `''`,桩用 `''`)+ 孤儿补桩 + 缺结果合成 + 重复结果丢弃,另加出站 developer→system 折叠。**修复前/后对比实测(真实 dmodel 上游,用当前代码起的临时网关)**:`A_call_content_null` / `B_orphan_plain` / `B_orphan_stub_null` / `C_tool_content_null` 四形态全部 400 → 200 出正文;`""`/文本/纯聊天对照行为不变。两个翻译网关(Qoder :3903、Trae :3902)共用该体检。复现与差分:`node scripts/probe-qoder-pairing.mjs --offline`(宿主真实序列化器,不花额度)、`node scripts/probe-qoder-null-content.mjs --models dmodel --gw-local`。
44
+ 2. **带内失败帧上抛**(0.9.9 早前提交):HTTP 200 里的业务错误不再被静默吞,流式给错误 chunk + `[DONE]`、非流式 502 `qoder_upstream_error` 带上游详情——这正是 Flash 故障的**可读表现**。
45
+ 3. **`Cosy-ClientType` 头保真**:官方 wasm 恒出 `5`,插件旧值 `'qoder'` 会在所有请求上留下第三方客户端指纹;已改为 `5`(与模型可用性无关,两种值下 Flash 都失败)。
46
+
47
+ ## 5. 未解与下一步
48
+
49
+ - **未解**:节点内部为什么失败(`Execution failed: null` 形似 Java NPE)、Flash 的节点指派是部署错误还是压根没上线、Qoder 是否另有 Flash 的服务端别名节点。三项都在服务端,客户端不可观测。
50
+ - **下一步**:节点恢复后跑 `node scripts/probe-qoder-flash-confirm.mjs`(3×Flash + 2×对照,自动落证据);确认恢复再跑 `node scripts/probe-qoder-matrix.mjs --suite flash` 复核全矩阵。
51
+ - **不建议**:在网关目录里把 qfmodel 摘掉"避免选到坏模型"——目录是上游状态的实时镜像,摘除会把**临时**故障固化成配置(且节点恢复后需人工加回);现有行为(选择器保留 + 可读 502 + 探测脚本可复核)更诚实。
52
+
53
+ ## 6. 证据索引
54
+
55
+ | 文件 | 内容 |
56
+ |---|---|
57
+ | `docs/probes/qoder-matrix-1790022926057.json` | qfmodel 变量矩阵(baseline/effort/max_tokens/tools/多轮/source) |
58
+ | `docs/probes/qoder-matrix-1790023075879.json` | 模型家族 × 孤儿 tool 拒绝面(dmodel/kmodel/mmodel 400,auto/qmodel/gmodel 容忍) |
59
+ | `docs/probes/qoder-matrix-1790023159490.json` | 修复策略验证(孤儿补桩 / 缺结果合成 / 重复结果) |
60
+ | `docs/probes/qoder-flash-confirm-*.json` | 3×Flash + 2×对照 的当前状态确认(可反复跑) |
61
+ | `docs/probes/qoder-qfmodel-node-map.json` | 节点归属图 + 头驱动路由证明(`X-Model-Key` 决定路由) |
62
+ | `docs/probes/qoder-qfmodel-key-sweep.json` | 11 个候选 key 的静默改派对照 |
63
+ | `docs/probes/qoder-qfmodel-agentshape.json` | agent 形态 body 复刻对照 |
64
+ | `docs/probes/qoder-qfmodel-clientmeta-sweep.json` | 官方 `QoderContext` × 9 组 clientMetadata |
65
+ | `docs/probes/qoder-qfmodel-minimal-version-scan.txt` | 客户端 bundle 无 `minimal_version` 读取点 |
66
+ | `docs/probes/qoder-qfmodel-client-timeline.txt` | 客户端 754 条 Flash 消息 + 最后成功时刻(transcript 扫描产物,含时间戳与正文摘录) |
67
+ | `docs/probes/qoder-qfmodel-final-confirm.json` | 子探测复核(第一次确认 0/3,含原始 SSE 字节) |