ldrouter 1.17.4 → 1.17.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,25 @@ All notable changes to this project are documented here. The format follows
4
4
  [Keep a Changelog](https://keepachangelog.com/) and the project adheres to
5
5
  [Semantic Versioning](https://semver.org/).
6
6
 
7
+ ## [1.17.6] - 2026-09-20
8
+
9
+ ### Fixed
10
+
11
+ - **Every Qoder request logged `Tokens in 0 · out 0 · cache 0`.** Two independent causes. Streaming: the Qoder branch pipes OpenAI-shaped chunks through the runner's shared chunk handler, which only parsed usage under `cfg.type === 'openai'`, so the accumulator stayed at zero for Qoder — live data confirmed 128 of 128 streaming Qoder requests carried no tokens. Non-streaming: a call asking for no stream can be answered with the whole completion as one plain JSON body, which the SSE-only envelope reader left in its unparsed buffer, so `usage` came back `undefined`. The reader now keeps non-`data:` lines rather than dropping them and the non-streaming path folds a plain completion in (a no-op when the origin did frame its answer as SSE). The usage shapes Qoder actually sends are read too: the object nested one level down as `{statusCodeValue, body}`, and OpenAI-compatible `prompt_tokens_details.cached_tokens`.
12
+ - **A fallback that succeeded was logged as a failure.** The failed first attempt's error was never cleared once a later candidate answered, so the row kept the first attempt's status: 17 requests were recorded `success=0` with `http_status` 502/504/429 while the client had received a 200 answer. The success rate and every per-model figure derived from the request log inherited that error.
13
+ - **The cache-hit rate double-counted the cache.** OpenAI-compatible providers report `cached_tokens` as a *subset* of `prompt_tokens`, so `input` already contains the cached prefix and the blanket `cacheRead / (input + cacheRead)` inflated the denominator — measured at 21.75% where the truth was 27.80% over real traffic. Anthropic is the exception (`input_tokens` *excludes* the cached prefix), so each request is now normalised on its own provider's semantics, on the server and in the live feed alike. The live feed also derives its rates from running totals instead of incrementing a ratio, which is meaningless.
14
+ - **Codex reported zero cached and zero reasoning tokens.** The Responses API nests them under `usage.input_tokens_details.cached_tokens` and `usage.output_tokens_details.reasoning_tokens`; only the flat names were read, so 1,381 requests (46% of traffic) contributed nothing to the cache-hit rate. Flat names still work for upstreams that send them.
15
+ - **A Codex account was retired on a single flaky refresh.** Every account kept turning `degraded` with `last_error=oauth_refresh_failed`, and because the gateway reads that code as "these credentials are dead" it also evicted the account (`enabled=0`). One failure code covered five unrelated things — a request timeout, a DNS error, a token-endpoint 5xx, a 200 carrying no usable token, and a rolled-back database write — and only a grant the endpoint actually rejected is a verdict. The outcomes are now split the way 9router splits them: `oauth_refresh_failed`, `credential_unavailable` and `account_not_found` are final, while `oauth_refresh_unavailable` and `invalid_refresh_response` are transient. A transient failure still moves the request to a sibling account but leaves this one enabled, keeps its health, and writes no `last_error`, so a status the operator reads always means what it says. The rejection is read out of the response body and only on a 4xx — a 5xx or a gateway page echoing `invalid_grant` is still transient — and the body is read whole, because OpenAI nests the code under `error.code` where a top-level read sees nothing. Live data: every account held `consecutive_failures = 0`, the counter meant to require a pattern before a verdict was never reached.
16
+ - **`+ Add` in the New combo dialog never added the model.** One handler served both dialogs and wrote to the *edit* form, so adding in the create dialog mutated the edit state while the create member list stayed empty — the button read as dead. The edit dialog worked, which is why the report was "sometimes". Each dialog now owns its member list.
17
+
18
+ ## [1.17.5] - 2026-09-19
19
+
20
+ ### Fixed
21
+
22
+ - **An account whose session was revoked kept being selected instead of leaving the Codex pool.** Auth0 answers `refresh_token_invalidated` ("Your session has ended. Please log in again.") for an account whose login was revoked; the account cannot serve anything again until it is re-imported. The credential layer surfaced that as a bare `oauth_refresh_failed`, which the router classified `unknown` — and `unknown` is not a routing decision, so `shouldFallback` answered false, the retry never advanced to a sibling account, and the request answered `502`. Live production data confirmed the shape: two accounts holding invalidated refresh tokens sat at the head of the pool, so every request landed on one of them, and 226 of 241 failed requests carried `attempts_count=1` while healthy accounts sat idle behind them. Credential failure is now its own failure class, checked before the error-type switch, and it always advances to the next candidate — with or without a combo plan, since a direct `codex1/...` model has no fallback triggers to consult.
23
+ - **The account with the dead credential is now disabled, not merely degraded.** Marking down *and* disabling (`enabled=0`) is what actually stops the churn, mirroring quota exhaustion: every selection path — `expandCodexAccountCandidates`, `getCodexAccountForProvider` — filters on `enabled`. Deliberately kept out of `isUpstreamHealthFailure`, because one re-imported account must not open the provider circuit breaker for its healthy siblings.
24
+ - **`codexCredentialError` discarded the raw error code.** Its wrapping message is deliberately generic, so without the cause the routing layer could not distinguish a dead account from any other `authentication_error`. The raw error is now carried as `cause`.
25
+
7
26
  ## [1.17.4] - 2026-09-18
8
27
 
9
28
  ### Fixed
@@ -261,6 +261,9 @@ export class GatewayRunner {
261
261
  });
262
262
  attempt.statusCode = out.statusCode ?? null;
263
263
  attempt.success = true;
264
+ // A fallback that succeeded answered the client 200: the earlier failed
265
+ // attempt must not keep the request logged (and reported) as a failure.
266
+ lastError = null;
264
267
  attempt.latencyMs = Date.now() - attemptStart;
265
268
  attempt.ttftMs = out.ttftMs;
266
269
  attempt.usage = out.usage;
@@ -296,6 +299,9 @@ export class GatewayRunner {
296
299
  catch (e) {
297
300
  const err = e instanceof GatewayError ? e : new GatewayError('upstream_error', e.message, { cause: e });
298
301
  const quotaFailure = isQuotaFailure(err);
302
+ // Only a final verdict retires the account (below); a transient refresh failure still moves
303
+ // the request to a sibling account but must leave this one enabled.
304
+ const credentialFailure = isPermanentCredentialFailure(err);
299
305
  debugUpstream(ctx.requestId, 'ATTEMPT ERROR', [
300
306
  `attempt=${i + 1}`,
301
307
  `provider=${provider.name}`,
@@ -329,6 +335,11 @@ export class GatewayRunner {
329
335
  // dead account on every request for hours (verified live) and answered "usage limited".
330
336
  if (quotaFailure)
331
337
  markQuotaExhausted(candidate, err.message);
338
+ // A dead refresh token also leaves the pool outright: the account cannot serve anything
339
+ // again until it is re-imported, so a retry may only be spent on a sibling account. Without
340
+ // this the retry landed on the same dead account and the request answered 502.
341
+ if (credentialFailure)
342
+ markCredentialsDead(candidate, err.message);
332
343
  if (isUpstreamHealthFailure(err)) {
333
344
  recordFailure(provider.id, provider.cbFailureThreshold, provider.cbCooldownSeconds);
334
345
  if (candidate.codexAccountId)
@@ -611,7 +622,9 @@ export class GatewayRunner {
611
622
  }
612
623
  try {
613
624
  const obj = JSON.parse(chunk.data);
614
- if (cfg.type === 'openai') {
625
+ // Qoder emits OpenAI-shaped chunks and pipes them through this same handler
626
+ // (see the qoder branch below), so its usage block is parsed here too.
627
+ if (cfg.type === 'openai' || cfg.type === 'qoder') {
615
628
  const choice = obj.choices?.[0];
616
629
  if (choice?.delta?.content)
617
630
  textBuf += choice.delta.content;
@@ -1032,6 +1045,53 @@ export function markQuotaExhausted(candidate, message) {
1032
1045
  if (candidate.qoderAccountId)
1033
1046
  setQoderAccountHealth(candidate.qoderAccountId, 'down', reason, false);
1034
1047
  }
1048
+ /**
1049
+ * A credential failure is a durable, per-account verdict — the same class as quota exhaustion.
1050
+ *
1051
+ * `withCodexCredentials` throws these as bare codes when the account's refresh token is dead
1052
+ * (Auth0 answers `refresh_token_invalidated`, "Your session has ended. Please log in again.").
1053
+ * The account cannot serve anything again until it is re-imported, so retrying it is pure waste
1054
+ * and leaving it in the pool answers 502 to every client.
1055
+ *
1056
+ * Live evidence (production, 2026-09-19): two accounts holding invalidated refresh tokens sat at
1057
+ * the head of the pool, so every request landed on one of them — `oauth_refresh_failed`,
1058
+ * classified `unknown`, `attempts_count=1`, 502 straight to the client (226 of 241 failed
1059
+ * requests had exactly one attempt while healthy accounts sat idle behind them).
1060
+ *
1061
+ * Deliberately NOT part of `isUpstreamHealthFailure`: one re-imported account must not open the
1062
+ * provider circuit breaker for its healthy siblings. The verdict is applied per account by
1063
+ * `markCredentialsDead` instead.
1064
+ *
1065
+ * Both spellings are checked because the code travels either raw (from `withCodexCredentials`
1066
+ * inside the gateway) or wrapped by `codexCredentialError`, which now keeps it as `cause`.
1067
+ */
1068
+ export function isCredentialFailure(err) {
1069
+ const codes = /^(oauth_refresh_failed|oauth_refresh_unavailable|invalid_refresh_response|credential_unavailable|account_not_found)$/;
1070
+ const cause = err.cause;
1071
+ return codes.test(err.message) || (typeof cause?.message === 'string' && codes.test(cause.message));
1072
+ }
1073
+ /**
1074
+ * Whether the verdict is final, i.e. the account cannot serve again until an operator re-imports it.
1075
+ *
1076
+ * A transient refresh failure (timeout, network error, token-endpoint 5xx, rolled-back write) is
1077
+ * still a credential failure for routing — the request should move to a sibling account — but it
1078
+ * must NOT retire this one. Treating both classes as final is what degraded healthy accounts and
1079
+ * disabled them (`enabled=0`) on a single flaky moment; every live account carried
1080
+ * `consecutive_failures = 0`, so no pattern was ever required before the verdict.
1081
+ */
1082
+ export function isPermanentCredentialFailure(err) {
1083
+ const codes = /^(oauth_refresh_failed|credential_unavailable|account_not_found)$/;
1084
+ const cause = err.cause;
1085
+ return codes.test(err.message) || (typeof cause?.message === 'string' && codes.test(cause.message));
1086
+ }
1087
+ /** Takes an account with a dead refresh token out of the pool. `enabled=0` is what stops the churn. */
1088
+ export function markCredentialsDead(candidate, message) {
1089
+ const reason = redactString(message);
1090
+ if (candidate.codexAccountId)
1091
+ setCodexAccountHealth(candidate.codexAccountId, 'down', reason, false);
1092
+ if (candidate.qoderAccountId)
1093
+ setQoderAccountHealth(candidate.qoderAccountId, 'down', reason, false);
1094
+ }
1035
1095
  /**
1036
1096
  * Whether a failed attempt should advance to the next candidate.
1037
1097
  *
@@ -1043,11 +1103,17 @@ export function markQuotaExhausted(candidate, message) {
1043
1103
  * plan means no fallback.
1044
1104
  */
1045
1105
  export function shouldRetryAttempt(comboPlan, err) {
1106
+ if (isCredentialFailure(err))
1107
+ return true;
1046
1108
  if (isQuotaFailure(err))
1047
1109
  return true;
1048
1110
  return comboPlan ? shouldFallback(comboPlan, { type: classifyFailure(err), status: err.status }) : false;
1049
1111
  }
1050
1112
  export function classifyFailure(err) {
1113
+ // Checked before the type switch for the same reason as quota: a dead account is not a routing
1114
+ // decision the combo's triggers were meant to gate, and `unknown` is not a routing decision at all.
1115
+ if (isCredentialFailure(err))
1116
+ return 'credential';
1051
1117
  // Checked first, before the type switch: a quota refusal arrives as `upstream_rate_limit` (a
1052
1118
  // 429) or `upstream_error` (Codex rewraps it as a 502 with a 429 cause; Qoder uses code 112) and
1053
1119
  // must not fall through to `http_status`/`unknown`, neither of which is a routing decision.
@@ -10,11 +10,24 @@ export function codexCredentialError(error) {
10
10
  const code = error instanceof Error ? error.message : '';
11
11
  if (code === 'account_not_found')
12
12
  return new GatewayError('invalid_request_error', 'Codex account not found', { status: 404 });
13
- if (code === 'oauth_refresh_failed' || code === 'invalid_refresh_response' || code === 'credential_unavailable') {
14
- return new GatewayError('authentication_error', 'Codex credentials could not be refreshed — re-import the account', { status: 401 });
13
+ if (code === 'oauth_refresh_unavailable' || code === 'invalid_refresh_response') {
14
+ // Transient: the token is probably still good, so the operator should retry, not re-import.
15
+ // Same wrapping as the final verdict (the routing layer sees one credential class either way)
16
+ // but the message must not send them to the import dialog over one flaky attempt.
17
+ return new GatewayError('authentication_error', 'Codex credentials could not be refreshed — try again', { status: 401, cause: error });
18
+ }
19
+ if (code === 'oauth_refresh_failed' || code === 'credential_unavailable') {
20
+ // `cause` carries the raw credential code: the wrapping message is deliberately generic, so
21
+ // without it the routing layer cannot tell a dead account from any other authentication_error.
22
+ return new GatewayError('authentication_error', 'Codex credentials could not be refreshed — re-import the account', { status: 401, cause: error });
15
23
  }
16
24
  return error;
17
25
  }
26
+ /** Only a final verdict may retire an account; a transient failure leaves it in the pool. */
27
+ const PERMANENT_REFRESH_FAILURES = new Set(['oauth_refresh_failed', 'credential_unavailable', 'account_not_found']);
28
+ export function isPermanentRefreshFailure(error) {
29
+ return PERMANENT_REFRESH_FAILURES.has(error);
30
+ }
18
31
  const REFRESH_LEAD_MS = 5 * 60 * 1000;
19
32
  const REFRESH_TIMEOUT_MS = 10_000;
20
33
  const flights = new Map();
@@ -37,8 +50,12 @@ async function defaultRefreshClient(input) {
37
50
  body: new URLSearchParams({ grant_type: 'refresh_token', refresh_token: input.refreshToken, client_id: CODEX_OAUTH.clientId }),
38
51
  signal: input.signal,
39
52
  });
40
- if (!response.ok)
41
- throw new Error(`oauth refresh http ${response.status}`);
53
+ if (!response.ok) {
54
+ // The RFC 6749 error body is what separates a dead grant from a bad moment; `safeError` classifies
55
+ // on it. Only the resulting code is ever persisted, never this text.
56
+ const detail = await response.text().catch(() => '');
57
+ throw Object.assign(new Error(`oauth refresh http ${response.status}`), { status: response.status, oauthError: detail });
58
+ }
42
59
  const value = await response.json();
43
60
  if (typeof value.access_token !== 'string' || !value.access_token)
44
61
  throw new Error('invalid refresh response');
@@ -50,9 +67,25 @@ async function defaultRefreshClient(input) {
50
67
  expiresAt: typeof value.expires_at === 'string' ? value.expires_at : undefined,
51
68
  };
52
69
  }
53
- function safeError(_error) {
54
- return 'oauth_refresh_failed';
70
+ /** Grant rejections the endpoint answers with; only these mean the refresh token is dead. */
71
+ const REJECTED_GRANT_MARKERS = ['invalid_grant', 'refresh_token_invalidated', 'refresh_token_reused', 'refresh_token_expired', 'session has ended'];
72
+ /**
73
+ * Decides whether a failed refresh is a verdict or just a failed attempt.
74
+ *
75
+ * A marker only counts when the endpoint itself rejected the grant (4xx): a 5xx or a gateway error
76
+ * page that happens to echo `invalid_grant` is still transient. Anything without a marker — a
77
+ * timeout, a DNS error, an aborted request, a rolled-back write — is transient by default, so an
78
+ * account is never retired on evidence this layer does not actually have.
79
+ */
80
+ function safeError(error) {
81
+ const status = error?.status;
82
+ if (typeof status === 'number' && status >= 500)
83
+ return 'oauth_refresh_unavailable';
84
+ const oauthError = error?.oauthError;
85
+ const text = `${error?.message ?? ''} ${typeof oauthError === 'string' ? oauthError : ''}`.toLowerCase();
86
+ return REJECTED_GRANT_MARKERS.some((marker) => text.includes(marker)) ? 'oauth_refresh_failed' : 'oauth_refresh_unavailable';
55
87
  }
88
+ /** Degrades an account for a final verdict only — a transient failure must leave health alone. */
56
89
  function markDegraded(accountId, error) {
57
90
  try {
58
91
  setCodexAccountHealth(accountId, 'degraded', error);
@@ -64,9 +97,11 @@ async function refreshOnce(accountId, now, force) {
64
97
  try {
65
98
  state = getCodexAccountRefreshState(accountId);
66
99
  }
67
- catch (error) {
68
- markDegraded(accountId, safeError(error));
69
- return { ok: false, error: 'oauth_refresh_failed' };
100
+ catch {
101
+ // A row that cannot be read is an operator problem, not a failed refresh: it needs a re-import,
102
+ // so the account leaves the pool instead of burning an attempt on every request.
103
+ markDegraded(accountId, 'credential_unavailable');
104
+ return { ok: false, error: 'credential_unavailable' };
70
105
  }
71
106
  if (!state)
72
107
  return { ok: false, error: 'account_not_found' };
@@ -77,7 +112,8 @@ async function refreshOnce(accountId, now, force) {
77
112
  try {
78
113
  const result = await refreshClient({ refreshToken: state.refreshToken, signal: controller.signal });
79
114
  if (typeof result.accessToken !== 'string' || !result.accessToken) {
80
- markDegraded(accountId, 'invalid_refresh_response');
115
+ // A 200 with nothing usable is suspicious, not proof the grant is dead: the next attempt may
116
+ // succeed, so the account keeps its health and stays in the pool.
81
117
  return { ok: false, error: 'invalid_refresh_response' };
82
118
  }
83
119
  const expiresAt = result.expiresAt ?? (typeof result.expiresIn === 'number' && Number.isFinite(result.expiresIn) && result.expiresIn > 0
@@ -85,7 +121,6 @@ async function refreshOnce(accountId, now, force) {
85
121
  : '');
86
122
  const expiresMs = Date.parse(expiresAt);
87
123
  if (!Number.isFinite(expiresMs) || expiresMs <= now.getTime()) {
88
- markDegraded(accountId, 'invalid_refresh_response');
89
124
  return { ok: false, error: 'invalid_refresh_response' };
90
125
  }
91
126
  try {
@@ -96,15 +131,17 @@ async function refreshOnce(accountId, now, force) {
96
131
  expiresAt,
97
132
  });
98
133
  }
99
- catch (error) {
100
- markDegraded(accountId, safeError(error));
101
- return { ok: false, error: 'oauth_refresh_failed' };
134
+ catch {
135
+ // The transaction rolled back, so the stored token is unchanged and still usable.
136
+ return { ok: false, error: 'oauth_refresh_unavailable' };
102
137
  }
103
138
  return { ok: true, expiresAt };
104
139
  }
105
140
  catch (error) {
106
- markDegraded(accountId, safeError(error));
107
- return { ok: false, error: 'oauth_refresh_failed' };
141
+ const failure = safeError(error);
142
+ if (isPermanentRefreshFailure(failure))
143
+ markDegraded(accountId, failure);
144
+ return { ok: false, error: failure };
108
145
  }
109
146
  finally {
110
147
  clearTimeout(timer);
@@ -127,8 +164,8 @@ export async function withCodexCredentials(accountId, fn, now = new Date()) {
127
164
  try {
128
165
  state = getCodexAccountRefreshState(accountId);
129
166
  }
130
- catch (error) {
131
- markDegraded(accountId, safeError(error));
167
+ catch {
168
+ markDegraded(accountId, 'credential_unavailable');
132
169
  throw new Error('credential_unavailable', { cause: new Error('credential_unavailable') });
133
170
  }
134
171
  if (!state)
@@ -142,8 +179,8 @@ export async function withCodexCredentials(accountId, fn, now = new Date()) {
142
179
  try {
143
180
  credentials = getCodexCredentials(accountId);
144
181
  }
145
- catch (error) {
146
- markDegraded(accountId, safeError(error));
182
+ catch {
183
+ markDegraded(accountId, 'credential_unavailable');
147
184
  throw new Error('credential_unavailable', { cause: new Error('credential_unavailable') });
148
185
  }
149
186
  try {
@@ -158,8 +195,8 @@ export async function withCodexCredentials(accountId, fn, now = new Date()) {
158
195
  try {
159
196
  credentials = getCodexCredentials(accountId);
160
197
  }
161
- catch (credentialError) {
162
- markDegraded(accountId, safeError(credentialError));
198
+ catch {
199
+ markDegraded(accountId, 'credential_unavailable');
163
200
  throw new Error('credential_unavailable', { cause: new Error('credential_unavailable') });
164
201
  }
165
202
  return fn(credentials);
@@ -151,7 +151,18 @@ function usage(value) {
151
151
  const u = (value && typeof value === 'object' ? value : {});
152
152
  const input = typeof u.input_tokens === 'number' ? u.input_tokens : 0;
153
153
  const output = typeof u.output_tokens === 'number' ? u.output_tokens : 0;
154
- return { input, output, total: input + output, cacheRead: typeof u.cached_input_tokens === 'number' ? u.cached_input_tokens : 0, cacheWrite: 0, reasoning: typeof u.reasoning_tokens === 'number' ? u.reasoning_tokens : 0 };
154
+ // The Responses API nests both of these (`usage.input_tokens_details.cached_tokens`,
155
+ // `usage.output_tokens_details.reasoning_tokens`). Reading only the flat names returned
156
+ // 0 for every Codex request — 1,381 of them, 46% of all traffic — so the statistics page
157
+ // reported a cache-hit rate missing its largest contributor. Flat names stay supported
158
+ // for compatibility with OpenAI-compatible upstreams that use them.
159
+ const inDetails = (u.input_tokens_details && typeof u.input_tokens_details === 'object' ? u.input_tokens_details : {});
160
+ const outDetails = (u.output_tokens_details && typeof u.output_tokens_details === 'object' ? u.output_tokens_details : {});
161
+ const cacheRead = typeof inDetails.cached_tokens === 'number' ? inDetails.cached_tokens
162
+ : typeof u.cached_input_tokens === 'number' ? u.cached_input_tokens : 0;
163
+ const reasoning = typeof outDetails.reasoning_tokens === 'number' ? outDetails.reasoning_tokens
164
+ : typeof u.reasoning_tokens === 'number' ? u.reasoning_tokens : 0;
165
+ return { input, output, total: input + output, cacheRead, cacheWrite: 0, reasoning };
155
166
  }
156
167
  export function codexResponseToCanonical(body, requestedModel) {
157
168
  let text = '';
@@ -209,6 +209,23 @@ async function drainEnvelope(reader, response, onChunk) {
209
209
  await upstream.cancel().catch(() => { });
210
210
  }
211
211
  }
212
+ /**
213
+ * Interpret a non-SSE body as a plain OpenAI chat completion and fold it into the reader.
214
+ * A no-op when the origin framed its answer as SSE frames, which is the usual case.
215
+ */
216
+ function applyPlainCompletion(reader) {
217
+ const raw = reader.takeUnparsed();
218
+ if (!raw)
219
+ return;
220
+ let body;
221
+ try {
222
+ body = JSON.parse(raw);
223
+ }
224
+ catch {
225
+ return;
226
+ }
227
+ reader.applyCompletion(body);
228
+ }
212
229
  /** Streaming: emit OpenAI-shaped chunks as they arrive, stop at the terminal frame. */
213
230
  export async function callQoderStreaming(cfg, req, onChunk, deps = {}) {
214
231
  const { response, reader, status, upstreamRequestId } = await openSignedStream(cfg, { ...req, stream: true }, deps);
@@ -219,6 +236,12 @@ export async function callQoderStreaming(cfg, req, onChunk, deps = {}) {
219
236
  export async function callQoderNonStreaming(cfg, req, deps = {}) {
220
237
  const { response, reader, status, upstreamRequestId } = await openSignedStream(cfg, { ...req, stream: false }, deps);
221
238
  await drainEnvelope(reader, response);
239
+ // The caller invoked this with stream disabled, so the origin may honour the stream
240
+ // preference and answer with the full completion in one body rather than SSE frames.
241
+ // The reader buffers those bytes as an unparsed trailing line (it only understands
242
+ // `data:` framing), which is why a non-streaming Qoder call logged zero tokens:
243
+ // `usage` came back `undefined` and the runner reported `in 0 · out 0 · cache 0`.
244
+ applyPlainCompletion(reader);
222
245
  return { status, upstreamRequestId, text: reader.text, toolCalls: reader.toolCalls, finishReason: reader.finishReason, usage: reader.usage ?? emptyUsage() };
223
246
  }
224
247
  export async function probeQoder(cfg) {
@@ -8,13 +8,18 @@ function numberOr(value, fallback) {
8
8
  }
9
9
  export function qoderUsageOf(value) {
10
10
  const usage = (value && typeof value === 'object' ? value : {});
11
- const details = (usage.prompt_tokens_details && typeof usage.prompt_tokens_details === 'object' ? usage.prompt_tokens_details : {});
12
- const input = numberOr(usage.prompt_tokens ?? usage.input_tokens, 0);
13
- const output = numberOr(usage.completion_tokens ?? usage.output_tokens, 0);
11
+ // The envelope nests the real OpenAI usage object one level down under `body`, so a
12
+ // few shapes are worth reading; the flat one stays authoritative when present.
13
+ const nested = (usage.usage && typeof usage.usage === 'object' ? usage.usage : {});
14
+ const asObject = (v) => (v && typeof v === 'object' ? v : {});
15
+ const details = asObject(usage.prompt_tokens_details ?? nested.prompt_tokens_details);
16
+ const input = numberOr(usage.prompt_tokens ?? usage.input_tokens ?? nested.prompt_tokens ?? nested.input_tokens, 0);
17
+ const output = numberOr(usage.completion_tokens ?? usage.output_tokens ?? nested.completion_tokens ?? nested.output_tokens, 0);
14
18
  return {
15
19
  input, output,
16
- total: numberOr(usage.total_tokens, input + output),
17
- cacheRead: numberOr(details.cached_tokens ?? usage.cached_tokens ?? usage.cache_read_input_tokens, 0),
20
+ total: numberOr(usage.total_tokens ?? nested.total_tokens, input + output),
21
+ cacheRead: numberOr(details.cached_tokens ?? usage.cached_tokens ?? usage.cache_read_input_tokens
22
+ ?? nested.cached_tokens, 0),
18
23
  cacheWrite: numberOr(details.cache_creation_tokens ?? usage.cache_creation_input_tokens, 0),
19
24
  reasoning: numberOr(usage.reasoning_tokens ?? usage.completion_tokens_details?.reasoning_tokens, 0),
20
25
  };
@@ -28,6 +33,8 @@ export class QoderEnvelopeReader {
28
33
  usage = null;
29
34
  error = null;
30
35
  buffer = '';
36
+ // Non-`data:` lines, in arrival order: a plain (unframed) JSON body. See takeUnparsed().
37
+ plain = [];
31
38
  ended = false;
32
39
  pendingFinish = null;
33
40
  pendingUsage = null;
@@ -39,6 +46,50 @@ export class QoderEnvelopeReader {
39
46
  }
40
47
  terminal() { return this.ended; }
41
48
  errorEnvelope() { return this.error; }
49
+ /**
50
+ * Take the buffered, never-`data:`-framed bytes (if any) and clear them. An origin
51
+ * answering a non-streaming call with a plain JSON completion leaves the whole
52
+ * response here — the SSE line scanner only recognises `data:` frames.
53
+ */
54
+ takeUnparsed() {
55
+ const raw = [...this.plain, this.buffer].join('\n').trim();
56
+ this.plain = [];
57
+ this.buffer = '';
58
+ return raw;
59
+ }
60
+ /**
61
+ * Fold an OpenAI-shaped non-streaming completion into this reader, so the caller
62
+ * gets text/toolCalls/finishReason/usage from one accessor set regardless of how
63
+ * the origin framed the answer. `usage` is only written when the body carries one:
64
+ * a zero would be indistinguishable from "the provider reported nothing".
65
+ */
66
+ applyCompletion(body) {
67
+ const choice = (Array.isArray(body.choices) ? body.choices[0] : null);
68
+ const message = (choice?.message ?? {});
69
+ if (typeof message.content === 'string')
70
+ this.text += message.content;
71
+ const calls = Array.isArray(message.tool_calls) ? message.tool_calls : [];
72
+ for (const call of calls) {
73
+ let input = {};
74
+ if (typeof call.function?.arguments === 'string' && call.function.arguments) {
75
+ try {
76
+ input = JSON.parse(call.function.arguments);
77
+ }
78
+ catch {
79
+ input = {};
80
+ }
81
+ }
82
+ this.toolCalls.push({ id: call.id ?? `call-${this.toolCalls.length}`, name: call.function?.name ?? 'unknown', input });
83
+ }
84
+ // Assigned directly, not via the pending machinery: this body is complete, and the
85
+ // stream has already ended (finish() set `ended`), which those paths return early on.
86
+ const finish = (choice?.finish_reason ?? (typeof body.status === 'string' ? body.status : null));
87
+ if (finish)
88
+ this.finishReason = finish;
89
+ if (body.usage)
90
+ this.usage = qoderUsageOf(body.usage);
91
+ this.ended = true;
92
+ }
42
93
  push(text) {
43
94
  if (this.ended)
44
95
  return;
@@ -63,9 +114,17 @@ export class QoderEnvelopeReader {
63
114
  }
64
115
  cancel() { this.ended = true; this.buffer = ''; }
65
116
  line(raw) {
66
- const trimmed = raw.replace(/\r$/, '').trim();
67
- if (!trimmed.startsWith('data:'))
117
+ // trim() also drops the trailing CR of a CRLF frame, so no separate strip is needed.
118
+ const trimmed = raw.trim();
119
+ if (!trimmed.startsWith('data:')) {
120
+ // An origin answering a non-streaming call may send the whole completion as one
121
+ // plain JSON body instead of SSE frames. Hold those lines rather than dropping
122
+ // them: finish() consumes the buffer, so this is the only place they survive.
123
+ // Blank lines and `:comment` keepalives are not content.
124
+ if (trimmed && !trimmed.startsWith(':'))
125
+ this.plain.push(trimmed);
68
126
  return;
127
+ }
69
128
  const data = trimmed.slice(5).trimStart();
70
129
  if (data === '[DONE]') {
71
130
  this.flushPending();
@@ -35,6 +35,7 @@ export function toSummary(r, maps) {
35
35
  finalModelPublicId: finalModel?.publicModelId ?? null,
36
36
  providerId: provider ? finalModel.providerId : null,
37
37
  providerName: provider?.name ?? null,
38
+ providerType: provider?.type ?? null,
38
39
  streaming: Boolean(r.streaming),
39
40
  httpStatus: r.httpStatus,
40
41
  success: Boolean(r.success),
@@ -18,6 +18,34 @@ const PRESETS = {
18
18
  return { from: new Date(now.getTime() - 30 * 24 * 3600 * 1000), to: now, bucket: 'day' };
19
19
  },
20
20
  };
21
+ /**
22
+ * Provider-cache hit rate for a window.
23
+ *
24
+ * The blanket `cacheRead / (input + cacheRead)` double-counts the cache: OpenAI-compatible
25
+ * providers report `cached_tokens` as a SUBSET of `prompt_tokens`, so the cached prefix is
26
+ * already inside `input` and adding it again inflates the denominator — understating the
27
+ * rate (measured on real traffic: 21.75% instead of 27.80%).
28
+ *
29
+ * Anthropic is the exception: its `input_tokens` EXCLUDES the cached prefix, so there the
30
+ * prompt really is `input + cache`. Each row is weighed by its provider's own `type` in a
31
+ * single pass.
32
+ */
33
+ function cacheHitRateFor(db, conds) {
34
+ const anthropic = sql `${schema.providers.type} = 'anthropic'`;
35
+ const r = db
36
+ .select({
37
+ cacheRead: sql `COALESCE(SUM(${schema.requests.cacheReadTokens}),0)`,
38
+ anthropicPrompt: sql `COALESCE(SUM(CASE WHEN ${anthropic} THEN ${schema.requests.inputTokens} + ${schema.requests.cacheReadTokens} ELSE 0 END),0)`,
39
+ otherPrompt: sql `COALESCE(SUM(CASE WHEN ${anthropic} THEN 0 ELSE ${schema.requests.inputTokens} END),0)`,
40
+ })
41
+ .from(schema.requests)
42
+ .leftJoin(schema.models, eq(schema.models.id, schema.requests.finalModelId))
43
+ .leftJoin(schema.providers, eq(schema.providers.id, schema.models.providerId))
44
+ .where(and(...conds))
45
+ .get();
46
+ const prompt = Number(r?.anthropicPrompt ?? 0) + Number(r?.otherPrompt ?? 0);
47
+ return prompt > 0 ? Number(r?.cacheRead ?? 0) / prompt : 0;
48
+ }
21
49
  function rangeFromQuery(q) {
22
50
  if (q.preset && PRESETS[q.preset])
23
51
  return PRESETS[q.preset]();
@@ -50,6 +78,7 @@ export async function registerStatsRoutes(app) {
50
78
  .from(schema.requests)
51
79
  .where(and(...conds))
52
80
  .get();
81
+ const cacheHitRate = cacheHitRateFor(db, conds);
53
82
  const latRows = db
54
83
  .select({ v: schema.requests.totalLatencyMs })
55
84
  .from(schema.requests)
@@ -84,7 +113,7 @@ export async function registerStatsRoutes(app) {
84
113
  p95LatencyMs: p95,
85
114
  averageTtftMs: avgTtft,
86
115
  p95TtftMs: p95Ttft,
87
- cacheHitRate: success ? Number(summary?.cacheRead ?? 0) > 0 ? Number(summary?.cacheRead ?? 0) / Math.max(1, Number(summary?.inputTokens ?? 0) + Number(summary?.cacheRead ?? 0)) : 0 : 0,
116
+ cacheHitRate,
88
117
  gatewayCacheHitRate: total ? Number(summary?.gatewayCacheHits ?? 0) / total : 0,
89
118
  fallbackRate: total ? Number(summary?.fallbacks ?? 0) / total : 0,
90
119
  };
@@ -241,6 +270,7 @@ function buildSummary(db, fromIso, toIso) {
241
270
  const success = Number(r?.success ?? 0);
242
271
  const cacheRead = Number(r?.cacheRead ?? 0);
243
272
  const inputTokens = Number(r?.inputTokens ?? 0);
273
+ const cacheHitRate = cacheHitRateFor(db, conds);
244
274
  return {
245
275
  totalRequests: total,
246
276
  successfulRequests: success,
@@ -256,7 +286,7 @@ function buildSummary(db, fromIso, toIso) {
256
286
  p95LatencyMs: 0, // caller fills via percentile query
257
287
  averageTtftMs: (r?.avgTtft ?? null),
258
288
  p95TtftMs: null,
259
- cacheHitRate: success ? cacheRead > 0 ? cacheRead / Math.max(1, inputTokens + cacheRead) : 0 : 0,
289
+ cacheHitRate,
260
290
  gatewayCacheHitRate: total ? Number(r?.gatewayCacheHits ?? 0) / total : 0,
261
291
  fallbackRate: total ? Number(r?.fallbacks ?? 0) / total : 0,
262
292
  };