@rikcodes/teamclaude 1.1.20-rik.13 → 1.1.20-rik.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -107,10 +107,12 @@ A Claude Code session can use OpenAI models alongside the Claude accounts. They
107
107
  **1. Install the sidecar and log it into your ChatGPT account** (one time):
108
108
 
109
109
  ```bash
110
- brew install raine/claude-code-proxy/claude-code-proxy
110
+ curl -fsSL https://raw.githubusercontent.com/rikbrown/claude-code-proxy/rik/main/scripts/install.sh | bash
111
111
  claude-code-proxy codex auth login
112
112
  ```
113
113
 
114
+ This installs [rikbrown/claude-code-proxy](https://github.com/rikbrown/claude-code-proxy), a fork of the reference sidecar. Take the fork rather than upstream's Homebrew build: it forwards Codex's quota headers, without which every bar on the account reads `unknown`, and it lets you raise the 60-second header timeout that otherwise kills a long reasoning turn. See [Timeouts](docs/openai.md#timeouts).
115
+
114
116
  **2. Connect it** — add four pieces to `~/.config/teamclaude.json`:
115
117
 
116
118
  ```json
@@ -161,11 +163,11 @@ Each request is routed by the model name in its body, so one session can freely
161
163
 
162
164
  1. Check that your sidecar build lists it: `curl -s http://127.0.0.1:18765/v1/models`. The sidecar has its own allow-list and rejects any id that it does not know, regardless of the TeamClaude configuration. Upgrade the sidecar if the id is missing.
163
165
  2. Add a `customModels` row. Codex publishes the window for each model as `context_window` in `~/.codex/models_cache.json`; copy it to `contextTokens`.
164
- 3. Start a new `teamclaude run` session. The rows are read at launch, so you do not need to restart the server. If you upgraded the sidecar binary, restart the server — or send `SIGTERM` to the sidecar process and let the supervisor restart it with the new binary.
166
+ 3. Start a new `teamclaude run` session. The rows are read at launch, so you do not need to restart the server. If you upgraded the sidecar binary, restart the server — or send `SIGTERM` to the sidecar process and let the supervisor restart it with the new binary. The first `SIGTERM` only starts a graceful shutdown, which a request in flight holds open; send it a second time to force the exit.
165
167
 
166
- Claude Code prints one `[claude-code:unrecognized_model]` line to stderr for each custom model. This is expected; suppressing it would lose the correct context window. The quota bars for the sidecar account show `unknown` unless the sidecar forwards Codex's rate-limit headers — see [Quota](docs/openai.md#quota). Keep the sidecar on loopback.
168
+ Claude Code prints one `[claude-code:unrecognized_model]` line to stderr for each custom model. This is expected; suppressing it would lose the correct context window. The quota bars for the sidecar account show `unknown` unless the sidecar forwards Codex's rate-limit headers, which the fork build in step 1 does and upstream's does not — see [Quota](docs/openai.md#quota). Keep the sidecar on loopback.
167
169
 
168
- The sidecar appears under the account table as a `⚙` line rather than a row because it holds no subscription, is the only account its route can use, and never rotates. The line also shows its supervised process state (`up pid 98018`, or `down (code 1) 3 restarts`).
170
+ The sidecar appears under the account table as a `⚙` line rather than a row because it holds no subscription, is the only account its route can use, and never rotates. The line also shows its supervised process state (`up pid 98018`, or `down (code 1) 3 restarts`), followed by what the sidecar reports about itself while it is up (`2 active 3 errors`, read from its own `/monitor` endpoint; each is omitted at zero, and both are omitted if that endpoint does not answer).
169
171
 
170
172
  #### Several ChatGPT accounts
171
173
 
@@ -221,7 +223,7 @@ Three details matter:
221
223
  - **`CCP_CODEX_TRANSPORT=http` is required.** A WebSocket upgrade is relayed with the caller's own headers and draws no account, so the WebSocket transport cannot be pooled.
222
224
  - **Do not reuse a name across providers.** Routes address accounts by name, so a shared name admits both — including the Claude account that cannot serve `gpt-*`, which outranks the sidecar on priority and wins. TeamClaude warns at startup when it sees one.
223
225
 
224
- Two things differ from the single-account setup: each turn appears **twice** in the activity list, once per hop, and tokens are booked against the sidecar account, so a ChatGPT account reads `N req · 0 tok`. Its quota bars are unaffected because they come from the `x-codex-*` headers on the second hop, where the subscription is.
226
+ One thing differs from the single-account setup: each turn appears **twice** in the activity list, once per hop. Tokens and quota bars both come from the second hop, where the subscription is, so they are booked against the ChatGPT account that served. The sidecar account holds a stub login and spends nothing of its own, so it reads `N req · 0 tok` — booking it as well would double every figure.
225
227
 
226
228
  Full details, including what happens to quota on each hop: [Several ChatGPT accounts behind one sidecar](docs/openai.md#several-chatgpt-accounts-behind-one-sidecar).
227
229
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@rikcodes/teamclaude",
3
- "version": "1.1.20-rik.13",
3
+ "version": "1.1.20-rik.15",
4
4
  "description": "Multi-account proxy for Claude Code and Codex: pools Claude Max, ChatGPT/Codex, API-key and third-party backend accounts, and rotates on quota",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -214,6 +214,13 @@ function makeAccount(acct, index) {
214
214
  hasClaudeMax: acct.hasClaudeMax ?? null,
215
215
  hasClaudePro: acct.hasClaudePro ?? null,
216
216
  priority: acct.priority || 0,
217
+ // Where this account sits in the list the TUI draws, arranged by the
218
+ // operator from the settings screen. Presentation only: nothing but
219
+ // _displayOrder in tui.js reads it, and it is emphatically NOT `priority`
220
+ // above — that one decides which account rotation spends next, and the two
221
+ // answer different questions about the same fleet. `null` means never
222
+ // placed, which sorts after every account that has been.
223
+ displayOrder: Number.isFinite(acct.displayOrder) ? acct.displayOrder : null,
217
224
  disabled: acct.disabled || false,
218
225
  maxUsage: acct.maxUsage ?? null,
219
226
  upstream: acct.upstream || null,
@@ -3485,7 +3492,9 @@ export class AccountManager {
3485
3492
  // threshold and every request fails, while a sibling sits at 0%.
3486
3493
  // A standalone sidecar (no Codex accounts here) is NOT a conduit: it holds
3487
3494
  // its own login, the forwarded numbers are its own, and they still apply.
3488
- if (!(isLocalUpstream(account) && this.accounts.some(a => providerOf(a) === 'codex'))) {
3495
+ // Shares its definition with the token counters, which drop a conduit hop
3496
+ // for the same reason: what it reports belongs to whoever served.
3497
+ if (!this._isCodexConduit(account)) {
3489
3498
  for (const window of ['primary', 'secondary']) {
3490
3499
  const used = parseFloat(headers[`x-codex-${window}-used-percent`]);
3491
3500
  const minutes = parseInt(headers[`x-codex-${window}-window-minutes`], 10);
@@ -3556,12 +3565,49 @@ export class AccountManager {
3556
3565
  }
3557
3566
  }
3558
3567
 
3568
+ /**
3569
+ * Whether `account` merely RELAYS to this fleet's own Codex pool instead of
3570
+ * holding a subscription of its own: a translating sidecar whose back leg is
3571
+ * pointed back at this proxy, so each turn crosses this process twice — once
3572
+ * inbound on `/v1/messages` and once outbound on `/backend-api/codex/*`
3573
+ * (docs/openai.md, "Several ChatGPT accounts behind one sidecar").
3574
+ *
3575
+ * Keyed on three things, all of which the documented setup has and no other
3576
+ * account does:
3577
+ *
3578
+ * a loopback upstream the sidecar runs on this machine.
3579
+ * the Anthropic wire a conduit is reached on `/v1/messages`, so it
3580
+ * carries no `provider` field — docs/openai.md warns
3581
+ * that giving one `"provider": "codex"` puts it in the
3582
+ * foreign-subscription partition and every `gpt-*`
3583
+ * request then fails to find an account. A pooled
3584
+ * ChatGPT account is therefore never a conduit,
3585
+ * whatever its upstream says; it IS the pool.
3586
+ * a pool to relay to a standalone sidecar holding its own ChatGPT login
3587
+ * is NOT a conduit: everything it reports is its own,
3588
+ * and it is the only hop there is.
3589
+ *
3590
+ * Still imprecise in one direction, deliberately: a fleet running BOTH a
3591
+ * self-hosting local Anthropic backend and Codex accounts reads the former as
3592
+ * a conduit. Narrowing that would mean correlating the two hops of one turn,
3593
+ * which nothing here can do — they are separate requests sharing no id — and
3594
+ * the cost of the false positive is a row that under-reports rather than one
3595
+ * that misroutes.
3596
+ *
3597
+ * @param {any} account
3598
+ */
3599
+ _isCodexConduit(account) {
3600
+ return isLocalUpstream(account)
3601
+ && providerOf(account) === DEFAULT_PROVIDER
3602
+ && this.accounts.some(a => providerOf(a) === 'codex');
3603
+ }
3604
+
3559
3605
  /**
3560
3606
  * Update cumulative token usage from response body data.
3561
3607
  */
3562
3608
  updateUsage(accountIndex, inputTokens, outputTokens) {
3563
3609
  const account = this.accounts[accountIndex];
3564
- if (!account) return;
3610
+ if (!account || this._isCodexConduit(account)) return;
3565
3611
  if (inputTokens) account.usage.totalInputTokens += inputTokens;
3566
3612
  if (outputTokens) account.usage.totalOutputTokens += outputTokens;
3567
3613
  }
@@ -3581,6 +3627,22 @@ export class AccountManager {
3581
3627
  */
3582
3628
  recordTokenUsage(accountIndex, sessionId, model, usage) {
3583
3629
  if (!usage) return;
3630
+ // A conduit hop is the SAME tokens, translated: the sidecar rebuilt this
3631
+ // report out of the Responses usage the pool sent it, and the pool's own hop
3632
+ // records the original a moment later. Both scopes would double — the two
3633
+ // account rows are distinct, but a session is one row and carries the same
3634
+ // id on both hops, so its context and spend would read twice the truth.
3635
+ //
3636
+ // Dropped rather than deduplicated because there is nothing to deduplicate
3637
+ // against: the two hops are separate requests that share no id, and the
3638
+ // conduit spends nothing of its own anyway (its login is a stub). Booking at
3639
+ // the hop that really spent is what keeps one turn one record.
3640
+ //
3641
+ // Only the ACCOUNT-scoped and SESSION-scoped counters stop here. Per-client
3642
+ // attribution still runs on the inbound hop, in the caller, because that is
3643
+ // the only hop that can see who asked: the outbound one comes from the
3644
+ // sidecar on loopback and carries no client identity at all.
3645
+ if (this._isCodexConduit(this.accounts[accountIndex])) return;
3584
3646
  // The same resolver routing uses, so a token total and a routing decision
3585
3647
  // agree about which family a request belonged to. Resolved here rather than
3586
3648
  // at the call sites: they parse a wire format and have no business knowing
package/src/claude-env.js CHANGED
@@ -76,6 +76,66 @@ export function mergeNoProxy(...inherited) {
76
76
  return out.join(',');
77
77
  }
78
78
 
79
+ /** The two ways a launched client can reach the proxy. */
80
+ export const CLIENT_MODES = ['mitm', 'base-url'];
81
+
82
+ /**
83
+ * Which mode `run` and `env` use: a `--mitm` or `--no-mitm` flag decides, else
84
+ * the config's `defaultClientMode`, else MITM.
85
+ *
86
+ * MITM routes every request of the launched client through the proxy, so the
87
+ * hard-coded api.anthropic.com endpoints and the Codex CLI (which honours only
88
+ * proxy variables) are covered. It is also the whole point of the setting: with
89
+ * `eval "$(teamclaude env)"` the proxy variables are shell-wide, and every other
90
+ * tool in that shell — gh, git, a package manager — follows them to a listener
91
+ * that only speaks to two hosts (#382). An operator who lives in such a shell
92
+ * sets `defaultClientMode: "base-url"` once and opts back in per launch.
93
+ *
94
+ * @param {{ defaultClientMode?: string }|null|undefined} config
95
+ * @param {string[]} flags the invocation's own arguments
96
+ * @returns {'mitm'|'base-url'}
97
+ */
98
+ export function resolveClientMode(config, flags) {
99
+ const mitm = flags.includes('--mitm');
100
+ const noMitm = flags.includes('--no-mitm');
101
+ if (mitm && noMitm) throw new Error('choose either --mitm or --no-mitm');
102
+ if (mitm) return 'mitm';
103
+ if (noMitm) return 'base-url';
104
+ return config?.defaultClientMode === 'base-url' ? 'base-url' : 'mitm';
105
+ }
106
+
107
+ const SHELL_PROXY_VARS = ['HTTP_PROXY', 'HTTPS_PROXY', 'ALL_PROXY', 'http_proxy', 'https_proxy', 'all_proxy'];
108
+
109
+ /**
110
+ * @param {unknown} value
111
+ * @param {number} port
112
+ */
113
+ function pointsAtLoopback(value, port) {
114
+ if (!value) return false;
115
+ try {
116
+ const url = new URL(String(value));
117
+ const host = url.hostname.replace(/^\[|\]$/g, '').toLowerCase();
118
+ return ['127.0.0.1', 'localhost', '::1'].includes(host) && Number(url.port || 80) === port;
119
+ } catch {
120
+ return false;
121
+ }
122
+ }
123
+
124
+ /**
125
+ * `unset` lines for proxy variables a previous MITM-mode eval left in the
126
+ * shell, so re-evaluating in base-URL mode takes the proxy back out of it. Only
127
+ * a value naming THIS proxy's loopback port is touched: a real corporate proxy
128
+ * in the same variables is the operator's and stays.
129
+ *
130
+ * @param {unknown} port
131
+ * @param {NodeJS.ProcessEnv} [env]
132
+ * @returns {string[]}
133
+ */
134
+ export function clearSelfProxyEnvLines(port, env = process.env) {
135
+ const checkedPort = validPort(port);
136
+ return SHELL_PROXY_VARS.filter(name => pointsAtLoopback(env[name], checkedPort)).map(name => `unset ${name}`);
137
+ }
138
+
79
139
  // Build the shell `export` lines that point Claude Code — or any tool that
80
140
  // spawns it, e.g. an agent multiplexer — at the proxy. This is the same
81
141
  // environment `teamclaude run` sets up, but emitted for `eval "$(teamclaude
package/src/config.js CHANGED
@@ -101,6 +101,7 @@ export function createDefaultConfig() {
101
101
  sessionTitles: { enabled: false, width: 18 },
102
102
  projection: { enabled: true, windowMinutes: 90, wasteFloor: 0.1 },
103
103
  eventLogging: 'hide',
104
+ defaultClientMode: 'mitm',
104
105
  activityLog: null,
105
106
  autoRestart: false,
106
107
  blockedModels: [],
package/src/index.js CHANGED
@@ -31,7 +31,7 @@ import { ensureCerts, mitmHosts } from './mitm.js';
31
31
  import { Prober } from './prober.js';
32
32
  import { Warmer } from './warmer.js';
33
33
  import { createRollingWarmupSchedule, formatWarmupScheduleConfirmation, resolveWarmupConfig, resolveWarmupSchedule } from './warmup-schedule.js';
34
- import { Sidecar } from './sidecar.js';
34
+ import { Sidecar, SIDECAR_DRAIN_GRACE_MS, SIDECAR_SHUTDOWN_GRACE_MS } from './sidecar.js';
35
35
  import { TUI } from './tui.js';
36
36
  import { SessionTitles } from './session-titles.js';
37
37
  import { RemoteControl, createAttachSession } from './tui-remote.js';
@@ -42,7 +42,7 @@ import { createVersionSource, UpdateWatcher } from './update-watch.js';
42
42
  import { renderStatus, formatPercent } from './status-renderer.js';
43
43
  import { sanitizeText } from './safe-text.js';
44
44
  import { ClientUsageTracker, UsageDimensionTracker } from './client-usage.js';
45
- import { buildClaudeEnvLines, buildCustomModelAgents, buildCustomModelSettings, buildCustomModelVars, bypassesAllHosts, encodePinComponent, mergeNoProxy } from './claude-env.js';
45
+ import { buildClaudeEnvLines, buildCustomModelAgents, buildCustomModelSettings, buildCustomModelVars, bypassesAllHosts, clearSelfProxyEnvLines, encodePinComponent, mergeNoProxy, resolveClientMode } from './claude-env.js';
46
46
  import { serviceKind, installService, uninstallService, serviceStatus, renderService, logPath } from './service.js';
47
47
  import { formatTerminalTitle, titleSequence, TITLE_STACK_PUSH, TITLE_STACK_POP } from './terminal-title.js';
48
48
  import { getUpstreamProxy, describeProxy, describeSelfProxy } from './upstream-proxy.js';
@@ -474,6 +474,8 @@ async function serverCommand() {
474
474
  // Both are read per request off this object (server.js) and the TUI already
475
475
  // persists them; without this a hand edit or another writer waited for a restart.
476
476
  config.eventLogging = diskConfig.eventLogging || 'hide';
477
+ // Read by `run`/`env` from disk, but the TUI settings screen shows it live.
478
+ config.defaultClientMode = diskConfig.defaultClientMode === 'base-url' ? 'base-url' : 'mitm';
477
479
  config.blockedModels = Array.isArray(diskConfig.blockedModels) ? diskConfig.blockedModels : [];
478
480
  config.projection = diskConfig.projection;
479
481
  accountManager.setProjection(config.projection);
@@ -537,6 +539,7 @@ async function serverCommand() {
537
539
  // screen too; the server reads them live from `config`, but without this
538
540
  // the edit never reached disk and was silently undone by the next start.
539
541
  if (config.eventLogging != null) diskConfig.eventLogging = config.eventLogging;
542
+ if (config.defaultClientMode != null) diskConfig.defaultClientMode = config.defaultClientMode;
540
543
  if (config.blockedModels != null) diskConfig.blockedModels = config.blockedModels;
541
544
  if (config.sessionTitles != null) diskConfig.sessionTitles = config.sessionTitles;
542
545
  // Persist the route table (edited from the TUI routes screen).
@@ -800,7 +803,20 @@ async function serverCommand() {
800
803
  prober?.stop();
801
804
  warmer?.stop();
802
805
  eventLoopMonitor.stop();
803
- sidecar?.stop();
806
+ // Started here, awaited at the end, so the sidecar's grace runs BESIDE the
807
+ // rest of this teardown instead of after it.
808
+ //
809
+ // Forcing, because this path already made that choice for everything else:
810
+ // it destroys live streaming connections a few lines down and hard-exits on
811
+ // a timer. The sidecar's first SIGTERM only BEGINS its shutdown, which an
812
+ // in-flight request holds open, so asking politely and leaving is how a
813
+ // copy of it outlived the server and kept its port — the exact orphan
814
+ // Sidecar._reapOrphan then had to clean up on the next start.
815
+ //
816
+ // Never rejects by construction; caught anyway, because an unhandled one
817
+ // between here and the await below would be this process's last word.
818
+ const sidecarStopped = Promise.resolve(
819
+ sidecar?.stop({ graceMs: SIDECAR_SHUTDOWN_GRACE_MS, force: true })).catch(() => {});
804
820
  if (quotaSaveInterval) clearInterval(quotaSaveInterval);
805
821
  await persistQuotaState();
806
822
  // Don't linger waiting on keep-alive / streaming connections: actively
@@ -808,7 +824,12 @@ async function serverCommand() {
808
824
  // short grace period in case anything still hangs.
809
825
  setTimeout(() => process.exit(0), 2000).unref?.();
810
826
  server.closeAllConnections?.();
811
- server.close(() => process.exit(0));
827
+ const closed = new Promise(resolve => server.close(() => resolve(undefined)));
828
+ // Both, not the first of the two. Exiting the moment the listener closed is
829
+ // what walked out on the sidecar mid-kill; the timer above is still what
830
+ // guarantees this ends.
831
+ await Promise.all([closed, sidecarStopped]);
832
+ process.exit(0);
812
833
  }
813
834
  process.on('SIGINT', shutdown);
814
835
  process.on('SIGTERM', shutdown);
@@ -865,7 +886,13 @@ async function serverCommand() {
865
886
  console.log(drained
866
887
  ? `[TeamClaude] Drained in ${(waitedMs / 1000).toFixed(1)}s. Restarting on the new build.`
867
888
  : `[TeamClaude] ${inFlight} request(s) still in flight after ${(waitedMs / 1000).toFixed(0)}s — restarting anyway.`);
868
- sidecar?.stop();
889
+ // Patient, and awaited: the drain above already proved nothing is in
890
+ // flight, so one SIGTERM is all this takes and the wait is only there to
891
+ // bound it. Awaited because the relaunch wants the port a moment later
892
+ // and this is the only place that can hand it over cleanly. No forcing
893
+ // signal — the whole point of this path is to cost the fleet nothing, and
894
+ // a sidecar that somehow outlasts the grace is what _reapOrphan is for.
895
+ await sidecar?.stop({ graceMs: SIDECAR_DRAIN_GRACE_MS });
869
896
  if (quotaSaveInterval) clearInterval(quotaSaveInterval);
870
897
  await persistQuotaState();
871
898
  } catch (err) {
@@ -1147,7 +1174,13 @@ async function envCommand() {
1147
1174
  process.exit(1);
1148
1175
  }
1149
1176
  const port = config.proxy.port;
1150
- const useMitm = !args.slice(1).includes('--no-mitm');
1177
+ let useMitm;
1178
+ try {
1179
+ useMitm = resolveClientMode(config, args.slice(1)) === 'mitm';
1180
+ } catch (/** @type {any} */ err) {
1181
+ process.stderr.write(`teamclaude env: ${err.message}\n`);
1182
+ process.exit(1);
1183
+ }
1151
1184
 
1152
1185
  let caPath = null;
1153
1186
  // The leaf has to name every host MITM will intercept, or the CONNECT for
@@ -1166,6 +1199,9 @@ async function envCommand() {
1166
1199
  inheritedNoProxy: [process.env.NO_PROXY, process.env.no_proxy].filter(Boolean).join(','),
1167
1200
  customModels: config.customModels,
1168
1201
  });
1202
+ // Base-URL mode takes the proxy back OUT of a shell an earlier MITM eval
1203
+ // put it into; only this proxy's own loopback address is unset.
1204
+ if (!useMitm) lines.push(...clearSelfProxyEnvLines(port));
1169
1205
  } catch (err) {
1170
1206
  // A bad proxy.port. Nothing reaches stdout: the shell is eval'ing it.
1171
1207
  process.stderr.write(`teamclaude env: ${err.message} (in ${getConfigPath()})\n`);
@@ -1183,7 +1219,9 @@ async function envCommand() {
1183
1219
  process.stderr.write(`# warning: no account named "${account}" in the config — the proxy will refuse this pin\n`);
1184
1220
  }
1185
1221
  }
1186
- process.stderr.write(`# apply to this shell: eval "$(teamclaude env${useMitm ? '' : ' --no-mitm'})"\n`);
1222
+ // The flag that reproduces this mode whatever the config's default says.
1223
+ process.stderr.write(`# apply to this shell: eval "$(teamclaude env${useMitm ? ' --mitm' : ' --no-mitm'})"\n`);
1224
+ process.stderr.write(`# default mode is ${config.defaultClientMode === 'base-url' ? 'base-URL' : 'MITM'} (config defaultClientMode; the TUI settings screen toggles it)\n`);
1187
1225
  if (!(await isProxyUp(port))) {
1188
1226
  process.stderr.write(`# note: proxy not running on port ${port} — start it with: teamclaude server\n`);
1189
1227
  }
@@ -1200,12 +1238,18 @@ async function runCommand() {
1200
1238
  // Args after 'run'. teamclaude flags (e.g. --no-mitm) are recognized only
1201
1239
  // before an optional `--` separator; everything after `--` goes verbatim to
1202
1240
  // claude. MITM forward-proxy mode is the default so hardcoded api.anthropic.com
1203
- // endpoints are intercepted too; --no-mitm opts back into base-URL-only routing.
1204
- // --mitm is still accepted (now a no-op) for backward compatibility.
1241
+ // endpoints are intercepted too; the config's `defaultClientMode` can make
1242
+ // base-URL the default instead, and --mitm / --no-mitm decide per launch.
1205
1243
  const rest = args.slice(1);
1206
1244
  const sep = rest.indexOf('--');
1207
1245
  const tcFlags = sep >= 0 ? rest.slice(0, sep) : rest;
1208
- const useMitm = !tcFlags.includes('--no-mitm');
1246
+ let useMitm;
1247
+ try {
1248
+ useMitm = resolveClientMode(config, tcFlags) === 'mitm';
1249
+ } catch (/** @type {any} */ err) {
1250
+ console.error(`[TeamClaude] ${err.message}`);
1251
+ process.exit(1);
1252
+ }
1209
1253
  const autoFallback = tcFlags.includes('--auto-fallback');
1210
1254
  const claudeArgs = sep >= 0
1211
1255
  ? rest.slice(sep + 1)
@@ -2283,10 +2327,12 @@ Commands:
2283
2327
  login OAuth login via browser
2284
2328
  login --token OAuth login via copy/paste (no local callback; for headless/remote)
2285
2329
  login --api Add an API key account
2286
- env [--no-mitm] Print export lines to point Claude Code at the proxy, for
2287
- 'eval "$(teamclaude env)"' (MITM forward-proxy by default;
2288
- --no-mitm for base-URL only). Handy for agent multiplexers
2289
- that spawn claude themselves instead of via 'teamclaude run'
2330
+ env [--mitm|--no-mitm]
2331
+ Print export lines to point Claude Code at the proxy, for
2332
+ 'eval "$(teamclaude env)"'. MITM forward-proxy unless the
2333
+ config's defaultClientMode is "base-url"; a flag decides
2334
+ per call. Handy for agent multiplexers that spawn claude
2335
+ themselves instead of via 'teamclaude run'
2290
2336
  run [--no-mitm] [--auto-fallback] [-- args...]
2291
2337
  Run Claude Code through the proxy (errors if it's down,
2292
2338
  unless --auto-fallback launches claude directly instead).
@@ -0,0 +1,131 @@
1
+ // Token usage as OpenAI's Responses API reports it, rewritten into the shape
2
+ // the rest of this proxy books.
3
+ //
4
+ // Every counter here — the account totals, the per-session ones, the per-client
5
+ // ones — speaks Anthropic's usage vocabulary, because that is the only shape
6
+ // this proxy had to read for its first year. Codex traffic is a passthrough, so
7
+ // it arrives in OpenAI's vocabulary instead, and the two disagree in a way that
8
+ // is easy to miss because they share field NAMES:
9
+ //
10
+ // Anthropic `input_tokens` is the UNCACHED input. The cached part is
11
+ // reported separately (`cache_read_input_tokens`,
12
+ // `cache_creation_input_tokens`) and the three are disjoint — the
13
+ // readers here add them up to get a context size or a total.
14
+ //
15
+ // Responses `input_tokens` is the WHOLE prompt, with the cached part named
16
+ // again underneath it as `input_tokens_details.cached_tokens`.
17
+ // Booking it as Anthropic's would count the cached prefix twice
18
+ // once the cache field is also read, and on Codex traffic that
19
+ // prefix is most of the prompt.
20
+ //
21
+ // So the subtraction below is the entire point of this file: the cached tokens
22
+ // move out of `input_tokens` and into `cache_read_input_tokens`, leaving three
23
+ // disjoint figures that every existing reader already knows how to add.
24
+ //
25
+ // `output_tokens_details.reasoning_tokens` is NOT added to the output side —
26
+ // it is a breakdown of `output_tokens`, not a second quantity beside it (a real
27
+ // turn: input 54904, output 132 of which 18 reasoning, total_tokens 55036 =
28
+ // 54904 + 132). `input_tokens_details.cache_write_tokens` is likewise left
29
+ // folded into the uncached remainder rather than mapped onto Anthropic's
30
+ // `cache_creation_input_tokens`: a cache write IS fresh input being processed,
31
+ // OpenAI does not bill it at a premium the way Anthropic does, and inventing a
32
+ // fourth disjoint bucket out of a field whose subset relationship is not stated
33
+ // anywhere would risk a total that no longer adds up.
34
+ //
35
+ // Shape confirmed against OpenAI's generated SDK types (`ResponseUsage`), the
36
+ // Codex CLI's own per-turn records, and the sidecar binary that translates this
37
+ // same body for the hop before us.
38
+
39
+ /** Terminal stream events, the only ones that carry a settled usage object.
40
+ *
41
+ * Named rather than inferred from "the event has a usage object": EVERY
42
+ * `response.*` event carries the whole response envelope, so `response.created`
43
+ * and `response.in_progress` have a `usage` key too — null today. Reading the
44
+ * key would make this proxy's accounting depend on that staying null, and an
45
+ * upstream that started reporting progress figures would be counted twice. A
46
+ * fixed set of terminal names makes a second report unrepresentable instead.
47
+ *
48
+ * `response.failed` is in the set because a failure that got far enough to
49
+ * report figures still spent them; when it carries none — the usual case — the
50
+ * normaliser below returns null and nothing is recorded.
51
+ */
52
+ const TERMINAL_EVENTS = new Set(['response.completed', 'response.incomplete', 'response.failed']);
53
+
54
+ /**
55
+ * A wire figure as a non-negative integer. Everything here is untrusted JSON:
56
+ * a string, a null, a NaN or a negative count must read as "nothing reported"
57
+ * rather than poison a cumulative total that is never recomputed.
58
+ *
59
+ * @param {unknown} v
60
+ * @returns {number}
61
+ */
62
+ function count(v) {
63
+ return typeof v === 'number' && Number.isFinite(v) && v > 0 ? Math.trunc(v) : 0;
64
+ }
65
+
66
+ /**
67
+ * A Responses `usage` object in Anthropic's disjoint shape, or null when there
68
+ * is nothing to record.
69
+ *
70
+ * Null for an absent, non-object or all-zero usage, because the recorders
71
+ * downstream treat a written record as an observation that happened: an
72
+ * all-zero one would claim a report arrived and count against `reports`, which
73
+ * is the one thing that distinguishes "upstream told us nothing" from "upstream
74
+ * told us zero".
75
+ *
76
+ * @param {any} usage
77
+ * @returns {{input_tokens: number, output_tokens: number, cache_read_input_tokens: number}|null}
78
+ */
79
+ export function normalizeResponsesUsage(usage) {
80
+ if (!usage || typeof usage !== 'object') return null;
81
+ const prompt = count(usage.input_tokens);
82
+ const output = count(usage.output_tokens);
83
+ // Clamped to the prompt it is a part of. Upstream cannot report more cached
84
+ // tokens than input tokens, but the uncached remainder is a subtraction, and
85
+ // an unclamped one would go negative and silently shrink a lifetime total.
86
+ const cached = Math.min(count(usage.input_tokens_details?.cached_tokens), prompt);
87
+ if (!prompt && !output) return null;
88
+ return {
89
+ input_tokens: prompt - cached,
90
+ output_tokens: output,
91
+ cache_read_input_tokens: cached,
92
+ };
93
+ }
94
+
95
+ /**
96
+ * Usage from one parsed SSE event of a Responses stream, or null when this
97
+ * event is not a terminal one (the common case — a stream is mostly deltas).
98
+ *
99
+ * @param {any} event
100
+ */
101
+ export function responsesEventUsage(event) {
102
+ if (!event || !TERMINAL_EVENTS.has(event.type)) return null;
103
+ return normalizeResponsesUsage(event.response?.usage);
104
+ }
105
+
106
+ /**
107
+ * Usage from a NON-streaming Responses body, or null when the body is not one.
108
+ *
109
+ * Two independent tells, either of which is enough, and neither of which an
110
+ * Anthropic message body can produce:
111
+ *
112
+ * `object: "response"` what the Responses envelope calls itself.
113
+ * `input_tokens_details` present exactly when a prompt-cache breakdown is
114
+ * reported, which is the case that actually needs
115
+ * rewriting.
116
+ *
117
+ * Both, rather than either alone, because the first is the honest discriminator
118
+ * while the second is the one that matters: a Responses body that reports no
119
+ * cache breakdown at all needs no rewriting — its `input_tokens` is already
120
+ * uncached input — so a body that somehow shows neither tell is read as
121
+ * Anthropic's and lands on exactly the same numbers it would have anyway.
122
+ *
123
+ * @param {any} json
124
+ */
125
+ export function responsesBodyUsage(json) {
126
+ if (!json || typeof json !== 'object') return null;
127
+ const usage = json.usage;
128
+ const tagged = json.object === 'response'
129
+ || (usage != null && typeof usage === 'object' && usage.input_tokens_details != null);
130
+ return tagged ? normalizeResponsesUsage(usage) : null;
131
+ }
package/src/server.js CHANGED
@@ -19,6 +19,7 @@ import { safeLine } from './safe-text.js';
19
19
  import { forwardRefusal, guardedLookup, FORBIDDEN_FORWARD } from './forward-target.js';
20
20
  import { renderDashboardHtml, dashboardCsp } from './dashboard.js';
21
21
  import { createUsageRecorder, resolveUsageDimensions, usageDimensionHeaderNames } from './client-usage.js';
22
+ import { responsesEventUsage, responsesBodyUsage } from './responses-usage.js';
22
23
  import { classificationPath } from './classification-path.js';
23
24
  /** @typedef {import('./types.js').CodedError} CodedError */
24
25
 
@@ -3041,8 +3042,24 @@ export function createSseLineScanner(onLine, maxChars = SSE_MAX_LINE_CHARS) {
3041
3042
  // over. One record per message is what makes double counting unrepresentable
3042
3043
  // rather than merely avoided.
3043
3044
  //
3044
- // Reads one `data:` line. Anthropic's events carry exactly one, so a line is an
3045
- // event for this purpose, and the scanner above never has to hold more.
3045
+ // A Responses stream (the Codex path) instead reports once, at the end, and in
3046
+ // OpenAI's own vocabulary — so it is rewritten into Anthropic's disjoint shape
3047
+ // before it reaches either counter (src/responses-usage.js explains why the two
3048
+ // disagree). It rides this function rather than a parser of its own because the
3049
+ // line is ALREADY parsed here: the branch costs a Set lookup on a string, not a
3050
+ // second pass over the stream. Nothing else would be cheap — a Responses stream
3051
+ // is mostly text deltas, and the settled figures arrive on one event near the end
3052
+ // with no header or marker to find it by.
3053
+ //
3054
+ // Reads one `data:` line. Both dialects carry exactly one per event, so a line
3055
+ // is an event for this purpose, and the scanner above never has to hold more.
3056
+ /**
3057
+ * @param {string} line
3058
+ * @param {number} accountIndex
3059
+ * @param {any} accountManager
3060
+ * @param {((inputTokens: number, outputTokens: number) => void)|null} [onUsage]
3061
+ * @param {Record<string, any>|null} [merged]
3062
+ */
3046
3063
  function parseSSEDataLine(line, accountIndex, accountManager, onUsage = null, merged = null) {
3047
3064
  if (!line.startsWith('data: ')) return;
3048
3065
 
@@ -3056,6 +3073,15 @@ function parseSSEDataLine(line, accountIndex, accountManager, onUsage = null, me
3056
3073
  accountManager.updateUsage(accountIndex, 0, data.usage.output_tokens);
3057
3074
  onUsage?.(0, data.usage.output_tokens || 0);
3058
3075
  if (merged) Object.assign(merged, data.usage);
3076
+ } else {
3077
+ // Both sides settle at once here, so unlike the Anthropic branches above
3078
+ // this is a single incremental update rather than one per side.
3079
+ const usage = responsesEventUsage(data);
3080
+ if (usage) {
3081
+ accountManager.updateUsage(accountIndex, usage.input_tokens, usage.output_tokens);
3082
+ onUsage?.(usage.input_tokens, usage.output_tokens);
3083
+ if (merged) Object.assign(merged, usage);
3084
+ }
3059
3085
  }
3060
3086
  } catch {
3061
3087
  // not valid JSON, skip
@@ -3066,9 +3092,16 @@ function extractUsageFromBody(buffer, accountIndex, accountManager, onUsage = nu
3066
3092
  try {
3067
3093
  const json = JSON.parse(buffer.toString());
3068
3094
  if (json.usage) {
3069
- accountManager.updateUsage(accountIndex, json.usage.input_tokens, json.usage.output_tokens);
3070
- onUsage?.(json.usage.input_tokens || 0, json.usage.output_tokens || 0);
3071
- accountManager.recordTokenUsage(accountIndex, sessionId, model, json.usage);
3095
+ // A buffered Responses body reports under the same two field NAMES with a
3096
+ // different meaning, so reading it as Anthropic's would book the cached
3097
+ // prefix as fresh input and never book it as a cache read at all. Only a
3098
+ // body that says it is one is rewritten; anything else — including a
3099
+ // Responses body carrying no figures to rewrite — falls through to the
3100
+ // reading this had before, unchanged.
3101
+ const usage = responsesBodyUsage(json) || json.usage;
3102
+ accountManager.updateUsage(accountIndex, usage.input_tokens, usage.output_tokens);
3103
+ onUsage?.(usage.input_tokens || 0, usage.output_tokens || 0);
3104
+ accountManager.recordTokenUsage(accountIndex, sessionId, model, usage);
3072
3105
  }
3073
3106
  } catch {
3074
3107
  // not JSON or no usage