@rikcodes/teamclaude 1.1.20-rik.13 → 1.1.20-rik.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +7 -5
- package/package.json +1 -1
- package/src/account-manager.js +64 -2
- package/src/claude-env.js +60 -0
- package/src/config.js +1 -0
- package/src/index.js +60 -14
- package/src/responses-usage.js +131 -0
- package/src/server.js +38 -5
- package/src/sidecar.js +266 -3
- package/src/sync-accounts.js +5 -0
- package/src/tui-remote.js +5 -0
- package/src/tui.js +213 -50
package/README.md
CHANGED
|
@@ -107,10 +107,12 @@ A Claude Code session can use OpenAI models alongside the Claude accounts. They
|
|
|
107
107
|
**1. Install the sidecar and log it into your ChatGPT account** (one time):
|
|
108
108
|
|
|
109
109
|
```bash
|
|
110
|
-
|
|
110
|
+
curl -fsSL https://raw.githubusercontent.com/rikbrown/claude-code-proxy/rik/main/scripts/install.sh | bash
|
|
111
111
|
claude-code-proxy codex auth login
|
|
112
112
|
```
|
|
113
113
|
|
|
114
|
+
This installs [rikbrown/claude-code-proxy](https://github.com/rikbrown/claude-code-proxy), a fork of the reference sidecar. Take the fork rather than upstream's Homebrew build: it forwards Codex's quota headers, without which every bar on the account reads `unknown`, and it lets you raise the 60-second header timeout that otherwise kills a long reasoning turn. See [Timeouts](docs/openai.md#timeouts).
|
|
115
|
+
|
|
114
116
|
**2. Connect it** — add four pieces to `~/.config/teamclaude.json`:
|
|
115
117
|
|
|
116
118
|
```json
|
|
@@ -161,11 +163,11 @@ Each request is routed by the model name in its body, so one session can freely
|
|
|
161
163
|
|
|
162
164
|
1. Check that your sidecar build lists it: `curl -s http://127.0.0.1:18765/v1/models`. The sidecar has its own allow-list and rejects any id that it does not know, regardless of the TeamClaude configuration. Upgrade the sidecar if the id is missing.
|
|
163
165
|
2. Add a `customModels` row. Codex publishes the window for each model as `context_window` in `~/.codex/models_cache.json`; copy it to `contextTokens`.
|
|
164
|
-
3. Start a new `teamclaude run` session. The rows are read at launch, so you do not need to restart the server. If you upgraded the sidecar binary, restart the server — or send `SIGTERM` to the sidecar process and let the supervisor restart it with the new binary.
|
|
166
|
+
3. Start a new `teamclaude run` session. The rows are read at launch, so you do not need to restart the server. If you upgraded the sidecar binary, restart the server — or send `SIGTERM` to the sidecar process and let the supervisor restart it with the new binary. The first `SIGTERM` only starts a graceful shutdown, which a request in flight holds open; send it a second time to force the exit.
|
|
165
167
|
|
|
166
|
-
Claude Code prints one `[claude-code:unrecognized_model]` line to stderr for each custom model. This is expected; suppressing it would lose the correct context window. The quota bars for the sidecar account show `unknown` unless the sidecar forwards Codex's rate-limit headers — see [Quota](docs/openai.md#quota). Keep the sidecar on loopback.
|
|
168
|
+
Claude Code prints one `[claude-code:unrecognized_model]` line to stderr for each custom model. This is expected; suppressing it would lose the correct context window. The quota bars for the sidecar account show `unknown` unless the sidecar forwards Codex's rate-limit headers, which the fork build in step 1 does and upstream's does not — see [Quota](docs/openai.md#quota). Keep the sidecar on loopback.
|
|
167
169
|
|
|
168
|
-
The sidecar appears under the account table as a `⚙` line rather than a row because it holds no subscription, is the only account its route can use, and never rotates. The line also shows its supervised process state (`up pid 98018`, or `down (code 1) 3 restarts`).
|
|
170
|
+
The sidecar appears under the account table as a `⚙` line rather than a row because it holds no subscription, is the only account its route can use, and never rotates. The line also shows its supervised process state (`up pid 98018`, or `down (code 1) 3 restarts`), followed by what the sidecar reports about itself while it is up (`2 active 3 errors`, read from its own `/monitor` endpoint; each is omitted at zero, and both are omitted if that endpoint does not answer).
|
|
169
171
|
|
|
170
172
|
#### Several ChatGPT accounts
|
|
171
173
|
|
|
@@ -221,7 +223,7 @@ Three details matter:
|
|
|
221
223
|
- **`CCP_CODEX_TRANSPORT=http` is required.** A WebSocket upgrade is relayed with the caller's own headers and draws no account, so the WebSocket transport cannot be pooled.
|
|
222
224
|
- **Do not reuse a name across providers.** Routes address accounts by name, so a shared name admits both — including the Claude account that cannot serve `gpt-*`, which outranks the sidecar on priority and wins. TeamClaude warns at startup when it sees one.
|
|
223
225
|
|
|
224
|
-
|
|
226
|
+
One thing differs from the single-account setup: each turn appears **twice** in the activity list, once per hop. Tokens and quota bars both come from the second hop, where the subscription is, so they are booked against the ChatGPT account that served. The sidecar account holds a stub login and spends nothing of its own, so it reads `N req · 0 tok` — booking it as well would double every figure.
|
|
225
227
|
|
|
226
228
|
Full details, including what happens to quota on each hop: [Several ChatGPT accounts behind one sidecar](docs/openai.md#several-chatgpt-accounts-behind-one-sidecar).
|
|
227
229
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@rikcodes/teamclaude",
|
|
3
|
-
"version": "1.1.20-rik.
|
|
3
|
+
"version": "1.1.20-rik.15",
|
|
4
4
|
"description": "Multi-account proxy for Claude Code and Codex: pools Claude Max, ChatGPT/Codex, API-key and third-party backend accounts, and rotates on quota",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "src/index.js",
|
package/src/account-manager.js
CHANGED
|
@@ -214,6 +214,13 @@ function makeAccount(acct, index) {
|
|
|
214
214
|
hasClaudeMax: acct.hasClaudeMax ?? null,
|
|
215
215
|
hasClaudePro: acct.hasClaudePro ?? null,
|
|
216
216
|
priority: acct.priority || 0,
|
|
217
|
+
// Where this account sits in the list the TUI draws, arranged by the
|
|
218
|
+
// operator from the settings screen. Presentation only: nothing but
|
|
219
|
+
// _displayOrder in tui.js reads it, and it is emphatically NOT `priority`
|
|
220
|
+
// above — that one decides which account rotation spends next, and the two
|
|
221
|
+
// answer different questions about the same fleet. `null` means never
|
|
222
|
+
// placed, which sorts after every account that has been.
|
|
223
|
+
displayOrder: Number.isFinite(acct.displayOrder) ? acct.displayOrder : null,
|
|
217
224
|
disabled: acct.disabled || false,
|
|
218
225
|
maxUsage: acct.maxUsage ?? null,
|
|
219
226
|
upstream: acct.upstream || null,
|
|
@@ -3485,7 +3492,9 @@ export class AccountManager {
|
|
|
3485
3492
|
// threshold and every request fails, while a sibling sits at 0%.
|
|
3486
3493
|
// A standalone sidecar (no Codex accounts here) is NOT a conduit: it holds
|
|
3487
3494
|
// its own login, the forwarded numbers are its own, and they still apply.
|
|
3488
|
-
|
|
3495
|
+
// Shares its definition with the token counters, which drop a conduit hop
|
|
3496
|
+
// for the same reason: what it reports belongs to whoever served.
|
|
3497
|
+
if (!this._isCodexConduit(account)) {
|
|
3489
3498
|
for (const window of ['primary', 'secondary']) {
|
|
3490
3499
|
const used = parseFloat(headers[`x-codex-${window}-used-percent`]);
|
|
3491
3500
|
const minutes = parseInt(headers[`x-codex-${window}-window-minutes`], 10);
|
|
@@ -3556,12 +3565,49 @@ export class AccountManager {
|
|
|
3556
3565
|
}
|
|
3557
3566
|
}
|
|
3558
3567
|
|
|
3568
|
+
/**
|
|
3569
|
+
* Whether `account` merely RELAYS to this fleet's own Codex pool instead of
|
|
3570
|
+
* holding a subscription of its own: a translating sidecar whose back leg is
|
|
3571
|
+
* pointed back at this proxy, so each turn crosses this process twice — once
|
|
3572
|
+
* inbound on `/v1/messages` and once outbound on `/backend-api/codex/*`
|
|
3573
|
+
* (docs/openai.md, "Several ChatGPT accounts behind one sidecar").
|
|
3574
|
+
*
|
|
3575
|
+
* Keyed on three things, all of which the documented setup has and no other
|
|
3576
|
+
* account does:
|
|
3577
|
+
*
|
|
3578
|
+
* a loopback upstream the sidecar runs on this machine.
|
|
3579
|
+
* the Anthropic wire a conduit is reached on `/v1/messages`, so it
|
|
3580
|
+
* carries no `provider` field — docs/openai.md warns
|
|
3581
|
+
* that giving one `"provider": "codex"` puts it in the
|
|
3582
|
+
* foreign-subscription partition and every `gpt-*`
|
|
3583
|
+
* request then fails to find an account. A pooled
|
|
3584
|
+
* ChatGPT account is therefore never a conduit,
|
|
3585
|
+
* whatever its upstream says; it IS the pool.
|
|
3586
|
+
* a pool to relay to a standalone sidecar holding its own ChatGPT login
|
|
3587
|
+
* is NOT a conduit: everything it reports is its own,
|
|
3588
|
+
* and it is the only hop there is.
|
|
3589
|
+
*
|
|
3590
|
+
* Still imprecise in one direction, deliberately: a fleet running BOTH a
|
|
3591
|
+
* self-hosting local Anthropic backend and Codex accounts reads the former as
|
|
3592
|
+
* a conduit. Narrowing that would mean correlating the two hops of one turn,
|
|
3593
|
+
* which nothing here can do — they are separate requests sharing no id — and
|
|
3594
|
+
* the cost of the false positive is a row that under-reports rather than one
|
|
3595
|
+
* that misroutes.
|
|
3596
|
+
*
|
|
3597
|
+
* @param {any} account
|
|
3598
|
+
*/
|
|
3599
|
+
_isCodexConduit(account) {
|
|
3600
|
+
return isLocalUpstream(account)
|
|
3601
|
+
&& providerOf(account) === DEFAULT_PROVIDER
|
|
3602
|
+
&& this.accounts.some(a => providerOf(a) === 'codex');
|
|
3603
|
+
}
|
|
3604
|
+
|
|
3559
3605
|
/**
|
|
3560
3606
|
* Update cumulative token usage from response body data.
|
|
3561
3607
|
*/
|
|
3562
3608
|
updateUsage(accountIndex, inputTokens, outputTokens) {
|
|
3563
3609
|
const account = this.accounts[accountIndex];
|
|
3564
|
-
if (!account) return;
|
|
3610
|
+
if (!account || this._isCodexConduit(account)) return;
|
|
3565
3611
|
if (inputTokens) account.usage.totalInputTokens += inputTokens;
|
|
3566
3612
|
if (outputTokens) account.usage.totalOutputTokens += outputTokens;
|
|
3567
3613
|
}
|
|
@@ -3581,6 +3627,22 @@ export class AccountManager {
|
|
|
3581
3627
|
*/
|
|
3582
3628
|
recordTokenUsage(accountIndex, sessionId, model, usage) {
|
|
3583
3629
|
if (!usage) return;
|
|
3630
|
+
// A conduit hop is the SAME tokens, translated: the sidecar rebuilt this
|
|
3631
|
+
// report out of the Responses usage the pool sent it, and the pool's own hop
|
|
3632
|
+
// records the original a moment later. Both scopes would double — the two
|
|
3633
|
+
// account rows are distinct, but a session is one row and carries the same
|
|
3634
|
+
// id on both hops, so its context and spend would read twice the truth.
|
|
3635
|
+
//
|
|
3636
|
+
// Dropped rather than deduplicated because there is nothing to deduplicate
|
|
3637
|
+
// against: the two hops are separate requests that share no id, and the
|
|
3638
|
+
// conduit spends nothing of its own anyway (its login is a stub). Booking at
|
|
3639
|
+
// the hop that really spent is what keeps one turn one record.
|
|
3640
|
+
//
|
|
3641
|
+
// Only the ACCOUNT-scoped and SESSION-scoped counters stop here. Per-client
|
|
3642
|
+
// attribution still runs on the inbound hop, in the caller, because that is
|
|
3643
|
+
// the only hop that can see who asked: the outbound one comes from the
|
|
3644
|
+
// sidecar on loopback and carries no client identity at all.
|
|
3645
|
+
if (this._isCodexConduit(this.accounts[accountIndex])) return;
|
|
3584
3646
|
// The same resolver routing uses, so a token total and a routing decision
|
|
3585
3647
|
// agree about which family a request belonged to. Resolved here rather than
|
|
3586
3648
|
// at the call sites: they parse a wire format and have no business knowing
|
package/src/claude-env.js
CHANGED
|
@@ -76,6 +76,66 @@ export function mergeNoProxy(...inherited) {
|
|
|
76
76
|
return out.join(',');
|
|
77
77
|
}
|
|
78
78
|
|
|
79
|
+
/** The two ways a launched client can reach the proxy. */
|
|
80
|
+
export const CLIENT_MODES = ['mitm', 'base-url'];
|
|
81
|
+
|
|
82
|
+
/**
|
|
83
|
+
* Which mode `run` and `env` use: a `--mitm` or `--no-mitm` flag decides, else
|
|
84
|
+
* the config's `defaultClientMode`, else MITM.
|
|
85
|
+
*
|
|
86
|
+
* MITM routes every request of the launched client through the proxy, so the
|
|
87
|
+
* hard-coded api.anthropic.com endpoints and the Codex CLI (which honours only
|
|
88
|
+
* proxy variables) are covered. It is also the whole point of the setting: with
|
|
89
|
+
* `eval "$(teamclaude env)"` the proxy variables are shell-wide, and every other
|
|
90
|
+
* tool in that shell — gh, git, a package manager — follows them to a listener
|
|
91
|
+
* that only speaks to two hosts (#382). An operator who lives in such a shell
|
|
92
|
+
* sets `defaultClientMode: "base-url"` once and opts back in per launch.
|
|
93
|
+
*
|
|
94
|
+
* @param {{ defaultClientMode?: string }|null|undefined} config
|
|
95
|
+
* @param {string[]} flags the invocation's own arguments
|
|
96
|
+
* @returns {'mitm'|'base-url'}
|
|
97
|
+
*/
|
|
98
|
+
export function resolveClientMode(config, flags) {
|
|
99
|
+
const mitm = flags.includes('--mitm');
|
|
100
|
+
const noMitm = flags.includes('--no-mitm');
|
|
101
|
+
if (mitm && noMitm) throw new Error('choose either --mitm or --no-mitm');
|
|
102
|
+
if (mitm) return 'mitm';
|
|
103
|
+
if (noMitm) return 'base-url';
|
|
104
|
+
return config?.defaultClientMode === 'base-url' ? 'base-url' : 'mitm';
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
const SHELL_PROXY_VARS = ['HTTP_PROXY', 'HTTPS_PROXY', 'ALL_PROXY', 'http_proxy', 'https_proxy', 'all_proxy'];
|
|
108
|
+
|
|
109
|
+
/**
|
|
110
|
+
* @param {unknown} value
|
|
111
|
+
* @param {number} port
|
|
112
|
+
*/
|
|
113
|
+
function pointsAtLoopback(value, port) {
|
|
114
|
+
if (!value) return false;
|
|
115
|
+
try {
|
|
116
|
+
const url = new URL(String(value));
|
|
117
|
+
const host = url.hostname.replace(/^\[|\]$/g, '').toLowerCase();
|
|
118
|
+
return ['127.0.0.1', 'localhost', '::1'].includes(host) && Number(url.port || 80) === port;
|
|
119
|
+
} catch {
|
|
120
|
+
return false;
|
|
121
|
+
}
|
|
122
|
+
}
|
|
123
|
+
|
|
124
|
+
/**
|
|
125
|
+
* `unset` lines for proxy variables a previous MITM-mode eval left in the
|
|
126
|
+
* shell, so re-evaluating in base-URL mode takes the proxy back out of it. Only
|
|
127
|
+
* a value naming THIS proxy's loopback port is touched: a real corporate proxy
|
|
128
|
+
* in the same variables is the operator's and stays.
|
|
129
|
+
*
|
|
130
|
+
* @param {unknown} port
|
|
131
|
+
* @param {NodeJS.ProcessEnv} [env]
|
|
132
|
+
* @returns {string[]}
|
|
133
|
+
*/
|
|
134
|
+
export function clearSelfProxyEnvLines(port, env = process.env) {
|
|
135
|
+
const checkedPort = validPort(port);
|
|
136
|
+
return SHELL_PROXY_VARS.filter(name => pointsAtLoopback(env[name], checkedPort)).map(name => `unset ${name}`);
|
|
137
|
+
}
|
|
138
|
+
|
|
79
139
|
// Build the shell `export` lines that point Claude Code — or any tool that
|
|
80
140
|
// spawns it, e.g. an agent multiplexer — at the proxy. This is the same
|
|
81
141
|
// environment `teamclaude run` sets up, but emitted for `eval "$(teamclaude
|
package/src/config.js
CHANGED
|
@@ -101,6 +101,7 @@ export function createDefaultConfig() {
|
|
|
101
101
|
sessionTitles: { enabled: false, width: 18 },
|
|
102
102
|
projection: { enabled: true, windowMinutes: 90, wasteFloor: 0.1 },
|
|
103
103
|
eventLogging: 'hide',
|
|
104
|
+
defaultClientMode: 'mitm',
|
|
104
105
|
activityLog: null,
|
|
105
106
|
autoRestart: false,
|
|
106
107
|
blockedModels: [],
|
package/src/index.js
CHANGED
|
@@ -31,7 +31,7 @@ import { ensureCerts, mitmHosts } from './mitm.js';
|
|
|
31
31
|
import { Prober } from './prober.js';
|
|
32
32
|
import { Warmer } from './warmer.js';
|
|
33
33
|
import { createRollingWarmupSchedule, formatWarmupScheduleConfirmation, resolveWarmupConfig, resolveWarmupSchedule } from './warmup-schedule.js';
|
|
34
|
-
import { Sidecar } from './sidecar.js';
|
|
34
|
+
import { Sidecar, SIDECAR_DRAIN_GRACE_MS, SIDECAR_SHUTDOWN_GRACE_MS } from './sidecar.js';
|
|
35
35
|
import { TUI } from './tui.js';
|
|
36
36
|
import { SessionTitles } from './session-titles.js';
|
|
37
37
|
import { RemoteControl, createAttachSession } from './tui-remote.js';
|
|
@@ -42,7 +42,7 @@ import { createVersionSource, UpdateWatcher } from './update-watch.js';
|
|
|
42
42
|
import { renderStatus, formatPercent } from './status-renderer.js';
|
|
43
43
|
import { sanitizeText } from './safe-text.js';
|
|
44
44
|
import { ClientUsageTracker, UsageDimensionTracker } from './client-usage.js';
|
|
45
|
-
import { buildClaudeEnvLines, buildCustomModelAgents, buildCustomModelSettings, buildCustomModelVars, bypassesAllHosts, encodePinComponent, mergeNoProxy } from './claude-env.js';
|
|
45
|
+
import { buildClaudeEnvLines, buildCustomModelAgents, buildCustomModelSettings, buildCustomModelVars, bypassesAllHosts, clearSelfProxyEnvLines, encodePinComponent, mergeNoProxy, resolveClientMode } from './claude-env.js';
|
|
46
46
|
import { serviceKind, installService, uninstallService, serviceStatus, renderService, logPath } from './service.js';
|
|
47
47
|
import { formatTerminalTitle, titleSequence, TITLE_STACK_PUSH, TITLE_STACK_POP } from './terminal-title.js';
|
|
48
48
|
import { getUpstreamProxy, describeProxy, describeSelfProxy } from './upstream-proxy.js';
|
|
@@ -474,6 +474,8 @@ async function serverCommand() {
|
|
|
474
474
|
// Both are read per request off this object (server.js) and the TUI already
|
|
475
475
|
// persists them; without this a hand edit or another writer waited for a restart.
|
|
476
476
|
config.eventLogging = diskConfig.eventLogging || 'hide';
|
|
477
|
+
// Read by `run`/`env` from disk, but the TUI settings screen shows it live.
|
|
478
|
+
config.defaultClientMode = diskConfig.defaultClientMode === 'base-url' ? 'base-url' : 'mitm';
|
|
477
479
|
config.blockedModels = Array.isArray(diskConfig.blockedModels) ? diskConfig.blockedModels : [];
|
|
478
480
|
config.projection = diskConfig.projection;
|
|
479
481
|
accountManager.setProjection(config.projection);
|
|
@@ -537,6 +539,7 @@ async function serverCommand() {
|
|
|
537
539
|
// screen too; the server reads them live from `config`, but without this
|
|
538
540
|
// the edit never reached disk and was silently undone by the next start.
|
|
539
541
|
if (config.eventLogging != null) diskConfig.eventLogging = config.eventLogging;
|
|
542
|
+
if (config.defaultClientMode != null) diskConfig.defaultClientMode = config.defaultClientMode;
|
|
540
543
|
if (config.blockedModels != null) diskConfig.blockedModels = config.blockedModels;
|
|
541
544
|
if (config.sessionTitles != null) diskConfig.sessionTitles = config.sessionTitles;
|
|
542
545
|
// Persist the route table (edited from the TUI routes screen).
|
|
@@ -800,7 +803,20 @@ async function serverCommand() {
|
|
|
800
803
|
prober?.stop();
|
|
801
804
|
warmer?.stop();
|
|
802
805
|
eventLoopMonitor.stop();
|
|
803
|
-
sidecar
|
|
806
|
+
// Started here, awaited at the end, so the sidecar's grace runs BESIDE the
|
|
807
|
+
// rest of this teardown instead of after it.
|
|
808
|
+
//
|
|
809
|
+
// Forcing, because this path already made that choice for everything else:
|
|
810
|
+
// it destroys live streaming connections a few lines down and hard-exits on
|
|
811
|
+
// a timer. The sidecar's first SIGTERM only BEGINS its shutdown, which an
|
|
812
|
+
// in-flight request holds open, so asking politely and leaving is how a
|
|
813
|
+
// copy of it outlived the server and kept its port — the exact orphan
|
|
814
|
+
// Sidecar._reapOrphan then had to clean up on the next start.
|
|
815
|
+
//
|
|
816
|
+
// Never rejects by construction; caught anyway, because an unhandled one
|
|
817
|
+
// between here and the await below would be this process's last word.
|
|
818
|
+
const sidecarStopped = Promise.resolve(
|
|
819
|
+
sidecar?.stop({ graceMs: SIDECAR_SHUTDOWN_GRACE_MS, force: true })).catch(() => {});
|
|
804
820
|
if (quotaSaveInterval) clearInterval(quotaSaveInterval);
|
|
805
821
|
await persistQuotaState();
|
|
806
822
|
// Don't linger waiting on keep-alive / streaming connections: actively
|
|
@@ -808,7 +824,12 @@ async function serverCommand() {
|
|
|
808
824
|
// short grace period in case anything still hangs.
|
|
809
825
|
setTimeout(() => process.exit(0), 2000).unref?.();
|
|
810
826
|
server.closeAllConnections?.();
|
|
811
|
-
server.close(() =>
|
|
827
|
+
const closed = new Promise(resolve => server.close(() => resolve(undefined)));
|
|
828
|
+
// Both, not the first of the two. Exiting the moment the listener closed is
|
|
829
|
+
// what walked out on the sidecar mid-kill; the timer above is still what
|
|
830
|
+
// guarantees this ends.
|
|
831
|
+
await Promise.all([closed, sidecarStopped]);
|
|
832
|
+
process.exit(0);
|
|
812
833
|
}
|
|
813
834
|
process.on('SIGINT', shutdown);
|
|
814
835
|
process.on('SIGTERM', shutdown);
|
|
@@ -865,7 +886,13 @@ async function serverCommand() {
|
|
|
865
886
|
console.log(drained
|
|
866
887
|
? `[TeamClaude] Drained in ${(waitedMs / 1000).toFixed(1)}s. Restarting on the new build.`
|
|
867
888
|
: `[TeamClaude] ${inFlight} request(s) still in flight after ${(waitedMs / 1000).toFixed(0)}s — restarting anyway.`);
|
|
868
|
-
|
|
889
|
+
// Patient, and awaited: the drain above already proved nothing is in
|
|
890
|
+
// flight, so one SIGTERM is all this takes and the wait is only there to
|
|
891
|
+
// bound it. Awaited because the relaunch wants the port a moment later
|
|
892
|
+
// and this is the only place that can hand it over cleanly. No forcing
|
|
893
|
+
// signal — the whole point of this path is to cost the fleet nothing, and
|
|
894
|
+
// a sidecar that somehow outlasts the grace is what _reapOrphan is for.
|
|
895
|
+
await sidecar?.stop({ graceMs: SIDECAR_DRAIN_GRACE_MS });
|
|
869
896
|
if (quotaSaveInterval) clearInterval(quotaSaveInterval);
|
|
870
897
|
await persistQuotaState();
|
|
871
898
|
} catch (err) {
|
|
@@ -1147,7 +1174,13 @@ async function envCommand() {
|
|
|
1147
1174
|
process.exit(1);
|
|
1148
1175
|
}
|
|
1149
1176
|
const port = config.proxy.port;
|
|
1150
|
-
|
|
1177
|
+
let useMitm;
|
|
1178
|
+
try {
|
|
1179
|
+
useMitm = resolveClientMode(config, args.slice(1)) === 'mitm';
|
|
1180
|
+
} catch (/** @type {any} */ err) {
|
|
1181
|
+
process.stderr.write(`teamclaude env: ${err.message}\n`);
|
|
1182
|
+
process.exit(1);
|
|
1183
|
+
}
|
|
1151
1184
|
|
|
1152
1185
|
let caPath = null;
|
|
1153
1186
|
// The leaf has to name every host MITM will intercept, or the CONNECT for
|
|
@@ -1166,6 +1199,9 @@ async function envCommand() {
|
|
|
1166
1199
|
inheritedNoProxy: [process.env.NO_PROXY, process.env.no_proxy].filter(Boolean).join(','),
|
|
1167
1200
|
customModels: config.customModels,
|
|
1168
1201
|
});
|
|
1202
|
+
// Base-URL mode takes the proxy back OUT of a shell an earlier MITM eval
|
|
1203
|
+
// put it into; only this proxy's own loopback address is unset.
|
|
1204
|
+
if (!useMitm) lines.push(...clearSelfProxyEnvLines(port));
|
|
1169
1205
|
} catch (err) {
|
|
1170
1206
|
// A bad proxy.port. Nothing reaches stdout: the shell is eval'ing it.
|
|
1171
1207
|
process.stderr.write(`teamclaude env: ${err.message} (in ${getConfigPath()})\n`);
|
|
@@ -1183,7 +1219,9 @@ async function envCommand() {
|
|
|
1183
1219
|
process.stderr.write(`# warning: no account named "${account}" in the config — the proxy will refuse this pin\n`);
|
|
1184
1220
|
}
|
|
1185
1221
|
}
|
|
1186
|
-
|
|
1222
|
+
// The flag that reproduces this mode whatever the config's default says.
|
|
1223
|
+
process.stderr.write(`# apply to this shell: eval "$(teamclaude env${useMitm ? ' --mitm' : ' --no-mitm'})"\n`);
|
|
1224
|
+
process.stderr.write(`# default mode is ${config.defaultClientMode === 'base-url' ? 'base-URL' : 'MITM'} (config defaultClientMode; the TUI settings screen toggles it)\n`);
|
|
1187
1225
|
if (!(await isProxyUp(port))) {
|
|
1188
1226
|
process.stderr.write(`# note: proxy not running on port ${port} — start it with: teamclaude server\n`);
|
|
1189
1227
|
}
|
|
@@ -1200,12 +1238,18 @@ async function runCommand() {
|
|
|
1200
1238
|
// Args after 'run'. teamclaude flags (e.g. --no-mitm) are recognized only
|
|
1201
1239
|
// before an optional `--` separator; everything after `--` goes verbatim to
|
|
1202
1240
|
// claude. MITM forward-proxy mode is the default so hardcoded api.anthropic.com
|
|
1203
|
-
// endpoints are intercepted too;
|
|
1204
|
-
//
|
|
1241
|
+
// endpoints are intercepted too; the config's `defaultClientMode` can make
|
|
1242
|
+
// base-URL the default instead, and --mitm / --no-mitm decide per launch.
|
|
1205
1243
|
const rest = args.slice(1);
|
|
1206
1244
|
const sep = rest.indexOf('--');
|
|
1207
1245
|
const tcFlags = sep >= 0 ? rest.slice(0, sep) : rest;
|
|
1208
|
-
|
|
1246
|
+
let useMitm;
|
|
1247
|
+
try {
|
|
1248
|
+
useMitm = resolveClientMode(config, tcFlags) === 'mitm';
|
|
1249
|
+
} catch (/** @type {any} */ err) {
|
|
1250
|
+
console.error(`[TeamClaude] ${err.message}`);
|
|
1251
|
+
process.exit(1);
|
|
1252
|
+
}
|
|
1209
1253
|
const autoFallback = tcFlags.includes('--auto-fallback');
|
|
1210
1254
|
const claudeArgs = sep >= 0
|
|
1211
1255
|
? rest.slice(sep + 1)
|
|
@@ -2283,10 +2327,12 @@ Commands:
|
|
|
2283
2327
|
login OAuth login via browser
|
|
2284
2328
|
login --token OAuth login via copy/paste (no local callback; for headless/remote)
|
|
2285
2329
|
login --api Add an API key account
|
|
2286
|
-
env [--no-mitm]
|
|
2287
|
-
|
|
2288
|
-
|
|
2289
|
-
|
|
2330
|
+
env [--mitm|--no-mitm]
|
|
2331
|
+
Print export lines to point Claude Code at the proxy, for
|
|
2332
|
+
'eval "$(teamclaude env)"'. MITM forward-proxy unless the
|
|
2333
|
+
config's defaultClientMode is "base-url"; a flag decides
|
|
2334
|
+
per call. Handy for agent multiplexers that spawn claude
|
|
2335
|
+
themselves instead of via 'teamclaude run'
|
|
2290
2336
|
run [--no-mitm] [--auto-fallback] [-- args...]
|
|
2291
2337
|
Run Claude Code through the proxy (errors if it's down,
|
|
2292
2338
|
unless --auto-fallback launches claude directly instead).
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
// Token usage as OpenAI's Responses API reports it, rewritten into the shape
|
|
2
|
+
// the rest of this proxy books.
|
|
3
|
+
//
|
|
4
|
+
// Every counter here — the account totals, the per-session ones, the per-client
|
|
5
|
+
// ones — speaks Anthropic's usage vocabulary, because that is the only shape
|
|
6
|
+
// this proxy had to read for its first year. Codex traffic is a passthrough, so
|
|
7
|
+
// it arrives in OpenAI's vocabulary instead, and the two disagree in a way that
|
|
8
|
+
// is easy to miss because they share field NAMES:
|
|
9
|
+
//
|
|
10
|
+
// Anthropic `input_tokens` is the UNCACHED input. The cached part is
|
|
11
|
+
// reported separately (`cache_read_input_tokens`,
|
|
12
|
+
// `cache_creation_input_tokens`) and the three are disjoint — the
|
|
13
|
+
// readers here add them up to get a context size or a total.
|
|
14
|
+
//
|
|
15
|
+
// Responses `input_tokens` is the WHOLE prompt, with the cached part named
|
|
16
|
+
// again underneath it as `input_tokens_details.cached_tokens`.
|
|
17
|
+
// Booking it as Anthropic's would count the cached prefix twice
|
|
18
|
+
// once the cache field is also read, and on Codex traffic that
|
|
19
|
+
// prefix is most of the prompt.
|
|
20
|
+
//
|
|
21
|
+
// So the subtraction below is the entire point of this file: the cached tokens
|
|
22
|
+
// move out of `input_tokens` and into `cache_read_input_tokens`, leaving three
|
|
23
|
+
// disjoint figures that every existing reader already knows how to add.
|
|
24
|
+
//
|
|
25
|
+
// `output_tokens_details.reasoning_tokens` is NOT added to the output side —
|
|
26
|
+
// it is a breakdown of `output_tokens`, not a second quantity beside it (a real
|
|
27
|
+
// turn: input 54904, output 132 of which 18 reasoning, total_tokens 55036 =
|
|
28
|
+
// 54904 + 132). `input_tokens_details.cache_write_tokens` is likewise left
|
|
29
|
+
// folded into the uncached remainder rather than mapped onto Anthropic's
|
|
30
|
+
// `cache_creation_input_tokens`: a cache write IS fresh input being processed,
|
|
31
|
+
// OpenAI does not bill it at a premium the way Anthropic does, and inventing a
|
|
32
|
+
// fourth disjoint bucket out of a field whose subset relationship is not stated
|
|
33
|
+
// anywhere would risk a total that no longer adds up.
|
|
34
|
+
//
|
|
35
|
+
// Shape confirmed against OpenAI's generated SDK types (`ResponseUsage`), the
|
|
36
|
+
// Codex CLI's own per-turn records, and the sidecar binary that translates this
|
|
37
|
+
// same body for the hop before us.
|
|
38
|
+
|
|
39
|
+
/** Terminal stream events, the only ones that carry a settled usage object.
|
|
40
|
+
*
|
|
41
|
+
* Named rather than inferred from "the event has a usage object": EVERY
|
|
42
|
+
* `response.*` event carries the whole response envelope, so `response.created`
|
|
43
|
+
* and `response.in_progress` have a `usage` key too — null today. Reading the
|
|
44
|
+
* key would make this proxy's accounting depend on that staying null, and an
|
|
45
|
+
* upstream that started reporting progress figures would be counted twice. A
|
|
46
|
+
* fixed set of terminal names makes a second report unrepresentable instead.
|
|
47
|
+
*
|
|
48
|
+
* `response.failed` is in the set because a failure that got far enough to
|
|
49
|
+
* report figures still spent them; when it carries none — the usual case — the
|
|
50
|
+
* normaliser below returns null and nothing is recorded.
|
|
51
|
+
*/
|
|
52
|
+
const TERMINAL_EVENTS = new Set(['response.completed', 'response.incomplete', 'response.failed']);
|
|
53
|
+
|
|
54
|
+
/**
|
|
55
|
+
* A wire figure as a non-negative integer. Everything here is untrusted JSON:
|
|
56
|
+
* a string, a null, a NaN or a negative count must read as "nothing reported"
|
|
57
|
+
* rather than poison a cumulative total that is never recomputed.
|
|
58
|
+
*
|
|
59
|
+
* @param {unknown} v
|
|
60
|
+
* @returns {number}
|
|
61
|
+
*/
|
|
62
|
+
function count(v) {
|
|
63
|
+
return typeof v === 'number' && Number.isFinite(v) && v > 0 ? Math.trunc(v) : 0;
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
/**
|
|
67
|
+
* A Responses `usage` object in Anthropic's disjoint shape, or null when there
|
|
68
|
+
* is nothing to record.
|
|
69
|
+
*
|
|
70
|
+
* Null for an absent, non-object or all-zero usage, because the recorders
|
|
71
|
+
* downstream treat a written record as an observation that happened: an
|
|
72
|
+
* all-zero one would claim a report arrived and count against `reports`, which
|
|
73
|
+
* is the one thing that distinguishes "upstream told us nothing" from "upstream
|
|
74
|
+
* told us zero".
|
|
75
|
+
*
|
|
76
|
+
* @param {any} usage
|
|
77
|
+
* @returns {{input_tokens: number, output_tokens: number, cache_read_input_tokens: number}|null}
|
|
78
|
+
*/
|
|
79
|
+
export function normalizeResponsesUsage(usage) {
|
|
80
|
+
if (!usage || typeof usage !== 'object') return null;
|
|
81
|
+
const prompt = count(usage.input_tokens);
|
|
82
|
+
const output = count(usage.output_tokens);
|
|
83
|
+
// Clamped to the prompt it is a part of. Upstream cannot report more cached
|
|
84
|
+
// tokens than input tokens, but the uncached remainder is a subtraction, and
|
|
85
|
+
// an unclamped one would go negative and silently shrink a lifetime total.
|
|
86
|
+
const cached = Math.min(count(usage.input_tokens_details?.cached_tokens), prompt);
|
|
87
|
+
if (!prompt && !output) return null;
|
|
88
|
+
return {
|
|
89
|
+
input_tokens: prompt - cached,
|
|
90
|
+
output_tokens: output,
|
|
91
|
+
cache_read_input_tokens: cached,
|
|
92
|
+
};
|
|
93
|
+
}
|
|
94
|
+
|
|
95
|
+
/**
|
|
96
|
+
* Usage from one parsed SSE event of a Responses stream, or null when this
|
|
97
|
+
* event is not a terminal one (the common case — a stream is mostly deltas).
|
|
98
|
+
*
|
|
99
|
+
* @param {any} event
|
|
100
|
+
*/
|
|
101
|
+
export function responsesEventUsage(event) {
|
|
102
|
+
if (!event || !TERMINAL_EVENTS.has(event.type)) return null;
|
|
103
|
+
return normalizeResponsesUsage(event.response?.usage);
|
|
104
|
+
}
|
|
105
|
+
|
|
106
|
+
/**
|
|
107
|
+
* Usage from a NON-streaming Responses body, or null when the body is not one.
|
|
108
|
+
*
|
|
109
|
+
* Two independent tells, either of which is enough, and neither of which an
|
|
110
|
+
* Anthropic message body can produce:
|
|
111
|
+
*
|
|
112
|
+
* `object: "response"` what the Responses envelope calls itself.
|
|
113
|
+
* `input_tokens_details` present exactly when a prompt-cache breakdown is
|
|
114
|
+
* reported, which is the case that actually needs
|
|
115
|
+
* rewriting.
|
|
116
|
+
*
|
|
117
|
+
* Both, rather than either alone, because the first is the honest discriminator
|
|
118
|
+
* while the second is the one that matters: a Responses body that reports no
|
|
119
|
+
* cache breakdown at all needs no rewriting — its `input_tokens` is already
|
|
120
|
+
* uncached input — so a body that somehow shows neither tell is read as
|
|
121
|
+
* Anthropic's and lands on exactly the same numbers it would have anyway.
|
|
122
|
+
*
|
|
123
|
+
* @param {any} json
|
|
124
|
+
*/
|
|
125
|
+
export function responsesBodyUsage(json) {
|
|
126
|
+
if (!json || typeof json !== 'object') return null;
|
|
127
|
+
const usage = json.usage;
|
|
128
|
+
const tagged = json.object === 'response'
|
|
129
|
+
|| (usage != null && typeof usage === 'object' && usage.input_tokens_details != null);
|
|
130
|
+
return tagged ? normalizeResponsesUsage(usage) : null;
|
|
131
|
+
}
|
package/src/server.js
CHANGED
|
@@ -19,6 +19,7 @@ import { safeLine } from './safe-text.js';
|
|
|
19
19
|
import { forwardRefusal, guardedLookup, FORBIDDEN_FORWARD } from './forward-target.js';
|
|
20
20
|
import { renderDashboardHtml, dashboardCsp } from './dashboard.js';
|
|
21
21
|
import { createUsageRecorder, resolveUsageDimensions, usageDimensionHeaderNames } from './client-usage.js';
|
|
22
|
+
import { responsesEventUsage, responsesBodyUsage } from './responses-usage.js';
|
|
22
23
|
import { classificationPath } from './classification-path.js';
|
|
23
24
|
/** @typedef {import('./types.js').CodedError} CodedError */
|
|
24
25
|
|
|
@@ -3041,8 +3042,24 @@ export function createSseLineScanner(onLine, maxChars = SSE_MAX_LINE_CHARS) {
|
|
|
3041
3042
|
// over. One record per message is what makes double counting unrepresentable
|
|
3042
3043
|
// rather than merely avoided.
|
|
3043
3044
|
//
|
|
3044
|
-
//
|
|
3045
|
-
//
|
|
3045
|
+
// A Responses stream (the Codex path) instead reports once, at the end, and in
|
|
3046
|
+
// OpenAI's own vocabulary — so it is rewritten into Anthropic's disjoint shape
|
|
3047
|
+
// before it reaches either counter (src/responses-usage.js explains why the two
|
|
3048
|
+
// disagree). It rides this function rather than a parser of its own because the
|
|
3049
|
+
// line is ALREADY parsed here: the branch costs a Set lookup on a string, not a
|
|
3050
|
+
// second pass over the stream. Nothing else would be cheap — a Responses stream
|
|
3051
|
+
// is mostly text deltas, and the settled figures arrive on one event near the end
|
|
3052
|
+
// with no header or marker to find it by.
|
|
3053
|
+
//
|
|
3054
|
+
// Reads one `data:` line. Both dialects carry exactly one per event, so a line
|
|
3055
|
+
// is an event for this purpose, and the scanner above never has to hold more.
|
|
3056
|
+
/**
|
|
3057
|
+
* @param {string} line
|
|
3058
|
+
* @param {number} accountIndex
|
|
3059
|
+
* @param {any} accountManager
|
|
3060
|
+
* @param {((inputTokens: number, outputTokens: number) => void)|null} [onUsage]
|
|
3061
|
+
* @param {Record<string, any>|null} [merged]
|
|
3062
|
+
*/
|
|
3046
3063
|
function parseSSEDataLine(line, accountIndex, accountManager, onUsage = null, merged = null) {
|
|
3047
3064
|
if (!line.startsWith('data: ')) return;
|
|
3048
3065
|
|
|
@@ -3056,6 +3073,15 @@ function parseSSEDataLine(line, accountIndex, accountManager, onUsage = null, me
|
|
|
3056
3073
|
accountManager.updateUsage(accountIndex, 0, data.usage.output_tokens);
|
|
3057
3074
|
onUsage?.(0, data.usage.output_tokens || 0);
|
|
3058
3075
|
if (merged) Object.assign(merged, data.usage);
|
|
3076
|
+
} else {
|
|
3077
|
+
// Both sides settle at once here, so unlike the Anthropic branches above
|
|
3078
|
+
// this is a single incremental update rather than one per side.
|
|
3079
|
+
const usage = responsesEventUsage(data);
|
|
3080
|
+
if (usage) {
|
|
3081
|
+
accountManager.updateUsage(accountIndex, usage.input_tokens, usage.output_tokens);
|
|
3082
|
+
onUsage?.(usage.input_tokens, usage.output_tokens);
|
|
3083
|
+
if (merged) Object.assign(merged, usage);
|
|
3084
|
+
}
|
|
3059
3085
|
}
|
|
3060
3086
|
} catch {
|
|
3061
3087
|
// not valid JSON, skip
|
|
@@ -3066,9 +3092,16 @@ function extractUsageFromBody(buffer, accountIndex, accountManager, onUsage = nu
|
|
|
3066
3092
|
try {
|
|
3067
3093
|
const json = JSON.parse(buffer.toString());
|
|
3068
3094
|
if (json.usage) {
|
|
3069
|
-
|
|
3070
|
-
|
|
3071
|
-
|
|
3095
|
+
// A buffered Responses body reports under the same two field NAMES with a
|
|
3096
|
+
// different meaning, so reading it as Anthropic's would book the cached
|
|
3097
|
+
// prefix as fresh input and never book it as a cache read at all. Only a
|
|
3098
|
+
// body that says it is one is rewritten; anything else — including a
|
|
3099
|
+
// Responses body carrying no figures to rewrite — falls through to the
|
|
3100
|
+
// reading this had before, unchanged.
|
|
3101
|
+
const usage = responsesBodyUsage(json) || json.usage;
|
|
3102
|
+
accountManager.updateUsage(accountIndex, usage.input_tokens, usage.output_tokens);
|
|
3103
|
+
onUsage?.(usage.input_tokens || 0, usage.output_tokens || 0);
|
|
3104
|
+
accountManager.recordTokenUsage(accountIndex, sessionId, model, usage);
|
|
3072
3105
|
}
|
|
3073
3106
|
} catch {
|
|
3074
3107
|
// not JSON or no usage
|