llm-switcher 1.1.5 → 1.1.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -2
- package/README.vi.md +17 -2
- package/contract.mjs +25 -7
- package/package.json +1 -1
- package/proxy.mjs +61 -8
- package/state.mjs +8 -0
- package/switch.mjs +8 -1
- package/tests/contract-lab.test.mjs +5 -2
- package/tests/datadir.test.mjs +14 -0
- package/tests/gateway.e2e.test.mjs +42 -2
- package/tests/switch.test.mjs +17 -0
package/README.md
CHANGED
|
@@ -191,7 +191,15 @@ flowchart LR
|
|
|
191
191
|
|
|
192
192
|
---
|
|
193
193
|
|
|
194
|
-
## Changes in 1.1.
|
|
194
|
+
## Changes in 1.1.6
|
|
195
|
+
|
|
196
|
+
- **Codex over WebSocket.** The gateway keeps the turns of each WebSocket session. A turn that sends `previous_response_id` gets the earlier turns back, so Codex no longer loses the task after the first tool call. An unknown id fails the turn with `previous_response_not_found`.
|
|
197
|
+
- **Codex warmup.** A `response.create` frame with `generate: false` gets a local answer. It no longer spends a model call.
|
|
198
|
+
- **Codex model name.** A profile without `publicModels` no longer sends `OpenAI-Model: main` in the handshake, and `switch codex` and `switch doctor` warn about it. Codex read `main` as a reroute and showed a false "high-risk cyber activity" warning. See "Codex-first setup".
|
|
199
|
+
- **Contract lab.** The gateway uploads each sample with the key that opened its trace. intact refused the 1.1.5 uploads with `HTTP 404 trace not found`.
|
|
200
|
+
- **Version stamp.** The stamp uses the last commit only when the checkout has no changes. Otherwise it uses the newest file time.
|
|
201
|
+
|
|
202
|
+
### Changes in 1.1.5
|
|
195
203
|
|
|
196
204
|
- **Contract lab privacy.** Masking now uses an allowlist. Every string value is masked except the enum values that intact reads. 1.1.4 masked a list of content keys and missed 16 fields (citations, document and web titles, web search queries, logprobs tokens, file names and URIs, stop sequences, participant names, error messages, tool descriptions).
|
|
197
205
|
|
|
@@ -342,6 +350,8 @@ The shim does not edit `~/.codex/config.toml`. When the gateway is active, it pa
|
|
|
342
350
|
| `review` | `review_model` | `publicModels[1]` |
|
|
343
351
|
| `subagent` | `agents.default_subagent_model` | `publicModels[2]` |
|
|
344
352
|
|
|
353
|
+
**A profile that serves Codex must have `publicModels`.** Without it, the gateway has no official name for Codex. It then writes no model catalog and sends no `OpenAI-Model` header, and Codex shows two false warnings: "Model metadata for `<model>` not found" and "Your account was flagged for potentially high-risk cyber activity". `switch codex` and `switch doctor` warn when the Codex profile has no `publicModels`.
|
|
354
|
+
|
|
345
355
|
**Codex never receives an internal name.** The slot aliases `main`, `review` and `subagent` stay inside the gateway. The CLI receives the official model names from `publicModels`, and `mapModel` resolves each one back to its slot. Set `codexRoles` in the profile when you want a different pairing than the order of that list.
|
|
346
356
|
|
|
347
357
|
The shim also passes `model_catalog_json`. That file is generated from `publicModels` on every profile change, and the `/model` picker reads it. The picker never calls `/v1/models`. `/v1/models` serves the same entries, and each window follows `model1M` for its slot.
|
|
@@ -544,7 +554,12 @@ re-opened from a shell where the shim is on `PATH`.
|
|
|
544
554
|
"sonnet": true,
|
|
545
555
|
"haiku": false,
|
|
546
556
|
"fable": true
|
|
547
|
-
}
|
|
557
|
+
},
|
|
558
|
+
// Codex keys. A profile that serves Codex MUST have publicModels (see Codex-first setup).
|
|
559
|
+
"publicModels": ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"], // official names for main, review, subagent
|
|
560
|
+
"codexRoles": { "review": "gpt-5.6-sol" }, // optional: pair one role with another public name
|
|
561
|
+
"blindfold": false, // true: Codex keeps its official endpoint (see Codex blindfold mode)
|
|
562
|
+
"blindfoldPort": 3457 // interceptor port when blindfold is true
|
|
548
563
|
}
|
|
549
564
|
},
|
|
550
565
|
"debug": false,
|
package/README.vi.md
CHANGED
|
@@ -191,7 +191,15 @@ flowchart LR
|
|
|
191
191
|
|
|
192
192
|
---
|
|
193
193
|
|
|
194
|
-
## Thay đổi trong bản 1.1.
|
|
194
|
+
## Thay đổi trong bản 1.1.6
|
|
195
|
+
|
|
196
|
+
- **Codex qua WebSocket.** Gateway giữ các lượt của mỗi phiên WebSocket. Lượt nào gửi `previous_response_id` sẽ nhận lại các lượt trước, nên Codex không còn mất nhiệm vụ sau lần gọi tool đầu tiên. Id không tồn tại làm lượt đó lỗi với `previous_response_not_found`.
|
|
197
|
+
- **Warmup của Codex.** Frame `response.create` có `generate: false` được trả lời ngay tại máy. Frame này không còn tốn một lần gọi model.
|
|
198
|
+
- **Tên model cho Codex.** Profile không có `publicModels` không còn gửi `OpenAI-Model: main` trong handshake, và `switch codex` cùng `switch doctor` cảnh báo trường hợp này. Codex đọc `main` là bị chuyển model và hiện cảnh báo sai "high-risk cyber activity". Xem mục "Cấu hình ưu tiên Codex".
|
|
199
|
+
- **Contract lab.** Gateway gửi mỗi mẫu bằng đúng key đã mở trace của mẫu đó. intact từ chối các lần gửi của bản 1.1.5 với `HTTP 404 trace not found`.
|
|
200
|
+
- **Dấu phiên bản.** Dấu phiên bản chỉ dùng commit cuối khi bản checkout không có thay đổi. Nếu có thay đổi, dấu dùng thời gian file mới nhất.
|
|
201
|
+
|
|
202
|
+
### Thay đổi trong bản 1.1.5
|
|
195
203
|
|
|
196
204
|
- **Bảo mật contract lab.** Việc che giờ dùng allowlist. Mọi giá trị string đều bị che, trừ các giá trị enum mà intact đọc. Bản 1.1.4 che theo danh sách key nội dung và bỏ sót 16 field (trích dẫn, tiêu đề tài liệu và trang web, câu truy vấn web search, token logprobs, tên và URI file, stop sequence, tên người tham gia, thông báo lỗi, mô tả tool).
|
|
197
205
|
|
|
@@ -344,6 +352,8 @@ Shim không sửa `~/.codex/config.toml`. Khi gateway hoạt động, shim truy
|
|
|
344
352
|
| `review` | `review_model` | `publicModels[1]` |
|
|
345
353
|
| `subagent` | `agents.default_subagent_model` | `publicModels[2]` |
|
|
346
354
|
|
|
355
|
+
**Profile phục vụ Codex bắt buộc có `publicModels`.** Nếu thiếu khóa này, gateway không có tên chính thức nào để đưa cho Codex. Khi đó gateway không ghi model catalog và không gửi header `OpenAI-Model`, và Codex hiện hai cảnh báo sai: "Model metadata for `<model>` not found" và "Your account was flagged for potentially high-risk cyber activity". `switch codex` và `switch doctor` cảnh báo khi profile Codex không có `publicModels`.
|
|
356
|
+
|
|
347
357
|
**Codex không bao giờ nhận tên nội bộ.** Alias `main`, `review`, `subagent` chỉ tồn tại bên trong gateway. CLI nhận tên model chính thức từ `publicModels`, và `mapModel` phân giải ngược từng tên về đúng slot. Đặt `codexRoles` trong profile nếu muốn ghép khác thứ tự danh sách đó.
|
|
348
358
|
|
|
349
359
|
Shim còn truyền `model_catalog_json`. File này được sinh lại từ `publicModels` mỗi lần đổi profile, và màn `/model` đọc chính nó. Màn `/model` không gọi `/v1/models`. `/v1/models` trả cùng các mục đó, và context window của mỗi mục theo `model1M` của slot.
|
|
@@ -544,7 +554,12 @@ trong `PATH`.
|
|
|
544
554
|
"sonnet": true,
|
|
545
555
|
"haiku": false,
|
|
546
556
|
"fable": true
|
|
547
|
-
}
|
|
557
|
+
},
|
|
558
|
+
// Khóa cho Codex. Profile phục vụ Codex BẮT BUỘC có publicModels (xem Cấu hình ưu tiên Codex).
|
|
559
|
+
"publicModels": ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"], // tên chính thức cho main, review, subagent
|
|
560
|
+
"codexRoles": { "review": "gpt-5.6-sol" }, // không bắt buộc: ghép một vai trò với tên public khác
|
|
561
|
+
"blindfold": false, // true: Codex giữ endpoint chính thức (xem chế độ blindfold)
|
|
562
|
+
"blindfoldPort": 3457 // cổng interceptor khi blindfold là true
|
|
548
563
|
}
|
|
549
564
|
},
|
|
550
565
|
"debug": false,
|
package/contract.mjs
CHANGED
|
@@ -42,6 +42,26 @@ let cachedVersion = '';
|
|
|
42
42
|
|
|
43
43
|
// package version + the commit time, the shape the half route demands. Without git (an unpacked
|
|
44
44
|
// copy) the mtime of package.json stands in, so the value always parses as a past UTC time.
|
|
45
|
+
// The time of the code that runs. The last commit counts only when the tree matches it: files copied
|
|
46
|
+
// over an old checkout (a ZIP) keep the old commit, so a dirty tree uses its newest code file instead.
|
|
47
|
+
export function versionStamp(rootDir) {
|
|
48
|
+
const git = (...a) => execFileSync('git', ['-C', rootDir, ...a], { encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'], timeout: 5000 }).trim();
|
|
49
|
+
try {
|
|
50
|
+
if (git('status', '--porcelain', '--untracked-files=no') === '') {
|
|
51
|
+
const out = git('log', '-1', '--format=%ct');
|
|
52
|
+
if (/^\d{1,12}$/.test(out)) return Number(out) * 1000;
|
|
53
|
+
}
|
|
54
|
+
} catch {}
|
|
55
|
+
let newest = 0;
|
|
56
|
+
try {
|
|
57
|
+
for (const name of fs.readdirSync(rootDir)) {
|
|
58
|
+
if (!/\.(mjs|json|html)$/.test(name)) continue;
|
|
59
|
+
newest = Math.max(newest, fs.statSync(path.join(rootDir, name)).mtimeMs);
|
|
60
|
+
}
|
|
61
|
+
} catch {}
|
|
62
|
+
return newest || Date.now();
|
|
63
|
+
}
|
|
64
|
+
|
|
45
65
|
export function switcherVersion() {
|
|
46
66
|
if (cachedVersion) return cachedVersion;
|
|
47
67
|
let version = '0.0.0';
|
|
@@ -50,17 +70,12 @@ export function switcherVersion() {
|
|
|
50
70
|
const pkgPath = path.join(ROOT_DIR, 'package.json');
|
|
51
71
|
const parsed = /^(\d{1,5}\.\d{1,5}\.\d{1,5})/.exec(JSON.parse(fs.readFileSync(pkgPath, 'utf8')).version || '');
|
|
52
72
|
if (parsed) version = parsed[1];
|
|
53
|
-
stampMs = fs.statSync(pkgPath).mtimeMs;
|
|
54
|
-
} catch {}
|
|
55
|
-
try {
|
|
56
|
-
const out = execFileSync('git', ['-C', ROOT_DIR, 'log', '-1', '--format=%ct'], { encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] }).trim();
|
|
57
|
-
if (/^\d{1,12}$/.test(out)) stampMs = Number(out) * 1000;
|
|
58
73
|
} catch {}
|
|
74
|
+
stampMs = versionStamp(ROOT_DIR);
|
|
59
75
|
cachedVersion = `${version}+${utcStamp(stampMs)}`;
|
|
60
76
|
return cachedVersion;
|
|
61
77
|
}
|
|
62
78
|
|
|
63
|
-
/** Copies what the gateway writes to the client. Past the cap the side is dropped, not cut. */
|
|
64
79
|
// ---------------- masking of client content ----------------
|
|
65
80
|
// A half leaves this machine, so every string value is masked with "x" of the same byte length,
|
|
66
81
|
// except the enum values that intact's reducer reads: the same key list, value shape and exclusions as
|
|
@@ -119,6 +134,7 @@ export function maskHalf(text) {
|
|
|
119
134
|
}).join('\n');
|
|
120
135
|
}
|
|
121
136
|
|
|
137
|
+
/** Copies what the gateway writes to the client. Past the cap the side is dropped, not cut. */
|
|
122
138
|
export function createHalfTap(limit = MAX_HALF_BYTES) {
|
|
123
139
|
const chunks = [];
|
|
124
140
|
let size = 0;
|
|
@@ -273,7 +289,9 @@ export function createContractLab(options = {}) {
|
|
|
273
289
|
try {
|
|
274
290
|
const res = await call(`${s.url}/api/contracts/traces/${encodeURIComponent(traceId)}/half`, {
|
|
275
291
|
method: 'POST',
|
|
276
|
-
|
|
292
|
+
// intact takes a half only from the key that opened the trace, the one this request used;
|
|
293
|
+
// contractLab.apiKey stays for the policy, findings and fixtures.
|
|
294
|
+
headers: { 'content-type': 'application/json', authorization: `Bearer ${half.openerKey || s.apiKey}` },
|
|
277
295
|
body
|
|
278
296
|
});
|
|
279
297
|
// 409 means intact already holds this half: the work is done, a retry would only repeat it.
|
package/package.json
CHANGED
package/proxy.mjs
CHANGED
|
@@ -861,7 +861,8 @@ async function handleConvert(clientFormat, req, res, bodyBuffer, opts = {}) {
|
|
|
861
861
|
toolResponse: halfTap?.text() || '',
|
|
862
862
|
toolVersion: toolVersionFromUA(req.headers['user-agent']),
|
|
863
863
|
inFormat: clientFormat,
|
|
864
|
-
outFormat
|
|
864
|
+
outFormat,
|
|
865
|
+
openerKey: profile.apiKey || ''
|
|
865
866
|
});
|
|
866
867
|
}
|
|
867
868
|
}
|
|
@@ -1461,14 +1462,44 @@ function tapWsEvent(tap, event, data) {
|
|
|
1461
1462
|
// Terminal failure for the WS (responses-ws) transport: Codex ends a turn only on
|
|
1462
1463
|
// response.completed / response.failed, so a bare {type:'error'} frame leaves the turn
|
|
1463
1464
|
// hanging. Emit the full created -> in_progress -> failed sequence instead.
|
|
1464
|
-
function sendWsFailed(socket, model, message, status = 500, tap = null) {
|
|
1465
|
+
function sendWsFailed(socket, model, message, status = 500, tap = null, code = null) {
|
|
1465
1466
|
if (!socket.writable) return;
|
|
1466
1467
|
const renderer = createResponsesStream((e, d) => {
|
|
1467
1468
|
tapWsEvent(tap, e, d);
|
|
1468
1469
|
if (socket.writable) socket.write(encodeWsFrame(JSON.stringify(d)));
|
|
1469
1470
|
}, model || 'main');
|
|
1470
1471
|
renderer.start();
|
|
1471
|
-
renderer.error(message, responsesErrorCode(status));
|
|
1472
|
+
renderer.error(message, code || responsesErrorCode(status));
|
|
1473
|
+
}
|
|
1474
|
+
|
|
1475
|
+
// Codex on WebSocket sends only the new items of a turn and names the turn before in
|
|
1476
|
+
// previous_response_id; the server keeps the rest. The gateway keeps it per socket, bounded.
|
|
1477
|
+
const WS_HISTORY_MAX = 8;
|
|
1478
|
+
// Only these output items are valid input for the next turn; reasoning items are not replayed.
|
|
1479
|
+
const WS_REPLAY_TYPES = new Set(['message', 'function_call', 'custom_tool_call', 'local_shell_call']);
|
|
1480
|
+
|
|
1481
|
+
function wsInputItems(input) {
|
|
1482
|
+
if (typeof input === 'string') return input ? [{ type: 'message', role: 'user', content: [{ type: 'input_text', text: input }] }] : [];
|
|
1483
|
+
return Array.isArray(input) ? input : [];
|
|
1484
|
+
}
|
|
1485
|
+
|
|
1486
|
+
function rememberWsTurn(history, response, input) {
|
|
1487
|
+
if (!history || !response?.id) return;
|
|
1488
|
+
const output = (Array.isArray(response.output) ? response.output : []).filter(i => i && WS_REPLAY_TYPES.has(i.type));
|
|
1489
|
+
history.set(response.id, [...input, ...output]);
|
|
1490
|
+
while (history.size > WS_HISTORY_MAX) history.delete(history.keys().next().value);
|
|
1491
|
+
}
|
|
1492
|
+
|
|
1493
|
+
// A generate:false frame warms the session up (Codex sends its whole tool list in it); it needs no model call.
|
|
1494
|
+
function answerWsWarmup(socket, payload, history, input) {
|
|
1495
|
+
let completed = null;
|
|
1496
|
+
const renderer = createResponsesStream((e, d) => {
|
|
1497
|
+
if (e === 'response.completed') completed = d.response;
|
|
1498
|
+
if (socket.writable) socket.write(encodeWsFrame(JSON.stringify(d)));
|
|
1499
|
+
}, typeof payload?.model === 'string' ? payload.model : 'main');
|
|
1500
|
+
renderer.start();
|
|
1501
|
+
renderer.finish('stop', { completion: 0, prompt: 0, cached: 0, reasoning: 0, hasTools: false });
|
|
1502
|
+
rememberWsTurn(history, completed, input);
|
|
1472
1503
|
}
|
|
1473
1504
|
|
|
1474
1505
|
// 9Router forwards OpenAI-format tools to Gemini/Vertex for ag/* models, which accept only
|
|
@@ -1483,7 +1514,7 @@ function geminiSafeTools(tools) {
|
|
|
1483
1514
|
}
|
|
1484
1515
|
|
|
1485
1516
|
// One turn of the Codex WS transport. The socket loop runs turns one at a time and owns `ac`.
|
|
1486
|
-
async function handleWsResponseCreate(socket, payload, req, ac) {
|
|
1517
|
+
async function handleWsResponseCreate(socket, payload, req, ac, history = null) {
|
|
1487
1518
|
const clientFormat = 'responses';
|
|
1488
1519
|
const { profileKey, profile, error: profileError } = getActiveProfile(clientFormat, req);
|
|
1489
1520
|
if (!loadConfig()) {
|
|
@@ -1495,6 +1526,22 @@ async function handleWsResponseCreate(socket, payload, req, ac) {
|
|
|
1495
1526
|
return;
|
|
1496
1527
|
}
|
|
1497
1528
|
|
|
1529
|
+
let input = wsInputItems(payload.input);
|
|
1530
|
+
if (payload.previous_response_id) {
|
|
1531
|
+
const earlier = history?.get(payload.previous_response_id);
|
|
1532
|
+
if (!earlier) {
|
|
1533
|
+
sendWsFailed(socket, payload?.model || 'main', `Previous response ${String(payload.previous_response_id).slice(0, 80)} is not known on this connection.`, 400, null, 'previous_response_not_found');
|
|
1534
|
+
return;
|
|
1535
|
+
}
|
|
1536
|
+
input = [...earlier, ...input];
|
|
1537
|
+
}
|
|
1538
|
+
if (payload.generate === false) {
|
|
1539
|
+
answerWsWarmup(socket, payload, history, input);
|
|
1540
|
+
return;
|
|
1541
|
+
}
|
|
1542
|
+
const { previous_response_id: _prev, ...rest } = payload;
|
|
1543
|
+
payload = { ...rest, input };
|
|
1544
|
+
|
|
1498
1545
|
let ir;
|
|
1499
1546
|
try {
|
|
1500
1547
|
ir = parseToIR('responses', payload);
|
|
@@ -1546,8 +1593,10 @@ async function handleWsResponseCreate(socket, payload, req, ac) {
|
|
|
1546
1593
|
const normalize = createUpstreamNormalizer(outFormat);
|
|
1547
1594
|
const col = createCollector();
|
|
1548
1595
|
|
|
1596
|
+
let completedResponse = null;
|
|
1549
1597
|
const renderer = createResponsesStream((e, d) => {
|
|
1550
1598
|
tapWsEvent(halfTap, e, d);
|
|
1599
|
+
if (e === 'response.completed') completedResponse = d.response;
|
|
1551
1600
|
if (socket.writable) {
|
|
1552
1601
|
socket.write(encodeWsFrame(JSON.stringify(d)));
|
|
1553
1602
|
}
|
|
@@ -1567,6 +1616,7 @@ async function handleWsResponseCreate(socket, payload, req, ac) {
|
|
|
1567
1616
|
} else {
|
|
1568
1617
|
renderer.finish(col.finish, { completion, prompt: col.prompt, cached: col.cached, reasoning: col.reasoning, hasTools: col.tools.size > 0 });
|
|
1569
1618
|
answered = true;
|
|
1619
|
+
rememberWsTurn(history, completedResponse, input);
|
|
1570
1620
|
}
|
|
1571
1621
|
log({
|
|
1572
1622
|
status: streamError ? 502 : 200, stream: true,
|
|
@@ -1588,7 +1638,8 @@ async function handleWsResponseCreate(socket, payload, req, ac) {
|
|
|
1588
1638
|
toolResponse: halfTap?.text() || '',
|
|
1589
1639
|
toolVersion: toolVersionFromUA(req.headers['user-agent']),
|
|
1590
1640
|
inFormat: clientFormat,
|
|
1591
|
-
outFormat
|
|
1641
|
+
outFormat,
|
|
1642
|
+
openerKey: profile.apiKey || ''
|
|
1592
1643
|
});
|
|
1593
1644
|
}
|
|
1594
1645
|
}
|
|
@@ -1633,14 +1684,15 @@ server.on('upgrade', (req, socket) => {
|
|
|
1633
1684
|
// The official name or the slot, never the upstream ID. The value goes into a raw header line.
|
|
1634
1685
|
const { profile } = getActiveProfile('responses', req);
|
|
1635
1686
|
const publicMain = profile ? codexPublicModel(profile, 'main') : '';
|
|
1636
|
-
|
|
1687
|
+
// With no public name, no header: Codex compares it with the model it asked for, and "main" reads as a reroute.
|
|
1688
|
+
const modelHeader = isSafeModelName(publicMain) ? `OpenAI-Model: ${publicMain}\r\n` : '';
|
|
1637
1689
|
|
|
1638
1690
|
socket.write(
|
|
1639
1691
|
'HTTP/1.1 101 Switching Protocols\r\n' +
|
|
1640
1692
|
'Upgrade: websocket\r\n' +
|
|
1641
1693
|
'Connection: Upgrade\r\n' +
|
|
1642
1694
|
`Sec-WebSocket-Accept: ${accept}\r\n` +
|
|
1643
|
-
|
|
1695
|
+
modelHeader +
|
|
1644
1696
|
'x-reasoning-included: true\r\n' +
|
|
1645
1697
|
'x-codex-turn-state: ready\r\n\r\n'
|
|
1646
1698
|
);
|
|
@@ -1649,6 +1701,7 @@ server.on('upgrade', (req, socket) => {
|
|
|
1649
1701
|
// Turns run one at a time: two at once interleave their events on one socket. Every turn,
|
|
1650
1702
|
// queued or running, holds a controller here, so a cancel or a close stops all of them.
|
|
1651
1703
|
const turns = new Set();
|
|
1704
|
+
const history = new Map(); // response id -> the input items of that turn plus its output
|
|
1652
1705
|
let turnChain = Promise.resolve();
|
|
1653
1706
|
const abortTurns = () => { for (const ac of turns) ac.abort(); };
|
|
1654
1707
|
|
|
@@ -1657,7 +1710,7 @@ server.on('upgrade', (req, socket) => {
|
|
|
1657
1710
|
const ac = new AbortController();
|
|
1658
1711
|
turns.add(ac);
|
|
1659
1712
|
turnChain = turnChain
|
|
1660
|
-
.then(() => (ac.signal.aborted ? null : handleWsResponseCreate(socket, msg, req, ac)))
|
|
1713
|
+
.then(() => (ac.signal.aborted ? null : handleWsResponseCreate(socket, msg, req, ac, history)))
|
|
1661
1714
|
.catch(err => {
|
|
1662
1715
|
// A throw before the handler's own try: Codex ends a turn only on response.failed.
|
|
1663
1716
|
console.error('[llm-switcher:ws] Unhandled turn error:', err);
|
package/state.mjs
CHANGED
|
@@ -285,6 +285,14 @@ export function isSafeModelName(name) {
|
|
|
285
285
|
return typeof name === 'string' && SAFE_MODEL_NAME.test(name) && !TOML_NON_STRING.test(name);
|
|
286
286
|
}
|
|
287
287
|
|
|
288
|
+
// Codex compares the handshake model with the one it asked for, and reads its /model catalog from publicModels.
|
|
289
|
+
export function codexPublicModelsWarning(key, profile) {
|
|
290
|
+
if (!profile || isSafeModelName(codexPublicModel(profile, 'main'))) return '';
|
|
291
|
+
return `Profile "${key}" serves Codex but has no publicModels. Codex gets no model catalog and shows false `
|
|
292
|
+
+ `"model metadata not found" and "high-risk cyber activity" warnings. Add for example `
|
|
293
|
+
+ `"publicModels": ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"], then run \`switch codex ${key}\` again.`;
|
|
294
|
+
}
|
|
295
|
+
|
|
288
296
|
/**
|
|
289
297
|
* Does this leaf certificate cover the host the interceptor will present it for?
|
|
290
298
|
* Changing blindfoldHost without rebuilding the leaf produces a TLS error that reads
|
package/switch.mjs
CHANGED
|
@@ -9,7 +9,7 @@ import {
|
|
|
9
9
|
applyLaunchState, clearLaunchState, computeLaunchState,
|
|
10
10
|
modelSlotsForProfile, modelForSlot, model1MForSlot, readAdminToken, adminTokenPath, openLog,
|
|
11
11
|
probeGateway, probeBlindfold, blindfoldPreflight, stopRecordedBlindfold, writeDashboardLauncher,
|
|
12
|
-
contractLabSettings
|
|
12
|
+
contractLabSettings, codexPublicModelsWarning
|
|
13
13
|
} from './state.mjs';
|
|
14
14
|
import { runProbe, runCheck } from './contract.mjs';
|
|
15
15
|
import {
|
|
@@ -436,6 +436,9 @@ async function turnOn(profileName, cliTarget) {
|
|
|
436
436
|
printTargets(getActiveMap(planned));
|
|
437
437
|
console.log(`\nClaude 1M: ${describeClaude1M(st)}`);
|
|
438
438
|
console.log(`Codex 1M: ${st.codex1M ? 'ACTIVE (1,000,000 tokens)' : 'OFF'}`);
|
|
439
|
+
const codexKey = getActiveMap(planned).responses;
|
|
440
|
+
const codexWarning = codexPublicModelsWarning(show(codexKey), planned.profiles[codexKey]);
|
|
441
|
+
if (codexWarning) console.warn(`\n[WARN] ${codexWarning}`);
|
|
439
442
|
}
|
|
440
443
|
|
|
441
444
|
async function turnOff(targetArg) {
|
|
@@ -770,6 +773,10 @@ async function runDoctor() {
|
|
|
770
773
|
const p = config.profiles[key];
|
|
771
774
|
if (!p) warn(`[WARN] Target ${t} points to missing profile "${show(key)}".`);
|
|
772
775
|
else if (!p.baseURL || /YOUR-|REPLACE-ME/i.test(`${p.baseURL} ${p.apiKey}`)) warn(`[WARN] Profile "${show(key)}" (${t}) still has placeholder baseURL/apiKey.`);
|
|
776
|
+
if (p && t === 'responses') {
|
|
777
|
+
const codexWarning = codexPublicModelsWarning(show(key), p);
|
|
778
|
+
if (codexWarning) warn(`[WARN] ${codexWarning}`);
|
|
779
|
+
}
|
|
773
780
|
}
|
|
774
781
|
|
|
775
782
|
// 3. Launcher flags match proxy state
|
|
@@ -394,7 +394,8 @@ test('the half of a sampled request reaches intact with the exchange and the ver
|
|
|
394
394
|
}, 'a sampled request to reach intact');
|
|
395
395
|
const last = halves.at(-1);
|
|
396
396
|
assert.match(last.traceId, TRACE_ID_RE);
|
|
397
|
-
|
|
397
|
+
// intact accepts a half only from the key that opened the trace: the key of the profile that made the call.
|
|
398
|
+
assert.equal(last.headers.authorization, 'Bearer sk-secret-ant');
|
|
398
399
|
assert.deepEqual(last.body.converter, { inFormat: 'anthropic', outFormat: 'anthropic' });
|
|
399
400
|
assert.equal(last.body.toolVersion, '1.2.3');
|
|
400
401
|
assert.match(last.body.switcherVersion, SWITCHER_VERSION_RE);
|
|
@@ -557,7 +558,9 @@ test('a sampled Codex WS turn is tagged and its half reaches intact', async () =
|
|
|
557
558
|
assert.match(body.switcherVersion, SWITCHER_VERSION_RE);
|
|
558
559
|
const asked = JSON.parse(body.toolRequest);
|
|
559
560
|
assert.equal(asked.type, 'response.create');
|
|
560
|
-
|
|
561
|
+
// The half holds the request as the upstream saw it: the turn's input as items (with any earlier
|
|
562
|
+
// turns restored from previous_response_id), its text masked.
|
|
563
|
+
assert.equal(asked.input[0].content[0].text, 'x'.repeat('hello over ws'.length), 'the WS request is masked');
|
|
561
564
|
assert.ok(body.toolResponse.includes('event: response.completed'), 'the half holds the events the turn wrote');
|
|
562
565
|
assert.ok(!body.toolResponse.includes('ws ok'), 'the answer text is masked in the half');
|
|
563
566
|
});
|
package/tests/datadir.test.mjs
CHANGED
|
@@ -49,3 +49,17 @@ test('make-certs.sh runs with an openssl that lacks LibreSSL-missing options', {
|
|
|
49
49
|
execFileSync('bash', [script, 'chatgpt.com', out], { env: { ...process.env, PATH: `${bin}:${process.env.PATH}` }, stdio: 'pipe' });
|
|
50
50
|
for (const f of ['ca.pem', 'leaf.pem', 'leaf.key']) assert.ok(fs.existsSync(path.join(out, f)), f);
|
|
51
51
|
});
|
|
52
|
+
|
|
53
|
+
// LS-4: files copied over an old checkout (a ZIP) must not borrow the time of the old commit.
|
|
54
|
+
test('the version stamp uses the commit time only when the working tree is clean', async () => {
|
|
55
|
+
const { versionStamp } = await import('../contract.mjs');
|
|
56
|
+
const dir = tmp();
|
|
57
|
+
const git = (...a) => execFileSync('git', ['-C', dir, ...a], { stdio: 'pipe', env: { ...process.env, GIT_AUTHOR_DATE: '2026-01-01T00:00:00Z', GIT_COMMITTER_DATE: '2026-01-01T00:00:00Z' } });
|
|
58
|
+
fs.writeFileSync(path.join(dir, 'package.json'), '{"version":"1.2.3"}');
|
|
59
|
+
fs.writeFileSync(path.join(dir, 'a.mjs'), 'export {};\n');
|
|
60
|
+
git('init', '-q'); git('-c', 'user.email=t@t', '-c', 'user.name=t', 'add', '.'); git('-c', 'user.email=t@t', '-c', 'user.name=t', 'commit', '-qm', 'old');
|
|
61
|
+
const commitMs = Date.parse('2026-01-01T00:00:00Z');
|
|
62
|
+
assert.equal(versionStamp(dir), commitMs, 'clean tree: commit time');
|
|
63
|
+
fs.writeFileSync(path.join(dir, 'a.mjs'), 'export const changed = 1;\n');
|
|
64
|
+
assert.ok(versionStamp(dir) > commitMs + 60_000, 'dirty tree: the time of the newest file, not of the old commit');
|
|
65
|
+
});
|
|
@@ -784,9 +784,10 @@ test('Codex WS: a frame larger than the body cap closes the socket before it is
|
|
|
784
784
|
assert.equal(close.payload.readUInt16BE(0), 1009);
|
|
785
785
|
});
|
|
786
786
|
|
|
787
|
-
test('Codex WS: 101 reply names the public main model or
|
|
787
|
+
test('Codex WS: 101 reply names the public main model, or no model, never the upstream id', async () => {
|
|
788
788
|
const hidden = await rawWs({ 'x-llm-profile': 'agmock' });
|
|
789
|
-
|
|
789
|
+
// With no public name the header is left out: Codex reads a different name as a reroute (LS-2).
|
|
790
|
+
assert.ok(!/\r\nOpenAI-Model:/i.test(hidden.reply), hidden.reply);
|
|
790
791
|
assert.ok(!hidden.reply.includes('ag/mock-flash'));
|
|
791
792
|
hidden.socket.destroy();
|
|
792
793
|
const published = await rawWs({ 'x-llm-profile': 'pub' });
|
|
@@ -805,6 +806,45 @@ test('Codex WS: overlapping response.create turns run one after the other', asyn
|
|
|
805
806
|
ws.socket.destroy();
|
|
806
807
|
});
|
|
807
808
|
|
|
809
|
+
// LS-1: from turn 2 Codex sends previous_response_id and only the new items; the gateway must put the
|
|
810
|
+
// earlier turns back, or the model loses the task and loops.
|
|
811
|
+
test('Codex WS: previous_response_id brings back the earlier turns', async () => {
|
|
812
|
+
const ws = await rawWs();
|
|
813
|
+
ws.socket.write(clientFrame(1, JSON.stringify({ type: 'response.create', model: 'main', input: 'remember the code word ZEBRA-42' })));
|
|
814
|
+
assert.ok(await until(() => ws.messages.some(m => m.type === 'response.completed')), 'turn 1 ends');
|
|
815
|
+
const firstId = ws.messages.find(m => m.type === 'response.completed').response.id;
|
|
816
|
+
const before = received.length;
|
|
817
|
+
ws.socket.write(clientFrame(1, JSON.stringify({ type: 'response.create', model: 'main', previous_response_id: firstId,
|
|
818
|
+
input: [{ type: 'function_call_output', call_id: 'call_1', output: 'tool says ok' }] })));
|
|
819
|
+
assert.ok(await until(() => ws.messages.filter(m => m.type === 'response.completed' || m.type === 'response.failed').length === 2), 'turn 2 ends');
|
|
820
|
+
const sent = JSON.stringify(received.slice(before).at(-1).body);
|
|
821
|
+
assert.ok(sent.includes('ZEBRA-42'), `turn 2 upstream lost turn 1: ${sent.slice(0, 400)}`);
|
|
822
|
+
assert.ok(sent.includes('tool says ok'), 'turn 2 keeps its own new item');
|
|
823
|
+
ws.socket.destroy();
|
|
824
|
+
});
|
|
825
|
+
|
|
826
|
+
test('Codex WS: an unknown previous_response_id fails the turn without an upstream call', async () => {
|
|
827
|
+
const ws = await rawWs();
|
|
828
|
+
const before = received.length;
|
|
829
|
+
ws.socket.write(clientFrame(1, JSON.stringify({ type: 'response.create', model: 'main', previous_response_id: 'resp_unknown', input: [] })));
|
|
830
|
+
assert.ok(await until(() => ws.messages.some(m => m.type === 'response.failed')), 'the turn fails');
|
|
831
|
+
assert.equal(ws.messages.find(m => m.type === 'response.failed').response.error.code, 'previous_response_not_found');
|
|
832
|
+
assert.equal(received.length, before, 'no upstream call');
|
|
833
|
+
ws.socket.destroy();
|
|
834
|
+
});
|
|
835
|
+
|
|
836
|
+
// LS-3: Codex opens a session with generate:false. It carries the whole tool list and needs no answer.
|
|
837
|
+
test('Codex WS: a generate:false warmup is answered locally, never sent upstream', async () => {
|
|
838
|
+
const ws = await rawWs();
|
|
839
|
+
const before = received.length;
|
|
840
|
+
ws.socket.write(clientFrame(1, JSON.stringify({ type: 'response.create', model: 'main', generate: false, input: [], tools: [] })));
|
|
841
|
+
assert.ok(await until(() => ws.messages.some(m => m.type === 'response.completed')), 'the warmup completes');
|
|
842
|
+
const done = ws.messages.find(m => m.type === 'response.completed');
|
|
843
|
+
assert.deepEqual(done.response.output, []);
|
|
844
|
+
assert.equal(received.length, before, 'no upstream call for a warmup');
|
|
845
|
+
ws.socket.destroy();
|
|
846
|
+
});
|
|
847
|
+
|
|
808
848
|
test('Codex WS: a mid-stream error is logged with its text', async () => {
|
|
809
849
|
const ws = await rawWs();
|
|
810
850
|
ws.socket.write(clientFrame(1, JSON.stringify({ type: 'response.create', model: 'main', input: 'MID_STREAM_ERROR over ws' })));
|
package/tests/switch.test.mjs
CHANGED
|
@@ -154,3 +154,20 @@ test('switch port stops the new gateway it started when it rolls back', { skip:
|
|
|
154
154
|
});
|
|
155
155
|
assert.equal(probe, 'free', 'nothing listens on the abandoned port');
|
|
156
156
|
});
|
|
157
|
+
|
|
158
|
+
// Without publicModels the gateway has no official name to give Codex: no model catalog, no OpenAI-Model
|
|
159
|
+
// header, and Codex prints false "metadata not found" and "high-risk cyber activity" warnings.
|
|
160
|
+
test('switch codex and switch doctor warn when the Codex profile has no publicModels', { skip: !POSIX && 'posix' }, async (t) => {
|
|
161
|
+
const base = { mode: 'convert', inFormat: 'auto', baseURL: 'http://127.0.0.1:9/v1', apiKey: 'k', defaultModels: { opus: 'o' } };
|
|
162
|
+
const ws = workspace(t, await freePort(), { bare: { name: 'Bare', ...base }, named: { name: 'Named', ...base, publicModels: ['gpt-5.6-sol'] } });
|
|
163
|
+
const on = await run(ws, ['codex', 'bare']);
|
|
164
|
+
assert.equal(on.status, 0, on.stderr);
|
|
165
|
+
assert.match(on.stdout + on.stderr, /\[WARN\].*"bare".*publicModels/);
|
|
166
|
+
const doctor = await run(ws, ['doctor']);
|
|
167
|
+
assert.match(doctor.stdout, /\[WARN\].*"bare".*publicModels/);
|
|
168
|
+
|
|
169
|
+
const named = await run(ws, ['codex', 'named']);
|
|
170
|
+
assert.equal(named.status, 0, named.stderr);
|
|
171
|
+
assert.doesNotMatch(named.stdout + named.stderr, /publicModels/);
|
|
172
|
+
assert.doesNotMatch((await run(ws, ['doctor'])).stdout, /publicModels/);
|
|
173
|
+
});
|