llm-switcher 1.1.11 → 1.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +63 -0
- package/README.md +202 -257
- package/README.vi.md +200 -256
- package/blindfold/blindfold.mjs +200 -53
- package/blindfold/make-certs.sh +26 -7
- package/catalog.mjs +246 -0
- package/classifier.mjs +238 -0
- package/config.example.json +12 -34
- package/docs/codex-blindfold.md +28 -17
- package/docs/cross-platform.md +16 -7
- package/docs/diagrams/ir-healer-pipeline.mmd +16 -0
- package/docs/diagrams/ir-healer-pipeline.png +0 -0
- package/docs/diagrams/ir-healer-pipeline.svg +90 -0
- package/docs/diagrams/ir-translation-pipeline.html +14925 -0
- package/docs/diagrams/ir-translation-pipeline.sequence.json +31 -0
- package/docs/diagrams/ir-translation-pipeline.svg +5128 -0
- package/docs/diagrams/system-architecture.architecture.json +76 -0
- package/docs/diagrams/system-architecture.html +14978 -0
- package/docs/diagrams/system-architecture.svg +5147 -0
- package/docs/diagrams/system-topology.mmd +30 -0
- package/docs/diagrams/system-topology.png +0 -0
- package/docs/diagrams/system-topology.svg +125 -0
- package/ensure-ca-bundle.mjs +28 -0
- package/formats.mjs +43 -156
- package/icons/antigravity.png +0 -0
- package/icons/claude.png +0 -0
- package/icons/codex.png +0 -0
- package/icons/deepseek.png +0 -0
- package/icons/gemini.png +0 -0
- package/icons/github.png +0 -0
- package/icons/groq.png +0 -0
- package/icons/intact.svg +1 -0
- package/icons/ollama.png +0 -0
- package/icons/openai.png +0 -0
- package/icons/openrouter.png +0 -0
- package/icons/qwen.png +0 -0
- package/icons/vertex.png +0 -0
- package/mcp.mjs +39 -11
- package/package.json +1 -1
- package/proxy.mjs +114 -37
- package/shim.mjs +200 -57
- package/skills/llm-switcher/SKILL.md +15 -10
- package/state.mjs +1100 -191
- package/switch.cmd +2 -2
- package/switch.mjs +228 -53
- package/tests/blindfold-e2e.test.mjs +380 -0
- package/tests/blindfold-task5.test.mjs +429 -0
- package/tests/blindfold.task3.test.mjs +700 -0
- package/tests/blindfold.test.mjs +10 -5
- package/tests/catalog.test.mjs +147 -0
- package/tests/classifier.test.mjs +210 -0
- package/tests/contract-lab.test.mjs +22 -7
- package/tests/formats.test.mjs +63 -46
- package/tests/gateway.e2e.test.mjs +136 -36
- package/tests/lifecycle.test.mjs +16 -10
- package/tests/mcp.test.mjs +78 -2
- package/tests/real-user-sim.test.mjs +464 -0
- package/tests/shim.test.mjs +159 -66
- package/tests/state.test.mjs +975 -193
- package/tests/switch.test.mjs +446 -2
- package/ui.html +1710 -1726
package/README.md
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
<p align="center">
|
|
4
4
|
<b>Zero-dependency, multi-protocol edge gateway & provider switcher</b><br>
|
|
5
5
|
Seamlessly bridge <b>Claude Code</b>, <b>Codex</b>, OpenAI, and Gemini SDKs to any upstream LLM API.<br>
|
|
6
|
-
Full bi-directional protocol conversion,
|
|
6
|
+
Full bi-directional protocol conversion, official-model context windows, thinking protocol extraction, and edge message healing.
|
|
7
7
|
</p>
|
|
8
8
|
|
|
9
9
|
<p align="center">
|
|
@@ -13,162 +13,118 @@
|
|
|
13
13
|
<p align="center">
|
|
14
14
|
<img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?logo=node.js&logoColor=white" alt="Node.js 18+">
|
|
15
15
|
<img src="https://img.shields.io/badge/Dependencies-Zero-38bdf8" alt="Zero Dependencies">
|
|
16
|
-
<img src="https://img.shields.io/badge/Context-
|
|
16
|
+
<img src="https://img.shields.io/badge/Context-window_follows_the_model-6366f1" alt="Context window follows the model">
|
|
17
17
|
<img src="https://img.shields.io/badge/Multi--Active-Concurrent_CLIs-f59e0b" alt="Multi-Active">
|
|
18
18
|
<img src="https://img.shields.io/badge/License-MIT-gray" alt="License MIT">
|
|
19
19
|
</p>
|
|
20
20
|
|
|
21
21
|
---
|
|
22
22
|
|
|
23
|
-
> ###
|
|
23
|
+
> ### 🛡️ Zero-Loss Native Emulation for Coding Agents
|
|
24
24
|
>
|
|
25
|
-
>
|
|
26
|
-
> - **Anthropic** requires dedicated `thinking` blocks (`thinking_delta` + `signature_delta`), strict alternating turn rules, and typed `tool_use` input schemas.
|
|
27
|
-
> - **OpenAI** streams reasoning as `reasoning_content` delta chunks or `reasoning_details[]`, and formats tools as `tool_calls` with JSON string arguments.
|
|
28
|
-
> - **Google Vertex AI** places reasoning in `candidates[0].content.parts[{thought: true, text, thoughtSignature}]` and tool arguments as raw objects.
|
|
29
|
-
> - **Open-source models (DeepSeek, Qwen, GLM)** often dump chain-of-thought directly into `content` or duplicate fields under conflicting keys.
|
|
25
|
+
> Generic proxies mangle response protocols: Anthropic loses `thinking_delta` reasoning blocks, tool arguments split, and prompt caches desynchronize.
|
|
30
26
|
>
|
|
31
|
-
> **
|
|
32
|
-
>
|
|
33
|
-
>
|
|
34
|
-
>
|
|
35
|
-
>
|
|
36
|
-
> Developers often blame the model for "getting dumber" when in reality **their proxy mangled the response protocol.**
|
|
37
|
-
>
|
|
38
|
-
> ### 🛡️ The Solution: Zero-Loss Native Emulation (Subscription-Grade Quality)
|
|
39
|
-
>
|
|
40
|
-
> **LLM Switcher solves this by acting as a high-precision, zero-loss protocol emulator.**
|
|
41
|
-
>
|
|
42
|
-
> It normalizes whatever your upstream provider emits (9Router, OpenRouter, Vertex, DeepSeek) and re-synthesizes it into the **exact native event stream the client agent was built to consume**:
|
|
43
|
-
> - **Claude Code** receives 100% genuine Anthropic SSE events (`message_start` ➔ `thinking_delta` ➔ `signature_delta` ➔ `content_block_start: tool_use` ➔ `message_delta`), performing **identically to an official Anthropic subscription**.
|
|
44
|
-
> - **Codex** receives 100% genuine Responses API events (`response.created` ➔ `output_text.delta` ➔ `function_call` ➔ `response.completed`).
|
|
45
|
-
>
|
|
46
|
-
> **You get the freedom and cost savings of 3rd-party APIs while maintaining 100% official subscription-grade agent intelligence.**
|
|
27
|
+
> **LLM Switcher solves this at the local network edge:**
|
|
28
|
+
> - **100% Native Emulation:** Normalizes upstream APIs (intact, 9Router, Vertex, DeepSeek) into genuine Anthropic SSE (`thinking_delta` + `tool_use`) for Claude Code, and genuine Responses API events for Codex.
|
|
29
|
+
> - **Client-Side Edge Companion:** Intentionally offloads heavy account pooling and key rotation to **[intact](https://github.com/louisphamdev/intact)** (recommended) or 9Router (basic alternative), keeping LLM Switcher zero-dependency and bloat-free.
|
|
30
|
+
> - **Targeted Tool Scope:** Built specifically for **Claude Code** and **OpenAI Codex** (OpenCode natively supports custom models without shims; refer to intact for account pooling; and Antigravity isn't worth building for 😏).
|
|
47
31
|
|
|
48
32
|
---
|
|
49
33
|
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
>
|
|
55
|
-
>
|
|
56
|
-
|
|
57
|
-
>
|
|
58
|
-
>
|
|
59
|
-
>
|
|
60
|
-
>
|
|
34
|
+
## Architecture & Interactive Diagrams
|
|
35
|
+
|
|
36
|
+
LLM Switcher runs locally on your workstation (`127.0.0.1:3456`) as a transparent edge interceptor and protocol bridge.
|
|
37
|
+
|
|
38
|
+
<p align="center">
|
|
39
|
+
<a href="docs/diagrams/system-architecture.html">
|
|
40
|
+
<img src="docs/diagrams/system-topology.svg" alt="LLM Switcher System Topology & Architecture" width="100%">
|
|
41
|
+
</a>
|
|
42
|
+
<br>
|
|
43
|
+
<sub><i>🎨 Themed with Pretty-Mermaid (Tokyo Night). Click diagram to open interactive Archify viewer (zoom, pan, tracing).</i></sub>
|
|
44
|
+
</p>
|
|
45
|
+
|
|
46
|
+
### 1. Interactive Archify Visual Library
|
|
47
|
+
|
|
48
|
+
All architecture maps and execution sequences are authored with **[Archify](https://github.com/tt-a1i/archify)** and rendered with **[Pretty-Mermaid](https://github.com/imxv/Pretty-mermaid-skills)**:
|
|
49
|
+
|
|
50
|
+
| Diagram | Description | Interactive Visual | Scalable Vector |
|
|
51
|
+
|---|---|---|---|
|
|
52
|
+
| **System Topology** | Complete edge architecture: Clients ➔ Optimizers ➔ Gateway & Healer Core ➔ Upstream Providers | [📊 Open Interactive View](docs/diagrams/system-architecture.html) | [SVG](docs/diagrams/system-topology.svg) • [PNG](docs/diagrams/system-topology.png) |
|
|
53
|
+
| **IR Healer Pipeline** | Inbound request normalization, schema healing, streaming synthesis, and abort propagation | [🔄 Open Interactive View](docs/diagrams/ir-translation-pipeline.html) | [SVG](docs/diagrams/ir-healer-pipeline.svg) • [PNG](docs/diagrams/ir-healer-pipeline.png) |
|
|
54
|
+
| **Codex Blindfold Routing** | TLS CONNECT proxy sequence, credential scrubbing, and upstream routing | [🛡️ Open Interactive View](docs/diagrams/blindfold-request-routing.html) | [HTML](docs/diagrams/blindfold-request-routing.html) |
|
|
55
|
+
| **Switch Lifecycle** | Zero-downtime tool toggle, CAS configuration writes, and interceptor sync | [⚡ Open Interactive View](docs/diagrams/blindfold-switch-lifecycle.html) | [HTML](docs/diagrams/blindfold-switch-lifecycle.html) |
|
|
61
56
|
|
|
62
57
|
---
|
|
63
58
|
|
|
64
|
-
|
|
59
|
+
### 2. Request Lifecycle & Healer Pipeline
|
|
65
60
|
|
|
66
|
-
|
|
61
|
+
<p align="center">
|
|
62
|
+
<a href="docs/diagrams/ir-translation-pipeline.html">
|
|
63
|
+
<img src="docs/diagrams/ir-healer-pipeline.svg" alt="Bi-Directional IR Healer Pipeline" width="100%">
|
|
64
|
+
</a>
|
|
65
|
+
<br>
|
|
66
|
+
<sub><i>💡 Click above to inspect the interactive IR Healer lifecycle sequence.</i></sub>
|
|
67
|
+
</p>
|
|
67
68
|
|
|
68
|
-
###
|
|
69
|
+
### 3. High-Level Flow
|
|
69
70
|
|
|
70
71
|
```mermaid
|
|
71
|
-
flowchart
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
72
|
+
flowchart LR
|
|
73
|
+
classDef client fill:#1e293b,stroke:#38bdf8,stroke-width:2px,color:#f8fafc;
|
|
74
|
+
classDef edge fill:#0f172a,stroke:#6366f1,stroke-width:2px,color:#f8fafc;
|
|
75
|
+
classDef healer fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#f8fafc;
|
|
76
|
+
classDef upstream fill:#2e1065,stroke:#a855f7,stroke-width:2px,color:#f8fafc;
|
|
77
|
+
classDef opt fill:#1e1b4b,stroke:#818cf8,stroke-dasharray: 4 4,color:#e0e7ff;
|
|
78
|
+
|
|
79
|
+
subgraph Clients[" 💻 Dev Clients & Coding CLIs "]
|
|
80
|
+
CC["Claude Code CLI\n(/v1/messages)"]:::client
|
|
81
|
+
CDX["OpenAI Codex CLI\n(/v1/responses)"]:::client
|
|
77
82
|
end
|
|
78
83
|
|
|
79
|
-
subgraph
|
|
80
|
-
OPT["
|
|
84
|
+
subgraph Middle[" ⚡ Optional Middle-Layer "]
|
|
85
|
+
OPT["Token Optimizers\n(Headroom / RTK)"]:::opt
|
|
81
86
|
end
|
|
82
87
|
|
|
83
|
-
subgraph
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
M1M["1M Context Unlocker\n& Auto-Compact Thresholds"]
|
|
89
|
-
LOGS["Live Inspector\n(In-Memory Ring Buffer)"]
|
|
90
|
-
ROUTER --> HEALER --> IR --> M1M --> LOGS
|
|
88
|
+
subgraph Gateway[" 🛡️ LLM Switcher Edge Gateway (:3456) "]
|
|
89
|
+
ROUTER["Edge Router\n(Zero-Mutation)"]:::edge
|
|
90
|
+
HEALER["Healer Engine\n(Auto-Fix Schemas)"]:::healer
|
|
91
|
+
IR["Bi-Directional IR\n(Event Synth)"]:::healer
|
|
92
|
+
ROUTER --> HEALER --> IR
|
|
91
93
|
end
|
|
92
94
|
|
|
93
|
-
subgraph
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
ANT["Anthropic Native API"]
|
|
97
|
-
GCP["Google Vertex AI / Gemini"]
|
|
95
|
+
subgraph Upstreams[" ☁️ Upstream Providers "]
|
|
96
|
+
INTACT["intact Gateway\n(Recommended Pooler)"]:::upstream
|
|
97
|
+
OTHER["9Router / Vertex / Other"]:::upstream
|
|
98
98
|
end
|
|
99
99
|
|
|
100
|
-
CC -->|
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
CDX -.->|
|
|
104
|
-
|
|
105
|
-
VTX --> ROUTER
|
|
106
|
-
OPT -->|Forward to Switcher| ROUTER
|
|
107
|
-
|
|
108
|
-
LOGS -->|Clean Outbound| R9
|
|
109
|
-
LOGS -->|Clean Outbound| OR
|
|
110
|
-
LOGS -->|Clean Outbound| ANT
|
|
111
|
-
LOGS -->|Clean Outbound| GCP
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
---
|
|
115
|
-
|
|
116
|
-
### 2. Bi-Directional IR (Intermediate Representation) Pipeline
|
|
100
|
+
CC -->|direct| ROUTER
|
|
101
|
+
CDX -->|direct| ROUTER
|
|
102
|
+
CC -.->|prune| OPT
|
|
103
|
+
CDX -.->|prune| OPT
|
|
104
|
+
OPT -->|forward| ROUTER
|
|
117
105
|
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
autonumber
|
|
121
|
-
actor CLI as Client (Claude Code / Codex / SDK)
|
|
122
|
-
participant GW as LLM Switcher (:3456)
|
|
123
|
-
participant IR as IR & Healer Engine
|
|
124
|
-
participant UP as Upstream (9Router / Anthropic / Vertex)
|
|
125
|
-
|
|
126
|
-
CLI->>GW: Inbound Request (Anthropic, Responses, Chat, or Vertex)
|
|
127
|
-
Note over GW,IR: Normalize to Canonical IR
|
|
128
|
-
GW->>IR: parseToIR(clientFormat, payload)
|
|
129
|
-
Note over IR: Healer checks:<br/>1. Repair orphaned tool_results<br/>2. Restore stripped thinking params<br/>3. Reconcile role alternations<br/>4. Apply 1M context limits
|
|
130
|
-
IR->>GW: emitUpstreamBody(outFormat, healedIR)
|
|
131
|
-
GW->>UP: Outbound API Call (fetch with AbortSignal)
|
|
132
|
-
UP-->>GW: Upstream Streaming SSE / JSON Chunks
|
|
133
|
-
Note over GW: normalizeUpstream(chunk)<br/>Extract reasoning_content, <think> tags, usage
|
|
134
|
-
GW->>CLI: Render client-native SSE (e.g. Anthropic thinking_delta + text_delta)
|
|
135
|
-
Note over CLI,GW: Client connection closes (Ctrl+C) -> GW aborts UP instantly!
|
|
106
|
+
IR -->|contract & pool| INTACT
|
|
107
|
+
IR -->|standard call| OTHER
|
|
136
108
|
```
|
|
137
109
|
|
|
138
110
|
---
|
|
139
111
|
|
|
140
|
-
###
|
|
112
|
+
### 4. How It Actually Works: Transparent Request Interception
|
|
141
113
|
|
|
142
|
-
|
|
114
|
+
LLM Switcher acts as a transparent man-in-the-middle without ever touching client configuration files:
|
|
143
115
|
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
C2["Codex CLI\n(/v1/responses)"]
|
|
149
|
-
C3["OpenAI SDK\n(/v1/chat/completions)"]
|
|
150
|
-
C4["Vertex SDK\n(/v1beta/models/*)"]
|
|
151
|
-
end
|
|
116
|
+
1. **Ephemeral Shim Activation:** When you invoke `claude` or `codex`, a lightweight shim at the front of your `PATH` executes first. It injects `HTTPS_PROXY=http://127.0.0.1:3457` and custom CA certs *only into that process's in-memory environment*, leaving `~/.claude/settings.json` and `~/.codex/config.toml` completely untouched.
|
|
117
|
+
2. **Network Interception (`:3457`):** The tool sends standard TLS requests to official hosts (`api.anthropic.com` or `api.openai.com`). The local Blindfold interceptor terminates TLS, strips client credentials, and transparently re-routes API calls (`/v1/messages`, `/v1/responses`, `/v1/models`) locally to the Gateway (`:3456`). All other traffic (OAuth logins, GitHub, web searches) tunnels untouched to the real internet.
|
|
118
|
+
3. **Protocol Conversion & Upstream Call (`:3456`):** The gateway loads your active profile from `config.json`, runs the Healer Engine (repairing empty `{}` schemas and restoring reasoning budgets), and converts the request into the upstream provider's native dialect (e.g. Gemini, OpenAI Chat, or Anthropic) with your configured credentials and base URL.
|
|
119
|
+
4. **Native Event Stream Synthesis:** When the upstream provider streams its response, the gateway captures reasoning tokens and tool calls, re-synthesizing them into genuine Anthropic SSE (`thinking_delta` + `tool_use`) or Codex Responses events. The tool receives genuine native events and believes it is communicating directly with the official provider!
|
|
152
120
|
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
subgraph Egress["Upstream Targets"]
|
|
161
|
-
U1["9Router (Opus 1M context)"]
|
|
162
|
-
U2["OpenRouter (Sonnet thinking)"]
|
|
163
|
-
U3["Local OpenAI Server (:8000)"]
|
|
164
|
-
U4["Official Google Endpoint"]
|
|
165
|
-
end
|
|
166
|
-
|
|
167
|
-
C1 --> SLOT1 --> U1
|
|
168
|
-
C2 --> SLOT2 --> U2
|
|
169
|
-
C3 --> SLOT3 --> U3
|
|
170
|
-
C4 --> SLOT4 --> U4
|
|
171
|
-
```
|
|
121
|
+
> #### 🔒 CA Security & Origin: Where does the CA come from and how safe is it?
|
|
122
|
+
>
|
|
123
|
+
> - **100% Locally Minted:** The CA certificate (`ca.pem`) and private key (`ca.key`) are generated entirely on your own machine using your local OpenSSL (`blindfold/make-certs.sh`). No keys are downloaded from the internet, and the private key is stored locally with strict `0600` permissions.
|
|
124
|
+
> - **Zero OS Trust Store Tampering:** Unlike tools like Charles or Fiddler, LLM Switcher **NEVER installs anything into your system or OS root certificate store** (no Windows Certificate Store, no macOS Keychain, no Linux `/etc/ssl/certs`). It requires **zero Administrator or sudo privileges**.
|
|
125
|
+
> - **Process-Scoped Trust Only:** The certificate is loaded ephemerally into the memory of `claude` (via `NODE_EXTRA_CA_CERTS`) and `codex` (via `CODEX_CA_CERTIFICATE`). Your browsers, banking apps, git, and other terminal sessions never trust this CA.
|
|
126
|
+
> - **Cryptographic Name Constraints:** The CA is minted with explicit X.509 `nameConstraints` strictly permitting only three domains: `api.anthropic.com`, `api.openai.com`, and `chatgpt.com`. Even if the local private key were compromised, standard TLS verifiers will reject it for any other domain (Google, GitHub, your bank).
|
|
127
|
+
> - **VPN & Corporate CA Preservation:** If your workstation already has a company CA in `NODE_EXTRA_CA_CERTS`, `ensure-ca-bundle.mjs` merges both into a combined bundle so internal corporate proxies never break.
|
|
172
128
|
|
|
173
129
|
---
|
|
174
130
|
|
|
@@ -184,68 +140,21 @@ flowchart LR
|
|
|
184
140
|
- Fixes orphaned `tool_result` blocks caused by aggressive prompt pruners (RTK, Headroom, Ponytail) before sending to Anthropic/OpenAI upstream.
|
|
185
141
|
- Automatically restores thinking parameters if an intermediary tool stripped them.
|
|
186
142
|
- Merges consecutive same-role turns to enforce strict alternating turn requirements.
|
|
187
|
-
- **
|
|
188
|
-
- **Zero Config Mutation:** Never
|
|
143
|
+
- **Context windows come from the model:** the switcher no longer forces a 1M window or an auto-compact limit, and it writes no model name into your environment. Claude Code sizes its own session from the window of the official model you pick, and a backend with a smaller window than that model can overflow in a long session. `model1M` now only decides what `/v1/models` reports.
|
|
144
|
+
- **Zero Config Mutation:** Never reads or writes `~/.claude/settings.json` or `~/.codex/config.toml`, and writes no environment variable and no `--config` argument that a coding tool reads as configuration. The tool reaches the gateway only through the interceptor, so no provider warning banner appears.
|
|
189
145
|
- **Live Request / Response Inspector:** Built-in dashboard tab displaying real-time requests, latency, token consumption, prompt previews, and thinking blocks.
|
|
190
146
|
- **Native Background Service:** Install and run as an OS background daemon on Windows (Task Scheduler), macOS (launchd), or Linux (systemd).
|
|
191
147
|
|
|
192
148
|
---
|
|
193
149
|
|
|
194
|
-
##
|
|
195
|
-
|
|
196
|
-
- **README.** A new section, "Self-improvement with intact", explains how this gateway and intact correct their own faults. intact is now public and on npm as `intact-gateway`.
|
|
197
|
-
|
|
198
|
-
### Changes in 1.1.9
|
|
150
|
+
## Recent Highlights (v1.2.0)
|
|
199
151
|
|
|
200
|
-
- **
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
- **
|
|
205
|
-
-
|
|
206
|
-
|
|
207
|
-
### Changes in 1.1.7
|
|
208
|
-
|
|
209
|
-
- **Codex tools.** With `publicModels` set, Codex lost its tools and ended after one answer. The model catalog copied the metadata of a real OpenAI model, which puts Codex in the "Responses Lite" form. The catalog now keeps Codex in its direct tool mode, and the gateway also reads tools that arrive as an `additional_tools` input item.
|
|
210
|
-
|
|
211
|
-
### Changes in 1.1.6
|
|
212
|
-
|
|
213
|
-
- **Codex over WebSocket.** The gateway keeps the turns of each WebSocket session. A turn that sends `previous_response_id` gets the earlier turns back, so Codex no longer loses the task after the first tool call. An unknown id fails the turn with `previous_response_not_found`.
|
|
214
|
-
- **Codex warmup.** A `response.create` frame with `generate: false` gets a local answer. It no longer spends a model call.
|
|
215
|
-
- **Codex model name.** A profile without `publicModels` no longer sends `OpenAI-Model: main` in the handshake, and `switch codex` and `switch doctor` warn about it. Codex read `main` as a reroute and showed a false "high-risk cyber activity" warning. See "Codex-first setup".
|
|
216
|
-
- **Contract lab.** The gateway uploads each sample with the key that opened its trace. intact refused the 1.1.5 uploads with `HTTP 404 trace not found`.
|
|
217
|
-
- **Version stamp.** The stamp uses the last commit only when the checkout has no changes. Otherwise it uses the newest file time.
|
|
218
|
-
|
|
219
|
-
### Changes in 1.1.5
|
|
220
|
-
|
|
221
|
-
- **Contract lab privacy.** Masking now uses an allowlist. Every string value is masked except the enum values that intact reads. 1.1.4 masked a list of content keys and missed 16 fields (citations, document and web titles, web search queries, logprobs tokens, file names and URIs, stop sequences, participant names, error messages, tool descriptions).
|
|
222
|
-
|
|
223
|
-
### Changes in 1.1.4
|
|
224
|
-
|
|
225
|
-
- **Contract lab privacy.** A sample that the gateway sends to intact carries no client content. Prompts, answers, tool arguments and results, files and user ids are masked on this machine before the upload. The fields that intact needs for analysis stay.
|
|
226
|
-
|
|
227
|
-
### Changes in 1.1.3
|
|
228
|
-
|
|
229
|
-
- **Dashboard.** Opened without its access token, the page no longer waits on "Checking status...". It says that it is locked and names `switch ui`, which opens it with the token.
|
|
230
|
-
- **Dashboard.** The Codex model names (session, review, subagent) are on the **Models** tab, next to the other model settings. The **Blindfold** tab holds only the interceptor settings.
|
|
231
|
-
|
|
232
|
-
### Changes in 1.1.2
|
|
233
|
-
|
|
234
|
-
- **npm package.** Install with `npm install -g llm-switcher` and run `switch`. An npm install keeps its data in `~/.llm-switcher`, so an upgrade does not erase your configuration. A git checkout keeps its data next to the code, as before.
|
|
235
|
-
- **Contract lab.** The gateway can send a small sample of complete exchanges to an [intact](https://github.com/louisphamdev/intact) server, which finds fields that the converter loses. It is off by default. See "Contract lab" below.
|
|
236
|
-
- **macOS.** `blindfold/make-certs.sh` now runs with LibreSSL, the default `openssl` on macOS.
|
|
237
|
-
- **Upgrade from 1.1.0 or older.** A running gateway older than 1.1.1 cannot prove its identity. `switch` now names it and does not stop it. Stop it by hand once, then run `switch on`.
|
|
238
|
-
- **Tests.** `npm test` runs only `tests/**/*.test.mjs`, also on Node.js 18 and 20.
|
|
239
|
-
|
|
240
|
-
### Earlier changes
|
|
241
|
-
|
|
242
|
-
- The desktop dashboard now uses a compact developer-tool layout. It has clearer route controls, keyboard-accessible tabs, labeled model fields, and no decorative emoji.
|
|
243
|
-
- Codex profiles now use three documented roles: `main`, `review`, and `subagent`.
|
|
244
|
-
- The Codex shim passes official configuration overrides for `model`, `review_model`, `agents.default_subagent_model`, `model_context_window`, and `model_auto_compact_token_limit`.
|
|
245
|
-
- Legacy profile keys remain readable. An explicit empty role now clears its legacy fallback.
|
|
246
|
-
- Claude Opus model IDs use version-independent reasoning detection.
|
|
247
|
-
|
|
248
|
-
The Codex shim no longer relies on `CODEX_MODEL`, `CODEX_MAX_CONTEXT_TOKENS`, or `CODEX_AUTO_COMPACT_WINDOW`. Codex does not document these environment variables. See the official [configuration reference](https://developers.openai.com/codex/config-reference/) and [advanced configuration guide](https://developers.openai.com/codex/config-advanced/).
|
|
152
|
+
- **Zero-Mutation Interceptor:** All traffic routes via `HTTPS_PROXY` without editing client configs (`~/.claude/settings.json`, `~/.codex/config.toml`).
|
|
153
|
+
- **Concurrent Multi-Tool Support:** Simultaneously configures `{ claude, codex }` profiles with dynamic, zero-downtime switching (`POST /_control/active-tools`).
|
|
154
|
+
- **Auto-Discovery & Dynamic Model Catalog:** Discovers official models for Claude Code & Codex; detects tool version updates and refreshes mappings on the fly (`switch models`).
|
|
155
|
+
- **Self-Healing Schemas:** Auto-repairs `{}` empty schemas for Gemini/Vertex and restores stripped `thinking` tokens.
|
|
156
|
+
- **Windows Reliability:** Strict CRLF `.cmd` shims, CA path resolution, and zero-drift subroutine routing.
|
|
157
|
+
- *For older releases (v1.1.2 – v1.1.10), see [CHANGELOG.md](CHANGELOG.md).*
|
|
249
158
|
|
|
250
159
|
---
|
|
251
160
|
|
|
@@ -302,20 +211,28 @@ Open the Web Dashboard at: **[http://127.0.0.1:3456/ui](http://127.0.0.1:3456/ui
|
|
|
302
211
|
|
|
303
212
|
## CLI Integration
|
|
304
213
|
|
|
305
|
-
###
|
|
214
|
+
### The environment comes from the shims, not from your shell
|
|
215
|
+
|
|
216
|
+
Do **not** add a `source env.sh` or `call env.cmd` line to `~/.bashrc`, `~/.zshrc` or a wrapper
|
|
217
|
+
script. Those two files are neutral stubs now: a comment line and nothing else, so an old rc line
|
|
218
|
+
keeps running and can never re-introduce a base URL.
|
|
306
219
|
|
|
307
|
-
|
|
220
|
+
Each tool gets its own file instead, and only the matching shim loads it:
|
|
308
221
|
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
```bash
|
|
315
|
-
source ~/.llm-switcher/env.sh
|
|
316
|
-
```
|
|
222
|
+
| File | Loaded by |
|
|
223
|
+
| --- | --- |
|
|
224
|
+
| `env-claude.sh` / `env-claude.cmd` | the `claude` shim |
|
|
225
|
+
| `env-codex.sh` / `env-codex.cmd` | the `codex` shim |
|
|
226
|
+
| `env.sh` / `env.cmd` | nobody. Neutral stub, kept only so an old rc line stays silent |
|
|
317
227
|
|
|
318
|
-
|
|
228
|
+
An empty per-tool file means that tool is off: the shim then leaves the environment alone and the
|
|
229
|
+
tool reaches its official endpoint. Before it loads anything, the shim also scrubs a stale
|
|
230
|
+
`ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL` or `ANTHROPIC_DEFAULT_<TIER>_MODEL` inherited from an older
|
|
231
|
+
release or from your own shell, so `switch off` really is off.
|
|
232
|
+
|
|
233
|
+
In practice you never call these files. `switch shim install` puts `~/.llm-switcher/bin` on `PATH`,
|
|
234
|
+
and every `claude` and `codex` invocation — including `claude --resume` in a brand-new terminal —
|
|
235
|
+
runs the shim, which injects the variables into that one process.
|
|
319
236
|
|
|
320
237
|
---
|
|
321
238
|
|
|
@@ -327,21 +244,11 @@ For a checkout, use the same files in the checkout folder.
|
|
|
327
244
|
node "path\to\llm-switcher\switch.mjs" %*
|
|
328
245
|
```
|
|
329
246
|
|
|
330
|
-
2.
|
|
247
|
+
2. Install the shim and put its directory **before** the real Claude Code directory in your User `PATH` (System Properties → Environment Variables), then open a new terminal:
|
|
331
248
|
```cmd
|
|
332
|
-
|
|
333
|
-
IF EXIST "%USERPROFILE%\.llm-switcher\active.flag" (
|
|
334
|
-
SET "ANTHROPIC_BASE_URL=http://127.0.0.1:3456"
|
|
335
|
-
SET "CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1"
|
|
336
|
-
)
|
|
337
|
-
IF EXIST "%USERPROFILE%\.llm-switcher\1m.flag" (
|
|
338
|
-
SET /P M1M=<"%USERPROFILE%\.llm-switcher\1m.flag"
|
|
339
|
-
IF "!M1M!"=="" SET "M1M=opus[1m]"
|
|
340
|
-
SET "ANTHROPIC_MODEL=!M1M!"
|
|
341
|
-
SET "CLAUDE_CODE_AUTO_COMPACT_WINDOW=900000"
|
|
342
|
-
)
|
|
249
|
+
switch shim install
|
|
343
250
|
```
|
|
344
|
-
>
|
|
251
|
+
> Do **not** patch `claude.cmd` in your npm global directory: npm rewrites it on every update, and the shim is what injects the gateway URL. `settings.json` is never touched, so no "custom API" banner appears.
|
|
345
252
|
|
|
346
253
|
---
|
|
347
254
|
|
|
@@ -359,21 +266,7 @@ codex
|
|
|
359
266
|
|
|
360
267
|
On Windows, add `%USERPROFILE%\.llm-switcher\bin` before the real Codex directory in `PATH`. Open a new terminal after the change.
|
|
361
268
|
|
|
362
|
-
The shim
|
|
363
|
-
|
|
364
|
-
| Profile role | Codex configuration key | Name the CLI receives |
|
|
365
|
-
|---|---|---|
|
|
366
|
-
| `main` | `model` | `publicModels[0]` |
|
|
367
|
-
| `review` | `review_model` | `publicModels[1]` |
|
|
368
|
-
| `subagent` | `agents.default_subagent_model` | `publicModels[2]` |
|
|
369
|
-
|
|
370
|
-
**A profile that serves Codex must have `publicModels`.** Without it, the gateway has no official name for Codex. It then writes no model catalog and sends no `OpenAI-Model` header, and Codex shows two false warnings: "Model metadata for `<model>` not found" and "Your account was flagged for potentially high-risk cyber activity". `switch codex` and `switch doctor` warn when the Codex profile has no `publicModels`.
|
|
371
|
-
|
|
372
|
-
**Codex never receives an internal name.** The slot aliases `main`, `review` and `subagent` stay inside the gateway. The CLI receives the official model names from `publicModels`, and `mapModel` resolves each one back to its slot. Set `codexRoles` in the profile when you want a different pairing than the order of that list.
|
|
373
|
-
|
|
374
|
-
The shim also passes `model_catalog_json`. That file is generated from `publicModels` on every profile change, and the `/model` picker reads it. The picker never calls `/v1/models`. `/v1/models` serves the same entries, and each window follows `model1M` for its slot.
|
|
375
|
-
|
|
376
|
-
It also passes `openai_base_url` for local routing. If 1M context is enabled for `main`, it passes `model_context_window=1000000` and `model_auto_compact_token_limit=900000`. Command-line overrides have higher precedence than user and project configuration. Re-run `switch shim install` after upgrading an older checkout.
|
|
269
|
+
The shim never edits `~/.codex/config.toml`. Codex reaches the gateway cleanly through `HTTPS_PROXY` and the interceptor, retaining its own model names and context windows. Model catalog entries are served dynamically at `/v1/models`.
|
|
377
270
|
|
|
378
271
|
### Blindfold mode (optional)
|
|
379
272
|
|
|
@@ -386,11 +279,25 @@ base URL is overridden to http://127.0.0.1:3456/v1. Selecting models may not be
|
|
|
386
279
|
Blindfold mode removes that line. Codex keeps its official endpoint, and the switcher intercepts the network hop instead. It needs no administrator rights, no certificate in a system trust store, and no change to `~/.codex/config.toml`.
|
|
387
280
|
|
|
388
281
|
```bash
|
|
389
|
-
bash "$(npm root -g)/llm-switcher/blindfold/make-certs.sh"
|
|
390
|
-
#
|
|
282
|
+
bash "$(npm root -g)/llm-switcher/blindfold/make-certs.sh" # once; a checkout runs blindfold/make-certs.sh
|
|
283
|
+
# the port is already top-level in config.json: "blindfold": { "port": 3457 }
|
|
391
284
|
switch codex <profile> # the gateway starts the interceptor
|
|
392
285
|
```
|
|
393
286
|
|
|
287
|
+
One interceptor serves both tools. It routes by the host of the CONNECT request and by the path,
|
|
288
|
+
and by nothing else:
|
|
289
|
+
|
|
290
|
+
| CONNECT host | Paths to this switcher's gateway | Everything else |
|
|
291
|
+
| --- | --- | --- |
|
|
292
|
+
| `api.anthropic.com` | `/v1/messages`, `/v1/messages/...` | to `api.anthropic.com`, unchanged |
|
|
293
|
+
| `api.openai.com` | `/v1/responses`, `/v1/responses/...`, `/v1/models`, `/v1/models/...` | to `api.openai.com`, unchanged |
|
|
294
|
+
| `chatgpt.com` | `/backend-api/codex/...`, forwarded as `/v1` | to `chatgpt.com`, unchanged |
|
|
295
|
+
|
|
296
|
+
There is no `blindfoldHost` and no `blindfoldPrefix` any more: that table is the routing, and it is
|
|
297
|
+
not a profile setting. A CONNECT host outside the table is tunneled untouched; a request whose
|
|
298
|
+
`Host` header names a different host than its CONNECT target gets `421` and opens no upstream
|
|
299
|
+
connection. See the [full table and the certificate rules](docs/cross-platform.md).
|
|
300
|
+
|
|
394
301
|
The gateway owns the interceptor: it starts it at boot and after every change, and `switch off` stops it. If the certificates are missing, or another process holds the gateway or interceptor port, `switch` refuses the activation and writes no file.
|
|
395
302
|
|
|
396
303
|
Read [📖 `docs/codex-blindfold.md`](docs/codex-blindfold.md) before you turn it on. The guide explains the interception scope, the risk of holding a private CA, and how to go back. It opens with three diagrams:
|
|
@@ -399,6 +306,19 @@ Read [📖 `docs/codex-blindfold.md`](docs/codex-blindfold.md) before you turn i
|
|
|
399
306
|
- [Model name resolution](docs/diagrams/codex-model-name-resolution.html) — which name the CLI sees, and where it resolves
|
|
400
307
|
- [Lifecycle under switch](docs/diagrams/blindfold-switch-lifecycle.html) — activation, refusal, and shutdown
|
|
401
308
|
|
|
309
|
+
### What you give up
|
|
310
|
+
|
|
311
|
+
- **1M context follows the official model's window.** The switcher no longer forces a 1M window
|
|
312
|
+
or an auto-compact limit, and it writes no model name into your environment. Claude Code sizes
|
|
313
|
+
its own session from the window of the model you pick. A backend whose window is smaller than
|
|
314
|
+
that model can overflow in a long session. `model1M` now only decides what `/v1/models` reports.
|
|
315
|
+
- **Codex needs certificates once.** Blindfold mode is what keeps Codex on its official endpoint,
|
|
316
|
+
and it needs a private CA plus a leaf naming the three hosts above. Skip it and Codex shows the
|
|
317
|
+
`base URL is overridden` line on its `/model` screen instead.
|
|
318
|
+
- **The interceptor decrypts the three hosts it serves.** It refuses a CONNECT to a local or
|
|
319
|
+
private address, but every process on the machine can reach it. `docs/codex-blindfold.md` opens
|
|
320
|
+
with the scope, the risk of holding a private CA, and how to go back.
|
|
321
|
+
|
|
402
322
|
---
|
|
403
323
|
|
|
404
324
|
## Co-existence with Token Optimizers (RTK, Headroom, Ponytail)
|
|
@@ -409,8 +329,8 @@ If you use prompt-trimming tools like **Headroom**, **Ponytail**, or **RTK (Rust
|
|
|
409
329
|
3. **LLM Switcher** serves as the protective final gateway before the internet:
|
|
410
330
|
- **Repairs Broken Schemas:** Fixes orphaned `tool_result` turns and consecutive same-role turns caused by aggressive history pruning.
|
|
411
331
|
- **Restores Stripped Thinking:** Detects reasoning models and restores thinking parameters if an optimizer stripped them to save tokens.
|
|
412
|
-
- **
|
|
413
|
-
- **Converts Protocols:** Bridges 2-way traffic to your target upstream (9Router, OpenRouter, Vertex, etc.).
|
|
332
|
+
- **Context windows follow the model:** reports the official window of each model; nothing is injected.
|
|
333
|
+
- **Converts Protocols:** Bridges 2-way traffic to your target upstream (intact, 9Router, OpenRouter, Vertex, etc.).
|
|
414
334
|
|
|
415
335
|
### Interoperability Test Report & Benchmark
|
|
416
336
|
|
|
@@ -442,8 +362,8 @@ To guarantee that AI coding agents (Claude Code, Cursor, Windsurf, Opencode) and
|
|
|
442
362
|
### 1. The Agent Skill (`skills/llm-switcher/SKILL.md`)
|
|
443
363
|
A standardized Agent Skill teaching the LLM:
|
|
444
364
|
- **Mandatory routing:** All LLM traffic and token compression tools (Headroom, RTK, Ponytail) MUST target `http://127.0.0.1:3456`.
|
|
445
|
-
- **Zero-mutation policy:** Prohibits the agent from editing `~/.claude/settings.json`
|
|
446
|
-
- **Sub-process safety:**
|
|
365
|
+
- **Zero-mutation policy:** Prohibits the agent from editing `~/.claude/settings.json` — the switcher never reads or writes that file either.
|
|
366
|
+
- **Sub-process safety:** Never tells a sub-agent to `source env.sh`; the shims in `~/.llm-switcher/bin` inject the proxy variables into the tool process itself and scrub anything stale first.
|
|
447
367
|
|
|
448
368
|
Install globally for Opencode / Claude:
|
|
449
369
|
```bash
|
|
@@ -456,7 +376,7 @@ cp -r skills/llm-switcher ~/.claude/skills/
|
|
|
456
376
|
|
|
457
377
|
### 2. The MCP Server (`mcp.mjs`)
|
|
458
378
|
A zero-dependency Model Context Protocol (MCP) server communicating over `stdio`:
|
|
459
|
-
- `switcher_status`: Read live active profiles and
|
|
379
|
+
- `switcher_status`: Read live active profiles and gateway state.
|
|
460
380
|
- `switcher_audit`: Audit the environment to detect rogue direct outbound calls or unrouted token compressors.
|
|
461
381
|
- `switcher_switch_profile`: Programmatically switch a CLI's active profile.
|
|
462
382
|
- `switcher_recent_logs`: Inspect recent request logs, token usage, and thinking extraction.
|
|
@@ -482,36 +402,42 @@ Add the server to your MCP configuration (for example `opencode.jsonc`, `claude_
|
|
|
482
402
|
switch ui # Open the Web UI dashboard in your browser
|
|
483
403
|
switch status # Display status for all active CLI targets
|
|
484
404
|
switch doctor # Audit environment, settings & routing
|
|
485
|
-
switch on [profile] # Start the gateway and activate a profile
|
|
486
|
-
switch <profile> # Activate profile for
|
|
487
|
-
switch claude <profile> # Set active profile
|
|
488
|
-
switch codex <profile> # Set active profile
|
|
489
|
-
switch openai <profile> # Set active profile specifically for OpenAI Chat
|
|
490
|
-
switch vertex <profile> # Set active profile specifically for Vertex / Gemini
|
|
405
|
+
switch on [profile] # Start the gateway and activate a profile
|
|
406
|
+
switch <profile> # Activate a profile for both tools
|
|
407
|
+
switch claude <profile> # Set the active profile for Claude Code only
|
|
408
|
+
switch codex <profile> # Set the active profile for Codex only
|
|
491
409
|
switch port <number> # Change the gateway port (restarts it if running)
|
|
492
410
|
switch service install # Install OS background autostart service (Windows / macOS / Linux)
|
|
493
411
|
switch service uninstall # Remove background autostart service
|
|
494
412
|
switch shim install # Route new Claude and Codex sessions through the gateway
|
|
495
413
|
switch shim status # Verify shims + detect running sessions that bypass the gateway
|
|
496
414
|
switch shim uninstall # Remove the launcher shims
|
|
497
|
-
switch off
|
|
415
|
+
switch off # Stop everything and restore the official endpoints
|
|
416
|
+
switch off claude # Turn Claude Code off; Codex keeps running
|
|
417
|
+
switch off codex # Turn Codex off; Claude Code keeps running
|
|
498
418
|
switch contract-probe [--model m] # Drive the contract-lab variants through the gateway
|
|
499
419
|
switch contract-check # Turn the open contract findings into failing tests
|
|
500
420
|
```
|
|
501
421
|
|
|
422
|
+
Targets are `claude` and `codex`, and they are the only two. A profile is one tool: `tool` is
|
|
423
|
+
either `"claude"` or `"codex"` (or `null` for a profile that is switched off), so `switch claude` and
|
|
424
|
+
`switch codex` can never point at the same profile by accident. There is no `openai` or `vertex`
|
|
425
|
+
target any more — the input routes those names stood for are gone.
|
|
426
|
+
|
|
502
427
|
The service runs without your shell. `switch service install` therefore copies `CLAUDE_CONFIG_DIR`, `LLM_SWITCHER_CONFIG`, `LLM_SWITCHER_STATE_DIR` and `LLM_SWITCHER_BLINDFOLD_CERTS` into the systemd unit or the launchd plist when they are set. The Windows task cannot carry them; set them as User environment variables instead. If the installed definition differs from the new one, for example after a hand edit, the old file is kept as `<file>.bak`. On Windows the task is created from an XML definition, so paths with spaces need no extra quoting and the task has no run-time limit. This Windows path is not tested on Windows yet.
|
|
503
428
|
|
|
504
429
|
### Resumed sessions & the shim (important)
|
|
505
430
|
|
|
506
|
-
`switch on` writes `env
|
|
507
|
-
`~/.claude/settings.json` —
|
|
508
|
-
|
|
431
|
+
`switch on` writes `env-claude.*` and `env-codex.*` and writes nothing into
|
|
432
|
+
`~/.claude/settings.json` — the switcher never reads or writes that file, which is what keeps
|
|
433
|
+
Claude Code from showing its "custom API" banner.
|
|
434
|
+
The side effect: a CLI started without the shim has **no**
|
|
509
435
|
`ANTHROPIC_BASE_URL`, so it talks to the provider directly and skips the gateway
|
|
510
|
-
(no Healer, no
|
|
436
|
+
(no Healer, no pooled quota). `claude --resume` in a fresh terminal is the
|
|
511
437
|
classic case.
|
|
512
438
|
|
|
513
|
-
The shim closes that hole. It installs tiny wrappers in `~/.llm-switcher/bin` that
|
|
514
|
-
|
|
439
|
+
The shim closes that hole. It installs tiny wrappers in `~/.llm-switcher/bin` that inject the
|
|
440
|
+
tool's own environment and then `exec` the real binary:
|
|
515
441
|
|
|
516
442
|
```bash
|
|
517
443
|
switch shim install
|
|
@@ -521,12 +447,15 @@ switch shim status # verify
|
|
|
521
447
|
|
|
522
448
|
Behaviour:
|
|
523
449
|
|
|
524
|
-
- **
|
|
525
|
-
is routed through the gateway.
|
|
526
|
-
-
|
|
527
|
-
`--config`. It does not depend on unsupported `CODEX_*` variables.
|
|
528
|
-
- **Gateway OFF** (no `active.flag`) → the wrapper is fully transparent and runs the real
|
|
450
|
+
- **Tool active** → the shim loads that tool's own env file, so every invocation (including
|
|
451
|
+
`--resume`) is routed through the gateway.
|
|
452
|
+
- **Tool off** (its env file is empty) → the shim is fully transparent and runs the real
|
|
529
453
|
binary untouched; it never forces routing.
|
|
454
|
+
- Before it loads anything, the shim scrubs a stale `ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL` or
|
|
455
|
+
`ANTHROPIC_DEFAULT_<TIER>_MODEL` inherited from an older release or from your shell, so a
|
|
456
|
+
variable cannot survive a `switch off`.
|
|
457
|
+
- No `--config` arguments are passed to Codex. The shim changes nothing on Codex's side but the
|
|
458
|
+
environment.
|
|
530
459
|
- The real binary is located with the shim directory stripped from `PATH`, so it can never
|
|
531
460
|
call itself recursively. If no real binary is found it exits `127` with a clear message
|
|
532
461
|
instead of failing silently.
|
|
@@ -544,21 +473,19 @@ re-opened from a shell where the shim is on `PATH`.
|
|
|
544
473
|
```jsonc
|
|
545
474
|
{
|
|
546
475
|
"port": 3456,
|
|
547
|
-
"activeProfile": "9router",
|
|
548
476
|
"activeProfiles": {
|
|
549
|
-
"
|
|
550
|
-
"
|
|
551
|
-
"openai-chat": "9router", // Active profile for OpenAI Chat
|
|
552
|
-
"vertex": "gemini-profile" // Active profile for Vertex / Gemini
|
|
477
|
+
"claude": "claude-default", // Active profile for Claude Code (/v1/messages)
|
|
478
|
+
"codex": "codex-default" // Active profile for Codex (/v1/responses)
|
|
553
479
|
},
|
|
480
|
+
"blindfold": { "port": 3457 }, // Interceptor port. Top-level, optional, default 3457
|
|
554
481
|
"profiles": {
|
|
555
|
-
"
|
|
556
|
-
"name": "
|
|
482
|
+
"claude-default": {
|
|
483
|
+
"name": "Intact Gateway",
|
|
557
484
|
"mode": "convert", // hybrid | convert | direct
|
|
558
|
-
"
|
|
485
|
+
"tool": "claude", // claude | codex | null (profile is off)
|
|
559
486
|
"outFormat": "openai-chat", // openai-chat | anthropic | vertex
|
|
560
487
|
"thinkingMode": "auto", // auto | native | off (see Advanced Options)
|
|
561
|
-
"baseURL": "https://api.9router.com/v1
|
|
488
|
+
"baseURL": "https://intact.example.com/v1", // or https://api.9router.com/v1
|
|
562
489
|
"apiKey": "sk-...",
|
|
563
490
|
"defaultModels": {
|
|
564
491
|
"opus": "ag/claude-opus-4-6-thinking",
|
|
@@ -571,12 +498,20 @@ re-opened from a shell where the shim is on `PATH`.
|
|
|
571
498
|
"sonnet": true,
|
|
572
499
|
"haiku": false,
|
|
573
500
|
"fable": true
|
|
574
|
-
}
|
|
575
|
-
|
|
501
|
+
}
|
|
502
|
+
},
|
|
503
|
+
"codex-default": {
|
|
504
|
+
"name": "Codex via router",
|
|
505
|
+
"mode": "convert",
|
|
506
|
+
"tool": "codex",
|
|
507
|
+
"outFormat": "vertex",
|
|
508
|
+
"baseURL": "https://YOUR-GATEWAY/v1",
|
|
509
|
+
"apiKey": "sk-...",
|
|
510
|
+
// A profile that serves Codex MUST have publicModels (see Codex-first setup).
|
|
576
511
|
"publicModels": ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"], // official names for main, review, subagent
|
|
577
512
|
"codexRoles": { "review": "gpt-5.6-sol" }, // optional: pair one role with another public name
|
|
578
|
-
"
|
|
579
|
-
"
|
|
513
|
+
"defaultModels": { "main": "gemini-3.8-flash", "review": "gemini-3.7-flash-medium", "subagent": "gemini-3.6-flash-low" },
|
|
514
|
+
"model1M": { "main": false, "review": false, "subagent": false }
|
|
580
515
|
}
|
|
581
516
|
},
|
|
582
517
|
"debug": false,
|
|
@@ -588,6 +523,16 @@ re-opened from a shell where the shim is on `PATH`.
|
|
|
588
523
|
}
|
|
589
524
|
```
|
|
590
525
|
|
|
526
|
+
`tool` replaces `inFormat`: it says which tool a profile serves, not what its upstream speaks.
|
|
527
|
+
There is no top-level `activeProfile`, no `blindfold`, `blindfoldPort`, `blindfoldHost` or
|
|
528
|
+
`blindfoldPrefix` inside a profile, and no `openai-chat` or `vertex` key in `activeProfiles`.
|
|
529
|
+
|
|
530
|
+
An older `config.json` is rewritten once, on load, through a compare-and-swap that refuses to
|
|
531
|
+
touch a file somebody else changed first. When two profiles would collapse onto the same key the
|
|
532
|
+
migration stops, names the clashing keys on the CLI and in the dashboard, and leaves the file
|
|
533
|
+
exactly as it was: every mutating `switch` command then exits non-zero, `switch off` and
|
|
534
|
+
`switch doctor` still work, and fixing the file makes everything work again.
|
|
535
|
+
|
|
591
536
|
### Contract lab
|
|
592
537
|
|
|
593
538
|
The contract lab finds fields that the converter loses. It is off by default.
|
|
@@ -617,7 +562,7 @@ Because of this split, a provider fingerprint is never a rule in this gateway. I
|
|
|
617
562
|
| `LLM_SWITCHER_CONFIG=/path/config.json` | Use a config file outside the data folder (the proxy, `switch` and `mcp.mjs` all honour it). |
|
|
618
563
|
| `--port <n>` / `LLM_SWITCHER_PORT` | Override the listening port (priority: flag > env > `config.port`). |
|
|
619
564
|
| `x-llm-profile: <key>` header (alias `x-profile`) or `?profile=<key>` | Route a single request through a specific profile. An unknown key returns HTTP 400 instead of silently falling back. |
|
|
620
|
-
| `profile.thinkingMode` | `auto` (default, for gateways like 9Router): restore stripped thinking, inject a `<think>` guide for non-reasoning models, send `thinking` + `reasoning_effort`. `native` (strict OpenAI APIs): send only `reasoning_effort` when the client asks, never touch the prompt, use `max_completion_tokens`. `off`: never send reasoning parameters. |
|
|
565
|
+
| `profile.thinkingMode` | `auto` (default, for gateways like intact or 9Router): restore stripped thinking, inject a `<think>` guide for non-reasoning models, send `thinking` + `reasoning_effort`. `native` (strict OpenAI APIs): send only `reasoning_effort` when the client asks, never touch the prompt, use `max_completion_tokens`. `off`: never send reasoning parameters. |
|
|
621
566
|
| `profile.endpoints.countTokens` | Override the Anthropic `count_tokens` URL. |
|
|
622
567
|
| `profile.endpoints` | Override upstream URLs per format: `{ "openai-chat": "...", "anthropic": "...", "vertex": "https://.../models/{model}:{action}" }`. |
|
|
623
568
|
| `CLAUDE_CONFIG_DIR` | Respected when locating Claude Code's `settings.json`. |
|