llm-switcher 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/.gitattributes +16 -0
  2. package/LICENSE +21 -0
  3. package/README.md +587 -0
  4. package/README.vi.md +585 -0
  5. package/blindfold/blindfold.mjs +633 -0
  6. package/blindfold/make-certs.sh +88 -0
  7. package/blindfold/wsframe.mjs +176 -0
  8. package/codex-catalog-template.json +1 -0
  9. package/config.example.json +84 -0
  10. package/contract-exclusions.json +41 -0
  11. package/contract.mjs +561 -0
  12. package/docs/LLM-RESPONSE-MATRIX.md +165 -0
  13. package/docs/TOKEN-OPTIMIZER-INTEROP.md +110 -0
  14. package/docs/codex-blindfold.md +214 -0
  15. package/docs/cross-platform.md +136 -0
  16. package/docs/diagrams/blindfold-request-routing.html +14972 -0
  17. package/docs/diagrams/blindfold-request-routing.sequence.json +175 -0
  18. package/docs/diagrams/blindfold-switch-lifecycle.html +14958 -0
  19. package/docs/diagrams/blindfold-switch-lifecycle.lifecycle.json +159 -0
  20. package/docs/diagrams/codex-model-name-resolution.html +15005 -0
  21. package/docs/diagrams/codex-model-name-resolution.workflow.json +71 -0
  22. package/docs/response-matrix.json +1131 -0
  23. package/formats.mjs +2308 -0
  24. package/mcp.mjs +340 -0
  25. package/package.json +36 -0
  26. package/proxy.mjs +1743 -0
  27. package/service.mjs +132 -0
  28. package/shim.mjs +292 -0
  29. package/skills/llm-switcher/SKILL.md +88 -0
  30. package/state.mjs +978 -0
  31. package/switch +5 -0
  32. package/switch.cmd +2 -0
  33. package/switch.mjs +930 -0
  34. package/tests/blindfold.test.mjs +307 -0
  35. package/tests/blindfold.wire.test.mjs +170 -0
  36. package/tests/contract/run.test.mjs +214 -0
  37. package/tests/contract-check.test.mjs +458 -0
  38. package/tests/contract-lab.test.mjs +755 -0
  39. package/tests/datadir.test.mjs +37 -0
  40. package/tests/formats.test.mjs +794 -0
  41. package/tests/gateway.e2e.test.mjs +999 -0
  42. package/tests/helpers.mjs +24 -0
  43. package/tests/lifecycle.test.mjs +416 -0
  44. package/tests/live-optimizer-interop.mjs +205 -0
  45. package/tests/mcp.test.mjs +91 -0
  46. package/tests/service.test.mjs +69 -0
  47. package/tests/shim.test.mjs +228 -0
  48. package/tests/state.test.mjs +675 -0
  49. package/tests/switch.test.mjs +156 -0
  50. package/tests/wsframe.test.mjs +154 -0
  51. package/ui.html +2234 -0
package/.gitattributes ADDED
@@ -0,0 +1,16 @@
1
+ # Line endings are a cross-platform correctness problem here, not a preference.
2
+ # A shell script checked out with CRLF fails on Linux and macOS with
3
+ # "bad interpreter: /usr/bin/env sh^M", so everything the repository executes is
4
+ # stored and checked out with LF. Only the Windows batch launchers need CRLF.
5
+ * text=auto eol=lf
6
+
7
+ switch text eol=lf
8
+ *.sh text eol=lf
9
+ *.mjs text eol=lf
10
+
11
+ *.cmd text eol=crlf
12
+ *.bat text eol=crlf
13
+
14
+ *.png binary
15
+ *.webp binary
16
+ *.pfx binary
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Louis Phạm
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,587 @@
1
+ # LLM Switcher
2
+
3
+ <p align="center">
4
+ <b>Zero-dependency, multi-protocol edge gateway & provider switcher</b><br>
5
+ Seamlessly bridge <b>Claude Code</b>, <b>Codex</b>, OpenAI, and Gemini SDKs to any upstream LLM API.<br>
6
+ Full bi-directional protocol conversion, 1M context unlock, thinking protocol extraction, and edge message healing.
7
+ </p>
8
+
9
+ <p align="center">
10
+ <b>English</b> • <a href="README.vi.md">Tiếng Việt</a>
11
+ </p>
12
+
13
+ <p align="center">
14
+ <img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?logo=node.js&logoColor=white" alt="Node.js 18+">
15
+ <img src="https://img.shields.io/badge/Dependencies-Zero-38bdf8" alt="Zero Dependencies">
16
+ <img src="https://img.shields.io/badge/Context-1%2C000%2C000_tokens-6366f1" alt="1M Context">
17
+ <img src="https://img.shields.io/badge/Multi--Active-Concurrent_CLIs-f59e0b" alt="Multi-Active">
18
+ <img src="https://img.shields.io/badge/License-MIT-gray" alt="License MIT">
19
+ </p>
20
+
21
+ ---
22
+
23
+ > ### 🎯 The Core Problem: Why Generic Proxies Cripple Your AI Coding Tools
24
+ >
25
+ > Every LLM provider uses a **subtly or drastically different API response standard**:
26
+ > - **Anthropic** requires dedicated `thinking` blocks (`thinking_delta` + `signature_delta`), strict alternating turn rules, and typed `tool_use` input schemas.
27
+ > - **OpenAI** streams reasoning as `reasoning_content` delta chunks or `reasoning_details[]`, and formats tools as `tool_calls` with JSON string arguments.
28
+ > - **Google Vertex AI** places reasoning in `candidates[0].content.parts[{thought: true, text, thoughtSignature}]` and tool arguments as raw objects.
29
+ > - **Open-source models (DeepSeek, Qwen, GLM)** often dump chain-of-thought directly into `content` or duplicate fields under conflicting keys.
30
+ >
31
+ > **When coding tools like Claude Code or Codex receive non-native or partially converted responses, they don't just look wrong — the agent's performance degrades catastrophically:**
32
+ > 1. **Lost Chain-of-Thought:** If Claude Code does not receive native `thinking_delta` blocks, it **completely misses the model's internal reasoning**. The agent acts prematurely, skips architectural planning, and produces buggy code.
33
+ > 2. **Broken Tool Execution:** Mismatched stop reasons (`tool_calls` vs `tool_use`) and split argument chunks cause tool execution failures and infinite retries.
34
+ > 3. **Token & Cache Miscounting:** Non-standard usage accounting breaks prompt cache alignment and premature context compaction.
35
+ >
36
+ > Developers often blame the model for "getting dumber" when in reality **their proxy mangled the response protocol.**
37
+ >
38
+ > ### 🛡️ The Solution: Zero-Loss Native Emulation (Subscription-Grade Quality)
39
+ >
40
+ > **LLM Switcher solves this by acting as a high-precision, zero-loss protocol emulator.**
41
+ >
42
+ > It normalizes whatever your upstream provider emits (9Router, OpenRouter, Vertex, DeepSeek) and re-synthesizes it into the **exact native event stream the client agent was built to consume**:
43
+ > - **Claude Code** receives 100% genuine Anthropic SSE events (`message_start` ➔ `thinking_delta` ➔ `signature_delta` ➔ `content_block_start: tool_use` ➔ `message_delta`), performing **identically to an official Anthropic subscription**.
44
+ > - **Codex** receives 100% genuine Responses API events (`response.created` ➔ `output_text.delta` ➔ `function_call` ➔ `response.completed`).
45
+ >
46
+ > **You get the freedom and cost savings of 3rd-party APIs while maintaining 100% official subscription-grade agent intelligence.**
47
+
48
+ ---
49
+
50
+ > ### 💡 Design Philosophy: The Client-Side Edge Companion to 9Router
51
+ >
52
+ > **LLM Switcher intentionally does NOT implement multi-account pooling, key rotation, quota tracking, or provider load balancing.**
53
+ >
54
+ > That heavy lifting belongs to server-side AI routing gateways like **[9Router](https://github.com/decolua/9router)**, which handle centralized account rotation, rate-limit retries, and quota management far more reliably and securely at the server layer.
55
+ >
56
+ > **LLM Switcher is specifically engineered as the optimal client-side edge extension to pair with 9Router (or similar gateways):**
57
+ > - **At the Local Workstation (LLM Switcher):** Translates coding tool protocols (Claude Code `/v1/messages`, Codex `/v1/responses`, Vertex `/v1beta/...`, OpenAI Chat), injects 1M context windows, manages local multi-CLI profiles, and runs the Healer Engine to fix mangled payloads from local prompt compressors (RTK, Headroom, Ponytail).
58
+ > - **At the Server Gateway (9Router):** Manages account pools, API key rotation, load balancing, billing quotas, and global provider failover.
59
+ >
60
+ > This clean division of responsibility keeps LLM Switcher **ultra-lightweight, zero-dependency, and bloat-free** while giving you an unbeatable developer setup.
61
+
62
+ ---
63
+
64
+ ## Architecture & Workflow
65
+
66
+ LLM Switcher sits locally on your workstation (`127.0.0.1:3456`). It acts as the **outermost edge gatekeeper** before requests leave for the internet.
67
+
68
+ ### 1. End-to-End System Topology
69
+
70
+ ```mermaid
71
+ flowchart TD
72
+ subgraph Clients["Dev Clients & Coding CLIs"]
73
+ CC["Claude Code CLI\n(/v1/messages)"]
74
+ CDX["OpenAI Codex CLI\n(/v1/responses)"]
75
+ OAI["OpenAI SDKs / Cursor\n(/v1/chat/completions)"]
76
+ VTX["Gemini / Vertex SDKs\n(/v1beta/models/*)"]
77
+ end
78
+
79
+ subgraph Optimizers["Optional Middle-Layer (Installed in CLI)"]
80
+ OPT["Prompt Optimizers & Trimmers\n(Headroom / RTK / Ponytail)\n[Configured upstream: :3456]"]
81
+ end
82
+
83
+ subgraph Switcher["LLM Switcher (:3456) — Outermost Edge Gatekeeper"]
84
+ direction TB
85
+ ROUTER["Protocol Auto-Detection & Multi-Active Routing"]
86
+ HEALER["Healer Engine\n• Heal orphaned tool_results\n• Restore stripped thinking\n• Merge consecutive turns"]
87
+ IR["Bi-Directional IR Translation\n(4 Client Formats ⟷ 3 Upstream Formats)"]
88
+ M1M["1M Context Unlocker\n& Auto-Compact Thresholds"]
89
+ LOGS["Live Inspector\n(In-Memory Ring Buffer)"]
90
+ ROUTER --> HEALER --> IR --> M1M --> LOGS
91
+ end
92
+
93
+ subgraph Upstream["Internet / Upstream Providers"]
94
+ R9["9Router / Selfhost Gateway"]
95
+ OR["OpenRouter / Together / Groq"]
96
+ ANT["Anthropic Native API"]
97
+ GCP["Google Vertex AI / Gemini"]
98
+ end
99
+
100
+ CC -->|Direct| ROUTER
101
+ CC -.->|Optional| OPT
102
+ CDX -->|Direct| ROUTER
103
+ CDX -.->|Optional| OPT
104
+ OAI --> ROUTER
105
+ VTX --> ROUTER
106
+ OPT -->|Forward to Switcher| ROUTER
107
+
108
+ LOGS -->|Clean Outbound| R9
109
+ LOGS -->|Clean Outbound| OR
110
+ LOGS -->|Clean Outbound| ANT
111
+ LOGS -->|Clean Outbound| GCP
112
+ ```
113
+
114
+ ---
115
+
116
+ ### 2. Bi-Directional IR (Intermediate Representation) Pipeline
117
+
118
+ ```mermaid
119
+ sequenceDiagram
120
+ autonumber
121
+ actor CLI as Client (Claude Code / Codex / SDK)
122
+ participant GW as LLM Switcher (:3456)
123
+ participant IR as IR & Healer Engine
124
+ participant UP as Upstream (9Router / Anthropic / Vertex)
125
+
126
+ CLI->>GW: Inbound Request (Anthropic, Responses, Chat, or Vertex)
127
+ Note over GW,IR: Normalize to Canonical IR
128
+ GW->>IR: parseToIR(clientFormat, payload)
129
+ Note over IR: Healer checks:<br/>1. Repair orphaned tool_results<br/>2. Restore stripped thinking params<br/>3. Reconcile role alternations<br/>4. Apply 1M context limits
130
+ IR->>GW: emitUpstreamBody(outFormat, healedIR)
131
+ GW->>UP: Outbound API Call (fetch with AbortSignal)
132
+ UP-->>GW: Upstream Streaming SSE / JSON Chunks
133
+ Note over GW: normalizeUpstream(chunk)<br/>Extract reasoning_content, <think> tags, usage
134
+ GW->>CLI: Render client-native SSE (e.g. Anthropic thinking_delta + text_delta)
135
+ Note over CLI,GW: Client connection closes (Ctrl+C) -> GW aborts UP instantly!
136
+ ```
137
+
138
+ ---
139
+
140
+ ### 3. Multi-Active CLI Independent Routing
141
+
142
+ You can run **multiple active profiles concurrently** — one profile dedicated to each CLI, without collision:
143
+
144
+ ```mermaid
145
+ flowchart LR
146
+ subgraph Inbound["Incoming Client Calls"]
147
+ C1["Claude Code\n(/v1/messages)"]
148
+ C2["Codex CLI\n(/v1/responses)"]
149
+ C3["OpenAI SDK\n(/v1/chat/completions)"]
150
+ C4["Vertex SDK\n(/v1beta/models/*)"]
151
+ end
152
+
153
+ subgraph Core["LLM Switcher Core (:3456)"]
154
+ SLOT1["Slot: Anthropic\nActive: [9Router]"]
155
+ SLOT2["Slot: Responses\nActive: [OpenRouter]"]
156
+ SLOT3["Slot: OpenAI\nActive: [Local LLM]"]
157
+ SLOT4["Slot: Vertex\nActive: [Off / Official]"]
158
+ end
159
+
160
+ subgraph Egress["Upstream Targets"]
161
+ U1["9Router (Opus 1M context)"]
162
+ U2["OpenRouter (Sonnet thinking)"]
163
+ U3["Local OpenAI Server (:8000)"]
164
+ U4["Official Google Endpoint"]
165
+ end
166
+
167
+ C1 --> SLOT1 --> U1
168
+ C2 --> SLOT2 --> U2
169
+ C3 --> SLOT3 --> U3
170
+ C4 --> SLOT4 --> U4
171
+ ```
172
+
173
+ ---
174
+
175
+ ## Core Features
176
+
177
+ - **Zero-Dependency Architecture:** Built 100% on Node.js standard libraries (`http`, `fs`, `os`, `path`, `fetch`). No npm dependencies, no bundled runtime bloat, cold-start under 50ms.
178
+ - **Bi-Directional Protocol Conversion:**
179
+ - **4 Client Inbound Formats:** Anthropic Messages, OpenAI Chat Completions, Codex Responses API, Vertex `generateContent`.
180
+ - **3 Upstream Outbound Formats:** OpenAI Chat, Anthropic Native, Vertex Native.
181
+ - **Multi-Active CLI Routing:** Run Claude Code on Profile A, Codex on Profile B, and Cursor on Profile C simultaneously on a single gateway instance.
182
+ - **Deep Thinking & Reasoning Extraction:** Tested on 48 live response combinations. Accurately extracts `reasoning_content`, `<think>` tags, Vertex `thought` parts, and signatures into native `thinking_delta` blocks.
183
+ - **Edge Healer Engine (Anti-Collision for Token Optimizers):**
184
+ - Fixes orphaned `tool_result` blocks caused by aggressive prompt pruners (RTK, Headroom, Ponytail) before sending to Anthropic/OpenAI upstream.
185
+ - Automatically restores thinking parameters if an intermediary tool stripped them.
186
+ - Merges consecutive same-role turns to enforce strict alternating turn requirements.
187
+ - **1M Context Window Unlocker:** Follows the profile's `model1M` map per tier: every tier marked 1M gets `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]` (so `/model sonnet`, tier switches and subagents keep 1M, while unmarked tiers stay at 200K), plus auto-compact at `900000`, with built-in visual risk warnings for unsupported models.
188
+ - **Zero Config Mutation:** Never writes endpoints or keys into `~/.claude/settings.json` (it removes only values it wrote itself: `ANTHROPIC_BASE_URL` for its own port and `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]`). Uses launcher flags and environment injection to prevent annoying provider warning banners.
189
+ - **Live Request / Response Inspector:** Built-in dashboard tab displaying real-time requests, latency, token consumption, prompt previews, and thinking blocks.
190
+ - **Native Background Service:** Install and run as an OS background daemon on Windows (Task Scheduler), macOS (launchd), or Linux (systemd).
191
+
192
+ ---
193
+
194
+ ## Changes in this update
195
+
196
+ - The desktop dashboard now uses a compact developer-tool layout. It has clearer route controls, keyboard-accessible tabs, labeled model fields, and no decorative emoji.
197
+ - Codex profiles now use three documented roles: `main`, `review`, and `subagent`.
198
+ - The Codex shim passes official configuration overrides for `model`, `review_model`, `agents.default_subagent_model`, `model_context_window`, and `model_auto_compact_token_limit`.
199
+ - Legacy profile keys remain readable. An explicit empty role now clears its legacy fallback.
200
+ - Claude Opus model IDs use version-independent reasoning detection.
201
+
202
+ The Codex shim no longer relies on `CODEX_MODEL`, `CODEX_MAX_CONTEXT_TOKENS`, or `CODEX_AUTO_COMPACT_WINDOW`. Codex does not document these environment variables. See the official [configuration reference](https://developers.openai.com/codex/config-reference/) and [advanced configuration guide](https://developers.openai.com/codex/config-advanced/).
203
+
204
+ ---
205
+
206
+ ## Quick Start
207
+
208
+ ### 1. Requirements
209
+ - Node.js 18.17 or later.
210
+ - The gateway has no npm dependencies.
211
+
212
+ ### 2. Install and configure
213
+
214
+ **Option A: npm (recommended)**
215
+ ```bash
216
+ npm install -g llm-switcher
217
+
218
+ # Copy the example configuration into your data folder.
219
+ mkdir -p ~/.llm-switcher
220
+ cp "$(npm root -g)/llm-switcher/config.example.json" ~/.llm-switcher/config.json
221
+ ```
222
+
223
+ An npm install keeps `config.json`, `admin.token` and the launch files in `~/.llm-switcher`. An upgrade replaces the package folder only, so your configuration stays.
224
+
225
+ **Option B: git clone**
226
+ ```bash
227
+ git clone https://github.com/louisphamdev/llm-switcher.git
228
+ cd llm-switcher
229
+
230
+ # Copy example config (config.json is git-ignored for safety)
231
+ cp config.example.json config.json
232
+ ```
233
+
234
+ A checkout keeps its data next to the code, as before. To use another folder in either case, set `LLM_SWITCHER_HOME`.
235
+
236
+ Edit `config.json` with your provider base URLs and API keys.
237
+
238
+ With the npm install, run `switch <command>`. With a checkout, run `node switch.mjs <command>` or put the checkout on `PATH`.
239
+
240
+ ### 3. Start the Gateway
241
+ ```bash
242
+ # Start in foreground or background
243
+ node switch.mjs on
244
+
245
+ # Or run directly:
246
+ node proxy.mjs
247
+ ```
248
+
249
+ The repository ships two launchers for the same script: `switch` for Linux and macOS,
250
+ `switch.cmd` for Windows. Put the repository directory on PATH and `switch <command>`
251
+ works the same on all three. Platform differences, and the two features that are not
252
+ available everywhere, are in [📖 `docs/cross-platform.md`](docs/cross-platform.md).
253
+
254
+ Open the Web Dashboard at: **[http://127.0.0.1:3456/ui](http://127.0.0.1:3456/ui)**
255
+
256
+ ---
257
+
258
+ ## CLI Integration
259
+
260
+ ### Universal Environment Loader (`env.cmd` / `env.sh`)
261
+
262
+ Every time you switch profiles, LLM Switcher writes ready-to-use environment loaders:
263
+
264
+ - **Windows (Command Prompt / PowerShell wrapper):**
265
+ ```cmd
266
+ call "path\to\llm-switcher\env.cmd"
267
+ ```
268
+ - **macOS / Linux (Bash / Zsh):**
269
+ ```bash
270
+ source "path/to/llm-switcher/env.sh"
271
+ ```
272
+
273
+ ---
274
+
275
+ ### Claude Code Setup (Windows)
276
+
277
+ 1. Create a quick wrapper in your PATH (e.g. `cc-switch.cmd`):
278
+ ```cmd
279
+ @echo off
280
+ node "path\to\llm-switcher\switch.mjs" %*
281
+ ```
282
+
283
+ 2. Patch your global Claude Code launcher (`claude.cmd` in your npm global directory):
284
+ ```cmd
285
+ SETLOCAL EnableDelayedExpansion
286
+ IF EXIST "path\to\llm-switcher\active.flag" (
287
+ SET "ANTHROPIC_BASE_URL=http://127.0.0.1:3456"
288
+ SET "CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1"
289
+ )
290
+ IF EXIST "path\to\llm-switcher\1m.flag" (
291
+ SET /P M1M=<"path\to\llm-switcher\1m.flag"
292
+ IF "!M1M!"=="" SET "M1M=opus[1m]"
293
+ SET "ANTHROPIC_MODEL=!M1M!"
294
+ SET "CLAUDE_CODE_AUTO_COMPACT_WINDOW=900000"
295
+ )
296
+ ```
297
+ > `SETLOCAL EnableDelayedExpansion` is required for `!M1M!`. npm rewrites `claude.cmd` on every update, so prefer a separate wrapper that runs `call "path\to\llm-switcher\env.cmd"` and then `claude %*`. Only `env.cmd` / `env.sh` carry the per-tier `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]` variables.
298
+
299
+ ---
300
+
301
+ ### Codex-first setup
302
+
303
+ Install the generated shim, put its directory first in `PATH`, and activate a Codex-compatible profile:
304
+
305
+ ```bash
306
+ switch shim install
307
+ export PATH="$HOME/.llm-switcher/bin:$PATH" # Bash or Zsh
308
+ switch codex <profile>
309
+ switch shim status
310
+ codex
311
+ ```
312
+
313
+ On Windows, add `%USERPROFILE%\.llm-switcher\bin` before the real Codex directory in `PATH`. Open a new terminal after the change.
314
+
315
+ The shim does not edit `~/.codex/config.toml`. When the gateway is active, it passes these official command-line overrides to the real Codex binary:
316
+
317
+ | Profile role | Codex configuration key | Name the CLI receives |
318
+ |---|---|---|
319
+ | `main` | `model` | `publicModels[0]` |
320
+ | `review` | `review_model` | `publicModels[1]` |
321
+ | `subagent` | `agents.default_subagent_model` | `publicModels[2]` |
322
+
323
+ **Codex never receives an internal name.** The slot aliases `main`, `review` and `subagent` stay inside the gateway. The CLI receives the official model names from `publicModels`, and `mapModel` resolves each one back to its slot. Set `codexRoles` in the profile when you want a different pairing than the order of that list.
324
+
325
+ The shim also passes `model_catalog_json`. That file is generated from `publicModels` on every profile change, and the `/model` picker reads it. The picker never calls `/v1/models`. `/v1/models` serves the same entries, and each window follows `model1M` for its slot.
326
+
327
+ It also passes `openai_base_url` for local routing. If 1M context is enabled for `main`, it passes `model_context_window=1000000` and `model_auto_compact_token_limit=900000`. Command-line overrides have higher precedence than user and project configuration. Re-run `switch shim install` after upgrading an older checkout.
328
+
329
+ ### Blindfold mode (optional)
330
+
331
+ A base URL override makes Codex print one line on its own `/model` screen:
332
+
333
+ ```
334
+ base URL is overridden to http://127.0.0.1:3456/v1. Selecting models may not be supported or work properly.
335
+ ```
336
+
337
+ Blindfold mode removes that line. Codex keeps its official endpoint, and the switcher intercepts the network hop instead. It needs no administrator rights, no certificate in a system trust store, and no change to `~/.codex/config.toml`.
338
+
339
+ ```bash
340
+ bash blindfold/make-certs.sh chatgpt.com # once
341
+ # then set "blindfold": true in the Codex profile
342
+ switch codex <profile> # the gateway starts the interceptor
343
+ ```
344
+
345
+ The gateway owns the interceptor: it starts it at boot and after every change, and `switch off` stops it. If the certificates are missing, or another process holds the gateway or interceptor port, `switch` refuses the activation and writes no file.
346
+
347
+ Read [📖 `docs/codex-blindfold.md`](docs/codex-blindfold.md) before you turn it on. The guide explains the interception scope, the risk of holding a private CA, and how to go back. It opens with three diagrams:
348
+
349
+ - [Request routing](docs/diagrams/blindfold-request-routing.html) — one request, from CONNECT to the provider
350
+ - [Model name resolution](docs/diagrams/codex-model-name-resolution.html) — which name the CLI sees, and where it resolves
351
+ - [Lifecycle under switch](docs/diagrams/blindfold-switch-lifecycle.html) — activation, refusal, and shutdown
352
+
353
+ ---
354
+
355
+ ## Co-existence with Token Optimizers (RTK, Headroom, Ponytail)
356
+
357
+ If you use prompt-trimming tools like **Headroom**, **Ponytail**, or **RTK (Rust Token Killer)**:
358
+ 1. Configure your CLI (Claude Code / Codex) to point to the optimizer proxy (e.g. `http://127.0.0.1:8787`).
359
+ 2. Configure that optimizer's upstream endpoint to point to **LLM Switcher** (`http://127.0.0.1:3456`).
360
+ 3. **LLM Switcher** serves as the protective final gateway before the internet:
361
+ - **Repairs Broken Schemas:** Fixes orphaned `tool_result` turns and consecutive same-role turns caused by aggressive history pruning.
362
+ - **Restores Stripped Thinking:** Detects reasoning models and restores thinking parameters if an optimizer stripped them to save tokens.
363
+ - **Enforces 1M Context Windows:** Injects local 1M context flags and auto-compact thresholds.
364
+ - **Converts Protocols:** Bridges 2-way traffic to your target upstream (9Router, OpenRouter, Vertex, etc.).
365
+
366
+ ### Interoperability Test Report & Benchmark
367
+
368
+ | Failure Scenario Caused by Optimizers | Direct Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
369
+ |---|---|---|
370
+ | **Orphaned `tool_result` turn** (Headroom prunes tool_use turn) | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block |
371
+ | **Consecutive `user` turns** (Optimizer drops assistant turns) | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns seamlessly |
372
+ | **Stripped `thinking` parameters** (Optimizer removes reasoning) | ⚠️ **Degraded AI**: Reasoning disabled, shallow single-liners | ✅ **HTTP 200 OK**: Automatically restores thinking budget |
373
+ | **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool into user context |
374
+ | **Custom optimizer headers** (`x-rtk-*`, `traceparent`) | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Clean transparent header passthrough |
375
+
376
+ Run the automated verification suites:
377
+ ```bash
378
+ # Offline (mock upstream, no API key needed): protocol conversion, healer, streaming, security
379
+ npm test
380
+
381
+ # Live (requires a running gateway and a real upstream; consumes tokens)
382
+ node tests/live-optimizer-interop.mjs
383
+ ```
384
+
385
+ See the full research report in [📖 `docs/TOKEN-OPTIMIZER-INTEROP.md`](docs/TOKEN-OPTIMIZER-INTEROP.md).
386
+
387
+ ---
388
+
389
+ ## Agent Skill & MCP Server Integration
390
+
391
+ To guarantee that AI coding agents (Claude Code, Cursor, Windsurf, Opencode) and spawned sub-processes **never bypass LLM Switcher**, this repository provides two control-plane assets:
392
+
393
+ ### 1. The Agent Skill (`skills/llm-switcher/SKILL.md`)
394
+ A standardized Agent Skill teaching the LLM:
395
+ - **Mandatory routing:** All LLM traffic and token compression tools (Headroom, RTK, Ponytail) MUST target `http://127.0.0.1:3456`.
396
+ - **Zero-mutation policy:** Prohibits the agent from editing `~/.claude/settings.json` directly.
397
+ - **Sub-process safety:** Automatically sources `env.cmd` or `env.sh` when launching sub-agents.
398
+
399
+ Install globally for Opencode / Claude:
400
+ ```bash
401
+ # For Opencode:
402
+ cp -r skills/llm-switcher ~/.config/opencode/skills/
403
+
404
+ # For Claude Code:
405
+ cp -r skills/llm-switcher ~/.claude/skills/
406
+ ```
407
+
408
+ ### 2. The MCP Server (`mcp.mjs`)
409
+ A zero-dependency Model Context Protocol (MCP) server communicating over `stdio`:
410
+ - `switcher_status`: Read live active profiles and 1M flags.
411
+ - `switcher_audit`: Audit the environment to detect rogue direct outbound calls or unrouted token compressors.
412
+ - `switcher_switch_profile`: Programmatically switch a CLI's active profile.
413
+ - `switcher_recent_logs`: Inspect recent request logs, token usage, and thinking extraction.
414
+
415
+ NOTE: `switcher_recent_logs` returns the first 150 characters of each recent prompt, from every client that used the gateway. The agent that calls the tool can read them.
416
+
417
+ Add to your MCP configuration (e.g. `opencode.jsonc`, `claude_desktop_config.json`, or Cursor):
418
+ ```json
419
+ "mcp": {
420
+ "llm-switcher": {
421
+ "type": "local",
422
+ "command": ["node", "path/to/llm-switcher/mcp.mjs"],
423
+ "enabled": true
424
+ }
425
+ }
426
+ ```
427
+
428
+ ---
429
+
430
+ ## CLI Reference
431
+
432
+ ```bash
433
+ switch ui # Open the Web UI dashboard in your browser
434
+ switch status # Display status for all active CLI targets
435
+ switch doctor # Audit environment, settings & routing
436
+ switch on [profile] # Start the gateway and activate a profile for all compatible targets
437
+ switch <profile> # Activate profile for all compatible targets
438
+ switch claude <profile> # Set active profile specifically for Claude Code
439
+ switch codex <profile> # Set active profile specifically for Codex
440
+ switch openai <profile> # Set active profile specifically for OpenAI Chat
441
+ switch vertex <profile> # Set active profile specifically for Vertex / Gemini
442
+ switch port <number> # Change the gateway port (restarts it if running)
443
+ switch service install # Install OS background autostart service (Windows / macOS / Linux)
444
+ switch service uninstall # Remove background autostart service
445
+ switch shim install # Route new Claude and Codex sessions through the gateway
446
+ switch shim status # Verify shims + detect running sessions that bypass the gateway
447
+ switch shim uninstall # Remove the launcher shims
448
+ switch off [target] # Deactivate gateway (or specific target) and restore official
449
+ switch contract-probe [--model m] # Drive the contract-lab variants through the gateway
450
+ switch contract-check # Turn the open contract findings into failing tests
451
+ ```
452
+
453
+ The service runs without your shell. `switch service install` therefore copies `CLAUDE_CONFIG_DIR`, `LLM_SWITCHER_CONFIG`, `LLM_SWITCHER_STATE_DIR` and `LLM_SWITCHER_BLINDFOLD_CERTS` into the systemd unit or the launchd plist when they are set. The Windows task cannot carry them; set them as User environment variables instead. If the installed definition differs from the new one, for example after a hand edit, the old file is kept as `<file>.bak`. On Windows the task is created from an XML definition, so paths with spaces need no extra quoting and the task has no run-time limit. This Windows path is not tested on Windows yet.
454
+
455
+ ### Resumed sessions & the shim (important)
456
+
457
+ `switch on` writes `env.sh` / `env.cmd` and deliberately **removes** proxy variables from
458
+ `~/.claude/settings.json` — that keeps Claude Code from showing its "custom API" banner.
459
+ The side effect: a CLI started from a shell that never sourced `env.sh` has **no**
460
+ `ANTHROPIC_BASE_URL`, so it talks to the provider directly and skips the gateway
461
+ (no Healer, no 1M unlock, no pooled quota). `claude --resume` in a fresh terminal is the
462
+ classic case.
463
+
464
+ The shim closes that hole. It installs tiny wrappers in `~/.llm-switcher/bin` that source
465
+ `env.sh` and then `exec` the real binary:
466
+
467
+ ```bash
468
+ switch shim install
469
+ export PATH="$HOME/.llm-switcher/bin:$PATH" # add to ~/.zshrc or ~/.bashrc
470
+ switch shim status # verify
471
+ ```
472
+
473
+ Behaviour:
474
+
475
+ - **Gateway ON** → the wrapper injects the env, so every invocation (including `--resume`)
476
+ is routed through the gateway.
477
+ - For Codex, the wrapper also passes the documented model-role and context settings with
478
+ `--config`. It does not depend on unsupported `CODEX_*` variables.
479
+ - **Gateway OFF** (no `active.flag`) → the wrapper is fully transparent and runs the real
480
+ binary untouched; it never forces routing.
481
+ - The real binary is located with the shim directory stripped from `PATH`, so it can never
482
+ call itself recursively. If no real binary is found it exits `127` with a clear message
483
+ instead of failing silently.
484
+ - `settings.json` is left alone, so **no warning banner** appears.
485
+
486
+ `switch on` installs the shims automatically and warns when `PATH` still needs the export
487
+ line. `switch doctor` and `switch shim status` additionally scan running `claude`/`codex`
488
+ processes and flag any that lack `ANTHROPIC_BASE_URL` — those sessions must be quit and
489
+ re-opened from a shell where the shim is on `PATH`.
490
+
491
+ ---
492
+
493
+ ## Configuration Schema (`config.json`)
494
+
495
+ ```jsonc
496
+ {
497
+ "port": 3456,
498
+ "activeProfile": "9router",
499
+ "activeProfiles": {
500
+ "anthropic": "9router", // Active profile for Claude Code (/v1/messages)
501
+ "responses": "codex-profile", // Active profile for Codex (/v1/responses)
502
+ "openai-chat": "9router", // Active profile for OpenAI Chat
503
+ "vertex": "gemini-profile" // Active profile for Vertex / Gemini
504
+ },
505
+ "profiles": {
506
+ "9router": {
507
+ "name": "9Router Cloud",
508
+ "mode": "convert", // hybrid | convert | direct
509
+ "inFormat": "auto", // auto | anthropic | openai-chat | responses | vertex
510
+ "outFormat": "openai-chat", // openai-chat | anthropic | vertex
511
+ "thinkingMode": "auto", // auto | native | off (see Advanced Options)
512
+ "baseURL": "https://api.9router.com/v1",
513
+ "apiKey": "sk-...",
514
+ "defaultModels": {
515
+ "opus": "ag/claude-opus-4-6-thinking",
516
+ "sonnet": "ag/gemini-3.7-flash",
517
+ "haiku": "ag/gemini-3.6-flash-medium",
518
+ "fable": "ag/gemini-3.8-flash"
519
+ },
520
+ "model1M": {
521
+ "opus": true,
522
+ "sonnet": true,
523
+ "haiku": false,
524
+ "fable": true
525
+ }
526
+ }
527
+ },
528
+ "debug": false,
529
+ "contractLab": {
530
+ "url": "https://intact.example.com",
531
+ "apiKey": "sk-...",
532
+ "enabled": false
533
+ }
534
+ }
535
+ ```
536
+
537
+ ### Contract lab
538
+
539
+ The contract lab finds fields that the converter loses. It is off by default.
540
+
541
+ - Set `contractLab: {url, apiKey, enabled}` in `config.json`. If `enabled` is `true`, the gateway sends a sample of complete exchanges to intact.
542
+ - `switch contract-probe [--model m]` sends six test requests per model and format through the gateway.
543
+ - `switch contract-check` gets the open findings from intact and writes one test file for each lost field.
544
+
545
+ ---
546
+
547
+ ## Advanced Options
548
+
549
+ | Option | Description |
550
+ |---|---|
551
+ | `LLM_SWITCHER_CONFIG=/path/config.json` | Use a config file outside the repo (the proxy, `switch` and `mcp.mjs` all honour it). |
552
+ | `--port <n>` / `LLM_SWITCHER_PORT` | Override the listening port (priority: flag > env > `config.port`). |
553
+ | `x-llm-profile: <key>` header (alias `x-profile`) or `?profile=<key>` | Route a single request through a specific profile. An unknown key returns HTTP 400 instead of silently falling back. |
554
+ | `profile.thinkingMode` | `auto` (default, for gateways like 9Router): restore stripped thinking, inject a `<think>` guide for non-reasoning models, send `thinking` + `reasoning_effort`. `native` (strict OpenAI APIs): send only `reasoning_effort` when the client asks, never touch the prompt, use `max_completion_tokens`. `off`: never send reasoning parameters. |
555
+ | `profile.endpoints.countTokens` | Override the Anthropic `count_tokens` URL. |
556
+ | `profile.endpoints` | Override upstream URLs per format: `{ "openai-chat": "...", "anthropic": "...", "vertex": "https://.../models/{model}:{action}" }`. |
557
+ | `CLAUDE_CONFIG_DIR` | Respected when locating Claude Code's `settings.json`. |
558
+ | `LLM_SWITCHER_STATE_DIR` | Move the launch files and the logs out of the checkout. The tests use it; the shims read the directory that was set when they were installed. |
559
+
560
+ ## Security Model
561
+
562
+ - The gateway binds to `127.0.0.1` only and rejects requests whose `Host` is not a loopback name (DNS-rebinding protection) or whose `Origin` is not the dashboard itself (CSRF protection).
563
+ - The admin API (`/api/*`) requires the `x-llm-switcher-token` header. The gateway creates the token in `admin.token`, next to `config.json`, with mode 0600. `switch ui` opens the dashboard with this token, and the MCP server reads the file. The dashboard keeps the token in `sessionStorage` of that tab only, so a page that another account serves on the same port while the gateway is off cannot read it; open the dashboard with `switch ui` again after the browser restarts. `/v1/*` and `/health` need no token.
564
+ - API keys are never sent to the browser: `/api/status` returns redacted profiles and the dashboard keeps the stored key unless you type a new one. The stored key goes only to the stored `baseURL` and `endpoints` of that profile. A save that changes either one must carry the key again.
565
+ - Every dashboard change carries the config revision that the page loaded. If another tab, the CLI or the MCP server saved in the meantime, the gateway answers 409 and the page reloads instead of overwriting that change.
566
+ - Client credentials such as `x-api-key`, `authorization` or `x-goog-api-key` are **not** forwarded to upstreams. Other `x-*` headers, `traceparent` and `tracestate` pass through. The gateway drops its own control headers (`x-profile`, `x-llm-profile`) and the network identity headers (`x-forwarded-*`, `x-real-ip`).
567
+ - `config.json` is written atomically with mode 0600. `~/.claude/settings.json` is rewritten only to remove values that the switcher wrote itself. `ANTHROPIC_AUTH_TOKEN`, `*_MODEL_NAME` and your own model or URL values stay, and `switch` prints the name of every value it removes.
568
+
569
+ ## Compatibility Notes
570
+
571
+ - **Direct Anthropic passthrough:** valid requests are forwarded byte-for-byte (thinking signatures, `cache_control`, documents stay intact). Malformed ones are repaired in place by a native Anthropic healer: orphaned `tool_result` → text, missing `tool_result` → placeholder, results moved to the start of the user turn.
572
+ - **Thinking signatures:** thinking blocks produced by conversion carry a gateway signature (`reasoning-sig`, or a foreign signature prefixed with `lsw1.`). They are stripped before a request reaches Anthropic. If that leaves an in-progress tool loop without the thinking block Anthropic requires, thinking is disabled for that single request instead of failing.
573
+ - **Codex tools:** `custom`/freeform tools (e.g. `apply_patch` with its Lark grammar), `namespace` tools and `local_shell` are exposed to upstreams as function tools and converted back to `custom_tool_call` / namespaced `function_call` / `local_shell_call` items. Hosted tools (`web_search`, `file_search`, `tool_search`, image generation) execute on OpenAI's servers, so other upstreams cannot provide them and they are omitted.
574
+ - **Gemini 3 thought signatures:** signatures returned with function calls are cached in memory by tool call id (last 5,000 calls) and replayed on the matching `functionCall` part, including through Gemini's OpenAI-compatible `extra_content`. After a gateway restart, unknown calls in the current turn get Google's documented `skip_thought_signature_validator` value, which Google notes may reduce quality.
575
+ - **`/v1/messages/count_tokens`:** exact when the active Claude Code profile uses a native Anthropic upstream; otherwise an estimate (other providers have no equivalent endpoint).
576
+
577
+ ## Research & Protocol Matrices
578
+
579
+ Detailed response research and live-tested format matrices are documented in:
580
+ - 📖 [`docs/LLM-RESPONSE-MATRIX.md`](docs/LLM-RESPONSE-MATRIX.md) — 48 live sample variations across 8 model families.
581
+ - 📊 [`docs/response-matrix.json`](docs/response-matrix.json) — Machine-readable response signature schema.
582
+
583
+ ---
584
+
585
+ ## License
586
+
587
+ MIT © 2026 LLM Switcher Contributors.