llm-switcher 1.1.11 → 1.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. package/CHANGELOG.md +63 -0
  2. package/README.md +202 -257
  3. package/README.vi.md +200 -256
  4. package/blindfold/blindfold.mjs +200 -53
  5. package/blindfold/make-certs.sh +26 -7
  6. package/catalog.mjs +246 -0
  7. package/classifier.mjs +238 -0
  8. package/config.example.json +12 -34
  9. package/docs/codex-blindfold.md +28 -17
  10. package/docs/cross-platform.md +16 -7
  11. package/docs/diagrams/ir-healer-pipeline.mmd +16 -0
  12. package/docs/diagrams/ir-healer-pipeline.png +0 -0
  13. package/docs/diagrams/ir-healer-pipeline.svg +90 -0
  14. package/docs/diagrams/ir-translation-pipeline.html +14925 -0
  15. package/docs/diagrams/ir-translation-pipeline.sequence.json +31 -0
  16. package/docs/diagrams/ir-translation-pipeline.svg +5128 -0
  17. package/docs/diagrams/system-architecture.architecture.json +76 -0
  18. package/docs/diagrams/system-architecture.html +14978 -0
  19. package/docs/diagrams/system-architecture.svg +5147 -0
  20. package/docs/diagrams/system-topology.mmd +30 -0
  21. package/docs/diagrams/system-topology.png +0 -0
  22. package/docs/diagrams/system-topology.svg +125 -0
  23. package/ensure-ca-bundle.mjs +28 -0
  24. package/formats.mjs +43 -156
  25. package/icons/antigravity.png +0 -0
  26. package/icons/claude.png +0 -0
  27. package/icons/codex.png +0 -0
  28. package/icons/deepseek.png +0 -0
  29. package/icons/gemini.png +0 -0
  30. package/icons/github.png +0 -0
  31. package/icons/groq.png +0 -0
  32. package/icons/intact.svg +1 -0
  33. package/icons/ollama.png +0 -0
  34. package/icons/openai.png +0 -0
  35. package/icons/openrouter.png +0 -0
  36. package/icons/qwen.png +0 -0
  37. package/icons/vertex.png +0 -0
  38. package/mcp.mjs +39 -11
  39. package/package.json +1 -1
  40. package/proxy.mjs +114 -37
  41. package/shim.mjs +200 -57
  42. package/skills/llm-switcher/SKILL.md +15 -10
  43. package/state.mjs +1100 -191
  44. package/switch.cmd +2 -2
  45. package/switch.mjs +228 -53
  46. package/tests/blindfold-e2e.test.mjs +380 -0
  47. package/tests/blindfold-task5.test.mjs +429 -0
  48. package/tests/blindfold.task3.test.mjs +700 -0
  49. package/tests/blindfold.test.mjs +10 -5
  50. package/tests/catalog.test.mjs +147 -0
  51. package/tests/classifier.test.mjs +210 -0
  52. package/tests/contract-lab.test.mjs +22 -7
  53. package/tests/formats.test.mjs +63 -46
  54. package/tests/gateway.e2e.test.mjs +136 -36
  55. package/tests/lifecycle.test.mjs +16 -10
  56. package/tests/mcp.test.mjs +78 -2
  57. package/tests/real-user-sim.test.mjs +464 -0
  58. package/tests/shim.test.mjs +159 -66
  59. package/tests/state.test.mjs +975 -193
  60. package/tests/switch.test.mjs +446 -2
  61. package/ui.html +1710 -1726
package/README.md CHANGED
@@ -3,7 +3,7 @@
3
3
  <p align="center">
4
4
  <b>Zero-dependency, multi-protocol edge gateway & provider switcher</b><br>
5
5
  Seamlessly bridge <b>Claude Code</b>, <b>Codex</b>, OpenAI, and Gemini SDKs to any upstream LLM API.<br>
6
- Full bi-directional protocol conversion, 1M context unlock, thinking protocol extraction, and edge message healing.
6
+ Full bi-directional protocol conversion, official-model context windows, thinking protocol extraction, and edge message healing.
7
7
  </p>
8
8
 
9
9
  <p align="center">
@@ -13,162 +13,118 @@
13
13
  <p align="center">
14
14
  <img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?logo=node.js&logoColor=white" alt="Node.js 18+">
15
15
  <img src="https://img.shields.io/badge/Dependencies-Zero-38bdf8" alt="Zero Dependencies">
16
- <img src="https://img.shields.io/badge/Context-1%2C000%2C000_tokens-6366f1" alt="1M Context">
16
+ <img src="https://img.shields.io/badge/Context-window_follows_the_model-6366f1" alt="Context window follows the model">
17
17
  <img src="https://img.shields.io/badge/Multi--Active-Concurrent_CLIs-f59e0b" alt="Multi-Active">
18
18
  <img src="https://img.shields.io/badge/License-MIT-gray" alt="License MIT">
19
19
  </p>
20
20
 
21
21
  ---
22
22
 
23
- > ### 🎯 The Core Problem: Why Generic Proxies Cripple Your AI Coding Tools
23
+ > ### 🛡️ Zero-Loss Native Emulation for Coding Agents
24
24
  >
25
- > Every LLM provider uses a **subtly or drastically different API response standard**:
26
- > - **Anthropic** requires dedicated `thinking` blocks (`thinking_delta` + `signature_delta`), strict alternating turn rules, and typed `tool_use` input schemas.
27
- > - **OpenAI** streams reasoning as `reasoning_content` delta chunks or `reasoning_details[]`, and formats tools as `tool_calls` with JSON string arguments.
28
- > - **Google Vertex AI** places reasoning in `candidates[0].content.parts[{thought: true, text, thoughtSignature}]` and tool arguments as raw objects.
29
- > - **Open-source models (DeepSeek, Qwen, GLM)** often dump chain-of-thought directly into `content` or duplicate fields under conflicting keys.
25
+ > Generic proxies mangle response protocols: Anthropic loses `thinking_delta` reasoning blocks, tool arguments split, and prompt caches desynchronize.
30
26
  >
31
- > **When coding tools like Claude Code or Codex receive non-native or partially converted responses, they don't just look wrong — the agent's performance degrades catastrophically:**
32
- > 1. **Lost Chain-of-Thought:** If Claude Code does not receive native `thinking_delta` blocks, it **completely misses the model's internal reasoning**. The agent acts prematurely, skips architectural planning, and produces buggy code.
33
- > 2. **Broken Tool Execution:** Mismatched stop reasons (`tool_calls` vs `tool_use`) and split argument chunks cause tool execution failures and infinite retries.
34
- > 3. **Token & Cache Miscounting:** Non-standard usage accounting breaks prompt cache alignment and premature context compaction.
35
- >
36
- > Developers often blame the model for "getting dumber" when in reality **their proxy mangled the response protocol.**
37
- >
38
- > ### 🛡️ The Solution: Zero-Loss Native Emulation (Subscription-Grade Quality)
39
- >
40
- > **LLM Switcher solves this by acting as a high-precision, zero-loss protocol emulator.**
41
- >
42
- > It normalizes whatever your upstream provider emits (9Router, OpenRouter, Vertex, DeepSeek) and re-synthesizes it into the **exact native event stream the client agent was built to consume**:
43
- > - **Claude Code** receives 100% genuine Anthropic SSE events (`message_start` ➔ `thinking_delta` ➔ `signature_delta` ➔ `content_block_start: tool_use` ➔ `message_delta`), performing **identically to an official Anthropic subscription**.
44
- > - **Codex** receives 100% genuine Responses API events (`response.created` ➔ `output_text.delta` ➔ `function_call` ➔ `response.completed`).
45
- >
46
- > **You get the freedom and cost savings of 3rd-party APIs while maintaining 100% official subscription-grade agent intelligence.**
27
+ > **LLM Switcher solves this at the local network edge:**
28
+ > - **100% Native Emulation:** Normalizes upstream APIs (intact, 9Router, Vertex, DeepSeek) into genuine Anthropic SSE (`thinking_delta` + `tool_use`) for Claude Code, and genuine Responses API events for Codex.
29
+ > - **Client-Side Edge Companion:** Intentionally offloads heavy account pooling and key rotation to **[intact](https://github.com/louisphamdev/intact)** (recommended) or 9Router (basic alternative), keeping LLM Switcher zero-dependency and bloat-free.
30
+ > - **Targeted Tool Scope:** Built specifically for **Claude Code** and **OpenAI Codex** (OpenCode natively supports custom models without shims; refer to intact for account pooling; and Antigravity isn't worth building for 😏).
47
31
 
48
32
  ---
49
33
 
50
- > ### 💡 Design Philosophy: The Client-Side Edge Companion to 9Router
51
- >
52
- > **LLM Switcher intentionally does NOT implement multi-account pooling, key rotation, quota tracking, or provider load balancing.**
53
- >
54
- > That heavy lifting belongs to server-side AI routing gateways like **[9Router](https://github.com/decolua/9router)**, which handle centralized account rotation, rate-limit retries, and quota management far more reliably and securely at the server layer.
55
- >
56
- > **LLM Switcher is specifically engineered as the optimal client-side edge extension to pair with 9Router (or similar gateways):**
57
- > - **At the Local Workstation (LLM Switcher):** Translates coding tool protocols (Claude Code `/v1/messages`, Codex `/v1/responses`, Vertex `/v1beta/...`, OpenAI Chat), injects 1M context windows, manages local multi-CLI profiles, and runs the Healer Engine to fix mangled payloads from local prompt compressors (RTK, Headroom, Ponytail).
58
- > - **At the Server Gateway (9Router):** Manages account pools, API key rotation, load balancing, billing quotas, and global provider failover.
59
- >
60
- > This clean division of responsibility keeps LLM Switcher **ultra-lightweight, zero-dependency, and bloat-free** while giving you an unbeatable developer setup.
34
+ ## Architecture & Interactive Diagrams
35
+
36
+ LLM Switcher runs locally on your workstation (`127.0.0.1:3456`) as a transparent edge interceptor and protocol bridge.
37
+
38
+ <p align="center">
39
+ <a href="docs/diagrams/system-architecture.html">
40
+ <img src="docs/diagrams/system-topology.svg" alt="LLM Switcher System Topology & Architecture" width="100%">
41
+ </a>
42
+ <br>
43
+ <sub><i>🎨 Themed with Pretty-Mermaid (Tokyo Night). Click diagram to open interactive Archify viewer (zoom, pan, tracing).</i></sub>
44
+ </p>
45
+
46
+ ### 1. Interactive Archify Visual Library
47
+
48
+ All architecture maps and execution sequences are authored with **[Archify](https://github.com/tt-a1i/archify)** and rendered with **[Pretty-Mermaid](https://github.com/imxv/Pretty-mermaid-skills)**:
49
+
50
+ | Diagram | Description | Interactive Visual | Scalable Vector |
51
+ |---|---|---|---|
52
+ | **System Topology** | Complete edge architecture: Clients ➔ Optimizers ➔ Gateway & Healer Core ➔ Upstream Providers | [📊 Open Interactive View](docs/diagrams/system-architecture.html) | [SVG](docs/diagrams/system-topology.svg) • [PNG](docs/diagrams/system-topology.png) |
53
+ | **IR Healer Pipeline** | Inbound request normalization, schema healing, streaming synthesis, and abort propagation | [🔄 Open Interactive View](docs/diagrams/ir-translation-pipeline.html) | [SVG](docs/diagrams/ir-healer-pipeline.svg) • [PNG](docs/diagrams/ir-healer-pipeline.png) |
54
+ | **Codex Blindfold Routing** | TLS CONNECT proxy sequence, credential scrubbing, and upstream routing | [🛡️ Open Interactive View](docs/diagrams/blindfold-request-routing.html) | [HTML](docs/diagrams/blindfold-request-routing.html) |
55
+ | **Switch Lifecycle** | Zero-downtime tool toggle, CAS configuration writes, and interceptor sync | [⚡ Open Interactive View](docs/diagrams/blindfold-switch-lifecycle.html) | [HTML](docs/diagrams/blindfold-switch-lifecycle.html) |
61
56
 
62
57
  ---
63
58
 
64
- ## Architecture & Workflow
59
+ ### 2. Request Lifecycle & Healer Pipeline
65
60
 
66
- LLM Switcher sits locally on your workstation (`127.0.0.1:3456`). It acts as the **outermost edge gatekeeper** before requests leave for the internet.
61
+ <p align="center">
62
+ <a href="docs/diagrams/ir-translation-pipeline.html">
63
+ <img src="docs/diagrams/ir-healer-pipeline.svg" alt="Bi-Directional IR Healer Pipeline" width="100%">
64
+ </a>
65
+ <br>
66
+ <sub><i>💡 Click above to inspect the interactive IR Healer lifecycle sequence.</i></sub>
67
+ </p>
67
68
 
68
- ### 1. End-to-End System Topology
69
+ ### 3. High-Level Flow
69
70
 
70
71
  ```mermaid
71
- flowchart TD
72
- subgraph Clients["Dev Clients & Coding CLIs"]
73
- CC["Claude Code CLI\n(/v1/messages)"]
74
- CDX["OpenAI Codex CLI\n(/v1/responses)"]
75
- OAI["OpenAI SDKs / Cursor\n(/v1/chat/completions)"]
76
- VTX["Gemini / Vertex SDKs\n(/v1beta/models/*)"]
72
+ flowchart LR
73
+ classDef client fill:#1e293b,stroke:#38bdf8,stroke-width:2px,color:#f8fafc;
74
+ classDef edge fill:#0f172a,stroke:#6366f1,stroke-width:2px,color:#f8fafc;
75
+ classDef healer fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#f8fafc;
76
+ classDef upstream fill:#2e1065,stroke:#a855f7,stroke-width:2px,color:#f8fafc;
77
+ classDef opt fill:#1e1b4b,stroke:#818cf8,stroke-dasharray: 4 4,color:#e0e7ff;
78
+
79
+ subgraph Clients[" 💻 Dev Clients & Coding CLIs "]
80
+ CC["Claude Code CLI\n(/v1/messages)"]:::client
81
+ CDX["OpenAI Codex CLI\n(/v1/responses)"]:::client
77
82
  end
78
83
 
79
- subgraph Optimizers["Optional Middle-Layer (Installed in CLI)"]
80
- OPT["Prompt Optimizers & Trimmers\n(Headroom / RTK / Ponytail)\n[Configured upstream: :3456]"]
84
+ subgraph Middle[" ⚡ Optional Middle-Layer "]
85
+ OPT["Token Optimizers\n(Headroom / RTK)"]:::opt
81
86
  end
82
87
 
83
- subgraph Switcher["LLM Switcher (:3456) — Outermost Edge Gatekeeper"]
84
- direction TB
85
- ROUTER["Protocol Auto-Detection & Multi-Active Routing"]
86
- HEALER["Healer Engine\n• Heal orphaned tool_results\n• Restore stripped thinking\n• Merge consecutive turns"]
87
- IR["Bi-Directional IR Translation\n(4 Client Formats ⟷ 3 Upstream Formats)"]
88
- M1M["1M Context Unlocker\n& Auto-Compact Thresholds"]
89
- LOGS["Live Inspector\n(In-Memory Ring Buffer)"]
90
- ROUTER --> HEALER --> IR --> M1M --> LOGS
88
+ subgraph Gateway[" 🛡️ LLM Switcher Edge Gateway (:3456) "]
89
+ ROUTER["Edge Router\n(Zero-Mutation)"]:::edge
90
+ HEALER["Healer Engine\n(Auto-Fix Schemas)"]:::healer
91
+ IR["Bi-Directional IR\n(Event Synth)"]:::healer
92
+ ROUTER --> HEALER --> IR
91
93
  end
92
94
 
93
- subgraph Upstream["Internet / Upstream Providers"]
94
- R9["9Router / Selfhost Gateway"]
95
- OR["OpenRouter / Together / Groq"]
96
- ANT["Anthropic Native API"]
97
- GCP["Google Vertex AI / Gemini"]
95
+ subgraph Upstreams[" ☁️ Upstream Providers "]
96
+ INTACT["intact Gateway\n(Recommended Pooler)"]:::upstream
97
+ OTHER["9Router / Vertex / Other"]:::upstream
98
98
  end
99
99
 
100
- CC -->|Direct| ROUTER
101
- CC -.->|Optional| OPT
102
- CDX -->|Direct| ROUTER
103
- CDX -.->|Optional| OPT
104
- OAI --> ROUTER
105
- VTX --> ROUTER
106
- OPT -->|Forward to Switcher| ROUTER
107
-
108
- LOGS -->|Clean Outbound| R9
109
- LOGS -->|Clean Outbound| OR
110
- LOGS -->|Clean Outbound| ANT
111
- LOGS -->|Clean Outbound| GCP
112
- ```
113
-
114
- ---
115
-
116
- ### 2. Bi-Directional IR (Intermediate Representation) Pipeline
100
+ CC -->|direct| ROUTER
101
+ CDX -->|direct| ROUTER
102
+ CC -.->|prune| OPT
103
+ CDX -.->|prune| OPT
104
+ OPT -->|forward| ROUTER
117
105
 
118
- ```mermaid
119
- sequenceDiagram
120
- autonumber
121
- actor CLI as Client (Claude Code / Codex / SDK)
122
- participant GW as LLM Switcher (:3456)
123
- participant IR as IR & Healer Engine
124
- participant UP as Upstream (9Router / Anthropic / Vertex)
125
-
126
- CLI->>GW: Inbound Request (Anthropic, Responses, Chat, or Vertex)
127
- Note over GW,IR: Normalize to Canonical IR
128
- GW->>IR: parseToIR(clientFormat, payload)
129
- Note over IR: Healer checks:<br/>1. Repair orphaned tool_results<br/>2. Restore stripped thinking params<br/>3. Reconcile role alternations<br/>4. Apply 1M context limits
130
- IR->>GW: emitUpstreamBody(outFormat, healedIR)
131
- GW->>UP: Outbound API Call (fetch with AbortSignal)
132
- UP-->>GW: Upstream Streaming SSE / JSON Chunks
133
- Note over GW: normalizeUpstream(chunk)<br/>Extract reasoning_content, <think> tags, usage
134
- GW->>CLI: Render client-native SSE (e.g. Anthropic thinking_delta + text_delta)
135
- Note over CLI,GW: Client connection closes (Ctrl+C) -> GW aborts UP instantly!
106
+ IR -->|contract & pool| INTACT
107
+ IR -->|standard call| OTHER
136
108
  ```
137
109
 
138
110
  ---
139
111
 
140
- ### 3. Multi-Active CLI Independent Routing
112
+ ### 4. How It Actually Works: Transparent Request Interception
141
113
 
142
- You can run **multiple active profiles concurrently** — one profile dedicated to each CLI, without collision:
114
+ LLM Switcher acts as a transparent man-in-the-middle without ever touching client configuration files:
143
115
 
144
- ```mermaid
145
- flowchart LR
146
- subgraph Inbound["Incoming Client Calls"]
147
- C1["Claude Code\n(/v1/messages)"]
148
- C2["Codex CLI\n(/v1/responses)"]
149
- C3["OpenAI SDK\n(/v1/chat/completions)"]
150
- C4["Vertex SDK\n(/v1beta/models/*)"]
151
- end
116
+ 1. **Ephemeral Shim Activation:** When you invoke `claude` or `codex`, a lightweight shim at the front of your `PATH` executes first. It injects `HTTPS_PROXY=http://127.0.0.1:3457` and custom CA certs *only into that process's in-memory environment*, leaving `~/.claude/settings.json` and `~/.codex/config.toml` completely untouched.
117
+ 2. **Network Interception (`:3457`):** The tool sends standard TLS requests to official hosts (`api.anthropic.com` or `api.openai.com`). The local Blindfold interceptor terminates TLS, strips client credentials, and transparently re-routes API calls (`/v1/messages`, `/v1/responses`, `/v1/models`) locally to the Gateway (`:3456`). All other traffic (OAuth logins, GitHub, web searches) tunnels untouched to the real internet.
118
+ 3. **Protocol Conversion & Upstream Call (`:3456`):** The gateway loads your active profile from `config.json`, runs the Healer Engine (repairing empty `{}` schemas and restoring reasoning budgets), and converts the request into the upstream provider's native dialect (e.g. Gemini, OpenAI Chat, or Anthropic) with your configured credentials and base URL.
119
+ 4. **Native Event Stream Synthesis:** When the upstream provider streams its response, the gateway captures reasoning tokens and tool calls, re-synthesizing them into genuine Anthropic SSE (`thinking_delta` + `tool_use`) or Codex Responses events. The tool receives genuine native events and believes it is communicating directly with the official provider!
152
120
 
153
- subgraph Core["LLM Switcher Core (:3456)"]
154
- SLOT1["Slot: Anthropic\nActive: [9Router]"]
155
- SLOT2["Slot: Responses\nActive: [OpenRouter]"]
156
- SLOT3["Slot: OpenAI\nActive: [Local LLM]"]
157
- SLOT4["Slot: Vertex\nActive: [Off / Official]"]
158
- end
159
-
160
- subgraph Egress["Upstream Targets"]
161
- U1["9Router (Opus 1M context)"]
162
- U2["OpenRouter (Sonnet thinking)"]
163
- U3["Local OpenAI Server (:8000)"]
164
- U4["Official Google Endpoint"]
165
- end
166
-
167
- C1 --> SLOT1 --> U1
168
- C2 --> SLOT2 --> U2
169
- C3 --> SLOT3 --> U3
170
- C4 --> SLOT4 --> U4
171
- ```
121
+ > #### 🔒 CA Security & Origin: Where does the CA come from and how safe is it?
122
+ >
123
+ > - **100% Locally Minted:** The CA certificate (`ca.pem`) and private key (`ca.key`) are generated entirely on your own machine using your local OpenSSL (`blindfold/make-certs.sh`). No keys are downloaded from the internet, and the private key is stored locally with strict `0600` permissions.
124
+ > - **Zero OS Trust Store Tampering:** Unlike tools like Charles or Fiddler, LLM Switcher **NEVER installs anything into your system or OS root certificate store** (no Windows Certificate Store, no macOS Keychain, no Linux `/etc/ssl/certs`). It requires **zero Administrator or sudo privileges**.
125
+ > - **Process-Scoped Trust Only:** The certificate is loaded ephemerally into the memory of `claude` (via `NODE_EXTRA_CA_CERTS`) and `codex` (via `CODEX_CA_CERTIFICATE`). Your browsers, banking apps, git, and other terminal sessions never trust this CA.
126
+ > - **Cryptographic Name Constraints:** The CA is minted with explicit X.509 `nameConstraints` strictly permitting only three domains: `api.anthropic.com`, `api.openai.com`, and `chatgpt.com`. Even if the local private key were compromised, standard TLS verifiers will reject it for any other domain (Google, GitHub, your bank).
127
+ > - **VPN & Corporate CA Preservation:** If your workstation already has a company CA in `NODE_EXTRA_CA_CERTS`, `ensure-ca-bundle.mjs` merges both into a combined bundle so internal corporate proxies never break.
172
128
 
173
129
  ---
174
130
 
@@ -184,68 +140,21 @@ flowchart LR
184
140
  - Fixes orphaned `tool_result` blocks caused by aggressive prompt pruners (RTK, Headroom, Ponytail) before sending to Anthropic/OpenAI upstream.
185
141
  - Automatically restores thinking parameters if an intermediary tool stripped them.
186
142
  - Merges consecutive same-role turns to enforce strict alternating turn requirements.
187
- - **1M Context Window Unlocker:** Follows the profile's `model1M` map per tier: every tier marked 1M gets `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]` (so `/model sonnet`, tier switches and subagents keep 1M, while unmarked tiers stay at 200K), plus auto-compact at `900000`, with built-in visual risk warnings for unsupported models.
188
- - **Zero Config Mutation:** Never writes endpoints or keys into `~/.claude/settings.json` (it removes only values it wrote itself: `ANTHROPIC_BASE_URL` for its own port and `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]`). Uses launcher flags and environment injection to prevent annoying provider warning banners.
143
+ - **Context windows come from the model:** the switcher no longer forces a 1M window or an auto-compact limit, and it writes no model name into your environment. Claude Code sizes its own session from the window of the official model you pick, and a backend with a smaller window than that model can overflow in a long session. `model1M` now only decides what `/v1/models` reports.
144
+ - **Zero Config Mutation:** Never reads or writes `~/.claude/settings.json` or `~/.codex/config.toml`, and writes no environment variable and no `--config` argument that a coding tool reads as configuration. The tool reaches the gateway only through the interceptor, so no provider warning banner appears.
189
145
  - **Live Request / Response Inspector:** Built-in dashboard tab displaying real-time requests, latency, token consumption, prompt previews, and thinking blocks.
190
146
  - **Native Background Service:** Install and run as an OS background daemon on Windows (Task Scheduler), macOS (launchd), or Linux (systemd).
191
147
 
192
148
  ---
193
149
 
194
- ## Changes in 1.1.10
195
-
196
- - **README.** A new section, "Self-improvement with intact", explains how this gateway and intact correct their own faults. intact is now public and on npm as `intact-gateway`.
197
-
198
- ### Changes in 1.1.9
150
+ ## Recent Highlights (v1.2.0)
199
151
 
200
- - **Dashboard.** Open `http://127.0.0.1:3456/ui` directly. The page no longer needs the link from `switch ui`: the gateway puts the admin token in the page. Pages of other sites still cannot read it.
201
-
202
- ### Changes in 1.1.8
203
-
204
- - **Claude Code tools.** A tool schema value of `0`, `false`, `""` or `null` (for example `minimum: 0`) became an empty object schema. Gemini refused every Claude Code request with HTTP 400 "Starting an object on a scalar field". These values now stay as they are.
205
- - **Shim PATH order.** If the shim folder is on `PATH` but after the real `claude` or `codex`, `switch shim status` and `switch doctor` now tell you to put the export line last in your shell files.
206
-
207
- ### Changes in 1.1.7
208
-
209
- - **Codex tools.** With `publicModels` set, Codex lost its tools and ended after one answer. The model catalog copied the metadata of a real OpenAI model, which puts Codex in the "Responses Lite" form. The catalog now keeps Codex in its direct tool mode, and the gateway also reads tools that arrive as an `additional_tools` input item.
210
-
211
- ### Changes in 1.1.6
212
-
213
- - **Codex over WebSocket.** The gateway keeps the turns of each WebSocket session. A turn that sends `previous_response_id` gets the earlier turns back, so Codex no longer loses the task after the first tool call. An unknown id fails the turn with `previous_response_not_found`.
214
- - **Codex warmup.** A `response.create` frame with `generate: false` gets a local answer. It no longer spends a model call.
215
- - **Codex model name.** A profile without `publicModels` no longer sends `OpenAI-Model: main` in the handshake, and `switch codex` and `switch doctor` warn about it. Codex read `main` as a reroute and showed a false "high-risk cyber activity" warning. See "Codex-first setup".
216
- - **Contract lab.** The gateway uploads each sample with the key that opened its trace. intact refused the 1.1.5 uploads with `HTTP 404 trace not found`.
217
- - **Version stamp.** The stamp uses the last commit only when the checkout has no changes. Otherwise it uses the newest file time.
218
-
219
- ### Changes in 1.1.5
220
-
221
- - **Contract lab privacy.** Masking now uses an allowlist. Every string value is masked except the enum values that intact reads. 1.1.4 masked a list of content keys and missed 16 fields (citations, document and web titles, web search queries, logprobs tokens, file names and URIs, stop sequences, participant names, error messages, tool descriptions).
222
-
223
- ### Changes in 1.1.4
224
-
225
- - **Contract lab privacy.** A sample that the gateway sends to intact carries no client content. Prompts, answers, tool arguments and results, files and user ids are masked on this machine before the upload. The fields that intact needs for analysis stay.
226
-
227
- ### Changes in 1.1.3
228
-
229
- - **Dashboard.** Opened without its access token, the page no longer waits on "Checking status...". It says that it is locked and names `switch ui`, which opens it with the token.
230
- - **Dashboard.** The Codex model names (session, review, subagent) are on the **Models** tab, next to the other model settings. The **Blindfold** tab holds only the interceptor settings.
231
-
232
- ### Changes in 1.1.2
233
-
234
- - **npm package.** Install with `npm install -g llm-switcher` and run `switch`. An npm install keeps its data in `~/.llm-switcher`, so an upgrade does not erase your configuration. A git checkout keeps its data next to the code, as before.
235
- - **Contract lab.** The gateway can send a small sample of complete exchanges to an [intact](https://github.com/louisphamdev/intact) server, which finds fields that the converter loses. It is off by default. See "Contract lab" below.
236
- - **macOS.** `blindfold/make-certs.sh` now runs with LibreSSL, the default `openssl` on macOS.
237
- - **Upgrade from 1.1.0 or older.** A running gateway older than 1.1.1 cannot prove its identity. `switch` now names it and does not stop it. Stop it by hand once, then run `switch on`.
238
- - **Tests.** `npm test` runs only `tests/**/*.test.mjs`, also on Node.js 18 and 20.
239
-
240
- ### Earlier changes
241
-
242
- - The desktop dashboard now uses a compact developer-tool layout. It has clearer route controls, keyboard-accessible tabs, labeled model fields, and no decorative emoji.
243
- - Codex profiles now use three documented roles: `main`, `review`, and `subagent`.
244
- - The Codex shim passes official configuration overrides for `model`, `review_model`, `agents.default_subagent_model`, `model_context_window`, and `model_auto_compact_token_limit`.
245
- - Legacy profile keys remain readable. An explicit empty role now clears its legacy fallback.
246
- - Claude Opus model IDs use version-independent reasoning detection.
247
-
248
- The Codex shim no longer relies on `CODEX_MODEL`, `CODEX_MAX_CONTEXT_TOKENS`, or `CODEX_AUTO_COMPACT_WINDOW`. Codex does not document these environment variables. See the official [configuration reference](https://developers.openai.com/codex/config-reference/) and [advanced configuration guide](https://developers.openai.com/codex/config-advanced/).
152
+ - **Zero-Mutation Interceptor:** All traffic routes via `HTTPS_PROXY` without editing client configs (`~/.claude/settings.json`, `~/.codex/config.toml`).
153
+ - **Concurrent Multi-Tool Support:** Simultaneously configures `{ claude, codex }` profiles with dynamic, zero-downtime switching (`POST /_control/active-tools`).
154
+ - **Auto-Discovery & Dynamic Model Catalog:** Discovers official models for Claude Code & Codex; detects tool version updates and refreshes mappings on the fly (`switch models`).
155
+ - **Self-Healing Schemas:** Auto-repairs `{}` empty schemas for Gemini/Vertex and restores stripped `thinking` tokens.
156
+ - **Windows Reliability:** Strict CRLF `.cmd` shims, CA path resolution, and zero-drift subroutine routing.
157
+ - *For older releases (v1.1.2 – v1.1.10), see [CHANGELOG.md](CHANGELOG.md).*
249
158
 
250
159
  ---
251
160
 
@@ -302,20 +211,28 @@ Open the Web Dashboard at: **[http://127.0.0.1:3456/ui](http://127.0.0.1:3456/ui
302
211
 
303
212
  ## CLI Integration
304
213
 
305
- ### Universal Environment Loader (`env.cmd` / `env.sh`)
214
+ ### The environment comes from the shims, not from your shell
215
+
216
+ Do **not** add a `source env.sh` or `call env.cmd` line to `~/.bashrc`, `~/.zshrc` or a wrapper
217
+ script. Those two files are neutral stubs now: a comment line and nothing else, so an old rc line
218
+ keeps running and can never re-introduce a base URL.
306
219
 
307
- Every time you switch profiles, LLM Switcher writes ready-to-use environment loaders into the data folder:
220
+ Each tool gets its own file instead, and only the matching shim loads it:
308
221
 
309
- - **Windows (Command Prompt / PowerShell wrapper):**
310
- ```cmd
311
- call "%USERPROFILE%\.llm-switcher\env.cmd"
312
- ```
313
- - **macOS / Linux (Bash / Zsh):**
314
- ```bash
315
- source ~/.llm-switcher/env.sh
316
- ```
222
+ | File | Loaded by |
223
+ | --- | --- |
224
+ | `env-claude.sh` / `env-claude.cmd` | the `claude` shim |
225
+ | `env-codex.sh` / `env-codex.cmd` | the `codex` shim |
226
+ | `env.sh` / `env.cmd` | nobody. Neutral stub, kept only so an old rc line stays silent |
317
227
 
318
- For a checkout, use the same files in the checkout folder.
228
+ An empty per-tool file means that tool is off: the shim then leaves the environment alone and the
229
+ tool reaches its official endpoint. Before it loads anything, the shim also scrubs a stale
230
+ `ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL` or `ANTHROPIC_DEFAULT_<TIER>_MODEL` inherited from an older
231
+ release or from your own shell, so `switch off` really is off.
232
+
233
+ In practice you never call these files. `switch shim install` puts `~/.llm-switcher/bin` on `PATH`,
234
+ and every `claude` and `codex` invocation — including `claude --resume` in a brand-new terminal —
235
+ runs the shim, which injects the variables into that one process.
319
236
 
320
237
  ---
321
238
 
@@ -327,21 +244,11 @@ For a checkout, use the same files in the checkout folder.
327
244
  node "path\to\llm-switcher\switch.mjs" %*
328
245
  ```
329
246
 
330
- 2. Patch your global Claude Code launcher (`claude.cmd` in your npm global directory):
247
+ 2. Install the shim and put its directory **before** the real Claude Code directory in your User `PATH` (System Properties → Environment Variables), then open a new terminal:
331
248
  ```cmd
332
- SETLOCAL EnableDelayedExpansion
333
- IF EXIST "%USERPROFILE%\.llm-switcher\active.flag" (
334
- SET "ANTHROPIC_BASE_URL=http://127.0.0.1:3456"
335
- SET "CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1"
336
- )
337
- IF EXIST "%USERPROFILE%\.llm-switcher\1m.flag" (
338
- SET /P M1M=<"%USERPROFILE%\.llm-switcher\1m.flag"
339
- IF "!M1M!"=="" SET "M1M=opus[1m]"
340
- SET "ANTHROPIC_MODEL=!M1M!"
341
- SET "CLAUDE_CODE_AUTO_COMPACT_WINDOW=900000"
342
- )
249
+ switch shim install
343
250
  ```
344
- > `SETLOCAL EnableDelayedExpansion` is required for `!M1M!`. npm rewrites `claude.cmd` on every update, so prefer a separate wrapper that runs `call "%USERPROFILE%\.llm-switcher\env.cmd"` and then `claude %*`. For a checkout, replace `%USERPROFILE%\.llm-switcher` with the checkout folder. Only `env.cmd` / `env.sh` carry the per-tier `ANTHROPIC_DEFAULT_<TIER>_MODEL=<tier>[1m]` variables.
251
+ > Do **not** patch `claude.cmd` in your npm global directory: npm rewrites it on every update, and the shim is what injects the gateway URL. `settings.json` is never touched, so no "custom API" banner appears.
345
252
 
346
253
  ---
347
254
 
@@ -359,21 +266,7 @@ codex
359
266
 
360
267
  On Windows, add `%USERPROFILE%\.llm-switcher\bin` before the real Codex directory in `PATH`. Open a new terminal after the change.
361
268
 
362
- The shim does not edit `~/.codex/config.toml`. When the gateway is active, it passes these official command-line overrides to the real Codex binary:
363
-
364
- | Profile role | Codex configuration key | Name the CLI receives |
365
- |---|---|---|
366
- | `main` | `model` | `publicModels[0]` |
367
- | `review` | `review_model` | `publicModels[1]` |
368
- | `subagent` | `agents.default_subagent_model` | `publicModels[2]` |
369
-
370
- **A profile that serves Codex must have `publicModels`.** Without it, the gateway has no official name for Codex. It then writes no model catalog and sends no `OpenAI-Model` header, and Codex shows two false warnings: "Model metadata for `<model>` not found" and "Your account was flagged for potentially high-risk cyber activity". `switch codex` and `switch doctor` warn when the Codex profile has no `publicModels`.
371
-
372
- **Codex never receives an internal name.** The slot aliases `main`, `review` and `subagent` stay inside the gateway. The CLI receives the official model names from `publicModels`, and `mapModel` resolves each one back to its slot. Set `codexRoles` in the profile when you want a different pairing than the order of that list.
373
-
374
- The shim also passes `model_catalog_json`. That file is generated from `publicModels` on every profile change, and the `/model` picker reads it. The picker never calls `/v1/models`. `/v1/models` serves the same entries, and each window follows `model1M` for its slot.
375
-
376
- It also passes `openai_base_url` for local routing. If 1M context is enabled for `main`, it passes `model_context_window=1000000` and `model_auto_compact_token_limit=900000`. Command-line overrides have higher precedence than user and project configuration. Re-run `switch shim install` after upgrading an older checkout.
269
+ The shim never edits `~/.codex/config.toml`. Codex reaches the gateway cleanly through `HTTPS_PROXY` and the interceptor, retaining its own model names and context windows. Model catalog entries are served dynamically at `/v1/models`.
377
270
 
378
271
  ### Blindfold mode (optional)
379
272
 
@@ -386,11 +279,25 @@ base URL is overridden to http://127.0.0.1:3456/v1. Selecting models may not be
386
279
  Blindfold mode removes that line. Codex keeps its official endpoint, and the switcher intercepts the network hop instead. It needs no administrator rights, no certificate in a system trust store, and no change to `~/.codex/config.toml`.
387
280
 
388
281
  ```bash
389
- bash "$(npm root -g)/llm-switcher/blindfold/make-certs.sh" chatgpt.com # once; a checkout runs blindfold/make-certs.sh
390
- # then set "blindfold": true in the Codex profile
282
+ bash "$(npm root -g)/llm-switcher/blindfold/make-certs.sh" # once; a checkout runs blindfold/make-certs.sh
283
+ # the port is already top-level in config.json: "blindfold": { "port": 3457 }
391
284
  switch codex <profile> # the gateway starts the interceptor
392
285
  ```
393
286
 
287
+ One interceptor serves both tools. It routes by the host of the CONNECT request and by the path,
288
+ and by nothing else:
289
+
290
+ | CONNECT host | Paths to this switcher's gateway | Everything else |
291
+ | --- | --- | --- |
292
+ | `api.anthropic.com` | `/v1/messages`, `/v1/messages/...` | to `api.anthropic.com`, unchanged |
293
+ | `api.openai.com` | `/v1/responses`, `/v1/responses/...`, `/v1/models`, `/v1/models/...` | to `api.openai.com`, unchanged |
294
+ | `chatgpt.com` | `/backend-api/codex/...`, forwarded as `/v1` | to `chatgpt.com`, unchanged |
295
+
296
+ There is no `blindfoldHost` and no `blindfoldPrefix` any more: that table is the routing, and it is
297
+ not a profile setting. A CONNECT host outside the table is tunneled untouched; a request whose
298
+ `Host` header names a different host than its CONNECT target gets `421` and opens no upstream
299
+ connection. See the [full table and the certificate rules](docs/cross-platform.md).
300
+
394
301
  The gateway owns the interceptor: it starts it at boot and after every change, and `switch off` stops it. If the certificates are missing, or another process holds the gateway or interceptor port, `switch` refuses the activation and writes no file.
395
302
 
396
303
  Read [📖 `docs/codex-blindfold.md`](docs/codex-blindfold.md) before you turn it on. The guide explains the interception scope, the risk of holding a private CA, and how to go back. It opens with three diagrams:
@@ -399,6 +306,19 @@ Read [📖 `docs/codex-blindfold.md`](docs/codex-blindfold.md) before you turn i
399
306
  - [Model name resolution](docs/diagrams/codex-model-name-resolution.html) — which name the CLI sees, and where it resolves
400
307
  - [Lifecycle under switch](docs/diagrams/blindfold-switch-lifecycle.html) — activation, refusal, and shutdown
401
308
 
309
+ ### What you give up
310
+
311
+ - **1M context follows the official model's window.** The switcher no longer forces a 1M window
312
+ or an auto-compact limit, and it writes no model name into your environment. Claude Code sizes
313
+ its own session from the window of the model you pick. A backend whose window is smaller than
314
+ that model can overflow in a long session. `model1M` now only decides what `/v1/models` reports.
315
+ - **Codex needs certificates once.** Blindfold mode is what keeps Codex on its official endpoint,
316
+ and it needs a private CA plus a leaf naming the three hosts above. Skip it and Codex shows the
317
+ `base URL is overridden` line on its `/model` screen instead.
318
+ - **The interceptor decrypts the three hosts it serves.** It refuses a CONNECT to a local or
319
+ private address, but every process on the machine can reach it. `docs/codex-blindfold.md` opens
320
+ with the scope, the risk of holding a private CA, and how to go back.
321
+
402
322
  ---
403
323
 
404
324
  ## Co-existence with Token Optimizers (RTK, Headroom, Ponytail)
@@ -409,8 +329,8 @@ If you use prompt-trimming tools like **Headroom**, **Ponytail**, or **RTK (Rust
409
329
  3. **LLM Switcher** serves as the protective final gateway before the internet:
410
330
  - **Repairs Broken Schemas:** Fixes orphaned `tool_result` turns and consecutive same-role turns caused by aggressive history pruning.
411
331
  - **Restores Stripped Thinking:** Detects reasoning models and restores thinking parameters if an optimizer stripped them to save tokens.
412
- - **Enforces 1M Context Windows:** Injects local 1M context flags and auto-compact thresholds.
413
- - **Converts Protocols:** Bridges 2-way traffic to your target upstream (9Router, OpenRouter, Vertex, etc.).
332
+ - **Context windows follow the model:** reports the official window of each model; nothing is injected.
333
+ - **Converts Protocols:** Bridges 2-way traffic to your target upstream (intact, 9Router, OpenRouter, Vertex, etc.).
414
334
 
415
335
  ### Interoperability Test Report & Benchmark
416
336
 
@@ -442,8 +362,8 @@ To guarantee that AI coding agents (Claude Code, Cursor, Windsurf, Opencode) and
442
362
  ### 1. The Agent Skill (`skills/llm-switcher/SKILL.md`)
443
363
  A standardized Agent Skill teaching the LLM:
444
364
  - **Mandatory routing:** All LLM traffic and token compression tools (Headroom, RTK, Ponytail) MUST target `http://127.0.0.1:3456`.
445
- - **Zero-mutation policy:** Prohibits the agent from editing `~/.claude/settings.json` directly.
446
- - **Sub-process safety:** Automatically sources `env.cmd` or `env.sh` when launching sub-agents.
365
+ - **Zero-mutation policy:** Prohibits the agent from editing `~/.claude/settings.json` — the switcher never reads or writes that file either.
366
+ - **Sub-process safety:** Never tells a sub-agent to `source env.sh`; the shims in `~/.llm-switcher/bin` inject the proxy variables into the tool process itself and scrub anything stale first.
447
367
 
448
368
  Install globally for Opencode / Claude:
449
369
  ```bash
@@ -456,7 +376,7 @@ cp -r skills/llm-switcher ~/.claude/skills/
456
376
 
457
377
  ### 2. The MCP Server (`mcp.mjs`)
458
378
  A zero-dependency Model Context Protocol (MCP) server communicating over `stdio`:
459
- - `switcher_status`: Read live active profiles and 1M flags.
379
+ - `switcher_status`: Read live active profiles and gateway state.
460
380
  - `switcher_audit`: Audit the environment to detect rogue direct outbound calls or unrouted token compressors.
461
381
  - `switcher_switch_profile`: Programmatically switch a CLI's active profile.
462
382
  - `switcher_recent_logs`: Inspect recent request logs, token usage, and thinking extraction.
@@ -482,36 +402,42 @@ Add the server to your MCP configuration (for example `opencode.jsonc`, `claude_
482
402
  switch ui # Open the Web UI dashboard in your browser
483
403
  switch status # Display status for all active CLI targets
484
404
  switch doctor # Audit environment, settings & routing
485
- switch on [profile] # Start the gateway and activate a profile for all compatible targets
486
- switch <profile> # Activate profile for all compatible targets
487
- switch claude <profile> # Set active profile specifically for Claude Code
488
- switch codex <profile> # Set active profile specifically for Codex
489
- switch openai <profile> # Set active profile specifically for OpenAI Chat
490
- switch vertex <profile> # Set active profile specifically for Vertex / Gemini
405
+ switch on [profile] # Start the gateway and activate a profile
406
+ switch <profile> # Activate a profile for both tools
407
+ switch claude <profile> # Set the active profile for Claude Code only
408
+ switch codex <profile> # Set the active profile for Codex only
491
409
  switch port <number> # Change the gateway port (restarts it if running)
492
410
  switch service install # Install OS background autostart service (Windows / macOS / Linux)
493
411
  switch service uninstall # Remove background autostart service
494
412
  switch shim install # Route new Claude and Codex sessions through the gateway
495
413
  switch shim status # Verify shims + detect running sessions that bypass the gateway
496
414
  switch shim uninstall # Remove the launcher shims
497
- switch off [target] # Deactivate gateway (or specific target) and restore official
415
+ switch off # Stop everything and restore the official endpoints
416
+ switch off claude # Turn Claude Code off; Codex keeps running
417
+ switch off codex # Turn Codex off; Claude Code keeps running
498
418
  switch contract-probe [--model m] # Drive the contract-lab variants through the gateway
499
419
  switch contract-check # Turn the open contract findings into failing tests
500
420
  ```
501
421
 
422
+ Targets are `claude` and `codex`, and they are the only two. A profile is one tool: `tool` is
423
+ either `"claude"` or `"codex"` (or `null` for a profile that is switched off), so `switch claude` and
424
+ `switch codex` can never point at the same profile by accident. There is no `openai` or `vertex`
425
+ target any more — the input routes those names stood for are gone.
426
+
502
427
  The service runs without your shell. `switch service install` therefore copies `CLAUDE_CONFIG_DIR`, `LLM_SWITCHER_CONFIG`, `LLM_SWITCHER_STATE_DIR` and `LLM_SWITCHER_BLINDFOLD_CERTS` into the systemd unit or the launchd plist when they are set. The Windows task cannot carry them; set them as User environment variables instead. If the installed definition differs from the new one, for example after a hand edit, the old file is kept as `<file>.bak`. On Windows the task is created from an XML definition, so paths with spaces need no extra quoting and the task has no run-time limit. This Windows path is not tested on Windows yet.
503
428
 
504
429
  ### Resumed sessions & the shim (important)
505
430
 
506
- `switch on` writes `env.sh` / `env.cmd` and deliberately **removes** proxy variables from
507
- `~/.claude/settings.json` — that keeps Claude Code from showing its "custom API" banner.
508
- The side effect: a CLI started from a shell that never sourced `env.sh` has **no**
431
+ `switch on` writes `env-claude.*` and `env-codex.*` and writes nothing into
432
+ `~/.claude/settings.json` — the switcher never reads or writes that file, which is what keeps
433
+ Claude Code from showing its "custom API" banner.
434
+ The side effect: a CLI started without the shim has **no**
509
435
  `ANTHROPIC_BASE_URL`, so it talks to the provider directly and skips the gateway
510
- (no Healer, no 1M unlock, no pooled quota). `claude --resume` in a fresh terminal is the
436
+ (no Healer, no pooled quota). `claude --resume` in a fresh terminal is the
511
437
  classic case.
512
438
 
513
- The shim closes that hole. It installs tiny wrappers in `~/.llm-switcher/bin` that source
514
- `env.sh` and then `exec` the real binary:
439
+ The shim closes that hole. It installs tiny wrappers in `~/.llm-switcher/bin` that inject the
440
+ tool's own environment and then `exec` the real binary:
515
441
 
516
442
  ```bash
517
443
  switch shim install
@@ -521,12 +447,15 @@ switch shim status # verify
521
447
 
522
448
  Behaviour:
523
449
 
524
- - **Gateway ON** → the wrapper injects the env, so every invocation (including `--resume`)
525
- is routed through the gateway.
526
- - For Codex, the wrapper also passes the documented model-role and context settings with
527
- `--config`. It does not depend on unsupported `CODEX_*` variables.
528
- - **Gateway OFF** (no `active.flag`) → the wrapper is fully transparent and runs the real
450
+ - **Tool active** → the shim loads that tool's own env file, so every invocation (including
451
+ `--resume`) is routed through the gateway.
452
+ - **Tool off** (its env file is empty) → the shim is fully transparent and runs the real
529
453
  binary untouched; it never forces routing.
454
+ - Before it loads anything, the shim scrubs a stale `ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL` or
455
+ `ANTHROPIC_DEFAULT_<TIER>_MODEL` inherited from an older release or from your shell, so a
456
+ variable cannot survive a `switch off`.
457
+ - No `--config` arguments are passed to Codex. The shim changes nothing on Codex's side but the
458
+ environment.
530
459
  - The real binary is located with the shim directory stripped from `PATH`, so it can never
531
460
  call itself recursively. If no real binary is found it exits `127` with a clear message
532
461
  instead of failing silently.
@@ -544,21 +473,19 @@ re-opened from a shell where the shim is on `PATH`.
544
473
  ```jsonc
545
474
  {
546
475
  "port": 3456,
547
- "activeProfile": "9router",
548
476
  "activeProfiles": {
549
- "anthropic": "9router", // Active profile for Claude Code (/v1/messages)
550
- "responses": "codex-profile", // Active profile for Codex (/v1/responses)
551
- "openai-chat": "9router", // Active profile for OpenAI Chat
552
- "vertex": "gemini-profile" // Active profile for Vertex / Gemini
477
+ "claude": "claude-default", // Active profile for Claude Code (/v1/messages)
478
+ "codex": "codex-default" // Active profile for Codex (/v1/responses)
553
479
  },
480
+ "blindfold": { "port": 3457 }, // Interceptor port. Top-level, optional, default 3457
554
481
  "profiles": {
555
- "9router": {
556
- "name": "9Router Cloud",
482
+ "claude-default": {
483
+ "name": "Intact Gateway",
557
484
  "mode": "convert", // hybrid | convert | direct
558
- "inFormat": "auto", // auto | anthropic | openai-chat | responses | vertex
485
+ "tool": "claude", // claude | codex | null (profile is off)
559
486
  "outFormat": "openai-chat", // openai-chat | anthropic | vertex
560
487
  "thinkingMode": "auto", // auto | native | off (see Advanced Options)
561
- "baseURL": "https://api.9router.com/v1",
488
+ "baseURL": "https://intact.example.com/v1", // or https://api.9router.com/v1
562
489
  "apiKey": "sk-...",
563
490
  "defaultModels": {
564
491
  "opus": "ag/claude-opus-4-6-thinking",
@@ -571,12 +498,20 @@ re-opened from a shell where the shim is on `PATH`.
571
498
  "sonnet": true,
572
499
  "haiku": false,
573
500
  "fable": true
574
- },
575
- // Codex keys. A profile that serves Codex MUST have publicModels (see Codex-first setup).
501
+ }
502
+ },
503
+ "codex-default": {
504
+ "name": "Codex via router",
505
+ "mode": "convert",
506
+ "tool": "codex",
507
+ "outFormat": "vertex",
508
+ "baseURL": "https://YOUR-GATEWAY/v1",
509
+ "apiKey": "sk-...",
510
+ // A profile that serves Codex MUST have publicModels (see Codex-first setup).
576
511
  "publicModels": ["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna"], // official names for main, review, subagent
577
512
  "codexRoles": { "review": "gpt-5.6-sol" }, // optional: pair one role with another public name
578
- "blindfold": false, // true: Codex keeps its official endpoint (see Codex blindfold mode)
579
- "blindfoldPort": 3457 // interceptor port when blindfold is true
513
+ "defaultModels": { "main": "gemini-3.8-flash", "review": "gemini-3.7-flash-medium", "subagent": "gemini-3.6-flash-low" },
514
+ "model1M": { "main": false, "review": false, "subagent": false }
580
515
  }
581
516
  },
582
517
  "debug": false,
@@ -588,6 +523,16 @@ re-opened from a shell where the shim is on `PATH`.
588
523
  }
589
524
  ```
590
525
 
526
+ `tool` replaces `inFormat`: it says which tool a profile serves, not what its upstream speaks.
527
+ There is no top-level `activeProfile`, no `blindfold`, `blindfoldPort`, `blindfoldHost` or
528
+ `blindfoldPrefix` inside a profile, and no `openai-chat` or `vertex` key in `activeProfiles`.
529
+
530
+ An older `config.json` is rewritten once, on load, through a compare-and-swap that refuses to
531
+ touch a file somebody else changed first. When two profiles would collapse onto the same key the
532
+ migration stops, names the clashing keys on the CLI and in the dashboard, and leaves the file
533
+ exactly as it was: every mutating `switch` command then exits non-zero, `switch off` and
534
+ `switch doctor` still work, and fixing the file makes everything work again.
535
+
591
536
  ### Contract lab
592
537
 
593
538
  The contract lab finds fields that the converter loses. It is off by default.
@@ -617,7 +562,7 @@ Because of this split, a provider fingerprint is never a rule in this gateway. I
617
562
  | `LLM_SWITCHER_CONFIG=/path/config.json` | Use a config file outside the data folder (the proxy, `switch` and `mcp.mjs` all honour it). |
618
563
  | `--port <n>` / `LLM_SWITCHER_PORT` | Override the listening port (priority: flag > env > `config.port`). |
619
564
  | `x-llm-profile: <key>` header (alias `x-profile`) or `?profile=<key>` | Route a single request through a specific profile. An unknown key returns HTTP 400 instead of silently falling back. |
620
- | `profile.thinkingMode` | `auto` (default, for gateways like 9Router): restore stripped thinking, inject a `<think>` guide for non-reasoning models, send `thinking` + `reasoning_effort`. `native` (strict OpenAI APIs): send only `reasoning_effort` when the client asks, never touch the prompt, use `max_completion_tokens`. `off`: never send reasoning parameters. |
565
+ | `profile.thinkingMode` | `auto` (default, for gateways like intact or 9Router): restore stripped thinking, inject a `<think>` guide for non-reasoning models, send `thinking` + `reasoning_effort`. `native` (strict OpenAI APIs): send only `reasoning_effort` when the client asks, never touch the prompt, use `max_completion_tokens`. `off`: never send reasoning parameters. |
621
566
  | `profile.endpoints.countTokens` | Override the Anthropic `count_tokens` URL. |
622
567
  | `profile.endpoints` | Override upstream URLs per format: `{ "openai-chat": "...", "anthropic": "...", "vertex": "https://.../models/{model}:{action}" }`. |
623
568
  | `CLAUDE_CONFIG_DIR` | Respected when locating Claude Code's `settings.json`. |