llm-switcher 1.2.8 → 1.2.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,9 @@
1
1
  # Changelog — LLM Switcher
2
2
 
3
+ ## Release 1.2.9
4
+
5
+ - **Bifrost:** Claude Code that reaches a Claude Code account on intact now goes through unchanged. Only the key changes. Before, a `convert` profile, or a `hybrid` profile with a `claude/...` model, changed the request to OpenAI Chat. The `direct` route also ran the healer and `thinkingMode`. Now the gateway asks intact for `bifrost_ua` of the mapped model. If the `User-Agent` of the client starts with this value, the gateway sends every client header, the body bytes, and the query string. There is no setting. intact 0.1.10 or newer gives `bifrost_ua`.
6
+
3
7
  ## Release 1.2.8
4
8
 
5
9
  - **Launch notice:** A person opened `claude` or `codex`, and nothing on screen said that the switcher took the traffic. The shim now raises a notice at launch. On Windows it is a desktop toast, because both tools can claim the whole screen and hide a printed line. On Linux and macOS it is one line on stderr. The notice names the profile, the host, the model, and the 1M window.
@@ -42,7 +46,7 @@ set through the page's own JavaScript. Each item below has a browser test that r
42
46
  - **A second instance had no port for the interceptor:** `--port` and `LLM_SWITCHER_PORT` move the gateway through `resolvePort`, but `computeLaunchState` read the interceptor port from `config.json` alone. `LLM_SWITCHER_BLINDFOLD_PORT` reached a hand-started `blindfold.mjs` and nothing else. A gateway sent to 3457 also landed on the port of the interceptor. The interceptor port now follows the same precedence as the gateway port, and it moves off the gateway port when the two meet. The README names both variables together.
43
47
  - **A fish user got a line that fish cannot run:** The PATH hint printed `export PATH="..."`, which is not fish syntax, and it named `~/.profile`, which fish does not read. The hint now names `~/.config/fish/config.fish` and gives `fish_add_path -m`. zsh and bash keep what they had.
44
48
  - **`switch on` installs the shims:** The README gave `switch shim install` as a separate step. `switch on` already installs the shims and prints what it installed. The README now says so, and it names the one case that still needs the command: a set `LLM_SWITCHER_STATE_DIR`.
45
- - **What the package says it serves:** The description and both README taglines gave OpenAI and Gemini as clients. This gateway accepts two clients, Claude Code and Codex, and `toIR` accepts no other input format. OpenAI, Anthropic and Vertex are upstreams. All three texts now say that.
49
+ - **What the package says it serves:** The description, both README taglines and the dashboard subtitle gave OpenAI and Gemini as clients. This gateway accepts two clients, Claude Code and Codex, and `toIR` accepts no other input format. OpenAI, Anthropic and Vertex are upstreams. All four texts now say that.
46
50
  - **A test that failed on a loaded machine:** The lock test released the lock 400 ms after it started the child, then asserted that the child waited 100 ms or more. A spawn plus an import measured 163 to 541 ms on Node 18, so the child sometimes found the lock gone. The test now counts the 400 ms from the moment the child reports its start, the same way the test below it does.
47
51
 
48
52
  ## Release 1.2.7
package/README.md CHANGED
@@ -403,6 +403,19 @@ Add the server to your MCP configuration (for example `opencode.jsonc`, `claude_
403
403
 
404
404
  ---
405
405
 
406
+ ## Bifrost: Claude Code to a Claude Code account on intact
407
+
408
+ Bifrost sends a Claude Code request to intact without a change. Only the key changes. The healer, `thinkingMode`, and the format conversion do not run.
409
+
410
+ Bifrost has no setting. The gateway turns it on for each request when two conditions are true:
411
+
412
+ 1. intact gives `bifrost_ua` for the mapped model in `GET /v1/models/{model}`. intact gives it only for a model of a Claude Code account.
413
+ 2. The `User-Agent` of the client starts with the `bifrost_ua` value.
414
+
415
+ The gateway keeps the answer from intact for 10 minutes for each model. If intact does not answer, the gateway keeps the result for 30 seconds and uses the normal route.
416
+
417
+ The gateway sends every client header, the body bytes, and the query string. It removes the credentials of the client and the hop-by-hop headers, and it sets `x-api-key` to the profile key. If the profile maps the model, the gateway also changes the `model` field. A different client, or a model that is not a Claude Code account, uses the normal route of the profile.
418
+
406
419
  ## CLI Reference
407
420
 
408
421
  ```bash
package/README.vi.md CHANGED
@@ -400,6 +400,19 @@ Thêm server vào cấu hình MCP (ví dụ `opencode.jsonc`, `claude_desktop_co
400
400
 
401
401
  ---
402
402
 
403
+ ## Bifrost: Claude Code tới tài khoản Claude Code trên intact
404
+
405
+ Bifrost gửi request của Claude Code tới intact nguyên vẹn. Chỉ key thay đổi. Healer, `thinkingMode` và bước chuyển định dạng không chạy.
406
+
407
+ Bifrost không có cấu hình. Gateway tự bật Bifrost cho từng request khi đủ hai điều kiện:
408
+
409
+ 1. intact trả `bifrost_ua` cho model đã map trong `GET /v1/models/{model}`. intact chỉ trả trường này cho model của tài khoản Claude Code.
410
+ 2. `User-Agent` của client bắt đầu bằng giá trị `bifrost_ua`.
411
+
412
+ Gateway giữ câu trả lời của intact 10 phút cho mỗi model. Nếu intact không trả lời, gateway giữ kết quả 30 giây và dùng đường thường.
413
+
414
+ Gateway gửi mọi header của client, nguyên byte body và query string. Gateway bỏ credential của client và các header hop-by-hop, rồi đặt `x-api-key` bằng key của profile. Nếu profile map model, gateway đổi thêm trường `model`. Client khác, hoặc model không thuộc tài khoản Claude Code, đi đường thường của profile.
415
+
403
416
  ## Bảng Tra cứu Lệnh CLI (`switch`)
404
417
 
405
418
  ```bash
@@ -1,110 +1,110 @@
1
- # Token Optimizer Interoperability & Failure Mode Report
2
-
3
- **How LLM Switcher acts as the protective outermost edge gateway for aggressive prompt/token optimizers (Headroom, RTK, Ponytail).**
4
-
5
- ---
6
-
7
- ## Executive Summary
8
-
9
- Third-party prompt optimizers and token compressors — such as **Headroom**, **RTK (Rust Token Killer)**, and **Ponytail** — attempt to reduce LLM input tokens by aggressively pruning message history, truncating command stdout, or forcing extreme prompt brevity.
10
-
11
- While these tools can reduce raw token counts in simple scenarios, **they frequently break complex agentic coding workflows** by corrupting message graphs, orphaning tool calls, and stripping reasoning parameters. When these pruned payloads hit upstream APIs directly (such as Anthropic, OpenAI, or 9Router), the provider immediately throws fatal `HTTP 400 Bad Request` errors or severely degrades reasoning depth.
12
-
13
- **LLM Switcher solves this by acting as the outermost edge gatekeeper (`127.0.0.1:3456`).** It intercepts the pruned payload before it leaves your machine, runs its built-in **Healer Engine** to repair message graphs and restore reasoning parameters, unlocks 1M context windows, and safely converts the protocol to your upstream provider.
14
-
15
- ---
16
-
17
- ## Tool Breakdown: What They Do & How They Break Payloads
18
-
19
- ### 1. Headroom (`headroomlabs-ai/headroom`)
20
- - **Mechanism:** Runs as a local proxy on `:8787` (or wraps CLI agents). Compresses conversation history, RAG chunks, and tool outputs using SmartCrusher (JSON), CodeCompressor (AST), and Kompress-v2-base. Also attempts "effort routing" to dial down thinking budgets.
21
- - **Critical Failure Points:**
22
- - **Orphaned `tool_result` blocks:** When pruning historical turns, Headroom often discards the `assistant` turn containing a `tool_use`, while retaining the subsequent `user` turn containing the `tool_result`. Anthropic's API strictly validates tool use IDs and crashes with:
23
- ```
24
- HTTP 400 invalid_request_error: "tool_use_id 'xxx' does not correspond to any tool_use"
25
- ```
26
- - **Consecutive `user` turns:** Dropping intermediary assistant turns causes multiple user messages to sit adjacent to each other. Anthropic strictly throws:
27
- ```
28
- HTTP 400 invalid_request_error: "roles must alternate between 'user' and 'assistant'"
29
- ```
30
- - **Reasoning Suppression:** Its "effort routing" dials down `thinking.budget_tokens` on routine tool turns. On complex models (Claude Opus, Gemini Flash), this prevents the model from formulating multi-step reasoning before acting.
31
-
32
- ### 2. RTK (`rtk-ai/rtk` - Rust Token Killer)
33
- - **Mechanism:** A single Rust binary that hooks into shell tool execution (e.g. `PreToolUse` in Claude Code / Cursor) and rewrites CLI commands (`git`, `ls`, `cat`, `grep`, `pytest`) to filter out noise, truncate lines, and inject recall tokens (`[full output: rtk recall xxx]`).
34
- - **Critical Failure Points:**
35
- - **Corrupted Structural Data:** When an agent invokes a tool expecting machine-readable JSON or exact AST formatting, RTK's heuristic summaries can alter structural delimiters, causing downstream tool call parsing errors.
36
- - **Custom Tracking Headers:** RTK and associated tracing proxies inject headers (`x-rtk-*`, `traceparent`, `x-optimizer-id`) that some strict upstream endpoints reject if not cleanly forwarded.
37
-
38
- ### 3. Ponytail (`DietrichGebert/ponytail`)
39
- - **Mechanism:** A behavioral prompt engineering plugin/ruleset that injects extreme conciseness instructions ("write one line, it works, YAGNI") into agent system prompts across 20+ coding tools.
40
- - **Critical Failure Points:**
41
- - **Premature Execution without Reasoning:** By commanding the model to be maximally brief and avoid planning, reasoning models are discouraged from spending thinking tokens. The model outputs untested single-liners that often fail type checks and test suites.
42
- - **System Prompt Prefix Invalidation:** Injected rules alter the leading system prompt bytes, busting provider prompt caches unless carefully aligned.
43
-
44
- ---
45
-
46
- ## The Healer Engine: How LLM Switcher Protects the Workflow
47
-
48
- LLM Switcher sits between the optimizer tool and the upstream LLM:
49
-
50
- ```
51
- [CLI Agent] ──> [Optimizer: Headroom / RTK] ──> [LLM Switcher :3456] ──> [Upstream / 9Router]
52
- │
53
- ├── 1. Heal Orphaned tool_results
54
- ├── 2. Merge Consecutive Turns
55
- ├── 3. Restore Stripped Thinking
56
- └── 4. Enforce 1M Context
57
- ```
58
-
59
- ### Protection Matrix: Before vs. After
60
-
61
- | Scenario | Direct to Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
62
- |---|---|---|
63
- | **Orphaned `tool_result` turn** | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block `[Tool Result (id)]: ...` |
64
- | **Consecutive `user` turns** | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns into a single valid turn seamlessly |
65
- | **Stripped `thinking` parameters** | ⚠️ **Degraded AI**: Reasoning disabled, model outputs shallow single-liners | ✅ **HTTP 200 OK**: Detects reasoning models and automatically restores safe thinking budget |
66
- | **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool message into user context |
67
- | **Custom tracking headers** | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Cleanly passes through `traceparent`, `x-request-id`, `x-rtk-*` |
68
-
69
- ---
70
-
71
- ## Test Methodology & Verification Suite
72
-
73
- We created an automated verification suite in [`tests/live-optimizer-interop.mjs`](../tests/live-optimizer-interop.mjs) that systematically replicates the failure modes of each tool:
74
-
75
- ### Running the Test Suite
76
- ```bash
77
- node tests/live-optimizer-interop.mjs
78
- ```
79
-
80
- ### Live Test Results
81
- ```text
82
- ================================================================
83
- LLM SWITCHER — TOKEN OPTIMIZER INTEROPERABILITY TEST SUITE
84
- Simulating failure modes from Headroom, RTK, and Ponytail
85
- ================================================================
86
-
87
- [TEST] Headroom Simulation: Orphaned tool_result turn... PASS (8840ms)
88
- ↳ Healed orphaned tool_result. Response HTTP 200: "It looks like you've shared a fragment of context ..."
89
- [TEST] Headroom Simulation: Consecutive User turns (Role alternation violation)... PASS (2683ms)
90
- ↳ Merged consecutive turns seamlessly. Stop reason: end_turn
91
- [TEST] Thinking Guard: Restoring stripped thinking parameter on reasoning models... PASS (5372ms)
92
- ↳ Automatically restored thinking: thinking_delta=true, signature_delta=true, text_delta=true
93
- [TEST] RTK Intermediary: Custom headers and traceparent passthrough... PASS (2453ms)
94
- ↳ Headers accepted cleanly with HTTP 200 OK
95
- [TEST] OpenAI Chat Healer: Orphaned role tool without preceding assistant tool_calls... PASS (1561ms)
96
- ↳ Chat Healer rescued orphaned tool role. Stop reason: length
97
-
98
- ================================================================
99
- TEST RESULTS: 5 PASSED / 0 FAILED
100
- ================================================================
101
- ```
102
-
103
- ---
104
-
105
- ## Recommended User Setup
106
-
107
- For developers using prompt optimization tools:
108
- 1. Keep the optimizer installed in your CLI tool as usual.
109
- 2. In the optimizer's configuration (e.g. `headroom.yaml` or RTK upstream settings), set the upstream target URL to **LLM Switcher** (`http://127.0.0.1:3456`).
110
- 3. Enjoy prompt compression savings without worrying about broken conversation graphs, HTTP 400 crashes, or lost reasoning depth.
1
+ # Token Optimizer Interoperability & Failure Mode Report
2
+
3
+ **How LLM Switcher acts as the protective outermost edge gateway for aggressive prompt/token optimizers (Headroom, RTK, Ponytail).**
4
+
5
+ ---
6
+
7
+ ## Executive Summary
8
+
9
+ Third-party prompt optimizers and token compressors — such as **Headroom**, **RTK (Rust Token Killer)**, and **Ponytail** — attempt to reduce LLM input tokens by aggressively pruning message history, truncating command stdout, or forcing extreme prompt brevity.
10
+
11
+ While these tools can reduce raw token counts in simple scenarios, **they frequently break complex agentic coding workflows** by corrupting message graphs, orphaning tool calls, and stripping reasoning parameters. When these pruned payloads hit upstream APIs directly (such as Anthropic, OpenAI, or 9Router), the provider immediately throws fatal `HTTP 400 Bad Request` errors or severely degrades reasoning depth.
12
+
13
+ **LLM Switcher solves this by acting as the outermost edge gatekeeper (`127.0.0.1:3456`).** It intercepts the pruned payload before it leaves your machine, runs its built-in **Healer Engine** to repair message graphs and restore reasoning parameters, unlocks 1M context windows, and safely converts the protocol to your upstream provider.
14
+
15
+ ---
16
+
17
+ ## Tool Breakdown: What They Do & How They Break Payloads
18
+
19
+ ### 1. Headroom (`headroomlabs-ai/headroom`)
20
+ - **Mechanism:** Runs as a local proxy on `:8787` (or wraps CLI agents). Compresses conversation history, RAG chunks, and tool outputs using SmartCrusher (JSON), CodeCompressor (AST), and Kompress-v2-base. Also attempts "effort routing" to dial down thinking budgets.
21
+ - **Critical Failure Points:**
22
+ - **Orphaned `tool_result` blocks:** When pruning historical turns, Headroom often discards the `assistant` turn containing a `tool_use`, while retaining the subsequent `user` turn containing the `tool_result`. Anthropic's API strictly validates tool use IDs and crashes with:
23
+ ```
24
+ HTTP 400 invalid_request_error: "tool_use_id 'xxx' does not correspond to any tool_use"
25
+ ```
26
+ - **Consecutive `user` turns:** Dropping intermediary assistant turns causes multiple user messages to sit adjacent to each other. Anthropic strictly throws:
27
+ ```
28
+ HTTP 400 invalid_request_error: "roles must alternate between 'user' and 'assistant'"
29
+ ```
30
+ - **Reasoning Suppression:** Its "effort routing" dials down `thinking.budget_tokens` on routine tool turns. On complex models (Claude Opus, Gemini Flash), this prevents the model from formulating multi-step reasoning before acting.
31
+
32
+ ### 2. RTK (`rtk-ai/rtk` - Rust Token Killer)
33
+ - **Mechanism:** A single Rust binary that hooks into shell tool execution (e.g. `PreToolUse` in Claude Code / Cursor) and rewrites CLI commands (`git`, `ls`, `cat`, `grep`, `pytest`) to filter out noise, truncate lines, and inject recall tokens (`[full output: rtk recall xxx]`).
34
+ - **Critical Failure Points:**
35
+ - **Corrupted Structural Data:** When an agent invokes a tool expecting machine-readable JSON or exact AST formatting, RTK's heuristic summaries can alter structural delimiters, causing downstream tool call parsing errors.
36
+ - **Custom Tracking Headers:** RTK and associated tracing proxies inject headers (`x-rtk-*`, `traceparent`, `x-optimizer-id`) that some strict upstream endpoints reject if not cleanly forwarded.
37
+
38
+ ### 3. Ponytail (`DietrichGebert/ponytail`)
39
+ - **Mechanism:** A behavioral prompt engineering plugin/ruleset that injects extreme conciseness instructions ("write one line, it works, YAGNI") into agent system prompts across 20+ coding tools.
40
+ - **Critical Failure Points:**
41
+ - **Premature Execution without Reasoning:** By commanding the model to be maximally brief and avoid planning, reasoning models are discouraged from spending thinking tokens. The model outputs untested single-liners that often fail type checks and test suites.
42
+ - **System Prompt Prefix Invalidation:** Injected rules alter the leading system prompt bytes, busting provider prompt caches unless carefully aligned.
43
+
44
+ ---
45
+
46
+ ## The Healer Engine: How LLM Switcher Protects the Workflow
47
+
48
+ LLM Switcher sits between the optimizer tool and the upstream LLM:
49
+
50
+ ```
51
+ [CLI Agent] ──> [Optimizer: Headroom / RTK] ──> [LLM Switcher :3456] ──> [Upstream / 9Router]
52
+ │
53
+ ├── 1. Heal Orphaned tool_results
54
+ ├── 2. Merge Consecutive Turns
55
+ ├── 3. Restore Stripped Thinking
56
+ └── 4. Enforce 1M Context
57
+ ```
58
+
59
+ ### Protection Matrix: Before vs. After
60
+
61
+ | Scenario | Direct to Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
62
+ |---|---|---|
63
+ | **Orphaned `tool_result` turn** | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block `[Tool Result (id)]: ...` |
64
+ | **Consecutive `user` turns** | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns into a single valid turn seamlessly |
65
+ | **Stripped `thinking` parameters** | ⚠️ **Degraded AI**: Reasoning disabled, model outputs shallow single-liners | ✅ **HTTP 200 OK**: Detects reasoning models and automatically restores safe thinking budget |
66
+ | **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool message into user context |
67
+ | **Custom tracking headers** | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Cleanly passes through `traceparent`, `x-request-id`, `x-rtk-*` |
68
+
69
+ ---
70
+
71
+ ## Test Methodology & Verification Suite
72
+
73
+ We created an automated verification suite in [`tests/live-optimizer-interop.mjs`](../tests/live-optimizer-interop.mjs) that systematically replicates the failure modes of each tool:
74
+
75
+ ### Running the Test Suite
76
+ ```bash
77
+ node tests/live-optimizer-interop.mjs
78
+ ```
79
+
80
+ ### Live Test Results
81
+ ```text
82
+ ================================================================
83
+ LLM SWITCHER — TOKEN OPTIMIZER INTEROPERABILITY TEST SUITE
84
+ Simulating failure modes from Headroom, RTK, and Ponytail
85
+ ================================================================
86
+
87
+ [TEST] Headroom Simulation: Orphaned tool_result turn... PASS (8840ms)
88
+ ↳ Healed orphaned tool_result. Response HTTP 200: "It looks like you've shared a fragment of context ..."
89
+ [TEST] Headroom Simulation: Consecutive User turns (Role alternation violation)... PASS (2683ms)
90
+ ↳ Merged consecutive turns seamlessly. Stop reason: end_turn
91
+ [TEST] Thinking Guard: Restoring stripped thinking parameter on reasoning models... PASS (5372ms)
92
+ ↳ Automatically restored thinking: thinking_delta=true, signature_delta=true, text_delta=true
93
+ [TEST] RTK Intermediary: Custom headers and traceparent passthrough... PASS (2453ms)
94
+ ↳ Headers accepted cleanly with HTTP 200 OK
95
+ [TEST] OpenAI Chat Healer: Orphaned role tool without preceding assistant tool_calls... PASS (1561ms)
96
+ ↳ Chat Healer rescued orphaned tool role. Stop reason: length
97
+
98
+ ================================================================
99
+ TEST RESULTS: 5 PASSED / 0 FAILED
100
+ ================================================================
101
+ ```
102
+
103
+ ---
104
+
105
+ ## Recommended User Setup
106
+
107
+ For developers using prompt optimization tools:
108
+ 1. Keep the optimizer installed in your CLI tool as usual.
109
+ 2. In the optimizer's configuration (e.g. `headroom.yaml` or RTK upstream settings), set the upstream target URL to **LLM Switcher** (`http://127.0.0.1:3456`).
110
+ 3. Enjoy prompt compression savings without worrying about broken conversation graphs, HTTP 400 crashes, or lost reasoning depth.