llm-switcher 1.2.11 → 1.2.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -0
- package/README.md +15 -1
- package/README.vi.md +16 -1
- package/docs/TOKEN-OPTIMIZER-INTEROP.md +110 -110
- package/docs/response-matrix.json +1130 -1130
- package/package.json +1 -1
- package/proxy.mjs +56 -2
- package/service.mjs +12 -6
- package/skills/llm-switcher/SKILL.md +93 -93
- package/switch +0 -0
- package/switch.cmd +2 -2
- package/switch.mjs +100 -15
- package/tests/helpers.mjs +24 -24
- package/tests/live-optimizer-interop.mjs +205 -205
- package/tests/service.test.mjs +27 -1
- package/tests/update.e2e.test.mjs +197 -0
- package/tests/update.test.mjs +145 -0
- package/ui.html +70 -1
- package/update.mjs +98 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,13 @@
|
|
|
1
1
|
# Changelog — LLM Switcher
|
|
2
2
|
|
|
3
|
+
## Release 1.2.12
|
|
4
|
+
|
|
5
|
+
- **Update now:** The update notice on the dashboard has an **Update now** button. The gateway installs the release, and a new gateway process starts on the new code. Then the dashboard reloads. `switch update` does the same from a terminal.
|
|
6
|
+
- **Safe for a git checkout:** A git checkout updates with `git pull --ff-only`. If the checkout has local changes, or commits that are not on the upstream branch, the update stops and changes nothing.
|
|
7
|
+
- **Open requests:** Open requests on the gateway finish on the old code. The Codex interceptor stops and the new gateway starts it again, so a Codex request that is open at that moment fails one time. The old process stays as the parent of the new process, so a service manager keeps the process that it started.
|
|
8
|
+
- **Update at logon:** `switch service install` adds `--autoupdate` to the service. The gateway then installs the newest release before it starts. `--no-autoupdate` installs the service without the flag. Run `switch service install` again to add the flag to a service that is already installed. `switch port` keeps the choice of the installed service.
|
|
9
|
+
- **macOS and Linux:** launchd and `systemd --user` start the gateway with a short `PATH`, and an npm from Homebrew or nvm is not on it. The update now runs npm with the directory of the running `node` first on `PATH`.
|
|
10
|
+
|
|
3
11
|
## Release 1.2.11
|
|
4
12
|
|
|
5
13
|
- **Test connection on a Bifrost model:** The test sent a plain `ping` request, and Anthropic answers that with a fake 429 for a Claude Code account. Now, when intact names `bifrost_ua` for the model, the test reports that the key works and the model is served, and it sends no message.
|
package/README.md
CHANGED
|
@@ -424,12 +424,13 @@ switch ui # Open the Web UI dashboard in your browser
|
|
|
424
424
|
switch status # Display status for all active CLI targets
|
|
425
425
|
switch version # Show the version and tell you when npm has a newer one
|
|
426
426
|
switch doctor # Audit environment, settings & routing
|
|
427
|
+
switch update # Install the newest release and restart the gateway on it
|
|
427
428
|
switch on [profile] # Start the gateway and activate a profile
|
|
428
429
|
switch <profile> # Activate a profile for both tools
|
|
429
430
|
switch claude <profile> # Set the active profile for Claude Code only
|
|
430
431
|
switch codex <profile> # Set the active profile for Codex only
|
|
431
432
|
switch port <number> # Change the gateway port (restarts it if running)
|
|
432
|
-
switch service install # Install OS background autostart service (Windows / macOS / Linux)
|
|
433
|
+
switch service install # Install OS background autostart service (Windows / macOS / Linux); updates at each logon
|
|
433
434
|
switch service uninstall # Remove background autostart service
|
|
434
435
|
switch shim install # Route new Claude and Codex sessions through the gateway
|
|
435
436
|
switch shim status # Verify shims + detect running sessions that bypass the gateway
|
|
@@ -453,6 +454,19 @@ The service runs without your shell. `switch service install` therefore copies `
|
|
|
453
454
|
|
|
454
455
|
On Linux the unit is a systemd *user* service: it starts when you log in. To start it at boot without a login (a server reached over SSH), run `loginctl enable-linger $USER` once. The unit records the absolute path of the current `node`; if you manage Node with nvm and remove that version, run `switch service install` again.
|
|
455
456
|
|
|
457
|
+
### Updates
|
|
458
|
+
|
|
459
|
+
When npm has a newer release, the dashboard shows a notice. Click **Update now** in this notice. The gateway installs the release and starts the new code. Then the dashboard reloads. `switch update` does the same from a terminal.
|
|
460
|
+
|
|
461
|
+
The gateway installs a release in one of two ways:
|
|
462
|
+
|
|
463
|
+
- A git checkout runs `git pull --ff-only`. If the checkout has local changes, or commits that are not on the upstream branch, the update stops and changes nothing.
|
|
464
|
+
- An npm install runs `npm install -g llm-switcher@<latest>`.
|
|
465
|
+
|
|
466
|
+
Requests that are open on the gateway when the update starts finish on the old code. New requests go to the new code. The Codex interceptor stops, and the new gateway starts it again. Thus a Codex request that is open at that moment fails one time. The old process stays as the parent of the new process until the gateway stops. Thus a service manager keeps the process that it started.
|
|
467
|
+
|
|
468
|
+
`switch service install` adds `--autoupdate` to the service command line. With this flag, the gateway installs the newest release before it starts, at each logon. To install the service without this flag, run `switch service install --no-autoupdate`. If a service from an older version is installed, run `switch service install` again to add the flag.
|
|
469
|
+
|
|
456
470
|
### Resumed sessions & the shim (important)
|
|
457
471
|
|
|
458
472
|
`switch on` writes `env-claude.*` and `env-codex.*` and writes nothing into
|
package/README.vi.md
CHANGED
|
@@ -421,12 +421,13 @@ switch ui # Mở giao diện Web UI trên trình duyệt
|
|
|
421
421
|
switch status # Xem trạng thái kích hoạt của tất cả các CLI
|
|
422
422
|
switch version # Xem version đang chạy và báo khi npm có bản mới
|
|
423
423
|
switch doctor # Quét & thanh tra toàn bộ môi trường, settings và định tuyến
|
|
424
|
+
switch update # Cài bản mới nhất và khởi động lại gateway trên bản đó
|
|
424
425
|
switch on [profile] # Khởi động gateway và kích hoạt một profile
|
|
425
426
|
switch <profile> # Kích hoạt một profile cho cả hai công cụ
|
|
426
427
|
switch claude <profile> # Đặt profile kích hoạt riêng cho Claude Code
|
|
427
428
|
switch codex <profile> # Đặt profile kích hoạt riêng cho Codex
|
|
428
429
|
switch port <number> # Đổi cổng gateway (tự restart nếu đang chạy)
|
|
429
|
-
switch service install # Cài đặt gateway thành service chạy ngầm tự bật cùng máy
|
|
430
|
+
switch service install # Cài đặt gateway thành service chạy ngầm tự bật cùng máy; tự update mỗi lần đăng nhập
|
|
430
431
|
switch service uninstall # Gỡ bỏ service chạy ngầm
|
|
431
432
|
switch shim install # Route phiên Claude và Codex mới qua gateway
|
|
432
433
|
switch shim status # Kiểm tra shim + phát hiện phiên đang chạy ngoài gateway
|
|
@@ -450,6 +451,20 @@ các route đầu vào mà chúng đại diện đã bị gỡ.
|
|
|
450
451
|
Service không chạy trong shell của bạn. Vì vậy `switch service install` chép `LLM_SWITCHER_HOME`, `CLAUDE_CONFIG_DIR`, `LLM_SWITCHER_CONFIG`, `LLM_SWITCHER_STATE_DIR` và `LLM_SWITCHER_BLINDFOLD_CERTS` vào unit systemd hoặc plist launchd khi các biến này có giá trị. Task Windows không mang được các biến này; hãy đặt chúng thành biến môi trường User. Nếu file định nghĩa đã cài khác file mới (ví dụ đã sửa tay), file cũ được giữ lại thành `<file>.bak`. Trên Windows, task được tạo từ file XML, nên đường dẫn có dấu cách không cần quote thêm và task không bị giới hạn thời gian chạy. Đường Windows này chưa được test trên Windows.
|
|
451
452
|
|
|
452
453
|
Trên Linux, unit là service systemd *của user*: nó khởi động khi bạn đăng nhập. Muốn nó chạy ngay khi máy boot mà không cần đăng nhập (ví dụ server truy cập qua SSH), chạy một lần `loginctl enable-linger $USER`. Unit ghi đường dẫn tuyệt đối của `node` hiện tại; nếu bạn quản lý Node bằng nvm và gỡ phiên bản đó, hãy chạy lại `switch service install`.
|
|
454
|
+
|
|
455
|
+
### Cập nhật
|
|
456
|
+
|
|
457
|
+
Khi npm có bản mới, dashboard hiện một thông báo. Bấm **Update now** trong thông báo này. Gateway cài bản mới và chạy code mới. Sau đó dashboard tự tải lại. `switch update` làm việc tương tự từ terminal.
|
|
458
|
+
|
|
459
|
+
Gateway cài bản mới theo một trong hai cách:
|
|
460
|
+
|
|
461
|
+
- Bản git checkout chạy `git pull --ff-only`. Nếu checkout có thay đổi cục bộ, hoặc có commit không nằm trên nhánh upstream, việc cập nhật dừng lại và không đổi gì.
|
|
462
|
+
- Bản cài bằng npm chạy `npm install -g llm-switcher@<latest>`.
|
|
463
|
+
|
|
464
|
+
Các request đang chạy trên gateway lúc cập nhật sẽ chạy xong trên code cũ. Request mới đi vào code mới. Interceptor của Codex dừng lại, và gateway mới khởi động nó lại. Vì vậy một request Codex đang chạy đúng lúc đó sẽ lỗi một lần. Tiến trình cũ ở lại làm tiến trình cha của tiến trình mới cho tới khi gateway dừng. Nhờ vậy service manager vẫn giữ đúng tiến trình mà nó đã khởi động.
|
|
465
|
+
|
|
466
|
+
`switch service install` thêm `--autoupdate` vào dòng lệnh của service. Với cờ này, gateway cài bản mới nhất trước khi khởi động, ở mỗi lần đăng nhập. Muốn cài service không có cờ này, chạy `switch service install --no-autoupdate`. Nếu service đã cài từ phiên bản cũ, chạy lại `switch service install` để thêm cờ.
|
|
467
|
+
|
|
453
468
|
### Phiên mở lại (`--resume`) và cơ chế shim — quan trọng
|
|
454
469
|
|
|
455
470
|
`switch on` ghi `env-claude.*` và `env-codex.*`, và **không ghi gì** vào
|
|
@@ -1,110 +1,110 @@
|
|
|
1
|
-
# Token Optimizer Interoperability & Failure Mode Report
|
|
2
|
-
|
|
3
|
-
**How LLM Switcher acts as the protective outermost edge gateway for aggressive prompt/token optimizers (Headroom, RTK, Ponytail).**
|
|
4
|
-
|
|
5
|
-
---
|
|
6
|
-
|
|
7
|
-
## Executive Summary
|
|
8
|
-
|
|
9
|
-
Third-party prompt optimizers and token compressors — such as **Headroom**, **RTK (Rust Token Killer)**, and **Ponytail** — attempt to reduce LLM input tokens by aggressively pruning message history, truncating command stdout, or forcing extreme prompt brevity.
|
|
10
|
-
|
|
11
|
-
While these tools can reduce raw token counts in simple scenarios, **they frequently break complex agentic coding workflows** by corrupting message graphs, orphaning tool calls, and stripping reasoning parameters. When these pruned payloads hit upstream APIs directly (such as Anthropic, OpenAI, or 9Router), the provider immediately throws fatal `HTTP 400 Bad Request` errors or severely degrades reasoning depth.
|
|
12
|
-
|
|
13
|
-
**LLM Switcher solves this by acting as the outermost edge gatekeeper (`127.0.0.1:3456`).** It intercepts the pruned payload before it leaves your machine, runs its built-in **Healer Engine** to repair message graphs and restore reasoning parameters, unlocks 1M context windows, and safely converts the protocol to your upstream provider.
|
|
14
|
-
|
|
15
|
-
---
|
|
16
|
-
|
|
17
|
-
## Tool Breakdown: What They Do & How They Break Payloads
|
|
18
|
-
|
|
19
|
-
### 1. Headroom (`headroomlabs-ai/headroom`)
|
|
20
|
-
- **Mechanism:** Runs as a local proxy on `:8787` (or wraps CLI agents). Compresses conversation history, RAG chunks, and tool outputs using SmartCrusher (JSON), CodeCompressor (AST), and Kompress-v2-base. Also attempts "effort routing" to dial down thinking budgets.
|
|
21
|
-
- **Critical Failure Points:**
|
|
22
|
-
- **Orphaned `tool_result` blocks:** When pruning historical turns, Headroom often discards the `assistant` turn containing a `tool_use`, while retaining the subsequent `user` turn containing the `tool_result`. Anthropic's API strictly validates tool use IDs and crashes with:
|
|
23
|
-
```
|
|
24
|
-
HTTP 400 invalid_request_error: "tool_use_id 'xxx' does not correspond to any tool_use"
|
|
25
|
-
```
|
|
26
|
-
- **Consecutive `user` turns:** Dropping intermediary assistant turns causes multiple user messages to sit adjacent to each other. Anthropic strictly throws:
|
|
27
|
-
```
|
|
28
|
-
HTTP 400 invalid_request_error: "roles must alternate between 'user' and 'assistant'"
|
|
29
|
-
```
|
|
30
|
-
- **Reasoning Suppression:** Its "effort routing" dials down `thinking.budget_tokens` on routine tool turns. On complex models (Claude Opus, Gemini Flash), this prevents the model from formulating multi-step reasoning before acting.
|
|
31
|
-
|
|
32
|
-
### 2. RTK (`rtk-ai/rtk` - Rust Token Killer)
|
|
33
|
-
- **Mechanism:** A single Rust binary that hooks into shell tool execution (e.g. `PreToolUse` in Claude Code / Cursor) and rewrites CLI commands (`git`, `ls`, `cat`, `grep`, `pytest`) to filter out noise, truncate lines, and inject recall tokens (`[full output: rtk recall xxx]`).
|
|
34
|
-
- **Critical Failure Points:**
|
|
35
|
-
- **Corrupted Structural Data:** When an agent invokes a tool expecting machine-readable JSON or exact AST formatting, RTK's heuristic summaries can alter structural delimiters, causing downstream tool call parsing errors.
|
|
36
|
-
- **Custom Tracking Headers:** RTK and associated tracing proxies inject headers (`x-rtk-*`, `traceparent`, `x-optimizer-id`) that some strict upstream endpoints reject if not cleanly forwarded.
|
|
37
|
-
|
|
38
|
-
### 3. Ponytail (`DietrichGebert/ponytail`)
|
|
39
|
-
- **Mechanism:** A behavioral prompt engineering plugin/ruleset that injects extreme conciseness instructions ("write one line, it works, YAGNI") into agent system prompts across 20+ coding tools.
|
|
40
|
-
- **Critical Failure Points:**
|
|
41
|
-
- **Premature Execution without Reasoning:** By commanding the model to be maximally brief and avoid planning, reasoning models are discouraged from spending thinking tokens. The model outputs untested single-liners that often fail type checks and test suites.
|
|
42
|
-
- **System Prompt Prefix Invalidation:** Injected rules alter the leading system prompt bytes, busting provider prompt caches unless carefully aligned.
|
|
43
|
-
|
|
44
|
-
---
|
|
45
|
-
|
|
46
|
-
## The Healer Engine: How LLM Switcher Protects the Workflow
|
|
47
|
-
|
|
48
|
-
LLM Switcher sits between the optimizer tool and the upstream LLM:
|
|
49
|
-
|
|
50
|
-
```
|
|
51
|
-
[CLI Agent] ──> [Optimizer: Headroom / RTK] ──> [LLM Switcher :3456] ──> [Upstream / 9Router]
|
|
52
|
-
│
|
|
53
|
-
├── 1. Heal Orphaned tool_results
|
|
54
|
-
├── 2. Merge Consecutive Turns
|
|
55
|
-
├── 3. Restore Stripped Thinking
|
|
56
|
-
└── 4. Enforce 1M Context
|
|
57
|
-
```
|
|
58
|
-
|
|
59
|
-
### Protection Matrix: Before vs. After
|
|
60
|
-
|
|
61
|
-
| Scenario | Direct to Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
|
|
62
|
-
|---|---|---|
|
|
63
|
-
| **Orphaned `tool_result` turn** | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block `[Tool Result (id)]: ...` |
|
|
64
|
-
| **Consecutive `user` turns** | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns into a single valid turn seamlessly |
|
|
65
|
-
| **Stripped `thinking` parameters** | ⚠️ **Degraded AI**: Reasoning disabled, model outputs shallow single-liners | ✅ **HTTP 200 OK**: Detects reasoning models and automatically restores safe thinking budget |
|
|
66
|
-
| **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool message into user context |
|
|
67
|
-
| **Custom tracking headers** | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Cleanly passes through `traceparent`, `x-request-id`, `x-rtk-*` |
|
|
68
|
-
|
|
69
|
-
---
|
|
70
|
-
|
|
71
|
-
## Test Methodology & Verification Suite
|
|
72
|
-
|
|
73
|
-
We created an automated verification suite in [`tests/live-optimizer-interop.mjs`](../tests/live-optimizer-interop.mjs) that systematically replicates the failure modes of each tool:
|
|
74
|
-
|
|
75
|
-
### Running the Test Suite
|
|
76
|
-
```bash
|
|
77
|
-
node tests/live-optimizer-interop.mjs
|
|
78
|
-
```
|
|
79
|
-
|
|
80
|
-
### Live Test Results
|
|
81
|
-
```text
|
|
82
|
-
================================================================
|
|
83
|
-
LLM SWITCHER — TOKEN OPTIMIZER INTEROPERABILITY TEST SUITE
|
|
84
|
-
Simulating failure modes from Headroom, RTK, and Ponytail
|
|
85
|
-
================================================================
|
|
86
|
-
|
|
87
|
-
[TEST] Headroom Simulation: Orphaned tool_result turn... PASS (8840ms)
|
|
88
|
-
↳ Healed orphaned tool_result. Response HTTP 200: "It looks like you've shared a fragment of context ..."
|
|
89
|
-
[TEST] Headroom Simulation: Consecutive User turns (Role alternation violation)... PASS (2683ms)
|
|
90
|
-
↳ Merged consecutive turns seamlessly. Stop reason: end_turn
|
|
91
|
-
[TEST] Thinking Guard: Restoring stripped thinking parameter on reasoning models... PASS (5372ms)
|
|
92
|
-
↳ Automatically restored thinking: thinking_delta=true, signature_delta=true, text_delta=true
|
|
93
|
-
[TEST] RTK Intermediary: Custom headers and traceparent passthrough... PASS (2453ms)
|
|
94
|
-
↳ Headers accepted cleanly with HTTP 200 OK
|
|
95
|
-
[TEST] OpenAI Chat Healer: Orphaned role tool without preceding assistant tool_calls... PASS (1561ms)
|
|
96
|
-
↳ Chat Healer rescued orphaned tool role. Stop reason: length
|
|
97
|
-
|
|
98
|
-
================================================================
|
|
99
|
-
TEST RESULTS: 5 PASSED / 0 FAILED
|
|
100
|
-
================================================================
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
---
|
|
104
|
-
|
|
105
|
-
## Recommended User Setup
|
|
106
|
-
|
|
107
|
-
For developers using prompt optimization tools:
|
|
108
|
-
1. Keep the optimizer installed in your CLI tool as usual.
|
|
109
|
-
2. In the optimizer's configuration (e.g. `headroom.yaml` or RTK upstream settings), set the upstream target URL to **LLM Switcher** (`http://127.0.0.1:3456`).
|
|
110
|
-
3. Enjoy prompt compression savings without worrying about broken conversation graphs, HTTP 400 crashes, or lost reasoning depth.
|
|
1
|
+
# Token Optimizer Interoperability & Failure Mode Report
|
|
2
|
+
|
|
3
|
+
**How LLM Switcher acts as the protective outermost edge gateway for aggressive prompt/token optimizers (Headroom, RTK, Ponytail).**
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Executive Summary
|
|
8
|
+
|
|
9
|
+
Third-party prompt optimizers and token compressors — such as **Headroom**, **RTK (Rust Token Killer)**, and **Ponytail** — attempt to reduce LLM input tokens by aggressively pruning message history, truncating command stdout, or forcing extreme prompt brevity.
|
|
10
|
+
|
|
11
|
+
While these tools can reduce raw token counts in simple scenarios, **they frequently break complex agentic coding workflows** by corrupting message graphs, orphaning tool calls, and stripping reasoning parameters. When these pruned payloads hit upstream APIs directly (such as Anthropic, OpenAI, or 9Router), the provider immediately throws fatal `HTTP 400 Bad Request` errors or severely degrades reasoning depth.
|
|
12
|
+
|
|
13
|
+
**LLM Switcher solves this by acting as the outermost edge gatekeeper (`127.0.0.1:3456`).** It intercepts the pruned payload before it leaves your machine, runs its built-in **Healer Engine** to repair message graphs and restore reasoning parameters, unlocks 1M context windows, and safely converts the protocol to your upstream provider.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Tool Breakdown: What They Do & How They Break Payloads
|
|
18
|
+
|
|
19
|
+
### 1. Headroom (`headroomlabs-ai/headroom`)
|
|
20
|
+
- **Mechanism:** Runs as a local proxy on `:8787` (or wraps CLI agents). Compresses conversation history, RAG chunks, and tool outputs using SmartCrusher (JSON), CodeCompressor (AST), and Kompress-v2-base. Also attempts "effort routing" to dial down thinking budgets.
|
|
21
|
+
- **Critical Failure Points:**
|
|
22
|
+
- **Orphaned `tool_result` blocks:** When pruning historical turns, Headroom often discards the `assistant` turn containing a `tool_use`, while retaining the subsequent `user` turn containing the `tool_result`. Anthropic's API strictly validates tool use IDs and crashes with:
|
|
23
|
+
```
|
|
24
|
+
HTTP 400 invalid_request_error: "tool_use_id 'xxx' does not correspond to any tool_use"
|
|
25
|
+
```
|
|
26
|
+
- **Consecutive `user` turns:** Dropping intermediary assistant turns causes multiple user messages to sit adjacent to each other. Anthropic strictly throws:
|
|
27
|
+
```
|
|
28
|
+
HTTP 400 invalid_request_error: "roles must alternate between 'user' and 'assistant'"
|
|
29
|
+
```
|
|
30
|
+
- **Reasoning Suppression:** Its "effort routing" dials down `thinking.budget_tokens` on routine tool turns. On complex models (Claude Opus, Gemini Flash), this prevents the model from formulating multi-step reasoning before acting.
|
|
31
|
+
|
|
32
|
+
### 2. RTK (`rtk-ai/rtk` - Rust Token Killer)
|
|
33
|
+
- **Mechanism:** A single Rust binary that hooks into shell tool execution (e.g. `PreToolUse` in Claude Code / Cursor) and rewrites CLI commands (`git`, `ls`, `cat`, `grep`, `pytest`) to filter out noise, truncate lines, and inject recall tokens (`[full output: rtk recall xxx]`).
|
|
34
|
+
- **Critical Failure Points:**
|
|
35
|
+
- **Corrupted Structural Data:** When an agent invokes a tool expecting machine-readable JSON or exact AST formatting, RTK's heuristic summaries can alter structural delimiters, causing downstream tool call parsing errors.
|
|
36
|
+
- **Custom Tracking Headers:** RTK and associated tracing proxies inject headers (`x-rtk-*`, `traceparent`, `x-optimizer-id`) that some strict upstream endpoints reject if not cleanly forwarded.
|
|
37
|
+
|
|
38
|
+
### 3. Ponytail (`DietrichGebert/ponytail`)
|
|
39
|
+
- **Mechanism:** A behavioral prompt engineering plugin/ruleset that injects extreme conciseness instructions ("write one line, it works, YAGNI") into agent system prompts across 20+ coding tools.
|
|
40
|
+
- **Critical Failure Points:**
|
|
41
|
+
- **Premature Execution without Reasoning:** By commanding the model to be maximally brief and avoid planning, reasoning models are discouraged from spending thinking tokens. The model outputs untested single-liners that often fail type checks and test suites.
|
|
42
|
+
- **System Prompt Prefix Invalidation:** Injected rules alter the leading system prompt bytes, busting provider prompt caches unless carefully aligned.
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## The Healer Engine: How LLM Switcher Protects the Workflow
|
|
47
|
+
|
|
48
|
+
LLM Switcher sits between the optimizer tool and the upstream LLM:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
[CLI Agent] ──> [Optimizer: Headroom / RTK] ──> [LLM Switcher :3456] ──> [Upstream / 9Router]
|
|
52
|
+
│
|
|
53
|
+
├── 1. Heal Orphaned tool_results
|
|
54
|
+
├── 2. Merge Consecutive Turns
|
|
55
|
+
├── 3. Restore Stripped Thinking
|
|
56
|
+
└── 4. Enforce 1M Context
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
### Protection Matrix: Before vs. After
|
|
60
|
+
|
|
61
|
+
| Scenario | Direct to Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| **Orphaned `tool_result` turn** | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block `[Tool Result (id)]: ...` |
|
|
64
|
+
| **Consecutive `user` turns** | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns into a single valid turn seamlessly |
|
|
65
|
+
| **Stripped `thinking` parameters** | ⚠️ **Degraded AI**: Reasoning disabled, model outputs shallow single-liners | ✅ **HTTP 200 OK**: Detects reasoning models and automatically restores safe thinking budget |
|
|
66
|
+
| **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool message into user context |
|
|
67
|
+
| **Custom tracking headers** | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Cleanly passes through `traceparent`, `x-request-id`, `x-rtk-*` |
|
|
68
|
+
|
|
69
|
+
---
|
|
70
|
+
|
|
71
|
+
## Test Methodology & Verification Suite
|
|
72
|
+
|
|
73
|
+
We created an automated verification suite in [`tests/live-optimizer-interop.mjs`](../tests/live-optimizer-interop.mjs) that systematically replicates the failure modes of each tool:
|
|
74
|
+
|
|
75
|
+
### Running the Test Suite
|
|
76
|
+
```bash
|
|
77
|
+
node tests/live-optimizer-interop.mjs
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
### Live Test Results
|
|
81
|
+
```text
|
|
82
|
+
================================================================
|
|
83
|
+
LLM SWITCHER — TOKEN OPTIMIZER INTEROPERABILITY TEST SUITE
|
|
84
|
+
Simulating failure modes from Headroom, RTK, and Ponytail
|
|
85
|
+
================================================================
|
|
86
|
+
|
|
87
|
+
[TEST] Headroom Simulation: Orphaned tool_result turn... PASS (8840ms)
|
|
88
|
+
↳ Healed orphaned tool_result. Response HTTP 200: "It looks like you've shared a fragment of context ..."
|
|
89
|
+
[TEST] Headroom Simulation: Consecutive User turns (Role alternation violation)... PASS (2683ms)
|
|
90
|
+
↳ Merged consecutive turns seamlessly. Stop reason: end_turn
|
|
91
|
+
[TEST] Thinking Guard: Restoring stripped thinking parameter on reasoning models... PASS (5372ms)
|
|
92
|
+
↳ Automatically restored thinking: thinking_delta=true, signature_delta=true, text_delta=true
|
|
93
|
+
[TEST] RTK Intermediary: Custom headers and traceparent passthrough... PASS (2453ms)
|
|
94
|
+
↳ Headers accepted cleanly with HTTP 200 OK
|
|
95
|
+
[TEST] OpenAI Chat Healer: Orphaned role tool without preceding assistant tool_calls... PASS (1561ms)
|
|
96
|
+
↳ Chat Healer rescued orphaned tool role. Stop reason: length
|
|
97
|
+
|
|
98
|
+
================================================================
|
|
99
|
+
TEST RESULTS: 5 PASSED / 0 FAILED
|
|
100
|
+
================================================================
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## Recommended User Setup
|
|
106
|
+
|
|
107
|
+
For developers using prompt optimization tools:
|
|
108
|
+
1. Keep the optimizer installed in your CLI tool as usual.
|
|
109
|
+
2. In the optimizer's configuration (e.g. `headroom.yaml` or RTK upstream settings), set the upstream target URL to **LLM Switcher** (`http://127.0.0.1:3456`).
|
|
110
|
+
3. Enjoy prompt compression savings without worrying about broken conversation graphs, HTTP 400 crashes, or lost reasoning depth.
|