llm-switcher 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.gitattributes +16 -0
- package/LICENSE +21 -0
- package/README.md +587 -0
- package/README.vi.md +585 -0
- package/blindfold/blindfold.mjs +633 -0
- package/blindfold/make-certs.sh +88 -0
- package/blindfold/wsframe.mjs +176 -0
- package/codex-catalog-template.json +1 -0
- package/config.example.json +84 -0
- package/contract-exclusions.json +41 -0
- package/contract.mjs +561 -0
- package/docs/LLM-RESPONSE-MATRIX.md +165 -0
- package/docs/TOKEN-OPTIMIZER-INTEROP.md +110 -0
- package/docs/codex-blindfold.md +214 -0
- package/docs/cross-platform.md +136 -0
- package/docs/diagrams/blindfold-request-routing.html +14972 -0
- package/docs/diagrams/blindfold-request-routing.sequence.json +175 -0
- package/docs/diagrams/blindfold-switch-lifecycle.html +14958 -0
- package/docs/diagrams/blindfold-switch-lifecycle.lifecycle.json +159 -0
- package/docs/diagrams/codex-model-name-resolution.html +15005 -0
- package/docs/diagrams/codex-model-name-resolution.workflow.json +71 -0
- package/docs/response-matrix.json +1131 -0
- package/formats.mjs +2308 -0
- package/mcp.mjs +340 -0
- package/package.json +36 -0
- package/proxy.mjs +1743 -0
- package/service.mjs +132 -0
- package/shim.mjs +292 -0
- package/skills/llm-switcher/SKILL.md +88 -0
- package/state.mjs +978 -0
- package/switch +5 -0
- package/switch.cmd +2 -0
- package/switch.mjs +930 -0
- package/tests/blindfold.test.mjs +307 -0
- package/tests/blindfold.wire.test.mjs +170 -0
- package/tests/contract/run.test.mjs +214 -0
- package/tests/contract-check.test.mjs +458 -0
- package/tests/contract-lab.test.mjs +755 -0
- package/tests/datadir.test.mjs +37 -0
- package/tests/formats.test.mjs +794 -0
- package/tests/gateway.e2e.test.mjs +999 -0
- package/tests/helpers.mjs +24 -0
- package/tests/lifecycle.test.mjs +416 -0
- package/tests/live-optimizer-interop.mjs +205 -0
- package/tests/mcp.test.mjs +91 -0
- package/tests/service.test.mjs +69 -0
- package/tests/shim.test.mjs +228 -0
- package/tests/state.test.mjs +675 -0
- package/tests/switch.test.mjs +156 -0
- package/tests/wsframe.test.mjs +154 -0
- package/ui.html +2234 -0
|
@@ -0,0 +1,165 @@
|
|
|
1
|
+
# LLM Provider Response Matrix
|
|
2
|
+
|
|
3
|
+
How different LLM APIs shape their responses — and how `llm-switcher`
|
|
4
|
+
normalizes them. Companion machine-readable file: [`response-matrix.json`](response-matrix.json).
|
|
5
|
+
|
|
6
|
+
## Method
|
|
7
|
+
|
|
8
|
+
- **Live-sampled (primary):** 48 responses from an OpenAI-compatible gateway
|
|
9
|
+
(9Router, 2026-09-12): **8 backend families × 6 variants**
|
|
10
|
+
(`base-stream`, `think-stream`, `think-nostream`, `tool-stream`,
|
|
11
|
+
`trunc-stream` with `max_tokens: 40`, `effort-stream` with
|
|
12
|
+
`reasoning_effort: high`). Reasoning prompts were used on purpose —
|
|
13
|
+
trivial prompts (e.g. "pong") often produce **no** thinking trace at all.
|
|
14
|
+
- **Doc-based (secondary):** OpenRouter reasoning docs, Vertex AI
|
|
15
|
+
`generateContent` REST shape, OpenAI/Anthropic API references — used to
|
|
16
|
+
cover field names this gateway never emits (e.g. `reasoning_details[]`,
|
|
17
|
+
Vertex `thought` parts, `thoughtSignature`).
|
|
18
|
+
|
|
19
|
+
No credentials or personal data appear in the samples.
|
|
20
|
+
|
|
21
|
+
## TL;DR
|
|
22
|
+
|
|
23
|
+
| # | Finding |
|
|
24
|
+
|---|---|
|
|
25
|
+
| 1 | Same gateway, different field names per backend: `reasoning_content` (claude/openai/nvidia/gemini-stream), `reasoning` + `reasoning_content` **duplicated** (zai), nothing at all (qwen/plain). |
|
|
26
|
+
| 2 | **Non-streaming drops hidden thinking** on most backends — only `content` (+ token counters) comes back. Streaming is the only reliable way to get thinking blocks. |
|
|
27
|
+
| 3 | Usage shapes differ: single total chunk (OpenAI-style), per-token deltas (Cloudflare-style `qwen`/`zai` with `neurons`), or **absent in stream** (`gpt-oss` family) — count deltas manually as fallback. |
|
|
28
|
+
| 4 | `finish_reason` may arrive in the **first** chunk (qwen) or repeat per chunk; always take the **last** non-null value. |
|
|
29
|
+
| 5 | Tool calls may stream split across chunks (qwen: 6 chunks) or whole; args may be `string` (OpenAI) or `object` (Vertex `functionCall.args`). |
|
|
30
|
+
| 6 | Request-side: thinking budgets `< 1024` are rejected (400); `reasoning_effort` is rejected by some backends (503 `INVALID_ARGUMENT` on `gpt-oss`); gateways cool down after bad requests (`reset after Ns`). |
|
|
31
|
+
|
|
32
|
+
## 1. Envelopes
|
|
33
|
+
|
|
34
|
+
### 1.1 OpenAI Chat Completions (most common)
|
|
35
|
+
|
|
36
|
+
Non-stream:
|
|
37
|
+
```json
|
|
38
|
+
{"id":"chatcmpl-...","object":"chat.completion","created":1789184460,"model":"...claude-opus-4-6-thinking",
|
|
39
|
+
"choices":[{"index":0,"message":{"role":"assistant","content":"pong"},"finish_reason":"stop"}],
|
|
40
|
+
"usage":{"prompt_tokens":2043,"completion_tokens":15,"total_tokens":2058}}
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
Stream (`data:` lines, terminal `[DONE]`):
|
|
44
|
+
```
|
|
45
|
+
data: {"...","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
|
|
46
|
+
data: {"...","choices":[{"index":0,"delta":{"reasoning_content":"p"},"finish_reason":null}]}
|
|
47
|
+
data: {"...","choices":[{"index":0,"delta":{"content":"p"},"finish_reason":null}]}
|
|
48
|
+
data: {"...","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{...}}
|
|
49
|
+
data: [DONE]
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
### 1.2 Vertex `generateContent` (native)
|
|
53
|
+
|
|
54
|
+
Non-stream:
|
|
55
|
+
```json
|
|
56
|
+
{"candidates":[{"content":{"role":"model","parts":[
|
|
57
|
+
{"text":"plan...","thought":true,"thoughtSignature":"..."},
|
|
58
|
+
{"text":"answer"},
|
|
59
|
+
{"functionCall":{"name":"calc","args":{"expr":"2+2"}}}]},
|
|
60
|
+
"finishReason":"STOP"}],
|
|
61
|
+
"usageMetadata":{"promptTokenCount":10,"candidatesTokenCount":20,"totalTokenCount":30}}
|
|
62
|
+
```
|
|
63
|
+
Stream (`:streamGenerateContent?alt=sse`): same objects as `data:` lines, incremental parts.
|
|
64
|
+
|
|
65
|
+
### 1.3 Anthropic Messages (native target)
|
|
66
|
+
|
|
67
|
+
```json
|
|
68
|
+
{"id":"msg_...","type":"message","role":"assistant","model":"...",
|
|
69
|
+
"content":[{"type":"thinking","thinking":"...","signature":"..."},
|
|
70
|
+
{"type":"text","text":"..."},
|
|
71
|
+
{"type":"tool_use","id":"...","name":"...","input":{}}],
|
|
72
|
+
"stop_reason":"end_turn",
|
|
73
|
+
"usage":{"input_tokens":0,"output_tokens":0}}
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
### 1.4 OpenRouter extensions (doc-based)
|
|
77
|
+
|
|
78
|
+
- `message.reasoning` / `delta.reasoning` (plain string; `reasoning_content` is an accepted alias).
|
|
79
|
+
- `message.reasoning_details[]` / `delta.reasoning_details[]` with typed items:
|
|
80
|
+
`reasoning.text {text, signature}`, `reasoning.summary {summary}`,
|
|
81
|
+
`reasoning.encrypted {data}` (+ `id`, `format`, `index`).
|
|
82
|
+
- Request: unified `reasoning: {effort | max_tokens, exclude, enabled}`;
|
|
83
|
+
Anthropic minimum budget **1024**; Gemini 3 uses `thinkingLevel`
|
|
84
|
+
(`minimal|low|medium|high`) instead of token budgets.
|
|
85
|
+
|
|
86
|
+
## 2. Reasoning / thinking — where it actually appears (live)
|
|
87
|
+
|
|
88
|
+
`R` = `reasoning_content`, `r` = `reasoning`, `–` = absent.
|
|
89
|
+
|
|
90
|
+
| Backend | base-stream | think-stream | think-nostream | tool-stream | trunc-stream |
|
|
91
|
+
|---|---|---|---|---|---|
|
|
92
|
+
| claude-budget (opus-4-6-thinking) | R | R | – (content only!) | R + tools | R |
|
|
93
|
+
| claude-adaptive (sonnet-4-6) | – | – | – | tools, no R | – |
|
|
94
|
+
| gemini (3.8-flash-high) | R (hard tasks only) | R | – (+`reasoning_tokens` counter) | tools, no R | R |
|
|
95
|
+
| openai (gpt-oss-120b) | R | R | – | R + tools | R, no content |
|
|
96
|
+
| qwen (qwq-32b) | – (CoT leaks into `content`) | – | – (long CoT as content) | tools split ×6 chunks | – |
|
|
97
|
+
| zai (glm-4.7-flash) | R **and** r (same text twice!) | R + r | R + r, `content: null` | R + r + tools | R + r |
|
|
98
|
+
| nvidia (nemotron) | R (no usage in stream!) | R | – | R + tools | R |
|
|
99
|
+
| plain (gpt-4o-mini) | – | – | – (+`padding` junk field) | tools | – |
|
|
100
|
+
|
|
101
|
+
Notes:
|
|
102
|
+
- Gemini emits `R` as **one big chunk** (not token-split); others stream it token by token.
|
|
103
|
+
- Easy prompts may yield zero thinking even on thinking-capable models, while
|
|
104
|
+
`reasoning_tokens` in usage still counts hidden work.
|
|
105
|
+
- `zai` duplicates every reasoning delta under both keys — parsers must take the
|
|
106
|
+
**first** match, never concatenate.
|
|
107
|
+
|
|
108
|
+
## 3. Tool calls (live)
|
|
109
|
+
|
|
110
|
+
| Backend | Shape |
|
|
111
|
+
|---|---|
|
|
112
|
+
| OpenAI-style | `tool_calls: [{id, type:"function", function:{name, arguments:"{...}"}}]`, whole or split across chunks |
|
|
113
|
+
| qwen | same shape, but split across ~6 delta chunks; `finish_reason:"stop"` repeats per chunk |
|
|
114
|
+
| zai | same shape + duplicated reasoning alongside |
|
|
115
|
+
| Vertex native | `parts: [{functionCall:{name, args:{...}}}]` (args is an **object**), no id |
|
|
116
|
+
| plain (gh) | same shape; stream interleaves `content:null` chunks |
|
|
117
|
+
|
|
118
|
+
Missing ids are synthesized as `call_<24 random chars>` (Vertex: `call_vtx_<n>_<name>`, healer: `call_heal_<n>`); args objects are stringified
|
|
119
|
+
for OpenAI-shaped outputs and parsed back to objects for Anthropic/Vertex outputs.
|
|
120
|
+
|
|
121
|
+
## 4. Usage accounting (live)
|
|
122
|
+
|
|
123
|
+
| Backend | Stream usage | Keys |
|
|
124
|
+
|---|---|---|
|
|
125
|
+
| claude/adaptive/gemini/plain | one final chunk | `prompt_tokens, completion_tokens, total_tokens` (+ `completion_tokens_details.reasoning_tokens` on gemini, `prompt_tokens_details.cached_tokens` where cached) |
|
|
126
|
+
| qwen/zai (Cloudflare) | **every** chunk | per-token deltas (`completion_tokens: 1` × N) + `neurons`, sometimes `estimated`; prompt total only on first chunk |
|
|
127
|
+
| nvidia | `usage: null` per chunk, one real total at end (`estimated` flag) | same keys as OpenAI-style |
|
|
128
|
+
| openai (gpt-oss) | **none in stream** | count deltas manually |
|
|
129
|
+
|
|
130
|
+
Rule used by the proxy: `prompt = MAX(seen)`. For per-chunk deltas, `completion = SUM(deltas)`.
|
|
131
|
+
For cumulative usage, `completion = LAST`. The shape is detected per stream. When upstream sends
|
|
132
|
+
nothing, the proxy counts the deltas itself. Gemini counts thoughts apart from candidates, so its
|
|
133
|
+
completion is `candidatesTokenCount + thoughtsTokenCount`.
|
|
134
|
+
|
|
135
|
+
## 5. Finish reasons, ids, errors (live)
|
|
136
|
+
|
|
137
|
+
- Observed `finish_reason`: `stop`, `tool_calls`, `length` (also `STOP`/`MAX_TOKENS`/`SAFETY…` on Vertex-native). Mapping: `stop→end_turn`, `length→max_tokens`, `tool_calls→tool_use` (Anthropic); `STOP/MAX_TOKENS/SAFETY` (Vertex output).
|
|
138
|
+
- If any tool call was seen, the final stop is forced to `tool_use` regardless.
|
|
139
|
+
- Id prefixes vary (`chatcmpl-`, `chatcmpl-msg_`, `chatcmpl-req_`, `id-…`); outputs are normalized to `msg_` / `chatcmpl-` / `resp_` per client protocol.
|
|
140
|
+
- Error envelope: `{"error":{"message":"[backend/model] [400]: {...} (reset after Ns)"}}` — note the cooldown hint; back off instead of retrying hot.
|
|
141
|
+
- Junk to ignore: `padding`, `copilot_usage`, `service_tier`, `system_fingerprint`, `logprobs`, `neurons`, `estimated`, `refusal`, `annotations`, `audio`, `token_ids`, empty `choices: []` chunks, doubled `[DONE]`.
|
|
142
|
+
|
|
143
|
+
## 6. Request-side gotchas (live)
|
|
144
|
+
|
|
145
|
+
- `thinking: {type:"enabled", budget_tokens:<1024}` → HTTP 400. Clamp to ≥ 1024.
|
|
146
|
+
- `reasoning_effort: "high"` → HTTP 503 `INVALID_ARGUMENT` on `gpt-oss`-class backends. Only send effort levels the backend accepts.
|
|
147
|
+
- `include_reasoning: true` (legacy) is accepted but changes nothing observable.
|
|
148
|
+
- Omitting `stream` defaults to **streaming** on some gateways — always send it explicitly.
|
|
149
|
+
- `stream_options: {include_usage:true}` is required by spec for stream usage, though some gateways send usage anyway.
|
|
150
|
+
|
|
151
|
+
## 7. How the proxy maps this (`formats.mjs`)
|
|
152
|
+
|
|
153
|
+
- `smartReasoning / smartText / smartToolCalls / smartUsage / smartFinish` — accept every field name above, first-match wins (no double-count).
|
|
154
|
+
- `normalizeUpstream(parsed, outFormat)` — one SSE chunk **or** one full JSON → canonical events `{think, text, tools, finish, usage}`; also walks native Anthropic event types and Vertex `candidates`.
|
|
155
|
+
- Renderers emit the client protocol: `createAnthropicStream` (thinking+signature → `thinking_delta`, `<think>`-tag fallback with split-tag buffering), `createChatStream` (`reasoning_content` extension), `createResponsesStream` (Codex reasoning summary + deltas), `createVertexStream` (`thought` parts + `thoughtSignature` passthrough when genuine).
|
|
156
|
+
- Known non-goals: hidden thinking that the backend never returns (gemini non-stream counters, qwen CoT-in-content) cannot be recovered; encrypted `reasoning.encrypted.data` blobs are skipped as thinking text.
|
|
157
|
+
|
|
158
|
+
## 8. Reproduce
|
|
159
|
+
|
|
160
|
+
Send the same prompt matrix (`base/think/nostream/tool/trunc/effort` × `stream on/off`)
|
|
161
|
+
against any OpenAI-compatible `/chat/completions` endpoint and record, per chunk,
|
|
162
|
+
`delta` key sets, reasoning/text presence, `finish_reason` sequence, usage key
|
|
163
|
+
sets, and id prefixes. A reasoning-heavy prompt is required — trivial prompts
|
|
164
|
+
often yield no thinking trace. Never commit keys; keep sampler scripts out of
|
|
165
|
+
the repo.
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
# Token Optimizer Interoperability & Failure Mode Report
|
|
2
|
+
|
|
3
|
+
**How LLM Switcher acts as the protective outermost edge gateway for aggressive prompt/token optimizers (Headroom, RTK, Ponytail).**
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Executive Summary
|
|
8
|
+
|
|
9
|
+
Third-party prompt optimizers and token compressors — such as **Headroom**, **RTK (Rust Token Killer)**, and **Ponytail** — attempt to reduce LLM input tokens by aggressively pruning message history, truncating command stdout, or forcing extreme prompt brevity.
|
|
10
|
+
|
|
11
|
+
While these tools can reduce raw token counts in simple scenarios, **they frequently break complex agentic coding workflows** by corrupting message graphs, orphaning tool calls, and stripping reasoning parameters. When these pruned payloads hit upstream APIs directly (such as Anthropic, OpenAI, or 9Router), the provider immediately throws fatal `HTTP 400 Bad Request` errors or severely degrades reasoning depth.
|
|
12
|
+
|
|
13
|
+
**LLM Switcher solves this by acting as the outermost edge gatekeeper (`127.0.0.1:3456`).** It intercepts the pruned payload before it leaves your machine, runs its built-in **Healer Engine** to repair message graphs and restore reasoning parameters, unlocks 1M context windows, and safely converts the protocol to your upstream provider.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Tool Breakdown: What They Do & How They Break Payloads
|
|
18
|
+
|
|
19
|
+
### 1. Headroom (`headroomlabs-ai/headroom`)
|
|
20
|
+
- **Mechanism:** Runs as a local proxy on `:8787` (or wraps CLI agents). Compresses conversation history, RAG chunks, and tool outputs using SmartCrusher (JSON), CodeCompressor (AST), and Kompress-v2-base. Also attempts "effort routing" to dial down thinking budgets.
|
|
21
|
+
- **Critical Failure Points:**
|
|
22
|
+
- **Orphaned `tool_result` blocks:** When pruning historical turns, Headroom often discards the `assistant` turn containing a `tool_use`, while retaining the subsequent `user` turn containing the `tool_result`. Anthropic's API strictly validates tool use IDs and crashes with:
|
|
23
|
+
```
|
|
24
|
+
HTTP 400 invalid_request_error: "tool_use_id 'xxx' does not correspond to any tool_use"
|
|
25
|
+
```
|
|
26
|
+
- **Consecutive `user` turns:** Dropping intermediary assistant turns causes multiple user messages to sit adjacent to each other. Anthropic strictly throws:
|
|
27
|
+
```
|
|
28
|
+
HTTP 400 invalid_request_error: "roles must alternate between 'user' and 'assistant'"
|
|
29
|
+
```
|
|
30
|
+
- **Reasoning Suppression:** Its "effort routing" dials down `thinking.budget_tokens` on routine tool turns. On complex models (Claude Opus, Gemini Flash), this prevents the model from formulating multi-step reasoning before acting.
|
|
31
|
+
|
|
32
|
+
### 2. RTK (`rtk-ai/rtk` - Rust Token Killer)
|
|
33
|
+
- **Mechanism:** A single Rust binary that hooks into shell tool execution (e.g. `PreToolUse` in Claude Code / Cursor) and rewrites CLI commands (`git`, `ls`, `cat`, `grep`, `pytest`) to filter out noise, truncate lines, and inject recall tokens (`[full output: rtk recall xxx]`).
|
|
34
|
+
- **Critical Failure Points:**
|
|
35
|
+
- **Corrupted Structural Data:** When an agent invokes a tool expecting machine-readable JSON or exact AST formatting, RTK's heuristic summaries can alter structural delimiters, causing downstream tool call parsing errors.
|
|
36
|
+
- **Custom Tracking Headers:** RTK and associated tracing proxies inject headers (`x-rtk-*`, `traceparent`, `x-optimizer-id`) that some strict upstream endpoints reject if not cleanly forwarded.
|
|
37
|
+
|
|
38
|
+
### 3. Ponytail (`DietrichGebert/ponytail`)
|
|
39
|
+
- **Mechanism:** A behavioral prompt engineering plugin/ruleset that injects extreme conciseness instructions ("write one line, it works, YAGNI") into agent system prompts across 20+ coding tools.
|
|
40
|
+
- **Critical Failure Points:**
|
|
41
|
+
- **Premature Execution without Reasoning:** By commanding the model to be maximally brief and avoid planning, reasoning models are discouraged from spending thinking tokens. The model outputs untested single-liners that often fail type checks and test suites.
|
|
42
|
+
- **System Prompt Prefix Invalidation:** Injected rules alter the leading system prompt bytes, busting provider prompt caches unless carefully aligned.
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## The Healer Engine: How LLM Switcher Protects the Workflow
|
|
47
|
+
|
|
48
|
+
LLM Switcher sits between the optimizer tool and the upstream LLM:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
[CLI Agent] ──> [Optimizer: Headroom / RTK] ──> [LLM Switcher :3456] ──> [Upstream / 9Router]
|
|
52
|
+
│
|
|
53
|
+
├── 1. Heal Orphaned tool_results
|
|
54
|
+
├── 2. Merge Consecutive Turns
|
|
55
|
+
├── 3. Restore Stripped Thinking
|
|
56
|
+
└── 4. Enforce 1M Context
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
### Protection Matrix: Before vs. After
|
|
60
|
+
|
|
61
|
+
| Scenario | Direct to Upstream (Without Switcher) | Through LLM Switcher (Healer Engine) |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| **Orphaned `tool_result` turn** | ❌ **HTTP 400 Crash**: `tool_use_id does not correspond to any tool_use` | ✅ **HTTP 200 OK**: Heals orphaned result into contextual text block `[Tool Result (id)]: ...` |
|
|
64
|
+
| **Consecutive `user` turns** | ❌ **HTTP 400 Crash**: `roles must alternate` | ✅ **HTTP 200 OK**: Merges consecutive turns into a single valid turn seamlessly |
|
|
65
|
+
| **Stripped `thinking` parameters** | ⚠️ **Degraded AI**: Reasoning disabled, model outputs shallow single-liners | ✅ **HTTP 200 OK**: Detects reasoning models and automatically restores safe thinking budget |
|
|
66
|
+
| **Orphaned `tool` role in Chat API** | ❌ **HTTP 400 Crash**: `tool role must respond to tool_calls` | ✅ **HTTP 200 OK**: Converts orphaned tool message into user context |
|
|
67
|
+
| **Custom tracking headers** | ⚠️ Connection dropped / unrecognized header warnings | ✅ **HTTP 200 OK**: Cleanly passes through `traceparent`, `x-request-id`, `x-rtk-*` |
|
|
68
|
+
|
|
69
|
+
---
|
|
70
|
+
|
|
71
|
+
## Test Methodology & Verification Suite
|
|
72
|
+
|
|
73
|
+
We created an automated verification suite in [`tests/live-optimizer-interop.mjs`](../tests/live-optimizer-interop.mjs) that systematically replicates the failure modes of each tool:
|
|
74
|
+
|
|
75
|
+
### Running the Test Suite
|
|
76
|
+
```bash
|
|
77
|
+
node tests/live-optimizer-interop.mjs
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
### Live Test Results
|
|
81
|
+
```text
|
|
82
|
+
================================================================
|
|
83
|
+
LLM SWITCHER — TOKEN OPTIMIZER INTEROPERABILITY TEST SUITE
|
|
84
|
+
Simulating failure modes from Headroom, RTK, and Ponytail
|
|
85
|
+
================================================================
|
|
86
|
+
|
|
87
|
+
[TEST] Headroom Simulation: Orphaned tool_result turn... PASS (8840ms)
|
|
88
|
+
↳ Healed orphaned tool_result. Response HTTP 200: "It looks like you've shared a fragment of context ..."
|
|
89
|
+
[TEST] Headroom Simulation: Consecutive User turns (Role alternation violation)... PASS (2683ms)
|
|
90
|
+
↳ Merged consecutive turns seamlessly. Stop reason: end_turn
|
|
91
|
+
[TEST] Thinking Guard: Restoring stripped thinking parameter on reasoning models... PASS (5372ms)
|
|
92
|
+
↳ Automatically restored thinking: thinking_delta=true, signature_delta=true, text_delta=true
|
|
93
|
+
[TEST] RTK Intermediary: Custom headers and traceparent passthrough... PASS (2453ms)
|
|
94
|
+
↳ Headers accepted cleanly with HTTP 200 OK
|
|
95
|
+
[TEST] OpenAI Chat Healer: Orphaned role tool without preceding assistant tool_calls... PASS (1561ms)
|
|
96
|
+
↳ Chat Healer rescued orphaned tool role. Stop reason: length
|
|
97
|
+
|
|
98
|
+
================================================================
|
|
99
|
+
TEST RESULTS: 5 PASSED / 0 FAILED
|
|
100
|
+
================================================================
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## Recommended User Setup
|
|
106
|
+
|
|
107
|
+
For developers using prompt optimization tools:
|
|
108
|
+
1. Keep the optimizer installed in your CLI tool as usual.
|
|
109
|
+
2. In the optimizer's configuration (e.g. `headroom.yaml` or RTK upstream settings), set the upstream target URL to **LLM Switcher** (`http://127.0.0.1:3456`).
|
|
110
|
+
3. Enjoy prompt compression savings without worrying about broken conversation graphs, HTTP 400 crashes, or lost reasoning depth.
|
|
@@ -0,0 +1,214 @@
|
|
|
1
|
+
# Hide the gateway from the Codex CLI
|
|
2
|
+
|
|
3
|
+
## Diagrams
|
|
4
|
+
|
|
5
|
+
Open these next to the text. Each one is a standalone HTML page.
|
|
6
|
+
|
|
7
|
+
| Diagram | Shows |
|
|
8
|
+
| --- | --- |
|
|
9
|
+
| [Request routing](diagrams/blindfold-request-routing.html) | One Codex request, from CONNECT to the provider |
|
|
10
|
+
| [Model name resolution](diagrams/codex-model-name-resolution.html) | Which name the CLI sees, and where it resolves |
|
|
11
|
+
| [Lifecycle under switch](diagrams/blindfold-switch-lifecycle.html) | Activation, refusal, and shutdown |
|
|
12
|
+
|
|
13
|
+
The JSON source of each diagram sits beside its HTML file. Edit the JSON and rebuild with the `archify` tool; do not edit the HTML.
|
|
14
|
+
|
|
15
|
+
## What this does
|
|
16
|
+
|
|
17
|
+
The normal route sends Codex to the gateway with `--config openai_base_url=...`. That works, but the CLI then prints one line on its own `/model` screen:
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
base URL is overridden to http://127.0.0.1:3456/v1. Selecting models may not be supported or work properly.
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
No configuration key removes that line. The line exists to report an overridden base URL, so the only way to remove it is to stop overriding the base URL.
|
|
24
|
+
|
|
25
|
+
Blindfold mode does that. Codex keeps its official endpoint. The switcher intercepts one layer lower, at the network hop. The configuration file of Codex stays unchanged.
|
|
26
|
+
|
|
27
|
+
## How it works
|
|
28
|
+
|
|
29
|
+
Codex reads the `HTTPS_PROXY` variable. In blindfold mode the switcher points that variable at `blindfold/blindfold.mjs`. Codex then sends `CONNECT chatgpt.com:443` to that process.
|
|
30
|
+
|
|
31
|
+
The process answers the CONNECT itself. It ends the TLS session with a leaf certificate for the target host, and it forwards the Codex API calls to the gateway. Three rules decide where each request goes:
|
|
32
|
+
|
|
33
|
+
| Request | Destination |
|
|
34
|
+
| --- | --- |
|
|
35
|
+
| Target host, path under `/backend-api/codex/` | The local gateway |
|
|
36
|
+
| Target host, any other path | The real host, over a new TLS session |
|
|
37
|
+
| Another public host | A raw tunnel. The process never reads the bytes. |
|
|
38
|
+
| A local or private address | Refused. |
|
|
39
|
+
|
|
40
|
+
The routing decision reads the **normalized** path, not the text the client sent. Node hands over the request target exactly as written, but the gateway resolves it with `new URL(...)`. A raw-text test would therefore accept a string that the gateway later reads as a different path: `/backend-api/codex/%2e%2e/api/logs` becomes `/api/logs`, which is the gateway's admin API. The decision also requires a segment boundary, so `/backend-api/codex-usage` stays with the host it belongs to.
|
|
41
|
+
|
|
42
|
+
Sign-in, token refresh and the usage page keep working, because they do not use the Codex API path.
|
|
43
|
+
|
|
44
|
+
## Why no system change is necessary
|
|
45
|
+
|
|
46
|
+
Codex reads a custom certificate authority from the `CODEX_CA_CERTIFICATE` variable. The private CA stays in the repository. Three things that other interception guides ask for are not necessary here:
|
|
47
|
+
|
|
48
|
+
- No certificate in the Windows or macOS trust store.
|
|
49
|
+
- No line in the `hosts` file, so no administrator rights.
|
|
50
|
+
- No change to `~/.codex/config.toml`.
|
|
51
|
+
|
|
52
|
+
To stop the interception, run `switch off`. It stops the interceptor and removes the launch files (`active.flag`, the `env*` files and the model catalog). These items stay until you remove them:
|
|
53
|
+
|
|
54
|
+
- The certificates in `blindfold/certs/`, including `ca.key`. Delete the directory when you do not use blindfold mode again.
|
|
55
|
+
- `blindfold.log` and `proxy.log`. Delete them by hand.
|
|
56
|
+
- The shims in `~/.llm-switcher/bin`. Run `switch shim uninstall` to remove them.
|
|
57
|
+
|
|
58
|
+
## Before you start
|
|
59
|
+
|
|
60
|
+
You need:
|
|
61
|
+
|
|
62
|
+
- Node.js 18.17 or later.
|
|
63
|
+
- `openssl`. Windows users have it with Git for Windows. Run the script from Git Bash.
|
|
64
|
+
- A gateway profile for Codex that already works in the normal route.
|
|
65
|
+
|
|
66
|
+
## Procedure
|
|
67
|
+
|
|
68
|
+
### 1. Build the certificates
|
|
69
|
+
|
|
70
|
+
Run this command in the repository root:
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
bash blindfold/make-certs.sh chatgpt.com
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
The script writes `blindfold/certs/ca.pem` and `blindfold/certs/leaf.pem`. It prints the extended key usage and the subject alternative name of the leaf. Make sure that the output contains `TLS Web Server Authentication`.
|
|
77
|
+
|
|
78
|
+
CAUTION: Do not build these certificates with `New-SelfSignedCertificate` in PowerShell. That command writes a leaf without the `serverAuth` extended key usage, and a strict client refuses it with `unsuitable certificate purpose`.
|
|
79
|
+
|
|
80
|
+
`blindfold/certs/` is in `.gitignore`. The private keys never enter the repository history.
|
|
81
|
+
|
|
82
|
+
### 2. Turn on blindfold mode in the profile
|
|
83
|
+
|
|
84
|
+
Add two keys to the Codex profile in `config.json`:
|
|
85
|
+
|
|
86
|
+
```json
|
|
87
|
+
"blindfold": true,
|
|
88
|
+
"blindfoldPort": 3457
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
`blindfoldPort` is optional. The default is 3457.
|
|
92
|
+
|
|
93
|
+
### 3. Activate the profile
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
switch codex <your-profile>
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
This one command starts the gateway and writes the environment files. The gateway then starts the interceptor on the configured port. `switch off` stops both and removes the generated files.
|
|
100
|
+
|
|
101
|
+
The gateway owns the interceptor. It brings the interceptor in line with `config.json` when it starts, after every change in the dashboard, and after every `switch` command. A service start, a `switch port`, and a change of host, prefix or port in the dashboard therefore never leave `HTTPS_PROXY` pointing at a port where nothing listens. When blindfold mode is turned off for the Codex target, the gateway stops the interceptor, and a running Codex session must be restarted.
|
|
102
|
+
|
|
103
|
+
`switch` refuses the activation and writes no file in these cases:
|
|
104
|
+
|
|
105
|
+
- The CA, the leaf certificate or the leaf key is missing, or the leaf does not cover the host.
|
|
106
|
+
- The leaf was not signed by the CA, or the leaf key does not match the leaf. This occurs when the files come from two different builds.
|
|
107
|
+
- Another process holds the interceptor port or the gateway port.
|
|
108
|
+
|
|
109
|
+
That refusal is deliberate: blindfold mode replaces the base URL override with `HTTPS_PROXY`, so a half-applied state would point Codex at a port where nothing listens, and Codex would then reach no host at all. If the interceptor fails to start after these checks pass, `switch` restores the previous `config.json` and launcher files and exits with an error.
|
|
110
|
+
|
|
111
|
+
The switcher trusts a port only after the process on it proves its identity with the token in `admin.token`. A process that only answers on the port, or replays an earlier answer, is treated as foreign.
|
|
112
|
+
|
|
113
|
+
The environment now holds `HTTPS_PROXY`, `NO_PROXY` and `CODEX_CA_CERTIFICATE` in `env-codex.cmd` and `env-codex.sh`. Only the Codex shim loads those two files. They are separate from `env.cmd` and `env.sh` on purpose: the `claude` shim loads the shared files, and a Codex-only proxy would otherwise capture every `claude` HTTPS call.
|
|
114
|
+
|
|
115
|
+
The shared files no longer hold `LLM_SWITCHER_CODEX_BASE_URL`, so the shim adds no base URL override.
|
|
116
|
+
|
|
117
|
+
### 4. Verify the route
|
|
118
|
+
|
|
119
|
+
Run Codex and open `/model`. The screen shows the official model names, and the override line is gone.
|
|
120
|
+
|
|
121
|
+
To test without Codex, ask the gateway for its model list through the proxy:
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
node -e "
|
|
125
|
+
const http=require('http'),tls=require('tls'),fs=require('fs');
|
|
126
|
+
const ca=fs.readFileSync('blindfold/certs/ca.pem');
|
|
127
|
+
const r=http.request({host:'127.0.0.1',port:3457,method:'CONNECT',path:'chatgpt.com:443'});
|
|
128
|
+
r.end();
|
|
129
|
+
r.on('connect',(res,socket)=>{
|
|
130
|
+
const s=tls.connect({socket,servername:'chatgpt.com',ca},()=>{
|
|
131
|
+
s.write('GET /backend-api/codex/models HTTP/1.1\r\nHost: chatgpt.com\r\nConnection: close\r\n\r\n');
|
|
132
|
+
});
|
|
133
|
+
s.on('data',d=>process.stdout.write(d));
|
|
134
|
+
});
|
|
135
|
+
"
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
The answer is the model list of the gateway. Change the path to `/backend-api/codex/%2e%2e/api/logs` and the answer comes from chatgpt.com instead, which proves the normalization.
|
|
139
|
+
|
|
140
|
+
NOTE: On Windows, `curl` uses the Schannel TLS backend. Schannel ignores `--cacert` for this test and reports error 60. Use the Node command above instead.
|
|
141
|
+
|
|
142
|
+
## Record what a client sends
|
|
143
|
+
|
|
144
|
+
Add `--capture <dir>` to write one JSON file per intercepted exchange:
|
|
145
|
+
|
|
146
|
+
```bash
|
|
147
|
+
node blindfold/blindfold.mjs --capture ./captures
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Each file holds the method, the URL, both header sets and both bodies, truncated
|
|
151
|
+
at 200000 characters. Use it to learn what a genuine CLI or IDE puts on the wire.
|
|
152
|
+
|
|
153
|
+
Both routes are recorded: the calls that go to the gateway, and the calls that are
|
|
154
|
+
re-originated to the real host. A compressed body is decoded first, because a
|
|
155
|
+
client asks for `gzip` and the bytes on the wire are not readable text. The
|
|
156
|
+
forwarded response keeps its original bytes; only the copy in the file is decoded.
|
|
157
|
+
|
|
158
|
+
To record a different tool, point the proxy at that tool's host and give it a
|
|
159
|
+
prefix that no path can match, so every request is re-originated and recorded:
|
|
160
|
+
|
|
161
|
+
```bash
|
|
162
|
+
bash blindfold/make-certs.sh api.anthropic.com ~/.llm-switcher/anthropic/certs
|
|
163
|
+
node blindfold/blindfold.mjs --host api.anthropic.com --prefix /no-gateway \
|
|
164
|
+
--port 3458 --certs ~/.llm-switcher/anthropic/certs --capture ~/.llm-switcher/anthropic/captures
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
Use directories that you own. `make-certs.sh` and `--capture` refuse a directory that another
|
|
168
|
+
account owns, because that account could read the key or the captures.
|
|
169
|
+
|
|
170
|
+
Then start the tool with `HTTPS_PROXY=http://127.0.0.1:3458` and the CA in the
|
|
171
|
+
variable that the tool reads. A Node client reads `NODE_EXTRA_CA_CERTS`. Set both
|
|
172
|
+
variables for that process only; a variable set for one process does not change a
|
|
173
|
+
process that already runs.
|
|
174
|
+
|
|
175
|
+
Credential headers never reach the file. `authorization`, `proxy-authorization`,
|
|
176
|
+
`cookie`, `set-cookie`, `x-api-key` and `api-key` are replaced with `<redacted>`,
|
|
177
|
+
and so are the account identifiers `chatgpt-account-id`, `openai-organization` and
|
|
178
|
+
`x-goog-user-project`. The header names stay, so the shape of the request is still
|
|
179
|
+
readable.
|
|
180
|
+
|
|
181
|
+
CAUTION: The redaction covers headers only. Both bodies are written as they
|
|
182
|
+
travelled, so a capture holds your prompts, your source code and the answers of
|
|
183
|
+
the model. A body can also hold a secret, for example a token that a tool printed.
|
|
184
|
+
`captures/` is in `.gitignore`. If you capture somewhere else, add that
|
|
185
|
+
path to `.gitignore` before you commit. Delete a capture when you do not need it.
|
|
186
|
+
|
|
187
|
+
The proxy creates the capture directory with mode 0700 and each file with mode 0600.
|
|
188
|
+
|
|
189
|
+
NOTE: WebSocket sessions are recorded too. Each session gives one file, with every
|
|
190
|
+
message in order, and the proxy updates the file while the session runs. Codex sends
|
|
191
|
+
its completions over a WebSocket, so these files hold the full conversation.
|
|
192
|
+
|
|
193
|
+
## How to go back
|
|
194
|
+
|
|
195
|
+
1. Set `"blindfold": false` in the profile.
|
|
196
|
+
2. Run `switch codex <your-profile>` again.
|
|
197
|
+
|
|
198
|
+
The switcher writes `LLM_SWITCHER_CODEX_BASE_URL` again, the gateway stops the interceptor, and Codex returns to the normal route.
|
|
199
|
+
|
|
200
|
+
## Limits
|
|
201
|
+
|
|
202
|
+
Read this section before you turn blindfold mode on.
|
|
203
|
+
|
|
204
|
+
**The CA is trusted for every host that Codex contacts.** `CODEX_CA_CERTIFICATE` adds this authority to the trust store that Codex uses for all of its HTTPS calls. A CA that `make-certs.sh` builds now carries a name constraint: it can sign only for the target host and its subdomains. A client that obeys name constraints refuses any other certificate from this CA. OpenSSL, rustls and the macOS and Windows verifiers obey them. A CA that an older version built has no constraint. Anybody who can read its `ca.key` can forge a certificate for any host. Run `make-certs.sh` again to replace it, and keep `blindfold/certs/` private in both cases.
|
|
205
|
+
|
|
206
|
+
**All traffic to the intercepted host is decrypted by this process.** That includes sign-in and token refresh, because the CONNECT for the whole host is terminated locally. Requests outside the Codex API path are re-originated to the real host over a new TLS session; they are forwarded, not tunneled. The `--verbose` flag prints the method and path of every such request.
|
|
207
|
+
|
|
208
|
+
**Directory permissions are weaker on Windows.** `make-certs.sh` calls `chmod` on the certificate directory, and that call does nothing under Git Bash for Windows. On that platform the keys are protected only by the account that owns the directory.
|
|
209
|
+
|
|
210
|
+
**Only one host is intercepted.** If your Codex uses API key authentication instead of ChatGPT authentication, the host is `api.openai.com`. Build the leaf for that host, and start the process with `--host api.openai.com --prefix /v1`.
|
|
211
|
+
|
|
212
|
+
**The interceptor is a proxy.** It binds to loopback, so a remote machine cannot use it, but every process on this machine can. It refuses a CONNECT to a local or private address, so it cannot be used to reach a service that listens only on this machine. The interceptor resolves the name first and checks every address in the answer, so a spelling such as `2130706433` or a name that resolves to `127.0.0.1` is also refused. It then connects to the address that it checked. It prints each refused target once, without `--verbose`. A VPN or split DNS can resolve a public name to a private address, and that message shows the cause.
|
|
213
|
+
|
|
214
|
+
**Some `--config` keys still travel.** Blindfold mode removes the base URL override only. The switcher still passes `model_catalog_json`, the three role names (`model`, `review_model`, `agents.default_subagent_model`) and the two context window keys, because those carry no address and no internal name.
|
|
@@ -0,0 +1,136 @@
|
|
|
1
|
+
# Run this repository on Linux, macOS and Windows
|
|
2
|
+
|
|
3
|
+
The gateway is plain Node.js with no dependencies. Nothing in the source holds an
|
|
4
|
+
absolute Windows path. Every command that differs between platforms has a branch for
|
|
5
|
+
each one. This guide states what to install, what changes per platform, and which two
|
|
6
|
+
features are not available everywhere.
|
|
7
|
+
|
|
8
|
+
## What you need
|
|
9
|
+
|
|
10
|
+
| Item | Why |
|
|
11
|
+
| --- | --- |
|
|
12
|
+
| Node.js 18.17 or later | The gateway, the launcher and the interceptor |
|
|
13
|
+
| `bash` and `openssl` | Only for blindfold mode, to build the certificates |
|
|
14
|
+
| Nothing else | `npm install` is not necessary. The package has no dependencies. |
|
|
15
|
+
|
|
16
|
+
Linux and macOS have `bash` and `openssl` already. On Windows both arrive with Git for
|
|
17
|
+
Windows, inside Git Bash.
|
|
18
|
+
|
|
19
|
+
## Install
|
|
20
|
+
|
|
21
|
+
The steps are the same on all three platforms.
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
git clone https://github.com/louisphamdev/llm-switcher.git
|
|
25
|
+
cd llm-switcher
|
|
26
|
+
cp config.example.json config.json
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Edit `config.json`, then start the gateway:
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
node switch.mjs on
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
## Put `switch` on PATH
|
|
36
|
+
|
|
37
|
+
The repository holds two launchers. Both run the same `switch.mjs`.
|
|
38
|
+
|
|
39
|
+
| Platform | File | How to reach it |
|
|
40
|
+
| --- | --- | --- |
|
|
41
|
+
| Linux, macOS | `switch` | `export PATH="/path/to/llm-switcher:$PATH"` in `~/.bashrc` or `~/.zshrc` |
|
|
42
|
+
| Windows | `switch.cmd` | Add the repository directory to the user PATH |
|
|
43
|
+
|
|
44
|
+
After that, `switch status`, `switch codex <profile>` and `switch off` behave the same
|
|
45
|
+
everywhere. If you prefer not to change PATH, run `node switch.mjs <command>` instead.
|
|
46
|
+
|
|
47
|
+
If `switch` reports "permission denied" on Linux or macOS, the executable bit was lost
|
|
48
|
+
in transfer. Restore it:
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
chmod +x switch
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
## What the switcher writes for each platform
|
|
55
|
+
|
|
56
|
+
`switch` writes both forms every time, so the same working copy serves a shell and a
|
|
57
|
+
command prompt:
|
|
58
|
+
|
|
59
|
+
| File | Read by |
|
|
60
|
+
| --- | --- |
|
|
61
|
+
| `env.sh`, `env-codex.sh` | The POSIX shims and any shell |
|
|
62
|
+
| `env.cmd`, `env-codex.cmd` | The Windows shims, Command Prompt and PowerShell |
|
|
63
|
+
|
|
64
|
+
The CLI shims follow the same rule. `switch shim install` writes
|
|
65
|
+
`~/.llm-switcher/bin/claude` and `~/.llm-switcher/bin/codex` on Linux and macOS, with
|
|
66
|
+
the executable bit set. On Windows it writes `claude.cmd` and `codex.cmd` in the same
|
|
67
|
+
directory. Put that directory before the real CLI in PATH:
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
export PATH="$HOME/.llm-switcher/bin:$PATH" # Linux, macOS
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
%USERPROFILE%\.llm-switcher\bin # Windows: add before the Codex directory
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
## Commands that differ inside the code
|
|
78
|
+
|
|
79
|
+
You do not call these yourself. They are listed so that a reader knows the platform
|
|
80
|
+
work is already done.
|
|
81
|
+
|
|
82
|
+
| Task | Linux, macOS | Windows |
|
|
83
|
+
| --- | --- | --- |
|
|
84
|
+
| Find the process on a port | `lsof` | `netstat -ano` |
|
|
85
|
+
| Stop a process | `process.kill` | `taskkill /F /PID` |
|
|
86
|
+
| Open the dashboard | `xdg-open`, `open` | `cmd /c start` |
|
|
87
|
+
|
|
88
|
+
## Two features are not available everywhere
|
|
89
|
+
|
|
90
|
+
**`switch doctor` cannot inspect running CLI sessions on Windows.** That check reads the
|
|
91
|
+
environment of a live process with `pgrep` and `ps`, which exist only on Linux and
|
|
92
|
+
macOS. On Windows the rest of `switch doctor` still runs, and this one section reports
|
|
93
|
+
that it is not supported. A Windows user who wants the same answer must close the CLI
|
|
94
|
+
and start it again from a shell where the shim is on PATH.
|
|
95
|
+
|
|
96
|
+
**`make-certs.sh` cannot protect the key directory on Windows.** The script calls
|
|
97
|
+
`chmod 700` on `blindfold/certs/`, and that call does nothing under Git Bash for
|
|
98
|
+
Windows. On that platform the private keys are protected only by the permissions of the
|
|
99
|
+
account that owns the directory. Keep the repository in a user-owned location.
|
|
100
|
+
|
|
101
|
+
Four tests in `tests/shim.test.mjs` run the generated POSIX shim as a real program, so
|
|
102
|
+
they skip on Windows. The suite reports them as skipped, not as failures.
|
|
103
|
+
|
|
104
|
+
## Blindfold mode on each platform
|
|
105
|
+
|
|
106
|
+
Blindfold mode works on all three platforms, because it needs no administrator rights,
|
|
107
|
+
no `hosts` file change and no certificate in a system trust store. See
|
|
108
|
+
[`codex-blindfold.md`](codex-blindfold.md) for what it does.
|
|
109
|
+
|
|
110
|
+
Build the certificates once. On Windows, run this line in Git Bash, not in PowerShell:
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
bash blindfold/make-certs.sh chatgpt.com
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
The subject of the certificate must match the host your Codex account calls:
|
|
117
|
+
|
|
118
|
+
| Codex sign-in | Host | Profile keys |
|
|
119
|
+
| --- | --- | --- |
|
|
120
|
+
| ChatGPT account | `chatgpt.com` | The defaults. Add nothing. |
|
|
121
|
+
| API key | `api.openai.com` | `"blindfoldHost": "api.openai.com"`, `"blindfoldPrefix": "/v1"` |
|
|
122
|
+
|
|
123
|
+
`switch` compares the host in the profile against the subject alternative names of the
|
|
124
|
+
leaf certificate. If they differ it refuses the activation, names the host, and prints
|
|
125
|
+
the command that rebuilds the certificate. No file is written on that path.
|
|
126
|
+
|
|
127
|
+
## Check that the platform work succeeded
|
|
128
|
+
|
|
129
|
+
Run the test suite:
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
npm test
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
On Linux and macOS every test runs. On Windows four shim tests skip. A failure on any
|
|
136
|
+
platform is a real defect; report it with the platform name and the Node version.
|