llm-switcher 1.2.0 → 1.2.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/CHANGELOG.md +23 -0
  2. package/README.md +1 -0
  3. package/blindfold/make-certs.sh +3 -1
  4. package/catalog.mjs +132 -42
  5. package/classifier.mjs +238 -0
  6. package/docs/TOKEN-OPTIMIZER-INTEROP.md +110 -110
  7. package/docs/response-matrix.json +1130 -1130
  8. package/formats.mjs +30 -1
  9. package/icons/antigravity.png +0 -0
  10. package/icons/claude.png +0 -0
  11. package/icons/codex.png +0 -0
  12. package/icons/deepseek.png +0 -0
  13. package/icons/gemini.png +0 -0
  14. package/icons/github.png +0 -0
  15. package/icons/groq.png +0 -0
  16. package/icons/intact.svg +1 -0
  17. package/icons/ollama.png +0 -0
  18. package/icons/openai.png +0 -0
  19. package/icons/openrouter.png +0 -0
  20. package/icons/qwen.png +0 -0
  21. package/icons/vertex.png +0 -0
  22. package/mcp.mjs +2 -2
  23. package/package.json +1 -1
  24. package/proxy.mjs +34 -20
  25. package/scripts/run-tests.mjs +3 -1
  26. package/skills/llm-switcher/SKILL.md +93 -93
  27. package/state.mjs +58 -4
  28. package/switch +0 -0
  29. package/switch.mjs +28 -4
  30. package/tests/blindfold-e2e.test.mjs +2 -2
  31. package/tests/blindfold-task5.test.mjs +1 -1
  32. package/tests/blindfold.task3.test.mjs +2 -2
  33. package/tests/blindfold.wire.test.mjs +3 -1
  34. package/tests/catalog.test.mjs +117 -19
  35. package/tests/classifier.test.mjs +210 -0
  36. package/tests/codex-daemon.test.mjs +92 -0
  37. package/tests/formats.test.mjs +30 -0
  38. package/tests/helpers.mjs +24 -24
  39. package/tests/lifecycle.test.mjs +49 -37
  40. package/tests/live-optimizer-interop.mjs +205 -205
  41. package/tests/make-certs.test.mjs +22 -0
  42. package/tests/real-user-sim.test.mjs +3 -0
  43. package/tests/shim.test.mjs +8 -3
  44. package/tests/switch.test.mjs +30 -24
  45. package/tests/ui.test.mjs +90 -0
  46. package/tests/version.test.mjs +78 -0
  47. package/ui.html +1895 -1616
  48. package/version.mjs +53 -0
package/CHANGELOG.md CHANGED
@@ -1,5 +1,28 @@
1
1
  # Changelog — LLM Switcher
2
2
 
3
+ ## Release 1.2.3
4
+
5
+ - **Update notice:** The dashboard sidebar shows the running version. When npm has a newer release, a notice shows the update command with a copy button. The new command `switch version` prints the same information. The switcher asks the npm registry at most once in 12 hours. Without an answer, no notice shows.
6
+ - **Codex daemon follows the route:** The interactive Codex TUI talks to a shared `codex app-server` daemon, and that daemon keeps the environment that it started with. As a result, the TUI bypassed the gateway after a switch. Now each change of the Codex route restarts the running daemon with the new variables. When no daemon runs, nothing starts.
7
+ - **API keys stay with their host:** A catalog refresh sent the key of the active profile to the official Anthropic and OpenAI model lists. Now a key goes only to the host of its own profile.
8
+ - **Real model catalog:** The catalog reads the model list that Codex and Claude Code keep on disk for the signed-in account. A failed refresh no longer marks the tool version as done, so the next request tries again.
9
+ - **Dashboard saves:** Every write action reported a failure after it succeeded, because the page read `ok` and the server sends `success`. Enter in a field no longer saves a half-edited profile. A text selection that ends outside the dialog no longer closes it.
10
+ - **Model slots:** Each slot is a searchable list of the provider models. The list loads when you open the Model Slots tab and sends the saved key. A profile without `outFormat` keeps it empty after a save.
11
+ - **Other fixes:** `make-certs.sh` works when the path contains a dot. `switch doctor` warns again about a Codex profile without `publicModels`. Three dashboard controls have an accessible name.
12
+ - **Tests:** The tests use the 1.2 configuration schema, never read the certificates of the checkout, and never reach the real Codex daemon, npm registry, or `claude` binary.
13
+
14
+ ## Release 1.2.2
15
+
16
+ - **Structured output reaches every upstream:** A request that asks for JSON that matches a schema now keeps that schema. Before this release, the gateway did not read `output_config.format` (Claude Code) or `text.format` (Codex), so the provider got a free-text request. The schema now goes to the provider as `response_format` (OpenAI Chat), `output_config.format` (Anthropic), or `responseMimeType` with `responseSchema` (Gemini and Vertex). A request for JSON without a schema reaches Anthropic as plain text, because Anthropic has no JSON mode without a schema.
17
+
18
+ ## Release 1.2.1
19
+
20
+ - **Intact Visual Architecture & Hash Routing:** Fully restructured dashboard using Intact design system, pure CSS tokens, and hash-based client routing (`#/routes`, `#/models`, `#/logs`, `#/doctor`).
21
+ - **Official Brand Icon Assets:** Integrated official brand icons for Anthropic Claude, OpenAI Codex, Google Gemini, Intact, OpenRouter, and Ollama, served statically under `/icons/*` with immutable caching and PNG/SVG whitelisting.
22
+ - **Prefetched Navigation Counters:** Navigation badges prefetch model discovery and request inspector counts immediately on page load, eliminating the delay where badges showed zero until the tab was selected.
23
+ - **Hardened Profile Slot Mapping:** Synchronized Claude model tiers (`sonnet`, `opus`, `haiku`, `fable`) and Codex model roles (`main`, `review`, `subagent`), ensuring slot persistence when editing profiles and removing redundant role tabs for Claude-targeted profiles.
24
+ - **CORS & Endpoint Fixes:** Added `x-llm-switcher-token` to preflight `Access-Control-Allow-Headers` and resolved routing precedence for `GET /api/catalog`.
25
+
3
26
  ## Release 1.2.0
4
27
 
5
28
  - **Zero-Mutation Interceptor Invariant:** The gateway now operates strictly in the network path via the blindfold interceptor, never modifying user configuration files like `~/.claude/settings.json` or `~/.codex/config.toml`. Loopback base URLs (`ANTHROPIC_BASE_URL` and `OPENAI_BASE_URL`) are scrubbed by shims so traffic routes through `HTTPS_PROXY` cleanly.
package/README.md CHANGED
@@ -401,6 +401,7 @@ Add the server to your MCP configuration (for example `opencode.jsonc`, `claude_
401
401
  ```bash
402
402
  switch ui # Open the Web UI dashboard in your browser
403
403
  switch status # Display status for all active CLI targets
404
+ switch version # Show the version and tell you when npm has a newer one
404
405
  switch doctor # Audit environment, settings & routing
405
406
  switch on [profile] # Start the gateway and activate a profile
406
407
  switch <profile> # Activate a profile for both tools
@@ -103,8 +103,10 @@ openssl req -x509 -new -key "$WORK/ca.key" -sha256 -days "$CA_DAYS" \
103
103
 
104
104
  openssl ecparam -name prime256v1 -genkey -noout -out "$WORK/leaf.key"
105
105
  openssl req -new -key "$WORK/leaf.key" -config "$WORK/leaf.cnf" -out "$WORK/leaf.csr"
106
+ # An explicit -CAserial: LibreSSL derives the default name by cutting the CA path at its first dot,
107
+ # so a home directory such as /Users/first.last sent the serial file to /Users/first.srl.
106
108
  openssl x509 -req -in "$WORK/leaf.csr" \
107
- -CA "$WORK/ca.pem" -CAkey "$WORK/ca.key" -CAcreateserial \
109
+ -CA "$WORK/ca.pem" -CAkey "$WORK/ca.key" -CAcreateserial -CAserial "$WORK/ca.srl" \
108
110
  -days "$LEAF_DAYS" -sha256 -extfile "$WORK/leaf.ext" \
109
111
  -out "$WORK/leaf.pem"
110
112
 
package/catalog.mjs CHANGED
@@ -7,6 +7,7 @@
7
7
  // ============================================================
8
8
 
9
9
  import fs from 'node:fs';
10
+ import os from 'node:os';
10
11
  import path from 'node:path';
11
12
 
12
13
  export const OFFICIAL_MODEL_URLS = {
@@ -69,6 +70,70 @@ export function loadCatalogCache(stateDir) {
69
70
  };
70
71
  }
71
72
 
73
+ function defaultSources() {
74
+ return {
75
+ codexHome: process.env.CODEX_HOME || path.join(os.homedir(), '.codex'),
76
+ claudeDir: process.env.CLAUDE_CONFIG_DIR || path.join(os.homedir(), '.claude')
77
+ };
78
+ }
79
+
80
+ /**
81
+ * Reads the model list that the tool itself keeps on disk for the signed-in account, so no key and
82
+ * no network call is needed. Codex writes models_cache.json; Claude Code writes cache/model-catalog.
83
+ * Returns null when the tool has no list.
84
+ */
85
+ export function readLocalToolModels(tool, sources = defaultSources()) {
86
+ try {
87
+ if (tool === 'codex') {
88
+ const d = JSON.parse(fs.readFileSync(path.join(sources.codexHome, 'models_cache.json'), 'utf8'));
89
+ const models = (Array.isArray(d?.models) ? d.models : [])
90
+ .filter(m => m && typeof m.slug === 'string' && m.slug)
91
+ .map(m => ({ id: m.slug, role: classifyCodexRole(m.slug), ...(Number.isInteger(m.context_window) ? { contextWindow: m.context_window } : {}) }));
92
+ return models.length ? { models, version: String(d.client_version || '') } : null;
93
+ }
94
+ if (tool === 'claude') {
95
+ // One file per account or surface; the newest fetch wins.
96
+ const dir = path.join(sources.claudeDir, 'cache', 'model-catalog');
97
+ let newest = null;
98
+ for (const f of fs.readdirSync(dir)) {
99
+ if (!f.endsWith('.json')) continue;
100
+ try {
101
+ const d = JSON.parse(fs.readFileSync(path.join(dir, f), 'utf8'));
102
+ const list = d?.catalog?.config?.models;
103
+ if (Array.isArray(list) && list.length && (!newest || (d.fetchedAt || 0) > newest.fetchedAt)) newest = { fetchedAt: d.fetchedAt || 0, list };
104
+ } catch {}
105
+ }
106
+ const models = (newest?.list || [])
107
+ .filter(m => m && typeof m.id === 'string' && m.id)
108
+ .map(m => ({ id: m.id, display_name: m.name || m.id, tier: classifyClaudeTier(m.id) }));
109
+ return models.length ? { models, version: '' } : null;
110
+ }
111
+ } catch {}
112
+ return null;
113
+ }
114
+
115
+ /**
116
+ * Copies the tools' own lists into the catalog when they changed. A tool can update and rewrite its
117
+ * list without any request through the gateway, so readers of the catalog call this first.
118
+ */
119
+ export function syncLocalCatalog(stateDir, sources = defaultSources()) {
120
+ const cache = loadCatalogCache(stateDir);
121
+ let changed = false;
122
+ for (const tool of ['claude', 'codex']) {
123
+ const local = readLocalToolModels(tool, sources);
124
+ if (!local) continue;
125
+ const ids = (cache[tool]?.models || []).map(m => m.id).join('\n');
126
+ if (ids === local.models.map(m => m.id).join('\n') && cache[tool]?.source === 'local') continue;
127
+ cache[tool] = { ...cache[tool], models: local.models, source: 'local' };
128
+ changed = true;
129
+ }
130
+ if (changed) {
131
+ cache.updatedAt = Date.now();
132
+ saveCatalogCache(stateDir, cache);
133
+ }
134
+ return cache;
135
+ }
136
+
72
137
  /** Atomically writes the model catalog cache to disk */
73
138
  export function saveCatalogCache(stateDir, catalog) {
74
139
  const p = catalogCachePath(stateDir);
@@ -141,19 +206,33 @@ export async function fetchToolModels(tool, { url, apiKey, timeout = 3000 } = {}
141
206
  }
142
207
  }
143
208
 
209
+ // A profile key is a credential for that profile's baseURL. Send it to the official model list only
210
+ // when the profile itself points at the official host, never to a third party.
211
+ function officialKey(profile, tool) {
212
+ try {
213
+ if (profile?.apiKey && new URL(profile.baseURL).host === new URL(OFFICIAL_MODEL_URLS[tool]).host) return profile.apiKey;
214
+ } catch {}
215
+ return undefined;
216
+ }
217
+
144
218
  /**
145
219
  * Refreshes the local catalog cache with models from both official endpoints
146
220
  * and saves to the state directory.
147
221
  */
148
- export async function refreshCatalog(stateDir, { claudeKey, codexKey } = {}) {
222
+ export async function refreshCatalog(stateDir, { claudeProfile, codexProfile, sources = defaultSources() } = {}) {
149
223
  const cache = loadCatalogCache(stateDir);
224
+ const local = { claude: readLocalToolModels('claude', sources), codex: readLocalToolModels('codex', sources) };
150
225
 
226
+ // A local list is the list of the account itself; the built-in names are guesses and are not added.
151
227
  const [claudeRes, codexRes] = await Promise.all([
152
- fetchToolModels('claude', { apiKey: claudeKey }),
153
- fetchToolModels('codex', { apiKey: codexKey })
228
+ local.claude || fetchToolModels('claude', { apiKey: officialKey(claudeProfile, 'claude') }),
229
+ local.codex || fetchToolModels('codex', { apiKey: officialKey(codexProfile, 'codex') })
154
230
  ]);
231
+ for (const tool of ['claude', 'codex']) {
232
+ if (local[tool]) cache[tool] = { ...cache[tool], models: local[tool].models, source: 'local' };
233
+ }
155
234
 
156
- if (claudeRes.ok && claudeRes.models.length > 0) {
235
+ if (!local.claude && claudeRes.ok && claudeRes.models.length > 0) {
157
236
  const existing = new Set(claudeRes.models.map(m => m.id));
158
237
  // Keep any baseline models that might be absent from the API
159
238
  for (const [tier, ids] of Object.entries(BASELINE_MODELS.claude)) {
@@ -161,17 +240,17 @@ export async function refreshCatalog(stateDir, { claudeKey, codexKey } = {}) {
161
240
  if (!existing.has(id)) claudeRes.models.push({ id, tier });
162
241
  }
163
242
  }
164
- cache.claude = { models: claudeRes.models };
243
+ cache.claude = { ...cache.claude, models: claudeRes.models, source: 'official' };
165
244
  }
166
245
 
167
- if (codexRes.ok && codexRes.models.length > 0) {
246
+ if (!local.codex && codexRes.ok && codexRes.models.length > 0) {
168
247
  const existing = new Set(codexRes.models.map(m => m.id));
169
248
  for (const [role, ids] of Object.entries(BASELINE_MODELS.codex)) {
170
249
  for (const id of ids) {
171
250
  if (!existing.has(id)) codexRes.models.push({ id, role });
172
251
  }
173
252
  }
174
- cache.codex = { models: codexRes.models };
253
+ cache.codex = { ...cache.codex, models: codexRes.models, source: 'official' };
175
254
  }
176
255
 
177
256
  cache.updatedAt = Date.now();
@@ -199,48 +278,59 @@ const refreshingTools = new Set();
199
278
 
200
279
  /**
201
280
  * Version-triggered auto-poll:
202
- * When a request arrives with a new tool version not yet seen in cache,
203
- * immediately records the new version and kicks off an asynchronous background refresh.
204
- * Subsequent requests with the same version do zero network calls.
281
+ * When a request arrives with a tool version that has no refreshed catalog yet, the catalog is
282
+ * refreshed from the tool's own list on disk, or else from the official list in the background.
283
+ * The version is recorded only after a refresh succeeds, so a failed refresh is tried again.
205
284
  */
206
- export function checkVersionAndRefresh(tool, headers, stateDir, apiKey) {
285
+ const RETRY_AFTER_MS = 10 * 60 * 1000;
286
+ const failedAttempts = new Map();
287
+
288
+ export function checkVersionAndRefresh(tool, headers, stateDir, apiKey, sources = defaultSources()) {
207
289
  const version = detectToolVersion(headers, tool);
208
290
  if (!version) return;
209
291
 
210
292
  const cache = loadCatalogCache(stateDir);
211
- const toolEntry = cache[tool] || {};
212
- const lastVersion = toolEntry.lastSeenVersion || '';
293
+ if (version === (cache[tool]?.lastSeenVersion || '')) return;
213
294
 
214
- if (version !== lastVersion) {
215
- toolEntry.lastSeenVersion = version;
216
- cache[tool] = toolEntry;
295
+ const local = readLocalToolModels(tool, sources);
296
+ if (local) {
297
+ cache[tool] = { ...cache[tool], lastSeenVersion: version, models: local.models, source: 'local' };
298
+ cache.updatedAt = Date.now();
217
299
  saveCatalogCache(stateDir, cache);
300
+ console.log(`[llm-switcher:catalog] Detected ${tool} ${version} -> model catalog read from the tool (${local.models.length} models)`);
301
+ return;
302
+ }
303
+
304
+ // Without a local list every request would call the network again; wait between failed attempts.
305
+ const failed = failedAttempts.get(tool);
306
+ if (failed && failed.version === version && Date.now() - failed.at < RETRY_AFTER_MS) return;
307
+ if (refreshingTools.has(tool)) return;
218
308
 
219
- if (!refreshingTools.has(tool)) {
220
- refreshingTools.add(tool);
221
- fetchToolModels(tool, { apiKey })
222
- .then(res => {
223
- if (res.ok && res.models.length > 0) {
224
- const fresh = loadCatalogCache(stateDir);
225
- const existing = new Set(res.models.map(m => m.id));
226
- const baseline = BASELINE_MODELS[tool] || {};
227
- for (const [, ids] of Object.entries(baseline)) {
228
- for (const id of ids) {
229
- if (!existing.has(id)) {
230
- res.models.push({ id, ...(tool === 'claude' ? { tier: classifyClaudeTier(id) } : { role: classifyCodexRole(id) }) });
231
- }
232
- }
233
- }
234
- fresh[tool] = { lastSeenVersion: version, models: res.models };
235
- fresh.updatedAt = Date.now();
236
- saveCatalogCache(stateDir, fresh);
237
- console.log(`[llm-switcher:catalog] Detected ${tool} version update to ${version} -> refreshed model catalog (${res.models.length} models)`);
309
+ refreshingTools.add(tool);
310
+ fetchToolModels(tool, { apiKey })
311
+ .then(res => {
312
+ if (!res.ok || res.models.length === 0) {
313
+ failedAttempts.set(tool, { version, at: Date.now() });
314
+ console.log(`[llm-switcher:catalog] ${tool} ${version}: no local model list, official list failed (${res.error || 'empty'})`);
315
+ return;
316
+ }
317
+ const fresh = loadCatalogCache(stateDir);
318
+ const existing = new Set(res.models.map(m => m.id));
319
+ const baseline = BASELINE_MODELS[tool] || {};
320
+ for (const [, ids] of Object.entries(baseline)) {
321
+ for (const id of ids) {
322
+ if (!existing.has(id)) {
323
+ res.models.push({ id, ...(tool === 'claude' ? { tier: classifyClaudeTier(id) } : { role: classifyCodexRole(id) }) });
238
324
  }
239
- })
240
- .catch(() => {})
241
- .finally(() => {
242
- refreshingTools.delete(tool);
243
- });
244
- }
245
- }
325
+ }
326
+ }
327
+ fresh[tool] = { ...fresh[tool], lastSeenVersion: version, models: res.models, source: 'official' };
328
+ fresh.updatedAt = Date.now();
329
+ saveCatalogCache(stateDir, fresh);
330
+ console.log(`[llm-switcher:catalog] Detected ${tool} version update to ${version} -> refreshed model catalog (${res.models.length} models)`);
331
+ })
332
+ .catch(() => {})
333
+ .finally(() => {
334
+ refreshingTools.delete(tool);
335
+ });
246
336
  }
package/classifier.mjs ADDED
@@ -0,0 +1,238 @@
1
+ /**
2
+ * Jev-powered Semantic Classifier & Router for llm-switcher
3
+ * Evaluates prompt complexity using TypeSafe System One (Jev) models
4
+ * and maps intent to optimal model tier (haiku, sonnet, opus).
5
+ */
6
+
7
+ export const DEFAULT_JEV_URL = 'https://api.typesafe.ai/v1/systemone';
8
+ export const DEFAULT_JEV_MODEL = 'jev-latest';
9
+ export const OPENROUTER_JEV_URL = 'https://openrouter.ai/api/alpha/decisions';
10
+ export const OPENROUTER_JEV_MODEL = 'typesafe/jev-1.13';
11
+
12
+ export const TIER_CRITERIA = {
13
+ haiku: 'Mechanical task, template fill, single fact lookup, extraction, binary decision, short formatting',
14
+ sonnet: 'Multi-step coding, refactor, search across files, general programming, document drafting',
15
+ opus: 'Architecture design, security analysis, deep reasoning, race conditions, migration plans, complex synthesis',
16
+ };
17
+
18
+ /**
19
+ * Fast synchronous heuristic fallback when offline or no API key is set.
20
+ */
21
+ export function heuristicClassify(prompt = '') {
22
+ const text = String(prompt || '').toLowerCase();
23
+
24
+ // Heavy reasoning or safety critical indicators -> opus
25
+ const opusKeywords = [
26
+ 'architecture', 'security audit', 'race condition', 'deadlock', 'concurrency',
27
+ 'threat model', 'penetration', 'migration plan', 'adversarial', 'formal verification'
28
+ ];
29
+ if (opusKeywords.some((kw) => text.includes(kw))) {
30
+ return 'opus';
31
+ }
32
+
33
+ // Trivial or mechanical indicators -> haiku
34
+ const haikuKeywords = [
35
+ 'translate', 'format json', 'extract', 'regex', 'rename', 'single word',
36
+ 'yes or no', 'spell check', 'fix typo', 'convert to csv'
37
+ ];
38
+ if (text.length < 120 && haikuKeywords.some((kw) => text.includes(kw))) {
39
+ return 'haiku';
40
+ }
41
+
42
+ // Default balanced tier
43
+ return 'sonnet';
44
+ }
45
+
46
+ /**
47
+ * Resolves Jev API key from environment.
48
+ */
49
+ export function findJevKey(env = process.env) {
50
+ if (typeof env?.TYPESAFE_API_KEY === 'string' && env.TYPESAFE_API_KEY.trim().length > 0) {
51
+ return { key: env.TYPESAFE_API_KEY.trim(), source: 'typesafe' };
52
+ }
53
+ if (typeof env?.JEV_API_KEY === 'string' && env.JEV_API_KEY.trim().length > 0) {
54
+ return { key: env.JEV_API_KEY.trim(), source: 'openrouter' };
55
+ }
56
+ return null;
57
+ }
58
+
59
+ /**
60
+ * Classifies prompt into optimal model tier using TypeSafe Jev.
61
+ * Falls back safely to heuristic on missing key, timeout, or network error.
62
+ */
63
+ export async function classifyPrompt({
64
+ prompt,
65
+ apiKey,
66
+ url,
67
+ model,
68
+ timeoutMs = 5000,
69
+ env = process.env,
70
+ fetchImpl = globalThis.fetch,
71
+ } = {}) {
72
+ const fallbackTier = heuristicClassify(prompt);
73
+
74
+ const keyInfo = apiKey ? { key: apiKey, source: 'custom' } : findJevKey(env);
75
+ if (!keyInfo) {
76
+ return {
77
+ tier: fallbackTier,
78
+ confidence: 0.5,
79
+ source: 'heuristic',
80
+ reason: 'no-key',
81
+ };
82
+ }
83
+
84
+ const endpointUrl = url || (keyInfo.source === 'openrouter' ? OPENROUTER_JEV_URL : DEFAULT_JEV_URL);
85
+ const modelName = model || (keyInfo.source === 'openrouter' ? OPENROUTER_JEV_MODEL : DEFAULT_JEV_MODEL);
86
+
87
+ const controller = new AbortController();
88
+ const timer = setTimeout(() => controller.abort(), timeoutMs);
89
+
90
+ try {
91
+ const payload = {
92
+ state: String(prompt || ''),
93
+ model: modelName,
94
+ questions: {
95
+ recommended_tier: {
96
+ type: 'choice',
97
+ instructions: 'Determine the optimal model tier for this task according to complexity and reasoning depth required.',
98
+ criteria: TIER_CRITERIA,
99
+ },
100
+ },
101
+ };
102
+
103
+ const res = await fetchImpl(endpointUrl, {
104
+ method: 'POST',
105
+ headers: {
106
+ Authorization: `Bearer ${keyInfo.key}`,
107
+ 'Content-Type': 'application/json',
108
+ },
109
+ body: JSON.stringify(payload),
110
+ signal: controller.signal,
111
+ });
112
+
113
+ clearTimeout(timer);
114
+
115
+ if (!res.ok) {
116
+ return {
117
+ tier: fallbackTier,
118
+ confidence: 0.5,
119
+ source: 'heuristic',
120
+ reason: `http-${res.status}`,
121
+ };
122
+ }
123
+
124
+ const data = await res.json();
125
+ const ans = data?.answers?.recommended_tier;
126
+ if (!ans || typeof ans.choice !== 'string') {
127
+ return {
128
+ tier: fallbackTier,
129
+ confidence: 0.5,
130
+ source: 'heuristic',
131
+ reason: 'bad-response',
132
+ };
133
+ }
134
+
135
+ const tier = ['haiku', 'sonnet', 'opus'].includes(ans.choice) ? ans.choice : fallbackTier;
136
+ const confidence = typeof ans.confidence === 'number' ? Number(ans.confidence.toFixed(3)) : 0.8;
137
+
138
+ return {
139
+ tier,
140
+ confidence,
141
+ distribution: ans.distribution || null,
142
+ source: 'jev',
143
+ model: data.model || modelName,
144
+ };
145
+ } catch (err) {
146
+ clearTimeout(timer);
147
+ const reason = err?.name === 'AbortError' ? 'timeout' : 'network';
148
+ return {
149
+ tier: fallbackTier,
150
+ confidence: 0.5,
151
+ source: 'heuristic',
152
+ reason,
153
+ };
154
+ }
155
+ }
156
+
157
+ /**
158
+ * Checks semantic equivalence between candidate text and cached target using Jev Noul.
159
+ * Optional gate: returns false immediately if Jev is not configured or unavailable.
160
+ */
161
+ export async function checkSemanticEquivalence({
162
+ candidate,
163
+ target,
164
+ apiKey,
165
+ url,
166
+ model,
167
+ timeoutMs = 3000,
168
+ env = process.env,
169
+ fetchImpl = globalThis.fetch,
170
+ } = {}) {
171
+ if (!candidate || !target) return { equivalent: false, confidence: 0, reason: 'empty-input' };
172
+ if (candidate.trim() === target.trim()) {
173
+ return { equivalent: true, confidence: 1.0, reason: 'exact-match' };
174
+ }
175
+
176
+ const keyInfo = apiKey ? { key: apiKey, source: 'custom' } : findJevKey(env);
177
+ if (!keyInfo) {
178
+ return { equivalent: false, confidence: 0, reason: 'no-key-fallback-bypass' };
179
+ }
180
+
181
+ const endpointUrl = url || (keyInfo.source === 'openrouter' ? OPENROUTER_JEV_URL : DEFAULT_JEV_URL);
182
+ const modelName = model || (keyInfo.source === 'openrouter' ? OPENROUTER_JEV_MODEL : DEFAULT_JEV_MODEL);
183
+
184
+ const controller = new AbortController();
185
+ const timer = setTimeout(() => controller.abort(), timeoutMs);
186
+
187
+ try {
188
+ const payload = {
189
+ state: `REQUEST_A:\n${candidate}\n\nREQUEST_B:\n${target}`,
190
+ model: modelName,
191
+ questions: {
192
+ is_equivalent: {
193
+ type: 'noul',
194
+ instructions: 'Do REQUEST_A and REQUEST_B ask for the exact same semantic task or answer, such that the response to B fully satisfies A?',
195
+ },
196
+ },
197
+ };
198
+
199
+ const res = await fetchImpl(endpointUrl, {
200
+ method: 'POST',
201
+ headers: {
202
+ Authorization: `Bearer ${keyInfo.key}`,
203
+ 'Content-Type': 'application/json',
204
+ },
205
+ body: JSON.stringify(payload),
206
+ signal: controller.signal,
207
+ });
208
+
209
+ clearTimeout(timer);
210
+ if (!res.ok) {
211
+ return { equivalent: false, confidence: 0, reason: `http-${res.status}` };
212
+ }
213
+
214
+ const data = await res.json();
215
+ const ans = data?.answers?.is_equivalent;
216
+ if (ans == null) {
217
+ return { equivalent: false, confidence: 0, reason: 'bad-response' };
218
+ }
219
+
220
+ // Noul returns either probability number directly or object with probability
221
+ const prob = typeof ans === 'number' ? ans : (typeof ans?.probability === 'number' ? ans.probability : 0);
222
+ const equivalent = prob >= 0.85;
223
+
224
+ return {
225
+ equivalent,
226
+ confidence: Number(prob.toFixed(3)),
227
+ reason: equivalent ? 'jev-confirmed' : 'below-threshold',
228
+ model: data.model || modelName,
229
+ };
230
+ } catch (err) {
231
+ clearTimeout(timer);
232
+ return {
233
+ equivalent: false,
234
+ confidence: 0,
235
+ reason: err?.name === 'AbortError' ? 'timeout' : 'network-error',
236
+ };
237
+ }
238
+ }