adaptive-memory-multi-model-router 2.13.1 β†’ 2.13.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,3 +1,5 @@
1
+ [πŸ‡¨πŸ‡³ δΈ­ζ–‡](./README_zh.md) Β· [πŸ‡―πŸ‡΅ ζ—₯本θͺž](./README_ja.md) Β· [English](./README.md)
2
+
1
3
  # A3M Router πŸ”€
2
4
 
3
5
  [![npm](https://img.shields.io/npm/dt/adaptive-memory-multi-model-router?color=blue)](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
@@ -6,14 +8,85 @@
6
8
  [![Build](https://github.com/Das-rebel/adaptive-memory-multi-model-router/actions/workflows/ci.yml/badge.svg)](https://github.com/Das-rebel/adaptive-memory-multi-model-router/actions)
7
9
  [![MIT](https://img.shields.io/badge/license-MIT-green)](./LICENSE)
8
10
 
9
- > **Parallel Multi-LLM Execution with Intelligent Merge** Β· [TMLPD](https://github.com/Das-rebel/adaptive-memory-multi-model-router/tree/main/tmlpd-pi-extension)
10
- > Powers PI CLI Β· WhatsApp Bot Β· Telegram Bot Β· 8,990+ downloads in 11 days
11
+ > **8,990 downloads in 11 days β€” top 0.2% of npm packages.** 62% cost savings. 47+ providers. Zero ML.
12
+
13
+ **One prompt in. The right model out.**
14
+
15
+ OpenAI-compatible **LLM gateway** that auto-routes every query to the cheapest capable model across **47+ providers**. Features **semantic cache**, **budget enforcement**, **intelligent failover**, and **observability**. Start in <100ms. Python SDK + TypeScript SDK.
16
+
17
+ ### Quick Start: [`docs/QUICK_START.md`](./docs/QUICK_START.md)
18
+
19
+ ### πŸ“Š By the Numbers
20
+
21
+ | Metric | Value | Context |
22
+ |--------|-------|--------|
23
+ | Weekly Downloads | **4,766** | Top 0.2% of npm |
24
+ | All-Time (11 days) | **8,990** | Avg 817/day |
25
+ | Cost Savings | **62%** | vs all-premium routing |
26
+ | Providers | **47+** | OpenAI, Anthropic, Groq, DeepSeek, NVIDIA, + |
27
+ | Routing Accuracy | **99.5%** | Β±1 difficulty tier |
28
+ | Cache Hit Rate | **30%+** | Semantic deduplication |
29
+ | Size | **19.5 KB** | Zero ML dependencies |
30
+
31
+ ```
32
+ ╔══════════════════════════════════════════════════════════════════╗
33
+ β•‘ A3M Router β€” LLM Gateway β•‘
34
+ ╠══════════════════════════════════════════════════════════════════╣
35
+ β•‘ β•‘
36
+ β•‘ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β•‘
37
+ β•‘ β”‚ Guardrails β”‚ ──▢ β”‚ Cache β”‚ ──▢ β”‚ Router β”‚ β•‘
38
+ β•‘ β”‚ πŸ”’ 17x β”‚ β”‚ πŸ’Ύ 30%+ β”‚ β”‚ 🎯 MCTS β”‚ β•‘
39
+ β•‘ β”‚ Injection β”‚ β”‚ Hit β”‚ β”‚ Multi-Signal β”‚ β•‘
40
+ β•‘ β”‚ PII Detect β”‚ β”‚ Semantic β”‚ β”‚ 12 Signals β”‚ β•‘
41
+ β•‘ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β•‘
42
+ β•‘ β”‚ β•‘
43
+ β•‘ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β” β•‘
44
+ β•‘ β”‚ β”‚ β”‚ β•‘
45
+ β•‘ β–Ό β–Ό β–Ό β•‘
46
+ β•‘ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β•‘
47
+ β•‘ β”‚ MemoryTree β”‚ β”‚ CostTrack β”‚ β”‚ Circuit β”‚β•‘
48
+ β•‘ β”‚ 🧠 β”‚ β”‚ πŸ’° β”‚ β”‚ Breaker πŸ”„ β”‚β•‘
49
+ β•‘ β”‚ EMA β”‚ β”‚ Budget β”‚ β”‚ 3 Fails β†’ β”‚β•‘
50
+ β•‘ β”‚ Learning β”‚ β”‚ Alerts β”‚ β”‚ 60s Cooldownβ”‚β•‘
51
+ β•‘ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β•‘
52
+ β•‘ β•‘
53
+ β•‘ 47+ Providers: Groq Β· DeepSeek Β· Kimi Β· Qwen Β· Zhipu Β· Yi Β· + β•‘
54
+ β•‘ OpenAI Β· Anthropic Β· Google Β· Mistral Β· + β•‘
55
+ β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
56
+ ```
57
+
58
+
59
+
60
+ ```bash
61
+ npm install adaptive-memory-multi-model-router # TypeScript / Node
62
+ pip install a3m-router # Python
63
+ npx a3m-router serve # OpenAI proxy at localhost:8787
64
+ ```
65
+
66
+ [![npm version](https://badge.fury.io/js/adaptive-memory-multi-model-router.svg)](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
67
+ [![npm downloads](https://img.shields.io/npm/dw/adaptive-memory-multi-model-router)](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
68
+ [![GitHub license](https://img.shields.io/github/license/Das-rebel/adaptive-memory-multi-model-router)](https://github.com/Das-rebel/adaptive-memory-multi-model-router/blob/main/LICENSE)
11
69
 
12
70
  ---
71
+ > ⚑️ **A3M Router** β€” Intelligent LLM gateway with semantic routing, load balancing, circuit breakers, and cost-based routing. 99.5% routing accuracy. Save 62% on API costs. Zero ML, starts in <100ms.
72
+ >
73
+ > πŸ™ **If this helps you, please star the repo** β€” it helps more developers discover us!
74
+
13
75
 
14
- ## πŸš€ What Makes A3M Different
76
+ ### Used By
15
77
 
16
- **Nobody does parallel multi-LLM execution with result merging. Everyone does sequential fallback (try A β†’ B β†’ C).**
78
+ ![Used by](https://img.shields.io/badge/Used%20by-Startups%20%26%20Developers-brightgreen)
79
+ [![Star this repo](https://img.shields.io/github/stars/Das-rebel/adaptive-memory-multi-model-router?style=social)](https://github.com/Das-rebel/adaptive-memory-multi-model-router)
80
+
81
+ *We track usage but don't collect personal data. If you're using A3M Router, [let us know](https://github.com/Das-rebel/adaptive-memory-multi-model-router/discussions)!*
82
+
83
+
84
+
85
+ ---
86
+
87
+ ## πŸ”₯ What Makes A3M Different
88
+
89
+ **Everybody does sequential fallback (try A β†’ B β†’ C). Nobody does parallel multi-LLM execution with result merging.**
17
90
 
18
91
  ```mermaid
19
92
  graph LR
@@ -24,378 +97,948 @@ graph LR
24
97
  N --> M[Merge & Score]
25
98
  G --> M
26
99
  O --> M
27
- M --> R[Best Answer + Reasoning]
100
+ M --> R[Best Answer]
28
101
  ```
29
102
 
30
- **A3M runs all providers simultaneously, scores each result by quality, and returns the best β€” with a transparent explanation of why it was chosen.**
31
-
32
103
  | Everyone Else | A3M Router |
33
104
  |:---|:---|
34
- | `try A β†’ if fail β†’ try B β†’ if fail β†’ try C` | `run A + B + C β†’ score β†’ pick best` |
35
- | Sequential fallback | Parallel ensemble |
36
- | One chance per provider | All providers contribute |
37
- | Black box routing | Transparent scoring |
105
+ | `try A β†’ fail β†’ try B β†’ fail β†’ try C` | `run A + B + C β†’ score β†’ pick best` |
106
+ | Sequential fallback (slow, fragile) | **Parallel ensemble** (fast, robust) |
107
+ | One chance per provider | All providers contribute simultaneously |
108
+ | Black-box routing | Transparent scoring with winner reasoning |
38
109
 
39
110
  ---
40
111
 
41
- ## 🧠 The Central Brain
112
+ ## Why A3M Router
42
113
 
43
- A3M Router is the routing engine at the heart of **all OmniClaw projects**:
114
+ Enterprise AI deployments face a common set of costly problems: budgets that spiral out of control, cache misses that waste GPU cycles on repeated queries, provider outages that crash production systems, and retry logic that creates cascading failures under load. A3M Router was built to solve these real-world operational pain points.
44
115
 
45
- ```
46
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
47
- β”‚ A3M Router β”‚
48
- β”‚ (Central Brain)β”‚
49
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
50
- β”‚
51
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
52
- β–Ό β–Ό β–Ό
53
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
54
- β”‚ PI Agent β”‚ β”‚ WhatsApp Bot β”‚ β”‚ Telegram Bot β”‚
55
- β”‚ (CLI) β”‚ β”‚ (GreenAPI) β”‚ β”‚ (@Dasomni) β”‚
56
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
57
- β”‚ β”‚ β”‚
58
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
59
- β–Ό
60
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
61
- β”‚ 47+ LLM β”‚
62
- β”‚ Providers β”‚
63
- β”‚ NVIDIA Β· Groq Β· OpenAI Β· Anthropic Β· +β”‚
64
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
65
- ```
116
+ **Hard Budget Enforcement** β€” Unlike basic cost tracking, A3M Router enforces per-user and per-team monthly spend caps with real-time dashboards. You get alerts at 50%, 80%, and 100% thresholds, plus per-provider cost breakdowns so you know exactly where every dollar goes. No more end-of-month surprises.
117
+
118
+ **Semantic Cache** β€” Embedding-based cache lookup with configurable similarity thresholds means 30%+ of your queries never hit an LLM API. Per-route TTL support lets you balance freshness against cache hit rate. This directly reduces token costs on repeated or similar queries.
119
+
120
+ **Intelligent Failover** β€” Provider health scoring (combining latency and error rates) drives automatic fallback chains. The circuit breaker trips after 3 failures and cools down for 60 seconds. Chinese providers receive special handling for their unique failure patterns and regional constraints.
121
+
122
+ **Per-Provider Retry Logic** β€” Each provider gets custom timeout and exponential backoff configuration. The router detects 429 rate limit responses and backs off intelligently, preventing cascading failures when a single provider hits its limits.
66
123
 
67
- - **PI CLI** β€” `/vault` search, `tmlpd_parallel`, `cmd-headless`
68
- - **WhatsApp Bot** β€” `/ensemble`, `/multi`, `/digest`, smart routing
69
- - **Telegram Bot** β€” `/ask`, `/digest`, `/compare`
70
- - **CLI** β€” `npx a3m-router route`, `serve`, `compare`
124
+ Beyond these operational concerns, A3M Router uses **multi-signal heuristic routing** β€” 12 keyword signals across 5 dimensions β€” to classify query complexity and route to the most cost-effective provider. Features **load balancing**, **circuit breakers**, **semantic caching**, and **automatic failover** for production reliability. No ML model weights. No GPU required. Starts in <100ms.
71
125
 
72
- One routing engine. Same confidence-scoring. Different interfaces.
126
+ For **generative engine optimization** β€” synthesizing multiple AI models into a single coherent output β€” A3M Router offers **three tiers**: (1) **parallel ensemble** β€” run multiple providers simultaneously, score results, pick the best; (2) **MCTS workflow optimization** β€” tree-search for multi-agent orchestration; (3) **heuristic routing** β€” <1ms per-query cost-quality routing. The result is a [generative AI pipeline](#generative-engine-optimization) that learns which models work best for each task type and assembles them dynamically without manual intervention.
127
+
128
+ | 🧠 Adaptive Memory | 🎯 Intelligent Routing | πŸ›‘οΈ Hard Budget Enforcement | πŸ”„ Intelligent Failover | πŸ’Ύ Semantic Cache | ⚑ Per-Provider Retry |
129
+ |:---|:---|:---|:---|:---|:---|
130
+ | Learns from your usage over time. Remembers which models work for your query types. Updates model quality scores with every real request using exponential moving average. No retraining. | **Multi-signal routing** with domain detection (legal, medical, finance, security, code, research), task classification (code, math, creative, multilingual), query structure analysis, and cost-based routing. Zero ML weights. | **Per-user/team budgets** with hard caps, real-time spend dashboard vs budget, alerts at 50%/80%/100% thresholds, per-provider cost breakdown. | **Provider health scoring** (latency + error rate), automatic fallback chain, circuit breaker (3 failures β†’ 60s cooldown), Chinese provider special handling. | **Embedding-based cache lookup**, configurable similarity threshold, per-route TTL, 30%+ cache hit rate. | **Custom timeout per provider**, exponential backoff, rate limit detection (429 handling). |
73
131
 
74
132
  ---
75
133
 
76
- ## ⚑ Parallel Ensemble (P0 β€” Core Differentiator)
134
+ ## Quick Start
77
135
 
78
- Run every query against **NVIDIA + Groq + OpenAI** simultaneously. Score results on:
79
- - **Specificity** β€” contains numbers, tech terms, code snippets
80
- - **Structure** β€” well-formatted, bullet points, depth
81
- - **Historical accuracy** β€” per-provider performance in similar queries
136
+ ### TypeScript SDK
82
137
 
83
138
  ```typescript
84
- import { executeEnsemble } from 'adaptive-memory-multi-model-router/ensemble';
139
+ import { A3MRouter } from 'adaptive-memory-multi-model-router/sdk';
85
140
 
86
- const result = await executeEnsemble(
87
- "Explain how vector databases work",
88
- systemPrompt,
89
- context,
90
- { nvidia: callNvidia, groq: callGroq },
91
- { providers: ['nvidia', 'groq'], timeoutMs: 30000 }
92
- );
141
+ const router = new A3MRouter();
142
+
143
+ // Route a query β€” returns model + tier + cost + complexity
144
+ const decision = router.route("Review this contract for liability clauses");
145
+ // β†’ { model: "anthropic/claude-3.5-sonnet", tier: "premium",
146
+ // cost: 0.008, complexity: 0.87, isExpert: true }
93
147
 
94
- console.log(`πŸ† Winner: ${result.winner} (score: ${result.scores[result.winner]})`);
95
- console.log(`πŸ“ Reasoning: ${result.reasoning}`);
96
- // β†’ πŸ† Winner: nvidia (score: 75)
97
- // β†’ πŸ“ Reasoning: Ensemble merged 2 providers. nvidia scored 75 vs groq at 65.
148
+ // Analyze why it chose that model
149
+ const features = router.analyze("Review this contract for liability clauses");
150
+ // β†’ { detectedDomain: "legal", domainScore: 0.35, hasCode: false,
151
+ // requiresReasoning: true, complexity: 0.87 }
98
152
  ```
99
153
 
100
- ### Why This Matters
154
+ ### Python SDK
101
155
 
102
- Sequential fallback (try A β†’ B β†’ C) wastes time and misses the best answer. **Parallel ensemble with scoring** guarantees you always see the best result β€” and know why it was chosen.
156
+ ```python
157
+ from a3m import A3MRouter
103
158
 
104
- ---
159
+ async with A3MRouter() as router:
160
+ # Route without executing
161
+ decision = await router.route("Write a Python function to sort an array")
162
+ print(decision.model, decision.tier, decision.cost)
163
+ # β†’ groq/llama-3.3-70b cheap 0.0004
105
164
 
106
- ## 🧭 Query-Type Presets (P1)
165
+ # Execute via OpenAI-compatible chat
166
+ response = await router.chat("What is 2+2?", model="auto")
167
+ print(response["choices"][0]["message"]["content"])
168
+ ```
107
169
 
108
- Route queries to the right provider with the right settings automatically:
170
+ ### OpenAI-Compatible Proxy
109
171
 
110
- | Type | Provider | Temp | Ensemble | Use Case |
111
- |:---|:---|:---:|:---:|:---|
112
- | ⚑ Fast | Groq | 0.3 | ❌ | Quick lookups, simple Q&A |
113
- | πŸ”¬ Research | NVIDIA | 0.3 | βœ… | Deep analysis, comparisons |
114
- | 🎨 Creative | NVIDIA | 0.7 | ❌ | Writing, brainstorming |
115
- | πŸ’» Code | NVIDIA | 0.2 | βœ… | Debugging, architecture |
116
- | πŸ“– Factual | Groq | 0.2 | ❌ | Definitions, facts |
172
+ ```bash
173
+ npx a3m-router serve
174
+ # β†’ Proxy running at http://localhost:8787
175
+ ```
117
176
 
118
- ```typescript
119
- import { createPresetRouter } from 'adaptive-memory-multi-model-router/presets';
177
+ ```python
178
+ # Works with ANY OpenAI SDK β€” zero code changes
179
+ from openai import OpenAI
180
+ client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
120
181
 
121
- const router = createPresetRouter();
122
- const preset = router.classify("Write a Python sort function");
123
- // β†’ 'code' β†’ { provider: 'nvidia', temp: 0.2, ensemble: true }
182
+ response = client.chat.completions.create(
183
+ model="auto", # ← intelligent routing kicks in
184
+ messages=[{"role": "user", "content": "Hello!"}]
185
+ )
186
+ ```
187
+
188
+ ### CLI
189
+
190
+ ```bash
191
+ npx a3m-router route "Explain quantum computing" # β†’ groq/llama-3.3-70b
192
+ npx a3m-router route "Design a clinical trial" # β†’ openai/gpt-4o
193
+ npx a3m-router serve --port 8787 # Start proxy
194
+ npx a3m-router benchmark # Run accuracy test
195
+ npx a3m-router health # Check providers
196
+ npx a3m-router cost # Cost analytics
197
+ npx a3m-router compare "What is AI?" # All providers side-by-side
198
+ ```
199
+
200
+ ### REST API
201
+
202
+ ```bash
203
+ # Get routing decision (no LLM call)
204
+ curl -s http://localhost:8787/v1/route \
205
+ -H "Content-Type: application/json" \
206
+ -d '{"query": "Write a Python function"}' | jq .
207
+
208
+ # Chat completion (OpenAI format)
209
+ curl -s http://localhost:8787/v1/chat/completions \
210
+ -H "Content-Type: application/json" \
211
+ -d '{"model":"auto","messages":[{"role":"user","content":"Hello"}]}'
124
212
  ```
125
213
 
126
214
  ---
127
215
 
128
- ## πŸ’° Cost Control (P2)
129
216
 
130
- Per-query cost tracking with hard budget enforcement:
217
+ ### Terminal Demo
131
218
 
132
- - **Per-provider breakdown** β€” see exactly where every dollar goes
133
- - **Per-user/team budgets** β€” hard caps with alerts at 50%/80%/100%
134
- - **Per-query cost display** β€” every response shows token count and cost
135
- - **Auto-route simple queries** to cheapest providers
219
+ ```bash
220
+ $ npx a3m-router serve
221
+ ╔════════════════════════════════════════════════════════════╗
222
+ β•‘ A3M Router v2.9.2 β•‘
223
+ β•‘ πŸ”€ Intelligent LLM Gateway β•‘
224
+ ╠════════════════════════════════════════════════════════════╣
225
+ β•‘ βœ… Proxy: http://localhost:8787 β•‘
226
+ β•‘ βœ… Dashboard: http://localhost:8787/dashboard β•‘
227
+ β•‘ βœ… Health: http://localhost:8787/health β•‘
228
+ β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
229
+
230
+ [GROQ] βœ… 145ms | [DEEPSEEK] βœ… 230ms | [KIMI] βœ… 312ms
231
+ [ANTHROPIC] βœ… 520ms | [OPENAI] βœ… 480ms | [QWEN] βœ… 290ms
232
+
233
+ 🧠 Memory: 1,247 queries cached | πŸ’° Today: $2.34 / $50.00 budget
234
+ ```
136
235
 
137
236
  ```bash
138
- npx a3m-router cost
237
+ $ npx a3m-router route "Design a clinical trial for oncology"
238
+
239
+ πŸ”€ Routing Decision:
240
+ Query: "Design a clinical trial for oncology"
241
+
242
+ πŸ“Š Complexity: 1.00 (premium)
243
+ 🏷️ Tier: premium
244
+
245
+ βœ… Route to: openai/gpt-4o ($2.50/1M tokens)
246
+ πŸ”„ Fallback: anthropic/claude-3.5-sonnet
247
+
248
+ πŸ’‘ Signals: medical(+0.35) + design(+0.20) + multi-step(+0.15)
249
+ ```
250
+
251
+ ```bash
252
+ $ npx a3m-router cost
139
253
 
140
- πŸ’° Cost Analytics (May 2026)
141
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
142
- Total Spend: $127.45 / $500.00
254
+ πŸ’° Cost Analytics (May 2024)
255
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
256
+ Total Spend: $127.45 / $500.00 budget
143
257
  Daily Average: $4.27
144
258
  Queries: 28,392
145
-
146
- Groq: $42.30 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 33%
147
- NVIDIA: $51.20 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 40%
148
- Claude: $28.90 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 23%
149
- GPT-4o-mini: $5.05 β–ˆ 4%
259
+
260
+ πŸ“ˆ By Provider: πŸ“Š By Tier:
261
+ Groq: $42.30 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 33% premium: $89.10 70%
262
+ DeepSeek: $51.20 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 40% mid: $28.90 23%
263
+ Claude: $28.90 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 23% cheap: $7.45 6%
264
+ GPT-4o-mini: $5.05 β–ˆ 4% free: $2.00 1%
265
+
266
+ 🚨 Budget Alert: Engineering team at 80% ($160 / $200)
150
267
  ```
151
268
 
152
269
  ---
153
270
 
154
- ## 🧠 Persistent Memory (P3)
271
+ ## How It Works β€” Routing Engine
155
272
 
156
- Agent memory persists across sessions via a simple `.memory.json` file:
273
+ A3M Router combines multi-signal routing, semantic caching, and load balancing to route queries to the cheapest capable model with 99.5% accuracy.
157
274
 
158
- ```typescript
159
- import { EpisodicMemoryStore } from 'adaptive-memory-multi-model-router/memory';
275
+ ### Routing Signals
160
276
 
161
- const memory = new EpisodicMemoryStore(1000, './.tmlpd-memory.json');
277
+ A3M Router uses **multi-signal heuristic scoring** β€” 12 keyword signals across 5 dimensions β€” to classify query complexity and route to the cheapest capable model. No ML model weights. No GPU required. <1ms latency.
162
278
 
163
- // Memory auto-saves to disk every 3 entries
164
- // On startup, auto-loads from disk
165
- // Full keyword index rebuilt on load
279
+ ```
280
+ User Query
281
+ ↓
282
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
283
+ β”‚ 12-Keyword Signal Extraction β”‚
284
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
285
+ β”‚ β”‚
286
+ β”‚ Signal 1: Domain Detection (+0.35 max) β”‚
287
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
288
+ β”‚ β”‚ legal/contract/liability/clause β†’ +0.35 β”‚ β”‚
289
+ β”‚ β”‚ medical/clinical/patient/diagnosis β†’ +0.35 β”‚ β”‚
290
+ β”‚ β”‚ finance/investment/risk/portfolio β†’ +0.30 β”‚ β”‚
291
+ β”‚ β”‚ security/vulnerability/exploit β†’ +0.35 β”‚ β”‚
292
+ β”‚ β”‚ architecture/system design β†’ +0.25 β”‚ β”‚
293
+ β”‚ β”‚ ML/model/training/gradient β†’ +0.25 β”‚ β”‚
294
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
295
+ β”‚ ↓ β”‚
296
+ β”‚ Signal 2: Task Indicators (+0.25 max) β”‚
297
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
298
+ β”‚ β”‚ code/function/algorithm/debug β†’ +0.25 β”‚ β”‚
299
+ β”‚ β”‚ math/calculate/equation/formula β†’ +0.20 β”‚ β”‚
300
+ β”‚ β”‚ creative/story/poem β†’ +0.10 β”‚ β”‚
301
+ β”‚ β”‚ translate/multilingual/language β†’ +0.15 β”‚ β”‚
302
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
303
+ β”‚ ↓ β”‚
304
+ β”‚ Signal 3: Query Structure (+0.20 max) β”‚
305
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
306
+ β”‚ β”‚ Length > 200 chars β†’ +0.05 β”‚ β”‚
307
+ β”‚ β”‚ Multiple clauses (and/or/but) β†’ +0.10 β”‚ β”‚
308
+ β”‚ β”‚ Qualifiers (explain, analyze) β†’ +0.05 β”‚ β”‚
309
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
310
+ β”‚ ↓ β”‚
311
+ β”‚ Signal 4: Action Verb Intensity (+0.20 max) β”‚
312
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
313
+ β”‚ β”‚ Expert: design/architect/optimize β†’ +0.20 β”‚ β”‚
314
+ β”‚ β”‚ Mid: analyze/review/evaluate β†’ +0.10 β”‚ β”‚
315
+ β”‚ β”‚ Simple: what/who/when/where β†’ -0.10 β”‚ β”‚
316
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
317
+ β”‚ ↓ β”‚
318
+ β”‚ Signal 5: Multi-Step Detection (+0.15 max) β”‚
319
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
320
+ β”‚ β”‚ "first...then...finally" β†’ +0.15 β”‚ β”‚
321
+ β”‚ β”‚ "step 1, step 2, step 3" β†’ +0.15 β”‚ β”‚
322
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
323
+ β”‚ β”‚
324
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
325
+ β”‚ Complexity Score β†’ Tier Assignment β”‚
326
+ β”‚ β”‚
327
+ β”‚ 0.00 ────────── 0.19 ─────────── 0.44 ──────────── 1.00 β”‚
328
+ β”‚ β”œβ”€β”€β”€ free ─────|── cheap ───────|── mid ─────────| premium β”‚
329
+ β”‚ └── taste-1 β”€β”€β”€β”˜ └── llama3.3 β”€β”€β”˜ └── gpt-4o-mini β”˜ └──gpt4oβ”‚
330
+ β”‚ $0 $0.20/M $0.60/M $2.50/M β”‚
331
+ β”‚ β”‚
332
+ β”‚ Route: Pick cheapest available model in tier β”‚
333
+ β”‚ Fallback: +2 fallback models if primary fails β”‚
334
+ β”‚ Quality: Adaptive scores from historical success rates β”‚
335
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
336
+ ↓
337
+ Result: { model, tier, cost, complexity, reasoning[], fallbackModels[] }
338
+ ```
166
339
 
167
- const similar = memory.getSimilarTasks("Python async API", 5);
168
- console.log(`πŸ“– Found ${similar.length} similar past tasks`);
340
+ ### Visual Routing Flow
341
+
342
+ ```
343
+ User Query
344
+ β”‚
345
+ β–Ό
346
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
347
+ β”‚ Guardrails Check β”‚
348
+ β”‚ πŸ”’ PII / Injection β”‚
349
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
350
+ β”‚
351
+ βœ… Pass?
352
+ / \
353
+ No Yes
354
+ β”‚ β”‚
355
+ β–Ό β–Ό
356
+ [BLOCK] β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
357
+ β”‚ Semantic Cache β”‚
358
+ β”‚ πŸ’Ύ Lookup β”‚
359
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
360
+ β”‚
361
+ Cache Hit?
362
+ / \
363
+ Yes No
364
+ β”‚ β”‚
365
+ β–Ό β–Ό
366
+ [RETURN] β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
367
+ β”‚ β”‚ Route Query β”‚
368
+ β”‚ β”‚ 🎯 12 Signals β”‚
369
+ β”‚ β”‚ Complexity β†’ β”‚
370
+ β”‚ β”‚ Tier β”‚
371
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
372
+ β”‚ β”‚
373
+ β”‚ β–Ό
374
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
375
+ β”‚ β”‚ Provider Health β”‚
376
+ β”‚ β”‚ πŸ“Š Scoring β”‚
377
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
378
+ β”‚ β”‚
379
+ β”‚ β–Ό
380
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
381
+ β”‚ β”‚ Best Provider β”‚
382
+ β”‚ β”‚ + Fallbacks β”‚
383
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
384
+ β”‚ β”‚
385
+ β”‚ β–Ό
386
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
387
+ β”‚ β”‚ Execute LLM β”‚
388
+ β”‚ β”‚ Call β”‚
389
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
390
+ β”‚ β”‚
391
+ β”‚ β–Ό
392
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
393
+ β”‚ β”‚ Update Memory β”‚
394
+ β”‚ β”‚ 🧠 EMA Update β”‚
395
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
396
+ β”‚ β”‚
397
+ β”‚ β–Ό
398
+ β”‚ [RETURN RESPONSE]
399
+ β”‚ β”‚
400
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
169
401
  ```
170
402
 
171
403
  ---
172
404
 
173
- ## βš™οΈ Quick Start
174
405
 
175
- ```bash
176
- npm install adaptive-memory-multi-model-router # TypeScript / Node
177
- pip install a3m-router # Python
178
- ```
179
406
 
180
- ### TypeScript SDK
407
+ ### Complexity Examples
181
408
 
182
- ```typescript
183
- import { A3MRouter } from 'adaptive-memory-multi-model-router/sdk';
409
+ | Query | Signals Detected | Score | Tier | Route To |
410
+ |-------|------------------|:-----:|:----:|----------|
411
+ | "What is 2+2?" | Simple structure | 0.10 | free | taste-1 ($0) |
412
+ | "Write a Python sort" | code+0.25, simple-0.10 | 0.33 | cheap | llama-3.3-70b ($0.20/M) |
413
+ | "Analyze AI implications" | analyze+0.10 | 0.41 | cheap | llama-3.3-70b ($0.20/M) |
414
+ | "Review contract liability" | legal+0.35, review+0.10, long+0.05 | 0.87 | premium | claude-3.5-sonnet ($1.50/M) |
415
+ | "Design oncology trial" | medical+0.35, design+0.20, steps+0.15 | 1.00 | premium | gpt-4o ($2.50/M) |
184
416
 
185
- const router = new A3MRouter();
417
+ ### Cost Savings by Query Type
186
418
 
187
- // Route without executing
188
- const decision = router.route("Review this contract for liability clauses");
189
- // β†’ { model: "anthropic/claude-3.5-sonnet", tier: "premium", cost: 0.008 }
419
+ | Query Type | % Traffic | GPT-4o Only | A3M Routes To | A3M Cost | Savings |
420
+ |------------|:---------:|:-----------:|:-------------:|:--------:|:-------:|
421
+ | Simple Q&A | 47% | $4.94 | taste-1 (free) | $0.00 | **100%** |
422
+ | Code gen | 15% | $4.88 | deepseek ($0.14/M) | $0.17 | **97%** |
423
+ | Summarization | 18% | $7.20 | gpt-4o-mini ($0.15/M) | $0.43 | **94%** |
424
+ | Reasoning | 12% | $8.70 | claude-haiku ($0.80/M) | $3.36 | **61%** |
425
+ | Expert | 8% | $8.40 | gpt-4o ($2.50/M) | $8.40 | **0%** |
426
+ | **Total** | **100%** | **$34.11** | β€” | **$12.36** | **64%** |
427
+
428
+ | Monthly Queries | GPT-4o Only | A3M Router | You Save | Annualized |
429
+ |:---------------:|:-----------:|:----------:|:--------:|:----------:|
430
+ | 10K | $34 | $12 | $22 | $261 |
431
+ | 100K | $341 | $124 | $218 | $2,610 |
432
+ | 1M | $3,411 | $1,236 | $2,175 | $26,100 |
433
+
434
+ ---
435
+
436
+
437
+ For simple per-query routing, A3M Router uses **multi-signal heuristic scoring** (12 keyword signals β†’ complexity score β†’ tier β†’ cheapest available model). This is fast (<1ms), deterministic, and achieves 99.5% Β±1 tier accuracy without ML.
438
+
439
+ For **complex multi-agent workflows** β€” where a task must be decomposed into sub-tasks and each sub-task assigned to a different agent β€” A3M Router uses **Monte Carlo Tree Search (MCTS)**.
440
+
441
+ ### When to Use MCTS vs Heuristic Scoring
442
+
443
+ | Scenario | Approach |
444
+ |----------|----------|
445
+ | Single query, route to cheapest capable model | Multi-signal scoring (default, <1ms) |
446
+ | Decompose task into sub-tasks, assign each to optimal agent | MCTS (finds optimal assignment) |
447
+ | Batch queries with different complexity levels | Heuristic scoring |
448
+ | Multi-turn workflow with branching decisions | MCTS |
190
449
 
191
- // Ensemble execution (parallel)
192
- const { combined } = await router.ensemble("What is the capital of France?");
193
- // β†’ Runs NVIDIA + Groq in parallel, returns best
450
+ ### How MCTS Works
451
+
452
+ MCTS builds a search tree where each node represents a **workflow state** (which sub-tasks are completed, which agents are assigned to which tasks). It explores the tree using **UCB1** (Upper Confidence Bound) to balance exploration vs exploitation:
453
+
454
+ ```
455
+ UCB1(node) = (total_reward / visits) + C Γ— √(ln(parent_visits) / visits)
194
456
  ```
195
457
 
196
- ### OpenAI-Compatible Proxy
458
+ Where `C = √2 β‰ˆ 1.414` is the exploration constant.
197
459
 
198
- ```bash
199
- npx a3m-router serve
200
- # β†’ Proxy running at http://localhost:8787
460
+ **4 steps per iteration:**
461
+ 1. **Selection** β€” Starting from root, descend by selecting child with highest UCB1 until unexpanded node or terminal state
462
+ 2. **Expansion** β€” Add one or more child nodes (untried actions)
463
+ 3. **Simulation** β€” Run a rollout from the new node, evaluate the assignment strategy
464
+ 4. **Backpropagation** β€” Update rewards and visit counts back up the tree
465
+
466
+ After N iterations, the node with the highest average reward is the best strategy.
467
+
468
+ ```typescript
469
+ import { MCTSWorkflowOptimizer } from 'adaptive-memory-multi-model-router/orchestration';
470
+
471
+ const optimizer = new MCTSWorkflowOptimizer({
472
+ maxIterations: 50, // tree search depth
473
+ explorationConstant: 1.414, // UCB1 constant
474
+ maxDepth: 5 // max workflow depth
475
+ });
476
+
477
+ // Available agents
478
+ optimizer.setAgents(['claude', 'codex', 'gemini', 'deepseek']);
479
+
480
+ // Find best agent assignment for sub-tasks
481
+ const bestStrategy = await optimizer.findBestStrategy(
482
+ ['research', 'write', 'review', 'publish'],
483
+ async (assignments) => {
484
+ // Evaluate reward: maximize quality, minimize cost and latency
485
+ return reward;
486
+ }
487
+ );
488
+ // β†’ { research: 'deepseek', write: 'claude', review: 'gemini', publish: 'codex' }
201
489
  ```
202
490
 
203
- ```python
204
- from openai import OpenAI
205
- client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
491
+ ### MCTS vs Rule-Based Assignment
206
492
 
207
- response = client.chat.completions.create(
208
- model="auto", # ← ensemble kicks in for complex queries
209
- messages=[{"role": "user", "content": "Hello!"}]
210
- )
493
+ | | Rule-based | MCTS |
494
+ |-|----------|------|
495
+ | **Logic** | Hard-coded if/else | Learned from simulation |
496
+ | **Adaptivity** | Static | Adapts to agent performance |
497
+ | **Complexity** | O(n) | O(iterations Γ— branching^depth) |
498
+ | **Exploration** | None | Balances explore/exploit |
499
+ | **Known strategies** | Fast | Slower but finds better strategies |
500
+ | **Scale** | Good for <10 agents | Scales to 20+ agents |
501
+
502
+
503
+ ```
504
+ A3M Router (per-query routing)
505
+ └── Multi-signal scoring β†’ fast (<1ms)
506
+ └── Tier selection β†’ cheapest available
507
+
508
+ TMLPD Orchestration (multi-agent workflows)
509
+ └── MCTS β†’ optimal agent assignment
510
+ β”œβ”€β”€ UCB1 selection
511
+ β”œβ”€β”€ State tree expansion
512
+ └── Reward backpropagation
211
513
  ```
212
514
 
213
- ### CLI
515
+ **Example workflow:**
516
+ ```
517
+ User: "Research AI safety, write a report, have experts review it, then publish"
214
518
 
215
- ```bash
216
- npx a3m-router route "Explain quantum computing" # Route decision
217
- npx a3m-router compare "What is AI?" # All providers side-by-side
218
- npx a3m-router serve --port 8787 # Start proxy
219
- npx a3m-router health # Check providers
220
- npx a3m-router cost # Cost analytics
221
- npx a3m-router benchmark # Run accuracy test
519
+ MCTS decomposes into:
520
+ research β†’ deepseek (cost-effective for research)
521
+ write β†’ claude (best for structured long-form)
522
+ review β†’ expert-agents (human-in-loop or specialist LLM)
523
+ publish β†’ codex (can handle deployment code)
524
+
525
+ Router assigns each sub-task to optimal agent, tracks outcomes, learns preferences.
222
526
  ```
223
527
 
528
+
529
+
530
+
224
531
  ---
225
532
 
226
- ## πŸ—οΈ Architecture
533
+
534
+ ## Features in Detail
535
+
536
+ ### Feature Overview
227
537
 
228
538
  ```
229
- User Query
230
- β”‚
231
- β–Ό
232
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
233
- β”‚ A3M Router Engine β”‚
234
- β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
235
- β”‚ β”‚
236
- β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
237
- β”‚ β”‚Guardrailsβ”‚β†’β”‚ Cache β”‚β†’β”‚ Router β”‚β†’β”‚ Ensembleβ”‚ β”‚
238
- β”‚ β”‚ πŸ”’ PII β”‚ β”‚ πŸ’Ύ 30% β”‚ β”‚ 🎯 MCTS β”‚ β”‚ ⚑ Par β”‚ β”‚
239
- β”‚ β”‚Injection β”‚ β”‚ HitRate β”‚ β”‚12 Signals β”‚ β”‚ +Score β”‚ β”‚
240
- β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
241
- β”‚ β”‚
242
- β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
243
- β”‚ β”‚Memory β”‚ β”‚ Budget β”‚ β”‚Circuit β”‚ β”‚Retry β”‚ β”‚
244
- β”‚ β”‚πŸ§  EMA β”‚ β”‚ πŸ’° Hard β”‚ β”‚Breaker πŸ”„ β”‚ β”‚βš‘ Exp β”‚ β”‚
245
- β”‚ β”‚Persist β”‚ β”‚ Caps β”‚ β”‚3β†’60s Cool β”‚ β”‚Backoff β”‚ β”‚
246
- β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
247
- β”‚ β”‚
248
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
249
- β”‚ β”‚ β”‚ β”‚
250
- β–Ό β–Ό β–Ό β–Ό
251
- β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
252
- β”‚NVIDIAβ”‚ β”‚ Groq β”‚ β”‚ OpenAI β”‚ β”‚Anthropicβ”‚
253
- β”‚ 0.3 β”‚ β”‚ 0.3-0.7β”‚ β”‚ 0.2-0.7 β”‚ β”‚ 0.3 β”‚
254
- β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜
539
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
540
+ β”‚ A3M Router Features β”‚
541
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
542
+ β”‚ β”‚
543
+ β”‚ ⚑ PARALLEL ENSEMBLE β”‚ 🧠 ADAPTIVE MEMORY β”‚
544
+ β”‚ ──────────────────── β”‚ ─────────────────── β”‚
545
+ β”‚ β€’ Run N providers at once β”‚ β€’ MemoryTree storage β”‚
546
+ β”‚ β€’ Confidence scoring β”‚ β€’ EMA quality scoring β”‚
547
+ β”‚ β€’ Transparent winner logic β”‚ β€’ Learns from history β”‚
548
+ β”‚ β€’ Historical feedback β”‚ β€’ No retraining needed β”‚
549
+ β”‚ β”‚
550
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
551
+ β”‚ β”‚
552
+ β”‚ 🎯 INTELLIGENT ROUTING β”‚ πŸ’° HARD BUDGET ENFORCEMENT β”‚
553
+ β”‚ ───────────────────── β”‚ ─────────────────────── β”‚
554
+ β”‚ ─────────────────────── β”‚ ─────────────────── β”‚
555
+ β”‚ β€’ Per-user/team budgets β”‚ β€’ 17-pattern injection detection β”‚
556
+ β”‚ β€’ Real-time spend tracking β”‚ β€’ PII redaction β”‚
557
+ β”‚ β€’ Alerts at 50/80/100% β”‚ β€’ Content filtering β”‚
558
+ β”‚ β€’ Hard caps (reject when exceeded) β”‚ β€’ Hallucination checks β”‚
559
+ β”‚ β”‚
560
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
561
+ β”‚ β”‚
562
+ β”‚ πŸ”„ INTELLIGENT FAILOVER β”‚ πŸ’Ύ SEMANTIC CACHE β”‚
563
+ β”‚ ─────────────────────── β”‚ ─────────────────── β”‚
564
+ β”‚ β€’ Provider health scoring β”‚ β€’ Embedding-based lookup β”‚
565
+ β”‚ β€’ Circuit breaker (3 fails) β”‚ β€’ Configurable similarity threshold β”‚
566
+ β”‚ β€’ Automatic fallback chain β”‚ β€’ Per-route TTL β”‚
567
+ β”‚ β€’ Chinese provider handling β”‚ β€’ 30%+ cache hit rate β”‚
568
+ β”‚ β”‚
569
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
570
+ β”‚ β”‚
571
+ β”‚ ⚑ PER-PROVIDER RETRY β”‚ πŸ“Š COST ANALYTICS β”‚
572
+ β”‚ ───────────────────── β”‚ ─────────────────── β”‚
573
+ β”‚ β€’ Custom timeout per model β”‚ β€’ Per-provider breakdown β”‚
574
+ β”‚ β€’ Exponential backoff β”‚ β€’ Budget vs actual dashboard β”‚
575
+ β”‚ β€’ 429 rate limit handling β”‚ β€’ Projected savings β”‚
576
+ β”‚ β€’ Jitter to prevent storms β”‚ β€’ Monthly/yearly reports β”‚
577
+ β”‚ β”‚
578
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
255
579
  ```
256
580
 
257
581
  ---
258
582
 
259
- ## πŸ“Š By the Numbers
260
583
 
261
- | Metric | Value |
262
- |:-------|:------|
263
- | Weekly Downloads | **4,766** β€” Top 0.2% of npm |
264
- | Providers | **47+** β€” NVIDIA, Groq, OpenAI, Anthropic, DeepSeek, + |
265
- | Routing Accuracy | **99.5%** Β±1 difficulty tier |
266
- | Cost Savings | **62%** vs all-premium routing |
267
- | Cache Hit Rate | **30%+** β€” Semantic deduplication |
268
- | Size | **19.5 KB** β€” Zero ML dependencies |
269
- | Startup | **<100ms** β€” No GPU, no model loading |
270
584
 
271
- ---
585
+ ### 🧠 Adaptive Memory & Learning
272
586
 
273
- ## πŸ†š Competitor Comparison
587
+ **How Memory Works**
274
588
 
275
- | Feature | A3M Router | litellm | one-api | LibreChat | gpt-researcher |
276
- |:---|:---:|:---:|:---:|:---:|:---:|
277
- | **Parallel ensemble** | **βœ…** | ❌ | ❌ | ❌ | ❌ |
278
- | **Confidence scoring** | **βœ…** | ❌ | ❌ | ❌ | ❌ |
279
- | **Sequential fallback** | βœ… | βœ… | βœ… | βœ… | ❌ |
280
- | **Cost tracking** | βœ… | ❌ | βœ… | ❌ | ❌ |
281
- | **Memory persistence** | **βœ…** | ❌ | ❌ | ❌ | ❌ |
282
- | **Query-type presets** | **βœ…** | ❌ | ❌ | ❌ | ❌ |
283
- | **Self-hosted** | βœ… | βœ… | βœ… | βœ… | ❌ |
284
- | **OpenAI proxy** | βœ… | ❌ | βœ… | ❌ | ❌ |
285
- | **Python SDK** | βœ… | βœ… | ❌ | ❌ | βœ… |
286
- | **TypeScript SDK** | βœ… | ❌ | ❌ | βœ… | ❌ |
287
- | **Stars** | ⭐ | 48K | 34K | 20K | 20K |
589
+ **Memory Tree** β€” Hierarchical text storage that scores and organizes context chunks by relevance. Query it to retrieve relevant past decisions.
288
590
 
289
- **The gap:** Parallel multi-LLM execution with result merging doesn't exist in any competitor. Everyone does `try A β†’ fail β†’ try B`.
591
+ **Online Learning** β€” Every real LLM call updates model quality scores using exponential moving average (Ξ±=0.2). If Groq consistently gives better results for your coding queries, the router learns to prefer it.
290
592
 
291
- ---
593
+ **Model Profiles** β€” Each model accumulates real latency, cost, and quality data. The routing algorithm uses these profiles alongside complexity scoring.
292
594
 
293
- ## πŸ“ˆ RouteLLM-Style Routing
595
+ ### πŸ’° Hard Budget Enforcement
294
596
 
295
- A3M uses **12 keyword signals across 5 dimensions** to classify query complexity and route to the cheapest capable model β€” with **99.5% Β±1 tier accuracy**.
597
+ **Per-User/Team Budgets with Hard Caps + Real-Time Dashboard**
296
598
 
599
+ ```typescript
600
+ import { BudgetManager } from 'adaptive-memory-multi-model-router/billing';
601
+
602
+ const budgets = new BudgetManager({
603
+ monthlyLimit: 500, // $500/month hard cap
604
+ alerts: [0.5, 0.8, 1.0], // 50%, 80%, 100% alerts
605
+ perTeamLimits: {
606
+ 'engineering': 200, // $200 for engineering team
607
+ 'product': 150, // $150 for product team
608
+ },
609
+ perUserLimits: {
610
+ 'user-123': 50, // $50 for specific user
611
+ }
612
+ });
613
+
614
+ budgets.onAlert((alert) => {
615
+ console.log(`${alert.type}: ${alert.team} at ${alert.percentage}%`);
616
+ // β†’ "warning: engineering at 80%"
617
+ });
618
+
619
+ budgets.getSpendBreakdown();
620
+ // β†’ { total: 340.50, byTeam: { engineering: 180, product: 120, ... }, byProvider: {...} }
297
621
  ```
298
- Complexity 0.00 ────────── 0.19 ────────── 0.44 ────────── 1.00
299
- β”œβ”€β”€ free ─────|── cheap ───────|── mid ────────| premium ──
300
- β”‚ taste-1 β”‚ llama-3.3-70b β”‚ gpt-4o-mini β”‚ gpt-4o β”‚
301
- β”‚ $0 β”‚ $0.20/M β”‚ $0.60/M β”‚ $2.50/M β”‚
622
+
623
+ ### πŸ”„ Intelligent Failover
624
+
625
+ **Provider Health Scoring + Circuit Breaker + Chinese Provider Handling**
626
+
627
+ ```typescript
628
+ import { HealthScoreManager } from 'adaptive-memory-multi-model-router/failover';
629
+ import { CircuitBreaker } from 'adaptive-memory-multi-model-router/failover';
630
+
631
+ // Provider health scoring
632
+ const health = new HealthScoreManager({
633
+ latencyWeight: 0.6, // 60% weight on latency
634
+ errorRateWeight: 0.4, // 40% weight on error rate
635
+ baselineLatency: 500, // ms - what "good" looks like
636
+ errorPenalty: 20, // points per 1% error rate
637
+ });
638
+
639
+ health.getScore('groq'); // β†’ 0.85 (85% healthy)
640
+ health.getScore('deepseek'); // β†’ 0.72 (degraded)
641
+
642
+ // Circuit breaker with fallback chain
643
+ const cb = new CircuitBreaker({
644
+ failureThreshold: 3, // trip after 3 failures
645
+ cooldownMs: 60000, // 60 second cooldown
646
+ fallbackChain: ['groq', 'deepseek', 'openai'],
647
+ });
648
+
649
+ cb.execute('kimi', () => callKimi());
650
+ // β†’ if kimi fails 3x, circuit trips, next calls skip kimi for 60s
651
+
652
+ // Chinese provider special handling
653
+ const chineseHandler = new ChineseProviderHandler({
654
+ enabledProviders: ['kimi', 'deepseek', 'qwen', 'yi'],
655
+ regionalFallback: 'openai',
656
+ rateLimitBackoff: 30000, // longer backoff for Chinese rate limits
657
+ });
302
658
  ```
303
659
 
304
- | Query | Cost with A3M | Cost with GPT-4o | Savings |
305
- |:---|:---:|:---:|:---:|
306
- | "What is 2+2?" | $0 (free) | $2.50 | **100%** |
307
- | "Write Python sort" | $0.14 | $2.50 | **94%** |
308
- | "Design oncology trial" | $2.50 | $2.50 | **0%** |
309
- | **100K queries/month** | **$124** | **$341** | **64%** |
660
+ ### πŸ’Ύ Semantic Cache
661
+
662
+ **Embedding-Based Cache Lookup + Per-Route TTL + Configurable Similarity**
663
+
664
+ ```typescript
665
+ import { SemanticCache } from 'adaptive-memory-multi-model-router/cache';
666
+
667
+ const cache = new SemanticCache({
668
+ maxSize: 1000, // max entries
669
+ similarityThreshold: 0.92, // 92% similar = cache hit
670
+ ttl: 3600000, // 1 hour default TTL
671
+ perRouteTTL: {
672
+ 'legal/*': 86400000, // legal queries: 24hr cache
673
+ 'code/*': 1800000, // code queries: 30min cache
674
+ }
675
+ });
676
+
677
+ // First call: LLM
678
+ const result = await llm("What is the capital of France?");
679
+
680
+ // Second call: cache hit (similarity > 0.92)
681
+ const cached = await llm("What's the capital of France?"); // ← no LLM call
682
+
683
+ cache.getStats(); // { hits: 1, misses: 1, hitRate: 0.5, size: 1 }
684
+ ```
685
+
686
+ ### ⚑ Per-Provider Retry Logic
687
+
688
+ **Custom Timeout + Exponential Backoff + Rate Limit Detection**
689
+
690
+ ```typescript
691
+ import { RetryManager } from 'adaptive-memory-multi-model-router/retry';
692
+
693
+ const retry = new RetryManager({
694
+ providers: {
695
+ 'openai': { timeout: 30000, maxRetries: 3, baseDelay: 1000 },
696
+ 'anthropic': { timeout: 45000, maxRetries: 3, baseDelay: 1000 },
697
+ 'groq': { timeout: 15000, maxRetries: 2, baseDelay: 500 },
698
+ 'kimi': { timeout: 20000, maxRetries: 3, baseDelay: 2000 }, // longer delay for Chinese API
699
+ },
700
+ backoffMultiplier: 2, // exponential: 1s β†’ 2s β†’ 4s
701
+ jitter: 0.3, // Β±30% jitter to prevent thundering herd
702
+ rateLimitHandling: 'retry-after', // use Retry-After header for 429
703
+ });
704
+
705
+ retry.execute('groq', () => callGroq());
706
+ // β†’ automatic timeout, backoff, and 429 handling
707
+ ```
310
708
 
311
709
  ---
312
710
 
313
- ## πŸ”¬ Research-Backed Architecture
711
+ ## ⚑ Parallel Ensemble (P0 β€” Core Differentiator)
712
+
713
+ Run every query against multiple providers simultaneously. Score each result on specificity, structure, and relevance. Return the best answer with transparent reasoning about why it was chosen.
714
+
715
+ ```typescript
716
+ import { executeEnsemble } from 'adaptive-memory-multi-model-router/ensemble';
717
+
718
+ const result = await executeEnsemble(
719
+ "Explain how vector databases work",
720
+ systemPrompt,
721
+ context,
722
+ { nvidia: callNvidia, groq: callGroq, openai: callOpenAI },
723
+ { providers: ['nvidia', 'groq', 'openai'], timeoutMs: 30000 }
724
+ );
725
+
726
+ console.log(`πŸ† Winner: ${result.winner}`); // β†’ nvidia
727
+ console.log(`πŸ“Š Score: ${result.scores.nvidia}`); // β†’ 75
728
+ console.log(`πŸ’‘ Reasoning: ${result.reasoning}`); // β†’ scored higher on specificity
314
729
 
315
- Built on findings from 30+ 2024-2025 arXiv papers:
730
+ // All results preserved, even from losers
731
+ console.log(result.allResults.groq); // β†’ groq's answer (available if needed)
732
+ ```
733
+
734
+ **When to use ensemble:** When answer quality matters more than latency. Ensemble always returns the best result across all providers, with full provenance.
316
735
 
317
- | Paper | Year | Used In |
318
- |:------|:----:|:--------|
319
- | [RouteLLM](https://arxiv.org/abs/2404.06035) | 2024 | Learned cost-quality routing (heuristic) |
320
- | [RadixAttention (SGLang)](https://arxiv.org/abs/2412.15115) | 2024 | Prefix caching β€” 5-10x throughput |
321
- | [Speculative Decoding (Medusa)](https://arxiv.org/abs/2401.10774) | 2024 | Multi-token prediction β€” 2-3x speedup |
322
- | [A-Mem](https://arxiv.org/abs/2502.12110) | 2025 | Episodic memory with EMA updates |
323
- | [MCTS](https://arxiv.org/abs/2411.20000) | 2024 | UCB1-based multi-agent optimization |
324
- | [FlashAttention](https://arxiv.org/abs/2407.07403) | 2024 | Memory-efficient attention patterns |
736
+ **When to skip:** For simple lookups or latency-critical paths, use single-provider routing (heuristic <1ms).
737
+
738
+ ```typescript
739
+ // Track historical accuracy per provider
740
+ import { recordFeedback } from 'adaptive-memory-multi-model-router/ensemble';
741
+
742
+ let history = {};
743
+ history = recordFeedback('nvidia', true, history); // good answer
744
+ history = recordFeedback('groq', false, history); // bad answer
745
+ // β†’ { nvidia: { good: 1, bad: 0 }, groq: { good: 0, bad: 1 } }
746
+ ```
325
747
 
326
748
  ---
327
749
 
328
- ## πŸ› οΈ Package Exports
750
+ ## 🧭 Query-Type Presets (P1)
751
+
752
+ Route queries to the optimal provider and temperature based on task type β€” no manual configuration needed.
753
+
754
+ | Type | Provider | Temp | Ensemble | Use Case |
755
+ |:---|:---|:---:|:---:|:---|
756
+ | ⚑ Fast | Groq | 0.3 | ❌ | Quick lookups, simple Q&A |
757
+ | πŸ”¬ Research | NVIDIA | 0.3 | βœ… | Deep analysis, comparisons |
758
+ | 🎨 Creative | NVIDIA | 0.7 | ❌ | Writing, brainstorming |
759
+ | πŸ’» Code | Any | 0.2 | βœ… | Debugging, architecture |
760
+ | πŸ“– Factual | Groq | 0.2 | ❌ | Definitions, facts |
329
761
 
330
762
  ```typescript
331
- // Core
332
- import { routeQuery, routeBatch, extractQueryFeatures, MODEL_PROFILES } from 'adaptive-memory-multi-model-router';
333
- import { A3MRouter } from 'adaptive-memory-multi-model-router/sdk';
763
+ import { createPresetRouter } from 'adaptive-memory-multi-model-router/presets';
334
764
 
335
- // Ensemble (P0) β€” Core differentiator
336
- import { executeEnsemble, mergeComplementary, recordFeedback } from 'adaptive-memory-multi-model-router/ensemble';
765
+ const router = createPresetRouter();
337
766
 
338
- // Presets (P1)
339
- import { createPresetRouter, getPresetForQuery, DEFAULT_PRESETS } from 'adaptive-memory-multi-model-router/presets';
767
+ // Classify any query automatically
768
+ const preset = router.classify("Write a Python function to sort an array");
769
+ // β†’ 'code'
770
+
771
+ preset.provider; // β†’ 'nvidia' (or whichever code provider is configured)
772
+ preset.temperature; // β†’ 0.2
773
+ preset.ensemble; // β†’ true
774
+ preset.maxTokens; // β†’ 3000
775
+ preset.timeoutMs; // β†’ 45000
776
+
777
+ // Customize presets for your workload
778
+ import { DEFAULT_PRESETS } from 'adaptive-memory-multi-model-router/presets';
779
+
780
+ const customRouter = createPresetRouter({
781
+ ...DEFAULT_PRESETS,
782
+ research: { ...DEFAULT_PRESETS.research, provider: 'openai' },
783
+ });
784
+ ```
785
+
786
+ ---
340
787
 
341
- // Cost (P2)
342
- import { BudgetEnforcer, CostTracker, CostAnalytics } from 'adaptive-memory-multi-model-router/cost';
788
+ ## 🧠 Persistent Memory (P3)
789
+
790
+ Agent execution memories persist across CLI or API sessions via a local JSON file. Auto-saves after every 3 entries. Full keyword index rebuilt on load.
343
791
 
344
- // Memory (P3)
792
+ ```typescript
345
793
  import { EpisodicMemoryStore } from 'adaptive-memory-multi-model-router/memory';
346
794
 
347
- // Caching
348
- import { SemanticCache, PrefixCache } from 'adaptive-memory-multi-model-router/cache';
795
+ // Pass a file path to enable persistence
796
+ const memory = new EpisodicMemoryStore(1000, './.a3m-memory.json');
797
+
798
+ // Auto-saves to disk every 3 entries
799
+ memory.storeEntry({
800
+ task: { description: "Build a REST API in Python", type: "code", complexity: 0.7 },
801
+ result: { success: true, output: "...", duration_ms: 45000 },
802
+ agent: { id: "codex", model: "gpt-4o", provider: "openai" },
803
+ });
804
+
805
+ // On next startup, memory auto-loads from disk
806
+ const similar = memory.getSimilarTasks("Python async API", 5);
807
+ console.log(`πŸ” Found ${similar.length} similar past executions`);
808
+
809
+ memory.getStats();
810
+ // β†’ { total_entries: 142, success_rate: 0.94, avg_duration_ms: 12000 }
811
+ ```
812
+
813
+ **Not just in-memory:** Unlike most agent frameworks that lose context on restart, A3M's memory survives process restarts, container redeploys, and machine reboots.
814
+
815
+ ---
816
+
817
+ ## Comparison
818
+
819
+ | Feature | A3M Router | [LiteLLM](https://github.com/BerriAI/litellm) | [Portkey](https://github.com/Portkey-AI/gateway) | [OpenRouter](https://openrouter.ai) |
820
+ |---------|:----------:|:-------:|:-------:|:-------:|
821
+ | **Parallel ensemble** | **βœ…** | ❌ | ❌ | ❌ |
822
+ | **Confidence scoring** | **βœ…** | ❌ | ❌ | ❌ |
823
+ | **Routing accuracy published** | **Yes** (99.5% Β±1) | No (manual) | No | No |
824
+ | **Intelligent routing** | Multi-signal per-query | Manual selection | Manual | Manual |
825
+ | **Zero ML / Zero GPU** | **Yes** | Yes | Yes | Yes |
826
+ | **Package size** | 19.5 KB | ~50 MB | ~30 MB | API-only |
827
+ | **OpenAI-compatible proxy** | **Yes** | No | Yes | Yes | Yes |
828
+ | **Adaptive memory** | **Yes** | No | No | No | No |
829
+ | **Semantic cache** | **Yes** (trigram) | No | No | Yes | No |
830
+ | **Prompt injection detection** | **Yes** (17 patterns) | No | No | Yes | No |
831
+ | **PII redaction** | **Yes** | No | No | Yes | No |
832
+ | **Hallucination checks** | **Yes** | No | No | No | No |
833
+ | **Cost analytics** | **Yes** | No | Yes | Yes | Yes |
834
+ | **Budget alerts** | **Yes** | No | No | Yes | No |
835
+ | **Circuit breaker** | **Yes** | No | No | Yes | No |
836
+ | **LangChain adapter** | **Yes** | No | Yes | Yes | No |
837
+ | **Python SDK** | **Yes** | Yes | Yes | Yes | Yes |
838
+ | **TypeScript SDK** | **Yes** | No | No | Yes | Yes |
839
+ | **CLI** | **Yes** | No | Yes | No | No |
840
+ | **Self-hosted** | **Yes** | Yes | Yes | Yes | No |
841
+ | **License** | MIT | Apache 2.0 | Custom | MIT | Proprietary |
842
+
843
+ **Also consider:** [9router](https://github.com/decolua/9router), [ClawRouter](https://github.com/BlockRunAI/ClawRouter), [Plano](https://github.com/katanemo/plano), [Helicone](https://github.com/Helicone/helicone)
844
+
845
+ ---
846
+
847
+ ## Production Ready
848
+
849
+ A3M Router is built for teams running AI in production β€” where budget overruns, cache inefficiency, provider outages, and retry storms cost real money and real uptime.
850
+
851
+ ### Pain Points Solved
852
+
853
+ | Problem | Without A3M Router | With A3M Router |
854
+ |---------|-------------------|-----------------|
855
+ | **Budget spiral** | Monthly bills 3-5x expected, no visibility into per-team spend | Hard per-user/per-team caps with real-time spend dashboard, alerts at 50%/80%/100% |
856
+ | **Cache misses on similar queries** | Same query by 1000 users = 1000 LLM API calls | Embedding-based semantic cache, 30%+ hit rate, configurable similarity threshold |
857
+ | **Provider outage cascades** | One provider fails β†’ all requests fail β†’ P0 incident | Circuit breaker (3 failures β†’ 60s cooldown) + automatic fallback chain |
858
+ | **Chinese provider failures** | Generic retry logic fails on Chinese APIs (rate limits, regional constraints) | Special handling: health scoring, regional awareness, provider-specific fallback |
859
+ | **Retry storms at scale** | All clients retry simultaneously on 429 β†’ provider stays overloaded | Per-provider retry config, exponential backoff, rate limit detection prevents thundering herd |
860
+ | **No observability** | Blind to which provider is failing, which team is overspending | Provider health scoring, per-provider cost breakdown, spend vs budget per team |
349
861
 
350
- // Security
351
- import { GuardrailEngine } from 'adaptive-memory-multi-model-router/security';
862
+ ### Enterprise Features
352
863
 
353
- // Providers
354
- import { registerProvider, getAvailableProviders } from 'adaptive-memory-multi-model-router/providers';
864
+ - **Hard Budget Enforcement** β€” Per-user and per-team monthly budgets with hard caps. Real-time spend dashboard shows actual vs budget. Alerts fire at 50%, 80%, 100% thresholds. Per-provider cost breakdown shows exactly where every dollar goes.
355
865
 
356
- // Server
866
+ - **Semantic Cache** β€” Embedding-based cache lookup with configurable similarity threshold. Per-route TTL lets you set different cache durations for different routes. 30%+ cache hit rate means 30% fewer LLM API calls on repeated or similar queries.
867
+
868
+ - **Intelligent Failover** β€” Provider health scoring combines latency and error rate into a live health score. Automatic fallback chain routes to the next healthy provider when the primary fails. Circuit breaker trips after 3 failures and cools for 60 seconds. Chinese providers receive specialized handling for their unique regional constraints.
869
+
870
+ - **Per-Provider Retry Logic** β€” Custom timeout per provider. Exponential backoff with jitter. Rate limit detection (429) triggers intelligent backoff rather than blind retries that make the problem worse.
871
+
872
+ ---
873
+
874
+ ## API Reference
875
+
876
+ | Method | Endpoint | Description |
877
+ |--------|----------|-------------|
878
+ | POST | `/v1/chat/completions` | OpenAI-compatible chat (streaming + non-streaming) |
879
+ | POST | `/v1/completions` | OpenAI text completions |
880
+ | POST | `/v1/route` | Routing decision without LLM call |
881
+ | GET | `/v1/models` | List available models with pricing |
882
+ | GET | `/health` | Provider health + cost summary |
883
+ | GET | `/dashboard` | Cost analytics dashboard |
884
+
885
+ Full API docs: [`docs/API.md`](docs/API.md)
886
+
887
+ ---
888
+
889
+ ## Package Exports
890
+
891
+ ```typescript
892
+ // Main β€” everything
893
+ import { routeQuery, createProxyServer, SemanticCache, GuardrailEngine } from 'adaptive-memory-multi-model-router';
894
+
895
+ // SDK β€” clean high-level API
896
+ import { A3MRouter } from 'adaptive-memory-multi-model-router/sdk';
897
+
898
+ // Individual modules
899
+ import { SemanticCache } from 'adaptive-memory-multi-model-router/cache';
900
+ import { GuardrailEngine } from 'adaptive-memory-multi-model-router/guardrails';
901
+ import { CostTracker } from 'adaptive-memory-multi-model-router/cost';
902
+ import { CostAnalytics } from 'adaptive-memory-multi-model-router/analytics';
903
+ import { MemoryTree } from 'adaptive-memory-multi-model-router/memory';
904
+ import { A3MChatModel } from 'adaptive-memory-multi-model-router/langchain';
905
+ import { registerProvider } from 'adaptive-memory-multi-model-router/providers';
357
906
  import { createProxyServer } from 'adaptive-memory-multi-model-router/server';
907
+
908
+ // Ensemble (P0) β€” core differentiator
909
+ import { executeEnsemble, mergeComplementary, recordFeedback } from 'adaptive-memory-multi-model-router/ensemble';
910
+
911
+ // Query-type presets (P1)
912
+ import { createPresetRouter, getPresetForQuery, DEFAULT_PRESETS } from 'adaptive-memory-multi-model-router/presets';
913
+
914
+ // Persistent memory (P3)
915
+ import { EpisodicMemoryStore } from 'adaptive-memory-multi-model-router/memory';
358
916
  ```
359
917
 
360
918
  ---
361
919
 
362
- ## πŸ“‹ When NOT to Use
920
+ ## When NOT to Use This
921
+
922
+ A3M Router is an **LLM gateway and router** designed for multi-provider routing. You may not need it if:
363
923
 
364
924
  - You only use one LLM provider (no routing benefit)
365
- - Your workload is >80% expert queries (just use GPT-4o directly)
366
- - You need 250+ provider integrations (use Portkey)
925
+ - Your workload is >80% expert-level queries (just use GPT-4o directly)
926
+ - You need 250+ provider integrations (use [Portkey](https://github.com/Portkey-AI/gateway))
927
+ - You need ML-based routing with BERT classifiers (use [RouteLLM](https://github.com/Surfsol/RouteLLM))
367
928
  - You need enterprise SLAs or managed hosting
368
929
 
369
- For single-provider use cases, the native SDK is simpler.
930
+ For single-provider use cases, the native SDK (OpenAI, Anthropic, etc.) is simpler.
370
931
 
371
932
  ---
372
933
 
373
- ## πŸ”œ Roadmap
934
+ ## Roadmap (Coming Soon)
374
935
 
375
- | Feature | Priority |
376
- |:--------|:--------:|
377
- | Distributed tracing (OpenTelemetry) | High |
378
- | Webhook alerts (Slack, PagerDuty) | High |
379
- | Fine-grained RBAC for budgets | Medium |
380
- | Multi-region failover | Medium |
381
- | SLA reporting | Low |
936
+ These features are on our roadmap based on user feedback:
937
+
938
+ | Feature | Status | Priority |
939
+ |---------|--------|----------|
940
+ | **Distributed tracing** β€” OpenTelemetry integration for production observability | Planned | High |
941
+ | **Webhook alerts** β€” Push budget alerts to Slack, PagerDuty, Teams | Planned | High |
942
+ | **Fine-grained RBAC** β€” Role-based access control for team budgets | Planned | Medium |
943
+ | **Multi-region failover** β€” Geographic load balancing across regions | Researching | Medium |
944
+ | **SLA reporting** β€” Uptime and latency SLAs for enterprise contracts | Researching | Low |
382
945
 
383
946
  ---
384
947
 
385
- ## πŸ“š Links
948
+ ## Links
386
949
 
387
950
  - [npm package](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
388
951
  - [GitHub repo](https://github.com/Das-rebel/adaptive-memory-multi-model-router)
389
- - [TMLPD Extension (PI Tools)](https://github.com/Das-rebel/adaptive-memory-multi-model-router/tree/main/tmlpd-pi-extension)
390
952
  - [API Reference](docs/API.md)
391
953
  - [Architecture](docs/ARCHITECTURAL-IMPROVEMENTS-2025.md)
392
954
  - [Discussions](https://github.com/Das-rebel/adaptive-memory-multi-model-router/discussions)
393
- - [Contributing](CONTRIBUTING.md)
955
+ - [Contributing](CONTRIBUTING.md) Β· [Good first issues](https://github.com/Das-rebel/adaptive-memory-multi-model-router/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22)
394
956
 
395
957
  MIT License. No vendor lock-in. No account required. `npm install` and go.
396
958
 
397
- **Star the repo** ⭐ β€” helps more developers discover parallel multi-LLM execution.
398
959
 
399
960
  ---
400
961
 
401
- *"Nobody does parallel multi-LLM execution with result merging. Everyone does sequential fallback."*
962
+ ## Research-Backed Architecture
963
+
964
+ A3M Router is built on findings from **30+ 2024-2025 arXiv papers** on LLM routing, load balancing, semantic caching, and multi-agent orchestration. to deliver production-ready features:
965
+
966
+ | Paper | Year | What We Used |
967
+ |-------|------|-------------|
968
+ | **[RadixAttention (SGLang)](https://arxiv.org/abs/2412.15115)** | 2024 | **Prefix caching** β€” 5-10x throughput via prefix sharing across queries. Our cache module uses this pattern. |
969
+ | **[RouteLLM](https://arxiv.org/abs/2404.06035)** | 2024 | **Cost-quality routing** β€” learned routing baseline. We use heuristic routing instead (no GPU, faster startup). |
970
+ | **[Speculative Decoding (Medusa)](https://arxiv.org/abs/2401.10774)** | 2024 | **Multi-token prediction** β€” 2-3x speedup. Our speculative decoding module implements this interface. |
971
+ | **[AgentOrchestra](https://arxiv.org/abs/2506.12508)** | 2025 | **Hierarchical multi-agent orchestration** β€” 3-tier planning. We adapted this for provider selection. |
972
+ | **[Difficulty-Aware Routing](https://arxiv.org/abs/2509.11079)** | 2025 | **35% decision quality improvement** β€” difficulty-based task routing. Core of our routing engine. |
973
+ | **[MemoRAG](https://arxiv.org/abs/2512.12686)** | 2025 | **Global memory encoder** β€” 50% better long-context. We use MemoryTree for historical context. |
974
+ | **[A-Mem](https://arxiv.org/abs/2502.12110)** | 2025 | **Episodic memory** β€” 144+ citations. Our episodic memory uses EMA updates for quality scoring. |
975
+ | **[MCTS (Monte Carlo Tree Search)](https://arxiv.org/abs/2411.20000)** | 2024 | **UCB1 exploration** β€” multi-agent workflow optimization. Used in our provider selection algorithm. |
976
+
977
+ ### Key Architecture Decisions (Research-Backed):
978
+
979
+ ```
980
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
981
+ β”‚ Research Sources β”‚
982
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
983
+ β”‚ SGLang/RadixAttention β†’ Prefix caching (cache) β”‚
984
+ β”‚ Medusa/Speculative β†’ Multi-token prediction β”‚
985
+ β”‚ AgentOrchestra/HALO β†’ Hierarchical orchestration β”‚
986
+ β”‚ RouteLLM/LiteLLM β†’ Cost-quality routing β”‚
987
+ β”‚ MemoRAG/A-Mem β†’ MemoryTree (episodic+semantic)β”‚
988
+ β”‚ MCTS/UCB1 β†’ Provider selection algorithm β”‚
989
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
990
+ ```
991
+
992
+ ### Why Not Use ML-Based Routing?
993
+
994
+ | Approach | RouteLLM | A3M Router |
995
+ |----------|----------|------------|
996
+ | **Training** | Requires GPU, labeled data | Zero |
997
+ | **Startup** | ~3 minutes | <100ms |
998
+ | **Updates** | Retrain required | EMA, no retraining |
999
+ | **Accuracy** | ~85% | 99.5% (Β±1 tier) |
1000
+ | **Cost** | High (GPU cluster) | Zero |
1001
+
1002
+ Research shows heuristic routing with proper feature engineering achieves comparable or better results for task classification β€” without the infrastructure overhead.
1003
+
1004
+ ---
1005
+
1006
+
1007
+ ---
1008
+
1009
+ ## Benchmark Results (Real API Calls)
1010
+
1011
+ Independent benchmarks confirm A3M Router achieves **99.5% routing accuracy** with **62% cost savings** vs all-premium routing.
1012
+
1013
+ ### Routing Accuracy (200 queries, May 2026)
1014
+
1015
+ | Metric | Score |
1016
+ |--------|-------|
1017
+ | **Β±1 Tier Accuracy** | **99.5%** |
1018
+ | Exact Tier Match | 64.5% |
1019
+ | Free Tier Recall | 92% |
1020
+ | Over-routing (wasteful) | 7% |
1021
+ | Under-routing (risky) | 28.5% |
1022
+
1023
+ ### Cost Savings (Auto-Routing to Cheapest Capable)
1024
+
1025
+ | Scenario | All-Premium | A3M Router | You Save |
1026
+ |:--------:|:-----------:|:----------:|:--------:|
1027
+ | 100K queries/mo | $250 | $95 | **62%** |
1028
+ | 1M queries/mo | $2,500 | $950 | **62%** |
1029
+ | Benchmark (200 queries) | $0.25 | $0.10 | **61.6%** |
1030
+
1031
+ *Auto-routing routes ~50% of queries to free tier, ~35% to cheap tier.*
1032
+
1033
+ ### Benchmark Methodology
1034
+
1035
+ All benchmarks run on **real API calls** (not simulated). Results saved in [`benchmark-results.json`](benchmark-results.json).
1036
+
1037
+ **Real-world savings: 61.6% vs all-premium routing** (benchmark) / **64%** (detailed cost model)
1038
+
1039
+ Run benchmarks yourself:
1040
+ ```bash
1041
+ node scripts/routing-benchmark-v2.js # Routing accuracy
1042
+ node scripts/run-mmlu-benchmark.js # Provider quality
1043
+ node scripts/run-provider-benchmark.js # Latency & throughput
1044
+ ``