adaptive-memory-multi-model-router 2.16.2 → 2.16.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.github/CODEOWNERS +2 -0
- package/.github/FUNDING.yml +7 -1
- package/.github/ISSUE_TEMPLATE/bug_report.md +56 -0
- package/.github/ISSUE_TEMPLATE/feature_request.md +41 -0
- package/.github/workflows/pages.yml +1 -1
- package/CHANGELOG.md +15 -8
- package/CONTRIBUTING.md +109 -25
- package/README.md +131 -219
- package/data/jev-distill.jsonl +320 -0
- package/dist/routing/jev/jevRouter.d.ts +49 -0
- package/dist/routing/jev/jevRouter.js +231 -0
- package/dist/routing/jev/jevRouter.js.map +1 -0
- package/dist/routing/jev/optionAttention.d.ts +63 -0
- package/dist/routing/jev/optionAttention.js +158 -0
- package/dist/routing/jev/optionAttention.js.map +1 -0
- package/dist/routing/jev/remote.d.ts +14 -0
- package/dist/routing/jev/remote.js +55 -0
- package/dist/routing/jev/remote.js.map +1 -0
- package/dist/routing/jev/types.d.ts +71 -0
- package/dist/routing/jev/types.js +19 -0
- package/dist/routing/jev/types.js.map +1 -0
- package/dist/routing/jev/weights/jev-router-weights.json +1 -0
- package/dist/server/modelMapper.js +18 -0
- package/dist/server/modelMapper.js.map +1 -1
- package/docs/assets/og-banner.svg +193 -0
- package/docs/index.html +1523 -465
- package/docs-site/index.html +1427 -563
- package/package.json +29 -20
- package/python/pyproject.toml +1 -1
- package/src/cli/setupWizard.ts +4 -2
- package/src/routing/jev/jevRouter.ts +251 -0
- package/src/routing/jev/optionAttention.ts +186 -0
- package/src/routing/jev/remote.ts +50 -0
- package/src/routing/jev/types.ts +79 -0
- package/src/routing/jev/weights/jev-router-weights.json +1 -0
- package/src/server/modelMapper.ts +17 -0
- package/tests/routing/jev.test.ts +110 -0
- package/tools/calibrate_temp.py +81 -0
- package/tools/distill.mjs +143 -0
- package/tools/train_jev.py +213 -0
- package/dist/cli/tui.d.ts +0 -6
- package/dist/cli/tui.js.map +0 -1
- package/dist/routing/shadowSampler.d.ts.map +0 -1
- /package/{ARCHITECTURE.md → archive/ARCHITECTURE.md} +0 -0
- /package/{ENTERPRISE_INTEGRATIONS.md → archive/ENTERPRISE_INTEGRATIONS.md} +0 -0
- /package/{MANIFESTO.md → archive/MANIFESTO.md} +0 -0
- /package/{README_ja.md → archive/README_ja.md} +0 -0
- /package/{README_zh.md → archive/README_zh.md} +0 -0
- /package/{SECURITY.md → archive/SECURITY.md} +0 -0
- /package/{TECHNICAL_README.md → archive/TECHNICAL_README.md} +0 -0
- /package/{TODO_BROWSER_AUTOMATION.md → archive/TODO_BROWSER_AUTOMATION.md} +0 -0
- /package/{AGENT_COUNCIL_FINDINGS.md → archive/campaign/AGENT_COUNCIL_FINDINGS.md} +0 -0
- /package/{AUDIT_REPORT.md → archive/campaign/AUDIT_REPORT.md} +0 -0
- /package/{CONTRIBUTORS.md → archive/campaign/CONTRIBUTORS.md} +0 -0
- /package/{IMPROVEMENT_PLAN.md → archive/campaign/IMPROVEMENT_PLAN.md} +0 -0
- /package/{INTEGRATION_PROGRESS.md → archive/campaign/INTEGRATION_PROGRESS.md} +0 -0
- /package/{CAMPAIGN_SUMMARY.md → archive/launch/CAMPAIGN_SUMMARY.md} +0 -0
- /package/{LANDING.md → archive/launch/LANDING.md} +0 -0
- /package/{LAUNCH-PAIN-DRIVEN.md → archive/launch/LAUNCH-PAIN-DRIVEN.md} +0 -0
- /package/{LAUNCH.md → archive/launch/LAUNCH.md} +0 -0
- /package/{LAUNCH_CHECKLIST.md → archive/launch/LAUNCH_CHECKLIST.md} +0 -0
- /package/{LAUNCH_SNAPSHOT.md → archive/launch/LAUNCH_SNAPSHOT.md} +0 -0
- /package/{REDESIGN.md → archive/launch/REDESIGN.md} +0 -0
- /package/{HEALTH_REPORT.md → archive/research/HEALTH_REPORT.md} +0 -0
- /package/{OPPORTUNITIES_100.md → archive/research/OPPORTUNITIES_100.md} +0 -0
- /package/{POPULARITY_BOOSTERS.md → archive/research/POPULARITY_BOOSTERS.md} +0 -0
- /package/{PR_STATUS_REPORT.md → archive/research/PR_STATUS_REPORT.md} +0 -0
- /package/{research-log.md → archive/research/research-log.md} +0 -0
- /package/{RELEASE_v2.16.0.md → archive/submissions/RELEASE_v2.16.0.md} +0 -0
- /package/{RUNKIT.md → archive/submissions/RUNKIT.md} +0 -0
- /package/{SUBMISSIONS.md → archive/submissions/SUBMISSIONS.md} +0 -0
- /package/{a3m-integrations-summary.md → archive/submissions/a3m-integrations-summary.md} +0 -0
- /package/{discoverability-diagnosis.md → archive/submissions/discoverability-diagnosis.md} +0 -0
package/README.md
CHANGED
|
@@ -1,312 +1,224 @@
|
|
|
1
|
-
#
|
|
1
|
+
# LLM Routing That Cuts Your AI Bill by 90%
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
**GPT-4o costs $0.03/run. A3M routes the same request to Groq/Mistral for $0.0001.**
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
A3M Router is the open-source answer. Built on 3 billion years of biological intelligence.
|
|
10
|
-
|
|
11
|
-
---
|
|
5
|
+
```python
|
|
6
|
+
# Before: Expensive and slow
|
|
7
|
+
response = openai.ChatCompletion.create(model="gpt-4o", messages=[...])
|
|
8
|
+
# $0.03 per request. Every time.
|
|
12
9
|
|
|
13
|
-
|
|
10
|
+
# After: Same API, 99.7% cheaper
|
|
11
|
+
from openai import OpenAI
|
|
12
|
+
client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
|
|
13
|
+
response = client.chat.completions.create(model="auto", messages=[...])
|
|
14
|
+
# Routes to cheapest capable provider. $0.0001 per request.
|
|
15
|
+
```
|
|
14
16
|
|
|
15
|
-
[](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
|
|
16
|
-
[](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
|
|
17
|
+
[](https://www.npmjs.com/npm/package/adaptive-memory-multi-model-router)
|
|
18
|
+
[](https://www.npmjs.com/npm/package/adaptive-memory-multi-model-router)
|
|
17
19
|
[](https://pypi.org/project/a3m-router/)
|
|
18
20
|
[](LICENSE)
|
|
19
|
-
[](https://github.com/Das-rebel/a3m-router/actions)
|
|
20
22
|
[](https://github.com/Das-rebel/a3m-router/stargazers)
|
|
21
23
|
|
|
22
|
-
**📦 Available on:**
|
|
23
|
-
[](https://www.npmjs.com/package/adaptive-memory-multi-model-router)
|
|
24
|
-
[](https://pypi.org/project/a3m-router/)
|
|
25
|
-
[](https://github.com/Das-rebel/a3m-router/stargazers)
|
|
26
|
-
|
|
27
|
-
---
|
|
28
|
-
|
|
29
|
-
## The Problem with Centralization
|
|
30
|
-
|
|
31
|
-
When one company controls the "dollar of intelligence," what happens to innovation?
|
|
32
|
-
|
|
33
|
-
History offers cautionary tales. When GitHub was acquired by Microsoft, forks proliferated. GitLab gained market share. The acquirer's brand became a liability for some users.
|
|
34
|
-
|
|
35
|
-
The same dynamic plays out here. A segment of OpenRouter's user base will start asking: **"Is there an open-source alternative?"**
|
|
36
|
-
|
|
37
|
-
**We're that alternative.** Not "better" — a different philosophy.
|
|
38
|
-
|
|
39
|
-
---
|
|
40
|
-
|
|
41
|
-
## Biology-Inspired Intelligence
|
|
42
|
-
|
|
43
|
-
Nature has been solving the routing problem for 3 billion years. Here's what we borrowed:
|
|
44
|
-
|
|
45
|
-
### 🐜 Swarm Intelligence → 99.99% Uptime
|
|
46
|
-
|
|
47
|
-
Ants never ask for directions. Yet colonies reliably find the shortest paths to food.
|
|
48
|
-
|
|
49
|
-
How? **Pheromone trails.** Each request leaves a trail. If a model fails, its trail weakens and requests avoid it. New paths emerge automatically.
|
|
50
|
-
|
|
51
|
-
This is how A3M Router achieves 99.99% uptime. Not one giant brain managing everything — millions of tiny smart decisions adding up to a resilient whole.
|
|
52
|
-
|
|
53
|
-
### 🧠 Neural Plasticity → Adaptive Learning
|
|
54
|
-
|
|
55
|
-
Your brain isn't static. It constantly rewires, strengthening used pathways and pruning unused ones.
|
|
56
|
-
|
|
57
|
-
A3M Router does the same: **time-decayed weights** prevent overfitting to outdated provider behavior. Recent performance matters more than old data.
|
|
58
|
-
|
|
59
|
-
Always learning. Always adapting. Never stuck in the past.
|
|
60
|
-
|
|
61
|
-
### 📊 Competitive Exclusion → Diversity
|
|
62
|
-
|
|
63
|
-
In nature, no species can dominate indefinitely. Success creates conditions for others to challenge it.
|
|
64
|
-
|
|
65
|
-
A3M Router implements **diversity penalty** (EXP3 algorithm). Higher market share = bigger penalty = natural equilibrium.
|
|
66
|
-
|
|
67
|
-
No monoculture. The plankton paradox solved.
|
|
68
|
-
|
|
69
|
-
### 🦚 Handicap Principle → Cost as Signal
|
|
70
|
-
|
|
71
|
-
Why does a peacock have an extravagant tail? Expensive signals are more credible. A peacock that survives despite its handicap must be truly exceptional.
|
|
72
|
-
|
|
73
|
-
A3M Router sees cost as a **credibility signal**. High cost = high computational investment = better quality for high-stakes queries.
|
|
74
|
-
|
|
75
|
-
Intelligent resource allocation based on task criticality.
|
|
76
|
-
|
|
77
24
|
---
|
|
78
25
|
|
|
79
|
-
## TL;DR — What Is This?
|
|
80
|
-
|
|
81
|
-
**Before:**
|
|
82
|
-
```python
|
|
83
|
-
# Pay GPT-4o prices for EVERY query
|
|
84
|
-
client = OpenAI(api_key="sk-...")
|
|
85
|
-
response = client.chat.completions.create(
|
|
86
|
-
model="gpt-4o",
|
|
87
|
-
messages=[{"role": "user", "content": "What is 2+2?"}]
|
|
88
|
-
) # Costs: $0.03
|
|
89
26
|
```
|
|
27
|
+
$ npx a3m-router serve
|
|
28
|
+
___ ___ ____ ____ _ _ __ ___
|
|
29
|
+
/ __)/ \/ ___)( __)( \/ )( )( _)
|
|
30
|
+
( (__( O ))__) ) _) ) ( )(/(
|
|
31
|
+
\___)\__/(____)(____)(_)\_)(____/
|
|
90
32
|
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
# A3M Router picks the right model automatically
|
|
94
|
-
client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
|
|
95
|
-
response = client.chat.completions.create(
|
|
96
|
-
model="auto", # ← Just change this
|
|
97
|
-
messages=[{"role": "user", "content": "What is 2+2?"}]
|
|
98
|
-
) # Routes to Groq/Mistral — costs: $0.0001
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
**Result:** Simple questions cost 300x less. Complex queries still go to premium models when needed.
|
|
102
|
-
|
|
103
|
-
---
|
|
33
|
+
A3M Router v2.16.3
|
|
34
|
+
Serving at http://localhost:8787/v1
|
|
104
35
|
|
|
105
|
-
|
|
36
|
+
Providers: 80+ | Mode: auto | Memory: enabled
|
|
106
37
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
| **Latency (P99)** | **162ms** | 189ms | **14% faster** |
|
|
112
|
-
| **Cost per 1K tokens** | **$0.00012** | $0.0015 | **92% cheaper** |
|
|
113
|
-
| **Quality Score** | **94%** | 92% | **2% better** |
|
|
114
|
-
| **Provider Coverage** | **80+** | 45 | **78% more** |
|
|
115
|
-
| **Uptime** | **99.99%** | Provider-dependent | **Always on** |
|
|
116
|
-
|
|
117
|
-
### Cost by Query Type
|
|
118
|
-
|
|
119
|
-
| Query Type | GPT-4o Cost | A3M Router Cost | Savings |
|
|
120
|
-
|------------|-------------|-----------------|---------|
|
|
121
|
-
| "What is 2+2?" | $0.03 | $0.0001 (Groq) | **99.7%** |
|
|
122
|
-
| "Write a Python function" | $0.05 | $0.002 (DeepSeek) | **96%** |
|
|
123
|
-
| "Design a database schema" | $0.15 | $0.008 (Mixed) | **95%** |
|
|
124
|
-
| "Complex reasoning" | $0.15 | $0.15 (GPT-4o) | **0%** (correctly routed) |
|
|
38
|
+
→ POST /v1/chat/completions
|
|
39
|
+
→ GET /v1/models
|
|
40
|
+
→ GET /health
|
|
41
|
+
```
|
|
125
42
|
|
|
126
43
|
---
|
|
127
44
|
|
|
128
|
-
##
|
|
45
|
+
## Get Started in 30 Seconds
|
|
129
46
|
|
|
130
47
|
```bash
|
|
131
|
-
# npm
|
|
132
48
|
npm install adaptive-memory-multi-model-router
|
|
133
49
|
npx a3m-router serve
|
|
134
50
|
|
|
135
|
-
#
|
|
136
|
-
pip install a3m-router
|
|
137
|
-
python -m a3m_router.serve
|
|
138
|
-
|
|
139
|
-
# Docker
|
|
140
|
-
docker run -p 8787:8787 ghcr.io/das-rebel/a3m-router
|
|
51
|
+
# Then use it like OpenAI:
|
|
141
52
|
```
|
|
142
53
|
|
|
143
|
-
Then use it like any OpenAI-compatible API:
|
|
144
|
-
|
|
145
54
|
```python
|
|
146
55
|
from openai import OpenAI
|
|
147
|
-
|
|
148
56
|
client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
|
|
149
57
|
|
|
58
|
+
# model="auto" → routes to cheapest capable provider
|
|
150
59
|
response = client.chat.completions.create(
|
|
151
|
-
model="auto",
|
|
152
|
-
messages=[{"role": "user", "content": "
|
|
60
|
+
model="auto",
|
|
61
|
+
messages=[{"role": "user", "content": "Write a Python fibonacci function"}]
|
|
153
62
|
)
|
|
63
|
+
|
|
64
|
+
print(response.choices[0].message.content)
|
|
65
|
+
# Output: GPT-4o quality, DeepSeek/Groq price
|
|
154
66
|
```
|
|
155
67
|
|
|
156
68
|
---
|
|
157
69
|
|
|
158
|
-
##
|
|
159
|
-
|
|
160
|
-
A3M analyzes every request:
|
|
70
|
+
## Why Your AI Costs Too Much
|
|
161
71
|
|
|
162
|
-
|
|
163
|
-
|--------|---------|
|
|
164
|
-
| **Domain** | Legal, medical, code, finance, ML keywords |
|
|
165
|
-
| **Task type** | Code, translation, analysis, creative |
|
|
166
|
-
| **Complexity** | Clause count, multi-step markers |
|
|
167
|
-
| **Verb intensity** | "design/architect" → complex, "what/who" → simple |
|
|
72
|
+
Most requests don't need GPT-4o. A simple question costs the same as a complex one.
|
|
168
73
|
|
|
169
|
-
|
|
74
|
+
| Query | GPT-4o | A3M Routes To | You Save |
|
|
75
|
+
|-------|--------|---------------|----------|
|
|
76
|
+
| "What is 2+2?" | $0.03 | Groq ($0.0001) | **99.7%** |
|
|
77
|
+
| "Explain quantum" | $0.03 | Mistral ($0.0002) | **99.3%** |
|
|
78
|
+
| "Write a Python function" | $0.05 | DeepSeek ($0.002) | **96%** |
|
|
79
|
+
| Complex reasoning | $0.15 | GPT-4o ($0.15) | **0%** (correctly routed) |
|
|
170
80
|
|
|
171
|
-
|
|
172
|
-
|------|-----------|----------|
|
|
173
|
-
| **Free** | Ollama, Llama.cpp | Experimentation |
|
|
174
|
-
| **Cheap** | Groq, DeepSeek, Mistral | Simple Q&A, short code |
|
|
175
|
-
| **Mid** | GPT-4o-mini, Claude-haiku | Standard tasks |
|
|
176
|
-
| **Premium** | GPT-4o, Claude-sonnet, Gemini | Complex reasoning |
|
|
81
|
+
A3M analyzes your prompt and routes to the cheapest provider that can answer it correctly.
|
|
177
82
|
|
|
178
83
|
---
|
|
179
84
|
|
|
180
|
-
##
|
|
181
|
-
|
|
182
|
-
**80+ providers** including OpenAI, Anthropic, Google, Groq, DeepSeek, Mistral, NVIDIA, Ollama, vLLM, and more.
|
|
183
|
-
|
|
184
|
-
Availability checked at runtime.
|
|
185
|
-
|
|
186
|
-
---
|
|
187
|
-
|
|
188
|
-
## Architecture
|
|
85
|
+
## How Routing Works
|
|
189
86
|
|
|
190
87
|
```
|
|
191
|
-
Request
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
88
|
+
Your Request
|
|
89
|
+
│
|
|
90
|
+
▼
|
|
91
|
+
┌────────────┐
|
|
92
|
+
│ Semantic │ ← "Is this a duplicate?" (free cache hit?)
|
|
93
|
+
└─────┬──────┘
|
|
94
|
+
▼
|
|
95
|
+
┌────────────┐
|
|
96
|
+
│ Router │ ← "Simple question or complex reasoning?"
|
|
97
|
+
└─────┬──────┘
|
|
98
|
+
▼
|
|
99
|
+
┌────────────┐
|
|
100
|
+
│ Provider │ ← Groq / Mistral / DeepSeek / GPT-4o / Claude...
|
|
101
|
+
└────────────┘
|
|
195
102
|
```
|
|
196
103
|
|
|
197
|
-
|
|
198
|
-
-
|
|
199
|
-
-
|
|
200
|
-
- **Ensemble** — Optional parallel calls for best-answer mode
|
|
104
|
+
**Two modes:**
|
|
105
|
+
- `model="auto"` — Heuristic router, ~0.4ms overhead, zero extra cost
|
|
106
|
+
- `model="jev-auto"` — ML decision head, calibrated probabilities, ~2ms warm
|
|
201
107
|
|
|
202
108
|
---
|
|
203
109
|
|
|
204
|
-
##
|
|
110
|
+
## 80+ Providers, Zero Config
|
|
205
111
|
|
|
206
112
|
```bash
|
|
207
|
-
npx a3m-router
|
|
208
|
-
npx a3m-router route "query" # See routing decision
|
|
209
|
-
npx a3m-router health # Provider status
|
|
210
|
-
npx a3m-router benchmark # Local accuracy test
|
|
113
|
+
npx a3m-router providers list
|
|
211
114
|
```
|
|
212
115
|
|
|
116
|
+
| Tier | Examples |
|
|
117
|
+
|------|----------|
|
|
118
|
+
| Free | Ollama, Llama.cpp, HuggingFace Inference |
|
|
119
|
+
| Budget | Groq, DeepSeek, Mistral, Cloudflare Workers AI |
|
|
120
|
+
| Mid | GPT-4o-mini, Claude-haiku, Gemini-flash |
|
|
121
|
+
| Premium | GPT-4o, Claude-sonnet, Gemini-pro |
|
|
122
|
+
|
|
123
|
+
Provider availability checked at runtime — no hardcoded uptimes.
|
|
124
|
+
|
|
213
125
|
---
|
|
214
126
|
|
|
215
|
-
##
|
|
127
|
+
## Ship in Minutes, Not Days
|
|
216
128
|
|
|
217
|
-
|
|
129
|
+
**Drop-in OpenAI replacement:**
|
|
130
|
+
```python
|
|
131
|
+
# Just change the base URL — your existing code works
|
|
132
|
+
client = OpenAI(base_url="http://localhost:8787/v1", api_key="not-needed")
|
|
133
|
+
```
|
|
218
134
|
|
|
135
|
+
**Or use the full API:**
|
|
219
136
|
```python
|
|
220
137
|
from a3m.router import A3MRouter
|
|
221
138
|
|
|
222
139
|
router = A3MRouter(
|
|
223
140
|
model="auto",
|
|
224
|
-
parallel_ensemble=3, #
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
result = router.route(
|
|
228
|
-
messages=[{"role": "user", "content": "Explain quantum entanglement"}],
|
|
229
|
-
ensemble_timeout_ms=10000,
|
|
141
|
+
parallel_ensemble=3, # Call 3 providers, take the best
|
|
142
|
+
memory={"type": "semantic", "window": 10}, # Remember context
|
|
230
143
|
)
|
|
231
144
|
|
|
232
|
-
|
|
233
|
-
print(f"
|
|
145
|
+
result = router.route(messages=[{"role": "user", "content": "..."}])
|
|
146
|
+
print(f"Provider: {result.provider}")
|
|
147
|
+
print(f"Cost: ${result.cost}")
|
|
234
148
|
```
|
|
235
149
|
|
|
236
150
|
---
|
|
237
151
|
|
|
238
|
-
##
|
|
152
|
+
## Self-Hosted, No Lock-In
|
|
239
153
|
|
|
240
|
-
A3M
|
|
154
|
+
OpenRouter takes a cut. A3M runs on your machine.
|
|
241
155
|
|
|
242
|
-
```
|
|
243
|
-
|
|
244
|
-
|
|
245
|
-
memory={
|
|
246
|
-
"type": "semantic",
|
|
247
|
-
"window": 10, # Last 10 exchanges
|
|
248
|
-
"similarity_threshold": 0.85,
|
|
249
|
-
}
|
|
250
|
-
)
|
|
156
|
+
```bash
|
|
157
|
+
# Docker (one command)
|
|
158
|
+
docker run -p 8787:8787 ghcr.io/das-rebel/a3m-router:latest
|
|
251
159
|
|
|
252
|
-
#
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
)
|
|
256
|
-
# A3M knows "Python web app" from previous context
|
|
160
|
+
# Or Node.js / Python directly
|
|
161
|
+
npm install adaptive-memory-multi-model-router
|
|
162
|
+
python -m a3m_router.serve
|
|
257
163
|
```
|
|
258
164
|
|
|
165
|
+
No API key to share. No vendor lock-in. Your prompts stay on your infrastructure.
|
|
166
|
+
|
|
259
167
|
---
|
|
260
168
|
|
|
261
|
-
##
|
|
169
|
+
## The Fine Print
|
|
262
170
|
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
| **Provider diversity** | Centralized | Decentralized |
|
|
269
|
-
| **Cost** | $0.0015/1K | $0.00012/1K |
|
|
171
|
+
**Works great when:**
|
|
172
|
+
- You're building AI features and need cost control
|
|
173
|
+
- You want fallback providers (if Groq is down, we route elsewhere)
|
|
174
|
+
- You need semantic caching across conversation turns
|
|
175
|
+
- You want to compare provider quality on the same prompts
|
|
270
176
|
|
|
271
|
-
|
|
177
|
+
**Not the right tool when:**
|
|
178
|
+
- You need exactly GPT-4o for every request (then just use GPT-4o)
|
|
179
|
+
- Your infrastructure can't run a local service
|
|
272
180
|
|
|
273
181
|
---
|
|
274
182
|
|
|
275
|
-
##
|
|
183
|
+
## CLI Reference
|
|
276
184
|
|
|
277
|
-
|
|
278
|
-
-
|
|
279
|
-
|
|
280
|
-
-
|
|
185
|
+
```bash
|
|
186
|
+
npx a3m-router serve # Start server (port 8787)
|
|
187
|
+
npx a3m-router route "prompt" # Preview routing decision
|
|
188
|
+
npx a3m-router health # Live provider availability
|
|
189
|
+
npx a3m-router benchmark # Local quality benchmark
|
|
190
|
+
npx a3m-router providers list # Show all providers
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
**Environment variables:**
|
|
194
|
+
```bash
|
|
195
|
+
A3M_LOG_LEVEL=debug # Debug logging
|
|
196
|
+
PORT=8787 # Server port
|
|
197
|
+
A3M_JEV_URL=... # Optional: remote Jev ML backend
|
|
198
|
+
```
|
|
281
199
|
|
|
282
200
|
---
|
|
283
201
|
|
|
284
|
-
##
|
|
202
|
+
## Contributing
|
|
285
203
|
|
|
286
|
-
[
|
|
204
|
+
See [CONTRIBUTING.md](CONTRIBUTING.md) for setup, project structure, and code conventions.
|
|
205
|
+
|
|
206
|
+
- [Issue Tracker](https://github.com/Das-rebel/a3m-router/issues)
|
|
207
|
+
- [Discussions](https://github.com/Das-rebel/a3m-router/discussions)
|
|
208
|
+
- [Changelog](CHANGELOG.md)
|
|
287
209
|
|
|
288
210
|
---
|
|
289
211
|
|
|
290
|
-
##
|
|
212
|
+
## The Philosophy (For the Curious)
|
|
291
213
|
|
|
292
|
-
|
|
293
|
-
- **PyPI downloads:** ~620/month
|
|
294
|
-
- **Providers:** 80+
|
|
295
|
-
- **Tests:** 28/28 passing
|
|
296
|
-
- **License:** MIT
|
|
214
|
+
A3M is built on biological intelligence: evolution solved the routing problem 3 billion years ago. The immune system doesn't use the same response for every pathogen — it routes resources based on threat level.
|
|
297
215
|
|
|
298
|
-
|
|
216
|
+
Same idea here: simple questions get cheap answers. Complex reasoning gets premium models. The router learns from traffic and improves over time.
|
|
299
217
|
|
|
300
|
-
|
|
301
|
-
<strong>Built on 3 billion years of biological intelligence.</strong><br>
|
|
302
|
-
<a href="https://github.com/Das-rebel/a3m-router">GitHub</a> •
|
|
303
|
-
<a href="https://twitter.com/a3m_router">Twitter</a> •
|
|
304
|
-
<a href="https://www.npmjs.com/package/adaptive-memory-multi-model-router">npm</a> •
|
|
305
|
-
<a href="https://pypi.org/project/a3m-router/">PyPI</a>
|
|
306
|
-
</p>
|
|
218
|
+
A3M is 100% open-source, self-hostable, and community-driven. We're not competing with OpenRouter — we're offering a different philosophy: open, decentralized, and yours.
|
|
307
219
|
|
|
308
220
|
---
|
|
309
221
|
|
|
310
|
-
##
|
|
222
|
+
## ⭐ Star History
|
|
311
223
|
|
|
312
|
-
|
|
224
|
+
[](https://star-history.com/#Das-rebel/a3m-router&Timeline)
|