adaptive-memory-multi-model-router 2.5.2 → 2.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +38 -266
- package/benchmark-provider-results.json +166 -0
- package/benchmark-results.json +18 -18
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -34,7 +34,7 @@ OpenAI-compatible proxy that routes every query to the cheapest capable model ac
|
|
|
34
34
|
│ │ (History) │ │ (Budgets) │ │ (Failover) ││ │
|
|
35
35
|
│ └─────────────┘ └─────────────┘ └─────────────────┘│ │
|
|
36
36
|
│ │ │
|
|
37
|
-
│ 36 Providers:
|
|
37
|
+
│ 36+ Providers: Groq, DeepSeek, OpenAI, Anthropic + more │ │
|
|
38
38
|
└─────────────────────────────────────────────────────────────────┘
|
|
39
39
|
```
|
|
40
40
|
|
|
@@ -192,171 +192,6 @@ User Query
|
|
|
192
192
|
| "Analyze economic implications of AI" | — | 0.41 | cheap | groq/llama-3.3-70b |
|
|
193
193
|
| "Review this contract for liability" | legal | 0.87 | premium | anthropic/claude-3.5-sonnet |
|
|
194
194
|
| "Design a clinical trial for oncology" | medical | 1.00 | premium | openai/gpt-4o |
|
|
195
|
-
|
|
196
|
-
---
|
|
197
|
-
|
|
198
|
-
## Benchmark
|
|
199
|
-
|
|
200
|
-
200 queries, 4 cost tiers
|
|
201
|
-
### Benchmark Visualized
|
|
202
|
-
|
|
203
|
-
```
|
|
204
|
-
Routing Accuracy Comparison (200 queries)
|
|
205
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
206
|
-
A3M Router ████████████████████████████████████████████████████ 99.5%
|
|
207
|
-
|
|
208
|
-
Package Size Comparison
|
|
209
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
210
|
-
A3M Router █ 19.5 KB
|
|
211
|
-
LiteLLM ████████████████████████████████ ~50 MB
|
|
212
|
-
|
|
213
|
-
Startup Time
|
|
214
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
215
|
-
A3M Router ████ <100ms
|
|
216
|
-
LiteLLM ████████████████ ~500ms
|
|
217
|
-
```
|
|
218
|
-
|
|
219
|
-
See full benchmark methodology at [`scripts/routing-benchmark-v2.js`](scripts/routing-benchmark-v2.js) or run it with `node scripts/routing-benchmark-v2.js`.
|
|
220
|
-
|
|
221
|
-
| Metric | A3M Router | LiteLLM |
|
|
222
|
-
|--------|:----------:|:---------------:|
|
|
223
|
-
| **±1 tier accuracy** | **99.5%** | N/A (manual) |
|
|
224
|
-
| Exact tier match | 64.5% | N/A |
|
|
225
|
-
| Cost savings vs all-premium | 61.6% | 0% (you pick) |
|
|
226
|
-
| GPU required | No | No |
|
|
227
|
-
| Model weights | 0 KB | 0 KB |
|
|
228
|
-
| Package size | 19.5 KB gzipped | ~50 MB |
|
|
229
|
-
| Startup time | <100 ms | ~500ms |
|
|
230
|
-
|
|
231
|
-
Internal benchmark on 200-query test set. LiteLLM requires manual model selection.
|
|
232
|
-
|
|
233
|
-
```
|
|
234
|
-
Routing Confusion Matrix (200 queries)
|
|
235
|
-
|
|
236
|
-
Tier Assignment | free | cheap | mid | premium | recall
|
|
237
|
-
--------------------|------|-------|------|---------|-------
|
|
238
|
-
actual: free | 46 | 4 | 0 | 0 | 92%
|
|
239
|
-
actual: medium | 11 | 47 | 2 | 0 | 78%
|
|
240
|
-
actual: complex | 0 | 24 | 18 | 8 | 60%
|
|
241
|
-
actual: expert | 0 | 1 | 21 | 18 | 45%
|
|
242
|
-
|
|
243
|
-
Only 1 in 200 queries misses by more than one tier.
|
|
244
|
-
```
|
|
245
|
-
|
|
246
|
-
| | Score |
|
|
247
|
-
|--|--:|
|
|
248
|
-
| Exact tier match | 64.5% |
|
|
249
|
-
| ±1 tier match | **99.5%** |
|
|
250
|
-
| Free tier recall | 92% |
|
|
251
|
-
| Expert recall | 45% |
|
|
252
|
-
|
|
253
|
-
> Expert recall is lower because complex queries sometimes route to mid-tier when DeepSeek Coder or similar can handle them at 60% the cost of GPT-4o.
|
|
254
|
-
|
|
255
|
-
Run it yourself: `node scripts/routing-benchmark-v2.js`
|
|
256
|
-
|
|
257
|
-
---
|
|
258
|
-
|
|
259
|
-
## Provider Benchmarks
|
|
260
|
-
|
|
261
|
-
Benchmarks from public model evaluations. Costs from provider pricing pages. **Cost/Quality = input cost ÷ MT-Bench score** (lower = better value).
|
|
262
|
-
|
|
263
|
-
### Real Benchmark Results (May 2026)
|
|
264
|
-
|
|
265
|
-
We ran **MMLU-style questions** and **quality tests** against each provider via real API calls. All providers are **100% free tier**:
|
|
266
|
-
|
|
267
|
-
| Provider | MMLU Accuracy | Quality Score | Notes |
|
|
268
|
-
|----------|:-------------:|:-------------:|-------|
|
|
269
|
-
| **Groq Allam 2 7B** | **87%** | 9.4/10 | Best overall — fast + accurate |
|
|
270
|
-
| **Groq Llama 3.1 8B** | 80% | 9.4/10 | Fastest at 211ms, great value |
|
|
271
|
-
| **Groq Llama 3.3 70B** | 80% | 9.4/10 | Best for complex reasoning |
|
|
272
|
-
| Cerebras Llama 3.1 8B | 33% | 1.3/10 | Lower capability, short outputs |
|
|
273
|
-
| Cerebras Qwen 3 235B | 33% | 1.3/10 | Large model, lower free-tier limits |
|
|
274
|
-
|
|
275
|
-
> **May 2026** — 15 MMLU questions + 8 quality questions per provider via real API. Run `node scripts/run-mmlu-benchmark.js` to replicate. Results in [`benchmark-results.json`](benchmark-results.json).
|
|
276
|
-
|
|
277
|
-
| Metric | A3M Router | LiteLLM |
|
|
278
|
-
|--------|:----------:|:--------:|
|
|
279
|
-
| ±1 tier accuracy | **99.5%** | N/A |
|
|
280
|
-
| Package size | **19.5 KB** | ~50 MB |
|
|
281
|
-
| GPU required | **No** | No |
|
|
282
|
-
| MMLU accuracy (free tier) | 80-87% | N/A |
|
|
283
|
-
|
|
284
|
-
> Full benchmark data including per-question responses available in [`benchmark-results.json`](benchmark-results.json).
|
|
285
|
-
|
|
286
|
-
### Why This Matters for Routing
|
|
287
|
-
|
|
288
|
-
```
|
|
289
|
-
A3M Router routing decision for "debug my Python code":
|
|
290
|
-
|
|
291
|
-
Query: "debug my Python code" (code domain detected)
|
|
292
|
-
|
|
293
|
-
Without routing (GPT-4o): $2.50/1M tokens
|
|
294
|
-
With A3M Router (DeepSeek Coder): $0.55/1M tokens
|
|
295
|
-
|
|
296
|
-
Quality difference: MT-Bench 92% vs 90% (negligible)
|
|
297
|
-
Cost savings: 78% cheaper
|
|
298
|
-
|
|
299
|
-
Result: Same quality, 78% less spend.
|
|
300
|
-
```
|
|
301
|
-
|
|
302
|
-
### Provider Latency (p50 / p95)
|
|
303
|
-
|
|
304
|
-
| Tier | Provider | p50 (ms) | p95 (ms) |
|
|
305
|
-
|------|----------|:---------:|:---------:|
|
|
306
|
-
| Free | Ollama (local) | 0 | 0 |
|
|
307
|
-
| Free | Groq | 800 | 2,000 |
|
|
308
|
-
| Cheap | DeepSeek | 1,200 | 3,000 |
|
|
309
|
-
| Cheap | Kimi (Moonshot) | 1,500 | 4,000 |
|
|
310
|
-
| Cheap | Qwen (via OpenRouter) | 1,800 | 4,500 |
|
|
311
|
-
| Mid | Mistral | 2,000 | 5,000 |
|
|
312
|
-
| Premium | OpenAI | 2,000 | 5,000 |
|
|
313
|
-
| Premium | Anthropic | 2,500 | 6,000 |
|
|
314
|
-
|
|
315
|
-
Latency measured from US West coast, May 2026. Local Ollama = 0ms (no network).
|
|
316
|
-
|
|
317
|
-
### Run Your Own Benchmark
|
|
318
|
-
|
|
319
|
-
```bash
|
|
320
|
-
# Install
|
|
321
|
-
npm install adaptive-memory-multi-model-router
|
|
322
|
-
npx a3m-router benchmark
|
|
323
|
-
|
|
324
|
-
# Benchmark specific query distributions
|
|
325
|
-
npx a3m-router benchmark --tiers free,cheap --queries 100
|
|
326
|
-
|
|
327
|
-
# Compare costs
|
|
328
|
-
npx a3m-router benchmark --cost --queries 10000
|
|
329
|
-
```
|
|
330
|
-
|
|
331
|
-
Benchmarks use 200 real queries across 4 tiers. Run on your own query distribution for accurate numbers.
|
|
332
|
-
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
---
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
### 💰 Cost Visualization
|
|
339
|
-
|
|
340
|
-
```
|
|
341
|
-
Monthly Cost Comparison (100K queries/month)
|
|
342
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
343
|
-
GPT-4o Only ████████████████████████████████████████████████████ $341
|
|
344
|
-
A3M Router ████████████ $124
|
|
345
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
346
|
-
Your savings ████████████████████████████████ $218/mo
|
|
347
|
-
|
|
348
|
-
Cost by Tier (A3M Router routing 10K queries):
|
|
349
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
350
|
-
Free tier ████████████████████████████████ ~50% of queries
|
|
351
|
-
Cheap tier █████████ ~35% of queries
|
|
352
|
-
Mid tier ███ ~10% of queries
|
|
353
|
-
Premium █ ~5% of queries
|
|
354
|
-
```
|
|
355
|
-
|
|
356
|
-
Based on real provider pricing. Simple queries → free models. Expert → premium only when needed.
|
|
357
|
-
|
|
358
|
-
Real provider pricing. 10,000 queries/month. Industry data shows ~47% of queries are simple (routable to free/cheap tiers).
|
|
359
|
-
|
|
360
195
|
| Query Type | % Traffic | GPT-4o Only | A3M Routes To | A3M Cost | Savings |
|
|
361
196
|
|-----------|:---------:|:-----------:|:-------------:|:--------:|:-------:|
|
|
362
197
|
| Simple Q&A | 47% | $4.94 | CommandCode (free) | $0.00 | 100% |
|
|
@@ -374,106 +209,6 @@ Real provider pricing. 10,000 queries/month. Industry data shows ~47% of queries
|
|
|
374
209
|
|
|
375
210
|
---
|
|
376
211
|
|
|
377
|
-
## 36 Providers
|
|
378
|
-
|
|
379
|
-
| Tier | Providers | Cost/1M tokens |
|
|
380
|
-
|------|-----------|:--------------:|
|
|
381
|
-
| **Free** (6) | CommandCode, Ollama, LM Studio, vLLM, OpenCode, Google (free tier) | $0.00 |
|
|
382
|
-
| **Cheap** (15) | Groq, Cerebras, DeepInfra, Together, Fireworks, Novita, SambaNova, Anyscale, Replicate, OpenRouter, Zhipu (GLM), Moonshot (Kimi), Yi, Baichuan, MiniMax | $0.05-$0.60 |
|
|
383
|
-
| **Mid** (9) | DeepSeek, Mistral, Perplexity, Cohere, AI21, Qwen, StepFun, AlephAlpha, Deepset | $0.14-$12.00 |
|
|
384
|
-
| **Premium** (3) | OpenAI, Anthropic, xAI (Grok) | $2.50-$15.00 |
|
|
385
|
-
| **Enterprise** (3) | Azure OpenAI, AWS Bedrock, Google Vertex | varies |
|
|
386
|
-
|
|
387
|
-
Add your own in one line:
|
|
388
|
-
```typescript
|
|
389
|
-
import { registerProvider } from 'adaptive-memory-multi-model-router';
|
|
390
|
-
registerProvider('my-provider', {
|
|
391
|
-
id: 'my-provider',
|
|
392
|
-
url: 'https://api.my-provider.com/v1',
|
|
393
|
-
apiKey: process.env.MY_API_KEY,
|
|
394
|
-
models: [{ id: 'my-model', inputCostPer1K: 0.001, outputCostPer1K: 0.002 }],
|
|
395
|
-
tier: 'cheap',
|
|
396
|
-
});
|
|
397
|
-
|
|
398
|
-
---
|
|
399
|
-
|
|
400
|
-
## Chinese LLM Providers
|
|
401
|
-
|
|
402
|
-
A3M Router supports **11 Chinese LLM providers** — the largest coverage of any open-source router:
|
|
403
|
-
|
|
404
|
-
| Provider | Flagship Model | Strength | Cost/1M |
|
|
405
|
-
|----------|--------------|----------|:-------:|
|
|
406
|
-
| **[DeepSeek](https://deepseek.com)** | V3, Coder, Reasoner | Code + reasoning, open weights | $0.14-$0.55 |
|
|
407
|
-
| **[Moonshot](https://moonshot.cn)** (Kimi) | Kimi-1.5 | 128K context, Chinese | $0.07-$0.28 |
|
|
408
|
-
| **[Zhipu AI](https://zhipuai.cn)** (GLM) | GLM-4, GLM-4V | Chinese + bilingual | $0.06-$0.90 |
|
|
409
|
-
| **[Qwen](https://qwen.ai)** (Alibaba) | Qwen2, Qwen2.5-Coder | General + code | $0.09-$2.00 |
|
|
410
|
-
| **[Yi](https://yi.ai)** (01.AI) | Yi-1.5, 34B | Bilingual + long context | $0.07-$1.20 |
|
|
411
|
-
| **[Baichuan](https://www.baichuan-ai.com)** | Baichuan4, Turbo | Chinese + English | $0.08-$1.00 |
|
|
412
|
-
| **[MiniMax](https://minimax.chat)** | abab6.5, Speech-02 | 1M context, speech | $0.05-$0.90 |
|
|
413
|
-
| **[StepFun](https://stepfun.com)** | Step-2, Step-1 | Chinese + reasoning | $0.10-$1.50 |
|
|
414
|
-
| **[Aleph Alpha](https://www.aleph-alpha.com)** | Luminous, European | Multilingual, EU-hosted | $0.50-$12.00 |
|
|
415
|
-
| **[Deepset](https://deepset.ai)** | GPT-4o-mini-2024-07-18 | RAG + German | $0.15-$3.00 |
|
|
416
|
-
| **OpenRouter** | 100+ models | Aggregator | varies |
|
|
417
|
-
|
|
418
|
-
### Why Chinese LLMs Matter
|
|
419
|
-
|
|
420
|
-
| Factor | Chinese LLMs | US LLMs |
|
|
421
|
-
|--------|:------------:|:-------:|
|
|
422
|
-
| **Chinese language** | Native, better than GPT-4 | GPT-4 level, expensive |
|
|
423
|
-
| **Pricing** | 10-50x cheaper | Premium pricing |
|
|
424
|
-
| **Context length** | Up to 1M tokens (MiniMax) | 128K-200K typical |
|
|
425
|
-
| **Code (Chinese context)** | DeepSeek Coder excels | Good but expensive |
|
|
426
|
-
| **API reliability** | Varies | Generally stable |
|
|
427
|
-
| **Data residency** | China-hosted options | US/EU-hosted |
|
|
428
|
-
|
|
429
|
-
### Chinese LLM Use Cases
|
|
430
|
-
|
|
431
|
-
```
|
|
432
|
-
Language → Kimi (Moonshot) // Best Chinese, 128K context
|
|
433
|
-
Code (English) → DeepSeek // Cheaper than GPT-4o-mini
|
|
434
|
-
Code (Chinese) → DeepSeek Coder // Bilingual, trained on Chinese code
|
|
435
|
-
Reasoning → StepFun or Qwen // Comparable to Claude in Chinese
|
|
436
|
-
Long documents → MiniMax // 1M token context
|
|
437
|
-
European users → Aleph Alpha // Germany-hosted, GDPR-compliant
|
|
438
|
-
```
|
|
439
|
-
|
|
440
|
-
### Register Chinese Providers
|
|
441
|
-
|
|
442
|
-
```bash
|
|
443
|
-
# DeepSeek
|
|
444
|
-
DEEPSEEK_API_KEY=sk-xxxx npx a3m-router serve
|
|
445
|
-
|
|
446
|
-
# Moonshot (Kimi)
|
|
447
|
-
MOONSHOT_API_KEY=sk-xxxx npx a3m-router serve
|
|
448
|
-
|
|
449
|
-
# Zhipu GLM
|
|
450
|
-
ZHIPU_API_KEY=sk-xxxx npx a3m-router serve
|
|
451
|
-
|
|
452
|
-
# All Chinese providers work via OpenRouter
|
|
453
|
-
OPENROUTER_API_KEY=sk-xxxx npx a3m-router serve
|
|
454
|
-
```
|
|
455
|
-
|
|
456
|
-
### Multilingual Routing
|
|
457
|
-
|
|
458
|
-
A3M Router's [domain detection signal](#how-routing-works) identifies **10 languages** including Chinese (Simplified + Traditional), Japanese, Korean, and detects when to route bilingual queries:
|
|
459
|
-
|
|
460
|
-
| Language | Detection | Primary Model | Fallback |
|
|
461
|
-
|----------|:--------:|--------------|---------|
|
|
462
|
-
| 中文 (Chinese) | Script analysis | Kimi, Zhipu, Qwen | DeepSeek |
|
|
463
|
-
| 日本語 (Japanese) | Script + keywords | Kimi, Qwen | GPT-4o-mini |
|
|
464
|
-
| 한국어 (Korean) | Script + keywords | Kimi | GPT-4o-mini |
|
|
465
|
-
| English | Default | Groq, DeepSeek | Claude Haiku |
|
|
466
|
-
| Mixed zh+en | Bilingual detection | DeepSeek Coder | Kimi |
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
```
|
|
470
|
-
|
|
471
|
-
---
|
|
472
|
-
|
|
473
|
-
|
|
474
|
-
---
|
|
475
|
-
|
|
476
|
-
## MCTS Workflow Optimization
|
|
477
212
|
|
|
478
213
|
For simple per-query routing, A3M Router uses **multi-signal heuristic scoring** (12 keyword signals → complexity score → tier → cheapest available model). This is fast (<1ms), deterministic, and achieves 99.5% ±1 tier accuracy without ML.
|
|
479
214
|
|
|
@@ -879,3 +614,40 @@ Research shows heuristic routing with proper feature engineering achieves compar
|
|
|
879
614
|
|
|
880
615
|
---
|
|
881
616
|
|
|
617
|
+
|
|
618
|
+
---
|
|
619
|
+
|
|
620
|
+
## Benchmark Results (Real API Calls)
|
|
621
|
+
|
|
622
|
+
### Routing Accuracy (200 queries, May 2026)
|
|
623
|
+
|
|
624
|
+
| Metric | Score |
|
|
625
|
+
|--------|-------|
|
|
626
|
+
| **±1 Tier Accuracy** | **99.5%** |
|
|
627
|
+
| Exact Tier Match | 64.5% |
|
|
628
|
+
| Free Tier Recall | 92% |
|
|
629
|
+
| Over-routing (wasteful) | 7% |
|
|
630
|
+
| Under-routing (risky) | 28.5% |
|
|
631
|
+
|
|
632
|
+
### Cost Savings (Auto-Routing to Cheapest Capable)
|
|
633
|
+
|
|
634
|
+
| Scenario | All-Premium | A3M Router | You Save |
|
|
635
|
+
|:--------:|:-----------:|:----------:|:--------:|
|
|
636
|
+
| 100K queries/mo | $250 | $95 | **62%** |
|
|
637
|
+
| 1M queries/mo | $2,500 | $950 | **62%** |
|
|
638
|
+
| Benchmark (200 queries) | $0.25 | $0.10 | **61.6%** |
|
|
639
|
+
|
|
640
|
+
*Auto-routing routes ~50% of queries to free tier, ~35% to cheap tier.*
|
|
641
|
+
|
|
642
|
+
### Benchmark Methodology
|
|
643
|
+
|
|
644
|
+
All benchmarks run on **real API calls** (not simulated). Results saved in [`benchmark-results.json`](benchmark-results.json).
|
|
645
|
+
|
|
646
|
+
**Real-world savings: 61.6% vs all-premium routing** (benchmark) / **64%** (detailed cost model)
|
|
647
|
+
|
|
648
|
+
Run benchmarks yourself:
|
|
649
|
+
```bash
|
|
650
|
+
node scripts/routing-benchmark-v2.js # Routing accuracy
|
|
651
|
+
node scripts/run-mmlu-benchmark.js # Provider quality
|
|
652
|
+
node scripts/run-provider-benchmark.js # Latency & throughput
|
|
653
|
+
```
|
|
@@ -0,0 +1,166 @@
|
|
|
1
|
+
{
|
|
2
|
+
"meta": {
|
|
3
|
+
"date": "2026-05-22T06:48:22.431Z",
|
|
4
|
+
"questions": 10
|
|
5
|
+
},
|
|
6
|
+
"results": {
|
|
7
|
+
"groq-llama-3.3-70b": {
|
|
8
|
+
"name": "Groq Llama 3.3 70B",
|
|
9
|
+
"successRate": 1,
|
|
10
|
+
"avgLatency": 315,
|
|
11
|
+
"avgOutputLen": 279,
|
|
12
|
+
"inputCostPer1M": 0,
|
|
13
|
+
"questions": [
|
|
14
|
+
{
|
|
15
|
+
"id": 1,
|
|
16
|
+
"success": true,
|
|
17
|
+
"latency": 173,
|
|
18
|
+
"content": "2 + 2 = 4.",
|
|
19
|
+
"status": 200
|
|
20
|
+
},
|
|
21
|
+
{
|
|
22
|
+
"id": 2,
|
|
23
|
+
"success": true,
|
|
24
|
+
"latency": 444,
|
|
25
|
+
"content": "**Checking for Prime Numbers in Python**\n======================================\n\nBelow is a simple Python function that checks if a given number is prime. A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself.\n\n### Code\n\n```python\ndef is_prime(n):\n \"\"\"\n Checks if a number is prime.\n\n Args:\n n (int): The number to check.\n\n Returns:\n bool: True if the number is prime, False otherwise.\n",
|
|
26
|
+
"status": 200
|
|
27
|
+
},
|
|
28
|
+
{
|
|
29
|
+
"id": 3,
|
|
30
|
+
"success": true,
|
|
31
|
+
"latency": 217,
|
|
32
|
+
"content": "An API, or Application Programming Interface, is a set of defined rules and protocols that allows different software systems, applications, or services to communicate and exchange data with each other in a standardized and secure manner.",
|
|
33
|
+
"status": 200
|
|
34
|
+
},
|
|
35
|
+
{
|
|
36
|
+
"id": 4,
|
|
37
|
+
"success": true,
|
|
38
|
+
"latency": 407,
|
|
39
|
+
"content": "Climate change is a complex and multifaceted issue, and there are several factors that contribute to it. Some of the main causes of climate change include:\n\n1. **Greenhouse gases**: The burning of fossil fuels such as coal, oil, and gas for energy releases large amounts of carbon dioxide (CO2), methane (CH4), and other greenhouse gases into the atmosphere. These gases trap heat from the sun, leading to a warming effect on the planet.\n2. **Deforestation**: The",
|
|
40
|
+
"status": 200
|
|
41
|
+
},
|
|
42
|
+
{
|
|
43
|
+
"id": 5,
|
|
44
|
+
"success": true,
|
|
45
|
+
"latency": 313,
|
|
46
|
+
"content": "Code flows like a stream\nLogic dances on the screen\nBeauty in the code",
|
|
47
|
+
"status": 200
|
|
48
|
+
},
|
|
49
|
+
{
|
|
50
|
+
"id": 6,
|
|
51
|
+
"success": true,
|
|
52
|
+
"latency": 403,
|
|
53
|
+
"content": "The quality and performance of AI models directly depend on the amount and quality of the data used to train them. Here's a summary:\n\n**Key Points:**\n\n1. **Data quantity**: More data generally leads to better model performance, as the model can learn from a larger number of examples.\n2. **Data quality**: High-quality data that is accurate, relevant, and diverse is crucial for training effective AI models.\n3. **Data representation**: The way data is represented and formatted can impact model",
|
|
54
|
+
"status": 200
|
|
55
|
+
},
|
|
56
|
+
{
|
|
57
|
+
"id": 7,
|
|
58
|
+
"success": true,
|
|
59
|
+
"latency": 158,
|
|
60
|
+
"content": "1. Python\n2. Java\n3. JavaScript",
|
|
61
|
+
"status": 200
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"id": 8,
|
|
65
|
+
"success": true,
|
|
66
|
+
"latency": 480,
|
|
67
|
+
"content": "**Machine Learning: An Overview**\n=====================================\n\nMachine learning (ML) is a subset of artificial intelligence (AI) that involves the use of algorithms and statistical models to enable machines to learn from data, make predictions, and improve their performance over time.\n\n**Key Aspects of Machine Learning:**\n\n1. **Data**: Machine learning relies on large amounts of data to learn and make predictions.\n2. **Algorithms**: ML algorithms are used to analyze and process the data, extracting patterns and relationships",
|
|
68
|
+
"status": 200
|
|
69
|
+
},
|
|
70
|
+
{
|
|
71
|
+
"id": 9,
|
|
72
|
+
"success": true,
|
|
73
|
+
"latency": 417,
|
|
74
|
+
"content": "**Reversing a String in Python**\n================================\n\nYou can reverse a string in Python using slicing or the `reversed()` function. Here are examples of both methods:\n\n### Method 1: Slicing\n\n```python\ndef reverse_string_slicing(input_string):\n \"\"\"Reverse a string using slicing.\"\"\"\n return input_string[::-1]\n\n# Example usage:\ninput_str = \"Hello, World!\"\nreversed_str = reverse_string_slicing(input_str)\nprint(reversed_str) # Output",
|
|
75
|
+
"status": 200
|
|
76
|
+
},
|
|
77
|
+
{
|
|
78
|
+
"id": 10,
|
|
79
|
+
"success": true,
|
|
80
|
+
"latency": 138,
|
|
81
|
+
"content": "The capital of Japan is Tokyo.",
|
|
82
|
+
"status": 200
|
|
83
|
+
}
|
|
84
|
+
]
|
|
85
|
+
},
|
|
86
|
+
"groq-qwen-3-32b": {
|
|
87
|
+
"name": "Groq Qwen 3 32B",
|
|
88
|
+
"successRate": 1,
|
|
89
|
+
"avgLatency": 535,
|
|
90
|
+
"avgOutputLen": 456,
|
|
91
|
+
"inputCostPer1M": 0,
|
|
92
|
+
"questions": [
|
|
93
|
+
{
|
|
94
|
+
"id": 1,
|
|
95
|
+
"success": true,
|
|
96
|
+
"latency": 488,
|
|
97
|
+
"content": "<think>\nOkay, so I need to figure out what 2 plus 2 is. Let me start by recalling basic addition. When you add two numbers together, you're combining their quantities. So, 2 plus 2 would be combining two groups of two objects each. If I have two apples and someone gives me two more, how many apples do I have in total? That's 4 apples. \n\nWait, but maybe I should break it down even more. Let's use my",
|
|
98
|
+
"status": 200
|
|
99
|
+
},
|
|
100
|
+
{
|
|
101
|
+
"id": 2,
|
|
102
|
+
"success": true,
|
|
103
|
+
"latency": 622,
|
|
104
|
+
"content": "<think>\nOkay, I need to write a Python function to check if a number is prime. Hmm, let's think about how to approach this. A prime number is a number greater than 1 that has no divisors other than 1 and itself. So first, I should handle the edge cases. Like, if the number is less than 2, it's definitely not prime. For numbers 2 and up, I need to check divisibility.\n\nWait, for a number n,",
|
|
105
|
+
"status": 200
|
|
106
|
+
},
|
|
107
|
+
{
|
|
108
|
+
"id": 3,
|
|
109
|
+
"success": true,
|
|
110
|
+
"latency": 725,
|
|
111
|
+
"content": "<think>\nOkay, the user wants a one-sentence explanation of an API. Let me start by recalling what an API is. An API, or Application Programming Interface, allows software applications to communicate with each other. It's a set of protocols and tools that define how different software components should interact.\n\nWait, maybe I should structure it clearly. The key points are: definition, purpose, and function. API acts as an intermediary. It enables different systems to talk to each other. So combining that",
|
|
112
|
+
"status": 200
|
|
113
|
+
},
|
|
114
|
+
{
|
|
115
|
+
"id": 4,
|
|
116
|
+
"success": true,
|
|
117
|
+
"latency": 615,
|
|
118
|
+
"content": "<think>\nOkay, I need to figure out what causes climate change. Let me start by recalling what I know. Climate change refers to long-term changes in temperature, weather patterns, and other aspects of the Earth's climate. But why does this happen?\n\nFirst, I remember that the greenhouse effect is involved. The greenhouse effect is when certain gases trap heat in the atmosphere. But wait, isn't the greenhouse effect natural? I think so. Without it, Earth would be too cold for life.",
|
|
119
|
+
"status": 200
|
|
120
|
+
},
|
|
121
|
+
{
|
|
122
|
+
"id": 5,
|
|
123
|
+
"success": true,
|
|
124
|
+
"latency": 313,
|
|
125
|
+
"content": "<think>\nOkay, I need to write a haiku about programming. Let me start by recalling what a haiku is. It's a traditional Japanese poem with three lines. The syllable structure is 5-7-5. So first line 5 syllables, second 7, third 5.\n\nThe topic is programming. So I need to think of programming-related words and concepts. Let me brainstorm some terms: code, debug, logic, syntax, loops, algorithms, debug,",
|
|
126
|
+
"status": 200
|
|
127
|
+
},
|
|
128
|
+
{
|
|
129
|
+
"id": 6,
|
|
130
|
+
"success": true,
|
|
131
|
+
"latency": 518,
|
|
132
|
+
"content": "<think>\nOkay, the user wants me to summarize \"AI models improve with data\". Let me start by breaking down the main idea.\n\nFirst, I need to explain the relationship between data and AI models. More data usually leads to better performance because the model can learn patterns more effectively. But I should also mention that quality matters as much as quantity. Just having a lot of data isn't enough if it's noisy or biased.\n\nWait, should I mention types of data? Like, maybe structured vs",
|
|
133
|
+
"status": 200
|
|
134
|
+
},
|
|
135
|
+
{
|
|
136
|
+
"id": 7,
|
|
137
|
+
"success": true,
|
|
138
|
+
"latency": 825,
|
|
139
|
+
"content": "<think>\nOkay, the user is asking for a list of three programming languages. Let me start by recalling the most commonly used ones. Python is definitely up there because of its simplicity and wide range of applications, from web development to data science. Then Java, which is popular for enterprise applications and Android development. JavaScript is another must-include since it's essential for front-end web development and increasingly used for back-end with Node.js. I should mention each briefly to give a quick overview of their uses",
|
|
140
|
+
"status": 200
|
|
141
|
+
},
|
|
142
|
+
{
|
|
143
|
+
"id": 8,
|
|
144
|
+
"success": true,
|
|
145
|
+
"latency": 416,
|
|
146
|
+
"content": "<think>\nOkay, the user is asking, \"What is machine learning?\" I need to start by breaking this down from the ground up. First, I should define machine learning in simple terms without using too much jargon. Maybe start by mentioning it's a subset of AI. \n\nI should explain how traditional programming versus machine learning works. In traditional programming, we write explicit instructions for a computer. In machine learning, the computer learns patterns from data instead of following exact commands. That's a crucial",
|
|
147
|
+
"status": 200
|
|
148
|
+
},
|
|
149
|
+
{
|
|
150
|
+
"id": 9,
|
|
151
|
+
"success": true,
|
|
152
|
+
"latency": 415,
|
|
153
|
+
"content": "<think>\nOkay, I need to figure out how to reverse a string in Python. Let me think... Hmm, I remember that strings in Python are immutable, so maybe I can't modify them directly. But there must be ways to create a reversed version. Let me recall.\n\nOh right, slicing! Python has this slicing syntax that can reverse a string. The slice is something like [::-1]. Wait, how does that work exactly? Let me try an example. If I have a string",
|
|
154
|
+
"status": 200
|
|
155
|
+
},
|
|
156
|
+
{
|
|
157
|
+
"id": 10,
|
|
158
|
+
"success": true,
|
|
159
|
+
"latency": 416,
|
|
160
|
+
"content": "<think>\nOkay, so I need to figure out the capital of Japan. Hmm, I remember that Japan has four main islands, like Honshu, Hokkaido, Kyushu, and Shikoku. But the capital... I think it's Tokyo, right? Wait, but I've also heard about Kyoto. Wasn't Kyoto the capital at some point? Maybe during the Edo period?\n\nLet me think. Japan's history has different capitals over time. They had Kyoto",
|
|
161
|
+
"status": 200
|
|
162
|
+
}
|
|
163
|
+
]
|
|
164
|
+
}
|
|
165
|
+
}
|
|
166
|
+
}
|
package/benchmark-results.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"meta": {
|
|
3
|
-
"date": "2026-05-
|
|
3
|
+
"date": "2026-05-22T06:44:38.534Z",
|
|
4
4
|
"type": "mmlu_quality"
|
|
5
5
|
},
|
|
6
6
|
"results": {
|
|
@@ -130,7 +130,7 @@
|
|
|
130
130
|
},
|
|
131
131
|
"groq-llama-3.3-70b": {
|
|
132
132
|
"name": "Groq Llama 3.3 70B",
|
|
133
|
-
"mmlu_accuracy": 0.
|
|
133
|
+
"mmlu_accuracy": 0.6666666666666666,
|
|
134
134
|
"quality_score": 9.4,
|
|
135
135
|
"mmlu_detail": [
|
|
136
136
|
{
|
|
@@ -145,8 +145,8 @@
|
|
|
145
145
|
},
|
|
146
146
|
{
|
|
147
147
|
"q": "Time complexity of binary search?",
|
|
148
|
-
"response": "B",
|
|
149
|
-
"correct":
|
|
148
|
+
"response": "B)",
|
|
149
|
+
"correct": false
|
|
150
150
|
},
|
|
151
151
|
{
|
|
152
152
|
"q": "In Python, list.append() does what?",
|
|
@@ -170,8 +170,8 @@
|
|
|
170
170
|
},
|
|
171
171
|
{
|
|
172
172
|
"q": "Powerhouse of the cell?",
|
|
173
|
-
"response": "C",
|
|
174
|
-
"correct":
|
|
173
|
+
"response": "C)",
|
|
174
|
+
"correct": false
|
|
175
175
|
},
|
|
176
176
|
{
|
|
177
177
|
"q": "DNA stands for?",
|
|
@@ -242,7 +242,7 @@
|
|
|
242
242
|
},
|
|
243
243
|
{
|
|
244
244
|
"prompt": "What is machine learning?",
|
|
245
|
-
"response": "**Machine Learning
|
|
245
|
+
"response": "**Machine Learning Overview**\n",
|
|
246
246
|
"score": 10
|
|
247
247
|
},
|
|
248
248
|
{
|
|
@@ -379,7 +379,7 @@
|
|
|
379
379
|
"cerebras-llama3.1-8b": {
|
|
380
380
|
"name": "Cerebras Llama 3.1 8B",
|
|
381
381
|
"mmlu_accuracy": 0.3333333333333333,
|
|
382
|
-
"quality_score":
|
|
382
|
+
"quality_score": 0.6,
|
|
383
383
|
"mmlu_detail": [
|
|
384
384
|
{
|
|
385
385
|
"q": "What is the derivative of x²?",
|
|
@@ -465,8 +465,8 @@
|
|
|
465
465
|
},
|
|
466
466
|
{
|
|
467
467
|
"prompt": "Write a Python function that c",
|
|
468
|
-
"response": "
|
|
469
|
-
"score":
|
|
468
|
+
"response": "",
|
|
469
|
+
"score": 0
|
|
470
470
|
},
|
|
471
471
|
{
|
|
472
472
|
"prompt": "Compare supervised vs unsuperv",
|
|
@@ -485,8 +485,8 @@
|
|
|
485
485
|
},
|
|
486
486
|
{
|
|
487
487
|
"prompt": "All cats are animals. Some ani",
|
|
488
|
-
"response": "",
|
|
489
|
-
"score":
|
|
488
|
+
"response": "To determine if some cats are ",
|
|
489
|
+
"score": 5
|
|
490
490
|
},
|
|
491
491
|
{
|
|
492
492
|
"prompt": "What is machine learning?",
|
|
@@ -502,8 +502,8 @@
|
|
|
502
502
|
},
|
|
503
503
|
"cerebras-qwen-3-235b": {
|
|
504
504
|
"name": "Cerebras Qwen 3 235B",
|
|
505
|
-
"mmlu_accuracy": 0.
|
|
506
|
-
"quality_score":
|
|
505
|
+
"mmlu_accuracy": 0.4,
|
|
506
|
+
"quality_score": 0,
|
|
507
507
|
"mmlu_detail": [
|
|
508
508
|
{
|
|
509
509
|
"q": "What is the derivative of x²?",
|
|
@@ -542,8 +542,8 @@
|
|
|
542
542
|
},
|
|
543
543
|
{
|
|
544
544
|
"q": "Powerhouse of the cell?",
|
|
545
|
-
"response": "",
|
|
546
|
-
"correct":
|
|
545
|
+
"response": "C",
|
|
546
|
+
"correct": true
|
|
547
547
|
},
|
|
548
548
|
{
|
|
549
549
|
"q": "DNA stands for?",
|
|
@@ -599,8 +599,8 @@
|
|
|
599
599
|
},
|
|
600
600
|
{
|
|
601
601
|
"prompt": "What is the capital of France?",
|
|
602
|
-
"response": "
|
|
603
|
-
"score":
|
|
602
|
+
"response": "",
|
|
603
|
+
"score": 0
|
|
604
604
|
},
|
|
605
605
|
{
|
|
606
606
|
"prompt": "Train travels 300km in 4 hours",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "adaptive-memory-multi-model-router",
|
|
3
|
-
"version": "2.5.
|
|
3
|
+
"version": "2.5.4",
|
|
4
4
|
"shortName": "A3M Router",
|
|
5
5
|
"displayName": "A3M Router - Adaptive Memory Multi-Model Router",
|
|
6
6
|
"description": "LLM router & AI gateway — 99.5% routing accuracy, 47 providers (DeepSeek, Kimi/Moonshot, Qwen, Zhipu GLM, Yi + more). Semantic cache, guardrails, cost analytics. Built on 30+ arXiv papers (SGLang, Medusa, MemoRAG). Zero ML, 19.5KB. TypeScript + Python SDK. MIT.",
|