@tokcalc/mcp-server 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +86 -0
- package/dist/index.js +21014 -0
- package/package.json +38 -0
package/README.md
ADDED
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# @tokcalc/mcp-server
|
|
2
|
+
|
|
3
|
+
**LLM serving capacity planner for AI agents.**
|
|
4
|
+
|
|
5
|
+
Open-source MCP (Model Context Protocol) server that lets AI agents (Cursor, Claude Desktop, Cline) estimate LLM serving capacity — model fit, KV cache, throughput, latency, multi-GPU topology, and cost.
|
|
6
|
+
|
|
7
|
+
## Tools
|
|
8
|
+
|
|
9
|
+
| Tool | What it does |
|
|
10
|
+
|---|---|
|
|
11
|
+
| `estimate_capacity` | VRAM/KV/throughput/latency/cost for one config |
|
|
12
|
+
| `compare_gpus` | Ranked GPU comparison for one workload |
|
|
13
|
+
| `recommend_topology` | TP/CP topology recommendation |
|
|
14
|
+
| `estimate_api_vs_self_host` | Break-even analysis |
|
|
15
|
+
| `list_models` | Discover supported model IDs (35 models) |
|
|
16
|
+
| `list_gpus` | Discover supported GPU IDs (30 GPUs) |
|
|
17
|
+
|
|
18
|
+
All tools are **read-only** — no side effects, no cloud credentials, no deployments.
|
|
19
|
+
|
|
20
|
+
## Install
|
|
21
|
+
|
|
22
|
+
### Claude Desktop
|
|
23
|
+
|
|
24
|
+
Add to `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or `%APPDATA%\Claude\claude_desktop_config.json` (Windows):
|
|
25
|
+
|
|
26
|
+
```json
|
|
27
|
+
{
|
|
28
|
+
"mcpServers": {
|
|
29
|
+
"tokcalc": {
|
|
30
|
+
"command": "npx",
|
|
31
|
+
"args": ["-y", "@tokcalc/mcp-server"]
|
|
32
|
+
}
|
|
33
|
+
}
|
|
34
|
+
}
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Restart Claude Desktop. The `tokcalc` server will be available as an MCP tool source.
|
|
38
|
+
|
|
39
|
+
### Cursor
|
|
40
|
+
|
|
41
|
+
Add to `.cursor/mcp.json` in your project:
|
|
42
|
+
|
|
43
|
+
```json
|
|
44
|
+
{
|
|
45
|
+
"mcpServers": {
|
|
46
|
+
"tokcalc": {
|
|
47
|
+
"command": "npx",
|
|
48
|
+
"args": ["-y", "@tokcalc/mcp-server"]
|
|
49
|
+
}
|
|
50
|
+
}
|
|
51
|
+
}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
### Cline (VS Code)
|
|
55
|
+
|
|
56
|
+
Add the same config to Cline's MCP settings.
|
|
57
|
+
|
|
58
|
+
## Example prompts
|
|
59
|
+
|
|
60
|
+
Ask your AI agent:
|
|
61
|
+
|
|
62
|
+
> "I need to serve Llama 3.3 70B at 32K context for 50 concurrent users. What GPU topology do you recommend, and how much will it cost per month?"
|
|
63
|
+
|
|
64
|
+
> "Compare H100 vs H200 for serving Qwen 2.5 72B in FP8 with continuous batching."
|
|
65
|
+
|
|
66
|
+
> "At what daily request volume does self-hosting Llama 70B on H200 beat the GPT-4o API?"
|
|
67
|
+
|
|
68
|
+
The agent calls `list_models` → `list_gpus` → `recommend_topology` → `estimate_capacity` and returns a structured plan with throughput ranges, latency, VRAM, cost, and confidence levels.
|
|
69
|
+
|
|
70
|
+
## Supported models (35)
|
|
71
|
+
|
|
72
|
+
Llama 3/3.1/3.3, Llama 4 Scout/Maverick, Mistral 7B, Mixtral 8x7B/8x22B, Mistral Large 3, Pixtral 12B, Codestral, Qwen 2/2.5/3 (incl. MoE + VL), DeepSeek V3/R1/Coder V2, Gemma 2, Phi-3/4, SmolLM2, Falcon 3, OLMo 2, BGE-M3, E5, GTE.
|
|
73
|
+
|
|
74
|
+
## Supported GPUs (30)
|
|
75
|
+
|
|
76
|
+
NVIDIA H100/H200/B200/B300, A100, L40S, L4, T4, V100, RTX 4090/3090/5090, RTX PRO 6000 Blackwell, AMD MI300X/MI325X, Intel Gaudi 3, Google TPU v5p/Trillium, Groq LPU, Cerebras CS-3, Apple M2/M3/M4 Ultra/Max.
|
|
77
|
+
|
|
78
|
+
## License
|
|
79
|
+
|
|
80
|
+
Apache 2.0 — same as the main tokcalc project.
|
|
81
|
+
|
|
82
|
+
## Links
|
|
83
|
+
|
|
84
|
+
- [Live calculator](https://tokcalc.vercel.app)
|
|
85
|
+
- [GitHub](https://github.com/stevecrates489-commits/tokcalc)
|
|
86
|
+
- [CONTRIBUTING](https://github.com/stevecrates489-commits/tokcalc/blob/main/CONTRIBUTING.md)
|