@tokcalc/mcp-server 0.1.3 → 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +86 -406
- package/dist/index.js +21285 -0
- package/package.json +32 -90
- package/.zscripts/build.sh +0 -175
- package/.zscripts/database-runtime-build.sh +0 -33
- package/.zscripts/dev.pid +0 -1
- package/.zscripts/dev.sh +0 -154
- package/.zscripts/mini-services-build.sh +0 -78
- package/.zscripts/mini-services-install.sh +0 -65
- package/.zscripts/mini-services-start.sh +0 -123
- package/.zscripts/python-runtime-build.sh +0 -120
- package/.zscripts/start.sh +0 -145
- package/CAPACITY_STUDY.md +0 -283
- package/CODE_OF_CONDUCT.md +0 -55
- package/CONTRIBUTING.md +0 -177
- package/Caddyfile +0 -23
- package/LICENSE +0 -204
- package/bun.lock +0 -1965
- package/components.json +0 -21
- package/db/custom.db +0 -0
- package/download/README.md +0 -1
- package/download/tokcalc-dark-calculator.png +0 -0
- package/download/tokcalc-dark-default.png +0 -0
- package/download/tokcalc-demo.webm +0 -0
- package/download/tokcalc-github-link.png +0 -0
- package/download/tokcalc-hydration-fixed.png +0 -0
- package/download/tokcalc-issue-resolved.png +0 -0
- package/download/tokcalc-light-mode.png +0 -0
- package/download/tokcalc-light-reference.png +0 -0
- package/download/tokcalc-long-context-qwen.png +0 -0
- package/download/tokcalc-long-context.png +0 -0
- package/download/tokcalc-og-image-preview.png +0 -0
- package/download/tokcalc-phase2-3.png +0 -0
- package/download/tokcalc-plain-english.png +0 -0
- package/download/tokcalc-preview.png +0 -0
- package/download/tokcalc-share-bvb.png +0 -0
- package/download/tokcalc-share-feature.png +0 -0
- package/download/tokcalc-tab-build-vs-buy.png +0 -0
- package/download/tokcalc-tab-calculator.png +0 -0
- package/download/tokcalc-tab-reference.png +0 -0
- package/eslint.config.mjs +0 -50
- package/examples/websocket/frontend.tsx +0 -196
- package/examples/websocket/server.ts +0 -138
- package/mini-services/.gitkeep +0 -0
- package/mini-services/mcp-server/README.md +0 -86
- package/mini-services/mcp-server/bun.lock +0 -202
- package/mini-services/mcp-server/index.ts +0 -504
- package/mini-services/mcp-server/package.json +0 -40
- package/next.config.ts +0 -12
- package/postcss.config.mjs +0 -5
- package/prisma/schema.prisma +0 -32
- package/public/google6f58ca6be85fa903.html +0 -1
- package/public/logo.svg +0 -29
- package/public/manifest.json +0 -51
- package/public/og-icon-256.png +0 -0
- package/public/og.png +0 -0
- package/public/robots.txt +0 -25
- package/public/sitemap.xml +0 -23
- package/public/tokcalc-demo.gif +0 -0
- package/scripts/og-template.html +0 -120
- package/scripts/render-og.mjs +0 -43
- package/server.json +0 -21
- package/src/app/api/pricing/aws/route.ts +0 -186
- package/src/app/api/pricing/azure/route.ts +0 -168
- package/src/app/api/pricing/gcp/route.ts +0 -230
- package/src/app/api/pricing/vast-ai/route.ts +0 -164
- package/src/app/api/route.ts +0 -5
- package/src/app/compare/h100-vs-h200/layout.tsx +0 -30
- package/src/app/compare/h100-vs-h200/page.tsx +0 -328
- package/src/app/globals.css +0 -122
- package/src/app/layout.tsx +0 -276
- package/src/app/page.tsx +0 -2670
- package/src/components/azure-live-pricing.tsx +0 -185
- package/src/components/benchmark-import.tsx +0 -340
- package/src/components/confidence-badge.tsx +0 -116
- package/src/components/live-pricing-comparison.tsx +0 -241
- package/src/components/theme-provider.tsx +0 -11
- package/src/components/theme-toggle.tsx +0 -55
- package/src/components/ui/accordion.tsx +0 -66
- package/src/components/ui/alert-dialog.tsx +0 -157
- package/src/components/ui/alert.tsx +0 -66
- package/src/components/ui/aspect-ratio.tsx +0 -11
- package/src/components/ui/avatar.tsx +0 -53
- package/src/components/ui/badge.tsx +0 -46
- package/src/components/ui/breadcrumb.tsx +0 -109
- package/src/components/ui/button.tsx +0 -59
- package/src/components/ui/calendar.tsx +0 -213
- package/src/components/ui/card.tsx +0 -92
- package/src/components/ui/carousel.tsx +0 -241
- package/src/components/ui/chart.tsx +0 -353
- package/src/components/ui/checkbox.tsx +0 -32
- package/src/components/ui/collapsible.tsx +0 -33
- package/src/components/ui/command.tsx +0 -184
- package/src/components/ui/context-menu.tsx +0 -252
- package/src/components/ui/dialog.tsx +0 -143
- package/src/components/ui/drawer.tsx +0 -135
- package/src/components/ui/dropdown-menu.tsx +0 -257
- package/src/components/ui/form.tsx +0 -167
- package/src/components/ui/hover-card.tsx +0 -44
- package/src/components/ui/input-otp.tsx +0 -77
- package/src/components/ui/input.tsx +0 -21
- package/src/components/ui/label.tsx +0 -24
- package/src/components/ui/menubar.tsx +0 -276
- package/src/components/ui/navigation-menu.tsx +0 -168
- package/src/components/ui/pagination.tsx +0 -127
- package/src/components/ui/popover.tsx +0 -48
- package/src/components/ui/progress.tsx +0 -31
- package/src/components/ui/radio-group.tsx +0 -45
- package/src/components/ui/resizable.tsx +0 -56
- package/src/components/ui/scroll-area.tsx +0 -58
- package/src/components/ui/select.tsx +0 -185
- package/src/components/ui/separator.tsx +0 -28
- package/src/components/ui/sheet.tsx +0 -139
- package/src/components/ui/sidebar.tsx +0 -726
- package/src/components/ui/skeleton.tsx +0 -13
- package/src/components/ui/slider.tsx +0 -63
- package/src/components/ui/sonner.tsx +0 -25
- package/src/components/ui/switch.tsx +0 -31
- package/src/components/ui/table.tsx +0 -116
- package/src/components/ui/tabs.tsx +0 -66
- package/src/components/ui/textarea.tsx +0 -18
- package/src/components/ui/toast.tsx +0 -129
- package/src/components/ui/toaster.tsx +0 -35
- package/src/components/ui/toggle-group.tsx +0 -73
- package/src/components/ui/toggle.tsx +0 -47
- package/src/components/ui/tooltip.tsx +0 -61
- package/src/components/vast-ai-live-pricing.tsx +0 -176
- package/src/hooks/use-mobile.ts +0 -19
- package/src/hooks/use-toast.ts +0 -194
- package/src/lib/benchmark-parser-sglang.ts +0 -150
- package/src/lib/benchmark-parser-tokcalc.ts +0 -247
- package/src/lib/benchmark-parser-trtllm.ts +0 -152
- package/src/lib/benchmark-parser-vllm.ts +0 -198
- package/src/lib/benchmark-schema.ts +0 -263
- package/src/lib/db.ts +0 -13
- package/src/lib/engine-presets.ts +0 -183
- package/src/lib/price-schema.ts +0 -141
- package/src/lib/token-calc.ts +0 -808
- package/src/lib/track.ts +0 -31
- package/src/lib/url-state.ts +0 -256
- package/src/lib/utils.ts +0 -6
- package/tailwind.config.ts +0 -64
- package/tests/database-runtime-build.sh +0 -75
- package/tests/python-runtime-build.sh +0 -64
- package/tests/python-runtime-container.sh +0 -31
- package/tool-results/bash_1789888171144_2c5381860539.txt +0 -161
- package/tool-results/bash_1789888175925_49c53ba3c61b.txt +0 -191
- package/tool-results/bash_1789888181202_49c53ba3c61b.txt +0 -191
- package/tool-results/bash_1789888195219_4a86a5c91411.txt +0 -200
- package/tool-results/bash_1789888203128_6cca13c71b47.txt +0 -199
- package/tool-results/bash_1789929256963_2a52aff0d0a8.txt +0 -160
- package/tool-results/read_1789888151021_69f58eec6a5b.txt +0 -653
- package/tool-results/read_1789888153837_1d3a8bfc2a94.txt +0 -653
- package/tool-results/read_1789888163087_ccc406d47505.txt +0 -122
- package/tool-results/read_1789888167347_67d1d7c9830a.txt +0 -122
- package/tool-results/read_1789929252529_d90e8f383a25.txt +0 -285
- package/tsconfig.json +0 -42
- package/upload/Pasted Content_1789887800864.txt +0 -652
- package/upload/Pasted Content_1789887909561.txt +0 -652
- package/upload/Pasted Content_1789887918428.txt +0 -652
- package/upload/Pasted Content_1789887959420.txt +0 -652
- package/upload/Pasted Content_1789888020485.txt +0 -652
- package/upload/Pasted Content_1789888058079.txt +0 -652
- package/upload/Pasted Content_1789888885033.txt +0 -686
- package/upload/Pasted Content_1789928912741.txt +0 -285
- package/upload/Pasted Content_1789928938402.txt +0 -285
- package/upload/Pasted Content_1789929160389.txt +0 -285
- package/upload/Pasted Content_1789929176660.txt +0 -285
- package/upload/issue_vision.json +0 -28
- package/upload/pasted_image_1789883175209.png +0 -0
- package/upload/pasted_image_1789899056690.png +0 -0
- package/upload/pasted_image_1789900371483.png +0 -0
- package/upload/pasted_image_1789900472823.png +0 -0
- package/upload/pasted_image_1789900490374.png +0 -0
- package/upload/pasted_image_1789900585552.png +0 -0
- package/upload/pasted_image_1789900606519.png +0 -0
- package/upload/pasted_image_1789901598705.png +0 -0
- package/upload/pasted_image_1789901613545.png +0 -0
- package/upload/pasted_image_1789978382674.png +0 -0
- package/upload/pasted_image_1789978392749.png +0 -0
- package/upload/pasted_image_1789978474879.png +0 -0
- package/upload/pasted_image_1789978523652.png +0 -0
- package/upload/pasted_image_1789984219089.png +0 -0
- package/upload/pasted_image_1789984491896.png +0 -0
- package/upload/pasted_image_1789985017950.png +0 -0
- package/upload/pasted_image_1789985036765.png +0 -0
- package/upload/pasted_image_1789985049848.png +0 -0
- package/upload/pasted_image_1790002427833.png +0 -0
- package/upload/pasted_image_1790002659944.png +0 -0
- package/upload/pasted_image_1790037038476.png +0 -0
- package/upload/screenshot_analysis.json +0 -28
- package/upload/vision_output.json +0 -28
|
@@ -1,686 +0,0 @@
|
|
|
1
|
-
# Research scope note
|
|
2
|
-
|
|
3
|
-
Your requested scope is broad enough to be a small data product: hundreds of model, accelerator, quantization, and price records—many of which are volatile, unavailable, proprietary, or not meaningfully comparable. The most useful implementation decision is therefore **not** to hard-code a giant static matrix, but to build tokcalc around versioned data adapters, confidence labels, and workload-specific calculators.
|
|
4
|
-
|
|
5
|
-
Also, several requested names are not valid current public model/hardware product families or do not expose the architectural data required for local-inference math:
|
|
6
|
-
|
|
7
|
-
- **Llama 4 Behemoth** was announced as a teacher model / not broadly released, so exact deployable weights/configuration should not be presented as a selectable local model without an official artifact.
|
|
8
|
-
- **OpenAI o1/o3-mini, Claude Thinking, Gemini Flash Thinking** are closed APIs; token pricing can be modeled, but layer count, hidden size, KV heads, and FLOPs cannot be reliably modeled.
|
|
9
|
-
- **“Llama 4 Vision”** is not a separate official model name in the way Qwen-VL or Pixtral are; multimodal capability needs to be represented per official model/config.
|
|
10
|
-
- **M4 Ultra** has not been a normal shipping Apple SKU in the same way M2 Ultra/M3 Ultra configurations have been; do not add it as available hardware until Apple publishes it.
|
|
11
|
-
- **V3.5 / R2** should not be assumed to exist solely because a model vendor has earlier V3/R1 naming.
|
|
12
|
-
- GPU rental pricing is market-, region-, commitment-, and instance-size-specific, particularly on Vast.ai and spot-like products. Store observations with timestamp, region, availability, billing granularity, and source URL rather than a single “truth” value.
|
|
13
|
-
|
|
14
|
-
The recommendations below are optimized for what tokcalc can calculate transparently and defensibly.
|
|
15
|
-
|
|
16
|
-
# SECTION A — TRENDING FEATURES (Sept 2025 - 2026)
|
|
17
|
-
|
|
18
|
-
## 1. Continuous batching and paged attention
|
|
19
|
-
|
|
20
|
-
| Item | Recommendation |
|
|
21
|
-
|---|---|
|
|
22
|
-
| What it is | Continuous batching admits and retires requests at token-generation boundaries instead of waiting for every request in a fixed batch to finish; paged attention stores KV-cache blocks non-contiguously, reducing fragmentation and raising usable cache capacity. |
|
|
23
|
-
| Why users want it | This is the difference between a toy “batch size × tokens/sec” estimate and production serving. Under concurrent mixed-length requests, static batching leaves GPU work idle or makes short requests wait for long ones. Published vLLM material describes large throughput gains over naïve Hugging Face generation—often cited as roughly **2–24×**, depending on model, request mix, latency target, and baseline. [vLLM paper](https://arxiv.org/abs/2309.06180) |
|
|
24
|
-
| Calculator inputs | Request arrival rate, prompt-length distribution, output-length distribution, maximum concurrent sequences, KV-block size, GPU memory utilization target, scheduler type, TTFT target, inter-token latency target. |
|
|
25
|
-
| Implementation | Model a queue simulator rather than one static formula. Start with selectable scheduler presets: “naïve static batch,” “continuous batching,” and “continuous batching + paged KV.” Output admitted concurrency, KV-cache capacity, aggregate output tok/s, TTFT estimate, and P50/P95 queue wait. |
|
|
26
|
-
| Complexity | **High** — ~40–80 hours for a useful transparent approximation; significantly more for a trace-driven discrete-event simulator. |
|
|
27
|
-
| Priority | **Must-have** |
|
|
28
|
-
|
|
29
|
-
**Implementation-ready approximation**
|
|
30
|
-
|
|
31
|
-
```ts
|
|
32
|
-
type SchedulerMode =
|
|
33
|
-
| "static_batch"
|
|
34
|
-
| "continuous_batch"
|
|
35
|
-
| "continuous_paged";
|
|
36
|
-
|
|
37
|
-
effectiveConcurrency =
|
|
38
|
-
min(
|
|
39
|
-
requestedConcurrency,
|
|
40
|
-
floor(availableKvBytes / kvBytesPerRequestAtTargetLength),
|
|
41
|
-
);
|
|
42
|
-
|
|
43
|
-
aggregateTokPerSec =
|
|
44
|
-
baseDecodeTokPerSec *
|
|
45
|
-
schedulerEfficiency[mode] *
|
|
46
|
-
concurrencyEfficiency(effectiveConcurrency, model, gpu);
|
|
47
|
-
|
|
48
|
-
queueDelay =
|
|
49
|
-
estimateQueueDelay(arrivalRate, requestServiceTime, effectiveConcurrency);
|
|
50
|
-
```
|
|
51
|
-
|
|
52
|
-
Do not expose “2–24×” as a universal multiplier. Present it as a **benchmark-derived range**, then show the selected multiplier and assumptions. Recent serving guidance commonly describes **2–4×** gains versus static batching in high-concurrency mixed-output settings, while the larger vLLM comparison figures depend heavily on the naïve baseline and workload. [vLLM paper](https://arxiv.org/abs/2309.06180) [RunPod vLLM overview](https://www.runpod.io/articles/guides/vllm-pagedattention-continuous-batching)
|
|
53
|
-
|
|
54
|
-
## 2. Prefix and prompt caching
|
|
55
|
-
|
|
56
|
-
| Item | Recommendation |
|
|
57
|
-
|---|---|
|
|
58
|
-
| What it is | Prefix/prompt caching reuses the computation and/or KV cache for a repeated prompt prefix—such as a system prompt, RAG corpus header, coding repository context, or shared chat history—rather than prefilling it again for every request. |
|
|
59
|
-
| Why users want it | It can sharply reduce time-to-first-token, prefill compute, GPU memory traffic, and API bills when prompts share long prefixes. Anthropic’s documented prompt-cache reads are priced at **10% of base input-token price**; its 5-minute cache writes cost **1.25×** base input and one-hour writes cost **2×** base input. [Anthropic prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) |
|
|
60
|
-
| OpenAI behavior | OpenAI prompt caching is automatic for supported models and matching prefixes; its pricing tables distinguish ordinary input from cached input. Treat eligibility, minimum prompt size, and supported model set as provider-configurable metadata because they change. [OpenAI prompt caching](https://platform.openai.com/docs/guides/prompt-caching) |
|
|
61
|
-
| Calculator inputs | Shared-prefix tokens, cache hit rate, TTL, cache-write multiplier, cached-read multiplier, number of requests, normal input price, local prefill tok/s, and KV-cache residency. |
|
|
62
|
-
| Implementation | Add a “Shared Prefix / Cache” panel: prefix size, hit rate, TTL, request count, provider mode (self-hosted/OpenAI/Anthropic/Gemini), and cache policy. Show cost avoided, prefill FLOPs avoided, saved TTFT, and cache break-even hit rate. |
|
|
63
|
-
| Complexity | **Medium** — ~16–30 hours. |
|
|
64
|
-
| Priority | **Must-have** |
|
|
65
|
-
|
|
66
|
-
**Useful equations**
|
|
67
|
-
|
|
68
|
-
\[
|
|
69
|
-
\text{API input cost} =
|
|
70
|
-
N \cdot \left[
|
|
71
|
-
(1-h) \cdot T_p \cdot P_{\text{write}}
|
|
72
|
-
+
|
|
73
|
-
h \cdot T_p \cdot P_{\text{read}}
|
|
74
|
-
+
|
|
75
|
-
T_u \cdot P_{\text{input}}
|
|
76
|
-
\right]
|
|
77
|
-
\]
|
|
78
|
-
|
|
79
|
-
Where \(N\) is request count, \(h\) is cache-hit rate, \(T_p\) is reusable prefix tokens, and \(T_u\) is uncached input tokens.
|
|
80
|
-
|
|
81
|
-
For self-hosting, show:
|
|
82
|
-
|
|
83
|
-
\[
|
|
84
|
-
\text{prefill tokens avoided} = N \cdot h \cdot T_p
|
|
85
|
-
\]
|
|
86
|
-
|
|
87
|
-
Anthropic documents cache-read pricing at 0.1× base input; consequently, a cache hit can represent a **90% reduction for those cached input tokens**, though the total request saving depends on how much of the request is cached. [Anthropic prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching)
|
|
88
|
-
|
|
89
|
-
## 3. Chunked prefill
|
|
90
|
-
|
|
91
|
-
| Item | Recommendation |
|
|
92
|
-
|---|---|
|
|
93
|
-
| What it is | Chunked prefill breaks a long prompt’s prefill work into token chunks and interleaves that work with decode steps for active requests. |
|
|
94
|
-
| Why users want it | A very long prompt can otherwise monopolize the GPU and harm inter-token latency for users who are already receiving generated output. Chunking improves fairness and tail latency rather than magically increasing every workload’s raw model throughput. |
|
|
95
|
-
| Implementation | Add `prefillChunkTokens`, toggle “prioritize decode,” and a chart that shows the TTFT vs inter-token-latency trade-off. Make it clear that chunking shifts scheduling behavior and may slightly reduce prefill efficiency. |
|
|
96
|
-
| Complexity | **Medium** — ~20–40 hours as part of the queue/scheduler model. |
|
|
97
|
-
| Priority | **Must-have** once continuous batching exists. |
|
|
98
|
-
|
|
99
|
-
vLLM documents chunked prefill as a scheduler feature that allows prefills to be chunked and scheduled alongside decode work; it is especially useful in mixed prompt-length and latency-sensitive serving. [vLLM optimization guide](https://docs.vllm.ai/en/latest/configuration/optimization/)
|
|
100
|
-
|
|
101
|
-
## 4. Prefill/decode disaggregation
|
|
102
|
-
|
|
103
|
-
| Item | Recommendation |
|
|
104
|
-
|---|---|
|
|
105
|
-
| What it is | PD disaggregation assigns prompt-prefill and autoregressive decode to separate GPU pools, connected by KV-cache transfer, because their bottlenecks and scaling behavior differ. |
|
|
106
|
-
| Why users want it | Prefill is generally compute-heavy and bursty; decode is memory-bandwidth- and latency-sensitive. Separating them helps operators independently scale TTFT capacity and sustained generation capacity instead of overprovisioning one mixed fleet. |
|
|
107
|
-
| Who uses it | This architecture is actively explored and implemented in production-oriented serving systems and research, including SGLang and vLLM ecosystem work; it is a central serving direction rather than a universally necessary default. [SGLang PD disaggregation docs](https://docs.sglang.ai/advanced_features/pd_disaggregation.html) [DistServe paper](https://arxiv.org/abs/2401.09670) |
|
|
108
|
-
| Throughput gain | Do **not** bake in one multiplier. Gains depend on workload mix, request skew, cache-transfer cost, network fabric, and the degree to which the monolithic fleet was imbalanced. Calculate fleet utilization and SLO capacity instead. |
|
|
109
|
-
| Implementation | Create a dual-pool calculator: number/type of prefill GPUs, number/type of decode GPUs, NIC bandwidth, KV-transfer bytes/request, prompt/output distributions, and SLOs. Return bottleneck, required replicas, capacity, cost/request, and required network throughput. |
|
|
110
|
-
| Complexity | **High** — ~60–120 hours for a credible planning calculator. |
|
|
111
|
-
| Priority | **Future**, but begin designing the schema now. |
|
|
112
|
-
|
|
113
|
-
A useful first-order KV-transfer estimate is:
|
|
114
|
-
|
|
115
|
-
\[
|
|
116
|
-
\text{KV bytes transferred} =
|
|
117
|
-
2 \cdot L \cdot T_{\text{prompt}} \cdot H_{\text{kv}} \cdot D_h \cdot B_{\text{KV}}
|
|
118
|
-
\]
|
|
119
|
-
|
|
120
|
-
where \(L\) is layers, \(T_{\text{prompt}}\) is prompt length, \(H_{\text{kv}}\) is KV-head count, \(D_h\) is head dimension, and \(B_{\text{KV}}\) is bytes per KV value.
|
|
121
|
-
|
|
122
|
-
## 5. Multi-LoRA serving
|
|
123
|
-
|
|
124
|
-
| Item | Recommendation |
|
|
125
|
-
|---|---|
|
|
126
|
-
| What it is | Multi-LoRA serving hosts one shared base model while dynamically applying many small low-rank adapters for tenants, customers, domains, or tasks. |
|
|
127
|
-
| Why users want it | It avoids loading or duplicating an entire base model for every fine-tune, making multi-tenant customization much cheaper and denser. Punica and S-LoRA specifically target efficient serving of many LoRA adapters. [Punica](https://arxiv.org/abs/2310.18547) [S-LoRA](https://arxiv.org/abs/2311.03285) |
|
|
128
|
-
| Throughput math | Adapter weight storage is roughly proportional to \(r(d_{in}+d_{out})\) per adapted matrix, times adapted modules and layers; runtime overhead depends on rank, batching mix, fused kernels, and adapter swapping/cache misses. It is not safely modeled by a single “LoRA = X% slower” constant. |
|
|
129
|
-
| Implementation | Add a Multi-LoRA tab: base model, adapter rank, target modules, number of adapters resident, adapter cache size, adapter popularity distribution, concurrent adapter mix, and engine. Show adapter VRAM, estimated cache-miss risk, and a throughput-overhead range. |
|
|
130
|
-
| Complexity | **Medium** — ~24–50 hours. |
|
|
131
|
-
| Priority | **Nice** |
|
|
132
|
-
|
|
133
|
-
## 6. Reasoning-model workloads
|
|
134
|
-
|
|
135
|
-
| Item | Recommendation |
|
|
136
|
-
|---|---|
|
|
137
|
-
| What it is | Reasoning-oriented models often spend substantially more generated tokens on hidden or visible deliberation, tool use, verification, and retries before producing a concise final response. |
|
|
138
|
-
| Why users want it | A calculator based only on user-visible output underestimates GPU time, API costs, queue time, and latency for reasoning applications. |
|
|
139
|
-
| Implementation | Add “reasoning multiplier” and “hidden reasoning-token budget” as explicit workload inputs rather than assigning undocumented internal-token assumptions to closed models. Support three presets: standard chat, bounded reasoning, and agentic reasoning. |
|
|
140
|
-
| Complexity | **Low** — ~8–16 hours. |
|
|
141
|
-
| Priority | **Must-have** |
|
|
142
|
-
|
|
143
|
-
For closed systems, distinguish **billed output tokens**, **visible output tokens**, and **unknown provider-internal compute**. OpenAI pricing and model documentation should be treated as the source of record for billing behavior; do not infer architecture from observed latency. [OpenAI pricing](https://openai.com/api/pricing/)
|
|
144
|
-
|
|
145
|
-
## 7. Agentic workloads
|
|
146
|
-
|
|
147
|
-
| Item | Recommendation |
|
|
148
|
-
|---|---|
|
|
149
|
-
| What it is | Agentic systems perform multiple model calls interleaved with tool invocations, retrieval, state updates, retries, and growing conversation context. |
|
|
150
|
-
| Why users want it | The cost and latency of an “answer” become a workflow total, not one request: a six-step tool-using agent can have six prefills, six decodes, tool latency, and repeated context. |
|
|
151
|
-
| Implementation | Add an agent-workflow calculator with turns, average tool calls/turn, tool latency, retry rate, context growth/turn, cache hit rate, input/output tokens/call, and model routing. Report end-to-end latency decomposition and dollars per completed task. |
|
|
152
|
-
| Complexity | **Medium** — ~20–40 hours. |
|
|
153
|
-
| Priority | **Must-have** |
|
|
154
|
-
|
|
155
|
-
## 8. Long-context serving
|
|
156
|
-
|
|
157
|
-
| Item | Recommendation |
|
|
158
|
-
|---|---|
|
|
159
|
-
| What it is | Long-context serving processes prompts from 128K to 1M+ tokens, making prefill time, KV-cache memory, and distributed attention communication dominate far more than ordinary chat workloads. |
|
|
160
|
-
| Why users want it | A model that “fits” in VRAM at 8K context may have very low concurrency or be infeasible at 128K because each live request retains a large KV cache. |
|
|
161
|
-
| RingAttention / YaRN | Ring Attention distributes long-sequence attention across devices; YaRN is a context-window extension approach based on RoPE scaling. They solve different issues: distributed execution versus positional extrapolation/extension. [Ring Attention](https://arxiv.org/abs/2310.01889) [YaRN](https://arxiv.org/abs/2309.00071) |
|
|
162
|
-
| Implementation | Promote context length from a simple input into a first-class workload distribution. Add KV-cache-by-context charts, maximum concurrency at 8K/32K/128K/256K/1M, prefill duration, and required TP/CP topology. |
|
|
163
|
-
| Complexity | **High** — ~30–70 hours, much of it data/model validation. |
|
|
164
|
-
| Priority | **Must-have** |
|
|
165
|
-
|
|
166
|
-
For decoder-only transformers:
|
|
167
|
-
|
|
168
|
-
\[
|
|
169
|
-
\text{KV bytes/request} =
|
|
170
|
-
2 \cdot L \cdot T \cdot H_{\text{kv}} \cdot D_h \cdot B
|
|
171
|
-
\]
|
|
172
|
-
|
|
173
|
-
This is the most important formula tokcalc should visibly expose. The factor of 2 stores both keys and values.
|
|
174
|
-
|
|
175
|
-
## 9. Embeddings calculator
|
|
176
|
-
|
|
177
|
-
| Item | Recommendation |
|
|
178
|
-
|---|---|
|
|
179
|
-
| What it is | An embeddings calculator estimates vectors/sec, documents/sec, tokens/sec, cost, and latency for encoder or embedding APIs rather than autoregressive generation. |
|
|
180
|
-
| Why users want it | Embedding workloads are often high-batch, fixed-output, encoder-style jobs where decode-token metrics are irrelevant; RAG teams need ingestion and re-indexing capacity planning. |
|
|
181
|
-
| Implementation | New mode: model, sequence length, batch size, embedding dimension, normalization, GPU, and target documents/day. Outputs throughput, GPU-hours, cost/M input tokens, index-vector storage, and vector-store egress estimate. |
|
|
182
|
-
| Complexity | **Medium** — ~24–48 hours. |
|
|
183
|
-
| Priority | **Must-have** |
|
|
184
|
-
|
|
185
|
-
Separate “embedding throughput” from “generation throughput.” Models such as BGE-M3 are designed for multilingual and long-text retrieval use cases, while API embedding models bill input tokens without autoregressive output-token pricing. [BGE-M3 model card](https://huggingface.co/BAAI/bge-m3) [OpenAI embeddings guide](https://platform.openai.com/docs/guides/embeddings)
|
|
186
|
-
|
|
187
|
-
## 10. Vision-language model serving
|
|
188
|
-
|
|
189
|
-
| Item | Recommendation |
|
|
190
|
-
|---|---|
|
|
191
|
-
| What it is | Vision-language model inference converts images or video frames into visual embeddings/tokens that are combined with text context before generation. |
|
|
192
|
-
| Why users want it | Image resolution, tiling, number of images, and video frames can dominate prompt length, memory use, and cost even when the typed text is short. |
|
|
193
|
-
| Implementation | Add modality inputs: image width, height, count, tile strategy, patch size, vision-token count override, video frames, and image pricing mode. Report effective prompt tokens, vision-encoder cost, prefill cost, and total latency. |
|
|
194
|
-
| Complexity | **High** — ~35–70 hours because tokenization rules vary by model. |
|
|
195
|
-
| Priority | **Nice** |
|
|
196
|
-
|
|
197
|
-
Do not use one universal “image = N tokens” constant. Model cards and provider docs define different image preprocessing, tiling, resolution, and token-equivalence rules. [Qwen2-VL model card](https://huggingface.co/Qwen/Qwen2-VL-72B-Instruct) [Pixtral 12B model card](https://huggingface.co/mistralai/Pixtral-12B-2409)
|
|
198
|
-
|
|
199
|
-
## 11. Training and fine-tuning calculator
|
|
200
|
-
|
|
201
|
-
| Item | Recommendation |
|
|
202
|
-
|---|---|
|
|
203
|
-
| What it is | A training calculator estimates memory, tokens/day, step time, GPU-hours, and cost for LoRA, QLoRA, full supervised fine-tuning, and possibly continued pretraining. |
|
|
204
|
-
| Why users want it | Teams deciding whether to prompt, RAG, fine-tune, or self-host need training cost next to inference cost. LoRA reduces trainable parameters by introducing low-rank adapters, and QLoRA combines quantized base weights with adapters to reduce memory requirements. [LoRA paper](https://arxiv.org/abs/2106.09685) [QLoRA paper](https://arxiv.org/abs/2305.14314) |
|
|
205
|
-
| Implementation | A separate calculator, not a mode bolted onto inference. Inputs: model architecture, sequence length, global batch, microbatch, precision, optimizer, activation checkpointing, ZeRO/FSDP level, LoRA rank, training tokens, GPU count. |
|
|
206
|
-
| Complexity | **High** — ~60–120 hours. |
|
|
207
|
-
| Priority | **Future** |
|
|
208
|
-
|
|
209
|
-
## 12. Build-vs-buy calculator
|
|
210
|
-
|
|
211
|
-
| Item | Recommendation |
|
|
212
|
-
|---|---|
|
|
213
|
-
| What it is | A build-vs-buy calculator compares self-hosted GPU capacity and operational overhead against token-priced API alternatives for the same workload. |
|
|
214
|
-
| Why users want it | It answers the practical question: “At our utilization and prompt/output mix, when does renting GPUs beat an API?” |
|
|
215
|
-
| Implementation | Take cloud GPU hourly price, usable aggregate tok/s, utilization, model/license, deployment overhead, prompt/output mix, cache hit rate, and API provider prices. Output break-even requests/day, cost/M input/output tokens, total monthly cost, and sensitivity ranges. |
|
|
216
|
-
| Complexity | **Medium** — ~24–45 hours once pricing and throughput data are normalized. |
|
|
217
|
-
| Priority | **Must-have** |
|
|
218
|
-
|
|
219
|
-
Use a utilization-sensitive formula:
|
|
220
|
-
|
|
221
|
-
\[
|
|
222
|
-
\text{self-hosted cost per million generated tokens} =
|
|
223
|
-
\frac{\text{GPU hourly price}}
|
|
224
|
-
{3600 \cdot \text{effective output tok/s} \cdot \text{utilization}}
|
|
225
|
-
\cdot 10^6
|
|
226
|
-
\]
|
|
227
|
-
|
|
228
|
-
The decisive term is **effective utilization**, not synthetic maximum throughput.
|
|
229
|
-
|
|
230
|
-
## 13. Energy and carbon
|
|
231
|
-
|
|
232
|
-
| Item | Recommendation |
|
|
233
|
-
|---|---|
|
|
234
|
-
| What it is | This estimates energy per request or per million tokens and converts it to location-dependent carbon emissions using a grid-emissions factor. |
|
|
235
|
-
| Why users want it | Procurement, enterprise reporting, and capacity planning increasingly require cost plus environmental impact; energy also reflects operating cost and cooling load. |
|
|
236
|
-
| Implementation | Inputs: GPU power draw or measured joules/token, PUE, utilization, region/grid intensity, and token counts. Output kWh/M tokens, kgCO₂e/M tokens, and uncertainty range. Use region-specific carbon data rather than a global constant. |
|
|
237
|
-
| Complexity | **Medium** — ~16–32 hours. |
|
|
238
|
-
| Priority | **Nice** |
|
|
239
|
-
|
|
240
|
-
Use:
|
|
241
|
-
|
|
242
|
-
\[
|
|
243
|
-
\text{kWh/M tokens} =
|
|
244
|
-
\frac{\text{average IT kW} \cdot \text{PUE}}
|
|
245
|
-
{\text{effective tok/s} \cdot 3600}
|
|
246
|
-
\cdot 10^6
|
|
247
|
-
\]
|
|
248
|
-
|
|
249
|
-
\[
|
|
250
|
-
\text{kgCO}_2\text{e/M tokens} =
|
|
251
|
-
\text{kWh/M tokens} \cdot \text{grid kgCO}_2\text{e/kWh}
|
|
252
|
-
\]
|
|
253
|
-
|
|
254
|
-
The Green Software Foundation’s SCI specification is a useful framing reference, but tokcalc should state that output is an estimate and disclose assumptions. [SCI Specification](https://sci.greensoftware.foundation/)
|
|
255
|
-
|
|
256
|
-
## 14. Serverless cold starts
|
|
257
|
-
|
|
258
|
-
| Item | Recommendation |
|
|
259
|
-
|---|---|
|
|
260
|
-
| What it is | Cold-start time is the delay to provision a worker, attach a GPU, pull/initialize the image and model, and become ready after a period of zero warm capacity. |
|
|
261
|
-
| Why users want it | It determines whether serverless is appropriate for interactive traffic, bursty batch work, or only asynchronous jobs. |
|
|
262
|
-
| Implementation | Add endpoint mode: warm, scale-to-zero, and provisioned minimum replicas. Inputs: model artifact size, download/cache state, GPU allocation time, container startup, engine initialization, and warm timeout. Report warm P50/P95 vs cold P50/P95 separately. |
|
|
263
|
-
| Complexity | **Medium** — ~16–28 hours for a planning model; production figures require provider-specific measurements. |
|
|
264
|
-
| Priority | **Nice** |
|
|
265
|
-
|
|
266
|
-
Do not present a universal cold-start number: provider, region, image/model cache state, model size, concurrency configuration, and persistent volumes materially change it. Modal documents cold-start behavior and configuration as application-level concerns. [Modal performance guide](https://modal.com/docs/guide/cold-start)
|
|
267
|
-
|
|
268
|
-
# SECTION B — MORE MODELS TO ADD
|
|
269
|
-
|
|
270
|
-
## Data-model recommendation
|
|
271
|
-
|
|
272
|
-
Store architectural records as versioned JSON, sourced from the model’s `config.json` where available, rather than manually maintained prose tables. A useful normalized schema is:
|
|
273
|
-
|
|
274
|
-
```ts
|
|
275
|
-
type ModelSpec = {
|
|
276
|
-
id: string;
|
|
277
|
-
officialName: string;
|
|
278
|
-
provider: string;
|
|
279
|
-
family: string;
|
|
280
|
-
modality: "text" | "vision-language" | "embedding" | "audio" | "closed-api";
|
|
281
|
-
status: "released" | "announced" | "unreleased" | "deprecated";
|
|
282
|
-
releaseDate?: string;
|
|
283
|
-
parametersTotalB?: number;
|
|
284
|
-
parametersActiveB?: number;
|
|
285
|
-
architecture?: {
|
|
286
|
-
numLayers?: number;
|
|
287
|
-
hiddenSize?: number;
|
|
288
|
-
numAttentionHeads?: number;
|
|
289
|
-
numKeyValueHeads?: number;
|
|
290
|
-
headDim?: number;
|
|
291
|
-
vocabSize?: number;
|
|
292
|
-
maxPositionEmbeddings?: number;
|
|
293
|
-
moeNumExperts?: number;
|
|
294
|
-
moeTopK?: number;
|
|
295
|
-
};
|
|
296
|
-
contextWindow?: number;
|
|
297
|
-
sourceUrls: string[];
|
|
298
|
-
confidence: "official-config" | "official-card" | "vendor-claim" | "estimated";
|
|
299
|
-
};
|
|
300
|
-
```
|
|
301
|
-
|
|
302
|
-
For sparse MoE models, distinguish **total stored parameters** from **active parameters/token**. This matters for weight loading, VRAM, and throughput.
|
|
303
|
-
|
|
304
|
-
## Recommended additions first
|
|
305
|
-
|
|
306
|
-
| Model family / representative release | Deployment type | Why add | Architecture data source |
|
|
307
|
-
|---|---:|---|---|
|
|
308
|
-
| Llama 4 Scout / Maverick | Open-weight multimodal MoE | High user demand and current family relevance; model cards/configs should drive exact selectable variants rather than copied architecture claims. | [Meta Llama collection](https://huggingface.co/meta-llama) |
|
|
309
|
-
| Qwen3 | Open-weight text / reasoning variants | Broad deployment use, multiple sizes, strong coding/reasoning interest. | [Qwen organization](https://huggingface.co/Qwen) |
|
|
310
|
-
| Qwen2.5-VL / Qwen2-VL | Open-weight VLM | Needed for image-token-aware workload planning. | [Qwen2-VL collection](https://huggingface.co/collections/Qwen/qwen2-vl-66e81a666a0dbd6a8321d75b) |
|
|
311
|
-
| DeepSeek-V3 / DeepSeek-R1 | Open-weight MoE / reasoning | Essential for MoE and reasoning workload modes. | [DeepSeek-V3](https://huggingface.co/deepseek-ai/DeepSeek-V3) / [DeepSeek-R1](https://huggingface.co/deepseek-ai/DeepSeek-R1) |
|
|
312
|
-
| Mistral Large 2 | Open-weight / research license as applicable | Large dense model benchmark point, high deployment interest. | [Mistral Large Instruct 2411](https://huggingface.co/mistralai/Mistral-Large-Instruct-2411) |
|
|
313
|
-
| Codestral / Devstral-class coding models | Code generation | Coding has distinct prompt lengths, tool loops, and output lengths. | [Mistral models](https://huggingface.co/mistralai) |
|
|
314
|
-
| Pixtral 12B | Open-weight VLM | Establishes a practical image-input calculator benchmark. | [Pixtral 12B](https://huggingface.co/mistralai/Pixtral-12B-2409) |
|
|
315
|
-
| Gemma 3 | Open-weight multimodal family | Widely used and a natural successor to Gemma 2. | [Google Gemma](https://huggingface.co/google) |
|
|
316
|
-
| Phi-4 / Phi-3.5 | Small-model deployment | Important for edge, cost-sensitive, and CPU/consumer-GPU scenarios. | [Microsoft Phi collection](https://huggingface.co/microsoft) |
|
|
317
|
-
| SmolLM2 | Small open models | Good low-end benchmark family. | [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct) |
|
|
318
|
-
| Falcon 3 | Open models | Regional/community deployment interest. | [Falcon 3](https://huggingface.co/tiiuae/Falcon3-10B-Instruct) |
|
|
319
|
-
| OLMo 2 | Fully open research model family | Strong fit for transparent benchmarking and reproducible configs. | [OLMo 2](https://huggingface.co/allenai/OLMo-2-1124-13B-Instruct) |
|
|
320
|
-
| DeepSeek-Coder-V2 | MoE code model | Must-have for code-specific comparison. | [DeepSeek-Coder-V2](https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Instruct) |
|
|
321
|
-
| Qwen2.5-Coder | Code model family | Widely used local coding family. | [Qwen2.5-Coder](https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct) |
|
|
322
|
-
| Code Llama | Older but still deployed | Add as “legacy / comparison” rather than a roadmap focus. | [Code Llama](https://huggingface.co/codellama) |
|
|
323
|
-
| QwQ | Open reasoning family | Useful for reasoning-budget presets and comparison. | [QwQ](https://huggingface.co/Qwen) |
|
|
324
|
-
| BGE-M3 | Embedding | Supports a dedicated embeddings mode, including multilingual retrieval. | [BGE-M3](https://huggingface.co/BAAI/bge-m3) |
|
|
325
|
-
| multilingual-e5-large | Embedding | Standard baseline for retrieval workloads. | [multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large) |
|
|
326
|
-
| GTE-large | Embedding | Widely used open retrieval baseline. | [gte-large](https://huggingface.co/thenlper/gte-large) |
|
|
327
|
-
| Nomic Embed | Embedding | Popular long-context embedding option. | [Nomic Embed Text](https://huggingface.co/nomic-ai/nomic-embed-text-v1.5) |
|
|
328
|
-
| Jina Embeddings v3 | Embedding | Multilingual/task-adapter embedding comparison. | [Jina Embeddings v3](https://huggingface.co/jinaai/jina-embeddings-v3) |
|
|
329
|
-
| InternVL | VLM | Strong multimodal benchmark family. | [InternVL collection](https://huggingface.co/OpenGVLab) |
|
|
330
|
-
| CogVLM2 | VLM | Add to compare open VLM architectures. | [CogVLM2](https://huggingface.co/THUDM/cogvlm2-llama3-chat-19B) |
|
|
331
|
-
|
|
332
|
-
## Model-family treatment
|
|
333
|
-
|
|
334
|
-
| Requested group | Tokcalc treatment | Priority | Notes |
|
|
335
|
-
|---|---|---:|---|
|
|
336
|
-
| Llama 4 Scout / Maverick / Behemoth | Add released public variants whose official config/model card is available; mark Behemoth as announced/unreleased unless a deployable official release appears | Must-have | Do not infer exact architecture from press reporting. [Meta Llama collection](https://huggingface.co/meta-llama) |
|
|
337
|
-
| Qwen3, Qwen2.5-VL, Qwen MoE, Qwen Coder | Add as separate text, VLM, MoE, and coding categories | Must-have | Pull config fields automatically in CI. [Qwen organization](https://huggingface.co/Qwen) |
|
|
338
|
-
| DeepSeek V3/R1 and newer | Add V3 and R1; only add a newer family when DeepSeek publishes official model/API documentation | Must-have | Never create speculative “V3.5” or “R2” records. [DeepSeek organization](https://huggingface.co/deepseek-ai) |
|
|
339
|
-
| Mistral Large 2, Codestral, Ministral, Pixtral | Add dense, code, small, and vision workload profiles | Must-have | Pixtral belongs in VLM math, not just a text-model dropdown. [Mistral organization](https://huggingface.co/mistralai) |
|
|
340
|
-
| Gemma 3 / Gemma 2 | Add current public released variants | Nice | Represent context and multimodal support per variant. [Google organization](https://huggingface.co/google) |
|
|
341
|
-
| Phi-4 / Phi-3.5 | Add as efficient/small-model benchmarks | Nice | Relevant for laptop/edge comparisons. [Microsoft organization](https://huggingface.co/microsoft) |
|
|
342
|
-
| WizardLM / Orca | Add only as legacy/reference, not primary roadmap | Future | Earlier families; source configs must still validate values. |
|
|
343
|
-
| BitNet b1.58 | Add as experimental architecture / quantization scenario | Future | Do not claim broad production engine parity. [BitNet b1.58 paper](https://arxiv.org/abs/2402.17764) |
|
|
344
|
-
| SmolLM2, TinyLlama, OLMo, Falcon 3, Yi 1.5, InternLM 2.5 | Create a “small/open/legacy” collection | Nice | Valuable for home-lab users and CPU/Apple scenarios. |
|
|
345
|
-
| OpenAI o1-mini/o3-mini, Gemini Thinking | API-only cost and workflow profiles | Must-have | No local architectural throughput claims. |
|
|
346
|
-
| OpenAI embeddings | API-cost-only embeddings records | Must-have | Use current official pricing and model docs. [OpenAI pricing](https://openai.com/api/pricing/) |
|
|
347
|
-
| Whisper / Parakeet / Distil-Whisper | Separate ASR calculator | Nice | Use audio duration, realtime factor, batch size, and encoder/decoder architecture rather than LLM token/sec. |
|
|
348
|
-
|
|
349
|
-
## Architecture-data pipeline
|
|
350
|
-
|
|
351
|
-
Do not hand-fill every requested tuple of parameters, layers, hidden size, heads, KV heads, head dimension, vocabulary, and context. Build an ingestion script that:
|
|
352
|
-
|
|
353
|
-
1. Fetches a Hugging Face model’s `config.json` and model card.
|
|
354
|
-
2. Maps common field aliases such as `num_hidden_layers`, `hidden_size`, `num_attention_heads`, `num_key_value_heads`, `head_dim`, `vocab_size`, and `max_position_embeddings`.
|
|
355
|
-
3. Saves raw config and normalized fields.
|
|
356
|
-
4. Flags missing data for manual review.
|
|
357
|
-
5. Pins revision SHA and retrieval date.
|
|
358
|
-
6. Displays a source link and confidence badge in the UI.
|
|
359
|
-
|
|
360
|
-
That makes your dataset maintainable and avoids stale copied values.
|
|
361
|
-
|
|
362
|
-
# SECTION C — MORE GPU OPTIONS
|
|
363
|
-
|
|
364
|
-
## Accelerator-data strategy
|
|
365
|
-
|
|
366
|
-
The requested fields cannot be uniformly compared:
|
|
367
|
-
|
|
368
|
-
- NVIDIA reports multiple tensor-performance modes: dense/sparse, FP16/BF16/FP8/FP4, sometimes per GPU and sometimes system-level.
|
|
369
|
-
- Google TPUs, Groq LPUs, Cerebras systems, and SambaNova systems are not drop-in “GPU” equivalents. They need separate accelerator schemas and benchmark-based serving records.
|
|
370
|
-
- Apple unified memory is system memory, not HBM VRAM, and practical inference bandwidth/thermal behavior differs.
|
|
371
|
-
- GB200 NVL72 and GB300 are **rack-scale systems**, not individual GPUs; represent them as topology products with nodes, accelerator count, fabric, and system price.
|
|
372
|
-
|
|
373
|
-
Use two tables:
|
|
374
|
-
|
|
375
|
-
1. `accelerator_specs`: immutable vendor specifications.
|
|
376
|
-
2. `rental_price_observations`: `provider`, `region`, `product`, `billing`, `price`, `timestamp`, `availability`, `source`.
|
|
377
|
-
|
|
378
|
-
## Must-add accelerator rows
|
|
379
|
-
|
|
380
|
-
| Product | Category | What tokcalc should model | Primary source |
|
|
381
|
-
|---|---|---|---|
|
|
382
|
-
| NVIDIA H200 | Datacenter GPU | HBM3e capacity/bandwidth, SXM vs PCIe form factor, NVLink topology | [NVIDIA H200](https://www.nvidia.com/en-us/data-center/h200/) |
|
|
383
|
-
| NVIDIA B200 | Datacenter GPU | FP4/FP8/FP16 modes, HBM3e, GB200/NVL system topology | [NVIDIA Blackwell](https://www.nvidia.com/en-us/data-center/blackwell-architecture/) |
|
|
384
|
-
| NVIDIA B300 / GB300 | New rack-scale / accelerator generation | Add only when official spec sheet and cloud offerings expose stable records | [NVIDIA data center products](https://www.nvidia.com/en-us/data-center/) |
|
|
385
|
-
| NVIDIA RTX PRO 6000 Blackwell | Workstation GPU | VRAM, bandwidth, consumer/workstation inference economics | [NVIDIA RTX PRO](https://www.nvidia.com/en-us/design-visualization/rtx-pro/) |
|
|
386
|
-
| RTX 5090 / 5080 | Consumer GPU | Local inference, PCIe constraints, power, VRAM | [NVIDIA GeForce RTX 50 series](https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/) |
|
|
387
|
-
| AMD MI300X | Datacenter GPU | HBM capacity/bandwidth, ROCm engine support, multi-GPU fabric | [AMD Instinct MI300X](https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html) |
|
|
388
|
-
| AMD MI325X / MI350 / MI355 | Datacenter GPU | Add released parts from official product datasheets; distinguish announced from available | [AMD Instinct](https://www.amd.com/en/products/accelerators/instinct.html) |
|
|
389
|
-
| Intel Gaudi 2 / Gaudi 3 | AI accelerator | Separate engine support and Ethernet scale-out assumptions | [Intel Gaudi](https://www.intel.com/content/www/us/en/products/details/processors/ai-accelerators/gaudi.html) |
|
|
390
|
-
| Google TPU v5e / v5p / Trillium | TPU | Pod/topology price and benchmark profile, not generic TFLOPS equivalence | [Google Cloud TPU](https://cloud.google.com/tpu/docs) |
|
|
391
|
-
| Groq LPU | Hosted inference system | Published model-specific tokens/sec and token API pricing | [Groq pricing](https://groq.com/pricing/) |
|
|
392
|
-
| Cerebras CS-3 / WSE-3 | Wafer-scale system | Model-specific latency/tokens/sec and API pricing where exposed | [Cerebras inference](https://www.cerebras.ai/inference) |
|
|
393
|
-
| SambaNova RDU | Dataflow system | Model-specific throughput rather than GPU FLOPS equivalence | [SambaNova](https://sambanova.ai/) |
|
|
394
|
-
| Apple M2 Ultra / M3 Ultra | Apple Silicon | Unified memory size/bandwidth, CPU/GPU shared-memory constraints, Metal backend profile | [Apple M2 Ultra](https://www.apple.com/newsroom/2023/06/apple-unveils-m2-ultra/) [Apple M3 Ultra](https://www.apple.com/newsroom/) |
|
|
395
|
-
| NVIDIA V100/P100/K80 | Legacy | Legacy comparison only, clearly marked as older architecture | [NVIDIA Tesla archive](https://www.nvidia.com/en-us/data-center/tesla/) |
|
|
396
|
-
|
|
397
|
-
## GPU specification fields
|
|
398
|
-
|
|
399
|
-
| Field | Why it belongs in tokcalc |
|
|
400
|
-
|---|---|
|
|
401
|
-
| `memoryBytes` | Determines weight and KV-cache fit |
|
|
402
|
-
| `memoryBandwidthGBps` | Strong predictor of decode-limited throughput |
|
|
403
|
-
| `bf16TfopsDense`, `fp16TfopsDense` | Helps model prefill / compute-bound work |
|
|
404
|
-
| `fp8Tfops`, `fp4Tfops` | Needed for supported low-precision kernels |
|
|
405
|
-
| `interconnectType`, `interconnectGBps` | Determines practical tensor-parallel and PD-disaggregation behavior |
|
|
406
|
-
| `pcieGeneration`, `pcieLanes` | Relevant for multi-GPU systems without NVLink |
|
|
407
|
-
| `tdpWatts` | Enables energy estimates |
|
|
408
|
-
| `formFactor` | PCIe, SXM, OAM, rack system, Apple SoC, TPU pod |
|
|
409
|
-
| `engineSupport` | vLLM, TensorRT-LLM, SGLang, llama.cpp, ROCm, Metal, Habana, TPU runtime |
|
|
410
|
-
| `priceObservation` | Separate, time-stamped, provider/region-specific object |
|
|
411
|
-
|
|
412
|
-
## Avoid a misleading universal GPU score
|
|
413
|
-
|
|
414
|
-
Decode usually tracks memory bandwidth and KV-cache access much more strongly than headline FLOPS; prefill may benefit more from compute throughput. Therefore expose separate scores:
|
|
415
|
-
|
|
416
|
-
- **Weight-streaming decode roofline**
|
|
417
|
-
- **Compute-bound prefill roofline**
|
|
418
|
-
- **KV-cache capacity / concurrency**
|
|
419
|
-
- **Multi-GPU communication penalty**
|
|
420
|
-
|
|
421
|
-
NVIDIA’s official data-center product pages are the appropriate primary source for current architecture specifications, but product marketing peak FLOPS should not be treated as achieved LLM tokens/sec. [NVIDIA data center](https://www.nvidia.com/en-us/data-center/)
|
|
422
|
-
|
|
423
|
-
# SECTION D — MORE QUANTIZATION FORMATS
|
|
424
|
-
|
|
425
|
-
## Quantization implementation rule
|
|
426
|
-
|
|
427
|
-
Do not show a single universal “accuracy loss” or “speedup” for a quantization format. It varies materially by:
|
|
428
|
-
|
|
429
|
-
- model family and scale;
|
|
430
|
-
- calibration dataset and quantization recipe;
|
|
431
|
-
- per-channel/per-group configuration;
|
|
432
|
-
- activation quantization;
|
|
433
|
-
- kernel and engine;
|
|
434
|
-
- prompt length, batch/concurrency, and GPU;
|
|
435
|
-
- benchmark metric.
|
|
436
|
-
|
|
437
|
-
Use **ranges plus source/model-specific benchmark records**, and always identify whether the number measures perplexity, MMLU, code, instruction-following, or a task metric.
|
|
438
|
-
|
|
439
|
-
## Format roadmap
|
|
440
|
-
|
|
441
|
-
| Format | What it is | Typical use | Engine support | Complexity | Priority |
|
|
442
|
-
|---|---|---|---|---:|---:|
|
|
443
|
-
| GGUF | `llama.cpp` model file format supporting many quantization schemes | CPU, Apple Silicon, consumer GPU local inference | llama.cpp and compatible runners | Medium | Must-have |
|
|
444
|
-
| GPTQ | Post-training weight quantization, often 4-bit | NVIDIA GPU local serving and older quantized model ecosystem | vLLM/TGI/TensorRT-LLM support varies by version; strong ecosystem presence | Medium | Must-have |
|
|
445
|
-
| AWQ | Activation-aware weight quantization, commonly W4A16 | GPU inference with preservation of salient weights | vLLM, TensorRT-LLM and others depending on release | Medium | Must-have |
|
|
446
|
-
| EXL2 | ExLlamaV2 mixed-bit quantization format | Consumer NVIDIA GPU inference | ExLlamaV2 primarily | Medium | Nice |
|
|
447
|
-
| SmoothQuant | Activation smoothing technique enabling weight-and-activation quantization | INT8-oriented serving | TensorRT-LLM / framework-dependent workflows | Medium | Nice |
|
|
448
|
-
| FP8 | 8-bit floating-point inference/training formats | H100/H200-class supported kernels | TensorRT-LLM, vLLM/backend-dependent | Medium | Must-have |
|
|
449
|
-
| NVFP4 | NVIDIA Blackwell FP4 inference format | Blackwell inference | NVIDIA Blackwell/TensorRT stack | Medium | Nice |
|
|
450
|
-
| MX / OCP microscaling | Block-scaled low-precision formats | Emerging hardware/software ecosystem | Hardware/software dependent | Future | Future |
|
|
451
|
-
| HQQ, EETQ, BitBLAS, QuaRot, QuIP#, SpinQuant | Research/tool-specific quantization methods | Experimental or specialized deployments | Uneven and volatile | High | Future |
|
|
452
|
-
| BitNet b1.58 | Ternary-native architecture/weights rather than a drop-in quantization of arbitrary models | Research and specialized inference | Limited general production support | High | Future |
|
|
453
|
-
|
|
454
|
-
Sources: [llama.cpp](https://github.com/ggerganov/llama.cpp) [GPTQ paper](https://arxiv.org/abs/2210.17323) [AWQ paper](https://arxiv.org/abs/2306.00978) [SmoothQuant paper](https://arxiv.org/abs/2211.10438) [NVFP4](https://developer.nvidia.com/blog/nvfp4-trains-with-precision-and-powers-efficient-inference/) [BitNet b1.58](https://arxiv.org/abs/2402.17764)
|
|
455
|
-
|
|
456
|
-
## GGUF variants
|
|
457
|
-
|
|
458
|
-
For GGUF, store the **exact file’s measured bytes/parameter** and benchmark metrics. Labels such as `Q4_K_M` identify quantization schemes but do not guarantee identical size/accuracy across model conversion versions and tensor mixes.
|
|
459
|
-
|
|
460
|
-
| GGUF family | Approximate nominal weight precision | Tokcalc UI treatment | Priority |
|
|
461
|
-
|---|---:|---|---:|
|
|
462
|
-
| Q2_K | ~2-bit class | Extreme-memory mode; warn strongly on quality loss | Nice |
|
|
463
|
-
| Q3_K_S / Q3_K_M | ~3-bit class | Low-memory mode; show quality-risk badge | Nice |
|
|
464
|
-
| Q4_0 / Q4_1 | ~4-bit legacy/simple schemes | Legacy compatibility choices | Nice |
|
|
465
|
-
| Q4_K_S / Q4_K_M | ~4-bit K-quant schemes | Recommended practical local-inference presets | Must-have |
|
|
466
|
-
| Q5_0 / Q5_1 | ~5-bit legacy/simple schemes | Mid-quality/local performance options | Nice |
|
|
467
|
-
| Q5_K_S / Q5_K_M | ~5-bit K-quant schemes | Higher-quality local presets | Must-have |
|
|
468
|
-
| Q6_K | ~6-bit class | Near-FP16-quality preference where memory allows | Nice |
|
|
469
|
-
| Q8_0 | ~8-bit class | High-quality compressed storage/reference point | Must-have |
|
|
470
|
-
|
|
471
|
-
The `llama.cpp` project and its associated conversion tooling are the primary practical source for supported GGUF formats; exact quality and size need to be benchmarked for the specific model artifact. [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
|
472
|
-
|
|
473
|
-
## Quantization calculator fields
|
|
474
|
-
|
|
475
|
-
```ts
|
|
476
|
-
type QuantizationProfile = {
|
|
477
|
-
id: string;
|
|
478
|
-
format: "fp16" | "bf16" | "int8" | "gptq" | "awq" | "gguf" | "exl2" | "fp8" | "nvfp4";
|
|
479
|
-
nominalWeightBits?: number;
|
|
480
|
-
activationBits?: number;
|
|
481
|
-
kvCacheBits?: number;
|
|
482
|
-
bytesPerParameterMeasured?: number;
|
|
483
|
-
groupSize?: number;
|
|
484
|
-
engineSupport: Record<string, "native" | "plugin" | "convert" | "unsupported">;
|
|
485
|
-
dequantKernel: "fused" | "unfused" | "native-low-precision";
|
|
486
|
-
qualityEvidence: BenchmarkRecord[];
|
|
487
|
-
throughputEvidence: BenchmarkRecord[];
|
|
488
|
-
};
|
|
489
|
-
```
|
|
490
|
-
|
|
491
|
-
## How to present speed and quality
|
|
492
|
-
|
|
493
|
-
Instead of this:
|
|
494
|
-
|
|
495
|
-
> “AWQ is 50% faster and loses 1% accuracy.”
|
|
496
|
-
|
|
497
|
-
Show this:
|
|
498
|
-
|
|
499
|
-
> “For **model X**, engine **Y**, GPU **Z**, workload **prompt/output/batch**, AWQ achieved **N output tok/s** versus FP16 **M output tok/s**. Quality score changed from **A** to **B** on benchmark **C**.”
|
|
500
|
-
|
|
501
|
-
This protects tokcalc’s credibility.
|
|
502
|
-
|
|
503
|
-
# SECTION E — CLOUD GPU PRICING (2025)
|
|
504
|
-
|
|
505
|
-
## Pricing-data architecture
|
|
506
|
-
|
|
507
|
-
Cloud pricing cannot responsibly be represented as one static 2025 value per accelerator/provider. Your pricing module should store price observations:
|
|
508
|
-
|
|
509
|
-
```ts
|
|
510
|
-
type PriceObservation = {
|
|
511
|
-
provider: string;
|
|
512
|
-
product: string;
|
|
513
|
-
accelerator: string;
|
|
514
|
-
acceleratorCount: number;
|
|
515
|
-
vramGB?: number;
|
|
516
|
-
region?: string;
|
|
517
|
-
billing: "on_demand" | "spot" | "reserved" | "serverless" | "per_second";
|
|
518
|
-
priceUSDPerHour: number;
|
|
519
|
-
minimumBillingSeconds?: number;
|
|
520
|
-
collectedAt: string;
|
|
521
|
-
sourceUrl: string;
|
|
522
|
-
availability?: "available" | "waitlist" | "not-listed";
|
|
523
|
-
notes?: string;
|
|
524
|
-
};
|
|
525
|
-
```
|
|
526
|
-
|
|
527
|
-
## Provider source registry
|
|
528
|
-
|
|
529
|
-
| Provider | Correct source category | Integration recommendation |
|
|
530
|
-
|---|---|---|
|
|
531
|
-
| RunPod | GPU Cloud pricing / API, serverless pricing | Separate secure/on-demand pods from serverless endpoint pricing. [RunPod pricing](https://www.runpod.io/gpu-instance/pricing) |
|
|
532
|
-
| Lambda | Cloud GPU pricing | Capture instance configuration and GPU count, not just a GPU label. [Lambda Cloud pricing](https://lambda.ai/service/gpu-cloud/pricing) |
|
|
533
|
-
| Modal | Serverless compute pricing | Model GPU price per second plus container/minimum/warm capacity behavior. [Modal pricing](https://modal.com/pricing) |
|
|
534
|
-
| Vast.ai | Marketplace / spot-like listings | Treat as live market observations, with availability and reliability uncertainty. [Vast.ai pricing](https://vast.ai/pricing) |
|
|
535
|
-
| Together AI | GPU cluster / service pricing | Distinguish hosted inference from dedicated GPU/rental offerings. [Together pricing](https://www.together.ai/pricing) |
|
|
536
|
-
| Replicate | Per-second prediction hardware pricing | Capture billed seconds and cold-start conditions. [Replicate pricing](https://replicate.com/pricing) |
|
|
537
|
-
| CoreWeave | Cloud/GPU pricing and enterprise quotation | Capture public rate only when publicly listed; otherwise mark quote-required. [CoreWeave](https://www.coreweave.com/) |
|
|
538
|
-
| TensorDock | GPU cloud pricing | Capture region/hosted node details. [TensorDock pricing](https://www.tensordock.com/pricing) |
|
|
539
|
-
| Hugging Face Inference Endpoints | Endpoint hardware pricing | Store endpoint hardware SKU and region. [Hugging Face endpoints pricing](https://huggingface.co/docs/inference-endpoints/pricing) |
|
|
540
|
-
| AWS EC2 | On-demand instance pricing | Map instance family to GPUs: p5, p5e, p4d, g6, g5; use region-specific price API or pricing pages. [AWS accelerated computing](https://aws.amazon.com/ec2/instance-types/accelerated-computing/) |
|
|
541
|
-
| Google Cloud | Accelerator/VM pricing | Map A3/A4 and TPU products separately. [Google Cloud GPU pricing](https://cloud.google.com/compute/gpus-pricing) |
|
|
542
|
-
| Azure | VM pricing | Map ND H100 v5 / H200 offerings and region. [Azure GPU VM pricing](https://azure.microsoft.com/pricing/details/virtual-machines/linux/) |
|
|
543
|
-
| Oracle Cloud | Compute GPU pricing | Map instance GPU configuration rather than isolated GPU. [Oracle GPU pricing](https://www.oracle.com/cloud/compute/gpu/pricing/) |
|
|
544
|
-
|
|
545
|
-
## UI behavior for unavailable pairs
|
|
546
|
-
|
|
547
|
-
The requested matrix includes many provider–accelerator combinations that will be blank because a provider does not offer that product. For example, a Mac M2 Ultra is normally purchased hardware rather than a mainstream cloud-GPU SKU; do not fabricate a cloud hourly rate. Render:
|
|
548
|
-
|
|
549
|
-
- **Listed price** — date, region, billing type, source
|
|
550
|
-
- **Not listed by provider** — no extrapolation
|
|
551
|
-
- **Marketplace range** — min / median / P90, timestamped
|
|
552
|
-
- **Quote required** — provider contact, no numerical claim
|
|
553
|
-
- **Deprecated** — historical reference only
|
|
554
|
-
|
|
555
|
-
## Recommended user-facing caveat
|
|
556
|
-
|
|
557
|
-
> Prices are observations, not guarantees. They vary by region, availability, spot/on-demand/reserved commitment, billing granularity, included CPU/RAM/storage/network, and date. Select a provider and region to see the source record.
|
|
558
|
-
|
|
559
|
-
# SECTION F — API PRICING FOR BUILD-VS-BUY COMPARISON
|
|
560
|
-
|
|
561
|
-
## Pricing integration design
|
|
562
|
-
|
|
563
|
-
API rates should be versioned just like hardware prices. Do not mix obsolete named models with current list prices without an “as of” date. Your application should source first-party price pages and retain historical snapshots.
|
|
564
|
-
|
|
565
|
-
| Provider | Primary pricing source | Caching fields to support |
|
|
566
|
-
|---|---|---|
|
|
567
|
-
| OpenAI | [OpenAI API pricing](https://openai.com/api/pricing/) | input, cached input, output, batch, reasoning/other applicable modes |
|
|
568
|
-
| Anthropic | [Anthropic pricing](https://www.anthropic.com/pricing) | base input, 5-minute cache write, 1-hour cache write, cache read, output |
|
|
569
|
-
| Google | [Gemini API pricing](https://ai.google.dev/pricing) | input/output, context tiers, cached-content pricing where offered |
|
|
570
|
-
| Together AI | [Together pricing](https://www.together.ai/pricing) | model-specific input/output rates and batch/endpoint distinctions |
|
|
571
|
-
| Groq | [Groq pricing](https://groq.com/pricing/) | model-specific input/output rates, hosted model availability |
|
|
572
|
-
| Fireworks AI | [Fireworks pricing](https://fireworks.ai/pricing) | model-specific serverless / on-demand distinctions |
|
|
573
|
-
| Cerebras | [Cerebras pricing](https://www.cerebras.ai/pricing) | model-specific API rate and service tier |
|
|
574
|
-
| DeepSeek | [DeepSeek pricing](https://api-docs.deepseek.com/quick_start/pricing) | cache miss/hit input rates, output rates, model/version |
|
|
575
|
-
| Mistral | [Mistral pricing](https://mistral.ai/pricing/) | model-specific input/output, batch, cached input if available |
|
|
576
|
-
|
|
577
|
-
## Prompt caching support
|
|
578
|
-
|
|
579
|
-
| Provider | Caching behavior | Tokcalc treatment |
|
|
580
|
-
|---|---|---|
|
|
581
|
-
| Anthropic | Explicit prompt caching with documented write/read rates and TTL choices | Full cache model: write multiplier, read multiplier, TTL, hit rate. [Anthropic prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) |
|
|
582
|
-
| OpenAI | Automatic prompt caching for eligible requests/models with matching prefixes; pricing includes cached-input categories where applicable | Model-specific cached input rate; no assumption that every prompt/model is eligible. [OpenAI prompt caching](https://platform.openai.com/docs/guides/prompt-caching) |
|
|
583
|
-
| Google | Supports cached content in supported Gemini workflows; pricing and eligibility depend on model/product | Use model-specific cached-content price records from Google’s official pricing docs. [Gemini API pricing](https://ai.google.dev/pricing) |
|
|
584
|
-
| DeepSeek | API pricing can distinguish cache-hit and cache-miss input handling | Store its published cache-hit/miss rates by model/version. [DeepSeek pricing](https://api-docs.deepseek.com/quick_start/pricing) |
|
|
585
|
-
| Other vendors | Varies by provider/model | Leave null until first-party documentation confirms it. |
|
|
586
|
-
|
|
587
|
-
## Build-vs-buy formula
|
|
588
|
-
|
|
589
|
-
For a workload with \(I\) input tokens and \(O\) output tokens:
|
|
590
|
-
|
|
591
|
-
\[
|
|
592
|
-
\text{API cost} =
|
|
593
|
-
\frac{I}{10^6}P_i +
|
|
594
|
-
\frac{I_c}{10^6}P_{ic} +
|
|
595
|
-
\frac{O}{10^6}P_o
|
|
596
|
-
\]
|
|
597
|
-
|
|
598
|
-
Where \(I_c\) is eligible cached input and \(P_{ic}\) is cached-input price.
|
|
599
|
-
|
|
600
|
-
For self-hosting:
|
|
601
|
-
|
|
602
|
-
\[
|
|
603
|
-
\text{self-host monthly cost} =
|
|
604
|
-
\left(\frac{\text{GPU price/hr} \cdot 730}{\text{utilization adjustment}}\right)
|
|
605
|
-
+ \text{storage} + \text{network} + \text{operational overhead}
|
|
606
|
-
\]
|
|
607
|
-
|
|
608
|
-
Keep operational overhead as an editable field rather than pretending it is zero.
|
|
609
|
-
|
|
610
|
-
# SECTION G — TOP 10 PRIORITIZED FEATURE EXPANSIONS
|
|
611
|
-
|
|
612
|
-
| Rank | Feature | Description | Why users want it | Complexity / estimate | Required data | Dependencies |
|
|
613
|
-
|---:|---|---|---|---|---|---|
|
|
614
|
-
| 1 | **Production serving mode** | Continuous batching, paged KV cache, concurrency, queueing, TTFT, and ITL estimates | Production throughput is dominated by scheduler behavior rather than single-request speed; vLLM introduced paged KV management and continuous batching for this exact reason. [vLLM paper](https://arxiv.org/abs/2309.06180) | **High, 50–90h** | Model architecture, GPU memory/bandwidth, engine profiles, workload distributions | Existing inference core |
|
|
615
|
-
| 2 | **Build-vs-buy calculator** | Compare self-hosted deployment with OpenAI/Anthropic/Google/hosted-model token pricing | This is the product decision most teams actually need: economic break-even under their prompt/output mix and utilization | **Medium, 24–45h** | GPU prices, API prices, model throughput, utilization settings | Pricing ingestion |
|
|
616
|
-
| 3 | **Prompt-cache calculator** | Compute self-hosted prefix-cache savings and provider API cache savings | Repeated system prompts, RAG context, and agent histories can yield large input-cost and TTFT reductions; Anthropic documents 0.1× cache-read pricing. [Anthropic prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) | **Medium, 16–30h** | Provider cache rules, workload/prompt profile | Build-vs-buy pricing |
|
|
617
|
-
| 4 | **Long-context and KV-cache planner** | Show fit, concurrency, latency, and cost through 1M-token contexts | KV-cache memory scales linearly with live context length and becomes the practical constraint in long-context serving | **High, 30–70h** | Layers, KV heads, head dim, KV precision, GPU VRAM | Architecture ingestion |
|
|
618
|
-
| 5 | **Live versioned data registry** | Automated model configs, accelerator specs, price observations, source URLs, dates, confidence | Every expansion request becomes sustainable only if data is refreshable, reviewable, and source-linked | **High, 45–80h** | Hugging Face configs, vendor specs, provider prices | None; start immediately |
|
|
619
|
-
| 6 | **Embedding calculator** | Estimate documents/sec, tokens/sec, indexing cost, vector storage, and API/self-host comparison | Embeddings are a distinct high-throughput encoder workload central to RAG, and do not fit decode-token math. [OpenAI embeddings guide](https://platform.openai.com/docs/guides/embeddings) | **Medium, 24–48h** | Embedding models, dimensions, encoder benchmarks, API prices | Data registry |
|
|
620
|
-
| 7 | **Reasoning and agent workflow planner** | Model hidden reasoning allowance, multi-call tools, context growth, retry rate, and end-to-end latency | Modern task completion often consists of several model calls plus tool latency, not a single chat completion | **Medium, 20–40h** | API rates, workload templates, cache model | Prompt cache, pricing |
|
|
621
|
-
| 8 | **Quantization explorer** | Compare GGUF/GPTQ/AWQ/EXL2/FP8/NVFP4 with measured size, quality, and engine support | Users need a defensible fit-quality-speed trade-off, especially on consumer hardware | **Medium–High, 35–65h** | Artifact sizes, model-specific quality benchmarks, engine/GPU throughput tests | Benchmark registry |
|
|
622
|
-
| 9 | **Vision workload mode** | Convert image/video inputs into model-specific visual-token/prefill estimates | Vision prompts can dominate effective context and cost; tokenization is model-specific. [Qwen2-VL](https://huggingface.co/Qwen/Qwen2-VL-72B-Instruct) | **High, 35–70h** | VLM configs, image-tokenization rules, vision benchmarks | Data registry |
|
|
623
|
-
| 10 | **Multi-LoRA capacity planner** | Estimate adapter VRAM, resident adapters, cache behavior, and throughput range | Multi-tenant customization is a major self-hosting use case; Punica/S-LoRA demonstrate the serving need. [Punica](https://arxiv.org/abs/2310.18547) [S-LoRA](https://arxiv.org/abs/2311.03285) | **Medium, 24–50h** | Adapter rank/modules, base model config, engine profiles | Production serving mode |
|
|
624
|
-
|
|
625
|
-
# SECTION H — KNOWN BENCHMARKS DATABASES
|
|
626
|
-
|
|
627
|
-
| Database / project | URL | What it benchmarks | Update pattern | API / downloadable data | Commercial-use notes |
|
|
628
|
-
|---|---|---|---|---|---|
|
|
629
|
-
| MLPerf Inference | [mlcommons.org/benchmarks/inference](https://mlcommons.org/benchmarks/inference/) | Standardized inference performance, latency, throughput, power in defined scenarios | Benchmark rounds / periodic releases | Public results downloads; review result-license terms | MLCommons materials have their own terms; do not assume unrestricted commercial redistribution. |
|
|
630
|
-
| Artificial Analysis | [artificialanalysis.ai](https://artificialanalysis.ai/) | Hosted-model quality, speed, latency, and pricing comparisons | Frequently updated | Website/API availability and terms must be checked before scraping | Treat as a licensed data source; do not scrape without permission. |
|
|
631
|
-
| LMSYS Chatbot Arena / FastChat | [lmarena.ai](https://lmarena.ai/) | Human preference / model quality; some ecosystem leaderboard data | Continuous / periodic snapshots | Public datasets may be released for specific artifacts | Check each dataset/card license; quality data is not throughput data. |
|
|
632
|
-
| Hugging Face Open LLM Leaderboard | [huggingface.co/open-llm-leaderboard](https://huggingface.co/open-llm-leaderboard) | Model-quality benchmarks | Continuously updated as submissions run | Hugging Face datasets/spaces; inspect individual licenses | License varies by dataset/model/result. |
|
|
633
|
-
| vLLM benchmarks | [vLLM benchmarks docs](https://docs.vllm.ai/en/latest/contributing/benchmarks/) | Serving throughput, latency, scheduler performance across configurations | Repository/version-driven | Scripts and results are available in project ecosystem | Open-source code license applies to code; validate result reuse. |
|
|
634
|
-
| TensorRT-LLM benchmarks | [TensorRT-LLM benchmarks](https://nvidia.github.io/TensorRT-LLM/performance/performance-tuning-guide/benchmarking-default.html) | NVIDIA inference performance by engine/configuration | Release-driven | Benchmark tooling/scripts; result availability varies | NVIDIA licensing/terms apply. |
|
|
635
|
-
| llmperf | [GitHub: ray-project/llmperf](https://github.com/ray-project/llmperf) | API serving latency, TTFT, throughput under load | Repository-driven | Open-source benchmark harness; generate your own CSV | Apache-2.0 repository license; inspect repository for current terms. |
|
|
636
|
-
| Hugging Face Optimum Benchmark | [GitHub: huggingface/optimum-benchmark](https://github.com/huggingface/optimum-benchmark) | Hardware/software benchmark harnesses for transformer inference | Repository-driven | Run locally and export your data | Open-source repository; verify its license/version. |
|
|
637
|
-
| Stanford HELM | [crfm.stanford.edu/helm](https://crfm.stanford.edu/helm/) | Model accuracy, robustness, fairness, efficiency-related scenarios | Release-driven | Public scenarios/results depending on release | Check repository/data terms before commercial use. |
|
|
638
|
-
| Papers with Code | [paperswithcode.com](https://paperswithcode.com/) | Research benchmark leaderboards, primarily quality | Community-maintained / variable | Web data/API policies vary | Do not scrape or redistribute without checking terms. |
|
|
639
|
-
|
|
640
|
-
## Benchmark ingestion recommendation
|
|
641
|
-
|
|
642
|
-
Use a three-tier evidence model:
|
|
643
|
-
|
|
644
|
-
1. **First-party standardized:** MLPerf, vendor benchmark submissions.
|
|
645
|
-
2. **Reproducible engine benchmark:** vLLM/TensorRT-LLM/llmperf command lines plus full configuration.
|
|
646
|
-
3. **Community observation:** clear provenance and lower-confidence label.
|
|
647
|
-
|
|
648
|
-
Every benchmark record should include:
|
|
649
|
-
|
|
650
|
-
```ts
|
|
651
|
-
type BenchmarkRecord = {
|
|
652
|
-
modelId: string;
|
|
653
|
-
modelRevision?: string;
|
|
654
|
-
quantizationId: string;
|
|
655
|
-
engine: string;
|
|
656
|
-
engineVersion?: string;
|
|
657
|
-
gpu: string;
|
|
658
|
-
gpuCount: number;
|
|
659
|
-
topology?: string;
|
|
660
|
-
promptTokens: number;
|
|
661
|
-
outputTokens: number;
|
|
662
|
-
concurrency: number;
|
|
663
|
-
requestRate?: number;
|
|
664
|
-
throughputOutputTokPerSec?: number;
|
|
665
|
-
throughputTotalTokPerSec?: number;
|
|
666
|
-
ttftMsP50?: number;
|
|
667
|
-
ttftMsP95?: number;
|
|
668
|
-
itlMsP50?: number;
|
|
669
|
-
itlMsP95?: number;
|
|
670
|
-
sourceUrl: string;
|
|
671
|
-
collectedAt: string;
|
|
672
|
-
evidenceTier: 1 | 2 | 3;
|
|
673
|
-
};
|
|
674
|
-
```
|
|
675
|
-
|
|
676
|
-
# TL;DR for the maintainer
|
|
677
|
-
|
|
678
|
-
1. **Build the data registry first.** Version model configs, GPU specs, benchmark records, API pricing, and rental-price observations with source URL, retrieval date, region, engine version, and confidence. This makes every later feature maintainable.
|
|
679
|
-
|
|
680
|
-
2. **Ship a production serving calculator next.** Add continuous batching, paged KV cache, concurrency, TTFT, inter-token latency, and queueing; this is the biggest credibility upgrade from a static throughput estimate. [vLLM paper](https://arxiv.org/abs/2309.06180)
|
|
681
|
-
|
|
682
|
-
3. **Add build-vs-buy plus prompt caching as one feature set.** Users can immediately compare self-hosting to APIs, model shared-prefix hit rates, and see the economics of their real workload. Anthropic explicitly documents 90%-discount cache reads relative to base input pricing. [Anthropic prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching)
|
|
683
|
-
|
|
684
|
-
4. **Make long-context/KV planning first-class.** Show KV memory, maximum concurrency, and cost at 8K, 32K, 128K, and beyond; this becomes essential as teams deploy long-context RAG and agent workflows.
|
|
685
|
-
|
|
686
|
-
5. **Expand the catalog strategically, not exhaustively.** Add Qwen3/Qwen-VL, DeepSeek-V3/R1, Llama 4 released variants, Mistral/Pixtral, Gemma 3, Phi-4, embedding models, and current hardware—but pull values from official model configs and vendor specs instead of maintaining manually copied architecture tables.
|