The bridge · prices from 2026-09-10 · throughput as dated per row
What a million tokens actually costs
Developers buy tokens; the market sells GPU-hours; nothing converts between them honestly. This page does: this week's cheapest tracked rental price × published, sourced inference throughput = $/million output tokens, per hardware config — with every assumption on the surface, because the assumptions are the story.
Read this box before the tables
- 100% utilization assumed. At 30% real utilization, multiply every $/M by ~3.3. This single assumption dominates everything below — and is exactly why serverless APIs win at low utilization.
- "Ceiling" vs "realistic" differ up to ~8× on identical hardware (short-context max-batch vs latency-constrained serving). Compare rows only within a mode.
- Software moves faster than silicon: the same 8×MI300X went 4,574 → 21,224 tok/s in 8 months from serving-stack tuning alone. Every row is stack-and-date-stamped; stale rows overstate cost.
- Precision is not quality-neutral: B200 headline numbers are FP4 — faster tokens are not identical tokens. This is why ASSAY-1 §4.3 reserves a quality dimension: a token is not yet a unit.
- Output tokens only; prices are this week's cheapest tracked walk-up (teasers excluded) and Grade A/B rates from the registry.
Llama 70B class
| Config | Output tok/s (aggregate) | Mode | Precision | Stack (date) | $/M tokens @ cheapest walk-up | @ Grade A/B | Src |
|---|---|---|---|---|---|---|---|
| 1xB200 | 11,264 | ceiling | FP4 | TensorRT-LLM (MLPerf Inference v4.1 offline, preview) (2024-08-28) | $0.15 | $0.18 | ↗ |
| 1xB200 (TP1) | 10,614 | ceiling | FP4 (NVFP4) | TensorRT-LLM v0.21 (2025) | $0.16 | $0.19 | ↗ |
| B200 (per-GPU, InferenceMAX) | 10,000 | ceiling | FP4 (NVFP4) | TensorRT-LLM (2025-10-09) | $0.17 | $0.20 | ↗ |
| 8xH200 SXM (700W) | 34,864 | ceiling | FP8 | TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28) | $0.25 | $0.27 | ↗ |
| 2xH100 SXM (TP2) | 6,092 | ceiling | FP8 | TensorRT-LLM v0.21 (2025) | $0.29 | $0.35 | ↗ |
| 2xH200 SXM (TP2) | 7,467 | ceiling | FP8 | TensorRT-LLM v0.21 (2025) | $0.30 | $0.32 | ↗ |
| 4xH100 SXM5 (NIM default TP4) | 7,000 | ceiling | BF16 | NVIDIA NIM (TensorRT-LLM) (2025-05-18) | $0.51 | $0.61 | ↗ |
| 8xH100 SXM (per-GPU plateau, InferenceMAX) | 900 | ceiling | FP8 | vLLM / TensorRT-LLM (InferenceMAX harness) (2025-10-09) | $0.99 | $1.19 | ↗ |
| 4xA100 80GB PCIe (NIM TP4) | 570 | ceiling | BF16 | NVIDIA NIM (TensorRT-LLM) (2025-05-18) | $1.77 | $4.48 | ↗ |
The bar to beat: serverless APIs sell Llama 70B class output at $1.04/M (Together AI (serverless)) · $0.32/M (DeepInfra) · $0.9/M (Fireworks AI (serverless, catch-all tier)). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.
Mixtral 8x7B
| Config | Output tok/s (aggregate) | Mode | Precision | Stack (date) | $/M tokens @ cheapest walk-up | @ Grade A/B | Src |
|---|---|---|---|---|---|---|---|
| 8xH100 SXM | 52,416 | ceiling | FP8 | TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28) | $0.14 | $0.16 | ↗ |
| 8xH200 SXM | 59,022 | ceiling | FP8 | TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28) | $0.15 | $0.16 | ↗ |
Llama 8B class
| Config | Output tok/s (aggregate) | Mode | Precision | Stack (date) | $/M tokens @ cheapest walk-up | @ Grade A/B | Src |
|---|---|---|---|---|---|---|---|
| 1xH100 SXM | 26,401 | ceiling | FP8 | TensorRT-LLM v0.21 (2025) | $0.03 | $0.04 | ↗ |
| 1xH200 SXM | 27,028 | ceiling | FP8 | TensorRT-LLM v0.21 (2025) | $0.04 | $0.04 | ↗ |
| 1xA100 80GB | 2,400 | ceiling | FP16 | vLLM (2024-06-05) | $0.10 | $0.27 | ↗ |
| 1xL40S | 325 | ceiling | unspecified (FP16 assumed) | unspecified (vLLM implied) (2026-09-06) | $1.17 | — | ↗ |
The bar to beat: serverless APIs sell Llama 8B class output at $0.14/M (Together AI (serverless)) · $0.04/M (DeepInfra). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.
DeepSeek-R1 / V3 (671B MoE)
| Config | Output tok/s (aggregate) | Mode | Precision | Stack (date) | $/M tokens @ cheapest walk-up | @ Grade A/B | Src |
|---|---|---|---|---|---|---|---|
| 1xRTX 4090 24GB | 2,769 | ceiling | FP16 | vLLM (2026-08-25) | $0.04 | — | ↗ |
| B200 (per-chip, InferenceX) | 4,852 | ceiling | FP8 | InferenceX harness (TRT-LLM/vLLM/SGLang best-of) (2026 (live dashboard, continuously updated)) | $0.34 | $0.41 | ↗ |
| 8xMI300X | 21,225 | ceiling | FP8 | Moreh MoAI (expert parallelism; SGLang-compatible serving) (2025-11-13) | $0.36 | — | ↗ |
| 8xB200 (DGX B200) | 30,000 | ceiling | FP4 (NVFP4) | TensorRT-LLM v0.17 (early Blackwell stack) (2025-03-18) | $0.44 | $0.53 | ↗ |
| 8xH200 SXM (TP8 + DP attention) | 19,024 | ceiling | FP8 (native) | SGLang v0.5.3 (2025-2026 (GPUStack 2.0 performance lab, undated)) | $0.47 | $0.50 | ↗ |
| H200 multi-node wide-EP (CoreWeave, InfiniBand CX-7) | 2,200 | ceiling | FP8 (native) | vLLM v0.11 (Wide-EP, DBO, EPLB, P/D disaggregation) (2025-12-17) | $0.50 | $0.54 | ↗ |
| H200 (per-chip, InferenceX) | 1,591 | ceiling | FP8 | InferenceX harness (TRT-LLM/vLLM/SGLang best-of) (2026 (live dashboard, continuously updated)) | $0.70 | $0.75 | ↗ |
| 8xMI300X | 4,574 | ceiling | FP8 | vLLM (ROCm, early 2025) (2025-03-18) | $1.68 | — | ↗ |
The bar to beat: serverless APIs sell DeepSeek-R1 / V3 (671B MoE) output at $2.15/M (DeepInfra) · $0.89/M (DeepInfra). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.
Why this page is hard to build honestly
Every $/M-token number you see quoted anywhere silently fixes five variables: utilization, serving mode, software version, precision, and context length. Change any one and the number moves by multiples — which is the deepest version of this site's thesis: compute has no kWh. A GPU-hour is not a unit, and neither, yet, is a token. The frame essay is the long version; the registry is the evidence.
Prices move weekly; stacks move monthly
The report tracks both — free, every row sourced.
Subscribe free