ComputeAssay

The bridge · prices from 2026-09-10 · throughput as dated per row

What a million tokens actually costs

Developers buy tokens; the market sells GPU-hours; nothing converts between them honestly. This page does: this week's cheapest tracked rental price × published, sourced inference throughput = $/million output tokens, per hardware config — with every assumption on the surface, because the assumptions are the story.

Read this box before the tables

Llama 70B class

ConfigOutput tok/s (aggregate)ModePrecisionStack (date)$/M tokens @ cheapest walk-up@ Grade A/BSrc
1xB20011,264ceilingFP4TensorRT-LLM (MLPerf Inference v4.1 offline, preview) (2024-08-28)$0.15$0.18
1xB200 (TP1)10,614ceilingFP4 (NVFP4)TensorRT-LLM v0.21 (2025)$0.16$0.19
B200 (per-GPU, InferenceMAX)10,000ceilingFP4 (NVFP4)TensorRT-LLM (2025-10-09)$0.17$0.20
8xH200 SXM (700W)34,864ceilingFP8TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28)$0.25$0.27
2xH100 SXM (TP2)6,092ceilingFP8TensorRT-LLM v0.21 (2025)$0.29$0.35
2xH200 SXM (TP2)7,467ceilingFP8TensorRT-LLM v0.21 (2025)$0.30$0.32
4xH100 SXM5 (NIM default TP4)7,000ceilingBF16NVIDIA NIM (TensorRT-LLM) (2025-05-18)$0.51$0.61
8xH100 SXM (per-GPU plateau, InferenceMAX)900ceilingFP8vLLM / TensorRT-LLM (InferenceMAX harness) (2025-10-09)$0.99$1.19
4xA100 80GB PCIe (NIM TP4)570ceilingBF16NVIDIA NIM (TensorRT-LLM) (2025-05-18)$1.77$4.48

The bar to beat: serverless APIs sell Llama 70B class output at $1.04/M (Together AI (serverless)) · $0.32/M (DeepInfra) · $0.9/M (Fireworks AI (serverless, catch-all tier)). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.

Mixtral 8x7B

ConfigOutput tok/s (aggregate)ModePrecisionStack (date)$/M tokens @ cheapest walk-up@ Grade A/BSrc
8xH100 SXM52,416ceilingFP8TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28)$0.14$0.16
8xH200 SXM59,022ceilingFP8TensorRT-LLM (MLPerf Inference v4.1 offline) (2024-08-28)$0.15$0.16

Llama 8B class

ConfigOutput tok/s (aggregate)ModePrecisionStack (date)$/M tokens @ cheapest walk-up@ Grade A/BSrc
1xH100 SXM26,401ceilingFP8TensorRT-LLM v0.21 (2025)$0.03$0.04
1xH200 SXM27,028ceilingFP8TensorRT-LLM v0.21 (2025)$0.04$0.04
1xA100 80GB2,400ceilingFP16vLLM (2024-06-05)$0.10$0.27
1xL40S325ceilingunspecified (FP16 assumed)unspecified (vLLM implied) (2026-09-06)$1.17

The bar to beat: serverless APIs sell Llama 8B class output at $0.14/M (Together AI (serverless)) · $0.04/M (DeepInfra). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.

DeepSeek-R1 / V3 (671B MoE)

ConfigOutput tok/s (aggregate)ModePrecisionStack (date)$/M tokens @ cheapest walk-up@ Grade A/BSrc
1xRTX 4090 24GB2,769ceilingFP16vLLM (2026-08-25)$0.04
B200 (per-chip, InferenceX)4,852ceilingFP8InferenceX harness (TRT-LLM/vLLM/SGLang best-of) (2026 (live dashboard, continuously updated))$0.34$0.41
8xMI300X21,225ceilingFP8Moreh MoAI (expert parallelism; SGLang-compatible serving) (2025-11-13)$0.36
8xB200 (DGX B200)30,000ceilingFP4 (NVFP4)TensorRT-LLM v0.17 (early Blackwell stack) (2025-03-18)$0.44$0.53
8xH200 SXM (TP8 + DP attention)19,024ceilingFP8 (native)SGLang v0.5.3 (2025-2026 (GPUStack 2.0 performance lab, undated))$0.47$0.50
H200 multi-node wide-EP (CoreWeave, InfiniBand CX-7)2,200ceilingFP8 (native)vLLM v0.11 (Wide-EP, DBO, EPLB, P/D disaggregation) (2025-12-17)$0.50$0.54
H200 (per-chip, InferenceX)1,591ceilingFP8InferenceX harness (TRT-LLM/vLLM/SGLang best-of) (2026 (live dashboard, continuously updated))$0.70$0.75
8xMI300X4,574ceilingFP8vLLM (ROCm, early 2025) (2025-03-18)$1.68

The bar to beat: serverless APIs sell DeepSeek-R1 / V3 (671B MoE) output at $2.15/M (DeepInfra) · $0.89/M (DeepInfra). If your utilization can't push rented $/M below that, rent the tokens, not the GPUs.

Why this page is hard to build honestly

Every $/M-token number you see quoted anywhere silently fixes five variables: utilization, serving mode, software version, precision, and context length. Change any one and the number moves by multiples — which is the deepest version of this site's thesis: compute has no kWh. A GPU-hour is not a unit, and neither, yet, is a token. The frame essay is the long version; the registry is the evidence.

Prices move weekly; stacks move monthly

The report tracks both — free, every row sourced.

Subscribe free