ComputeAssay

Model guide · prices from 2026-09-10 · API prices daily, full survey weekly

Cheapest way to run Llama 4 Maverick (400B MoE)

Llama 4 Maverick (400B MoE) needs ≈240 GB of VRAM at 4-bit — this week that starts at $2.72/hour (3× A100-SXM, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.

MoE note: Llama 4 Maverick (400B MoE) activates ~17B active per token, but the full 400B weights must fit in VRAM — memory scales with total parameters, speed with active ones.

FP16 — ≈960 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB12 (multi-node)$10.88/hr$27.60/hr
RTX509032 GB30 (multi-node)$12.03/hr
RTX409024 GB40 (multi-node)$14.96/hr
A100-PCIe80 GB12 (multi-node)$16.20/hr
MI300X192 GB5$17.25/hr
H100-PCIe80 GB12 (multi-node)$23.88/hr
L40S48 GB20 (multi-node)$27.40/hr
H200-SXM141 GB7$27.93/hr$30.03/hr

FP8/INT8 — ≈480 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB6$5.44/hr$13.80/hr
RTX509032 GB15 (multi-node)$6.02/hr
RTX409024 GB20 (multi-node)$7.48/hr
A100-PCIe80 GB6$8.10/hr
MI300X192 GB3$10.35/hr
H100-PCIe80 GB6$11.94/hr
L40S48 GB10 (multi-node)$13.70/hr
B300288 GB2$15.00/hr$35.60/hr

4-bit — ≈240 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB3$2.72/hr$6.90/hr
RTX509032 GB8$3.21/hr
RTX409024 GB10 (multi-node)$3.74/hr
A100-PCIe80 GB3$4.05/hr
H100-PCIe80 GB3$5.97/hr
L40S48 GB5$6.85/hr
MI300X192 GB2$6.90/hr
B300288 GB1$7.50/hr$17.80/hr

Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.

These prices move weekly

The Delivered Compute Report tracks them — free, every row sourced.

Subscribe free