ComputeAssay

Model guide · prices from 2026-09-10 · API prices daily, full survey weekly

Cheapest way to run Qwen2.5 72B

Qwen2.5 72B needs ≈43 GB of VRAM at 4-bit — this week that starts at $0.75/hour (2× RTX4090, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.

FP16 — ≈173 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB6$2.41/hr
A100-SXM80 GB3$2.72/hr$6.90/hr
RTX409024 GB8$2.99/hr
MI300X192 GB1$3.45/hr
A100-PCIe80 GB3$4.05/hr
L40S48 GB4$5.48/hr
H100-PCIe80 GB3$5.97/hr
B200192 GB1$6.00/hr$7.15/hr

FP8/INT8 — ≈86 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB3$1.20/hr
RTX409024 GB4$1.50/hr
A100-SXM80 GB2$1.81/hr$4.60/hr
A100-PCIe80 GB2$2.70/hr
L40S48 GB2$2.74/hr
MI300X192 GB1$3.45/hr
L424 GB4$3.48/hr
H100-PCIe80 GB2$3.98/hr

4-bit — ≈43 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX409024 GB2$0.75/hr
RTX509032 GB2$0.80/hr
A100-SXM80 GB1$0.91/hr$2.30/hr
A100-PCIe80 GB1$1.35/hr
L40S48 GB1$1.37/hr
L424 GB2$1.74/hr
H100-PCIe80 GB1$1.99/hr
A10080 GB1$2.70/hr

Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.

These prices move weekly

The Delivered Compute Report tracks them — free, every row sourced.

Subscribe free