Model guide · prices from 2026-09-10 · API prices daily, full survey weekly
Cheapest way to run Qwen2.5 32B
Qwen2.5 32B needs ≈19 GB of VRAM at 4-bit — this week that starts at $0.37/hour (1× RTX4090, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.
FP16 — ≈77 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| A100-SXM | 80 GB | 1 | $0.91/hr | $2.30/hr |
| RTX5090 | 32 GB | 3 | $1.20/hr | — |
| A100-PCIe | 80 GB | 1 | $1.35/hr | — |
| RTX4090 | 24 GB | 4 | $1.50/hr | — |
| H100-PCIe | 80 GB | 1 | $1.99/hr | — |
| A100 | 80 GB | 1 | $2.70/hr | — |
| L40S | 48 GB | 2 | $2.74/hr | — |
| H100-SXM | 80 GB | 1 | $3.20/hr | $3.85/hr |
FP8/INT8 — ≈38 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| RTX4090 | 24 GB | 2 | $0.75/hr | — |
| RTX5090 | 32 GB | 2 | $0.80/hr | — |
| A100-SXM | 80 GB | 1 | $0.91/hr | $2.30/hr |
| A100-PCIe | 80 GB | 1 | $1.35/hr | — |
| L40S | 48 GB | 1 | $1.37/hr | — |
| L4 | 24 GB | 2 | $1.74/hr | — |
| H100-PCIe | 80 GB | 1 | $1.99/hr | — |
| A100 | 80 GB | 1 | $2.70/hr | — |
4-bit — ≈19 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| RTX4090 | 24 GB | 1 | $0.37/hr | — |
| RTX5090 | 32 GB | 1 | $0.40/hr | — |
| L4 | 24 GB | 1 | $0.87/hr | — |
| A100-SXM | 80 GB | 1 | $0.91/hr | $2.30/hr |
| A100-PCIe | 80 GB | 1 | $1.35/hr | — |
| L40S | 48 GB | 1 | $1.37/hr | — |
| H100-PCIe | 80 GB | 1 | $1.99/hr | — |
| A100 | 80 GB | 1 | $2.70/hr | — |
Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.
These prices move weekly
The Delivered Compute Report tracks them — free, every row sourced.
Subscribe free