Model guide · prices from 2026-09-10 · API prices daily, full survey weekly
Cheapest way to run Kimi K2 (1T MoE)
Kimi K2 (1T MoE) needs ≈616 GB of VRAM at 4-bit — this week that starts at $7.26/hour (8× A100-SXM, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.
MoE note: Kimi K2 (1T MoE) activates ~32B active per token, but the full 1026B weights must fit in VRAM — memory scales with total parameters, speed with active ones.
FP16 — ≈2462 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| A100-SXM | 80 GB | 31 (multi-node) | $28.12/hr | $71.30/hr |
| RTX5090 | 32 GB | 77 (multi-node) | $30.88/hr | — |
| RTX4090 | 24 GB | 103 (multi-node) | $38.52/hr | — |
| A100-PCIe | 80 GB | 31 (multi-node) | $41.85/hr | — |
| MI300X | 192 GB | 13 (multi-node) | $44.85/hr | — |
| H100-PCIe | 80 GB | 31 (multi-node) | $61.69/hr | — |
| B300 | 288 GB | 9 (multi-node) | $67.50/hr | $160.22/hr |
| L40S | 48 GB | 52 (multi-node) | $71.24/hr | — |
FP8/INT8 — ≈1231 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| A100-SXM | 80 GB | 16 (multi-node) | $14.51/hr | $36.80/hr |
| RTX5090 | 32 GB | 39 (multi-node) | $15.64/hr | — |
| RTX4090 | 24 GB | 52 (multi-node) | $19.45/hr | — |
| A100-PCIe | 80 GB | 16 (multi-node) | $21.60/hr | — |
| MI300X | 192 GB | 7 | $24.15/hr | — |
| H100-PCIe | 80 GB | 16 (multi-node) | $31.84/hr | — |
| L40S | 48 GB | 26 (multi-node) | $35.62/hr | — |
| H200-SXM | 141 GB | 9 (multi-node) | $35.91/hr | $38.61/hr |
4-bit — ≈616 GB needed
| GPU | VRAM | GPUs needed | Cheapest tracked | Verifiable A/B |
|---|---|---|---|---|
| A100-SXM | 80 GB | 8 | $7.26/hr | $18.40/hr |
| RTX5090 | 32 GB | 20 (multi-node) | $8.02/hr | — |
| RTX4090 | 24 GB | 26 (multi-node) | $9.72/hr | — |
| A100-PCIe | 80 GB | 8 | $10.80/hr | — |
| MI300X | 192 GB | 4 | $13.80/hr | — |
| H100-PCIe | 80 GB | 8 | $15.92/hr | — |
| L40S | 48 GB | 13 (multi-node) | $17.81/hr | — |
| H200-SXM | 141 GB | 5 | $19.95/hr | $21.45/hr |
Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.
These prices move weekly
The Delivered Compute Report tracks them — free, every row sourced.
Subscribe free