ComputeAssay

Model guide · prices from 2026-09-10 · API prices daily, full survey weekly

Cheapest way to run Kimi K2 (1T MoE)

Kimi K2 (1T MoE) needs ≈616 GB of VRAM at 4-bit — this week that starts at $7.26/hour (8× A100-SXM, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.

MoE note: Kimi K2 (1T MoE) activates ~32B active per token, but the full 1026B weights must fit in VRAM — memory scales with total parameters, speed with active ones.

FP16 — ≈2462 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB31 (multi-node)$28.12/hr$71.30/hr
RTX509032 GB77 (multi-node)$30.88/hr
RTX409024 GB103 (multi-node)$38.52/hr
A100-PCIe80 GB31 (multi-node)$41.85/hr
MI300X192 GB13 (multi-node)$44.85/hr
H100-PCIe80 GB31 (multi-node)$61.69/hr
B300288 GB9 (multi-node)$67.50/hr$160.22/hr
L40S48 GB52 (multi-node)$71.24/hr

FP8/INT8 — ≈1231 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB16 (multi-node)$14.51/hr$36.80/hr
RTX509032 GB39 (multi-node)$15.64/hr
RTX409024 GB52 (multi-node)$19.45/hr
A100-PCIe80 GB16 (multi-node)$21.60/hr
MI300X192 GB7$24.15/hr
H100-PCIe80 GB16 (multi-node)$31.84/hr
L40S48 GB26 (multi-node)$35.62/hr
H200-SXM141 GB9 (multi-node)$35.91/hr$38.61/hr

4-bit — ≈616 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB8$7.26/hr$18.40/hr
RTX509032 GB20 (multi-node)$8.02/hr
RTX409024 GB26 (multi-node)$9.72/hr
A100-PCIe80 GB8$10.80/hr
MI300X192 GB4$13.80/hr
H100-PCIe80 GB8$15.92/hr
L40S48 GB13 (multi-node)$17.81/hr
H200-SXM141 GB5$19.95/hr$21.45/hr

Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.

These prices move weekly

The Delivered Compute Report tracks them — free, every row sourced.

Subscribe free