ComputeAssay

Model guide · prices from 2026-09-10 · API prices daily, full survey weekly

Cheapest way to run Mixtral 8x22B (141B MoE)

Mixtral 8x22B (141B MoE) needs ≈85 GB of VRAM at 4-bit — this week that starts at $1.20/hour (3× RTX5090, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.

MoE note: Mixtral 8x22B (141B MoE) activates ~39B active per token, but the full 141B weights must fit in VRAM — memory scales with total parameters, speed with active ones.

FP16 — ≈338 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB11 (multi-node)$4.41/hr
A100-SXM80 GB5$4.54/hr$11.50/hr
RTX409024 GB15 (multi-node)$5.61/hr
A100-PCIe80 GB5$6.75/hr
MI300X192 GB2$6.90/hr
H100-PCIe80 GB5$9.95/hr
L40S48 GB8$10.96/hr
H200-SXM141 GB3$11.97/hr$12.87/hr

FP8/INT8 — ≈169 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB6$2.41/hr
A100-SXM80 GB3$2.72/hr$6.90/hr
RTX409024 GB8$2.99/hr
MI300X192 GB1$3.45/hr
A100-PCIe80 GB3$4.05/hr
L40S48 GB4$5.48/hr
H100-PCIe80 GB3$5.97/hr
B200192 GB1$6.00/hr$7.15/hr

4-bit — ≈85 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB3$1.20/hr
RTX409024 GB4$1.50/hr
A100-SXM80 GB2$1.81/hr$4.60/hr
A100-PCIe80 GB2$2.70/hr
L40S48 GB2$2.74/hr
MI300X192 GB1$3.45/hr
L424 GB4$3.48/hr
H100-PCIe80 GB2$3.98/hr

Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.

These prices move weekly

The Delivered Compute Report tracks them — free, every row sourced.

Subscribe free