ComputeAssay

Model guide · prices from 2026-09-10 · API prices daily, full survey weekly

Cheapest way to run Qwen3 235B (MoE)

Qwen3 235B (MoE) needs ≈141 GB of VRAM at 4-bit — this week that starts at $1.81/hour (2× A100-SXM, cheapest tracked walk-up price). FP8 and FP16 below. VRAM calculators stop at the memory number; this connects it to live, graded prices across 20 providers.

MoE note: Qwen3 235B (MoE) activates ~22B active per token, but the full 235B weights must fit in VRAM — memory scales with total parameters, speed with active ones.

FP16 — ≈564 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB18 (multi-node)$7.22/hr
A100-SXM80 GB8$7.26/hr$18.40/hr
RTX409024 GB24 (multi-node)$8.98/hr
MI300X192 GB3$10.35/hr
A100-PCIe80 GB8$10.80/hr
B300288 GB2$15.00/hr$35.60/hr
H100-PCIe80 GB8$15.92/hr
H200-SXM141 GB4$15.96/hr$17.16/hr

FP8/INT8 — ≈282 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
RTX509032 GB9 (multi-node)$3.61/hr
A100-SXM80 GB4$3.63/hr$9.20/hr
RTX409024 GB12 (multi-node)$4.49/hr
A100-PCIe80 GB4$5.40/hr
MI300X192 GB2$6.90/hr
B300288 GB1$7.50/hr$17.80/hr
H100-PCIe80 GB4$7.96/hr
H200-SXM141 GB2$7.98/hr$8.58/hr

4-bit — ≈141 GB needed

GPUVRAMGPUs neededCheapest trackedVerifiable A/B
A100-SXM80 GB2$1.81/hr$4.60/hr
RTX509032 GB5$2.00/hr
RTX409024 GB6$2.24/hr
A100-PCIe80 GB2$2.70/hr
MI300X192 GB1$3.45/hr
H100-PCIe80 GB2$3.98/hr
H200-SXM141 GB1$3.99/hr$4.29/hr
L40S48 GB3$4.11/hr

Method: VRAM ≈ parameters × bytes/parameter × 1.2 (20% headroom for KV cache and activations at ~8K context). Longer context or high concurrency needs more; training needs far more (optimizer states: ~4–8× weights). GPU counts assume even sharding. Above one 8-GPU node, the interconnect becomes the constraint — restrict to Grade A/B listings. Grades measure what sellers publish — deliverability of the written claim, not measured performance. No fleet holds a measured (Verified) badge yet; that tier arrives with the verification harness. Prices are this week's cheapest tracked walk-up listings (teaser prices excluded); every listing with grades and sources on the GPU pages. Interactive version for any model size: Model Fit.

These prices move weekly

The Delivered Compute Report tracks them — free, every row sourced.

Subscribe free