ComputeAssay

Model guides · updated 2026-09-10

What it costs to run each model

Per-model guides connecting VRAM requirements to this week's cheapest graded GPU rentals. Pick your model; get GPU counts and live prices at every precision.

Llama 3.1 8B
≈5 GB at 4-bit · from $0.37/hr
Phi-4 (14B)
≈8 GB at 4-bit · from $0.37/hr
Mistral Small 3 (24B)
≈14 GB at 4-bit · from $0.37/hr
Gemma 2 27B
≈16 GB at 4-bit · from $0.37/hr
Qwen2.5 32B
≈19 GB at 4-bit · from $0.37/hr
Llama 3.3 70B
≈42 GB at 4-bit · from $0.75/hr
Qwen2.5 72B
≈43 GB at 4-bit · from $0.75/hr
Llama 4 Scout (109B MoE)
≈65 GB at 4-bit · from $0.91/hr
gpt-oss-120b (117B MoE)
≈70 GB at 4-bit · from $0.91/hr
Mistral Large 2 (123B)
≈74 GB at 4-bit · from $0.91/hr
Mixtral 8x22B (141B MoE)
≈85 GB at 4-bit · from $1.20/hr
Qwen3 235B (MoE)
≈141 GB at 4-bit · from $1.81/hr
Llama 4 Maverick (400B MoE)
≈240 GB at 4-bit · from $2.72/hr
Llama 3.1 405B
≈243 GB at 4-bit · from $3.21/hr
DeepSeek-R1 (671B MoE)
≈403 GB at 4-bit · from $5.21/hr
Kimi K2 (1T MoE)
≈616 GB at 4-bit · from $7.26/hr

Any other size: the interactive fit finder covers arbitrary models and precisions. Thinking in tokens instead of GPUs? What a million tokens actually costs.