AI Models & Memory Math

Every AI model needs enough GPU memory (VRAM) to hold its parameters, plus overhead for running. Rule of thumb: 1GB of VRAM per 1 billion parameters, plus about 20% extra for the model to actually run (KV cache, activations, etc).

Llama-3.1-405B

Meta's largest open-weight model — frontier-class general reasoning and chat.

405B params × 1GB per billion params = 405GB, plus 20% overhead for running the model (KV cache, activations, etc.) = 486GB needed → 486GB needed
Hardware VRAM per unit Units needed Total VRAM Fits?
RTX 4090 24GB 21× 504GB Needs 21× — not practical (no fast memory pooling between cards)
RTX 6000 Ada 48GB 11× 528GB Needs 11× — not practical (no fast memory pooling between cards)
H100 SXM 80GB 560GB Fits with 7×, pooled memory
DGX H100 640GB 1 640GB Fits on 1
GB200 NVL72 rack 13824GB 1 13,824GB Fits on 1

Qwen2.5-72B

A strong mid-size open model — the sweet spot for capability per dollar.

72B params × 1GB per billion params = 72GB, plus 20% overhead = 86GB needed → 86GB needed
Hardware VRAM per unit Units needed Total VRAM Fits?
RTX 4090 24GB 96GB Needs 4× — not practical (no fast memory pooling between cards)
RTX 6000 Ada 48GB 96GB Needs 2× — not practical (no fast memory pooling between cards)
H100 SXM 80GB 160GB Fits with 2×, pooled memory
DGX H100 640GB 1 640GB Fits on 1
GB200 NVL72 rack 13824GB 1 13,824GB Fits on 1

DeepSeek-V3

A massive mixture-of-experts model rivaling the biggest closed models.

671B params × 1GB per billion params = 671GB, plus 20% overhead = 805GB needed → 805GB needed
Hardware VRAM per unit Units needed Total VRAM Fits?
RTX 4090 24GB 34× 816GB Needs 34× — not practical (no fast memory pooling between cards)
RTX 6000 Ada 48GB 17× 816GB Needs 17× — not practical (no fast memory pooling between cards)
H100 SXM 80GB 11× 880GB Fits with 11×, pooled memory
DGX H100 640GB 1,280GB Fits with 2×, pooled memory
GB200 NVL72 rack 13824GB 1 13,824GB Fits on 1
Why "needs 11× RTX 4090" isn't a real answer: the RTX 4090 has no NVLink, so those cards can't pool memory into one fast pool the way data-center GPUs can. Piling up desktop cards for models this size means slow, clunky software workarounds, not a real solution — see the Clusters page for why.