Help Me Choose
Two common situations, two very different budgets. Pick the one that sounds like you.
Small Startup
- You want to self-host an open model for a chatbot or prototype.
- You need it to be genuinely capable, not a toy — so you target Qwen2.5-72B (needs 86GB VRAM).
- You want the cheapest hardware that can actually hold 86GB — a single RTX 6000 Ada (48GB) isn't enough, but two are.
Recommended build
2x NVIDIA RTX 6000 Ada
$13,600
96GB VRAM · 600W · runs Qwen2.5-72B (needs 86GB — this build has 96GB)
600W = 0.50 homes' worth of power,
or 0.0067 of an EV battery charge per hour.
Request a Quote for This Build
Mid-Size Company
- You need to run a frontier-class model like Llama-3.1-405B (needs 486GB VRAM) in production.
- You need headroom for real traffic — longer context windows and multiple users at once, not just a demo.
- One DGX H100 (640GB) fits the model with 154GB of headroom to spare, in a single supported server.
Recommended build
1x NVIDIA DGX H100
$200,000
640GB VRAM · 10,200W · runs Llama-3.1-405B (needs 486GB — this build has 640GB, 154GB of headroom)
10,200W = 8.50 homes' worth of power,
or 0.1133 of an EV battery charge per hour.
Request a Quote for This Build