// tools / llm gpu calculator
How many GPUs does your model need?
Weights plus KV cache, mapped to today's data-centre GPUs. Every formula is on the page.
70.6B params · 80 layers · 8 KV heads · head dim 128
Weight precision
KV cache precision
Requests in flight at the same moment, not total staff.
GPU memory needed
306 GB
- Weights
- 141.2 GB
- KV cache
- 134.2 GB
- KV per token
- 0.328 MB
- Runtime headroom
- 10%
| GPU | Memory | GPUs by memory | Deploy as |
|---|---|---|---|
| NVIDIA L40S | 48 GB GDDR6 | 7 | 8 GPUs, tensor parallel 8 |
| NVIDIA RTX PRO 6000 Blackwell | 96 GB GDDR7 | 4 | 4 GPUs, tensor parallel 4 |
| NVIDIA H100 SXM | 80 GB HBM3 | 4 | 4 GPUs, tensor parallel 4 |
| NVIDIA H100 NVL | 94 GB HBM3 | 4 | 4 GPUs, tensor parallel 4 |
| NVIDIA H200 | 141 GB HBM3e | 3 | 4 GPUs, tensor parallel 4 |
| NVIDIA B200 | 180 GB HBM3e | 2 | 2 GPUs, tensor parallel 2 |
| NVIDIA B300 | 288 GB HBM3e | 2 | 2 GPUs, tensor parallel 2 |
| AMD Instinct MI300X | 192 GB HBM3 | 2 | 2 GPUs, tensor parallel 2 |
| AMD Instinct MI355X | 288 GB HBM3E | 2 | 2 GPUs, tensor parallel 2 |
Memory is step one. Throughput, latency, power and the network decide the real build.
Size my actual workloadHow the calculator works
Weights = parameters × bytes per weight: 2 for FP16/BF16, 1 for FP8, about 0.55 for INT4 (0.5 plus group-wise scales).
KV cache = 2 (keys and values) × layers × KV heads × head dimension × bytes per value × context tokens × concurrent sequences. Models with grouped-query attention, like Llama 3, keep far fewer KV heads than attention heads, which is why their cache is small per token.
Required memory = (weights + KV cache) ÷ 0.9, leaving 10% for activations, CUDA graphs and fragmentation; vLLM’s default memory utilisation is 0.9.
Deploy as rounds up to a tensor-parallel size the model can actually be split by (it must divide the KV head count, so 1, 2, 4 or 8 per server). Past eight GPUs the model spans servers, which brings in the network.
This sizes capacity, not speed. Tokens per second depend on memory bandwidth, batch size, prompt-to-output ratio and the serving stack. Mixture-of-experts models count every expert, because all of them must be resident. GPU figures are vendor datasheet values; HGX B200 and B300 show the per-GPU share of the 8-GPU baseboard. Worked example: how many GPUs a 70B model needs for 100 concurrent users.
Need to deploy private AI?
We design it, validate it, source it, deploy it and help run it.
Give us your workload, users, budget and timeline. Within 48 hours we'll give you a validated cloud/on-prem architecture, BOM, expected costs and sourcing options.