// tools / llm gpu calculator

How many GPUs does your model need?

Weights plus KV cache, mapped to today's data-centre GPUs. Every formula is on the page.

70.6B params · 80 layers · 8 KV heads · head dim 128

Weight precision

KV cache precision

Requests in flight at the same moment, not total staff.

GPU memory needed

306 GB

Weights
141.2 GB
KV cache
134.2 GB
KV per token
0.328 MB
Runtime headroom
10%
GPUMemoryGPUs by memoryDeploy as
NVIDIA L40S48 GB GDDR678 GPUs, tensor parallel 8
NVIDIA RTX PRO 6000 Blackwell96 GB GDDR744 GPUs, tensor parallel 4
NVIDIA H100 SXM80 GB HBM344 GPUs, tensor parallel 4
NVIDIA H100 NVL94 GB HBM344 GPUs, tensor parallel 4
NVIDIA H200141 GB HBM3e34 GPUs, tensor parallel 4
NVIDIA B200180 GB HBM3e22 GPUs, tensor parallel 2
NVIDIA B300288 GB HBM3e22 GPUs, tensor parallel 2
AMD Instinct MI300X192 GB HBM322 GPUs, tensor parallel 2
AMD Instinct MI355X288 GB HBM3E22 GPUs, tensor parallel 2

Memory is step one. Throughput, latency, power and the network decide the real build.

Size my actual workload

How the calculator works

Weights = parameters × bytes per weight: 2 for FP16/BF16, 1 for FP8, about 0.55 for INT4 (0.5 plus group-wise scales).

KV cache = 2 (keys and values) × layers × KV heads × head dimension × bytes per value × context tokens × concurrent sequences. Models with grouped-query attention, like Llama 3, keep far fewer KV heads than attention heads, which is why their cache is small per token.

Required memory = (weights + KV cache) ÷ 0.9, leaving 10% for activations, CUDA graphs and fragmentation; vLLM’s default memory utilisation is 0.9.

Deploy as rounds up to a tensor-parallel size the model can actually be split by (it must divide the KV head count, so 1, 2, 4 or 8 per server). Past eight GPUs the model spans servers, which brings in the network.

This sizes capacity, not speed. Tokens per second depend on memory bandwidth, batch size, prompt-to-output ratio and the serving stack. Mixture-of-experts models count every expert, because all of them must be resident. GPU figures are vendor datasheet values; HGX B200 and B300 show the per-GPU share of the 8-GPU baseboard. Worked example: how many GPUs a 70B model needs for 100 concurrent users.

Need to deploy private AI?

We design it, validate it, source it, deploy it and help run it.

Get my 48-hour blueprint

Give us your workload, users, budget and timeline. Within 48 hours we'll give you a validated cloud/on-prem architecture, BOM, expected costs and sourcing options.