GPU Sizing
How Many GPUs Does a 70B LLM Need for 100 Concurrent Users?
Pankaj Kharkwal · Co-founder & CTO, Pankh.AIShort answer: at FP16 with a 4,096-token context, Llama 3.3 70B serving 100 concurrent users needs about 306 GB of GPU memory. That is four H100 80 GB GPUs, or two B200s, or two MI300Xs. Quantise weights and KV cache to FP8 and it drops to about 153 GB: two H100s, or a single B200 or MI300X.
That answer is only as good as its assumptions, and the biggest one, context length, can move it by an order of magnitude. Here is the arithmetic, so you can redo it for your own workload.
Assumptions
Change any of these and the result changes; the LLM GPU calculator lets you do that interactively.
- Model: Llama 3.3 70B, 70.6B parameters, 80 layers, 8 KV heads (grouped-query attention), head dimension 128
- 100 concurrent sequences: requests in flight at the same moment, not total users
- 4,096 tokens held in cache per sequence (prompt plus generated output)
- Serving with vLLM or similar, using 90% of GPU memory for weights and cache
Step 1: model weights
Weights take parameters × bytes per parameter. At FP16 or BF16 that is 70.6B × 2 bytes ≈ 141 GB. At FP8 it halves to about 71 GB; INT4 with group-wise scales lands near 39 GB. Weights are a fixed cost: they do not grow with users.
Step 2: KV cache
Every token a user has in context stores a key and a value vector in every layer. Per token: 2 × 80 layers × 8 KV heads × 128 dimensions × 2 bytes = 327,680 bytes, about 0.33 MB at FP16.
Multiply by context and concurrency: 0.33 MB × 4,096 tokens × 100 sequences ≈ 134 GB. That is almost as much as the weights, and unlike the weights it scales linearly with both users and context. FP8 KV cache halves it to about 67 GB.
Step 3: total memory and what fits
Add weights and cache, then divide by 0.9 for runtime headroom: (141 + 134) ÷ 0.9 ≈ 306 GB at FP16, or (71 + 67) ÷ 0.9 ≈ 153 GB at FP8. GPU counts are then rounded up to a tensor-parallel size the model can be split by: it must divide the 8 KV heads, so 1, 2, 4 or 8 per server.
| GPU | Memory | FP16 (306 GB) | FP8 (153 GB) |
|---|---|---|---|
| NVIDIA L40S | 48 GB | 8 | 4 |
| NVIDIA RTX PRO 6000 Blackwell | 96 GB | 4 | 2 |
| NVIDIA H100 SXM | 80 GB | 4 | 2 |
| NVIDIA H100 NVL | 94 GB | 4 | 2 |
| NVIDIA H200 | 141 GB | 4 (3 by memory) | 2 |
| NVIDIA B200 | 180 GB | 2 | 1 |
| NVIDIA B300 | 288 GB | 2 | 1 |
| AMD Instinct MI300X | 192 GB | 2 | 1 |
| AMD Instinct MI355X | 288 GB | 2 | 1 |
GPUs needed by memory, rounded to a valid tensor-parallel size. Vendor datasheet memory; HGX B200/B300 shown as the per-GPU share of the 8-GPU baseboard.
Context length matters more than model size
Raise the context from 4k to 32k tokens and the FP16 KV cache grows eightfold, to about 1,074 GB. The same 100 users now need roughly 1.35 TB of GPU memory: more than an 8-GPU H100 or H200 server holds, so the deployment spans servers and the network becomes part of the design.
This is why a RAG system that stuffs long documents into every prompt can need far more hardware than a chat assistant on the same model. Measure real prompt lengths before sizing.
What memory does not tell you
Fitting in memory is necessary, not sufficient. Throughput and latency depend on memory bandwidth, batch size, the ratio of prompt to output tokens, and the serving stack. Two deployments that fit identically can deliver very different tokens per second.
- Benchmark with your real prompts and output lengths, not a synthetic default
- Decide a latency target: time to first token and tokens per second per user
- Size for peak concurrency, then check average utilisation for the cost case
- Check rack power: an 8-GPU H100 or H200 server can draw 10 kW or more
Cloud or owned?
Four H100s busy all day is a different economic decision from four H100s busy two hours a day. Once you know the GPU count, run both through the cloud vs on-prem calculator, or send us the workload and we will cost both options with a bill of materials within 48 hours.
Working through this for your own team?