// gpu servers

GPU servers, sized to the model.

From a single AI workstation to 8-GPU HGX systems, specified from your workload and sourced through Arrow, TD SYNNEX, Dell, HPE, Lenovo and Supermicro.

// three form factors

Which kind of GPU server?

01

AI workstation

1–2 GPUs (e.g. RTX PRO 6000, 96 GB). Development, fine-tuning small models, a single team's inference. Runs under a desk on standard power.
02

PCIe GPU server

2–8 PCIe GPUs (L40S, H100 NVL, H200 NVL, RTX PRO 6000) in a 2U–4U rack server. Production inference for most 8B–70B models.
03

8-GPU HGX / OAM

8 SXM or OAM GPUs (H100, H200, B200, B300, MI300X, MI355X) linked by NVLink or Infinity Fabric. Large models, high concurrency, training. Needs serious rack power.

// accelerators

GPU memory decides what fits

Memory per GPU sets how big a model and how much context you can hold; bandwidth sets how fast tokens come out.

GPUMemoryBandwidthUsually ships in
NVIDIA L40S48 GB GDDR60.864 TB/sPCIe server
NVIDIA RTX PRO 6000 Blackwell96 GB GDDR71.79 TB/sWorkstation or PCIe server
NVIDIA H100 SXM80 GB HBM33.35 TB/s8-GPU HGX server
NVIDIA H100 NVL94 GB HBM33.9 TB/sPCIe server
NVIDIA H200141 GB HBM3e4.8 TB/s8-GPU HGX or PCIe (NVL) server
NVIDIA B200180 GB HBM3e8 TB/s8-GPU HGX server
NVIDIA B300288 GB HBM3e8 TB/s8-GPU HGX server
AMD Instinct MI300X192 GB HBM35.3 TB/s8-GPU OAM server
AMD Instinct MI355X288 GB HBM3E8 TB/s8-GPU OAM server

Vendor datasheet figures; HGX B200/B300 memory is the per-GPU share of the 8-GPU baseboard. Exact configurations, availability and lead times are confirmed on quote.

// what we specify

A GPU server quote is more than the GPUs

  • GPU model and count, from weights + KV cache + headroom
  • CPU and system memory that won't starve the GPUs
  • NVMe for model weights and local scratch
  • Network cards: front-end, storage and GPU fabric
  • Power supplies and the rack power they need
  • Warranty, GPU-specific support and RMA terms

Need a number before a quote?

Run your model through the LLM GPU calculator, or browse exact models in the AI server catalogue. For a full specification, BOM and price, send us the workload.

Need to deploy private AI?

We design it, validate it, source it, deploy it and help run it.

Get my 48-hour blueprint

Give us your workload, users, budget and timeline. Within 48 hours we'll give you a validated cloud/on-prem architecture, BOM, expected costs and sourcing options.