// on-premise ai
On-prem AI that fits your racks.
LLM inference and RAG on servers in your data centre or colo, sized to the workload and the power you actually have.
// when it makes sense
On-prem is the right call when…
01
Usage is steady
GPUs busy most of the working day, every day. Idle owned GPUs are expensive; busy ones are cheap per token.
Read more 02
Data can't leave
Regulated records, privileged documents, or contracts that forbid third-party processing.
03
You have the facility
Rack space with the power and cooling for 5–15 kW per GPU server, or a colo partner who does.
// what you need
An on-prem LLM deployment, layer by layer
| Layer | What we specify | Typical decision |
|---|---|---|
| GPU servers | GPU model and count, CPU, system memory, NVMe | PCIe server vs 8-GPU HGX; see GPU servers |
| Power & cooling | kW per rack, PDUs, airflow or liquid | Can the room take 10 kW+ per server? |
| Network | Front-end Ethernet, GPU fabric if multi-node | Single server: 25–100G is plenty. Multi-node training: 400G fabric |
| Storage | Model weights, vector index, logs | Local NVMe for weights; shared storage for RAG corpora |
| Software | OS, drivers, CUDA, Kubernetes or Slurm, vLLM / NIM | Pinned versions validated against the GPU support matrix |
| Operations | Monitoring, patching, warranty and RMA | Who gets paged when a GPU fails at 2 a.m.? |
// site readiness
Check before you buy
- Power per rack and per circuit, with headroom
- Cooling capacity for sustained full load, not average
- Physical space, weight limits and delivery access
- Network uplinks and where the GPU fabric terminates
- OS, driver and CUDA versions your models and tools require
- Support and RMA terms for GPUs specifically
Start from the workload
Size memory with the LLM GPU calculator, check the economics with the cloud vs on-prem calculator, then send us the workload. The assessment returns an on-prem architecture and BOM, with the cloud alternative priced next to it.
Need to deploy private AI?
We design it, validate it, source it, deploy it and help run it.
Get my 48-hour blueprint
Give us your workload, users, budget and timeline. Within 48 hours we'll give you a validated cloud/on-prem architecture, BOM, expected costs and sourcing options.