An on-premise AI server is a GPU server in your own data centre or colocation cage that runs language models, embeddings, retrieval and agents, so prompts and documents never travel to a cloud API. Size it by GPU memory: model weights plus the KV cache your concurrent users create. One 48–96 GB GPU covers a pilot; hundreds of active users need a multi-GPU node.
What an on-premise AI server has to run
An AI server is rarely one model behind one endpoint. In production it carries four workloads, and each leans on a different part of the machine.
| Workload | What it does | What it consumes |
|---|---|---|
| LLM inference | Answers questions, drafts, summarises and classifies | GPU memory for weights and KV cache; memory bandwidth for generation speed |
| Embeddings | Converts documents and queries into vectors | GPU time in bursts while documents are ingested; little at query time |
| Retrieval (RAG) | Searches the vector index, reranks results and builds the prompt | CPU, RAM and fast SSD; retrieved passages lengthen every prompt |
| Agents | Plan steps, call tools, check policy and retry | Several model calls per task, plus CPU for orchestration |
Two of these change the hardware answer more than buyers expect. Retrieval makes prompts longer, and a longer prompt means a larger KV cache for that conversation. Agents multiply model calls, because one agent run can issue several requests where a chat user issues one. Our reference architecture suggests starting the orchestration tier at 2–4 CPU cores and 8–16 GB of RAM per concurrent agent session, then correcting from measured traces.
Choose the models first, by task. The local LLM guide lists the 2026 open-weight options and the hardware tier each one needs. This guide picks up from there, at the server.
GPU options and VRAM for an LLM server
Three numbers on a GPU spec sheet decide most of the outcome. Memory capacity sets which models fit and how many conversations run at once. Memory bandwidth largely sets how quickly each conversation generates text. The interconnect decides how well a model split across several GPUs performs.
The figures below come from the vendors’ product pages (verified October 2026).
| Accelerator | Memory | Bandwidth | Max power | Form factor | Typical role |
|---|---|---|---|---|---|
| NVIDIA L40S | 48 GB GDDR6 with ECC | 864 GB/s | 350 W | PCIe dual-slot, passive cooling; no NVLink or MIG | Pilots, departmental RAG, 30B-class models at 4-bit |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 96 GB GDDR7 with ECC | 1,597 GB/s | Up to 600 W, configurable | PCIe, dual-slot air or single-slot liquid; MIG up to four instances | A 70B model at 4-bit on one card, or several small models |
| NVIDIA H100 SXM / H100 NVL | 80 GB / 94 GB | 3.35 / 3.9 TB/s | Up to 700 W / 350–400 W | SXM module / PCIe dual-slot | Production serving with continuous batching |
| NVIDIA H200 SXM / H200 NVL | 141 GB | 4.8 TB/s | Up to 700 W / up to 600 W | SXM module / PCIe dual-slot | A 70B model at FP8 on one GPU, with room left for KV cache |
| NVIDIA DGX B200 | Eight GPUs, 1,440 GB HBM3e in total | 64 TB/s in total | About 14.3 kW per system | 10 RU system | Many teams and models on one node; the largest open-weight models |
| AMD Instinct MI350X / MI355X | 288 GB HBM3E | Up to 8 TB/s | 1,000 W / 1,400 W | Air-cooled (MI350X) or direct liquid-cooled (MI355X) | Very large models per GPU on the ROCm stack |
| NVIDIA DGX Spark | 128 GB unified LPDDR5x | 273 GB/s | 240 W power supply | Desk-side system | Model evaluation and development, not shared serving |
A few notes behind the table:
- SXM or PCIe. H100 and H200 SXM modules sit on multi-GPU boards, as in NVIDIA’s DGX systems, and exchange data over NVLink. PCIe cards such as the L40S, the RTX PRO 6000 and the NVL parts go into ordinary rack servers and talk over the PCIe bus, which is much slower: on the H100, NVIDIA lists 900 GB/s for NVLink against 128 GB/s for PCIe Gen5.
- The L40S has no NVLink. Favour models that fit on one card, or give each card its own smaller model.
- AMD runs the same serving software. vLLM supports AMD Instinct MI300 and MI350 GPUs through ROCm, so the serving layer does not tie you to one GPU vendor.
- DGX Spark is a test bench. NVIDIA rates it for inference on models of up to 200 billion parameters, or 700 billion across four linked units, but every user shares one pool of 273 GB/s memory bandwidth.
Sizing an on-premise AI server by number of users
Size from concurrency rather than headcount. The reference architecture groups deployments into three tiers: 25–100 named users with 2–10 concurrent conversations, 100–1,000 users with 10–75 concurrent sessions, and thousands of users with 75 or more. Two formulas turn a tier into gigabytes:
- Weights ≈ parameters × bits per weight ÷ 8. A 70B model needs about 140 GB at 16-bit, 70 GB at FP8 and 35 GB at 4-bit, before format overhead.
- KV cache ≈ 2 × layers × key-value heads × head dimension × bytes per value × tokens in context, for every sequence in flight. NVIDIA’s inference guide derives it.
Meta’s Llama 3 70B has 80 layers and 8 key-value heads of dimension 128, so at 16-bit it stores 327,680 bytes of KV cache per token: about 2.7 GB for one 8,192-token conversation. Qwen3-32B has 64 layers with the same head layout, which works out at 262,144 bytes per token, or about 2.1 GB per conversation.
Memory by tier, at an 8,192-token context (add roughly 10% for the runtime):
| Tier | Sessions at once | Example model | Weights | KV cache | Server that fits |
|---|---|---|---|---|---|
| Pilot, 25–100 users | Up to 10 | 32B dense at FP8, 16-bit KV cache | ≈33 GB | ≈21 GB | One H100 80 GB or RTX PRO 6000 96 GB; one L40S if weights drop to 4-bit and the KV cache to FP8 (≈29 GB in total) |
| Department, 100–1,000 users | Up to 75 | 70B dense at FP8, FP8 KV cache | ≈70 GB | ≈101 GB | Two H200, or four H100 or RTX PRO 6000 cards |
| Enterprise, thousands of users | 75 or more, e.g. 200 | 70B dense at FP8, FP8 KV cache, plus other models | ≈70 GB | ≈268 GB at 200 sessions | An eight-GPU node such as a DGX H200 (1,128 GB) or DGX B200 (1,440 GB), with a second node for failover |
Context length moves these numbers more than headcount does. The same 75 department sessions at 32,768 tokens would need about 400 GB of FP8 KV cache, four times the figure in the table.
These estimates tell you whether the work fits in memory, not how fast it runs. Generation speed depends on memory bandwidth, batch size and the serving engine, so load-test a shortlisted server with your own prompts before you sign. The GPU sizing walkthrough works through 5, 50 and 200 users step by step.
Power, cooling and rack space
Ask facilities three questions before you pick a form factor: how many kilowatts each rack can draw, whether the room can remove that heat with air, and whether liquid cooling is available or planned. The vendor figures show the range (verified October 2026):
- Four L40S cards draw up to 1.4 kW between them (4 × 350 W), before the host’s CPUs, memory and fans.
- Four RTX PRO 6000 Server Edition cards can draw up to 2.4 kW, although their power limit is configurable.
- An eight-GPU DGX H100 is an 8U system rated at about 10.2 kW maximum, and NVIDIA specifies DGX H100 and H200 systems for room temperatures of 5–30 °C.
- A DGX B200 is a 10 RU system rated at about 14.3 kW maximum.
- AMD’s MI355X has a 1,400 W board power per GPU and is direct liquid-cooled; the 1,000 W MI350X is the air-cooled option.
If your server rooms were built for general-purpose racks, an eight-GPU system may need a power and cooling upgrade or a colocation cage. The GPU sourcing guide compares owning, colocation, leasing and dedicated hosting for exactly that situation.
Networking and storage
Inside the server. SXM systems connect their GPUs over NVLink, and PCIe cards talk across the PCIe bus. The difference matters when one model is split over several GPUs with tensor parallelism, because each generation step then exchanges data between them.
Between servers. A model that fits on one node needs only ordinary data-centre networking to reach its users, since requests and responses are small text payloads. Models split across nodes are different: they need a high-bandwidth fabric, which is why DGX H100 and B200 systems carry eight ConnectX-7 adapters at up to 400 Gb/s of InfiniBand or Ethernet.
Storage. Keep model weights on local NVMe so restarts and model swaps load quickly. A 70B model takes about 140 GB at 16-bit, and you will keep earlier versions for rollback. As a reference point, a DGX H100 ships with eight 3.84 TB NVMe drives for data cache and two 1.92 TB drives for the operating system. Vector indexes, document stores and audit logs need fast SSDs and the same backup schedule as any production database.
The software stack on an LLM server
From the metal up, a production AI server runs six layers:
- Operating system and drivers. A supported Linux distribution with the NVIDIA driver and CUDA, or ROCm on AMD Instinct.
- Containers and orchestration. Docker on a single host. Kubernetes for several nodes, with the NVIDIA GPU Operator managing drivers, the container toolkit, the device plugin and DCGM-based GPU monitoring.
- Serving engine. vLLM serves models over OpenAI-compatible Chat Completions and Embeddings APIs, started with
vllm serve <model>. Its--kv-cache-dtype fp8option stores the KV cache at 8 bits instead of 16, which halves it; check answer quality afterwards, because some attention types are more sensitive to KV-cache quantization. Ollama and llama.cpp suit single-user test machines. - Gateway and router. One endpoint for every application, which sends each request to the right model and applies quotas and request logging.
- Retrieval. An embedding model, a vector database and a reranker for permission-aware RAG.
- Identity, audit and monitoring. Single sign-on from your directory, role-based access, logs in your SIEM and GPU metrics in your monitoring system.
If the pilot fits on one server, run it there under Docker Compose. Move to Kubernetes when you need several GPU nodes, rolling upgrades or failover.
What an on-premise AI server costs
Data-centre servers are priced on quote. NVIDIA’s product pages for the H100, H200, L40S and DGX B200 show no prices (checked October 2026), so a budget starts from partner or OEM quotes.
A planning anchor. Our TCO guide puts a production server with four H100 GPUs, enough to serve a 70B model, at roughly €80,000–€140,000 per node at mid-2026 pricing. It estimates that a deployment for hundreds of users starts at two to four such nodes, plus redundancy.
Evaluation hardware has a list price. NVIDIA’s DGX Spark costs $4,699, after a February 2026 increase from $3,999 that NVIDIA attributed to industry-wide memory supply constraints (list, verified October 2026).
The server is one line of the budget. Add power and cooling, support contracts, spare capacity, operations staff and a refresh cycle, then compare the total with cloud API spend at your real volume. The on-premise LLM cost comparison shows where the breakeven falls. For VDF AI software, see current pricing.
Buy, lease or rent
Owning the server gives you custody and the lowest cost per request once it is busy. It also commits capital before you know your real concurrency and context lengths. A sensible sequence is to pilot on dedicated single-tenant GPUs or a single owned card, measure real usage, then buy the production server those measurements justify.
Leasing changes how the hardware is paid for, and colocation changes where it sits; neither has to change who controls it. Check lead times early, since GPU servers and data-hall power can take longer to arrive than the software.
How VDF AI fits
VDF AI is the software layer for the server you choose. It ships as containers for Docker or Kubernetes, and its platform services need CPU and memory rather than GPUs: the self-hosting documentation sizes a pilot host at 4–8 vCPU and 16–32 GB of RAM. Your GPUs serve the models through Ollama, vLLM or another OpenAI-compatible endpoint, and VDF AI Router sends routine requests to smaller local models so the large model’s capacity goes where it improves the answer.
Users sign in with Microsoft Entra ID single sign-on, which is built in. Okta, Keycloak and other SAML or OIDC providers connect through an SSO-aware reverse proxy in front of the platform. Roles and permission groups then decide who can reach which agents, tools and workspaces. The on-premises LLM page walks through the deployment pattern layer by layer.
Sources
- NVIDIA H100 specifications
- NVIDIA H200 specifications
- NVIDIA L40S specifications
- NVIDIA RTX PRO 6000 Blackwell Server Edition
- NVIDIA DGX B200 specifications
- NVIDIA DGX H100 datasheet
- NVIDIA DGX H100/H200 user guide: hardware specifications
- NVIDIA DGX Spark specifications
- NVIDIA DGX Spark price change, February 2026
- AMD Instinct MI350 Series architecture (ROCm docs)
- vLLM online serving
- vLLM quantized KV cache
- vLLM GPU platform support
- NVIDIA GPU Operator overview
- NVIDIA, Mastering LLM Techniques: Inference Optimization
- Meta, The Llama 3 Herd of Models
- Qwen3-32B model card