AI Infrastructure

On-Premise AI Server: How to Choose, Size and Cost an LLM Server (2026)

A buyer's guide to the on-premise AI server: what it has to run, which GPUs and how much VRAM, how to size it by number of users, the power, cooling, networking and storage it needs, the software stack, and what the hardware costs.

An on-premise AI server is a GPU server in your own data centre or colocation cage that runs language models, embeddings, retrieval and agents, so prompts and documents never travel to a cloud API. Size it by GPU memory: model weights plus the KV cache your concurrent users create. One 48–96 GB GPU covers a pilot; hundreds of active users need a multi-GPU node.

What an on-premise AI server has to run

An AI server is rarely one model behind one endpoint. In production it carries four workloads, and each leans on a different part of the machine.

WorkloadWhat it doesWhat it consumes
LLM inferenceAnswers questions, drafts, summarises and classifiesGPU memory for weights and KV cache; memory bandwidth for generation speed
EmbeddingsConverts documents and queries into vectorsGPU time in bursts while documents are ingested; little at query time
Retrieval (RAG)Searches the vector index, reranks results and builds the promptCPU, RAM and fast SSD; retrieved passages lengthen every prompt
AgentsPlan steps, call tools, check policy and retrySeveral model calls per task, plus CPU for orchestration

Two of these change the hardware answer more than buyers expect. Retrieval makes prompts longer, and a longer prompt means a larger KV cache for that conversation. Agents multiply model calls, because one agent run can issue several requests where a chat user issues one. Our reference architecture suggests starting the orchestration tier at 2–4 CPU cores and 8–16 GB of RAM per concurrent agent session, then correcting from measured traces.

Choose the models first, by task. The local LLM guide lists the 2026 open-weight options and the hardware tier each one needs. This guide picks up from there, at the server.

GPU options and VRAM for an LLM server

Three numbers on a GPU spec sheet decide most of the outcome. Memory capacity sets which models fit and how many conversations run at once. Memory bandwidth largely sets how quickly each conversation generates text. The interconnect decides how well a model split across several GPUs performs.

The figures below come from the vendors’ product pages (verified October 2026).

AcceleratorMemoryBandwidthMax powerForm factorTypical role
NVIDIA L40S48 GB GDDR6 with ECC864 GB/s350 WPCIe dual-slot, passive cooling; no NVLink or MIGPilots, departmental RAG, 30B-class models at 4-bit
NVIDIA RTX PRO 6000 Blackwell Server Edition96 GB GDDR7 with ECC1,597 GB/sUp to 600 W, configurablePCIe, dual-slot air or single-slot liquid; MIG up to four instancesA 70B model at 4-bit on one card, or several small models
NVIDIA H100 SXM / H100 NVL80 GB / 94 GB3.35 / 3.9 TB/sUp to 700 W / 350–400 WSXM module / PCIe dual-slotProduction serving with continuous batching
NVIDIA H200 SXM / H200 NVL141 GB4.8 TB/sUp to 700 W / up to 600 WSXM module / PCIe dual-slotA 70B model at FP8 on one GPU, with room left for KV cache
NVIDIA DGX B200Eight GPUs, 1,440 GB HBM3e in total64 TB/s in totalAbout 14.3 kW per system10 RU systemMany teams and models on one node; the largest open-weight models
AMD Instinct MI350X / MI355X288 GB HBM3EUp to 8 TB/s1,000 W / 1,400 WAir-cooled (MI350X) or direct liquid-cooled (MI355X)Very large models per GPU on the ROCm stack
NVIDIA DGX Spark128 GB unified LPDDR5x273 GB/s240 W power supplyDesk-side systemModel evaluation and development, not shared serving

A few notes behind the table:

  • SXM or PCIe. H100 and H200 SXM modules sit on multi-GPU boards, as in NVIDIA’s DGX systems, and exchange data over NVLink. PCIe cards such as the L40S, the RTX PRO 6000 and the NVL parts go into ordinary rack servers and talk over the PCIe bus, which is much slower: on the H100, NVIDIA lists 900 GB/s for NVLink against 128 GB/s for PCIe Gen5.
  • The L40S has no NVLink. Favour models that fit on one card, or give each card its own smaller model.
  • AMD runs the same serving software. vLLM supports AMD Instinct MI300 and MI350 GPUs through ROCm, so the serving layer does not tie you to one GPU vendor.
  • DGX Spark is a test bench. NVIDIA rates it for inference on models of up to 200 billion parameters, or 700 billion across four linked units, but every user shares one pool of 273 GB/s memory bandwidth.

Sizing an on-premise AI server by number of users

Size from concurrency rather than headcount. The reference architecture groups deployments into three tiers: 25–100 named users with 2–10 concurrent conversations, 100–1,000 users with 10–75 concurrent sessions, and thousands of users with 75 or more. Two formulas turn a tier into gigabytes:

  1. Weights ≈ parameters × bits per weight ÷ 8. A 70B model needs about 140 GB at 16-bit, 70 GB at FP8 and 35 GB at 4-bit, before format overhead.
  2. KV cache ≈ 2 × layers × key-value heads × head dimension × bytes per value × tokens in context, for every sequence in flight. NVIDIA’s inference guide derives it.

Meta’s Llama 3 70B has 80 layers and 8 key-value heads of dimension 128, so at 16-bit it stores 327,680 bytes of KV cache per token: about 2.7 GB for one 8,192-token conversation. Qwen3-32B has 64 layers with the same head layout, which works out at 262,144 bytes per token, or about 2.1 GB per conversation.

Memory by tier, at an 8,192-token context (add roughly 10% for the runtime):

TierSessions at onceExample modelWeightsKV cacheServer that fits
Pilot, 25–100 usersUp to 1032B dense at FP8, 16-bit KV cache≈33 GB≈21 GBOne H100 80 GB or RTX PRO 6000 96 GB; one L40S if weights drop to 4-bit and the KV cache to FP8 (≈29 GB in total)
Department, 100–1,000 usersUp to 7570B dense at FP8, FP8 KV cache≈70 GB≈101 GBTwo H200, or four H100 or RTX PRO 6000 cards
Enterprise, thousands of users75 or more, e.g. 20070B dense at FP8, FP8 KV cache, plus other models≈70 GB≈268 GB at 200 sessionsAn eight-GPU node such as a DGX H200 (1,128 GB) or DGX B200 (1,440 GB), with a second node for failover

Context length moves these numbers more than headcount does. The same 75 department sessions at 32,768 tokens would need about 400 GB of FP8 KV cache, four times the figure in the table.

These estimates tell you whether the work fits in memory, not how fast it runs. Generation speed depends on memory bandwidth, batch size and the serving engine, so load-test a shortlisted server with your own prompts before you sign. The GPU sizing walkthrough works through 5, 50 and 200 users step by step.

Power, cooling and rack space

Ask facilities three questions before you pick a form factor: how many kilowatts each rack can draw, whether the room can remove that heat with air, and whether liquid cooling is available or planned. The vendor figures show the range (verified October 2026):

  • Four L40S cards draw up to 1.4 kW between them (4 × 350 W), before the host’s CPUs, memory and fans.
  • Four RTX PRO 6000 Server Edition cards can draw up to 2.4 kW, although their power limit is configurable.
  • An eight-GPU DGX H100 is an 8U system rated at about 10.2 kW maximum, and NVIDIA specifies DGX H100 and H200 systems for room temperatures of 5–30 °C.
  • A DGX B200 is a 10 RU system rated at about 14.3 kW maximum.
  • AMD’s MI355X has a 1,400 W board power per GPU and is direct liquid-cooled; the 1,000 W MI350X is the air-cooled option.

If your server rooms were built for general-purpose racks, an eight-GPU system may need a power and cooling upgrade or a colocation cage. The GPU sourcing guide compares owning, colocation, leasing and dedicated hosting for exactly that situation.

Networking and storage

Inside the server. SXM systems connect their GPUs over NVLink, and PCIe cards talk across the PCIe bus. The difference matters when one model is split over several GPUs with tensor parallelism, because each generation step then exchanges data between them.

Between servers. A model that fits on one node needs only ordinary data-centre networking to reach its users, since requests and responses are small text payloads. Models split across nodes are different: they need a high-bandwidth fabric, which is why DGX H100 and B200 systems carry eight ConnectX-7 adapters at up to 400 Gb/s of InfiniBand or Ethernet.

Storage. Keep model weights on local NVMe so restarts and model swaps load quickly. A 70B model takes about 140 GB at 16-bit, and you will keep earlier versions for rollback. As a reference point, a DGX H100 ships with eight 3.84 TB NVMe drives for data cache and two 1.92 TB drives for the operating system. Vector indexes, document stores and audit logs need fast SSDs and the same backup schedule as any production database.

The software stack on an LLM server

From the metal up, a production AI server runs six layers:

  1. Operating system and drivers. A supported Linux distribution with the NVIDIA driver and CUDA, or ROCm on AMD Instinct.
  2. Containers and orchestration. Docker on a single host. Kubernetes for several nodes, with the NVIDIA GPU Operator managing drivers, the container toolkit, the device plugin and DCGM-based GPU monitoring.
  3. Serving engine. vLLM serves models over OpenAI-compatible Chat Completions and Embeddings APIs, started with vllm serve <model>. Its --kv-cache-dtype fp8 option stores the KV cache at 8 bits instead of 16, which halves it; check answer quality afterwards, because some attention types are more sensitive to KV-cache quantization. Ollama and llama.cpp suit single-user test machines.
  4. Gateway and router. One endpoint for every application, which sends each request to the right model and applies quotas and request logging.
  5. Retrieval. An embedding model, a vector database and a reranker for permission-aware RAG.
  6. Identity, audit and monitoring. Single sign-on from your directory, role-based access, logs in your SIEM and GPU metrics in your monitoring system.

If the pilot fits on one server, run it there under Docker Compose. Move to Kubernetes when you need several GPU nodes, rolling upgrades or failover.

What an on-premise AI server costs

Data-centre servers are priced on quote. NVIDIA’s product pages for the H100, H200, L40S and DGX B200 show no prices (checked October 2026), so a budget starts from partner or OEM quotes.

A planning anchor. Our TCO guide puts a production server with four H100 GPUs, enough to serve a 70B model, at roughly €80,000–€140,000 per node at mid-2026 pricing. It estimates that a deployment for hundreds of users starts at two to four such nodes, plus redundancy.

Evaluation hardware has a list price. NVIDIA’s DGX Spark costs $4,699, after a February 2026 increase from $3,999 that NVIDIA attributed to industry-wide memory supply constraints (list, verified October 2026).

The server is one line of the budget. Add power and cooling, support contracts, spare capacity, operations staff and a refresh cycle, then compare the total with cloud API spend at your real volume. The on-premise LLM cost comparison shows where the breakeven falls. For VDF AI software, see current pricing.

Buy, lease or rent

Owning the server gives you custody and the lowest cost per request once it is busy. It also commits capital before you know your real concurrency and context lengths. A sensible sequence is to pilot on dedicated single-tenant GPUs or a single owned card, measure real usage, then buy the production server those measurements justify.

Leasing changes how the hardware is paid for, and colocation changes where it sits; neither has to change who controls it. Check lead times early, since GPU servers and data-hall power can take longer to arrive than the software.

How VDF AI fits

VDF AI is the software layer for the server you choose. It ships as containers for Docker or Kubernetes, and its platform services need CPU and memory rather than GPUs: the self-hosting documentation sizes a pilot host at 4–8 vCPU and 16–32 GB of RAM. Your GPUs serve the models through Ollama, vLLM or another OpenAI-compatible endpoint, and VDF AI Router sends routine requests to smaller local models so the large model’s capacity goes where it improves the answer.

Users sign in with Microsoft Entra ID single sign-on, which is built in. Okta, Keycloak and other SAML or OIDC providers connect through an SSO-aware reverse proxy in front of the platform. Roles and permission groups then decide who can reach which agents, tools and workspaces. The on-premises LLM page walks through the deployment pattern layer by layer.

Sources

Frequently asked questions

What is an on-premise AI server?

It is a GPU-equipped server that sits in your own data centre, or in a colocation cage you control, and runs AI models locally. It hosts the language model, the embedding model used for retrieval and often the agent platform, so prompts, documents and answers stay inside your network. Most are standard rack servers with one to eight data-centre GPUs, chosen by how much GPU memory the models and the concurrent users need.

How much does an on-premise AI server cost?

Data-centre GPU servers rarely have public list prices, so budgets start from OEM or partner quotes. As a planning anchor, a production server with four NVIDIA H100 GPUs, enough for a 70B model, has been budgeted at roughly €80,000 to €140,000 per node at mid-2026 pricing. A desk-side NVIDIA DGX Spark for evaluation lists at $4,699. Add power, cooling, support, spare capacity and staff time before you compare the total with cloud API spend.

What GPU do I need for an on-premise LLM server?

Start from GPU memory. Weights need about parameters times bits per weight divided by eight, so a 70B model needs roughly 35 GB at 4-bit and 70 GB at 8-bit, and every concurrent conversation adds KV cache on top. A 48 GB L40S suits pilots with 30B-class models, a 96 GB RTX PRO 6000 or an 80 GB H100 runs a 70B model at 4-bit, and a 141 GB H200 holds a 70B model at FP8 with room for many sessions.

How many users can one AI server support?

It depends on how many people use it at the same moment and how long their conversations are, more than on headcount. Planning tiers assume about 2 to 10 simultaneous conversations for 25 to 100 named users, and 10 to 75 for up to 1,000 users. One 80 to 96 GB GPU holds a 32B model plus ten 8,000-token conversations, while 75 concurrent sessions on a 70B model need around 170 GB or more. Load-test with your own prompts to confirm speed.

Do I need Kubernetes to run an LLM server on premise?

No. A single server can run the serving engine, the vector database and the application containers under Docker or Docker Compose, which is the quickest route for a pilot. Kubernetes earns its place when you run several GPU nodes, need rolling upgrades and failover, or want the NVIDIA GPU Operator to manage drivers, the device plugin and GPU monitoring across a cluster. Many teams start on one host and move to Kubernetes for production.

Should we buy or rent GPU servers for on-prem AI?

Rent first if you do not yet know your concurrency, context lengths and model mix, because a pilot on dedicated single-tenant capacity produces the numbers a purchase needs. Buy when utilisation is steady and high enough that owned hardware beats rented hours, and when custody rules out a provider holding administrative access. Leasing and colocation sit in between: they change how the hardware is paid for and where it is hosted, not who controls it.

Filed under
on-premises AIAI infrastructurelocal AI infrastructurelocal AI hardwarelocal LLMAI platform TCO
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading