Hosting an LLM on Kubernetes means running an inference server such as vLLM as pods on GPU nodes, with the NVIDIA GPU Operator managing drivers and the device plugin, weights mounted from a volume or image, replicas scaled on request queue depth, and authentication and monitoring placed around the endpoint. The same design works on an air-gapped cluster once images and weights are mirrored inside.
A single GPU server running vLLM under Docker is the quickest way to serve a model on your own hardware. Kubernetes becomes worth its overhead once you have several GPU nodes, more than one model, or a service people depend on and that needs rolling upgrades and failover. This guide walks through each layer in the order you build it. Version numbers were checked against each project’s releases on 2 October 2026, and every project here ships often, so pin versions and re-check before you install.
If you are still choosing hardware, size it first with the GPU requirements method and the on-premise AI server guide. This guide assumes the servers exist and the question is how to run models on them.
The stack, layer by layer
| Layer | What it does | Common choices (verified October 2026) |
|---|---|---|
| GPU node software | Driver, container runtime hook, device plugin, GPU metrics | NVIDIA GPU Operator v26.7 |
| Scheduling | Keeps GPU nodes for GPU pods and places replicas | Node pools, taints and tolerations, nvidia.com/gpu limits, Dynamic Resource Allocation |
| Model storage | Gets weights into the pod | PersistentVolumeClaim, S3-compatible object storage, OCI images |
| Inference engine | Loads the model and serves an OpenAI-compatible API | vLLM, SGLang, TensorRT-LLM |
| Serving control plane | Rollout, routing and replica management | vLLM Helm chart or production stack, KServe, llm-d, NVIDIA NIM Operator, NVIDIA Dynamo |
| Autoscaling | Adds and removes replicas | KEDA on Prometheus metrics |
| Edge | TLS, identity, rate limits, model-aware routing | Gateway API with the Inference Extension, an identity-aware proxy |
| Observability | Metrics, logs and alerts | Prometheus, vLLM’s /metrics, NVIDIA DCGM exporter, Grafana |
The inference engine is the part people compare most, and the vLLM vs Ollama comparison covers that choice. On Kubernetes the usual answer is vLLM or SGLang, because both batch many concurrent requests on data-centre GPUs.
GPU nodes: the GPU Operator, node pools and taints
The NVIDIA GPU Operator automates the software a GPU node needs: the driver, the NVIDIA Container Toolkit, the Kubernetes device plugin, GPU Feature Discovery, DCGM-based monitoring with the DCGM exporter, the MIG Manager and a validator. It deploys Node Feature Discovery by default to label GPU nodes. Version 26.7.1 was released on 23 September 2026. If your nodes already carry a driver, install the operator with driver.enabled=false and let it manage the rest.
Keep GPU nodes for GPU work. The Kubernetes documentation on taints and tolerations uses exactly this case: nodes with special hardware should be reserved for pods that need it. Put the GPU nodes in their own pool, taint them, and give only the serving pods a matching toleration. GPUs are requested in the limits section of the container spec, as nvidia.com/gpu:
spec:
nodeSelector:
pool: gpu-h100 # your own node-pool label
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
volumes:
- name: weights
persistentVolumeClaim:
claimName: llama-3-3-70b-weights
readOnly: true
- name: shm
emptyDir:
medium: Memory
sizeLimit: 8Gi
containers:
- name: vllm
image: vllm/vllm-openai:<pinned-version>
command: ["/bin/sh", "-c"]
args: ["vllm serve /models/llama-3.3-70b-instruct --tensor-parallel-size 2"]
resources:
limits:
nvidia.com/gpu: 2
volumeMounts:
- name: weights
mountPath: /models
- name: shm
mountPath: /dev/shm
readinessProbe:
httpGet:
path: /health
port: 8000
The /dev/shm volume matters: vLLM’s Kubernetes guide notes that tensor-parallel inference needs the host’s shared memory, and its example mounts a memory-backed emptyDir there.
Two newer options are worth knowing about:
- GPU sharing. Time-slicing lets several pods share a GPU, but NVIDIA’s GPU sharing documentation warns that its replicas have no memory or fault isolation. MIG partitions supported GPUs into isolated instances. Both suit small models such as embedders and rerankers. A large model normally takes whole GPUs.
- Dynamic Resource Allocation. DRA reached general availability in Kubernetes 1.34. In NVIDIA’s DRA driver, full-GPU and existing MIG allocation are generally available. It needs Kubernetes 1.34.2 or later and GPU driver 580 or newer, and under the GPU Operator a cluster runs either DRA or the device plugin, not both (NVIDIA DRA guide).
Getting model weights to the pods
A 70B model at 16-bit precision is roughly 140 GB of files. How those bytes reach a new pod decides how long a restart or a scale-out takes.
| Option | How it works | Fits | Watch for |
|---|---|---|---|
| Download at start-up | The server pulls from Hugging Face into a cache volume | Demos | Needs internet access; slow and repeated cold starts |
| PersistentVolumeClaim | Load the weights once onto shared storage and mount them read-only | Most on-premises clusters | Storage throughput sets load time |
| S3-compatible object storage | vLLM’s Run:ai Model Streamer streams tensors straight to GPU memory, or KServe’s storage initializer copies them first | Clusters with MinIO or Ceph object storage | Network bandwidth to the nodes |
| OCI image | Package the weights as an image in your registry | Air-gapped sites, versioned releases | Large images; pre-pull them to the GPU nodes |
For OCI images, Kubernetes now mounts an image directly as a read-only volume. The image volume feature is stable in Kubernetes 1.36, provided your container runtime supports it. KServe offers a similar pattern called Modelcars, which is off by default.
Whatever you choose, version the weights like code: one immutable path or image tag per model revision, so a rollback means changing a reference rather than copying files back.
Choosing the serving layer
The vLLM documentation lists more than a dozen ways to run it on Kubernetes. Five cover most enterprise needs:
| Option | What it adds | Choose it when | State (verified October 2026) |
|---|---|---|---|
| Plain Deployment or vLLM’s Helm chart | One model behind a Service, with probes | One or two models and a small team | Maintained in the vLLM docs |
| vLLM production stack | Helm chart, a request router (round-robin or session-based, prefix-aware routing in progress), Prometheus and Grafana, LMCache for KV-cache offloading | Several vLLM models that need one entry point and dashboards | Chart 0.1.13, 29 September 2026 |
| KServe | InferenceService and the newer LLMInferenceService, built on llm-d, with a Hugging Face runtime that uses vLLM by default | You already run KServe for other models or want one serving API | v0.21.0, 25 September 2026; CNCF incubating |
| llm-d | Prefix-cache and load-aware routing through the Gateway API Inference Extension, prefill/decode disaggregation, wide expert parallelism, tiered KV offloading | Large models, long shared prompts, high concurrency | v0.10.0, 29 September 2026, pinning vLLM 0.30.0; CNCF sandbox |
| NVIDIA NIM Operator | NIMCache to download and store models, NIMService to run NVIDIA NIM microservices | You standardise on NVIDIA’s packaged containers | v3.1.2, 12 August 2026 |
The licence is different for the last row. NVIDIA states that using NIM in production requires an NVIDIA AI Enterprise licence, while research, development and experimentation are free on up to 16 GPUs (NIM documentation). NVIDIA Dynamo, a datacenter-scale inference stack over vLLM, SGLang or TensorRT-LLM, also installs on Kubernetes through its own operator and Helm chart.
Multi-GPU and multi-node models
Inside one node, tensor parallelism splits each layer across GPUs: set --tensor-parallel-size to the number of GPUs the pod requests. When a model needs more GPUs than one node holds, vLLM uses Ray across nodes by default, and the pods have to start, fail and restart as a group. LeaderWorkerSet is the Kubernetes API built for this: it treats a group of pods as one unit of replication for multi-host inference, and both vLLM and KServe document it for multi-node serving. If one node of the group fails, the whole replica goes down, so plan spare capacity for a full replica.
Autoscaling on queue depth
GPU utilisation is a poor scaling signal for LLM serving, because a busy GPU can still have spare batch capacity. vLLM exposes better signals on its /metrics endpoint (metrics reference):
vllm:num_requests_waiting: requests queued for a slot.vllm:num_requests_running: requests in the current batch.vllm:kv_cache_usage_perc: KV-cache occupancy, where 1 means full.vllm:time_to_first_token_secondsandvllm:inter_token_latency_seconds: the latency users feel.
KEDA scales a Deployment on any Prometheus query. Its Prometheus scaler takes a server address, a query and a threshold:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: llama-70b
spec:
scaleTargetRef:
name: llama-70b-vllm
minReplicaCount: 1
maxReplicaCount: 4
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring.svc:9090
query: sum(vllm:num_requests_waiting)
threshold: "8"
Three cautions apply. A new replica is not useful until its weights are loaded, which can take minutes for a large model, so scale before latency degrades rather than after. KEDA can scale to zero, which saves GPUs for rarely used models but turns the first request into a cold start. And on premises the GPU pool does not grow on demand: autoscaling moves fixed capacity between models, so set maximum replica counts that add up to the GPUs you actually have.
Ingress, authentication and observability
vLLM’s built-in --api-key check is a static credential set at launch, and its CLI reference warns that it only covers endpoints under the /v1, /v2 and /inference prefixes. Treat the pod as an internal service:
- Put identity in front. Route traffic through a gateway or an identity-aware proxy such as oauth2-proxy, a CNCF sandbox project that authenticates against OpenID Connect providers, so requests carry a user or service identity.
- Route by model. The Gateway API Inference Extension, an official Kubernetes project, adds an
InferencePoolresource that lets gateways such as Envoy Gateway, kgateway, Istio or NGINX Gateway Fabric pick a model-server pod using its load and cache state. Its v1 API was released in September 2025. - Restrict the network. Apply a NetworkPolicy so only the gateway reaches the serving pods, keep
/metricson the cluster network, and deny egress from the serving namespace.
For monitoring, scrape vLLM’s metrics alongside the DCGM exporter, which publishes GPU telemetry such as DCGM_FI_DEV_GPU_UTIL and DCGM_FI_DEV_FB_USED in Prometheus format. Alert on queue length and time to first token first; GPU temperature and memory errors come next.
Running it air-gapped
Every layer above assumes registries and model hubs are reachable. Inside an air gap, bring them with you:
- Mirror the images. NVIDIA’s air-gapped install guide requires all GPU Operator images in a local registry reachable from every node. Do the same for vLLM, the router and any operator you use.
- Plan the driver. The driver container needs a local package repository during installation, unless your nodes can run NVIDIA’s precompiled driver containers.
- Copy the weights. Transfer them once, verify each file’s SHA-256 digest against the source, and store them on a PVC, internal object storage or an image in your registry.
- Serve from a local path.
vllm serve /models/<model-dir>loads a model from a directory, andHF_HUB_OFFLINE=1stops the Hugging Face client from making network calls. - Switch off telemetry. vLLM collects anonymous usage data by default; set
VLLM_NO_USAGE_STATS=1orDO_NOT_TRACK=1. - Prove it. Start the whole stack with the uplink disconnected and watch for failed outbound connections. The air-gapped LLM pattern covers the wider architecture.
Deployment checklist
- GPU Operator installed and its validator passing on every GPU node.
- GPU nodes in their own pool, tainted, with only serving pods tolerating the taint.
- GPU requests in
limits, and/dev/shmmounted for tensor-parallel pods. - Weights on immutable, versioned storage, with load time measured.
- Readiness probe on
/health, with an initial delay longer than the model load time. - Image tags and Helm chart versions pinned, never
latest. - Authentication at the gateway; vLLM’s own key treated as a backstop.
- NetworkPolicy limiting ingress to the gateway and blocking egress.
- Prometheus scraping vLLM and DCGM, with alerts on queue length and time to first token.
- Autoscaling bounds that match the GPUs you own.
- A rolling-upgrade plan with spare capacity for one replica.
- An offline start-up test if the cluster is, or may become, air-gapped.
For the decisions around this list (which model, how many users, which deployment mode), the on-premises LLM overview and the self-hosted LLM page set out the options.
How VDF AI fits
VDF AI’s self-hosting documentation names Kubernetes as the recommended production target. Each package’s install guide carries the manifests or a Helm chart, and the images run on any conformant cluster, managed or self-managed. For air-gapped sites the documented path is the one described above: pull the images on a connected staging host, push them to an internal registry such as Harbor or Artifactory, and deploy from there.
The packages log JSON to stdout, expose HTTP health endpoints and publish Prometheus-compatible metrics, so they join the same monitoring as the model servers. VDF AI Router is a Docker-packaged service that runs on VMs, Kubernetes or bare metal next to the GPUs, giving applications one endpoint in front of the model deployments; in air-gap mode it routes to local models only.
Sources
- NVIDIA GPU Operator overview, release notes and air-gapped installation
- NVIDIA GPU sharing: time-slicing and MIG
- NVIDIA DRA driver for GPUs
- Kubernetes: taints and tolerations, scheduling GPUs and image volumes
- Kubernetes v1.34: DRA graduates to GA
- vLLM on Kubernetes, Helm chart, LWS, metrics, serve CLI, usage statistics and Run:ai Model Streamer
- vLLM production stack
- KServe LLMInferenceService, storage options and CNCF project page
- llm-d repository and CNCF project page
- Gateway API Inference Extension
- LeaderWorkerSet
- NVIDIA NIM Operator and NIM licensing
- NVIDIA Dynamo
- KEDA Prometheus scaler and KEDA concepts
- NVIDIA DCGM exporter
- oauth2-proxy
- Hugging Face Hub environment variables