AI Infrastructure

Hosting an LLM on Kubernetes: GPU Nodes, Model Storage, Serving and Autoscaling

A vendor-neutral guide to hosting a large language model on Kubernetes: preparing GPU nodes with the NVIDIA GPU Operator, getting weights to the pods, choosing between vLLM's Helm chart, KServe, llm-d and the NIM Operator, scaling on queue depth, and running the whole stack without internet access.

Hosting an LLM on Kubernetes means running an inference server such as vLLM as pods on GPU nodes, with the NVIDIA GPU Operator managing drivers and the device plugin, weights mounted from a volume or image, replicas scaled on request queue depth, and authentication and monitoring placed around the endpoint. The same design works on an air-gapped cluster once images and weights are mirrored inside.

A single GPU server running vLLM under Docker is the quickest way to serve a model on your own hardware. Kubernetes becomes worth its overhead once you have several GPU nodes, more than one model, or a service people depend on and that needs rolling upgrades and failover. This guide walks through each layer in the order you build it. Version numbers were checked against each project’s releases on 2 October 2026, and every project here ships often, so pin versions and re-check before you install.

If you are still choosing hardware, size it first with the GPU requirements method and the on-premise AI server guide. This guide assumes the servers exist and the question is how to run models on them.

The stack, layer by layer

LayerWhat it doesCommon choices (verified October 2026)
GPU node softwareDriver, container runtime hook, device plugin, GPU metricsNVIDIA GPU Operator v26.7
SchedulingKeeps GPU nodes for GPU pods and places replicasNode pools, taints and tolerations, nvidia.com/gpu limits, Dynamic Resource Allocation
Model storageGets weights into the podPersistentVolumeClaim, S3-compatible object storage, OCI images
Inference engineLoads the model and serves an OpenAI-compatible APIvLLM, SGLang, TensorRT-LLM
Serving control planeRollout, routing and replica managementvLLM Helm chart or production stack, KServe, llm-d, NVIDIA NIM Operator, NVIDIA Dynamo
AutoscalingAdds and removes replicasKEDA on Prometheus metrics
EdgeTLS, identity, rate limits, model-aware routingGateway API with the Inference Extension, an identity-aware proxy
ObservabilityMetrics, logs and alertsPrometheus, vLLM’s /metrics, NVIDIA DCGM exporter, Grafana

The inference engine is the part people compare most, and the vLLM vs Ollama comparison covers that choice. On Kubernetes the usual answer is vLLM or SGLang, because both batch many concurrent requests on data-centre GPUs.

GPU nodes: the GPU Operator, node pools and taints

The NVIDIA GPU Operator automates the software a GPU node needs: the driver, the NVIDIA Container Toolkit, the Kubernetes device plugin, GPU Feature Discovery, DCGM-based monitoring with the DCGM exporter, the MIG Manager and a validator. It deploys Node Feature Discovery by default to label GPU nodes. Version 26.7.1 was released on 23 September 2026. If your nodes already carry a driver, install the operator with driver.enabled=false and let it manage the rest.

Keep GPU nodes for GPU work. The Kubernetes documentation on taints and tolerations uses exactly this case: nodes with special hardware should be reserved for pods that need it. Put the GPU nodes in their own pool, taint them, and give only the serving pods a matching toleration. GPUs are requested in the limits section of the container spec, as nvidia.com/gpu:

spec:
  nodeSelector:
    pool: gpu-h100            # your own node-pool label
  tolerations:
    - key: nvidia.com/gpu
      operator: Exists
      effect: NoSchedule
  volumes:
    - name: weights
      persistentVolumeClaim:
        claimName: llama-3-3-70b-weights
        readOnly: true
    - name: shm
      emptyDir:
        medium: Memory
        sizeLimit: 8Gi
  containers:
    - name: vllm
      image: vllm/vllm-openai:<pinned-version>
      command: ["/bin/sh", "-c"]
      args: ["vllm serve /models/llama-3.3-70b-instruct --tensor-parallel-size 2"]
      resources:
        limits:
          nvidia.com/gpu: 2
      volumeMounts:
        - name: weights
          mountPath: /models
        - name: shm
          mountPath: /dev/shm
      readinessProbe:
        httpGet:
          path: /health
          port: 8000

The /dev/shm volume matters: vLLM’s Kubernetes guide notes that tensor-parallel inference needs the host’s shared memory, and its example mounts a memory-backed emptyDir there.

Two newer options are worth knowing about:

  • GPU sharing. Time-slicing lets several pods share a GPU, but NVIDIA’s GPU sharing documentation warns that its replicas have no memory or fault isolation. MIG partitions supported GPUs into isolated instances. Both suit small models such as embedders and rerankers. A large model normally takes whole GPUs.
  • Dynamic Resource Allocation. DRA reached general availability in Kubernetes 1.34. In NVIDIA’s DRA driver, full-GPU and existing MIG allocation are generally available. It needs Kubernetes 1.34.2 or later and GPU driver 580 or newer, and under the GPU Operator a cluster runs either DRA or the device plugin, not both (NVIDIA DRA guide).

Getting model weights to the pods

A 70B model at 16-bit precision is roughly 140 GB of files. How those bytes reach a new pod decides how long a restart or a scale-out takes.

OptionHow it worksFitsWatch for
Download at start-upThe server pulls from Hugging Face into a cache volumeDemosNeeds internet access; slow and repeated cold starts
PersistentVolumeClaimLoad the weights once onto shared storage and mount them read-onlyMost on-premises clustersStorage throughput sets load time
S3-compatible object storagevLLM’s Run:ai Model Streamer streams tensors straight to GPU memory, or KServe’s storage initializer copies them firstClusters with MinIO or Ceph object storageNetwork bandwidth to the nodes
OCI imagePackage the weights as an image in your registryAir-gapped sites, versioned releasesLarge images; pre-pull them to the GPU nodes

For OCI images, Kubernetes now mounts an image directly as a read-only volume. The image volume feature is stable in Kubernetes 1.36, provided your container runtime supports it. KServe offers a similar pattern called Modelcars, which is off by default.

Whatever you choose, version the weights like code: one immutable path or image tag per model revision, so a rollback means changing a reference rather than copying files back.

Choosing the serving layer

The vLLM documentation lists more than a dozen ways to run it on Kubernetes. Five cover most enterprise needs:

OptionWhat it addsChoose it whenState (verified October 2026)
Plain Deployment or vLLM’s Helm chartOne model behind a Service, with probesOne or two models and a small teamMaintained in the vLLM docs
vLLM production stackHelm chart, a request router (round-robin or session-based, prefix-aware routing in progress), Prometheus and Grafana, LMCache for KV-cache offloadingSeveral vLLM models that need one entry point and dashboardsChart 0.1.13, 29 September 2026
KServeInferenceService and the newer LLMInferenceService, built on llm-d, with a Hugging Face runtime that uses vLLM by defaultYou already run KServe for other models or want one serving APIv0.21.0, 25 September 2026; CNCF incubating
llm-dPrefix-cache and load-aware routing through the Gateway API Inference Extension, prefill/decode disaggregation, wide expert parallelism, tiered KV offloadingLarge models, long shared prompts, high concurrencyv0.10.0, 29 September 2026, pinning vLLM 0.30.0; CNCF sandbox
NVIDIA NIM OperatorNIMCache to download and store models, NIMService to run NVIDIA NIM microservicesYou standardise on NVIDIA’s packaged containersv3.1.2, 12 August 2026

The licence is different for the last row. NVIDIA states that using NIM in production requires an NVIDIA AI Enterprise licence, while research, development and experimentation are free on up to 16 GPUs (NIM documentation). NVIDIA Dynamo, a datacenter-scale inference stack over vLLM, SGLang or TensorRT-LLM, also installs on Kubernetes through its own operator and Helm chart.

Multi-GPU and multi-node models

Inside one node, tensor parallelism splits each layer across GPUs: set --tensor-parallel-size to the number of GPUs the pod requests. When a model needs more GPUs than one node holds, vLLM uses Ray across nodes by default, and the pods have to start, fail and restart as a group. LeaderWorkerSet is the Kubernetes API built for this: it treats a group of pods as one unit of replication for multi-host inference, and both vLLM and KServe document it for multi-node serving. If one node of the group fails, the whole replica goes down, so plan spare capacity for a full replica.

Autoscaling on queue depth

GPU utilisation is a poor scaling signal for LLM serving, because a busy GPU can still have spare batch capacity. vLLM exposes better signals on its /metrics endpoint (metrics reference):

  • vllm:num_requests_waiting: requests queued for a slot.
  • vllm:num_requests_running: requests in the current batch.
  • vllm:kv_cache_usage_perc: KV-cache occupancy, where 1 means full.
  • vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds: the latency users feel.

KEDA scales a Deployment on any Prometheus query. Its Prometheus scaler takes a server address, a query and a threshold:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llama-70b
spec:
  scaleTargetRef:
    name: llama-70b-vllm
  minReplicaCount: 1
  maxReplicaCount: 4
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring.svc:9090
        query: sum(vllm:num_requests_waiting)
        threshold: "8"

Three cautions apply. A new replica is not useful until its weights are loaded, which can take minutes for a large model, so scale before latency degrades rather than after. KEDA can scale to zero, which saves GPUs for rarely used models but turns the first request into a cold start. And on premises the GPU pool does not grow on demand: autoscaling moves fixed capacity between models, so set maximum replica counts that add up to the GPUs you actually have.

Ingress, authentication and observability

vLLM’s built-in --api-key check is a static credential set at launch, and its CLI reference warns that it only covers endpoints under the /v1, /v2 and /inference prefixes. Treat the pod as an internal service:

  • Put identity in front. Route traffic through a gateway or an identity-aware proxy such as oauth2-proxy, a CNCF sandbox project that authenticates against OpenID Connect providers, so requests carry a user or service identity.
  • Route by model. The Gateway API Inference Extension, an official Kubernetes project, adds an InferencePool resource that lets gateways such as Envoy Gateway, kgateway, Istio or NGINX Gateway Fabric pick a model-server pod using its load and cache state. Its v1 API was released in September 2025.
  • Restrict the network. Apply a NetworkPolicy so only the gateway reaches the serving pods, keep /metrics on the cluster network, and deny egress from the serving namespace.

For monitoring, scrape vLLM’s metrics alongside the DCGM exporter, which publishes GPU telemetry such as DCGM_FI_DEV_GPU_UTIL and DCGM_FI_DEV_FB_USED in Prometheus format. Alert on queue length and time to first token first; GPU temperature and memory errors come next.

Running it air-gapped

Every layer above assumes registries and model hubs are reachable. Inside an air gap, bring them with you:

  1. Mirror the images. NVIDIA’s air-gapped install guide requires all GPU Operator images in a local registry reachable from every node. Do the same for vLLM, the router and any operator you use.
  2. Plan the driver. The driver container needs a local package repository during installation, unless your nodes can run NVIDIA’s precompiled driver containers.
  3. Copy the weights. Transfer them once, verify each file’s SHA-256 digest against the source, and store them on a PVC, internal object storage or an image in your registry.
  4. Serve from a local path. vllm serve /models/<model-dir> loads a model from a directory, and HF_HUB_OFFLINE=1 stops the Hugging Face client from making network calls.
  5. Switch off telemetry. vLLM collects anonymous usage data by default; set VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1.
  6. Prove it. Start the whole stack with the uplink disconnected and watch for failed outbound connections. The air-gapped LLM pattern covers the wider architecture.

Deployment checklist

  1. GPU Operator installed and its validator passing on every GPU node.
  2. GPU nodes in their own pool, tainted, with only serving pods tolerating the taint.
  3. GPU requests in limits, and /dev/shm mounted for tensor-parallel pods.
  4. Weights on immutable, versioned storage, with load time measured.
  5. Readiness probe on /health, with an initial delay longer than the model load time.
  6. Image tags and Helm chart versions pinned, never latest.
  7. Authentication at the gateway; vLLM’s own key treated as a backstop.
  8. NetworkPolicy limiting ingress to the gateway and blocking egress.
  9. Prometheus scraping vLLM and DCGM, with alerts on queue length and time to first token.
  10. Autoscaling bounds that match the GPUs you own.
  11. A rolling-upgrade plan with spare capacity for one replica.
  12. An offline start-up test if the cluster is, or may become, air-gapped.

For the decisions around this list (which model, how many users, which deployment mode), the on-premises LLM overview and the self-hosted LLM page set out the options.

How VDF AI fits

VDF AI’s self-hosting documentation names Kubernetes as the recommended production target. Each package’s install guide carries the manifests or a Helm chart, and the images run on any conformant cluster, managed or self-managed. For air-gapped sites the documented path is the one described above: pull the images on a connected staging host, push them to an internal registry such as Harbor or Artifactory, and deploy from there.

The packages log JSON to stdout, expose HTTP health endpoints and publish Prometheus-compatible metrics, so they join the same monitoring as the model servers. VDF AI Router is a Docker-packaged service that runs on VMs, Kubernetes or bare metal next to the GPUs, giving applications one endpoint in front of the model deployments; in air-gap mode it routes to local models only.

Sources

Frequently asked questions

Can you host an LLM on Kubernetes?

Yes, and it is a common way to run models on your own GPUs. The inference server, usually vLLM, runs as a pod on a GPU node and exposes an OpenAI-compatible API. Kubernetes adds what a single server lacks: scheduling across several GPU nodes, rolling upgrades, health checks, restarts and autoscaling. The extra work sits underneath, in preparing GPU nodes with drivers and the device plugin and in getting tens or hundreds of gigabytes of weights to each pod quickly.

What is the best way to deploy vLLM on Kubernetes?

For one model on one GPU node, a plain Deployment or vLLM's Helm chart is enough and the easiest to debug. For several models, routing and Prometheus dashboards, vLLM's production stack adds a router and observability. For prefix-cache-aware routing, disaggregated prefill and decode or wide expert parallelism, look at llm-d or KServe's LLMInferenceService, which is built on llm-d. Choose by the problem you have today, since each step up adds components to operate.

Do I need the NVIDIA GPU Operator?

Not strictly, but most clusters use it. The GPU Operator installs and manages the driver, the NVIDIA Container Toolkit, the Kubernetes device plugin, GPU Feature Discovery and DCGM-based monitoring as containers, so GPU nodes are configured the same way every time. If your nodes already have drivers installed, the operator can manage the remaining components with its driver installation switched off. Without it, you install and upgrade each component on every node yourself.

How do you autoscale an LLM on Kubernetes?

Scale on the length of the request queue, not on GPU utilisation. vLLM exports the number of waiting requests and KV-cache usage as Prometheus metrics, and KEDA's Prometheus scaler can add replicas when the queue stays above a threshold. Expect a new replica to take minutes to become ready, because the pod has to load its weights into GPU memory, and on premises the number of GPU nodes is fixed, so autoscaling redistributes capacity you already own.

Can an LLM run on an air-gapped Kubernetes cluster?

Yes. Mirror every container image, including the GPU Operator's, into a registry inside the network, provide a local package repository for driver builds or use precompiled driver containers, and copy the model weights onto internal storage. Then point the server at the local model path, set the Hugging Face client to offline mode, turn off vLLM's anonymous usage statistics and block egress from the serving namespace. Test that the cluster starts cleanly with the uplink unplugged before you rely on it.

Filed under
KubernetesvLLMlocal LLMAI infrastructureair-gapped AIon-premises AI
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading