AI Infrastructure

Best Local LLM by VRAM: 8, 12, 16, 24 and 32 GB (2026)

Which open-weight model to run on an 8, 12, 16, 24 or 32 GB graphics card in October 2026, with the real size of each 4-bit download, the context window that still fits beside it, the licence and what each model is good at. Every fit is worked out with the same arithmetic as our LLM VRAM calculator.

The best local LLM for your VRAM is the strongest open-weight model whose 4-bit file plus KV cache fits in about 90% of the card. In October 2026 that means Qwen3.5-9B or Gemma 4 E4B on 8 GB, Gemma 4 12B on 12 GB, gpt-oss-20b on 16 GB, Qwen3.6-35B-A3B or Gemma 4 26B A4B on 24 GB, and Qwen3.6-35B-A3B with its full 256K window on a 32 GB RTX 5090.

Most “which model fits my GPU” answers stop at the weights. That is half the job. Every token you keep in the conversation also occupies memory, and two models of the same size can differ fivefold in how much. This guide picks models tier by tier and shows the context each one leaves you, so you can tell a model that loads from a model you can actually work with.

Every model, licence and parameter count below was checked on the vendor’s Hugging Face card, and every file size on the published GGUF repository, on 6 October 2026. GPU memory figures come from NVIDIA’s comparison page.

Quick picks by VRAM tier

VRAMExample cards (NVIDIA spec)Best pickAlso strongContext left for one chat
8 GBRTX 5060, RTX 4060, RTX 3070Qwen3.5-9B (Q4_K_M)Gemma 4 E4B, Qwen3.5-4B≈46K tokens
12 GBRTX 5070, RTX 4070, RTX 3060 12 GBGemma 4 12B (Q4_0)Qwen3.5-9B (Q6_K)≈198K tokens
16 GBRTX 5080, RTX 5070 Ti, RTX 5060 Ti 16 GB, RTX 4080gpt-oss-20b (MXFP4)Gemma 4 12B (Q8_0), Qwen3.5-9B (Q8_0)≈93K tokens
24 GBRTX 4090, RTX 3090, RTX 3090 TiQwen3.6-35B-A3B (Q4_K_M)Gemma 4 26B A4B, Qwen3.8-27B, GLM-4.7-Flash≈54K tokens
32 GBRTX 5090Qwen3.6-35B-A3B (Q4_K_M)Gemma 4 31B, Qwen3.8-27B, GLM-4.7-Flash256K (the model’s limit)

Context figures assume one session and a 16-bit KV cache. If you write code with these models, our local coding model comparison covers coder-specific checkpoints.

How the fit is calculated

The method is the one behind our VRAM and KV cache calculator and the GPU sizing guide:

  1. Usable memory is 90% of the card, leaving room for the runtime, activations and the desktop.
  2. Weights are the size of the file you actually load. We use the GGUF published by the llama.cpp project’s ggml-org account, or by the model vendor where it ships its own (Mistral and IBM), or LM Studio’s community build for the small Qwen3.5 models.
  3. KV cache per token is 2 × attention layers × KV heads × head dimension × 2 bytes at 16-bit. Sliding-window layers stop growing at their window. Hybrid models such as Qwen3.5 and Nemotron also hold a small fixed state per session, which we add separately.
  4. Context that fits is the memory left after steps 2 and 3, divided by the per-token cost, capped at the model’s own limit.

Real files rarely match the textbook arithmetic. Embeddings, output layers and vision encoders are often kept at higher precision, and different builders make different choices. Qwen3.6-27B at Q4_K_M is 19.1 GB from ggml-org but 17 GB in Ollama’s library, so check the file you download. The tables assume text only; loading a vision projector adds 0.6 to 1.2 GB.

Worked example: Gemma 4 31B on a 24 GB card

  • Usable: 24 × 0.9 = 21.6 GB.
  • Weights: the Q4_0 file is 17.99 GB.
  • Sliding-window layers: 50 layers × 2 × 16 KV heads × 256 × 2 bytes = 819,200 bytes per token, but only for the last 1,024 tokens, so a fixed 0.84 GB.
  • Global layers: 10 × 2 × 4 KV heads × 512 × 2 bytes = 81,920 bytes per token, for the whole conversation.
  • Room for context: (21.6 − 17.99 − 0.84) GB ÷ 81,920 bytes ≈ 33,800 tokens. An 8-bit KV cache stretches that to about 77,900.

The calculator’s Gemma 4 31B preset at GGUF Q4_K_M (4.89 bits per weight, 19.1 GB) returns about 20,000 tokens for the same card. The gap is the file: Q4_K_M stores more bits per weight than this Q4_0 build, and that difference comes straight out of your context.

Best LLM for 8 GB VRAM

Cards in this tier include the RTX 5060, the 8 GB RTX 5060 Ti, the RTX 4060 and 4060 Ti 8 GB, and the RTX 3070. Usable memory is about 7.2 GB.

ModelParametersQuantizationWeightsContext that fitsLicenceGood at
Qwen3.5-9B9.7B dense (with vision encoder)Q4_K_M5.6 GB≈46KApache 2.0Best quality here; Qwen reports 66.1 on BFCL-V4 and 63.0 on AA-LCR
Gemma 4 E4B8B with embeddings, 4.5B effectiveQ4_04.6 GB≈90KApache 2.0Text, image and audio input
Qwen3.5-4B4.7B denseQ4_K_M2.7 GB≈135KApache 2.0Long documents on a small card
Ministral 3 8B8.9B (with vision encoder)Q4_K_M5.2 GB≈14KApache 2.0Native function calling and JSON output
Granite 4.2 8B8.8B denseQ4_K_M5.4 GB≈11KApache 2.0Tool calling in 12 tested languages

The spread in the context column is architecture, not size. Qwen3.5 keeps a growing cache in only 8 of its 32 layers, 32,768 bytes per token, while Granite 4.2 8B caches all 40 layers at 163,840 bytes per token. Gemma 4 E4B’s figure is conservative: it assumes every layer keeps its own cache, although the last 18 layers are designed to share earlier ones. Qwen3.5 models think before answering by default, so long reasoning traces eat into the same budget.

Best LLM for 12 GB VRAM

Examples are the RTX 5070, the RTX 4070, 4070 SUPER and 4070 Ti, the 12 GB RTX 3060 and the RTX 3080 Ti. Usable memory is about 10.8 GB.

ModelParametersQuantizationWeightsContext that fitsLicenceGood at
Gemma 4 12B12.0B denseQ4_07.2 GB≈198KApache 2.0Most context per gigabyte; text, image and audio; Google reports 69.0% on Tau2
Qwen3.5-9B9.7B denseQ6_K7.4 GB≈103KApache 2.0Higher-precision weights with a long window
Gemma 4 E4B8B with embeddingsQ8_08.0 GB≈95KApache 2.0Near-lossless on-device model
Ministral 3 14B13.9B (with vision encoder)Q4_K_M8.2 GB≈16KApache 2.0Mistral calls it comparable to its Mistral Small 3.2 24B
Granite 4.2 8B8.8B denseQ4_K_M5.4 GB≈33KApache 2.0Enterprise tool use, 128K native window

Gemma 4 12B is the standout. Only 8 of its 48 layers attend over the whole conversation, each with a single 512-dimension KV head, so a 200K-token chat costs about 3.6 GB of cache, sliding-window layers included. gpt-oss-20b does not fit: its 12.1 GB file exceeds the 10.8 GB budget. llama.cpp can keep some mixture-of-experts weights in system RAM with --n-cpu-moe, at a cost in speed.

Best LLM for 16 GB VRAM

This is the most common enthusiast and workstation tier: RTX 5080, RTX 5070 Ti, the 16 GB RTX 5060 Ti, RTX 4080 and 4080 SUPER, RTX 4070 Ti SUPER and the 16 GB RTX 4060 Ti. Usable memory is about 14.4 GB.

ModelParametersQuantizationWeightsContext that fitsLicenceGood at
gpt-oss-20b21B total, 3.6B activeMXFP4 (as trained)12.1 GB≈93KApache 2.0 plus usage policyAdjustable reasoning effort, function calling, structured outputs
Gemma 4 12B12.0B denseQ8_012.7 GB≈85KApache 2.08-bit quality with a long window; or Q4_0 at the full 256K
Qwen3.5-9B9.7B denseQ8_09.5 GB≈147KApache 2.0Long context at 8-bit; image input
Ministral 3 14B13.9BQ4_K_M8.2 GB≈38KApache 2.0Function calling and JSON in 11 listed languages
Granite 4.2 8B8.8B denseQ4_K_M5.4 GB≈55KApache 2.0Thinking modes you can switch off per request

gpt-oss-20b was trained with MXFP4 weights, so there is no separate quantization step to second-guess; OpenAI’s card says it runs within 16 GB. OpenAI says the model must be used with its harmony response format; the chat template in the repository applies it, and llama.cpp lists a native handler for it.

What does not fit: Mistral Small 3.2 24B at Q4_K_M is 14.3 GB and leaves a few hundred tokens, Gemma 4 26B A4B’s 14.6 GB Q4_0 file is just over budget, and the 27B Qwen models need 19 GB. One trap applies to this whole tier: Ollama’s documentation says it defaults to a 4K context on GPUs with less than 24 GiB, so set OLLAMA_CONTEXT_LENGTH to the window you sized for.

Best LLM for 24 GB VRAM

The RTX 4090, RTX 3090 and RTX 3090 Ti sit here. Usable memory is about 21.6 GB, and the choice widens to 24B–36B models.

ModelParametersQuantizationWeightsContext that fitsLicenceGood at
Qwen3.6-35B-A3B36B total, 3B activeQ4_K_M20.4 GB≈54KApache 2.0Agent work; Qwen reports 62.8 on MCP-Atlas
Gemma 4 26B A4B25.8B total, 3.8B activeQ4_014.6 GB256K (model limit)Apache 2.0The full window on one card; long documents
Qwen3.8-27B27.8B denseQ4_K_M19.0 GB≈38KApache 2.0Newest Qwen dense model; 79.5 on IFBench per Qwen
Gemma 4 31B31.3B denseQ4_018.0 GB≈34KApache 2.0Strongest Gemma: 85.2% MMLU Pro, 76.9% Tau2 per Google
GLM-4.7-Flash31.2B total, about 3B activeQ4_K18.2 GB≈62KMITMulti-turn tool use: 79.5 on τ²-Bench per Z.ai
Mistral Small 3.2 24B24.0B denseQ4_K_M14.3 GB≈44KApache 2.0Image input; robust function calling
Nemotron 3.5 Lightning 30B-A3B31.6B total, 3B activeQ4_018.9 GB≈430KOpenMDW-1.1Very long inputs: only 6 of 52 layers are attention

Two notes on the dense Qwen models. Qwen’s Qwen3.6 card advises keeping at least 128K of context to preserve thinking quality, which a 24 GB card cannot give the 27B at 16-bit. An 8-bit KV cache roughly doubles the figure; turning thinking off for simple requests is the other lever. Qwen3.6-27B behaves like Qwen3.8-27B here, with a 19.1 GB file and about 36K tokens.

Best LLM for 32 GB VRAM (RTX 5090)

The GeForce RTX 5090 is the only 32 GB card in NVIDIA’s consumer line. NVIDIA lists 32 GB of GDDR7 on a 512-bit interface, 1,792 GB/s of memory bandwidth, 575 W of total graphics power and a 1,000 W system power requirement (verified October 2026). Usable memory is about 28.8 GB.

ModelParametersQuantizationWeightsContext that fitsLicenceGood at
Qwen3.6-35B-A3B36B total, 3B activeQ4_K_M20.4 GB256K (model limit)Apache 2.0Full window plus room for parallel sessions
Qwen3.8-27B27.8B denseQ4_K_M19.0 GB≈148KApache 2.0Meets Qwen’s 128K thinking advice
Gemma 4 31B31.3B denseQ4_018.0 GB≈122KApache 2.0Best dense Gemma with a long window
Gemma 4 26B A4B25.8B total, 3.8B activeQ8_026.9 GB≈84KApache 2.08-bit weights for quality-sensitive work
GLM-4.7-Flash31.2B totalQ4_K18.2 GB≈195KMITLong multi-turn tool sessions
Mistral Small 3.2 24B24.0B denseQ6_K19.4 GB≈58KApache 2.0Higher-precision Mistral
Nemotron 3.5 Lightning 30B-A3B31.6B totalQ4_018.9 GB1M (model limit)OpenMDW-1.1Million-token inputs on one card

The extra 8 GB over a 24 GB card mostly buys context and concurrency rather than bigger models. 8-bit 27B files (28.6 GB) load but leave almost nothing for the cache, and 70B dense models are out of reach: Ollama’s Llama 3.3 70B at q4_K_M is 43 GB. For that class, look at 48 GB-plus GPUs or the unified-memory machines in our DGX Spark, Mac Studio and Ryzen AI Max comparison.

When the model you want does not fit

Five levers, in the order we would pull them:

  1. Quantize the KV cache. llama.cpp accepts --cache-type-k q8_0 (and the matching V flag), Ollama exposes OLLAMA_KV_CACHE_TYPE=q8_0, and vLLM supports an FP8 KV cache. Ollama’s FAQ describes q8_0 as roughly half the memory of f16 with a very small loss in precision.
  2. Shorten the window. Size context for the task, not the model card’s maximum.
  3. Offload experts. For mixture-of-experts models, llama.cpp’s --n-cpu-moe N keeps the expert weights of the first N layers in system RAM. It runs, slower.
  4. Let llama.cpp fit it. Current llama-server builds adjust unset arguments to device memory by default (--fit), keeping a 1,024 MiB margin per device unless you change it.
  5. Drop a quantization level. Weight precision is the last lever, because quality loss varies by model and task. Our quantization trade-off guide explains how to test it.

If several people will share one GPU, every concurrent session needs its own cache, and single-user runners queue requests by default. Our vLLM, Ollama and llama.cpp comparison covers the batching servers built for that. For the same picks on Apple hardware, see the Mac memory-tier guide; for retrieval and agent workloads, see our picks for local RAG and tool calling. To pick the card itself, the GPU buyer’s guide compares them by budget, and the Gemma 4 setup guide walks through one family end to end.

How VDF AI fits

A single GPU under a desk serves one person well. When a team needs the same local models, VDF AI Router registers Ollama and custom on-premises deployments alongside any cloud models your policy allows, and its air-gap mode restricts routing to local models only. Local runtimes are probed continuously, so a request can move to another permitted model when one endpoint is down.

That lets you keep a small model on modest hardware for routine requests and send heavier work to a larger model on a server, behind one API. The local LLM overview covers the wider deployment picture.

Sources

Model cards, file listings and specifications verified 6 October 2026.


Moving a local model from one GPU to a shared service? See how VDF AI Router routes across on-premises models, or book a demo.

Frequently asked questions

What is the best LLM for 16GB VRAM?

For most people it is gpt-oss-20b. OpenAI ships it natively in MXFP4, the llama.cpp project's GGUF build is 12.1 GB, and a 16 GB card still has room for roughly 93,000 tokens of 16-bit KV cache beside it. Gemma 4 12B at 8-bit and Qwen3.5-9B at 8-bit are the alternatives if you want image input or a longer window. All three carry Apache 2.0 licences. Larger 24B to 27B models do not fit at 4-bit with useful context.

What is the best LLM for 24GB VRAM?

Qwen3.6-35B-A3B at Q4_K_M is the strongest all-round choice: the 20.4 GB file leaves room for about 54,000 tokens, and only 3 billion parameters are active per token. Choose Gemma 4 26B A4B when you need the full 256K window, which its 14.6 GB Q4_0 file allows, or Gemma 4 31B and Qwen3.8-27B when you prefer a dense model and can live with 34,000 to 38,000 tokens of context.

Can an 8GB GPU run a useful local LLM?

Yes, if you pick a model built to keep its KV cache small. Qwen3.5-9B at Q4_K_M takes 5.6 GB and still leaves about 46,000 tokens of context, and Gemma 4 E4B takes 4.6 GB with about 90,000 tokens to spare. Older dense 8B designs with a cache in every layer, such as Granite 4.2 8B or Ministral 3 8B, fit the weights but leave only 11,000 to 14,000 tokens.

How do I work out if a model fits my VRAM?

Take about 90 percent of the card's memory as usable. Subtract the size of the quantized file you will load. Divide what is left by the model's KV cache per token, which is 2 times the attention layers times the KV heads times the head dimension times 2 bytes at 16-bit. The result is the longest context one session can hold. Sliding-window layers only count up to their window, and an 8-bit KV cache doubles the result.

Is the RTX 5090 worth it for local LLMs?

It is the only GeForce card with 32 GB, and NVIDIA lists 1,792 GB/s of memory bandwidth on a 512-bit bus. That extra 8 GB over a 24 GB card is what lets Qwen3.6-35B-A3B run with its full 256K window, or Gemma 4 31B with about 122,000 tokens. It does not reach 70B dense models at 4-bit, which need 40 GB or more, and NVIDIA asks for a 1,000 W power supply.

Filed under
local LLMopen-weight modelsquantizationlocal AI hardwarelocal AI infrastructureon-premises AI
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading