The best local LLM for your VRAM is the strongest open-weight model whose 4-bit file plus KV cache fits in about 90% of the card. In October 2026 that means Qwen3.5-9B or Gemma 4 E4B on 8 GB, Gemma 4 12B on 12 GB, gpt-oss-20b on 16 GB, Qwen3.6-35B-A3B or Gemma 4 26B A4B on 24 GB, and Qwen3.6-35B-A3B with its full 256K window on a 32 GB RTX 5090.
Most “which model fits my GPU” answers stop at the weights. That is half the job. Every token you keep in the conversation also occupies memory, and two models of the same size can differ fivefold in how much. This guide picks models tier by tier and shows the context each one leaves you, so you can tell a model that loads from a model you can actually work with.
Every model, licence and parameter count below was checked on the vendor’s Hugging Face card, and every file size on the published GGUF repository, on 6 October 2026. GPU memory figures come from NVIDIA’s comparison page.
Quick picks by VRAM tier
| VRAM | Example cards (NVIDIA spec) | Best pick | Also strong | Context left for one chat |
|---|---|---|---|---|
| 8 GB | RTX 5060, RTX 4060, RTX 3070 | Qwen3.5-9B (Q4_K_M) | Gemma 4 E4B, Qwen3.5-4B | ≈46K tokens |
| 12 GB | RTX 5070, RTX 4070, RTX 3060 12 GB | Gemma 4 12B (Q4_0) | Qwen3.5-9B (Q6_K) | ≈198K tokens |
| 16 GB | RTX 5080, RTX 5070 Ti, RTX 5060 Ti 16 GB, RTX 4080 | gpt-oss-20b (MXFP4) | Gemma 4 12B (Q8_0), Qwen3.5-9B (Q8_0) | ≈93K tokens |
| 24 GB | RTX 4090, RTX 3090, RTX 3090 Ti | Qwen3.6-35B-A3B (Q4_K_M) | Gemma 4 26B A4B, Qwen3.8-27B, GLM-4.7-Flash | ≈54K tokens |
| 32 GB | RTX 5090 | Qwen3.6-35B-A3B (Q4_K_M) | Gemma 4 31B, Qwen3.8-27B, GLM-4.7-Flash | 256K (the model’s limit) |
Context figures assume one session and a 16-bit KV cache. If you write code with these models, our local coding model comparison covers coder-specific checkpoints.
How the fit is calculated
The method is the one behind our VRAM and KV cache calculator and the GPU sizing guide:
- Usable memory is 90% of the card, leaving room for the runtime, activations and the desktop.
- Weights are the size of the file you actually load. We use the GGUF published by the llama.cpp project’s
ggml-orgaccount, or by the model vendor where it ships its own (Mistral and IBM), or LM Studio’s community build for the small Qwen3.5 models. - KV cache per token is 2 × attention layers × KV heads × head dimension × 2 bytes at 16-bit. Sliding-window layers stop growing at their window. Hybrid models such as Qwen3.5 and Nemotron also hold a small fixed state per session, which we add separately.
- Context that fits is the memory left after steps 2 and 3, divided by the per-token cost, capped at the model’s own limit.
Real files rarely match the textbook arithmetic. Embeddings, output layers and vision encoders are often kept at higher precision, and different builders make different choices. Qwen3.6-27B at Q4_K_M is 19.1 GB from ggml-org but 17 GB in Ollama’s library, so check the file you download. The tables assume text only; loading a vision projector adds 0.6 to 1.2 GB.
Worked example: Gemma 4 31B on a 24 GB card
- Usable: 24 × 0.9 = 21.6 GB.
- Weights: the Q4_0 file is 17.99 GB.
- Sliding-window layers: 50 layers × 2 × 16 KV heads × 256 × 2 bytes = 819,200 bytes per token, but only for the last 1,024 tokens, so a fixed 0.84 GB.
- Global layers: 10 × 2 × 4 KV heads × 512 × 2 bytes = 81,920 bytes per token, for the whole conversation.
- Room for context: (21.6 − 17.99 − 0.84) GB ÷ 81,920 bytes ≈ 33,800 tokens. An 8-bit KV cache stretches that to about 77,900.
The calculator’s Gemma 4 31B preset at GGUF Q4_K_M (4.89 bits per weight, 19.1 GB) returns about 20,000 tokens for the same card. The gap is the file: Q4_K_M stores more bits per weight than this Q4_0 build, and that difference comes straight out of your context.
Best LLM for 8 GB VRAM
Cards in this tier include the RTX 5060, the 8 GB RTX 5060 Ti, the RTX 4060 and 4060 Ti 8 GB, and the RTX 3070. Usable memory is about 7.2 GB.
| Model | Parameters | Quantization | Weights | Context that fits | Licence | Good at |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | 9.7B dense (with vision encoder) | Q4_K_M | 5.6 GB | ≈46K | Apache 2.0 | Best quality here; Qwen reports 66.1 on BFCL-V4 and 63.0 on AA-LCR |
| Gemma 4 E4B | 8B with embeddings, 4.5B effective | Q4_0 | 4.6 GB | ≈90K | Apache 2.0 | Text, image and audio input |
| Qwen3.5-4B | 4.7B dense | Q4_K_M | 2.7 GB | ≈135K | Apache 2.0 | Long documents on a small card |
| Ministral 3 8B | 8.9B (with vision encoder) | Q4_K_M | 5.2 GB | ≈14K | Apache 2.0 | Native function calling and JSON output |
| Granite 4.2 8B | 8.8B dense | Q4_K_M | 5.4 GB | ≈11K | Apache 2.0 | Tool calling in 12 tested languages |
The spread in the context column is architecture, not size. Qwen3.5 keeps a growing cache in only 8 of its 32 layers, 32,768 bytes per token, while Granite 4.2 8B caches all 40 layers at 163,840 bytes per token. Gemma 4 E4B’s figure is conservative: it assumes every layer keeps its own cache, although the last 18 layers are designed to share earlier ones. Qwen3.5 models think before answering by default, so long reasoning traces eat into the same budget.
Best LLM for 12 GB VRAM
Examples are the RTX 5070, the RTX 4070, 4070 SUPER and 4070 Ti, the 12 GB RTX 3060 and the RTX 3080 Ti. Usable memory is about 10.8 GB.
| Model | Parameters | Quantization | Weights | Context that fits | Licence | Good at |
|---|---|---|---|---|---|---|
| Gemma 4 12B | 12.0B dense | Q4_0 | 7.2 GB | ≈198K | Apache 2.0 | Most context per gigabyte; text, image and audio; Google reports 69.0% on Tau2 |
| Qwen3.5-9B | 9.7B dense | Q6_K | 7.4 GB | ≈103K | Apache 2.0 | Higher-precision weights with a long window |
| Gemma 4 E4B | 8B with embeddings | Q8_0 | 8.0 GB | ≈95K | Apache 2.0 | Near-lossless on-device model |
| Ministral 3 14B | 13.9B (with vision encoder) | Q4_K_M | 8.2 GB | ≈16K | Apache 2.0 | Mistral calls it comparable to its Mistral Small 3.2 24B |
| Granite 4.2 8B | 8.8B dense | Q4_K_M | 5.4 GB | ≈33K | Apache 2.0 | Enterprise tool use, 128K native window |
Gemma 4 12B is the standout. Only 8 of its 48 layers attend over the whole conversation, each with a single 512-dimension KV head, so a 200K-token chat costs about 3.6 GB of cache, sliding-window layers included. gpt-oss-20b does not fit: its 12.1 GB file exceeds the 10.8 GB budget. llama.cpp can keep some mixture-of-experts weights in system RAM with --n-cpu-moe, at a cost in speed.
Best LLM for 16 GB VRAM
This is the most common enthusiast and workstation tier: RTX 5080, RTX 5070 Ti, the 16 GB RTX 5060 Ti, RTX 4080 and 4080 SUPER, RTX 4070 Ti SUPER and the 16 GB RTX 4060 Ti. Usable memory is about 14.4 GB.
| Model | Parameters | Quantization | Weights | Context that fits | Licence | Good at |
|---|---|---|---|---|---|---|
| gpt-oss-20b | 21B total, 3.6B active | MXFP4 (as trained) | 12.1 GB | ≈93K | Apache 2.0 plus usage policy | Adjustable reasoning effort, function calling, structured outputs |
| Gemma 4 12B | 12.0B dense | Q8_0 | 12.7 GB | ≈85K | Apache 2.0 | 8-bit quality with a long window; or Q4_0 at the full 256K |
| Qwen3.5-9B | 9.7B dense | Q8_0 | 9.5 GB | ≈147K | Apache 2.0 | Long context at 8-bit; image input |
| Ministral 3 14B | 13.9B | Q4_K_M | 8.2 GB | ≈38K | Apache 2.0 | Function calling and JSON in 11 listed languages |
| Granite 4.2 8B | 8.8B dense | Q4_K_M | 5.4 GB | ≈55K | Apache 2.0 | Thinking modes you can switch off per request |
gpt-oss-20b was trained with MXFP4 weights, so there is no separate quantization step to second-guess; OpenAI’s card says it runs within 16 GB. OpenAI says the model must be used with its harmony response format; the chat template in the repository applies it, and llama.cpp lists a native handler for it.
What does not fit: Mistral Small 3.2 24B at Q4_K_M is 14.3 GB and leaves a few hundred tokens, Gemma 4 26B A4B’s 14.6 GB Q4_0 file is just over budget, and the 27B Qwen models need 19 GB. One trap applies to this whole tier: Ollama’s documentation says it defaults to a 4K context on GPUs with less than 24 GiB, so set OLLAMA_CONTEXT_LENGTH to the window you sized for.
Best LLM for 24 GB VRAM
The RTX 4090, RTX 3090 and RTX 3090 Ti sit here. Usable memory is about 21.6 GB, and the choice widens to 24B–36B models.
| Model | Parameters | Quantization | Weights | Context that fits | Licence | Good at |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | 36B total, 3B active | Q4_K_M | 20.4 GB | ≈54K | Apache 2.0 | Agent work; Qwen reports 62.8 on MCP-Atlas |
| Gemma 4 26B A4B | 25.8B total, 3.8B active | Q4_0 | 14.6 GB | 256K (model limit) | Apache 2.0 | The full window on one card; long documents |
| Qwen3.8-27B | 27.8B dense | Q4_K_M | 19.0 GB | ≈38K | Apache 2.0 | Newest Qwen dense model; 79.5 on IFBench per Qwen |
| Gemma 4 31B | 31.3B dense | Q4_0 | 18.0 GB | ≈34K | Apache 2.0 | Strongest Gemma: 85.2% MMLU Pro, 76.9% Tau2 per Google |
| GLM-4.7-Flash | 31.2B total, about 3B active | Q4_K | 18.2 GB | ≈62K | MIT | Multi-turn tool use: 79.5 on τ²-Bench per Z.ai |
| Mistral Small 3.2 24B | 24.0B dense | Q4_K_M | 14.3 GB | ≈44K | Apache 2.0 | Image input; robust function calling |
| Nemotron 3.5 Lightning 30B-A3B | 31.6B total, 3B active | Q4_0 | 18.9 GB | ≈430K | OpenMDW-1.1 | Very long inputs: only 6 of 52 layers are attention |
Two notes on the dense Qwen models. Qwen’s Qwen3.6 card advises keeping at least 128K of context to preserve thinking quality, which a 24 GB card cannot give the 27B at 16-bit. An 8-bit KV cache roughly doubles the figure; turning thinking off for simple requests is the other lever. Qwen3.6-27B behaves like Qwen3.8-27B here, with a 19.1 GB file and about 36K tokens.
Best LLM for 32 GB VRAM (RTX 5090)
The GeForce RTX 5090 is the only 32 GB card in NVIDIA’s consumer line. NVIDIA lists 32 GB of GDDR7 on a 512-bit interface, 1,792 GB/s of memory bandwidth, 575 W of total graphics power and a 1,000 W system power requirement (verified October 2026). Usable memory is about 28.8 GB.
| Model | Parameters | Quantization | Weights | Context that fits | Licence | Good at |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | 36B total, 3B active | Q4_K_M | 20.4 GB | 256K (model limit) | Apache 2.0 | Full window plus room for parallel sessions |
| Qwen3.8-27B | 27.8B dense | Q4_K_M | 19.0 GB | ≈148K | Apache 2.0 | Meets Qwen’s 128K thinking advice |
| Gemma 4 31B | 31.3B dense | Q4_0 | 18.0 GB | ≈122K | Apache 2.0 | Best dense Gemma with a long window |
| Gemma 4 26B A4B | 25.8B total, 3.8B active | Q8_0 | 26.9 GB | ≈84K | Apache 2.0 | 8-bit weights for quality-sensitive work |
| GLM-4.7-Flash | 31.2B total | Q4_K | 18.2 GB | ≈195K | MIT | Long multi-turn tool sessions |
| Mistral Small 3.2 24B | 24.0B dense | Q6_K | 19.4 GB | ≈58K | Apache 2.0 | Higher-precision Mistral |
| Nemotron 3.5 Lightning 30B-A3B | 31.6B total | Q4_0 | 18.9 GB | 1M (model limit) | OpenMDW-1.1 | Million-token inputs on one card |
The extra 8 GB over a 24 GB card mostly buys context and concurrency rather than bigger models. 8-bit 27B files (28.6 GB) load but leave almost nothing for the cache, and 70B dense models are out of reach: Ollama’s Llama 3.3 70B at q4_K_M is 43 GB. For that class, look at 48 GB-plus GPUs or the unified-memory machines in our DGX Spark, Mac Studio and Ryzen AI Max comparison.
When the model you want does not fit
Five levers, in the order we would pull them:
- Quantize the KV cache. llama.cpp accepts
--cache-type-k q8_0(and the matching V flag), Ollama exposesOLLAMA_KV_CACHE_TYPE=q8_0, and vLLM supports an FP8 KV cache. Ollama’s FAQ describes q8_0 as roughly half the memory of f16 with a very small loss in precision. - Shorten the window. Size context for the task, not the model card’s maximum.
- Offload experts. For mixture-of-experts models, llama.cpp’s
--n-cpu-moe Nkeeps the expert weights of the first N layers in system RAM. It runs, slower. - Let llama.cpp fit it. Current
llama-serverbuilds adjust unset arguments to device memory by default (--fit), keeping a 1,024 MiB margin per device unless you change it. - Drop a quantization level. Weight precision is the last lever, because quality loss varies by model and task. Our quantization trade-off guide explains how to test it.
If several people will share one GPU, every concurrent session needs its own cache, and single-user runners queue requests by default. Our vLLM, Ollama and llama.cpp comparison covers the batching servers built for that. For the same picks on Apple hardware, see the Mac memory-tier guide; for retrieval and agent workloads, see our picks for local RAG and tool calling. To pick the card itself, the GPU buyer’s guide compares them by budget, and the Gemma 4 setup guide walks through one family end to end.
How VDF AI fits
A single GPU under a desk serves one person well. When a team needs the same local models, VDF AI Router registers Ollama and custom on-premises deployments alongside any cloud models your policy allows, and its air-gap mode restricts routing to local models only. Local runtimes are probed continuously, so a request can move to another permitted model when one endpoint is down.
That lets you keep a small model on modest hardware for routine requests and send heavier work to a larger model on a server, behind one API. The local LLM overview covers the wider deployment picture.
Sources
Model cards, file listings and specifications verified 6 October 2026.
- Qwen3.5-9B model card and config
- Qwen3.5-4B model card
- Qwen3.6-35B-A3B model card
- Qwen3.6-27B model card
- Qwen3.8-27B model card
- Gemma 4 31B model card (also covers E4B, 12B and 26B A4B)
- Gemma 4 licence page
- gpt-oss-20b model card
- Ministral 3 8B and Ministral 3 14B model cards
- Mistral Small 3.2 24B model card
- Granite 4.2 8B model card
- GLM-4.7-Flash model card
- Nemotron 3.5 Lightning model card and OpenMDW-1.1 licence
- GGUF file sizes: ggml-org Qwen3.6-35B-A3B, ggml-org Qwen3.8-27B, ggml-org Gemma 4 31B, ggml-org gpt-oss-20b, Mistral Ministral 3 14B GGUF, IBM Granite 4.2 8B GGUF, LM Studio Qwen3.5-9B GGUF
- Ollama library: Qwen3.6 tags and Llama 3.3 tags
- NVIDIA GeForce RTX 5090 specifications
- NVIDIA GeForce graphics card comparison
- llama.cpp server README
- Ollama: context length and Ollama FAQ
- vLLM quantized KV cache
Moving a local model from one GPU to a shared service? See how VDF AI Router routes across on-premises models, or book a demo.