The best local LLM for RAG answers only from the retrieved passages, cites them, and says so when they do not hold the answer. On October 2026 evidence, Gemma 4 26B A4B is the strongest single-GPU pick, the Qwen3.5 to 3.8 models lead long-document reasoning, Llama 3.3 70B hallucinates least among larger open models, and Command A+ adds native citations on server GPUs.
Retrieval-augmented generation splits the work. Embedding models and rerankers find the passages; the generator turns them into an answer. This guide is about the generator. If retrieval itself is the problem, start with our guide to embedders and rerankers for private RAG.
Model facts below come from each vendor’s Hugging Face card, hallucination rates from Vectara’s public leaderboard (last updated 22 September 2026), and file sizes from the published GGUF repositories, all checked on 6 October 2026.
Quick picks
| Situation | Pick | Why |
|---|---|---|
| One 24–32 GB GPU, many users | Gemma 4 26B A4B | 5.2% grounded-summary hallucination rate, 256K window, small KV cache, Apache 2.0 |
| Long reports and contracts | Qwen3.6-35B-A3B or Qwen3.8-27B | Highest long-context reasoning scores among single-GPU models; Apache 2.0 |
| 8–16 GB GPU or a 16 GB Mac | Granite 4.2 8B or Qwen3.5-9B | 128K and 262K native windows; Apache 2.0 |
| 64–128 GB of GPU or unified memory | Llama 3.3 70B | 4.1% hallucination rate, the lowest of the larger open models; Llama 3.3 Community License |
| Server GPUs, citations required | Command A+ | Built-in citation mode; Apache 2.0; 218B total, 25B active |
What a RAG generator has to get right
A model that tops a general leaderboard can still be a poor RAG generator. Score candidates on five behaviours:
- Faithfulness. Every claim in the answer is supported by a retrieved passage. Unsupported additions are the failure that reaches users.
- Abstention. When the passages do not contain the answer, the model says so instead of filling the gap from its training data.
- Long-context use. It finds and combines facts spread across many passages, not only the first or last one.
- Instruction following. It respects the output format, the citation style and the “answer only from the documents” rule in the system prompt.
- Citations. It points to the passage behind each statement, in a form your application can check.
Public benchmarks cover the first, third and fourth reasonably well. Abstention and citation accuracy you will mostly have to measure on your own corpus.
Grounding and long-context evidence, side by side
Vectara’s leaderboard gives each model a set of documents to summarise using only their facts, then scores the summaries with its HHEM-2.3 evaluation model at temperature 0. Long-context scores come from each vendor’s card unless noted, so treat them as indicative rather than directly comparable.
| Model | Hallucination rate (Vectara) | Long-context evidence | Native window | Licence |
|---|---|---|---|---|
| Llama 3.3 70B | 4.1% | 128K window | 128K | Llama 3.3 Community License |
| Qwen3-8B / Qwen3-14B | 4.8% / 5.4% | 32K native; 131,072 with YaRN scaling | 32K | Apache 2.0 |
| Gemma 4 26B A4B | 5.2% | 44.1% on MRCR v2 8-needle at 128K (Google); 57.56 on AA-LCR (NVIDIA’s harness) | 256K | Apache 2.0 |
| Granite 4.0 H Small (32B, 9B active) | 5.2% | 128K; IBM lists RAG among its capabilities | 128K | Apache 2.0 |
| Gemma 4 31B | 7.4% | 66.4% on MRCR v2 8-needle at 128K (Google) | 256K | Apache 2.0 |
| Qwen3.5-35B-A3B | 10.5% | 58.5 on AA-LCR (Qwen); Qwen3.6-35B-A3B 61.06 (NVIDIA’s harness) | 262K | Apache 2.0 |
| Qwen3.5-27B | 12.1% | 66.1 on AA-LCR, 60.6 on LongBench v2 (Qwen) | 262K | Apache 2.0 |
| gpt-oss-120b | 14.2% | 50.7 on AA-LCR (Qwen’s comparison table) | 128K | Apache 2.0 plus usage policy |
| Ministral 3 14B | 19.4% | 256K window | 256K | Apache 2.0 |
AA-LCR, from Artificial Analysis, asks 100 questions that require reasoning across documents of 10,000 to 100,000 tokens and grades answers as pass or fail. Granite 4.2 8B and 30B are too new for Vectara’s board; IBM reports 71.41 and 81.38 on RULER at 128K for them.
What the numbers say:
- The two rankings disagree. The Qwen3.5 generation leads the long-context reasoning tests, yet in grounded summaries it added unsupported content about twice as often as the older Qwen3 models from the same lab.
- Gemma 4 26B A4B sits near the top of both. It is the only 2026 open model with a hallucination rate in the leading group and a 256K window on a single consumer GPU.
- Bigger is not safer. gpt-oss-120b and Mistral’s December 2025 releases scored higher hallucination rates than several 8B to 14B models.
- Retrieval inside an agent is still hard. On τ³-bench’s knowledge-retrieval banking domain, NVIDIA’s single-harness run put every open model it tested between 7% and 14%.
Check the answer rate too. Vectara reports how often each model declined to summarise; GLM-4.7-Flash, for example, answered 91.6% of prompts, so its 9.3% rate covers fewer documents.
Best local LLM for RAG by memory tier
| Hardware | Pick | 4-bit file | Concurrent 16K-token prompts that fit |
|---|---|---|---|
| 8–16 GB GPU | Granite 4.2 8B; Qwen3.5-9B for longer documents | 5.4 GB; 5.6 GB | On 16 GB: 3 and 14 |
| 24 GB GPU | Gemma 4 26B A4B | 14.6 GB | 12 (25 with an 8-bit KV cache) |
| 32 GB GPU | Gemma 4 26B A4B; Qwen3.6-35B-A3B | 14.6 GB; 20.4 GB | 26; 20 |
| 128 GB unified memory | Llama 3.3 70B; Qwen3.5-122B-A10B | 39.7 GB; 69.6 GB (MLX) | 11; 58 |
| Data-centre GPUs | Command A+ | W4A4 build on one B200 or two H100s (Cohere) | Size with the calculator |
The concurrency column uses the method of our memory calculator: 90% of GPU memory or 80% of unified memory, minus the weights file, divided by the KV cache one 16K-token prompt needs. Our VRAM tier guide and Mac guide list the same models with single-session context.
Worked example: RAG prompts on a 24 GB GPU
Suppose each request carries a system prompt, eight retrieved chunks and some chat history, about 16,384 tokens in all, on a 24 GB card with 21.6 GB usable.
- Gemma 4 26B A4B, Q4_0 (14.62 GB). Five global layers cost 2 × 5 × 2 KV heads × 512 × 2 bytes = 20,480 bytes per token, or 0.34 GB per prompt. Its 25 sliding-window layers add a fixed 0.21 GB. That is 0.55 GB per request, so (21.6 − 14.62) ÷ 0.55 gives 12 concurrent prompts.
- Mistral Small 3.2 24B, Q4_K_M (14.33 GB). All 40 layers keep a full cache: 163,840 bytes per token, 2.68 GB per prompt. The same card holds 2 concurrent prompts.
The weights are almost the same size. The difference is attention design, and it decides whether one GPU serves a department or a single analyst. The same arithmetic, applied to your headcount and context, is in our GPU sizing guide.
Citations: native spans or prompt and verify
Native citations. Command A+ accepts enable_citations=True in its chat template. The model then wraps supported text in <co> tags that name the tool call and result indices behind it, so your application can link every span to a source. It is a 218B-parameter mixture-of-experts model with 25B active, a 128K input window and an Apache 2.0 licence. Cohere lists a W4A4 build that runs on one B200 or two H100 GPUs, and vLLM 0.21 or later with the cohere_command4 parser.
Prompt and verify. With every other model:
- Number each retrieved passage in the prompt and keep a map from number to document and section.
- Instruct the model to cite passage numbers after each claim, and to say it cannot answer when no passage supports one.
- Reject or flag answers whose citation numbers do not exist in the map.
- Score support automatically: Vectara publishes an open variant of its evaluator, HHEM-2.1-Open, and its Open-RAG-Eval framework adds retrieval, groundedness and citation metrics.
Serving settings that quietly break RAG
- A small default window. Ollama’s documentation sets the default context from GPU memory: 4K below 24 GiB, 32K up to 48 GiB, 256K above. A 16K RAG prompt will not fit the 4K default, so set
OLLAMA_CONTEXT_LENGTHexplicitly. - Parallel requests multiply memory. Ollama’s FAQ notes that memory scales with
OLLAMA_NUM_PARALLEL× context length, and the default is one request per model at a time. For many users, a batching server is the better fit; see our vLLM, Ollama and llama.cpp comparison. - Static YaRN on Qwen3. Qwen notes that open-source frameworks implement static YaRN, which can reduce quality on shorter texts. Enable scaling only where requests exceed 32K.
- Thinking on by default. Qwen3.5 and later models think before answering unless you pass
enable_thinking: false. For short extractive answers, test both settings; reasoning adds latency and output tokens. - Quantization drift. A 4-bit build can move faithfulness scores. Run your grounding checks on the exact file you deploy; our quantization guide explains the trade-off.
For agents that call search or database tools instead of a fixed retrieval step, see our tool-calling model guide.
How to choose between two candidates on your own documents
Public scores narrow the field to two or three models. Your corpus decides the winner. A test that takes a few days:
- Collect 100 to 200 real questions from the people who will use the system, including some whose answer is not in the corpus at all.
- Freeze retrieval. Run each question through your retriever once and store the passages, so every generator sees exactly the same evidence.
- Generate with each candidate using the same system prompt, citation instructions and the quantized file you would deploy.
- Score four things per answer: supported by the passages, correct, correctly cited, and an honest “not found” when the passages lack the answer. An automatic grounding checker handles the first; people should review a sample of the rest.
- Record the cost side as well: memory per request at your typical prompt length, and how many concurrent requests the hardware holds.
- Re-run the set whenever you change the model, the quantization or the prompt, and keep the results with the deployment record.
Models that tie on faithfulness usually separate on abstention. That is the behaviour users notice when a confident answer turns out to come from nowhere.
How VDF AI fits
VDF AI Chat runs the whole RAG path inside your perimeter. According to its product page, connectors carry each source’s access-control lists into the index, and every query is filtered against the requesting user’s identity before candidate chunks are scored. The model answers from the retrieved passages with inline citations back to the source document and section.
Each turn writes an append-only audit record with the requesting identity, the model version, the retrieved chunk IDs and any tool calls, so an answer can be traced to its evidence later. Model choice is a per-agent setting: open-weight models such as the Qwen, DeepSeek, Mistral and Llama families served inside your perimeter, or any OpenAI-compatible endpoint, which lets you test two generators from this list against the same corpus.
Sources
Leaderboards, model cards and file listings verified 6 October 2026.
- Vectara hallucination leaderboard (updated 22 September 2026)
- Vectara HHEM-2.1-Open
- Artificial Analysis: AA-LCR
- τ-bench repository and τ³-bench notes
- Nemotron 3.5 Lightning card with NVIDIA’s single-harness comparison
- Gemma 4 31B model card (benchmarks for all Gemma 4 sizes)
- Qwen3.5-27B model card and Qwen3.5-35B-A3B model card
- Qwen3-14B model card
- Granite 4.2 8B model card and Granite 4.0 H Small model card
- Llama 3.3 70B Instruct
- Command A+ model card
- gpt-oss-120b model card
- Ministral 3 14B model card
- GGUF file sizes: ggml-org Gemma 4 26B A4B, LM Studio Mistral Small 3.2, IBM Granite 4.2 8B GGUF
- Ollama: context length and Ollama FAQ
Building private RAG over sensitive documents? See how VDF AI Chat keeps retrieval permission-aware, or book a demo.