AI Infrastructure

Best Local LLM for RAG (2026): Models That Stay Grounded

Which open-weight model to use as the generator in a private retrieval-augmented generation system, judged on what RAG needs: staying faithful to retrieved passages, reasoning across long contexts, following format instructions and citing sources. Hallucination and long-context scores from named sources, memory sizing for concurrent RAG prompts, and picks for each hardware tier.

The best local LLM for RAG answers only from the retrieved passages, cites them, and says so when they do not hold the answer. On October 2026 evidence, Gemma 4 26B A4B is the strongest single-GPU pick, the Qwen3.5 to 3.8 models lead long-document reasoning, Llama 3.3 70B hallucinates least among larger open models, and Command A+ adds native citations on server GPUs.

Retrieval-augmented generation splits the work. Embedding models and rerankers find the passages; the generator turns them into an answer. This guide is about the generator. If retrieval itself is the problem, start with our guide to embedders and rerankers for private RAG.

Model facts below come from each vendor’s Hugging Face card, hallucination rates from Vectara’s public leaderboard (last updated 22 September 2026), and file sizes from the published GGUF repositories, all checked on 6 October 2026.

Quick picks

SituationPickWhy
One 24–32 GB GPU, many usersGemma 4 26B A4B5.2% grounded-summary hallucination rate, 256K window, small KV cache, Apache 2.0
Long reports and contractsQwen3.6-35B-A3B or Qwen3.8-27BHighest long-context reasoning scores among single-GPU models; Apache 2.0
8–16 GB GPU or a 16 GB MacGranite 4.2 8B or Qwen3.5-9B128K and 262K native windows; Apache 2.0
64–128 GB of GPU or unified memoryLlama 3.3 70B4.1% hallucination rate, the lowest of the larger open models; Llama 3.3 Community License
Server GPUs, citations requiredCommand A+Built-in citation mode; Apache 2.0; 218B total, 25B active

What a RAG generator has to get right

A model that tops a general leaderboard can still be a poor RAG generator. Score candidates on five behaviours:

  1. Faithfulness. Every claim in the answer is supported by a retrieved passage. Unsupported additions are the failure that reaches users.
  2. Abstention. When the passages do not contain the answer, the model says so instead of filling the gap from its training data.
  3. Long-context use. It finds and combines facts spread across many passages, not only the first or last one.
  4. Instruction following. It respects the output format, the citation style and the “answer only from the documents” rule in the system prompt.
  5. Citations. It points to the passage behind each statement, in a form your application can check.

Public benchmarks cover the first, third and fourth reasonably well. Abstention and citation accuracy you will mostly have to measure on your own corpus.

Grounding and long-context evidence, side by side

Vectara’s leaderboard gives each model a set of documents to summarise using only their facts, then scores the summaries with its HHEM-2.3 evaluation model at temperature 0. Long-context scores come from each vendor’s card unless noted, so treat them as indicative rather than directly comparable.

ModelHallucination rate (Vectara)Long-context evidenceNative windowLicence
Llama 3.3 70B4.1%128K window128KLlama 3.3 Community License
Qwen3-8B / Qwen3-14B4.8% / 5.4%32K native; 131,072 with YaRN scaling32KApache 2.0
Gemma 4 26B A4B5.2%44.1% on MRCR v2 8-needle at 128K (Google); 57.56 on AA-LCR (NVIDIA’s harness)256KApache 2.0
Granite 4.0 H Small (32B, 9B active)5.2%128K; IBM lists RAG among its capabilities128KApache 2.0
Gemma 4 31B7.4%66.4% on MRCR v2 8-needle at 128K (Google)256KApache 2.0
Qwen3.5-35B-A3B10.5%58.5 on AA-LCR (Qwen); Qwen3.6-35B-A3B 61.06 (NVIDIA’s harness)262KApache 2.0
Qwen3.5-27B12.1%66.1 on AA-LCR, 60.6 on LongBench v2 (Qwen)262KApache 2.0
gpt-oss-120b14.2%50.7 on AA-LCR (Qwen’s comparison table)128KApache 2.0 plus usage policy
Ministral 3 14B19.4%256K window256KApache 2.0

AA-LCR, from Artificial Analysis, asks 100 questions that require reasoning across documents of 10,000 to 100,000 tokens and grades answers as pass or fail. Granite 4.2 8B and 30B are too new for Vectara’s board; IBM reports 71.41 and 81.38 on RULER at 128K for them.

What the numbers say:

  • The two rankings disagree. The Qwen3.5 generation leads the long-context reasoning tests, yet in grounded summaries it added unsupported content about twice as often as the older Qwen3 models from the same lab.
  • Gemma 4 26B A4B sits near the top of both. It is the only 2026 open model with a hallucination rate in the leading group and a 256K window on a single consumer GPU.
  • Bigger is not safer. gpt-oss-120b and Mistral’s December 2025 releases scored higher hallucination rates than several 8B to 14B models.
  • Retrieval inside an agent is still hard. On τ³-bench’s knowledge-retrieval banking domain, NVIDIA’s single-harness run put every open model it tested between 7% and 14%.

Check the answer rate too. Vectara reports how often each model declined to summarise; GLM-4.7-Flash, for example, answered 91.6% of prompts, so its 9.3% rate covers fewer documents.

Best local LLM for RAG by memory tier

HardwarePick4-bit fileConcurrent 16K-token prompts that fit
8–16 GB GPUGranite 4.2 8B; Qwen3.5-9B for longer documents5.4 GB; 5.6 GBOn 16 GB: 3 and 14
24 GB GPUGemma 4 26B A4B14.6 GB12 (25 with an 8-bit KV cache)
32 GB GPUGemma 4 26B A4B; Qwen3.6-35B-A3B14.6 GB; 20.4 GB26; 20
128 GB unified memoryLlama 3.3 70B; Qwen3.5-122B-A10B39.7 GB; 69.6 GB (MLX)11; 58
Data-centre GPUsCommand A+W4A4 build on one B200 or two H100s (Cohere)Size with the calculator

The concurrency column uses the method of our memory calculator: 90% of GPU memory or 80% of unified memory, minus the weights file, divided by the KV cache one 16K-token prompt needs. Our VRAM tier guide and Mac guide list the same models with single-session context.

Worked example: RAG prompts on a 24 GB GPU

Suppose each request carries a system prompt, eight retrieved chunks and some chat history, about 16,384 tokens in all, on a 24 GB card with 21.6 GB usable.

  • Gemma 4 26B A4B, Q4_0 (14.62 GB). Five global layers cost 2 × 5 × 2 KV heads × 512 × 2 bytes = 20,480 bytes per token, or 0.34 GB per prompt. Its 25 sliding-window layers add a fixed 0.21 GB. That is 0.55 GB per request, so (21.6 − 14.62) ÷ 0.55 gives 12 concurrent prompts.
  • Mistral Small 3.2 24B, Q4_K_M (14.33 GB). All 40 layers keep a full cache: 163,840 bytes per token, 2.68 GB per prompt. The same card holds 2 concurrent prompts.

The weights are almost the same size. The difference is attention design, and it decides whether one GPU serves a department or a single analyst. The same arithmetic, applied to your headcount and context, is in our GPU sizing guide.

Citations: native spans or prompt and verify

Native citations. Command A+ accepts enable_citations=True in its chat template. The model then wraps supported text in <co> tags that name the tool call and result indices behind it, so your application can link every span to a source. It is a 218B-parameter mixture-of-experts model with 25B active, a 128K input window and an Apache 2.0 licence. Cohere lists a W4A4 build that runs on one B200 or two H100 GPUs, and vLLM 0.21 or later with the cohere_command4 parser.

Prompt and verify. With every other model:

  1. Number each retrieved passage in the prompt and keep a map from number to document and section.
  2. Instruct the model to cite passage numbers after each claim, and to say it cannot answer when no passage supports one.
  3. Reject or flag answers whose citation numbers do not exist in the map.
  4. Score support automatically: Vectara publishes an open variant of its evaluator, HHEM-2.1-Open, and its Open-RAG-Eval framework adds retrieval, groundedness and citation metrics.

Serving settings that quietly break RAG

  • A small default window. Ollama’s documentation sets the default context from GPU memory: 4K below 24 GiB, 32K up to 48 GiB, 256K above. A 16K RAG prompt will not fit the 4K default, so set OLLAMA_CONTEXT_LENGTH explicitly.
  • Parallel requests multiply memory. Ollama’s FAQ notes that memory scales with OLLAMA_NUM_PARALLEL × context length, and the default is one request per model at a time. For many users, a batching server is the better fit; see our vLLM, Ollama and llama.cpp comparison.
  • Static YaRN on Qwen3. Qwen notes that open-source frameworks implement static YaRN, which can reduce quality on shorter texts. Enable scaling only where requests exceed 32K.
  • Thinking on by default. Qwen3.5 and later models think before answering unless you pass enable_thinking: false. For short extractive answers, test both settings; reasoning adds latency and output tokens.
  • Quantization drift. A 4-bit build can move faithfulness scores. Run your grounding checks on the exact file you deploy; our quantization guide explains the trade-off.

For agents that call search or database tools instead of a fixed retrieval step, see our tool-calling model guide.

How to choose between two candidates on your own documents

Public scores narrow the field to two or three models. Your corpus decides the winner. A test that takes a few days:

  1. Collect 100 to 200 real questions from the people who will use the system, including some whose answer is not in the corpus at all.
  2. Freeze retrieval. Run each question through your retriever once and store the passages, so every generator sees exactly the same evidence.
  3. Generate with each candidate using the same system prompt, citation instructions and the quantized file you would deploy.
  4. Score four things per answer: supported by the passages, correct, correctly cited, and an honest “not found” when the passages lack the answer. An automatic grounding checker handles the first; people should review a sample of the rest.
  5. Record the cost side as well: memory per request at your typical prompt length, and how many concurrent requests the hardware holds.
  6. Re-run the set whenever you change the model, the quantization or the prompt, and keep the results with the deployment record.

Models that tie on faithfulness usually separate on abstention. That is the behaviour users notice when a confident answer turns out to come from nowhere.

How VDF AI fits

VDF AI Chat runs the whole RAG path inside your perimeter. According to its product page, connectors carry each source’s access-control lists into the index, and every query is filtered against the requesting user’s identity before candidate chunks are scored. The model answers from the retrieved passages with inline citations back to the source document and section.

Each turn writes an append-only audit record with the requesting identity, the model version, the retrieved chunk IDs and any tool calls, so an answer can be traced to its evidence later. Model choice is a per-agent setting: open-weight models such as the Qwen, DeepSeek, Mistral and Llama families served inside your perimeter, or any OpenAI-compatible endpoint, which lets you test two generators from this list against the same corpus.

Sources

Leaderboards, model cards and file listings verified 6 October 2026.


Building private RAG over sensitive documents? See how VDF AI Chat keeps retrieval permission-aware, or book a demo.

Frequently asked questions

What is the best local LLM for RAG?

On a single 24 to 32 GB GPU, Gemma 4 26B A4B is the best balance in October 2026: Apache 2.0, a 256K window, a 5.2 percent hallucination rate on Vectara's grounded-summary leaderboard, and a small KV cache that serves many concurrent prompts. For long-document reasoning, the Qwen3.5 to 3.8 models score higher on long-context tests. With 64 GB or more, Llama 3.3 70B posted the lowest hallucination rate of the larger open models on that leaderboard.

Which local LLM hallucinates least with retrieved documents?

On Vectara's leaderboard, which asks models to summarise documents using only their facts, the open-weight models with the lowest rates in September 2026 included Llama 3.3 70B at 4.1 percent, Qwen3-8B at 4.8 percent and Gemma 4 26B A4B and Granite 4.0 H Small at 5.2 percent. Several newer reasoning models scored higher, for example Qwen3.5-27B at 12.1 percent and gpt-oss-120b at 14.2 percent. Test on your own documents before deciding.

Do I need a long-context model for RAG?

You need enough context for the system prompt, the retrieved passages, the conversation and the answer, which for most enterprise RAG is 8,000 to 32,000 tokens per request. A larger window helps when one question needs whole contracts or reports. Remember that every concurrent user needs their own KV cache, so a model with a cheap per-token cache can serve far more people at 16,000 tokens than a model with a costly one.

Can a local LLM cite its sources?

Yes, in two ways. Cohere's Command A+ has a citation mode in its chat template that tags each supported span with the tool results it came from. With other models, number the retrieved passages, instruct the model to cite passage numbers, and check programmatically that every cited number exists and that the cited passage supports the sentence. A grounding checker such as Vectara's open HHEM model can automate the second check.

What is the best small local LLM for RAG on 8 to 16 GB?

Granite 4.2 8B is a strong default: Apache 2.0, a 128K native window, and IBM reports 71.41 on RULER at 128K. Qwen3-8B, an older Apache 2.0 model, had one of the lowest grounded-summary hallucination rates at 4.8 percent but only a 32K native window. Qwen3.5-9B scores higher on long-context reasoning tests and keeps a much smaller KV cache, which leaves more room for retrieved passages on a small card.

Filed under
private RAGRAGlocal LLMopen-weight modelson-premises AIlocal AI infrastructure
Private RAG & Search

Evaluate your knowledge stack

Find out how a private RAG and retrieval layer would perform on your data — accuracy, latency, governance, and what to fix before you scale.

Or start free — no credit card →

Keep reading