RAG

Best RAG Frameworks (2026): LlamaIndex, LangChain, Haystack and More

Open-source RAG frameworks checked against their GitHub repositories and documentation in October 2026: LlamaIndex, LangChain with LangGraph, Haystack, RAGFlow, DSPy, txtai, Kotaemon and LightRAG, compared on ingestion, retrieval, evaluation, agentic RAG and deployment, with the projects that are archived or in maintenance mode flagged.

The best RAG framework in 2026 depends on how much of the pipeline you want to own. Haystack gives explicit pipelines with built-in evaluators, LangChain with LangGraph adds durable agentic retrieval, LlamaIndex brings the widest data integrations, and RAGFlow is a ready self-hosted RAG engine with strong document parsing. All are open source.

Quick picks

Licences, release dates and maintenance status were checked on GitHub on 6 October 2026; features come from each project’s own documentation.

If you want…Start withWhy
Explicit, inspectable pipelines with evaluators built inHaystackComponents for indexing, retrieval, ranking and nine evaluators; 3.0 added agent hooks
Agentic RAG that can pause for approval and resumeLangChain with LangGraphDurable execution, interrupts and memory in LangGraph; RAG building blocks in LangChain
The widest set of data and model integrationsLlamaIndexOver 300 integration packages and a full evaluation module
A self-hosted RAG application rather than a libraryRAGFlowLayout analysis, OCR and table recognition, citations and a web UI in Docker Compose
Prompts optimised against a metric instead of hand-tunedDSPyOptimizers such as MIPROv2 compile a RAG program against your evaluation set
Graph-based retrieval over entities and relationshipsLightRAGKnowledge graph plus vector retrieval with several storage backends
One lightweight library for search, RAG and agentstxtaiEmbeddings database with sparse, dense, graph and SQL search
A document chat UI to put in front of users quicklyKotaemonHybrid retrieval, reranking and citations shown in a PDF viewer

Library, engine or platform: what you are choosing

“RAG framework” covers three different things, and mixing them up wastes evaluation time.

  • Libraries such as LlamaIndex, LangChain, Haystack, DSPy and txtai give you components to assemble in your own code. You own the service, the API, the deployment and the operations.
  • Engines and applications such as RAGFlow, Kotaemon and LightRAG’s server ship a running system with a UI and an API. You configure more than you code, and you accept their architecture.
  • Platforms add identity, permissions, audit and administration across many use cases. Our page on enterprise RAG describes what that layer has to do.

A framework also sits on top of a vector database, which is a separate decision. Our list of vector databases for RAG covers that layer.

RAG frameworks compared

Licence, release and maintenance status

FrameworkLicenceLatest release (as of 6 Oct 2026)Status
LlamaIndexMIT0.14.25 (21 Sep 2026)Active; company focus now on LlamaParse
LangChain / LangGraphMITlangchain 1.4.3 (28 Sep 2026); LangGraph 1.2.13 (5 Oct 2026)Active; both 1.0 since 17 Oct 2025
HaystackApache 2.03.3.0 (1 Oct 2026)Active; 3.0 released 20 Jul 2026
RAGFlowApache 2.01.0.0-rc1 (29 Sep 2026); 0.27.2 (10 Sep 2026)Active; 1.0 moves the server to Go
DSPyMIT3.4.0 (25 Sep 2026)Active
txtaiApache 2.09.13.0 (27 Aug 2026)Active
KotaemonApache 2.00.12.0 (31 May 2026)Slower; last commit 14 Jul 2026
LightRAGMIT1.5.7 (2 Sep 2026)Active
Microsoft GraphRAGMIT3.2.0 (24 Sep 2026)Largely in maintenance mode
R2RMIT3.6.5 (6 Jun 2025)No commits since 7 Nov 2025
Verba (Weaviate)BSD 3-Clause2.1.3 (14 Jul 2025)Archived
Cognita (TrueFoundry)Apache 2.0–Archived

Capabilities

A dash in this table means we found no documentation of the capability on the pages we read; check the project before ruling it out.

FrameworkIngestionRetrievalEvaluation hooksAgentic RAGDeployment
LlamaIndex300+ integration packages; LlamaParse (commercial) or LiteParse for parsingRetrievers and query engines over any supported vector storeFaithfulness, relevancy, correctness and guideline evaluators; hit rate and MRRAgent WorkflowsLibrary; llama-agents-server exposes workflows as REST
LangChain / LangGraphDocument loaders and text splittersRetrievers; 2-step, agentic and hybrid RAG patternsopenevals (MIT): correctness, groundedness, retrieval relevance; LangSmith (commercial)LangGraph with persistence and interruptsLibrary; LangSmith Deployment
HaystackIndexing components in explicit pipelinesRetrieval, ranking, filtering and routing componentsNine built-in evaluators plus Ragas and DeepEvalAgent with lifecycle hooksHayhooks: REST, MCP or OpenAI-compatible endpoints
RAGFlowDeepDoc layout analysis, OCR and tables; Word, slides, Excel, images, scansMultiple recall with fused re-ranking; compiled graphs, trees and wikis–Agentic RAG with four thinking modesDocker Compose
DSPyBring your owndspy.retrievers.Embeddings or your own retrieverdspy.Evaluate with metrics such as SemanticF1Agent loops as modulesLibrary
txtaiPipelines including transcription and summarisationSparse and dense indexes, graph traversal, SQL–AgentsLibrary, API server, MCP API, Docker
KotaemonPDF, HTML and XLSX; more via UnstructuredHybrid full-text and vector with reranking–ReAct and ReWOO agentsDocker images, one bundled with Ollama
LightRAGMultimodal parsing via MinerU or DoclingKnowledge graph plus vector, with rerankerRAGAS integration–Server with WebUI; Docker setup wizard

Notes on each framework

LlamaIndex

LlamaIndex is the most integration-heavy option: core plus more than 300 packages for models, embeddings and vector stores. Its evaluation module is the most complete of the libraries here, with LLM-based correctness, semantic similarity, faithfulness, context relevancy, answer relevancy and guideline evaluators, ranking metrics such as hit rate and MRR for retrievers, synthetic question generation, and integrations with Ragas, DeepEval and others. Know the company context, though. In September 2026 the README added a note that the company’s primary focus has shifted to LlamaParse, its commercial parsing, extraction and indexing platform, while the framework remains available as an open toolkit. Releases continue on a regular cadence, so this is a signal to watch rather than an exit sign. For agents, LlamaIndex now builds on Agent Workflows, covered in our agent framework list.

LangChain and LangGraph

LangChain’s documentation names five building blocks for retrieval: document loaders, text splitters, embedding models, vector stores and retrievers. It describes three architectures: 2-step RAG, where retrieval always runs first; agentic RAG, where an agent decides when and how to retrieve; and a hybrid of the two with validation steps. LangGraph supplies what agentic RAG needs in production, namely durable execution that resumes after failures, human-in-the-loop interrupts and short- and long-term memory. Both packages reached 1.0 on 17 October 2025. Evaluation can run through LangSmith, a commercial service, or the MIT-licensed openevals package, whose RAG evaluators cover correctness, helpfulness, groundedness and retrieval relevance.

Haystack

Haystack, from deepset, builds RAG and agents as explicit pipelines of components with loops, branches and conditional routing, which makes every step inspectable. Version 3.0, released on 20 July 2026, gave the Agent hooks such as before_llm and before_tool for guardrails and human-in-the-loop checkpoints, plus first-class skills and async serving. Evaluation is built in: statistical evaluators for exact match, MRR, MAP, NDCG and recall, and model-based ones for context relevance, faithfulness and semantic answer similarity, with Ragas and DeepEval wrappers. Hayhooks serves a pipeline as a REST API, an MCP server or an OpenAI-compatible chat endpoint. One setting to change for private deployments: Haystack sends anonymous component-usage telemetry by default, which HAYSTACK_TELEMETRY_ENABLED=False switches off.

RAGFlow

RAGFlow is a self-hosted RAG engine with a web interface, not a library. Its strength is parsing: the DeepDoc module handles layout analysis, OCR and table recognition across Word, slides, Excel, text, images, scans and web pages, and chunking is template-based with a visual view so people can correct chunks by hand. Retrieval uses multiple recall paths with fused re-ranking, and August 2026 releases added agentic RAG with four thinking modes and “knowledge compilation” into wikis, graphs, trees and timelines. Answers carry traceable citations. The 1.0.0-rc1 release of 29 September 2026 moves the server to Go. The README recommends at least 4 CPU cores, 16 GB of RAM and 50 GB of disk, with Elasticsearch or Infinity as the document engine. Our guide to OCR for scanned documents explains why parsing quality decides so much downstream.

DSPy

DSPy, from Stanford NLP, treats a RAG system as a program whose prompts and weights are optimised against a metric. Its RAG tutorial retrieves with dspy.retrievers.Embeddings, scores answers with dspy.Evaluate and a SemanticF1 metric, and compiles the program with the MIPROv2 optimizer, which in that tutorial lifted semantic F1 on the development set from about 42% to 61%. That is the project’s own example, not a general benchmark. DSPy brings no connectors or ingestion layer, so teams usually pair it with another framework or their own loaders, and a good evaluation set is a prerequisite.

txtai

txtai, from NeuML, centres on an embeddings database that combines sparse and dense vector indexes, graph networks and a relational store. Its RAG pipeline draws context from that database and generates with Hugging Face Transformers, llama.cpp or LiteLLM backends, and its examples cover citations and graph path traversal for multi-hop questions. It also runs agents, exposes web and MCP APIs, and installs with pip or Docker. It suits small teams that want one dependency instead of five.

Kotaemon and LightRAG

Kotaemon is a Gradio-based document chat interface with a hybrid full-text and vector retriever, reranking, question decomposition, ReAct and ReWOO agents, and citations shown with highlights in an in-browser PDF viewer. Docker images include a variant bundled with Ollama for fully local use. Activity has slowed, with the latest release on 31 May 2026. LightRAG, from the HKUDS group, adds a knowledge graph to vector retrieval, stores data in PostgreSQL, MongoDB, Neo4j or OpenSearch, and integrated RAGAS evaluation and Langfuse tracing in late 2025. When relationships across documents matter more than single passages, read our explainer on knowledge graph RAG first.

Archived and slowing projects

Four names still appear in older comparisons. Weaviate’s Verba and TrueFoundry’s Cognita are archived on GitHub. Microsoft’s GraphRAG README says the project is largely in maintenance mode and will not accept new features, and warns that its indexing can be expensive. R2R is not archived, but its latest release is from June 2025 and its last commit from November 2025. None is a sound base for a new production system.

How to choose a RAG framework

  1. Decide library or engine. If you have engineers who will own a service, pick a library. If you need working document Q&A next month, an engine such as RAGFlow gets you further faster.
  2. Test parsing on your worst documents. Scanned contracts, nested tables and slide decks break more pipelines than retrieval settings do. Our guide to chunking enterprise documents covers what to check after parsing.
  3. Check the retrieval you need is first-class. Hybrid keyword and vector search, metadata filters and reranking should be configuration, not custom code.
  4. Decide whether retrieval must be agentic. Multi-hop questions and tool use favour LangGraph, Haystack’s agent or LlamaIndex workflows; most FAQ-style assistants do not need it. Our comparison of agentic and traditional RAG sets out when it pays.
  5. Wire evaluation in before features. Pick a framework whose evaluators you will actually run in CI.
  6. Check maintenance, not stars. Look at the latest release, the last commit and any status note in the README.
  7. Audit the defaults. Telemetry, hosted tracing and cloud parsing services can send data out of a private deployment unless you switch them off.

Evaluating a RAG pipeline before production

Whichever framework you choose, accuracy is measured outside it. Build a labelled set of real user questions and the passages that answer them, and score retrieval and generation separately:

LayerWhat to measureBuilt-in options
RetrievalRecall, hit rate, MRR, NDCGHaystack document evaluators; LlamaIndex retrieval evaluation
GenerationFaithfulness to retrieved context, answer relevanceHaystack, LlamaIndex, openevals; Ragas and DeepEval via wrappers
CitationsWhether each cited passage supports the claimUsually custom, using an LLM judge plus spot checks
End to endSemantic overlap with reference answersDSPy SemanticF1; LlamaIndex semantic similarity

Run the set after every change to chunking, embeddings, retrieval settings, prompts or the model, and keep the results with the release. Our RAG accuracy evaluation framework describes the method and the thresholds regulated teams use.

How VDF AI fits

Frameworks give you the building blocks; VDF AI ships the retrieval layer as part of the product. In VDF AI Chat, connectors, chunking, embedding, a local vector store and a permission filter form one layer, and retrieval is filtered by the reader’s permissions at query time. Answers cite their sources, and the governance layer writes append-only logs for every turn, applies retention policy and exports to your SIEM.

Teams that want more control can build vector indexes in VDF AI Data with their own chunk size, overlap and embedding model, and restrict sensitive knowledge sources to named people and agents. In VDF AI Networks, a Data Source node carries retrieval settings, embedding models and query templates into multi-step workflows, so retrieval becomes one governed step in a larger workflow. Everything runs on-premises, in a private cloud or air-gapped.

Sources

Verified 6 October 2026.

Frequently asked questions

What is the best RAG framework in 2026?

It depends on how much of the pipeline you want to own. Haystack suits teams that want explicit, inspectable pipelines with built-in evaluators. LangChain with LangGraph suits teams building agentic RAG that must pause, resume and keep state. LlamaIndex has the broadest set of data integrations and evaluators. RAGFlow is the pick when you want a self-hosted RAG application with strong document parsing rather than a library, and DSPy when you want to optimise prompts against a metric.

What is the best open-source RAG framework?

All eight frameworks in this comparison carry permissive licences: LlamaIndex, LangChain, LangGraph, DSPy and LightRAG are MIT, while Haystack, RAGFlow, txtai and Kotaemon are Apache 2.0. Licence is therefore rarely the deciding factor. Maintenance is. Check the date of the latest release and the last commit before you commit, because several once-popular projects, including Verba and Cognita, are now archived, and R2R has had no commits since November 2025.

Should I use LlamaIndex or LangChain for RAG?

Both are MIT licensed, actively released and cover loading, splitting, embedding and retrieval. LlamaIndex ships more retrieval and response evaluators in its core package and a large catalogue of data integrations, but its company now focuses mainly on its commercial parsing platform. LangChain pairs with LangGraph for durable, interruptible agentic RAG and has 1.x releases since October 2025. If your RAG system will grow into agents with approvals, LangGraph is the stronger base.

Do I need a framework to build RAG?

Not always. A simple pipeline is a parser, a chunker, an embedding model, a vector database query and a prompt, and many teams write that in a few hundred lines. Frameworks pay off when you need many connectors, swappable models and stores, evaluation hooks, agentic retrieval or a deployment wrapper. The cost is an abstraction layer that changes with every major release, so pin versions and keep your own evaluation set as the arbiter.

How do I evaluate a RAG pipeline before production?

Build a labelled set of real questions with the passages that answer them, then measure two layers separately. For retrieval, track recall, hit rate or mean reciprocal rank. For generation, track faithfulness to the retrieved context, answer relevance and citation correctness. Haystack, LlamaIndex, DSPy and LangChain's openevals package all ship evaluators for these metrics. Re-run the set after every change to chunking, embeddings, retrieval settings or the model.

Which RAG frameworks are no longer maintained?

As of October 2026, Weaviate's Verba and TrueFoundry's Cognita repositories are archived on GitHub. Microsoft's GraphRAG states that it is largely in maintenance mode and will not add new features, though it still ships bug fixes. R2R from SciPhi is not archived, but its latest release dates from June 2025 and its last commit from November 2025. Kotaemon's release pace has slowed, with the latest release in May 2026.

Filed under
RAGprivate RAGRAG architectureAI evaluationopen-source AIon-premises AI
Private RAG & Search

Evaluate your knowledge stack

Find out how a private RAG and retrieval layer would perform on your data — accuracy, latency, governance, and what to fix before you scale.

Or start free — no credit card →

Keep reading