RAG

Best Vector Database for RAG (2026): Open-Source and Self-Hosted Options

Ten vector databases checked against their own documentation and GitHub repositories in October 2026: pgvector, Qdrant, Milvus, Weaviate, Chroma, OpenSearch, Elasticsearch, LanceDB, Redis and Vespa, plus Pinecone as the managed reference. Compared on licence, hybrid search, filtering, quantization, multitenancy and the work it takes to run each one yourself.

The best vector database for RAG is the one that fits your corpus size, filtering rules and operations team. In October 2026 that means pgvector for teams already on PostgreSQL, Qdrant or Weaviate for a dedicated self-hosted retrieval service, Milvus or Vespa for billions of vectors, and OpenSearch or Elasticsearch when a search cluster already exists.

Quick picks

Each pick below links to a fuller note further down.

If you…Start withWhy
Already run PostgreSQL and have up to a few million chunkspgvectorNo new system, and permissions, joins and deletes stay in SQL
Want a dedicated, self-hosted retrieval service with heavy filteringQdrantPayload filters, sparse and dense vectors, quantization down to 1 bit, tenant-aware sharding
Serve many customers or business units from one indexWeaviate or QdrantNative shard-per-tenant (Weaviate) or tenant-keyed partitions that can be promoted to their own shards (Qdrant)
Expect hundreds of millions to billions of vectorsMilvus or VespaDistributed architectures that scale storage and query nodes separately
Have a tuned OpenSearch or Elasticsearch clusterThe engine you runKeyword relevance and vector search in one query, one operations model
Are prototyping on one machineChroma or LanceDBEmbedded, installs with pip, no server to run
Have no operations capacity and data may leave your networkPineconeServerless, with a bring-your-own-cloud option on the Enterprise plan

If you are new to the category, our explainer on what a vector database does covers the concepts. This page is the buyer’s shortlist.

How to choose a vector database for RAG

Vendor benchmarks test each engine on its own settings, so they rarely settle the choice. Seven questions usually do.

  1. Where may the vectors live? Embeddings and their stored text are derived from your documents, so treat the index like the documents it came from. If those documents may not leave your network, the index cannot either, and a managed service is out unless it runs in your own cloud account.
  2. Do you need hybrid search? Dense vectors miss exact tokens such as contract numbers and error codes. Our piece on combining keyword and vector retrieval shows why. Check whether fusion happens inside the engine or in your application code.
  3. How do filters behave under load? RAG queries almost always carry filters for source, date, language or access rights. Some engines filter after the approximate search and can return fewer results than you asked for; others plan filters into the index scan. Our guide to metadata filters in private RAG explains the permission side.
  4. How many vectors, and how fast do they change? A nightly rebuild and a live sync from a ticketing system stress an engine differently.
  5. How much memory can you afford? Quantization trades a little recall for a large cut in RAM. Most engines now rescore with full-precision vectors to win the recall back.
  6. Do you have tenants? One index per tenant, one partition per tenant and one shared index with a tenant filter behave very differently once some tenants are a thousand times larger than others.
  7. Who runs it? A PostgreSQL extension, a single binary, an embedded library and a Kubernetes cluster with five dependencies are different jobs for an operations team.

The ten vector databases at a glance

Versions and features are as documented on 6 October 2026. Where a cell shows a dash, the pages we read did not document that capability, so treat it as unconfirmed rather than missing.

DatabaseLicenceLatest releaseHybrid searchQuantizationMultitenancyHow it runs
pgvectorPostgreSQL License0.8.7 (1 Oct 2026)Postgres full-text search fused with RRF or a cross-encoderHalf precision, binary via expression indexesList partitions or separate tablesExtension for PostgreSQL 13+
QdrantApache 2.01.19.2 (5 Oct 2026)Dense and sparse prefetches fused with RRF or DBSFScalar, binary (1, 1.5, 2-bit), product, TurboQuantis_tenant payload index; large tenants promoted to own shardsSingle binary; distributed with sharding and replication
MilvusApache 2.03.0.2 (20 Sep 2026)BM25 full-text and learned sparse vectors next to denseQuantized IVF, HNSW and DiskANN variants; GPU indexesDatabase, collection, partition or partition keyMilvus Lite, standalone, or distributed on Kubernetes
WeaviateBSD 3-Clause; Enterprise Edition features need a key from 1.401.39.9 (5 Oct 2026)BM25 plus vector, relativeScoreFusion by defaultRQ (recommended), PQ, BQ, SQOne shard per tenant; active, inactive or offloadedDocker or Kubernetes; Weaviate Cloud
ChromaApache 2.01.5.9 (5 May 2026)Sparse BM25 or SPLADE with RRF in the Search API–Tenants and databasesEmbedded, single-node server, or distributed
OpenSearchApache 2.03.9.0 (29 Sep 2026)Hybrid query with normalization or RRF processorsScalar and binary (1, 2 or 4 bits); 32x on-disk mode–Cluster with Faiss, Lucene or JVector engines
ElasticsearchAGPLv3, SSPL or Elastic License 2.09.5.4 (15 Sep 2026)RRF and linear retrievers over BM25 and kNNint8, int4, BBQ; DiskBBQ needs Enterprise–Cluster, or Elastic Cloud
LanceDBApache 2.00.39.0 (17 Sep 2026)Full-text index plus vectors, RRF reranker by defaultIVF_PQ, IVF_RQ (RaBitQ), IVF_HNSW_SQ–Embedded library over Lance files; Cloud and Enterprise
RedisRSALv2, SSPLv1 or AGPLv3 (Redis 8+)8.10.2 (17 Sep 2026)FT.HYBRID from 8.4 with RRF or linear fusionINT8 and UINT8 types; SVS-VAMANA compression–In-memory server or cluster
VespaApache 2.08.763.13 (5 Oct 2026)Text and nearestNeighbor in one query, multi-phase rankingBinarized vectors (32x) with full-precision re-ranking–Config, container and content clusters; Vespa Cloud

Notes on each database

pgvector and pgvectorscale

pgvector adds a vector type and two approximate indexes, HNSW and IVFFlat, to PostgreSQL 13 and later. Indexes cover up to 2,000 dimensions for vector, 4,000 for halfvec and 64,000 for bit, and sparsevec handles sparse embeddings. Its README is candid about filtering: with approximate indexes the filter is applied after the index scan, so version 0.8.0 added iterative index scans that keep reading until enough rows match. For tenants it recommends list partitioning or separate tables, because a shared index lets one tenant’s vectors affect another’s recall. Hybrid search uses Postgres full-text search, combined with reciprocal rank fusion or a cross-encoder in your query. Scale out with read replicas, or shard with Citus or PgDog. Timescale’s pgvectorscale, also under the PostgreSQL License, adds a StreamingDiskANN index, statistical binary quantization and label-based filtering.

Qdrant

Qdrant is written in Rust and licensed Apache 2.0. Points carry a JSON payload that can be filtered with keyword, full-text, numeric, geo and boolean conditions. Hybrid search runs as prefetch sub-queries over dense and sparse vectors, fused with RRF or distribution-based score fusion, and multivector models such as ColBERT can rescore candidates. The quantization guide lists scalar (4x), binary at 1, 1.5 and 2 bits (32x, 24x and 16x), product quantization (up to 64x) and TurboQuant (up to 32x). Multitenancy has three documented patterns: a tenant field marked is_tenant (since 1.11), custom sharding per tenant, and tiered multitenancy (since 1.16), where large tenants move to dedicated shards. Qdrant Edge runs the engine inside your application process.

Milvus

Milvus, an LF AI & Data Foundation project written in Go and C++, is the most scale-oriented engine on the list. Version 3.0.0 shipped on 29 July 2026 with external collections that index lakehouse files in place, online schema changes with backfill, a rebuilt sparse index and long-text fields. It supports BM25 full-text search and learned sparse embeddings such as SPLADE and BGE-M3 alongside dense vectors, with reranking across result sets. Multitenancy can be enforced at database, collection, partition or partition-key level. The price is operational: a cluster needs etcd for metadata, object storage such as MinIO or S3, a write-ahead log on Woodpecker, Kafka or Pulsar, plus coordinator, streaming, query and data nodes behind stateless proxies. Milvus Lite and standalone mode cover development and small deployments.

Weaviate

Weaviate, written in Go, combines BM25 and vector search in one hybrid query; relativeScoreFusion has been the default since 1.24, with rankedFusion as the alternative. Compression options are rotational quantization (marked recommended), PQ, BQ and SQ, and multi-vector embeddings can be encoded with MUVERA. Multitenancy is native: each tenant lives on its own shard and can be active, inactive or offloaded to cold storage. Read the licence carefully. Most of the repository is BSD 3-Clause, but version 1.40, in release candidate since September 2026, introduces an Enterprise Edition. Shard self-recovery, namespaces and deduplicated backups then require a commercial licence key, while existing features stay free to use.

Chroma

Chroma’s core is written in Rust, licensed Apache 2.0, and starts in memory with four API calls. Metadata filters and document filters ($contains, $regex and their negations) work out of the box. Its single-node performance guide tested collections up to about 7 million embeddings and gives a sizing rule of roughly 0.245 million embeddings per GB of RAM. Distributed Chroma is open source and is what Chroma Cloud runs. BM25 and SPLADE sparse vectors with RRF ranking are documented in the Search API, whose pages sit in the Cloud section of the docs, so confirm what your self-hosted version supports. The latest release on GitHub and PyPI is 1.5.9 from 5 May 2026, although commits continue.

OpenSearch

OpenSearch, licensed Apache 2.0, stores vectors in knn_vector fields. Faiss is the default engine, Lucene offers efficient filtering, NMSLIB is deprecated, and a JVector plugin adds DiskANN-style search. The on_disk mode, introduced in 2.17, compresses vectors 32x by default and rescores with full-precision vectors read from disk; scalar quantization became its default in 3.6, with binary quantization at 1, 2 or 4 bits as the alternative. Hybrid search combines a hybrid query with a search pipeline that normalizes scores or applies reciprocal rank fusion.

Elasticsearch

Elasticsearch’s default licence lets you choose AGPLv3, SSPL or the Elastic License 2.0, and code in the x-pack folder is Elastic License 2.0 only. For dense_vector fields it offers int8_hnsw (4x smaller), int4_hnsw (8x) and BBQ variants (32x); from 9.1, float fields of 384 or more dimensions default to bbq_hnsw. DiskBBQ (bbq_disk), a clustering index that keeps vectors on disk, requires an Enterprise subscription and becomes the default in 9.4 where licensed. Hybrid search uses RRF or linear retrievers over BM25 and kNN results.

LanceDB

LanceDB is an embedded database written in Rust over the Lance columnar format, so data lives as files on local disk or object storage and there is no server to operate. It indexes vectors with IVF_PQ, IVF_RQ (RaBitQ, about 1/32 of raw size) and IVF_HNSW_SQ, filters with SQL-style where clauses, and runs hybrid search by fusing a full-text index with vector results through RRFReranker by default, in the open-source edition. LanceDB Cloud and Enterprise add managed serving with no servers for you to run.

Redis

Redis 8 folded the Query Engine into Redis Open Source and added AGPLv3 to its RSALv2 and SSPLv1 licence options. Vector fields use FLAT, HNSW or, since 8.2, SVS-VAMANA indexes, with filters on tag, numeric, full-text and geo fields. FT.HYBRID, available since 8.4, fuses a BM25 text query with vector similarity using RRF or a linear combination. One caveat from the docs: Intel’s LVQ and LeanVec compression for SVS-VAMANA are not available in Redis Open Source, which falls back to 8-bit scalar quantization. Redis keeps data in memory, so RAM sets the cost of a large corpus.

Vespa

Vespa, Apache 2.0, is a full serving engine for text, vectors, tensors and structured data. A single query can combine keyword matching with the nearestNeighbor operator and filters, and multi-phase rank profiles re-rank the top hits with richer signals. Its binarization guide packs float vectors into int8 tensors for 32x compression, searches by Hamming distance and re-ranks with full-precision vectors in a later phase. The architecture has a configuration cluster, stateless container clusters and content clusters that rebalance data automatically when nodes change.

Managed options: Pinecone and the hosted editions

Pinecone is the reference managed vector database. New indexes default to document indexes that hold dense vectors, sparse vectors and full-text search in one schema. Namespaces partition an index, one per customer if you need isolation, with limits from 100 on the Starter plan to 1,000,000 on Enterprise. Metadata filters support operators such as $eq, $in and $or on up to 40 KB of metadata per record. For data residency, Pinecone BYOC runs the data plane inside your own AWS, GCP or Azure account and requires the Enterprise plan; the control plane stays with Pinecone.

Most open-source engines above also sell a hosted edition, from Zilliz Cloud for Milvus to Qdrant, Weaviate, Chroma, Elastic, LanceDB and Vespa clouds. Hosting removes the operations work and moves your embeddings onto the vendor’s infrastructure.

Which vector database by data size and team

Corpus size (vectors)TeamGood starting pointWatch for
Up to a few millionApplication team already on PostgreSQLpgvectorFiltered queries returning too few rows; enable iterative scans
Up to a few millionData scientists prototypingChroma or LanceDB, embeddedMoving to a server or cloud edition later
Tens of millionsPlatform team, many filters or tenantsQdrant or WeaviateQuantization settings and rescoring; Weaviate’s licence-gated features from 1.40
Tens of millionsSearch team with an existing clusterOpenSearch or ElasticsearchMemory per node; Elastic features that need a paid subscription
Hundreds of millions to billionsPlatform team on KubernetesMilvus or VespaEtcd, object storage and WAL for Milvus; ranking expertise for Vespa
Any size, latency-criticalTeam already operating RedisRedisRAM cost; licence choice
Any size, no operations capacityProduct teamPinecone or a hosted editionWhere the embeddings are stored

Memory arithmetic for an index

Raw vector storage is chunks × dimensions × bytes per value. For 10 million chunks of 1,024-dimension embeddings:

PrecisionBytes per valueRaw vectors
float324≈ 41 GB
float16 or bfloat162≈ 20.5 GB
int81≈ 10.2 GB
1-bit binary0.125≈ 1.3 GB

These figures exclude the HNSW graph or partition index, payload and metadata indexes, replicas and headroom. Binary and on-disk modes usually keep the full-precision vectors on disk for rescoring, so disk grows even as RAM shrinks. Measure recall on your own questions before you commit to a compression level.

Running a vector database inside a private network

A self-hosted index only keeps data private if everything around it does too.

  • Turn off phone-home telemetry. Qdrant’s open-source container sends anonymised usage statistics by default; set QDRANT__TELEMETRY_DISABLED=true. Weaviate sends telemetry every 24 hours, including a machine ID, version and object counts; set DISABLE_TELEMETRY=true. Chroma stopped collecting product telemetry in 1.5.4.
  • Embed inside the boundary. An on-premises index fed by a cloud embedding API still ships every chunk out at indexing time and every question out at query time.
  • Carry permissions into the index. Store access-control attributes with each chunk and filter on them at query time. Shared indexes need a tenant strategy from day one; our note on multi-tenant RAG data architecture compares the options.
  • Make deletion reach the vectors. When a source document is deleted or a person exercises an erasure right, the chunks, vectors and caches must go too. Our guide to erasure in private RAG covers the mechanics.
  • Pin versions. Licences in this market keep shifting: Redis 8 added AGPLv3, Elasticsearch offers AGPLv3 next to SSPL and its own licence, and Weaviate 1.40 adds licence-keyed features. Record the version and licence you approved, and read the release notes before every upgrade.

The database is one layer of the pipeline. The framework that loads, chunks and queries it is the other; our companion list of RAG frameworks compares those.

How VDF AI fits

In VDF AI Chat, retrieval is its own layer: connectors, chunking, embedding, the vector store and the permission filter. The product page describes a local vector store filtered by the reader’s permissions at query time, and every answer cites its sources. In VDF AI Data, teams can also build their own vector indexes from a database column, a curated feature list, a Confluence space, Jira project or GitHub repository, or a file collection, choosing chunk size, overlap and the embedding model. Once an index is ready, Chat, Agents and Networks can search it.

Indexes are snapshots that you rebuild on your own schedule, and admins can see that an index was built without seeing the chunks inside it. Sensitive knowledge sources can be restricted so that only named people can attach them to agents. VDF AI deploys on-premises, in a private cloud or air-gapped, with vector storage inside that perimeter. The product pages do not name the vector engine underneath, so if your team has standardised on one of the databases above, ask how it plugs in during evaluation.

Sources

Verified 6 October 2026.

Frequently asked questions

What is the best vector database for RAG?

There is no single winner; the right pick follows your data size and the systems you already run. Teams with PostgreSQL and a corpus of a few million chunks get far with pgvector. Qdrant and Weaviate suit a dedicated retrieval service with heavy filtering and many tenants. Milvus and Vespa are built for hundreds of millions to billions of vectors on a cluster. Teams with a tuned OpenSearch or Elasticsearch cluster can add vectors to the engine they already operate.

What is the best open-source vector database?

Judge it by licence first. Qdrant, Milvus, Chroma, OpenSearch, LanceDB and Vespa are Apache 2.0, and pgvector uses the permissive PostgreSQL License. Weaviate is BSD 3-Clause, but version 1.40 adds Enterprise Edition features that need a commercial licence key. Elasticsearch and Redis 8 let you choose AGPLv3 among their licences, which is OSI-approved but carries network copyleft obligations your legal team should review. Within the permissive group, choose by scale, filtering needs and the operations work your team can absorb.

Can PostgreSQL with pgvector replace a dedicated vector database?

For many RAG systems, yes. pgvector adds HNSW and IVFFlat indexes, half-precision, binary and sparse vectors, and iterative index scans that keep filtered queries from returning too few rows. Hybrid search combines Postgres full-text search with reciprocal rank fusion in SQL. The trade-off appears at large scale and high write rates, where you shard with Citus or similar tools yourself. A dedicated engine starts to pay off when the vector workload outgrows the database it shares.

How much memory does a vector database need?

Start from the raw vectors: number of chunks times dimensions times bytes per value. Ten million chunks of 1,024-dimension float32 embeddings take about 41 GB before any index structures or metadata. Half precision halves that, 8-bit quantization cuts it to about 10 GB, and 1-bit binary quantization to about 1.3 GB, usually with full-precision vectors kept on disk for rescoring. Add the graph or partition index, payload indexes and headroom, then measure on your own data.

Do I need hybrid search for RAG?

Most enterprise corpora benefit from it. Dense vectors find passages that mean the same thing in different words, but they often miss exact tokens such as policy numbers, part codes, error strings and names. Keyword or sparse retrieval catches those. Every engine in this comparison now documents some form of hybrid search, usually fusing keyword and vector rankings with reciprocal rank fusion, so test both modes on real questions from your users before you decide.

Is a self-hosted vector database safe for an air-gapped network?

It can be, with three checks. First, switch off phone-home telemetry: Qdrant and Weaviate send anonymous usage statistics by default and document how to disable them, while Chroma stopped collecting product telemetry in version 1.5.4. Second, generate embeddings with a model that runs inside the network, because a self-hosted index fed by a cloud embedding API still sends every chunk out. Third, mirror container images and packages to an internal registry.

Filed under
vector databasevector searchprivate RAGRAGopen-source AIon-premises AI
Private RAG & Search

Evaluate your knowledge stack

Find out how a private RAG and retrieval layer would perform on your data — accuracy, latency, governance, and what to fix before you scale.

Or start free — no credit card →

Keep reading