The best vector database for RAG is the one that fits your corpus size, filtering rules and operations team. In October 2026 that means pgvector for teams already on PostgreSQL, Qdrant or Weaviate for a dedicated self-hosted retrieval service, Milvus or Vespa for billions of vectors, and OpenSearch or Elasticsearch when a search cluster already exists.
Quick picks
Each pick below links to a fuller note further down.
| If you… | Start with | Why |
|---|---|---|
| Already run PostgreSQL and have up to a few million chunks | pgvector | No new system, and permissions, joins and deletes stay in SQL |
| Want a dedicated, self-hosted retrieval service with heavy filtering | Qdrant | Payload filters, sparse and dense vectors, quantization down to 1 bit, tenant-aware sharding |
| Serve many customers or business units from one index | Weaviate or Qdrant | Native shard-per-tenant (Weaviate) or tenant-keyed partitions that can be promoted to their own shards (Qdrant) |
| Expect hundreds of millions to billions of vectors | Milvus or Vespa | Distributed architectures that scale storage and query nodes separately |
| Have a tuned OpenSearch or Elasticsearch cluster | The engine you run | Keyword relevance and vector search in one query, one operations model |
| Are prototyping on one machine | Chroma or LanceDB | Embedded, installs with pip, no server to run |
| Have no operations capacity and data may leave your network | Pinecone | Serverless, with a bring-your-own-cloud option on the Enterprise plan |
If you are new to the category, our explainer on what a vector database does covers the concepts. This page is the buyer’s shortlist.
How to choose a vector database for RAG
Vendor benchmarks test each engine on its own settings, so they rarely settle the choice. Seven questions usually do.
- Where may the vectors live? Embeddings and their stored text are derived from your documents, so treat the index like the documents it came from. If those documents may not leave your network, the index cannot either, and a managed service is out unless it runs in your own cloud account.
- Do you need hybrid search? Dense vectors miss exact tokens such as contract numbers and error codes. Our piece on combining keyword and vector retrieval shows why. Check whether fusion happens inside the engine or in your application code.
- How do filters behave under load? RAG queries almost always carry filters for source, date, language or access rights. Some engines filter after the approximate search and can return fewer results than you asked for; others plan filters into the index scan. Our guide to metadata filters in private RAG explains the permission side.
- How many vectors, and how fast do they change? A nightly rebuild and a live sync from a ticketing system stress an engine differently.
- How much memory can you afford? Quantization trades a little recall for a large cut in RAM. Most engines now rescore with full-precision vectors to win the recall back.
- Do you have tenants? One index per tenant, one partition per tenant and one shared index with a tenant filter behave very differently once some tenants are a thousand times larger than others.
- Who runs it? A PostgreSQL extension, a single binary, an embedded library and a Kubernetes cluster with five dependencies are different jobs for an operations team.
The ten vector databases at a glance
Versions and features are as documented on 6 October 2026. Where a cell shows a dash, the pages we read did not document that capability, so treat it as unconfirmed rather than missing.
| Database | Licence | Latest release | Hybrid search | Quantization | Multitenancy | How it runs |
|---|---|---|---|---|---|---|
| pgvector | PostgreSQL License | 0.8.7 (1 Oct 2026) | Postgres full-text search fused with RRF or a cross-encoder | Half precision, binary via expression indexes | List partitions or separate tables | Extension for PostgreSQL 13+ |
| Qdrant | Apache 2.0 | 1.19.2 (5 Oct 2026) | Dense and sparse prefetches fused with RRF or DBSF | Scalar, binary (1, 1.5, 2-bit), product, TurboQuant | is_tenant payload index; large tenants promoted to own shards | Single binary; distributed with sharding and replication |
| Milvus | Apache 2.0 | 3.0.2 (20 Sep 2026) | BM25 full-text and learned sparse vectors next to dense | Quantized IVF, HNSW and DiskANN variants; GPU indexes | Database, collection, partition or partition key | Milvus Lite, standalone, or distributed on Kubernetes |
| Weaviate | BSD 3-Clause; Enterprise Edition features need a key from 1.40 | 1.39.9 (5 Oct 2026) | BM25 plus vector, relativeScoreFusion by default | RQ (recommended), PQ, BQ, SQ | One shard per tenant; active, inactive or offloaded | Docker or Kubernetes; Weaviate Cloud |
| Chroma | Apache 2.0 | 1.5.9 (5 May 2026) | Sparse BM25 or SPLADE with RRF in the Search API | – | Tenants and databases | Embedded, single-node server, or distributed |
| OpenSearch | Apache 2.0 | 3.9.0 (29 Sep 2026) | Hybrid query with normalization or RRF processors | Scalar and binary (1, 2 or 4 bits); 32x on-disk mode | – | Cluster with Faiss, Lucene or JVector engines |
| Elasticsearch | AGPLv3, SSPL or Elastic License 2.0 | 9.5.4 (15 Sep 2026) | RRF and linear retrievers over BM25 and kNN | int8, int4, BBQ; DiskBBQ needs Enterprise | – | Cluster, or Elastic Cloud |
| LanceDB | Apache 2.0 | 0.39.0 (17 Sep 2026) | Full-text index plus vectors, RRF reranker by default | IVF_PQ, IVF_RQ (RaBitQ), IVF_HNSW_SQ | – | Embedded library over Lance files; Cloud and Enterprise |
| Redis | RSALv2, SSPLv1 or AGPLv3 (Redis 8+) | 8.10.2 (17 Sep 2026) | FT.HYBRID from 8.4 with RRF or linear fusion | INT8 and UINT8 types; SVS-VAMANA compression | – | In-memory server or cluster |
| Vespa | Apache 2.0 | 8.763.13 (5 Oct 2026) | Text and nearestNeighbor in one query, multi-phase ranking | Binarized vectors (32x) with full-precision re-ranking | – | Config, container and content clusters; Vespa Cloud |
Notes on each database
pgvector and pgvectorscale
pgvector adds a vector type and two approximate indexes, HNSW and IVFFlat, to PostgreSQL 13 and later. Indexes cover up to 2,000 dimensions for vector, 4,000 for halfvec and 64,000 for bit, and sparsevec handles sparse embeddings. Its README is candid about filtering: with approximate indexes the filter is applied after the index scan, so version 0.8.0 added iterative index scans that keep reading until enough rows match. For tenants it recommends list partitioning or separate tables, because a shared index lets one tenant’s vectors affect another’s recall. Hybrid search uses Postgres full-text search, combined with reciprocal rank fusion or a cross-encoder in your query. Scale out with read replicas, or shard with Citus or PgDog. Timescale’s pgvectorscale, also under the PostgreSQL License, adds a StreamingDiskANN index, statistical binary quantization and label-based filtering.
Qdrant
Qdrant is written in Rust and licensed Apache 2.0. Points carry a JSON payload that can be filtered with keyword, full-text, numeric, geo and boolean conditions. Hybrid search runs as prefetch sub-queries over dense and sparse vectors, fused with RRF or distribution-based score fusion, and multivector models such as ColBERT can rescore candidates. The quantization guide lists scalar (4x), binary at 1, 1.5 and 2 bits (32x, 24x and 16x), product quantization (up to 64x) and TurboQuant (up to 32x). Multitenancy has three documented patterns: a tenant field marked is_tenant (since 1.11), custom sharding per tenant, and tiered multitenancy (since 1.16), where large tenants move to dedicated shards. Qdrant Edge runs the engine inside your application process.
Milvus
Milvus, an LF AI & Data Foundation project written in Go and C++, is the most scale-oriented engine on the list. Version 3.0.0 shipped on 29 July 2026 with external collections that index lakehouse files in place, online schema changes with backfill, a rebuilt sparse index and long-text fields. It supports BM25 full-text search and learned sparse embeddings such as SPLADE and BGE-M3 alongside dense vectors, with reranking across result sets. Multitenancy can be enforced at database, collection, partition or partition-key level. The price is operational: a cluster needs etcd for metadata, object storage such as MinIO or S3, a write-ahead log on Woodpecker, Kafka or Pulsar, plus coordinator, streaming, query and data nodes behind stateless proxies. Milvus Lite and standalone mode cover development and small deployments.
Weaviate
Weaviate, written in Go, combines BM25 and vector search in one hybrid query; relativeScoreFusion has been the default since 1.24, with rankedFusion as the alternative. Compression options are rotational quantization (marked recommended), PQ, BQ and SQ, and multi-vector embeddings can be encoded with MUVERA. Multitenancy is native: each tenant lives on its own shard and can be active, inactive or offloaded to cold storage. Read the licence carefully. Most of the repository is BSD 3-Clause, but version 1.40, in release candidate since September 2026, introduces an Enterprise Edition. Shard self-recovery, namespaces and deduplicated backups then require a commercial licence key, while existing features stay free to use.
Chroma
Chroma’s core is written in Rust, licensed Apache 2.0, and starts in memory with four API calls. Metadata filters and document filters ($contains, $regex and their negations) work out of the box. Its single-node performance guide tested collections up to about 7 million embeddings and gives a sizing rule of roughly 0.245 million embeddings per GB of RAM. Distributed Chroma is open source and is what Chroma Cloud runs. BM25 and SPLADE sparse vectors with RRF ranking are documented in the Search API, whose pages sit in the Cloud section of the docs, so confirm what your self-hosted version supports. The latest release on GitHub and PyPI is 1.5.9 from 5 May 2026, although commits continue.
OpenSearch
OpenSearch, licensed Apache 2.0, stores vectors in knn_vector fields. Faiss is the default engine, Lucene offers efficient filtering, NMSLIB is deprecated, and a JVector plugin adds DiskANN-style search. The on_disk mode, introduced in 2.17, compresses vectors 32x by default and rescores with full-precision vectors read from disk; scalar quantization became its default in 3.6, with binary quantization at 1, 2 or 4 bits as the alternative. Hybrid search combines a hybrid query with a search pipeline that normalizes scores or applies reciprocal rank fusion.
Elasticsearch
Elasticsearch’s default licence lets you choose AGPLv3, SSPL or the Elastic License 2.0, and code in the x-pack folder is Elastic License 2.0 only. For dense_vector fields it offers int8_hnsw (4x smaller), int4_hnsw (8x) and BBQ variants (32x); from 9.1, float fields of 384 or more dimensions default to bbq_hnsw. DiskBBQ (bbq_disk), a clustering index that keeps vectors on disk, requires an Enterprise subscription and becomes the default in 9.4 where licensed. Hybrid search uses RRF or linear retrievers over BM25 and kNN results.
LanceDB
LanceDB is an embedded database written in Rust over the Lance columnar format, so data lives as files on local disk or object storage and there is no server to operate. It indexes vectors with IVF_PQ, IVF_RQ (RaBitQ, about 1/32 of raw size) and IVF_HNSW_SQ, filters with SQL-style where clauses, and runs hybrid search by fusing a full-text index with vector results through RRFReranker by default, in the open-source edition. LanceDB Cloud and Enterprise add managed serving with no servers for you to run.
Redis
Redis 8 folded the Query Engine into Redis Open Source and added AGPLv3 to its RSALv2 and SSPLv1 licence options. Vector fields use FLAT, HNSW or, since 8.2, SVS-VAMANA indexes, with filters on tag, numeric, full-text and geo fields. FT.HYBRID, available since 8.4, fuses a BM25 text query with vector similarity using RRF or a linear combination. One caveat from the docs: Intel’s LVQ and LeanVec compression for SVS-VAMANA are not available in Redis Open Source, which falls back to 8-bit scalar quantization. Redis keeps data in memory, so RAM sets the cost of a large corpus.
Vespa
Vespa, Apache 2.0, is a full serving engine for text, vectors, tensors and structured data. A single query can combine keyword matching with the nearestNeighbor operator and filters, and multi-phase rank profiles re-rank the top hits with richer signals. Its binarization guide packs float vectors into int8 tensors for 32x compression, searches by Hamming distance and re-ranks with full-precision vectors in a later phase. The architecture has a configuration cluster, stateless container clusters and content clusters that rebalance data automatically when nodes change.
Managed options: Pinecone and the hosted editions
Pinecone is the reference managed vector database. New indexes default to document indexes that hold dense vectors, sparse vectors and full-text search in one schema. Namespaces partition an index, one per customer if you need isolation, with limits from 100 on the Starter plan to 1,000,000 on Enterprise. Metadata filters support operators such as $eq, $in and $or on up to 40 KB of metadata per record. For data residency, Pinecone BYOC runs the data plane inside your own AWS, GCP or Azure account and requires the Enterprise plan; the control plane stays with Pinecone.
Most open-source engines above also sell a hosted edition, from Zilliz Cloud for Milvus to Qdrant, Weaviate, Chroma, Elastic, LanceDB and Vespa clouds. Hosting removes the operations work and moves your embeddings onto the vendor’s infrastructure.
Which vector database by data size and team
| Corpus size (vectors) | Team | Good starting point | Watch for |
|---|---|---|---|
| Up to a few million | Application team already on PostgreSQL | pgvector | Filtered queries returning too few rows; enable iterative scans |
| Up to a few million | Data scientists prototyping | Chroma or LanceDB, embedded | Moving to a server or cloud edition later |
| Tens of millions | Platform team, many filters or tenants | Qdrant or Weaviate | Quantization settings and rescoring; Weaviate’s licence-gated features from 1.40 |
| Tens of millions | Search team with an existing cluster | OpenSearch or Elasticsearch | Memory per node; Elastic features that need a paid subscription |
| Hundreds of millions to billions | Platform team on Kubernetes | Milvus or Vespa | Etcd, object storage and WAL for Milvus; ranking expertise for Vespa |
| Any size, latency-critical | Team already operating Redis | Redis | RAM cost; licence choice |
| Any size, no operations capacity | Product team | Pinecone or a hosted edition | Where the embeddings are stored |
Memory arithmetic for an index
Raw vector storage is chunks × dimensions × bytes per value. For 10 million chunks of 1,024-dimension embeddings:
| Precision | Bytes per value | Raw vectors |
|---|---|---|
| float32 | 4 | ≈ 41 GB |
| float16 or bfloat16 | 2 | ≈ 20.5 GB |
| int8 | 1 | ≈ 10.2 GB |
| 1-bit binary | 0.125 | ≈ 1.3 GB |
These figures exclude the HNSW graph or partition index, payload and metadata indexes, replicas and headroom. Binary and on-disk modes usually keep the full-precision vectors on disk for rescoring, so disk grows even as RAM shrinks. Measure recall on your own questions before you commit to a compression level.
Running a vector database inside a private network
A self-hosted index only keeps data private if everything around it does too.
- Turn off phone-home telemetry. Qdrant’s open-source container sends anonymised usage statistics by default; set
QDRANT__TELEMETRY_DISABLED=true. Weaviate sends telemetry every 24 hours, including a machine ID, version and object counts; setDISABLE_TELEMETRY=true. Chroma stopped collecting product telemetry in 1.5.4. - Embed inside the boundary. An on-premises index fed by a cloud embedding API still ships every chunk out at indexing time and every question out at query time.
- Carry permissions into the index. Store access-control attributes with each chunk and filter on them at query time. Shared indexes need a tenant strategy from day one; our note on multi-tenant RAG data architecture compares the options.
- Make deletion reach the vectors. When a source document is deleted or a person exercises an erasure right, the chunks, vectors and caches must go too. Our guide to erasure in private RAG covers the mechanics.
- Pin versions. Licences in this market keep shifting: Redis 8 added AGPLv3, Elasticsearch offers AGPLv3 next to SSPL and its own licence, and Weaviate 1.40 adds licence-keyed features. Record the version and licence you approved, and read the release notes before every upgrade.
The database is one layer of the pipeline. The framework that loads, chunks and queries it is the other; our companion list of RAG frameworks compares those.
How VDF AI fits
In VDF AI Chat, retrieval is its own layer: connectors, chunking, embedding, the vector store and the permission filter. The product page describes a local vector store filtered by the reader’s permissions at query time, and every answer cites its sources. In VDF AI Data, teams can also build their own vector indexes from a database column, a curated feature list, a Confluence space, Jira project or GitHub repository, or a file collection, choosing chunk size, overlap and the embedding model. Once an index is ready, Chat, Agents and Networks can search it.
Indexes are snapshots that you rebuild on your own schedule, and admins can see that an index was built without seeing the chunks inside it. Sensitive knowledge sources can be restricted so that only named people can attach them to agents. VDF AI deploys on-premises, in a private cloud or air-gapped, with vector storage inside that perimeter. The product pages do not name the vector engine underneath, so if your team has standardised on one of the databases above, ask how it plugs in during evaluation.
Sources
Verified 6 October 2026.
- pgvector repository and README
- pgvector changelog
- pgvectorscale repository
- Qdrant repository
- Qdrant hybrid queries
- Qdrant quantization
- Qdrant multitenancy
- Qdrant usage statistics
- Milvus repository
- Milvus 3.0.0 release notes
- Milvus architecture overview
- Weaviate repository and licence
- Weaviate releases
- Weaviate hybrid search
- Weaviate vector compression
- Weaviate multi-tenancy
- Weaviate telemetry
- Chroma repository
- Chroma open source and telemetry
- Chroma single-node performance
- Chroma full-text and regex search
- Chroma hybrid search with RRF
- OpenSearch repository
- OpenSearch methods and engines
- OpenSearch disk-based vector search
- OpenSearch hybrid search
- Elasticsearch licence
- Elasticsearch dense vector field
- Elasticsearch BBQ and DiskBBQ
- Elasticsearch reciprocal rank fusion
- LanceDB repository
- LanceDB vector indexes
- LanceDB hybrid search
- Redis licence
- Redis 8 general availability
- Redis vector search concepts
- Redis FT.HYBRID command
- Vespa repository
- Vespa architecture overview
- Vespa nearest neighbor search
- Vespa binarizing vectors
- Pinecone indexing overview
- Pinecone hybrid search
- Pinecone BYOC