pgvector, Pinecone, Qdrant, Weaviate, Milvus or OpenSearch? The criteria that actually matter for RAG, a sizing example and how to benchmark on your own data.
In this article
- 01What a vector database does in RAG
- 02The main options
- 03Start with the question: do you already run PostgreSQL?
- 04Criterion 1: Filtering
- 05Criterion 2: Hybrid search and reranking
- 06Criterion 3: Index type and recall
- 07Criterion 4: Memory, quantization and sizing
- 08Criterion 5: Updates, deletes and multi-tenancy
- 09Criterion 6: Operations, residency and cost
- 10Benchmark on your own data
- 11Common mistakes when choosing
- 12A practical decision guide
What a vector database does in RAG
In a retrieval-augmented generation system, documents are split into chunks, each chunk is turned into an embedding vector and the vectors are stored for similarity search. At query time, the question is embedded and the database returns the chunks whose vectors are closest. A vector database makes that nearest-neighbor search fast on millions of vectors, usually through approximate indexes, and stores the metadata needed to filter results.
Choosing one is less dramatic than vendor marketing suggests. Retrieval quality depends far more on parsing, chunking, the embedding model and reranking than on the database. The database decision is mostly about scale, filtering, operations and cost.
The main options
The market falls into three groups:
- Extensions to databases you already run: pgvector for PostgreSQL, vector search in MongoDB Atlas, Redis and the k-NN features of OpenSearch and Elasticsearch
- Dedicated open-source vector databases you can self-host or use as a managed service: Qdrant, Weaviate, Milvus and Chroma
- Fully managed services: Pinecone and the managed editions of the open-source databases above
- Embedded libraries for local or small-scale use: LanceDB, or FAISS when you only need an in-memory index
Start with the question: do you already run PostgreSQL?
For many products, pgvector is the right first choice. Your vectors live next to your application data, transactions keep them consistent, permissions and tenant filters use ordinary SQL and there is no new system to secure, back up and monitor. It supports HNSW and IVFFlat approximate indexes and handles collections into the millions of vectors comfortably on suitable hardware.
Dedicated databases earn their place when the collection grows very large, when filtered search performance becomes a bottleneck, when you need features such as built-in hybrid search, multi-vector or advanced quantization, or when the vector workload would compete with your transactional database for memory.
Criterion 1: Filtering
Real RAG queries are filtered: by tenant, by permission group, by document type, by date. How a database combines filters with approximate search matters. Post-filtering, which finds the nearest vectors and then removes those that fail the filter, can return too few results when the filter is selective. Good engines apply filters during the index search. Test your most selective real filters, such as a single small tenant, not just unfiltered queries.
Criterion 2: Hybrid search and reranking
Vector search alone misses exact matches such as product codes, error codes and names. Hybrid search combines keyword ranking such as BM25 with vector similarity. Some databases, including Weaviate, Qdrant and OpenSearch, support this natively; with pgvector you can combine PostgreSQL full-text search with vector search in one query. Either way, plan for a reranking step after retrieval, which usually improves answer quality more than switching databases.
Criterion 3: Index type and recall
HNSW is the most common approximate index. It builds a layered graph that gives high recall and fast queries at the cost of memory and slower inserts. Its main settings are the number of connections per node (often called m), the build-time search width (ef_construction) and the query-time search width (ef_search or similar). Higher values improve recall and cost speed and memory.
Recall means how many of the true nearest neighbors the approximate search returns. Measure it by comparing approximate results with exact search on a sample of your queries, and tune until recall is high enough that it is no longer the weak link in your retrieval.
Criterion 4: Memory, quantization and sizing
Vectors are large. A float32 vector uses four bytes per dimension, so one million 1,536-dimension vectors take about 6 GB before any index overhead, and HNSW indexes perform best when they fit in memory. That arithmetic drives cost more than anything else.
Quantization reduces it. Scalar quantization stores each dimension in fewer bits, binary quantization in a single bit, and product quantization compresses groups of dimensions. Many systems search quantized vectors first and rescore the top results with full precision, keeping most of the quality at a fraction of the memory. Smaller embedding dimensions, which some embedding models offer, help too.
Criterion 5: Updates, deletes and multi-tenancy
Company documents change. Check how the database handles frequent upserts and deletes, whether deleted vectors are cleaned up efficiently and how re-indexing works when you change embedding models, which requires re-embedding everything. For SaaS products, check how tenants are isolated: a tenant field with filtering, separate collections or namespaces per tenant, or separate deployments for customers who require them.
Criterion 6: Operations, residency and cost
Self-hosting gives control and can be cheaper at scale, but someone must run upgrades, backups, monitoring and capacity planning. Managed services remove that work and price by storage, compute or queries, which can grow quickly. Check where data is hosted if you have residency requirements, whether the service supports private networking and encryption with your keys, and how backups and point-in-time recovery work.
Benchmark on your own data
Published benchmarks rarely match your workload. Run a small benchmark of your own before committing:
- Load a realistic sample of your chunks with real metadata
- Replay 100 or more real questions with your typical filters
- Measure recall against exact search, p95 latency and ingestion time
- Repeat with your most selective filters and concurrent queries
- Estimate monthly cost at your expected scale in a year
Common mistakes when choosing
Teams often regret vector database choices for reasons that have little to do with raw speed:
- Choosing on benchmark headlines measured with no filters, then struggling with real filtered queries
- Storing vectors without the metadata needed for permissions and tenant isolation
- Ignoring the cost of re-embedding everything when a better embedding model arrives
- Adding a new database to the stack when PostgreSQL with pgvector would have been enough
- Skipping backups because the index can be rebuilt, then discovering how long rebuilding takes
A practical decision guide
If you run PostgreSQL and expect up to a few million chunks, start with pgvector. If you already run OpenSearch or Elasticsearch for search, use their vector features. If you expect very large collections, need advanced filtering or quantization or want vector search isolated from your main database, pick a dedicated database such as Qdrant, Weaviate or Milvus, self-hosted or managed. Choose a fully managed service like Pinecone when you want zero operations and accept usage-based pricing.
Keep the retrieval layer behind your own interface so you can switch later. Nexzem's RAG development work includes this kind of benchmark, and our RAG vs fine-tuning comparison helps confirm RAG is the right approach first.


