Choosing a Vector Database

Every RAG system, semantic search feature, or recommendation engine needs somewhere to store and search embeddings. The market is crowded: extensions to databases you already run (pgvector), search engines with vector support (Elasticsearch, OpenSearch), dedicated vector databases (Pinecone, Qdrant, Weaviate, Milvus), serverless object-storage-backed stores, and embedded libraries (LanceDB, Chroma, FAISS).

There's no universally best option. The right choice depends on scale, filtering and hybrid search needs, the operational capacity of your team, existing infrastructure, latency and cost targets, and data governance. For many teams, the best first vector database is the database they already have.

TL;DR

Quick Example

A pragmatic evaluation harness comparing candidates on your queries:

Decisions based on measured recall, filtered latency at realistic concurrency, and estimated cost are far better than decisions based on marketing benchmarks.

Core Concepts

Categories of Vector Stores

Evaluation Criteria

  1. Scale: vectors now and in two years, dimensions, and growth rate.
  2. Query profile: QPS, latency targets (p95/p99), and top-k.
  3. Filtering: selectivity, multi-tenancy, and permission filters. See metadata filtering.
  4. Hybrid search: BM25 plus vectors, sparse vectors, and reranking integrations.
  5. Freshness and ingestion: write throughput, update and delete handling, and time to searchable.
  6. Memory efficiency: quantization, disk-based indexes, and tiering. See vector quantization.
  7. Operations: managed vs self-hosted, backups, upgrades, monitoring, HA, and the team's expertise.
  8. Security and compliance: tenant isolation, encryption, data residency, and access control.
  9. Cost: infrastructure, replicas, egress, and pricing model (pods vs serverless vs self-hosted).
  10. Ecosystem: client libraries, LangChain and LlamaIndex integrations, and community maturity.

When pgvector Is Enough

PostgreSQL with pgvector handles millions (sometimes tens of millions) of vectors well, with HNSW, halfvec, iterative scans for filters, and partitioning. Benefits: transactional updates alongside business data, SQL joins and filters, row-level security for tenancy, and no new infrastructure. It's the default choice for many SaaS RAG features.

When a Dedicated Engine Pays Off

When a Search Engine Is Best

If keyword relevance, faceting, analyzers, and hybrid ranking matter as much as vectors (e-commerce search, documentation search, log search with semantic features), Elasticsearch, OpenSearch, or Vespa combine both in one query path.

Decision Framework

  1. Prototype: embedded (Chroma, LanceDB) or pgvector locally.
  2. Production, moderate scale, existing Postgres → pgvector.
  3. Production, text relevance plus hybrid ranking is central → Elasticsearch, OpenSearch, or Vespa.
  4. Large scale, filter-heavy, or cost-sensitive vector workloads → Qdrant, Milvus, Weaviate (self-hosted or cloud), or a serverless store.
  5. Minimal operations, willing to pay for managed → Pinecone, Qdrant Cloud, Weaviate Cloud, Zilliz, or your cloud provider's offering.

Re-evaluate as scale and requirements change. The migration path matters more than getting the first choice perfect.

Best Practices

Benchmark With Your Embeddings and Filters

Public benchmarks rarely match your dimensions, data distribution, filters, and concurrency. Run a small, fair bake-off on real queries before committing.

Abstract the Retrieval Layer

Keep vector store calls behind an internal interface (upsert, delete, search with filters), and store source text and metadata in your system of record. Switching engines later then becomes a migration, not a rewrite.

Factor in Operations Honestly

A self-hosted cluster needs monitoring, backups, upgrades, capacity planning, and on-call. Managed services cost more per GB, but may be cheaper overall for small teams.

Plan for Hybrid and Reranking

Most production RAG systems end up with hybrid retrieval plus reranking. Prefer stores that support hybrid natively, or make it easy to combine with a keyword engine.

Common Mistakes

Adopting a New Database for a Small Corpus

Standing up a dedicated vector cluster for 100k chunks adds operational burden with little benefit over pgvector or an embedded store.

Choosing on Unfiltered Benchmarks

Real workloads filter by tenant, language, or permissions. An engine that's fastest unfiltered may degrade badly with selective filters. Test filtered queries.

Locking Data Into One Vendor Format

Without source text and a portable export path, re-embedding and migration become painful. Keep canonical data in your own storage.

FAQ

Do I need a dedicated vector database?

Often not at first. If you already run PostgreSQL, Elasticsearch/OpenSearch, MongoDB Atlas, or Redis, their vector capabilities can handle many workloads with less operational overhead. Dedicated vector databases become compelling at large scale, with complex filtering at low latency, or when you need advanced vector-specific features.

pgvector vs Pinecone?

pgvector keeps vectors alongside relational data with SQL, transactions, and row-level security, at no extra infrastructure cost if you run Postgres. It's strong up to millions to tens of millions of vectors. Pinecone is a fully managed, vector-first service that scales further with less tuning, but it's a separate system and a paid service. Choose by scale, operational preferences, and integration needs.

Which vector database is fastest?

It depends on dataset, dimensions, filters, recall target, hardware, and configuration. Rankings on public benchmarks shift frequently. Measure recall-latency trade-offs on your own data and query patterns, including filtered queries and concurrent ingestion.

Can Elasticsearch replace a vector database?

For many use cases, yes. It supports dense and sparse vectors, approximate kNN with filters, quantization, and hybrid ranking (RRF) combined with mature text search and aggregations. Dedicated vector databases may still win on memory efficiency or vector-specific features at very large scale.

Related Topics

References