Scaling Vector Search
A prototype with 50,000 chunks in a local vector store behaves nothing like production with 200 million vectors, thousands of queries per second, continuous ingestion, strict tenant isolation, and a budget. Scaling vector databases is mostly about a few resources: memory (ANN indexes want RAM), ingestion throughput (embedding and indexing new content), query throughput and latency, and the operational lifecycle of embeddings, including the day you change embedding models and must re-embed everything.
Good capacity planning, the right index and quantization choices, sharding and replication, and solid ingestion pipelines keep vector search fast and affordable as it grows.
TL;DR
- Estimate memory: vectors (N × dims × bytes) plus index overhead (HNSW links), per replica.
- Cut memory with quantization, smaller or Matryoshka dimensions, and disk-based indexes with rescoring.
- Shard for data size and write throughput; replicate for query throughput and availability.
- Build idempotent ingestion pipelines (chunk → embed → upsert) with batching, retries, and change detection.
- Plan model upgrades: version embeddings, re-embed in the background, and swap aliases or collections.
- Monitor latency, recall, freshness, and cost, not just uptime.
Quick Example
Back-of-the-envelope sizing for 50M chunks with 1,024-dim embeddings (see estimation):
An ingestion worker with batching and idempotent upserts:
Core Concepts
Memory and Storage Sizing
Main components:
- Vectors: N × d × bytes per value (4 for float32, 2 for float16, 1 for int8, 1/8 for binary).
- Index structures: HNSW links (about N × M × 2 × 4–8 bytes, plus upper layers), IVF centroids and lists, PQ codebooks.
- Metadata/payloads: often on disk, with indexed fields in memory.
- Replicas multiply everything.
Reduce it with quantization, fewer dimensions, disk-backed indexes (DiskANN, on-disk vectors with in-RAM quantized copies), and storing full text elsewhere (object storage or a database), keeping only IDs and needed metadata in the vector store.
Sharding and Replication
- Sharding splits vectors across nodes. Queries fan out to all shards, and results merge. It scales data size and ingestion, but adds fan-out latency (the slowest shard dominates).
- Tenant- or key-based sharding routes queries to one shard when filters are selective, such as per-tenant data. See metadata filtering.
- Replication copies shards to handle more QPS and survive node failures.
- Managed and serverless offerings (Pinecone, Qdrant Cloud, Weaviate Cloud, Zilliz, Turbopuffer, and the vector features of Elasticsearch, OpenSearch, and MongoDB Atlas) handle much of this. Self-hosted engines need shard and replica planning. See database sharding.
Ingestion Pipelines
Production ingestion is a data pipeline:
- Detect changes (webhooks, CDC, crawls, and content hashes) to avoid re-embedding unchanged content.
- Chunk deterministically, with stable IDs. See RAG chunking.
- Embed in batches, respecting provider rate limits, with retries and backoff. It's often the throughput and cost bottleneck.
- Upsert idempotently by ID, including metadata and the embedding model version.
- Delete chunks for removed or updated documents (by document ID), so stale content doesn't linger.
Use queues and workers (Kafka or SQS consumers) for continuous ingestion, and bulk jobs for backfills. Track freshness lag: the time from a source change until it's searchable.
Embedding Model Upgrades
Vectors from different models aren't comparable. Upgrading models means re-embedding the whole corpus:
- Store the model name and version with each vector.
- Build a new collection or index (
chunks_v4) in the background, dual-writing new content to both. - Evaluate retrieval quality on the new index. See RAG evaluation.
- Switch queries via an alias or config flag, and keep the old index briefly for rollback.
- Budget for embedding costs and time. Keep source text available so re-embedding is always possible.
Updates, Deletes, and Compaction
Frequent updates and deletes fragment ANN indexes: HNSW graphs accumulate deleted nodes, and segment-based engines accumulate tombstones. Engines compact or optimize segments in the background. Monitor segment counts, schedule optimization, and rebuild indexes periodically for heavy-churn collections.
Query Performance
- Tune ANN parameters (
ef_search,nprobe) for your recall target. See ANN indexes. - Batch queries where possible (offline jobs), and use GPUs for very high-throughput batch search.
- Cache embeddings of frequent queries, and results for popular searches.
- Keep payloads small in responses: return IDs and snippets, and fetch full documents separately.
- Colocate the vector store with the application and embedding services, to cut network latency.
Cost Control
Main cost drivers: RAM-heavy nodes, replicas, embedding API calls, and storage. Levers: quantization and dimension reduction, tiered storage (hot tenants in RAM, cold on disk), removing duplicate and stale chunks, choosing embedding models by cost-quality trade-off, and serverless or usage-based offerings for spiky workloads. See cloud costs.
Monitoring
Tie these into LLMOps dashboards alongside generation metrics.
Best Practices
Start With What You Have
pgvector or your existing search engine (Elasticsearch/OpenSearch) often handles millions of vectors well, with simpler operations. Move to a dedicated vector database when scale, filtering performance, or features demand it. See vector database selection.
Make Ingestion Idempotent and Observable
Deterministic chunk IDs, content hashes, and upserts let you safely retry and replay. Metrics on lag and failures catch silent staleness.
Version Everything
Record embedding model, chunking strategy version, and index parameters with each collection. Retrieval regressions are then traceable.
Load Test With Realistic Queries and Filters
Test with production-like query distributions, filters, concurrency, and ongoing ingestion. Unfiltered, idle benchmarks overstate performance.
Common Mistakes
Discarding Source Text
Keeping only vectors makes re-embedding impossible when you upgrade models. Always retain source text, or a way to regenerate it.
Ignoring Deletes
Updated documents create new chunks while the old ones remain searchable, so answers cite outdated content. Delete by document ID on update.
Mixing Embedding Models in One Index
Querying an index containing vectors from two different models produces meaningless similarities. Use one model per collection, and migrate fully.
FAQ
How much memory does a vector database need?
Roughly vector count × dimensions × bytes per value, plus index overhead (for HNSW, links per node) and replicas. For example, 10M vectors at 1,024 float32 dimensions need about 41 GB for raw vectors alone. Quantization and disk-based indexes can reduce RAM needs by 4–32×.
How do I change embedding models without downtime?
Build a new index with the new model in parallel (backfilling existing content and dual-writing new content), evaluate retrieval quality, then switch query traffic via an alias or configuration flag. Keep the old index briefly for rollback, then retire it.
When should I shard a vector index?
When a single node can't hold the index in memory (or on fast disk), or can't sustain ingestion throughput. Replicas address query throughput and availability. Tenant-aware sharding helps when queries naturally target one tenant.
What drives vector search costs?
Primarily memory-heavy compute for indexes (multiplied by replicas), embedding generation for ingestion and re-embedding, and storage. Quantization, fewer dimensions, deduplication, tiering, and right-sizing replicas are the main levers.
Related Topics
- Vector Databases — Engines and architectures
- Vector Quantization — Reducing memory per vector
- ANN Indexes — Index choices and tuning
- Vector Database Selection — Picking an engine
- RAG Chunking — Deterministic chunks for ingestion
- LLMOps — Operating retrieval-augmented systems