Scaling Vector Search

A prototype with 50,000 chunks in a local vector store behaves nothing like production with 200 million vectors, thousands of queries per second, continuous ingestion, strict tenant isolation, and a budget. Scaling vector databases is mostly about a few resources: memory (ANN indexes want RAM), ingestion throughput (embedding and indexing new content), query throughput and latency, and the operational lifecycle of embeddings, including the day you change embedding models and must re-embed everything.

Good capacity planning, the right index and quantization choices, sharding and replication, and solid ingestion pipelines keep vector search fast and affordable as it grows.

TL;DR

Quick Example

Back-of-the-envelope sizing for 50M chunks with 1,024-dim embeddings (see estimation):

An ingestion worker with batching and idempotent upserts:

Core Concepts

Memory and Storage Sizing

Main components:

Reduce it with quantization, fewer dimensions, disk-backed indexes (DiskANN, on-disk vectors with in-RAM quantized copies), and storing full text elsewhere (object storage or a database), keeping only IDs and needed metadata in the vector store.

Sharding and Replication

Ingestion Pipelines

Production ingestion is a data pipeline:

  1. Detect changes (webhooks, CDC, crawls, and content hashes) to avoid re-embedding unchanged content.
  2. Chunk deterministically, with stable IDs. See RAG chunking.
  3. Embed in batches, respecting provider rate limits, with retries and backoff. It's often the throughput and cost bottleneck.
  4. Upsert idempotently by ID, including metadata and the embedding model version.
  5. Delete chunks for removed or updated documents (by document ID), so stale content doesn't linger.

Use queues and workers (Kafka or SQS consumers) for continuous ingestion, and bulk jobs for backfills. Track freshness lag: the time from a source change until it's searchable.

Embedding Model Upgrades

Vectors from different models aren't comparable. Upgrading models means re-embedding the whole corpus:

Updates, Deletes, and Compaction

Frequent updates and deletes fragment ANN indexes: HNSW graphs accumulate deleted nodes, and segment-based engines accumulate tombstones. Engines compact or optimize segments in the background. Monitor segment counts, schedule optimization, and rebuild indexes periodically for heavy-churn collections.

Query Performance

Cost Control

Main cost drivers: RAM-heavy nodes, replicas, embedding API calls, and storage. Levers: quantization and dimension reduction, tiered storage (hot tenants in RAM, cold on disk), removing duplicate and stale chunks, choosing embedding models by cost-quality trade-off, and serverless or usage-based offerings for spiky workloads. See cloud costs.

Monitoring

Tie these into LLMOps dashboards alongside generation metrics.

Best Practices

Start With What You Have

pgvector or your existing search engine (Elasticsearch/OpenSearch) often handles millions of vectors well, with simpler operations. Move to a dedicated vector database when scale, filtering performance, or features demand it. See vector database selection.

Make Ingestion Idempotent and Observable

Deterministic chunk IDs, content hashes, and upserts let you safely retry and replay. Metrics on lag and failures catch silent staleness.

Version Everything

Record embedding model, chunking strategy version, and index parameters with each collection. Retrieval regressions are then traceable.

Load Test With Realistic Queries and Filters

Test with production-like query distributions, filters, concurrency, and ongoing ingestion. Unfiltered, idle benchmarks overstate performance.

Common Mistakes

Discarding Source Text

Keeping only vectors makes re-embedding impossible when you upgrade models. Always retain source text, or a way to regenerate it.

Ignoring Deletes

Updated documents create new chunks while the old ones remain searchable, so answers cite outdated content. Delete by document ID on update.

Mixing Embedding Models in One Index

Querying an index containing vectors from two different models produces meaningless similarities. Use one model per collection, and migrate fully.

FAQ

How much memory does a vector database need?

Roughly vector count × dimensions × bytes per value, plus index overhead (for HNSW, links per node) and replicas. For example, 10M vectors at 1,024 float32 dimensions need about 41 GB for raw vectors alone. Quantization and disk-based indexes can reduce RAM needs by 4–32×.

How do I change embedding models without downtime?

Build a new index with the new model in parallel (backfilling existing content and dual-writing new content), evaluate retrieval quality, then switch query traffic via an alias or configuration flag. Keep the old index briefly for rollback, then retire it.

When should I shard a vector index?

When a single node can't hold the index in memory (or on fast disk), or can't sustain ingestion throughput. Replicas address query throughput and availability. Tenant-aware sharding helps when queries naturally target one tenant.

What drives vector search costs?

Primarily memory-heavy compute for indexes (multiplied by replicas), embedding generation for ingestion and re-embedding, and storage. Quantization, fewer dimensions, deduplication, tiering, and right-sizing replicas are the main levers.

Related Topics

References