Vector Quantization

Embeddings are big. A 1,024-dimensional float32 vector takes 4 KB, so 100 million of them need about 400 GB just for the raw vectors, before any index overhead. In-memory ANN indexes like HNSW want that data in RAM, which gets expensive fast. Quantization compresses vectors, often 4× to 32× or more, with a small, controllable loss in search accuracy, cutting memory, storage, and cost, and often speeding up search.

Modern vector search stacks combine several techniques: half precision, scalar (int8) and binary quantization, product quantization, and Matryoshka embeddings that can be truncated to fewer dimensions. The key trick is rescoring: search with compressed vectors, then re-rank the top candidates using full-precision vectors to recover accuracy.

TL;DR

Quick Example

Binary quantization with rescoring in Qdrant:

Truncating a Matryoshka embedding and storing half precision in pgvector:

Core Concepts

Memory Math

Index structures (HNSW links, IVF lists) add overhead on top. See vector search scaling.

Half Precision

Storing vectors as float16 or bfloat16 halves memory with almost no effect on similarity rankings for typical embeddings. It's the easiest first step (pgvector halfvec, and most engines' storage options).

Scalar Quantization (int8)

Each dimension is mapped from a float range to 256 integer levels, using per-dimension or global min and max (often clipped at quantiles to ignore outliers). Distances are computed on int8 values with SIMD. It gives 4× compression, often with around 99% recall, especially with rescoring. It's the safe default for large collections.

Binary Quantization

Each dimension becomes a single bit (positive becomes 1, otherwise 0), and similarity uses Hamming distance (XOR plus popcount), which is extremely fast. It gives 32× compression. Quality depends heavily on the model: high-dimensional embeddings (1,024+ dims) from models trained or evaluated with binary quantization retain surprisingly good rankings. Always use oversampling plus rescoring. Variants include 1.5-bit, 2-bit, and asymmetric schemes, where the query stays full precision.

Product Quantization (PQ)

PQ splits a d-dimensional vector into m subvectors, learns a codebook of (typically) 256 centroids per subspace with k-means, and stores each subvector as a 1-byte centroid ID. Distances are computed with precomputed lookup tables (asymmetric distance). Very high compression, but lower accuracy, and it requires training on representative data. It's classic in IVF-PQ indexes (FAISS, Milvus) for billion-scale search. OPQ rotates the data first to improve accuracy.

Matryoshka Embeddings and Dimension Reduction

Matryoshka Representation Learning trains models so the first n dimensions form a useful embedding on their own. You can truncate from 3,072 to 1,024 or 256 dimensions and renormalize, with gradual quality loss. Many current embedding models (OpenAI text-embedding-3, Nomic, Jina, Voyage, and others) support it. Shorter vectors cut memory and compute linearly, and they combine with quantization.

Oversampling and Rescoring

The standard pattern for accurate compressed search:

  1. Search the compressed index for k × oversampling candidates (for example 3× or 4×).
  2. Load full-precision (or int8) vectors for those candidates, from disk or RAM.
  3. Recompute exact distances, and return the top k.

Compressed vectors live in RAM, full vectors can sit on cheaper disk, and recall approaches that of uncompressed search.

Choosing a Strategy

Best Practices

Evaluate on Your Data and Model

Quantization impact varies by embedding model and domain. Measure recall@k and downstream quality (RAG answer quality, search relevance) before and after compression. See RAG evaluation.

Keep Full-Precision Vectors Somewhere

Store originals, on disk or in object storage, for rescoring and for re-indexing with a different quantization later. Re-embedding millions of documents because you discarded originals is expensive.

Normalize Before Quantizing

For cosine similarity, normalize vectors first. Quantization ranges then become stable, and dot product equals cosine.

Combine With Reranking

In retrieval pipelines, a cross-encoder reranker over the top candidates compensates for small losses from aggressive quantization, often at lower total cost than uncompressed indexes.

Common Mistakes

Binary Quantization Without Rescoring

Hamming-distance rankings alone can be noticeably worse than full precision. Oversample and rescore.

Truncating Non-Matryoshka Embeddings

Cutting dimensions from a model not trained for it can severely damage quality. Only truncate models documented to support it, and renormalize after truncation.

Training PQ on Unrepresentative Data

Codebooks trained on a small or different sample generalize poorly, and recall drops for real queries. Train on a large, representative sample, and retrain when data shifts.

FAQ

What is vector quantization?

Compressing embedding vectors by representing their values with fewer bits: 8-bit integers, single bits, or compact codes from learned codebooks. It reduces memory and storage costs, and speeds up distance computations, at the cost of some precision in similarity search.

How much accuracy do I lose with int8 quantization?

Usually very little. Recall often stays around 97–99% or higher, and rescoring with full-precision vectors recovers most of the remainder. The exact impact depends on the model and data, so measure it.

When does binary quantization work well?

With high-dimensional embeddings (roughly 1,024 dimensions and up) from models that tolerate it, combined with oversampling and full-precision rescoring. It gives 32× memory savings and very fast search. It's less suitable for low-dimensional embeddings or models not evaluated with binary quantization.

What are Matryoshka embeddings?

Embeddings trained so that prefixes of the vector (the first 256, 512, 1,024 dimensions…) are themselves good embeddings. You can shorten them to save memory and compute, trading a little accuracy, without re-embedding with a different model.

Related Topics

References