Hybrid Search
Hybrid search combines two retrieval methods: lexical (keyword) search, typically BM25, which matches exact terms, and semantic (vector) search, which matches meaning via embeddings. Each fails in different ways. Vector search misses exact identifiers, product codes, rare names, and acronyms, while keyword search misses paraphrases and synonyms. Running both and fusing their results consistently beats either alone, and it's now the default recommendation for RAG retrieval.
Most search engines and vector databases support hybrid search natively: Elasticsearch, OpenSearch, Weaviate, Qdrant, Pinecone, Milvus, Vespa, and PostgreSQL with pgvector plus full-text search. The main design questions are how to fuse results, how many candidates to fetch from each side, and what to do after fusion (reranking).
TL;DR
- BM25 excels at exact terms (error codes, SKUs, names, jargon); dense vectors excel at meaning and paraphrase.
- Hybrid search runs both, then fuses the ranked lists.
- Reciprocal Rank Fusion (RRF) combines by rank, with no score normalization needed. It's a robust default.
- Weighted score fusion (for example 0.7 × vector + 0.3 × keyword) needs normalized scores and tuning.
- Sparse learned models (SPLADE, BGE-M3 sparse) add term expansion with keyword-style indexes.
- Apply metadata filters to both retrievers, then rerank the fused top candidates for best quality.
Quick Example
Hybrid search in PostgreSQL with pgvector and full-text search, fused with RRF:
The top 20 fused results then go to a reranker, and the best 5–8 go to the LLM.
Core Concepts
Lexical Search (BM25)
BM25 scores documents by query-term frequency, inverse document frequency (rare terms matter more), and document length normalization. It runs on inverted indexes and is fast, explainable, and precise for exact tokens. Weaknesses: no understanding of synonyms ("laptop" vs "notebook"), paraphrase, or cross-language matches, and sensitivity to spelling and tokenization. See PostgreSQL full-text search and Elasticsearch.
Dense Vector Search
Embedding models map queries and passages into vectors, where semantic similarity becomes geometric closeness (cosine or dot product), found with approximate nearest-neighbor (ANN) indexes like HNSW. Strengths: meaning, paraphrase, multilingual matching, and natural-language questions. Weaknesses:
- Exact identifiers and rare terms (
ERR_CONN_RESET_4012, part numbers, new product names) often embed poorly. - Domain jargon unseen during embedding training.
- Hard to explain why something matched.
- ANN indexes interact awkwardly with restrictive metadata filters.
Fusion Methods
Reciprocal Rank Fusion scores each document by summing 1 / (k + rank) across result lists (k is typically 60):
- It uses only ranks, so BM25 scores and cosine similarities don't need to be on the same scale.
- It rewards documents that appear high in both lists.
- There's one parameter, and it's robust across domains, which makes it the common default.
Weighted score fusion (convex combination) normalizes each retriever's scores (min-max or z-score) and combines them with weights (α × semantic + (1−α) × lexical). It can outperform RRF when tuned on labeled data, but the scores are fragile across queries and corpora.
Some engines offer learned fusion or query-dependent weighting, for example weighting keywords higher for queries containing codes or quoted phrases.
Sparse Learned Retrieval
Models like SPLADE and the sparse output of BGE-M3 produce sparse vectors over the vocabulary with learned term weights and expansion (adding related terms not in the text). They keep keyword-style efficiency and exact matching while capturing some semantics, and can serve as the "lexical" side of hybrid search or as a third signal.
Candidate Depth
Fetch more candidates from each retriever than you'll finally use (for example 50 each, fuse to 20, rerank to 5). Too few candidates means a relevant document found by only one retriever never reaches the reranker.
Hybrid Search in Common Engines
Best Practices
Start With RRF, Then Measure
RRF with k=60 and 25–100 candidates per retriever is a strong baseline. Only move to tuned weighted fusion if evaluation data shows a clear gain. See RAG evaluation.
Filter Consistently on Both Sides
Apply tenant, permission, language, and date filters to both the keyword and the vector query. Filtering only after fusion can leak unauthorized content into candidates, or empty the results.
Index Contextualized Text for BM25 Too
The same contextual chunk headers or generated context you add for embeddings (see chunking) improve BM25 matching, since titles and section names add important keywords.
Always Rerank the Fused Set
Fusion decides which candidates to consider; a cross-encoder reranker decides their final order far more accurately. Hybrid plus rerank is the standard high-quality pipeline.
Common Mistakes
Adding Raw Scores From Different Retrievers
Vector-Only Retrieval for Technical Content
Support documentation, logs, code, and catalogs are full of identifiers that embeddings handle poorly. Users searching for SQLSTATE 40001 or SKU-88213 get vaguely related results. Add lexical search.
Too Few Candidates Before Fusion
Taking only the top 5 from each retriever means documents ranked 6th by both, which may be the best overall, are never considered. Retrieve deeper, then narrow.
FAQ
Why not just use vector search?
Embeddings capture meaning but struggle with exact matches (product codes, error messages, names, acronyms), and with domain vocabulary they weren't trained on. Keyword search handles those precisely. Combining the two covers both failure modes, and benchmarks consistently show hybrid retrieval outperforming either alone.
What is Reciprocal Rank Fusion?
A method for merging ranked lists: each document's fused score is the sum over lists of 1/(k + its rank in that list), with k usually 60. Because it uses ranks rather than raw scores, it needs no normalization, and it favors documents ranked well by multiple retrievers.
Do I need a separate search engine for hybrid search?
Not necessarily. PostgreSQL can do both with pgvector and full-text search, and most vector databases support sparse and keyword search. A dedicated engine like Elasticsearch or OpenSearch makes sense when you need advanced text analysis, faceting, or already run one.
How should I weight keyword vs semantic results?
Start with RRF, which weights them equally by rank. If you have labeled evaluation queries, tune weighted fusion (or a per-query-type weighting) and compare. Query types differ: identifier-heavy queries benefit from keyword emphasis, and conversational questions from semantic.
Related Topics
- RAG — Retrieval-augmented generation overview
- Reranking — Reordering fused candidates
- Embeddings — Dense vector representations
- Vector Databases — Engines with hybrid support
- PostgreSQL Full-Text Search — Lexical search in Postgres
- Elasticsearch — BM25 and kNN in one engine