Elasticsearch Query DSL
Elasticsearch searches are written in the Query DSL, a JSON language for combining full-text relevance queries, exact filters, ranges, geospatial conditions, and vector similarity. A product search box, a log explorer, and an autocomplete are all built from the same small set of queries (match, term, range, bool), composed differently.
The two ideas that matter most are the difference between query context (does it match, and how well?) and filter context (does it match, yes or no), and how bool queries combine them. Getting these right gives you relevant results and fast, cacheable filters.
TL;DR
- Query context calculates a relevance score; filter context is yes or no, cached, and faster.
- Use
match,multi_match, andmatch_phrasefor full-text ontextfields. - Use
term,terms,range, andexistsfor exact values onkeyword, numeric, and date fields. boolcombines clauses:must(scored, required),should(scored, optional boosts),filter(required, unscored),must_not(excluded).- Relevance uses BM25; tune with field boosts,
function_score, and recency or popularity signals. - Paginate deep results with
search_afterplus a point in time, not largefromoffsets.
Quick Example
A product search: full-text relevance, exact filters, boosts, facets, and highlighting.
Core Concepts
Query vs Filter Context
Put anything that shouldn't affect ranking (status, dates, permissions, categories, price ranges) in filter. It's faster, and it keeps scores focused on text relevance.
Full-Text Queries
fuzziness: "AUTO" tolerates typos. Full-text queries apply the field's search analyzer to the input, so analyzers must be consistent (see analyzers).
Term-Level Queries
term, terms, range, exists, prefix, wildcard, regexp, and ids match exact, unanalyzed values. Use them on keyword, numeric, date, and boolean fields. A term query against a text field usually fails to match, because indexed tokens are lowercased and stemmed while the query term isn't.
The bool Query
must: required, contributes to the score.should: optional, adds to the score when matched. It's required only if there's nomustorfilter, or withminimum_should_match.filter: required, no scoring.must_not: excluded, no scoring.
bool queries nest, so complex search logic stays composable.
Relevance Scoring
Elasticsearch scores with BM25: matches on rarer terms and in shorter fields score higher, with diminishing returns for repeated terms. Tuning levers:
- Field boosts (
title^3) andshouldclauses for phrase matches. function_scoreorscript_scoreto blend in business signals: recency decay (gausson dates), popularity (field_value_factoron sales or ratings), or availability.rank_featurefields for efficient signal boosting.- Synonyms and analyzers to improve recall.
explain: trueshows why a document scored as it did.
Pagination
from+sizeworks for shallow pages; the defaultindex.max_result_windowis 10,000, and deep pages get expensive.search_afterwith a sort tiebreaker, plus a point in time (PIT) for a consistent snapshot, is the scalable approach for deep pagination and exports.- The Scroll API is legacy for bulk export; prefer PIT with
search_after.
Vector and Hybrid Search
Elasticsearch supports approximate kNN search on dense_vector fields, combinable with filters and BM25 queries, fused via RRF retrievers. See hybrid search and RAG.
Best Practices
Filter Everything That Isn't Relevance
Status, visibility, tenancy, dates, and categories go in filter. That improves speed (caching) and relevance (scores reflect text only). For multi-tenant search, a mandatory tenant filter is also an access control measure.
Use search_template or Server-Side Query Building
Build queries server-side from validated parameters (or stored search templates), rather than passing raw user input into query_string, which can produce expensive or error-throwing queries.
Measure Relevance
Keep a set of real queries with expected results, and evaluate changes with the Ranking Evaluation API (_rank_eval) or offline metrics (nDCG, precision@k). Relevance tuning without measurement is guesswork. See RAG evaluation.
Profile Slow Queries
Use "profile": true and slow logs to find expensive clauses: leading wildcards, regexes, deep pagination, large terms aggregations, and scripts.
Common Mistakes
term Queries on Analyzed Text
Scoring Filters
Putting in_stock: true inside must adds it to the score and prevents caching. Move it to filter.
Leading Wildcards
"wildcard": {"sku": "*123"} scans every term in the field. Use wildcard field types, n-gram analyzers, or reversed tokens for suffix search.
FAQ
What's the difference between must and filter in a bool query?
Both require matches. must clauses contribute to the relevance score; filter clauses don't affect scoring and can be cached, which makes them faster. Use filter for yes/no conditions and must for relevance-bearing text queries.
Why doesn't my term query return results?
It's probably targeting a text field. The indexed tokens were analyzed (lowercased, stemmed), but term doesn't analyze the query. Query the keyword version of the field, or use match for full-text search.
How do I paginate beyond 10,000 results?
Use search_after with a unique sort tiebreaker, and a point in time for consistency. Raising max_result_window makes deep from pagination possible, but increasingly expensive.
How can I boost newer or more popular documents?
Wrap the query in function_score with a decay function on the date field (for example gauss with an origin of now and a scale of 30d), and/or a field_value_factor on popularity, or use rank_feature fields. Keep text relevance as the main signal and business signals as moderate boosts.
Related Topics
- Elasticsearch — The search engine overview
- Elasticsearch Mappings — Field types that queries depend on
- Elasticsearch Analyzers — How text is tokenized
- Elasticsearch Aggregations — Facets and analytics alongside search
- Hybrid Search — BM25 plus vectors
- PostgreSQL Full-Text Search — A lighter-weight alternative