Elasticsearch Analyzers & Text Analysis

When a document is indexed into a text field, Elasticsearch doesn't store the string as-is for searching. It runs it through an analyzer that breaks it into tokens (terms) and normalizes them: lowercasing, removing accents, stemming "running" to "run", and optionally adding synonyms. At query time the search input goes through the same, or a compatible, analyzer, and matching happens on tokens. Analysis is why a search for "Café" finds "cafe", and why "shoe" matches "Running Shoes".

Analyzers are the biggest lever for search recall (finding relevant documents) and a major factor in precision. They're set in mappings, so changing them requires reindexing, and it pays to design and test them early.

TL;DR

Quick Example

A custom analyzer setup for product search with autocomplete:

(The synonyms_set and english_possessive_stemmer filter would be defined via the synonyms API and a stemmer filter, respectively.)

Core Concepts

Anatomy of an Analyzer

  1. Character filters transform the raw string: strip HTML (html_strip), map characters (mapping: & → and), or apply regex replacements.
  2. Tokenizer splits text into tokens: standard (Unicode word boundaries), whitespace, keyword (the whole string as one token), pattern, ngram/edge_ngram, path_hierarchy, uax_url_email, or language-specific tokenizers (ICU, kuromoji for Japanese, smartcn for Chinese).
  3. Token filters modify, add, or remove tokens: lowercase, asciifolding (é → e), stop, stemmers (porter_stem, snowball, language stemmer), synonym_graph, word_delimiter_graph (splits "WiFi-6E" into parts), shingle (word pairs), unique, length.

The _analyze API shows the tokens at every step with explain: true.

Built-In Analyzers

Stemming and Stop Words

Stemming reduces words to a root ("running", "runs" → "run"), which increases recall but can conflate distinct words ("university"/"universe" with aggressive stemmers). Light stemmers or lemmatization (via plugins) are gentler. Stop words (the, a, of) were traditionally removed to save space. Modern BM25 handles them reasonably well, and removing them can break phrase queries like "to be or not to be", so use them selectively.

Synonyms

Autocomplete and Partial Matching

Use a different, non-n-gram search analyzer at query time, so the query "sho" isn't itself exploded into n-grams.

Normalizers

keyword fields aren't analyzed, but a normalizer (restricted to character-level filters like lowercase and asciifolding) makes exact matching and aggregations case- and accent-insensitive: emails, tags, usernames.

Multilingual Text

Use separate fields per language (title.en, title.de) with the matching language analyzers, detect language at ingest (the inference processor or an external library), and use ICU plugins for proper Unicode handling. CJK languages need dedicated tokenizers, since whitespace tokenization doesn't work.

Best Practices

Start Simple, Then Tune With Real Queries

Begin with standard or a language analyzer, collect real search queries and failures ("zero results" searches, poor top results), and add synonyms, stemming changes, or decompounding based on evidence.

Keep Index and Search Analyzers Compatible

They may differ (n-grams at index time only, synonyms at search time), but both must produce tokens in the same normalized form: the same lowercasing, folding, and stemming.

Make Synonyms Search-Time and Reloadable

Search-time synonym sets can be updated through the API without reindexing, so merchandisers and search teams can iterate quickly.

Test Analyzers in CI

Add _analyze assertions for tricky inputs (product codes, hyphenated names, accents, plurals) to your mapping tests. Analyzer regressions silently degrade search quality.

Common Mistakes

Stemming Identifiers

Running SKUs, model numbers, or codes ("A-100-X") through text analyzers splits and mangles them. Map identifiers as keyword (with a normalizer), or use word_delimiter_graph deliberately.

Index-Time Synonym Expansion

Adding synonyms to the index analyzer means every synonym change requires a full reindex, and it inflates the index. Prefer search-time synonym_graph.

N-grams on Large Fields

Indexing ngram tokens (min 1, max 20) on long descriptions multiplies index size many times over, and slows indexing. Restrict n-grams to short fields, like names and titles.

FAQ

What is an analyzer in Elasticsearch?

A pipeline that converts text into tokens for the inverted index. It consists of optional character filters (clean the text), exactly one tokenizer (split it into tokens), and optional token filters (normalize, remove, or add tokens). The same process is applied to query text so that queries and documents match on normalized terms.

Should synonyms be applied at index time or search time?

Search time, in most cases. Search-time synonym_graph handles multi-word synonyms correctly, keeps the index smaller, and lets you update synonyms without reindexing. Index-time expansion is rarely worth its costs.

How do I implement autocomplete?

Use a search_as_you_type field, or a subfield analyzed with an edge_ngram filter at index time and a standard analyzer at search time, queried with match_bool_prefix or multi_match of type bool_prefix. For very fast suggestions from a fixed list, the completion suggester works well.

How do I make keyword searches case-insensitive?

Add a normalizer with lowercase (and asciifolding if needed) to the keyword field. Both indexed values and term queries on that field are then normalized, making exact matches and aggregations case-insensitive.

Related Topics

References