Elasticsearch Mappings
A mapping is an Elasticsearch index's schema. It defines each field's type and how it's indexed: whether a string is analyzed for full-text search (text) or stored as an exact value for filtering and aggregations (keyword), whether numbers are integers or scaled floats, and how nested objects are represented. Mappings determine what queries are possible, how relevant results are, how much disk and memory an index uses, and how fast it is.
Elasticsearch can infer mappings automatically (dynamic mapping), which is convenient for experiments and dangerous in production: guessed types are often wrong, and many field types can't be changed once data is indexed. Defining explicit mappings, usually through index templates, is one of the most important Elasticsearch best practices.
TL;DR
- Every field has a type;
textis analyzed for full-text search, andkeywordis exact-match for filters, sorting, and aggregations. - Use multi-fields (
titleastextplustitle.keyword) when you need both. - Dynamic mapping guesses types from the first document. Control it with
dynamic: strictor dynamic templates to avoid mapping explosions. - Arrays of objects lose relationships unless mapped as
nested, which costs extra at query time. - Index templates apply mappings and settings to new indices (time-series, rollover).
- Existing field types can't be changed. Reindex into a new index and switch an alias.
Quick Example
Applications read and write through the products alias, so the underlying index can be replaced later without code changes.
Core Concepts
Core Field Types
text vs keyword
A text field passes through an analyzer: "Running Shoes" becomes the tokens run and shoe, which matches searches for "shoe" or "running". A keyword field stores the exact string "Running Shoes", which works for term filters, sorting, and aggregations. You can't efficiently sort or aggregate on text. Multi-fields give you both from one source value.
Dynamic Mapping
When a document contains an unknown field, Elasticsearch adds it automatically (strings become text plus a .keyword subfield, numbers become long or float, and ISO strings become date). Problems:
- Wrong types: a first document with
"zip": "02134"maps as text, and"price": 5maps price aslong, so later decimals are truncated. - Mapping explosion: user-generated keys (for example
attributes.<anything>) create thousands of fields, bloating cluster state and slowing everything. There's a default limit of 1,000 fields per index.
Control it with "dynamic": "strict" (reject unknown fields), "dynamic": false (store but don't index unknown fields), "runtime", dynamic templates that map patterns to types, or flattened for arbitrary key-value data.
Objects vs Nested
Elasticsearch flattens object arrays. The document variants: [{color: red, size: S}, {color: blue, size: L}] is indexed as variants.color: [red, blue] and variants.size: [S, L], so a query for "red AND L" wrongly matches. The nested type indexes each object as a hidden sub-document and requires nested queries, which is correct but costs more at index and query time. Alternatives are denormalizing, or modeling variants as separate documents.
Index Templates and Aliases
- Index templates (composed from component templates) automatically apply settings and mappings to indices whose names match patterns. They're essential for time-series data with rollover and data streams.
- Aliases point to one or more indices. Apps query the alias, and you swap indices behind it atomically.
Changing Mappings
You can add new fields to an existing mapping, but you generally can't change an existing field's type or analyzer, because the data was already indexed that way. The standard procedure:
- Create
products-v2with the new mapping. POST _reindexfromproducts-v1toproducts-v2(or re-ingest from the source of truth).- Atomically move the
productsalias toproducts-v2. - Delete
products-v1after verification.
Best Practices
Define Explicit Mappings for Production Indices
Use templates with explicit types, and dynamic: strict or carefully scoped dynamic templates. Treat mapping changes like database migrations: versioned, reviewed, and tested.
Map Fields for How You Query Them
Filters, facets, and sorting need keyword, numeric, or date types; free-text search needs text with the right analyzer. Map identifiers as keyword even if they look numeric, since you never do math on them.
Disable What You Don't Need
Set "index": false on fields that are only returned, never searched, "doc_values": false on fields never sorted or aggregated, and "enabled": false on objects stored only for display. Each saves disk and indexing time.
Always Front Indices With Aliases
Aliases make reindexing, rollover, and blue-green index swaps invisible to applications.
Common Mistakes
Letting Dynamic Mapping Pick Types
Map amount explicitly as scaled_float or double before indexing data.
Aggregating on text Fields
Terms aggregations on a text field fail, or require enabling fielddata, which is memory-hungry and aggregates tokens rather than values. Aggregate on the .keyword subfield or a keyword field.
Using Nested Everywhere
Nested fields multiply the internal document count and slow queries. Use them only where cross-field matching within array elements actually matters.
FAQ
What's the difference between text and keyword in Elasticsearch?
text fields are analyzed into tokens for full-text search, with relevance scoring, stemming, and partial matches. keyword fields store the exact value, for exact-match filtering, sorting, and aggregations. Many string fields should be mapped as both using a multi-field.
Can I change a field's type in Elasticsearch?
Not in place. Create a new index with the corrected mapping, reindex the data (with the Reindex API or from your source database), and switch an alias to the new index. Adding new fields to an existing index is allowed.
What is a mapping explosion?
An uncontrolled growth in the number of mapped fields, usually from dynamic mapping of arbitrary keys (user attributes, log labels). It bloats cluster state, consumes memory, and slows the whole cluster. Prevent it with strict mappings, dynamic templates, flattened fields, and field-count limits.
When should I use nested fields?
When documents contain arrays of objects and queries must match conditions within the same object, such as "a variant that is red and size L". If you don't need that correlation, plain objects are cheaper; if arrays are huge or updated independently, consider separate documents instead.
Related Topics
- Elasticsearch — The search engine overview
- Elasticsearch Analyzers — How text fields are tokenized
- Elasticsearch Query DSL — Querying mapped fields
- Elasticsearch Aggregations — Keyword and numeric fields in analytics
- Elasticsearch Scaling — Shards, templates, and rollover
- Hybrid Search — Combining text and vector fields