Elasticsearch Mappings

A mapping is an Elasticsearch index's schema. It defines each field's type and how it's indexed: whether a string is analyzed for full-text search (text) or stored as an exact value for filtering and aggregations (keyword), whether numbers are integers or scaled floats, and how nested objects are represented. Mappings determine what queries are possible, how relevant results are, how much disk and memory an index uses, and how fast it is.

Elasticsearch can infer mappings automatically (dynamic mapping), which is convenient for experiments and dangerous in production: guessed types are often wrong, and many field types can't be changed once data is indexed. Defining explicit mappings, usually through index templates, is one of the most important Elasticsearch best practices.

TL;DR

Quick Example

Applications read and write through the products alias, so the underlying index can be replaced later without code changes.

Core Concepts

Core Field Types

text vs keyword

A text field passes through an analyzer: "Running Shoes" becomes the tokens run and shoe, which matches searches for "shoe" or "running". A keyword field stores the exact string "Running Shoes", which works for term filters, sorting, and aggregations. You can't efficiently sort or aggregate on text. Multi-fields give you both from one source value.

Dynamic Mapping

When a document contains an unknown field, Elasticsearch adds it automatically (strings become text plus a .keyword subfield, numbers become long or float, and ISO strings become date). Problems:

Control it with "dynamic": "strict" (reject unknown fields), "dynamic": false (store but don't index unknown fields), "runtime", dynamic templates that map patterns to types, or flattened for arbitrary key-value data.

Objects vs Nested

Elasticsearch flattens object arrays. The document variants: [{color: red, size: S}, {color: blue, size: L}] is indexed as variants.color: [red, blue] and variants.size: [S, L], so a query for "red AND L" wrongly matches. The nested type indexes each object as a hidden sub-document and requires nested queries, which is correct but costs more at index and query time. Alternatives are denormalizing, or modeling variants as separate documents.

Index Templates and Aliases

Changing Mappings

You can add new fields to an existing mapping, but you generally can't change an existing field's type or analyzer, because the data was already indexed that way. The standard procedure:

  1. Create products-v2 with the new mapping.
  2. POST _reindex from products-v1 to products-v2 (or re-ingest from the source of truth).
  3. Atomically move the products alias to products-v2.
  4. Delete products-v1 after verification.

Best Practices

Define Explicit Mappings for Production Indices

Use templates with explicit types, and dynamic: strict or carefully scoped dynamic templates. Treat mapping changes like database migrations: versioned, reviewed, and tested.

Map Fields for How You Query Them

Filters, facets, and sorting need keyword, numeric, or date types; free-text search needs text with the right analyzer. Map identifiers as keyword even if they look numeric, since you never do math on them.

Disable What You Don't Need

Set "index": false on fields that are only returned, never searched, "doc_values": false on fields never sorted or aggregated, and "enabled": false on objects stored only for display. Each saves disk and indexing time.

Always Front Indices With Aliases

Aliases make reindexing, rollover, and blue-green index swaps invisible to applications.

Common Mistakes

Letting Dynamic Mapping Pick Types

Map amount explicitly as scaled_float or double before indexing data.

Aggregating on text Fields

Terms aggregations on a text field fail, or require enabling fielddata, which is memory-hungry and aggregates tokens rather than values. Aggregate on the .keyword subfield or a keyword field.

Using Nested Everywhere

Nested fields multiply the internal document count and slow queries. Use them only where cross-field matching within array elements actually matters.

FAQ

What's the difference between text and keyword in Elasticsearch?

text fields are analyzed into tokens for full-text search, with relevance scoring, stemming, and partial matches. keyword fields store the exact value, for exact-match filtering, sorting, and aggregations. Many string fields should be mapped as both using a multi-field.

Can I change a field's type in Elasticsearch?

Not in place. Create a new index with the corrected mapping, reindex the data (with the Reindex API or from your source database), and switch an alias to the new index. Adding new fields to an existing index is allowed.

What is a mapping explosion?

An uncontrolled growth in the number of mapped fields, usually from dynamic mapping of arbitrary keys (user attributes, log labels). It bloats cluster state, consumes memory, and slows the whole cluster. Prevent it with strict mappings, dynamic templates, flattened fields, and field-count limits.

When should I use nested fields?

When documents contain arrays of objects and queries must match conditions within the same object, such as "a variant that is red and size L". If you don't need that correlation, plain objects are cheaper; if arrays are huge or updated independently, consider separate documents instead.

Related Topics

References