MongoDB Indexes
Without an index, MongoDB answers a query by scanning every document in the collection (a COLLSCAN). With the right one, it walks a B-tree straight to the matching documents. Indexes are the single biggest performance lever in MongoDB, and as in any database they aren't free: each index slows writes and consumes memory, which is precious because MongoDB performs best when indexes fit in RAM.
MongoDB's index types map to its document model: multikey indexes for arrays, indexes on embedded fields, TTL indexes that expire documents, and partial indexes for subsets. The core design skill is the same as in relational databases: build compound indexes that match your queries' filters and sorts. See database indexing for the general theory.
TL;DR
- Every collection has a unique index on
_id; add indexes for your frequent query shapes. - Compound indexes follow the ESR rule: Equality fields first, then Sort fields, then Range fields.
- Compound indexes serve queries on their prefixes:
{a, b, c}supportsa,a+b, anda+b+c. - Indexing an array field creates a multikey index covering each element.
- Special types: TTL (auto-expire), partial (subset of documents), unique, text, wildcard, 2dsphere (geo).
- Verify with
explain("executionStats"): look forIXSCANand a low examined-to-returned ratio.
Quick Example
Core Concepts
Single-Field and Compound Indexes
A single-field index sorts documents by one field. A compound index sorts by several fields in order, like a phone book sorted by last name, then first name. Field order matters enormously:
- The index supports queries on any prefix of its fields.
- It can satisfy a
sort()without an in-memory sort if the sort matches the index order (or its exact reverse) after the equality fields.
The ESR Rule
For compound indexes, order fields as:
- Equality fields (exact matches) first, since they narrow the search most precisely.
- Sort fields next, so results come out of the index already ordered.
- Range fields (
$gt,$lt,$inwith many values, regex) last.
Putting a range field before a sort field forces a blocking in-memory sort; putting it before an equality field makes the index scan far more keys.
Multikey Indexes
Indexing a field that holds an array (tags: ["redis", "cache"]) creates an entry per element, so find({ tags: "redis" }) uses the index. A compound index can include at most one array field per document. Multikey indexes on large arrays create many index entries, one more reason to avoid unbounded arrays (see schema design).
Special Index Types
Covered Queries
If every field the query filters on and returns is in the index, MongoDB answers from the index alone without fetching documents. Project only indexed fields and exclude _id (unless it's in the index) to get a covered query, visible in explain output as zero documents examined.
Reading explain()
Key things to check in executionStats:
- Winning plan stage:
IXSCAN(good) vsCOLLSCAN(a full scan). totalKeysExaminedandtotalDocsExaminedvsnReturned: ideally close to 1:1. Examining 50,000 documents to return 20 means the index isn't selective for this query.SORTstage: an in-memory sort; consider putting the sort field in the index per ESR.executionTimeMillis: the end-to-end time for the query.
The Performance Advisor (Atlas) and the database profiler (db.setProfilingLevel(1, { slowms: 100 })) surface slow queries and suggest indexes.
Best Practices
Index for Your Actual Query Shapes
Collect the frequent queries (filters plus sorts) and design a small set of compound indexes that serve them, preferring one well-ordered compound index over several single-field ones.
Remove Redundant and Unused Indexes
An index on {a} is redundant if {a, b} exists. $indexStats shows how often each index is used. Every unused index costs write throughput and RAM.
Build Indexes Safely in Production
Since MongoDB 4.2, index builds hold exclusive locks only briefly at the start and end, but they still consume I/O and CPU. Build large indexes during quieter periods and watch replication lag. Rolling builds across replica set members are an option for very large collections.
Keep the Working Set in Memory
Monitor index sizes (db.collection.stats()) against available RAM. When frequently used indexes no longer fit in the WiredTiger cache, performance drops sharply.
Common Mistakes
Wrong Field Order in Compound Indexes
Low-Selectivity Indexes Alone
An index on a boolean or a three-value status field matches huge fractions of the collection and often loses to a collection scan. Combine such fields with selective ones in a compound index, or use a partial index for the rare value you query.
Regexes That Can't Use the Index
{ name: /smith/i } (case-insensitive, unanchored) scans every index key. Anchored, case-sensitive prefixes (/^Smi/) use the index efficiently. For real search, use a collation-based index for case-insensitivity or Atlas Search.
FAQ
How many indexes should a collection have?
As few as serve your query patterns well, often somewhere between a handful and a dozen for a busy collection. There's a hard limit of 64 per collection, but write amplification and memory use become problems long before that.
Does index direction (1 vs -1) matter?
For single-field indexes, no: MongoDB can traverse in either direction. For compound indexes used for sorting on multiple fields with mixed directions ({ a: 1, b: -1 }), the directions must match the sort or its exact inverse.
Can a query use more than one index?
MongoDB can use index intersection in limited cases, but it rarely beats a well-designed compound index. Design compound indexes for important queries rather than relying on intersection.
How do TTL indexes work?
A background task runs about every 60 seconds and deletes documents whose indexed date field is older than expireAfterSeconds. Deletion isn't instant, so documents may linger for up to a minute or more under load. Don't rely on TTL for exact-time security expiry; check timestamps in the application too.
Related Topics
- MongoDB — The database overview
- Database Indexing — B-trees and indexing fundamentals
- MongoDB Aggregation — Index use in pipelines
- MongoDB Schema Design — Designing documents and indexes together
- PostgreSQL Index Types — The relational counterpart
- Query Optimization — Reading execution plans