MongoDB Aggregation Pipeline
The aggregation pipeline is MongoDB's tool for anything beyond simple finds: grouping and totals, reshaping documents, joining collections, computing running averages, and building reports. A pipeline is an array of stages. Documents flow through them in order, and each stage filters, transforms, groups, or joins its input and passes the result to the next.
It plays the role SQL's SELECT … WHERE … GROUP BY … JOIN … ORDER BY plays in relational databases, expressed as composable steps instead of one declarative statement. Written well, with filtering first and indexes used, pipelines are fast enough for production API endpoints, not just offline analytics.
TL;DR
- A pipeline is
db.collection.aggregate([ stage1, stage2, … ]); each stage transforms the document stream. - Core stages:
$match(filter),$project/$set(reshape),$group(aggregate),$sort,$limit,$unwind(flatten arrays),$lookup(join). - Put
$matchand$sortfirst so they can use indexes and shrink the data early. $facetruns several sub-pipelines at once (results plus counts);$setWindowFieldscomputes running totals and rankings.- Blocking stages have a 100 MB memory limit per stage;
allowDiskUse(on by default in recent versions) spills to disk. $merge/$outwrite results to a collection for materialized views.
Quick Example
Monthly revenue per product category for paid orders in 2026, top five categories:
Core Concepts
Essential Stages
Expressions
Inside stages, expression operators compute values: arithmetic ($add, $multiply), conditionals ($cond, $switch, $ifNull), strings ($concat, $toLower, $regexMatch), dates ($dateTrunc, $dateToString), arrays ($filter, $map, $reduce, $size), and type conversion ($toDecimal, $toObjectId). Field paths are referenced as "$field.subfield", and variables as "$var".
Joins With $lookup
Index the foreignField in the joined collection; otherwise every input document triggers a collection scan. Heavy $lookup use on hot paths is a sign the schema design may need more embedding.
Pagination With $facet
One round trip returns a page and the total count. For deep pagination, prefer range-based (keyset) queries over large $skip values.
Window Functions
Performance
- Filter early. A leading
$match(and a$sortdirectly after it) can use indexes. Once documents pass through$unwind,$group, or$project, later stages can't use collection indexes. - Project late, but trim when documents are large. The optimizer handles many projections automatically; explicitly dropping big fields before memory-heavy stages still helps.
$sort+$limitis coalesced into a top-k sort that keeps only k documents in memory.- Check the plan with
db.collection.explain("executionStats").aggregate([...]). Look forIXSCANrather thanCOLLSCANin the first stage, and for documents examined vs returned. - Memory: stages like
$groupand$sortare limited to 100 MB each before spilling to disk. Spilling works but is slower; reduce data earlier if you hit it often.
See MongoDB indexes for designing indexes that support pipelines.
Best Practices
Build Pipelines Incrementally
Develop stage by stage, inspecting output after each one (MongoDB Compass's aggregation builder is excellent for this). Complex pipelines are much easier to debug when each stage's output is known.
Materialize Expensive Aggregations
For dashboards that recompute the same heavy aggregation on every request, run it on a schedule or on change and $merge the results into a summary collection. Reads then become simple indexed finds. This is the document-database version of a materialized view.
Offload Analytics From the Primary
Run heavy reporting pipelines against a secondary (readPreference: "secondaryPreferred") or a dedicated analytics node, so they don't compete with transactional traffic. For large-scale analytics, export to a warehouse; see data warehousing.
Common Mistakes
$match After $unwind or $group
Unindexed $lookup
Joining 100,000 orders to a customers collection without an index on the join field means 100,000 collection scans. Always index foreignField.
Huge $push Accumulators
$group with $push: "$ROOT" builds arrays of entire documents and can exceed the 16 MB document limit or the stage memory limit. Push only the fields you need, cap with $topN/$firstN, or restructure the query.
FAQ
Is the aggregation pipeline slower than find?
For the same filter and projection, no: a pipeline starting with $match and $project uses indexes just like find. Pipelines get expensive when they process many documents through grouping, unwinding, or joins, which find can't do at all.
What replaced map-reduce?
The aggregation pipeline. Map-reduce is deprecated; pipelines, including $function and $accumulator for custom JavaScript logic when truly needed, cover its use cases and run much faster.
Can I update documents with an aggregation?
Yes, in two ways: $merge writes pipeline output into a collection (insert, replace, or merge into matching documents), and update commands accept an aggregation pipeline as the update (updateMany({}, [ { $set: { total: { $sum: "$items.price" } } } ])) to compute new values from existing fields.
How do I do full-text or vector search in a pipeline?
On MongoDB Atlas, $search (Lucene-based full-text) and $vectorSearch stages run as the first stage of a pipeline, followed by any normal stages. Self-managed deployments have basic $text search, and recent versions add search and vector search through the community mongot component.
Related Topics
- MongoDB — The database overview
- MongoDB Indexes — Indexes that make pipelines fast
- MongoDB Schema Design — Shaping documents to avoid expensive joins
- Query Optimization — Reading plans and reducing work
- SQL — The relational equivalent of these operations