MongoDB Schema Design
MongoDB is often called "schemaless", but every real application has a schema. The difference is that the schema lives in how you shape documents, not in CREATE TABLE statements. And the rules are almost the reverse of relational modeling. In SQL you normalize first and join at query time. In MongoDB you design documents around how the application reads and writes data, often keeping related data together in one document so a single read returns everything a page needs.
Good document design makes MongoDB fast and simple. Poor design — relational tables translated one-to-one into collections, or unbounded arrays growing inside a document — produces slow $lookup-heavy queries and documents that hit the 16 MB limit.
TL;DR
- Model for your access patterns: list the queries first, then shape documents so the most frequent ones touch one document.
- Embed data that's read together and owned by the parent (addresses, order line items).
- Reference data that's large, shared across many documents, unbounded, or updated independently.
- Never let arrays grow without bound; documents have a 16 MB limit, and huge arrays hurt performance long before that.
- Use proven patterns: bucket, subset, extended reference, computed, outlier.
- Enforce structure with
$jsonSchemavalidation and a typed ODM or driver layer.
Quick Example
An e-commerce order, embedding what's read together and referencing what's shared:
Rendering an order page is one indexed read. The full customer record lives in customers and is referenced by customer._id. Only the fields the order view needs are copied.
Core Concepts
Embedding vs Referencing
A single-document write is always atomic, so embedding also gives you transactional updates for free. See MongoDB transactions for multi-document cases.
Modeling Relationships
- One-to-few (a user's 3 addresses): embed an array.
- One-to-many (a product's 500 reviews): reference from the child (
review.productId), and perhaps embed a small subset (the latest 10) in the parent. - One-to-squillions (a server's billions of log lines): reference from the child only. Never keep an array of IDs in the parent.
- Many-to-many (students and courses): arrays of references on one or both sides, or a separate linking collection when the relationship has attributes of its own.
Duplication Is a Tool
Copying a customer's name into orders (the extended reference pattern) avoids a lookup on every order read. The cost is updating copies when the name changes, or deciding that historical orders should keep the name at purchase time. Duplicate fields that are read often and change rarely.
Common Patterns
Schema Validation
MongoDB can enforce structure server-side:
Combine database validation with application-level types (Mongoose schemas, Pydantic/Beanie models, TypeScript types plus Zod) so bad data is rejected at both layers.
Best Practices
Start From Queries, Not Entities
Write down the application's operations and how often each runs (for example "show order: 10k/s; update order status: 50/s"). Optimize document shape for the frequent reads, then check the writes remain reasonable.
Keep Documents Reasonably Sized
Aim for documents in the kilobytes, not megabytes. Large documents waste cache memory and network bandwidth even when you need only a few fields.
Index for Your Shape
Embedded fields and array elements are indexable (items.sku), and multikey indexes cover arrays. Design indexes alongside the schema; see MongoDB indexes.
Use Proper Types
Store money as Decimal128, timestamps as Date, and IDs as ObjectId or consistent strings. Mixed types for the same field (a number in some documents, a string in others) break queries and indexes quietly.
Common Mistakes
Unbounded Arrays
Porting a Normalized SQL Schema Directly
One collection per table with joins everywhere means every read needs several $lookup stages. MongoDB can join, but it's optimized for reading whole documents. Denormalize around access patterns.
Massive Numbers of Collections
A collection per user or per tenant multiplies files, indexes, and metadata overhead. Use a tenant field and compound indexes instead.
FAQ
Is MongoDB really schemaless?
The database doesn't require a fixed schema, which makes iteration and polymorphic data easy. But your application depends on document shapes, so you should define and enforce them with validation rules and typed models. "Flexible schema" is more accurate than "schemaless".
When should I use $lookup?
For occasional queries, reporting, and relationships where embedding would duplicate too much. If a $lookup sits on your hottest read path, that's often a sign to embed or copy fields instead.
How do I migrate document shapes?
Common approaches: a schemaVersion field with code that handles both versions and upgrades documents lazily on read or write, plus background scripts that migrate the rest in batches. Avoid a big-bang rewrite of large collections during peak traffic.
When is a relational database a better fit?
When data is highly relational with many-to-many queries across entities, when you rely on ad hoc joins and complex reporting, or when strict cross-entity constraints matter most. See SQL vs NoSQL.
Related Topics
- MongoDB — The database overview
- MongoDB Indexes — Indexing embedded fields and arrays
- MongoDB Aggregation — Querying and reshaping documents
- MongoDB Transactions — When single-document atomicity isn't enough
- SQL vs NoSQL — Choosing a data model
- Domain-Driven Design — Aggregates map naturally to documents