Data Management in Microservices
Data is where microservices get genuinely hard. The core principle, each service owns its data, is what makes services independently deployable and scalable, but it gives up conveniences a monolith takes for granted: ACID transactions across entities, joins across tables, and a single source of truth for reporting. Placing an order that reserves inventory and charges a card now spans three services and three databases.
Microservice architectures replace those conveniences with patterns: sagas for multi-service workflows, the transactional outbox for reliable event publishing, API composition and CQRS read models for cross-service queries, and deliberate data duplication with eventual consistency. Getting service boundaries right, so most operations stay within one service, matters more than any of these patterns.
TL;DR
- Database per service: only the owning service reads and writes its data; others use its API or events.
- Shared databases couple schemas, deployments, and scaling. Avoid them, or treat them as a transitional state.
- Replace distributed transactions (2PC) with sagas: sequences of local transactions with compensating actions.
- Publish events reliably with the transactional outbox (plus polling or CDC), never dual writes.
- Query across services with API composition (aggregate at request time) or CQRS read models (pre-joined views built from events).
- Accept eventual consistency where the business allows it, and design UX and reconciliation for it.
Quick Example
Order placement across Orders, Inventory, and Payments with a saga and outbox:
Each step is a local ACID transaction. The overall business transaction is eventually consistent, and failures trigger compensations rather than rollbacks.
Core Concepts
Database per Service
Each service owns its schema and storage, and may choose the database that fits (PostgreSQL for orders, Redis for sessions, Elasticsearch for search, a graph DB for recommendations). Other services never query it directly; they use the owner's API, or subscribe to its events. Benefits:
- Schemas evolve without coordinating with other teams.
- Services scale and fail independently.
- Clear ownership and encapsulation of business rules.
"Database per service" can mean separate servers, separate databases on a shared server, or separate schemas with strict access control. The key is logical ownership and no cross-service table access.
The Shared Database Anti-Pattern
Multiple services reading and writing the same tables creates hidden coupling: schema changes break other services, deployments must be coordinated, and one service's heavy queries slow others down. It's common as a transitional state while decomposing a monolith (see strangler fig), but plan to move each table to a single owner.
Distributed Transactions and Sagas
Two-phase commit (2PC) across services is fragile, blocks under failures, and isn't supported by most modern databases and brokers. Sagas replace it:
- A sequence of local transactions, each publishing an event or message that triggers the next.
- On failure, compensating transactions undo prior steps semantically (refund the payment, release the stock).
- Choreography: services react to each other's events. Orchestration: a coordinator tells each service what to do. See saga pattern and choreography vs orchestration.
Sagas lack isolation: intermediate states are visible, so use countermeasures like semantic locks (a PENDING status), commutative updates, and re-checking conditions.
Reliable Event Publishing: The Outbox
A service that updates its database and then publishes an event can crash in between, losing the event, or publish an event for a transaction that rolls back. The transactional outbox writes the event to an outbox table in the same transaction as the state change, and a separate process (a polling publisher, or change data capture with Debezium) publishes it to the broker. Consumers must be idempotent, since at-least-once delivery means duplicates. See outbox pattern.
Querying Across Services
See CQRS. For search across domains, a dedicated search index fed by events (Elasticsearch) is common.
Eventual Consistency in Practice
- Show pending states to users ("Order received, confirming payment…").
- Design idempotent handlers, and handle out-of-order events (versions, timestamps).
- Add reconciliation jobs that detect and repair drift between services.
- Reserve strong consistency for invariants that truly require it, and keep those within a single service boundary.
Reporting and Analytics
Operational services shouldn't serve heavy analytical queries. Stream changes (events or CDC) into a data warehouse or lakehouse, where data from all services can be joined for reporting, BI, and ML, without coupling services' schemas.
Best Practices
Draw Boundaries Around Transactions
If two entities must be updated atomically most of the time, they probably belong in the same service. Use domain-driven design (bounded contexts, aggregates) to find boundaries where consistency needs are local.
Publish Events From the Outbox, Always
Never dual-write (database plus broker) in application code. Outbox plus CDC or relay is the standard, reliable way to publish state changes.
Version Your Events and APIs
Other services depend on your events. Evolve schemas backward-compatibly, use schema registries, and document event contracts. See event schema evolution.
Own Data Quality and Lineage
Each service owns the correctness of its data, including events it publishes. Monitor event lag, dead-letter queues, and reconciliation discrepancies.
Common Mistakes
Splitting Services by Entity
Creating separate Customer, Order, and Product services that constantly need each other's data for every operation produces chatty synchronous calls and distributed transactions everywhere. Split by business capability instead.
Dual Writes
Use the transactional outbox.
Assuming Events Arrive Once and in Order
Brokers deliver at least once, and ordering holds only per partition or key. Handlers that aren't idempotent or order-aware corrupt data under retries and rebalances.
FAQ
Why should each microservice have its own database?
So services can evolve schemas, choose storage technologies, scale, and deploy independently, and so business rules around data are enforced in one place. Shared databases create hidden coupling that undermines those microservice benefits.
How do you handle transactions across microservices?
With sagas: a series of local transactions coordinated through events (choreography) or a coordinator (orchestration), where failures trigger compensating actions. Reliable messaging via the transactional outbox, and idempotent consumers, make sagas robust.
How do I join data from multiple microservices?
Either compose at query time (call each service and merge results) or maintain a read model: a denormalized view built from events published by the owning services. For analytics, replicate data into a warehouse. Direct cross-database joins break service ownership.
Is eventual consistency acceptable?
For many business processes, yes, since real-world processes are often eventually consistent anyway (payments settle, stock is counted). Identify the invariants requiring immediate consistency, keep them within one service, and design UX and reconciliation around the rest.
Related Topics
- Microservices — The architecture overview
- Saga Pattern — Distributed business transactions
- Outbox Pattern — Reliable event publishing
- CQRS — Separate read models
- Change Data Capture — Streaming database changes
- Domain-Driven Design — Finding service boundaries