Data Management in Microservices

Data is where microservices get genuinely hard. The core principle, each service owns its data, is what makes services independently deployable and scalable, but it gives up conveniences a monolith takes for granted: ACID transactions across entities, joins across tables, and a single source of truth for reporting. Placing an order that reserves inventory and charges a card now spans three services and three databases.

Microservice architectures replace those conveniences with patterns: sagas for multi-service workflows, the transactional outbox for reliable event publishing, API composition and CQRS read models for cross-service queries, and deliberate data duplication with eventual consistency. Getting service boundaries right, so most operations stay within one service, matters more than any of these patterns.

TL;DR

Quick Example

Order placement across Orders, Inventory, and Payments with a saga and outbox:

Each step is a local ACID transaction. The overall business transaction is eventually consistent, and failures trigger compensations rather than rollbacks.

Core Concepts

Database per Service

Each service owns its schema and storage, and may choose the database that fits (PostgreSQL for orders, Redis for sessions, Elasticsearch for search, a graph DB for recommendations). Other services never query it directly; they use the owner's API, or subscribe to its events. Benefits:

"Database per service" can mean separate servers, separate databases on a shared server, or separate schemas with strict access control. The key is logical ownership and no cross-service table access.

The Shared Database Anti-Pattern

Multiple services reading and writing the same tables creates hidden coupling: schema changes break other services, deployments must be coordinated, and one service's heavy queries slow others down. It's common as a transitional state while decomposing a monolith (see strangler fig), but plan to move each table to a single owner.

Distributed Transactions and Sagas

Two-phase commit (2PC) across services is fragile, blocks under failures, and isn't supported by most modern databases and brokers. Sagas replace it:

Sagas lack isolation: intermediate states are visible, so use countermeasures like semantic locks (a PENDING status), commutative updates, and re-checking conditions.

Reliable Event Publishing: The Outbox

A service that updates its database and then publishes an event can crash in between, losing the event, or publish an event for a transaction that rolls back. The transactional outbox writes the event to an outbox table in the same transaction as the state change, and a separate process (a polling publisher, or change data capture with Debezium) publishes it to the broker. Consumers must be idempotent, since at-least-once delivery means duplicates. See outbox pattern.

Querying Across Services

See CQRS. For search across domains, a dedicated search index fed by events (Elasticsearch) is common.

Eventual Consistency in Practice

Reporting and Analytics

Operational services shouldn't serve heavy analytical queries. Stream changes (events or CDC) into a data warehouse or lakehouse, where data from all services can be joined for reporting, BI, and ML, without coupling services' schemas.

Best Practices

Draw Boundaries Around Transactions

If two entities must be updated atomically most of the time, they probably belong in the same service. Use domain-driven design (bounded contexts, aggregates) to find boundaries where consistency needs are local.

Publish Events From the Outbox, Always

Never dual-write (database plus broker) in application code. Outbox plus CDC or relay is the standard, reliable way to publish state changes.

Version Your Events and APIs

Other services depend on your events. Evolve schemas backward-compatibly, use schema registries, and document event contracts. See event schema evolution.

Own Data Quality and Lineage

Each service owns the correctness of its data, including events it publishes. Monitor event lag, dead-letter queues, and reconciliation discrepancies.

Common Mistakes

Splitting Services by Entity

Creating separate Customer, Order, and Product services that constantly need each other's data for every operation produces chatty synchronous calls and distributed transactions everywhere. Split by business capability instead.

Dual Writes

Use the transactional outbox.

Assuming Events Arrive Once and in Order

Brokers deliver at least once, and ordering holds only per partition or key. Handlers that aren't idempotent or order-aware corrupt data under retries and rebalances.

FAQ

Why should each microservice have its own database?

So services can evolve schemas, choose storage technologies, scale, and deploy independently, and so business rules around data are enforced in one place. Shared databases create hidden coupling that undermines those microservice benefits.

How do you handle transactions across microservices?

With sagas: a series of local transactions coordinated through events (choreography) or a coordinator (orchestration), where failures trigger compensating actions. Reliable messaging via the transactional outbox, and idempotent consumers, make sagas robust.

How do I join data from multiple microservices?

Either compose at query time (call each service and merge results) or maintain a read model: a denormalized view built from events published by the owning services. For analytics, replicate data into a warehouse. Direct cross-database joins break service ownership.

Is eventual consistency acceptable?

For many business processes, yes, since real-world processes are often eventually consistent anyway (payments settle, stock is counted). Identify the invariants requiring immediate consistency, keep them within one service, and design UX and reconciliation around the rest.

Related Topics

References