Transactional Outbox Pattern
A service saves an order to its database and then publishes OrderPlaced to Kafka. What happens if the database commit succeeds but the process crashes before publishing? The order exists, but no other service ever hears about it. Reverse the order, and you can publish an event for an order that was rolled back. This is the dual-write problem: you can't atomically write to two separate systems (a database and a message broker) without distributed transactions, which most brokers don't support and most teams don't want.
The transactional outbox solves it with a simple idea: write the event to an outbox table in the same database transaction as the business data. A separate relay process then reads the outbox and publishes the events to the broker. Either both the state change and the event are committed, or neither is. It's one of the foundational reliability patterns in event-driven microservices.
TL;DR
- Dual writes (DB + broker) can't be made atomic. Crashes cause lost or phantom events.
- Outbox: insert event rows into an
outboxtable in the same transaction as the state change. - A relay publishes outbox rows to the broker, by polling or change data capture (e.g. Debezium).
- Delivery is at-least-once, so consumers must be idempotent (the inbox pattern or dedupe keys).
- Preserve per-aggregate ordering with the broker partition key = aggregate ID.
- Clean up published rows, and monitor relay lag.
Quick Example
Writing business data and the event atomically (PostgreSQL + Node.js):
A polling relay:
Core Concepts
The Dual-Write Problem
Two-phase commit (XA) across the DB and broker is rarely supported or practical. The outbox reduces the problem to a single local transaction.
Relay Options
Polling publisher: a background job queries unpublished rows, publishes them, and marks them published. It's simple and works anywhere. Trade-offs: polling latency and DB load. FOR UPDATE SKIP LOCKED enables concurrent relays.
Change data capture (CDC): a connector tails the database's transaction log (PostgreSQL WAL, MySQL binlog) and turns outbox inserts into broker messages. Debezium's outbox event router is the common implementation, running on Kafka Connect. Benefits: low latency, no polling load, and ordering by commit. Trade-off: operating CDC infrastructure. See change data capture.
With CDC, rows can be deleted right after insertion (or the table kept tiny), since the log captures them.
Delivery Guarantees
The relay can publish and then crash before marking rows published, so it republishes on restart. The outbox therefore gives at-least-once delivery. Consumers must handle duplicates using the event ID. See idempotency.
Ordering
Consumers often need events for the same aggregate in order (OrderPlaced before OrderCancelled). Preserve it by:
- Publishing with the aggregate ID as the partition key, so a Kafka partition keeps order per key. See Kafka partitions.
- Relaying in commit order per aggregate (CDC does this naturally. With polling, order by a sequence, and beware concurrent transactions committing out of
created_atorder). - Including a per-aggregate version in events, so consumers can detect gaps or stale updates.
The Inbox Pattern
The consumer-side counterpart: record processed event IDs in an inbox table in the same transaction as the consumer's state changes. If a duplicate arrives, the insert conflicts and the handler skips it. Outbox + inbox give effectively-once processing across services.
Event Payload Design
Outbox events are integration events: stable, versioned contracts for other services. Include the event ID, type, aggregate ID, version, timestamp, and a payload with the data consumers need (a "fat" event), or just identifiers (a "thin" event), and consumers call back for details. See domain events and event schema evolution.
Alternatives
- Listen to yourself: publish the event only, and have the service consume its own event to update its database. It's simple, but the service's state becomes eventually consistent with its own writes.
- Event sourcing: the event store is the database, and subscriptions publish events. See event sourcing.
- Transactional producers within the same system (such as Kafka transactions for Kafka-to-Kafka processing). They don't cover database writes.
- Durable execution engines persist workflow steps and retry them reliably. See durable execution.
Best Practices
Keep the Relay Boring and Observable
Monitor the unpublished row count, the oldest unpublished age, and publish errors. Alert when lag grows, since a stuck relay silently stalls every downstream process.
Clean Up Published Rows
Delete or archive published rows on a schedule (or immediately, with CDC). An ever-growing outbox slows polling and bloats the database.
Use One Outbox per Database
The outbox belongs to the service's own database and schema. Don't share outbox tables across services.
Make Event IDs Deterministic and Unique
Generate event IDs when writing to the outbox, and carry them through to message headers, so consumers can deduplicate reliably across republishing.
Common Mistakes
Publishing Inside the Transaction "Just to Be Safe"
Calling the broker before commit still produces phantom events on rollback, and holds DB locks during network calls. Write to the outbox only.
Assuming Exactly-Once Delivery
Relays republish after crashes. Without consumer idempotency, duplicates cause double charges, emails, or inventory decrements.
Losing Ordering With Parallel Relays
Multiple relay workers publishing the same aggregate's events concurrently can reorder them. Partition relay work by aggregate, or use CDC.
FAQ
What is the transactional outbox pattern?
A way to reliably publish messages when data changes: the service writes the business data and an event record to an outbox table in one local database transaction, and a separate relay process later publishes outbox records to the message broker. It avoids lost or phantom events from dual writes.
Polling or CDC for the outbox relay?
Polling is simpler to start with and needs no extra infrastructure, and it works well at moderate volumes. CDC (like Debezium) offers lower latency, less database load, and strict commit ordering, and it's preferable at scale, or when you already run Kafka Connect.
Does the outbox pattern guarantee exactly-once delivery?
No, it guarantees at-least-once delivery: events are never lost, but may be published more than once. Combine it with idempotent consumers or the inbox pattern to achieve effectively-once processing.
Do I need the outbox if I use event sourcing?
Usually not in the same form. With event sourcing, the event store is the source of truth, and publishing is done by subscribing to it. If events are stored in a relational database, the event table effectively serves as the outbox.
Related Topics
- Event-Driven Architecture — Pillar overview
- Change Data Capture — Log-based outbox relays
- Idempotency — Handling duplicate deliveries
- Microservices Data Management — Data ownership across services
- Saga Pattern — Multi-step processes using events
- Domain Events — What to put in events