Transactional Outbox Pattern

A service saves an order to its database and then publishes OrderPlaced to Kafka. What happens if the database commit succeeds but the process crashes before publishing? The order exists, but no other service ever hears about it. Reverse the order, and you can publish an event for an order that was rolled back. This is the dual-write problem: you can't atomically write to two separate systems (a database and a message broker) without distributed transactions, which most brokers don't support and most teams don't want.

The transactional outbox solves it with a simple idea: write the event to an outbox table in the same database transaction as the business data. A separate relay process then reads the outbox and publishes the events to the broker. Either both the state change and the event are committed, or neither is. It's one of the foundational reliability patterns in event-driven microservices.

TL;DR

Quick Example

Writing business data and the event atomically (PostgreSQL + Node.js):

A polling relay:

Core Concepts

The Dual-Write Problem

Two-phase commit (XA) across the DB and broker is rarely supported or practical. The outbox reduces the problem to a single local transaction.

Relay Options

Polling publisher: a background job queries unpublished rows, publishes them, and marks them published. It's simple and works anywhere. Trade-offs: polling latency and DB load. FOR UPDATE SKIP LOCKED enables concurrent relays.

Change data capture (CDC): a connector tails the database's transaction log (PostgreSQL WAL, MySQL binlog) and turns outbox inserts into broker messages. Debezium's outbox event router is the common implementation, running on Kafka Connect. Benefits: low latency, no polling load, and ordering by commit. Trade-off: operating CDC infrastructure. See change data capture.

With CDC, rows can be deleted right after insertion (or the table kept tiny), since the log captures them.

Delivery Guarantees

The relay can publish and then crash before marking rows published, so it republishes on restart. The outbox therefore gives at-least-once delivery. Consumers must handle duplicates using the event ID. See idempotency.

Ordering

Consumers often need events for the same aggregate in order (OrderPlaced before OrderCancelled). Preserve it by:

The Inbox Pattern

The consumer-side counterpart: record processed event IDs in an inbox table in the same transaction as the consumer's state changes. If a duplicate arrives, the insert conflicts and the handler skips it. Outbox + inbox give effectively-once processing across services.

Event Payload Design

Outbox events are integration events: stable, versioned contracts for other services. Include the event ID, type, aggregate ID, version, timestamp, and a payload with the data consumers need (a "fat" event), or just identifiers (a "thin" event), and consumers call back for details. See domain events and event schema evolution.

Alternatives

Best Practices

Keep the Relay Boring and Observable

Monitor the unpublished row count, the oldest unpublished age, and publish errors. Alert when lag grows, since a stuck relay silently stalls every downstream process.

Clean Up Published Rows

Delete or archive published rows on a schedule (or immediately, with CDC). An ever-growing outbox slows polling and bloats the database.

Use One Outbox per Database

The outbox belongs to the service's own database and schema. Don't share outbox tables across services.

Make Event IDs Deterministic and Unique

Generate event IDs when writing to the outbox, and carry them through to message headers, so consumers can deduplicate reliably across republishing.

Common Mistakes

Publishing Inside the Transaction "Just to Be Safe"

Calling the broker before commit still produces phantom events on rollback, and holds DB locks during network calls. Write to the outbox only.

Assuming Exactly-Once Delivery

Relays republish after crashes. Without consumer idempotency, duplicates cause double charges, emails, or inventory decrements.

Losing Ordering With Parallel Relays

Multiple relay workers publishing the same aggregate's events concurrently can reorder them. Partition relay work by aggregate, or use CDC.

FAQ

What is the transactional outbox pattern?

A way to reliably publish messages when data changes: the service writes the business data and an event record to an outbox table in one local database transaction, and a separate relay process later publishes outbox records to the message broker. It avoids lost or phantom events from dual writes.

Polling or CDC for the outbox relay?

Polling is simpler to start with and needs no extra infrastructure, and it works well at moderate volumes. CDC (like Debezium) offers lower latency, less database load, and strict commit ordering, and it's preferable at scale, or when you already run Kafka Connect.

Does the outbox pattern guarantee exactly-once delivery?

No, it guarantees at-least-once delivery: events are never lost, but may be published more than once. Combine it with idempotent consumers or the inbox pattern to achieve effectively-once processing.

Do I need the outbox if I use event sourcing?

Usually not in the same form. With event sourcing, the event store is the source of truth, and publishing is done by subscribing to it. If events are stored in a relational database, the event table effectively serves as the outbox.

Related Topics

References