Event Schema Evolution

In an event-driven system, events are contracts between teams that often don't coordinate deployments: a producer publishes OrderPlaced, and a dozen consumers (billing, analytics, email, search) read it. Events are also durable: they sit in Kafka topics for days or forever, and in event-sourced systems they live as permanent history. When the producer changes the event's shape, every consumer, and every stored event, must keep working.

Schema evolution is the discipline of changing event formats safely. It borrows ideas from API versioning, but it's stricter: consumers may read old and new events side by side, you can't "force upgrade" data already written, and there's no request-response negotiation.

TL;DR

Quick Example

An Avro schema evolved in a backward- and forward-compatible way:

Compatibility enforced by the registry, and checked in CI:

Core Concepts

Compatibility Modes

For shared event streams with independent teams and long retention, full transitive compatibility is the safest default.

Safe and Breaking Changes

Format-Specific Rules

Schema Registries

A schema registry (Confluent Schema Registry, Apicurio, AWS Glue Schema Registry) stores versioned schemas per subject (usually per topic value), enforces compatibility on registration, and lets serializers embed a small schema ID in each message. Consumers fetch schemas by ID to deserialize. Combined with producer settings that forbid auto-registration in production, it stops incompatible changes before they reach the topic.

Handling Breaking Changes

When a change can't be made compatible:

  1. New event type or version: OrderPlacedV2 (or com.acme.orders.v2.OrderPlaced), published alongside v1 during a migration window. Consumers migrate, then v1 is retired.
  2. New topic: orders.v2, with a bridge that translates v1 ↔ v2 while consumers move.
  3. Expand and contract: add the new field and populate both, migrate consumers to the new field, then stop populating and remove the old one (after retention expires).

Communicate deprecation timelines, and track which consumers still read old versions.

Upcasting Stored Events

In event-sourced systems, old events stay in the store forever. Upcasters transform old versions into the current shape when events are loaded (OrderPlaced v1 → v2: split name into first_name/last_name, and default currency), so domain code only handles the latest version. Keep upcasters small, tested, and chained (v1 → v2 → v3).

Envelopes and Metadata

Standardize an envelope with metadata separate from the payload: event ID, type, schema version, source, timestamp, correlation and causation IDs, and the partition key. CloudEvents is a vendor-neutral specification for this metadata, with bindings for Kafka, HTTP, and other transports. See domain events.

Best Practices

Treat Events as Public APIs

Review schema changes like API changes: owners, changelogs, deprecation policies, and documentation. The producer team owns the contract, and consumers depend on it.

Be a Tolerant Reader

Consumers should ignore unknown fields, handle unknown enum values gracefully, and read only the fields they need. Tolerant consumers make forward-compatible evolution possible.

Test Compatibility in CI

Run registry compatibility checks (or tools like Buf's breaking for Protobuf) on every pull request that changes a schema. Add consumer-driven contract tests so consumers declare the fields they rely on.

Never Change Meaning Silently

Changing units, semantics, or the interpretation of a field is the most dangerous change, because schemas can't detect it. Add a new field with a new name instead.

Common Mistakes

Schemaless JSON Events

Without schemas, every change is a guess about what consumers parse, and breakage shows up in production. Even with JSON, publish JSON Schemas and validate at produce time.

Checking Compatibility Only Against the Latest Version

With long retention or replays, consumers may read events from several versions back. Use transitive compatibility.

Reusing Protobuf Field Numbers

Reusing a removed field's number makes old data deserialize into the wrong field. Mark removed numbers and names as reserved.

FAQ

What is the difference between backward and forward compatibility?

Backward compatibility means new readers can read old data, so you upgrade consumers first. Forward compatibility means old readers can read new data, so you upgrade producers first. Full compatibility supports both, which lets producers and consumers deploy in any order.

Should I use Avro, Protobuf, or JSON Schema for events?

Avro is popular in the Kafka ecosystem, for compact encoding and rich schema resolution. Protobuf suits teams already using gRPC, and it offers strong tooling (Buf) and code generation. JSON Schema is human-readable and easy to adopt, but larger and looser. Any works well with a schema registry and discipline.

How do I rename a field in an event?

Don't rename it in place. Add the new field, populate both during a migration period, move consumers to the new field, then deprecate and later remove the old one. Avro aliases can help readers map old names, but coordinate carefully.

Do I need a schema registry?

For small systems with few consumers, schemas in a shared repository with CI checks can suffice. As the number of teams, topics, and consumers grows, a registry that enforces compatibility at produce time prevents incompatible events from ever being published.

Related Topics

References