Kafka Topics & Partitions

In Kafka, a topic is a named stream of records, such as orders or page-views. Each topic is split into partitions: ordered, append-only logs spread across brokers. Partitions are the unit of parallelism (more partitions allow more consumers), ordering (order is guaranteed only within a partition), and replication (each partition is copied to several brokers for durability).

Nearly every important Kafka design decision comes back to partitions: which key to use, how many partitions to create, how to keep related events in order, and how much data to retain. Getting them right early matters, because some choices, like partition count with keyed data, are painful to change later.

TL;DR

Quick Example

Producing with a key keeps each customer's events ordered:

Core Concepts

Partitions and Offsets

A partition is a log of records, each identified by a sequential offset (0, 1, 2…). Records are never modified; new records are appended at the end. Consumers track their position by offset, so different consumer groups can read the same data at different speeds, rewind to reprocess, or start from the beginning. On disk, each partition is a series of segment files, which makes retention cheap: old segments are simply deleted.

Keys and Ordering

The producer's partitioner chooses a partition for each record:

Choose the key based on what must stay ordered: customerId for per-customer event order, orderId for per-order state changes, deviceId for telemetry. Kafka gives no global ordering across partitions. If you truly need a total order, you need one partition, and that caps throughput.

Choosing a Partition Count

Partitions set the maximum parallelism: a consumer group can have at most one active consumer per partition. Guidelines:

Replication, Leaders, and ISR

Each partition has a replication factor (usually 3). One replica is the leader, which handles reads and writes, and the others are followers that copy from it. The in-sync replica set (ISR) is the replicas currently caught up. If the leader fails, a new leader is elected from the ISR.

Durability comes from combining producer and topic settings:

With RF=3, min.insync.replicas=2, and acks=all, an acknowledged write survives the loss of any single broker. See Kafka producers.

Retention and Compaction

Best Practices

Pick Keys That Spread Load Evenly

A key with low cardinality (country, status), or one dominated by a few values (one giant tenant), creates hot partitions: one partition, and so one consumer, does most of the work. Use high-cardinality keys, or add a suffix to spread very large entities while keeping order where it matters.

Size Partitions for Peak and Growth

Base partition counts on peak throughput and expected consumer parallelism over the next year or two. Over-partitioning is cheaper to live with than re-partitioning keyed data later.

Use Naming Conventions and Governance

Adopt consistent topic names (<domain>.<entity>.<event>, for example sales.order.created), disable automatic topic creation in production, and manage topics as code (Terraform providers, GitOps operators such as Strimzi). Schema registries enforce the data contract. See Kafka Connect.

Monitor Under-Replicated Partitions

UnderReplicatedPartitions and UnderMinIsrPartitionCount are key health signals: they mean durability is degraded, and producers with acks=all may start failing.

Common Mistakes

Expecting Global Order

Design consumers to rely only on per-key order.

Adding Partitions to a Keyed Topic Casually

Going from 6 to 12 partitions sends cust-42 to a different partition than before. For a while, its new events and old events live in different partitions and can be processed out of order. Plan capacity up front, or migrate to a new topic deliberately.

Replication Factor 1 in Production

A single replica means one broker failure loses data, or makes partitions unavailable until it returns. Use RF 3 for anything important.

FAQ

How many partitions should a topic have?

Enough to support your target consumer parallelism and throughput with headroom, typically a small multiple of the number of consumers you expect. Measure per-consumer throughput, divide target throughput by it, and round up generously for growth. Avoid extremes: one partition caps scaling, and thousands add overhead.

Can I reduce the number of partitions?

No. Kafka doesn't support decreasing partitions. You'd have to create a new topic with fewer partitions and migrate producers and consumers to it.

What's the difference between retention and compaction?

Retention deletes records older than a time or size limit, regardless of key. Compaction keeps the most recent record per key indefinitely and removes superseded ones, so the topic acts like a changelog of current state. Compacted topics can also delete keys via tombstones.

What happens if a broker goes down?

Partitions whose leader was on that broker elect a new leader from the in-sync replicas, and clients automatically redirect. With RF 3 and min.insync.replicas=2, writes continue while one broker is down. If too many replicas are out of sync, acks=all producers get errors rather than silently losing durability.

Related Topics

References