kafka idempotent producer exactly once setup
Sets up a Kafka idempotent producer for exactly-once delivery semantics. Use when duplicate messages from producer retries corrupt downstream state, when you need the producer side of exactly-once, or when configuring enable.idempotence and transactional producers. Not for consumer lag monitoring, for message size errors, or for fixing rebalance storms.
TL;DR
Set enable.idempotence=true (which implies acks=all and retries unlimited with max.in.flight.requests.per.connection<=5) for retry-safe dedup; add a transactional.id and the init/send/commit calls for true exactly-once across partitions. Idempotence dedupes retries, transactions make multi-partition writes atomic.
kafka idempotent producer exactly once setupUse this when
- Producer retries create duplicate messages downstream
- You need exactly-once semantics on the produce path
- You are writing to multiple partitions atomically
Not for this skill when
- You are monitoring consumer lag
- Messages exceed size limits
- Consumers rebalance constantly
Steps
- Enable idempotence. Three settings, one concept:
producer = Producer({
"bootstrap.servers": "YOUR_HOST:9092",
"enable.idempotence": True,
"acks": "all",
"retries": 2147483647,
"max.in.flight.requests.per.connection": 5,
})Expected output: the producer assigns each message a sequence number per partition, and the broker rejects duplicates from retries. Modern clients set the dependent values automatically when idempotence is on; the explicit listing is for older clients.
- Understand what this does and does not guarantee:
Guaranteed: no duplicates from producer retries within one producer session.
NOT guaranteed: atomicity across partitions, or dedup across producer restarts
(that needs transactions, step 3).Expected output: correct expectations. Idempotence alone is enough when each message is independent and retries are the only dup source.
- For atomic multi-partition writes, add transactions:
producer.init_transactions()
producer.begin_transaction()
producer.produce("orders", payload, "42")
producer.produce("order-events", event, "42")
producer.commit_transaction()Expected output: both messages commit atomically or neither does. Each producer instance needs a unique transactional.id so the broker can fence zombies after a restart.
- Make the consumer side cooperate with transactions:
consumer = Consumer({
"bootstrap.servers": "YOUR_HOST:9092",
"group.id": "orders-consumer",
"isolation.level": "read_committed",
})Expected output: consumers skip aborted transactional messages. Without read_committed, consumers see uncommitted messages and exactly-once is broken on the read path.
- Handle the failure modes explicitly:
try:
producer.commit_transaction()
except KafkaException:
producer.abort_transaction()Expected output: failed transactions abort cleanly instead of half-committing. Abort on any error during the transaction, then retry the whole unit of work.
Variant phrasings
kafka exactly once semantics producer
Idempotence plus transactions on the produce side, read_committed on the consume side (steps 1, 3, 4). All three pieces or it is not exactly-once.
kafka enable.idempotence vs transactions
Idempotence dedupes retries cheaply; transactions add atomicity across partitions and producer fencing. Start with idempotence, add transactions when you write multiple records that must land together.
kafka duplicate messages on retry
The producer retried a batch whose ack was lost. enable.idempotence=true makes the broker drop the duplicate by sequence number.
Why it happens
Without idempotence, a produce request can succeed on the broker while the ack is lost, so the client retries and the message lands twice. The broker cannot tell retry from new message without the producer's sequence numbers. Transactions extend the same idea with a transaction coordinator so multi-partition writes share one atomic fate.
Edge cases
transactional.idmust be stable per logical producer across restarts, random IDs break fencing.- Transactional message throughput is lower than plain idempotent produce, benchmark before enabling everywhere.
- The transaction coordinator adds latency to the first message of each transaction (
init_transactionsis slow, call it once). - Consumers with
read_committedsee slightly higher end-to-end latency because they wait for the transaction markers.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_jMs3dqRsRvQpQ0cUz0iDyg