Kafka Exactly-Once Broken by Producer Retries — Fix
Duplicates despite idempotence plus transactions? Learn fencing, read_committed gaps, and retry storms that break EOS in production..
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓A transactional producer you can configure
- ✓Consumer groups and offset commit basics
- ✓Access to producer logs and downstream counts
- EOS needs three parts: idempotent producer, fenced transactions, and read_committed consumers — miss one and duplicates flow
- enable.idempotence dedups retried sends per partition; keep max.in.flight at 5 or below or ordering breaks
- Unique transactional.id per instance lets the broker fence zombies; shared IDs make writers kill each other
- Unbounded retries replay past broker dedup windows after blips — bound them and keep handlers idempotent
Imagine a bank where tellers number deposit slips in order and the vault rejects numbers already filed — that is idempotence. Two-account transfers go in sealed envelopes filed whole or rejected whole — those are transactions. Junior clerks see only filed envelopes, never drafts — that is read_committed. But shared ID badges, out-of-order numbering in a rush, or clerks reading the trash move money twice.
Your ledger shows 4,800 duplicate payouts. The producer config says enable.idempotence=true. Transactions are on. The dashboard claimed exactly-once. And yet finance is holding a spreadsheet where 312 merchants got paid twice — $214,000 duplicated — because exactly-once is a contract with four signatories, and one of yours never signed.
Kafka's EOS story is genuinely strong but narrow: idempotent producers dedup retried sends, transactions make multi-partition writes atomic, and read_committed consumers hide aborted records. Break any one joint — in-flight requests too high, transactional.id shared across replicas, a downstream reader on read_uncommitted, a retry storm that outruns dedup windows — and duplicates flow while every individual setting looks correct.
This advanced guide traces all four joints with production failure modes. You'll learn how sequence numbers and epochs actually work, why fencing exceptions are protection (not errors to suppress), how retry storms defeat broker-side dedup, and how to verify EOS end to end with chaos tests instead of config reviews. By the end you'll treat exactly-once as a tested property, not a flag.
What Exactly-Once Promises (and What It Doesn't)
Exactly-once semantics promises each record takes effect once, even with retries, crashes, and rebalances. Kafka implements it as three cooperating mechanisms, and the promise holds only when all three are configured. Miss one and you silently downgrade to at-least-once while the config still says EOS — the most expensive kind of wrong.
The honest scope matters as much as the machinery. Kafka's EOS covers the produce path (broker dedups retried sends) and the consume-transform-produce loop (transactions make writes plus offset commits atomic). It does not cover your downstream database: if your handler applies a record twice because it replays, Kafka can't retract the second write. True end-to-end EOS always pairs broker guarantees with idempotent handlers — dedup keys, upserts, conditional writes.
Treat the flag as the start of verification, not the end. enable.idempotence=true with transactions enabled passes every config audit and still duplicates under shared transactional.ids, read_uncommitted readers, or retry storms past dedup windows. The sections below dissect each joint with its production failure mode, so you can check the machinery instead of trusting the label. Bookmark this scope before any incident: when duplicates appear, you will know exactly which of the four joints to interrogate first.
Idempotent Producers: enable.idempotence Deep Dive
The idempotent producer attaches a sequence number to every batch per partition (PID plus epoch plus sequence). The broker tracks the highest sequence seen and drops any resend with a sequence it already applied — turning at-least-once retries into exactly-once writes for that partition. This survives transient network errors transparently: the client retries, the broker dedups, your app never knows.
Two boundaries limit the magic. First, ordering: dedup assumes batches arrive in sequence order, which holds only when max.in.flight.requests.per.connection stays at 5 or below (the protocol's safe window). Raise it for throughput and retried batches interleave out of order — sequence gaps the broker can't reconcile, duplicates it can't catch. Second, session scope: PID state resets on producer restart, so pre-restart retries dedup but post-restart replays don't — that's the gap transactions and epochs close.
Keep the baseline boring: enable.idempotence=true, acks=all, retries at max, in-flight at 5, delivery timeout bounded at 120s. Verify the live config with grep, not the wiki — one 'performance tuning' commit raising in-flight to 10 silently voids the guarantee. Idempotence is the foundation; the next sections build fencing and atomicity on top of it.
Transactions: transactional.id Fencing Explained
Transactions extend dedup across restarts and across partitions. Each transactional.id gets an epoch from the broker, bumped whenever a new producer instance claims the ID. The broker rejects writes carrying stale epochs with ProducerFencedException — which means a zombie instance (network-partitioned old primary, resurrected standby) can't silently duplicate the new primary's writes. Fencing is split-brain protection, and fencing exceptions are it working.
The identity discipline is absolute: one live writer per transactional.id, and the ID must be stable across restarts of the same logical writer (statefulset ordinal, stable pod name) but unique across distinct writers. Share one ID between primary and hot standby and they fence each other forever — abort rates spike past 30% while both instances believe they're the victim. Generate IDs from hostname or ordinal and the problem class vanishes.
Structure every loop as begin, produce, sendOffsetsToTransaction, commit — with abort on any exception. The offset commit inside the transaction is what makes consume-transform-produce atomic: either the outputs and the progress marker land together, or neither does. Keep transaction.timeout.ms above your slowest batch so long-but-healthy transactions aren't reaped mid-work, and alert on ProducerFencedException rate — a trickle means failovers, a flood means shared IDs.
read_committed: the Consumer Half Everyone Forgets
Transactions write abort markers for rolled-back work, and consumers choose whether to respect them. read_uncommitted (the default) delivers everything including aborted records — fast, but it exposes ghosts: records visible now, gone on re-read, reconciled as real by anyone downstream. read_committed filters to committed data only, holding back uncommitted reads until the transaction resolves.
The failure mode is organizational, not technical. Your team sets read_committed on the ledger consumer; the analytics team plugs the default console consumer into the same topic for a dashboard; finance reconciles ghosts against the ledger and pages you for 'lost' records that were correctly aborted. One default-config reader anywhere downstream re-exposes every abort the pipeline paid to hide.
Audit readers the way you audit producers: grep isolation.level across every properties file, Connect worker config, and stream topology touching transactional topics. Encode read_committed in shared consumer templates so new services inherit it, and add it to onboarding checklists for topic access — 'which isolation level and why' should gate every new subscription the way schema review gates new producers.
How Retry Storms Break EOS (and How to Bound Them)
Retry storms break EOS at the buffer, not the protocol. With unbounded retries, a 2-minute broker blip queues tens of minutes of unsent records in producer memory. When connectivity returns, the flood replays — exceeding the broker's sequence tracking window, arriving out of order, and duplicating everything the pre-blip instance already sent through a fenced predecessor. The broker dedups what fits its window; the storm doesn't fit.
Bound the blast radius with delivery.timeout.ms (120s is sane for most pipelines) so doomed sends fail into your error handling instead of replaying forever. Add retry.backoff.ms=100 so retries don't hammer a recovering broker into a second outage. Then accept the residual risk structurally: make downstream handlers idempotent on the business key (payout_id upserts, ON CONFLICT DO NOTHING), so any duplicate that escapes the broker becomes a no-op at the sink.
This layered posture — broker dedup first, bounded retries second, idempotent sinks third — is what 'exactly-once' means in production. No single layer holds under all failures; the three together hold under every failure you've tested. And testing is the operative word, which the final section covers.
Verifying EOS End to End With Chaos Tests
Config reviews can't prove EOS — only adversarial tests can. Build a nightly gate that kills producers mid-transaction (kill -9, no flush), partitions networks during commits, and restarts brokers mid-batch, then counts duplicates downstream with a DISTINCT check on the business key. Zero means the guarantee holds; anything else names the leaking joint with a number attached.
Cover the full matrix over time: kill -9 during begin, during send, during commit; broker bounce during commit; standby promotion while primary is partitioned (fencing check); consumer rebalance mid-transaction (offset atomicity check). Each scenario maps to one joint, so failures diagnose themselves. Run the suite after every producer config change — 'performance tuning' commits are the leading cause of silently voided EOS.
Track three rates as leading indicators between chaos runs: transaction abort rate (spikes mean fencing fights or timeouts), ProducerFencedException rate (trickle means failovers, flood means shared IDs), and downstream duplicate rate (any nonzero means a joint is already leaking). Dashboards on these three turn EOS from a launch-day claim into a continuously verified property — which is the only kind worth putting in front of finance.
The $214,000 Duplicate Payout Behind enable.idempotence=true
- A shared transactional.id turns failover into mutual fencing — unique stable IDs are what make fencing protect instead of destroy.
- One read_uncommitted consumer anywhere downstream re-exposes every abort you paid transactions to hide; audit all readers, not just yours.
- Infinite retries plus finite dedup windows guarantee duplicate storms after blips — bound delivery time and keep handlers idempotent regardless.
| File | Command / Code | Purpose |
|---|---|---|
| producer.properties | enable.idempotence=true | Idempotent Producers |
| txn_producer.py | from kafka import KafkaProducer | Transactions |
| consumer.properties | isolation.level=read_committed | read_committed |
| retry_bounds.properties | delivery.timeout.ms=120000 | How Retry Storms Break EOS (and How to Bound Them) |
| chaos_eos.sh | TXN_PRODUCER_PID=$(pgrep -f txn_producer.py) | Verifying EOS End to End With Chaos Tests |
Key takeaways
Common mistakes to avoid
5 patternsRaising max.in.flight.requests past 5 with idempotence on
Sharing one transactional.id across standby replicas
Forgetting read_committed on downstream consumers
Retrying forever with no delivery timeout
Committing consumer offsets outside the transaction
Interview Questions on This Topic
What are the three pieces of Kafka exactly-once?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Kafka. Mark it forged?
5 min read · try the examples if you haven't