Kafka Leader Not Available After Broker Restart
Leader -1 after a broker restart is usually election in flight.
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓A Kafka cluster you can describe topics on
- ✓Broker restart and replication basics
- ✓Access to broker and client logs
- Leader -1 after a restart usually means an election is in flight — verify ISR trends before restarting anything else
- Only an in-sync replica can take over without data loss; unclean election restores writes by discarding history
- acks=all with min.insync.replicas=2 on RF=3 stalls writes (safely) when replicas drop — that's protection working
- Recover with preferred-replica-election, then force clients to refresh metadata before declaring victory
Think of each partition as a classroom needing exactly one teacher (leader). When the teacher steps out (broker restart), the office picks a qualified substitute from trained aides (in-sync replicas) — class pauses briefly. But with no trained aide available, the office must leave class unsupervised or hand it to an untrained volunteer (unclean election) who loses what yesterday covered.
You restart one broker for a routine patch, and within seconds the alerts start: LeaderNotAvailableException across three topics, producers throwing, consumers stalling. Your first instinct is that the restart broke something. In most cases it didn't — you're watching leader election do its job, and the scariest minute of Kafka operations is also the most normal.
Still, 'normal' covers a wide range. A healthy election finishes in under two minutes and nobody pages. A stuck one — no surviving ISR replica, a durability config that can't be satisfied, clients caching the leaderless state — sits at leader -1 for twenty minutes while you restart brokers that were fine, making everything worse.
This guide teaches you to tell those two apart fast. You'll learn what the leader field and ISR list actually say during an outage, when unclean leader election is (rarely) justified, how min.insync.replicas and acks interact under failure, and the exact recovery sequence — preferred-replica-election included — that restores balanced leadership instead of a fragile pile-up on survivors.
What LeaderNotAvailableException Actually Means
Every partition has exactly one leader — the replica that accepts produces and serves reads — plus followers that replicate. Clients learn the leader from metadata responses, and LeaderNotAvailableException is the client's way of saying the metadata listed no leader for that partition. It's a routing failure, not a data verdict: the records may be perfectly safe on disk with nobody currently authorized to serve them.
Two very different situations produce the same exception, and confusing them causes most of the damage in these incidents. In situation one, an election is in flight: the old leader is gone, a new one is being chosen from the in-sync replica set (ISR), and the leader field will repopulate in seconds. In situation two, no ISR member is alive, so no legitimate election can complete and the field sits at -1 indefinitely. The first needs patience; the second needs a replica restored.
Your first command decides which world you're in: topic describe shows Leader and Isr per partition. Leader -1 with live ISR members means wait and re-check. Leader -1 with all replicas offline means waiting is futile. Teams that skip this read and restart brokers instead reliably convert situation one into situation two, because each bounce kills the replicas that were about to win the election.
Broker Restarts and the Leader Election Timeline
A broker restart fires a precise sequence: the controller notices the missing heartbeat, marks its led partitions leaderless, and starts elections among ISR members. Followers that were in sync campaign, the controller picks winners, and the new leaders fetch metadata updates to clients. On a healthy cluster with replication factor 3, this completes in 30-120 seconds — the window where LeaderNotAvailable is expected and harmless.
The timeline stretches when each stage degrades. Slow log recovery on restart (replaying unflushed segments) delays the broker's return to ISR. A second restart mid-election kills candidates and restarts the whole cycle — this is how 90 seconds becomes 20 minutes. Network partitions between controller and replicas add detection delay on top, and clients with long metadata.max.age.ms keep erroring minutes after brokers recovered because nobody told them the news.
Operate restarts like landings: one broker down at a time, verify under-replicated partitions return to zero, then proceed. Automate the verification — a deploy pipeline that restarts broker 2 while 40 partitions are still under-replicated from broker 1 is just an outage with extra steps. The election window is normal; overlapping windows are self-inflicted.
Unclean Leader Election: When No ISR Survives
When no ISR replica is alive, Kafka faces a genuine dilemma: keep the partition offline (durable but unavailable) or elect an out-of-sync replica (available but lossy). The unclean.leader.election.enable flag picks the answer. False — the default and the correct choice for anything valuable — keeps leader at -1 until an ISR member returns. True elects whoever is alive, and whatever committed records they lack are gone forever.
Understand exactly what 'gone' means. The out-of-sync replica missed writes the old leader acknowledged — including records producers received success for under acks=all. After unclean election those records never come back; consumers see the log jump backward and downstream systems inherit a gap no retry can fill. One real incident lost 41 seconds of committed payments this way, and reconciliation took finance a full day.
Reserve true for genuinely expendable data: clickstreams, metrics, debug logs — topics where a gap is a shrug. Document the choice per topic in your topic registry, because the engineer flipping the flag at 3 AM won't otherwise know payments-authorized and clickstream-raw deserve opposite answers. For money topics, the correct response to leader -1 with no ISR is restoring a replica and accepting the downtime, painful as that feels mid-incident.
min.insync.replicas vs acks: the Durability Contract
acks and min.insync.replicas are a pair, and misconfiguring one wastes the other. acks=all tells the producer to wait for every in-sync replica to acknowledge. min.insync.replicas sets the broker-side floor: if the ISR shrinks below it, acks=all produces fail with NotEnoughReplicas instead of risking durability. Together on RF=3 with min.insync.replicas=2, you survive one replica loss with full guarantees and fail loudly (safely) on two.
The classic misconfiguration is acks=all against min.insync.replicas=1, which pays full latency for single-replica durability — the worst of both worlds. The mirror mistake is lowering min.insync.replicas mid-incident to silence NotEnoughReplicas errors, converting a safe stall into silent single-replica writes that the next failure turns into loss.
Treat NotEnoughReplicas during an outage as protection working, not breakage. The correct response is restoring replicas (restart the downed broker, fix its disk, let it catch up) while producers retry with backoff. Monitor ISR shrink as a leading indicator: alert when under-replicated partitions stay nonzero for 5 minutes, so you start the day with a warning instead of ending it with an outage. When in doubt, page the topic owner before flipping the flag — a two-minute discussion beats a two-day reconciliation.
preferred-replica-election: Restoring Balance
After an outage, leadership piles onto survivors. Broker 1 holds 100% of leaders while brokers 2 and 3 sit idle but healthy — and the next routine restart of broker 1 becomes a full outage because every leader lives there. Preferred-replica-election fixes this by moving each partition's leadership back to its preferred replica (the first in the replica list) once that replica is caught up.
The operation is safe by design: it only elects preferred replicas that are in sync, so no data risk attaches. Run it after every ISR recovery, not just incidents — any restart skews leadership a little, and skew compounds across months of patches until one broker secretly leads everything. The targeted JSON-file form lets you move leadership gradually during business hours if a full rebalance feels spicy.
Automate what humans forget. A 10-minute scheduler tick running preferred election (or your operator's auto-rebalance) means post-maintenance skew self-heals overnight. Verify with the leader histogram: roughly equal counts per broker is healthy, 90%-on-one is a pager waiting to happen. One command, run routinely, removes an entire class of 'routine restart becomes outage' surprises. Rehearse single-broker loss in staging quarterly so the first time you see ISR shrink is not during a real outage.
Verifying ISR Recovery and Client Recovery
Broker-side recovery and client-visible recovery are two different events, and incidents stay open in the gap. Producers and consumers cache metadata — including which broker leads each partition — and refresh it on metadata.max.age.ms (default 5 minutes) or on demand after errors. A client that cached the leaderless state can keep throwing LeaderNotAvailable for minutes after the partition is fully healthy.
Close the gap deliberately. Shorten metadata.max.age.ms to 30 seconds on critical clients so post-recovery refresh happens fast, and after any leader incident run an end-to-end probe: produce one record to each affected topic and consume it back. That single round trip proves more than any dashboard — it exercises metadata fetch, leader routing, produce path, ISR write, and fetch path in one shot.
Make the probe part of your runbook's definition of done. 'Topic describe shows leaders' is a broker claim; 'produce plus consume round-trips' is client proof. If the probe fails while brokers look healthy, suspect stale metadata first (restart the client or force refresh) before reopening the broker investigation. Declaring victory from broker metrics alone is how teams close incidents that are still burning.
The Patch Restart That Ate 41 Seconds of Committed Payments
- One broker down at a time, always — the second restart turned a 90-second election into a 23-minute outage plus a data gap.
- Unclean leader election on a money topic is never the shortcut; 41 seconds of committed writes vanished the moment an out-of-sync replica took over.
- Verify ISR recovery per partition and rebalance leadership afterward, or the next maintenance hits the same overloaded survivors.
| File | Command / Code | Purpose |
|---|---|---|
| election_watch.sh | bin/kafka-topics.sh --bootstrap-server $BROKERS \ | Broker Restarts and the Leader Election Timeline |
| server.properties | unclean.leader.election.enable=false | Unclean Leader Election |
| durability_check.sh | bin/kafka-configs.sh --bootstrap-server $BROKERS \ | min.insync.replicas vs acks |
| preferred_election.sh | bin/kafka-preferred-replica-election.sh --bootstrap-server $BROKERS | preferred-replica-election |
| client_verify.sh | bin/kafka-console-producer.sh --bootstrap-server $BROKERS --topic payments-autho... | Verifying ISR Recovery and Client Recovery |
Key takeaways
Common mistakes to avoid
5 patternsRestarting more brokers during the election window
Enabling unclean leader election to 'reduce downtime'
Running acks=all against min.insync.replicas=1
Never running preferred-replica-election after recovery
Trusting stale client metadata after ISR recovery
Interview Questions on This Topic
What does LeaderNotAvailableException mean?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Kafka. Mark it forged?
5 min read · try the examples if you haven't