Home › Data Engineering › Kafka Consumer Lag Growing but CPU Idle — Fix It
Intermediate 5 min · September 23, 2026

Kafka Consumer Lag Growing but CPU Idle — Fix It

Lag climbs while brokers sit idle because your consumer is slow.

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 22 min
  • ✓A running Kafka consumer group you can describe
  • ✓Basic consumer configs (poll records, intervals)
  • ✓Access to consumer logs and JMX metrics
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Lag with idle brokers means your consumer is the bottleneck — brokers can't push faster than you poll and process
  • Cut max.poll.records to 50-200 so each batch finishes well inside max.poll.interval.ms instead of timing out
  • Profile the handler first: one synchronous DB call per record turns 500 records into a batch that never finishes
  • Confirm with kafka-consumer-groups.sh --describe plus GC logs and JMX fetch metrics before scaling anything
✦ Definition~90s read
What is Kafka Consumer Lag Growing but CPU Idle?

Consumer lag is the distance between the newest record on a partition (log-end-offset) and the last offset your group committed. If a partition ends at offset 900,000 and your group committed 850,000, lag is 50,000 records waiting. Sum across partitions for the group total.

★
Picture a warehouse conveyor belt feeding packing tables.

Lag is normal in bursts — producers spike, consumers catch up. Lag is an incident when it grows monotonically for 15+ minutes while brokers look healthy.

The consumer group protocol explains the idle-broker paradox. Consumers pull (poll) instead of receiving pushes, and the broker does bounded work per fetch — a sequential read, usually from page cache. Your consumer's poll loop does unbounded work per batch: deserialize, validate, call downstream services, write to databases.

When that loop averages slower than the produce rate, committed offsets fall behind end offsets and lag grows, no matter how bored the brokers look.

Four consumer-side limits decide the outcome: handler speed (seconds per record), batch sizing (max.poll.records against max.poll.interval.ms), JVM liveness (GC pauses vs heartbeat deadlines), and group stability (rebalances replaying uncommitted work). Brokers only matter when fetch latency or throttling shows they're struggling — and then CPU, disk, and request metrics will say so loudly.

Green brokers plus red lag is the consumer telling on itself.

Plain-English First

Picture a warehouse conveyor belt feeding packing tables. The belt (Kafka broker) runs smooth and quiet, but boxes pile up because one packer opens every box, phones the supplier, waits on hold, then packs it. The belt isn't broken — the packer is slow. Speeding up the belt changes nothing. You fix it by giving the packer smaller stacks of boxes, a faster phone routine, and extra hands — that's exactly what tuning max.poll.records, handlers, and consumer count does.

It's 10 AM and the lag dashboard is a ski slope pointing the wrong way: 2.4 million messages behind and climbing 40,000 every minute. You pull up the broker dashboards expecting fire and find nothing — CPU at 15%, disks calm, network flat. Somebody says the cluster is fine and closes the tab. They're half right, and that half will cost you the afternoon.

Idle brokers with growing lag is Kafka's classic misdirection. Brokers only store and serve; they can't force your consumer to poll faster. When lag grows and brokers yawn, the bottleneck lives in your consumer process: a handler that takes 800 ms per record, a batch sized for a benchmark you never ran, a garbage collector freezing the world for 30 seconds, or a rebalance loop replaying the same records forever.

This guide walks the exact diagnosis order that separates those four causes in minutes. You'll learn to read per-partition lag instead of totals, size max.poll.records from measured timings, spot GC stalls in logs, and confirm rebalance storms with two greps. By the end, idle brokers won't fool you again — you'll know precisely which consumer-side knob to turn.

Why Lag Grows While Brokers Look Idle

Consumer lag is the gap between what producers appended and what your group committed: LAG = log-end-offset minus committed offset, per partition. Brokers track both numbers, but neither number measures broker effort. Serving a fetch is cheap — a sequential disk read the page cache usually handles. Processing the fetched records is where real work happens, and that work lives entirely in your consumer process.

That's why idle brokers prove nothing about consumer health. A broker at 15% CPU happily serves 50 MB/s of fetches while your consumer chews one record per second through a blocking HTTP call. The end offset races ahead, commits crawl, and the subtraction grows. Every minute of slow processing adds 60 seconds of catch-up debt, and the debt compounds because new records keep arriving while you grind through old ones.

The mental shift that fixes this class of incident: treat brokers and consumers as separate systems with separate dashboards. Broker CPU, disk, and request latency tell you about the cluster. Per-partition lag, batch processing time, and commit rate tell you about the consumer. When the first set is green and the second is red, stop touching the cluster — you'll find the answer in poll-loop timings, handler profiles, and GC logs.

📊 Production Insight
A team doubled broker sizes during a 2.4M lag incident and changed nothing, because fetches were already served from page cache in 8 ms. Rule: green broker dashboards plus red lag means hands off the cluster.
🎯 Key Takeaway
Lag measures consumer progress, not broker effort — idle brokers with rising lag always point at the consumer side.

max.poll.records Too High: Big Batches, Slow Commits

max.poll.records caps how many records one poll() call returns — default 500. max.poll.interval.ms is the deadline to call poll() again before the coordinator declares you dead — default 300,000 ms (5 minutes). These two knobs form a contract: everything you fetched must be processed and the next poll() issued before the interval expires. Break the contract and your partitions get revoked mid-batch.

Do the arithmetic openly. If your handler needs 1.2 seconds per record (a realistic number with one downstream call), 500 records need 600 seconds — double the default interval. The consumer looks busy the whole time, brokers look idle, and the group rebalances every 5 minutes forever. Cut max.poll.records to 100 and the same work needs 120 seconds, comfortably inside the limit with room for GC wobble.

Size from measurements, not vibes. Time your slowest realistic batch on production-shaped data — big records, cold caches, slow downstream — then pick max.poll.records so the worst case lands under half the interval. That 2x headroom absorbs traffic spikes and minor GC pauses without a revoke. Re-measure after every handler change, because one new API call per record can silently double batch time and restart the whole cycle.

consumer.propertiesPROPERTIES
1
2
3
4
5
6
7
8
9
10
11
# consumer.properties — size batches from measured timings
max.poll.records=100
max.poll.interval.ms=300000
max.poll.interval.ms=300000
session.timeout.ms=10000
heartbeat.interval.ms=3000
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor

# Rule of thumb baked in:
# worst_batch_seconds = max.poll.records * slowest_seconds_per_record
# 100 records * 1.2s = 120s, safely under half the 300s interval
⚠ Don't Widen the Timeout to Hide Slow Batches
Never raise max.poll.interval.ms to paper over slow batches. A 30-minute interval just lets a stuck consumer hold partitions hostage for 30 minutes before anyone notices. Fix the batch first, then tune the timeout.
📊 Production Insight
A consumer with max.poll.records=1000 and 800 ms records needed 800s per batch against a 300s limit — revoked every cycle. Rule: worst-case batch under 50% of max.poll.interval.ms.
🎯 Key Takeaway
Batch size times seconds-per-record must land under half the poll interval — measure the worst case, then set max.poll.records.

Slow Consumer Processing: Finding the Hot Path

Slow processing is the most common root cause and the hardest to see from Kafka dashboards, because the consumer never errors — it just takes forever. The usual villain is blocking I/O inside the record loop: a pricing API call, a Postgres upsert, an S3 read, each innocent alone and lethal at 500x. One 300 ms call per record caps you at ~3 records/sec; a 1,000-record batch then needs over 5 minutes and the interval kills you first.

Profile before optimizing. Log batch start, batch end, and record count on every poll cycle — three lines that immediately reveal per-record cost. If per-record time dominates, attack it structurally: batch the writes (200-row inserts instead of 200 single upserts), cache the hot reads (a 5-minute TTL on pricing data hit 94% in one real incident), or fan CPU-bound work to a bounded thread pool while keeping offset commits strictly ordered after the full batch.

Keep one correctness rule sacred: commit offsets only after the entire batch is durably handled. Parallelizing handlers tempts people into committing early, which trades duplicates (annoying, recoverable) for data loss (silent, permanent). Slow-but-correct beats fast-and-lossy every time — then make correct fast with batching and caching.

consumer_poll_loop.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import time
import logging
from kafka import KafkaConsumer

consumer = KafkaConsumer(
    "recs",
    bootstrap_servers="kafka:9092",
    group_id="recs-group",
    max_poll_records=100,
    enable_auto_commit=False,
)
log = logging.getLogger("poll-loop")

while True:
    batch = consumer.poll(timeout_ms=5000, max_records=100)
    if not batch:
        continue
    start = time.monotonic()
    count = sum(len(r) for r in batch.values())
    # process whole batch FIRST (batched writes, cached lookups)
    process_batch(batch)
    consumer.commit_sync()  # commit only after full batch succeeds
    elapsed = time.monotonic() - start
    log.info("batch=%d took=%.1fs per_record=%.3fs", count, elapsed, elapsed / max(count, 1))
📊 Production Insight
A 94%-hit-rate cache on one lookup took batches from 800s to 90s without touching Kafka config. Rule: the handler is guilty until batch timings prove otherwise.
🎯 Key Takeaway
Log batch timings to expose per-record cost, then batch writes, cache reads, and parallelize CPU work — committing only after full batches.

GC Pauses That Stall the Poll Loop

The JVM can stall your consumer without executing a single line of your code. A full garbage collection freezes all threads — including the background heartbeat thread that tells the coordinator you're alive. Miss enough heartbeats past session.timeout.ms and you're evicted: partitions revoke, lag jumps, and the next consumer replays everything you processed but never committed.

These stalls are invisible in application logs because nothing runs during the pause — the log just resumes 30 seconds later as if nothing happened. The evidence lives in GC logs, which most teams don't ship. A consumer on a default 1 GB heap doing heavy JSON parsing can easily hit 20-60 second full GCs under load, and each one is a guaranteed rebalance.

Treat GC observability as mandatory consumer infrastructure. Run every consumer JVM with unified GC logging to a shipped file, set -Xms equal to -Xmx so the heap never resizes under load, and prefer G1GC or ZGC for pause-sensitive poll loops. Alert on any single pause over 5 seconds — that's half a session timeout gone in one freeze. When heartbeat failures correlate with GC pauses in the same minute, you've found your culprit and no handler rewrite will help until the heap is fixed.

gc_check.shBASH
1
2
3
4
5
6
7
8
9
10
11
# Run consumers with GC logging always on (Java 11+ unified logging)
export KAFKA_HEAP_OPTS="-Xms2g -Xmx2g"
export KAFKA_JVM_PERFORMANCE_OPTS="-server -XX:+UseG1GC \
  -Xlog:gc*:file=/var/log/app/gc.log:time,uptime,level,tags \
  -XX:MaxGCPauseMillis=200"

# Spot the killer pauses (anything over 5s risks a revoke)
grep -E 'Pause Full|pause.*[0-9]{5,}ms' /var/log/app/gc.log | tail -20

# Correlate with heartbeat failures in the same window
grep -E 'Heartbeat failed|Revoked partitions' /var/log/app/consumer.log | tail -20
📊 Production Insight
30-second full GCs on a 1 GB heap evicted consumers every few minutes with zero app-log evidence. Rule: correlate heartbeat failures with GC timestamps before rewriting handlers.
🎯 Key Takeaway
Full GCs freeze heartbeats and trigger evictions — ship GC logs, size heaps fixed, and alert on pauses over 5 seconds.

Rebalance Storms That Reset Your Progress

Rebalances turn slow consumers into stuck consumers. Every revoke discards processed-but-uncommitted work, and the next owner replays it from the last commit — paying the slow-handler cost all over again. With the default eager assignor, one member leaving pauses the entire group; with rolling deploys on 12 pods, the group can spend more time rebalancing than processing, and lag climbs in the famous sawtooth: drains a little, jumps back, drains, jumps.

Two settings end most storms. Static membership (group.instance.id set to a stable pod-unique value) lets a restarted member reclaim its partitions without a full reshuffle — the coordinator recognizes it instead of treating it as a stranger. The cooperative sticky assignor migrates only the partitions that must move instead of revoking everything, so the group keeps processing through membership changes.

Verify calm the same way you diagnosed the storm: count 'Revoked partitions' lines per hour (steady groups should show near zero) and watch per-partition lag drain monotonically after a deploy. If a rolling restart still causes a group-wide stop, one consumer is missing its group.instance.id — grep the configs, because a single dynamic member can force eager behavior for everyone.

membership_fix.shBASH
1
2
3
4
5
6
7
8
9
10
11
# Stable identity + cooperative assignment ends deploy storms
# consumer.properties additions:
# group.instance.id=recs-consumer-${HOSTNAME}
# partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor

# Count rebalance damage over the last hour (want near zero on steady groups)
grep -cE 'Revoked partitions|Attempt to heartbeat failed' /var/log/app/consumer.log

# Verify lag drains monotonically after the fix (no sawtooth jumps)
kafka-consumer-groups.sh --bootstrap-server $BROKERS \
  --describe --group recs-group | awk 'NR>1 {print $2, $5}'
📊 Production Insight
14 deploy-triggered revokes per day replayed hours of processed records until group.instance.id was set. Rule: rolling restarts must cause zero group-wide stops.
🎯 Key Takeaway
Static membership plus cooperative assignment lets restarts reclaim partitions quietly — verify with revoke counts near zero.

Scaling Out: Partitions, Threads, and Fetch Tuning

Scaling out is the last step, not the first — but when the handler is tuned and batches fit, it's the right one. The hard ceiling is partition count: a group can't assign one partition to two consumers, so the 7th consumer on a 6-partition topic idles forever while you pay for its pod. Check kafka-topics --describe first; if consumers already equal partitions, add partitions (accepting the key-ordering change) before adding pods.

Use JMX fetch metrics to confirm headroom exists before scaling. records-lag-max rising with a healthy fetch-rate means records arrive fine and processing is the limit — more consumers help. Near-zero fetch-rate means records aren't arriving (throttling, network, fetch config), and extra consumers will idle just like the current ones. Scale the bottleneck you measured, not the one you assumed.

Roll out gradually and watch the drain rate, not the absolute lag. Double consumers, confirm records/sec roughly doubles and per-partition lag falls evenly. If one partition stays hot while others drain, you've got key skew — no consumer count fixes a single key holding 80% of traffic. Fix the partition key or split the hot entity before burning more pods. Start with one extra consumer per hot partition, confirm the drain rate climbs linearly, and stop when per-partition lag falls evenly — linear gains mean you found the true ceiling.

scale_check.shBASH
1
2
3
4
5
6
7
8
9
10
# You cannot use more consumers than partitions — check first
kafka-topics.sh --bootstrap-server $BROKERS --describe --topic recs

# Per-partition lag: find hot partitions averages hide
kafka-consumer-groups.sh --bootstrap-server $BROKERS \
  --describe --group recs-group

# JMX fetch health per client (jmxterm one-liner)
echo "get -b kafka.consumer:type=consumer-fetch-manager-metrics,client-id=* records-lag-max,fetch-rate" \
  | java -jar jmxterm.jar -l localhost:9999 -n
📊 Production Insight
A team added 6 pods to a 6-partition topic and 6 idled; adding partitions then pods drained 2.4M in 47 minutes. Rule: partitions first, consumers second, skew check always.
🎯 Key Takeaway
Consumers can't exceed partitions — confirm JMX headroom, scale gradually, and watch drain rate plus per-partition spread.
● Production incidentPOST-MORTEMseverity: high

The 800 ms Lookup That Built a 2.4M Lag Behind Idle Brokers

Symptom
LAG on the recommendations group climbed past 2.4 million by 11:30 AM, growing ~40,000/min while broker CPU sat at 12-15% and P99 fetch latency stayed under 8 ms. Consumer logs showed 'Attempt to heartbeat failed' and 'Revoked partitions [recs-0..11]' every 5-6 minutes like clockwork. Personalized carousels went stale for 31% of afternoon traffic.
Assumption
The team assumed lag meant the cluster was under-provisioned, so they doubled the broker instance sizes first. When that changed nothing, they assumed the topic needed more partitions and raised it from 12 to 24 — but the consumer count stayed at 6, so half the new partitions sat idle and lag kept climbing. Nobody timed a single batch for the first three hours.
Root cause
The recommendation consumer fetched max.poll.records=1,000 and called the pricing API synchronously per record at ~800 ms each — 800 seconds of work per batch against max.poll.interval.ms=300000 (300 seconds). The coordinator revoked the partitions mid-batch every single cycle, the consumer recommitted nothing, and all 1,000 records replayed on the next assignment. Broker CPU never passed 15% because the cluster was barely touched; the consumer simply never finished a batch.
Fix
They cut max.poll.records from 1,000 to 100, replaced the per-record pricing lookup with a 5-minute TTL cache (hit rate 94%), and batched Postgres writes 200 rows at a time. They also set group.instance.id to the pod name and switched to CooperativeStickyAssignor. Lag drained from 2.4M to zero in 47 minutes at ~850 records/sec, and deploy-triggered revokes dropped from 14 per day to zero.
Key lesson
  • Time one batch before changing any infrastructure — 1,000 records at 800 ms each can never fit a 5-minute poll interval, and no broker size fixes arithmetic.
  • Per-record external calls are the usual killer; a cache with 94% hit rate did more than doubling broker hardware ever could.
  • Set static membership on every consumer from day one — without group.instance.id, each deploy replays hours of already-processed records.
Production debug guideFive checks that separate slow handlers, oversized batches, GC stalls, fetch problems, and rebalance loops.5 entries
Symptom · 01
Total lag climbs but you don't know if it's skew or uniform slowness
→
Fix
Run kafka-consumer-groups.sh --bootstrap-server $BROKERS --describe --group $GROUP and look at per-partition LAG, not the total. One partition at 800,000 with the rest near zero means key skew or a single-threaded handler — don't scale, fix the hot path. Evenly high lag across all partitions means every consumer is uniformly slow, so move to batch sizing next.
Symptom · 02
Suspected oversized batches vs slow handlers — prove which one
→
Fix
Time a real batch end to end: log timestamps around poll() and around the commit, then compute records * seconds-per-record. If 500 records at 1.2s each equals 600s against max.poll.interval.ms=300000 (300s), the batch can never finish. Cut max.poll.records to 50 and re-measure until the worst-case batch lands under 150s — half the interval.
Symptom · 03
Processing looks fast but members keep getting evicted
→
Fix
Check GC directly: grep -E 'Pause Full|pause.ms' /var/log/app/gc.log | tail -20 and look for any pause over 5000 ms. Then confirm heartbeat starvation with grep -E 'Heartbeat failed|Member . sending LeaveGroup|Revoked partitions' /var/log/app/consumer.log | tail -20. Pauses plus heartbeat failures mean the JVM, not the handler, is your bottleneck — fix heap and collector before touching code.
Symptom · 04
Can't tell fetch stalls from processing stalls
→
Fix
Pull JMX fetch metrics to separate fetch stalls from processing stalls: check records-lag-max and fetch-rate per consumer (via jmxterm or your agent: kafka.consumer:type=consumer-fetch-manager-metrics,client-id=*). Steady fetch-rate with rising records-lag-max proves records arrive fine but aren't processed. Near-zero fetch-rate instead points at fetch.max.wait.ms, network, or broker-side throttling.
Symptom · 05
Lag moves in sawtooth steps — drains, then jumps back up
→
Fix
Count rebalance damage: grep -cE 'Revoked partitions|Attempt to heartbeat failed|rebalance' /var/log/app/consumer.log over the last hour, and check group.instance.id is actually set via grep -E 'group.instance.id|partition.assignment.strategy' /etc/app/consumer.properties. More than 2-3 revokes per hour on a steady group means membership churn — set a stable group.instance.id and CooperativeStickyAssignor, then watch the revoke count drop to near zero.
Consumer Lag With Idle Brokers: Diagnose at a Glance
Root CauseHow to ConfirmFixPrevention
Slow per-record processingTime one batch: records * seconds-per-record exceeds poll interval; handler spans show DB/HTTP waitsBatch writes, cache hot lookups, bound worker pool for CPU workLoad-test handlers with production-shaped records in CI
max.poll.records too highkafka-consumer-groups --describe shows sawtooth lag; logs show Revoked partitions mid-batchCut max.poll.records to 50-200 to fit half the poll intervalSize batches from measured timings, not defaults
JVM GC pausesGC logs show 10s+ pauses; heartbeat thread starves while processing looks fastSize heap correctly, switch to G1GC/ZGC, fix allocation churnShip GC logging by default; alert on pauses over 5s
Rebalance stormsLogs show frequent Revoked/Assigned cycles on every deploy or restartStatic membership (group.instance.id) plus CooperativeStickyAssignorRolling deploys only; test a deploy causes zero full-group stops
⚙ Quick Reference
5 commands from this guide
FileCommand / CodePurpose
consumer.propertiesmax.poll.records=100max.poll.records Too High
consumer_poll_loop.pyfrom kafka import KafkaConsumerSlow Consumer Processing
gc_check.shexport KAFKA_HEAP_OPTS="-Xms2g -Xmx2g"GC Pauses That Stall the Poll Loop
membership_fix.shgrep -cE 'Revoked partitions|Attempt to heartbeat failed' /var/log/app/consumer....Rebalance Storms That Reset Your Progress
scale_check.shkafka-topics.sh --bootstrap-server $BROKERS --describe --topic recsScaling Out

Key takeaways

1
Idle brokers plus growing lag means the consumer is slow
never start by scaling the cluster.
2
Size max.poll.records from measured batch timings so the worst case fits in half the poll interval.
3
Per-record blocking I/O is the top throughput killer; batch, cache, or parallelize it.
4
GC pauses over 5 seconds starve heartbeats and cause revokes
ship GC logs by default.
5
Static membership plus cooperative rebalancing ends deploy-triggered rebalance storms.
6
You can't run more useful consumers than partitions
add partitions before adding pods.

Common mistakes to avoid

5 patterns
×

Leaving max.poll.records at the 500 default with slow per-record work

Symptom
Batches randomly exceed max.poll.interval.ms, the group rebalances, and lag jumps in sawtooth steps instead of draining steadily.
Fix
Cut max.poll.records to 50-200 so the slowest realistic batch finishes in under half of max.poll.interval.ms. Measure seconds-per-record on production-shaped data first, then set batch size = (interval * 0.5) / seconds-per-record.
×

Doing synchronous HTTP or DB calls per record inside the poll loop

Symptom
One 300 ms downstream call times 1,000 records equals a 5-minute batch. Throughput caps at a few records per second while every dashboard blames Kafka.
Fix
Move blocking I/O out of the poll loop: batch DB writes, cache lookups, and parallelize record handling with a bounded worker pool. Commit only after the whole batch is done.
×

Running consumers on default 1 GB heaps with CMS or no GC logging

Symptom
Mysterious 20-60 second stalls with no application log output, followed by a rebalance. GC logs later show full collections you never watched.
Fix
Size the heap with GC logs first (-Xlog:gc*), then set -Xmx/-Xms to the same value and switch to G1GC or ZGC. Alert on any single pause over 5 seconds on consumer JVMs.
×

Redeploying consumers without static membership

Symptom
Every deploy revokes all partitions, lag spikes to the previous high-water mark, and the team learns to fear releases instead of trusting them.
Fix
Set group.instance.id to a stable value (pod name or host ID), use CooperativeStickyAssignor, and keep deploys to rolling restarts. Verify with a deploy that causes zero group-wide stops.
×

Adding consumer instances when partitions are the real ceiling

Symptom
New consumers join the group but stay idle with no assigned partitions. Lag doesn't move and the extra pods just burn money.
Fix
Add partitions first (you can't use more consumers than partitions), then scale consumers to match. Monitor records-lag-max per partition so one hot partition can't hide behind a healthy average.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Lag climbs while broker CPU sits at 15%. What does that tell you?
Q02SENIOR
How do max.poll.records and max.poll.interval.ms interact?
Q03SENIOR
Walk me through your lag diagnosis order.
Q04SENIOR
Each record does a 200 ms DB lookup. How do you fix throughput without l...
Q05SENIOR
How does a GC pause turn into duplicate processing?
Q01 of 05JUNIOR

Lag climbs while broker CPU sits at 15%. What does that tell you?

ANSWER
LAG = log-end-offset minus committed offset per partition, summed across the group. Brokers look idle because produce/fetch load is light — the consumer simply polls and commits slower than producers append, so the gap widens on the client side.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
If brokers are idle, doesn't that mean Kafka is healthy?
02
Should I watch total lag or per-partition lag?
03
What's a good max.poll.records value?
04
When should I just add more consumers?
05
Lag keeps rising for a while after my fix — did it fail?
06
What does a dash (-) in the lag column mean?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Kafka. Mark it forged?

5 min read · try the examples if you haven't

1 / 4 · Kafka
Next
Kafka Leader Not Available After Broker Restart
→