Kafka RecordTooLargeException — Fix the Size Chain
RecordTooLarge means one of four size limits bit.
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
- ✓A Kafka topic you can produce to
- ✓Basic producer and broker config familiarity
- ✓Sample payloads near your size limits
- RecordTooLarge means your payload beat one of four size limits — find which link before touching any knob
- The chain: broker max.message.bytes, topic message.max.bytes, replica.fetch.max.bytes, consumer fetch.max.bytes
- Don't forget producer max.request.size: legal-size records can still fail as an oversized batch
- Durable fix is usually chunking or claim-check (blob in storage, reference in Kafka), not bigger knobs
Imagine mailing a package through four checkpoints: post office counter, sorting machine, delivery truck door, and your mailbox slot. Widening just the counter changes nothing if the truck door stays narrow — the package gets stuck at the next stop. Kafka's four size settings work the same way: every checkpoint must fit the package. And the smarter move is usually splitting the shipment into small boxes (chunking) or mailing a pickup slip instead of the piano itself (claim-check).
Your producer throws RecordTooLargeException on a Tuesday afternoon. Someone suggests raising max.message.bytes, so you do — and the error moves to the replicas. You raise that too, and now the consumers choke. Three config changes later the original 12 MB payload flows, and so does a quiet tax on every other topic on the cluster.
This is the size-chain trap: Kafka enforces message size in four places, and widening one link just pushes the failure to the next. Worse, each increase has a real cost — broker memory, replication bandwidth, consumer heap — paid by every topic, not just yours.
This beginner-friendly guide maps the whole chain: max.message.bytes, message.max.bytes, replica.fetch.max.bytes, and consumer fetch.max.bytes, what each guards, and how to check all four in minutes. More importantly, it shows why the durable fix is usually a design change — chunking or claim-check references — instead of bigger knobs. You'll leave knowing exactly when raising a limit is fine and when it's a loan your cluster will collect.
What RecordTooLargeException Tells You
RecordTooLargeException is the producer client telling you the broker refused your batch: at least one record (or the batch envelope) exceeded a configured limit. The broker answers with MESSAGE_TOO_LARGE, the client surfaces it as RecordTooLargeException, and the send fails — usually retrying pointlessly, since retries don't shrink payloads. It's a config-and-design error, not a transient one.
Beginners misread it as a single knob because the first search result says 'raise max.message.bytes.' That advice is a quarter of the story. Kafka checks size at four gates: the broker append limit, the topic-level override, the replica fetch limit followers use, and the consumer fetch limit your readers use. Clearing gate one with a 12 MB payload just schedules a failure at gate two or three — each louder and weirder than the last.
The right reflex is measurement before mutation. Log the serialized byte size of the failing payload, read all four caps in one pass, and identify every link smaller than your payload. Then — and only then — decide whether this payload deserves bigger gates or a smaller shape. Most oversize payloads are accidents (full PDFs, raw images, unbounded JSON arrays) that no commit log should carry at any limit.
The Four-Knob Size Chain, End to End
The chain has four links, and the narrowest wins. Broker max.message.bytes (default ~1 MB) caps single records at append time. Topic-level message.max.bytes overrides it per topic — same meaning, narrower scope. Replica replica.fetch.max.bytes caps what followers may fetch for replication; smaller than your records means followers starve and under-replicated partitions grow. Consumer fetch.max.bytes and max.partition.fetch.bytes cap what readers pull; smaller than your records wedges consumers at one offset forever.
Two producer-side settings complete the picture. max.request.size (default 1 MB) caps the entire produce request — batch envelope, not single record. With batch.size at 16 KB and linger.ms batching aggressively, fifty modest records can breach it together. buffer.memory and compression.type (lz4 or zstd) shape how often you hit the ceiling: compression shrinks JSON payloads 4-10x and is the cheapest headroom you'll ever buy.
Audit the chain as one unit with the script above. Write down all six numbers (four gates plus batch/request caps) next to your largest real payload before changing anything. If any gate is smaller than the payload, that gate is your current error — and every other smaller-or-equal gate is your next error. Change them as a reviewed set or don't change them at all.
Broker vs Topic Config: Where to Set What
Broker defaults apply to every topic, so raising them globally taxes the whole cluster for one producer's appetite. Topic-level overrides are the containment vessel: kafka-configs.sh --alter sets max.message.bytes on invoices-pdf alone, leaving the other 200 topics at 1 MB. When the PDF team later migrates to claim-check, you revert one topic instead of renegotiating cluster-wide policy.
Pair the topic override with a dedicated producer config. Big-payload producers want small batches (32 KB keeps the request envelope predictable), short linger.ms (don't accumulate a mountain), lz4 compression (halves most document payloads for near-zero CPU), and acks=all (large records deserve full durability since each one hurts more to lose). Don't share this config with your clickstream producer — their tuning fights yours.
Document the exception where everyone can find it: topic name, approved max size, owning team, expiry date for review. Size exceptions without owners become permanent, and permanent exceptions become the new default when the next team copies them. A registry entry with a 90-day review turns 'temporary 8 MB allowance' into a migration deadline instead of a tradition. Review every exception quarterly and retire the ones whose payloads migrated to claim-check — exceptions should shrink, never grow.
Why Raising Limits Is Usually the Wrong Fix
Raising limits feels like a fix and behaves like a loan. A 12 MB record at replication factor 3 pushes ~36 MB across brokers before any consumer reads it. Fetches for that partition stall behind the giant, page cache fills with bytes read once, and consumers need heap headroom for the largest record plus its deserialized form — a 12 MB JSON blob easily becomes 60 MB of objects. Multiply by partition count and the cluster carries your payload everywhere.
Growth finishes the argument. Payloads that hit limits once almost always grow: 2 MB exports become 12 MB PDFs become 40 MB bundles, because nothing upstream constrains them. Each raise buys weeks while the tax compounds permanently. Teams that raise quarterly end up running a blob store with commit-log prices — slow brokers, giant disks, fragile consumers.
Apply the 1 MB rule of thumb: payloads under ~1 MB with stable schemas are fine with tuned knobs. Anything bigger or still growing gets chunking (ordered, bounded pieces) or claim-check (blob outside, reference inside). The afternoon you spend on plumbing pays back the first time upstream doubles payload size and your pipeline doesn't notice. Show the growth chart at review: payloads that doubled twice will double again, and each raise mortgages every topic's latency.
Chunking Large Payloads Without Breaking Gates
Chunking splits one big payload into ordered pieces that each fit the gates, with metadata for reassembly. Key every chunk identically so they land on one partition in order; attach headers (blob id, sequence, total count, checksum) so consumers can detect gaps. Keep chunks comfortably under the smallest gate — 900 KB against 1 MB caps leaves headroom for headers and envelope overhead.
Consumers reassemble with a small state machine: buffer chunks per blob id, and when sequence count reaches total, concatenate, verify the SHA-256, and process. Add a timeout for incomplete blobs (a missing chunk shouldn't pin memory forever) and a max-buffer cap per blob so a corrupt manifest can't OOM the consumer. Idempotent replays stay safe because reassembly is deterministic — same chunks, same bytes, same checksum.
Chunking fits when you must keep bytes inside Kafka: strict ordering, log-compaction semantics, or no external store available. It's honest plumbing with real edge cases (partial blobs, version skew on header format), so prefer claim-check when an object store exists. But against 12 MB PDFs on a 1 MB topic, 14 chunks of 900 KB beats any knob raise — bounded, reviewable, and invisible to other topics.
Storing References Instead of Blobs: Claim-Check
Claim-check stores the bulky bytes outside Kafka and sends a small reference event through the log. The producer uploads the PDF to object storage, then publishes an event with bucket, key, byte size, and SHA-256 checksum — typically under 1 KB. Consumers download the blob, verify the checksum, and process. Ordering, partitioning, and replay all keep working because the log still carries one event per PDF in sequence.
This pattern wins on every axis that matters. Broker traffic drops 1000x per message, replication stays cheap, consumers need no special heap, and retention policies apply to tiny events while lifecycle rules expire the blobs independently. Checksums make corruption detectable instead of silent, and the object store's versioning gives you an audit trail the commit log was never designed to provide.
Handle the two edge cases up front: blob lifecycle (7-day expiry with a tombstone event for late consumers) and missing blobs (dead-letter with the reference intact so ops can re-upload). With those covered, claim-check scales from 12 MB PDFs to 500 MB videos without another Kafka config change — which is precisely why it's the default answer for growing payloads. Start new large-payload topics on claim-check by default so the easy path is also the scalable one.
The 12 MB PDF That Moved Through Three Knobs and Killed 24 Pods
- The four knobs are one atomic change or none at all — raising max.message.bytes alone just relocates the failure to replicas, then consumers.
- Payloads that grow (2 MB to 12 MB in 6 weeks) will outrun every limit; claim-check ends the ratchet permanently.
- Validate size at the producer and auto-route oversize payloads — humans shouldn't hand-tune broker configs per PDF.
| File | Command / Code | Purpose |
|---|---|---|
| size_chain_audit.sh | grep -E 'max.message.bytes|message.max.bytes|replica.fetch.max.bytes' /etc/kafka... | The Four-Knob Size Chain, End to End |
| topic_override.sh | bin/kafka-configs.sh --bootstrap-server $BROKERS \ | Broker vs Topic Config |
| chunked_producer.py | from kafka import KafkaProducer | Chunking Large Payloads Without Breaking Gates |
| claim_check.py | from kafka import KafkaProducer, KafkaConsumer | Storing References Instead of Blobs |
Key takeaways
Common mistakes to avoid
5 patternsRaising max.message.bytes alone and calling it fixed
Forgetting max.request.size on the producer
Letting any producer send unbounded payloads
Raising limits to 50 MB instead of fixing the payload
Fixing produce but never testing the consumer fetch path
Interview Questions on This Topic
What does RecordTooLargeException mean?
Frequently Asked Questions
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
That's Kafka. Mark it forged?
5 min read · try the examples if you haven't