Elasticsearch Circuit Breaking: Data Too Large Fix
Elasticsearch circuit_breaking_exception? Read the breaker name, drop fielddata on text, shrink pages, and size heap right..
20+ years shipping production backend systems. Written from production experience, not tutorials.
- ✓An Elasticsearch cluster with cat and nodes-stats API access
- ✓A slow log or query log showing heavy searches
- ✓Ability to update mappings and index settings
- The breaker name picks the fix: request means oversized queries, fielddata means text aggregation, in_flight means concurrency
- Aggregate on keyword sub-fields with doc_values instead of enabling fielddata on analyzed text ever
- Shrink pages to 1000 or less and split deep aggregations since one giant request trips where sliced ones pass
- Size heap at half of RAM under 30GB and alert on JVM pressure past 75% for days of warning
Think of the node as a kitchen with a strict fire code: each dish (query) gets a size estimate, and anything over the limit is refused at the door instead of burning everything down. fielddata on text is a banquet order filling the whole kitchen; giant pages are stacked trays blocking every walkway. The fix isn't bribing the inspector with higher limits — it's splitting orders, using the pantry (disk) instead of the counter (heap), and expanding the kitchen only after the menu is sane.
The dashboard that worked for a year starts failing at 10 AM sharp: circuit_breaking_exception, Data too large, rejected execution. Retry and it sometimes passes. Add data and it fails daily. The cluster is green, disk is fine, CPU is bored — yet queries die with 429s.
The breaker tripped because some request's estimated heap cost breached its budget. Usual suspects: fielddata uninverting analyzed text into gigabytes of heap, a size-10000 page stapled to deep aggregations, or a heap sized so generously there's no page cache left for Lucene. The error names the breaker — request, fielddata, in_flight — and that name is the diagnosis.
This guide reads the breaker name, finds the offending query in the slow log, and shrinks it with keyword mappings, smaller pages, and composite aggregations. You'll learn heap sizing that survives growth, which limits are safe to tune, and pressure alerts that warn days before users see 429s. The breaker name in your error message picks the section to read first, so start there and skip the rest for now.
Reading the Breaker Name Like a Diagnosis
Every trip names its breaker in brackets: request for query and aggregation heap, fielddata for uninverted text structures, in_flight_requests for concurrent request bytes, accounting for tracking overhead. The name is a routing slip — request trips go to query shape, fielddata trips go to mappings, in_flight trips go to concurrency and client fan-out.
Confirm with nodes stats. The breaker section shows estimated versus limit per breaker; the JVM section shows old-gen pressure and GC times. One breaker pinned near 100% with low overall pressure convicts a single query shape. All breakers high with 90% pressure convicts sizing — heap, caches, or node count.
Defaults encode hard-won ratios: request near 60% of heap, fielddata near 40-60%, total capping combined estimators. These aren't suggestions to outgrow — they're the guardrails that keep one bad dashboard from OOMing shared nodes. Tune only with measured headroom, never to silence an unexamined query.
The 429 response carries machine-readable guidance most clients ignore. Retry-After hints when to try again; the error body names the breaker and the byte counts on both sides of the limit. Well-behaved clients back off with jitter on 429 instead of hammering — retry storms convert one trip into a cascade that 429s innocent queries. Track tripped counts per breaker from nodes stats to separate chronic pressure (rising estimates) from single-query spikes (one estimator pinned). Dashboards that graph estimates against limits show tomorrow's trips today. Read the numbers the breaker already computed for you.
fielddata on Text: the 11 GB Foot-Gun
Analyzed text fields can't aggregate directly — their tokens live in the inverted index, built for search, not grouping. Enabling fielddata uninverts those tokens into heap structures at query time: gigabytes, unbounded, per segment, reloaded on every eviction. It's the single most expensive mapping choice in Elasticsearch.
The fix is architectural: aggregate on keyword sub-fields backed by doc_values, disk-based columnar structures built at index time with tiny heap footprints. Most indices already have a .keyword multi-field — point terms aggs and sorts there for instant relief. Where it's missing, add the multi-field and reindex; the one-time cost buys permanent headroom.
Guard the mapping forever. Lint index templates in CI to reject fielddata=true on text, review every new sortable or aggregatable field for keyword backing, and alert on fielddata memory crossing 10% of heap. Mappings are capacity decisions wearing schema clothes.
Keyword mappings deserve deliberate design, not defaults. Add normalizers for case-insensitive terms aggs, split fields for both analyzed search and exact aggregation, and ignore_above to cap monster tokens before they bloat ordinals. Numeric fields want the narrowest type that fits: scaled_float for prices, half_float for metrics, unsigned_long for counters — narrower types mean smaller doc_values and faster aggs. Date fields need explicit formats so mixed producer shapes stop guessing games. Every mapping choice is a capacity choice; review new fields with the same gravity as new queries.
Oversized Requests: Pages, Buckets, and Deep Aggs
The request breaker estimates aggregations, bucket counts, and page payloads together — and size-10000 pages with nested aggs blow the budget alone. Each hit loads _source, each bucket holds ordinals and counts, and deep levels multiply: 5 nested terms aggs at size 1000 can materialize millions of buckets before returning ten rows.
Shrink in three moves. Cap size at 100-1000 and paginate deep reads with search_after instead of from-offsets. Split multi-aggregation monsters across sequential requests the client merges. Convert high-cardinality terms aggs to composite pagination or sampler-scoped variants that bound bucket construction.
Find the offenders in the slow log, not by guessing. Trip timestamps plus slow-log entries with huge took values and sizes name the exact dashboard and query. Fix the client code behind it — breaker trips from page sizes are client bugs, and the client is where the permanent fix ships.
Composite aggregations turn unbounded terms into pageable streams. Instead of one giant terms agg, composite pages buckets with after_key cursors the client follows — memory stays flat while coverage stays complete. Sampler and diversified_sampler aggregations scope expensive aggs to representative subsets for exploratory dashboards that do not need exact counts. The cardinality agg itself estimates distinct counts cheaply for sizing decisions. For exports, scroll slices parallelize across shards while search_after keeps ordering stable. Big answers come from small requests composed well.
Heap Sizing: Half for JVM, Half for Lucene
Heap sizing follows one rule: 50% of RAM, capped near 26-30 GB so compressed ordinary object pointers stay enabled. The other half serves Lucene's page cache — the OS-level cache queries need for segment reads. All-heap nodes starve the page cache, forcing disk reads that inflate heap caches in turn: slower queries tripping sooner.
Crossing the ~30 GB cap silently disables compressed pointers, inflating every object and effectively shrinking usable heap. A 40 GB heap can hold less than a 28 GB one while GC runs longer. The cap isn't modesty — it's physics.
Bound the caches explicitly. indices.fielddata.cache.size caps the most dangerous cache; query and request caches get modest percentages reviewed against hit rates. Then watch JVM pressure (old-gen used ratio) as the master gauge: sustained 75% warns, 85% pages, 95% means trips are already firing.
JVM internals explain why balanced nodes win. Young-gen collections should finish in milliseconds; sustained old-gen growth between full GCs signals leaks or caches without bounds. Nodes stats expose pool usage and collection times — graph both, and alert when old-gen exceeds 75 percent after major collections. G1GC (the modern default) favors predictable pauses over maximum throughput, matching search workloads. Heap dumps on OOM are postmortems, not monitoring; pressure trends are the monitoring. Tune GC flags rarely, size heaps correctly always.
Concurrency and the in_flight Avalanche
Concurrency trips the in_flight breaker when too many heavy requests overlap: bulk exports at 9 AM, dashboard refresh storms, retry avalanches after one timeout. Each request is legal alone; together they breach the combined budget and the node starts 429ing innocents.
Smooth the overlap. Stagger scheduled exports, lengthen dashboard refresh intervals, cap client concurrency with backoff and jitter, and bulkhead export traffic onto coordinating-only nodes or off-peak windows. Queue rejections (429) are the signal to slow down, not retry harder — retry storms convert one trip into a cascade.
Scale only after smoothing. More data nodes spread shards and heap; dedicated coordinating nodes absorb aggregation fan-in. But nodes added under unexamined queries just host bigger trips — shrink first, smooth second, scale third. That order keeps every added node productive instead of flammable.
Routing and cancellation multiply effective capacity. Coordinating-only nodes absorb aggregation fan-in so data nodes keep serving shards; dedicated master nodes stay out of query paths entirely. The task management API lists running searches with runtimes — cancel the runaway export instead of watching it trip breakers for everyone. Search thread-pool rejections in _cat/thread_pool signal concurrency saturation before breakers fire. Bulkhead export workloads onto off-peak windows or separate clusters. Capacity is topology plus discipline, not just heap.
Guardrails That Outlive the Incident
Close the loop with guardrails that outlive the incident. Mapping CI rejects fielddata on text; client review caps page sizes and bucket counts; dashboards default to sane refresh intervals. Each rule is a one-line check that prevents a recurrence class permanently.
Alert on pressure, not just trips. JVM pressure above 75% sustained warns days ahead; breaker trip counters page immediately; p99 latency trends confirm user impact. The three together separate warning (act Tuesday) from emergency (act now).
Verify for a full business cycle before closing: a week with zero trips, pressure under 70%, and the previously failing dashboard green at peak. Breaker incidents recur on schedules — weekly exports, monthly reports — and only a full cycle proves the fix.
Index maintenance pays steady dividends against breaker pressure. Force-merge read-only indices to one segment so aggregations scan less; raise refresh_interval on write-heavy indices to cut segment churn; drop _source on metrics-like indices where stored fields suffice. Rollover on size or age keeps shards uniformly small so no single shard dominates request heap. Schedule maintenance in low-traffic windows and watch pressure graphs flatten over days. Small shards, few segments, calm breakers — the virtuous cycle compounds.
The Dashboard That Ate 11 GB of Heap Every Morning
- Raised limits convert loud rejections into silent GC death — the breaker was the messenger, not the problem.
- One dashboard with fielddata on text can hold a whole cluster's heap hostage; mapping review is capacity work.
- Slow-log evidence beats limit knobs: the offending query was visible for months before anyone looked.
| File | Command / Code | Purpose |
|---|---|---|
| GET /_nodes/stats/breaker | Reading the Breaker Name Like a Diagnosis | |
| POST /marketing/_search | fielddata on Text | |
| POST /orders/_search | Oversized Requests | |
| PUT /_cluster/settings | Heap Sizing |
Key takeaways
Common mistakes to avoid
5 patternsEnabling fielddata=true on text to make a dashboard work
Raising breaker limits instead of shrinking queries
Giving Elasticsearch all server RAM as heap
Paginating with size 10000 plus deep aggregations
Watching only trip errors while ignoring memory pressure
Interview Questions on This Topic
What are circuit breakers protecting?
Frequently Asked Questions
20+ years shipping production backend systems. Written from production experience, not tutorials.
That's Elasticsearch. Mark it forged?
5 min read · try the examples if you haven't