Home › Observability › Elasticsearch Circuit Breaking: Data Too Large Fix
Advanced 5 min · September 23, 2026

Elasticsearch Circuit Breaking: Data Too Large Fix

Elasticsearch circuit_breaking_exception? Read the breaker name, drop fielddata on text, shrink pages, and size heap right..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 14 min
  • ✓An Elasticsearch cluster with cat and nodes-stats API access
  • ✓A slow log or query log showing heavy searches
  • ✓Ability to update mappings and index settings
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • The breaker name picks the fix: request means oversized queries, fielddata means text aggregation, in_flight means concurrency
  • Aggregate on keyword sub-fields with doc_values instead of enabling fielddata on analyzed text ever
  • Shrink pages to 1000 or less and split deep aggregations since one giant request trips where sliced ones pass
  • Size heap at half of RAM under 30GB and alert on JVM pressure past 75% for days of warning
✦ Definition~90s read
What is Elasticsearch Circuit Breaking Exception?

Elasticsearch circuit breakers are pre-execution heap budgets that reject requests before they can crash the JVM. Before running a query, the node estimates its heap cost — aggregation buckets, fielddata structures, page payloads, in-flight bytes — and compares against per-breaker limits: request near 60% of heap, fielddata near 40-60%, with a parent total capping combined estimators.

★
Think of the node as a kitchen with a strict fire code: each dish (query) gets a size estimate, and anything over the limit is refused at the door instead of burning everything down.

Estimates over budget fail fast with circuit_breaking_exception and HTTP 429 instead of OutOfMemory death.

The two memory systems behind most trips work very differently. fielddata uninverts analyzed text tokens into heap at query time — gigabytes, unbounded, the most expensive structure in the system. doc_values are disk-based columnar stores built at index time for keyword, numeric, date, and boolean fields, aggregating with tiny heap footprints. Queries against keyword sub-fields ride doc_values; queries against raw text demand fielddata.

Request sizing completes the picture. Page size, aggregation depth, and bucket counts multiply into the request estimate, so one size-20000 export with nested aggs breaches budgets that a thousand sliced requests sail under. Heap sizing (50% of RAM, capped near 30GB for compressed pointers) and page-cache room decide how much budget exists at all.

Fix the query shape first, tune limits only with measured headroom, and the breakers return to their proper job: silent insurance that never fires.

Plain-English First

Think of the node as a kitchen with a strict fire code: each dish (query) gets a size estimate, and anything over the limit is refused at the door instead of burning everything down. fielddata on text is a banquet order filling the whole kitchen; giant pages are stacked trays blocking every walkway. The fix isn't bribing the inspector with higher limits — it's splitting orders, using the pantry (disk) instead of the counter (heap), and expanding the kitchen only after the menu is sane.

The dashboard that worked for a year starts failing at 10 AM sharp: circuit_breaking_exception, Data too large, rejected execution. Retry and it sometimes passes. Add data and it fails daily. The cluster is green, disk is fine, CPU is bored — yet queries die with 429s.

The breaker tripped because some request's estimated heap cost breached its budget. Usual suspects: fielddata uninverting analyzed text into gigabytes of heap, a size-10000 page stapled to deep aggregations, or a heap sized so generously there's no page cache left for Lucene. The error names the breaker — request, fielddata, in_flight — and that name is the diagnosis.

This guide reads the breaker name, finds the offending query in the slow log, and shrinks it with keyword mappings, smaller pages, and composite aggregations. You'll learn heap sizing that survives growth, which limits are safe to tune, and pressure alerts that warn days before users see 429s. The breaker name in your error message picks the section to read first, so start there and skip the rest for now.

Reading the Breaker Name Like a Diagnosis

Every trip names its breaker in brackets: request for query and aggregation heap, fielddata for uninverted text structures, in_flight_requests for concurrent request bytes, accounting for tracking overhead. The name is a routing slip — request trips go to query shape, fielddata trips go to mappings, in_flight trips go to concurrency and client fan-out.

Confirm with nodes stats. The breaker section shows estimated versus limit per breaker; the JVM section shows old-gen pressure and GC times. One breaker pinned near 100% with low overall pressure convicts a single query shape. All breakers high with 90% pressure convicts sizing — heap, caches, or node count.

Defaults encode hard-won ratios: request near 60% of heap, fielddata near 40-60%, total capping combined estimators. These aren't suggestions to outgrow — they're the guardrails that keep one bad dashboard from OOMing shared nodes. Tune only with measured headroom, never to silence an unexamined query.

The 429 response carries machine-readable guidance most clients ignore. Retry-After hints when to try again; the error body names the breaker and the byte counts on both sides of the limit. Well-behaved clients back off with jitter on 429 instead of hammering — retry storms convert one trip into a cascade that 429s innocent queries. Track tripped counts per breaker from nodes stats to separate chronic pressure (rising estimates) from single-query spikes (one estimator pinned). Dashboards that graph estimates against limits show tomorrow's trips today. Read the numbers the breaker already computed for you.

JSON
1
2
3
4
5
6
7
8
# Which breaker trips, and how close each runs to its limit
GET /_nodes/stats/breaker
# breaker.request.estimated_size_in_bytes vs .limit_size_in_bytes
# breaker.fielddata.estimated_size_in_bytes vs .limit_size_in_bytes

# JVM pressure — sustained >75% warns days before trips
GET /_nodes/stats/jvm
# nodes.*.jvm.mem.pools.old.used_in_bytes / .max_in_bytes
📊 Production Insight
A team raised every limit chasing unnamed trips for a month. The first engineer to read the brackets found fielddata — a mapping fix ended it that afternoon. The name had been in every error all along.
🎯 Key Takeaway
Breaker name plus nodes-stats estimates route every trip to mappings, queries, or sizing.

fielddata on Text: the 11 GB Foot-Gun

Analyzed text fields can't aggregate directly — their tokens live in the inverted index, built for search, not grouping. Enabling fielddata uninverts those tokens into heap structures at query time: gigabytes, unbounded, per segment, reloaded on every eviction. It's the single most expensive mapping choice in Elasticsearch.

The fix is architectural: aggregate on keyword sub-fields backed by doc_values, disk-based columnar structures built at index time with tiny heap footprints. Most indices already have a .keyword multi-field — point terms aggs and sorts there for instant relief. Where it's missing, add the multi-field and reindex; the one-time cost buys permanent headroom.

Guard the mapping forever. Lint index templates in CI to reject fielddata=true on text, review every new sortable or aggregatable field for keyword backing, and alert on fielddata memory crossing 10% of heap. Mappings are capacity decisions wearing schema clothes.

Keyword mappings deserve deliberate design, not defaults. Add normalizers for case-insensitive terms aggs, split fields for both analyzed search and exact aggregation, and ignore_above to cap monster tokens before they bloat ordinals. Numeric fields want the narrowest type that fits: scaled_float for prices, half_float for metrics, unsigned_long for counters — narrower types mean smaller doc_values and faster aggs. Date fields need explicit formats so mixed producer shapes stop guessing games. Every mapping choice is a capacity choice; review new fields with the same gravity as new queries.

JSON
1
2
3
4
5
6
7
8
9
# Relief now: aggregate the keyword sub-field (doc_values backed)
POST /marketing/_search
{ "size": 0, "aggs": { "by_campaign": { "terms": { "field": "campaign.keyword", "size": 100 } } } }

# Permanent: add keyword multi-field, then reindex
PUT /marketing_v2
{ "mappings": { "properties": {
  "campaign": { "type": "text", "fields": { "keyword": { "type": "keyword" } } } } } }
POST /_reindex { "source": { "index": "marketing" }, "dest": { "index": "marketing_v2" } }
📊 Production Insight
One marketing dashboard sorting analyzed text loaded 11 GB into a 16 GB heap every morning. Reindexing with a keyword multi-field cut that heap to 40 MB — a 275x reduction from a mapping change.
🎯 Key Takeaway
Aggregate keyword sub-fields on doc_values; never enable fielddata on text.

Oversized Requests: Pages, Buckets, and Deep Aggs

The request breaker estimates aggregations, bucket counts, and page payloads together — and size-10000 pages with nested aggs blow the budget alone. Each hit loads _source, each bucket holds ordinals and counts, and deep levels multiply: 5 nested terms aggs at size 1000 can materialize millions of buckets before returning ten rows.

Shrink in three moves. Cap size at 100-1000 and paginate deep reads with search_after instead of from-offsets. Split multi-aggregation monsters across sequential requests the client merges. Convert high-cardinality terms aggs to composite pagination or sampler-scoped variants that bound bucket construction.

Find the offenders in the slow log, not by guessing. Trip timestamps plus slow-log entries with huge took values and sizes name the exact dashboard and query. Fix the client code behind it — breaker trips from page sizes are client bugs, and the client is where the permanent fix ships.

Composite aggregations turn unbounded terms into pageable streams. Instead of one giant terms agg, composite pages buckets with after_key cursors the client follows — memory stays flat while coverage stays complete. Sampler and diversified_sampler aggregations scope expensive aggs to representative subsets for exploratory dashboards that do not need exact counts. The cardinality agg itself estimates distinct counts cheaply for sizing decisions. For exports, scroll slices parallelize across shards while search_after keeps ordering stable. Big answers come from small requests composed well.

JSON
1
2
3
4
5
6
7
8
9
10
11
# Before: one giant request that trips the breaker
# { "size": 20000, "aggs": { "l1": { "terms": { "field": "a", "aggs": { "l2": ... } } } } }

# After: sliced reads the client stitches together
POST /orders/_search
{
  "size": 1000,
  "sort": [{ "_shard_doc": "asc" }],
  "search_after": ["042", "orders", 48291],
  "aggs": { "by_status": { "terms": { "field": "status.keyword", "size": 20 } } }
}
📊 Production Insight
Size-20000 CSV exports with 5-level aggs tripped the request breaker hourly. Slicing to 1000-row search_after pages kept every export working while peak request heap fell 80%.
🎯 Key Takeaway
Small pages, split aggs, search_after — sliced requests pass where giants trip.

Heap Sizing: Half for JVM, Half for Lucene

Heap sizing follows one rule: 50% of RAM, capped near 26-30 GB so compressed ordinary object pointers stay enabled. The other half serves Lucene's page cache — the OS-level cache queries need for segment reads. All-heap nodes starve the page cache, forcing disk reads that inflate heap caches in turn: slower queries tripping sooner.

Crossing the ~30 GB cap silently disables compressed pointers, inflating every object and effectively shrinking usable heap. A 40 GB heap can hold less than a 28 GB one while GC runs longer. The cap isn't modesty — it's physics.

Bound the caches explicitly. indices.fielddata.cache.size caps the most dangerous cache; query and request caches get modest percentages reviewed against hit rates. Then watch JVM pressure (old-gen used ratio) as the master gauge: sustained 75% warns, 85% pages, 95% means trips are already firing.

JVM internals explain why balanced nodes win. Young-gen collections should finish in milliseconds; sustained old-gen growth between full GCs signals leaks or caches without bounds. Nodes stats expose pool usage and collection times — graph both, and alert when old-gen exceeds 75 percent after major collections. G1GC (the modern default) favors predictable pauses over maximum throughput, matching search workloads. Heap dumps on OOM are postmortems, not monitoring; pressure trends are the monitoring. Tune GC flags rarely, size heaps correctly always.

JSON
1
2
3
4
5
6
7
# Cap the most dangerous cache via cluster settings
PUT /_cluster/settings
{ "persistent": { "indices.fielddata.cache.size": "20%" } }

# jvm.options per node: 50% of 64 GB RAM, under the ~30 GB cap
-Xms31g
-Xmx31g
📊 Production Insight
A 64 GB node with 60 GB heap GC-paused for 8 seconds while tripping constantly. Dropping heap to 30 GB freed the page cache, cut p99 from 12s to 400ms, and ended trips — less heap, more performance.
🎯 Key Takeaway
50% RAM under ~30GB for compressed pointers; the rest feeds the page cache queries live on.

Concurrency and the in_flight Avalanche

Concurrency trips the in_flight breaker when too many heavy requests overlap: bulk exports at 9 AM, dashboard refresh storms, retry avalanches after one timeout. Each request is legal alone; together they breach the combined budget and the node starts 429ing innocents.

Smooth the overlap. Stagger scheduled exports, lengthen dashboard refresh intervals, cap client concurrency with backoff and jitter, and bulkhead export traffic onto coordinating-only nodes or off-peak windows. Queue rejections (429) are the signal to slow down, not retry harder — retry storms convert one trip into a cascade.

Scale only after smoothing. More data nodes spread shards and heap; dedicated coordinating nodes absorb aggregation fan-in. But nodes added under unexamined queries just host bigger trips — shrink first, smooth second, scale third. That order keeps every added node productive instead of flammable.

Routing and cancellation multiply effective capacity. Coordinating-only nodes absorb aggregation fan-in so data nodes keep serving shards; dedicated master nodes stay out of query paths entirely. The task management API lists running searches with runtimes — cancel the runaway export instead of watching it trip breakers for everyone. Search thread-pool rejections in _cat/thread_pool signal concurrency saturation before breakers fire. Bulkhead export workloads onto off-peak windows or separate clusters. Capacity is topology plus discipline, not just heap.

📊 Production Insight
Nine-AM export storms overlapped dashboard refreshes and 429'd the whole floor. Staggering exports by 20 minutes and halving refresh rates cleared in_flight trips permanently — zero new hardware.
🎯 Key Takeaway
Stagger heavies, back off on 429, bulkhead exports — scale only after queries are sane.

Guardrails That Outlive the Incident

Close the loop with guardrails that outlive the incident. Mapping CI rejects fielddata on text; client review caps page sizes and bucket counts; dashboards default to sane refresh intervals. Each rule is a one-line check that prevents a recurrence class permanently.

Alert on pressure, not just trips. JVM pressure above 75% sustained warns days ahead; breaker trip counters page immediately; p99 latency trends confirm user impact. The three together separate warning (act Tuesday) from emergency (act now).

Verify for a full business cycle before closing: a week with zero trips, pressure under 70%, and the previously failing dashboard green at peak. Breaker incidents recur on schedules — weekly exports, monthly reports — and only a full cycle proves the fix.

Index maintenance pays steady dividends against breaker pressure. Force-merge read-only indices to one segment so aggregations scan less; raise refresh_interval on write-heavy indices to cut segment churn; drop _source on metrics-like indices where stored fields suffice. Rollover on size or age keeps shards uniformly small so no single shard dominates request heap. Schedule maintenance in low-traffic windows and watch pressure graphs flatten over days. Small shards, few segments, calm breakers — the virtuous cycle compounds.

⚠ Limits Are Guardrails, Not Quotas
Raising breaker limits without shrinking the guilty queries converts loud, survivable 429s into silent GC death. The node stops refusing work it can't afford and starts freezing for seconds per GC cycle — slower for everyone, tripping later at higher heap, and far harder to diagnose.
📊 Production Insight
After mapping lint plus pressure alerts, one fleet caught the next fielddata attempt in CI and the next pressure rise 9 days before trips. Both fixes shipped as calm tickets — the 3-week morning outage never repeated.
🎯 Key Takeaway
Lint mappings, cap pages, alert pressure at 75% — verify a full cycle before closing.
● Production incidentPOST-MORTEMseverity: high

The Dashboard That Ate 11 GB of Heap Every Morning

Symptom
Marketing dashboards failed with fielddata breaker trips every morning for 3 weeks, then all queries degraded to 12s p99s after the limit raise. Two nodes showed 95% JVM pressure with 8-second GC pauses during business hours.
Assumption
The team assumed data growth demanded bigger breakers and raised request limit from 60% to 90% plus heap from 16 to 28 GB. Trips paused 6 days, then returned with multi-second GC pauses and 12-second query p99s across all indices.
Root cause
A text field with fielddata=true loaded 11 GB of uninverted terms into a 16 GB heap on every marketing-dashboard load, while CSV exports pulled size-20000 pages with nested aggregations through the request breaker. Raising limits let both allocate deeper until old-GC pauses froze the node for seconds per cycle.
Fix
They found a marketing dashboard sorting on a text field with fielddata enabled plus size-20000 exports with 5-level aggregations. They reindexed with keyword multi-fields, rewrote exports to search_after slices of 1000, and restored breaker defaults with heap at 50% of 64 GB. Trips hit zero the same day; p99 fell from 12s to 400ms.
Key lesson
  • Raised limits convert loud rejections into silent GC death — the breaker was the messenger, not the problem.
  • One dashboard with fielddata on text can hold a whole cluster's heap hostage; mapping review is capacity work.
  • Slow-log evidence beats limit knobs: the offending query was visible for months before anyone looked.
Production debug guideFive steps from 429s to headroom without raising limits blindly.5 entries
Symptom · 01
Queries failing with circuit_breaking_exception
→
Fix
Copy the exact breaker name from the error: [request], [fielddata], or [in_flight_requests]. Request means oversized queries or aggs; fielddata means analyzed-text aggregation; in_flight means too many concurrent heavies. The name picks your section below — don't tune limits until it names one.
Symptom · 02
Breaker named but offending query unknown
→
Fix
Run GET /_nodes/stats/breaker and GET /_nodes/stats/jvm. Compare estimated_size against limit for the named breaker and read mem pressure. Pressure above 85% with modest queries means heap sizing or cache bounds are wrong; one breaker near 100% with low pressure means a single query shape is guilty.
Symptom · 03
fielddata breaker named
→
Fix
Check the mapping of the aggregated field: GET /index/_mapping and look for type text without a keyword sub-field behind the tripping agg. If found, rewrite the agg onto the keyword sub-field immediately for relief, then schedule a reindex adding keyword multi-fields. Never set fielddata=true on text to stop the pain.
Symptom · 04
Request breaker named
→
Fix
Pull the slow log around trip times for size, aggs, and bucket counts. Cap size at 100-1000, convert deep pagination to search_after, split multi-agg monsters across requests, and convert high-cardinality terms aggs to composite or sampler variants. Re-run the exact failing dashboard to confirm green.
Symptom · 05
Trips stopped; keep them stopped
→
Fix
Set heap to 50% RAM capped near 30GB, bound fielddata cache, alert JVM pressure above 75% sustained, and add mapping-lint plus page-size caps to code review. Verify a week of zero trips with pressure under 70% before closing.
Breaker-Trip Causes Compared
Root CauseHow to ConfirmFixPrevention
fielddata on analyzed textTrips name fielddata breaker; mapping shows textAggregate on keyword sub-field; reindexLint mappings; never enable fielddata on text
Oversized request (aggs + pages)Trips name request breaker; size 10000 in slow logShrink pages; split aggs; search_afterCap client page sizes; review dashboard aggs
Undersized heap / no page cacheHigh GC + pressure with modest queriesHeap 50% RAM under 30GB; bound cachesCapacity review on data growth; pressure alerts
Cardinality-heavy global aggsTerms agg on high-cardinality keyword tripsPartition terms, composite aggs, samplerPre-aggregate; cap bucket counts in reviews
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
GET /_nodes/stats/breakerReading the Breaker Name Like a Diagnosis
POST /marketing/_searchfielddata on Text
POST /orders/_searchOversized Requests
PUT /_cluster/settingsHeap Sizing

Key takeaways

1
The breaker name in the error is the diagnosis
request, fielddata, or in_flight each point at a fix.
2
Never enable fielddata on text
aggregate on keyword sub-fields backed by doc_values.
3
Shrink pages and split aggregations before touching breaker limits.
4
Size heap at 50% of RAM under ~30GB and leave page cache for Lucene.
5
Find offenders in the slow log; cap client page sizes in code review.
6
Alert on JVM pressure above 75% for days of warning before trips.

Common mistakes to avoid

5 patterns
×

Enabling fielddata=true on text to make a dashboard work

Symptom
One aggregation loads gigabytes of uninverted terms into heap and trips the breaker for the whole node.
Fix
Keep fielddata disabled on text and aggregate on keyword sub-fields. If sort or terms on analyzed text is truly needed, add a keyword multi-field and reindex — one mapping change beats every breaker tuning.
×

Raising breaker limits instead of shrinking queries

Symptom
Trips stop for a week, then return at higher heap with longer GC pauses and slower everything.
Fix
Raise indices.breaker.request.limit modestly (toward 40%) only with measured headroom, and pair it with smaller pages and scroll slices. The limit is a guardrail — moving it without shrinking queries just moves the crash.
×

Giving Elasticsearch all server RAM as heap

Symptom
No page cache left for Lucene, GC runs for seconds, and the breaker trips on queries that fit easily elsewhere.
Fix
Size heap at 50% of RAM capped near 26-30 GB for compressed pointers, and keep fielddata cache bounded with indices.fielddata.cache.size. Monitor JVM pressure, not just heap bytes.
×

Paginating with size 10000 plus deep aggregations

Symptom
Single dashboard loads trip the request breaker and queue rejections cascade to innocent queries.
Fix
Cap page sizes at 100-1000, use search_after for deep pagination, and slice scrolls. Rewrite the client first — breaker trips from size-10000 requests are client bugs, not cluster bugs.
×

Watching only trip errors while ignoring memory pressure

Symptom
The first sign of trouble is user-facing 429s, with no warning trend to act on beforehand.
Fix
Alert on JVM memory pressure above 75% sustained and on breaker trip counters. Pressure trends give days of warning; trip counts give minutes. Both belong on the same dashboard as query latency.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
What are circuit breakers protecting?
Q02SENIOR
Why does fielddata on text trip the breaker?
Q03SENIOR
A dashboard trips the request breaker. Your moves?
Q04SENIOR
Why is heap capped near 30GB instead of maximized?
Q05SENIOR
Why does raising limits only delay the next trip?
Q01 of 05JUNIOR

What are circuit breakers protecting?

ANSWER
Breakers estimate each request's heap cost and reject with 429 when it would breach budgets (request ~60%, fielddata ~40-60%, total near heap). They trade failed queries for a live node — without them, one giant aggregation OOMs the JVM and takes all shards with it.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
What does circuit_breaking_exception actually mean?
02
fielddata vs doc_values — what's the difference?
03
How large should Elasticsearch heap be?
04
Can I keep the dashboard and stop the trips?
05
Is raising breaker limits safe?
06
Which metric warns before trips start?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Elasticsearch. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
Elasticsearch Cluster Health Red: Unassigned Shards
2 / 3 · Elasticsearch
Next
Elasticsearch Mapper Parsing Exception on Field Type Conflict
→