Home › Observability › Elasticsearch Unassigned Shards: Red to Green Fix
Intermediate 6 min · September 23, 2026

Elasticsearch Unassigned Shards: Red to Green Fix

Elasticsearch red with unassigned shards? Read allocation explain, fix disk watermarks or replica math, then reroute safely..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 13 min
  • ✓An Elasticsearch cluster you can query (even yellow counts)
  • ✓curl access with manage-cluster or viewer privileges
  • ✓Basic JSON reading for API responses
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Red means primaries are unassigned so writes fail; list them with the shards API and treat primaries as the paging half
  • Ask allocation explain per stuck shard since it names the blocking decider: disk, node loss, or replica math
  • Free disk below 85% or restore the lost node first, then call reroute with retry_failed exactly once
  • Force-allocate stale primaries only for permanently dead nodes with sign-off, since rejoining good copies get deleted
✦ Definition~90s read
What is Elasticsearch Cluster Health Red?

Elasticsearch splits each index into shards — primary shards holding authoritative data and replica shards copying them for redundancy and read scale. The cluster continuously assigns every shard copy to a node; assigned means serving, unassigned means waiting.

★
Think of the cluster as a library with every book shelved twice in different wings.

Cluster health summarizes the assignment: green when all copies are placed, yellow when replicas wait, red when any primary waits and its index can't take writes.

Shards go unassigned for concrete reasons enforced by allocation deciders. Disk watermarks (85% low, 90% high, 95% flood_stage) refuse placement on full nodes and eventually block writes. Lost nodes take their primaries with them until restored. Replica counts exceeding placeable nodes or zones can never be satisfied.

Disabled allocation freezes everything, and shards failing 5 consecutive attempts wait for manual retry.

Recovery follows the deciders, not willpower. The allocation explain API names the blocking decider per shard; the fix removes that blocker — freeing disk, restoring nodes, fitting replica counts, re-enabling allocation — and a single reroute with retry_failed lets placement proceed.

Force-allocation commands bypass deciders at the cost of data, reserved for permanently destroyed nodes with explicit sign-off.

Plain-English First

Think of the cluster as a library with every book shelved twice in different wings. Unassigned shards are books with no shelf: the wing flooded (disk full), burned down (node lost), or was never built (too few nodes for the replica count). The librarian's log names the exact problem per book — read it before reshelving, because forcing a damaged copy onto the shelf destroys the good copy when the real one turns up.

The status page flips red at the worst hour: cluster health red, 14 unassigned shards, writes failing on the orders index. Nobody deployed, nobody deleted a node knowingly — but one data node ran out of disk overnight, and the primaries it held have nowhere to go.

Unassigned shards are Elasticsearch telling you exactly what's wrong, if you ask correctly. Every stuck shard carries an allocation explanation naming the decider that blocked it: disk watermark, lost node, impossible replica count, disabled allocation. The incident is won or lost on whether you read that explanation before acting.

This guide takes you from red to green in order: triage with health and shard listings, interrogate with the explain API, fix the four blockers with exact commands, and use reroute only where it's safe. You'll leave with replica math, watermark thresholds, and the discipline that prevents force-allocating your way into data loss. Keep this guide open during the incident: the command sequence below takes most red clusters green without data loss.

Reading Red vs Yellow Before Touching Anything

Cluster health compresses the whole story into one word. Green means every primary and replica is assigned. Yellow means primaries are fine but some replicas sit unassigned — degraded redundancy, working writes. Red means at least one primary is unassigned, and writes to that index fail now.

Scope before treating. The health endpoint gives counts; the shards listing names names — which indices, which shard numbers, primaries or replicas, and the unassigned reason per shard. Primaries are the paging half: each unassigned primary is an index that can't take writes. Replicas are the worry half: the cluster works but one more failure cascades.

Record the baseline: health JSON, full unassigned list, node count, and disk figures. These four snapshots frame every later decision and prove recovery at the end. Incidents worked without baselines drift — engineers fix three shards, declare victory, and miss the fourth.

The _cat APIs compress triage into one-liners worth memorizing. _cat/health shows status, node counts, and pending tasks in a single row — red with relocating shards means recovery in progress, red with zero movement means stuck. _cat/pending_tasks reveals cluster-level blockers like stuck mappings updates that shard listings never show. _cat/thread_pool exposes rejected rejections when masters are overloaded. Run the trio before any shard-level digging; cluster problems masquerade as shard problems constantly. Five seconds of cats saves fifty minutes of wrong-tree barking.

JSON
1
2
3
4
5
6
# Triage: health, then the unassigned list
GET /_cluster/health
GET /_cat/shards?v&h=index,shard,prirep,state,unassigned.reason,node

# Scope it: primaries (p) unassigned = red, writes failing
# GET /_cat/shards/orders-2026.09?v&h=shard,prirep,state,unassigned.reason
📊 Production Insight
A team treated yellow like red and force-allocated replicas during a routine restart, forking an index. Five minutes of scoping would have shown zero unassigned primaries — a wait-and-watch case, not a reroute case.
🎯 Key Takeaway
Health names the severity; the shards list names the patients — snapshot both before acting.

Allocation Explain: Interrogating Each Stuck Shard

The allocation explain API is the single highest-value call in this incident class. For one index and shard it returns every node's verdict with deciders: disk threshold, awareness, replica-count, throttling — each YES or NO with the numbers. The NO votes are your todo list, ordered by authority.

Learn its three signatures. Disk NO with used-percent figures means watermarks — free space, don't reroute. Awareness NO across a zone means shard-awareness rules can't be satisfied — fix topology or filters. Every node NO on replica placement with too few nodes means the replica count exceeds the topology — lower replicas or add nodes.

Explain one shard per pattern, not every shard. Shards on the same index with the same reason share the fix; ten identical explanations waste incident time. One representative per index-reason pair, then act on the pattern.

Two explain flags sharpen the interrogation. include_disk_info adds per-node disk figures to the verdict so watermark math is visible inline. include_yes_decisions lists nodes that WOULD accept the shard, proving the blocker is selective rather than global — an empty yes-list means impossible placement everywhere. Query with ?include_yes_decisions=true once per incident to calibrate: if some nodes say yes, fix the no-voters; if none do, fix the requirement. The API answers allocation questions faster than any dashboard because it runs the deciders directly. Trust its arithmetic over intuition.

JSON
1
2
3
4
5
6
7
8
9
# Interrogate one stuck shard — the decider list is the diagnosis
GET /_cluster/allocation/explain
{
  "index": "orders-2026.09",
  "shard": 2,
  "primary": true
}
# Read: node_allocation_decisions[].deciders[]
# "disk threshold exceeded" -> watermarks | "awareness" -> zone rules
📊 Production Insight
An engineer ran explain on 40 shards serially while writes failed. The first shard's disk-threshold NO already contained the whole diagnosis — 30 minutes burned reading 39 duplicates of the same answer.
🎯 Key Takeaway
Explain names the blocking decider per shard — fix that blocker instead of overriding placement.

Disk Watermarks: the 85/90/95 Tripwire

Disk watermarks strand more shards than all other causes combined. Low at 85% stops new allocations to the node, high at 90% starts relocating shards away, flood_stage at 95% block-writes indices with index.blocks.write. A node past flood_stage looks dead from the outside while its data sits perfectly intact.

Confirm with df -h on data nodes alongside explain's disk figures — both must agree before you act. Then free space decisively: delete expired indices per retention, snapshot-and-delete old ones, force-merge read-only indices, or add nodes. Partial frees that land at 89% still block; target below 85% with margin.

Only then retry. One POST /_cluster/reroute?retry_failed re-attempts shards that exhausted their 5 retries, and the shards listing should show movement within a minute. Retrying into a still-full disk succeeds at the API and changes nothing — the ritual teams repeat for hours when they skip the disk check.

Watermark settings reward deliberate tuning. The defaults (85/90/95 percent) suit general fleets; dense logging clusters often raise them a few points with faster ILM deletes as compensation. Set them via persistent cluster settings so restarts keep the policy, and keep transient overrides for incident-time experiments only. Relocation throttling (cluster.routing.allocation.node_concurrent_recoveries) paces the stampede after a fix so recovery does not saturate networks. Tune watermarks and throttles together: aggressive watermarks with slow recovery still page, just later. Document both numbers where on-call can find them blindfolded.

BASH
1
2
3
4
5
# Disk pressure per data node, then one retry after freeing below 85%
df -h /var/lib/elasticsearch
curl -s 'localhost:9200/_cat/allocation?v&h=shards,disk.used_percent,node'
curl -s -XPOST 'localhost:9200/_cluster/reroute?retry_failed'
curl -s 'localhost:9200/_cat/shards?v&h=index,shard,prirep,state,node' | grep -c UNASSIGNED
📊 Production Insight
A cluster sat red for 5 hours across repeated retry_failed calls — disk at 93% the whole time. One 400 GB ILM delete plus a single retry greened it in 8 minutes. The retries were never the fix; the space was.
🎯 Key Takeaway
Free below 85% first, then one retry_failed — retrying into full disks changes nothing.

Replica Math: Copies Need Somewhere to Live

Replica math is topology arithmetic with no mercy. Replicas can't share a node with their primary, and awareness rules can require distinct zones — so replicas: 1 needs 2 nodes, replicas: 2 with zone awareness needs 3 zones. Any index whose replica count exceeds placeable slots stays yellow forever, patiently, through every retry.

Diagnose by counting. Compare number_of_replicas on the stuck index against placeable nodes per the explain output. Single-node clusters with default replica 1 are the classic case — yellow from birth, alarming every new engineer once. Autoscaling down node counts recreates it in production clusters that were green yesterday.

Fix the arithmetic, not the allocator. Lower replicas to nodes-minus-one for the affected indices, or add nodes to fit the count. Then set index templates so new indices inherit topology-fitting defaults — the permanent fix that stops each morning's new yellow index.

Awareness and filtering complete the placement picture. cluster.routing.allocation.awareness.attributes (rack, zone) forces replicas onto distinct failure domains — correct and the reason single-zone dev clusters stay yellow. Index-level allocation filtering (require/include/exclude by node attributes) pins hot indices to fast nodes but strands shards when filters outlive the nodes they named. Audit filters after every hardware refresh; stale require rules are silent shard killers. Total-shards-per-node caps hot-spotting on small clusters. Placement is policy, and policy rots without audits.

JSON
1
2
3
4
5
6
7
8
9
10
# Fit replicas to a 3-node topology
PUT /orders-2026.09/_settings
{ "index": { "number_of_replicas": 1 } }

# Template so new indices inherit sane defaults
PUT /_index_template/orders_template
{
  "index_patterns": ["orders-*"],
  "template": { "settings": { "number_of_replicas": 1 } }
}
📊 Production Insight
A scale-in from 5 nodes to 3 left replica-2 indices yellow across 60 indices. One template change plus a settings update greened everything in minutes — the allocator was innocent throughout.
🎯 Key Takeaway
Replicas need distinct nodes; fit counts to topology and template the defaults.

Reroute Commands: Safe Verbs and Loaded Guns

Reroute commands move, cancel, and force-allocate individual shards — power tools with permanent consequences. Move and allocate_replica are safe placement hints the balancer then rationalizes. allocate_stale_primary and allocate_empty_primary destroy data by design: stale promotes an old copy, empty creates a blank one, and both require accept_data_loss: true as a speed bump.

The catastrophic sequence is well documented: force a stale primary, the lost node rejoins with the good copy, the cluster deletes the good copy as divergent. Recovery means reindexing from queues or snapshots — hours of writes lost for a command that took seconds. This is why stale allocation needs written stakeholder sign-off and a snapshot attempt first.

Prefer the safe verbs. Fix disk, restore nodes, fit replicas — then a single retry_failed lets the allocator finish honestly. Reserve force-allocation for permanently destroyed nodes where the alternative is red forever, and even then snapshot anything reachable first.

Two reroute verbs handle the everyday cases without drama. move relocates a started shard between named nodes for balancing or draining; allocate_replica assigns an unassigned replica copy to a node that holds it. Both are hints the balancer rationalizes afterward, so follow-up moves by the cluster are normal, not failure. Append ?metric=none to reroute calls for terse responses during incidents when seconds matter. Reserve cancel for shards stuck in broken recoveries, and only with allow_primary when you fully accept primary cancellation semantics. Safe verbs first, always.

JSON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# Safe: move a replica to a named node, balancer rationalizes after
POST /_cluster/reroute
{
  "commands": [
    { "allocate_replica": { "index": "orders-2026.09", "shard": 2, "node": "data-03" } }
  ]
}

# LAST RESORT: stale primary for a permanently dead node only
# Requires accept_data_loss and stakeholder sign-off
POST /_cluster/reroute
{
  "commands": [
    { "allocate_stale_primary": {
        "index": "orders-2026.09", "shard": 0,
        "node": "data-02", "accept_data_loss": true } }
  ]
}
📊 Production Insight
A premature stale-primary call forked 3 indices and forced a 6-hour reindex of 2.1M orders. The disk fix that followed would have recovered everything losslessly — 20 minutes of patience was worth 6 hours of writes.
🎯 Key Takeaway
Move and retry are safe; stale and empty primaries destroy data — sign-off and snapshot first.

Staying Green: Templates, ILM, and Precursor Alerts

Green is the start of prevention, not the end of the incident. Set index templates with replica counts that fit your topology so new indices arrive placeable. Schedule ILM delete phases that actually run — audit them quarterly, because silent ILM failures fill disks on a timer.

Alert on the precursors, not the color. Disk at 75% warns, 80% pages; unassigned-shard count growth pages before red does; allocation-disabled settings page the moment maintenance forgets to re-enable. These three alerts convert 4 AM red pages into Tuesday-afternoon tasks.

Close with proof: health green, zero unassigned, zero relocating, disk below 80% everywhere, snapshots current. Write the timeline down — what blocked, what fixed, what guardrail was added. The next red cluster at 4 AM inherits this one instead of repeating it.

Snapshots and lifecycle policies make the next incident boring. Snapshot lifecycle management (SLM) with hourly snapshots and tested restores converts force-allocation from a gamble into a recoverable choice. Index lifecycle management moves indices through hot, warm, and cold phases before deleting on schedule — the delete phase is the disk-pressure vaccine. Flood-stage blocks auto-release when usage falls back below the high watermark, so freeing space genuinely unblocks writes without manual unblock calls. Verify SLM snapshots restore quarterly; untested backups are rumors. Green clusters are maintained, not lucky.

💡Force-Allocation Destroys the Good Copy
Never force-allocate a stale or empty primary while the lost node might return. The rejoining good copy gets deleted as divergent, converting a recoverable outage into permanent data loss. Restore or rule out the node first — the command waits, the data doesn't come back.
📊 Production Insight
After adding 75/80% disk alerts and quarterly ILM audits, one fleet went 18 months without a red cluster. The two yellow episodes in that span self-resolved before business hours — the alerts had fired days earlier.
🎯 Key Takeaway
Template replicas, audit ILM deletes, alert on disk and unassigned growth — prove green with numbers.
● Production incidentPOST-MORTEMseverity: high

The Full Disk That Faked Node Death for 6 Hours

Symptom
Orders index went red at 4 AM with 14 unassigned shards and failed writes. Checkout errors hit 100% for 6 hours while the team force-allocated, freed disk, and reindexed 2.1M order documents from the queue.
Assumption
The team assumed a node had crashed and force-allocated the primaries to surviving nodes within 20 minutes. The node hadn't crashed — its disk filled past flood_stage at 95%, which block-wrote its indices and froze it out of allocation. The force-allocate created stale primaries beside the good copies.
Root cause
Log indices without a working ILM delete phase filled data nodes to 95% flood_stage, which set index.blocks.write and stalled allocation. Fourteen shards (6 primaries) went unassigned. The premature allocate_stale_primary calls forked 3 indices, requiring reindexing from the order queue after disk was freed.
Fix
They freed 400 GB by deleting expired log indices per ILM ages that had never run, dropped nodes below 85%, and called retry_failed. The duplicates required closing the stale-forced indices and reindexing 6 hours of orders from the queue. Then they enabled ILM deletes, set disk alerts at 75% and 80%, and banned force-allocation without two-person sign-off.
Key lesson
  • Flood_stage looks like node death from the outside — check disk before declaring nodes dead.
  • Force-allocation under pressure converts a disk incident into a data-loss incident; explain first, always.
  • ILM policies only help if they actually run — audit delete phases quarterly, not after the flood.
Production debug guideFive steps from red status to guarded green without data loss.5 entries
Symptom · 01
Cluster red or yellow with unknown scope
→
Fix
Run GET /_cluster/health and GET /_cat/shards?v with grep UNASSIGNED. Count primaries (p) vs replicas (r): any unassigned p means red and failed writes — that's the paging half. Note the indices and shard numbers; everything downstream keys off this list.
Symptom · 02
Unassigned shards listed but cause unknown
→
Fix
Call GET /_cluster/allocation/explain with the index, shard, and primary flag for one stuck shard. Read the decider list: disk threshold NO means watermarks, awareness NO means zone rules, same-shard NO with one node means replica math. Fix the named blocker — never reroute blind.
Symptom · 03
Explain blames the disk decider
→
Fix
Check df -h on data nodes against 85% low, 90% high, 95% flood_stage. If breached, delete aged indices or snapshots, force-merge, or add nodes until below 85%. Only then call POST /_cluster/reroute?retry_failed once and confirm shards move via the shards listing.
Symptom · 04
Node lost holding primaries
→
Fix
Compare current node count to the missing names in cluster state. If the node is recoverable, restore it or its data path and let primaries recover locally. Only for permanently destroyed nodes, with written sign-off, use allocate_stale_primary with accept_data_loss: true — and snapshot first if any copy is reachable.
Symptom · 05
Green again; keep it green
→
Fix
Set index templates with topology-fitting replica defaults, alert disk at 75% and 80%, alert on status red and unassigned-count growth, and alert if cluster.routing.allocation.enable drifts from all. Verify green plus zero relocating shards before closing the incident.
Unassigned-Shard Causes Compared
Root CauseHow to ConfirmFixPrevention
Disk watermark breachedExplain shows decider NO with disk threshold; df confirmsFree disk or add nodes; retry_failedDisk alerts at 75/80%; ILM deletes on schedule
Node lost with primariesNodes count dropped; primaries unassignedRestore node; stale-primary only if gone foreverAwareness attributes; snapshot before topology work
Replica count exceeds nodesShards want nodes that don't exist per explainLower replicas or add nodesIndex templates with topology-aware replica counts
Allocation disabled or max retriesSetting none; explain shows repeated failures = 5Re-enable; fix cause then retry_failedMaintenance checklist; alert on the setting
⚙ Quick Reference
5 commands from this guide
FileCommand / CodePurpose
GET /_cluster/healthReading Red vs Yellow Before Touching Anything
GET /_cluster/allocation/explainAllocation Explain
df -h /var/lib/elasticsearchDisk Watermarks
PUT /orders-2026.09/_settingsReplica Math
POST /_cluster/rerouteReroute Commands

Key takeaways

1
Red means unassigned primaries
triage with health, shards listing, then explain, in that order.
2
Disk watermarks (85/90/95%) strand shards; free disk before any reroute.
3
Read allocation explain before rerouting
it names the exact blocking decider.
4
Force-allocate stale primaries only for permanently dead nodes with sign-off.
5
Replica counts must fit node topology; recompute after every scaling event.
6
Alert on disk at 75%, red status, and allocation-disabled settings.

Common mistakes to avoid

5 patterns
×

Retrying reroute while the disk watermark still blocks allocation

Symptom
Every retry_failed returns success but shards stay unassigned, because the decider rejects each attempt identically.
Fix
Free disk below the low watermark before retrying: delete old indices, force-merge, or add nodes. Then call POST /_cluster/reroute?retry_failed once and watch allocation explain confirm movement.
×

Force-allocating primaries while the lost node might return

Symptom
The node rejoins with the good copy and the cluster deletes it in favor of the stale copy you forced — permanent silent data loss.
Fix
Restore the lost node or its data path first, then let the cluster recover. Use allocate_stale_primary only when the node is permanently gone and stakeholders accept the loss window.
×

Leaving number_of_replicas above available node count

Symptom
Yellow status forever: the cluster wants copies it has nowhere to put, and every new index inherits the impossible setting.
Fix
Set replicas to nodes-minus-one for critical indices (1 replica on 2+ nodes, 2 on 3+). Recompute after every scaling event — autoscaling changes the math silently.
×

Running reroute commands without reading allocation explain first

Symptom
Manual allocations fight the balancer, shards flap between nodes, and the original blocker (disk, awareness) still jams everything else.
Fix
Call the explain API for the exact index and shard before any reroute. The decider list names the blocker; fix that blocker instead of overriding placement.
×

Forgetting allocation was disabled during maintenance

Symptom
Nodes healthy, disk free, shards unassigned for hours — because the cluster was told not to allocate anything.
Fix
Re-enable allocation after the maintenance that disabled it, and alert on cluster.routing.allocation.enable != all. A forgotten none setting freezes recovery indefinitely.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Primaries vs replicas, and what red vs yellow mean.
Q02SENIOR
How do disk watermarks strand shards?
Q03SENIOR
Why read allocation explain before rerouting?
Q04SENIOR
What does allocate_stale_primary risk?
Q05SENIOR
A shard exhausted max_retries. What's the recovery?
Q01 of 05JUNIOR

Primaries vs replicas, and what red vs yellow mean.

ANSWER
Primaries hold the authoritative copy and accept writes; replicas copy primaries and serve reads. Red means an unassigned primary (writes fail); yellow means only replicas unassigned (reduced redundancy). Replica math needs distinct nodes — replicas can't share a node with their primary.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Red vs yellow cluster health: which pages?
02
Is allocate_stale_primary ever safe?
03
How many times does Elasticsearch retry allocation?
04
Why keep 15% disk free on data nodes?
05
Single-node cluster stuck yellow — normal?
06
When is unassigned-shard noise vs emergency?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Elasticsearch. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Grafana Datasource Proxy Error: Connection Refused
1 / 3 · Elasticsearch
Next
Elasticsearch Circuit Breaking Exception: Data Too Large
→