Elasticsearch Unassigned Shards: Red to Green Fix
Elasticsearch red with unassigned shards? Read allocation explain, fix disk watermarks or replica math, then reroute safely..
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓An Elasticsearch cluster you can query (even yellow counts)
- ✓curl access with manage-cluster or viewer privileges
- ✓Basic JSON reading for API responses
- Red means primaries are unassigned so writes fail; list them with the shards API and treat primaries as the paging half
- Ask allocation explain per stuck shard since it names the blocking decider: disk, node loss, or replica math
- Free disk below 85% or restore the lost node first, then call reroute with retry_failed exactly once
- Force-allocate stale primaries only for permanently dead nodes with sign-off, since rejoining good copies get deleted
Think of the cluster as a library with every book shelved twice in different wings. Unassigned shards are books with no shelf: the wing flooded (disk full), burned down (node lost), or was never built (too few nodes for the replica count). The librarian's log names the exact problem per book — read it before reshelving, because forcing a damaged copy onto the shelf destroys the good copy when the real one turns up.
The status page flips red at the worst hour: cluster health red, 14 unassigned shards, writes failing on the orders index. Nobody deployed, nobody deleted a node knowingly — but one data node ran out of disk overnight, and the primaries it held have nowhere to go.
Unassigned shards are Elasticsearch telling you exactly what's wrong, if you ask correctly. Every stuck shard carries an allocation explanation naming the decider that blocked it: disk watermark, lost node, impossible replica count, disabled allocation. The incident is won or lost on whether you read that explanation before acting.
This guide takes you from red to green in order: triage with health and shard listings, interrogate with the explain API, fix the four blockers with exact commands, and use reroute only where it's safe. You'll leave with replica math, watermark thresholds, and the discipline that prevents force-allocating your way into data loss. Keep this guide open during the incident: the command sequence below takes most red clusters green without data loss.
Reading Red vs Yellow Before Touching Anything
Cluster health compresses the whole story into one word. Green means every primary and replica is assigned. Yellow means primaries are fine but some replicas sit unassigned — degraded redundancy, working writes. Red means at least one primary is unassigned, and writes to that index fail now.
Scope before treating. The health endpoint gives counts; the shards listing names names — which indices, which shard numbers, primaries or replicas, and the unassigned reason per shard. Primaries are the paging half: each unassigned primary is an index that can't take writes. Replicas are the worry half: the cluster works but one more failure cascades.
Record the baseline: health JSON, full unassigned list, node count, and disk figures. These four snapshots frame every later decision and prove recovery at the end. Incidents worked without baselines drift — engineers fix three shards, declare victory, and miss the fourth.
The _cat APIs compress triage into one-liners worth memorizing. _cat/health shows status, node counts, and pending tasks in a single row — red with relocating shards means recovery in progress, red with zero movement means stuck. _cat/pending_tasks reveals cluster-level blockers like stuck mappings updates that shard listings never show. _cat/thread_pool exposes rejected rejections when masters are overloaded. Run the trio before any shard-level digging; cluster problems masquerade as shard problems constantly. Five seconds of cats saves fifty minutes of wrong-tree barking.
Allocation Explain: Interrogating Each Stuck Shard
The allocation explain API is the single highest-value call in this incident class. For one index and shard it returns every node's verdict with deciders: disk threshold, awareness, replica-count, throttling — each YES or NO with the numbers. The NO votes are your todo list, ordered by authority.
Learn its three signatures. Disk NO with used-percent figures means watermarks — free space, don't reroute. Awareness NO across a zone means shard-awareness rules can't be satisfied — fix topology or filters. Every node NO on replica placement with too few nodes means the replica count exceeds the topology — lower replicas or add nodes.
Explain one shard per pattern, not every shard. Shards on the same index with the same reason share the fix; ten identical explanations waste incident time. One representative per index-reason pair, then act on the pattern.
Two explain flags sharpen the interrogation. include_disk_info adds per-node disk figures to the verdict so watermark math is visible inline. include_yes_decisions lists nodes that WOULD accept the shard, proving the blocker is selective rather than global — an empty yes-list means impossible placement everywhere. Query with ?include_yes_decisions=true once per incident to calibrate: if some nodes say yes, fix the no-voters; if none do, fix the requirement. The API answers allocation questions faster than any dashboard because it runs the deciders directly. Trust its arithmetic over intuition.
Disk Watermarks: the 85/90/95 Tripwire
Disk watermarks strand more shards than all other causes combined. Low at 85% stops new allocations to the node, high at 90% starts relocating shards away, flood_stage at 95% block-writes indices with index.blocks.write. A node past flood_stage looks dead from the outside while its data sits perfectly intact.
Confirm with df -h on data nodes alongside explain's disk figures — both must agree before you act. Then free space decisively: delete expired indices per retention, snapshot-and-delete old ones, force-merge read-only indices, or add nodes. Partial frees that land at 89% still block; target below 85% with margin.
Only then retry. One POST /_cluster/reroute?retry_failed re-attempts shards that exhausted their 5 retries, and the shards listing should show movement within a minute. Retrying into a still-full disk succeeds at the API and changes nothing — the ritual teams repeat for hours when they skip the disk check.
Watermark settings reward deliberate tuning. The defaults (85/90/95 percent) suit general fleets; dense logging clusters often raise them a few points with faster ILM deletes as compensation. Set them via persistent cluster settings so restarts keep the policy, and keep transient overrides for incident-time experiments only. Relocation throttling (cluster.routing.allocation.node_concurrent_recoveries) paces the stampede after a fix so recovery does not saturate networks. Tune watermarks and throttles together: aggressive watermarks with slow recovery still page, just later. Document both numbers where on-call can find them blindfolded.
Replica Math: Copies Need Somewhere to Live
Replica math is topology arithmetic with no mercy. Replicas can't share a node with their primary, and awareness rules can require distinct zones — so replicas: 1 needs 2 nodes, replicas: 2 with zone awareness needs 3 zones. Any index whose replica count exceeds placeable slots stays yellow forever, patiently, through every retry.
Diagnose by counting. Compare number_of_replicas on the stuck index against placeable nodes per the explain output. Single-node clusters with default replica 1 are the classic case — yellow from birth, alarming every new engineer once. Autoscaling down node counts recreates it in production clusters that were green yesterday.
Fix the arithmetic, not the allocator. Lower replicas to nodes-minus-one for the affected indices, or add nodes to fit the count. Then set index templates so new indices inherit topology-fitting defaults — the permanent fix that stops each morning's new yellow index.
Awareness and filtering complete the placement picture. cluster.routing.allocation.awareness.attributes (rack, zone) forces replicas onto distinct failure domains — correct and the reason single-zone dev clusters stay yellow. Index-level allocation filtering (require/include/exclude by node attributes) pins hot indices to fast nodes but strands shards when filters outlive the nodes they named. Audit filters after every hardware refresh; stale require rules are silent shard killers. Total-shards-per-node caps hot-spotting on small clusters. Placement is policy, and policy rots without audits.
Reroute Commands: Safe Verbs and Loaded Guns
Reroute commands move, cancel, and force-allocate individual shards — power tools with permanent consequences. Move and allocate_replica are safe placement hints the balancer then rationalizes. allocate_stale_primary and allocate_empty_primary destroy data by design: stale promotes an old copy, empty creates a blank one, and both require accept_data_loss: true as a speed bump.
The catastrophic sequence is well documented: force a stale primary, the lost node rejoins with the good copy, the cluster deletes the good copy as divergent. Recovery means reindexing from queues or snapshots — hours of writes lost for a command that took seconds. This is why stale allocation needs written stakeholder sign-off and a snapshot attempt first.
Prefer the safe verbs. Fix disk, restore nodes, fit replicas — then a single retry_failed lets the allocator finish honestly. Reserve force-allocation for permanently destroyed nodes where the alternative is red forever, and even then snapshot anything reachable first.
Two reroute verbs handle the everyday cases without drama. move relocates a started shard between named nodes for balancing or draining; allocate_replica assigns an unassigned replica copy to a node that holds it. Both are hints the balancer rationalizes afterward, so follow-up moves by the cluster are normal, not failure. Append ?metric=none to reroute calls for terse responses during incidents when seconds matter. Reserve cancel for shards stuck in broken recoveries, and only with allow_primary when you fully accept primary cancellation semantics. Safe verbs first, always.
Staying Green: Templates, ILM, and Precursor Alerts
Green is the start of prevention, not the end of the incident. Set index templates with replica counts that fit your topology so new indices arrive placeable. Schedule ILM delete phases that actually run — audit them quarterly, because silent ILM failures fill disks on a timer.
Alert on the precursors, not the color. Disk at 75% warns, 80% pages; unassigned-shard count growth pages before red does; allocation-disabled settings page the moment maintenance forgets to re-enable. These three alerts convert 4 AM red pages into Tuesday-afternoon tasks.
Close with proof: health green, zero unassigned, zero relocating, disk below 80% everywhere, snapshots current. Write the timeline down — what blocked, what fixed, what guardrail was added. The next red cluster at 4 AM inherits this one instead of repeating it.
Snapshots and lifecycle policies make the next incident boring. Snapshot lifecycle management (SLM) with hourly snapshots and tested restores converts force-allocation from a gamble into a recoverable choice. Index lifecycle management moves indices through hot, warm, and cold phases before deleting on schedule — the delete phase is the disk-pressure vaccine. Flood-stage blocks auto-release when usage falls back below the high watermark, so freeing space genuinely unblocks writes without manual unblock calls. Verify SLM snapshots restore quarterly; untested backups are rumors. Green clusters are maintained, not lucky.
The Full Disk That Faked Node Death for 6 Hours
- Flood_stage looks like node death from the outside — check disk before declaring nodes dead.
- Force-allocation under pressure converts a disk incident into a data-loss incident; explain first, always.
- ILM policies only help if they actually run — audit delete phases quarterly, not after the flood.
| File | Command / Code | Purpose |
|---|---|---|
| GET /_cluster/health | Reading Red vs Yellow Before Touching Anything | |
| GET /_cluster/allocation/explain | Allocation Explain | |
| df -h /var/lib/elasticsearch | Disk Watermarks | |
| PUT /orders-2026.09/_settings | Replica Math | |
| POST /_cluster/reroute | Reroute Commands |
Key takeaways
Common mistakes to avoid
5 patternsRetrying reroute while the disk watermark still blocks allocation
Force-allocating primaries while the lost node might return
Leaving number_of_replicas above available node count
Running reroute commands without reading allocation explain first
Forgetting allocation was disabled during maintenance
Interview Questions on This Topic
Primaries vs replicas, and what red vs yellow mean.
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Elasticsearch. Mark it forged?
6 min read · try the examples if you haven't