Databricks Spot Reclaimed? Fall Back, Checkpoint
Spot reclaim killed your cluster mid-job? Keep the driver on-demand, set spot fallback, checkpoint to durable storage, and use pools for fast recovery..
20+ years shipping production backend systems. Written from production experience, not tutorials.
- ✓Basic Databricks clusters: drivers, workers, and jobs
- ✓What cloud spot instances are (no deep pricing knowledge needed)
- ✓Batch vs streaming jobs at a conceptual level
- Spot instances save up to ~90% but can be reclaimed anytime with ~2 minutes notice — a reclaimed worker takes its shuffle data and cache with it
- Keep the driver on on-demand (a lost driver kills the whole job) and enable spot-to-on-demand fallback so replacements arrive even when spare spot capacity is gone
- Checkpoint long work to durable storage (S3/ADLS): Structured Streaming checkpoints plus Delta's transaction log make restarts resume instead of repeat
- Use instance pools for minute-fast replacement nodes, and pick cost-optimized fleets for dev but reliability-optimized policies for deadline jobs
Spot instances are standby airline seats: up to 90% cheaper, but you're bumped when full-fare passengers appear. If the bumped traveler carried the group's luggage (a worker's shuffle data), the trip stalls. Databricks lets you buy smart: seat the tour guide full-fare (driver on on-demand — lose the guide, lose the tour), rebook bumped travelers on any airline (fallback to on-demand), photograph luggage at every stop (checkpointing), and keep a shuttle waiting (instance pools).
The page reads 'cluster terminated,' the job was 5 hours into a 6-hour run, and the postmortem says one word: reclaimed. Spot instances — spare cloud capacity sold at up to ~90% off — can be taken back with about 2 minutes of warning whenever demand spikes. Databricks handles the mechanics (request replacements, retry tasks), but your job's design decides whether recovery takes 4 minutes or 4 hours.
Three design choices separate cheap-and-safe from cheap-and-sorry. Driver placement: workers are expendable, the driver isn't — lose the driver and the entire application dies regardless of how healthy workers are. Fallback policy: without spot-to-on-demand fallback, a reclaim during a capacity crunch means no replacements at any price and a dead cluster. Recovery posture: jobs that checkpoint to durable storage resume near the reclaim point, while jobs holding everything in memory or local disk restart from zero.
This beginner-friendly article makes spot eviction boring. You'll learn which node types go where, which fleet policy matches which workload, how checkpointing and Delta make restarts cheap, and how instance pools shrink replacement time from 10 minutes to 1 — so your 90% discount never costs a deadline.
What Reclamation Really Takes From You
A reclaim gives about 2 minutes of warning, then the VM is gone: its memory, its local disks, its shuffle files, its cached blocks — everything not on durable storage evaporates. For a worker, Spark can recover: the DAGScheduler re-submits lost tasks, the external shuffle service (if enabled) may even preserve shuffle blocks beyond the executor, and sibling executors recompute missing partitions from lineage. Recovery costs time proportional to what was lost — minutes for narrow transformations, hours for wide shuffles with no persisted intermediate.
For the driver, there is no recovery: the driver holds the application master, the DAGScheduler state, and all task bookkeeping. A reclaimed driver is a dead application, full stop — executors orphan, the job restarts from zero, and streaming queries restart from their last checkpoint (or from earliest, if none exists). This asymmetry (workers = delay, driver = death) is the single fact that organizes all spot strategy.
Reclaims also correlate: demand spikes reclaim many spot VMs at once, often across jobs sharing a family/AZ. The night you lose 2 workers is the night replacements are scarcest — which is exactly when spot-only fleets (no fallback) discover they can't buy capacity at any price. Design for correlated reclaims during scarcity, not single reclaims during plenty.
Driver on On-Demand: The $3 Rule
Databricks launches the driver on the cluster's first node — so making the first node on-demand guarantees driver safety in one setting. The cost is trivially small (one node of dozens, ~$3/night in our incident) against the downside (total run loss plus SLA penalties). There is no workload where a spot driver is rational for scheduled production jobs: the expected value is negative the first time a driver reclaim lands inside a long run, and reclaims are a when, not an if, on horizons that matter.
Enforce it so nobody re-learns it: a workspace cluster policy requiring first-node on-demand (or driver-on-demand where the UI exposes it) rejects all-spot configs at launch time with a clear message, not at 5:40 AM with a dead application. Extend the rule to long-running interactive clusters used for production-adjacent work — a reclaimed driver kills the notebook state data scientists spent hours building, and their retry is manual.
The same logic covers the single-point-of-failure siblings: the external shuffle service node where used, and any singleton coordinator your pipeline depends on. Spot is for the replaceable many; on-demand is for the irreplaceable few. One line of policy, permanent immunity from the fatal class.
Fleet Policies and Fallback: Degrade Cost, Not Capacity
Databricks fleet options trade savings against reclaim risk on a spectrum. All-spot maximizes discount and maximizes correlated-reclaim exposure (no replacements during crunches). Spot-with-fallback keeps the discount in normal times and buys on-demand automatically when spot vanishes — crunches degrade your bill, not your capacity. Mixed fleets (N on-demand workers + rest spot) add a guaranteed base that survives total spot droughts: the job slows instead of dying.
Match the point on the spectrum to the stakes. Dev, experiments, and retry-friendly ETL: all-spot with fallback (cheap, and failures cost minutes). SLA-bound production: mixed base (4+ on-demand of 24 in our fix) plus fallback — the base holds shuffle-critical mass while spot churns around it. Regulatory/deadline jobs: consider reliability-optimized or all-on-demand during the critical window; the 90% discount isn't worth any penalty clause.
Diversify within spot: multiple instance families and AZs in the fleet spread reclaim correlation (one family spiking doesn't take the whole fleet), and excluding chronic-reclaim families (tracked per family/AZ from your own event logs) removes repeat offenders. Fallback plus diversification is the belt and suspenders: one handles scarcity, the other handles spikes.
Checkpointing: Resume Instead of Repeat
Checkpointing converts driver death from total loss to partial loss: work recorded on durable storage (S3/ADLS, never local disk) survives the cluster. Structured Streaming mandates it — checkpointLocation plus write-ahead logs give exactly-once resume from the last committed micro-batch; without it, a driver reclaim replays from earliest (days of duplication) or fails outright. Set it on durable paths at creation: moving a checkpoint later breaks the resume chain.
Long batch jobs need the same thinking in manual form: stage outputs to Delta per phase (ingest → clean → aggregate as separate committed tables or partitions), so a 5:40 AM death resumes from the last staged commit, not from input. RDD/Dataset checkpointing (rdd.checkpoint() with a durable checkpoint dir) truncates lineage for iterative algorithms — without it, recovery replays the full 5-hour lineage graph instead of one checkpoint read.
Delta's transaction log is the quiet partner: every commit is atomic and versioned, so resumed writers pick up at version N+1 with no partial state — and idempotent sinks (keyed MERGEs, batch-ID-tracked foreachBatch) make the resumed replay converge instead of duplicate. Checkpoint the progress, commit atomically, write idempotently: the trio that makes any restart boring.
Instance Pools: Replacements in a Minute, Not Ten
Replacement latency decides whether worker reclaims are invisible: tasks retry, but if new nodes take 10 minutes, stages stall, streams lag, and correlated reclaims cascade into timeouts. Instance pools keep pre-warmed (idle) nodes ready, cutting replacement time from ~9 minutes (cold cloud launch + Databricks setup) to ~1 minute (claim idle + join). Our fix's 10-idle pool cost pennies per night against the 9-minute stall it erased.
Size pools from reclaim math, not hope: idle count covering your p95 simultaneous-reclaim burst (2-4 nodes for typical fleets), max capacity above peak fleet size, and the same families/AZs as production fleets so replacements match. Pools also accelerate everyday elasticity (autoscaling grabs warm nodes) and dev cluster startup — one pool serves many clusters, amortizing idle cost across the workspace.
Pre-warm correctly: pool nodes should carry the Databricks runtime version and disk image your jobs use, or the 'warm' node still pays setup time. Tag pool costs separately and review monthly — an oversized pool of idle GPU nodes is its own budget incident. Right-sized pools are the rare optimization that's simultaneously cheaper (less stall waste) and safer (faster recovery).
Matching Policy to Stakes: A Spot Playbook
One fleet policy for every workload is how teams get both wasted money and missed SLAs: dev overpays on on-demand while production gambles on all-spot. Split into three lanes. Lane 1 (dev/experiments): all-spot with fallback, pools shared, checkpoints optional — failures cost minutes and teach. Lane 2 (routine production ETL): on-demand driver, spot-majority workers with fallback, durable checkpoints mandatory, pool-attached — the discount with guardrails. Lane 3 (SLA/regulatory): mixed on-demand base sized to shuffle-critical mass, reliability-optimized spot remainder, staged Delta commits, rehearsed recovery — savings where safe, certainty where owed.
Govern the lanes with policy, not tribal knowledge: workspace cluster policies per lane (instance rules, fallback flags, pool attachment, autotermination) so the safe shape is the default shape and the risky shape requires breaking glass. Showback completes it: dashboard spot-share of worker-hours and reclaim counts per lane so finance sees the discount and engineering sees the risk — our fixed setup held ~70% spot worker-hours with zero SLA misses, the number that ended the all-spot debate permanently.
Revisit per quarter or per growth spurt: new instance families shift reclaim rates, new AZs shift correlation, and doubled volumes shift which stages are shuffle-critical. The playbook is a living doc with one owner, one dashboard, and one game-day — boring by design, because spot excitement is always bad news.
Reclaimed at Hour 5 of 6: The $40 Saving That Cost a Deadline
- Price the driver correctly: $3/night of on-demand insurance against an $18K late penalty is the cheapest line item in the budget — all-spot fleets gamble the entire run to save the smallest node.
- Worker reclaims are routine; driver reclaims are fatal. Design for the fatal case (on-demand driver, durable checkpoints, Delta's resumable log) and the routine case handles itself through retries and pools.
- Fallback converts capacity crunches from outages into overcharges: spot-to-on-demand replacement kept the fixed job green through 3 subsequent crunch nights at slightly higher cost instead of another missed SLA.
| File | Command / Code | Purpose |
|---|---|---|
| spot_safe_cluster.py | cluster = { | Driver on On-Demand |
| resumable_stream.py | from pyspark.sql import functions as F | Checkpointing |
| pool_provision.sh | databricks instance-pools create --json '{ | Instance Pools |
| lane_policies.sh | databricks clusters create --json '{ | Matching Policy to Stakes |
Key takeaways
Common mistakes to avoid
5 patternsRunning the driver on spot to save ~$3/night
Spot-only fleets with no on-demand fallback
Checkpointing to local disk or not at all
Cold replacements taking 9+ minutes per reclaim
Same fleet policy for dev toys and SLA jobs
Interview Questions on This Topic
Why is losing a spot driver fatal while losing a spot worker is routine?
Frequently Asked Questions
20+ years shipping production backend systems. Written from production experience, not tutorials.
That's Databricks. Mark it forged?
5 min read · try the examples if you haven't