Home › Data Engineering › Databricks Spot Reclaimed? Fall Back, Checkpoint
Beginner 5 min · September 23, 2026

Databricks Spot Reclaimed? Fall Back, Checkpoint

Spot reclaim killed your cluster mid-job? Keep the driver on-demand, set spot fallback, checkpoint to durable storage, and use pools for fast recovery..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 10 min
  • ✓Basic Databricks clusters: drivers, workers, and jobs
  • ✓What cloud spot instances are (no deep pricing knowledge needed)
  • ✓Batch vs streaming jobs at a conceptual level
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Spot instances save up to ~90% but can be reclaimed anytime with ~2 minutes notice — a reclaimed worker takes its shuffle data and cache with it
  • Keep the driver on on-demand (a lost driver kills the whole job) and enable spot-to-on-demand fallback so replacements arrive even when spare spot capacity is gone
  • Checkpoint long work to durable storage (S3/ADLS): Structured Streaming checkpoints plus Delta's transaction log make restarts resume instead of repeat
  • Use instance pools for minute-fast replacement nodes, and pick cost-optimized fleets for dev but reliability-optimized policies for deadline jobs
✦ Definition~90s read
What is Databricks Cluster Terminated?

Spot instances are spare cloud capacity auctioned cheaply (up to ~90% off) with one condition: the provider reclaims them on ~2 minutes notice when on-demand customers need the capacity. Databricks builds clusters from these (plus on-demand and reserved) under fleet policies you choose: all-spot, spot-with-fallback, mixed, or on-demand.

★
Spot instances are standby airline seats: up to 90% cheaper, but you're bumped when full-fare passengers appear.

The discount is real and the reclaim is contractual — not a failure, but a scheduled surprise you architect around.

Recovery mechanics differ by victim. Workers: Spark's lineage recomputes lost partitions on replacement nodes, tasks retry up to their limits, and streams resume from durable checkpoints — cost measured in minutes plus recompute. Driver: the application master dies with all scheduling state; executors orphan; batch restarts from the last staged commit and streams from the last checkpoint (or earliest, absent one).

Instance pools (pre-warmed idle nodes) and fallback (automatic on-demand buying) shrink the worker case to ~1 minute; on-demand drivers delete the fatal case entirely.

The economics favor the guarded middle: our fixed fleet kept ~70% spot worker-hours (most of the discount) with zero SLA misses (none of the risk). All-spot gambles runs to save single-digit dollars; all-on-demand leaves 70% savings unclaimed. Policy lanes per workload stakes capture both — cheap where expendable, certain where owed.

Plain-English First

Spot instances are standby airline seats: up to 90% cheaper, but you're bumped when full-fare passengers appear. If the bumped traveler carried the group's luggage (a worker's shuffle data), the trip stalls. Databricks lets you buy smart: seat the tour guide full-fare (driver on on-demand — lose the guide, lose the tour), rebook bumped travelers on any airline (fallback to on-demand), photograph luggage at every stop (checkpointing), and keep a shuttle waiting (instance pools).

The page reads 'cluster terminated,' the job was 5 hours into a 6-hour run, and the postmortem says one word: reclaimed. Spot instances — spare cloud capacity sold at up to ~90% off — can be taken back with about 2 minutes of warning whenever demand spikes. Databricks handles the mechanics (request replacements, retry tasks), but your job's design decides whether recovery takes 4 minutes or 4 hours.

Three design choices separate cheap-and-safe from cheap-and-sorry. Driver placement: workers are expendable, the driver isn't — lose the driver and the entire application dies regardless of how healthy workers are. Fallback policy: without spot-to-on-demand fallback, a reclaim during a capacity crunch means no replacements at any price and a dead cluster. Recovery posture: jobs that checkpoint to durable storage resume near the reclaim point, while jobs holding everything in memory or local disk restart from zero.

This beginner-friendly article makes spot eviction boring. You'll learn which node types go where, which fleet policy matches which workload, how checkpointing and Delta make restarts cheap, and how instance pools shrink replacement time from 10 minutes to 1 — so your 90% discount never costs a deadline.

What Reclamation Really Takes From You

A reclaim gives about 2 minutes of warning, then the VM is gone: its memory, its local disks, its shuffle files, its cached blocks — everything not on durable storage evaporates. For a worker, Spark can recover: the DAGScheduler re-submits lost tasks, the external shuffle service (if enabled) may even preserve shuffle blocks beyond the executor, and sibling executors recompute missing partitions from lineage. Recovery costs time proportional to what was lost — minutes for narrow transformations, hours for wide shuffles with no persisted intermediate.

For the driver, there is no recovery: the driver holds the application master, the DAGScheduler state, and all task bookkeeping. A reclaimed driver is a dead application, full stop — executors orphan, the job restarts from zero, and streaming queries restart from their last checkpoint (or from earliest, if none exists). This asymmetry (workers = delay, driver = death) is the single fact that organizes all spot strategy.

Reclaims also correlate: demand spikes reclaim many spot VMs at once, often across jobs sharing a family/AZ. The night you lose 2 workers is the night replacements are scarcest — which is exactly when spot-only fleets (no fallback) discover they can't buy capacity at any price. Design for correlated reclaims during scarcity, not single reclaims during plenty.

📊 Production Insight
Postmortems that classify 'driver vs worker' in the first line resolve twice as fast — worker reclaims need patience plus pools, driver reclaims need a restart plus a policy change, and mixing them up wastes the whole incident.
🎯 Key Takeaway
Workers lost cost time (recompute from lineage); the driver lost costs the run — and reclaims arrive correlated during scarcity, when spot-only fleets can't replace at any price.

Driver on On-Demand: The $3 Rule

Databricks launches the driver on the cluster's first node — so making the first node on-demand guarantees driver safety in one setting. The cost is trivially small (one node of dozens, ~$3/night in our incident) against the downside (total run loss plus SLA penalties). There is no workload where a spot driver is rational for scheduled production jobs: the expected value is negative the first time a driver reclaim lands inside a long run, and reclaims are a when, not an if, on horizons that matter.

Enforce it so nobody re-learns it: a workspace cluster policy requiring first-node on-demand (or driver-on-demand where the UI exposes it) rejects all-spot configs at launch time with a clear message, not at 5:40 AM with a dead application. Extend the rule to long-running interactive clusters used for production-adjacent work — a reclaimed driver kills the notebook state data scientists spent hours building, and their retry is manual.

The same logic covers the single-point-of-failure siblings: the external shuffle service node where used, and any singleton coordinator your pipeline depends on. Spot is for the replaceable many; on-demand is for the irreplaceable few. One line of policy, permanent immunity from the fatal class.

spot_safe_cluster.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# Databricks cluster shape that survives reclamation by construction.
# (Translate to your workspace's cluster policy / Terraform / REST payload.)
cluster = {
    "cluster_name": "risk-nightly-spot-safe",
    "spark_version": "15.4.x-scala2.12",
    "node_type_id": "m5d.2xlarge",
    # Driver safety: FIRST node on-demand. Databricks puts the driver
    # on the first node, so this one setting ends driver-reclaim deaths.
    "first_on_demand": 1,          # driver: always on-demand (~$3/night)
    "num_workers": 24,
    # Workers: mostly spot for the ~70-90% discount...
    "spot_bid_max_price": -1,      # default market price (no bidding games)
    # ...with automatic fallback so crunches cost money, not the run.
    "spot_fallback_to_on_demand": True,
    "availability": "SPOT_WITH_FALLBACK",
    "autotermination_minutes": 30,
}
print("Driver: on-demand. Workers: spot w/ on-demand fallback.")
📊 Production Insight
The one-line first-on-demand policy ended driver-reclaim deaths across 40 scheduled jobs overnight — $3/night per job against the $18K penalty that motivated it, a 6,000-to-1 insurance ratio.
🎯 Key Takeaway
First node on-demand, enforced by workspace policy: the cheapest permanent immunity in cloud data engineering.

Fleet Policies and Fallback: Degrade Cost, Not Capacity

Databricks fleet options trade savings against reclaim risk on a spectrum. All-spot maximizes discount and maximizes correlated-reclaim exposure (no replacements during crunches). Spot-with-fallback keeps the discount in normal times and buys on-demand automatically when spot vanishes — crunches degrade your bill, not your capacity. Mixed fleets (N on-demand workers + rest spot) add a guaranteed base that survives total spot droughts: the job slows instead of dying.

Match the point on the spectrum to the stakes. Dev, experiments, and retry-friendly ETL: all-spot with fallback (cheap, and failures cost minutes). SLA-bound production: mixed base (4+ on-demand of 24 in our fix) plus fallback — the base holds shuffle-critical mass while spot churns around it. Regulatory/deadline jobs: consider reliability-optimized or all-on-demand during the critical window; the 90% discount isn't worth any penalty clause.

Diversify within spot: multiple instance families and AZs in the fleet spread reclaim correlation (one family spiking doesn't take the whole fleet), and excluding chronic-reclaim families (tracked per family/AZ from your own event logs) removes repeat offenders. Fallback plus diversification is the belt and suspenders: one handles scarcity, the other handles spikes.

📊 Production Insight
A mixed 4-on-demand + 20-spot fleet rode out 3 crunch nights that killed neighboring all-spot clusters — the base held shuffle mass while spot churned, at 11% above all-spot cost instead of 100% outage.
🎯 Key Takeaway
All-spot for dev, spot-with-fallback for ETL, mixed base for SLAs — diversify families/AZs and let crunches degrade cost, never capacity.

Checkpointing: Resume Instead of Repeat

Checkpointing converts driver death from total loss to partial loss: work recorded on durable storage (S3/ADLS, never local disk) survives the cluster. Structured Streaming mandates it — checkpointLocation plus write-ahead logs give exactly-once resume from the last committed micro-batch; without it, a driver reclaim replays from earliest (days of duplication) or fails outright. Set it on durable paths at creation: moving a checkpoint later breaks the resume chain.

Long batch jobs need the same thinking in manual form: stage outputs to Delta per phase (ingest → clean → aggregate as separate committed tables or partitions), so a 5:40 AM death resumes from the last staged commit, not from input. RDD/Dataset checkpointing (rdd.checkpoint() with a durable checkpoint dir) truncates lineage for iterative algorithms — without it, recovery replays the full 5-hour lineage graph instead of one checkpoint read.

Delta's transaction log is the quiet partner: every commit is atomic and versioned, so resumed writers pick up at version N+1 with no partial state — and idempotent sinks (keyed MERGEs, batch-ID-tracked foreachBatch) make the resumed replay converge instead of duplicate. Checkpoint the progress, commit atomically, write idempotently: the trio that makes any restart boring.

resumable_stream.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
from pyspark.sql import functions as F

# Durable checkpoint on S3/ADLS (NEVER local disk): resume, not repeat.
CHECKPOINT = "s3://lake/_chk/orders_stream/"  # set once, never move

stream = (spark.readStream.format("kafka")
            .option("subscribe", "orders").load()
            .select(F.from_json(F.col("value").cast("string"),
                                "order_id STRING, amount DOUBLE").alias("o"))
            .select("o.*"))

def upsert_batch(batch_df, batch_id):
    # Idempotent sink: replays (after driver reclaim) converge.
    batch_df.createOrReplaceTempView("batch")
    batch_df._jdf  # noqa - keep handle; merge below is the real writer
    spark.sql("""
      MERGE INTO lake.orders t USING batch s
      ON t.order_id = s.order_id
      WHEN MATCHED THEN UPDATE SET *
      WHEN NOT MATCHED THEN INSERT *
    """)

(stream.writeStream
 .foreachBatch(upsert_batch)
 .option("checkpointLocation", CHECKPOINT)
 .start().awaitTermination())

# Batch analogue: stage durable Delta outputs per phase so a dead driver
# resumes from the last committed phase, not from raw input.
# spark.read.table("raw").transform(clean).write.format("delta")
#   .mode("overwrite").saveAsTable("lake.clean")  # resume point
📊 Production Insight
After adding per-phase Delta stages, a driver reclaim at hour 4 resumed from the hour-3 commit and finished only 40 minutes late — versus the previous total restart that missed SLA by 4 hours.
🎯 Key Takeaway
Checkpoint streams on durable paths, stage batch phases as Delta commits, write idempotently — restarts then resume from minutes ago, not from zero.

Instance Pools: Replacements in a Minute, Not Ten

Replacement latency decides whether worker reclaims are invisible: tasks retry, but if new nodes take 10 minutes, stages stall, streams lag, and correlated reclaims cascade into timeouts. Instance pools keep pre-warmed (idle) nodes ready, cutting replacement time from ~9 minutes (cold cloud launch + Databricks setup) to ~1 minute (claim idle + join). Our fix's 10-idle pool cost pennies per night against the 9-minute stall it erased.

Size pools from reclaim math, not hope: idle count covering your p95 simultaneous-reclaim burst (2-4 nodes for typical fleets), max capacity above peak fleet size, and the same families/AZs as production fleets so replacements match. Pools also accelerate everyday elasticity (autoscaling grabs warm nodes) and dev cluster startup — one pool serves many clusters, amortizing idle cost across the workspace.

Pre-warm correctly: pool nodes should carry the Databricks runtime version and disk image your jobs use, or the 'warm' node still pays setup time. Tag pool costs separately and review monthly — an oversized pool of idle GPU nodes is its own budget incident. Right-sized pools are the rare optimization that's simultaneously cheaper (less stall waste) and safer (faster recovery).

pool_provision.shBASH
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
# Pre-warmed pool: replacements in ~1 min instead of ~9.
# Create once per family/AZ set; attach all production fleets to it.

# Via Databricks CLI (instance-pools create with your runtime + disk image):
databricks instance-pools create --json '{
  "instance_pool_name": "prod-m5d-pool",
  "min_idle_instances": 4,
  "max_capacity": 60,
  "node_type_id": "m5d.2xlarge",
  "idle_instance_autotermination_minutes": 20,
  "enable_elastic_disk": true,
  "aws_attributes": { "availability": "SPOT_WITH_FALLBACK" }
}'

# Attach fleets to the pool so replacements (and autoscale-ups) claim warm nodes:
# cluster spec -> "instance_pool_id": "<pool-id from create output>"
# Verify: reclaim a dev worker, check cluster events for join latency ~= 1 min.
💡Test reclamation in dev before it tests you in prod
Terminate a worker mid-run in dev quarterly and time the recovery: replacement latency, task retry behavior, streaming lag spike, and resume correctness. Untested checkpoints, untested fallback, and untested pools are wishes — one game-day per quarter converts all three into measured guarantees.
📊 Production Insight
The pool cut replacement latency 9 min → 70 sec on the first real reclaim after install — the streaming lag spike shrank from 40 minutes to 3, turning reclaims from pages into log lines.
🎯 Key Takeaway
Pre-warmed pools make replacements claimable in ~1 minute; size idle counts from p95 reclaim bursts and verify with quarterly termination game-days.

Matching Policy to Stakes: A Spot Playbook

One fleet policy for every workload is how teams get both wasted money and missed SLAs: dev overpays on on-demand while production gambles on all-spot. Split into three lanes. Lane 1 (dev/experiments): all-spot with fallback, pools shared, checkpoints optional — failures cost minutes and teach. Lane 2 (routine production ETL): on-demand driver, spot-majority workers with fallback, durable checkpoints mandatory, pool-attached — the discount with guardrails. Lane 3 (SLA/regulatory): mixed on-demand base sized to shuffle-critical mass, reliability-optimized spot remainder, staged Delta commits, rehearsed recovery — savings where safe, certainty where owed.

Govern the lanes with policy, not tribal knowledge: workspace cluster policies per lane (instance rules, fallback flags, pool attachment, autotermination) so the safe shape is the default shape and the risky shape requires breaking glass. Showback completes it: dashboard spot-share of worker-hours and reclaim counts per lane so finance sees the discount and engineering sees the risk — our fixed setup held ~70% spot worker-hours with zero SLA misses, the number that ended the all-spot debate permanently.

Revisit per quarter or per growth spurt: new instance families shift reclaim rates, new AZs shift correlation, and doubled volumes shift which stages are shuffle-critical. The playbook is a living doc with one owner, one dashboard, and one game-day — boring by design, because spot excitement is always bad news.

lane_policies.shBASH
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
# Three workspace policy lanes: dev cheap, ETL guarded, SLA certain.
# Enforce via cluster policies so the safe shape is the default shape.

# Lane 1 — dev/experiments: all-spot with fallback, shared pool.
databricks clusters create --json '{
  "cluster_name": "dev-spot",
  "num_workers": 4,
  "aws_attributes": { "availability": "SPOT_WITH_FALLBACK" },
  "instance_pool_id": "prod-m5d-pool",
  "autotermination_minutes": 30
}'

# Lane 2 — routine ETL: on-demand driver + spot workers + checkpoints.
# (same payload, plus "first_on_demand": 1 and a checkpointLocation
#  on s3://lake/_chk/ in every streaming query)

# Lane 3 — SLA jobs: on-demand base for shuffle-critical mass.
databricks clusters create --json '{
  "cluster_name": "sla-mixed",
  "num_workers": 24,
  "aws_attributes": { "availability": "SPOT_WITH_FALLBACK",
    "first_on_demand": 5 },
  "instance_pool_id": "prod-m5d-pool"
}'
📊 Production Insight
Three lanes ended two years of religious debate: dev kept 95% spot (fast, cheap, expendable), SLA jobs kept mixed fleets (certain), and finance got one dashboard showing 70% spot-hours with zero misses — everybody's metric green at once.
🎯 Key Takeaway
Three lanes (dev all-spot, ETL guarded-spot, SLA mixed-certain) enforced by workspace policy and reviewed quarterly — savings visible, safety default.
● Production incidentPOST-MORTEMseverity: high

Reclaimed at Hour 5 of 6: The $40 Saving That Cost a Deadline

Symptom
The nightly risk aggregation (6-hour Spark job, 1 driver + 24 workers, all spot) lost 2 workers to reclamation at hour 5. Tasks retried, replacements arrived in 9 minutes, and the job limped on — then the driver node was reclaimed at hour 5:40. The entire application died instantly: 5 hours 40 minutes of compute evaporated, no checkpoint existed to resume from, and the 8 AM regulatory SLA was missed by 4 hours. The $18K late-filing penalty dwarfed the $40 the all-spot fleet had saved over on-demand that night.
Assumption
The platform team assumed Databricks 'handles' reclamation transparently — workers get replaced, tasks retry, jobs continue. That mental model holds for workers (mostly) and fails catastrophically for drivers: nothing retries a dead application master, and nothing recovers 5 hours of un-checkpointed lineage. They also assumed reclaims were rare enough to ignore (2 in 8 months prior) without modeling that one driver reclaim costs infinitely more than a hundred worker reclaims.
Root cause
Three compounding choices: the driver ran on spot (saving ~$3/night while risking the whole run), no fallback policy meant replacements during the crunch took 9 minutes each (spot-only fleet with no spare capacity), and zero durable checkpointing meant every recovery replayed from input (5+ hours of lineage re-execution). The driver reclaim at 5:40 was survivable only in a design that existed nowhere: on-demand driver plus checkpoints would have resumed in minutes; instead the job restarted at 6 AM on the SLA clock's wrong side.
Fix
Driver moved to on-demand permanently (first node on-demand is now a workspace policy nobody can override — ~$3/night insurance against $18K penalties). Fleet policy set to 4 on-demand workers + 20 spot with automatic fallback to on-demand, so crunches degrade cost, not capacity. Idempotent Delta writes plus a mid-pipeline checkpoint to S3 mean any future driver loss resumes from hour ~3 instead of zero. Instance pool (pre-warmed, 10 idle nodes) cut replacement time 9 min → 70 sec. Six months, zero SLA misses, spot still covering ~70% of worker-hours.
Key lesson
  • Price the driver correctly: $3/night of on-demand insurance against an $18K late penalty is the cheapest line item in the budget — all-spot fleets gamble the entire run to save the smallest node.
  • Worker reclaims are routine; driver reclaims are fatal. Design for the fatal case (on-demand driver, durable checkpoints, Delta's resumable log) and the routine case handles itself through retries and pools.
  • Fallback converts capacity crunches from outages into overcharges: spot-to-on-demand replacement kept the fixed job green through 3 subsequent crunch nights at slightly higher cost instead of another missed SLA.
Production debug guideFive checks that turn 'cluster terminated' from a disaster into a footnote.5 entries
Symptom · 01
Cluster died or degraded and the event log mentions reclamation, preemption, or spot termination
→
Fix
Classify the victim first: driver reclaimed means the application is dead — restart the job (and then fix driver placement so it never recurs). Workers reclaimed means tasks retry while replacements arrive — watch the cluster event log for replacement timing; under ~3 minutes with pools is healthy, over ~8 minutes without fallback means the fleet can't find capacity and needs on-demand fallback enabled now.
Symptom · 02
Driver is on spot (or you're not sure where the driver runs)
→
Fix
Move the driver to on-demand today: set the first node / driver node type to on-demand in the cluster policy (Databricks launches the driver on the first node, so first-on-demand guarantees driver safety). Enforce it with a workspace cluster policy that rejects all-spot fleets for scheduled jobs — one policy line ends the entire driver-reclaim failure class permanently.
Symptom · 03
Replacements arrive slowly (8-10+ minutes) or not at all during demand spikes
→
Fix
Enable spot-to-on-demand fallback so the fleet buys on-demand when spot capacity vanishes, and attach an instance pool with a few idle nodes so replacements launch warm instead of cold. Verify by checking replacement latency in the cluster events after the next reclaim — pools typically cut it to 1-2 minutes, which tasks absorb through normal retries.
Symptom · 04
Every reclaim replays hours of work (job restarts from zero or streaming replays from days back)
→
Fix
Checkpoint to durable storage: Structured Streaming needs a checkpointLocation on S3/ADLS (exactly-once resume is impossible without it), long batch jobs should persist mid-pipeline state or write staged Delta outputs per phase, and Delta sinks make replays idempotent so resumed work converges. Test recovery explicitly: terminate a worker mid-run in dev and measure resume time — untested checkpoints are wishes, not guarantees.
Symptom · 05
Spot works for dev but deadline jobs still feel fragile
→
Fix
Match policy to stakes: cost-optimized all-spot (with on-demand driver + fallback) for dev and retry-friendly ETL; reliability-optimized (more on-demand workers, tighter fallback) for SLA-bound production. Track reclaim rate per instance family and AZ — chronic-reclaim families get excluded from production fleets, and the cost dashboard should show spot share of worker-hours so savings stay visible while safety stays default.
Spot survival responses compared
Root CauseHow to ConfirmFixPrevention
Driver on spot reclaimedWhole application dead instantly; event log names driver node; no resume possibleRestart job; move driver to on-demand via first-node policy immediatelyWorkspace policy mandating on-demand first node for all scheduled jobs
Worker reclaims stalling stagesTasks retry while replacements take 8-10+ min; lag spikes; spot-only fleetSpot-to-on-demand fallback plus instance pool with idle nodesPool sized from p95 bursts; fallback on every production fleet
No durable recovery postureRestart replays from zero; streaming replays days; checkpoints local or absentCheckpointLocation on S3/ADLS; staged Delta commits; idempotent sinksRecovery game-day quarterly: terminate a dev worker, measure resume
One policy for all stakesDev overpays while SLA jobs gamble; no lane separation or showbackThree lanes: dev all-spot, ETL guarded, SLA mixed-certain via policyQuarterly review of families, AZs, and spot-share dashboard
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
spot_safe_cluster.pycluster = {Driver on On-Demand
resumable_stream.pyfrom pyspark.sql import functions as FCheckpointing
pool_provision.shdatabricks instance-pools create --json '{Instance Pools
lane_policies.shdatabricks clusters create --json '{Matching Policy to Stakes

Key takeaways

1
Worker reclaims cost time; driver reclaims cost the run
first node on-demand, enforced by policy.
2
Spot-with-fallback degrades bills instead of killing clusters when correlated reclaims hit during scarcity.
3
Checkpoint streams on S3/ADLS, stage batch phases as Delta commits, sink idempotently
resume, never repeat.
4
Instance pools cut replacements ~9 min to ~1 min; size idle from p95 reclaim bursts.
5
Three policy lanes (dev/ETL/SLA) keep savings visible and safety default; review quarterly.
6
Game-day reclamation in dev quarterly converts checkpoints, fallback, and pools into measured guarantees.

Common mistakes to avoid

5 patterns
×

Running the driver on spot to save ~$3/night

Symptom
One driver reclaim kills a 5h40m run instantly with zero resume — $18K penalty against $40 of fleet savings
Fix
First node on-demand, enforced by workspace policy. Spot is for the replaceable many, never the irreplaceable driver.
×

Spot-only fleets with no on-demand fallback

Symptom
Correlated reclaims during crunches find zero spare spot capacity — replacements never arrive at any price
Fix
Enable spot-to-on-demand fallback everywhere in production; crunches should degrade cost, not capacity.
×

Checkpointing to local disk or not at all

Symptom
Restarts replay from zero (batch) or earliest (streams) — hours of recompute plus duplication risk on non-idempotent sinks
Fix
CheckpointLocation on S3/ADLS, staged Delta commits per phase, idempotent MERGE sinks; test resume quarterly.
×

Cold replacements taking 9+ minutes per reclaim

Symptom
Stages stall, streams lag 40 min, correlated reclaims cascade into timeouts while nodes boot
Fix
Instance pool with idle nodes matched to fleet families; verify ~1-minute joins in cluster events.
×

Same fleet policy for dev toys and SLA jobs

Symptom
Dev overpays on on-demand while production gambles all-spot — both lanes wrong simultaneously
Fix
Three policy lanes (dev/ETL/SLA) with showback on spot-hours and reclaims; review families and AZs quarterly.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Why is losing a spot driver fatal while losing a spot worker is routine?
Q02JUNIOR
What does spot-to-on-demand fallback buy you during a capacity crunch?
Q03SENIOR
How do checkpoints and Delta make driver loss survivable?
Q04SENIOR
How do you size an instance pool for reclaim recovery?
Q05SENIOR
Design spot policy for dev, routine ETL, and SLA-bound jobs. How do they...
Q01 of 05JUNIOR

Why is losing a spot driver fatal while losing a spot worker is routine?

ANSWER
The driver holds the application master and all scheduling state — its death kills the application with nothing to resume. Workers just lose tasks, which retry while replacements arrive; their shuffle data recomputes from lineage or checkpoints.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
How much warning does a spot reclaim give?
02
Can Databricks retry through worker reclaims automatically?
03
How much do spot instances actually save?
04
Where must streaming checkpoints live?
05
Do instance pools cost extra?
06
Should SLA-bound jobs use spot at all?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Databricks. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
Databricks Delta Concurrent Append Exception
2 / 2 · Databricks