Home › Data Engineering › Delta Concurrent Append Conflicts — Retry Smart
Intermediate 5 min · September 23, 2026

Delta Concurrent Append Conflicts — Retry Smart

Append rejected? Delta's optimistic concurrency blocks overlapping writes.

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 14 min
  • ✓Delta Lake basics: tables, versions, and time travel
  • ✓Batch and streaming writes (MERGE, append, foreachBatch)
  • ✓Partitioning concepts and basic SQL predicate logic
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Delta Lake guards tables with optimistic concurrency: overlapping commits on the same data get rejected with concurrent-append or commit-conflict errors, not silent corruption
  • Retry idempotently: deterministic job logic plus merge/upsert keys (or streaming transaction IDs) make re-runs safe instead of duplicating rows
  • Cut contention by partitioning so writers touch different partitions, and by separating OPTIMIZE/VACUUM windows from write traffic
  • Pick isolation deliberately: Serializable (default) fails on any overlapping change, WriteSerializable allows compatible appends — then alert on retry rates per table
✦ Definition~90s read
What is Databricks Delta Concurrent Append Exception?

Delta Lake turns cloud object storage into ACID tables with a transaction log: every write is an atomic commit appending a versioned log entry, readers get snapshot isolation (a consistent version, never partial writes), and writers get optimistic concurrency (proceed without locks, validate at commit). The log records which files each version adds and removes, so time travel, auditing, and conflict detection all read the same source of truth.

★
Think of a Delta table as a shared ledger where writers check 'did anyone edit since I looked?' at commit time.

No sidecar database, no lock manager — the log in storage is the concurrency control.

Conflicts arise exactly where the log's guarantees meet overlapping writers. Each commit declares its read version; if a newer version changed overlapping data first, the commit fails rather than silently overwriting. Serializable flags any overlapping data change; WriteSerializable permits concurrent blind appends.

Streaming micro-batches, batch backfills, OPTIMIZE compactions, and VACUUM expirations are all log writers — any pair overlapping in time and data can collide, and all collisions resolve the same way: loser retries on the fresh snapshot.

This design converts the scariest distributed failure (silent lost writes) into the most manageable one (loud, retryable, logged errors). Pipelines that embrace it — idempotent writers, fenced partitions, scheduled maintenance, per-table isolation — get exactly-once semantics on object storage. Pipelines that fight it with blind retries get duplication, lag, and midnight pages.

Plain-English First

Think of a Delta table as a shared ledger where writers check 'did anyone edit since I looked?' at commit time. If someone touched the same pages, your commit is rejected and you retry on the fresh version. That rejection is the concurrent-append error — protection, not corruption. Real ledgers fix it the same way: split the book so clerks rarely collide (partitioning), agree pure additions never conflict (isolation levels), and make entries re-writable without duplication (idempotent retries).

It starts as an occasional flake: a ConcurrentAppendException at 3 AM, a retried Airflow task, green by morning. Then traffic doubles, the streaming sink and the hourly backfill start colliding daily, OPTIMIZE jobs trip over writers, and 'transient Delta errors' become the top on-call category — each retry re-reading the table, each collision wasting a full commit's work. Teams discover that Delta's guarantees are strict in exactly the ways concurrent pipelines violate.

Delta Lake uses optimistic concurrency control: each commit reads the latest table version, does its work, then commits only if no conflicting transaction landed meanwhile. Serializable isolation (the default) rejects any overlapping data change; WriteSerializable permits concurrent appends that don't touch the same files. The errors look scary but they're the system working — the alternative (last-writer-wins on parquet directories) is silent data loss, which is strictly worse.

This article turns conflicts from pages into plumbing. You'll learn what each conflict error actually guards, how to write idempotent jobs that retry safely, how partitioning and scheduling dissolve contention, when to relax isolation (and when never to), and how to monitor conflict rates so growth never surprises you again.

Optimistic Concurrency: The Ledger That Checks at Commit

Delta commits are atomic log appends: each transaction reads the latest version N, computes new files, then appends version N+1 — but only if the table is still at N. If another writer committed N+1 first with conflicting changes, your commit is rejected and you retry against the fresh snapshot. No locks, no waiting, no silent overwrites: contention surfaces as explicit errors instead of lost data. This is optimistic concurrency — assume no conflict, verify at commit, retry on rejection.

Serializable (default) treats any overlapping data change as conflict: concurrent appends to the same partition, MERGE vs overwrite on shared files, OPTIMIZE rewriting files a writer read. WriteSerializable relaxes this for blind appends — writers that add files without reading existing ones (pure inserts, streaming appends) commit alongside each other freely. The distinction is semantic, not just performance: Serializable preserves full snapshot equivalence for mixed workloads; WriteSerializable trades snapshot strictness for append throughput where business logic tolerates it.

Read the error as the guard working: ConcurrentAppendException names the overlapping versions, commit-conflict messages name the files. Both tell you who (operation type), when (versions), and where (partitions) — the full collision report. Log these fields per table and you get contention analytics for free; swallow them into generic retries and you get mystery lag.

collision_forensics.sqlSQL
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
-- Name the contenders: overlapping commits on overlapping partitions.
DESCRIBE HISTORY lake.orders LIMIT 25;

-- Contention analytics: failed vs successful commits per operation.
-- (Wrap writers to log attempts; history shows commits that landed.)
SELECT version, timestamp, operation,
       operationParameters['predicate'] AS predicate,
       operationMetrics['numFiles'] AS files
FROM (DESCRIBE HISTORY lake.orders)
ORDER BY version DESC
LIMIT 25;

-- Per-table isolation: strict money paths stay Serializable...
ALTER TABLE lake.orders SET TBLPROPERTIES (
  'delta.isolationLevel' = 'Serializable'
);
-- ...while pure-append sidecars allow compatible concurrent appends.
ALTER TABLE lake.order_events SET TBLPROPERTIES (
  'delta.isolationLevel' = 'WriteSerializable'
);
📊 Production Insight
Teams that log contender operation-types from history resolve new conflict pairs in minutes (the pattern names the fix); teams with generic retries average days because every collision looks like the last flake.
🎯 Key Takeaway
Optimistic concurrency rejects at commit instead of corrupting silently — read contender, version, and partition from history, and set isolation per table semantics.

Idempotent Writers: Retries That Converge

A retry is only safe if re-running can't duplicate: that's idempotency, and non-idempotent retries caused our incident's $410K double-apply. The failure mode is 'committed but unacknowledged' — the commit lands, the acknowledgment dies in the conflict, the orchestrator reports failure, and the retry applies everything again. Blind appends and full overwrites re-applied are duplication machines; keyed MERGEs re-applied converge to the same rows.

Build idempotency per writer shape. Batch corrections: MERGE on business keys (order_id + updated_at) so re-runs update the same rows to the same values. Streaming sinks: foreachBatch with transaction tracking (streaming query ID + batch ID recorded in the target, or Delta's built-in exactly-once append handling) so replays skip committed batches. Partition overwrites: dynamic overwrite of the exact same partition predicate — re-running writes identical files, and Delta's log makes the second commit a no-op equivalent.

Prove it, don't assert it: the idempotency test runs every writer twice against fixed input and diffs row counts plus checksums. Two runs, identical table — merge that test and ambiguous failures become boring. Without it, every Airflow retry bump is a loaded duplication gun pointed at the money table.

idempotent_merge.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
from delta.tables import DeltaTable

# Keyed MERGE: re-runs converge to identical rows — retries can't duplicate.
orders = DeltaTable.forName(spark, "lake.orders")
corrections = spark.read.parquet("s3://lake/corrections/dt=2026-09-22/")

(orders.alias("t").merge(
    corrections.alias("s"),
    "t.order_id = s.order_id AND t.updated_at <= s.updated_at")
 .whenMatchedUpdateAll()
 .whenNotMatchedInsertAll()
 .execute())

# Idempotency proof: run twice against fixed input, diff counts + sums.
# before = spark.table("lake.orders").agg({"order_id": "count"}).collect()
# ... re-run the merge above ...
# after = spark.table("lake.orders").agg({"order_id": "count"}).collect()
# assert before == after, "Writer is not idempotent — do NOT auto-retry"

# Streaming exactly-once: checkpoint + foreachBatch with batchId tracking.
# (spark.writeStream.foreachBatch(write_with_batch_id)
#   .option("checkpointLocation", "s3://lake/_chk/orders/").start())
📊 Production Insight
The run-twice assertion caught 3 non-idempotent writers in review the quarter after the duplication — including one 'harmless' backfill whose retry would have triple-applied fee adjustments.
🎯 Key Takeaway
Key every writer (merge keys, batch IDs, exact predicates) and prove it with run-twice tests — safe retries are engineered, never assumed.

Partition Fencing: Writers That Never Meet

The best conflict is the one structurally impossible: partition fencing gives each writer exclusive territory so snapshots never overlap. Streaming owns the fresh tail (last 6 hours), backfills own settled history (older than 6 hours), enforced by write predicates both jobs share from one config — not by schedule luck. Overlapping predicates are the collision; disjoint predicates are the cure, regardless of timing.

Granularity sets fence strength. Date-partitioned tables fence by day (coarse but simple); hourly sub-partitioning fences tighter for hot tails. When two writers must touch one partition (late-arriving corrections to today's data), fence by key range within it instead — corrections MERGE on order_id ranges the streaming append never rewrites. Fences compose with isolation: fenced Serializable writers never meet, so they pay no conflict cost while keeping full guarantees.

Maintenance gets fenced too: OPTIMIZE and VACUUM rewrite and expire files, making them contenders like any writer. Run them in writer-free windows (4 AM, not midnight alongside backfills), scope OPTIMIZE to settled partitions (WHERE dt < today), and never let VACUUM's retention window (default 7 days) approach streaming lateness — a long stall plus aggressive retention deletes files a restarting reader still needs, converting conflicts into data loss.

fenced_writers.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
# One shared fence config: disjoint predicates per writer, enforced in code.
FRESH_HOURS = 6  # streaming owns the tail; backfill owns settled history

stream_pred = f"event_time >= current_timestamp() - INTERVAL {FRESH_HOURS} HOURS"
backfill_pred = f"event_time < current_timestamp() - INTERVAL {FRESH_HOURS} HOURS"

# Streaming micro-batch: hard predicate keeps it inside its territory.
(spark.readStream.table("lake.staging")
 .where(stream_pred).writeStream
 .foreachBatch(lambda df, bid: df.write.format("delta")
               .mode("append").saveAsTable("lake.orders"))
 .option("checkpointLocation", "s3://lake/_chk/orders/").start())

# Backfill: keyed MERGE fenced to settled partitions — retries converge.
# MERGE INTO lake.orders t USING corrections s
#   ON t.order_id = s.order_id AND s.event_time < now() - INTERVAL 6 HOURS
# WHEN MATCHED THEN UPDATE SET * WHEN NOT MATCHED THEN INSERT *
📊 Production Insight
Fencing alone took the incident from 47 conflicts a night to zero for 6 weeks — no isolation change, no retry change, just disjoint predicates plus a 4 AM maintenance window.
🎯 Key Takeaway
Disjoint write predicates per writer (plus writer-free maintenance windows) make conflicts structurally impossible — fence first, tune second.

Isolation Levels: Serializable vs WriteSerializable

Choose isolation per table semantics, never cluster-wide. Serializable (default): any overlapping committed change aborts yours — required wherever updates, deletes, or MERGEs mix with other writers (money tables, dimension rebuilds), because snapshot equivalence is the correctness property. WriteSerializable: concurrent blind appends commit freely, conflicts only on file-level modification overlap — right for append-only event sidecars and high-throughput ingest where writers never update each other's rows.

The cost of wrong choice is asymmetric. Serializable on pure-append ingest buys conflicts you didn't need (throughput collapses under fan-in). WriteSerializable on mixed workloads buys anomalies you can't see (lost updates that pass every row-count check). When unsure, keep Serializable and fence harder — its errors are loud and retryable, while relaxed anomalies are silent and forensic.

Snapshot isolation (Databricks: the Serializable implementation detail) adds one more subtlety: readers see consistent snapshots regardless, so long-running readers never block writers — but writers with stale snapshots fail more as table churn rises. Shorten batch windows and commit granularly (per-partition commits over whole-table rewrites) to keep snapshots fresh: the smaller your commit's footprint, the less surface conflicts can strike.

IsolationChoice.scalaSCALA
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
// Per-table isolation, documented next to the table it protects.
// Money path: mixed MERGE + backfill writers -> Serializable stays.
spark.sql("""ALTER TABLE lake.orders SET TBLPROPERTIES (""" +
  """'delta.isolationLevel' = 'Serializable')""")

// Append-only sidecar: pure inserts, no updates -> WriteSerializable.
spark.sql("""ALTER TABLE lake.order_events SET TBLPROPERTIES (""" +
  """'delta.isolationLevel' = 'WriteSerializable')""")

// Keep commit footprints small: per-partition MERGE beats whole-table
// overwrite — less snapshot surface, fewer invalidations, faster retries.
// MERGE INTO lake.orders t USING corrections s
//   ON t.order_id = s.order_id AND t.dt = s.dt
// WHEN MATCHED THEN UPDATE SET *
// WHEN NOT MATCHED THEN INSERT *
⚠ Never relax isolation to fix a fencing problem
WriteSerializable on writers that update shared rows trades loud retryable conflicts for silent lost updates — anomalies no row-count check catches. Fence first (disjoint predicates, maintenance windows), keep Serializable on mixed tables, and relax only pure-append paths where blind inserts are the entire workload.
📊 Production Insight
An audit-sidecar team relaxed isolation correctly (pure appends, 12x ingest fan-in) and cut conflicts to zero; a sibling team copied the setting onto a MERGE-heavy table and lost 3 weeks of corrections silently — same knob, opposite semantics.
🎯 Key Takeaway
Serializable for mixed writers (loud conflicts beat silent anomalies), WriteSerializable for pure appends only — and small commit footprints beat both settings combined.

Scheduling and Maintenance Windows That Don't Collide

Conflicts are timetable failures as often as semantic ones: the midnight stack (streaming + backfill + OPTIMIZE + VACUUM) collides because everything runs when engineers sleep, not when tables are free. Draw the contention calendar per table — every writer and maintenance job with its partitions and duration — and the overlaps glow. Stagger heavy commits (backfill at :15, never :00 with the streamer), move OPTIMIZE to 4 AM writer-free windows, and run VACUUM weekly against 7-day retention with readers' max lateness modeled in.

Streaming deserves special scheduling respect because it never stops: its micro-batch commits (every 2 minutes in our incident) make any overlapping batch writer a continuous collision source. Either fence the streamer off the batch's partitions entirely or pause-and-catch-up around bounded backfills (stop stream, run backfill, resume — lag spikes once instead of conflicting for hours). The math favors the pause: one 20-minute catch-up beats 47 conflicted retries by orders of magnitude.

Orchestrator config finishes the job: retries with exponential backoff plus jitter (thundering herds re-collide deterministically), conflict-aware alerting (page on sustained conflict rate, not single flakes), and SLA-budgeted backfill windows that shrink as volume grows — the 6-hour window that 'worked' at half volume is the midnight page at full volume.

📊 Production Insight
Staggering two jobs by 15 minutes plus moving OPTIMIZE to 4 AM removed 90% of one fleet's conflicts before any code changed — timetables, not settings, were the contention.
🎯 Key Takeaway
Draw per-table contention calendars, stagger heavy commits, fence the never-stopping streamer, and retry with backoff plus jitter — collisions are scheduled before they're coded.

Monitoring Conflict Health as You Scale

Conflict rate per table is the leading indicator volume growth shows up in first — weeks before lag pages. Derive it from the transaction log: commits landed (DESCRIBE HISTORY) versus attempts made (writer-side attempt logging), per operation type, per day. A table drifting 0 → 5 → 20 daily conflicts is 2 quarters from a storm; the fix (re-fence, split tables, graduate to WriteSerializable append paths) is cheap at 5 and structural at 47.

Dashboard the full contention surface, not just errors: streaming lag p99, backfill duration trend, OPTIMIZE overlap minutes with writers, and retry-count histograms per job. Retries hide conflicts from task-green dashboards while lag and cost grow — the incident's Airflow showed green for weeks while streaming lag climbed 2 → 40 minutes. Alert on the lag and the rate, never on single flakes (one conflict is the system working).

Review on onboarding, not on fire: every new writer gets a contention review (partitions touched, isolation of target, idempotency proof, schedule vs existing calendar) before its first production commit. Growth re-collides fenced writers within 2-3 quarters as volumes shift predicates and schedules drift — the review is the fence maintenance, and the conflict-rate graph proves it holds.

📊 Production Insight
A 0→5-conflicts-per-day drift caught in a quarterly review prompted one re-fencing afternoon; the unfenced sibling table hit 47/night and needed a 3-week project — same growth, 20x cost difference from one graph.
🎯 Key Takeaway
Track conflict rate per table as the growth leading indicator, dashboard lag over task-green, and review every new writer before its first commit.
● Production incidentPOST-MORTEMseverity: high

The Midnight Collision: Streaming Sink vs Backfill, 47 Conflicts a Night

Symptom
The orders Delta table (11 TB, partitioned by dt) served a Structured Streaming upsert sink (2-min micro-batches) plus an hourly idempotency backfill that rewrote the last 6 hours. As volume doubled in Q3, commit conflicts rose 2 → 47 per night: each conflict aborted a writer that had already done minutes of work, streaming lag grew 2 min → 40 min by morning, and the backfill window stretched 45 min → 6 hours. Then a manual retry of a 'failed' backfill double-applied $410K of adjustments — the retry wasn't idempotent, and the first attempt had actually committed.
Assumption
The team treated conflicts as infra flakes: bump Airflow retries 3→10, widen the backfill window, and page less. Retries helped the dashboard (tasks eventually green) while hurting the system (every retry re-did full partition rewrites, raising collision odds). The duplication was blamed on 'Airflow double-trigger' for a week — nobody checked whether the first attempt had committed, because the error text said 'failed' and failed means didn't-happen, right?
Root cause
Both writers touched the same 4 recent partitions with Serializable isolation: the streaming MERGE rewrote files the backfill's overwrite was also rewriting, so one's commit always invalidated the other's snapshot. Optimistic concurrency correctly rejected the loser — 47 times a night at peak. The duplication came from non-idempotent retry: the backfill's overwrite wasn't keyed, so re-running a half-ambiguous failure applied adjustments twice. OPTIMIZE on the same partitions at midnight added a third contender, and VACUUM's default 7-day retention window nearly uncut the streaming reader's time travel during one long stall.
Fix
Contention dissolved structurally: the backfill switched from partition overwrite to keyed MERGE (order_id + updated_at), so retries became idempotent and concurrent appends stopped invalidating each other; streaming and backfill were partition-fenced (stream owns <6h old, backfill owns older, enforced by predicates); OPTIMIZE moved to a 4 AM window with no writers; isolation dropped to WriteSerializable for the append-only audit sidecar table only. Conflicts fell 47 → 0-1 per night, streaming lag capped at 3 min, and the backfill holds 38 min. Conflict-rate alerting per table pages before lag does.
Key lesson
  • Retries without idempotency convert conflicts into duplication: the $410K double-apply happened because 'failed' can mean 'committed but unacknowledged'. Key every writer (merge keys, transaction IDs) so re-runs converge instead of multiply.
  • Bumping retries treats the dashboard while feeding the disease — each retry re-does full rewrites and raises collision odds. Dissolve contention structurally (partition fencing, keyed merges, separated maintenance windows) instead of out-retrying it.
  • Serializable is a default, not a law: audit-style append-only tables run fine (and faster) on WriteSerializable, while money tables keep Serializable plus fencing. Choose per table, document why, and alert on conflict rates to catch drift.
Production debug guideIdentify the contenders, make retries safe, then separate them — in that order.5 entries
Symptom · 01
ConcurrentAppendException or commit-conflict errors mentioning overlapping files or versions
→
Fix
Name the contenders from the table history: DESCRIBE HISTORY shows every commit's timestamp, operation (WRITE, MERGE, OPTIMIZE), and partitions touched — overlapping timestamps on overlapping partitions are your collision pair. Note each contender's operation type and isolation level; a MERGE vs overwrite on the same partition is the classic Serializable rejection, and OPTIMIZE in the mix makes three.
Symptom · 02
Two writers provably touching the same partitions at the same time
→
Fix
Fence them by partition and time before changing any setting: give streaming the fresh tail (predicates on recent hours) and backfills the settled past, move OPTIMIZE/VACUUM to writer-free windows, and stagger schedules so heavy commits don't align. Verify with history that commit timestamps stop overlapping — most conflict storms end here with zero config changes.
Symptom · 03
Retries exist but you're unsure they're safe (duplication risk on ambiguous failures)
→
Fix
Make every writer idempotent: keyed MERGE on business keys (order_id + updated_at) instead of blind overwrites/appends for correction paths, deterministic output for batch jobs (same input → same rows), and streaming transaction IDs (appId + version tracking) for exactly-once sinks. Then test it: run the writer twice against the same input and assert identical row counts — a retry that can't duplicate is a retry you can automate.
Symptom · 04
Contenders fenced and idempotent, but append-only tables still conflict under load
→
Fix
Consider WriteSerializable for pure-append tables: it permits concurrent appends that don't modify the same files, which is exactly the high-throughput ingest shape. Keep Serializable on tables with updates/deletes/merges (money paths), set isolation per table (not cluster-wide), and document the choice next to the table definition. Never relax isolation to fix a fencing problem — separation first, settings second.
Symptom · 05
Conflicts rare today but volume grows every quarter
→
Fix
Monitor conflict rate per table (failed-commit count from history over time) with pages at sustained elevation, not just errors — retries hide conflicts from task dashboards while lag and cost grow. Track streaming lag, backfill duration, and OPTIMIZE overlap as leading indicators, and re-fence partitions whenever a new writer onboards. Growth re-collides fenced writers within 2-3 quarters without this.
Delta conflict responses compared
Root CauseHow to ConfirmFixPrevention
Overlapping writers, same partitionsHistory shows concurrent commits touching shared partitions; lag grows with volumePartition fencing with disjoint predicates; stagger heavy commit timesContention calendar per table; review every new writer pre-launch
Non-idempotent retries duplicatingRe-run applies twice; counts differ run-to-run on fixed inputKeyed MERGE, batch-ID tracking, exact predicates; run-twice proof testIdempotency test mandatory in CI for every writer
Maintenance colliding with writersOPTIMIZE/VACUUM commits interleave with writes in history; midnight overlapWriter-free maintenance windows; scope OPTIMIZE to settled partitionsScheduled windows documented; retention modeled against reader lateness
Over-strict isolation for append shapePure-append writers conflicting under fan-in despite disjoint dataWriteSerializable on append-only tables; Serializable stays on mixed tablesIsolation choice documented per table; conflict-rate alerts catch drift
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
collision_forensics.sqlDESCRIBE HISTORY lake.orders LIMIT 25;Optimistic Concurrency
idempotent_merge.pyfrom delta.tables import DeltaTableIdempotent Writers
fenced_writers.pyFRESH_HOURS = 6 # streaming owns the tail; backfill owns settled historyPartition Fencing
IsolationChoice.scalaspark.sql("""ALTER TABLE lake.orders SET TBLPROPERTIES (""" +Isolation Levels

Key takeaways

1
Optimistic concurrency rejects stale snapshots at commit
read contender, version, and partition from history.
2
Fence writers with disjoint predicates and writer-free maintenance windows before touching any setting.
3
Key every writer and prove run-twice idempotency
'failed' can mean committed-but-unacknowledged.
4
Serializable for mixed tables, WriteSerializable for pure appends
never relax to fix a fencing gap.
5
Stagger heavy commits, pause-and-catch-up around bounded backfills, retry with backoff plus jitter.
6
Track conflict rate per table as the growth leading indicator; review every new writer pre-launch.

Common mistakes to avoid

5 patterns
×

Bumping orchestrator retries instead of fencing writers

Symptom
Tasks eventually green while lag and cost climb — every retry re-does full rewrites and raises the next collision's odds
Fix
Fence by partition and time first; reserve retries for proven-transient single flakes with backoff plus jitter.
×

Retrying non-idempotent writers on ambiguous failures

Symptom
'Failed' attempt actually committed, retry double-applies — $410K duplicated adjustments discovered a week later
Fix
Keyed MERGEs, batch-ID tracking, exact predicates; run-twice proof tests before any auto-retry.
×

Running OPTIMIZE/VACUUM alongside writers at midnight

Symptom
Three-way collisions on hot partitions plus near-miss retention deletes against stalled readers
Fix
Writer-free maintenance windows (4 AM), OPTIMIZE scoped to settled partitions, retention modeled on max lateness.
×

Relaxing isolation to fix overlapping writers

Symptom
Conflicts hush while lost updates accumulate silently on mixed tables — anomalies row-counts never catch
Fix
Fence first; WriteSerializable for pure appends only; Serializable stays wherever updates mix.
×

Alerting on task failure instead of conflict rate and lag

Symptom
Green dashboards for weeks while conflicts climb 2→47/night and streaming lag 2→40 min — pages arrive at the storm, not the drift
Fix
Page on sustained conflict rate per table plus lag p99; single conflicts are the system working, trends are the warning.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
What does a ConcurrentAppendException actually protect against?
Q02SENIOR
How do you make a Delta writer safe to retry?
Q03SENIOR
When is WriteSerializable appropriate, and when is it dangerous?
Q04SENIOR
A streaming sink and hourly backfill collide nightly. How do you separat...
Q05SENIOR
What metrics warn of conflict storms quarters early?
Q01 of 05JUNIOR

What does a ConcurrentAppendException actually protect against?

ANSWER
Silent overwrites: optimistic concurrency rejects commits whose snapshot went stale on overlapping data. The error is the guard working — last-writer-wins on raw files would lose data quietly instead.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
What causes ConcurrentAppendException in Delta Lake?
02
Should I just increase retries for Delta conflicts?
03
What's the difference between Serializable and WriteSerializable?
04
Can OPTIMIZE cause commit conflicts?
05
How do streaming sinks achieve exactly-once into Delta?
06
Is VACUUM dangerous for concurrent readers?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Written from production experience, not tutorials.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Databricks. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
Spark Small Files Problem Destroys Read Performance
1 / 2 · Databricks
Next
Databricks Cluster Terminated: Spot Instance Reclaimed
→