Home › Observability › Prometheus Out-of-Order Samples: Fix Dropped Writes
Intermediate 6 min · September 23, 2026

Prometheus Out-of-Order Samples: Fix Dropped Writes

Prometheus drops late samples with out-of-order errors.

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 12 min
  • ✓A running Prometheus server you can query and restart
  • ✓Basic PromQL for checking up and rate() queries
  • ✓Access to exporter hosts for clock checks
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Prometheus rejects a sample older than the newest stored sample for that series, logging out-of-order or too-old errors with a dropped count
  • Set storage.tsdb.out_of_order_time_window to 10m so samples up to 10 minutes late still ingest instead of dropping
  • Give each HA writer a unique replica external label and dedupe at the receiver so two servers stop racing on one series
  • Fix exporter clock drift with chrony or NTP, since skew past a few seconds guarantees late samples on every scrape
✦ Definition~90s read
What is Prometheus Error on Ingesting Out-of-Order Samples?

Prometheus stores every metric as a time-ordered series in its TSDB head, an in-memory structure optimized for append-only writes. Each new sample for a series must carry a timestamp newer than the last one stored; anything older is out-of-order and gets rejected.

★
Think of each metric as a diary with numbered pages in order.

When you've configured storage.tsdb.out_of_order_time_window above zero, samples within the window behind the head's maximum timestamp are still accepted, and anything older fails with a too-old error instead.

The window setting, available since Prometheus 2.39, lives under the storage key in prometheus.yml with a default of 0s (strict ordering). A sample is ingestible when its timestamp is at or after head-max-time minus the window. This covers honest lateness — slow scrapes, brief network delays, slightly skewed clocks — without paying the cost of fully random inserts into sealed chunks.

Three failure modes produce these errors in practice. Duplicate writers send two interleaved timestamp streams for one series, so every other sample loses the race. Clock skew stamps samples in the past on every scrape from the drifting host. Remote-write retry queues replay stale batches after receiver outages, colliding with live samples.

All three look identical in dashboards (gappy graphs) but need completely different fixes, which is why naming the path and measuring lateness come before any config change.

Plain-English First

Think of each metric as a diary with numbered pages in order. Prometheus only lets you write on the next blank page — you can't go back and squeeze a sentence onto page 12 once page 15 is full. When two people share one diary, or someone's watch runs slow, entries arrive for pages that are already full, and Prometheus throws them away. The fix is giving each writer their own diary, fixing the slow watch, and allowing a small grace period for slightly late entries.

Your dashboards developed gaps overnight, but every target shows up on the Targets page. The server log tells a different story: thousands of lines reading Error on ingesting out-of-order samples with a num_dropped count climbing every minute. Nothing was deployed, no config changed, and the metrics simply stopped landing.

Prometheus stores each series as a time-ordered append log, and that design choice is the whole story. A sample whose timestamp is older than the newest sample already stored for that series can't be appended cheaply, so the TSDB rejects it. The usual triggers are painfully ordinary: two servers writing the same series, a VM clock that drifted 40 seconds, or a remote-write queue replaying yesterday's backlog after an outage.

This guide shows you how to confirm which trigger you've got, fix it without guessing, and size the out_of_order_time_window setting from real measurements. You'll leave with queries that quantify your sample lateness, configs that stop duplicate writes, and alerts that catch clock skew before it eats your data.

How the TSDB Append Path Rejects Late Samples

Every sample Prometheus stores travels one path: scrape, append to the in-memory head, then cut into two-hour blocks. The head keeps each series as an append-only sequence, so a sample with a timestamp older than the last one stored has nowhere cheap to go. The server rejects it and logs which path failed: Error on ingesting out-of-order samples for scrapes, Out of order sample from remote write for remote-write receivers, or too old sample when you've configured a window and the sample is older than head-max minus that window.

This strictness is a performance bargain, not a bug. Ordered appends let the head compress samples into tiny chunks at millions of samples per second. Random inserts into sealed chunks would cost far more CPU and memory for a case that's usually caused by broken clocks or duplicate writers. The window added in 2.39 keeps the bargain while tolerating real-world lateness: slightly late samples buffer in the head, badly late ones still drop.

Your first job in any incident is naming the path. Scrape-path errors name a scrape_pool and target, so you investigate exporters and scrape config. Remote-write errors name a series and timestamp, so you investigate senders, queues, and receiver topology. Mixing the two up wastes hours because the fixes live on opposite ends of the pipeline.

PROMQL
1
2
3
4
5
6
7
8
9
10
# Error on ingesting out-of-order samples (scrape path, Prometheus 2.54+)
# {caller=scrape.go, component="scrape manager", scrape_pool="cadvisor"}
# Out of order sample from remote write (remote-write path)
# too old sample (window configured, sample older than head-max minus window)

# Count drops per minute by job on the receiving server
sum by (job) (rate(prometheus_remote_storage_samples_dropped_total[5m]) * 60)

# Retry volume that usually precedes a post-outage drop spike
sum by (remote_name) (rate(prometheus_remote_storage_queue_retry_samples_total[5m]) * 60)
📊 Production Insight
A team chased exporter clocks for two days before noticing their errors said from remote write. The exporters were fine — an HA sender pair was racing. Reading the error prefix first would have cut the incident to an hour.
🎯 Key Takeaway
Name the failing path first: scrape-path errors point at exporters, remote-write errors point at senders and receiver topology.

Configuring out_of_order_time_window Without Regret

The window lives under the storage key in prometheus.yml, not under global or scrape configs. Setting storage.tsdb.out_of_order_time_window to 10m tells the head to accept samples with timestamps back to head-max minus ten minutes. The default of 0s keeps strict ordering, which is right for fleets with healthy clocks and one writer per series.

Pick the value from your lateness data, not from a blog post. If 99% of your late samples land within 8 minutes, 10m covers them with margin. Values above 30m need justification because the cost lands on every series in the head, and values in hours are almost always masking duplicates or retries that deserve a real fix.

Roll it out like any storage change: apply to staging first, reload with kill -HUP, and watch the dropped-sample rate plus head memory for a full compaction cycle. Keep the old value in a comment so the next engineer knows what changed and why.

Reloading applies the window without a restart: send SIGHUP or POST to /-/reload and the head keeps serving queries throughout. The new value governs samples appended after reload; in-flight scrapes finish under the old rule, so watch the dropped counter across two full scrape intervals before judging. Pair the window with retention thinking: a 10m window on 15d retention changes head behavior only, while disk blocks stay untouched. Record the chosen value beside retention.time in your runbook with the p99 lateness number that justified it, so the next capacity review inherits evidence instead of folklore.

prometheus.ymlYAML
1
2
3
4
5
6
7
8
9
10
11
12
13
# prometheus.yml — storage window sized from measured lateness
global:
  scrape_interval: 15s
  evaluation_interval: 15s

storage:
  tsdb:
    out_of_order_time_window: 10m
    retention.time: 15d

# Reload without restart after editing:
# kill -HUP $(pidof prometheus)
# Confirm with: prometheus --version and the config reload log line
📊 Production Insight
One fleet set the window to 2h on a 3-million-series server and added 9 GB of head memory overnight. Their lateness data later showed 99% of late samples arrived within 6 minutes — 10m would have fixed the drops at a fraction of the cost.
🎯 Key Takeaway
Put the window under storage.tsdb, size it from measured lateness, and watch head memory for a full compaction cycle.

Duplicate Writers: HA Pairs and Retry Replays

Duplicate writers are the most common cause and the least responsive to window tuning. An HA pair scraping the same targets with identical external labels sends two timestamp streams for every series. Each stream is individually ordered, but interleaved at the receiver they're a coin flip — the second arrival per timestamp loses and drops as out-of-order.

The fix is topological. Give each writer a unique external label such as replica: a versus replica: b, then dedupe at the receiver layer with Thanos Receive or Mimir's dedupe. The dedupe keeps one ordered stream per series and the drops vanish regardless of window size. Remote-write retries create the same pattern after an outage: the queue replays stale batches that collide with live samples, so bound the retry queue and backfill history through a proper backfill path instead.

Confirm duplicates before you reconfigure. If count by (job, instance) (up) shows two up series per expected target, or drops flip between senders after failover, you've found it. Window changes can't win a race between two clocks — only one writer per series can.

Beyond HA pairs, the same race hides in overlapping scrape jobs: two jobs targeting one exporter with different intervals interleave timestamps on shared series. Audit with count by (job, instance) (up) and merge duplicates into a single job per target set. One writer per series is a topology invariant, not a tuning knob.

prometheus-ha.ymlYAML
1
2
3
4
5
6
7
8
9
10
11
12
13
# Give each HA writer a unique identity so the receiver can dedupe
global:
  external_labels:
    replica: a        # second server uses replica: b
    region: us-east-1

remote_write:
  - url: "https://receive.example.com/api/v1/receive"
    queue_config:
      capacity: 10000
      max_shards: 50
      max_samples_per_send: 2000
      batch_send_deadline: 5s
📊 Production Insight
After adding replica labels and Thanos dedupe, a billing pipeline went from 12,000 drops per minute to zero in two scrape intervals — with the window actually reduced from 30m back to 10m.
🎯 Key Takeaway
Two writers on one series always race; unique replica labels plus receiver dedupe end the race permanently.

Clock Skew: the Drift That Outgrows Every Window

Exporter clocks drift more than teams expect. VMs slip 30 to 90 seconds after live migration, containers inherit host drift, and edge devices with dead NTP daemons wander minutes per week. Every scrape from a skewed host stamps samples in the past, and no window survives skew that grows without bound.

Diagnose with two commands: compare date +%s on the server and exporter in the same minute, and query node_timex_offset_seconds where the node exporter runs. Offsets above 5 seconds deserve a clock fix, not a config change. Standardize on chrony, point every host at the same NTP tier, and alert on offset so drift pages someone before samples start dropping.

The honor_timestamps setting deserves caution here. Flipping it to false stamps arrival time instead of exporter time, which silences errors while quietly shifting your data in time. Keep the default true, fix the clock, and your rate() windows and alert evaluations stay honest.

Standardize time sync with chrony on every host: two or three stratum sources, makestep for fast correction after boot, and maxslewrate to bound drift between polls. Virtual machines should use the hypervisor clocksource (kvm-clock on KVM, xen on Xen) so live migrations do not inject jumps. Containers inherit the host clock, which makes host NTP a fleet-wide dependency worth monitoring. Alert with node_timex_offset_seconds above 5 seconds plus a max-error variant for systemd-timesyncd hosts. When an alert fires, fix the daemon and verify with chronyc tracking before touching any Prometheus setting.

📊 Production Insight
A fleet widened its window three times in a month chasing drops that moved between hosts. The real cause was a dead NTP daemon on one hypervisor; 40 VMs drifted 70 seconds. One chrony fix ended all three incidents at once.
🎯 Key Takeaway
Skew above 5 seconds is a clock incident, not a lateness budget — fix NTP and alert on offset.

Measuring Lateness So the Window Fits the Data

Guessing the window size is how fleets end up at 2h with 9 GB of extra head memory. Measure instead. The dropped-sample and retry counters tell you volume, and timestamp comparisons tell you distribution: how late are the late samples, and does the 99th percentile sit at 6 minutes or 60?

Build a small dashboard before changing config: drops per minute by job, retry volume by remote name, head series count, and resident memory. Record a week of baseline so you can attribute every change. Then set the window just above your measured p99 lateness and verify the drops hit zero within two scrape intervals.

Keep measuring after the fix. Lateness drifts as fleets grow, new regions add network delay, and receivers get slower. A quarterly glance at the lateness panel tells you whether 10m still fits or needs a nudge — a two-minute check that prevents the next gap incident.

A practical lateness proxy needs no new instrumentation: max by (job) (time() - timestamp(up)) measures seconds since each target's latest sample landed. Graph it per job for a week; the p99 of that series is your honest lateness budget, including scrape, network, and remote-write delay. Set the window just above that p99 — 10m covers fleets measuring 6 to 8 minutes — and re-check quarterly as regions and receivers change. Alert when the proxy exceeds the window: it means lateness outgrew the budget and drops are imminent. This single panel turns window sizing from guesswork into arithmetic, and it doubles as an early-warning gauge long after the incident closes.

PROMQL
1
2
3
4
5
6
7
8
9
10
# Drops per minute by job — must fall to zero after the fix
sum by (job) (rate(prometheus_remote_storage_samples_dropped_total[5m]) * 60)

# Head pressure — compare before and after any window change
prometheus_tsdb_head_series
process_resident_memory_bytes

# Alert: page when drops persist for 10 minutes
sum by (job) (rate(prometheus_remote_storage_samples_dropped_total[5m])) > 0
for: 10m
📊 Production Insight
A team that measured first found p99 lateness of 6 minutes on a 2-million-series fleet. The 10m window killed 100% of drops while head memory rose only 400 MB — versus 9 GB when they'd previously guessed 2h.
🎯 Key Takeaway
Dashboard drops, retries, head series, and memory first; set the window just above measured p99 lateness.

When Not to Widen the Window

A bigger window isn't free. Late inserts fragment head chunks, raise resident memory, and slow the compactions that cut two-hour blocks. On million-series servers the cost is gigabytes, and it persists 24/7 — you pay it during quiet nights for lateness that happens at peak.

There are also cases the window can't help at all. Samples older than head-max minus the window still drop as too-old, so multi-hour backlogs need a backfill path, not a bigger number. Federation and recording rules that re-stamp timestamps can manufacture disorder no window should excuse — fix the rule instead.

Treat 30m as a smell threshold. If your measurements demand more, something upstream is broken: a receiver outage replaying history, a region with chronic delay, or duplicates masquerading as lateness. Fix that layer and the window usually shrinks back to 10m with drops still at zero.

The memory math explains the smell threshold. Each out-of-order sample forces the head to hold overlapping chunk ranges instead of one clean append stream, and the overhead lands on every active series whether it receives late samples or not. Past 30m the cost typically exceeds a gigabyte per million series while buying tolerance for delays that indicate broken pipelines. Genuine exceptions exist: cross-region federation with known multi-minute lag, or edge buffers that flush hourly by design. For those, isolate the late streams in a dedicated Prometheus with a wider window instead of taxing the main fleet. Keep the primary window tight, quarantine the exotic lateness, and both systems stay cheap.

⚠ Don't Buy Time With a Bigger Window
Never widen the window past 30m to cover clock skew or duplicate writers. The memory and compaction cost hits every series on the server, while the underlying race keeps dropping samples anyway. Fix writers and clocks first, then size the window from data.
📊 Production Insight
An edge fleet demanded a 4h window until someone traced the lateness to a single congested NAT gateway. Moving remote-write to a regional receiver cut p99 lateness from 3h to 4 minutes, and the window dropped to 10m with memory savings of 7 GB.
🎯 Key Takeaway
Past 30m you're usually masking an upstream bug — fix writers, clocks, or backfill paths instead of paying head memory forever.
● Production incidentPOST-MORTEMseverity: high

The HA Pair That Ate 18% of Billing Metrics for 6 Days

Symptom
Billing dashboards showed 18% missing samples for 6 days while both Prometheus servers reported all targets up. Logs on the receiver showed Out of order sample from remote write at roughly 12,000 drops per minute, and the gap pattern flipped between servers after each failover.
Assumption
The team assumed the drops were normal late data from edge sites with slow links, so they raised the window from 0s to 30m and moved on. Drops fell for a day, then returned worse than before, and head memory climbed 6 GB in a week.
Root cause
Both HA servers scraped the same 400 targets and remote-wrote to one receiver with identical external labels. Whichever server's sample landed second lost the race on every series. The 30m window only masked it: samples more than minutes apart still collided, and the wider window inflated head memory by 6 GB across 1.8 million series.
Fix
They added replica: a and replica: b external labels to the two servers, placed Thanos Receive with dedupe in front of long-term storage, and dropped the window back to 10m. They also pinned NTP with chrony on all collectors and added an alert on node_timex_offset_seconds above 5 seconds. Drops went to zero within two scrape intervals and head memory fell back 5 GB overnight.
Key lesson
  • Duplicate writers with identical labels are a topology bug, not a lateness problem — no window size fixes two clocks racing on one series.
  • A window increase that helps briefly then fades is the signature of duplicates or retries, not genuine lateness. Measure the lateness distribution before touching config.
  • Every HA pair needs unique replica labels and a dedupe layer from day one, plus a clock-offset alert, or this incident repeats after every failover.
Production debug guideFive checks that separate duplicate writers, clock skew, retry storms, and an undersized window.5 entries
Symptom · 01
Logs show ingestion errors but you don't know which path drops samples
→
Fix
Run journalctl -u prometheus | grep -iE 'out-of-order|too old|out of order' | tail -50. Note whether the error names a scrape_pool and target (scrape path) or says from remote write (remote-write path). The path decides which of the next steps applies — don't change config until you know the path.
Symptom · 02
Drops cluster on one host or one job while others are clean
→
Fix
Run date +%s on the Prometheus server and on the exporter host within the same minute, or query node_timex_offset_seconds on the exporter. An offset above 5 seconds confirms skew. Fix chrony/NTP on the drifting host and restart the time service — don't widen the window to cover a growing clock error.
Symptom · 03
Every target looks fine yet drops continue across many jobs
→
Fix
Query count by (job, instance) (up) and compare against your inventory. Two up series for one expected target means duplicate scraping. Check for overlapping static_configs, duplicated Kubernetes service discovery roles, or an HA pair scraping the same endpoints. Remove one, or add unique external labels per writer.
Symptom · 04
Drops started right after a remote-write receiver outage recovered
→
Fix
Query rate(prometheus_remote_storage_queue_retry_samples_total[5m]) and rate(prometheus_remote_storage_samples_dropped_total[5m]). A retry spike that precedes the drops means the sender replayed stale batches after a receiver outage. Tighten queue_config retries and use a backfill path for history instead of replaying it through the live queue.
Symptom · 05
You've ruled out duplicates, skew, and retries but late samples still drop
→
Fix
Query histogram_quantile(0.99, rate(prometheus_target_scrape_pool_late_samples_total[15m])) if exposed, or compare sample timestamps against head max time for the affected job. If 99% of late samples land within 8 minutes, set out_of_order_time_window: 10m, reload, and watch the dropped counter fall to zero over two scrape intervals before declaring victory.
Out-of-Order Causes Compared — Diagnose in Minutes
Root CauseHow to ConfirmFixPrevention
Duplicate writers (HA pair, same labels)Two scrape_pool targets report up for one job; series has samples from two instance IPsAdd unique replica external label and dedupe at receiverOne writer per series, or dedupe layer in front of storage
Clock skew on exporter hostnode_timex_offset_seconds above 5s; late samples cluster on one hostFix NTP/chrony, replace drifting VM clock sourceAlert on clock offset; run chrony everywhere
Remote-write retry replay after outageprometheus_remote_storage_queue_retry_samples_total spikes before dropsBound queue_config retries; backfill with backfill API insteadAlert on retry volume; keep receivers highly available
Window too small for normal latenessLate samples land within minutes of head max time but still dropRaise out_of_order_time_window to 10m, measure againTrack lateness distribution; size window from data
⚙ Quick Reference
3 commands from this guide
FileCommand / CodePurpose
sum by (job) (rate(prometheus_remote_storage_samples_dropped_total[5m]) * 60)How the TSDB Append Path Rejects Late Samples
prometheus.ymlglobal:Configuring out_of_order_time_window Without Regret
prometheus-ha.ymlglobal:Duplicate Writers

Key takeaways

1
Out-of-order drops mean a sample arrived older than a sample already stored
check writers, clocks, and retries first.
2
One writer per series plus unique replica labels beats any window size for duplicate-write conflicts.
3
Size out_of_order_time_window from measured lateness; 10m covers most fleets without heavy memory cost.
4
Fix NTP drift instead of masking it
skew grows and returns as bigger gaps later.
5
Bound remote-write retries so an outage recovery can't replay stale batches into the head.
6
Alert on dropped-sample and queue-retry counters so the next incident pages you before graphs go gappy.

Common mistakes to avoid

5 patterns
×

Running two HA Prometheus servers with identical external labels

Symptom
Half your samples vanish with out-of-order errors even though each server scrapes cleanly on its own.
Fix
Give each writer a unique external label like replica: A/B and dedupe at the receiver with Thanos or Mimir. Don't run two identical Prometheus servers writing the same series to one endpoint.
×

Setting honor_timestamps: false to silence timestamp errors

Symptom
Errors stop but graphs shift in time and alerts evaluate stale data, hiding the real clock problem.
Fix
Use honor_timestamps: true (the default) unless the exporter is known to emit bad timestamps. Don't mask clock bugs with honor_timestamps: false in production.
×

Setting out_of_order_time_window to hours on day one

Symptom
Memory climbs steadily, compactions take longer, and query latency grows across the whole server.
Fix
Start with 10m and only raise it after measuring lateness. Widening the window raises head memory and slows compaction for every series, not just late ones.
×

Assuming clocks are fine because the servers are in one cloud region

Symptom
VMs drift 30-90 seconds after live migration, and out-of-order errors spike on random targets each week.
Fix
Sync clocks with chrony or NTP on every exporter host. Alert when node_time_offset_seconds exceeds 5 seconds so skew never silently returns.
×

Leaving remote-write retry queues unbounded during receiver outages

Symptom
After a 20-minute receiver outage, the reconnecting sender replays huge backlogged batches that arrive out of order and get rejected.
Fix
Cap retries with sensible queue_config values (max_shards, max_samples_per_send, capacity) and alert on prometheus_remote_storage_queue_retry_samples_total so a retry storm pages you before it corrupts ingestion.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Why does Prometheus reject out-of-order samples instead of inserting the...
Q02JUNIOR
What does out_of_order_time_window do, and what's its default?
Q03SENIOR
You run HA Prometheus pairs and see out-of-order drops. What's the root ...
Q04SENIOR
How do you prove clock skew is the cause rather than duplicates?
Q05SENIOR
Why is raising the window the wrong first fix for duplicate-writer drops...
Q01 of 05JUNIOR

Why does Prometheus reject out-of-order samples instead of inserting them?

ANSWER
The head keeps samples per series in timestamp order for fast appends. Accepting older timestamps would force expensive inserts into sealed chunks, so Prometheus rejects them. Version 2.39 added a bounded window so slightly late samples can still land in the head without paying the full cost of random inserts.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
What's the difference between out-of-order, out-of-bounds, and too-old?
02
Can Prometheus ingest late data at all?
03
Does a bigger window cost anything?
04
Do out-of-order drops trigger the storage size limits?
05
How do I backfill last month's data without hitting this?
06
Should both HA Prometheus servers remote-write the same series?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Prometheus. Mark it forged?

6 min read · try the examples if you haven't

1 / 4 · Prometheus
Next
Prometheus Query Returns No Data but Target Is Up
→