Prometheus Out-of-Order Samples: Fix Dropped Writes
Prometheus drops late samples with out-of-order errors.
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓A running Prometheus server you can query and restart
- ✓Basic PromQL for checking up and
rate()queries - ✓Access to exporter hosts for clock checks
- Prometheus rejects a sample older than the newest stored sample for that series, logging out-of-order or too-old errors with a dropped count
- Set storage.tsdb.out_of_order_time_window to 10m so samples up to 10 minutes late still ingest instead of dropping
- Give each HA writer a unique replica external label and dedupe at the receiver so two servers stop racing on one series
- Fix exporter clock drift with chrony or NTP, since skew past a few seconds guarantees late samples on every scrape
Think of each metric as a diary with numbered pages in order. Prometheus only lets you write on the next blank page — you can't go back and squeeze a sentence onto page 12 once page 15 is full. When two people share one diary, or someone's watch runs slow, entries arrive for pages that are already full, and Prometheus throws them away. The fix is giving each writer their own diary, fixing the slow watch, and allowing a small grace period for slightly late entries.
Your dashboards developed gaps overnight, but every target shows up on the Targets page. The server log tells a different story: thousands of lines reading Error on ingesting out-of-order samples with a num_dropped count climbing every minute. Nothing was deployed, no config changed, and the metrics simply stopped landing.
Prometheus stores each series as a time-ordered append log, and that design choice is the whole story. A sample whose timestamp is older than the newest sample already stored for that series can't be appended cheaply, so the TSDB rejects it. The usual triggers are painfully ordinary: two servers writing the same series, a VM clock that drifted 40 seconds, or a remote-write queue replaying yesterday's backlog after an outage.
This guide shows you how to confirm which trigger you've got, fix it without guessing, and size the out_of_order_time_window setting from real measurements. You'll leave with queries that quantify your sample lateness, configs that stop duplicate writes, and alerts that catch clock skew before it eats your data.
How the TSDB Append Path Rejects Late Samples
Every sample Prometheus stores travels one path: scrape, append to the in-memory head, then cut into two-hour blocks. The head keeps each series as an append-only sequence, so a sample with a timestamp older than the last one stored has nowhere cheap to go. The server rejects it and logs which path failed: Error on ingesting out-of-order samples for scrapes, Out of order sample from remote write for remote-write receivers, or too old sample when you've configured a window and the sample is older than head-max minus that window.
This strictness is a performance bargain, not a bug. Ordered appends let the head compress samples into tiny chunks at millions of samples per second. Random inserts into sealed chunks would cost far more CPU and memory for a case that's usually caused by broken clocks or duplicate writers. The window added in 2.39 keeps the bargain while tolerating real-world lateness: slightly late samples buffer in the head, badly late ones still drop.
Your first job in any incident is naming the path. Scrape-path errors name a scrape_pool and target, so you investigate exporters and scrape config. Remote-write errors name a series and timestamp, so you investigate senders, queues, and receiver topology. Mixing the two up wastes hours because the fixes live on opposite ends of the pipeline.
Configuring out_of_order_time_window Without Regret
The window lives under the storage key in prometheus.yml, not under global or scrape configs. Setting storage.tsdb.out_of_order_time_window to 10m tells the head to accept samples with timestamps back to head-max minus ten minutes. The default of 0s keeps strict ordering, which is right for fleets with healthy clocks and one writer per series.
Pick the value from your lateness data, not from a blog post. If 99% of your late samples land within 8 minutes, 10m covers them with margin. Values above 30m need justification because the cost lands on every series in the head, and values in hours are almost always masking duplicates or retries that deserve a real fix.
Roll it out like any storage change: apply to staging first, reload with kill -HUP, and watch the dropped-sample rate plus head memory for a full compaction cycle. Keep the old value in a comment so the next engineer knows what changed and why.
Reloading applies the window without a restart: send SIGHUP or POST to /-/reload and the head keeps serving queries throughout. The new value governs samples appended after reload; in-flight scrapes finish under the old rule, so watch the dropped counter across two full scrape intervals before judging. Pair the window with retention thinking: a 10m window on 15d retention changes head behavior only, while disk blocks stay untouched. Record the chosen value beside retention.time in your runbook with the p99 lateness number that justified it, so the next capacity review inherits evidence instead of folklore.
Duplicate Writers: HA Pairs and Retry Replays
Duplicate writers are the most common cause and the least responsive to window tuning. An HA pair scraping the same targets with identical external labels sends two timestamp streams for every series. Each stream is individually ordered, but interleaved at the receiver they're a coin flip — the second arrival per timestamp loses and drops as out-of-order.
The fix is topological. Give each writer a unique external label such as replica: a versus replica: b, then dedupe at the receiver layer with Thanos Receive or Mimir's dedupe. The dedupe keeps one ordered stream per series and the drops vanish regardless of window size. Remote-write retries create the same pattern after an outage: the queue replays stale batches that collide with live samples, so bound the retry queue and backfill history through a proper backfill path instead.
Confirm duplicates before you reconfigure. If count by (job, instance) (up) shows two up series per expected target, or drops flip between senders after failover, you've found it. Window changes can't win a race between two clocks — only one writer per series can.
Beyond HA pairs, the same race hides in overlapping scrape jobs: two jobs targeting one exporter with different intervals interleave timestamps on shared series. Audit with count by (job, instance) (up) and merge duplicates into a single job per target set. One writer per series is a topology invariant, not a tuning knob.
Clock Skew: the Drift That Outgrows Every Window
Exporter clocks drift more than teams expect. VMs slip 30 to 90 seconds after live migration, containers inherit host drift, and edge devices with dead NTP daemons wander minutes per week. Every scrape from a skewed host stamps samples in the past, and no window survives skew that grows without bound.
Diagnose with two commands: compare date +%s on the server and exporter in the same minute, and query node_timex_offset_seconds where the node exporter runs. Offsets above 5 seconds deserve a clock fix, not a config change. Standardize on chrony, point every host at the same NTP tier, and alert on offset so drift pages someone before samples start dropping.
The honor_timestamps setting deserves caution here. Flipping it to false stamps arrival time instead of exporter time, which silences errors while quietly shifting your data in time. Keep the default true, fix the clock, and your rate() windows and alert evaluations stay honest.
Standardize time sync with chrony on every host: two or three stratum sources, makestep for fast correction after boot, and maxslewrate to bound drift between polls. Virtual machines should use the hypervisor clocksource (kvm-clock on KVM, xen on Xen) so live migrations do not inject jumps. Containers inherit the host clock, which makes host NTP a fleet-wide dependency worth monitoring. Alert with node_timex_offset_seconds above 5 seconds plus a max-error variant for systemd-timesyncd hosts. When an alert fires, fix the daemon and verify with chronyc tracking before touching any Prometheus setting.
Measuring Lateness So the Window Fits the Data
Guessing the window size is how fleets end up at 2h with 9 GB of extra head memory. Measure instead. The dropped-sample and retry counters tell you volume, and timestamp comparisons tell you distribution: how late are the late samples, and does the 99th percentile sit at 6 minutes or 60?
Build a small dashboard before changing config: drops per minute by job, retry volume by remote name, head series count, and resident memory. Record a week of baseline so you can attribute every change. Then set the window just above your measured p99 lateness and verify the drops hit zero within two scrape intervals.
Keep measuring after the fix. Lateness drifts as fleets grow, new regions add network delay, and receivers get slower. A quarterly glance at the lateness panel tells you whether 10m still fits or needs a nudge — a two-minute check that prevents the next gap incident.
A practical lateness proxy needs no new instrumentation: max by (job) (time() - timestamp(up)) measures seconds since each target's latest sample landed. Graph it per job for a week; the p99 of that series is your honest lateness budget, including scrape, network, and remote-write delay. Set the window just above that p99 — 10m covers fleets measuring 6 to 8 minutes — and re-check quarterly as regions and receivers change. Alert when the proxy exceeds the window: it means lateness outgrew the budget and drops are imminent. This single panel turns window sizing from guesswork into arithmetic, and it doubles as an early-warning gauge long after the incident closes.
When Not to Widen the Window
A bigger window isn't free. Late inserts fragment head chunks, raise resident memory, and slow the compactions that cut two-hour blocks. On million-series servers the cost is gigabytes, and it persists 24/7 — you pay it during quiet nights for lateness that happens at peak.
There are also cases the window can't help at all. Samples older than head-max minus the window still drop as too-old, so multi-hour backlogs need a backfill path, not a bigger number. Federation and recording rules that re-stamp timestamps can manufacture disorder no window should excuse — fix the rule instead.
Treat 30m as a smell threshold. If your measurements demand more, something upstream is broken: a receiver outage replaying history, a region with chronic delay, or duplicates masquerading as lateness. Fix that layer and the window usually shrinks back to 10m with drops still at zero.
The memory math explains the smell threshold. Each out-of-order sample forces the head to hold overlapping chunk ranges instead of one clean append stream, and the overhead lands on every active series whether it receives late samples or not. Past 30m the cost typically exceeds a gigabyte per million series while buying tolerance for delays that indicate broken pipelines. Genuine exceptions exist: cross-region federation with known multi-minute lag, or edge buffers that flush hourly by design. For those, isolate the late streams in a dedicated Prometheus with a wider window instead of taxing the main fleet. Keep the primary window tight, quarantine the exotic lateness, and both systems stay cheap.
The HA Pair That Ate 18% of Billing Metrics for 6 Days
- Duplicate writers with identical labels are a topology bug, not a lateness problem — no window size fixes two clocks racing on one series.
- A window increase that helps briefly then fades is the signature of duplicates or retries, not genuine lateness. Measure the lateness distribution before touching config.
- Every HA pair needs unique replica labels and a dedupe layer from day one, plus a clock-offset alert, or this incident repeats after every failover.
| File | Command / Code | Purpose |
|---|---|---|
| sum by (job) (rate(prometheus_remote_storage_samples_dropped_total[5m]) * 60) | How the TSDB Append Path Rejects Late Samples | |
| prometheus.yml | global: | Configuring out_of_order_time_window Without Regret |
| prometheus-ha.yml | global: | Duplicate Writers |
Key takeaways
Common mistakes to avoid
5 patternsRunning two HA Prometheus servers with identical external labels
Setting honor_timestamps: false to silence timestamp errors
Setting out_of_order_time_window to hours on day one
Assuming clocks are fine because the servers are in one cloud region
Leaving remote-write retry queues unbounded during receiver outages
Interview Questions on This Topic
Why does Prometheus reject out-of-order samples instead of inserting them?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Prometheus. Mark it forged?
6 min read · try the examples if you haven't