Prometheus rate() vs increase(): Fix Wrong Alerts
Prometheus rate() vs increase() explained: per-second alerts use rate(), window totals use increase().
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
- ✓Basic PromQL: metric names, labels, and range selectors
- ✓One counter metric you can query in the expression browser
- ✓An alert rule file you've read at least once
- rate() returns per-second average growth for alerts, while increase() returns total growth over the window for reports
- Both handle counter resets automatically, so restarts never corrupt the math the way raw counters would
- Keep windows at 4x the scrape interval since short windows on slow jobs flap from single-sample noise
- Page with rate() plus for-durations, and show increase() totals on dashboards where window scaling is visible
Think of a water meter. rate() tells you how fast water flows right now — liters per second — which is what you need to detect a burst pipe. increase() tells you how many liters filled the bucket in the last hour — what you need for the bill. Setting a burst-pipe alarm using the bucket total means the alarm's meaning changes every time you swap bucket sizes. Match the measurement to the job: speed for alarms, totals for bills.
The alert looked airtight in review: error budget burn measured with increase() over 5 minutes against a threshold the team agreed on. Then it paged 14 times in one night during normal traffic, stayed silent through a real 20-minute outage the next week, and nobody trusts it anymore.
The bug is a unit mismatch hiding in plain sight. rate() answers per-second — how fast. increase() answers total — how many. An alert threshold written in one unit and evaluated in the other misfires by the exact ratio of the window length, and counter resets plus short windows add two more ways to be wrong.
This guide makes the semantics concrete with numbers you can verify in the expression browser. You'll learn which function pages and which reports, how resets are handled for you, why windows must respect scrape intervals, and how to fix the three alert patterns that mix them up. Eight minutes here ends the flapping-alert era for your team. Every example here runs in the expression browser, so you can prove each claim before touching alerts.
rate() Is Speed, increase() Is Distance
rate() and increase() read the same counter but answer different questions. rate() divides growth by window seconds, returning events per second — a speed. increase() returns the raw growth over the window, a total count. For steady traffic at 14 errors/sec, rate(m[5m]) reads 14 while increase(m[5m]) reads about 4,200.
The relationship is exact: increase over a window roughly equals rate times the window length in seconds. That identity is your debugging superpower — evaluate both side by side and the ratio should match the window. When it doesn't, something interesting happened inside the window, usually a reset.
Choose by consumer. Pagers need speeds with stable thresholds, so alerts use rate(). Humans need totals for standups and postmortems, so dashboards use increase(). Mixing them means thresholds silently scale with window edits, which is how teams end up paged by their own config.
The implementation detail behind rate explains its edge behavior. It takes the first and last samples in the window, adjusts for resets between them, then extrapolates the slope to the full window boundaries. Extrapolation assumes samples near both edges; a scrape gap at the window start inflates the result because the slope stretches over missing time. This is why wider windows beat narrow ones on gappy jobs — more real samples anchor both ends. When values look 10 percent high on scrape-edge cases, suspect extrapolation before suspecting traffic. The function is honest about what it saw; sparse windows just show it little.
increase() > 1000 while thinking per-second. Normal 14/sec traffic read as 4,200 and paged nightly. One function swap to rate() with a 25/sec threshold ended months of false pages.rate(); humans read window-total increase() — never swap them.Counter Resets Are Handled — Trust the Functions
Counters reset to zero on every exporter restart, and both functions expect it. When they see the value drop, they treat it as a reset: they add the growth before the drop to the growth after, instead of reporting a negative plunge. Your alert math survives deploys without special cases.
This is why alerting on raw counters is always wrong. A raw errors_total graph cliffs to zero at every restart, and any threshold on it pages on deploys and misses real growth between restarts. rate() and increase() exist precisely to absorb resets — let them.
Verify it yourself once and you'll trust it forever. Graph a raw counter across a restart, then overlay increase(m[5m]): the raw line cliffs while the increase line sails through unchanged. That ten-second experiment ends every debate about whether restarts corrupt alert math.
Resets interact with federation and recording rules in ways worth knowing. Federated counters reset when the scraping server restarts, independent of exporter restarts — two reset sources, both handled the same way. Recording rules that sum counters before rating change reset visibility: sum first and individual resets hide inside the total, which is usually what you want for availability math. Long-term-storage downsampling preserves the compensation because it keeps enough resolution. The practical rule stays simple: rate the rawest counter you have, as close to the source as possible, and let every layer absorb the resets that occur there.
rate()-based alerts stayed silent because reset compensation absorbed every restart.Window Size vs Scrape Interval: Stop the Flapping
Range windows must respect the scrape interval. rate() needs at least two samples in the window to compute slope; with a 1m window on a 60s scrape job, most evaluations see one sample and extrapolate wildly. The result flaps between zero and huge on every evaluation cycle.
The rule is simple: windows at 4x the interval or wider. Fifteen-second scrapes pair with 5m windows; 60s scrapes need 5m minimum. Wider windows also smooth legitimate burstiness so alerts fire on sustained shifts rather than single bad scrapes.
Add for-durations as the second half of the fix. for: 5m requires the breach to persist across evaluations before paging, which filters deploy spikes and traffic bursts while catching real regressions within minutes. Window width and for-duration together are what make alerts trustworthy.
The for-duration is a state machine with three states: inactive, pending, firing. A breach starts pending; recovery before the duration ends returns to inactive with no page; sustained breach fires and stays firing until resolution plus a resolve timeout. This design filters spikes exactly because most spikes die in pending. Pair for with evaluation alignment: a 5m for on a 1m group needs five consecutive breaching evaluations, while the same for on a 5m group needs one. Shorter groups plus moderate for-durations catch real shifts fastest with fewest false pages. Tune the pair together, never alone.
Gauges Are Not Counters: Don't rate() Them
Gauges break rate() because the function assumes counter semantics: any decrease looks like a reset and gets added back as phantom growth. A queue-depth gauge dipping from 100 to 40 reads as a reset of 60, inflating the result with traffic that never existed.
Use the right tool per type. Counters (requests, errors, bytes sent) get rate() and increase(). Gauges (temperature, queue depth, memory) get raw-value alerts or delta() for change-over-window. The metric exposition or docs declare the type — check before wrapping.
Negative outputs are the tell. A rate() that returns negative values on a supposedly counter-like metric means it's actually a gauge, or the instrumentation resets it arbitrarily. Either way the function is wrong for the data, and the fix is at the query, not the threshold.
Two more gauge functions complete the toolkit. deriv() returns the per-second derivative of a gauge over the window — useful for disk-fill rates and temperature slopes where resets never occur. predict_linear() extends the derivative into the future, powering classic disk-full-in-N-hours alerts from filesystem free bytes. Both assume gauge semantics, so counter resets corrupt them exactly as rate() corrupts gauges. Match function family to metric type as a review checklist: counters get rate and increase, gauges get delta, deriv, and predict_linear, info metrics get no math at all. Type discipline eliminates the whole bug class.
rate() alert paged on every normal drain because each dip read as a reset. Switching to a raw-value threshold ended the pages and actually caught the one real backlog the rate version had masked.delta() or raw values for up-and-down metrics.Rewriting Alerts You'll Actually Trust
Rewrite broken alerts in a fixed pattern. Paging rules use rate() with per-second thresholds and for-durations: rate(http_errors_total[5m]) > 25 for 5m. Info rules and dashboards use increase() totals: increase(http_errors_total[1h]) tracking errors per hour alongside business metrics.
Backtest before enabling paging. Replay the new expression against a week of history in the expression browser and compare pages against known incidents. Every false page in backtest is a free bug caught before it wakes someone; every missed real incident is a threshold to tune.
Name alerts with their units. CheckoutErrorsPerSecondHigh reads differently from CheckoutErrorsHigh, and reviewers catch mismatches when the unit is in the name. This naming habit is the cheapest prevention in the guide — one word that stops the whole bug class at review time.
Test alert expressions before they page anyone. promtool test rules evaluates rule files against scripted series and asserts firing behavior — write a test per alert with a normal-traffic case and an incident case. Store tests beside the rules file and run them in CI; a threshold edit that breaks the incident case fails the build instead of the on-call. Backtest on real history too: replay the expression over the last incident week in the browser and confirm it would have paged. Alerts with tests earn shorter for-durations because false pages stop being scary. Untested alerts deserve long durations and low trust.
rate() with units in the name; report totals with increase(); backtest a week before paging.Units Are the Whole Game
The deeper lesson is that PromQL functions carry units, and ignoring them breaks alerts as surely as ignoring units breaks physics. rate() yields per-second, increase() yields per-window, and thresholds must speak the same unit. Code review for alerts should check units first, logic second.
This discipline compounds. Teams that name units, backtest windows, and respect scrape intervals build alert suites that page rarely and correctly. Teams that don't accumulate muted alerts until they're blind during real outages — the exact failure mode of this guide's incident.
Start with your noisiest alert today. Check its function against its threshold unit, widen its window to 4x the interval, and add a for-duration. That single rewrite usually silences the worst pager while making the next real incident louder.
Burn-rate alerting applies these lessons at scale. Instead of one threshold, pair a fast-burn alert (high rate, short window, paging) with a slow-burn alert (lower rate, long window, ticket). The fast leg catches cliffs in minutes; the slow leg catches leaks in hours without paging for noise. Both legs use rate() with per-second thresholds, both carry for-durations, and both reset cleanly across deploys. Error budgets computed with increase() over 30 days sit on dashboards beside them, translating pages into budget language managers understand. This pattern ends the single-threshold flapping era permanently.
rate(), window-total threshold means increase(). Getting this one decision right prevents every unit-mismatch page in this guide.The increase() Alert That Paged 14 Times, Then Slept
increase() over 5 minutes, so normal traffic at 14 errors/sec read as 4,200.increase() values near 4,200, paging nightly. During the real outage, a counter reset inside the window plus alert fatigue meant the page was missed for 20 minutes.increase() total to a dashboard panel tracking errors per deploy. False pages dropped from 14 per night to zero, and the next real outage paged in 3 minutes with a value the team could read at a glance.- Write the unit in the alert name (errors_per_second) so reviewers catch function-threshold mismatches before merge.
- Every window edit on an
increase()alert requires a threshold edit too —rate()thresholds survive window changes. - Verify new alerts against a week of history before enabling paging; backtesting catches unit bugs for free.
rate() means the threshold must be per-second (like 10 errors/sec), increase() means it must be a window total (like 3000 errors/5m). If the threshold's unit doesn't match the function's unit, you've found the bug — convert one side before touching anything else.rate() queryrate() wraps a gauge like temperature or queue depth, replace it with delta() or alert on the raw value. Negative rate outputs on a gauge confirm the misuse instantly.rate() with per-second thresholds plus for-durations, and move totals into increase() dashboard panels and info-level rules. Deploy, then verify with a week of history: pages should match real incidents, and totals should track business metrics like order counts.| File | Command / Code | Purpose |
|---|---|---|
| rate(http_errors_total[5m]) | rate() Is Speed, increase() Is Distance | |
| http_errors_total | Counter Resets Are Handled | |
| rate(queue_depth[5m]) | Gauges Are Not Counters | |
| alerts | - alert: CheckoutErrorsHigh | Rewriting Alerts You'll Actually Trust |
Key takeaways
increase() is window total for dashboards and reports.increase() break when windows change; rate() thresholds survive.rate() around gaugesdelta() or the raw value.Common mistakes to avoid
5 patternsMixing rate() and increase() in one alert without converting
rate() for per-second alert thresholds and increase() for totals over the window. If the threshold reads per-second, the function must be rate(); if it reads total events, it must be increase().Using 1m windows on 60s-scrape jobs
rate() extrapolates from a single sample pair.Wrapping rate() around gauges or already-aggregated values
rate() or increase() directly to the raw counter, never to an already-rated gauge. If you need per-series-then-sum math, sum the counters first and rate the sum.Paging on increase() totals that scale with window length
Alerting on instant rate spikes with no for-duration
Interview Questions on This Topic
What do rate() and increase() each return?
increase() returns total growth over the window — business-friendly for events counted. For steady traffic, increase()[1h] is roughly rate()[1h] times 3600.Frequently Asked Questions
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
That's Prometheus. Mark it forged?
6 min read · try the examples if you haven't