Home › Observability › Prometheus rate() vs increase(): Fix Wrong Alerts
Beginner 6 min · September 23, 2026

Prometheus rate() vs increase(): Fix Wrong Alerts

Prometheus rate() vs increase() explained: per-second alerts use rate(), window totals use increase().

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 8 min
  • ✓Basic PromQL: metric names, labels, and range selectors
  • ✓One counter metric you can query in the expression browser
  • ✓An alert rule file you've read at least once
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • rate() returns per-second average growth for alerts, while increase() returns total growth over the window for reports
  • Both handle counter resets automatically, so restarts never corrupt the math the way raw counters would
  • Keep windows at 4x the scrape interval since short windows on slow jobs flap from single-sample noise
  • Page with rate() plus for-durations, and show increase() totals on dashboards where window scaling is visible
✦ Definition~90s read
What is Prometheus rate() vs increase()?

Prometheus counters track cumulative totals that only rise — HTTP requests served, errors counted, bytes sent — resetting to zero when the exporter restarts. PromQL's rate() and increase() turn these ever-growing totals into useful signals over a range window like [5m]. rate() returns the per-second average growth rate, answering how fast events happen. increase() returns the total growth over the window, answering how many events happened.

★
Think of a water meter.

Both functions share reset compensation: when the counter value drops, they assume a restart and add pre-drop growth to post-drop growth instead of reporting negative values. Both need at least two samples inside the window, which is why window width must respect the scrape interval — a 1m window on a 60s job often holds a single sample and extrapolates noise.

The practical split follows from the units. rate() yields events-per-second, stable across window edits, ideal for paging thresholds like 25 errors/sec. increase() yields events-per-window, scaling with window length, ideal for dashboards, burn accounting, and business reports. Alerts that use increase() with fixed thresholds break on every window tweak; dashboards that use rate() for totals confuse every reader.

Match the function to the consumer and both stay correct.

Plain-English First

Think of a water meter. rate() tells you how fast water flows right now — liters per second — which is what you need to detect a burst pipe. increase() tells you how many liters filled the bucket in the last hour — what you need for the bill. Setting a burst-pipe alarm using the bucket total means the alarm's meaning changes every time you swap bucket sizes. Match the measurement to the job: speed for alarms, totals for bills.

The alert looked airtight in review: error budget burn measured with increase() over 5 minutes against a threshold the team agreed on. Then it paged 14 times in one night during normal traffic, stayed silent through a real 20-minute outage the next week, and nobody trusts it anymore.

The bug is a unit mismatch hiding in plain sight. rate() answers per-second — how fast. increase() answers total — how many. An alert threshold written in one unit and evaluated in the other misfires by the exact ratio of the window length, and counter resets plus short windows add two more ways to be wrong.

This guide makes the semantics concrete with numbers you can verify in the expression browser. You'll learn which function pages and which reports, how resets are handled for you, why windows must respect scrape intervals, and how to fix the three alert patterns that mix them up. Eight minutes here ends the flapping-alert era for your team. Every example here runs in the expression browser, so you can prove each claim before touching alerts.

rate() Is Speed, increase() Is Distance

rate() and increase() read the same counter but answer different questions. rate() divides growth by window seconds, returning events per second — a speed. increase() returns the raw growth over the window, a total count. For steady traffic at 14 errors/sec, rate(m[5m]) reads 14 while increase(m[5m]) reads about 4,200.

The relationship is exact: increase over a window roughly equals rate times the window length in seconds. That identity is your debugging superpower — evaluate both side by side and the ratio should match the window. When it doesn't, something interesting happened inside the window, usually a reset.

Choose by consumer. Pagers need speeds with stable thresholds, so alerts use rate(). Humans need totals for standups and postmortems, so dashboards use increase(). Mixing them means thresholds silently scale with window edits, which is how teams end up paged by their own config.

The implementation detail behind rate explains its edge behavior. It takes the first and last samples in the window, adjusts for resets between them, then extrapolates the slope to the full window boundaries. Extrapolation assumes samples near both edges; a scrape gap at the window start inflates the result because the slope stretches over missing time. This is why wider windows beat narrow ones on gappy jobs — more real samples anchor both ends. When values look 10 percent high on scrape-edge cases, suspect extrapolation before suspecting traffic. The function is honest about what it saw; sparse windows just show it little.

PROMQL
1
2
3
4
5
6
7
8
9
10
# Per-second speed — for alerts with per-second thresholds
rate(http_errors_total[5m])
# > 25 means 25 errors/sec sustained

# Window total — for dashboards and burn accounting
increase(http_errors_total[5m])
# ~ rate * 300 during steady traffic (300s window)

# Wrong: threshold in one unit, function in the other
increase(http_errors_total[5m]) > 1000   # 1000 what? total, not per-second
📊 Production Insight
A team paged on increase() > 1000 while thinking per-second. Normal 14/sec traffic read as 4,200 and paged nightly. One function swap to rate() with a 25/sec threshold ended months of false pages.
🎯 Key Takeaway
Alerts page on per-second rate(); humans read window-total increase() — never swap them.

Counter Resets Are Handled — Trust the Functions

Counters reset to zero on every exporter restart, and both functions expect it. When they see the value drop, they treat it as a reset: they add the growth before the drop to the growth after, instead of reporting a negative plunge. Your alert math survives deploys without special cases.

This is why alerting on raw counters is always wrong. A raw errors_total graph cliffs to zero at every restart, and any threshold on it pages on deploys and misses real growth between restarts. rate() and increase() exist precisely to absorb resets — let them.

Verify it yourself once and you'll trust it forever. Graph a raw counter across a restart, then overlay increase(m[5m]): the raw line cliffs while the increase line sails through unchanged. That ten-second experiment ends every debate about whether restarts corrupt alert math.

Resets interact with federation and recording rules in ways worth knowing. Federated counters reset when the scraping server restarts, independent of exporter restarts — two reset sources, both handled the same way. Recording rules that sum counters before rating change reset visibility: sum first and individual resets hide inside the total, which is usually what you want for availability math. Long-term-storage downsampling preserves the compensation because it keeps enough resolution. The practical rule stays simple: rate the rawest counter you have, as close to the source as possible, and let every layer absorb the resets that occur there.

PROMQL
1
2
3
4
5
6
7
8
9
10
# Raw counter cliffs to zero at every restart — never alert on this
http_errors_total

# increase() sails through the same restart unchanged
increase(http_errors_total[5m])

# Correct paging alert: per-second rate + sustained duration
- alert: CheckoutErrorsHigh
  expr: rate(http_errors_total[5m]) > 25
  for: 5m
📊 Production Insight
During a rolling restart of 60 API pods, raw-counter alerts would have paged 60 times. The rate()-based alerts stayed silent because reset compensation absorbed every restart.
🎯 Key Takeaway
Both functions add pre-reset and post-reset growth; raw counters cliff, so never alert on them.

Window Size vs Scrape Interval: Stop the Flapping

Range windows must respect the scrape interval. rate() needs at least two samples in the window to compute slope; with a 1m window on a 60s scrape job, most evaluations see one sample and extrapolate wildly. The result flaps between zero and huge on every evaluation cycle.

The rule is simple: windows at 4x the interval or wider. Fifteen-second scrapes pair with 5m windows; 60s scrapes need 5m minimum. Wider windows also smooth legitimate burstiness so alerts fire on sustained shifts rather than single bad scrapes.

Add for-durations as the second half of the fix. for: 5m requires the breach to persist across evaluations before paging, which filters deploy spikes and traffic bursts while catching real regressions within minutes. Window width and for-duration together are what make alerts trustworthy.

The for-duration is a state machine with three states: inactive, pending, firing. A breach starts pending; recovery before the duration ends returns to inactive with no page; sustained breach fires and stays firing until resolution plus a resolve timeout. This design filters spikes exactly because most spikes die in pending. Pair for with evaluation alignment: a 5m for on a 1m group needs five consecutive breaching evaluations, while the same for on a 5m group needs one. Shorter groups plus moderate for-durations catch real shifts fastest with fewest false pages. Tune the pair together, never alone.

📊 Production Insight
A 1m-window alert on a 60s job flapped 30 times a day. Widening to 5m with for: 5m cut it to zero false pages while still catching a real 8-minute regression the next month.
🎯 Key Takeaway
Windows at 4x the interval plus for-durations turn flapping into trustworthy pages.

Gauges Are Not Counters: Don't rate() Them

Gauges break rate() because the function assumes counter semantics: any decrease looks like a reset and gets added back as phantom growth. A queue-depth gauge dipping from 100 to 40 reads as a reset of 60, inflating the result with traffic that never existed.

Use the right tool per type. Counters (requests, errors, bytes sent) get rate() and increase(). Gauges (temperature, queue depth, memory) get raw-value alerts or delta() for change-over-window. The metric exposition or docs declare the type — check before wrapping.

Negative outputs are the tell. A rate() that returns negative values on a supposedly counter-like metric means it's actually a gauge, or the instrumentation resets it arbitrarily. Either way the function is wrong for the data, and the fix is at the query, not the threshold.

Two more gauge functions complete the toolkit. deriv() returns the per-second derivative of a gauge over the window — useful for disk-fill rates and temperature slopes where resets never occur. predict_linear() extends the derivative into the future, powering classic disk-full-in-N-hours alerts from filesystem free bytes. Both assume gauge semantics, so counter resets corrupt them exactly as rate() corrupts gauges. Match function family to metric type as a review checklist: counters get rate and increase, gauges get delta, deriv, and predict_linear, info metrics get no math at all. Type discipline eliminates the whole bug class.

PROMQL
1
2
3
4
5
6
# Wrong: rate() on a gauge invents phantom growth on every dip
rate(queue_depth[5m])

# Right: alert on the gauge value or its window change
delta(queue_depth[5m]) > 500
queue_depth > 10000
📊 Production Insight
A queue-depth rate() alert paged on every normal drain because each dip read as a reset. Switching to a raw-value threshold ended the pages and actually caught the one real backlog the rate version had masked.
🎯 Key Takeaway
rate() on gauges invents reset growth; use delta() or raw values for up-and-down metrics.

Rewriting Alerts You'll Actually Trust

Rewrite broken alerts in a fixed pattern. Paging rules use rate() with per-second thresholds and for-durations: rate(http_errors_total[5m]) > 25 for 5m. Info rules and dashboards use increase() totals: increase(http_errors_total[1h]) tracking errors per hour alongside business metrics.

Backtest before enabling paging. Replay the new expression against a week of history in the expression browser and compare pages against known incidents. Every false page in backtest is a free bug caught before it wakes someone; every missed real incident is a threshold to tune.

Name alerts with their units. CheckoutErrorsPerSecondHigh reads differently from CheckoutErrorsHigh, and reviewers catch mismatches when the unit is in the name. This naming habit is the cheapest prevention in the guide — one word that stops the whole bug class at review time.

Test alert expressions before they page anyone. promtool test rules evaluates rule files against scripted series and asserts firing behavior — write a test per alert with a normal-traffic case and an incident case. Store tests beside the rules file and run them in CI; a threshold edit that breaks the incident case fails the build instead of the on-call. Backtest on real history too: replay the expression over the last incident week in the browser and confirm it would have paged. Alerts with tests earn shorter for-durations because false pages stop being scary. Untested alerts deserve long durations and low trust.

alerts/checkout.ymlYAML
1
2
3
4
5
6
7
8
9
# Page on speed, report on totals — one rule each
- alert: CheckoutErrorsHigh
  expr: rate(http_errors_total[5m]) > 25
  for: 5m
  labels:
    severity: page

# Dashboard panel tracks the business total beside it:
# increase(http_errors_total[1h])
📊 Production Insight
After rewriting six alerts to this pattern and backtesting, a team went from 40 false pages a month to 2 — both during real incidents. Alert trust recovered enough to shorten for-durations further.
🎯 Key Takeaway
Page on rate() with units in the name; report totals with increase(); backtest a week before paging.

Units Are the Whole Game

The deeper lesson is that PromQL functions carry units, and ignoring them breaks alerts as surely as ignoring units breaks physics. rate() yields per-second, increase() yields per-window, and thresholds must speak the same unit. Code review for alerts should check units first, logic second.

This discipline compounds. Teams that name units, backtest windows, and respect scrape intervals build alert suites that page rarely and correctly. Teams that don't accumulate muted alerts until they're blind during real outages — the exact failure mode of this guide's incident.

Start with your noisiest alert today. Check its function against its threshold unit, widen its window to 4x the interval, and add a for-duration. That single rewrite usually silences the worst pager while making the next real incident louder.

Burn-rate alerting applies these lessons at scale. Instead of one threshold, pair a fast-burn alert (high rate, short window, paging) with a slow-burn alert (lower rate, long window, ticket). The fast leg catches cliffs in minutes; the slow leg catches leaks in hours without paging for noise. Both legs use rate() with per-second thresholds, both carry for-durations, and both reset cleanly across deploys. Error budgets computed with increase() over 30 days sit on dashboards beside them, translating pages into budget language managers understand. This pattern ends the single-threshold flapping era permanently.

💡Match Function to Threshold Unit First
Pick the function by the threshold's unit before writing anything else. Per-second threshold means rate(), window-total threshold means increase(). Getting this one decision right prevents every unit-mismatch page in this guide.
📊 Production Insight
One team's noisiest alert needed only a function swap plus for: 5m. False pages fell 95% in a week, and the on-call rotation stopped dreading nights — from a 3-line diff.
🎯 Key Takeaway
Treat PromQL units like physical units: mismatched function and threshold always pages wrong.
● Production incidentPOST-MORTEMseverity: high

The increase() Alert That Paged 14 Times, Then Slept

Symptom
Checkout error alert paged 14 times between midnight and 6 AM with no customer impact, then failed to page during a genuine 20-minute checkout outage the following week. The team muted the alert and flew blind for a month.
Assumption
The team assumed error traffic had genuinely spiked because the alert value read 4,200 against a 1,000 threshold. Nobody converted units: the threshold was written as errors-per-second in the design doc but implemented with increase() over 5 minutes, so normal traffic at 14 errors/sec read as 4,200.
Root cause
The alert used increase(http_errors_total[5m]) > 1000 with a threshold meant as 1,000 errors per second. Normal traffic of 14 errors/sec produced increase() values near 4,200, paging nightly. During the real outage, a counter reset inside the window plus alert fatigue meant the page was missed for 20 minutes.
Fix
They rewrote the paging alert with rate(errors_total[5m]) against a 25/sec threshold with for: 5m, and moved the increase() total to a dashboard panel tracking errors per deploy. False pages dropped from 14 per night to zero, and the next real outage paged in 3 minutes with a value the team could read at a glance.
Key lesson
  • Write the unit in the alert name (errors_per_second) so reviewers catch function-threshold mismatches before merge.
  • Every window edit on an increase() alert requires a threshold edit too — rate() thresholds survive window changes.
  • Verify new alerts against a week of history before enabling paging; backtesting catches unit bugs for free.
Production debug guideFive checks that separate unit mismatches, resets, windows, and gauge misuse.5 entries
Symptom · 01
Alert fires at wrong times with no code change
→
Fix
Read the alert expression and name the function: rate() means the threshold must be per-second (like 10 errors/sec), increase() means it must be a window total (like 3000 errors/5m). If the threshold's unit doesn't match the function's unit, you've found the bug — convert one side before touching anything else.
Symptom · 02
Values look impossible after a deploy or restart
→
Fix
Evaluate increase(errors_total[5m]) and rate(errors_total[5m]) side by side in the expression browser during steady traffic. The increase value should roughly equal rate times 300. If the ratio is wildly off, a reset happened inside the window — check raw counter graphs for drops to zero and confirm restarts in deploy logs.
Symptom · 03
Alert flaps on and off every few minutes
→
Fix
Compare the job's scrape_interval in prometheus.yml against the alert's window. If the window is under 2x the interval (like 1m on 60s scrapes), widen it to 5m and add for: 5m. Re-evaluate for 15 minutes — flapping that stops was single-sample extrapolation, not real traffic.
Symptom · 04
Negative or absurd values from a rate() query
→
Fix
Check the metric type in exposition or docs: counters only rise (plus resets), gauges move both ways. If rate() wraps a gauge like temperature or queue depth, replace it with delta() or alert on the raw value. Negative rate outputs on a gauge confirm the misuse instantly.
Symptom · 05
Team ignores the alert after too many false pages
→
Fix
Rewrite paging alerts as rate() with per-second thresholds plus for-durations, and move totals into increase() dashboard panels and info-level rules. Deploy, then verify with a week of history: pages should match real incidents, and totals should track business metrics like order counts.
rate() vs increase() Faults Compared
Root CauseHow to ConfirmFixPrevention
Used increase() where rate() belongedThreshold is per-second but value grows with windowSwitch alert to rate(); keep increase() for dashboardsName rules with unit: _rps alerts use rate()
Window shorter than 2 scrape intervalsFlapping clears when window widens to 5mWindow >= 4x scrape intervalLint rule windows against job intervals
Counter reset misread as dropRaw counter falls to zero at reset timeTrust rate()/increase() reset handling; don't alert on rawDocument resets in runbooks; test with restarts
rate() applied to a gaugeNegative values on a should-be rateUse the gauge directly or delta() for gaugesType-check metrics: counters rise, gauges vary
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
rate(http_errors_total[5m])rate() Is Speed, increase() Is Distance
http_errors_totalCounter Resets Are Handled
rate(queue_depth[5m])Gauges Are Not Counters
alertscheckout.yml- alert: CheckoutErrorsHighRewriting Alerts You'll Actually Trust

Key takeaways

1
rate() is per-second speed for alerts; increase() is window total for dashboards and reports.
2
Both handle counter resets automatically
never alert on raw counter values.
3
Keep windows at 4x the scrape interval or wider to stop flapping.
4
Fixed thresholds on increase() break when windows change; rate() thresholds survive.
5
Never wrap rate() around gauges
use delta() or the raw value.
6
Add for-durations so spikes from deploys don't page but sustained shifts do.

Common mistakes to avoid

5 patterns
×

Mixing rate() and increase() in one alert without converting

Symptom
Thresholds that look right in review misfire in production because one side is per-second and the other is a window total.
Fix
Use rate() for per-second alert thresholds and increase() for totals over the window. If the threshold reads per-second, the function must be rate(); if it reads total events, it must be increase().
×

Using 1m windows on 60s-scrape jobs

Symptom
Alerts flap on every evaluation because rate() extrapolates from a single sample pair.
Fix
Keep range windows at 4x the scrape interval or wider. For 15s scrapes use 5m; for 60s scrapes use 5m minimum. Short windows on slow jobs compute over one sample and lie.
×

Wrapping rate() around gauges or already-aggregated values

Symptom
Negative or nonsensical results that trigger alerts on metrics that only move in one direction.
Fix
Apply rate() or increase() directly to the raw counter, never to an already-rated gauge. If you need per-series-then-sum math, sum the counters first and rate the sum.
×

Paging on increase() totals that scale with window length

Symptom
Every window tweak re-pages the team because doubling the window doubles the value against a fixed threshold.
Fix
Alert on the per-second rate with a separate total-growth info rule for context. The page carries the rate; the dashboard shows the totals that explain it.
×

Alerting on instant rate spikes with no for-duration

Symptom
Deploys and traffic bursts page nightly while real sustained regressions hide between the spikes.
Fix
Add explicit for: durations and require the condition across evaluations. A single-window spike from a deploy shouldn't page; a sustained rate shift should.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
What do rate() and increase() each return?
Q02JUNIOR
What happens to rate() when a counter resets?
Q03SENIOR
Why does an increase() alert misfire after a window edit?
Q04SENIOR
Why does rate(m[1m]) flap on a 60s scrape interval?
Q05SENIOR
Why does rate() on a gauge produce garbage?
Q01 of 05JUNIOR

What do rate() and increase() each return?

ANSWER
rate() returns per-second average growth over the window — dashboard-friendly for traffic speed. increase() returns total growth over the window — business-friendly for events counted. For steady traffic, increase()[1h] is roughly rate()[1h] times 3600.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Is irate() just a faster rate()?
02
Do counter resets break these functions?
03
Does changing the window change increase() values?
04
What do I use for gauges that go up and down?
05
My rate() alert flaps on a slow-scrape job. What now?
06
Can I convert increase() to per-second manually?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Prometheus. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Prometheus Cardinality Explosion From a Label
4 / 4 · Prometheus
Next
Grafana Panel Shows No Data Despite Valid Query
→