Prometheus No Data but Target Up: 4 Fast Fixes
Prometheus target is up but graphs are empty? Check metric typos, label matchers, stale markers, and range vs scrape interval first..
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓A Prometheus server with at least one working target
- ✓Access to the expression browser in the Prometheus UI
- ✓Ability to curl a target's /metrics endpoint
- up == 1 only proves the scrape succeeded, so verify the metric name itself with curl on the target's /metrics endpoint
- Strip label matchers one at a time since a single wrong matcher like a missing port narrows results to nothing
- Keep graph ranges above 4x the scrape interval because a 1m range on a 60s job often holds zero samples
- Bridge restarts with offset 5m for display, but alert on up == 0 so real outages still page you
Imagine calling a restaurant that picks up the phone — that's up equals 1, the line works. But if you ask for a dish by the wrong name, the kitchen sends nothing back. Your query is the order: one wrong word, one extra condition like gluten-free when the dish isn't labeled that way, and you get an empty plate from a working kitchen. Fix the order, not the phone line.
It's the most confusing Prometheus moment for newcomers: the Targets page glows green, up equals 1, yet your panel says No Data. You refresh, widen the time range, restart the exporter — nothing. The instinct is to blame scraping, but scraping is the one thing working fine.
Nine times out of ten the problem lives in the query or in timing, not in collection. A mistyped metric name, a label matcher that matches nothing, a graph range shorter than your scrape interval, or stale markers left by a restarted target — each produces a perfectly empty graph above a perfectly healthy target.
This guide walks you through the four checks in order, from fastest to slowest. You'll learn to verify names against the real exposition, strip matchers scientifically, size ranges to your scrape interval, and bridge renames without corrupting alerts. Ten minutes here saves hours of restarting things that were never broken. The expression browser plus one curl command resolves most cases without touching any config.
up Means the Scrape Worked — Nothing More
The up series is the most misunderstood metric in Prometheus. It's generated by the server itself after each scrape: 1 means the HTTP request worked, 0 means it failed. It records nothing about which metrics came back in the body, so an exporter serving HTTP 200 with an empty page still reports up 1.
That makes up a scrape health signal, not a data signal. When your panel is empty and up is 1, the scrape succeeded and your investigation belongs in the query layer: names, matchers, ranges, and staleness. Touching scrape config at this point only risks breaking the part that works.
Build the habit of checking up first every time. If up is 1, say out loud that scraping works, then move to the expression browser with your exact panel query. This single discipline eliminates the most expensive mistake in this incident class: restarting infrastructure over a typo.
Learn to read the Targets page like a log file. Each target row shows last scrape duration, the error string from the most recent failure, and labels discovered before relabeling. A target flapping between up and down with context deadline exceeded points at scrape_timeout shorter than the endpoint's response time — raise the timeout rather than the interval. Discovered labels reveal service-discovery surprises: duplicate targets, stale endpoints, unexpected ports. During any empty-panel incident, screenshot the target row first; it timestamps the evidence and ends debates about whether scraping worked.
Metric Names Lie: Typos and Silent Renames
Metric names change more often than you'd think. Client libraries add _total suffixes, exporters rename seconds to duration, and major version bumps restructure whole families. Dashboards written last year query names that no longer exist, and each one fails silently as No Data.
The fix starts at the source of truth: the /metrics endpoint on the target. Curl it, grep for your metric's prefix, and copy the exact name including suffixes. Then confirm storage agrees by listing label values for __name__ with a match selector on your job — if the name isn't there, no query can return it.
Protect yourself going forward. Diff /metrics output in CI when upgrading exporters, and create recording rules that alias old names to new ones during migrations. Aliases keep history continuous and let dashboards migrate at your pace instead of during an incident.
OpenMetrics suffixes explain a whole class of renames. Counters gain _total, gauges stay bare, info metrics end _info, and created timestamps end _created. Client libraries append these automatically, so dashboards written against bare names break on upgrade. The label-values API settles every dispute: GET /api/v1/label/__name__/values with a match[] selector for your job lists exactly what storage holds. When the name exists on /metrics but not in storage, a metric_relabel_configs drop or honor_labels clash sits between — check both before blaming the exporter. Keep a short list of canonical names per job in your runbook and diff it after upgrades.
Label Matchers That Match Nothing Together
Label matchers areANDed: every matcher must match the same series. One wrong matcher — a missing port on instance, a version label that changed, a team label the series never had — narrows the result to the empty set even when each matcher looks plausible alone.
Debug by subtraction. Run the bare metric name first; if that returns data, add matchers back one at a time until the result empties. The last matcher added is the liar. This binary search takes under a minute and beats staring at four matchers trying to spot the wrong one.
The instance label causes most of these. Prometheus sets it to host:port from the scrape target, so web-01 never matches web-01:8080. When in doubt, query the series without instance, read the actual instance value from the result, and paste it back exactly.
Regex matchers deserve respect and suspicion in equal measure. The =~ operator matches RE2 patterns fully anchored, so status=~"5.." catches every 500-level code while status=~"5" matches nothing — anchoring surprises almost everyone once. Negative matchers (!=, !~) select series lacking the label too, which quietly widens results when labels are absent. Whitespace inside label values is significant: a trailing space in an annotation becomes part of the value and defeats exact matchers. When subtraction debugging stalls, render candidates with count by for each label value and read the actual strings. One glance at real values beats ten minutes of guessing at patterns.
Scrape Interval vs Graph Range Mismatch
Timing faults look exactly like missing data. A job scraped every 60s drops one sample per minute, so a 1-minute graph range frequently contains zero samples and range functions like rate() that need two points return nothing. The target is healthy; the window is just too small to see it.
Size the view to the scrape. Keep the panel range at 4x the interval or wider, and set min step at or above the interval so each evaluation step contains a sample. For 60s jobs that means 5m ranges and 60s steps as a floor. Grafana's auto interval usually handles this, but custom min intervals override it and cause these gaps.
Evaluation alignment matters too. Prometheus evaluates range queries at step boundaries that rarely align with scrape times, so tight ranges flicker between one sample and none on each refresh. Wider ranges absorb the jitter and the flicker disappears without touching any config.
Two timing rules govern every range query. First, Prometheus marks series stale after 5 minutes without samples, so ranges crossing a restart show honest gaps no window can fill. Second, rule evaluation adds its own jitter: groups evaluate on their interval boundaries, and a 15s group reading a 60s job sees single-sample windows half the time. Size range windows for the slowest job in the query, not the fastest — mixed-interval dashboards gap on their slowest leg. When a panel mixes jobs, consider splitting it: one panel per interval keeps each range honest instead of compromising all of them.
Stale Markers After Restarts and Renames
When a target restarts, Prometheus marks its old series stale and starts fresh ones, often with a changed instance port or pod name. Queries spanning the restart see two half-series with a gap between them, and dashboards look broken even though collection never stopped.
The offset modifier bridges the gap for display. Appending offset 5m shifts the evaluation window into the past where the old series still has samples, drawing a continuous line across restarts and renames. It's a viewing lens, not a data fix — the underlying series are still distinct.
Keep offsets out of alerting rules. An offset alert evaluates minutes-old data, fires late on real outages, and can miss fresh series entirely. Alert on up == 0 for dead targets and on fresh counter rates for behavior, and let dashboards use offsets for pretty continuity.
Stale markers behave differently by cause, and the difference guides the fix. A target going down stamps staleness after missed scrapes, and the series resumes cleanly when it returns — gaps are truthful history. A restart with new labels (fresh pod name, new port) orphans the old series permanently; the old line ends and a new one begins, and no offset reunites them for alerting. Kubernetes rollouts produce the second kind in bulk, which is why service-level labels outlive pod labels in every durable query. For business continuity, aggregate by stable labels in recording rules so deploys never redraw history. Offsets stay a display convenience for humans, never an input to machines that page.
The Five-Minute Checklist That Ends the Incident
Put the checks in order and most incidents resolve in minutes. First confirm up equals 1 so scrape is eliminated. Then verify the exact metric name against /metrics and the label-values API. Then subtract matchers until data returns. Then widen the range past 4x the interval. Only then consider staleness and offsets.
This order is deliberate: each step is cheaper than the next and eliminates a whole class of causes. Name checks take seconds, matcher subtraction takes a minute, range sizing takes a glance at scrape config. Engineers who jump to restarts skip all four cheap steps and pay with WAL replays and extended gaps.
Write the sequence on your runbook. During an incident nobody thinks clearly, and a five-line checklist beats memory. Teams that follow it resolve empty-panel pages in under 10 minutes; teams that freestyle average over an hour with at least one unnecessary restart.
Turn the checklist into a runbook page with copy-paste commands for your fleet: the up query, the /metrics curl, the matcher-subtraction order, the range floor per scrape interval. Link it from the alert so the paged engineer lands on procedure, not a blank search box. Rehearse it in game days by renaming a metric in staging and timing the team — ten-minute resolution should feel routine before production demands it. After each real incident, append the cause and the step that caught it; runbooks compound like interest. The checklist that ends incidents in minutes is built on quiet afternoons, not during them.
The Renamed Counter That Emptied Every Checkout Dashboard
- Client library upgrades rename metrics silently — diff /metrics output before and after every upgrade.
- Restarting Prometheus for an empty panel converts a query bug into a real ingestion gap for all jobs.
- Recording-rule aliases preserve history across renames and cost almost nothing to maintain.
rate() need at least two samples inside the window, and short ranges on slow jobs often contain zero.| File | Command / Code | Purpose |
|---|---|---|
| up{job="node"} | up Means the Scrape Worked | |
| curl -s localhost:8080/metrics | grep -E '^http_(server_)?requests' | head | Metric Names Lie | |
| rate(http_requests_total[1m]) | Scrape Interval vs Graph Range Mismatch | |
| rate(http_requests_total[5m] offset 5m) | Stale Markers After Restarts and Renames |
Key takeaways
Common mistakes to avoid
5 patternsTyping metric names from memory instead of copying them
Stacking label matchers that match nothing together
Using a graph range shorter than two scrape intervals
Hiding gaps with or vector(0) in alerting rules
rate() math.Restarting Prometheus before checking the query
Interview Questions on This Topic
Target shows up but the graph is empty. Where do you look first?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Prometheus. Mark it forged?
6 min read · try the examples if you haven't