Home › Observability › Prometheus No Data but Target Up: 4 Fast Fixes
Beginner 6 min · September 23, 2026

Prometheus No Data but Target Up: 4 Fast Fixes

Prometheus target is up but graphs are empty? Check metric typos, label matchers, stale markers, and range vs scrape interval first..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 9 min
  • ✓A Prometheus server with at least one working target
  • ✓Access to the expression browser in the Prometheus UI
  • ✓Ability to curl a target's /metrics endpoint
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • up == 1 only proves the scrape succeeded, so verify the metric name itself with curl on the target's /metrics endpoint
  • Strip label matchers one at a time since a single wrong matcher like a missing port narrows results to nothing
  • Keep graph ranges above 4x the scrape interval because a 1m range on a 60s job often holds zero samples
  • Bridge restarts with offset 5m for display, but alert on up == 0 so real outages still page you
✦ Definition~90s read
What is Prometheus Query Returns No Data but Target Is Up?

Prometheus collects metrics by scraping HTTP endpoints: each target exposes a /metrics page, the server fetches it on a schedule called the scrape interval, and the results become queryable series. The special series up records whether each scrape succeeded — 1 for yes, 0 for no.

★
Imagine calling a restaurant that picks up the phone — that's up equals 1, the line works.

An empty graph with up equal to 1 therefore means collection works and the problem sits between storage and your screen.

Four gaps open up in that space. Metric names drift when exporters and client libraries rename counters between versions, so old queries address series that no longer exist. Label matchers combine with AND logic, so one wrong matcher empties the whole result.

Graph ranges shorter than the scrape interval contain too few samples for range functions that need at least two points. Stale markers appear when targets restart and old series are retired, splitting history across names.

The debugging strategy follows from that map. Verify the name against the live /metrics exposition and the label-values API. Subtract matchers until data returns. Widen ranges past four times the interval. Bridge restarts with offset for display while alerting on fresh signals. Each check isolates one layer, and together they cover every common cause of the healthy-target empty graph.

Plain-English First

Imagine calling a restaurant that picks up the phone — that's up equals 1, the line works. But if you ask for a dish by the wrong name, the kitchen sends nothing back. Your query is the order: one wrong word, one extra condition like gluten-free when the dish isn't labeled that way, and you get an empty plate from a working kitchen. Fix the order, not the phone line.

It's the most confusing Prometheus moment for newcomers: the Targets page glows green, up equals 1, yet your panel says No Data. You refresh, widen the time range, restart the exporter — nothing. The instinct is to blame scraping, but scraping is the one thing working fine.

Nine times out of ten the problem lives in the query or in timing, not in collection. A mistyped metric name, a label matcher that matches nothing, a graph range shorter than your scrape interval, or stale markers left by a restarted target — each produces a perfectly empty graph above a perfectly healthy target.

This guide walks you through the four checks in order, from fastest to slowest. You'll learn to verify names against the real exposition, strip matchers scientifically, size ranges to your scrape interval, and bridge renames without corrupting alerts. Ten minutes here saves hours of restarting things that were never broken. The expression browser plus one curl command resolves most cases without touching any config.

up Means the Scrape Worked — Nothing More

The up series is the most misunderstood metric in Prometheus. It's generated by the server itself after each scrape: 1 means the HTTP request worked, 0 means it failed. It records nothing about which metrics came back in the body, so an exporter serving HTTP 200 with an empty page still reports up 1.

That makes up a scrape health signal, not a data signal. When your panel is empty and up is 1, the scrape succeeded and your investigation belongs in the query layer: names, matchers, ranges, and staleness. Touching scrape config at this point only risks breaking the part that works.

Build the habit of checking up first every time. If up is 1, say out loud that scraping works, then move to the expression browser with your exact panel query. This single discipline eliminates the most expensive mistake in this incident class: restarting infrastructure over a typo.

Learn to read the Targets page like a log file. Each target row shows last scrape duration, the error string from the most recent failure, and labels discovered before relabeling. A target flapping between up and down with context deadline exceeded points at scrape_timeout shorter than the endpoint's response time — raise the timeout rather than the interval. Discovered labels reveal service-discovery surprises: duplicate targets, stale endpoints, unexpected ports. During any empty-panel incident, screenshot the target row first; it timestamps the evidence and ends debates about whether scraping worked.

PROMQL
1
2
3
4
5
6
7
8
# 1. Does the scrape work? (1 = yes, stop touching scrape config)
up{job="node"}

# 2. Does the exact series exist in storage?
http_requests_total{job="api", instance="web-01:8080"}

# 3. List every metric name this job actually stores
# GET /api/v1/label/__name__/values?match[]={job="api"}
📊 Production Insight
A team restarted Prometheus for an empty panel and caused a 20-minute WAL replay gap across all jobs. The cause was a renamed counter. Checking up first and then the query would have found it in two minutes with zero downtime.
🎯 Key Takeaway
up == 1 closes the scrape investigation; the bug is in names, matchers, ranges, or staleness.

Metric Names Lie: Typos and Silent Renames

Metric names change more often than you'd think. Client libraries add _total suffixes, exporters rename seconds to duration, and major version bumps restructure whole families. Dashboards written last year query names that no longer exist, and each one fails silently as No Data.

The fix starts at the source of truth: the /metrics endpoint on the target. Curl it, grep for your metric's prefix, and copy the exact name including suffixes. Then confirm storage agrees by listing label values for __name__ with a match selector on your job — if the name isn't there, no query can return it.

Protect yourself going forward. Diff /metrics output in CI when upgrading exporters, and create recording rules that alias old names to new ones during migrations. Aliases keep history continuous and let dashboards migrate at your pace instead of during an incident.

OpenMetrics suffixes explain a whole class of renames. Counters gain _total, gauges stay bare, info metrics end _info, and created timestamps end _created. Client libraries append these automatically, so dashboards written against bare names break on upgrade. The label-values API settles every dispute: GET /api/v1/label/__name__/values with a match[] selector for your job lists exactly what storage holds. When the name exists on /metrics but not in storage, a metric_relabel_configs drop or honor_labels clash sits between — check both before blaming the exporter. Keep a short list of canonical names per job in your runbook and diff it after upgrades.

BASH
1
2
3
4
5
6
7
# On the target host — the source of truth for names
curl -s localhost:8080/metrics | grep -E '^http_(server_)?requests' | head

# recording rule aliasing old name to new during migration
# rules/checkout.yml
- record: job:http_requests_total:sum
  expr: sum by (job) (http_server_requests_total)
📊 Production Insight
A Go client upgrade renamed http_requests_total overnight. The team found it by curling /metrics during the incident — a CI diff of exposition output would have caught it before deploy.
🎯 Key Takeaway
Copy names from /metrics output and confirm via the label-values API; alias renames with recording rules.

Label Matchers That Match Nothing Together

Label matchers areANDed: every matcher must match the same series. One wrong matcher — a missing port on instance, a version label that changed, a team label the series never had — narrows the result to the empty set even when each matcher looks plausible alone.

Debug by subtraction. Run the bare metric name first; if that returns data, add matchers back one at a time until the result empties. The last matcher added is the liar. This binary search takes under a minute and beats staring at four matchers trying to spot the wrong one.

The instance label causes most of these. Prometheus sets it to host:port from the scrape target, so web-01 never matches web-01:8080. When in doubt, query the series without instance, read the actual instance value from the result, and paste it back exactly.

Regex matchers deserve respect and suspicion in equal measure. The =~ operator matches RE2 patterns fully anchored, so status=~"5.." catches every 500-level code while status=~"5" matches nothing — anchoring surprises almost everyone once. Negative matchers (!=, !~) select series lacking the label too, which quietly widens results when labels are absent. Whitespace inside label values is significant: a trailing space in an annotation becomes part of the value and defeats exact matchers. When subtraction debugging stalls, render candidates with count by for each label value and read the actual strings. One glance at real values beats ten minutes of guessing at patterns.

📊 Production Insight
An alert queried instance="db-01" while the target was db-01:9100. It never fired for 4 months, including through a real disk-full event. Subtraction debugging found it in 30 seconds.
🎯 Key Takeaway
Remove matchers one at a time until data returns — the last removed matcher was wrong.

Scrape Interval vs Graph Range Mismatch

Timing faults look exactly like missing data. A job scraped every 60s drops one sample per minute, so a 1-minute graph range frequently contains zero samples and range functions like rate() that need two points return nothing. The target is healthy; the window is just too small to see it.

Size the view to the scrape. Keep the panel range at 4x the interval or wider, and set min step at or above the interval so each evaluation step contains a sample. For 60s jobs that means 5m ranges and 60s steps as a floor. Grafana's auto interval usually handles this, but custom min intervals override it and cause these gaps.

Evaluation alignment matters too. Prometheus evaluates range queries at step boundaries that rarely align with scrape times, so tight ranges flicker between one sample and none on each refresh. Wider ranges absorb the jitter and the flicker disappears without touching any config.

Two timing rules govern every range query. First, Prometheus marks series stale after 5 minutes without samples, so ranges crossing a restart show honest gaps no window can fill. Second, rule evaluation adds its own jitter: groups evaluate on their interval boundaries, and a 15s group reading a 60s job sees single-sample windows half the time. Size range windows for the slowest job in the query, not the fastest — mixed-interval dashboards gap on their slowest leg. When a panel mixes jobs, consider splitting it: one panel per interval keeps each range honest instead of compromising all of them.

PROMQL
1
2
3
4
5
6
# Gappy: 1m range on a 60s scrape interval (often 0-1 samples)
rate(http_requests_total[1m])

# Fixed: 5m range with a step >= the 60s scrape interval
rate(http_requests_total[5m])
# Grafana panel: Min interval = 60s, Range = last 15m or wider
📊 Production Insight
A 5-minute executive dashboard on a 60s-scrape job flickered to No Data every other refresh. Changing the range from 1m to 15m fixed it permanently — zero config changes, zero restarts.
🎯 Key Takeaway
Range at 4x the scrape interval and step above the interval; tight ranges on slow jobs flicker to empty.

Stale Markers After Restarts and Renames

When a target restarts, Prometheus marks its old series stale and starts fresh ones, often with a changed instance port or pod name. Queries spanning the restart see two half-series with a gap between them, and dashboards look broken even though collection never stopped.

The offset modifier bridges the gap for display. Appending offset 5m shifts the evaluation window into the past where the old series still has samples, drawing a continuous line across restarts and renames. It's a viewing lens, not a data fix — the underlying series are still distinct.

Keep offsets out of alerting rules. An offset alert evaluates minutes-old data, fires late on real outages, and can miss fresh series entirely. Alert on up == 0 for dead targets and on fresh counter rates for behavior, and let dashboards use offsets for pretty continuity.

Stale markers behave differently by cause, and the difference guides the fix. A target going down stamps staleness after missed scrapes, and the series resumes cleanly when it returns — gaps are truthful history. A restart with new labels (fresh pod name, new port) orphans the old series permanently; the old line ends and a new one begins, and no offset reunites them for alerting. Kubernetes rollouts produce the second kind in bulk, which is why service-level labels outlive pod labels in every durable query. For business continuity, aggregate by stable labels in recording rules so deploys never redraw history. Offsets stay a display convenience for humans, never an input to machines that page.

PROMQL
1
2
3
4
5
6
7
8
# Bridge a restart gap for display (dashboards only)
rate(http_requests_total[5m] offset 5m)

# Alert on the live signal, never on offset data
# alerts/targets.yml
- alert: TargetDown
  expr: up{job="api"} == 0
  for: 5m
📊 Production Insight
After a Kubernetes rollout renamed every pod, 40 dashboards gapped at once. Display offsets restored them in minutes while the team migrated queries to stable service labels — no data was ever lost.
🎯 Key Takeaway
offset 5m bridges restart gaps for display; alerts stay on fresh series and up == 0.

The Five-Minute Checklist That Ends the Incident

Put the checks in order and most incidents resolve in minutes. First confirm up equals 1 so scrape is eliminated. Then verify the exact metric name against /metrics and the label-values API. Then subtract matchers until data returns. Then widen the range past 4x the interval. Only then consider staleness and offsets.

This order is deliberate: each step is cheaper than the next and eliminates a whole class of causes. Name checks take seconds, matcher subtraction takes a minute, range sizing takes a glance at scrape config. Engineers who jump to restarts skip all four cheap steps and pay with WAL replays and extended gaps.

Write the sequence on your runbook. During an incident nobody thinks clearly, and a five-line checklist beats memory. Teams that follow it resolve empty-panel pages in under 10 minutes; teams that freestyle average over an hour with at least one unnecessary restart.

Turn the checklist into a runbook page with copy-paste commands for your fleet: the up query, the /metrics curl, the matcher-subtraction order, the range floor per scrape interval. Link it from the alert so the paged engineer lands on procedure, not a blank search box. Rehearse it in game days by renaming a metric in staging and timing the team — ten-minute resolution should feel routine before production demands it. After each real incident, append the cause and the step that caught it; runbooks compound like interest. The checklist that ends incidents in minutes is built on quiet afternoons, not during them.

⚠ Don't Restart the Server Over a Query Bug
Never restart Prometheus because a panel is empty. Restarts replay the WAL, interrupt all scrapes, and stamp fresh stale markers — converting a 2-minute query typo into a 20-minute ingestion gap for every job on the server.
📊 Production Insight
After adopting this checklist, one team's mean time to resolve empty-panel pages dropped from 70 minutes to 8. Restarts during these incidents went from common to zero in a quarter.
🎯 Key Takeaway
up, then name, then matchers, then range, then staleness — in that order, every time.
● Production incidentPOST-MORTEMseverity: high

The Renamed Counter That Emptied Every Checkout Dashboard

Symptom
All checkout dashboards went empty at 2 PM right after a routine app deploy. Targets showed up, CPU and memory looked fine, and the team burned 3 hours restarting exporters and then Prometheus before checking the metric name.
Assumption
The team assumed the exporter was crashing under load because gaps matched traffic peaks. They restarted the exporter fleet twice and then restarted Prometheus itself, which replayed the WAL for 20 minutes and extended every gap.
Root cause
The Go client upgrade renamed http_requests_total to http_server_requests_total with new labels. Every dashboard and alert still queried the old name, so all panels went empty while up stayed 1. The exporter was healthy the entire time — only the query was stale.
Fix
One engineer curled /metrics and found http_requests_total had become http_server_requests_total after the client library upgrade — a rename in the changelog nobody read. They updated dashboards to the new name, added a recording rule aliasing the old name for history, and put query checks into CI so renamed metrics fail the build instead of production graphs.
Key lesson
  • Client library upgrades rename metrics silently — diff /metrics output before and after every upgrade.
  • Restarting Prometheus for an empty panel converts a query bug into a real ingestion gap for all jobs.
  • Recording-rule aliases preserve history across renames and cost almost nothing to maintain.
Production debug guideFive checks from fastest to slowest that isolate query, timing, and staleness faults.5 entries
Symptom · 01
Panel says No Data but target shows up
→
Fix
Open the expression browser and run up{job="your-job"}. If it's 1, scraping works — stop touching scrape config. Then run your exact panel query there. An empty result in the browser proves the query or timing is wrong, not the target, and saves you from restarting anything.
Symptom · 02
Query returns empty in the expression browser too
→
Fix
Curl the target directly: curl -s target:9100/metrics | grep your_metric. If grep finds nothing, copy the closest real name from the output and paste it into your query. Exporter upgrades often rename metrics (adding _total or _seconds), so never trust the name in an old dashboard.
Symptom · 03
Metric exists on target but the query with labels is empty
→
Fix
Remove matchers one at a time, starting with instance, and re-run after each removal. When data appears, the matcher you just removed was wrong — usually a missing port (db-01 vs db-01:9100) or a label the series doesn't carry. Re-add only matchers you've verified against real series.
Symptom · 04
Query works for some jobs but gaps on slow-scrape jobs
→
Fix
Check the panel range against the job's scrape_interval in prometheus.yml. For a 60s interval, set the range to 5m or wider and the min step to 60s or above. Range functions like rate() need at least two samples inside the window, and short ranges on slow jobs often contain zero.
Symptom · 05
Gaps line up exactly with deploys or restarts
→
Fix
Run your counter query with offset 5m appended inside the range selector, e.g. rate(http_requests_total[5m] offset 5m). If the line returns, stale markers from a restart or rename are the cause. Use the offset for display continuity only — keep alerts on up == 0 and fresh series so real outages still page.
No-Data Causes Compared — Find Yours Fast
Root CauseHow to ConfirmFixPrevention
Metric name typo or rename/metrics on target lacks the name; label values API agreesCopy exact name; alias renames with recording rulesDashboards reference rule names; test queries in CI
Over-strict label matchersRemoving one matcher returns data instantlyMatch only labels the series carriesCopy matchers from working queries; review label changes
Range shorter than scrape intervalWidening range or step restores the lineRange >= 4x interval; step >= intervalTemplate min-step from interval; document per-job ranges
Stale markers after target restartup flaps 1/0; gaps align with deploysKeep old series with or; alert on up, not absenceStable instance labels; rolling restarts; alert on up == 0
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
up{job="node"}up Means the Scrape Worked
curl -s localhost:8080/metrics | grep -E '^http_(server_)?requests' | headMetric Names Lie
rate(http_requests_total[1m])Scrape Interval vs Graph Range Mismatch
rate(http_requests_total[5m] offset 5m)Stale Markers After Restarts and Renames

Key takeaways

1
Up == 1 only proves the scrape worked
verify the metric name itself next.
2
Copy metric names from /metrics output; never type them from memory.
3
Strip label matchers one at a time until data returns.
4
Keep graph ranges above 4x the scrape interval and steps above the interval.
5
Stale markers after restarts are normal
alert on up, bridge renames for display only.
6
Never restart Prometheus for an empty panel; restarts turn typos into real data gaps.

Common mistakes to avoid

5 patterns
×

Typing metric names from memory instead of copying them

Symptom
Panels fail after every exporter upgrade because counters gained suffixes like _total or _seconds that nobody rechecked.
Fix
Copy the exact metric name from /metrics output or query prometheus_api with match[]. Keep a dashboard annotation of rename commits, and use recording rules to alias old names during migrations.
×

Stacking label matchers that match nothing together

Symptom
Each matcher works alone but the combined query is empty, usually because instance includes a port you forgot.
Fix
Write matchers that select exactly what exists: {job="node", instance="db-01:9100"}. Test each matcher in the expression browser and remove labels that don't exist on the series.
×

Using a graph range shorter than two scrape intervals

Symptom
Single-stat panels and 5-minute graphs flicker between values and no-data on every refresh.
Fix
Set the panel range to at least 4x the scrape interval and the min step above the interval. For 60s scrapes, never use a 1m graph with a 15s step.
×

Hiding gaps with or vector(0) in alerting rules

Symptom
Alerts stop firing during real outages because the query now always returns a value.
Fix
Use or vector(0) only for display continuity on low-traffic counters, and never inside alerting or recording rules where fake zeros corrupt rate() math.
×

Restarting Prometheus before checking the query

Symptom
Restarts wipe the head and WAL replay, turning a 2-minute query typo into a 20-minute data gap for everyone.
Fix
Check the API response for status success with empty result, then check target labels and __name__ via /api/v1/label/__name__/values. The query layer and the scrape layer are different systems — debug them separately.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Target shows up but the graph is empty. Where do you look first?
Q02JUNIOR
How do you tell a typo apart from staleness?
Q03SENIOR
Why does a 60s-scrape job show gaps on a 1-minute graph?
Q04SENIOR
Your query has four label matchers and returns nothing. What's the metho...
Q05SENIOR
Explain what offset 5m does for a renamed counter, and its cost.
Q01 of 05JUNIOR

Target shows up but the graph is empty. Where do you look first?

ANSWER
Up is a synthetic series recording whether the last scrape succeeded. It says nothing about which metrics the endpoint exposes. An exporter can return HTTP 200 with renamed or missing metrics while up stays 1, so empty graphs with up == 1 point at the query or the exposition, not the scrape.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Does up == 1 guarantee my metric exists?
02
How long after a restart until stale gaps fill in?
03
Should I use or vector(0) to fill gaps?
04
Can federation or relabeling hide my series?
05
What lookahead should offset queries use?
06
How do I list metric names Prometheus really has?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Prometheus. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Prometheus Error on Ingesting Out-of-Order Samples
2 / 4 · Prometheus
Next
Prometheus Cardinality Explosion From a Label
→