Home › Observability › Grafana Datasource Connection Refused: 4 Fixes
Beginner 6 min · September 23, 2026

Grafana Datasource Connection Refused: 4 Fixes

Grafana datasource connection refused? Check proxy mode, server-side URL, firewall egress, and auth headers from Grafana's network..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 8 min
  • ✓Grafana admin access to edit datasources
  • ✓Shell access to the Grafana server or container
  • ✓Backend URL, port, and credentials at hand
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Use proxy mode so queries travel from the Grafana server with server-side secrets instead of browsers fighting CORS
  • Verify the URL from inside Grafana's host or container since localhost there means the container itself
  • Confirm firewall egress from the server with nc because laptop success never proves the server's path
  • Read the test message literally and fix auth headers or TLS trust, then lock the working config in provisioning
✦ Definition~90s read
What is Grafana Datasource Proxy Error?

A Grafana datasource plugin connects the Grafana backend (or the viewer's browser) to a datastore like Prometheus, Loki, or Elasticsearch. In proxy mode the Grafana server issues every query: it resolves the configured URL from its own network, attaches stored credentials, enforces TLS, and returns results to the browser.

★
Think of Grafana as a receptionist placing calls for you.

In direct mode the browser calls the datastore itself, inheriting the viewer's network, DNS, and certificate trust instead.

Connection failures therefore split across four layers. Access mode picks the wrong network path entirely. URL and DNS fail from the Grafana server's vantage even when laptops resolve fine. Firewalls allow office IPs while dropping the server's egress, producing timeouts only Grafana sees.

Auth headers expire or TLS trust breaks at the backend, returning 401s and x509 errors after the TCP connection succeeds.

Test-and-save exercises the full chain and reports the first failing layer, which is why its message is diagnostic gold: refused means TCP, timeout means firewall or routing, 401 means credentials, x509 means trust. Debugging means reproducing each layer from the correct vantage point — the Grafana host for proxy, the browser for direct — and fixing layers in order without batching blind changes.

Plain-English First

Think of Grafana as a receptionist placing calls for you. You can call the warehouse from your own phone (laptop curl) all day, but the receptionist's phone has a different line, a different blocked-numbers list (firewall), and a different speed-dial book (URLs). Direct mode asks visitors to call the warehouse themselves from the lobby — often blocked. Fix the receptionist's phone line and contacts, not yours.

You click Save and Test and get the red banner: Connection refused. The backend is definitely running — you just curled it from your laptop. You tweak the URL, toggle TLS, re-paste the token. Still red, and the dashboards stay dark.

Datasource connections fail in exactly four places: access mode, URL reachability from Grafana's network, firewall rules on Grafana's path, and auth or TLS at the backend. Your laptop's success proves nothing because it walks a different network path with different credentials than the Grafana server does.

This guide tests each layer from the correct vantage point in order. You'll confirm proxy mode, verify the URL from inside Grafana's network, check the firewall from the server, and fix auth headers with test-and-save as your scoreboard. Most red banners turn green in under fifteen minutes once you test from where Grafana stands. Each layer gets one decisive test from the correct vantage point, so bring shell access and expect a green banner soon. Most red banners turn green within fifteen minutes.

Proxy vs Direct: Whose Network Path Counts

Access mode decides whose network path queries travel. Proxy mode sends them from the Grafana server: secrets stay server-side, browsers never touch the backend, and CORS disappears as a concept. Direct mode sends them from each viewer's browser: every laptop, phone, and VPN path must independently reach the backend with working DNS and certificates.

Production backends want proxy, nearly always. Browsers roam across networks Grafana can't control, while the server-to-backend path is one stable, firewallable, monitorable hop. Direct mode survives for local debugging against permissive backends — and causes a disproportionate share of production incidents wherever it escapes that role.

When test-and-save fails, mode is the first toggle because it's the cheapest experiment with the biggest signal. A CORS or browser-fetch flavored error that becomes a clean server-side message under proxy convicts the mode instantly. Set proxy, retest, and let the new message guide the next step.

Browser devtools turn mode confusion into evidence. Open the network tab, reload the dashboard, and watch where queries go: requests to the backend host mean direct mode, requests to /api/ds/query mean proxy. Failed direct requests show CORS or certificate errors verbatim; failed proxy requests show Grafana's error JSON with the server-side cause. Mixed-content blocks (https page calling http backend) appear only in direct mode and vanish under proxy. One devtools screenshot settles every proxy-vs-direct debate in seconds.

📊 Production Insight
An org-wide direct-mode default worked on office wifi and failed on every VPN and phone. Switching 40 datasources to proxy fixed months of unreproducible complaints in an afternoon.
🎯 Key Takeaway
Proxy for production server paths; direct only for local debugging — toggle mode first, it's the cheapest test.

URLs That Resolve From the Server, Not Your Laptop

The URL must resolve from the Grafana server's network, not yours. Hostnames need server-visible DNS, ports must be the backend's listening port (not a laptop port-forward), and schemes must match what the backend serves. Containerized Grafana adds the classic trap: localhost is the Grafana container, where your backend almost certainly isn't listening.

Verify from inside. Exec into the Grafana container or SSH the Grafana host, then curl -v the exact configured URL. Verbose output separates DNS failure (could not resolve), refused (nothing on the port), and TLS mismatch (certificate errors) — three fixes for three messages. Your laptop's result is inadmissible evidence; only the server's path counts.

In clusters, use service DNS names like prometheus.monitoring.svc:9090 that survive rescheduling. In VMs, use stable private DNS or IPs from infrastructure code. Whatever you choose, write it in a provisioning file so the working URL is versioned instead of remembered.

Cluster DNS adds its own traps beyond localhost. CoreDNS resolves service names only inside the cluster network, so external Grafana needs ingress or an exposed endpoint, not a .svc name. Headless services return multiple A records where clients expect one — use the ClusterIP service for datasource URLs. Port names and numbers drift across chart versions; verify the listening port with the backend's own status endpoint, not the chart README. When DNS works from some pods but not Grafana's, compare resolv.conf and ndots settings across namespaces. DNS bugs wear many masks; curl -v unmasks all of them.

YAML
1
2
3
4
5
6
7
8
9
# provisioning/datasources/prometheus.yml — proxy is the default to keep
apiVersion: 1
datasources:
  - name: Prometheus-Prod
    uid: prod-prometheus
    type: prometheus
    access: proxy
    url: http://prometheus.monitoring.svc:9090
    isDefault: true
📊 Production Insight
A team debugged refused errors for 4 hours from laptops before exec'ing into the Grafana container. The first curl there showed localhost resolving to the container itself — diagnosis took 30 seconds from the right shell.
🎯 Key Takeaway
curl -v from the Grafana host; DNS, refused, and TLS messages each name their fix.

Firewall Egress on Grafana's Path

Firewalls distinguish connection refused from connection timeout, and the difference directs the fix. Instant refused means packets arrived and nothing listens — wrong port or dead backend. A hang ending in timeout means packets were silently dropped — firewall or security group on the Grafana-to-backend path.

Test with nc from the Grafana host to isolate the layer. Then compare against your laptop: laptop-open plus server-blocked is the signature of egress rules scoped to office IPs that never learned the Grafana server's address. Cloud moves, cluster migrations, and new NAT gateways all silently change the server's egress identity.

Fix in infrastructure code, not clicks. Open the backend port to the Grafana server's egress IPs or security group, commit the rule, and re-verify with nc before touching Grafana again. Firewall fixes made in consoles during incidents get lost in the next Terraform apply — the commit is the fix.

Cloud networking layers each get a vote. Security groups are stateful and usually innocent; NACLs are stateless and block return traffic when misconfigured; egress gateways and NAT instances rewrite source IPs that allowlists must then include. Service meshes add mutual TLS that rejects plaintext scrapes with connection resets instead of helpful errors — check for sidecars when refused appears only on meshed namespaces. PrivateLink and peering routes change silently during account migrations. Map the full path from Grafana's pod to the backend port, and test each hop instead of assuming one.

BASH
1
2
3
4
5
6
7
# From inside the Grafana container/host — the only test that counts
curl -v http://prometheus.monitoring.svc:9090/-/healthy
nc -vz prometheus.monitoring.svc 9090

# Instant refused = wrong port or nothing listening
# Hang then timeout = firewall drop (next section)
# Could not resolve = DNS gap on the server network
📊 Production Insight
After a NAT gateway change, Grafana's egress IP shifted and every datasource timed out while laptops worked fine. One security-group rule to the new egress range restored all panels — found by nc timing out from the server in 60 seconds.
🎯 Key Takeaway
Refused is wrong port; timeout is firewall — fix in infra code and re-verify with nc.

Auth Headers and TLS Trust at the Backend

With the network clear, auth and TLS get their turn. Expired bearer tokens produce 401s, trimmed secrets produce 403s, and backend CA rotations produce x509 unknown-authority errors — each message naming its layer. Read test-and-save literally instead of retrying blindly.

Store secrets where rotation is possible. Grafana's secure JSON data fields keep tokens out of plain JSON and dashboard exports; better setups inject them from a vault at provision time. Whatever the store, the rotation calendar matters more than the store — 90-day expiries cause 90-day incidents without reminders.

For TLS, match trust chains deliberately. Internal CAs need their bundle on the Grafana server; public backends need valid unexpired chains. Skip-verify exists for scoped debugging only — a permanent skip on a production datasource trades an error message for silent exposure.

Authentication stacks deeper than one header. Forward-auth proxies (oauth2-proxy, cloud IAP) sit in front of backends and return 403s that look like backend rejections — check for auth-request annotations on the ingress. TLS SNI must match the certificate when backends serve multiple names; wrong SNI yields handshake failures identical to trust errors. Custom CA bundles need mounting into the Grafana container and referencing in provisioning, not just present on the host. Rotate certificates with overlap periods and monitor expiry with blackbox probes. Auth layers fail closed and quiet; log each one during setup.

YAML
1
2
3
4
5
6
7
8
9
10
11
12
13
# provisioning with secure fields — token never in plain JSON
apiVersion: 1
datasources:
  - name: Prometheus-Prod
    uid: prod-prometheus
    type: prometheus
    access: proxy
    url: https://prometheus.example.com
    secureJsonData:
      httpHeaderValue1: $__env{PROM_BEARER_TOKEN}
    jsonData:
      httpHeaderName1: Authorization
      tlsAuthWithCACert: true
📊 Production Insight
A 90-day-old service token expired on a Saturday and darkened every dashboard. Secure-field storage plus a rotation alert would have made it a Tuesday-morning task instead of a weekend page.
🎯 Key Takeaway
401 is credentials, x509 is trust — store secrets securely and rotate on a calendar.

Test-and-Save Diagnostics as a Scoreboard

Test-and-save is your scoreboard — learn to read it like one. Connection refused names the TCP layer, i/o timeout names routing or firewall, 401/403 name credentials, x509 names TLS trust, and 404 on query paths names a wrong URL suffix. Each message maps to exactly one section of this guide.

Work the messages in order without skipping. Fix mode, verify URL from the server, clear the firewall, then repair auth — retesting after each change. Random multi-knob tweaking between tests corrupts nearly-correct configs and destroys the evidence trail.

After green, instrument. A synthetic panel alerting on datasource error rates catches the next expiry or firewall change before wallboards darken. The alert costs one query and pays for itself the first time a token ages out quietly.

Instrument datasource health instead of trusting green badges. A synthetic panel querying vector(1) through each production datasource with an alert on No Data pages when any backend path breaks. Blackbox probes from outside the cluster verify the same URLs on a schedule and record history the badge never keeps. Track test-and-save outcomes in CI after every provisioning change: apply, test, alert on failure. These three checks convert silent expiries and firewall edits into daytime tickets with names attached. Monitoring the monitoring is not paranoia; it is the only monitoring that cannot lie to you.

BASH
1
2
3
4
# From the Grafana host: backend health, then Grafana's view of the datasource
curl -s http://prometheus.monitoring.svc:9090/-/healthy
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  http://grafana:3000/api/datasources/uid/prod-prometheus | head -c 400
📊 Production Insight
A team that retested after every single change found their fault in 3 tests (mode, then localhost URL, then firewall). A parallel team that batched changes needed 11 attempts and a config restore — sequencing beats speed.
🎯 Key Takeaway
One fix per test, read each message literally, then alert on error rates to stay green.

Locking In a Green Test-and-Save

Lock the working state so the incident can't recur. Commit the provisioning file with mode, URL, and secret references; record rotation dates in the datasource description; and add the server-network curl to onboarding docs. Future moves re-run the same verification instead of rediscovering it.

Review access modes quarterly. Direct-mode datasources creep back through clones and experiments, each one a future CORS ticket. A provisioning lint that rejects direct mode on production UIDs keeps the standard automatic.

Finally, treat every red banner as a message, not a mood. The five tests in this guide resolve nearly every case in order, and the discipline of testing from Grafana's network generalizes to every backend Grafana will ever proxy — Loki, Tempo, Elasticsearch alike.

Treat dashboards and datasources as code with the same rigor as apps. Provisioning files live in version control beside the dashboards they serve; grizzly or similar tooling syncs both on merge. Pull-request previews spin ephemeral Grafana with the branch's datasources so reviewers click panels instead of reading JSON. Rollbacks revert one commit and restore the last green state in minutes. Teams running dashboards-as-code stop fearing migrations because every move is rehearsed, reviewed, and reversible. Clicks are for exploring; commits are for keeping.

⚠ Test From Where Grafana Stands
Never debug datasource connectivity from your laptop alone. Laptops and Grafana servers live on different networks with different DNS, firewalls, and credentials. Every test that matters runs from the Grafana host or container — the laptop result is inadmissible.
📊 Production Insight
After provisioning all datasources with pinned URLs and proxy mode, one org's datasource pages dropped from monthly to zero in a year. The single recurrence was caught by the error-rate alert two days before token expiry.
🎯 Key Takeaway
Provision the working config, lint modes quarterly, and keep server-network checks in onboarding.
● Production incidentPOST-MORTEMseverity: high

localhost Died When Grafana Moved Into a Container

Symptom
Every Prometheus panel failed at 10 AM after Grafana's move to Kubernetes, all with connection refused. Prometheus was healthy, laptop curls worked, and two backend rollbacks changed nothing over 4 hours.
Assumption
The team assumed the Prometheus migration had broken the backend because Grafana failed at the same hour. They rolled the backend forward and back twice while the real fault sat in Grafana's config: the migration runbook said localhost, which was true on the old host and false in the new container.
Root cause
The datasource URL pointed at localhost:9090, valid when Grafana ran on the monitoring host beside Prometheus. After Grafana moved into a container, localhost meant the Grafana container itself, where nothing listens on 9090. Direct access mode compounded it by pushing queries to browsers that couldn't reach the cluster network either.
Fix
They changed the URL to the in-cluster service DNS name, switched the datasource to proxy mode, and moved the bearer token into secure JSON data. Green test in 10 minutes. Then they provisioned the datasource from a file with the stable URL and added a CI curl from the Grafana network namespace.
Key lesson
  • Runbook URLs must be topology-aware: localhost dies the moment Grafana containerizes.
  • Test-and-save from the wrong mental model wastes hours — always verify from the Grafana host.
  • Provisioned datasource files make working connections reproducible instead of tribal knowledge.
Production debug guideFive tests from the correct vantage point, mode first and lockdown last.5 entries
Symptom · 01
Save and Test fails with browser or CORS flavored errors
→
Fix
Open the datasource config and set access to proxy (server) mode. Save and test. If the error changes from a browser CORS or fetch failure to a server-side message, mode was the bug — keep proxy and continue down this list only if it still fails.
Symptom · 02
Backend works from laptop but Grafana reports refused
→
Fix
Exec into the Grafana host or container and run curl -v for the exact backend URL plus nc -vz host port. Success here clears URL, DNS, and routing; failure here (while your laptop works) proves the server network path is broken. Fix the URL to a server-reachable host:port — service DNS names in clusters, never container localhost.
Symptom · 03
curl from Grafana host fails or hangs
→
Fix
From the Grafana host, time the connection: instant refused means nothing listens or the port is wrong; a hang then timeout means firewall drop. Ask for the backend port opened to the Grafana server's egress IPs, then re-test with nc before touching Grafana config again.
Symptom · 04
Network checks pass but tests return 401 or TLS errors
→
Fix
Move tokens into secure JSON data fields, re-paste the current token, and save-and-test while watching backend auth logs. A 401 means expired or wrong credentials; x509 means TLS trust — match the backend CA on the Grafana server. Confirm one green test, then set a rotation reminder and an error-rate alert.
Symptom · 05
Green test achieved; keep it green
→
Fix
Record the working URL, mode, and rotation date in the datasource description and provisioning file. Add a synthetic panel alerting on datasource error rates so the next expiry or firewall change pages before wallboards go dark.
Connection-Refused Causes Compared
Root CauseHow to ConfirmFixPrevention
Proxy vs direct mode mismatchError mentions CORS or browser fetch failureUse proxy mode for server backendsStandardize proxy; document the exception
Wrong URL or DNS from server netcurl from Grafana host fails; laptop worksFix URL to reachable host:port; add DNSHealth check from server net in CI
Firewall blocking Grafana egressnc timeout from server; telnet hangsOpen backend port to Grafana IPsInfra-as-code firewall rules; review on move
Expired auth header or TLS mismatch401/403 or x509 errors in test messageRotate token; fix CA/skip-verify scopeRotation calendar; alert on error rate
⚙ Quick Reference
3 commands from this guide
FileCommand / CodePurpose
apiVersion: 1URLs That Resolve From the Server, Not Your Laptop
curl -v http://prometheus.monitoring.svc:9090/-/healthyFirewall Egress on Grafana's Path
curl -s http://prometheus.monitoring.svc:9090/-/healthyTest-and-Save Diagnostics as a Scoreboard

Key takeaways

1
Test from Grafana's network path, never your laptop
different networks, different rules.
2
Use proxy mode for production; direct mode belongs to local debugging.
3
Read the test-and-save message literally
refused, timeout, 401, and x509 each name a layer.
4
Localhost inside containers is the container itself
use service DNS names.
5
Store secrets in secure fields, rotate on schedule, alert on error rates.
6
Verify with curl and nc from the Grafana host before changing config.

Common mistakes to avoid

5 patterns
×

Choosing direct access for a production datasource

Symptom
Dashboards work on office wifi and fail everywhere else, with CORS and mixed-content errors nobody can reproduce consistently.
Fix
Use proxy mode for all server-side backends and reserve direct mode for local debugging only. Document the one sanctioned mode in onboarding so nobody rediscovers the browser path in production.
×

Using localhost URLs inside containerized Grafana

Symptom
Test-and-save fails with connection refused while the backend is demonstrably running — on a different network namespace.
Fix
Point the URL at the in-cluster service name and port, and verify from inside the Grafana container with curl. Localhost in a container is the container, not your laptop.
×

Pasting tokens into plain custom headers and forgetting them

Symptom
Everything breaks 90 days later when the token expires, and nobody remembers which header holds the secret.
Fix
Store tokens in Grafana's secure JSON data and rotate them on schedule. Test-and-save after every rotation and alert on datasource error-rate spikes.
×

Testing connectivity from a laptop instead of the Grafana server

Symptom
Curl works locally while Grafana fails, because the firewall allows your IP and blocks the server's.
Fix
Open the backend security group or firewall to the Grafana server's egress IPs for the backend port. Verify with nc from the Grafana host, not your laptop.
×

Retrying test-and-save blindly instead of reading the error

Symptom
Ten retries with random tweaks corrupt a nearly-correct config into a fully broken one.
Fix
Read the test-and-save message literally — timeout, refused, TLS, and 401 each name the layer. Match the message to this guide's row before changing config.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Explain proxy vs direct access mode.
Q02JUNIOR
Test-and-save fails. What does the message tell you?
Q03SENIOR
Why does localhost fail in containerized Grafana?
Q04SENIOR
Laptop reaches the backend but Grafana can't. Now what?
Q05SENIOR
Proxy mode connects but queries return 401. How do you isolate it?
Q01 of 05JUNIOR

Explain proxy vs direct access mode.

ANSWER
Proxy sends queries from the Grafana server, keeping secrets server-side and avoiding browser CORS. Direct sends them from the viewer's browser, exposing network paths and requiring the browser to reach the backend. Production uses proxy; direct is for local debugging.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
When should I use proxy vs direct mode?
02
What does connection refused actually mean?
03
Is curl from my laptop enough to verify?
04
Where should datasource passwords live?
05
Can expired credentials cause refused errors?
06
How do I fix x509 certificate errors?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Grafana. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Grafana Panel Shows No Data Despite Valid Query
2 / 2 · Grafana
Next
Elasticsearch Cluster Health Red: Unassigned Shards
→