Grafana Datasource Connection Refused: 4 Fixes
Grafana datasource connection refused? Check proxy mode, server-side URL, firewall egress, and auth headers from Grafana's network..
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓Grafana admin access to edit datasources
- ✓Shell access to the Grafana server or container
- ✓Backend URL, port, and credentials at hand
- Use proxy mode so queries travel from the Grafana server with server-side secrets instead of browsers fighting CORS
- Verify the URL from inside Grafana's host or container since localhost there means the container itself
- Confirm firewall egress from the server with nc because laptop success never proves the server's path
- Read the test message literally and fix auth headers or TLS trust, then lock the working config in provisioning
Think of Grafana as a receptionist placing calls for you. You can call the warehouse from your own phone (laptop curl) all day, but the receptionist's phone has a different line, a different blocked-numbers list (firewall), and a different speed-dial book (URLs). Direct mode asks visitors to call the warehouse themselves from the lobby — often blocked. Fix the receptionist's phone line and contacts, not yours.
You click Save and Test and get the red banner: Connection refused. The backend is definitely running — you just curled it from your laptop. You tweak the URL, toggle TLS, re-paste the token. Still red, and the dashboards stay dark.
Datasource connections fail in exactly four places: access mode, URL reachability from Grafana's network, firewall rules on Grafana's path, and auth or TLS at the backend. Your laptop's success proves nothing because it walks a different network path with different credentials than the Grafana server does.
This guide tests each layer from the correct vantage point in order. You'll confirm proxy mode, verify the URL from inside Grafana's network, check the firewall from the server, and fix auth headers with test-and-save as your scoreboard. Most red banners turn green in under fifteen minutes once you test from where Grafana stands. Each layer gets one decisive test from the correct vantage point, so bring shell access and expect a green banner soon. Most red banners turn green within fifteen minutes.
Proxy vs Direct: Whose Network Path Counts
Access mode decides whose network path queries travel. Proxy mode sends them from the Grafana server: secrets stay server-side, browsers never touch the backend, and CORS disappears as a concept. Direct mode sends them from each viewer's browser: every laptop, phone, and VPN path must independently reach the backend with working DNS and certificates.
Production backends want proxy, nearly always. Browsers roam across networks Grafana can't control, while the server-to-backend path is one stable, firewallable, monitorable hop. Direct mode survives for local debugging against permissive backends — and causes a disproportionate share of production incidents wherever it escapes that role.
When test-and-save fails, mode is the first toggle because it's the cheapest experiment with the biggest signal. A CORS or browser-fetch flavored error that becomes a clean server-side message under proxy convicts the mode instantly. Set proxy, retest, and let the new message guide the next step.
Browser devtools turn mode confusion into evidence. Open the network tab, reload the dashboard, and watch where queries go: requests to the backend host mean direct mode, requests to /api/ds/query mean proxy. Failed direct requests show CORS or certificate errors verbatim; failed proxy requests show Grafana's error JSON with the server-side cause. Mixed-content blocks (https page calling http backend) appear only in direct mode and vanish under proxy. One devtools screenshot settles every proxy-vs-direct debate in seconds.
URLs That Resolve From the Server, Not Your Laptop
The URL must resolve from the Grafana server's network, not yours. Hostnames need server-visible DNS, ports must be the backend's listening port (not a laptop port-forward), and schemes must match what the backend serves. Containerized Grafana adds the classic trap: localhost is the Grafana container, where your backend almost certainly isn't listening.
Verify from inside. Exec into the Grafana container or SSH the Grafana host, then curl -v the exact configured URL. Verbose output separates DNS failure (could not resolve), refused (nothing on the port), and TLS mismatch (certificate errors) — three fixes for three messages. Your laptop's result is inadmissible evidence; only the server's path counts.
In clusters, use service DNS names like prometheus.monitoring.svc:9090 that survive rescheduling. In VMs, use stable private DNS or IPs from infrastructure code. Whatever you choose, write it in a provisioning file so the working URL is versioned instead of remembered.
Cluster DNS adds its own traps beyond localhost. CoreDNS resolves service names only inside the cluster network, so external Grafana needs ingress or an exposed endpoint, not a .svc name. Headless services return multiple A records where clients expect one — use the ClusterIP service for datasource URLs. Port names and numbers drift across chart versions; verify the listening port with the backend's own status endpoint, not the chart README. When DNS works from some pods but not Grafana's, compare resolv.conf and ndots settings across namespaces. DNS bugs wear many masks; curl -v unmasks all of them.
Firewall Egress on Grafana's Path
Firewalls distinguish connection refused from connection timeout, and the difference directs the fix. Instant refused means packets arrived and nothing listens — wrong port or dead backend. A hang ending in timeout means packets were silently dropped — firewall or security group on the Grafana-to-backend path.
Test with nc from the Grafana host to isolate the layer. Then compare against your laptop: laptop-open plus server-blocked is the signature of egress rules scoped to office IPs that never learned the Grafana server's address. Cloud moves, cluster migrations, and new NAT gateways all silently change the server's egress identity.
Fix in infrastructure code, not clicks. Open the backend port to the Grafana server's egress IPs or security group, commit the rule, and re-verify with nc before touching Grafana again. Firewall fixes made in consoles during incidents get lost in the next Terraform apply — the commit is the fix.
Cloud networking layers each get a vote. Security groups are stateful and usually innocent; NACLs are stateless and block return traffic when misconfigured; egress gateways and NAT instances rewrite source IPs that allowlists must then include. Service meshes add mutual TLS that rejects plaintext scrapes with connection resets instead of helpful errors — check for sidecars when refused appears only on meshed namespaces. PrivateLink and peering routes change silently during account migrations. Map the full path from Grafana's pod to the backend port, and test each hop instead of assuming one.
Auth Headers and TLS Trust at the Backend
With the network clear, auth and TLS get their turn. Expired bearer tokens produce 401s, trimmed secrets produce 403s, and backend CA rotations produce x509 unknown-authority errors — each message naming its layer. Read test-and-save literally instead of retrying blindly.
Store secrets where rotation is possible. Grafana's secure JSON data fields keep tokens out of plain JSON and dashboard exports; better setups inject them from a vault at provision time. Whatever the store, the rotation calendar matters more than the store — 90-day expiries cause 90-day incidents without reminders.
For TLS, match trust chains deliberately. Internal CAs need their bundle on the Grafana server; public backends need valid unexpired chains. Skip-verify exists for scoped debugging only — a permanent skip on a production datasource trades an error message for silent exposure.
Authentication stacks deeper than one header. Forward-auth proxies (oauth2-proxy, cloud IAP) sit in front of backends and return 403s that look like backend rejections — check for auth-request annotations on the ingress. TLS SNI must match the certificate when backends serve multiple names; wrong SNI yields handshake failures identical to trust errors. Custom CA bundles need mounting into the Grafana container and referencing in provisioning, not just present on the host. Rotate certificates with overlap periods and monitor expiry with blackbox probes. Auth layers fail closed and quiet; log each one during setup.
Test-and-Save Diagnostics as a Scoreboard
Test-and-save is your scoreboard — learn to read it like one. Connection refused names the TCP layer, i/o timeout names routing or firewall, 401/403 name credentials, x509 names TLS trust, and 404 on query paths names a wrong URL suffix. Each message maps to exactly one section of this guide.
Work the messages in order without skipping. Fix mode, verify URL from the server, clear the firewall, then repair auth — retesting after each change. Random multi-knob tweaking between tests corrupts nearly-correct configs and destroys the evidence trail.
After green, instrument. A synthetic panel alerting on datasource error rates catches the next expiry or firewall change before wallboards darken. The alert costs one query and pays for itself the first time a token ages out quietly.
Instrument datasource health instead of trusting green badges. A synthetic panel querying vector(1) through each production datasource with an alert on No Data pages when any backend path breaks. Blackbox probes from outside the cluster verify the same URLs on a schedule and record history the badge never keeps. Track test-and-save outcomes in CI after every provisioning change: apply, test, alert on failure. These three checks convert silent expiries and firewall edits into daytime tickets with names attached. Monitoring the monitoring is not paranoia; it is the only monitoring that cannot lie to you.
Locking In a Green Test-and-Save
Lock the working state so the incident can't recur. Commit the provisioning file with mode, URL, and secret references; record rotation dates in the datasource description; and add the server-network curl to onboarding docs. Future moves re-run the same verification instead of rediscovering it.
Review access modes quarterly. Direct-mode datasources creep back through clones and experiments, each one a future CORS ticket. A provisioning lint that rejects direct mode on production UIDs keeps the standard automatic.
Finally, treat every red banner as a message, not a mood. The five tests in this guide resolve nearly every case in order, and the discipline of testing from Grafana's network generalizes to every backend Grafana will ever proxy — Loki, Tempo, Elasticsearch alike.
Treat dashboards and datasources as code with the same rigor as apps. Provisioning files live in version control beside the dashboards they serve; grizzly or similar tooling syncs both on merge. Pull-request previews spin ephemeral Grafana with the branch's datasources so reviewers click panels instead of reading JSON. Rollbacks revert one commit and restore the last green state in minutes. Teams running dashboards-as-code stop fearing migrations because every move is rehearsed, reviewed, and reversible. Clicks are for exploring; commits are for keeping.
localhost Died When Grafana Moved Into a Container
- Runbook URLs must be topology-aware: localhost dies the moment Grafana containerizes.
- Test-and-save from the wrong mental model wastes hours — always verify from the Grafana host.
- Provisioned datasource files make working connections reproducible instead of tribal knowledge.
| File | Command / Code | Purpose |
|---|---|---|
| apiVersion: 1 | URLs That Resolve From the Server, Not Your Laptop | |
| curl -v http://prometheus.monitoring.svc:9090/-/healthy | Firewall Egress on Grafana's Path | |
| curl -s http://prometheus.monitoring.svc:9090/-/healthy | Test-and-Save Diagnostics as a Scoreboard |
Key takeaways
Common mistakes to avoid
5 patternsChoosing direct access for a production datasource
Using localhost URLs inside containerized Grafana
Pasting tokens into plain custom headers and forgetting them
Testing connectivity from a laptop instead of the Grafana server
Retrying test-and-save blindly instead of reading the error
Interview Questions on This Topic
Explain proxy vs direct access mode.
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Grafana. Mark it forged?
6 min read · try the examples if you haven't