SSRF and Cloud Metadata — Block Server-Side Forgery
SSRF tricks your server into fetching attacker URLs, exposing cloud metadata.
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓An app feature that fetches remote URLs (previews, imports, webhooks)
- ✓Basic knowledge of cloud IAM roles and networking
- ✓Access to instance and subnet configuration for hardening
- SSRF happens when your server fetches a URL an attacker chose: webhooks, previews, file imports, and PDF renderers are classic entry points
- Cloud metadata at 169.254.169.254 can hand out temporary credentials, so an SSRF hole on EC2 or GCE can become full account takeover
- Allowlist the exact destination hosts your feature needs and reject everything else, including resolved IPs in private ranges
- Layer egress controls: security groups, proxies, and DNS filtering that block instance-metadata access from application subnets
- Require IMDSv2 with a low hop limit so stolen requests from containers can't reach the metadata service
Imagine you ask a hotel concierge to pick up a package from an address you wrote on a card — and the concierge goes without checking where it is. An attacker slips in a card reading 'the hotel safe room' and the helpful concierge brings back the contents. That's SSRF: your server is the concierge, faithfully fetching whatever URL it's handed. In the cloud, one of those addresses is a special internal desk that hands out master keys, which turns a fetching trick into a building takeover.
Your app has a harmless feature: paste a URL, get a link preview. Or import a file from a URL. Or receive webhooks from a partner. Behind each of these sits server-side code that fetches an arbitrary address — and if any part of that address comes from user input, you've built an SSRF hole: Server-Side Request Forgery.
SSRF lets an outsider aim your server's network access at targets they choose. Scan internal ports. Hit admin panels that listen only on localhost. And in the cloud, query the instance metadata service — a special internal address that happily describes your infrastructure and, worst of all, dispenses temporary credentials for the instance's IAM role. A single crafted URL can escalate from 'fetch my avatar' to 'keys to the kingdom'.
What makes SSRF stubborn is that every fix teams try first is bypassable. Block the metadata IP and attackers use DNS rebinding, decimal IP encoding, or redirects to reach it anyway. Check the URL string but not the resolved address and redirects walk right past. This guide covers the entry points, the metadata service mechanics, and the layered defense — allowlists, egress controls, IMDSv2 — that actually holds when any single layer fails.
How a URL Field Becomes a Server-Side Request
SSRF starts with a feature that seems user-friendly: paste a link, get a preview card. Import your avatar from a URL instead of uploading. Register a webhook and we'll call you back. Each of these hands part of an outbound request to the user — the host, the path, sometimes headers — and the server executes it with the server's network position. Your backend can reach the internal admin panel on localhost:8080, the database on its private IP, and the metadata service. The attacker can't reach those directly, so they borrow your server's legs.
The entry points cluster in a few shapes. Fetchers (previews, imports, converters) take full URLs. Webhook registrations store a callback the server will later call, sometimes with secrets attached. Document pipelines (PDF renderers, SVG processors, XML parsers) resolve embedded external references during conversion — a file upload becomes a request engine. And open redirectors elsewhere in your app become accomplices, laundering attacker destinations through your own domain.
Map these features before you harden them. For each one, answer: which URL parts does the user control, does the fetcher follow redirects, and is DNS re-resolved at fetch time? The snippet below is the hardened fetch helper every one of these features should share — allowlisted hosts, DNS resolved and checked, private ranges rejected, redirects off. One choke point beats ten ad-hoc validators.
The Cloud Metadata Address and Why Attackers Love It
Every major cloud runs an instance metadata service at the link-local address 169.254.169.254 — reachable only from inside the instance, serving data about the machine itself. That data includes instance identity, network config, user-data scripts (which sometimes embed secrets), and IAM role credentials. The credentials are the prize: temporary, auto-rotating, and usable from anywhere on the internet until they expire. Read them once through SSRF and you wear the instance's permissions like a costume.
The shape is identical on every platform: query a well-known URL, get JSON back, extract the temporary credentials. Attackers automate the walk — role name first, credentials second — in seconds. From there the blast radius equals the role's permissions: storage buckets, queues, other instances, sometimes the whole account. The mining-fleet incident earlier is the standard monetization; data theft is the quieter one.
This is why metadata access deserves its own defensive layer regardless of application fixes. Application code changes weekly; the metadata service is forever. Enforce the hardened metadata version, restrict which network tiers can reach the address at all, and keep instance roles minimal so that even a successful read yields little. The fetcher fix stops the knock; the metadata hardening decides what's behind the door.
Allowlists and Egress Controls That Actually Hold
Denylists fail against SSRF because attacker input has infinite spellings: decimal IPs (2130706433 for localhost), octal variants, DNS rebinding that resolves safe-then-evil, and redirect chains that start trusted and end anywhere. Every denylist is a bet you've enumerated all evil; allowlists flip the bet to enumerating the small set of legitimate destinations, which for most features is a handful of hosts.
Build the allowlist at the right layer. In code, check the parsed hostname against an explicit set — never a regex on the raw string — then resolve DNS yourself and reject private, loopback, link-local, and metadata ranges on every resolved address. Handle redirects by re-running the full check on each hop or by refusing to follow them; most preview and import features work fine fetching only the first response. Set aggressive timeouts so attackers can't use your fetcher for slow port scans.
Then assume the code check has a bug and add network controls that don't depend on it. Egress proxy rules that permit only approved external hosts, security groups that deny the metadata address from application subnets, and DNS filtering that sinkholes metadata-looking queries all keep working when the parser slips. The YAML below shows the Kubernetes-flavored version: a NetworkPolicy that confines the app tier's egress so even a compromised pod can't wander the network freely.
IMDSv2 and Hop Limits as Defense in Depth
IMDSv2 (Instance Metadata Service version 2) fixes the protocol mismatch that made SSRF so rewarding: v1 answers plain GETs, which is all SSRF can easily produce, while v2 demands a PUT to mint a session token before any data is served. URL fetchers, XML resolvers, and PDF renderers issue GETs — they can't complete the PUT handshake, so the metadata stays silent even when the request reaches it. Enforcing v2-only (HttpTokens=required) converts most SSRF from credential theft back into mere reconnaissance.
The hop limit (HttpPutResponseHopLimit) handles the container wrinkle. Responses from the metadata service carry a TTL that decrements per network hop; containers add hops between the pod and the host network, so defaults were historically raised to 2 for container hosts — which also lets SSRF payloads running in containers reach the service. Setting the limit to 1 on hosts that don't need container metadata access closes that path; where containers legitimately need it, prefer distinct task roles (like ECS task roles) over instance roles so stolen credentials carry less power.
Roll this out as infrastructure, not advice: set the requirement in launch templates and organization policies so new instances inherit it, then audit existing fleets region by region. The commands below show the verification habit — check what's actually enforced, don't trust the ticket that said it was done. Pair the rollout with a smoke test of every agent (monitoring, config management) that reads metadata legitimately, since v2-only breaks clients too old to send the PUT.
Validating and Sandboxing Outbound Fetches
Some features genuinely need the open web — a feed reader, a link preview for arbitrary sites, a migration importer. Allowlists can't cover 'any URL the user pastes,' so sandbox the fetch instead. Run retrievals in an isolated worker with no cloud credentials in its environment, no route to internal networks, and strict output limits: cap response size, cap time, strip active content from what you store. The fetcher becomes a disposable glove — useful, and safe to contaminate.
Design the sandbox in concentric rings. The worker runs with a dedicated minimal IAM role (or none), in a subnet whose route table has no path to internal tiers and an explicit deny for the metadata address. DNS goes through a resolver that refuses to answer private ranges. The fetched bytes are treated as hostile: images re-encoded, HTML parsed for text only, archives scanned before extraction. What returns to your main app is a sanitized artifact, never the raw response.
This architecture also simplifies incident response. When (not if) someone finds a bypass in your validation, the blast radius is a credential-less sandbox that can only talk to the public internet — annoying, not catastrophic. Log every fetch with source feature, target host, resolved IP, and redirect chain so the bypass announces itself in your telemetry instead of hiding in a generic HTTP client log.
Detecting SSRF Probes in Logs Before They Succeed
SSRF campaigns have a recognizable shape in telemetry: bursts of requests to unusual internal-looking targets, sequential paths or ports from few accounts, and redirect chains that terminate at link-local addresses. Your fetch helper should log the full chain — original URL, each redirect hop, final resolved IP — so analysts see the laundering, not just the innocent first hop. Without chain logging, the attack looks like normal traffic to a trusted domain.
Correlate application logs with network and cloud telemetry. VPC flow logs showing app hosts connecting to the metadata address are a finding on their own — legitimate apps rarely need it per-request. CloudTrail showing role credentials used from unexpected networks means probing already succeeded and the response is now revocation plus rotation. Set the alert threshold on patterns (sequential sweeping, metadata paths, private-range targets), not just volume, since low-and-slow probing stays under volumetric radar.
Rehearse the response so detection converts to containment in minutes: feature-flag the fetcher off, revoke and rotate the exposed role's sessions, and scope with logs before announcing. The 11-hour mining bill in the earlier incident wasn't a detection failure of exotic kind — billing was simply the only alert watching. A metadata-access alert would have fired on the first credential read.
A Link Preview Fetched Credentials and Cost $214,000 in 11 Hours
- String denylists don't stop SSRF: redirects, DNS rebinding, and IP encoding walk past them. Allowlist destinations and re-validate after DNS resolution.
- Over-broad IAM roles turn SSRF into account takeover. Least privilege on instance roles is the control that decides what stolen credentials are worth.
- Billing alerts aren't intrusion detection. Alert on metadata access patterns and credential use from unexpected networks, not just on spend.
| File | Command / Code | Purpose |
|---|---|---|
| app | from urllib.parse import urlparse | How a URL Field Becomes a Server-Side Request |
| if curl -s --max-time 3 http://169.254.169.254/latest/meta-data/ami-id; then | The Cloud Metadata Address and Why Attackers Love It | |
| k8s | apiVersion: networking.k8s.io/v1 | Allowlists and Egress Controls That Actually Hold |
| aws ec2 describe-instances \ | IMDSv2 and Hop Limits as Defense in Depth |
Key takeaways
Common mistakes to avoid
5 patternsBlocking the metadata IP as a string and stopping there
Following redirects without re-validating each hop
Leaving IMDSv1 enabled 'until the migration is scheduled'
Attaching over-broad IAM roles to web instances
Putting fetchers in subnets with internal routes and full credentials
Interview Questions on This Topic
What is SSRF and why is it worse in the cloud?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Cloud. Mark it forged?
6 min read · try the examples if you haven't