Log4Shell JNDI RCE — Patch Log4j and Scan Dependencies
Log4Shell turned log messages into remote code execution via JNDI lookups.
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
- ✓Java services (or any stack) with third-party dependencies
- ✓Access to build pipelines and artifact storage
- ✓Basic familiarity with vulnerability scanners and WAFs
- Log4Shell (CVE-2021-44228) let text written to logs trigger JNDI lookups in vulnerable Log4j 2 versions, turning user input into server-side requests
- Patch first: upgrade Log4j 2 to a fixed release in the 2.17 line or later, and treat every copy — transitive and shaded — as in scope
- Use WAF rules that block the lookup indicator only as a stopgap while you patch; filters don't fix the library
- Keep a software bill of materials (SBOM) and scan dependencies on every build so the next logging-library flaw finds you fast
- Hunt historically: search archived logs for the lookup indicator to scope whether anyone probed you before the patch landed
Imagine a librarian who files every note handed to her — but whenever a note contains the phrase 'please fetch,' she leaves the building to fetch whatever's named. Attackers started handing in notes that said exactly that, naming addresses they controlled. That's Log4Shell: a logging library that treated message text as instructions and performed lookups it never should have. The fix: file notes without reading them as orders — and check every branch got the memo.
In December 2021, a flaw in one of the world's most-used Java logging libraries redrew everyone's weekend plans. Log4Shell meant that a string — typed into a chat box, a username field, a user-agent header, anything an application logged — could make a vulnerable server perform a network lookup and load remote code. Logging, the thing you do to understand incidents, had become the incident.
The vulnerability sat in Log4j 2's message substitution: when log text contained a lookup expression, the library resolved it, including directory lookups that reached across the network. Because applications log untrusted input constantly — usernames, URLs, form fields, headers — the attack surface was essentially 'anything your app writes down.' Security teams spent weeks finding every Java service they owned, because the library hides inside frameworks, shaded jars, and vendor products.
This guide stays firmly on the defensive side: how the mechanism worked (so you recognize its shape), which versions were affected and the patch discipline that ends it, why WAF rules are only a stopgap, and the SBOM-plus-scanning habit that turns the next library emergency from a forensic dig into a database query. No exploit details — just the response playbook.
How a Log Line Became Code Execution
Logging libraries sit in a uniquely trusted spot: every layer of your app hands them untrusted text — usernames, URLs, error details, headers — and they handle it millions of times a day. Log4j 2 added a feature called message substitution, where special expressions in log text were resolved at render time: dates formatted, environment values inserted, directory service names looked up. That last capability, JNDI lookups, is where routine logging crossed into remote code execution.
JNDI is Java's directory-access API: given a name, it can consult remote servers to resolve it. When a vulnerable Log4j rendered a message containing a lookup expression naming an attacker-controlled server, it connected out and loaded what it was given. The attacker never needed an account, a special endpoint, or even a request the app considered important — anything the app wrote to its logs was a delivery channel, including fields like usernames that get logged on every failed login.
The defensive takeaway is architectural, not anecdotal. Any component that interprets data as instructions — log renderers, template engines, expression languages, deserializers — must default to treating untrusted input as inert text. When you evaluate libraries, ask what their rendering path can trigger: network calls, class loading, and command execution have no business in a log statement. The safest log pipeline is a dumb one: format strings, append, ship.
Affected Versions and the Patch Discipline That Ends It
The vulnerable range ran from Log4j 2.0-beta9 through 2.14.1, with the initial fix in 2.15.0 and further hardening in 2.16.0 and the 2.17 line — each follow-up closing bypasses and disabling the dangerous capability by default. The practical rule for your fleet is blunt: anything on the Log4j 2 line older than the fixed releases gets upgraded, and you verify the upgrade by inspecting what's actually deployed, not what the ticket says.
Version discipline fails in three predictable places. Transitive dependencies pull old copies through frameworks you didn't choose directly — your pom looks clean while the dependency tree quietly includes the vulnerable jar. Shaded jars bundle renamed copies inside vendor SDKs where manifest scanners can't see them. And running processes keep serving the old classes until restarted, so file replacement without restart validation is theater. Your patch runbook must cover all three: tree, contents, and restarts.
Make the fix durable with dependency governance. Pin logging versions in a managed BOM (bill of materials) pom, fail builds that resolve banned versions, and regenerate SBOMs on every build so the next emergency starts with a query instead of a treasure hunt. The snippet below shows the Maven-side habit: enforce the floor version so no transitive downgrade can sneak the vulnerable line back in.
WAF Rules as a Stopgap, Never the Fix
When Log4Shell broke, WAF rules matching the lookup indicator were the fastest relief available: deploy once at the edge, block the obvious probes, buy the weekend to patch properly. That role — time buyer — is the only honest way to describe them. A filter that recognizes one spelling of malicious input can't remediate a library that executes whatever it receives.
Know the blind spots by name so nobody mistakes coverage for closure. Encoded and obfuscated variants slip past literal string matches. Encrypted traffic (TLS to your backends, internal mTLS) is opaque to edge inspection. Internal paths — queue consumers, batch jobs, service-to-service calls carrying logged user content — never cross the WAF at all. And legitimate-looking traffic that the filter waves through still reaches the vulnerable code with full effect.
Use stopgaps with stopgap hygiene: document what each rule covers and what it can't see, alert on rule hits as threat intelligence (probe volume tells you you're being targeted), set an expiry review so temporary rules don't fossilize into false confidence, and report them separately from remediation metrics. The rule count going up is not the vulnerable-host count going down. Patching is the fix; everything else is weatherproofing while the roofers drive over.
SBOM and Dependency Scanning as a Daily Habit
The teams that answered Log4Shell fastest weren't the best patchers — they were the ones who already knew where every library lived. An SBOM (software bill of materials) is that knowledge in machine-readable form: every dependency, every version, every service, regenerated on every build. When the next logging-library CVE lands, the first question ('are we affected and where?') becomes a query instead of a multi-week dig through repos, containers, and vendor bundles.
Build the habit in CI, not in a wiki. Generate the SBOM at build time so it reflects what's actually shipped, scan it against vulnerability feeds on every pipeline run, and fail builds on critical issues in reachable code. Scan contents, not just manifests: match on classes and file hashes so shaded and vendored copies can't hide. Store SBOMs per release so incident response can ask historical questions ('which customers got the build containing version X?') without rebuilding the past.
Extend the inventory past your own code. Vendor SDKs, base images, build plugins, and infrastructure agents all ship libraries — require SBOMs from suppliers where you can, and scan what they deliver where you can't. The pipeline snippet below shows the shape: generate, scan, gate. Ten lines that would have saved the 19-day shaded-jar embarrassment in the incident above.
Hunting the Lookup Indicator in Your Logs
After patching, you still owe yourself an answer: did anyone exploit this before the fix? The lookup indicator — the distinctive lookup-expression marker the vulnerable feature resolved — is your detection anchor. Search current and archived logs for that marker in attacker-controlled fields: usernames, user-agent strings, search boxes, form inputs, header values. A hit means someone knocked; a hit followed by suspicious outbound connections means they may have entered.
Search defensively and completely. Cover archived and cold storage, not just the last 7 days — Log4Shell was probed within hours of disclosure and mass-scanned for months. Include encoded shapes of the marker, since scanners obfuscated it to dodge the very WAF rules everyone deployed. Correlate ruthlessly: a lone probe with no follow-on traffic is reconnaissance, while a probe adjacent to an unknown egress connection escalates the host to isolation and forensic review.
Treat the hunt as a scheduled operation, not a one-off. Keep the detection search running for at least 30 days post-patch to catch slow-burn exploitation and straggler systems (un-rebuilt containers, restored backups, forgotten regions). Document negative results too — 'searched 6TB across these sources, found only blocked probes' is what lets you responsibly close the incident instead of anxiously assuming.
A Patch Rollout Playbook for Logging Libraries
Library RCEs compress into hours what normally takes sprints, so run them from a playbook, not from memory. Split the response into parallel tracks from minute one: patching (upgrade direct, transitive, and shaded copies), stopgaps (indicator filters with documented blind spots), detection (historical indicator hunt plus egress correlation), and communications (status updates that distinguish 'blocked at edge' from 'remediated'). Assign an owner to each track so nothing queues behind anything else.
Sequence the patching track for maximum risk reduction per hour. Internet-facing Java services first, then internal consumers of logged user content (queue workers, analytics pipelines), then everything else including admin tools and test environments — attackers don't respect your environment tiers. Validate each wave by confirming running processes loaded the fixed jars, not just that files changed on disk. A deploy without restart validation is a hope, not a fix.
Close the incident with evidence, not exhaustion. Require the content-based scan to return zero across all artifacts, the SBOM query to agree, the 30-day detection watch to be scheduled, and communications to state residual risk honestly (which stragglers remain, what covers them). Then hold the retrospective while memories are fresh: which inventory gap cost the most time, and what CI gate would have closed it? The playbook you improve today is the weekend you get back next time.
One Shaded Jar Kept 340 Servers Vulnerable for 19 Days
- Scan jar contents, not just manifests: shaded and vendored copies hide the same vulnerable class under different names.
- WAF indicator-blocking is a time buyer with explicit blind spots (encrypted and internal traffic), not a remediation you can close on.
- Don't declare a library incident closed until the SBOM query for the flaw returns zero across every service, including vendor SDKs.
| File | Command / Code | Purpose |
|---|---|---|
| mvn dependency:list -Dincludes=org.apache.logging.log4j:log4j-core \ | Affected Versions and the Patch Discipline That Ends It | |
| .github | name: supply-chain-gate | SBOM and Dependency Scanning as a Daily Habit |
| MARKER='jndi:' | Hunting the Lookup Indicator in Your Logs | |
| JAR=$(ls -t /opt/app/lib/log4j-core-*.jar | head -1) | A Patch Rollout Playbook for Logging Libraries |
Key takeaways
Common mistakes to avoid
5 patternsTrusting manifest-only scans to find every vulnerable copy
Counting WAF blocks as remediated hosts
Replacing files without verifying running processes reloaded them
Searching only recent logs for pre-patch exploitation
Patching production but forgetting admin tools and test environments
Interview Questions on This Topic
In one paragraph, what made Log4Shell so severe?
Frequently Asked Questions
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
That's Supply Chain. Mark it forged?
5 min read · try the examples if you haven't