JWT Expired Still Accepted: Clock Skew Fix Guide
Expired JWTs get accepted through missing exp checks and loose leeway.
20+ years shipping production backend systems. Written from production experience, not tutorials.
- ✓JWT header, payload, and exp basics
- ✓How API auth headers work
- ✓Refresh token concepts
- A token accepted after expiry usually means hand-rolled code never validated exp, or leeway is set far too wide
- Clock skew between issuers and verifiers is real but small; cap leeway at 60 to 120 seconds, not minutes or hours
- Always validate exp, iat, iss, and aud after the signature passes, using the library's built-in claim checks
- Short-lived access tokens (5 to 15 minutes) plus rotating refresh tokens shrink replay and leak windows fast
- Log accept-after-expiry events per service so silent misconfiguration can't hide for months
Think of a JWT like a concert wristband stamped with an expiry time. The bouncer should check the stamp under good light and turn away old bands. But if the bouncer never looks at the stamp, or allows bands from last week because their watch might be off, anyone keeps getting in. That's this bug. The fix is checking every stamp with a good clock, allowing only a minute of wiggle room for watch differences, and handing out wristbands that expire quickly so lost ones stop working fast.
Nothing erodes trust in authentication like learning that expired tokens still open doors. Logout feels broken, offboarding stops working, and leaked tokens stay useful far past their stamped lifetime. Teams usually discover this during an audit or after a stolen token gets replayed days later, when the damage is already done.
Two causes explain nearly every case. Hand-rolled verification decodes the payload and checks the signature but never looks at exp at all. Or the library is configured with a generous leeway meant to cover clock skew, stretched to tens of minutes by someone chasing away occasional failures. Both leave the door open long after it should have closed.
The fix combines strict validation with short lifetimes. Validate exp and the surrounding claims on every request, cap skew tolerance near a minute, and issue access tokens that live for minutes while refresh tokens handle continuity. Clock discipline with NTP keeps servers agreeing on time.
This guide shows how to confirm the hole, tighten validation without breaking honest clients, and build the access-plus-refresh pattern that makes expiry meaningful again.
Why Expired Tokens Get Accepted: the Two Usual Gaps
Accept-after-expiry almost always traces to one of two gaps. The first is hand-rolled verification that decodes the payload, checks the signature, reads the role, and never glances at exp. It feels complete because every line looks security-related, yet the lifetime check is simply absent. Any signed token then works forever, including tokens from deactivated users.
The second gap is library verification with expiry disabled or leeway stretched beyond reason. Someone fighting sporadic failures sets leeway to 600 seconds or flips verify_exp off to calm the dashboard, and the change survives because nothing tests stale tokens. The stamped 30-minute lifetime becomes 40 minutes or infinity while the config comment still says clock skew.
Both gaps share a root habit: trusting issuance instead of verifying consumption. Setting exp at login is only half the job; each consumer must enforce it on every request. Audit the consume path, not the login path, when this bug is suspected.
Confirm with behavior, not reading. Mint a short-lived token, let it expire, and replay it. A 200 means the gap is real regardless of what anyone remembers about the config. That single test converts a vague suspicion into a concrete failing case the team can fix.
Clock Skew Is Real but Small: NTP and Sane Leeway
Servers genuinely disagree about time. Virtual machines pause during migration, containers inherit host clocks, and NTP takes minutes to converge after boot. A token minted at second 100 on the issuer may arrive at a verifier that believes it is second 97. Without tolerance, honest tokens fail at the edges and engineers get paged for phantom outages.
The answer is small, explicit leeway plus healthy clocks. Cap tolerance at 60 to 120 seconds, which absorbs real drift while adding negligible replay risk. Then fix the clocks themselves: run chrony or systemd-timesyncd on every host, monitor offset with chronyc tracking, and block deploys on hosts with unsynchronized clocks. Containers should inherit from synced hosts, not carry their own drift.
Resist the urge to widen leeway when failures appear. A 10-minute tolerance that silences alerts also grants every stolen token 10 bonus minutes on every request. Investigate the skew first with date comparisons and NTP status across issuer and verifier; the fix is usually one unsynced host, not a config change.
Document the chosen value once and enforce it centrally. A shared verify helper with leeway pinned at 90 seconds beats five services with five opinions. Alert when any consumer overrides it, since widening tolerance is a security decision disguised as reliability tuning.
Strict Claim Validation: exp, iat, iss, and aud Together
Expiry alone isn't enough; the surrounding claims stop token misuse across services and time. Validate exp so dead tokens die, iat to catch tokens minted in the future beyond skew, iss to confirm the expected login service issued it, and aud to confirm it was meant for your API. A token minted for the mobile API shouldn't open the admin panel just because the signature verifies.
Libraries do this well when asked. PyJWT's decode with require options and audience checks, java-jwt's acceptExpiresAt plus acceptIssuer, and jose equivalents all enforce the set in a few lines. The failure mode is always the same: defaults left permissive, audience unset because two services share tokens, or issuer unchecked after a migration added a second login path.
Standardize one strict helper per language in your repo and route every consumer through it. The helper pins algorithms, requires the claim set, caps leeway, and rejects unknown kids. New services import it instead of writing their own, which ends hand-rolled drift permanently.
Test the helper negatively: expired tokens, future-dated tokens, wrong issuer, and wrong audience must all receive 401s. Those four tests run in milliseconds and catch the exact misconfigurations that cause silent acceptance in production.
Short Access Plus Refresh: Making Expiry Mean Something
Long-lived access tokens make every leak a long incident. A 24-hour token stolen at 9 AM works until tomorrow morning, surviving logout, password change, and even deactivation when expiry isn't checked. Shortening access to 5 to 15 minutes flips the math: the same leak buys minutes, while legitimate users never notice because refresh happens silently.
The refresh side carries the continuity. Refresh tokens live longer, travel only to the token endpoint, are stored securely (httpOnly cookies for browsers, secure storage for apps), rotate on each use, and revoke as a family on logout or suspicion. The server tracks refresh state, so revocation actually works, while stateless access tokens stay fast and small.
Rollout needs care around concurrency and reuse. Mobile apps with parallel requests can race refresh; accept the previous refresh briefly or serialize renewal. Treat refresh reuse as a theft signal: invalidate the whole family when a used token reappears, since only a cloned token could do that.
Measure the result directly. Plot access-token age at use time; after migration nearly every accepted token should be minutes old. The nightly stale-replay test then guards the property forever: any consumer answering 200 to yesterday's token pages immediately.
Finding Every Consumer That Skips the Check
JWT verification sprawls. The web API checks strictly, the billing worker decodes loosely, and a legacy cron job splits segments by hand. Each consumer makes its own decision about exp, so fleet-wide safety needs an inventory, not assumptions. List every service that accepts tokens before changing anything.
Hunt mechanically. Search for decode, verify, and split calls across repos, flag hand-rolled segment parsing, and record each consumer's leeway and required claims in one table. The stale-replay test then runs against each entry: mint, expire, replay, record 401 or 200. The resulting map shows exactly which services need the strict helper.
Migrate in waves with monitoring. Convert the highest-traffic consumer first, watch 401 rates for honest-client breakage, then proceed down the list. Keep a dashboard of post-expiry accept events per service; it should read zero everywhere and page otherwise.
Lock the win with standards. New services import the shared verifier, CI runs the four negative claim tests, and config changes to leeway need security review. The inventory stays current because the nightly replay proves it, not because someone remembers to update a wiki. Record the decision per consumer so future migrations inherit the map instead of rebuilding it.
Monitoring and Offboarding: Proving Dead Means Dead
Expiry enforcement pays off in offboarding and incident response. When an employee leaves or a password changes, short access plus revoked refresh ends sessions within minutes instead of someday. But only monitoring proves it: log token age at use, alert on post-expiry accepts, and chart the maximum accepted age per service daily.
Build the logout story explicitly. Access tokens can't be individually revoked at scale, so logout revokes the refresh family and lets access die naturally within minutes. Document that window honestly for support and auditors rather than claiming instant global logout. For high-risk events like role removal, pair revocation with a short deny-list of recent access tokens.
Run the numbers after migration. The incident team above watched accepted-token age collapse from 26 days to under 10 minutes on deploy day. Deactivated-account access attempts started returning 401s immediately, and the quarterly audit replay became a non-event.
Keep the evidence flowing: nightly stale replays, per-service accept-age dashboards, and alerts on unknown-kid or future-dated spikes. Dead meaning dead is a property you demonstrate continuously, not a box you check once. Evidence beats memory during audits.
Leaked Tokens Worked 26 Days Past Their Stamped Expiry
- Stamped expiry means nothing without verification. A nightly job that replays a stale token against every consumer proves exp is checked, which config reviews never do.
- Hand-rolled JWT code is where expiry goes to die. Use the library's claim validation on every consumer, and delete custom verifiers instead of patching them.
- Short access plus revocable refresh beats long access every time. Ten-minute access tokens make leaks boring, while refresh rotation keeps users signed in cleanly.
| File | Command / Code | Purpose |
|---|---|---|
| stale_replay_test.py | SECRET = 'test-only-secret' | Why Expired Tokens Get Accepted |
| clock_skew_check.sh | set -euo pipefail | Clock Skew Is Real but Small |
| strict_claims.py | SECRET = 'test-only-secret' | Strict Claim Validation |
| refresh_flow.py | ACCESS_LIFE = 600 # 10 minutes | Short Access Plus Refresh |
Key takeaways
Common mistakes to avoid
5 patternsWriting a custom JWT verifier instead of using the library
Setting leeway to 10 minutes to silence edge failures
Issuing day-long access tokens for convenience
Skipping audience and issuer checks across services
Assuming issuance config proves consumer enforcement
Interview Questions on This Topic
Why would an expired JWT still be accepted?
Frequently Asked Questions
20+ years shipping production backend systems. Written from production experience, not tutorials.
That's Auth. Mark it forged?
5 min read · try the examples if you haven't