Terraform State Lock Error: Unlock Stuck S3 Applies
Copy the Lock ID from the error and run terraform force-unlock, then re-plan.
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓You've run terraform init and apply against an S3 backend at least once
- ✓Basic comfort with the AWS CLI and DynamoDB concepts
- ✓Access to a non-production state file where you can practice safely
- A stuck lock almost always means a previous apply crashed or timed out, leaving its DynamoDB lease behind — not proof someone else is applying now.
- Copy the Lock ID from the error's Lock Info block and run
terraform force-unlockin the same directory and workspace. - Before unlocking, confirm the holder is dead: check CI logs,
ps aux | grep terraform, and the lock's Created timestamp. - Prevent repeats with
-lock-timeouton applies, CI step timeouts, and exactly one writer job per state path.
Think of Terraform state like a library's paper sign-out sheet for a popular book. Before changing anything, you clip your name tag to the sheet so nobody else writes on it at the same time. If you walk out mid-edit and leave your tag clipped on, the next person can't do their work — that's a stuck lock. The fix isn't tearing up the sheet; it's checking you actually left the building, unclipping your tag properly, and making sure the sheet still reads correctly before anyone writes again.
It's Friday at 4 PM and your deploy pipeline just turned red with a message you've never seen: Error acquiring the state lock. Nobody on the team knowingly changed anything about state, the last apply finished an hour ago, and the error helpfully suggests running a command called force-unlock that sounds like it'll either save your evening or end your career.
That fear is healthy, because the lock is load-bearing. Terraform stores your infrastructure as a single JSON state file with no transactions, so before every write it takes a short-lived lease — for S3 backends, a row in a DynamoDB table. If two processes wrote at once, the last writer would win and the loser's resources would silently vanish from state while still running and billing you. The lock is the only thing standing between you and that outcome.
Most stuck locks come from something boring: a CI runner that timed out mid-apply, a laptop lid closed during a long apply, a spot instance reclaimed under the process. The lease outlives its owner and the next run trips over it. This article shows you how to tell a dead lock from a live one, how to release it without corrupting state, and how to set up your backend and pipelines so you stop seeing this error at 4 PM on Fridays.
What the 'Error Acquiring the State Lock' Message Actually Means
When Terraform prints Error acquiring the state lock, it's telling you a write lease it needs is currently held. The Lock Info block that follows is the whole investigation in miniature: an ID (a UUID identifying this exact lease), a Path (which state file is locked), an Operation (usually Apply or Plan), a Who field (user@host of the holder), a Created timestamp, and a Version. Every one of those lines earns its place, and skipping them is how teams unlock the wrong state or interrupt a live deploy.
The reason the lease exists at all is that object storage has no transactions. Your state file is one JSON document; two simultaneous writers would each read it, each write their version back, and the loser would silently erase the winner's resources from state while those resources kept running and billing you. For S3 backends the lease is a row in a DynamoDB table holding the lock ID plus metadata. Terraform writes that row before any state mutation and deletes it when the operation completes or fails cleanly.
So the error never means your configuration is broken — it means someone (or some leftover process) holds the pen. Your first move is always reading, not unlocking: note the ID, check the Path against your backend key, look at Who and Created, and decide whether the holder could still be alive. Only then do you act. Engineers who jump straight to force-unlock are playing roulette with the state file. Take notes as you go — the Lock ID is long and mistyping it wastes a whole cycle.
Lease vs Stuck Lock: Telling a Live Apply From a Dead One
A lease and a stuck lock look identical in the error message — the difference is whether the holder is still alive. A lease is healthy: the Created timestamp is minutes old, the Who field names a teammate whose apply is visibly running, or your CI dashboard shows an in-progress job for the same state path. The correct response is patience, not force: wait, or rerun with -lock-timeout so Terraform queues behind the holder.
A stuck lock is a leftover: the timestamp is old, the named host is a CI runner that no longer exists, the pipeline shows killed or timed out, and no terraform process answers on the machine. That's the signature of a crashed apply — laptop lid closed, spot instance reclaimed, step timeout fired mid-write. The lease outlived its owner and nothing will release it except you.
Build the habit of checking all three signals — process, pipeline, timestamp — before deciding. The DynamoDB item's Info document carries the same ID, operation, and creator as the error, so you can confirm the row matches the complaint and isn't residue from an older incident. Five minutes of verification here is what separates a routine unlock from a state-corruption postmortem. Write down what you find — the next unlock goes twice as fast.
Reading the Lock ID and DynamoDB Lock Row Before You Touch Anything
Before you release anything, collect two artifacts: the Lock ID from the error and the matching row in the DynamoDB table. The ID is a UUID in the Lock Info block — that's the only value force-unlock accepts, and it must match the lease exactly. Copy it carefully; retyping a 36-character UUID by hand is asking for a mismatch error at the worst moment.
Then look at the lock table itself. Its partition key is LockID, and each item carries an Info document with the operation type, creator, creation time, and version — the same fields the error shows. Comparing the two confirms you're looking at the current lease and not residue from an older incident, and the table name in your backend block tells you which table to query. If your backend key is prod/network/terraform.tfstate, the locked state is exactly that path, and the fix must run from the configuration that manages it.
This is also where you catch the classic trap: the lock belongs to a different workspace or environment than your current directory. Staging and production have separate state paths and separate leases. Unlocking from the wrong directory either errors out on ID mismatch or — worse with a forced flag — touches the wrong state. Match the Path line to your backend key first, always.
Running terraform force-unlock Safely in Production
force-unlock is safe when the holder is provably dead and dangerous otherwise — the command itself can't tell the difference, so you provide the judgment. Run it from the same directory and workspace that created the lock, passing the exact Lock ID from the error. Terraform asks for confirmation; in interactive use, type yes rather than bypassing the prompt, because that pause is your last chance to catch a wrong-directory mistake.
In CI, where there's no human to confirm, use the non-interactive flag only inside a job that already verified the holder is dead — never as a blind pre-step before every apply. A scheduled unlock that fires while a real apply runs will let two writers onto one state file, which is precisely the corruption the lock exists to prevent.
After the unlock, always run plan before apply. The crashed run may have created half its resources, and only a fresh plan shows which half survived. Review that plan the way you'd review any production change: unexpected creates suggest orphaned writes, unexpected destroys suggest the crash landed mid-delete. If the plan looks clean, proceed; if it doesn't, investigate before the next write. The unlock fixed the lease, not the interrupted work. Patience here is cheap; state surgery is not.
S3 Backend and DynamoDB Lock Table Settings That Prevent Stuck Locks
The backend settings that make stuck locks rare are unglamorous: a dedicated DynamoDB lock table with on-demand capacity so bursts of applies don't get throttled, server-side encryption on the state bucket, and versioning enabled so a corrupted state file can be restored from history. None of this is exotic, but teams that set state up once and never revisit it usually lack at least one of the three.
On the client side, two flags do most of the work. -lock-timeout tells Terraform to wait for a live holder instead of failing instantly — a few minutes of patience absorbs the normal overlap when two pipelines fire close together. Sensible CI step timeouts do the rest: a timeout shorter than the runner's kill timeout lets Terraform exit cleanly and release its lease instead of being murdered mid-write.
Finally, treat backend configuration as reviewed code. The bucket, key, region, and lock table name should live in version control and change through pull requests like anything else. When region typos or IAM trims masquerade as stuck locks — every command failing, not just concurrent ones — a reviewed backend block gives you a known-good baseline to diff against instead of a mystery. Revisit that baseline quarterly — teams grow, regions multiply, and yesterday's correct backend silently becomes tomorrow's puzzle.
CI Pipelines: Timeouts, Retries, and Lock Hygiene in Automation
Automation is where locks go to multiply, because CI happily runs two applies against one state file if nothing stops it. The fix is a single-writer rule per state path: a concurrency group in GitHub Actions, a resource group in GitLab, or a deploy queue in front of your pipeline. Additional runs wait their turn instead of colliding, which eliminates the most common lock source — overlapping schedules — without any human intervention.
Pair serialization with timeouts on both sides. The CI step timeout should be shorter than the runner's hard kill timeout so Terraform can shut down cleanly and release its lease. The -lock-timeout flag on plan and apply covers the residual overlap when a previous run overruns by minutes rather than hours. Together they convert most would-be incidents into slightly slower green builds.
Last, make failures loud. Alert on every non-zero Terraform exit in CI and include the state path in the alert, so a stuck lock pages someone with context instead of sitting until the next deploy discovers it. Log the Lock ID and the job URL together — future-you, clearing the lock at midnight, will be grateful for the archaeology head start. Review these alerts monthly; a quiet lock table means the hygiene is working.
The Friday Deploy That Left a Phantom Lock — and the Manual Delete That Made It Worse
terraform apply held the DynamoDB lease, so the lease was never released — a classic stuck lock. Then the manual DynamoDB delete removed the lease while the automatic retry was already mid-write, allowing two writers on one state file. The loser's resources dropped out of state while still running in AWS, including a NAT gateway that kept billing until the audit found it.terraform plan -refresh-only to reconcile state with reality and listed every resource whose real-world existence didn't match the file. Then they imported the orphaned resources back into state and destroyed one accidental duplicate NAT gateway the double-write had created. Finally they replaced the console-delete habit with a runbook: verify the holder is dead, force-unlock with the Lock ID, re-plan, and require a second engineer to approve any unlock during business hours.- A killed CI job absolutely leaves side effects — the lock outlives the runner. Every pipeline that applies Terraform needs a step timeout shorter than the infrastructure kill timeout, plus an alert on non-zero exits so a human inspects for stuck locks before retrying.
- Never clear a lock by deleting the DynamoDB row. force-unlock exists because it validates the lock ID against live state; a console delete can't tell a dead lock from a live apply, and guessing wrong corrupts the state file.
- A retried apply after a crash must start with plan, not apply. The crashed run may have created half its resources, and only a fresh plan shows which half — re-applying blind is how you get duplicates.
ps aux | grep '[t]erraform' and by reading the CI job log — a killed or timed-out job means the holder is dead. Then run terraform force-unlock <LOCK_ID> in the same directory and workspace, answer yes, and follow with terraform plan to confirm state health.terraform plan -lock-timeout=10m or terraform apply -lock-timeout=10m. Watch the other run finish, then proceed. If your CI hits this often, put both jobs in one concurrency group so they queue instead of colliding.terraform workspace show and compare your backend key with the Path line in the error. Switch with terraform workspace select <name> or cd into the right directory, then rerun terraform force-unlock <LOCK_ID>. Backend keys differ per environment, so staging and production locks are released independently.TF_LOG=DEBUG terraform plan 2>&1 | grep -i -E 'denied|dynamodb|LockID' to capture the real failure. Then verify IAM with aws dynamodb describe-table --table-name <lock-table> and confirm your role has dynamodb:GetItem, PutItem, and DeleteItem on the lock table. Fix the policy, rerun terraform init, and try again.aws dynamodb scan --table-name <lock-table> --projection-expression 'LockID' to see which states hold leases. Then serialize writers with a CI concurrency group and add -lock-timeout so short overlaps self-heal.| File | Command / Code | Purpose |
|---|---|---|
| backend.tf | terraform { | What the 'Error Acquiring the State Lock' Message Actually M |
| check-lock-holder.sh | ps aux | grep '[t]erraform' | Lease vs Stuck Lock |
| inspect-lock-table.sh | aws dynamodb scan \ | Reading the Lock ID and DynamoDB Lock Row Before You Touch A |
| release-the-lock.sh | terraform workspace show | Running terraform force-unlock Safely in Production |
| deploy.yml | concurrency: | CI Pipelines |
Key takeaways
Common mistakes to avoid
5 patternsDeleting the DynamoDB lock row by hand instead of running force-unlock
terraform force-unlock <ID> in the same directory and workspace. That command removes the lease through Terraform's own backend code, so the state file and the lock table stay consistent.Passing -lock=false to dodge the error in CI
-lock-timeout=10m so Terraform waits for a live holder to finish, or confirm the holder is dead and force-unlock. Reserve -lock=false for read-only experiments on throwaway state, never for applies.Running force-unlock from the wrong directory or workspace
terraform workspace show and compare the backend key in your config with the Path in the error message. Re-run force-unlock from the exact directory and workspace where the original apply ran.Letting two CI pipelines write the same state file
resource_group in GitLab, or a deploy queue. Queue additional runs instead of running them side by side.Killing hung applies with no timeout or retry policy
-lock-timeout to applies, and alert when an apply exits non-zero so a human checks for a stuck lock before the next run starts.Interview Questions on This Topic
Why does Terraform lock state, and why can a lock outlive the process that took it?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Terraform. Mark it forged?
6 min read · try the examples if you haven't