Home › Cloud › Terraform State Lock Error: Unlock Stuck S3 Applies
Intermediate 6 min · September 23, 2026

Terraform State Lock Error: Unlock Stuck S3 Applies

Copy the Lock ID from the error and run terraform force-unlock, then re-plan.

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 14 min
  • ✓You've run terraform init and apply against an S3 backend at least once
  • ✓Basic comfort with the AWS CLI and DynamoDB concepts
  • ✓Access to a non-production state file where you can practice safely
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • A stuck lock almost always means a previous apply crashed or timed out, leaving its DynamoDB lease behind — not proof someone else is applying now.
  • Copy the Lock ID from the error's Lock Info block and run terraform force-unlock in the same directory and workspace.
  • Before unlocking, confirm the holder is dead: check CI logs, ps aux | grep terraform, and the lock's Created timestamp.
  • Prevent repeats with -lock-timeout on applies, CI step timeouts, and exactly one writer job per state path.
✦ Definition~90s read
What is Terraform Error Acquiring the State Lock?

Terraform state locking is a mutual-exclusion lease that protects the state file from concurrent writes. Because a state file is a single JSON document with no transaction support, two processes writing at once would each read the same version, each write back their own, and the last writer would silently erase the other's resources from state — while those resources kept running in the cloud.

★
Think of Terraform state like a library's paper sign-out sheet for a popular book.

The lock makes writes serialized: one holder at a time, everyone else waits or fails fast.

On S3 backends the lease lives in DynamoDB. Your backend block names the state location (bucket plus key) and a lock table; before any operation that could write state, Terraform inserts a row keyed by the state path containing a lock ID, the operation type, the creator, and a timestamp.

When the operation finishes — success or clean failure — Terraform deletes the row. Plans take the lock too when they might persist state, which is why even a read-looking command can trip over a stuck lease.

The tradeoff of this design is that the lease can outlive its owner. DynamoDB has no idea your CI runner was killed; the row sits there until something removes it. That's why force-unlock exists as an explicit, auditable human action rather than an automatic timeout — only you can verify the holder is actually dead, and getting that judgment wrong corrupts state.

Understand the lease lifecycle and the error becomes routine instead of terrifying.

Plain-English First

Think of Terraform state like a library's paper sign-out sheet for a popular book. Before changing anything, you clip your name tag to the sheet so nobody else writes on it at the same time. If you walk out mid-edit and leave your tag clipped on, the next person can't do their work — that's a stuck lock. The fix isn't tearing up the sheet; it's checking you actually left the building, unclipping your tag properly, and making sure the sheet still reads correctly before anyone writes again.

It's Friday at 4 PM and your deploy pipeline just turned red with a message you've never seen: Error acquiring the state lock. Nobody on the team knowingly changed anything about state, the last apply finished an hour ago, and the error helpfully suggests running a command called force-unlock that sounds like it'll either save your evening or end your career.

That fear is healthy, because the lock is load-bearing. Terraform stores your infrastructure as a single JSON state file with no transactions, so before every write it takes a short-lived lease — for S3 backends, a row in a DynamoDB table. If two processes wrote at once, the last writer would win and the loser's resources would silently vanish from state while still running and billing you. The lock is the only thing standing between you and that outcome.

Most stuck locks come from something boring: a CI runner that timed out mid-apply, a laptop lid closed during a long apply, a spot instance reclaimed under the process. The lease outlives its owner and the next run trips over it. This article shows you how to tell a dead lock from a live one, how to release it without corrupting state, and how to set up your backend and pipelines so you stop seeing this error at 4 PM on Fridays.

What the 'Error Acquiring the State Lock' Message Actually Means

When Terraform prints Error acquiring the state lock, it's telling you a write lease it needs is currently held. The Lock Info block that follows is the whole investigation in miniature: an ID (a UUID identifying this exact lease), a Path (which state file is locked), an Operation (usually Apply or Plan), a Who field (user@host of the holder), a Created timestamp, and a Version. Every one of those lines earns its place, and skipping them is how teams unlock the wrong state or interrupt a live deploy.

The reason the lease exists at all is that object storage has no transactions. Your state file is one JSON document; two simultaneous writers would each read it, each write their version back, and the loser would silently erase the winner's resources from state while those resources kept running and billing you. For S3 backends the lease is a row in a DynamoDB table holding the lock ID plus metadata. Terraform writes that row before any state mutation and deletes it when the operation completes or fails cleanly.

So the error never means your configuration is broken — it means someone (or some leftover process) holds the pen. Your first move is always reading, not unlocking: note the ID, check the Path against your backend key, look at Who and Created, and decide whether the holder could still be alive. Only then do you act. Engineers who jump straight to force-unlock are playing roulette with the state file. Take notes as you go — the Lock ID is long and mistyping it wastes a whole cycle.

backend.tfHCL
1
2
3
4
5
6
7
8
9
terraform {
  backend "s3" {
    bucket         = "acme-terraform-state"
    key            = "prod/network/terraform.tfstate"
    region         = "us-east-1"
    dynamodb_table = "acme-terraform-locks"
    encrypt        = true
  }
}
📊 Production Insight
In the Friday incident, the Path line immediately showed the lock belonged to prod/network while the engineer was sitting in staging. That one line saved them from unlocking the wrong environment's state.
🎯 Key Takeaway
The Lock Info block identifies exactly which lease, state file, holder, and operation you're dealing with — read every line before acting.

Lease vs Stuck Lock: Telling a Live Apply From a Dead One

A lease and a stuck lock look identical in the error message — the difference is whether the holder is still alive. A lease is healthy: the Created timestamp is minutes old, the Who field names a teammate whose apply is visibly running, or your CI dashboard shows an in-progress job for the same state path. The correct response is patience, not force: wait, or rerun with -lock-timeout so Terraform queues behind the holder.

A stuck lock is a leftover: the timestamp is old, the named host is a CI runner that no longer exists, the pipeline shows killed or timed out, and no terraform process answers on the machine. That's the signature of a crashed apply — laptop lid closed, spot instance reclaimed, step timeout fired mid-write. The lease outlived its owner and nothing will release it except you.

Build the habit of checking all three signals — process, pipeline, timestamp — before deciding. The DynamoDB item's Info document carries the same ID, operation, and creator as the error, so you can confirm the row matches the complaint and isn't residue from an older incident. Five minutes of verification here is what separates a routine unlock from a state-corruption postmortem. Write down what you find — the next unlock goes twice as fast.

check-lock-holder.shBASH
1
2
3
4
5
6
7
8
9
# Is any terraform process alive on this machine?
ps aux | grep '[t]erraform'

# Inspect the lease row: Info.Created tells you when it was taken
aws dynamodb get-item \
  --table-name acme-terraform-locks \
  --key '{"LockID": {"S": "acme-terraform-state/prod/network/terraform.tfstate"}}' \
  --projection-expression 'LockID, Info' \
  --output json
📊 Production Insight
One team nearly unlocked a live 40-minute RDS apply because the error 'looked stuck.' The CI log showed it still running — that single check prevented an interrupted database migration.
🎯 Key Takeaway
Fresh timestamp plus a live process means wait; old timestamp plus a dead job means unlock — confirm with process, pipeline, and lock-row evidence.

Reading the Lock ID and DynamoDB Lock Row Before You Touch Anything

Before you release anything, collect two artifacts: the Lock ID from the error and the matching row in the DynamoDB table. The ID is a UUID in the Lock Info block — that's the only value force-unlock accepts, and it must match the lease exactly. Copy it carefully; retyping a 36-character UUID by hand is asking for a mismatch error at the worst moment.

Then look at the lock table itself. Its partition key is LockID, and each item carries an Info document with the operation type, creator, creation time, and version — the same fields the error shows. Comparing the two confirms you're looking at the current lease and not residue from an older incident, and the table name in your backend block tells you which table to query. If your backend key is prod/network/terraform.tfstate, the locked state is exactly that path, and the fix must run from the configuration that manages it.

This is also where you catch the classic trap: the lock belongs to a different workspace or environment than your current directory. Staging and production have separate state paths and separate leases. Unlocking from the wrong directory either errors out on ID mismatch or — worse with a forced flag — touches the wrong state. Match the Path line to your backend key first, always.

inspect-lock-table.shBASH
1
2
3
4
5
6
7
8
9
10
11
# List every outstanding lease so you can match the error's Lock ID
aws dynamodb scan \
  --table-name acme-terraform-locks \
  --projection-expression 'LockID' \
  --output table

# Pull the full Info doc for the state in the error's Path line
aws dynamodb get-item \
  --table-name acme-terraform-locks \
  --key '{"LockID": {"S": "acme-terraform-state/prod/network/terraform.tfstate"}}' \
  --output json | python3 -m json.tool
📊 Production Insight
An engineer once spent an hour fighting 'lock ID mismatch' before noticing the error's Path said prod while their shell was in staging. The Path line is the cheapest diagnostic in the whole message.
🎯 Key Takeaway
Copy the UUID from the error, confirm it against the DynamoDB lock row, and verify the Path matches your backend key and workspace.

Running terraform force-unlock Safely in Production

force-unlock is safe when the holder is provably dead and dangerous otherwise — the command itself can't tell the difference, so you provide the judgment. Run it from the same directory and workspace that created the lock, passing the exact Lock ID from the error. Terraform asks for confirmation; in interactive use, type yes rather than bypassing the prompt, because that pause is your last chance to catch a wrong-directory mistake.

In CI, where there's no human to confirm, use the non-interactive flag only inside a job that already verified the holder is dead — never as a blind pre-step before every apply. A scheduled unlock that fires while a real apply runs will let two writers onto one state file, which is precisely the corruption the lock exists to prevent.

After the unlock, always run plan before apply. The crashed run may have created half its resources, and only a fresh plan shows which half survived. Review that plan the way you'd review any production change: unexpected creates suggest orphaned writes, unexpected destroys suggest the crash landed mid-delete. If the plan looks clean, proceed; if it doesn't, investigate before the next write. The unlock fixed the lease, not the interrupted work. Patience here is cheap; state surgery is not.

release-the-lock.shBASH
1
2
3
4
5
# From the SAME directory and workspace that created the lock:
terraform workspace show
terraform force-unlock 3f9a8c1e-7b2d-4a5f-9c0e-1d2b3a4c5e6f
# Type 'yes' at the prompt, then verify state health:
terraform plan -out /tmp/verify.plan
📊 Production Insight
The Friday team now requires unlocks during business hours to be approved by a second engineer in chat. The approval takes two minutes and has caught two wrong-workspace attempts since.
🎯 Key Takeaway
Same directory, same workspace, exact Lock ID, confirm deliberately — then plan before you apply to account for the crashed run's partial work.

S3 Backend and DynamoDB Lock Table Settings That Prevent Stuck Locks

The backend settings that make stuck locks rare are unglamorous: a dedicated DynamoDB lock table with on-demand capacity so bursts of applies don't get throttled, server-side encryption on the state bucket, and versioning enabled so a corrupted state file can be restored from history. None of this is exotic, but teams that set state up once and never revisit it usually lack at least one of the three.

On the client side, two flags do most of the work. -lock-timeout tells Terraform to wait for a live holder instead of failing instantly — a few minutes of patience absorbs the normal overlap when two pipelines fire close together. Sensible CI step timeouts do the rest: a timeout shorter than the runner's kill timeout lets Terraform exit cleanly and release its lease instead of being murdered mid-write.

Finally, treat backend configuration as reviewed code. The bucket, key, region, and lock table name should live in version control and change through pull requests like anything else. When region typos or IAM trims masquerade as stuck locks — every command failing, not just concurrent ones — a reviewed backend block gives you a known-good baseline to diff against instead of a mystery. Revisit that baseline quarterly — teams grow, regions multiply, and yesterday's correct backend silently becomes tomorrow's puzzle.

⚠ Don't Bypass the Lock — Release It
Never delete the DynamoDB lock row from the console and never ship -lock=false in CI to dodge the error. Both remove the only guard against concurrent writes, and the resulting corruption is silent until the next plan proposes creating resources that already exist.
📊 Production Insight
After the incident, the team enabled bucket versioning and found they'd needed it months earlier — a corrupted state from an unrelated crash took a full day to rebuild that versioning would have reduced to a restore.
🎯 Key Takeaway
Versioned encrypted bucket, dedicated lock table, -lock-timeout on commands, and backend config in code review stop most locks before they form.

CI Pipelines: Timeouts, Retries, and Lock Hygiene in Automation

Automation is where locks go to multiply, because CI happily runs two applies against one state file if nothing stops it. The fix is a single-writer rule per state path: a concurrency group in GitHub Actions, a resource group in GitLab, or a deploy queue in front of your pipeline. Additional runs wait their turn instead of colliding, which eliminates the most common lock source — overlapping schedules — without any human intervention.

Pair serialization with timeouts on both sides. The CI step timeout should be shorter than the runner's hard kill timeout so Terraform can shut down cleanly and release its lease. The -lock-timeout flag on plan and apply covers the residual overlap when a previous run overruns by minutes rather than hours. Together they convert most would-be incidents into slightly slower green builds.

Last, make failures loud. Alert on every non-zero Terraform exit in CI and include the state path in the alert, so a stuck lock pages someone with context instead of sitting until the next deploy discovers it. Log the Lock ID and the job URL together — future-you, clearing the lock at midnight, will be grateful for the archaeology head start. Review these alerts monthly; a quiet lock table means the hygiene is working.

deploy.ymlBASH
1
2
3
4
5
6
7
8
9
# Serialize writers: one deploy at a time per state path
concurrency:
  group: terraform-prod-network
  cancel-in-progress: false

# Then apply with patience instead of force
terraform init -input=false
terraform plan -input=false -lock-timeout=10m -out=tfplan
terraform apply -input=false -auto-approve tfplan
📊 Production Insight
A team that added concurrency groups plus -lock-timeout went from a stuck lock every few weeks to zero in six months — the cheapest reliability win in their whole Terraform setup.
🎯 Key Takeaway
One writer per state path, step timeouts that let Terraform exit cleanly, -lock-timeout for short overlaps, and alerts that name the state.
● Production incidentPOST-MORTEMseverity: high

The Friday Deploy That Left a Phantom Lock — and the Manual Delete That Made It Worse

Symptom
At 4:12 PM the deploy pipeline failed with Error acquiring the state lock, pointing at a Lock ID created at 3:38 PM by a job the dashboard showed as timed out. Three subsequent retries failed identically. After the manual lock-table delete, the next apply succeeded — but the following plan wanted to create a NAT gateway, two subnets, and a route table that already existed in the AWS console.
Assumption
The pipeline had a 30-minute step timeout, and the runner's infrastructure killed the job at 30 minutes without running cleanup. The team assumed a killed job couldn't leave side effects, so nobody checked the lock table. The retry an hour later failed on the same lock, and the on-call engineer — under pressure from a waiting release — deleted the DynamoDB item from the console to unblock the queue.
Root cause
Two failures stacked. The CI runner was killed by a step timeout while terraform apply held the DynamoDB lease, so the lease was never released — a classic stuck lock. Then the manual DynamoDB delete removed the lease while the automatic retry was already mid-write, allowing two writers on one state file. The loser's resources dropped out of state while still running in AWS, including a NAT gateway that kept billing until the audit found it.
Fix
The team restored order in three steps. First they ran terraform plan -refresh-only to reconcile state with reality and listed every resource whose real-world existence didn't match the file. Then they imported the orphaned resources back into state and destroyed one accidental duplicate NAT gateway the double-write had created. Finally they replaced the console-delete habit with a runbook: verify the holder is dead, force-unlock with the Lock ID, re-plan, and require a second engineer to approve any unlock during business hours.
Key lesson
  • A killed CI job absolutely leaves side effects — the lock outlives the runner. Every pipeline that applies Terraform needs a step timeout shorter than the infrastructure kill timeout, plus an alert on non-zero exits so a human inspects for stuck locks before retrying.
  • Never clear a lock by deleting the DynamoDB row. force-unlock exists because it validates the lock ID against live state; a console delete can't tell a dead lock from a live apply, and guessing wrong corrupts the state file.
  • A retried apply after a crash must start with plan, not apply. The crashed run may have created half its resources, and only a fresh plan shows which half — re-applying blind is how you get duplicates.
Production debug guideFive Terraform state lock patterns on S3 backends, with the exact commands that resolve each.5 entries
Symptom · 01
Lock error appears right after a cancelled, timed-out, or crashed apply
→
Fix
Copy the Lock ID from the Lock Info block. Check for a live writer with ps aux | grep '[t]erraform' and by reading the CI job log — a killed or timed-out job means the holder is dead. Then run terraform force-unlock <LOCK_ID> in the same directory and workspace, answer yes, and follow with terraform plan to confirm state health.
Symptom · 02
Lock error while a teammate's apply (or another pipeline) is visibly running
→
Fix
Don't unlock — you'd interrupt a real deployment. Rerun your command with a wait: terraform plan -lock-timeout=10m or terraform apply -lock-timeout=10m. Watch the other run finish, then proceed. If your CI hits this often, put both jobs in one concurrency group so they queue instead of colliding.
Symptom · 03
force-unlock says the lock ID doesn't match or that no lock exists
→
Fix
Run terraform workspace show and compare your backend key with the Path line in the error. Switch with terraform workspace select <name> or cd into the right directory, then rerun terraform force-unlock <LOCK_ID>. Backend keys differ per environment, so staging and production locks are released independently.
Symptom · 04
Every Terraform command fails on locking, even plain plan, with access errors
→
Fix
Run TF_LOG=DEBUG terraform plan 2>&1 | grep -i -E 'denied|dynamodb|LockID' to capture the real failure. Then verify IAM with aws dynamodb describe-table --table-name <lock-table> and confirm your role has dynamodb:GetItem, PutItem, and DeleteItem on the lock table. Fix the policy, rerun terraform init, and try again.
Symptom · 05
The lock returns minutes after every unlock, always around the same hour
→
Fix
List recent runs for that state path and look for overlapping start times — two schedules or a retry storm. Run aws dynamodb scan --table-name <lock-table> --projection-expression 'LockID' to see which states hold leases. Then serialize writers with a CI concurrency group and add -lock-timeout so short overlaps self-heal.
Terraform State Lock Errors — Root Cause Comparison for S3 Backends
Root CauseHow to ConfirmFixPrevention
Crashed or killed apply left its lease behindLock Info Created timestamp is old, no terraform process is alive, and the CI job shows killed or timed outRun terraform force-unlock with the Lock ID, then plan to verify stateSet CI step timeouts and keep one writer job per state path
A second apply started while the first still holds the lockLock Info Who and Created are fresh, and ps or the CI log shows an apply still runningWait for it to finish, or rerun with -lock-timeout=10mSerialize deploys with a concurrency group or deploy queue
force-unlock ran in the wrong workspace or directoryTerraform reports a lock ID mismatch or no lock, while applies keep failing on the same stateSwitch to the right workspace and rerun from the original directoryPin the working directory and workspace in CI and document the backend key per environment
Backend can't reach the DynamoDB lock tableEvery command fails including plan, with AccessDenied or region errors in TF_LOG=DEBUG outputFix the IAM policy and region, then rerun terraform initCheck the least-privilege backend policy into code review like any other change
⚙ Quick Reference
5 commands from this guide
FileCommand / CodePurpose
backend.tfterraform {What the 'Error Acquiring the State Lock' Message Actually M
check-lock-holder.shps aux | grep '[t]erraform'Lease vs Stuck Lock
inspect-lock-table.shaws dynamodb scan \Reading the Lock ID and DynamoDB Lock Row Before You Touch A
release-the-lock.shterraform workspace showRunning terraform force-unlock Safely in Production
deploy.ymlconcurrency:CI Pipelines

Key takeaways

1
A stuck lock usually means a crashed apply left its DynamoDB lease behind
verify the holder is dead before touching it.
2
Always release with terraform force-unlock and the Lock ID from the error, never by deleting the DynamoDB row.
3
Run force-unlock in the same directory and workspace that created the lock, then plan to confirm state health.
4
A fresh timestamp plus a live process means wait; an old timestamp plus a dead job means unlock.
5
Keep one writer per state path with CI concurrency groups and add -lock-timeout to applies.
6
Treat backend IAM and lock table config as reviewed code so permission failures don't mimic stuck locks.

Common mistakes to avoid

5 patterns
×

Deleting the DynamoDB lock row by hand instead of running force-unlock

Symptom
The lock disappears but the next apply behaves strangely — Terraform reports a state serial mismatch or tries to recreate resources that already exist, because the manual delete skipped the consistency checks force-unlock performs.
Fix
Copy the Lock ID from the error's Lock Info block and run terraform force-unlock <ID> in the same directory and workspace. That command removes the lease through Terraform's own backend code, so the state file and the lock table stay consistent.
×

Passing -lock=false to dodge the error in CI

Symptom
Two applies write the same state file at once. The last writer wins, the loser's resources vanish from state while still running in AWS, and you've created orphaned infrastructure that bills you monthly.
Fix
Rerun with -lock-timeout=10m so Terraform waits for a live holder to finish, or confirm the holder is dead and force-unlock. Reserve -lock=false for read-only experiments on throwaway state, never for applies.
×

Running force-unlock from the wrong directory or workspace

Symptom
Terraform says the lock ID doesn't match or that no lock exists, yet every apply still fails with the same lock error — you're talking to a different state file than the one that's locked.
Fix
Run terraform workspace show and compare the backend key in your config with the Path in the error message. Re-run force-unlock from the exact directory and workspace where the original apply ran.
×

Letting two CI pipelines write the same state file

Symptom
Lock errors appear randomly, roughly when two schedules overlap. Each team insists nobody else was applying, and both are right — the overlap window is only a few minutes wide.
Fix
Give each state path exactly one writer job: concurrency groups in GitHub Actions, resource_group in GitLab, or a deploy queue. Queue additional runs instead of running them side by side.
×

Killing hung applies with no timeout or retry policy

Symptom
Every cancelled pipeline leaves a lock behind. Monday's deploy queue is three stuck locks deep, and nobody knows which ones are safe to clear because nothing recorded why each apply died.
Fix
Set a CI step timeout shorter than the runner's kill timeout, add -lock-timeout to applies, and alert when an apply exits non-zero so a human checks for a stuck lock before the next run starts.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Why does Terraform lock state, and why can a lock outlive the process th...
Q02JUNIOR
Walk me through safely clearing an 'Error acquiring the state lock' fail...
Q03SENIOR
How do you distinguish a live lease from a stuck lock?
Q04SENIOR
How would you design CI so stuck locks rarely happen?
Q05SENIOR
An engineer deleted the DynamoDB lock row during a live apply. What brea...
Q01 of 05JUNIOR

Why does Terraform lock state, and why can a lock outlive the process that took it?

ANSWER
State files are single JSON documents with no transactions, so two writers would overwrite each other and orphan real resources. The backend takes a short-lived lease (a DynamoDB row for S3 backends) before any write. If the process dies without releasing it, the lease stays until a human clears it with force-unlock.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Can I reuse a lock ID from a previous incident?
02
What if the engineer who caused the lock is offline?
03
Will Terraform wait for a lock instead of failing immediately?
04
Does force-unlock roll back the crashed apply's changes?
05
Is it safe to run plan while a lock is held?
06
How do I move off local state before this bites my team?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Terraform. Mark it forged?

6 min read · try the examples if you haven't

1 / 5 · Terraform
Next
Terraform Resource Already Exists — Import, Don't Recreate
→