Terraform count vs for_each: Stop Index Shifts Eating VMs
Count addresses by position, so reorders destroy survivors.
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓You've used count to create multiple similar resources
- ✓Comfort with maps, sets, and basic for expressions in HCL
- ✓A non-production workspace where you can practice a migration
- count numbers instances by position (web[0], web[1]), so reordering or deleting mid-list renumbers survivors into replacements.
- for_each keys instances by stable identity (web["a"]), so only genuinely added, removed, or changed keys are touched.
- Migrate safely with moved blocks mapping each index to its key, and demand an empty plan before applying.
- Key by attributes that never change (hostname, username); keep mutable settings in the value.
Imagine a parking lot where cars are tracked by space number. When the first car leaves and everyone shuffles forward, the attendant's log says cars 1 through 10 are all 'different cars' — even though only one left. That's count. for_each is license plates instead: each car keeps its identity no matter where it parks, so the log only changes for cars that truly arrived or left. Moved blocks are the afternoon the lot switches systems, carefully recording that space 3's car is really plate XYZ.
Someone alphabetizes a list of server names — a cosmetic cleanup, the kind of commit that shouldn't even need a review — and Terraform responds by proposing to destroy half your fleet. Not because anything about those servers changed, but because count addresses instances by position: web[0], web[1], web[2]. Rename the first entry and every survivor gets renumbered, and renumbering looks exactly like replacement.
This is the count trap, and nearly every Terraform team walks into it. Count is the first loop you learn, so it becomes the default for everything: servers, DNS records, IAM users. It works beautifully until the day the list changes shape — a reorder, a mid-list deletion, a sort — and then it bills you in destroyed databases and rotated IP addresses.
for_each is the exit: it addresses instances by stable keys instead of positions, so edits touch only what actually changed. This article shows you why index shifts destroy, how to migrate live infrastructure with moved blocks and zero downtime, and how to pick keys that stay stable for the lifetime of your resources.
Why Reordering a List Destroys Servers Under count
Count addresses instances by position in a resource array: web[0] is whatever item zero currently holds, not a specific server. When the backing list reorders — an alphabetical sort, a new item inserted at the top, a mid-list deletion — every affected index points at a different item than before. Terraform compares desired addresses against state, finds every shifted address holding the 'wrong' object, and plans destroy-plus-create for each one. The servers never changed; their numbers did.
Mid-list deletion is the cruelest variant. Remove item zero from a five-item list and items one through four slide down a slot: four instances destroyed and recreated to delete one server. Appends at the end are the only safe mutation, and even those survive only until someone sorts the list or inserts above them. Every count-backed list is one well-meaning edit away from a fleet replacement plan.
This is why experienced teams treat count as a specialist tool for fixed-size, interchangeable sets — three identical zone workers, two NAT gateways — and never for named things with identity. If members have names, meanings, or individual lifecycles, positional addressing is a mismatch that will eventually bill you. The plan output is your early warning: destroys on resources you didn't touch mean position just masqueraded as identity.
for_each Keys: Identity That Survives Reorders and Deletions
for_each replaces positions with keys: web["web-a"] is the server named web-a regardless of how the collection is ordered, what was added, or what was removed. Adding a key creates one instance; removing a key destroys exactly one; editing a value updates one in place. Neighbors are never renumbered because there are no numbers — identity travels with the key, immune to every edit that devastates count.
The key insight is the split between identity and configuration. The key declares which object this is; the value declares how it's configured. Changing the value (instance type, tags, AMI) updates the same object in place, while changing the key addresses a different object entirely. This is why key choice matters so much: keys must be the stable, immutable identity (hostname, username, bucket name), and everything mutable belongs in the value.
Conversion starts at the input: lists must become maps or sets before for_each accepts them. The standard transform is { for u in var.users : u.name => u }, keying each element by its stable attribute. Keep that transform beside the resource so reviewers see the key source, and prefer maps over sets when values carry per-instance settings — sets of strings work for uniform fleets, maps of objects for heterogeneous ones.
Migrating Live Fleets With moved Blocks and Zero Downtime
Migrating live infrastructure from count to for_each without moved blocks is a self-inflicted outage: Terraform sees indexed addresses vanish and keyed addresses appear, and plans destroy-all plus create-all. Moved blocks prevent that by declaring continuity — the object at aws_instance.web[0] is the same object as aws_instance.web["web-a"] — so Terraform remaps state instead of replacing infrastructure. One block per instance, each mapping the old index to its new key.
The migration ritual has four steps. First, convert the resource to for_each with keys matching current reality (index zero's actual name becomes the first key — verify against state, not assumptions). Second, write the moved blocks. Third, run plan and demand emptiness: a clean plan proves every mapping is correct and the migration touches zero infrastructure. Fourth, apply the pure remap, then remove the moved blocks in a later cleanup once all branches and environments have applied them.
If the plan shows any destroy or create, stop — a mapping is wrong. Compare each from address against terraform state list output and each to key against the live object's real name. Common culprits: assuming index order matches name order (verify!), or keying by an attribute that differs from reality. The empty plan is the entire safety case; never waive it.
Fixing Downstream References That Still Think Positionally
The migration doesn't end at the resource — every downstream reference must convert from positional to keyed thinking. Splat expressions and numeric indexes over the old count resource (aws_subnet.main[count.index]) silently follow positions, not machines; after a reorder they attach to the wrong instances with no error. Convert each one: iterate the keyed map directly (for_each = aws_instance.web) or look up by key (aws_instance.web["web-a"].id).
Outputs need the same treatment. A list output built positionally (web_ids[0]) becomes meaningless once keys exist — downstream modules indexing it positionally inherit the original fragility. Convert outputs to maps keyed identically ({ for k, inst : k => inst.id }), and update every consumer to look up by key. This is the tedious half of the migration and the half most often skipped, which is why 'migrated' fleets sometimes keep their old ordering bugs in downstream attachments.
Audit mechanically: grep for the old resource name across all files and convert every hit, then plan and read downstream diffs for attachments landing on unexpected instances. The rule going forward is absolute — never index a for_each resource positionally. Keys exist precisely so positions don't matter; any code that reintroduces positional assumptions rebuilds the trap you just escaped.
Choosing Keys That Stay Stable for the Life of the Resource
Key selection is the decision that outlives the migration, because keys become permanent identity: changing a key later replaces the instance. The rule is simple to state and requires judgment to apply — key by the attribute that will never change over the object's lifetime. Hostnames for servers, usernames for IAM users, bucket names for buckets. These are the names the real world already uses as identity, which is exactly why they make stable keys.
The failure mode is keying by something mutable: an environment tag that gets renamed, a role that evolves, a size label that changes. Every such key is a future replacement hiding in plain sight — the day the attribute changes, Terraform sees a new key and destroys the instance. Review key sources with replacement-colored glasses: if this string could ever legitimately change, it must live in the value, not the key.
Enforce this at review time with two questions for every for_each: what is the key, and what happens when it changes? If the answer to the second is 'replace,' confirm that's intended (sometimes it is — immutable infrastructure welcomes it). Write the key-source rationale in a comment beside non-obvious transforms, so the engineer editing values in two years doesn't accidentally promote a mutable attribute into identity.
Count Still Has a Job: Fixed Interchangeable Sets
After five sections on count's dangers, a defense is overdue: count isn't broken, it's specialized. Its home turf is fixed-size sets of interchangeable things — two NAT gateways for zone redundancy, three identical workers spread across availability zones, a replica number that scales as one value. Here no member has individual meaning: gateway zero and gateway one differ only by position, and if one is replaced, no identity is lost. Positional addressing fits because position is all there is.
The dividing line is meaning. Ask whether members have names that outlive the deployment: servers with hostnames, users with logins, buckets holding data — all carry identity, all deserve for_each. Ask whether the size changes by editing a list humans reorder: any yes means keys, not indexes. Count survives these questions only when the answer is genuinely three of the same and you don't care which is which. NAT gateways pass; application servers don't.
Keeping count healthy in its niche takes two habits. First, never let humans edit the backing structure positionally — derive counts from numbers (count = 3) or length() of machine-generated lists, not from hand-ordered name lists. Second, keep downstream references positional only where the resource is positional: subnet lookups by zone index are fine when zones are fixed, but the moment instances gain names, the whole chain should go keyed. Respect the boundary and both constructs behave; cross it and you're back in index-shift territory.
The Alphabetical Sort That Tried to Destroy Twelve Servers
- Variable-file edits are infrastructure changes when count is involved — require plans on every pull request that touches lists backing count, no matter how cosmetic.
- Catch destroy plans in review, not in apply: the team's rule that any destroy in a plan needs explicit human approval is what saved the fleet.
- Sort-proof your configs proactively: any count-backed list a human might reorder is a migration candidate, not a stable design.
terraform plan — iterate until the plan is empty, proving pure state remap with zero infrastructure change. Then apply.terraform plan before applying — the diff should name only the removed key.terraform plan and confirm the only changes are address remaps (listed as moved, not replaced). If any resource shows destroy/create, its moved block is missing or mistyped — compare each from address against terraform state list output.aws_instance.web["a"].id or iterate with for k, inst in aws_instance.web. Run terraform validate then terraform plan to confirm downstream diffs reference the intended instances.grep -rn 'for_each' --include='*.tf' . to find every keyed resource, and keep mutable settings in the value. Re-plan: pure updates mean the keys are stable.| File | Command / Code | Purpose |
|---|---|---|
| count-trap.tf | variable "server_names" { | Why Reordering a List Destroys Servers Under count |
| for-each-stable.tf | variable "servers" { | for_each Keys |
| migration.tf | moved { | Migrating Live Fleets With moved Blocks and Zero Downtime |
| keyed-references.tf | resource "aws_eip" "web" { | Fixing Downstream References That Still Think Positionally |
| count-proper.tf | resource "aws_nat_gateway" "main" { | Count Still Has a Job |
Key takeaways
Common mistakes to avoid
5 patternsReaching for count by default for every multi-instance resource
Keying for_each on an attribute that itself changes
Feeding for_each a raw list instead of a map or set
toset() or build the map explicitly: { for u in var.users : u.name => u }. Keep the transformation beside the resource so reviewers can see the key source at a glance.Migrating count to for_each without moved blocks
Indexing into for_each resources positionally like a list
Interview Questions on This Topic
What's the practical difference between count and for_each?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Terraform. Mark it forged?
5 min read · try the examples if you haven't