AKS ImagePullBackOff: ACR Access Denied Fix
Fix AKS ImagePullBackOff by granting AcrPull to the kubelet identity: verify with check-acr, then assign the role at the registry scope..
20+ years shipping production backend systems. Written from production experience, not tutorials.
- ✓An AKS cluster with kubectl access configured
- ✓An Azure Container Registry with a pushed image
- ✓Azure CLI installed with az login completed
- ImagePullBackOff from ACR means the node kubelet identity lacks permission: same subscription alone grants nothing
- Confirm the path with
az aks check-acr, then compare the live kubelet identity against the role assignment - Fix it with one grant:
az role assignment create --role AcrPullscoped to the registry - Keep imagePullSecrets only as a fallback, and rule out missing tags plus ACR firewall blocks first
Picture the ACR as a private warehouse and your AKS nodes as delivery drivers. ImagePullBackOff means the driver arrived but the door won't open — their badge (the kubelet identity) isn't on the access list (the AcrPull role). The door checks the list, not the org chart. Add the badge to the list. Rebuilding the depot or yelling at the driver changes nothing until the badge is authorized.
You deploy to AKS on a Monday morning and every pod sits in ImagePullBackOff. The image exists — you just pushed it. The registry is in the same subscription, maybe the same resource group. Yet the nodes act like they've never heard of your container. Someone suggests rebuilding the cluster. Someone else blames a typo. Both cost more than the actual fix, which is usually one role assignment.
ImagePullBackOff against Azure Container Registry is a permissions problem wearing a networking costume. The kubelet identity on your nodes must hold the AcrPull role on the registry, and nothing about proximity grants it automatically. Same subscription, same tenant, same team — none of that matters to Azure RBAC. Without the explicit grant, every pull is denied and Kubernetes retries with growing delays.
This error loves to appear right after cluster rebuilds, identity rotations, and subscription moves, because all three silently change which identity pulls. The assignment that worked Friday points at an object ID that no longer exists Monday.
This guide gives you the exact diagnosis order: describe the pod, check the pull path with az aks check-acr, compare identities, and assign the role at the right scope. You'll see a real outage caused by a stale identity, plus the infrastructure-as-code pattern that makes this failure nearly impossible.
What ImagePullBackOff From ACR Actually Signals
ImagePullBackOff is Kubernetes telling you the kubelet tried to fetch the container image, failed, and is waiting longer between retries. The BackOff part is just exponential delay — the useful signal is the pull failure underneath. Against ACR that failure is usually authorization: the node presented its managed identity, ACR checked Azure RBAC, and found no AcrPull grant. The registry answers 401, the kubelet records Failed to pull image, and the pod waits.
This error misleads because it looks environmental. Same subscription, same resource group, recently working — everything suggests the plumbing broke. But Azure RBAC never infers permission from proximity. An AKS cluster and an ACR can share an owner, a subscription, and a VNet while the pull still fails, because the only question ACR asks is whether this specific object ID holds a pull grant on this specific registry scope. Everything else is scenery.
Two other failures wear the same costume. A misspelled tag or an unpushed image produces pull errors with no permission problem at all. An ACR firewall in Deny mode drops node traffic before authentication, which reads as timeouts rather than denials. The discipline is checking all three in order — identity grant, image existence, network path — instead of assuming the first theory. The debug guide below runs exactly that order.
The Kubelet Identity: Finding the Real Puller
The kubelet identity is the managed identity your node pools use for Azure API calls, including registry pulls. It is not the cluster identity — that one belongs to the control plane and manages load balancers and disks. Mixing them up is the most common reason a correct-looking role assignment fails to fix anything: the grant sits on an identity that never pulls. Resolve the real puller first, every time, with the query below.
Object IDs are the trap inside the trap. The kubelet identity's object ID changes when node pools are recreated, when the cluster is rebuilt, and sometimes during upgrades. Any assignment pinned to the old ID becomes a grant for a ghost. The portal doesn't wave a flag; it shows a valid-looking assignment whose principal no longer exists. The only reliable check compares the live identity from az aks show against the assignee on the role assignment, as the snippet demonstrates.
Treat that comparison as the heart of the diagnosis. If the IDs match and AcrPull is present, move on to tags and firewall with confidence. If they differ, you've found the outage: assign AcrPull to the live identity at the registry scope and watch the pods recover within a minute or two. Then fix the process that hardcoded the stale ID so the next rebuild doesn't repeat the incident.
Validating the Pull Path With az aks check-acr
az aks check-acr is the single most underused command in this space. It asks Azure to validate the pull path using the cluster's actual identity and reports success or the precise failure — no pod required, no deploy needed. Run it after cluster creation, after any identity operation, after ACR firewall edits, and inside your pipeline as a post-apply gate. A green result means RBAC and network both pass; a red result names which side failed.
Use it as the branch point in your diagnosis. If check-acr fails, stay on the platform side: compare identities, inspect assignments, check the firewall. If check-acr passes while pods still back off, leave RBAC alone — the problem is in the pod spec, the image reference, or a stale imagePullSecret. This one branch saves the hour teams otherwise spend re-granting roles for a typo.
Note what it doesn't cover. It validates the cluster's default pull identity against the registry, not per-pod imagePullSecrets or cross-tenant scenarios. For secret-based pulls, test the credential directly with a docker login from a throwaway host. For private endpoints and firewall rules, pair check-acr with a look at the effective network rules. The command is a gate, not the whole audit — but as gates go, it's the cheapest thirty seconds in AKS operations.
imagePullSecrets Fallback: When and How to Use It
imagePullSecrets are the legitimate fallback: a Kubernetes docker-registry secret holding registry credentials, referenced from the pod spec. They shine for third-party registries and cross-tenant pulls where managed identity can't reach. They fail as a primary strategy because every credential rotation is manual — renew the password, update the secret, restart the pods — and whatever step you forget becomes the next outage. The 3 AM version of this bug is always a rotated admin password nobody synced.
If you must use secrets, run them properly. Disable nothing you need, but scope the secret per namespace, name it consistently, and attach it via service accounts so individual pod specs stay clean. Store the actual credential in Key Vault and sync it with an operator rather than pasting passwords into YAML. Most importantly, alert on secret age: a secret older than your rotation period is an incident waiting for a trigger.
Prefer managed identity everywhere Azure-to-Azure. AcrPull on the kubelet identity rotates itself, scopes to the registry, and leaves no password to leak in a manifest. Keep one documented runbook for the secret path — cross-tenant ACR, emergency fallback — and let the identity path carry daily traffic. When pods back off, check which path they're even using before debugging the other one.
When the Firewall, Not RBAC, Blocks the Pull
Networking produces the cruelest variant of this failure because it mimics RBAC while defeating RBAC fixes. An ACR with public access disabled or a default-Deny firewall drops node traffic before authentication happens. The pod events show timeouts and retries rather than clean 401s, so teams keep re-granting roles while packets die at the firewall. Always inspect the registry's network posture when denials don't behave like denials.
The usual setups each have their failure mode. Private endpoints require DNS resolution from the nodes to the registry's private IP — a VNet link or private DNS zone misconfiguration sends pulls into the void. Firewall allowlists require the cluster's egress IPs, which change when NAT gateways or load balancer outbound rules change. Service endpoints require the subnet registered on both sides. Any of these drifting breaks pulls that RBAC can't fix.
Diagnose network-first when check-acr fails but assignments look right, or when events show timeouts instead of unauthorized. List the ACR network rules, confirm the node subnet or egress IP is present, and test DNS from a debug pod. Re-run check-acr after each change — it's the fastest confirmation that the path reopened. And record the network design next to the RBAC design in your runbooks, because the next debugger will otherwise assume the half you documented is the whole story.
Making AKS-to-ACR Auth Self-Healing With IaC
The durable fix puts the AcrPull assignment in the same infrastructure module as the cluster, resolving the kubelet identity dynamically. In Terraform that means referencing the kubelet object ID output rather than a literal; in Bicep it means chaining the role assignment resource to the cluster's identity profile. Either way, rebuilds re-grant automatically because the assignment follows the identity instead of fossilizing it. Hardcoded principal IDs should fail code review on sight.
Layer verification around the module. A post-apply pipeline step runs az aks check-acr and fails the apply on red. A nightly drift check diffs live assignments against desired state and pages on unexpected change. Image references pin to digests so tags can't drift under running manifests. Together these turn pull auth from tribal knowledge into tested infrastructure.
Keep the fallback documented but distinct. One runbook covers managed-identity pulls (the default path); a separate, shorter runbook covers imagePullSecrets for the registries identity can't reach. Onboard new services to the identity path by default and require justification for secrets. When the next ImagePullBackOff lands, the responder opens the right runbook in the first minute instead of the wrong one in the thirtieth.
The Upgrade That Quietly Replaced the Pull Identity
az role assignment create --assignee <NEW_KUBELET_ID> --role AcrPull --scope <ACR_ID>, verified with az aks check-acr. Pods started within a minute. The lasting fix moved the AcrPull assignment into the same Terraform module as the cluster, resolving the kubelet identity as a data output instead of a hardcoded object ID, plus a post-apply check-acr gate in the pipeline.- Never hardcode managed-identity object IDs in role assignments. Cluster rebuilds mint new identities silently, and a hardcoded ID keeps pointing at a ghost while the portal display looks reassuring.
- Run
az aks check-acrin the pipeline after every cluster change. A thirty-second gate would have caught this before any workload was scheduled onto the new nodes. - Treat identity changes as breaking changes. Rebuilds, rotations, and subscription moves should trigger the same verification rigor as a database migration.
kubectl describe pod <POD> -n <NS> and read the Events section for Failed to pull image plus the reason. Then run kubectl get events -n <NS> --sort-by=.lastTimestamp to see whether every pod or just one node pool is affected — single-pool failures point at node identity, fleet-wide at the registry.az aks check-acr --name <CLUSTER> --resource-group <RG> --acr <ACR_NAME>. It tests the pull path with the cluster's real identity. A failure here confirms the platform path is broken; success means the problem is in your pod spec (wrong image name or secret).az aks show --name <CLUSTER> --resource-group <RG> --query "{kubelet:identityProfile.kubeletidentity.objectId, cluster:identity.principalId}" -o table. Then run az role assignment list --assignee <KUBELET_ID> --scope <ACR_ID> -o table. If the kubelet ID has no AcrPull row, create it with az role assignment create.az acr repository show-tags --name <ACR> --repository <REPO> -o table and compare letter-for-letter with the image in your manifest. If the tag is missing, the registry never had it — check CI push logs. Pin digests in production manifests so tags can't drift under you.az acr show --name <ACR> --query "{fw:networkRuleSet.defaultAction, public:publicNetworkAccess}" and az acr network-rule list --name <ACR>. If defaultAction is Deny, confirm the cluster's egress IPs or VNet are allowlisted. Re-run az aks check-acr after any firewall edit.| File | Command / Code | Purpose |
|---|---|---|
| grant-acrpull.sh | az aks show --name <CLUSTER> --resource-group <RG> \ | The Kubelet Identity |
| check-acr-path.sh | az aks check-acr --name <CLUSTER> --resource-group <RG> --acr <ACR_NAME> | Validating the Pull Path With az aks check-acr |
| acrpull-secret-fallback.sh | kubectl create secret docker-registry acr-pull-secret \ | imagePullSecrets Fallback |
| check-acr-network.sh | az acr show --name <ACR> --query "{defaultAction:networkRuleSet.defaultAction, p... | When the Firewall, Not RBAC, Blocks the Pull |
Key takeaways
az aks check-acr after every cluster, identity, or firewall change.Common mistakes to avoid
5 patternsAssuming same-subscription ACR access works without a role assignment
az role assignment create --assignee <KUBELET_OBJECT_ID> --role AcrPull --scope <ACR_ID>. Then verify with az role assignment list --assignee <ID> --scope <ACR_ID>. Never assume same-subscription means same-permissions.Granting AcrPull to the cluster identity instead of the kubelet identity
az aks show --query identityProfile.kubeletidentity.objectId and assign the role to that object ID. After any cluster identity change, re-run the assignment before redeploying workloads.Using admin credentials in imagePullSecrets and forgetting they expire
az acr credential renew, store them in Key Vault, and sync them to the Kubernetes secret through an operator. Alert on secret age so rotation happens before expiry, not after an outage.Blaming RBAC when the image tag simply doesn't exist
az acr repository show-tags and pin deployments to immutable digests. Keep imagePullPolicy explicit so a missing tag surfaces as InvalidImageName instead of a misleading pull-backoff loop.Locking down ACR networking without an exception for the cluster
az aks check-acr after any network change.Interview Questions on This Topic
What does ImagePullBackOff mean when the registry is ACR?
kubectl describe pod for the event, then az aks check-acr to test the path.Frequently Asked Questions
20+ years shipping production backend systems. Written from production experience, not tutorials.
That's Azure. Mark it forged?
5 min read · try the examples if you haven't