GCP Quota Exceeded: CPUS_ALL_REGIONS Fix
Fix GCP CPUS_ALL_REGIONS quota: check global vs regional usage, free idle VMs, request increases early, alert at 70% before launch day..
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
- ✓A GCP project with Compute Engine API enabled
- ✓Google Cloud SDK installed with gcloud auth login done
- ✓Basic familiarity with VM instances and regions
- CPUS_ALL_REGIONS caps total vCPUs across every region: regional headroom means nothing when the sum is spent
- Find the blocker with project-info describe globally plus regions describe locally, then free idle VMs
- Stopped instances still consume quota: snapshot disks you need, then delete to release CPUs
- File Edit Quotas increases days ahead, and alert at 70 percent on both levels so launches never stall
Think of GCP regions as branch offices and CPUs as company cars. Each branch has its own parking limit (regional quota), but the company also caps total cars across all branches (CPUS_ALL_REGIONS). Your branch has empty spaces, yet you can't get a car — because forgotten cars sit in other branches' lots, and the company-wide cap is hit. The fix: collect the forgotten cars (delete idle VMs), or ask headquarters for a bigger fleet (quota increase) — days before you need it, since approvals lag.
Launch day arrives, you run the scale-up, and GCP answers: quota exceeded, CPUS_ALL_REGIONS. Not your region — all regions. Your target zone has free CPUs on paper, yet every instance create fails. The launch waits while you learn that GCP caps your total CPUs summed across the planet, and you've been spending that global budget in regions you forgot about.
CPUS_ALL_REGIONS is the aggregate ceiling over every regional CPU quota. Teams manage regions one by one and never watch the sum — until the sum stops them. Old dev VMs in three regions, a batch job in a fourth, and suddenly production can't grow in the fifth. The error names the global metric, but tired eyes read it as just another regional limit and request the wrong increase.
Worse, quota increases take days to approve. Filing the morning of the launch guarantees a delay no workaround fixes.
This guide shows the full play: read both quota levels from the CLI, find and free idle CPUs across regions, file the increase correctly with lead time, and set alerts so the next crunch pages you at 70 percent instead of failing you at 100. You'll see a real incident where forgotten dev instances ate a launch, plus the governance that keeps quota ahead of growth.
Two Ceilings: Regional Quotas and the Global Sum
GCP enforces CPU quotas at two levels and both can stop you. Each region has its own CPUs limit capping vCPUs in that region, and CPUS_ALL_REGIONS caps the sum of vCPUs across every region in the project. Launching an instance checks both: the region must have room and the global aggregate must have room. Either ceiling blocks the create, and the error names whichever one tripped — read it carefully, because teams routinely fix the wrong one.
The global aggregate surprises people because nothing in daily work shows the sum. You watch europe-west1 at 40 percent and feel safe while us-central1, asia-east1, and three dev regions quietly spend the shared budget. The Quotas console page can show the aggregate, but only if you filter for it; the default regional view hides the number that matters. Make the aggregate a first-class metric in your capacity reviews, not an error message you meet on launch day.
Internalize the mental math: regional quotas divide the budget, the global quota sizes it. Growing in one region while shrinking in another keeps the sum flat. Growing everywhere — or forgetting idle machines everywhere — pushes the sum into the ceiling. Every capacity decision is global whether you intended it or not, so check the sum before any launch that adds machines.
Reading Both Quota Levels From the CLI
Reading quotas from the CLI takes seconds and settles every debate about which ceiling binds. project-info describe with a flatten filter shows the CPUS_ALL_REGIONS limit and usage for the whole project; regions describe shows the regional CPUs numbers for the target region. Run both before any launch, any increase request, and any incident call — the pair answers which level to fix and how much headroom exists.
Learn to read the output skeptically. Usage lags reality slightly during rapid scaling, so a reading of 95 percent during a burst means effectively full. Limits differ per machine family in some projects — N2, C2, and GPU families carry their own quotas alongside the general CPUs pool. If the general numbers look fine while creates fail, check the family-specific quota for the machine type you're launching.
Record both readings in launch checklists and incident notes. A launch ticket that states global 62 percent and regional 45 percent gives every approver the same facts; an incident note with both readings stops the next responder from re-running the same commands. Quota numbers are cheap to capture and expensive to re-derive under pressure — write them down where the team looks.
Freeing Idle CPUs Across Every Region
Idle instances are quota burned for nothing, and every project accumulates them: dev VMs from finished features, stopped test rigs, forgotten proof-of-concepts in faraway regions. Stopped instances still reserve their CPUs against quota — only deletion returns them. The cleanup pattern is mechanical: list everything with status, identify what nobody claims, snapshot disks worth keeping, delete the instances, and watch the global usage drop within minutes.
Ownership is what makes cleanup safe and repeatable. Require labels (owner, expiry, purpose) on every non-production instance and enforce them with policy constraints that deny unlabeled creates. A weekly report of instances past expiry goes to owners; unclaimed ones get stopped, then deleted after a grace period. Automation handles the schedule; humans handle the judgment calls about what matters.
Treat static IPs and disks as part of the same sweep. Detached static IPs cost money while holding nothing useful, and retained disks are cheap compared to the CPUs their instances reserved. Snapshot first, delete second, release IPs third. The few minutes of snapshot cost buy the confidence to delete aggressively — which is the only way the pool stays healthy.
Taming Autoscaler Bursts That Eat the Aggregate
Autoscalers and batch jobs create the dramatic version of this failure: steady-state fits comfortably, then everything bursts at once. A GKE node pool scaling to max during a deploy, a MIG handling traffic, and a nightly ML batch starting early can collectively exceed the aggregate even though each fits alone. The failure strikes exactly when elasticity was supposed to save you, which makes it feel like betrayal rather than arithmetic.
Size for maximums, not averages. Add up the max replicas of every autoscaler, MIG, and job queue that can fire simultaneously, add steady-state base load, add 20 percent headroom, and compare against both quota levels. If the sum exceeds either ceiling, cap the scalers, stagger the windows, or raise the quota — before launch, not during. Load-test the combined burst; individual component tests never reveal aggregate overshoot.
Cap every scaler deliberately. An uncapped autoscaler with quota headroom is a incident that scales itself: it bursts, hits the ceiling, and starves every other workload of creates. Set max nodes per pool, max replicas per MIG, and concurrency limits per batch queue, each chosen from the burst math. Caps convert unbounded failure into bounded degradation — some requests wait instead of everything failing.
Requesting More Quota the Right Way
When headroom is genuinely insufficient, request more through the proper channel with proper lead time. In the console, IAM & Admin > Quotas, filter for CPUS_ALL_REGIONS (and the regional CPUs metric if that's tight), select, and Edit Quotas with a clear business justification: what launches, when, how many CPUs, why existing quota can't cover it. Vague requests wait; specific justified ones move. Expect days, not minutes — capacity planning happens on Google's side too.
File at 70 percent, never at 100 percent the day before launch. Track each quota's approval lead time in your runbook so forecasts convert to filing dates automatically. A quota request is capacity procurement with a human in the loop; treat it with the same seriousness as hardware orders, because the lead-time dynamics are identical.
While waiting, buy room with the cleanup and capping moves from earlier sections — they work in minutes and often cover the gap. Never treat an increase request as the only plan; pair every filing with immediate reclamation so the launch has two paths to green. And record the new limits in capacity docs the day they're approved, or the next planner starts from stale numbers and repeats the whole cycle.
Staying Ahead of Quota Forever
Prevention is a dashboard plus a calendar. The dashboard shows global and regional CPU utilization against limits with alert policies at 70 percent (plan) and 90 percent (act now), reviewed in the weekly ops meeting. The calendar holds quarterly cleanup reviews, pre-launch burst-math sign-offs, and quota-filing deadlines derived from tracked lead times. Together they move quota from emergency to routine.
Wire alerts to the team that can file increases, not a general channel where everyone assumes someone else acts. Each alert links the capacity runbook: current readings, how to free idle CPUs in minutes, who approves filings, and the justification template. An alert without an owner and a playbook is just a notification of future failure.
Close the loop with launch checklists that require quota evidence: both readings pasted, burst math attached, increase case numbers referenced. Launches that can't show headroom don't ship until they can — a rule that feels bureaucratic exactly until the first time it saves a launch. Quota governance is capacity planning made visible, and visible planning rarely fails at midnight. Keep the dashboard green and launches stay boring — boring launches are the goal.
The Launch That 60 Forgotten Dev VMs Ate
- Stopped does not mean freed. The team treated stopped VMs as returned capacity for months — the launch taught them that only deletion releases quota, and the lesson now sits in onboarding docs.
- Request the increase the forecast demands, not the error names. The regional approval felt like progress while the global ceiling stayed fixed — always fix the level that's actually exhausted.
- Quota is capacity planning, not paperwork. Treating increases as launch-week admin guarantees delays; treating them as forecasted capacity with lead times keeps launches boring.
gcloud compute project-info describe --project=<PROJECT> --flatten='quotas[metric=CPUS_ALL_REGIONS]' --format='table(limit, usage)' for the global ceiling, then gcloud compute regions describe <REGION> --format='json(quotas)' for the regional side. Whichever shows usage at limit is your blocker — fix that level.gcloud compute instances list --format='table(name, zone, status, machineType)' and sort by status. Terminated and long-stopped instances are pure quota waste — snapshot anything worth keeping with gcloud compute disks snapshot, then delete the instances to release CPUs.gcloud container clusters list --format='table(name, location, currentNodeCount)' and check node-pool autoscaling maximums, plus gcloud compute instance-groups managed list for MIG sizes. Sum the maximums: if autoscalers can collectively burst past quota, cap them before the next scale event.| File | Command / Code | Purpose |
|---|---|---|
| check-cpu-quotas.sh | gcloud compute project-info describe --project=<PROJECT_ID> \ | Reading Both Quota Levels From the CLI |
| free-idle-cpus.sh | gcloud compute instances list --project=<PROJECT_ID> --filter='status:(TERMINATE... | Freeing Idle CPUs Across Every Region |
| cap-autoscaler-burst.sh | gcloud container clusters describe <CLUSTER> --region=<REGION> --project=<PROJEC... | Taming Autoscaler Bursts That Eat the Aggregate |
| quota-increase-evidence.sh | gcloud compute project-info describe --project=<PROJECT_ID> --flatten='quotas[me... | Requesting More Quota the Right Way |
Key takeaways
Common mistakes to avoid
5 patternsRequesting regional quota when the global aggregate is the blocker
gcloud compute project-info describe --flatten='quotas[metric=CPUS_ALL_REGIONS]' for the global ceiling and per-region CPU quotas for the region you target. Request headroom at whichever level is tightest, not just the one the error named.Hoarding idle VMs across regions until launch day
gcloud compute instances list and a label-based ownership report. Delete or stop what nobody claims, snapshot disks you might need, and release the static IPs attached to dead instances.Sizing quota for steady state instead of peak burst
Filing the quota increase the day before launch
Running production with no quota utilization alerts
Interview Questions on This Topic
Your launches fail everywhere with CPUS_ALL_REGIONS exceeded. What does it mean?
Frequently Asked Questions
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
That's GCP. Mark it forged?
5 min read · try the examples if you haven't