Home › Cloud › GCP Quota Exceeded: CPUS_ALL_REGIONS Fix
Beginner 5 min · September 23, 2026

GCP Quota Exceeded: CPUS_ALL_REGIONS Fix

Fix GCP CPUS_ALL_REGIONS quota: check global vs regional usage, free idle VMs, request increases early, alert at 70% before launch day..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Production
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 9 min
  • ✓A GCP project with Compute Engine API enabled
  • ✓Google Cloud SDK installed with gcloud auth login done
  • ✓Basic familiarity with VM instances and regions
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • CPUS_ALL_REGIONS caps total vCPUs across every region: regional headroom means nothing when the sum is spent
  • Find the blocker with project-info describe globally plus regions describe locally, then free idle VMs
  • Stopped instances still consume quota: snapshot disks you need, then delete to release CPUs
  • File Edit Quotas increases days ahead, and alert at 70 percent on both levels so launches never stall
✦ Definition~90s read
What is GCP Quota Exceeded?

CPUS_ALL_REGIONS is a GCP Compute Engine quota metric capping the total number of virtual CPUs your project may use summed across all regions. It sits above the per-region CPUs quotas: each region limits its own vCPUs, while the aggregate limits the project's planetary total.

★
Think of GCP regions as branch offices and CPUs as company cars.

Creating a VM checks both — room in the region and room in the global sum — and fails naming whichever ceiling tripped. The metric counts running and stopped (but not deleted) instances against both levels.

Exhaustion has three usual authors. Idle accumulation: stopped and forgotten VMs in many regions each reserve CPUs that sum into the ceiling — the most common cause, and the cheapest to fix. Correlated bursts: autoscalers, MIGs, and batch jobs peaking together exceed a sum that fits each workload alone.

And genuine growth: the business outgrew the default limits, which start modestly and grow by request. All three present identically — failed creates naming CPUS_ALL_REGIONS — with different correct responses.

Resolution follows a fixed order. First read both levels: project-info describe for the global aggregate, regions describe for the local side. Then reclaim: snapshot worthy disks and delete idle instances everywhere, watching global usage fall. Then right-size the burst: cap autoscalers and stagger heavy jobs against the burst math.

Finally procure: file an Edit Quotas increase for the headroom growth genuinely needs.

Prevention converts this from incident to routine: utilization alerts at 70 and 90 percent on both levels, quarterly idle-instance reviews with ownership labels, burst math in every launch checklist, and quota filings tracked like any procurement with lead times. Quota is capacity you manage visibly — or capacity that manages your launches for you.

Plain-English First

Think of GCP regions as branch offices and CPUs as company cars. Each branch has its own parking limit (regional quota), but the company also caps total cars across all branches (CPUS_ALL_REGIONS). Your branch has empty spaces, yet you can't get a car — because forgotten cars sit in other branches' lots, and the company-wide cap is hit. The fix: collect the forgotten cars (delete idle VMs), or ask headquarters for a bigger fleet (quota increase) — days before you need it, since approvals lag.

Launch day arrives, you run the scale-up, and GCP answers: quota exceeded, CPUS_ALL_REGIONS. Not your region — all regions. Your target zone has free CPUs on paper, yet every instance create fails. The launch waits while you learn that GCP caps your total CPUs summed across the planet, and you've been spending that global budget in regions you forgot about.

CPUS_ALL_REGIONS is the aggregate ceiling over every regional CPU quota. Teams manage regions one by one and never watch the sum — until the sum stops them. Old dev VMs in three regions, a batch job in a fourth, and suddenly production can't grow in the fifth. The error names the global metric, but tired eyes read it as just another regional limit and request the wrong increase.

Worse, quota increases take days to approve. Filing the morning of the launch guarantees a delay no workaround fixes.

This guide shows the full play: read both quota levels from the CLI, find and free idle CPUs across regions, file the increase correctly with lead time, and set alerts so the next crunch pages you at 70 percent instead of failing you at 100. You'll see a real incident where forgotten dev instances ate a launch, plus the governance that keeps quota ahead of growth.

Two Ceilings: Regional Quotas and the Global Sum

GCP enforces CPU quotas at two levels and both can stop you. Each region has its own CPUs limit capping vCPUs in that region, and CPUS_ALL_REGIONS caps the sum of vCPUs across every region in the project. Launching an instance checks both: the region must have room and the global aggregate must have room. Either ceiling blocks the create, and the error names whichever one tripped — read it carefully, because teams routinely fix the wrong one.

The global aggregate surprises people because nothing in daily work shows the sum. You watch europe-west1 at 40 percent and feel safe while us-central1, asia-east1, and three dev regions quietly spend the shared budget. The Quotas console page can show the aggregate, but only if you filter for it; the default regional view hides the number that matters. Make the aggregate a first-class metric in your capacity reviews, not an error message you meet on launch day.

Internalize the mental math: regional quotas divide the budget, the global quota sizes it. Growing in one region while shrinking in another keeps the sum flat. Growing everywhere — or forgetting idle machines everywhere — pushes the sum into the ceiling. Every capacity decision is global whether you intended it or not, so check the sum before any launch that adds machines.

📊 Production Insight
One company ran regional quota dashboards for two years and still got stopped by the global sum during expansion into a second region. Their new capacity review opens with the aggregate number in giant type — the metric that actually gates growth leads the meeting.
🎯 Key Takeaway
Every create checks region and global sum — watch the aggregate, not just your home region.

Reading Both Quota Levels From the CLI

Reading quotas from the CLI takes seconds and settles every debate about which ceiling binds. project-info describe with a flatten filter shows the CPUS_ALL_REGIONS limit and usage for the whole project; regions describe shows the regional CPUs numbers for the target region. Run both before any launch, any increase request, and any incident call — the pair answers which level to fix and how much headroom exists.

Learn to read the output skeptically. Usage lags reality slightly during rapid scaling, so a reading of 95 percent during a burst means effectively full. Limits differ per machine family in some projects — N2, C2, and GPU families carry their own quotas alongside the general CPUs pool. If the general numbers look fine while creates fail, check the family-specific quota for the machine type you're launching.

Record both readings in launch checklists and incident notes. A launch ticket that states global 62 percent and regional 45 percent gives every approver the same facts; an incident note with both readings stops the next responder from re-running the same commands. Quota numbers are cheap to capture and expensive to re-derive under pressure — write them down where the team looks.

check-cpu-quotas.shBASH
1
2
3
4
5
6
7
8
9
# Global aggregate: limit vs usage for the whole project
 gcloud compute project-info describe --project=<PROJECT_ID> \
  --flatten='quotas[metric=CPUS_ALL_REGIONS]' --format='table(quotas.limit, quotas.usage)'

# Regional side: CPUs limit vs usage in the target region
gcloud compute regions describe <REGION> --project=<PROJECT_ID> --format='json(quotas)'

# What is actually holding CPUs? List every instance with status
 gcloud compute instances list --project=<PROJECT_ID> --format='table(name, zone, status, machineType)'
📊 Production Insight
A launch checklist that requires pasting both quota readings cut quota incidents to zero in one org — not because readings prevent exhaustion, but because writing the numbers forces someone to notice 91 percent before it becomes 100 percent at the worst moment.
🎯 Key Takeaway
project-info for the global sum, regions describe for local — run both, believe the tighter.

Freeing Idle CPUs Across Every Region

Idle instances are quota burned for nothing, and every project accumulates them: dev VMs from finished features, stopped test rigs, forgotten proof-of-concepts in faraway regions. Stopped instances still reserve their CPUs against quota — only deletion returns them. The cleanup pattern is mechanical: list everything with status, identify what nobody claims, snapshot disks worth keeping, delete the instances, and watch the global usage drop within minutes.

Ownership is what makes cleanup safe and repeatable. Require labels (owner, expiry, purpose) on every non-production instance and enforce them with policy constraints that deny unlabeled creates. A weekly report of instances past expiry goes to owners; unclaimed ones get stopped, then deleted after a grace period. Automation handles the schedule; humans handle the judgment calls about what matters.

Treat static IPs and disks as part of the same sweep. Detached static IPs cost money while holding nothing useful, and retained disks are cheap compared to the CPUs their instances reserved. Snapshot first, delete second, release IPs third. The few minutes of snapshot cost buy the confidence to delete aggressively — which is the only way the pool stays healthy.

free-idle-cpus.shBASH
1
2
3
4
5
6
7
8
# Find the waste: stopped and terminated instances across all regions
gcloud compute instances list --project=<PROJECT_ID> --filter='status:(TERMINATED,SUSPENDED)' --format='table(name, zone, status, machineType)'

# Snapshot a disk worth keeping before deleting its instance
gcloud compute disks snapshot <DISK> --zone=<ZONE> --snapshot-names=<DISK>-pre-cleanup --project=<PROJECT_ID>

# Delete the idle instance and release its CPUs back to quota
gcloud compute instances delete <INSTANCE> --zone=<ZONE> --project=<PROJECT_ID> --quiet
📊 Production Insight
One quarterly cleanup freed 40 percent of a project's global CPU quota — two years of forgotten dev machines. The team now runs expiry labels with auto-stop, and cleanups take minutes instead of archaeology. Quota hygiene compounds like any other hygiene.
🎯 Key Takeaway
Stopped still counts — snapshot disks, delete instances, watch global usage drop.

Taming Autoscaler Bursts That Eat the Aggregate

Autoscalers and batch jobs create the dramatic version of this failure: steady-state fits comfortably, then everything bursts at once. A GKE node pool scaling to max during a deploy, a MIG handling traffic, and a nightly ML batch starting early can collectively exceed the aggregate even though each fits alone. The failure strikes exactly when elasticity was supposed to save you, which makes it feel like betrayal rather than arithmetic.

Size for maximums, not averages. Add up the max replicas of every autoscaler, MIG, and job queue that can fire simultaneously, add steady-state base load, add 20 percent headroom, and compare against both quota levels. If the sum exceeds either ceiling, cap the scalers, stagger the windows, or raise the quota — before launch, not during. Load-test the combined burst; individual component tests never reveal aggregate overshoot.

Cap every scaler deliberately. An uncapped autoscaler with quota headroom is a incident that scales itself: it bursts, hits the ceiling, and starves every other workload of creates. Set max nodes per pool, max replicas per MIG, and concurrency limits per batch queue, each chosen from the burst math. Caps convert unbounded failure into bounded degradation — some requests wait instead of everything failing.

cap-autoscaler-burst.shBASH
1
2
3
4
5
6
7
8
# GKE: current nodes vs autoscaling maximums per pool
gcloud container clusters describe <CLUSTER> --region=<REGION> --project=<PROJECT_ID> --format='json(nodePools)' | python3 -c "import json,sys; [print(p['name'], p.get('initialNodeCount'), p.get('autoscaling',{})) for p in json.load(sys.stdin)]"

# MIGs: current vs max replicas that could burst at once
gcloud compute instance-groups managed list --project=<PROJECT_ID> --format='table(name, location, targetSize)'

# Cap a MIG so one workload cannot starve the quota pool
gcloud compute instance-groups managed set-autoscaling <MIG> --region=<REGION> --max-num-replicas=<CAP> --project=<PROJECT_ID>
📊 Production Insight
A Black Friday postmortem found three autoscalers whose combined maximum exceeded global quota by 3x — each sized sensibly alone. Capped maximums plus staggered batch windows turned the next peak into a non-event. Burst math belongs in every launch review with autoscaling.
🎯 Key Takeaway
Sum the maximums plus headroom; cap every scaler against the burst math.

Requesting More Quota the Right Way

When headroom is genuinely insufficient, request more through the proper channel with proper lead time. In the console, IAM & Admin > Quotas, filter for CPUS_ALL_REGIONS (and the regional CPUs metric if that's tight), select, and Edit Quotas with a clear business justification: what launches, when, how many CPUs, why existing quota can't cover it. Vague requests wait; specific justified ones move. Expect days, not minutes — capacity planning happens on Google's side too.

File at 70 percent, never at 100 percent the day before launch. Track each quota's approval lead time in your runbook so forecasts convert to filing dates automatically. A quota request is capacity procurement with a human in the loop; treat it with the same seriousness as hardware orders, because the lead-time dynamics are identical.

While waiting, buy room with the cleanup and capping moves from earlier sections — they work in minutes and often cover the gap. Never treat an increase request as the only plan; pair every filing with immediate reclamation so the launch has two paths to green. And record the new limits in capacity docs the day they're approved, or the next planner starts from stale numbers and repeats the whole cycle.

quota-increase-evidence.shBASH
1
2
3
4
5
6
# Document the filing with exact numbers: current usage, limit, requested limit
 gcloud compute project-info describe --project=<PROJECT_ID> --flatten='quotas[metric=CPUS_ALL_REGIONS]' --format='json(quotas.limit, quotas.usage)' > /tmp/quota-evidence-$(date +%F).json
cat /tmp/quota-evidence-*.json

# After approval lands, confirm the new ceiling before launching
gcloud compute project-info describe --project=<PROJECT_ID> --flatten='quotas[metric=CPUS_ALL_REGIONS]' --format='table(quotas.limit, quotas.usage)'
📊 Production Insight
Teams that file quota requests with pasted CLI evidence get approvals faster — reviewers see exact usage, exact need, and exact timeline instead of round numbers and urgency. Evidence converts a plea into procurement.
🎯 Key Takeaway
Console Edit Quotas with justification at 70 percent — approvals take days, so file early.

Staying Ahead of Quota Forever

Prevention is a dashboard plus a calendar. The dashboard shows global and regional CPU utilization against limits with alert policies at 70 percent (plan) and 90 percent (act now), reviewed in the weekly ops meeting. The calendar holds quarterly cleanup reviews, pre-launch burst-math sign-offs, and quota-filing deadlines derived from tracked lead times. Together they move quota from emergency to routine.

Wire alerts to the team that can file increases, not a general channel where everyone assumes someone else acts. Each alert links the capacity runbook: current readings, how to free idle CPUs in minutes, who approves filings, and the justification template. An alert without an owner and a playbook is just a notification of future failure.

Close the loop with launch checklists that require quota evidence: both readings pasted, burst math attached, increase case numbers referenced. Launches that can't show headroom don't ship until they can — a rule that feels bureaucratic exactly until the first time it saves a launch. Quota governance is capacity planning made visible, and visible planning rarely fails at midnight. Keep the dashboard green and launches stay boring — boring launches are the goal.

⚠ Fresh Quota Readings for Every Launch
Never launch headcount-scale capacity on quota you haven't verified the same week. Quota gets consumed by other teams' bursts, cleanups get reverted, and approvals expire. A reading from last quarter is a rumor — paste fresh project-info and regions output into every launch ticket.
📊 Production Insight
Orgs that review quota utilization weekly report the same arc: three months of boring green dashboards, then one early catch that would have been a launch-stopper. The boring meetings are the product — uneventful launches are manufactured, not lucky.
🎯 Key Takeaway
Dashboard plus calendar plus checklist — quota becomes routine instead of emergency.
● Production incidentPOST-MORTEMseverity: high

The Launch That 60 Forgotten Dev VMs Ate

Symptom
Every instance create failed project-wide with quota exceeded CPUS_ALL_REGIONS, while each regional dashboard showed free CPUs. The launch runbook scale-up step went red in all regions at once — the signature of a global ceiling, not a regional shortage.
Assumption
The team assumed the target region was out of capacity and waited for GCP to free space, losing two hours. Then they assumed the quota page was stale and requested a regional increase — which was approved quickly and changed nothing, because the regional limit was never the blocker. Each theory cost half the launch window.
Root cause
Sixty stopped dev VMs across three regions still reserved their CPUs against the global aggregate, compounded by a batch job bursting that same hour. Regional quotas all showed headroom, but summed usage had hit CPUS_ALL_REGIONS. Nobody watched the aggregate because every dashboard was regional.
Fix
The fix came in two parts. First, they deleted 60 stopped dev VMs across three regions (after snapshotting two disks worth keeping), freeing enough global CPUs to launch immediately. Second, they filed a CPUS_ALL_REGIONS increase with headroom for growth plus per-region alerts at 70 percent. The launch slipped four hours; the next three launches used quota headroom booked weeks ahead.
Key lesson
  • Stopped does not mean freed. The team treated stopped VMs as returned capacity for months — the launch taught them that only deletion releases quota, and the lesson now sits in onboarding docs.
  • Request the increase the forecast demands, not the error names. The regional approval felt like progress while the global ceiling stayed fixed — always fix the level that's actually exhausted.
  • Quota is capacity planning, not paperwork. Treating increases as launch-week admin guarantees delays; treating them as forecasted capacity with lead times keeps launches boring.
Production debug guideFive checks, in order, from quota levels to alerts that prevent repeats.5 entries
Symptom · 01
Creates fail and you don't know which quota level bit
→
Fix
Run gcloud compute project-info describe --project=<PROJECT> --flatten='quotas[metric=CPUS_ALL_REGIONS]' --format='table(limit, usage)' for the global ceiling, then gcloud compute regions describe <REGION> --format='json(quotas)' for the regional side. Whichever shows usage at limit is your blocker — fix that level.
Symptom · 02
Quota is exhausted but nobody knows by what
→
Fix
Run gcloud compute instances list --format='table(name, zone, status, machineType)' and sort by status. Terminated and long-stopped instances are pure quota waste — snapshot anything worth keeping with gcloud compute disks snapshot, then delete the instances to release CPUs.
Symptom · 03
Usage spikes unpredictably with autoscaling
→
Fix
Run gcloud container clusters list --format='table(name, location, currentNodeCount)' and check node-pool autoscaling maximums, plus gcloud compute instance-groups managed list for MIG sizes. Sum the maximums: if autoscalers can collectively burst past quota, cap them before the next scale event.
Symptom · 04
You genuinely need a higher ceiling
→
Fix
Open IAM & Admin > Quotas, filter metric CPUS_ALL_REGIONS (and the regional CPUs metric), select the binding quota, and click Edit Quotas with a business justification and the needed limit. Note the case number and expected timeline — then plan the launch around it, not through it.
Symptom · 05
Nobody saw the crunch coming
→
Fix
Build the utilization view the error should have come with: global and regional usage versus limit on one dashboard, with alert policies at 70 and 90 percent. Review it in the weekly ops meeting while headroom still exists to act on.
CPUS_ALL_REGIONS Quota Failures Compared
Root CauseHow to ConfirmFixPrevention
Global aggregate exhaustedproject-info quotas show CPUS_ALL_REGIONS at limit while regions look freeFree CPUs globally or file an Edit Quotas increase requestAlert at 70 percent on the global metric; forecast peaks
Regional quota exhaustedRegion describe shows CPUs at limit; creates fail in one region onlyClean up in that region or request a regional increasePer-region dashboards; keep lone-region workloads balanced
Idle resources hoarding quotainstances list shows stopped and forgotten VMs holding CPUsDelete or stop idle instances; snapshot disks firstQuarterly ownership audits; auto-stop policies for dev
Autoscaler burst overshootMIG/GKE max sizes sum above quota; failures coincide with scale eventsCap max replicas; stagger batch and deploy windowsSum maximums when sizing; load-test burst before launch
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
check-cpu-quotas.shgcloud compute project-info describe --project=<PROJECT_ID> \Reading Both Quota Levels From the CLI
free-idle-cpus.shgcloud compute instances list --project=<PROJECT_ID> --filter='status:(TERMINATE...Freeing Idle CPUs Across Every Region
cap-autoscaler-burst.shgcloud container clusters describe <CLUSTER> --region=<REGION> --project=<PROJEC...Taming Autoscaler Bursts That Eat the Aggregate
quota-increase-evidence.shgcloud compute project-info describe --project=<PROJECT_ID> --flatten='quotas[me...Requesting More Quota the Right Way

Key takeaways

1
CPUS_ALL_REGIONS caps vCPUs summed across all regions
regional headroom means nothing when the sum is spent.
2
Check both levels with project-info describe plus regions describe before requesting anything.
3
Stopped VMs still consume quota; only deletion releases CPUs
snapshot disks, then delete.
4
File Edit Quotas increases days ahead with justification; approvals can't be rushed.
5
Size for combined burst maximums plus 20 percent, and cap every autoscaler.
6
Alert at 70 and 90 percent on regional and global quotas; review in weekly ops.

Common mistakes to avoid

5 patterns
×

Requesting regional quota when the global aggregate is the blocker

Symptom
The region shows free CPUs while every create still fails. The team raises the regional limit twice and nothing changes, because CPUS_ALL_REGIONS — the sum across regions — was the actual ceiling all along.
Fix
Read both levels before launching: gcloud compute project-info describe --flatten='quotas[metric=CPUS_ALL_REGIONS]' for the global ceiling and per-region CPU quotas for the region you target. Request headroom at whichever level is tightest, not just the one the error named.
×

Hoarding idle VMs across regions until launch day

Symptom
Dev instances from three quarters ago consume the global quota while production can't scale. The cleanup that should have happened in peacetime happens during the launch, with everyone watching.
Fix
Audit quarterly with gcloud compute instances list and a label-based ownership report. Delete or stop what nobody claims, snapshot disks you might need, and release the static IPs attached to dead instances.
×

Sizing quota for steady state instead of peak burst

Symptom
Everything fits on a normal Tuesday, then a deploy, a batch job, and an autoscaling event coincide and exhaust the aggregate. The failure lands exactly when elasticity mattered most.
Fix
Sum maximums, not currents: GKE autoscalers, MIGs, and batch jobs can all burst simultaneously. Request quota for the combined peak plus 20 percent headroom, and cap autoscalers so one workload can't starve the rest.
×

Filing the quota increase the day before launch

Symptom
The request sits in review while the launch waits. Quota approvals involve capacity planning on Google's side and can't be rushed, so last-minute filings become launch delays with no workaround.
Fix
File quota increases the moment forecasts show pressure — approvals take days, not minutes. Track lead time per quota in your runbook and trigger requests at 70 percent utilization, never at launch minus one day.
×

Running production with no quota utilization alerts

Symptom
Utilization climbs silently for months until an ordinary scale-up fails. The postmortem finds the trend was visible for half a year to nobody, because no chart existed and no alert fired.
Fix
Alert at 70 and 90 percent on both regional and global CPU quotas, routed to the team that can file increases. Review the dashboard in the weekly ops meeting while there's still time to act on it.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Your launches fail everywhere with CPUS_ALL_REGIONS exceeded. What does ...
Q02JUNIOR
How do you free CPU quota without harming anything?
Q03SENIOR
Regional quota looks fine but creates still fail. What's your order of c...
Q04SENIOR
How do you size quota for autoscaling workloads?
Q05SENIOR
Design quota governance so launches never stall on CPUs.
Q01 of 05JUNIOR

Your launches fail everywhere with CPUS_ALL_REGIONS exceeded. What does it mean?

ANSWER
CPUS_ALL_REGIONS caps total vCPUs summed across every region, separate from each region's own CPU limit. Launching fails when the sum is exhausted even if the target region looks free. I'd confirm with project-info describe, free idle resources globally, and file an Edit Quotas increase if headroom is genuinely needed.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Regional CPU quota vs CPUS_ALL_REGIONS: what's the difference?
02
How do I request a quota increase?
03
How do I check current quota usage from the CLI?
04
Do stopped VMs still consume CPU quota?
05
What usually causes sudden quota exhaustion?
06
How do I get warned before quota runs out?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Verified
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
🔥

That's GCP. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
GCP Cloud Run Container Failed to Start and Listen on PORT
3 / 3 · GCP