Home › Cloud › Cloud Run PORT Failure: Container Won't Start
Beginner 5 min · September 23, 2026

Cloud Run PORT Failure: Container Won't Start

Fix Cloud Run PORT failure by listening on $PORT bound to 0.0.0.0: read revision logs, fix the bind, verify with Docker first..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Production
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 10 min
  • ✓A containerized app you can build with Docker
  • ✓A GCP project with Cloud Run API enabled
  • ✓Google Cloud SDK installed with gcloud auth login done
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Cloud Run injects $PORT (default 8080) and probes exactly that: silence there fails the revision
  • Read revision logs first, then reproduce locally with docker run -p plus -e PORT
  • Fix the listen call: read $PORT with an 8080 fallback and bind 0.0.0.0, never localhost
  • Listen before heavy init, use exec-form CMD, and log the bound port at startup
  • Use CPU boost and timeout as breathing room while slow init moves after listen
✦ Definition~90s read
What is GCP Cloud Run Container Failed to Start and Listen on PORT?

Cloud Run runs your container image as a revision and routes HTTP traffic to it only after a startup probe confirms something is listening. The probe target comes from the PORT environment variable, which Cloud Run injects with a default of 8080 (overridable per service).

★
Think of Cloud Run as a hotel that tells every guest their room number on a card (the PORT variable, usually 8080) and then knocks on that door to check you arrived.

Your container must open a TCP listener on that port, on all network interfaces (0.0.0.0), reasonably soon after starting. Miss any part — wrong port, localhost-only bind, too-slow startup — and the revision fails with the container-failed-to-listen error.

The localhost trap deserves emphasis because it fools experienced developers. Inside the container, 127.0.0.1 works for in-container curls and looks alive in logs. But the Cloud Run prober connects from outside the container instance, where the container's loopback is unreachable.

Only 0.0.0.0 (or a specific external interface) exposes the port to the probe. The same code passes every laptop test and fails every Cloud Run deploy — the environments differ in where the client sits.

The Dockerfile's CMD form controls shutdown behavior that compounds startup issues. Shell-form CMD wraps the app in a shell that swallows SIGTERM, so revisions hang during traffic shifts and scale-to-zero. Exec-form CMD (JSON array) makes the app PID 1, delivering signals directly.

Combined with a SIGTERM handler that drains quickly, exec form keeps the whole revision lifecycle — start, serve, stop — clean.

Diagnosis always starts at the revision logs: they record the framework banner, the bound address, and the probe outcome in order. Compare expected PORT against the actual bind, reproduce locally with injected PORT and host-side curl, fix the listen call, and redeploy. The loop takes minutes once the contract is understood — and seconds once it's templated.

Plain-English First

Think of Cloud Run as a hotel that tells every guest their room number on a card (the PORT variable, usually 8080) and then knocks on that door to check you arrived. PORT failure means you went to a different room (hardcoded 3000), hid under the bed where knocks can't reach (localhost bind), or spent an hour unpacking before opening the door (slow init). The hotel isn't broken — you missed the instructions on the card. Read the card, open that door promptly, and answer when it knocks.

You containerize the app, push the image, deploy to Cloud Run, and get the dreaded revision error: container failed to start and then listen on the port. The image runs perfectly on your laptop. The Dockerfile looks right. Yet Cloud Run insists nothing is listening, kills the revision, and leaves you staring at a port number you never chose.

That port number is the whole story. Cloud Run injects a PORT environment variable (8080 by default) and probes your container on exactly that port. If your app listens on a hardcoded 3000, binds only to localhost, or spends three minutes downloading a model before listening, the probe finds silence and the revision fails. Your laptop never enforces this contract, so the bug ships undetected every time.

Beginners hit this on their first Cloud Run deploy; veterans hit it when a framework default changes or a second Dockerfile drifts. Either way the fix is small and the lesson permanent.

This guide walks the diagnosis in order: read the revision logs, compare expected versus actual port, fix the listen call, correct the Dockerfile, and verify locally with the same contract. You'll see a real launch-day incident caused by a hardcoded port, plus the template that makes this failure nearly impossible.

The PORT Contract: What Cloud Run Actually Probes

Cloud Run's contract with your container is short: listen on the port in the PORT environment variable, on all interfaces, promptly after starting. Cloud Run sets PORT to 8080 unless you configure otherwise, then probes that port to decide whether the revision is healthy. Answer the probe and traffic flows; miss it and the revision fails with the container-failed-to-listen error. There is no negotiation, no fallback scan of other ports, and no credit for listening beautifully on the wrong number.

Three violations cover nearly every failure. A hardcoded port (3000 in Node tutorials, 5000 in Flask defaults, 8000 in Django habits) ignores PORT entirely. A localhost bind answers only inside the container while the prober knocks from outside. Slow initialization — model downloads, migrations, cache warmups — delays the bind past the probe window. All three produce the identical revision error, which is why the message alone never identifies the cause.

Your laptop enforces none of this. Docker run without -e PORT leaves your hardcoded default working; curl from inside the container reaches localhost binds; no prober times your startup. The gap between localdocker and Cloud Run is exactly the set of behaviors this error reports. Close the gap deliberately — reproduce with the platform's terms — and the error becomes a pre-deploy catch instead of a launch-day surprise.

📊 Production Insight
Launch-day PORT failures follow a pattern: the demo ran on localhost:3000 all through development, and nobody rehearsed the platform contract until users waited. One team now deploys to a staging Cloud Run service from the first week — the contract gets tested before anything matters.
🎯 Key Takeaway
PORT env plus all interfaces plus prompt bind — miss any of the three and the revision fails.

Reading Revision Logs: What the Container Really Did

Revision logs are the ground truth because they show what the container actually did, not what the Dockerfile intended. Cloud Run streams container stdout and stderr per revision, including framework startup banners that usually print the bound address. A line like listening on 0.0.0.0:8080 is the all-clear; listening on 127.0.0.1:3000 names both bugs at once; no listen line at all means the app crashed or is still initializing.

Read them with intent. Filter to the failed revision and scan from the container start marker: dependency installs, framework banner, listen line, then probe results. The time gap between start and listen is your startup duration — if it spans minutes, initialization order is the problem regardless of the port. Errors above the listen line (missing modules, failed migrations) are startup crashes wearing a PORT costume; fix them first and the port complaint often vanishes.

Make the logs work for you permanently by logging the effective port and address at startup in every service. One line — serving on 0.0.0.0:8080 — turns every future PORT incident into a ten-second diagnosis. The snippet below pulls the right logs for the failed revision so you stop guessing from the error message alone.

read-revision-logs.shBASH
1
2
3
4
5
6
# What does Cloud Run expect? Check service env and container config
gcloud run services describe <SERVICE> --region <REGION> --format="json(spec.template.spec.containers)"

# What did the container actually do? Read this revision's logs
REVISION=$(gcloud run revisions list --service <SERVICE> --region <REGION> --limit 1 --format="value(name)")
gcloud logging read "resource.type=cloud_run_revision AND resource.labels.revision_name=${REVISION}" --limit=50 --format=json | python3 -c "import json,sys; [print(e.get('textPayload','')) for e in json.load(sys.stdin)]"
📊 Production Insight
A team once debugged a PORT failure for an hour before discovering they were reading the previous revision's logs. Listing revisions first and filtering to the latest failure is now step zero in their runbook — check you're reading the right revision before reading anything else.
🎯 Key Takeaway
Filter to the failed revision; the listen line (or its absence) names the bug.

Fixing the Listen Call in Any Stack

The listen call is where the fix lands, and every stack expresses it slightly differently. Node's app.listen(process.env.PORT || 8080, '0.0.0.0'), Python's app.run(host='0.0.0.0', port=int(os.environ.get('PORT', 8080))), Go's net.Listen on ':' + os.Getenv('PORT') — same contract, different syntax. The two non-negotiables: the port comes from the environment with an 8080 fallback for local runs, and the host is the all-interfaces address, never localhost.

Framework defaults actively work against you here. Express tutorials hardcode 3000, Flask defaults to 127.0.0.1:5000, and several frameworks bind localhost when no host is given. Scaffolding that runs beautifully on a laptop ships all three violations at once. Treat every framework default as guilty until proven bound to $PORT on 0.0.0.0 — a one-line review that prevents the most common Cloud Run failure in existence.

Verify the fix where Cloud Run can't argue with it: a local container run with the platform's terms. Inject PORT, publish the port, and curl from the host (outside the container). Host-reachable means prober-reachable; in-container-only success means nothing. Keep this exact command in the service README so every contributor rehearses the contract, not just the deployer.

verify-port-locally.shBASH
1
2
3
4
5
6
7
# Reproduce the Cloud Run contract locally: inject PORT, publish it, curl from HOST
 docker build -t portcheck . && docker run -d --name portcheck -p 8080:8080 -e PORT=8080 portcheck
sleep 5; curl -s -o /dev/null -w "HTTP %{http_code}\n" http://localhost:8080/ || echo "FAILED from host: same silence Cloud Run hears"

# Suspect a localhost bind? Compare in-container vs host reachability
docker exec portcheck sh -c "wget -q -O /dev/null http://127.0.0.1:8080/ && echo reachable-inside" || echo "dead-inside-too"
docker stop portcheck && docker rm portcheck
📊 Production Insight
Flask's default bind (127.0.0.1:5000) has caused more Cloud Run PORT failures than any other single default. Teams standardizing on Python now ship a run snippet with host and PORT preset — the framework default never reaches production because the template overwrites it.
🎯 Key Takeaway
$PORT with 8080 fallback on 0.0.0.0 — then prove it with host-side curl.

Dockerfile Discipline: Exec CMD and Fast Starts

The Dockerfile decides process behavior that code can't override. Exec-form CMD (JSON array: CMD ["node", "server.js"]) runs your app as PID 1, delivering SIGTERM straight to it for graceful shutdown. Shell-form CMD (CMD node server.js) wraps your app in /bin/sh, which swallows SIGTERM — deploys hang, scale-to-zero stalls, and revisions linger in half-dead states. Cloud Run sends SIGTERM on traffic shifts and instance shutdowns constantly, so signal handling isn't edge-case hygiene, it's the normal path.

Port exposure deserves a matching review. EXPOSE documents intent but enforces nothing — Cloud Run probes PORT regardless of EXPOSE declarations. Still, keep EXPOSE aligned with your default port for local tooling that does honor it, and never rely on it as the fix. The listen call is the fix; the Dockerfile is the delivery mechanism.

Order inside the Dockerfile matters for startup speed too. Dependency installation belongs in early cached layers; the listen-first application code goes last. Multi-stage builds keep production images lean, which shortens cold starts and narrows the probe window pressure. Review Dockerfiles with the same seriousness as the listen call — the container that starts fast, binds right, and shuts down cleanly is built, not wished for.

audit-container-startup.shBASH
1
2
3
4
5
6
7
8
# Inspect the image's startup command and exposed ports before deploying
docker inspect <IMAGE> --format='CMD={{json .Config.Cmd}} Entrypoint={{json .Config.Entrypoint}} Ports={{json .Config.ExposedPorts}}'

# Confirm graceful shutdown: stop should exit in seconds, not hang
docker run -d --name shutdown-test <IMAGE> && sleep 3 && time docker stop shutdown-test && docker rm shutdown-test

# Deploy with explicit CPU boost while optimizing slow starters
gcloud run deploy <SERVICE> --image <IMAGE> --region <REGION> --cpu-boost
📊 Production Insight
Shell-form CMD once stalled a Black Friday scale-down: instances ignored SIGTERM, held connections open, and burned budget for hours. The postmortem standardized exec-form CMD org-wide — a one-line Dockerfile convention that paid for itself the next traffic spike.
🎯 Key Takeaway
Exec-form CMD, lean layers, SIGTERM handled — the container must start fast and stop clean.

Startup Order: Listen First, Initialize After

Startup order bugs hide behind correct ports. The container binds 8080 properly — after downloading a 2 GB model, running migrations, and warming caches for four minutes. The prober's patience expires first, and a perfectly addressed server still fails the revision. The fix is sequencing: bind the port and answer health checks first, then complete heavy initialization in the background or ahead of time at build.

Several patterns keep listen-first honest. Bake models and assets into the image at build time so startup reads local disk instead of the network. Run migrations as a separate Cloud Run job or Cloud Build step rather than inside the serving container's boot. Warm caches lazily after first response, accepting a slower first request instead of risking the whole revision. Each pattern trades a little runtime elegance for startup certainty — the right trade on a probed platform.

Measure what you sequence. Startup duration (container start to first listen) belongs on a dashboard per service, with alerts when it approaches the probe budget. Teams that watch this metric catch dependency bloat — the SDK that added forty seconds, the model that doubled — while it's still a graph trend, not a failed deploy. Fast startup is a feature you monitor, not luck you hope for.

measure-startup-time.shBASH
1
2
3
4
5
6
# Time your startup: how long from container start to first listen?
REVISION=$(gcloud run revisions list --service <SERVICE> --region <REGION> --limit 1 --format="value(name)")
gcloud logging read "resource.type=cloud_run_revision AND resource.labels.revision_name=${REVISION}" --limit=100 --format="table(timestamp, textPayload)" | head -40

# Give legitimate slow starters room while you optimize (breathing room, not the fix)
gcloud run services update <SERVICE> --region <REGION> --cpu-boost --timeout=300
📊 Production Insight
An ML service once downloaded weights at every cold start — four minutes per instance, every scale event. Baking weights into the image cut startup to nine seconds and halved the bill, since fewer instances idled while downloading. Build-time work is startup time saved.
🎯 Key Takeaway
Bind and answer health checks first; heavy init goes to build time or background.

Preventing PORT Failures Across Every Service

Lasting prevention is a template, not a document. A shared service skeleton — Dockerfile with exec CMD, starter code that reads PORT and binds 0.0.0.0, a startup log line with the effective address, and a README with the local verify command — makes the right behavior the default for every new service. New projects inherit the contract; reviewers check deviations instead of re-teaching basics.

Automation guards the template. A CI check greps for hardcoded listen ports and localhost binds and fails the build with the fix attached. A startup-duration dashboard per service trends the probe budget margin. A staging Cloud Run service per team absorbs first-deploys from day one, so the contract gets exercised while nothing matters. Each control is cheap; together they delete the failure class.

Onboard people through the contract too. Every Cloud Run onboarding exercise should include deliberately breaking the port — hardcode 3000, deploy, read the failure, fix it — so the lesson arrives as muscle memory instead of launch-day panic. Engineers who have seen the probe fail once diagnose it in seconds forever. The cheapest incident is always the rehearsed one, never the launch-day surprise.

⚠ Fix the Code, Not the Variable
Never set Cloud Run's PORT variable to match a hardcoded app port as the permanent fix. It papers over code that will break in every other environment — staging, local Docker, the next service that copies it. Fix the code to read $PORT; the variable is the contract, not the number.
📊 Production Insight
Orgs with a shared Cloud Run template report PORT failures dropping to near zero within a quarter — the remaining cases are all non-template legacy services. The template doesn't just prevent bugs; it concentrates them where migration effort pays most.
🎯 Key Takeaway
Template plus CI checks plus staging deploys — make the contract automatic.
● Production incidentPOST-MORTEMseverity: high

The Launch-Day Deploy That Listened on the Wrong Port

Symptom
Every revision failed with container failed to start and then listen on the port, while local docker run worked flawlessly. Container logs showed the server happily started on port 3000. Cloud Run's prober knocked on 8080, heard silence, and killed each revision.
Assumption
The team assumed the container registry was serving a stale image, so they retagged and redeployed twice. Then they assumed Cloud Run was unhealthy in the region and considered moving regions — during a launch. Nobody read the revision logs for twenty minutes because the error message looked self-explanatory and pointed nowhere.
Root cause
The application called app.listen(3000) with the port hardcoded, ignoring the PORT environment variable Cloud Run injects (8080). The container started cleanly — on a port nothing probed — so the startup probe timed out waiting on 8080 and marked every revision failed.
Fix
The fix was a three-line code change: read process.env.PORT with an 8080 fallback and bind 0.0.0.0, verified locally with docker run plus host curl, then redeployed. The revision went green in four minutes. The lasting fix was a shared service template with the PORT contract built in, plus a CI check that fails builds containing hardcoded listen ports.
Key lesson
  • Read the revision logs before forming theories. The logs showed no bind on 8080 within the first minute — every alternative theory cost a redeploy cycle against evidence already available.
  • Local success proves little without the platform contract. Docker on a laptop doesn't inject PORT or probe it; reproducing with -e PORT plus host curl closes that gap for every future deploy.
  • Templates beat documentation for contracts like this. Nobody re-reads the PORT docs before each deploy, but everyone inherits the template that already implements them.
Production debug guideFive checks, in order, from revision logs to local reproduction.5 entries
Symptom · 01
The revision failed and you don't know what the container did
→
Fix
Run gcloud run services describe <SVC> --region <REGION> --format="json(env, containers)" to see the expected PORT and container config. Then run gcloud logging read 'resource.type=cloud_run_revision AND resource.labels.service_name=<SVC>' --limit=50 --format=json and search for listen lines, errors, and the bound address.
Symptom · 02
You need to separate image bugs from platform config
→
Fix
Run the image exactly as Cloud Run would: docker run -p 8080:8080 -e PORT=8080 <IMAGE> and curl http://localhost:8080 from the host. If it fails here, the bug is in the image — fix the listen call before touching any Cloud Run setting.
Symptom · 03
The app may ignore $PORT entirely
→
Fix
Grep the codebase for hardcoded ports: grep -rn 'listen(3000)\|listen(8080)\|:3000' --include='.js' --include='.py' src/ (adjust for your stack). Every hardcoded listen is a suspect — replace with the PORT env read plus an 8080 fallback and a startup log line.
Symptom · 04
Logs show started, yet the prober hears nothing
→
Fix
Check the bind address in code and logs: look for 127.0.0.1 or localhost in listen calls. From outside the container test with docker run plus host curl as above — localhost binds pass in-container curls and fail host curls, which is exactly the Cloud Run symptom.
Symptom · 05
The container needs more startup time legitimately
→
Fix
Run gcloud run services update <SVC> --region <REGION> --cpu-boost for slow starters and raise --timeout for long requests, but treat these as breathing room while you move init after listen. Re-check revision startup duration in logs after each change.
Cloud Run PORT Failure Compared
Root CauseHow to ConfirmFixPrevention
Hardcoded port ignoring $PORTCode shows listen(3000); revision logs never show a bind on 8080Read process.env.PORT or os.environ PORT with 8080 fallbackLint for hardcoded ports; log the bound port at startup
Bound to localhost onlyLogs show 127.0.0.1:8080 but the prober reports nothing listeningBind 0.0.0.0 explicitly in the listen callReview bind addresses; test with docker run -p from another host
Port bound after slow initLogs show minutes of downloads before any listen lineListen first; move heavy init to build time or backgroundAlert on startup duration; keep listen as the first action
Shell-form CMD swallowing signalsdocker stop hangs; revisions stall on shutdownUse exec-form CMD and handle SIGTERM gracefullyTest docker stop locally; standardize exec-form Dockerfiles
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
read-revision-logs.shgcloud run services describe <SERVICE> --region <REGION> --format="json(spec.tem...Reading Revision Logs
verify-port-locally.shdocker build -t portcheck . && docker run -d --name portcheck -p 8080:8080 -e PO...Fixing the Listen Call in Any Stack
audit-container-startup.shdocker inspect <IMAGE> --format='CMD={{json .Config.Cmd}} Entrypoint={{json .Con...Dockerfile Discipline
measure-startup-time.shREVISION=$(gcloud run revisions list --service <SERVICE> --region <REGION> --lim...Startup Order

Key takeaways

1
Cloud Run probes $PORT (default 8080)
the container must listen there, not on a hardcoded favorite.
2
Bind 0.0.0.0, never localhost
the prober connects from outside the container.
3
Listen first, initialize after
heavy startup work belongs at build time or in the background.
4
Use exec-form CMD so the app gets SIGTERM directly and shuts down cleanly.
5
Reproduce locally with docker run -p plus -e PORT before changing Cloud Run config.
6
Log the bound port at startup and alert on first-response latency per service.

Common mistakes to avoid

5 patterns
×

Hardcoding 8080 or 3000 instead of reading $PORT

Symptom
Works locally on 3000, dies in Cloud Run where PORT is 8080. The container listens faithfully — on a port nobody probes — and the revision fails with the container-started-but-silent signature.
Fix
Read PORT with a fallback: os.environ.get('PORT', '8080') in Python, process.env.PORT || 8080 in Node. Log the bound port at startup so the revision logs prove which value the container used. Never trust a number you can't see in the logs.
×

Listening on localhost instead of all interfaces

Symptom
Logs show the server started on 127.0.0.1:8080, yet Cloud Run reports nothing listening. Localhost inside the container is invisible to the prober outside it — started and reachable are different things.
Fix
Bind 0.0.0.0 explicitly in every framework's listen call and assert it in code review. Add a container-structure test or startup self-check that fails the build when the bind address is localhost.
×

Using shell-form CMD that swallows shutdown signals

Symptom
Revisions hang on deploy, scale-to-zero stalls, and traffic shifts drag. The shell parent eats SIGTERM, the app never shuts down cleanly, and Cloud Run eventually kills it mid-request.
Fix
Use exec-form CMD so the app is PID 1 and receives SIGTERM directly, and handle SIGTERM with a fast graceful shutdown. Test locally with docker stop and confirm the container exits within seconds, not after a kill timeout.
×

Doing minutes of init work before binding the port

Symptom
The container downloads a model for three minutes, then binds — but the prober already gave up. Startup logs show heroic preparation followed by a failed revision that never served anything.
Fix
Move model downloads, migrations, and cache warmups out of the request path: run them at build time, on a schedule, or lazily after first response. Keep the listen-first order sacred and measure startup time in the revision logs.
×

Testing one Dockerfile locally while deploying another

Symptom
Local runs use an old Dockerfile with the right port while the deployed one exposes something else. The team debugs the wrong file for an hour because two truths exist and only one ships.
Fix
Keep one canonical Dockerfile per service at the repo root, build it in CI with the same flags Cloud Run uses, and log the image digest on every deploy. Diff the local and deployed Dockerfiles the moment behavior diverges.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
Cloud Run says the container failed to start and listen on PORT. What ha...
Q02JUNIOR
Why must the container bind 0.0.0.0 instead of localhost?
Q03SENIOR
How do you debug a PORT failure before touching production config?
Q04SENIOR
Shell-form vs exec-form CMD: why does it matter on Cloud Run?
Q05SENIOR
How do you prevent PORT failures across a hundred Cloud Run services?
Q01 of 05JUNIOR

Cloud Run says the container failed to start and listen on PORT. What happened?

ANSWER
Cloud Run injects a PORT variable (default 8080) and probes the container on it; nothing answered. The usual causes are a hardcoded different port, a localhost-only bind, or slow init before listening. I'd check the revision logs for the bind line, then fix the listen call to use $PORT on 0.0.0.0.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Which port does Cloud Run expect by default?
02
Why does localhost work locally but fail on Cloud Run?
03
How do I see what port Cloud Run expects vs what my app bound?
04
Can I reproduce the failure locally with Docker?
05
Should I change Cloud Run's PORT or my code?
06
How long can container startup take on Cloud Run?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Verified
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
🔥

That's GCP. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
GCP Permission Denied on Service Account Impersonation
2 / 3 · GCP
Next
GCP Quota Exceeded: CPUS_ALL_REGIONS
→