Cloud Run PORT Failure: Container Won't Start
Fix Cloud Run PORT failure by listening on $PORT bound to 0.0.0.0: read revision logs, fix the bind, verify with Docker first..
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓A containerized app you can build with Docker
- ✓A GCP project with Cloud Run API enabled
- ✓Google Cloud SDK installed with gcloud auth login done
- Cloud Run injects $PORT (default 8080) and probes exactly that: silence there fails the revision
- Read revision logs first, then reproduce locally with docker run -p plus -e PORT
- Fix the listen call: read $PORT with an 8080 fallback and bind 0.0.0.0, never localhost
- Listen before heavy init, use exec-form CMD, and log the bound port at startup
- Use CPU boost and timeout as breathing room while slow init moves after listen
Think of Cloud Run as a hotel that tells every guest their room number on a card (the PORT variable, usually 8080) and then knocks on that door to check you arrived. PORT failure means you went to a different room (hardcoded 3000), hid under the bed where knocks can't reach (localhost bind), or spent an hour unpacking before opening the door (slow init). The hotel isn't broken — you missed the instructions on the card. Read the card, open that door promptly, and answer when it knocks.
You containerize the app, push the image, deploy to Cloud Run, and get the dreaded revision error: container failed to start and then listen on the port. The image runs perfectly on your laptop. The Dockerfile looks right. Yet Cloud Run insists nothing is listening, kills the revision, and leaves you staring at a port number you never chose.
That port number is the whole story. Cloud Run injects a PORT environment variable (8080 by default) and probes your container on exactly that port. If your app listens on a hardcoded 3000, binds only to localhost, or spends three minutes downloading a model before listening, the probe finds silence and the revision fails. Your laptop never enforces this contract, so the bug ships undetected every time.
Beginners hit this on their first Cloud Run deploy; veterans hit it when a framework default changes or a second Dockerfile drifts. Either way the fix is small and the lesson permanent.
This guide walks the diagnosis in order: read the revision logs, compare expected versus actual port, fix the listen call, correct the Dockerfile, and verify locally with the same contract. You'll see a real launch-day incident caused by a hardcoded port, plus the template that makes this failure nearly impossible.
The PORT Contract: What Cloud Run Actually Probes
Cloud Run's contract with your container is short: listen on the port in the PORT environment variable, on all interfaces, promptly after starting. Cloud Run sets PORT to 8080 unless you configure otherwise, then probes that port to decide whether the revision is healthy. Answer the probe and traffic flows; miss it and the revision fails with the container-failed-to-listen error. There is no negotiation, no fallback scan of other ports, and no credit for listening beautifully on the wrong number.
Three violations cover nearly every failure. A hardcoded port (3000 in Node tutorials, 5000 in Flask defaults, 8000 in Django habits) ignores PORT entirely. A localhost bind answers only inside the container while the prober knocks from outside. Slow initialization — model downloads, migrations, cache warmups — delays the bind past the probe window. All three produce the identical revision error, which is why the message alone never identifies the cause.
Your laptop enforces none of this. Docker run without -e PORT leaves your hardcoded default working; curl from inside the container reaches localhost binds; no prober times your startup. The gap between localdocker and Cloud Run is exactly the set of behaviors this error reports. Close the gap deliberately — reproduce with the platform's terms — and the error becomes a pre-deploy catch instead of a launch-day surprise.
Reading Revision Logs: What the Container Really Did
Revision logs are the ground truth because they show what the container actually did, not what the Dockerfile intended. Cloud Run streams container stdout and stderr per revision, including framework startup banners that usually print the bound address. A line like listening on 0.0.0.0:8080 is the all-clear; listening on 127.0.0.1:3000 names both bugs at once; no listen line at all means the app crashed or is still initializing.
Read them with intent. Filter to the failed revision and scan from the container start marker: dependency installs, framework banner, listen line, then probe results. The time gap between start and listen is your startup duration — if it spans minutes, initialization order is the problem regardless of the port. Errors above the listen line (missing modules, failed migrations) are startup crashes wearing a PORT costume; fix them first and the port complaint often vanishes.
Make the logs work for you permanently by logging the effective port and address at startup in every service. One line — serving on 0.0.0.0:8080 — turns every future PORT incident into a ten-second diagnosis. The snippet below pulls the right logs for the failed revision so you stop guessing from the error message alone.
Fixing the Listen Call in Any Stack
The listen call is where the fix lands, and every stack expresses it slightly differently. Node's app.listen(process.env.PORT || 8080, '0.0.0.0'), Python's app.run(host='0.0.0.0', port=int(os.environ.get('PORT', 8080))), Go's net.Listen on ':' + os.Getenv('PORT') — same contract, different syntax. The two non-negotiables: the port comes from the environment with an 8080 fallback for local runs, and the host is the all-interfaces address, never localhost.
Framework defaults actively work against you here. Express tutorials hardcode 3000, Flask defaults to 127.0.0.1:5000, and several frameworks bind localhost when no host is given. Scaffolding that runs beautifully on a laptop ships all three violations at once. Treat every framework default as guilty until proven bound to $PORT on 0.0.0.0 — a one-line review that prevents the most common Cloud Run failure in existence.
Verify the fix where Cloud Run can't argue with it: a local container run with the platform's terms. Inject PORT, publish the port, and curl from the host (outside the container). Host-reachable means prober-reachable; in-container-only success means nothing. Keep this exact command in the service README so every contributor rehearses the contract, not just the deployer.
Dockerfile Discipline: Exec CMD and Fast Starts
The Dockerfile decides process behavior that code can't override. Exec-form CMD (JSON array: CMD ["node", "server.js"]) runs your app as PID 1, delivering SIGTERM straight to it for graceful shutdown. Shell-form CMD (CMD node server.js) wraps your app in /bin/sh, which swallows SIGTERM — deploys hang, scale-to-zero stalls, and revisions linger in half-dead states. Cloud Run sends SIGTERM on traffic shifts and instance shutdowns constantly, so signal handling isn't edge-case hygiene, it's the normal path.
Port exposure deserves a matching review. EXPOSE documents intent but enforces nothing — Cloud Run probes PORT regardless of EXPOSE declarations. Still, keep EXPOSE aligned with your default port for local tooling that does honor it, and never rely on it as the fix. The listen call is the fix; the Dockerfile is the delivery mechanism.
Order inside the Dockerfile matters for startup speed too. Dependency installation belongs in early cached layers; the listen-first application code goes last. Multi-stage builds keep production images lean, which shortens cold starts and narrows the probe window pressure. Review Dockerfiles with the same seriousness as the listen call — the container that starts fast, binds right, and shuts down cleanly is built, not wished for.
Startup Order: Listen First, Initialize After
Startup order bugs hide behind correct ports. The container binds 8080 properly — after downloading a 2 GB model, running migrations, and warming caches for four minutes. The prober's patience expires first, and a perfectly addressed server still fails the revision. The fix is sequencing: bind the port and answer health checks first, then complete heavy initialization in the background or ahead of time at build.
Several patterns keep listen-first honest. Bake models and assets into the image at build time so startup reads local disk instead of the network. Run migrations as a separate Cloud Run job or Cloud Build step rather than inside the serving container's boot. Warm caches lazily after first response, accepting a slower first request instead of risking the whole revision. Each pattern trades a little runtime elegance for startup certainty — the right trade on a probed platform.
Measure what you sequence. Startup duration (container start to first listen) belongs on a dashboard per service, with alerts when it approaches the probe budget. Teams that watch this metric catch dependency bloat — the SDK that added forty seconds, the model that doubled — while it's still a graph trend, not a failed deploy. Fast startup is a feature you monitor, not luck you hope for.
Preventing PORT Failures Across Every Service
Lasting prevention is a template, not a document. A shared service skeleton — Dockerfile with exec CMD, starter code that reads PORT and binds 0.0.0.0, a startup log line with the effective address, and a README with the local verify command — makes the right behavior the default for every new service. New projects inherit the contract; reviewers check deviations instead of re-teaching basics.
Automation guards the template. A CI check greps for hardcoded listen ports and localhost binds and fails the build with the fix attached. A startup-duration dashboard per service trends the probe budget margin. A staging Cloud Run service per team absorbs first-deploys from day one, so the contract gets exercised while nothing matters. Each control is cheap; together they delete the failure class.
Onboard people through the contract too. Every Cloud Run onboarding exercise should include deliberately breaking the port — hardcode 3000, deploy, read the failure, fix it — so the lesson arrives as muscle memory instead of launch-day panic. Engineers who have seen the probe fail once diagnose it in seconds forever. The cheapest incident is always the rehearsed one, never the launch-day surprise.
The Launch-Day Deploy That Listened on the Wrong Port
- Read the revision logs before forming theories. The logs showed no bind on 8080 within the first minute — every alternative theory cost a redeploy cycle against evidence already available.
- Local success proves little without the platform contract. Docker on a laptop doesn't inject PORT or probe it; reproducing with -e PORT plus host curl closes that gap for every future deploy.
- Templates beat documentation for contracts like this. Nobody re-reads the PORT docs before each deploy, but everyone inherits the template that already implements them.
gcloud run services describe <SVC> --region <REGION> --format="json(env, containers)" to see the expected PORT and container config. Then run gcloud logging read 'resource.type=cloud_run_revision AND resource.labels.service_name=<SVC>' --limit=50 --format=json and search for listen lines, errors, and the bound address.docker run -p 8080:8080 -e PORT=8080 <IMAGE> and curl http://localhost:8080 from the host. If it fails here, the bug is in the image — fix the listen call before touching any Cloud Run setting.grep -rn 'listen(3000)\|listen(8080)\|:3000' --include='.js' --include='.py' src/ (adjust for your stack). Every hardcoded listen is a suspect — replace with the PORT env read plus an 8080 fallback and a startup log line.docker run plus host curl as above — localhost binds pass in-container curls and fail host curls, which is exactly the Cloud Run symptom.gcloud run services update <SVC> --region <REGION> --cpu-boost for slow starters and raise --timeout for long requests, but treat these as breathing room while you move init after listen. Re-check revision startup duration in logs after each change.| File | Command / Code | Purpose |
|---|---|---|
| read-revision-logs.sh | gcloud run services describe <SERVICE> --region <REGION> --format="json(spec.tem... | Reading Revision Logs |
| verify-port-locally.sh | docker build -t portcheck . && docker run -d --name portcheck -p 8080:8080 -e PO... | Fixing the Listen Call in Any Stack |
| audit-container-startup.sh | docker inspect <IMAGE> --format='CMD={{json .Config.Cmd}} Entrypoint={{json .Con... | Dockerfile Discipline |
| measure-startup-time.sh | REVISION=$(gcloud run revisions list --service <SERVICE> --region <REGION> --lim... | Startup Order |
Key takeaways
Common mistakes to avoid
5 patternsHardcoding 8080 or 3000 instead of reading $PORT
os.environ.get('PORT', '8080') in Python, process.env.PORT || 8080 in Node. Log the bound port at startup so the revision logs prove which value the container used. Never trust a number you can't see in the logs.Listening on localhost instead of all interfaces
Using shell-form CMD that swallows shutdown signals
Doing minutes of init work before binding the port
Testing one Dockerfile locally while deploying another
Interview Questions on This Topic
Cloud Run says the container failed to start and listen on PORT. What happened?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's GCP. Mark it forged?
5 min read · try the examples if you haven't