Home› Cloud› Complete Guide
Complete Guide

Complete Cloud Tutorial

This complete guide covers all 12 Cloud tutorials on TheCodeForge, organised by topic.

Learning Roadmap
Beginner → Build a strong foundation
Intermediate → Deepen your understanding with practical topics
Advanced → Master advanced concepts and real-world applications
12
Topics
5
Beginner
7
Intermediate
0
Advanced
Jump to section
Terraform (5)Azure (4)GCP (3)

Cloud errors are rarely about the cloud. They are about three things that happen to live there: state, identity, and quota. Terraform's state lock is state. Azure's AADSTS700016 is identity. GCP's CPUS_ALL_REGIONS is quota. Recognise which of the three you are looking at and the fix is usually obvious; miss it and you can spend an afternoon reading the wrong documentation.

The hardest of the three is identity, because cloud providers have converged on the same layered model — a principal, a role, a scope, a resource policy — and a request fails if any layer says no while the error names only one. The productive habit is to stop reading the denial and start asking the provider to explain the decision, which all three major clouds will now do.

Terraform state is the product; the config is just input

Most Terraform pain comes from treating the .tf files as the source of truth. They are not — the state file is. State maps your resource addresses to real cloud object IDs, and every confusing Terraform error is state disagreeing with either your config or reality.

Error acquiring the state lock means another operation holds the lock, or a previous one died without releasing it. Resource already exists means the object is real but absent from state, so Terraform plans a create that the provider rejects. The answer is almost never to delete the object.

ErrorState saysAction
Error acquiring the state lockSomeone else is mid-apply, or a crashed run left the lockWait, then verify the run is really gone before force-unlock with the reported lock ID. Never force-unlock a run that may still be writing
Resource already existsNothing — the object is real but unmanagedterraform import, or an import block, to adopt it rather than recreate it
Provider configuration not presentA resource still exists in state whose provider block was removed or renamed during a refactorRe-add the provider long enough to destroy or move the resource, or state rm and re-import under the new address
Cycle in dependenciesTwo resources reference each other, directly or through a module outputBreak the loop with a separate association resource, or pass a known value instead of a computed attribute
In practicecount addresses resources by index, so removing the second item of five renumbers everything after it and Terraform plans three destroys and three creates. for_each addresses them by a stable key, so removing one item touches exactly one resource. For anything with identity — users, DNS records, buckets — for_each is the safe default.

Identity errors: stop reading the denial, ask for the decision

AADSTS700016 — application not found in the directory — usually means the client ID is correct but for a different tenant, or the app registration exists while the service principal in this tenant does not. GCP's impersonation failures usually mean the caller lacks roles/iam.serviceAccountTokenCreator on the target service account, which is a permission on the account itself, not on the project.

Both clouds provide a way to ask why, and it is far quicker than inference.

bash
# Azure: does this app registration exist in THIS tenant?
az account show --query tenantId -o tsv
az ad app show --id <client-id> -o table            # registration
az ad sp show  --id <client-id> -o table            # service principal in tenant

# GCP: who may impersonate this service account, and what can it do?
gcloud iam service-accounts get-iam-policy target@proj.iam.gserviceaccount.com
gcloud projects get-iam-policy my-project \
  --flatten="bindings[].members" \
  --filter="bindings.members:caller@proj.iam.gserviceaccount.com"

# Prove the impersonation independently of your application
gcloud auth print-access-token \
  --impersonate-service-account=target@proj.iam.gserviceaccount.com

Two habits save the most time here. First, always confirm which tenant, subscription or project the credential actually belongs to before debugging the permission — a surprising share of identity errors are a correct credential pointed at the wrong place. Second, test the identity in isolation, outside your application, so you learn whether the problem is the grant or your code's use of it.

Startup contracts: the platform is waiting for something specific

Cloud Run's container failed to start and listen on the port defined by the PORT environment variable and Azure App Service's 502.5 process failure are the same class of error: the platform started your process, waited for it to satisfy a contract, and gave up. Cloud Run's contract is explicit — bind 0.0.0.0 on $PORT within the startup window. Hard-coding 8080 works until the platform picks something else; binding 127.0.0.1 fails always, because the health probe arrives from outside the container.

Azure Functions' 230-second ceiling belongs to the same family. It is not a timeout you can raise on the consumption plan — it is the load balancer's limit on a synchronous HTTP response. Work that genuinely takes longer needs an asynchronous shape: accept the request, return an operation id, do the work in a durable function or a queue-triggered worker, and let the client poll.

python
import os

# Wrong on Cloud Run: fixed port, loopback interface
# app.run(host="127.0.0.1", port=8080)

# Right: honour $PORT, bind all interfaces so the health probe can reach you
port = int(os.environ.get("PORT", 8080))
app.run(host="0.0.0.0", port=port)

# And do slow setup lazily - the startup probe has a deadline
_model = None
def get_model():
    global _model
    if _model is None:
        _model = load_model()      # after the port is listening, not before
    return _model

Quota is a design constraint, not an obstacle

Quota exceeded: CPUS_ALL_REGIONS surprises people because it is a global limit that no single region's dashboard shows. It is worth treating quota as part of the architecture rather than as paperwork: every cloud enforces limits per region, per project, per account and sometimes globally, and an autoscaling group that cannot obtain instances fails in a way that looks like an application outage.

The engineering response is to know your ceilings before you need them, request increases ahead of a launch rather than during one, and make scaling behaviour degrade visibly when the ceiling is hit — an alert on quota utilisation is worth more than an alert on the failure it eventually causes.

Frequently Asked Questions

Is it ever safe to run terraform force-unlock?
Only after you have confirmed the holding operation is genuinely dead — a crashed CI job, a killed local run — and you are certain no apply is still writing. Force-unlocking a live apply risks two processes mutating the same state, which is how state files get corrupted. Use the lock ID from the error rather than unlocking blind, and prefer waiting whenever waiting is possible.
count or for_each?
for_each for anything with identity, which is most things. count is fine for genuinely interchangeable, homogeneous resources and for a simple zero-or-one conditional. The deciding question is what happens when you remove a middle element: with count everything after it is renumbered and replaced; with for_each only that one resource is destroyed.
Why does my Cloud Run container work locally but fail to start?
Because locally nothing enforces the startup contract. The three usual causes are binding to 127.0.0.1 instead of 0.0.0.0, ignoring $PORT in favour of a hard-coded number, and doing slow work — loading a model, warming a cache, running migrations — before the listener opens, so the startup probe times out. Move that work behind the listener or into a startup probe with a longer deadline.
Can I raise the Azure Functions 230-second limit?
Not for a synchronous HTTP response on the consumption plan — the ceiling belongs to the load balancer in front of the function, not to the function host. Restructure instead: return 202 with an operation id and poll, use Durable Functions for orchestration, or move the work to a queue-triggered function where the execution timeout is configurable.
Which cloud should I learn if I already know AWS?
Azure if you are aiming at enterprise and Microsoft-heavy organisations, GCP if you are aiming at data and ML platforms. The transferable part is not the service names but the three models in this track — identity layering, quota, and declarative state — and those are close enough across providers that the second cloud takes a fraction of the time the first did.
Should Terraform state live in the cloud or in git?
In the cloud, in a remote backend with locking and versioning — never in git. State contains resource IDs and often secrets in plain text, and git provides no locking, so two concurrent applies will produce a conflict after the damage rather than preventing it. A versioned object store with a lock table is the baseline.

Terraform

Azure

GCP

Also Explore
DevOps 304 tutorials → System Design 145 tutorials → Security 16 tutorials → C# / .NET 60 tutorials → Observability 10 tutorials → Data Engineering 13 tutorials →
Start from the beginning

Every tutorial starts with a plain-English analogy — then real code, then interview questions.

Browse Cloud Tutorials →