Azure Function 230s Timeout: Beat the Limit
Beat the Azure Function 230-second limit with async patterns: return 202 fast, fan out with Durable Functions, and poll status..
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
- ✓An Azure Function app with an HTTP trigger deployed
- ✓Azure CLI installed with az login completed
- ✓Basic familiarity with Application Insights queries
- Azure caps every HTTP-triggered response at 230 seconds on all plans, whatever functionTimeout says
- Prove it with App Insights durations clustering near 230s while clients report 502s
- Fix it async: return 202 immediately, finish via Durable orchestration or queue, let clients poll status
- Fan-out/fan-in parallelizes batch work, and plan choice (Consumption vs Premium) covers background execution
Think of the 230-second limit as a restaurant rule: the waiter must bring something to your table within four minutes or the order is cancelled — even if the kitchen is still cooking. Your function is the kitchen; the HTTP response is the waiter. The fix isn't a faster kitchen, it's a different system: the waiter immediately brings a buzzer (HTTP 202 plus a status URL) and the kitchen delivers the meal when ready.
Your report endpoint works beautifully in testing. Then the quarterly report runs against real data, the browser spins for nearly four minutes, and dies with a 502. The function logs show the work nearly finished. Rerun it and it dies at the same wall: 230 seconds. Nothing in your code mentions that number, because it isn't your number — it's Azure's.
Every HTTP-triggered Azure Function must respond within 230 seconds, no matter the hosting plan or the timeout setting. The front-end load balancer enforces it uniformly. Your function can legally run longer in the background, but the HTTP response holding the client's connection gets cut. Local testing never shows this because localhost has no load balancer.
Teams usually discover the limit with their most important endpoint: the big export, the batch onboard, the end-of-month calculation. The failure arrives as a 502 with no function-side error, which sends debugging toward networking instead of architecture.
This guide explains the two limits that interact here, shows the async patterns that dissolve them — Durable orchestrations, fan-out/fan-in, queue offloading — and walks a real incident where a report endpoint died every quarter-end. You'll leave with endpoints that answer in milliseconds and work that finishes on its own schedule.
The Two Clocks: functionTimeout vs the 230-Second Wall
Two different clocks govern an HTTP-triggered function, and confusing them causes this entire incident class. The first clock is functionTimeout in host.json: the maximum runtime of one execution. It defaults to 5 minutes on Consumption (raisable to 10) and 30 minutes on Premium and Dedicated (raisable to unbounded). The second clock is the front-end load balancer's 230-second cap on an open HTTP response, which applies on every plan with no setting to raise it. Your execution may legally run nine minutes while its HTTP response dies at three minutes fifty.
Local development hides both clocks. The Functions Core Tools runtime applies generous local defaults and, crucially, there is no load balancer between your curl and your function. Code that responds in eight minutes on localhost fails in Azure at 230 seconds with zero code changes. This is why the bug always debuts in staging or production, never on a laptop.
The diagnostic signature is distinctive: client-side 502s or timeouts at ~230 seconds paired with function executions that continue past the client failure in Application Insights. If the function itself errored, you'd see exceptions; here the function looks healthy and the client looks broken. That split — healthy execution, dead response — is the fingerprint of the front-end wall. Confirm it with duration percentiles before redesigning anything.
The Async HTTP Pattern: 202, Orchestrate, Poll
The async HTTP API pattern is the canonical escape: never hold an HTTP response open across long work. An HTTP starter function receives the request, starts a Durable orchestration, and returns 202 Accepted with management URLs — including statusQueryGetUri — in under a second. The orchestrator runs activities that do the real work over minutes. The client polls the status URL with backoff until the orchestration reports completion, then fetches the result.
This inverts the timeout math completely. No execution holds a client connection, so neither the 230-second wall nor client-side proxy timeouts matter. The orchestration itself can run for days, surviving function restarts and replays, with each activity independently retried on failure. Corporate proxies that kill connections at 60 seconds become irrelevant because every response completes in milliseconds.
Adopting it means changing the API contract, which is the real work. Clients must handle 202, poll with exponential backoff, and render progress states. Document the status schema — runtimeStatus, output, customStatus for progress percentages — so every consumer polls the same way. The snippet below shows the starter contract and the status poll loop you can rehearse from any shell before writing client code.
Fan-Out/Fan-In: Turning Linear Time Into Parallel Time
Fan-out/fan-in attacks the duration itself instead of just hiding it. The orchestrator takes a batch — regions, files, accounts — and starts one activity function per item. All activities run in parallel across the plan's instances. The orchestrator waits for every activity (fan-in), aggregates the outputs, and completes. Eight minutes of sequential work becomes one minute of parallel work, and each activity enjoys its own full execution budget.
Sizing matters. Activities should be coarse enough that orchestration overhead stays trivial but fine enough that no single activity nears a timeout — minutes each, not seconds and not tens of minutes. Sub-orchestrators split enormous batches hierarchically: one parent fans out to ten children, each fanning out to a hundred activities. Durable's replay mechanics make this deterministic, so retries and restarts resume cleanly instead of duplicating side effects.
Mind the determinism rules that make replay safe: orchestrators must not do I/O, sleep, or random numbers directly — those belong in activities. Violating this corrupts replay history in ways that surface as bizarre stuck orchestrations weeks later. Keep orchestrators pure coordination: start activities, wait, aggregate. Everything with a side effect lives in an activity with its own retry policy.
Picking the Plan: Consumption vs Premium vs Dedicated
Choosing the plan is the second half of the fix, governing background execution rather than HTTP responses. Consumption fits short, spiky HTTP work: five-minute default timeout, scale-to-zero, pay-per-execution. Premium fits longer orchestration activities and steady throughput: thirty-minute default timeout, pre-warmed instances, VNet integration. Dedicated (App Service plan) fits workloads that need the full thirty minutes plus cohabitation with web apps. None of them moves the 230-second HTTP wall — plan choice buys execution budget, not response time.
Verify the current reality before recommending a move. The plan SKU, the deployed functionTimeout, and the measured duration percentiles together dictate the answer. A function peaking at eight minutes needs Premium or Dedicated (or chunking); a function peaking at four minutes against an HTTP client needs the async pattern on any plan. Upgrading the plan for an HTTP-wall problem wastes money with surgical precision.
Cost follows the same logic. Consumption looks cheapest until timeouts force rewrites under incident pressure. Premium's always-ready instances cost more hourly but absorb orchestration fan-out without cold-start cliffs. Model the decision on measured P95 durations plus growth headroom, revisit quarterly, and remember the cheapest plan is the one whose limits your architecture respects.
Queue Offloading: The Lighter Async Alternative
Queue offloading is the lighter alternative when full Durable orchestration feels heavy. The HTTP function validates the request, drops a message on a queue (or Service Bus, or Event Hubs), and returns 202 with a job ID. A queue-triggered function picks up the message and does the minutes-long work without any HTTP response to hold open. Status lives in table storage or Cosmos DB, which the client polls. Fewer moving parts than orchestration, same immunity to the 230-second wall.
This pattern also absorbs load spikes gracefully. A thousand simultaneous requests become a thousand queue messages processed at the plan's steady throughput instead of a thousand concurrent executions racing the wall. Poison-message handling and dead-letter queues convert permanent failures into inspectable records rather than lost requests. For workloads that are naturally message-shaped — file processing, notifications, ETL rows — queues are the idiomatic answer.
Choose between queues and orchestration by coordination needs. Independent items with no aggregation step belong on queues. Workflows needing fan-in aggregation, human-approval pauses, or durable timers belong in orchestrations. Many systems use both: queues ingest, orchestrations coordinate. Either way the HTTP layer stays thin, fast, and far from every timeout.
Keeping Endpoints Fast Forever: Contracts and Dashboards
Prevention is a contract plus a dashboard. The contract: every HTTP endpoint documents its response-time budget, returns 202 for anything over thirty seconds of work, and exposes a status URL with a stable schema. New endpoints get reviewed against this contract like any API guideline — sync responses for minutes-long work fail review before they fail in production. Client libraries ship polling helpers so consumers default to the async path.
The dashboard tracks the leading indicators. Max and P95 execution duration per function, plotted against the 200-second warning line. Client-side 502 rates per HTTP endpoint, which spike the day the wall starts biting. Queue depths and orchestration runtimes for the async paths, proving the escape valves have headroom. One screen, reviewed weekly, replaces every surprise.
Close the loop with load tests that use production-size inputs on a schedule, not just at launch. Data grows; today's two-minute report is next year's four-minute outage. A quarterly load test with current data volumes catches the march toward the wall while there's still room to refactor. Timeout resilience isn't a one-time migration — it's a budget you defend against every release that adds work to an endpoint.
The Quarterly Report That Always Died at Four Minutes
- Measure duration percentiles before choosing timeouts or plans. The team tuned everything except the actual constraint because nobody had charted execution time against the 230-second wall.
- Synchronous HTTP is the wrong contract for minutes-long work. A 202-plus-polling API absorbs arbitrary durations, retries, and client timeouts by design instead of by luck.
- Load-test with production-size payloads on a schedule. The endpoint passed every test with sample data and failed only on real quarterly volume — the one input nobody rehearsed.
az functionapp show --name <APP> --resource-group <RG> --query "{plan:serverFarmId, kind:kind}" -o tsv and az functionapp plan show --name <PLAN> --resource-group <RG> --query "{sku:sku.name, tier:sku.tier}" -o table. A Consumption (Y1/Dynamic) plan means a 5-minute default execution ceiling on top of the 230-second HTTP cap.https://<APP>.scm.azurewebsites.net/api/vfs/site/wwwroot/host.json) and read functionTimeout. Then run az functionapp config appsettings list --name <APP> --resource-group <RG> --query "[?name=='FUNCTIONS_EXTENSIONBUNDLE_VERSION']" to confirm the runtime. A missing functionTimeout means defaults apply silently.az monitor app-insights query --app <AI> --analytics-query "requests | where timestamp > ago(7d) | summarize max(duration), percentile(duration, 95) by name | order by max_duration desc". Functions whose max parks near 230,000 ms are hitting the front-end wall.az monitor metrics list --resource <APP_ID> --metric HttpResponseTime --interval PT1H and correlate spikes with 502 counts. Client-side 502s with healthy function executions confirm the front end cut the response while work continued — the signature of this exact limit.| File | Command / Code | Purpose |
|---|---|---|
| durable-async-http.sh | curl -i -X POST "https://<APP>.azurewebsites.net/api/orchestrators/ReportOrchest... | The Async HTTP Pattern |
| inspect-orchestrations.sh | curl -s "$STATUS_URL" | python3 -m json.tool | Fan-Out/Fan-In |
| audit-function-plan.sh | az functionapp show --name <APP> --resource-group <RG> --query "{plan:serverFarm... | Picking the Plan |
| verify-queue-offload.sh | az storage message peek --queue-name <QUEUE> --account-name <STORAGE> --num-mess... | Queue Offloading |
Key takeaways
Common mistakes to avoid
5 patternsTreating an HTTP trigger like a batch job that may run for minutes
az functionapp show and az functionapp plan show, then design every HTTP handler to finish in seconds, not minutes.Assuming the default timeout covers the workload on every plan
az functionapp config appsettings list plus the deployed host.json, and assert functionTimeout explicitly. Better: stop depending on the timeout entirely by going async — timeouts are guardrails, not schedules.Looping over hundreds of items inside one HTTP-triggered execution
Making the client block on the HTTP response for the whole job
Shipping long functions without duration monitoring
Interview Questions on This Topic
Your HTTP-triggered function dies at 230 seconds. What happened?
Frequently Asked Questions
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
That's Azure. Mark it forged?
5 min read · try the examples if you haven't