Airflow Task Stuck in Queued Forever — Fix It
Airflow tasks queued for hours means slots, scheduler, or stuck state.
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓An Airflow deployment with Grid view access
- ✓Pools, executors, and scheduler basics
- ✓CLI access to pools list and task logs
- Queued means ready but homeless: no worker slot, pool slot, or concurrency limit available — not broken code
- Check in order: scheduler heartbeat, pool occupancy, parallelism, then greedy-neighbor DAGs hogging slots
- One uncapped backfill on shared default_pool can starve every other DAG for hours past your SLA
- Fix structurally (private pools, max_active_tasks caps, deferrable waits); clear only truly stuck instances
Picture an airport where planes (tasks) wait at the gate for a runway slot. Grey queued planes aren't broken — they're homeless: the runway (pool slots) is full, the control tower (scheduler) went home sick, or one jumbo jet (greedy DAG) booked every slot. Clearing the departure board doesn't build runways. You count free slots, wake the tower, park the jumbo in its own lane, and only then rebook the flights that truly missed their window.
Friday 6:40 AM, and the 7 AM SLA is already dead. The Grid view shows a sea of grey: 40 tasks queued, some for 4 hours, workers half-idle behind failed upstreams. Somebody suggests restarting everything. Somebody else already cleared the whole DAG twice, and the grey came back both times.
Queued-forever is Airflow's most misdiagnosed state because grey squares all look alike while their causes differ completely: pool slots exhausted by a greedy neighbor, executor parallelism maxed out, a scheduler that died Tuesday and nobody noticed, or a single task instance wedged in stale metadata. Each needs a different fix, and the wrong one (bulk-clearing a healthy queue behind a dead scheduler) burns hours.
This guide uses the lifecycle vocabulary from the task-states primer — scheduled, queued, running, deferred, and friends — and turns grey squares into specific diagnoses. You'll check pools, parallelism, scheduler heartbeats, and stuck state in order, with the exact commands at each step. By the end, queued age will be your favorite leading indicator instead of your Friday surprise.
What Queued Forever Actually Means
In the lifecycle vocabulary, a task instance moves scheduled to queued to running to a terminal state (success, failed, skipped). Scheduled means dependencies aren't met yet — upstream tasks unfinished, trigger rules unsatisfied. Queued means the task is ready and waiting for a slot: a worker willing to run it, a pool with a free place, and no concurrency ceiling blocking it. That distinction decides your whole response: scheduled-forever is an upstream problem, queued-forever is a capacity or scheduler problem.
Grey squares deceive because every flavor of homelessness looks identical in Grid view. A task queued behind an exhausted pool, one queued behind a dead scheduler, and one wedged in stale metadata all render the same grey. Teams that treat grey as one problem apply one fix — usually bulk-clearing — and watch the grey return, because clearing re-queues tasks into the same slotless world that stranded them.
Build the habit of naming the landlord before evicting anyone. Pool occupancy (pools list), scheduler liveness (heartbeat ticks), executor pressure (parallelism vs running counts), and per-DAG greed (max_active_tasks unbounded) each answer in seconds. Queued age — how long the oldest task has waited — is the metric that turns this from archaeology into alerting: anything past 10 minutes deserves a page, not a shrug.
Pool Slots Exhausted: the Shared-Bucket Starvation
Pools are named slot buckets tasks must pass through — every task needs one pool slot plus a worker to run. The default_pool ships with 128 slots shared by every DAG that doesn't declare otherwise, which works until one backfill fans out 200 tasks and holds 121 slots for 6 hours. Every polite DAG behind it queues, workers idle for lack of admittable tasks, and the dashboard shows a healthy cluster doing nothing.
Diagnose with pools list (occupancy per pool) crossed against states-for-dag-run per DAG (who waits). The signature is unmistakable: one pool at 100% used, one DAG holding most slots, everyone else grey. The structural fix has two halves — raise slots deliberately where the ceiling is genuinely low, and isolate where one tenant misbehaves: private pools for critical DAGs (etl_pool=32 for warehouse_sync) so backfills can never eat SLA traffic again.
Treat pool design as capacity planning, not firefighting. Review occupancy weekly as DAG counts grow; defaults that fit 20 DAGs strangle 200. Set pool_slots: 1 on ordinary tasks and reserve higher values for genuinely heavy ones, because a task holding 4 slots runs at 4x the queue cost. Queues drain when admission is managed — not when workers multiply into a full pool.
Executor Pressure and max_active_tasks: Capping the Greedy
Beyond pools, two scheduler-wide ceilings shape every queue: parallelism (total task instances the scheduler admits cluster-wide) and max_active_tasks or max_active_runs per DAG. Parallelism at 32 with 200 runnable tasks means 168 wait no matter how many workers you own. An uncapped hourly DAG fanning out 200 tasks per run fills that ceiling alone — polite DAGs queue behind a neighbor they'd never suspect from their own logs.
The polite-DAG pattern above shows the full stack in one file: max_active_tasks=8 caps the DAG's footprint, pool etl_pool isolates its admission, pool_slots=1 keeps per-task cost minimal, and execution_timeout bounds runaways so a hung task can't squat a slot forever. Retries with backoff ride along so transient blips heal instead of failing into manual clears. Copy this header onto every new DAG and the greedy-neighbor class shrinks to legacy DAGs you migrate on sight.
Defaults age badly — re-tune quarterly as DAG counts grow. What fit 20 DAGs (parallelism 32, everything on default_pool) strangles 200. Capacity planning here is scheduling work, and the review takes an hour: list every DAG without max_active_tasks, cap them, and watch queued age fall across the fleet the same day.
Zombie Scheduler: the Tower Went Home Sick
The scheduler is the control tower: it moves tasks from scheduled to queued, admits them to slots, and reaps the dead. When it dies — OOMKill, bad deploy, disk-full on its log volume — everything freezes in place. New runs never leave scheduled, queued tasks never get adopted, and every task-level fix fails because no living process reads the queue. One real scheduler flatlined Tuesday; the team debugged tasks until Friday's SLA breach.
Check liveness before anything else: jobs check for the SchedulerJob plus fresh heartbeat lines in scheduler logs (seconds old, not days). Flat heartbeat means restart or fail over to standby in HA setups — and in HA, verify only one scheduler leads, since dual-active schedulers double-run tasks into duplicate writes. After recovery, confirm with list-runs and states-for-dag-run that fresh tasks flow scheduled to queued to running within a minute.
Zombie schedulers deserve special suspicion after K8s churn: liveness-probe restarts and node scale-downs leave scheduler processes half-alive — heartbeating just enough to block failover, dead enough to schedule nothing. Alert on heartbeat age directly (stale past 60s pages), not on queued counts that lag the cause by hours. The tower must be awake before any runway discussion matters.
Priority, Pools, and Placement: Designing Who Waits
Priority and pool design decide who waits when slots run short. priority_weight orders admission inside a pool's queue — interactive pipelines at 5+ jump ahead of backfills at 3, so stale history can't block fresh SLA traffic. Separate pools make the ordering structural: backfill_pool with modest slots bounds history replay regardless of weights, while etl_pool stays reserved for daily promises. Heavy tasks declare pool_slots=2 so the scheduler accounts their real cost instead of admitting four heavies into four 1-slot places.
max_active_runs=1 on backfill DAGs stops the stampede at the source: one run at a time means catchup processes serially instead of fanning 90 runs at once. Combined with catchup=False on interactive DAGs, the fleet stops competing with its own history. These are design decisions, not incident responses — set them when the DAG is born, because retrofitting priority during an outage means editing the DAGs you most fear touching.
Review the matrix quarterly: list every DAG's pool, weight, max_active_tasks, and max_active_runs in one table. Anything on default_pool with weight 1 and no caps is a future incident wearing a trench coat. Thirty minutes of review beats another Friday of grey squares.
Clearing Stuck State Without Replaying the World
Sometimes one instance is genuinely wedged: metadata stale after a kill, a zombie the heartbeat already reaped, a failure whose cause is fixed but whose state never moved. Clearing returns it to the queue with try_number bumped and per-try history intact — logs stay inspectable, lineage survives. The ritual matters: read the log tail first (the fatal line names the cause), fix the cause, clear just that instance, then verify it flows to running within a minute.
Restraint is the whole skill. Bulk-clearing a DAG run replays every downstream write — duplicate warehouse loads, double-fired webhooks, re-sent emails — while the original failure recurs unfixed. The incident that motivates this section featured two full-DAG clears that re-queued healthy tasks into the same slotless pool and duplicated a day's loads. Scope every clear to the stuck instance (-t stuck_task), and snapshot list-runs output before any bulk operation so you can account for what replayed.
Close the loop with prevention: execution_timeout on long tasks bounds squatters, queued-age alerts catch the next stall at minute 10 instead of hour 4, and deferrable operators move hour-long waits to the triggerer where they cost nothing. Clears should be rare, surgical, and followed by a config change that makes that clear unnecessary next time.
40 Tasks Queued 4 Hours While Doubled Workers Idled
- Queued age past 10 minutes is the capacity alarm — alert on it and you'll catch stalls at 3 AM instead of at the 7 AM SLA.
- Shared default_pool lets one backfill starve everything; private pools plus max_active_tasks caps are structural, restarts are cosmetic.
- Check the scheduler heartbeat before any task surgery — clearing tasks behind a dead scheduler just re-queues them.
| File | Command / Code | Purpose |
|---|---|---|
| pool_triage.sh | AIRFLOW_HOME=~/airflow airflow pools list | Pool Slots Exhausted |
| dags | from datetime import timedelta | Executor Pressure and max_active_tasks |
| scheduler_check.sh | AIRFLOW_HOME=~/airflow airflow jobs check --job-type SchedulerJob --hostname $SC... | Zombie Scheduler |
| dags | from datetime import timedelta | Priority, Pools, and Placement |
| clear_stuck.sh | AIRFLOW_HOME=~/airflow airflow tasks logs warehouse_sync stuck_task 2026-09-03 -... | Clearing Stuck State Without Replaying the World |
Key takeaways
Common mistakes to avoid
5 patternsRunning every DAG on default_pool with 128 slots
Leaving max_active_tasks unbounded on hourly DAGs
Clearing tasks when the scheduler is dead
Setting Celery worker_concurrency sky-high
Bulk-clearing entire DAG runs to 'unstick' one task
Interview Questions on This Topic
A task sits in queued for 3 hours. What does queued mean?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Airflow. Mark it forged?
5 min read · try the examples if you haven't