Home › Data Engineering › Airflow Task Stuck in Queued Forever — Fix It
Intermediate 5 min · September 23, 2026

Airflow Task Stuck in Queued Forever — Fix It

Airflow tasks queued for hours means slots, scheduler, or stuck state.

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Production
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 21 min
  • ✓An Airflow deployment with Grid view access
  • ✓Pools, executors, and scheduler basics
  • ✓CLI access to pools list and task logs
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Queued means ready but homeless: no worker slot, pool slot, or concurrency limit available — not broken code
  • Check in order: scheduler heartbeat, pool occupancy, parallelism, then greedy-neighbor DAGs hogging slots
  • One uncapped backfill on shared default_pool can starve every other DAG for hours past your SLA
  • Fix structurally (private pools, max_active_tasks caps, deferrable waits); clear only truly stuck instances
✦ Definition~90s read
What is Airflow Task Stuck in Queued Forever?

In the task lifecycle, queued is the waiting room between scheduled (dependencies unmet) and running (a worker owns it). A queued task is fully ready — upstream green, trigger rule satisfied — but homeless: no worker slot, pool slot, or concurrency allowance is free.

★
Picture an airport where planes (tasks) wait at the gate for a runway slot.

Brief queuing is healthy; it means the scheduler Works through bursts in order. Queued-forever (age past 10-15 minutes on a normally fluid system) means the waiting room has no exit.

Four landlords control the exits. Pools grant named slots (default_pool's 128 shared by everyone unless a DAG declares otherwise). Parallelism and max_active_tasks cap how many tasks the scheduler admits globally and per DAG. The executor (Celery, Kubernetes, Local) provides actual workers to run admitted tasks.

The scheduler itself moves tasks between every state — dead scheduler, frozen world. A task clears queued only when all four agree.

Zombie tasks and stuck metadata form the fifth, rarer cause: workers killed by OOM or reschedules leave instances the scheduler must reap via heartbeat timeouts, and wedged metadata needs a surgical tasks clear. The sections below walk all five in diagnosis order — pools, caps, scheduler, placement, clearing — so grey squares become a checklist instead of a mystery.

Plain-English First

Picture an airport where planes (tasks) wait at the gate for a runway slot. Grey queued planes aren't broken — they're homeless: the runway (pool slots) is full, the control tower (scheduler) went home sick, or one jumbo jet (greedy DAG) booked every slot. Clearing the departure board doesn't build runways. You count free slots, wake the tower, park the jumbo in its own lane, and only then rebook the flights that truly missed their window.

Friday 6:40 AM, and the 7 AM SLA is already dead. The Grid view shows a sea of grey: 40 tasks queued, some for 4 hours, workers half-idle behind failed upstreams. Somebody suggests restarting everything. Somebody else already cleared the whole DAG twice, and the grey came back both times.

Queued-forever is Airflow's most misdiagnosed state because grey squares all look alike while their causes differ completely: pool slots exhausted by a greedy neighbor, executor parallelism maxed out, a scheduler that died Tuesday and nobody noticed, or a single task instance wedged in stale metadata. Each needs a different fix, and the wrong one (bulk-clearing a healthy queue behind a dead scheduler) burns hours.

This guide uses the lifecycle vocabulary from the task-states primer — scheduled, queued, running, deferred, and friends — and turns grey squares into specific diagnoses. You'll check pools, parallelism, scheduler heartbeats, and stuck state in order, with the exact commands at each step. By the end, queued age will be your favorite leading indicator instead of your Friday surprise.

What Queued Forever Actually Means

In the lifecycle vocabulary, a task instance moves scheduled to queued to running to a terminal state (success, failed, skipped). Scheduled means dependencies aren't met yet — upstream tasks unfinished, trigger rules unsatisfied. Queued means the task is ready and waiting for a slot: a worker willing to run it, a pool with a free place, and no concurrency ceiling blocking it. That distinction decides your whole response: scheduled-forever is an upstream problem, queued-forever is a capacity or scheduler problem.

Grey squares deceive because every flavor of homelessness looks identical in Grid view. A task queued behind an exhausted pool, one queued behind a dead scheduler, and one wedged in stale metadata all render the same grey. Teams that treat grey as one problem apply one fix — usually bulk-clearing — and watch the grey return, because clearing re-queues tasks into the same slotless world that stranded them.

Build the habit of naming the landlord before evicting anyone. Pool occupancy (pools list), scheduler liveness (heartbeat ticks), executor pressure (parallelism vs running counts), and per-DAG greed (max_active_tasks unbounded) each answer in seconds. Queued age — how long the oldest task has waited — is the metric that turns this from archaeology into alerting: anything past 10 minutes deserves a page, not a shrug.

📊 Production Insight
A team bulk-cleared grey tasks twice behind a dead scheduler and the grey returned both times. Rule: scheduled means upstream trouble, queued means capacity trouble — diagnose accordingly.
🎯 Key Takeaway
Queued is ready-but-homeless; name which slot is missing (pool, worker, scheduler) before clearing anything.

Pool Slots Exhausted: the Shared-Bucket Starvation

Pools are named slot buckets tasks must pass through — every task needs one pool slot plus a worker to run. The default_pool ships with 128 slots shared by every DAG that doesn't declare otherwise, which works until one backfill fans out 200 tasks and holds 121 slots for 6 hours. Every polite DAG behind it queues, workers idle for lack of admittable tasks, and the dashboard shows a healthy cluster doing nothing.

Diagnose with pools list (occupancy per pool) crossed against states-for-dag-run per DAG (who waits). The signature is unmistakable: one pool at 100% used, one DAG holding most slots, everyone else grey. The structural fix has two halves — raise slots deliberately where the ceiling is genuinely low, and isolate where one tenant misbehaves: private pools for critical DAGs (etl_pool=32 for warehouse_sync) so backfills can never eat SLA traffic again.

Treat pool design as capacity planning, not firefighting. Review occupancy weekly as DAG counts grow; defaults that fit 20 DAGs strangle 200. Set pool_slots: 1 on ordinary tasks and reserve higher values for genuinely heavy ones, because a task holding 4 slots runs at 4x the queue cost. Queues drain when admission is managed — not when workers multiply into a full pool.

pool_triage.shBASH
1
2
3
4
5
6
7
8
9
# Pool triage: who holds the slots?
AIRFLOW_HOME=~/airflow airflow pools list
AIRFLOW_HOME=~/airflow airflow pools get default_pool
AIRFLOW_HOME=~/airflow airflow tasks states-for-dag-run warehouse_sync 2026-09-03

# Structural fix: private pool for the critical DAG (via UI or API),
# then point its tasks at etl_pool:
# default_args={"pool": "etl_pool", "pool_slots": 1, ...}
# Cap the hog: max_active_tasks=8 on the backfill DAG
⚠ Don't Scale Workers Into a Full Pool
Doubling workers while default_pool stays at 128 shared slots just buys idle workers. New capacity only helps when the pool has free places to admit tasks — check pools list before any scaling decision.
📊 Production Insight
A backfill held 121 of 128 default_pool slots for 6 hours while doubled workers idled. Rule: private pools for SLA traffic, caps for backfills.
🎯 Key Takeaway
Cross pools list with per-DAG queued counts, then isolate critical DAGs in private pools and cap the hogs.

Executor Pressure and max_active_tasks: Capping the Greedy

Beyond pools, two scheduler-wide ceilings shape every queue: parallelism (total task instances the scheduler admits cluster-wide) and max_active_tasks or max_active_runs per DAG. Parallelism at 32 with 200 runnable tasks means 168 wait no matter how many workers you own. An uncapped hourly DAG fanning out 200 tasks per run fills that ceiling alone — polite DAGs queue behind a neighbor they'd never suspect from their own logs.

The polite-DAG pattern above shows the full stack in one file: max_active_tasks=8 caps the DAG's footprint, pool etl_pool isolates its admission, pool_slots=1 keeps per-task cost minimal, and execution_timeout bounds runaways so a hung task can't squat a slot forever. Retries with backoff ride along so transient blips heal instead of failing into manual clears. Copy this header onto every new DAG and the greedy-neighbor class shrinks to legacy DAGs you migrate on sight.

Defaults age badly — re-tune quarterly as DAG counts grow. What fit 20 DAGs (parallelism 32, everything on default_pool) strangles 200. Capacity planning here is scheduling work, and the review takes an hour: list every DAG without max_active_tasks, cap them, and watch queued age fall across the fleet the same day.

dags/polite_hourly.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
from datetime import timedelta
import pendulum
from airflow.sdk import dag, task

@dag(
    schedule="@hourly",
    start_date=pendulum.datetime(2026, 1, 1, tz="UTC"),
    catchup=False,
    max_active_tasks=8,  # cap this DAG so it can't starve neighbors
    default_args={
        "owner": "platform",
        "pool": "etl_pool",
        "pool_slots": 1,
        "retries": 3,
        "retry_delay": timedelta(minutes=5),
        "execution_timeout": timedelta(hours=1),
    },
    tags=["capped"],
)
def polite_hourly():
    @task
    def extract() -> str:
        return "partition-2026-09-23"

    @task
    def load(partition: str) -> None:
        print(f"loading {partition}")

    load(extract())

polite_hourly()
📊 Production Insight
One uncapped hourly DAG filled parallelism=32 alone and starved 10 polite DAGs. Rule: every DAG declares max_active_tasks; no exceptions.
🎯 Key Takeaway
Stack parallelism, per-DAG caps, and private pools — and re-tune quarterly as DAG counts grow past old defaults.

Zombie Scheduler: the Tower Went Home Sick

The scheduler is the control tower: it moves tasks from scheduled to queued, admits them to slots, and reaps the dead. When it dies — OOMKill, bad deploy, disk-full on its log volume — everything freezes in place. New runs never leave scheduled, queued tasks never get adopted, and every task-level fix fails because no living process reads the queue. One real scheduler flatlined Tuesday; the team debugged tasks until Friday's SLA breach.

Check liveness before anything else: jobs check for the SchedulerJob plus fresh heartbeat lines in scheduler logs (seconds old, not days). Flat heartbeat means restart or fail over to standby in HA setups — and in HA, verify only one scheduler leads, since dual-active schedulers double-run tasks into duplicate writes. After recovery, confirm with list-runs and states-for-dag-run that fresh tasks flow scheduled to queued to running within a minute.

Zombie schedulers deserve special suspicion after K8s churn: liveness-probe restarts and node scale-downs leave scheduler processes half-alive — heartbeating just enough to block failover, dead enough to schedule nothing. Alert on heartbeat age directly (stale past 60s pages), not on queued counts that lag the cause by hours. The tower must be awake before any runway discussion matters.

scheduler_check.shBASH
1
2
3
4
5
6
7
8
9
# Is the scheduler alive? heartbeat first, surgery never
AIRFLOW_HOME=~/airflow airflow jobs check --job-type SchedulerJob --hostname $SCHED_HOST

# Recent scheduler ticks in logs (want seconds-ago, not Tuesday)
grep -iE 'heartbeat|scheduler.*tick|adopting|queued' ~/airflow/logs/scheduler/*.log | tail -10

# After failover/restart: confirm tasks leave scheduled state
AIRFLOW_HOME=~/airflow airflow dags list-runs -d warehouse_sync --limit 5
AIRFLOW_HOME=~/airflow airflow tasks states-for-dag-run warehouse_sync 2026-09-03
📊 Production Insight
A scheduler dead since Tuesday froze all queues until Friday's SLA breach. Rule: alert on heartbeat age past 60s, not on queue length.
🎯 Key Takeaway
Heartbeat first, always — a flat scheduler freezes everything, and no task fix works until failover or restart.

Priority, Pools, and Placement: Designing Who Waits

Priority and pool design decide who waits when slots run short. priority_weight orders admission inside a pool's queue — interactive pipelines at 5+ jump ahead of backfills at 3, so stale history can't block fresh SLA traffic. Separate pools make the ordering structural: backfill_pool with modest slots bounds history replay regardless of weights, while etl_pool stays reserved for daily promises. Heavy tasks declare pool_slots=2 so the scheduler accounts their real cost instead of admitting four heavies into four 1-slot places.

max_active_runs=1 on backfill DAGs stops the stampede at the source: one run at a time means catchup processes serially instead of fanning 90 runs at once. Combined with catchup=False on interactive DAGs, the fleet stops competing with its own history. These are design decisions, not incident responses — set them when the DAG is born, because retrofitting priority during an outage means editing the DAGs you most fear touching.

Review the matrix quarterly: list every DAG's pool, weight, max_active_tasks, and max_active_runs in one table. Anything on default_pool with weight 1 and no caps is a future incident wearing a trench coat. Thirty minutes of review beats another Friday of grey squares.

dags/capped_backfill.pyPYTHON
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
from datetime import timedelta
import pendulum
from airflow.sdk import dag, task

@dag(
    schedule="@daily",
    start_date=pendulum.datetime(2026, 1, 1, tz="UTC"),
    catchup=False,
    max_active_runs=1,  # one run at a time: backfills can't stampede
    default_args={
        "owner": "platform",
        "pool": "backfill_pool",
        "priority_weight": 3,  # below interactive pipelines (5+)
        "retries": 2,
        "retry_delay": timedelta(minutes=10),
    },
    tags=["backfill"],
)
def capped_backfill():
    @task(pool_slots=2)  # heavy task declares its real cost
    def heavy() -> None:
        print("heavy lifting")

    heavy()

capped_backfill()
📊 Production Insight
Backfills at weight 1 on default_pool blocked SLA pipelines weekly until pools plus weights isolated them. Rule: history replays in its own pool, never in SLA lanes.
🎯 Key Takeaway
Weights order the queue, private pools bound the tenants, max_active_runs serializes backfills — set all three at DAG birth.

Clearing Stuck State Without Replaying the World

Sometimes one instance is genuinely wedged: metadata stale after a kill, a zombie the heartbeat already reaped, a failure whose cause is fixed but whose state never moved. Clearing returns it to the queue with try_number bumped and per-try history intact — logs stay inspectable, lineage survives. The ritual matters: read the log tail first (the fatal line names the cause), fix the cause, clear just that instance, then verify it flows to running within a minute.

Restraint is the whole skill. Bulk-clearing a DAG run replays every downstream write — duplicate warehouse loads, double-fired webhooks, re-sent emails — while the original failure recurs unfixed. The incident that motivates this section featured two full-DAG clears that re-queued healthy tasks into the same slotless pool and duplicated a day's loads. Scope every clear to the stuck instance (-t stuck_task), and snapshot list-runs output before any bulk operation so you can account for what replayed.

Close the loop with prevention: execution_timeout on long tasks bounds squatters, queued-age alerts catch the next stall at minute 10 instead of hour 4, and deferrable operators move hour-long waits to the triggerer where they cost nothing. Clears should be rare, surgical, and followed by a config change that makes that clear unnecessary next time.

clear_stuck.shBASH
1
2
3
4
5
6
7
8
# Clear ONE stuck instance: read tail, fix cause, clear, verify
AIRFLOW_HOME=~/airflow airflow tasks logs warehouse_sync stuck_task 2026-09-03 --tail 50
AIRFLOW_HOME=~/airflow airflow tasks clear warehouse_sync -t stuck_task --yes
AIRFLOW_HOME=~/airflow airflow tasks states-for-dag-run warehouse_sync 2026-09-03

# Bulk operations ONLY with a snapshot + intent to replay downstream:
# AIRFLOW_HOME=~/airflow airflow dags list-runs -d warehouse_sync --limit 5
# AIRFLOW_HOME=~/airflow airflow tasks clear warehouse_sync --start-date 2026-09-03 --yes
📊 Production Insight
Two bulk clears replayed a day's warehouse loads without fixing the full pool behind them. Rule: surgical clears plus a prevention change, never blind replays.
🎯 Key Takeaway
Read tail, fix cause, clear one instance, verify — and follow every clear with the config change that prevents the next.
● Production incidentPOST-MORTEMseverity: high

40 Tasks Queued 4 Hours While Doubled Workers Idled

Symptom
By 3:20 AM the Grid showed 40 grey queued squares across 3 DAGs sharing default_pool, queued age climbing past 240 minutes. The 7 AM SLA breached with dashboards stale; on-call woke at 6:40 AM to a queue that hadn't moved in 4 hours. Doubling the worker fleet at 5 AM changed nothing — new workers sat idle with zero tasks assigned.
Assumption
The team assumed queued meant broken task code, so two engineers spent 3 hours reading SQL that was fine. They then assumed workers were too few and doubled the Celery fleet — but default_pool stayed at 128 shared slots, so the new workers idled while the queue sat. Nobody checked the scheduler heartbeat or the pool dashboard until the SLA had already breached.
Root cause
Three layers failed together. All DAGs shared default_pool (128 slots), and an uncapped backfill DAG grabbed 121 of them at 2 AM. The scheduler had gone zombie Tuesday after an OOMKill — its heartbeat flatlined, so nothing new scheduled while old tasks sat. Doubling Celery workers changed nothing because no pool slots were free to admit tasks to them. Queued age passed 240 minutes with workers at 30% while the SLA burned.
Fix
They gave warehouse_sync its own pool (etl_pool=32), capped the backfill DAG at max_active_tasks=8, and restarted the zombie scheduler that had been flat since Tuesday. They added a queued-age alert at 10 minutes and moved the 2-hour sensor waits to deferrable operators. Queued age dropped from 240 minutes to under 3, and the 7 AM SLA has held for 6 straight weeks.
Key lesson
  • Queued age past 10 minutes is the capacity alarm — alert on it and you'll catch stalls at 3 AM instead of at the 7 AM SLA.
  • Shared default_pool lets one backfill starve everything; private pools plus max_active_tasks caps are structural, restarts are cosmetic.
  • Check the scheduler heartbeat before any task surgery — clearing tasks behind a dead scheduler just re-queues them.
Production debug guideFive checks — pools, scheduler, greedy neighbors, zombies, stuck state — that drain the queue.5 entries
Symptom · 01
Tasks queued for hours across multiple DAGs
→
Fix
Run AIRFLOW_HOME=~/airflow airflow pools list and AIRFLOW_HOME=~/airflow airflow tasks states-for-dag-run sales_daily 2026-09-03. If the pool shows 100% used with queues everywhere, raise that pool's slots deliberately or lower the greedy DAG's max_active_tasks, then watch queued age fall on the next runs. Don't add workers until pools have headroom — workers can't serve tasks no pool will admit.
Symptom · 02
Nothing leaves queued no matter what you clear
→
Fix
Check the scheduler is alive before anything else: read its heartbeat metric and logs for recent ticks (airflow jobs check --job-type SchedulerJob --hostname $SCHED_HOST). Flat heartbeat means a dead or zombie scheduler — restart it or fail over to standby, then confirm new tasks leave scheduled state. Clearing tasks behind a dead scheduler just re-queues them.
Symptom · 03
One DAG starves every other pipeline
→
Fix
Find the greedy neighbor: compare per-DAG queued counts from states-for-dag-run across DAGs sharing the pool. The DAG with hundreds queued and unbounded max_active_tasks is the hog — cap it (max_active_tasks=8) or move it to a private pool. Queued age on polite DAGs should recover within one schedule interval after the cap lands.
Symptom · 04
Slots free but a task still sits queued
→
Fix
Suspect zombies when slots look free but tasks sit: check worker logs for OOMKilled, node reschedules, or liveness restarts around the stuck task's start time, and confirm heartbeat expiry reaps them. Tune task_instance_heartbeat_sec for your K8s churn rate, and move hour-long sensor waits to deferrable operators so they park in the triggerer instead of dying on workers.
Symptom · 05
One wedged instance blocks everything downstream
→
Fix
Read the tail first (AIRFLOW_HOME=~/airflow airflow tasks logs my_dag my_task 2026-09-03 --tail 50), fix the underlying cause, then clear just that instance: AIRFLOW_HOME=~/airflow airflow tasks clear my_dag -t my_task --yes. Verify it reruns with a bumped try_number and per-try history intact. Never bulk-clear whole DAG runs unless you intend to replay every downstream write.
Stuck in Queued: Diagnose at a Glance
Root CauseHow to ConfirmFixPrevention
Pool slots exhaustedairflow pools list shows 100% used; states-for-dag-run shows queued widelyRaise pool slots or move greedy DAGs to private poolsDedicated pools for critical DAGs; weekly occupancy review
Executor / parallelism saturatedQueued across all pools; scheduler CPU pegged; parallelism at ceilingRaise parallelism deliberately or cut max_active_tasks on greedy DAGsAlert on queued age over 10 minutes as capacity signal
Dead or zombie schedulerScheduler heartbeat flat; new runs never leave scheduled stateRestart scheduler / fail over to standby; reap zombiesHA schedulers plus heartbeat alerting
Stuck task state needing clearSlots free but task sits queued; log shows old failure; metadata staleClear just that instance with tasks clear; verify rerunexecution_timeout plus queued-age alerts catch it early
⚙ Quick Reference
5 commands from this guide
FileCommand / CodePurpose
pool_triage.shAIRFLOW_HOME=~/airflow airflow pools listPool Slots Exhausted
dagspolite_hourly.pyfrom datetime import timedeltaExecutor Pressure and max_active_tasks
scheduler_check.shAIRFLOW_HOME=~/airflow airflow jobs check --job-type SchedulerJob --hostname $SC...Zombie Scheduler
dagscapped_backfill.pyfrom datetime import timedeltaPriority, Pools, and Placement
clear_stuck.shAIRFLOW_HOME=~/airflow airflow tasks logs warehouse_sync stuck_task 2026-09-03 -...Clearing Stuck State Without Replaying the World

Key takeaways

1
Queued means waiting for a slot
diagnose capacity (pools, parallelism, scheduler) before touching task code.
2
Check scheduler heartbeat first; clearing tasks behind a dead scheduler just re-queues them.
3
One greedy DAG starves ten polite ones
cap max_active_tasks and give critical DAGs private pools.
4
Alert on queued age past 10 minutes, not just on failures, to catch stalls hours before SLAs.
5
Zombie tasks come from OOMKills and reschedules
heartbeat timeouts reap them; deferrable waits avoid them.
6
Clear single stuck instances (history preserved, try_number bumps)
never bulk-clear healthy DAGs.

Common mistakes to avoid

5 patterns
×

Running every DAG on default_pool with 128 slots

Symptom
One greedy backfill eats all 128 slots and every other DAG queues for hours. The pool dashboard shows 100% used but nobody knows by whom.
Fix
Give the critical DAG its own pool with reserved slots (e.g. etl_pool=16) instead of sharing default_pool. Review pool occupancy weekly as DAG counts grow.
×

Leaving max_active_tasks unbounded on hourly DAGs

Symptom
A single DAG fans out 200 tasks per run and starves ten polite DAGs. Queued age climbs past the SLA while workers look busy.
Fix
Set max_active_tasks on every greedy DAG (start with 8-16) and cap parallelism deliberately. Monitor queued age, not just queue length.
×

Clearing tasks when the scheduler is dead

Symptom
Cleared tasks return to queued and sit again. Hours burn on task-level surgery while the scheduler heartbeat has been flat since Tuesday.
Fix
Check scheduler heartbeats and logs first; restart the scheduler (or fail over to standby in HA) before clearing anything. A dead scheduler makes every slot look full.
×

Setting Celery worker_concurrency sky-high

Symptom
Workers OOMKill under load, tasks return to queued after each death, and the queue never drains despite 'plenty' of workers.
Fix
Keep Celery worker concurrency matched to real CPU/RAM (8-16 per worker), monitor broker queue depth, and autoscale on queued age. Overconcurrency causes OOMKills that look like queue stalls.
×

Bulk-clearing entire DAG runs to 'unstick' one task

Symptom
Thirty healthy tasks rerun, downstream systems get duplicate writes, and the original failure recurs because its cause was never read.
Fix
Read the task log tail for the fatal line, fix the cause, then clear just that task instance (try_number bumps, history preserved). Never bulk-clear a whole DAG without a state snapshot.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
A task sits in queued for 3 hours. What does queued mean?
Q02SENIOR
Walk me through your queued-diagnosis order.
Q03SENIOR
Cleared tasks return to queued and sit again. Now what?
Q04SENIOR
What creates zombie tasks and how does the scheduler handle them?
Q05SENIOR
Design concurrency for 200 DAGs sharing one warehouse.
Q01 of 05JUNIOR

A task sits in queued for 3 hours. What does queued mean?

ANSWER
queued means the task is ready but waits for a worker slot, pool slot, or concurrency limit — homelessness, not brokenness. Short queues are normal; queued age past 10 minutes signals capacity trouble, and I confirm with pools list plus states-for-dag-run.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Queued vs scheduled — what's the difference?
02
Can I prioritize one DAG over others?
03
How do I stop one greedy DAG starving the rest?
04
Is clearing a task safe? Does it lose history?
05
Scheduler vs executor — which do I blame first?
06
Do deferrable operators help with queued pressure?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Drawn from code that ran under real load.

Follow
✓ Verified
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
🔥

That's Airflow. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
Kafka Exactly-Once Semantics Broken by Producer Retries
1 / 1 · Airflow
Next
dbt Compilation Error: Model Depends on a Node Not Found
→