Home › Mobile › Kotlin CancellationException: Scopes Leaking After Exit
Intermediate 5 min · September 23, 2026
Kotlin Coroutine JobCancellationException — Scope Leaks

Kotlin CancellationException: Scopes Leaking After Exit

Rethrow CancellationException and scope to viewModelScope: swallowed cancellation leaks work, and GlobalScope outlives every screen..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 15 min
  • ✓Basic Kotlin coroutines: launch, async, suspend functions
  • ✓An Android app using ViewModel or Lifecycle components
  • ✓Familiarity with logcat and basic unit testing
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Coroutines cancel by throwing CancellationException into the job, and JobCancellationException is its normal shutdown signal
  • Swallowing that exception in catch-all blocks keeps dead work alive, draining battery and data in the background
  • Scope every coroutine to viewModelScope or lifecycleScope so teardown cancels work automatically
  • Use SupervisorJob or supervisorScope when one child's failure shouldn't kill its siblings
  • You'll fix most issues by rethrowing cancellation, scoping by owner, and adding ensureActive() to loops
✦ Definition~90s read
What is Kotlin Coroutine JobCancellationException?

JobCancellationException is the exception coroutines use to implement cancellation. When you cancel a Job — directly, or indirectly by destroying its scope — the runtime throws CancellationException (JobCancellationException for job-backed coroutines) inside the coroutine at its next suspension point or explicit checkpoint. finally blocks execute, resources unwind, and the job completes as cancelled rather than failed.

★
Picture a restaurant kitchen where every order belongs to a table.

Parent scopes observe child cancellation as normal teardown, not as an error, so viewModelScope can cancel dozens of children on rotation silently.

Structured concurrency is the design that makes this safe: coroutines form a parent-child tree rooted in a scope, children cannot outlive their parent, and failures propagate upward by default. viewModelScope roots work in the ViewModel lifecycle (cancelled in onCleared), lifecycleScope roots it in UI lifecycles, and coroutineScope/supervisorScope builders create sub-hierarchies with all-or-nothing or independent failure semantics. SupervisorJob variants change one rule — child failure doesn't propagate — for work items that stand alone.

Everything breaks when code escapes the tree or fights the signal. GlobalScope and leaked custom scopes create parentless coroutines that outlive their screens, capturing views and contexts for the process lifetime. catch-all handlers that swallow CancellationException convert teardown into zombie work that retries forever.

Non-cooperative loops without ensureActive or suspending checkpoints ignore cancellation entirely. Cooperative cancellation (ensureActive, cancellable suspending calls, NonCancellable-only cleanup) plus lifecycle-rooted scopes plus correct supervisor choice is the complete defense — each piece verifiable with tests and the debugger.

Plain-English First

Picture a restaurant kitchen where every order belongs to a table. When a table leaves, the chef shouts 'stop their dishes' and cooks drop that work instantly — that's a scope cancelling. Bugs happen when a cook covers their ears and keeps cooking the cancelled dish (swallowed CancellationException), or takes orders from nobody and grills after closing (GlobalScope leak). Fix both by letting cancellations through and attaching every cook to a table.

JobCancellationException is the exception you're supposed to see — until it shows up where it shouldn't, or never shows up because something swallowed it. Coroutines cancel by throwing: when a scope dies, every child gets a CancellationException that unwinds its stack. Navigate away from a screen and viewModelScope cancels its children; that's the system working. The trouble starts when a catch-all block eats that exception and the 'cancelled' work keeps running, or when a GlobalScope launch escapes every scope and leaks for the rest of the process.

In production these bugs wear two masks. The loud one is the crash: a leaked coroutine touches a destroyed view and throws IllegalStateException minutes after the user left. The quiet one is the drain: cancelled downloads keep burning data and battery, duplicate jobs pile up on every rotation, and one failed thumbnail cancels an entire screen because the scope type was wrong. Both trace back to scope discipline, not coroutine magic.

This article gives you that discipline. You'll learn structured concurrency in practical terms, when viewModelScope and lifecycleScope do the work for you, how SupervisorJob isolates failures, and how ensureActive makes loops cancellable. You'll leave able to read any coroutine crash and know which scope broke its promise.

Structured Concurrency and What Cancellation Really Means

Structured concurrency is a simple contract: every coroutine is a child of some scope, children die with their parent, and failures travel upward unless you say otherwise. When viewModelScope is cancelled in onCleared, each child coroutine gets a CancellationException thrown at its next suspension point, its finally blocks run, and the job ends. No orphans, no strays, no work billed to a screen that's gone. The hierarchy is the entire memory-safety story for coroutines.

JobCancellationException is the concrete type you'll meet — a subclass of CancellationException (which extends IllegalStateException, confusingly) thrown when a Job is cancelled. Treat its appearance during teardown as a success log, not a crash: it proves the scope did its job. The danger is purely in mishandling. Catch it in a generic handler and show an error toast, and users see phantom failures on every navigation. Swallow it into a retry loop, and dead work runs forever.

The rule that prevents both: catch CancellationException first and rethrow it, always, before handling real failures. Write that ordering so consistently it becomes muscle memory, and review any catch (e: Exception) in suspend code as a suspected cancellation bug until proven otherwise. Your future self, reading a 2 AM crash log, will thank you for the discipline.

FeedViewModel.ktKOTLIN
1
2
3
4
5
6
7
8
9
10
11
12
13
class FeedViewModel(private val repo: FeedRepo) : ViewModel() {
    // viewModelScope cancels automatically in onCleared()
    fun refresh() = viewModelScope.launch {
        try {
            val items = repo.load() // suspending, cancellable
            _state.value = State.Ok(items)
        } catch (e: CancellationException) {
            throw e // MUST rethrow: cancellation is not failure
        } catch (e: IOException) {
            _state.value = State.Error("Offline — showing cache")
        }
    }
}
📊 Production Insight
Most 'mystery background work' incidents end at a catch-all that logged cancellation as an error and retried it.
🎯 Key Takeaway
Children die with their scope via CancellationException. Rethrow it first in every handler; handle only real failures after.

viewModelScope, lifecycleScope, and Escaping GlobalScope

viewModelScope and lifecycleScope exist so you rarely manage jobs by hand. viewModelScope lives as long as the ViewModel and cancels in onCleared — perfect for repositories, flows, and anything that should survive rotation but die with the screen's data owner. lifecycleScope ties to a LifecycleOwner and cancels on destroy; use the viewLifecycleOwner variant in fragments so collection stops when the view dies, not when the fragment instance does. Together they cover nearly all Android work.

The anti-pattern is anything that escapes them. GlobalScope launches live for the whole process and hold every reference they capture — views, contexts, callbacks — until the work finishes or the app dies. Hand-rolled CoroutineScopes are fine only when you manage their lifecycle explicitly: store them, cancel in the matching teardown, and never create one per click without cancelling the last. Audit for GlobalScope with a grep; each hit is a leak report waiting for a user.

One subtlety: withContext doesn't change your scope, it only switches dispatchers — cancellation still flows through. That's good: a withContext(Dispatchers.IO) network call inside viewModelScope still cancels with the ViewModel. But launching a new coroutine inside withContext creates a child of the same scope, not of the call — keep track of parentage when nesting, and prefer structured builders (coroutineScope, supervisorScope) over bare launch for dependent work.

Scopes.ktKOTLIN
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
class DetailViewModel(repo: DetailRepo, id: String) : ViewModel() {
    init {
        // UI-bound reload: dies with the ViewModel
        viewModelScope.launch { repo.watch(id).collect { _detail.value = it } }
    }
}

// In a Fragment: dies with the view lifecycle
fun Fragment.loadImage(url: String) {
    viewLifecycleOwner.lifecycleScope.launch {
        val bitmap = withContext(Dispatchers.IO) { fetch(url) }
        imageView.setImageBitmap(bitmap) // safe: scope outlives this line
    }
}
📊 Production Insight
Leaked coroutines touching destroyed views are a top-five Android crash; scoping to the lifecycle owner deletes the whole category.
🎯 Key Takeaway
Bind work to its owner's scope, use viewLifecycleOwner in fragments, and treat every GlobalScope as a bug.

SupervisorJob and supervisorScope: Isolating Failures

Default scopes propagate failure upward: one child throws, the parent cancels, siblings die. That's correct for transactions — a checkout whose payment step failed shouldn't still ship the order. It's wrong for independent work — a dashboard whose news tile failed should still show stats and alerts. SupervisorJob and supervisorScope flip the rule: children fail alone, siblings continue, and the parent only fails if it throws directly.

Pick the semantics per screen, not per app. Feeds, dashboards, and prefetch batches want supervision because each item stands alone. Multi-step writes, migrations, and anything with commit/rollback wants plain coroutineScope so partial completion is impossible. Mixing them is fine and common: a supervisorScope for the independent fetches feeding into a coroutineScope commit phase, each chosen deliberately and commented.

Handle child results explicitly under supervision, since exceptions no longer cancel the batch — they arrive at await(). Wrap each await in its own try/catch (rethrowing CancellationException first, always) and decide per child: default value, skip, or aggregated error state. Unhandled child exceptions under a supervisor still crash if never awaited, so never fire-and-forget an async you intend to supervise — collect every deferred.

Dashboard.ktKOTLIN
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
fun ViewModel.loadDashboard(repo: DashRepo) = viewModelScope.launch {
    // Independent tiles: one failure must not kill the rest
    supervisorScope {
        val stats = async { repo.stats() }
        val news = async { repo.news() }
        val alerts = async { repo.alerts() }
        _state.value = State.Ok(stats.safe(), news.safe(), alerts.safe())
    }
}

// Deferred<T>.safe(): null on failure instead of cancelling siblings
suspend fun <T> Deferred<T>.safe(): T? = try {
    await()
} catch (e: CancellationException) {
    throw e // never swallow cancellation, even here
} catch (e: Exception) {
    null
}
📊 Production Insight
Full-screen error pages caused by one bad tile are the signature of a missing supervisor — users forgive a missing widget, not a missing screen.
🎯 Key Takeaway
Supervise independent tiles, use plain scopes for transactions, and handle every awaited child explicitly.

Cooperative Cancellation With ensureActive and finally

Cancellation is cooperative — it lands only where the coroutine lets it. Suspending functions like delay, withContext, and retrofit awaits check for cancellation automatically, so typical network code cancels cleanly. Pure computation doesn't: a tight loop parsing ten thousand records never suspends, never checks, and sails past cancel() as if nothing happened. The job claims cancelled while the CPU burns to completion, and the replacement work queues behind it.

ensureActive() is the checkpoint you place yourself. It throws CancellationException immediately if the job was cancelled, costs nothing when it wasn't, and belongs at the top of every loop body plus between heavy phases. yield() does the same while also giving other coroutines a scheduling turn — useful in CPU-bound batches on limited dispatcher threads. Replace Thread.sleep with delay everywhere inside coroutines; the former blocks a thread cancellation can't reach, the latter suspends cooperatively.

Cleanup during cancellation needs finally, which runs as the stack unwinds. Close files, release locks, and discard partial state there — but remember the coroutine is still cancelled inside finally, so any suspending cleanup must use withContext(NonCancellable). Keep that block minimal and fast: it's a graceful exit, not a second chance to finish the work. Anything long-lived there converts cancellation back into the delay you were trying to eliminate.

Sync.ktKOTLIN
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
suspend fun syncAll(items: List<Record>, api: SyncApi) {
    for ((i, item) in items.withIndex()) {
        ensureActive() // checkpoint: throws if the job was cancelled
        api.upload(item) // cancellable suspend call
        if (i % 50 == 0) yield() // extra checkpoint on CPU-heavy batches
    }
}

// finally still runs during cancellation: clean up here
suspend fun download(url: String): File = try {
    fetchToTemp(url)
} finally {
    withContext(NonCancellable) { tempFiles.deleteStale() }
}
📊 Production Insight
Sync loops without ensureActive are the classic 'cancel does nothing' report — one line per loop ends the whole complaint class.
🎯 Key Takeaway
Checkpoint loops with ensureActive, prefer delay over sleep, and restrict NonCancellable cleanup to fast essentials.

Proving It: Tests, Debugger, and CI Gates

Testing cancellation is straightforward and almost nobody does it — which is why these bugs reach production. The pattern: launch the work in a test scope, advance or delay briefly, cancel the job, then assert the aftermath. Assert no further network calls occur (verify your fake received nothing more), assert partial state was cleaned up, and assert siblings under supervision still completed. kotlinx-coroutines-test with StandardTestDispatcher gives you deterministic control without real waiting.

The coroutine debugger covers the production side. Run the app, reproduce the leak, and inspect live coroutines: each shows its name, state, stack, and creation trace. Filter to your feature's scope and look for jobs alive past their owner's destruction. The creation trace names the exact launch site — no guessing, no logging archaeology. Make this check part of feature QA for anything long-running: start it, leave the screen, confirm it dies.

For CI, add two permanent gates. First, a GlobalScope ban — a grep or lint check that fails the build on unstructured launches. Second, cancellation tests for every long-running job: sync loops, downloads, collectors. These tests run in milliseconds with virtual time and catch every regression where someone wraps your careful code in a catch-all or moves it to a wider scope. Cancellation you don't test is cancellation you'll lose.

⚠ Cancelled Parents Eat Their Children
Never launch coroutines from a finally block or from a cancelled scope's context without a fresh Job — children of a cancelled parent cancel instantly, and the failure looks like a haunted no-op.
📊 Production Insight
Cancellation tests with virtual time run in milliseconds — there's no cost excuse for shipping untested long-running work.
🎯 Key Takeaway
Test cancel-then-assert-quiet for every long job, confirm leaks die in the debugger, and ban GlobalScope in CI.

Keeping Scopes Honest as the App Grows

Long-term health comes from making the right scope the easiest scope. Provide base ViewModels or use-case helpers that expose the correct scope so feature code never chooses. Document the team's two scope types — supervised for independent tiles, plain for transactions — with one example each in the repo wiki. New engineers copy the example; the examples encode the policy, and the policy stops being tribal knowledge.

Review with a checklist, not vibes. Every suspend function with try/catch: does it rethrow CancellationException first? Every launch: which scope owns it, and when does that scope die? Every loop over unbounded work: where's the ensureActive checkpoint? Every async: is it awaited, and under which failure semantics? Five questions, thirty seconds each, and the review catches what tests can't — intent mismatches where the code works but means the wrong thing.

Finally, watch the signals that predict the next incident. Rising background battery in vitals, server request growth without user growth, and IllegalStateException crashes naming destroyed views all point at scope discipline slipping. Treat any of them as a trigger for a coroutine audit: dump the debugger, grep GlobalScope and bare catch blocks, re-run cancellation tests. You'll catch the next leak while it's a metric wiggle, not a weekend incident.

📊 Production Insight
Battery and server-load anomalies surface scope leaks weeks before users report them — instrument first, debug second.
🎯 Key Takeaway
Encode scope choice in helpers and examples, review five questions per coroutine PR, and audit on the first metric wiggle.
● Production incidentPOST-MORTEMseverity: high

A Swallowed Cancellation Retried Sync Forever and Tripled Server Load

Symptom
Server request volume tripled over two days with no release and no traffic spike. Android vitals showed rising background battery use, and a few users reported the app felt hot. Backend logs showed the same devices retrying the same sync endpoint hundreds of times.
Assumption
The team blamed the backend because error rates rose on the server dashboard too — more requests were arriving. Nobody connected it to the client retrying, since each retry looked like legitimate user behavior in the logs.
Root cause
The sync ran in a hand-rolled scope with catch (e: Exception) that logged and retried — including on CancellationException. Cancelling never unwound the coroutine; each retry re-entered the loop, and every rotation launched another immortal copy. The retry storm presented as backend overload.
Fix
The catch block was split to rethrow CancellationException immediately, the sync loop gained ensureActive() per batch, and the whole job moved into viewModelScope so rotation cancels it. A test now cancels mid-sync and asserts no further network calls happen. Server load dropped back the same day.
Key lesson
  • A catch-all around suspend work must rethrow CancellationException first — swallowing it turns every cancel into a leak.
  • Client-side leaks look like server-side load; check cancellation behavior before scaling the backend.
  • Rotation without scoped jobs multiplies work: every config change spawned a sync that never died.
Production debug guideFive checks that separate swallowed cancellation from true leaks in minutes.5 entries
Symptom · 01
Cancelled work keeps running after the user leaves the screen
→
Fix
Add a temporary log in the catch block printing e::class.simpleName. If you see JobCancellationException or CancellationException being handled as failure, split the catch: catch CancellationException { throw it } first, then your real error handling. Re-run and confirm background work actually stops.
Symptom · 02
Memory grows and old screens seem to stay alive
→
Fix
Open the coroutine debugger (Android Studio: View > Tool Windows > Coroutine Debugger) while reproducing. Look for coroutines whose owner (Activity/ViewModel) is destroyed but whose Job is still active. The creation stack trace points at the leaking launch — usually GlobalScope or a forgotten custom scope.
Symptom · 03
One failed item cancels the entire screen's load
→
Fix
Write a test that launches the batch, fails one child (throw IOException in one async), and asserts the siblings complete. If siblings die, wrap the batch in supervisorScope or give the scope a SupervisorJob, then re-run. Keep coroutineScope only where all-or-nothing is the requirement.
Symptom · 04
Cancelling a sync or loop has no effect
→
Fix
In a test, launch the loop, cancel the job after 100ms, and assert it stops promptly (join with timeout). If it runs to completion, insert ensureActive() at the top of the loop body and replace Thread.sleep with delay. Re-run until cancellation lands within milliseconds.
Symptom · 05
You need to prove the scope wiring is correct before changing code
→
Fix
Log scope identity (this.coroutineContext[Job]) at launch and at teardown (onCleared/onDestroy). If teardown fires without the job completing, the scope wiring is correct and the bug is elsewhere. If the job never hears about teardown, the launch escaped its scope — move it inside.
Cancellation and Leak Causes, Checks, and Fixes
Root CauseHow to ConfirmFixPrevention
CancellationException swallowed by catch-allBreakpoint the catch; log e::class.simpleName on every catchRethrow CancellationException; catch only real failuresLint: no bare catch(Exception) in suspend code
Coroutine launched outside lifecycle scopeDump coroutine debugger; owner destroyed but job activeMove to viewModelScope/lifecycleScope child scopeBan GlobalScope in review; scope by owner
Sibling failure cancels whole batchReproduce one child failure; observe siblings diesupervisorScope or SupervisorJob for independent workChoose scope type per failure semantics
Non-cooperative loop ignores cancelCancel job; loop still completes in testensureActive() in loops; cancellable suspend callsUnit-test cancellation of every long loop
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
FeedViewModel.ktclass FeedViewModel(private val repo: FeedRepo) : ViewModel() {Structured Concurrency and What Cancellation Really Means
Scopes.ktclass DetailViewModel(repo: DetailRepo, id: String) : ViewModel() {viewModelScope, lifecycleScope, and Escaping GlobalScope
Dashboard.ktfun ViewModel.loadDashboard(repo: DashRepo) = viewModelScope.launch {SupervisorJob and supervisorScope
Sync.ktsuspend fun syncAll(items: List<Record>, api: SyncApi) {Cooperative Cancellation With ensureActive and finally

Key takeaways

1
Cancellation throws CancellationException into the coroutine
it's control flow, not failure, until you mishandle it.
2
Never swallow CancellationException
rethrow it and catch only the failures you can handle.
3
Launch everything in viewModelScope, lifecycleScope, or an explicit child scope
never GlobalScope.
4
Use SupervisorJob/supervisorScope for independent work, plain scopes for all-or-nothing transactions.
5
Make loops cooperative with ensureActive() and cancellable suspending calls.
6
Confirm leaks with the coroutine debugger, then prove cancellation with a unit test per long-running job.

Common mistakes to avoid

5 patterns
×

Swallowing CancellationException in a generic catch block

Symptom
Cancelled screens keep consuming battery and network in the background, and navigation behaves erratically because dead coroutines never die.
Fix
Catch CancellationException separately and rethrow it; handle only real failures (IOException, HttpException) in the generic branch. Never swallow cancellation into a fallback that pretends the work succeeded.
×

Launching coroutines in GlobalScope or a hand-rolled scope

Symptom
Work outlives screens, leaks View references, and crashes with IllegalStateException when it touches destroyed views minutes later.
Fix
Scope every launch/async to a lifecycle-aware scope (viewModelScope, lifecycleScope) or an explicit child scope. Reserve GlobalScope for nothing — there is no production use that survives review.
×

Using coroutineScope where one failing child shouldn't kill siblings

Symptom
A single failed image load cancels the entire screen's data load, showing a full error page for one bad thumbnail.
Fix
Wrap each child in supervisorScope or give the parent a SupervisorJob so one failure cancels only its own subtree. Keep plain coroutineScope for transactions where all-or-nothing is correct.
×

Writing non-cooperative coroutines that ignore cancellation

Symptom
Cancelling the job does nothing — the loop spins to completion anyway, wasting CPU and delaying the replacement work.
Fix
Call ensureActive() in loops and between blocking segments, use cancellable suspending functions (delay, withContext), and never block threads with Thread.sleep inside coroutines.
×

Forgetting to cancel long-running jobs when the owner dies

Symptom
Duplicate network calls pile up on every rotation, and the app gets slower the longer the user navigates.
Fix
Cache the Job from launch, cancel it in onCleared/onDestroy, and guard re-entry (if (job?.isActive == true) return). Structured concurrency handles the rest when scopes are right.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
What is JobCancellationException and when is it normal?
Q02SENIOR
Explain structured concurrency and how it prevents scope leaks.
Q03SENIOR
How do viewModelScope and lifecycleScope fix most Android leaks?
Q04SENIOR
Why does cancelling a CPU-heavy loop sometimes do nothing?
Q05SENIOR
Design cancellation for a sync that fetches, merges, and commits.
Q01 of 05JUNIOR

What is JobCancellationException and when is it normal?

ANSWER
JobCancellationException extends CancellationException and is thrown to unwind a coroutine when its Job is cancelled. It's normal control flow — scopes cancel children on teardown (screen exit, ViewModel cleared). It becomes a bug when swallowed by catch-all handlers, which keeps dead work running, or when unexpected because a scope was misconfigured.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Is JobCancellationException always a bug?
02
What's the difference between a regular Job and a SupervisorJob?
03
Why does code after my cancelled network call never run?
04
When should I use NonCancellable?
05
Does cancelling a collector stop the underlying Flow?
06
How does the coroutine debugger in Android Studio help?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Kotlin. Mark it forged?

5 min read · try the examples if you haven't

←
Previous
Kotlin lateinit Property Has Not Been Initialized
3 / 5 · Kotlin
Next
Kotlin suspend Function Called From Non-Coroutine Context
→