Kotlin CancellationException: Scopes Leaking After Exit
Rethrow CancellationException and scope to viewModelScope: swallowed cancellation leaks work, and GlobalScope outlives every screen..
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
- ✓Basic Kotlin coroutines: launch, async, suspend functions
- ✓An Android app using ViewModel or Lifecycle components
- ✓Familiarity with logcat and basic unit testing
- Coroutines cancel by throwing CancellationException into the job, and JobCancellationException is its normal shutdown signal
- Swallowing that exception in catch-all blocks keeps dead work alive, draining battery and data in the background
- Scope every coroutine to viewModelScope or lifecycleScope so teardown cancels work automatically
- Use SupervisorJob or supervisorScope when one child's failure shouldn't kill its siblings
- You'll fix most issues by rethrowing cancellation, scoping by owner, and adding ensureActive() to loops
Picture a restaurant kitchen where every order belongs to a table. When a table leaves, the chef shouts 'stop their dishes' and cooks drop that work instantly — that's a scope cancelling. Bugs happen when a cook covers their ears and keeps cooking the cancelled dish (swallowed CancellationException), or takes orders from nobody and grills after closing (GlobalScope leak). Fix both by letting cancellations through and attaching every cook to a table.
JobCancellationException is the exception you're supposed to see — until it shows up where it shouldn't, or never shows up because something swallowed it. Coroutines cancel by throwing: when a scope dies, every child gets a CancellationException that unwinds its stack. Navigate away from a screen and viewModelScope cancels its children; that's the system working. The trouble starts when a catch-all block eats that exception and the 'cancelled' work keeps running, or when a GlobalScope launch escapes every scope and leaks for the rest of the process.
In production these bugs wear two masks. The loud one is the crash: a leaked coroutine touches a destroyed view and throws IllegalStateException minutes after the user left. The quiet one is the drain: cancelled downloads keep burning data and battery, duplicate jobs pile up on every rotation, and one failed thumbnail cancels an entire screen because the scope type was wrong. Both trace back to scope discipline, not coroutine magic.
This article gives you that discipline. You'll learn structured concurrency in practical terms, when viewModelScope and lifecycleScope do the work for you, how SupervisorJob isolates failures, and how ensureActive makes loops cancellable. You'll leave able to read any coroutine crash and know which scope broke its promise.
Structured Concurrency and What Cancellation Really Means
Structured concurrency is a simple contract: every coroutine is a child of some scope, children die with their parent, and failures travel upward unless you say otherwise. When viewModelScope is cancelled in onCleared, each child coroutine gets a CancellationException thrown at its next suspension point, its finally blocks run, and the job ends. No orphans, no strays, no work billed to a screen that's gone. The hierarchy is the entire memory-safety story for coroutines.
JobCancellationException is the concrete type you'll meet — a subclass of CancellationException (which extends IllegalStateException, confusingly) thrown when a Job is cancelled. Treat its appearance during teardown as a success log, not a crash: it proves the scope did its job. The danger is purely in mishandling. Catch it in a generic handler and show an error toast, and users see phantom failures on every navigation. Swallow it into a retry loop, and dead work runs forever.
The rule that prevents both: catch CancellationException first and rethrow it, always, before handling real failures. Write that ordering so consistently it becomes muscle memory, and review any catch (e: Exception) in suspend code as a suspected cancellation bug until proven otherwise. Your future self, reading a 2 AM crash log, will thank you for the discipline.
viewModelScope, lifecycleScope, and Escaping GlobalScope
viewModelScope and lifecycleScope exist so you rarely manage jobs by hand. viewModelScope lives as long as the ViewModel and cancels in onCleared — perfect for repositories, flows, and anything that should survive rotation but die with the screen's data owner. lifecycleScope ties to a LifecycleOwner and cancels on destroy; use the viewLifecycleOwner variant in fragments so collection stops when the view dies, not when the fragment instance does. Together they cover nearly all Android work.
The anti-pattern is anything that escapes them. GlobalScope launches live for the whole process and hold every reference they capture — views, contexts, callbacks — until the work finishes or the app dies. Hand-rolled CoroutineScopes are fine only when you manage their lifecycle explicitly: store them, cancel in the matching teardown, and never create one per click without cancelling the last. Audit for GlobalScope with a grep; each hit is a leak report waiting for a user.
One subtlety: withContext doesn't change your scope, it only switches dispatchers — cancellation still flows through. That's good: a withContext(Dispatchers.IO) network call inside viewModelScope still cancels with the ViewModel. But launching a new coroutine inside withContext creates a child of the same scope, not of the call — keep track of parentage when nesting, and prefer structured builders (coroutineScope, supervisorScope) over bare launch for dependent work.
SupervisorJob and supervisorScope: Isolating Failures
Default scopes propagate failure upward: one child throws, the parent cancels, siblings die. That's correct for transactions — a checkout whose payment step failed shouldn't still ship the order. It's wrong for independent work — a dashboard whose news tile failed should still show stats and alerts. SupervisorJob and supervisorScope flip the rule: children fail alone, siblings continue, and the parent only fails if it throws directly.
Pick the semantics per screen, not per app. Feeds, dashboards, and prefetch batches want supervision because each item stands alone. Multi-step writes, migrations, and anything with commit/rollback wants plain coroutineScope so partial completion is impossible. Mixing them is fine and common: a supervisorScope for the independent fetches feeding into a coroutineScope commit phase, each chosen deliberately and commented.
Handle child results explicitly under supervision, since exceptions no longer cancel the batch — they arrive at await(). Wrap each await in its own try/catch (rethrowing CancellationException first, always) and decide per child: default value, skip, or aggregated error state. Unhandled child exceptions under a supervisor still crash if never awaited, so never fire-and-forget an async you intend to supervise — collect every deferred.
Cooperative Cancellation With ensureActive and finally
Cancellation is cooperative — it lands only where the coroutine lets it. Suspending functions like delay, withContext, and retrofit awaits check for cancellation automatically, so typical network code cancels cleanly. Pure computation doesn't: a tight loop parsing ten thousand records never suspends, never checks, and sails past cancel() as if nothing happened. The job claims cancelled while the CPU burns to completion, and the replacement work queues behind it.
ensureActive() is the checkpoint you place yourself. It throws CancellationException immediately if the job was cancelled, costs nothing when it wasn't, and belongs at the top of every loop body plus between heavy phases. yield() does the same while also giving other coroutines a scheduling turn — useful in CPU-bound batches on limited dispatcher threads. Replace Thread.sleep with delay everywhere inside coroutines; the former blocks a thread cancellation can't reach, the latter suspends cooperatively.
Cleanup during cancellation needs finally, which runs as the stack unwinds. Close files, release locks, and discard partial state there — but remember the coroutine is still cancelled inside finally, so any suspending cleanup must use withContext(NonCancellable). Keep that block minimal and fast: it's a graceful exit, not a second chance to finish the work. Anything long-lived there converts cancellation back into the delay you were trying to eliminate.
Proving It: Tests, Debugger, and CI Gates
Testing cancellation is straightforward and almost nobody does it — which is why these bugs reach production. The pattern: launch the work in a test scope, advance or delay briefly, cancel the job, then assert the aftermath. Assert no further network calls occur (verify your fake received nothing more), assert partial state was cleaned up, and assert siblings under supervision still completed. kotlinx-coroutines-test with StandardTestDispatcher gives you deterministic control without real waiting.
The coroutine debugger covers the production side. Run the app, reproduce the leak, and inspect live coroutines: each shows its name, state, stack, and creation trace. Filter to your feature's scope and look for jobs alive past their owner's destruction. The creation trace names the exact launch site — no guessing, no logging archaeology. Make this check part of feature QA for anything long-running: start it, leave the screen, confirm it dies.
For CI, add two permanent gates. First, a GlobalScope ban — a grep or lint check that fails the build on unstructured launches. Second, cancellation tests for every long-running job: sync loops, downloads, collectors. These tests run in milliseconds with virtual time and catch every regression where someone wraps your careful code in a catch-all or moves it to a wider scope. Cancellation you don't test is cancellation you'll lose.
Keeping Scopes Honest as the App Grows
Long-term health comes from making the right scope the easiest scope. Provide base ViewModels or use-case helpers that expose the correct scope so feature code never chooses. Document the team's two scope types — supervised for independent tiles, plain for transactions — with one example each in the repo wiki. New engineers copy the example; the examples encode the policy, and the policy stops being tribal knowledge.
Review with a checklist, not vibes. Every suspend function with try/catch: does it rethrow CancellationException first? Every launch: which scope owns it, and when does that scope die? Every loop over unbounded work: where's the ensureActive checkpoint? Every async: is it awaited, and under which failure semantics? Five questions, thirty seconds each, and the review catches what tests can't — intent mismatches where the code works but means the wrong thing.
Finally, watch the signals that predict the next incident. Rising background battery in vitals, server request growth without user growth, and IllegalStateException crashes naming destroyed views all point at scope discipline slipping. Treat any of them as a trigger for a coroutine audit: dump the debugger, grep GlobalScope and bare catch blocks, re-run cancellation tests. You'll catch the next leak while it's a metric wiggle, not a weekend incident.
A Swallowed Cancellation Retried Sync Forever and Tripled Server Load
- A catch-all around suspend work must rethrow CancellationException first — swallowing it turns every cancel into a leak.
- Client-side leaks look like server-side load; check cancellation behavior before scaling the backend.
- Rotation without scoped jobs multiplies work: every config change spawned a sync that never died.
| File | Command / Code | Purpose |
|---|---|---|
| FeedViewModel.kt | class FeedViewModel(private val repo: FeedRepo) : ViewModel() { | Structured Concurrency and What Cancellation Really Means |
| Scopes.kt | class DetailViewModel(repo: DetailRepo, id: String) : ViewModel() { | viewModelScope, lifecycleScope, and Escaping GlobalScope |
| Dashboard.kt | fun ViewModel.loadDashboard(repo: DashRepo) = viewModelScope.launch { | SupervisorJob and supervisorScope |
| Sync.kt | suspend fun syncAll(items: List<Record>, api: SyncApi) { | Cooperative Cancellation With ensureActive and finally |
Key takeaways
Common mistakes to avoid
5 patternsSwallowing CancellationException in a generic catch block
Launching coroutines in GlobalScope or a hand-rolled scope
Using coroutineScope where one failing child shouldn't kill siblings
Writing non-cooperative coroutines that ignore cancellation
Forgetting to cancel long-running jobs when the owner dies
Interview Questions on This Topic
What is JobCancellationException and when is it normal?
Frequently Asked Questions
20+ years shipping production backend systems. Lessons pulled from things that broke in production.
That's Kotlin. Mark it forged?
5 min read · try the examples if you haven't