Rust Error Handling: anyhow vs thiserror Done Right
Use anyhow in apps for context and bail, thiserror in libraries for typed errors.
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓Comfortable with Rust ownership, Result, and Option basics
- ✓Built a small Cargo binary with external dependencies
- ✓Familiarity with traits, derives, and the From conversion trait
- The ? operator propagates errors by early-returning Err with an automatic From conversion, so mixed error types unify into one function-level error without manual mapping
- Option chains replace unwrap with combinators: map for total transforms, and_then for fallible lookups, unwrap_or_else for lazy defaults, and ok_or_else to enter ? pipelines
- Use anyhow in applications: anyhow::Result plus with_context, bail, and ensure give every failure a causal chain with zero conversion boilerplate across the whole binary
- Use thiserror in libraries: derive typed error enums with #[error] messages, #[from] conversions, and #[source] linkage so downstream users can match reliably on failure modes
- Reserve panic, unwrap, and expect for violated invariants that mean programmer error, never for runtime input; enforce the policy with deny-level clippy lints in CI
- Keep anyhow out of public APIs and convert library errors into app context at the binary boundary with .context(); print production errors with {:#} under RUST_BACKTRACE=1
Think of a restaurant kitchen during dinner rush. Every station — grill, sauces, pastry — can hit a problem: missing ingredients, a broken burner, a dropped tray. A good kitchen doesn't make the waiter sprint back to each station to ask what went wrong; each station reports failures up a chain with a note attached saying where it happened and what was cooking. Rust errors work the same way. The ? operator is the runner carrying the failure upstairs, anyhow is the incident report with full context attached for the manager, and thiserror is the standardized form each station fills out so reports stay consistent. Nobody panics and shuts down the restaurant over one burnt steak — unless the gas main is leaking, which is the only time panic is the right call.
You've seen the two failure modes in every Rust codebase that grew past a prototype. One team unwraps everywhere and ships a binary that panics on the first malformed config in production. The other team defines fourteen error enums by hand, maps each one manually, and drowns the actual logic in conversion boilerplate nobody reviews.
Both teams are paying for the same misunderstanding. Rust doesn't have exceptions, so errors are values — and values need a transport strategy. Without one, your codebase drifts toward panic-driven ops or boilerplate paralysis, and both hurt at 3 AM.
The ecosystem already settled this debate. Applications use anyhow for ergonomic context-rich errors, libraries use thiserror for typed public error contracts, and the ? operator plus Option combinators connect everything in between.
But the boundary between the two crates is where seniors earn their keep. Leak anyhow into a library API and downstream users can't match on your failures. Hand-roll conversions in an app and you'll waste days writing From impls anyhow derives for free.
When you finish this guide, you'll propagate with ?, chain Options without a single unwrap, pick anyhow or thiserror correctly every time, and enforce a panic policy your on-call rotation will thank you for.
The ? Operator: Propagation, Conversion, and Early Return
The question-mark operator is three operations wearing a trench coat, and seeing all three ends the mystery around it. Applied to Result<T, E> inside a function returning Result<U, F>, it unwraps Ok values inline, converts Err values with From::from into the function's error type, and early-returns the converted error — all in one keystroke. The same shape works for Option inside Option-returning functions. That conversion step is the load-bearing one: ? only compiles when the error types connect through From, which is exactly why library authors derive #[from] variants and app authors standardize on anyhow::Result.
The desugar makes error-type mismatches readable instead of magical. expr? expands to a match that returns value on Ok and executes return Err(From::from(err)) on Err — the return keyword matters because it exits your function, not just the closure or block. Juniors get burned placing ? inside map closures or main functions returning () where no compatible return type exists; the compiler's E0277 then names the missing From impl precisely. Reading that error as a wiring diagram — found type, required type, missing bridge — turns a confusing failure into a five-second fix.
Main functions and tests deserve explicit treatment because their return types set the error ceiling.fn main() -> anyhow::Result<()> lets every ? in main propagate with context instead of ceremony, and test fns returning Result<(), E> fail gracefully on Err rather than panicking — with anyhow or a custom Debug error type. The old habit of fn main() with bare unwraps survives in tutorials but dies fast in production review, because a staging panic with no context is how you spend a Friday evening.
The discipline around ? is about where you add context, not whether you propagate. Bare ? at every layer produces flat chains that name the root cause but not the journey — file opened, config parsed, connection dialed. Seniors place .with_context() at trust boundaries (filesystem, network, subprocess, parsing user input) and bare ? for internal plumbing between already-contextualized frames. Three to five context frames per production error is the healthy range; one frame means you're under-contextualizing, twelve means you're wrapping internal calls that add no information.
The Try-trait machinery underneath explains why Option and Result don't mix freely. The ? operator works in any function implementing the Try protocol — Result-returning functions for Result values, Option-returning functions for Option values — but crossing the streams needs an explicit bridge. Using ? on an Option inside a Result-returning function fails to compile, full stop; the conversion is .ok_or(Error)? or .context(...)? first, then ?. This strictness is deliberate: silent None-to-Err conversions would hide which absence mattered, while the explicit bridge names the error at the exact line it enters the Result world.
Closures get their own ? rules because each closure is a separate function for return-type purposes. A map closure returning Result can use ? internally against its own return type, which is how fallible transforms compose inside adapters — .map(|s| s.parse::<u64>().map_err(...)?) works when the closure returns Result. Immediately-invoked closures extend this to blocks: let cfg = (|| -> Result<Config, LoadError> { ... })()? runs a multi-step fallible computation inline and propagates outward. Between Try-protocol awareness, explicit Option bridging, and closure-scoped propagation, the operator stays predictable in every position it appears.
Boxed-error transport follows the same conversion logic with dynamic dispatch. Functions returning Result<T, Box<dyn std::error::Error>> accept ? from any 'static error type automatically, which makes boxed errors the pragmatic default for examples, tests, and dependency-free tools. The cost is one allocation per error plus erased matchability — acceptable where errors are rare and consumers are human. Where consumers are code, the typed enum earns its keep; where the reader is an operator scanning logs, the box suffices. Choose by consumer, not by habit.
Option Combinators: map, and_then, unwrap_or, and ok_or
Option handling has a mechanical ladder from worst to best, and code review should enforce climbing it. The bottom rung is match with two arms for a simple transform — six lines for one idea. Above it sits if let, fine for side effects but awkward for value pipelines. The top rung is combinators: map for transforming the inner value, and_then for chaining fallible lookups that return Option, unwrap_or and unwrap_or_else for defaults, ok_or and ok_or_else for converting absence into a typed error. Each rung up deletes branches and makes the happy path linear.
Map versus and_then is the distinction that unlocks everything else. Map takes T to U and wraps automatically — opt.map(|s| s.len()) turns Option<String> into Option<usize>. And_then takes T to Option<U> without double-wrapping — opt.and_then(|s| s.parse::<u64>().ok()) chains lookups that can each fail. Using map where and_then belongs yields Option<Option<T>> nesting that juniors flatten with unwrap; using and_then where map belongs overcomplicates a total function. The one-line test is whether your closure can fail: fallible means and_then, total means map.
Unwrap_or versus unwrap_or_else versus unwrap_or_default looks like trivia until a hot path proves otherwise. Unwrap_or evaluates its default eagerly, so .unwrap_or(expensive_fallback()) pays the cost even when the Option is Some — a real latency line-item in request handlers. Unwrap_or_else takes a closure evaluated lazily only on None. Unwrap_or_default uses Default::default for zero-cost empties like Vec::new or 0. The same laziness split governs ok_or (eager) versus ok_or_else (lazy) when converting to Result, and expect-message quality governs the rare cases where panicking accessors survive review.
Ok_or is the bridge from Option-land to Result-land and therefore the bridge into ? pipelines. Environment lookups, HashMap gets, and first-match searches all return Option; appending .ok_or(Error::MissingKey)? folds them into the function's error flow in one breath. Combined with filter, copied, and cloned adapters on Option::iter, these combinators handle defaults, fallbacks, and validation chains with zero matches and zero unwraps. When a reviewer spots a match on Option doing a pure transform, the comment writes itself: rewrite with map or and_then.
Three more combinators round out daily use. Transpose flips Option<Result<T, E>> into Result<Option<T>, E>, which routes the first parse failure out through ? while keeping clean absence as Ok(None) — the exact shape of optional config values that must still validate when present. Flatten collapses Option<Option<T>> from nested lookups, and zip pairs two Options into one, yielding Some only when both sides exist — parallel optional lookups without nested matches. Filter keeps the value only when a predicate holds, turning guard clauses into chain links.
Map_or versus map_or_else replays the eager-lazy lesson from defaults. Map_or computes its fallback eagerly on every call, map_or_else defers to a closure on None only — same latency trap as unwrap_or in hot paths, same one-word fix. Is_some_and (stabilized for Option) replaces the .map(...).unwrap_or(false) dance with a single predicate check that short-circuits cleanly. Between transpose for validation, zip for pairing, and lazy fallbacks for defaults, Option pipelines stay branch-free from lookup to use — and every branch you delete is a path coverage obligation that vanishes with it.
Reference adapters keep borrowed pipelines zero-copy. As_ref converts &Option<String> to Option<&str> so chains borrow instead of cloning, as_deref goes further into Option<&str> from smart-pointer payloads, and copied/cloned lift Copy and Clone values out only at the final step. The pattern is borrow through the chain, materialize at the boundary: lookups, validation, and matching all run on references, with a single to_owned where ownership truly begins. Chains that clone at every stage allocate geometrically; chains that borrow allocate once — the flame graph difference is unmistakable past a thousand requests per second.
compute_default()) where the default opened a config file and parsed it — 4 ms of filesystem IO on every request even when the env var was set, adding 4 ms to p50 across 12K RPS. Switching to unwrap_or_else moved the file read to the None path only and p50 dropped 3.8 ms overnight. Rule: eager defaults in hot paths are silent latency; lazy closures are free insurance.anyhow for Applications: context, bail, and ensure
Anyhow exists because application code has different error needs than libraries: one binary, many failure sources, zero downstream matchers. Its core type erases the concrete error behind a stable anyhow::Error handle while preserving the full causal chain, so reqwest failures, IO failures, and parsing failures all flow through a single anyhow::Result<T> alias. Functions stop declaring five-variant error enums and start declaring what they return, and the ? operator converts any std::error::Error automatically. The boilerplate savings are immediate — teams typically delete hundreds of conversion lines in the migration week.
Context is the feature that justifies the crate. The Context trait adds .context() for static messages and .with_context() for lazily computed ones to any Result, attaching breadcrumbs that render as an ordered chain: failed to bind 0.0.0.0:8080, caused by address in use. Without context, production logs show only the root cause — connection refused — with no record of which of nine connection sites failed. The convention that works is context at every trust boundary (filesystem paths, URLs, subprocess names, config keys) with the dynamic values operators need: format!("reading {}", path) beats a static string every time.
Bail and ensure handle the early-return shapes that ? can't express. bail!("unsupported scheme {s}") returns Err immediately with a formatted message, replacing return Err(anyhow!(...)) ceremony. ensure!(port > 0, "port must be nonzero") checks invariants in one line, replacing three-line if-not-return blocks. Both read as intent rather than plumbing, and both keep validation preambles — the first forty lines of many handlers — compact enough that reviewers actually read them.
Debug formatting completes the story: {:#} renders the alternate chain view with each context on its own line, which is what you print in logs, while {:?} includes the backtrace when captured. Downcasting with downcast_ref recovers concrete types for the rare test or retry branch that needs them, but reaching for downcast in business logic is a smell — it means a library boundary wants thiserror instead. Anyhow owns the binary interior; typed contracts own the edges.
Rendering choices decide what operators actually see. Display ({}) prints the outermost message only — right for user-facing CLI output where chains confuse. Alternate Debug ({:#}) prints each context on its own line — right for logs where the journey matters. Full Debug ({:?}) adds the captured backtrace — right for triage artifacts attached to incident tickets. Standardizing per sink (CLI gets {}, logs get {:#}, tickets get {:?}) removes the recurring argument about noisy versus thin errors, because each audience receives exactly the depth it needs.
The anyhow! macro and Context-on-Option complete daily fluency. anyhow!("port {p} out of range") builds ad-hoc errors with format syntax for validation sites that need no type. Context works on Option too — opt.context("database url missing")? converts None into an anyhow error with your message, unifying absence and failure into one chain. Ensure! covers invariants, bail! covers early exits, context covers boundaries: three macros plus one trait method handle every app-side error shape without a single custom type. Option-to-anyhow bridges (opt.context("config value missing")?) deserve the same boundary treatment as Results: name the absent key, not just the absence.
Ensure versus assert encodes a release-behavior contract reviewers must understand. Assert panics unconditionally and stays active in release unless explicitly disabled — appropriate for invariants, dangerous for input validation that should return errors. Ensure returns Err through normal channels, keeping validation failures recoverable and testable. The one-line migration of input asserts to ensure calls converts crash sites into handled errors, and clippy's panic-group lints flag the asserts hiding in request paths. Invariants assert, inputs ensure — no exceptions in reviewed code.
thiserror for Libraries: Typed Contracts Downstream Can Match
Libraries answer to a different master than applications: downstream code that must react differently to different failures. A storage crate's users retry on timeouts, abort on corruption, and refresh credentials on auth expiry — which requires matchable variants, stable across versions, carrying the data each arm needs. Thiserror generates those contracts from annotated enums: #[derive(Error)] plus #[error("...")] Display strings per variant, #[from] for automatic foreign-error conversions, and #[source] or transparent passthrough preserving causal chains. The derive output is exactly the hundred-line hand implementation nobody wants to write or review.
The attribute vocabulary is small and each item earns its place. #[error("failed to read {path}")] generates Display with inline field interpolation, keeping messages beside the variants they describe. #[from] on a field of type io::Error generates the From impl that makes ? convert automatically at every call site — the wiring diagram from section one, solved declaratively. #[source] marks the underlying cause when the field name alone doesn't imply it, and #[error(transparent)] delegates Display and source entirely to a wrapped error for newtype passthroughs. Struct variants carry context data (path: PathBuf, retries: u8) that match arms and logs both consume.
Designing the variant list is the actual senior work, because variants are a public API with semver weight. One variant per actionable failure mode, named for the condition not the layer — Timeout and CorruptFrame rather than InnerError and OtherError. Group unrecoverable foreign errors behind #[from] variants; expose fields users match on (which key, how many retries) and hide the rest. Adding a variant is a minor release, removing or reshaping one is major — so start narrow with a hidden escape hatch only if you must, and document which variants are stable.
The anyhow boundary rule is absolute: libraries export thiserror types, binaries consume them into anyhow with ?. A library returning anyhow::Error forces every downstream user into string matching and downcasting, which breaks across versions and can't be documented. Conversion flows one way — library error into app context via .context() — and reviewers should reject any public signature exposing anyhow. Crate-level enforcement is a public-API test asserting exported error types implement std::error::Error without type erasure.
Transparent passthrough handles the newtype case without ceremony. A wrapper like struct DbError(#[from] sqlx::Error) with #[error(transparent)] delegates Display and source entirely to the inner error — useful when your crate adds namespacing without adding information. Backtrace capture integrates through the standard Backtrace type: a #[backtrace] field records construction-site frames automatically, giving library errors first-class traces without depending on anyhow. These two attributes cover the long tail of enum design that hand implementations always fumbled.
Semver discipline turns variant design from taste into process. Mark the enum #[non_exhaustive] when downstream matches must not be exhaustive — adding variants then stays non-breaking even for external matchers, at the cost of requiring a wildcard arm. Document which variants are stable and which are provisional; group experimental failures behind a single Unstable(String) variant rather than proliferating public commitments. Review every new variant with the question callers will ask: what do I do differently when I see this? No distinct action means no distinct variant — fold it into an existing one with a data field.
Field interpolation in #[error] strings keeps messages beside the data they describe. Named fields render inline — #[error("connection to {host} timed out after {ms}ms")] — so message reviews happen at the variant definition, not in a distant Display impl. Unit variants suit flag-like failures, struct variants carry context, and tuple variants wrap foreign errors with #[from]. Deriving PartialEq alongside Error (where all fields allow it) unlocks assert_eq on errors in tests, turning failure-mode assertions from matches! gymnastics into direct equality — a small ergonomic win that compounds across hundreds of error-path tests.
Custom Error Enums by Hand: What the Derive Does for You
Every senior should hand-write one error enum in their career — not for production, but because the exercise makes the derive's value concrete and the trait obligations explicit. A manual implementation needs four pieces: the enum itself with context-carrying variants, Display mapping each variant to a human message, std::error::Error with source() returning the underlying cause per variant, and From impls for each foreign error the ? operator must convert. That's roughly sixty lines for three variants, all of it mechanical, all of it review surface for typos in messages nobody tests.
Display is where hand implementations quietly rot. Each arm formats its fields, and message quality depends on whoever wrote that arm at midnight — some variants name the key, others don't, and consistency drifts with every contributor. Source linkage rots faster: returning Some(&cause) requires borrowing through the match correctly, and skipped source() impls sever causal chains that debugging depends on. Thiserror's attributes fix both by construction — the message template sits on the variant, source follows from field types and annotations, and every variant gets identical treatment regardless of author or hour.
From impls are the highest-churn piece and the strongest argument for derivation. Each foreign error source needs its own impl mapping into the right variant, and every new dependency version or call site adds another. Hand-written impls also invite the from-string antipattern — mapping everything into a catch-all String variant that destroys matchability and source chains in one move. The derive's #[from] generates these impls mechanically and keeps the typed variants intact, which is why the style guide for libraries is one line: derive, don't implement.
Keep the hand-rolled skill for interviews, code archaeology, and no-dependency crates where adding thiserror genuinely costs more than sixty lines — embedded targets with vendoring constraints, or single-file tools. Everywhere else, the derive wins on consistency, review cost, and semver hygiene. When you inherit a hand-rolled enum, the migration is mechanical: annotate variants, delete the manual impls, and add a test per variant asserting Display output so the messages survive the translation byte-identical.
Object safety and trait bounds complete the mental model behind the boilerplate. Std errors must implement Debug plus Display — Debug for the {:?} triage rendering, Display for human messages — and Send plus Sync when they cross thread boundaries, which production errors always do. A compile-time assertion fn assert_error<T: std::error::Error + Send + Sync>() {} instantiated on your enum pins these bounds so a future variant with an Rc field fails the build instead of failing async callers. Box<dyn Error> transport requires 'static, which rules out borrowed payloads — errors own their data, another reason variants carry Strings and PathBufs rather than &str references.
Error::provide and the request-value API extend source chains into typed context, letting consumers query structured data (retry hints, spans) without downcasting — advanced machinery most teams never need, but worth knowing exists before you invent it. The practical takeaway stays simple: derive the four obligations, pin the thread-safety bounds with a test, and keep every payload owned. Hand-rolling taught you what the machine does; the derive guarantees it does it identically on every variant you will ever add.
Foreign-error mapping strategy decides whether causes survive translation. One distinct variant per foreign source (Io(#[from] io::Error), Parse(#[from] ParseIntError)) preserves identity, enables targeted matching, and keeps source chains intact. Collapsing multiple sources into a single External(String) variant destroys all three — matchability, identity, and linkage — for the price of one fewer enum arm. The savings are illusory and the debugging cost is real: every collapsed variant becomes a formatting site that future triage cannot see past. Distinct variants per source, always, with tests asserting each conversion preserves its cause.
source(), and two mapped distinct IO failures into one String catch-all. Support tickets citing those messages took 2.3x longer to resolve because operators couldn't tell which key or disk had failed. Deriving with thiserror and adding per-variant Display tests standardized all 22 messages in one sprint. Rule: untested Display arms rot; pin them with assertions.panic, unwrap, and expect: A Policy Your On-Call Will Thank You For
Panic handling starts from a blunt mechanical fact: panicking aborts the current thread, runs destructors during unwinding, and in a binary without a catch boundary takes down the process. In libraries it poisons mutexes and crashes hosts that never agreed to die; in servers it drops every in-flight request on that thread. The ? operator and combinators exist precisely so runtime failures — bad input, missing files, refused connections — never need this path. Panic is reserved for one category: violated invariants where continuing would corrupt data or lie to the caller, meaning programmer error rather than environmental failure.
The policy that survives contact with production has three tiers. Tier one is deny: application and library crates enable #![deny(clippy::unwrap_used, clippy::expect_used)] so panicking accessors fail the build, with targeted #[allow] only where a proof of infallibility sits in a comment beside it. Tier two is expect with evidence: where infallibility is provable — a regex compiled from a literal, a lock unpoisoned by construction — expect carries a message naming the invariant, not the hope: .expect("startup regex literals always compile"). Tier three is explicit panic! for unreachable states, carrying the values that prove reachability was assumed: panic!("negative retry count after clamp: {n}").
Unwrap in tests is the sanctioned exception, and even it has limits. Test code panicking on setup failure is correct — a test that can't arrange its fixtures has nothing to assert. But assertions about production error paths must exercise the Err values, not unwrap past them: assert!(matches!(...)) on thiserror variants, downcast checks on anyhow chains. Teams that unwrap through error-path tests discover their coverage was theatrical the week a real failure arrives untested.
Debug versus release behavior sharpens the argument. Debug builds overflow-check, panic on arithmetic wrap in some configurations, and run slower — surfacing invariant violations early. Release builds strip those checks for speed, meaning an invariant you relied on debug to catch silently wraps in production. The policy implication is direct: never use runtime checks as substitutes for validation, and never ship expect where user input flows. Validate at the boundary with Result, assert invariants inside with documented expects, and let the deny lints prove you did.
Unwind boundaries matter where Rust meets foreign code. Panics cannot cross FFI frames — a panic escaping into C is undefined behavior — so every extern fn boundary needs catch_unwind translating panics into error codes, with AssertUnwindSafe documented where the closure's safety case holds. Plugin systems and WASM hosts impose the same requirement: host boundaries catch, guest code never panics across. Auditing these perimeters with ripgrep for extern fn blocks missing catch_unwind belongs in the same quarterly pass as the swallowed-error audit.
Build profiles tune the panic machinery itself. Debug profiles keep overflow checks and full unwind tables; release profiles strip checks for speed, and panic = "abort" trades stack unwinding for smaller binaries and immediate core dumps — the right call for embedded and some edge fleets, the wrong call where destructors guard consistency (locks, temp files, transactions). Overflow-checks = true in release catches arithmetic invariants production would otherwise wrap silently. These are Cargo.toml lines with incident-scale consequences, which is why they belong in reviewed config rather than tribal knowledge.
The panic-adjacent macros need their own discipline because they look harmless in isolation. Unreachable! documents states your logic excludes — with the values that prove exclusion in the message — while unimplemented! and todo! mark unfinished work that panics when reached. A todo! surviving into a release build is an incident with a paper trail; grep for it in CI and fail the build on any hit outside explicitly experimental modules. Debug_assert! covers dev-only checks stripped from release — cheap invariant enforcement during testing with zero production cost. Each macro has exactly one legitimate habitat; outside it, they are bugs wearing syntax.
Errors in main and Tests: Return Result Instead of Panicking
Main functions returning Result are the highest-leverage error habit in Rust, because main is where context goes to die in most codebases. fn main() -> anyhow::Result<()> upgrades every ? in the startup path from a panic risk into a propagated, contextualized, printable error — and the runtime Debug-prints the returned Err, chain included, before exiting non-zero. The alternative — fn main() with unwraps scattered through argument parsing, file loading, and client construction — converts each startup dependency into a crash site with a riddle message. Staging environments punish this within the first week.
The startup sequence pattern is worth standardizing across every binary you own: parse args into a Config struct with contexts naming each flag, load files with paths in every message, build clients with endpoint URLs attached, then run. Each stage gets .context() describing the stage in operator language — "loading server config", "connecting to postgres at {url}" — so the 3 AM reader sees a story, not a stack of conversions. Exit codes deserve one line of thought: Err from main exits 1, which suffices for most services; CLIs needing distinct codes match on a thiserror-typed core error before printing.
Tests returning Result fix the awkwardness of testing fallible code. A #[test] fn returning Result<(), anyhow::Error> lets the body use ? throughout, failing the test gracefully with the error value instead of panicking mid-setup. This shines in integration tests that spin up fixtures, write temp files, and query test doubles — five fallible steps that otherwise nest in expects. The limitation is real but narrow: returned Err must implement Debug, which anyhow::Error and well-formed custom errors satisfy, and async test runtimes support Result returns identically.
Assertion strategy completes the picture. Happy paths assert values; error paths assert shapes — assert!(matches!(err, ConfigError::MissingKey(_))) for typed cores, snapshot assertions on anyhow chain strings for binary-level behavior. Property tests generating malformed inputs catch the parse-and-validate gaps that hand-written cases miss, and each regression input from an incident becomes a permanent fixture. The config outage that opens this article would have been a fifteen-line test returning Result; that asymmetry between prevention cost and incident cost is the whole argument.
The Termination trait governs what main may return and how exit codes propagate. Beyond () and Result, custom types implementing Termination control process exit codes directly — the mechanism behind CLIs distinguishing usage errors (exit 2) from runtime failures (exit 1) without calling process::exit, which skips destructors and poisons drop guarantees. The cleaner pattern extracts run() -> anyhow::Result<ExitCode> holding all logic, leaving main as a three-line reporter that prints {:#} chains on failure. Testing targets run() instead of main, keeping exit-code assertions in-process and hermetic.
Report rendering is the last mile most teams ignore until users complain. Raw Debug chains serve operators; end users deserve miette-style reports with snippets, highlights, and fix suggestions — a conversion applied at the CLI boundary, never in library code. Colored output belongs behind a tty check so piped logs stay parseable. Between Termination-aware mains, extracted run() cores, and audience-appropriate rendering, the binary edge becomes as carefully designed as the error types behind it — which is exactly the standard production incidents grade you on.
Doctests need special handling because the harness wraps examples in fn main() -> (). The ? operator inside a doctest therefore fails to compile unless the example defines its own Result-returning function and calls it — a two-line wrapper most authors discover through the compiler error. Explicitly typed examples (asserting Ok values and matches! on Err variants) double as executable documentation that CI verifies on every build. Teams that doctest their error constructors catch message regressions the same day they land, because the documentation is the test suite.
Backtraces and RUST_BACKTRACE: From Riddle to Root Cause
Backtraces answer the question context cannot: not what failed, but which code path got there. A contextualized anyhow chain says loading server config: reading /etc/svc.toml: connection refused — which names the journey through layers. The backtrace names the exact frames: main at main.rs:42, load at config.rs:118, dial at net.rs:77. Together they compress incident triage from tens of minutes to single digits, which is why backtrace configuration belongs in the production image definition rather than in anyone's memory.
RUST_BACKTRACE is the runtime switch with three positions you must know cold. Unset or 0 captures nothing — panics print messages without frames, anyhow errors carry no trace. RUST_BACKTRACE=1 captures the trace on panic and, with anyhow's backtrace feature enabled, on error construction — the setting every production deployment wants. Full captures all frames including runtime internals, useful for compiler and FFI debugging but noisy for app triage. The operational default is 1 in staging and production images, full available on demand for the weird ones.
Anyhow's relationship with backtraces has one wrinkle: capturing requires the backtrace feature and adds construction cost per error, which is negligible on error paths (cold by definition) and irrelevant on happy paths (no error constructed). Errors printed with {:?} include the captured trace; {:#} renders the message chain operators read first. Standardize log lines to print both — chain for humans, trace for triage — and confirm once per service that traces actually reach the aggregator, because container runtimes that swallow stderr past 16 KB will truncate exactly the frames you needed.
The tracing ecosystem extends this from errors to request causality. tracing-error with its SpanTrace attaches the current span context to errors, so a failure inside request 7f3a carries the request ID, route, and user tier through the chain automatically. Combined with anyhow context at boundaries and RUST_BACKTRACE=1 underneath, production failures arrive as self-triaging artifacts: what happened, where in code, inside which request. Building this stack takes an afternoon; every incident after that pays dividends.
The std Backtrace API offers direct control where anyhow's capture isn't available. std::backtrace::Backtrace::capture() snapshots frames at any point — library code can attach one to a thiserror variant's backtrace field, custom reporters can sample at construction, and tests can assert capture status. Short versus full formats trade readability against completeness: short frames skip runtime internals, full includes them for FFI and codegen mysteries. Knowing both exists matters the week a failure hides inside a macro expansion that short mode elides.
Symbol quality decides whether traces name lines or hex addresses. Release profiles strip debug info by default, so [profile.release] debug = true (line tables without full debuginfo weight) belongs in every service manifest — a few percent larger binary for traces that name file:line instead of 0x7f3a. Verify symbolization in staging by triggering one handled error per deploy and reading the trace end to end. Traces without symbols are archaeology; traces with symbols are directions.
Frame filtering keeps traces readable at a glance. RUST_LIB_BACKTRACE=0 hides standard-library and runtime frames, leaving only application frames — the difference between a 60-frame dump and the 8 frames that matter. Full backtraces stay one env var away for the FFI and codegen mysteries that need runtime internals. Standardize the filtered default in every deployment manifest and document the full-trace override in the runbook, so triage starts clean and escalates deliberately rather than drowning in frames from the first page. Confirm the filtered default survives container log pipelines: JSON log shippers that truncate long lines will cut full traces mid-frame, so assert one complete trace per service per deploy in staging before promoting to production.
anyhow vs thiserror: The Boundary Rule and Crossing It Safely
The decision rule fits in one sentence and ends ninety percent of team arguments: libraries expose thiserror, binaries run on anyhow, and conversion flows one way at the boundary. Libraries need matchable contracts because strangers depend on them; applications need ergonomic chains because one team owns the whole binary. Every debate about which crate a piece of code uses resolves by asking who consumes the error: downstream matchers mean thiserror, internal operators mean anyhow. Workspace members that are both — a shared core used by your CLI and imported by others — expose thiserror and let each binary wrap it.
Crossing the boundary safely is a three-line pattern worth memorizing. Library functions return Result<T, LibError>; the binary calls them with .with_context(|| "doing X for {id}")? which converts LibError into anyhow::Error automatically (anyhow implements From for all std errors) while attaching the operational frame. Retry and fallback logic that needs the typed variant matches before conversion — match on the library error, decide, then convert the surviving failures. Converting first and downcasting later works but surrenders the compiler's exhaustiveness checking, which is the entire value of typed errors.
Boxed dyn Error is the std-only middle path for code that can't take dependencies: Result<T, Box<dyn std::error::Error + Send + Sync>> transports any error with ? conversions and dynamic dispatch, at the cost of an allocation per error and no downcasting ergonomics. It's the right choice for examples, embedded-vendored crates, and std-only libraries — and the wrong choice wherever anyhow or thiserror is available, because both dominate it on ergonomics without meaningful overhead. Know it exists, reach for it rarely.
Migration direction matters when inheriting code. Hand-rolled enums migrate to thiserror by annotating variants and deleting manual impls, keeping Display output byte-identical with per-variant tests. Stringly-typed app errors (Result<T, String>) migrate to anyhow by aliasing the return type and adding context at boundaries — an afternoon's work that typically deletes more lines than it adds. anyhow-in-library migrations go the other way: introduce the enum, convert internals, and publish the contract restoration as a semver-noted fix. The comparison table below compresses all of this into the reference your team bookmarks.
The concrete workspace layout ends layout debates permanently. A core crate exposes thiserror types and pure logic; a cli crate depends on core, wraps calls with context, and ships anyhow::Result from main; integration tests assert chains end to end. Re-exports keep the surface tidy — pub use core::{Error, Config} — so consumers depend on one path. This shape scales from side projects to hundred-crate workspaces because the boundary rule is fractal: every lib-to-bin edge converts once, with context, in the same direction.
Versioning follows from the layout without special cases. Adding anyhow context frames never breaks semver — chains are behavior, not API. Adding thiserror variants is minor under #[non_exhaustive], major without it, which is why new libraries start non-exhaustive and stabilize deliberately. Removing Stringly catch-alls counts as a fix, not a break, when the messages survive byte-identical under per-variant tests. Publish the contract, version the contract, test the contract — the error enum is a public API and deserves the full API treatment.
Re-export patterns keep the consumer surface tidy as workspaces grow. The binary crate re-exports core error types (pub use core::LoadError) so downstream code paths reference one canonical location even as internals reorganize. Conversion into anyhow needs no manual From impls — the blanket impl covers all std errors, and .context() attaches frames during the same ? step. For the rare deliberate construction, anyhow::Error::new(typed_err).context("stage") builds chains explicitly. One import, one conversion direction, zero boilerplate: the boundary stays boring, which is precisely the goal.
Silent Killers: Swallowed Errors, Stringly Types, and Lost Sources
The most expensive error bugs aren't type errors — the compiler catches those — but handling bugs that compile perfectly and destroy information at runtime. Swallowed errors top the list: map_err(|_| ()), unwrap_or_default on failures that deserved escalation, and let _ = fallible_call() silencing the must_use warning. Each one converts a diagnosable failure into a mystery — the empty dashboard, the zeroed counter, the request that returned default data nobody questioned. The ripgrep audit is one command — rg -n 'let _ =|unwrap_or_default\(\)|map_err\(\|_\|' — and every hit needs either a log line, a metric, or a justification comment.
Stringly-typed errors are the second killer, and Result<T, String> is their flagship. Strings can't be matched reliably, can't carry structured fields like retry_after or offending_key, can't link sources, and change wording between releases — breaking every downstream string match silently. The migration is mechanical: apps move to anyhow::Error (keeping messages, gaining chains), libraries move to thiserror enums (keeping messages, gaining variants). The intermediate step of defining error structs with message fields preserves information while the team migrates call sites — but String as the error type should never survive review.
Lost source chains are the subtlest of the three because the code looks correct. Mapping foreign errors into flat message variants — BadValue(format!("{e}")) — preserves the text but severs source(), killing backtrace-linked causal walks and breaking tools that traverse chains. The fix is structural: #[source] fields in thiserror variants, context wrappers in anyhow instead of message reformatting, and tests asserting source().is_some() on converted errors. A chain is only as debuggable as its weakest conversion, so audit conversions, not just origins.
The meta-fix for all three is making invisible handling visible. Deny let_underscore_must_use lints where they matter, require error-path tests for every new variant, and add a review checklist item: does this error reach the operator with its cause, its context, and its data intact? Teams that ask that question on every PR stop generating silent killers; teams that don't keep discovering them in quarterly audits with dollar signs attached.
Double logging is the operational twin of swallowing. Returning an error up the stack while also logging it at the origin produces duplicate alerts with diverging context — the origin line lacks request scope, the boundary line lacks internals, and on-call chases two tickets for one failure. The convention is single-point reporting: libraries return, binaries log once at the edge with the full {:#} chain, and middleware (tracing spans, request IDs) enriches rather than duplicates. Audit alert rules for pairs firing on the same incident ID and merge them.
Metrics per variant close the observability loop. Incrementing a counter labeled by error variant at the reporting edge turns failure modes into dashboards — timeout spikes versus corruption trickles become visible without reading a single log line. Anyhow chains need explicit classification points (downcast once at the edge, or classify in the typed core before conversion) because erased types can't label themselves. Between single-point logging, variant metrics, and quarterly discard audits, errors stay visible from occurrence through triage — which is the entire job of an error system, and the standard silent killers fail.
Documentation lints turn error contracts into reviewed artifacts. Clippy's missing_errors_doc requires a documented Errors section on every public function returning Result — forcing authors to state which failures callers should expect before merging. Combined with per-variant Display tests and source-linkage assertions, the error surface gains three independent guards: docs describe it, tests pin it, lints enforce it. New variants then arrive with messages, causes, and documentation in the same PR, instead of as undocumented surprises discovered during the next incident review.
The Unwrap in Config Parsing That Crashed 214 Edge Nodes at Once
- Fail fast means fail descriptively: an expect with a static string is a riddle, while a validated loader returning one anyhow error per bad key is a runbook. Startup code deserves richer errors than hot paths, not poorer ones, because its failures page humans directly.
- Correlated config changes turn one panic into a fleet-wide outage. Validate the entire config surface in a single pass and report all failures at once, so operators see the full blast radius in the first log line instead of discovering it across 38 minutes of diffs.
- Deny unwrap and expect at the crate level with clippy lints and enforce it in CI. Review culture cannot catch every expect in a growing codebase, but a deny attribute catches all of them on every build, forever.
err.to_string() snapshots for anyhow paths. Better fix: push matchable logic into a thiserror-typed library core and assert on variants with assert!(matches!(err, LoadError::MissingKey(_))), leaving anyhow for the binary wrapper. This keeps unit tests precise while integration tests assert full context chains.source() linkage so future variants stay consistent.| File | Command / Code | Purpose |
|---|---|---|
| src | use std::fs; | The ? Operator |
| src | use std::collections::HashMap; | Option Combinators |
| src | use anyhow::{bail, ensure, Context, Result}; | anyhow for Applications |
| src | use std::path::PathBuf; | thiserror for Libraries |
| src | use std::fmt; | Custom Error Enums by Hand |
| src | fn retries_from_env(raw: Option<&str>) -> Result<u32, String> { | panic, unwrap, and expect |
| src | use anyhow::{Context, Result}; | anyhow vs thiserror |
| src | use std::num::ParseIntError; | Silent Killers |
Key takeaways
Common mistakes to avoid
7 patternsReturning anyhow::Error from a public library API
Using expect or unwrap on runtime input like env vars and file contents
Propagating with bare ? at every layer and no context
Typing errors as String and formatting away the source
source().is_some().Swallowing errors with let _ =, map_err discards, or silent defaults
Shipping production images without RUST_BACKTRACE=1
Matching on anyhow errors with downcast in business logic
Interview Questions on This Topic
What three things does the ? operator do, and when does it fail to compile?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Core. Mark it forged?
25 min read · try the examples if you haven't