Rust Smart Pointers: Box, Rc, Arc and Interior Mutability
Box owns data on the heap, Rc shares it on one thread, Arc across threads.
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓Comfortable writing Rust with ownership, borrowing, and cargo
- ✓Hands-on experience with threads, closures, and generic types
- ✓Basic familiarity with trait objects and collection-based designs
- Box
is single ownership on the heap: use it for recursive types, large values you want to move cheaply, and trait objects whose size the compiler cannot know - Rc
adds reference-counted sharing on one thread with clone-to-share semantics, while Arc is the atomic, thread-safe twin required once values cross thread boundaries - Deref coercion lets &Box
and &Rc act as &str at function boundaries, so borrow-first APIs rarely need to know which pointer any caller chose - RefCell
moves borrow checking to runtime for single-threaded interior mutability and panics on conflicting borrows, while Cell covers Copy values without ever panicking - Mutex
and RwLock are the threaded equivalents: blocking locks that can poison and deadlock, so keep critical sections short and never hold two locks in inconsistent order - Reference cycles between Rc pointers leak silently because each count keeps the other alive: break them with Weak
, and pick the pointer from the decision guide before building the graph
Imagine a shared workshop manual that several mechanics need to read. You could photocopy it for everyone, but updates become chaos. You could keep one copy on a clipboard with a sign-out sheet that only works inside one building — that's Rc, cheap sharing for a single-threaded crew. You could lock it in a fireproof cabinet with a key log honored across every site — that's Arc, heavier but safe everywhere. A Box is simpler: one mechanic owns the manual outright and keeps it in a labeled crate nobody else touches. RefCell and Mutex are the rules taped to the clipboard about who may scribble corrections while others are reading, enforced by a supervisor watching in real time rather than by the clipboard's printed instructions.
Rust's ownership rules insist every value has exactly one owner. Smart pointers don't break that rule. They give you controlled, well-lit ways to share, borrow, and mutate when single ownership alone can't express your design.
Most developers meet Box first, usually while fighting a recursive-type error. Then Rc appears in a GUI tree or an AST, Arc shows up the moment threads enter the picture, and RefCell or Mutex arrives when shared data also needs to change. Each pointer solves one problem well and creates new failure modes you'll want to recognize early.
The confusion usually centers on choosing between them. Rc and Arc look interchangeable until a thread boundary rejects your code. RefCell and Mutex both offer interior mutability, but one panics on misuse while the other blocks. Trait objects add another axis: static generics versus dynamic dispatch, each with different pointer requirements.
This guide works through every major pointer in order: Box and deref coercion, the Rc versus Arc thread-safety split, RefCell's runtime borrow rules, Mutex and RwLock discipline, trait objects and object safety, reference cycles with Weak, and leak and drop-order hazards. It closes with a decision guide you can apply to any design review.
You'll leave able to pick the cheapest pointer that satisfies your threading and mutation needs, to break cycles before they leak, and to read pointer-related compiler errors as directions rather than obstacles.
Box: Recursive Types and the End of Infinite Size Errors
Rust must know every type's size at compile time, and recursive types break that requirement. An enum like List with a variant holding another List inline would nest forever: each List contains a List contains a List, and no finite size satisfies the equation. The compiler reports E0072 with the recursive cycle spelled out. Box resolves it by adding indirection: a Box<T> is always one pointer wide regardless of T's size, so Node(T, Box<List>) has a fixed layout the compiler can measure and lay out.
The same indirection serves large values and cheap moves. Moving a 4 KB struct by value copies 4 KB; moving a Box<Big> copies 8 bytes and transfers heap ownership. Constructors returning Box<Config> hand the caller a stable heap address without exposing allocation details, and dropping the box frees exactly once through its Drop implementation. None of this involves reference counting or locking: Box is sole ownership with deterministic destruction, the cheapest pointer in the language and the right default whenever one owner suffices.
Trait objects need Box for the same sizing reason. A bare dyn Trait has no size because any implementor could back it, so functions return Box<dyn Trait> to erase the concrete type behind a uniform pointer. The vtable pointer travels with the data pointer in the fat-pointer layout, and method calls dispatch dynamically. This is the standard way to build plugin registries and heterogeneous collections where the concrete types aren't known at compile time.
Unboxing is explicit and total. Dereferencing with * moves or borrows the contents, Box::leak converts to a permanent reference for genuine singletons, and into_inner-style patterns recover values. Because ownership is unique, passing &Box<T> where &T is expected works through deref coercion, which the next section covers in depth with call-site examples.
Sizedness also governs generics in ways Box quietly fixes. A generic function over T: Sized rejects trait objects and unsized slices, while one over T: ?Sized accepts them behind references and boxes. Returning Box<dyn Trait> from a factory keeps the factory's own signature concrete while its products vary, a combination plain generics cannot express. Whenever the type varies but the handle must stay uniform, Box is the adapter.
Reach for Box when the problem is size, recursion, or type erasure on a single owner. The moment two owners need the same value, Box's uniqueness becomes the obstacle, and the Rc versus Arc decision takes over. Keep the progression in mind: own first with Box, share second with counted pointers, mutate third with interior-ownership wrappers.
Cow<T> sits beside Box for the borrow-or-own choice. An enum-like Cow<'a, str> holds either borrowed or owned text, letting parsers return zero-copy slices on the fast path and allocated Strings only when escapes demand it. Functions accepting impl Into<Cow<'a, str>> take literals, owned Strings, and slices uniformly, and callers pay allocation solely for inputs that need it. When a Box return forces allocation unconditionally, consider whether Cow expresses the real contract: borrow when possible, own when necessary.
Unsized coercions extend Box past sized types. A Box<String> converts to Box<str> and a Box<Vec<T>> to Box<[T]>, trimming the capacity metadata and freezing the allocation at its exact size. These coercions run through the same unsized machinery as trait objects, producing fat pointers the compiler tracks statically. Long-lived configuration buffers and interned strings benefit: allocate dynamically during setup, freeze into the exact-size form, and share the immutable result without further bookkeeping.
Allocation profiling keeps Box usage honest. Tracking bytes allocated per request before and after introducing Box indirection confirms the move actually helped: fewer copies should show as lower allocator traffic, not merely tidier types. Jemalloc stats or massif snapshots attribute the heap to the owning types, and a Box that merely relocated copying without removing it shows up as unchanged totals. Apply the same measurement to every pointer choice in this guide, because each one trades one cost for another and only profiles reveal the balance.
Deref Coercion: Writing Code That Accepts Every Pointer
Smart pointers implement Deref, which lets &Box<String>, &Rc<String>, and &Arc<String> all coerce to &str where the target type is known. A function declared fn greet(name: &str) accepts all three without generics, because the compiler inserts deref steps automatically at coercion sites: &Box<String> derefs to &String, which derefs to &str. Method resolution follows the same chain, so boxed and counted values call str methods directly with no ceremony.
The coercion rules are deliberately narrow. Deref coercion applies to references, never to owned values: passing a Box<String> where String is expected still requires explicit * unboxing. It fires at function arguments, method receivers, and let bindings with explicit types, but not through generic inference, which is why fn f<T: AsRef<str>> and fn f(&str) behave differently under pointer arguments. Understanding these sites prevents the common surprise of a coercion working in one position and failing in another.
DerefMut extends the chain to mutable access for Box, and auto-ref on method calls makes &mut Box<T> behave like &mut T. Rc and Arc deliberately omit DerefMut, because shared ownership plus mutation would alias: two owners observing different values through the same allocation breaks the contract the borrow checker enforces. Mutation of shared data goes through Cell, RefCell, Mutex, or clone-on-write instead, each covered in its own section.
Design APIs around borrowed forms to stay pointer-agnostic. Accept &str rather than &String, &[u8] rather than &Vec<u8>, &dyn Trait rather than &Box<dyn Trait>. Callers with any pointer coerce for free, tests pass literals directly, and the signature documents the minimal access it needs. Reserve pointer-typed parameters for functions that actually manage ownership: taking Box<T> to store, Rc<T> to share, Arc<T> to send across threads.
AsRef and Borrow complement coercion for generic code. A bound like T: AsRef<str> accepts String, &str, Box<str>, and PathBuf uniformly where a &str parameter would force the caller to coerce first. Borrow goes further for map lookups, letting HashMap<String, V>::get accept &str keys. Together these traits extend the borrow-first philosophy into generics without sacrificing the ergonomics deref coercion provides for concrete signatures.
Custom smart pointers participate through their own Deref impls, with the Target associated type naming the borrowed view. Guard types like MutexGuard and RefMut deref to the protected data, which is why locked mutexes feel like direct references. One coherent mechanism underlies the whole ecosystem, and writing borrow-first signatures plugs your code into it permanently.
Method receivers exploit deref chains invisibly. Calling text.len() on a Box<String> auto-refs through two deref steps to str::len without a single explicit borrow, and the same call works on Rc<String> and Arc<String> identically. This uniformity is why pointer-heavy code reads like value code: the receiver position coerces through every layer. Explicit reborrows appear only when the target type is ambiguous, such as generic functions with multiple candidate Deref targets, where a targeted as_str() or &** clarifies intent. Trust the chain for method calls, and annotate only where inference reports ambiguity.
Deref hides costs that reviewers should still see. Cloning through an Rc deref target deep-copies the inner value while Rc::clone shares it, and method-call syntax cannot distinguish the two. String concatenation via + on derefed values allocates fresh buffers that look like lightweight operations. Senior reviewers read deref chains for hidden allocations the way they read loop nests for complexity: the syntax is quiet, so the cost analysis must be explicit. Comment the expensive steps, and prefer sharing or borrowing spellings where the cheap path exists.
Coercion sites deserve a review checklist of their own. Function arguments coerce, method receivers coerce, and explicitly typed let bindings coerce, while generic inference, struct field initialization, and return-position inference do not. Marking each boundary in a new API with a comment about which coercions apply saves every future caller an experiment. When a coercion fails where one succeeded nearby, the checklist names the differing site in seconds rather than through trial and error.
Rc: Single-Threaded Sharing With Honest Costs
Rc<T> enables multiple ownership on one thread through a non-atomic reference count. Cloning an Rc bumps the count instead of copying the data, and the allocation frees when the last owner drops. Graphs, ASTs with shared subtrees, and GUI widget trees all use Rc where several parents or caches reference one value whose lifetime no single owner controls. The pointer is explicit about its trade: sharing without threads, at the cost of one counter update per clone.
The clone semantics deserve emphasis because they surprise newcomers. Rc::clone(&a) shares the allocation; a.clone() on the inner value would deep-copy it. Clippy's redundant_clone lint flags the expensive mistake of cloning through the Rc when sharing was intended. Strong counts are inspectable via Rc::strong_count, which lets tests assert sharing structure directly: after cloning, the count reads 2, and after dropping one owner it returns to 1.
Rc pairs with interior mutability for shared-mutable designs. Rc<RefCell<Node>> is the standard shape for a mutable graph on one thread: the Rc shares ownership, the RefCell gates mutation with runtime borrow checks. Each layer's failure mode stays distinct: leaked cycles from the Rc layer, borrow panics from the RefCell layer. Debugging stays tractable because only one layer can be at fault for any given symptom.
The hard boundary is Send. Rc is !Send and !Sync because non-atomic counts would tear under concurrent clones from two threads. The compiler rejects any attempt to move an Rc into thread::spawn or share it through channels. This is not bureaucracy: two threads cloning simultaneously could corrupt the count and free live data or leak it silently. The rejection is the soundness guarantee working exactly as designed.
Weak counts ride alongside strong counts for cycle breaking, covered in full later. An Rc allocation tracks both, freeing the value when strong hits zero while keeping the control block until weak hits zero too. Understanding the two counters explains upgrade expiry and the small residual cost of outstanding Weak handles after teardown.
Use Rc when all sharing stays on one thread and the value's owner is genuinely plural. Prefer borrowed references where one owner with a clear scope suffices, since borrows carry no count overhead. Graduate to Arc the moment any roadmap includes threads, because retrofitting atomicity after the design assumes non-atomic counts is the incident described at the top of this article.
Rc::new_cyclic constructs self-referential graphs soundly. The constructor hands the initializer a Weak to the not-yet-complete allocation, letting a node register itself with its own parent or observer list during creation. Without it, setup code juggles temporaries and upgrade failures during the half-built phase. Cyclic construction plus Weak back-edges gives graphs that build in one expression and tear down deterministically, with no unsafe code and no initialization-order panics.
Identity comparison uses Rc::ptr_eq rather than value equality. Two handles to one allocation compare pointer-identical even when the values also compare equal to a third copy, which distinguishes sharing from coincidence in tests and caches. Canonicalization tables exploit this: interning returns the existing Rc on content match, and ptr_eq verifies the deduplication actually shared. Value equality answers whether contents match; pointer equality answers whether the sharing design worked.
Drop order inside Rc graphs follows field declaration order, which makes struct layout load-bearing for teardown. A node declaring child before parent drops its strong child edge first, cascading correctly; reversing the fields can delay release until the outer scope ends. Tests asserting counts after each drop stage catch ordering regressions the same way they catch cycles. Treat field order as part of the sharing design, not as formatting. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
Arc: Paying the Atomic Price for Thread-Safe Sharing
Arc<T> is Rc's atomic twin: same shared-ownership API, but reference counts updated with synchronized instructions so clones from different threads never tear. Moving an Arc into thread::spawn is sound because every count transition is atomic, and the allocation frees exactly once when the last owner across all threads drops. The price is small but real: each clone and drop pays for atomicity, and the type requires its contents to be Send + Sync before it will itself be Send.
The standard pattern clones before the move. A worker pool holding Arc<Config> calls Arc::clone(&cfg) per thread, giving each thread its own counted handle to one allocation. Clones are cheap enough to do per task at moderate rates, and the shared data stays immutable so no locking is needed. When mutation enters, Arc pairs with Mutex or RwLock: Arc<Mutex<State>> shares both ownership and synchronized access, with each layer doing exactly one job.
Weak has an atomic counterpart in sync::Weak, and the upgrade-downgrade dance works identically across threads. Caches keyed by thread-shared entries use Arc for values and Weak for back-references or eviction callbacks, preserving the cycle-breaking discipline from the single-threaded world. Strong counts remain inspectable for tests, which can assert that worker shutdown returns counts to their baseline.
Performance reasoning stays grounded with numbers. Atomic increments cost single-digit nanoseconds on modern hardware, invisible against any I/O or parsing work. The cost only matters in tight numeric loops cloning per iteration, where restructuring to borrow the Arc's contents once outside the loop removes it. Never choose Rc over Arc for imagined speed in threaded code: the compiler forbids it, and workarounds with thread-local copies usually cost more than the atomics.
Channels and thread pools compose naturally with Arc. Sending Arc<Job> through crossbeam or std channels shares the payload without copying, and scoped threads borrow where spawn demands ownership. The rule stays simple: Arc at every ownership boundary between threads, plain borrows everywhere within a single thread's execution.
The decision rule is binary and deserves to be stated plainly. One thread and shared ownership means Rc. Any thread boundary anywhere in the value's lifetime means Arc. Mixed designs put Arc at the boundary and borrow inside each thread, keeping atomic operations at spawn time rather than in hot loops.
Atomics pair with Arc where counters and flags need lock-free mutation. An Arc<AtomicU64> shared across threads increments without any guard, at the cost of reasoning about memory ordering instead of borrow scopes. Relaxed ordering suffices for statistics; acquire-release suits flags coordinating shutdown. Reserve atomics for Copy-sized state with simple transitions, and reach for Mutex when invariants span multiple fields. The two tools divide by invariant width: one field with one transition means atomics, anything wider means locks.
Downgrade discipline extends Weak across threads through sync::Weak. Caches holding Arc values hand out Weak handles to readers that must tolerate eviction, with upgrade failure driving a clean refetch path. Shutdown sequences drop strong owners first and join workers before asserting weak counts reach zero, which converts teardown races into orderly expiry. Tests mirror the single-threaded asserts: spawn workers, drop the root, join everything, then verify counts at baseline.
Executor awareness completes the Arc story for async code. Multi-threaded runtimes move futures between threads at await points, demanding Send futures and therefore Arc-shared state. Single-threaded or pinned-task executors permit Rc with RefCell exactly like sync code. The runtime's thread model answers the Rc-versus-Arc question before any pointer is written, so read the executor docs first and let the Send bounds confirm the choice at compile time. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
RefCell and Cell: Runtime Borrows on One Thread
RefCell<T> brings borrowing rules to runtime for single-threaded interior mutability. Code holding &Container can mutate through a RefCell field by calling borrow_mut, which checks at runtime that no conflicting borrow is live. Compliant patterns cost only a flag check. Violations panic with already borrowed or already mutably borrowed instead of corrupting memory. The panic message plus a backtrace names the conflict, which makes these failures loud and local rather than silent and distant.
The borrow guard discipline mirrors the compile-time rules. A Ref guard from borrow() acts like a shared reference until dropped; a RefMut from borrow_mut() acts like an exclusive one. Holding a Ref across a borrow_mut call panics, exactly as the static checker would reject the overlap. Scoping guards tightly, mapping them with Ref::map for field projections, and dropping them before reentrant calls keeps programs panic-free. Temporary try_borrow during debugging converts panics into logged sites while you narrow the overlap.
Cell<T> covers the Copy case with no failure mode at all. A Cell<u64> counter behind a shared reference supports get, set, and update operations without guards, panics, or counts. Metrics, generation counters, and lazy-initialized flags fit Cell precisely because the values copy in and out atomically from the single thread's perspective. The restriction to Copy types is the price of infallibility, and it rules out Strings and Vecs.
The reentrancy hazard is the one that bites in production. Calling user callbacks or event handlers while a RefMut guard is live lets the callback re-borrow the same cell and panic. Designs that separate mutation phases from notification phases avoid the trap: compute under the guard, drop it, then invoke callbacks. Recursive data-structure traversals that borrow parent and child cells simultaneously need the same phase discipline applied consistently.
OnceCell and LazyCell extend the family to write-once state. A value computed on first access through shared references needs no locking on one thread and no Option dance at every site. Configuration caches and lazily built lookup tables fit here, with get_or_init expressing the pattern directly instead of manual flag checks.
Choose RefCell when shared ownership genuinely requires mutation on one thread and restructuring cannot scope borrows statically. Prefer plain &mut where a single owner with clear phases suffices, Cell for Copy state, and Mutex once threads appear. Runtime checking is a tool for designs whose borrow patterns depend on data, not an excuse to skip designing the phases.
Ref::map and RefMut::map project guards onto fields without releasing the borrow. A Ref<Config> maps to Ref<DbUrl> for the sub-view a function needs, preserving the runtime check across the projection. Chained maps compose, and the mapped guard drops exactly like the original. This avoids cloning large fields merely to satisfy function signatures, keeping zero-copy structure inside runtime-checked cells the way subslicing does for compile-time borrows.
try_borrow and try_borrow_mut convert panics into recoverable paths. Hot code that may legitimately contend, such as reentrant event dispatch, attempts the borrow and degrades to queuing or logging instead of crashing. The pattern pairs with metrics counting contention events, which distinguishes rare reentrancy from systemic overlap. Reserve unwrap-style borrowing for phases the design guarantees exclusive, and instrument every place the guarantee depends on data rather than structure.
Testing runtime-borrowed state needs contention scenarios, not just happy paths. Unit tests that borrow, reenter, and release in adversarial order shake out guard-scope mistakes single-pass tests miss. Fuzzing event sequences through a RefCell-guarded router found overlaps that deterministic tests never constructed. Treat RefCell coverage like lock coverage: the interleavings matter more than the states. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
Mutex and RwLock: Blocking Discipline for Threads
Mutex<T> is RefCell's threaded counterpart with blocking instead of panicking as the contention response. Locking returns a guard derefing to the protected data; a second thread locking the same mutex waits until the first guard drops. The rules are checked by the operating system's synchronization primitives rather than by flags, which makes Mutex sound across threads at the cost of potential blocking. Poisoning adds the failure mode: if a thread panics while holding the lock, later locks return PoisonError until explicitly recovered.
Critical sections deserve to be short and obviously scoped. Compute inputs before locking, hold the guard across the minimal mutation, and drop it before I/O, callbacks, or further locking. The pattern of cloning needed data out from under the guard and releasing before slow work keeps contention proportional to mutation rather than to latency. Locking around a network call serializes all threads on the slowest dependency and shows up immediately as collapsed throughput in every load test.
RwLock refines the contract for read-heavy data. Many readers hold shared guards concurrently while writers exclude everyone, which suits configuration and routing tables read 10,000 times per mutation. The trade-offs are writer starvation under continuous read load and slightly higher per-operation cost than Mutex. Benchmark before assuming RwLock wins: below roughly 80% reads, a plain Mutex often performs equivalently with far simpler reasoning about fairness.
Deadlock is the discipline failure both primitives share. Two locks acquired in opposite order by different paths wedge permanently, as the incident at this article's opening demonstrated with 12 frozen workers. One global lock order, documented where the locks are declared, prevents it. Try_lock with timeouts on request paths converts residual risk into logged degradation instead of full stalls.
Lock granularity deserves the same care as lock order. One giant mutex around a whole server state serializes everything and wastes 12 cores; per-shard or per-field locks let independent work proceed. Splitting hot counters into per-thread cells aggregated on read removes contention entirely for statistics paths. Match granularity to access patterns, and re-measure after each split.
Poisoning policy should be explicit. Where protected data stays valid after a panic, recover with into_inner and log the event. Where invariants may be broken, propagate the poison and fail loudly. Either choice beats unwrap in request paths, where one historical panic shouldn't cascade into perpetual failures.
Condition variables pair with Mutex for event-driven waiting. A worker holding Arc<(Mutex<Queue>, Condvar)> sleeps without spinning until producers notify, replacing poll loops that burn cores. The wait atomically releases the guard and reacquires it on wakeup, which keeps the handoff sound. Spurious wakeups mandate predicate loops: while queue.is_empty() { guard = cvar.wait(guard).unwrap() }. Timeout variants bound the wait for shutdown paths. Queues, barriers, and completion signals all build on this pair rather than on sleep-retry loops.
Poison recovery deserves an explicit policy per lock. Where protected data stays valid after a panic, lock().unwrap_or_else(|e| e.into_inner()) recovers and logs the event for later triage. Where invariants may span mutations, propagate the poison and fail loudly rather than serving torn state. Request paths should never bare-unwrap: one historical panic must not cascade into perpetual failures. Decide the policy where the lock is declared, beside the comment stating what the lock protects.
Benchmarks settle Mutex-versus-RwLock debates faster than reasoning. Contention profiles under realistic read-write ratios show whether reader parallelism actually materializes or collapses into cache-line bouncing. Below the crossover point the simpler Mutex wins on both throughput and reviewability. Record the measured ratio beside the lock choice so future traffic shifts trigger reevaluation instead of silent degradation.
Trait Objects: dyn Dispatch and the Object Safety Line
Trait objects erase concrete types behind a uniform interface: Box<dyn Handler> holds any implementor, dispatching method calls through a vtable at runtime. Heterogeneous collections, plugin registries, and test doubles all rely on this erasure where generics would force a single concrete type. The cost is one indirection per call plus the loss of inlining, negligible against I/O but measurable in tight numeric loops where monomorphized generics stay faster.
Object safety draws the boundary of what can be erased. Methods returning Self, taking generic parameters, or requiring Sized cannot dispatch through a vtable, because the erased object cannot name the concrete type at runtime. The compiler rejects non-object-safe traits in dyn position with precise notes. Standard repairs include boxing return values as Box<dyn Trait>, adding where Self: Sized to exclude offending methods from the object, or splitting the trait into an object-safe core plus a generic extension.
The object lifetime is the second axis, worth restating here because it bites pointer code constantly. Box<dyn Trait> defaults to Box<dyn Trait + 'static>, demanding fully owned implementors. Handlers borrowing configuration need Box<dyn Trait + 'reg> naming the registry's region. Borrowing the object as &dyn Trait sidesteps ownership entirely for call-scoped dispatch, and Arc<dyn Trait + Send + Sync> is the threaded registry shape for worker pools.
Sizing choices compose with dispatch. Box<dyn Trait> owns, &dyn Trait borrows, Rc<dyn Trait> shares on one thread, Arc<dyn Trait> shares across threads. Each composes the pointer's ownership contract with dynamic dispatch, and deref coercion keeps call sites uniform. Pick the pointer by ownership needs first, then add dyn where heterogeneity requires it, never the reverse.
Auto traits gate threaded dispatch explicitly. A registry shared across threads needs dyn Trait + Send + Sync, and the compiler checks every implementor at the boundary where the object is constructed. Forgetting Sync on a handler holding RefCell fails fast with a clear note instead of a data race. State the auto-trait bounds where the registry is declared so implementors see the requirement beside the trait.
Prefer generics where the type set is closed and performance matters; prefer trait objects where the set is open or heterogeneous. A parser with three known backends stays generic. A plugin system loading unknown handlers at runtime needs dyn. The decision is about openness, not fashion, and reviewers should ask which unknown type tomorrow's code must accommodate.
impl Trait in return position offers static dispatch without naming types. A function returning impl Iterator<Item = &str> keeps monomorphized speed and full inlining while hiding the adapter chain from callers. Trait objects trade that speed for heterogeneity: Box<dyn Handler> holds mixed implementors one vector cannot otherwise contain. The two compose across API layers: generic cores for speed, dyn boundaries for plugin points. Choosing per boundary rather than per codebase gives both performance where measured and openness where required.
Fat pointers explain dyn costs mechanically. A &dyn Trait is two words, data plus vtable, versus one word for a concrete reference. Method calls index the vtable instead of inlining, and the branch predictor absorbs the steady targets well. Thin-pointer generics avoid all of this at the cost of code bloat per instantiation. Neither dominates universally: profile the dispatch site, and prefer generics in numeric loops with dyn at architectural seams.
Versioning object-safe traits keeps plugin ecosystems compiling. Adding a method with a generic parameter or a Self return breaks every dyn consumer at once, so new capability ships as a supertrait or an extension trait with Self: Sized. The object-safe core stays frozen while generics evolve beside it. Review trait changes for object-safety impact the way API changes are reviewed for semver impact.
Reference Cycles: How Rc Graphs Leak and How Weak Breaks Them
Reference cycles defeat counting. When node A holds Rc<B> and B holds Rc<A>, each strong count stays above zero as long as the other lives, so dropping the external handles leaves the pair alive but unreachable. The memory leaks silently: no panic, no error, just RSS climbing under load. Long-lived graphs, doubly-linked structures, and parent-child widget trees all construct cycles by default unless back-edges are designed as non-owning from the start.
Weak<T> is the non-owning counterpart. Downgrading an Rc produces a Weak that observes without incrementing the strong count; upgrading returns Option<Rc<T>>, yielding None once the value is gone. The discipline is directional: children own parents strongly or parents own children strongly, and the reverse edge stays Weak. Dropping the root then cascades correctly, because no count reaches back up to keep ancestors alive.
The upgrade pattern needs care around expiry. Code holding only a Weak must handle None as a normal outcome, typically by rebuilding, skipping, or reporting absence. Unwrapping upgrades converts orderly teardown into panics during shutdown sequences when parents drop before children's callbacks run. Treat upgrade failure as information about destruction order, and structure teardown so owners outlive observers' final callbacks.
Detecting cycles starts with counts. Logging Rc::strong_count at insert and evict points shows values that never return to baseline; a count stuck at 2 after eviction names the retained edge. Leak-checking tests assert baseline counts after dropping test graphs, and massif profiles attribute growth to the leaking type. Suspect cycles whenever memory grows with graph churn but request volume stays flat.
Observer lists and caches are the everyday cycle sources. A subject holding strong handles to observers that each hold the subject strongly cycles on every subscription. Event buses, UI signal graphs, and self-registering plugins all need Weak subscriber edges or explicit unsubscribe tokens consumed on drop. Design the deregistration path before the first subscription ships.
The threaded world mirrors all of this through sync::Weak. Arc cycles leak identically, and the same directional discipline with Weak back-edges applies. Caches deserve special mention: values referencing their own cache key or eviction handle cycle trivially, so cache entries should borrow cache state weakly or through explicit eviction tokens rather than strong handles.
Rc::new_cyclic deserves emphasis as the constructor that makes self-observation sound. Passing |weak| Node { parent: weak, .. } into the constructor wires the back-edge during allocation, before any strong handle escapes. The resulting graph needs no post-construction fixup phase where invariants are temporarily broken. Combined with RefCell fields for link mutation and Weak edges for cycles, cyclic construction completes a toolkit where safe graph code reads linearly from allocation to teardown.
Count-based tests make leaks deterministic. Building a representative graph, dropping roots, and asserting strong and weak counts at zero catches regressions the same commit that introduces them. Heap profiling under churn catches the structural cases tests miss: flat traffic with climbing RSS names the leaking type, and count logging at insert and evict points names the retained edge. Both layers matter because unit tests cover shapes while profiles cover lifetimes.
Shutdown sequencing turns upgrade expiry from hazard into routine. Dropping strong owners before joining observer threads guarantees every Weak upgrade resolves deterministically during teardown. Tests encoding this order fail loudly when refactors reorder destruction. Document the teardown sequence beside the graph type so the invariant survives team turnover. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
Leak, Drop Order, and the Hazards Nobody Tests
Not every leak is a cycle. Box::leak, Rc::into_raw, and mem::forget all intentionally relinquish destruction, trading permanent memory for simplified lifetimes or FFI stability. Used once for a global config, the cost is bounded and documented. Used per request, the cost compounds exactly like the leak in the companion lifetimes article: kilobytes per message at thousands of messages per second until the OOM killer audits your design. Every intentional leak deserves a comment stating its bound and a test asserting the bound holds.
Drop order is the subtler hazard. Struct fields drop in declaration order, locals drop in reverse declaration order, and temporaries drop at statement end. Code depending on teardown sequencing, such as guards releasing locks or files flushing before a directory handle closes, must encode the order in field layout rather than in comments. Reordering fields for aesthetics can reorder destruction with observable consequences for lock release and file durability.
ManuallyDrop and reference counting interact dangerously. Forgetting a guard keeps a mutex locked forever; forgetting an Rc keeps its allocation alive with no owner to free it. Explicit forget calls should be rare, reviewed, and paired with a documented reclaim path. Prefer scoped guards and structured teardown so destruction stays automatic and ordered by the compiler.
Raw-pointer round trips need matching discipline. Rebuilding a Box with Box::from_raw must happen exactly once per into_raw, under the same allocator, with no live references in between. Reference-counted raw handles add the count protocol on top. Encapsulate every raw round trip in a small module with debug assertions on the pairing, and never let from_raw sites proliferate across call sites.
Cyclic Rc graphs deserve a testing strategy proportionate to their silence. Unit tests that build representative graphs, drop the roots, and assert strong and weak counts at zero catch leaks deterministically. Heap profiling under realistic churn catches the rest: flat request volume with climbing RSS is the signature, and attributing growth to one type names the cycle's members.
Treat destruction as part of the API contract. Document field drop order where it matters, bound every intentional leak, and test teardown with the same rigor as construction. Memory safety guarantees memory won't corrupt; only design discipline guarantees it gets freed.
Explicit drop() beats scope tricks for readability. Calling drop(guard) at the exact statement a lock should release documents intent better than an extra block level, and reviewers see the release point without tracking braces. Shadowing a variable to end a borrow works but reads as accident; prefer the explicit call. Guard structs extend the pattern to custom resources: a TempDir guard removing its directory on drop, or a metrics timer recording on drop, both convert cleanup from convention into construction.
needs_drop and Drop-scope reasoning guide generic containers. Generic code holding T must assume destruction has side effects unless bounded otherwise, which constrains unsafe teardown paths most application code never writes. For safe code the lesson is simpler: store guards and buffers as ordinary fields and let declaration order define teardown. Reordering fields for aesthetics reorders destruction, so treat field layout as load-bearing wherever guards, files, or locks are involved.
Sanitizer and profiler passes belong in CI for pointer-heavy services. AddressSanitizer catches raw round-trip mistakes, ThreadSanitizer flags missing synchronization around shared mutation, and heap profiles attribute growth to leaking types. Each tool covers a failure mode silent in normal tests. Schedule the slow suites nightly if per-commit cost bites, but never leave cycles and races to production telemetry alone. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
Choosing the Pointer: A Decision Guide With Thread Table
Start from ownership cardinality. One owner with known size means plain values or references; no pointer is the cheapest pointer. One owner needing indirection for size, recursion, or type erasure means Box. Several owners on one thread mean Rc. Several owners across threads mean Arc. Each step adds exactly one capability and its cost, so stopping at the first pointer that fits keeps designs minimal and reviewable.
Add mutation as the second axis. No mutation needs nothing beyond the ownership pointer. Single-threaded mutation through sharing needs Cell for Copy values or RefCell for the rest. Threaded mutation needs Mutex for general use or RwLock for read-heavy access. Never combine sharing and mutation in one mental step: pick the sharing pointer first, then wrap only the mutable fields, keeping immutable data directly under the Arc or Rc.
The thread table collapses the whole guide into one lookup. Box, Rc, RefCell, and Cell are !Send: confined to one thread by construction. Arc, Mutex, and RwLock are Send when their contents are: the crossing guards of the ecosystem. Trait objects add Send + Sync bounds at the dyn boundary for threaded registries. When the compiler rejects a crossing, it is reading this table aloud; the fix is moving up to the threaded row, not casting around it.
Erasure is the third axis. Closed type sets stay generic for speed and inlining. Open or heterogeneous sets take dyn behind whichever pointer ownership selected: Box<dyn T> to own, Rc<dyn T> to share singly-threaded, Arc<dyn T> to share across threads. Object safety and object lifetimes constrain the trait side while the pointer side follows the same sharing rules as concrete types.
Cost awareness sharpens each choice. Box adds one allocation. Rc adds a counter bump per clone. Arc adds atomic bumps. RefCell adds flag checks. Mutex adds potential blocking. None of these matter beside I/O, and only hot-loop measurements should overturn the default pick. Profile before downgrading safety for speed you haven't measured.
Apply the guide in order at every design review: cardinality, then mutation, then threads, then erasure. Document the answers beside the type alias so the next engineer inherits the reasoning. Most pointer misuse is not ignorance of any single type but skipping one axis: sharing chosen without considering threads, mutation added without scoping guards, erasure assumed without checking object safety.
The async axis extends the table without changing it. Futures holding Rc across await points stay !Send, so spawn_blocking and multi-threaded executors demand Arc there too. Single-threaded runtimes permit Rc-shared state with RefCell mutation, mirroring the sync rules exactly. Channel choice follows the same split: non-Sync payloads stay on local channels, Send payloads cross threads. Threading questions answer identically in async code once the executor's thread model is known.
Anti-patterns complete the guide by naming what the table forbids. Arc<Mutex<T>> as a default wraps immutable data in needless synchronization; clone-on-write or plain Arc sharing fits better. Rc in request handlers breaks at the first thread pool; borrowing or Arc fits from the start. Mutex around whole servers serializes independent work; sharded or field-level locks restore parallelism. Each anti-pattern is one skipped axis from the decision order, and the review checklist catches all three before merge.
Documentation standards close the loop on pointer choices. Every shared type alias carries a one-line rationale naming cardinality, mutation, threads, and erasure answers. New team members inherit the reasoning instead of reverse-engineering it, and reviewers check the rationale rather than re-deriving the design. The four answers fit in one comment line and prevent an entire class of drive-by regressions. Review checklists that ask for the sharing story before the code story keep these designs honest: every counted pointer in review should point at a written reason it is counted rather than borrowed.
An Rc Cache Crossed a Thread Boundary and Blocked the 2 AM Deploy
- Rc and RefCell are single-threaded by construction, not by inconvenience. The compiler's Send rejection at 01:47 was the last guardrail; design the threading story before building the cache, and default to Arc the moment a second thread exists on any roadmap.
- Every second mutex doubles your deadlock surface. Split immutable data out of locks so only the 16-byte mutable counter needs protection, document one global lock order, and enforce it with try_lock timeouts that degrade to misses instead of wedges.
- Single-threaded tests cannot validate threaded sharing. Gate every shared-pointer change on a 12-thread stress test with 200,000 mixed operations, because the interleaving that deadlocks in 90 seconds of staging will never appear in 400 unit tests.
borrow() with try_borrow() temporarily and log the failure site with cargo run 2>&1 | grep -B 5 panicked. Fix by scoping guards, splitting the RefCell, or removing reentrant borrows.lock().unwrap_or_else(|e| e.into_inner()) where stale data is safe, or propagate where it is not. Harden with cargo test --release stress_ to shake out the underlying panic under load.| File | Command / Code | Purpose |
|---|---|---|
| src | enum List { | Box |
| src | fn greet(name: &str) -> String { | Deref Coercion |
| src | use std::rc::Rc; | Rc |
| src | use std::sync::Arc; | Arc |
| src | use std::cell::{Cell, RefCell}; | RefCell and Cell |
| src | use std::sync::{Arc, Mutex}; | Mutex and RwLock |
| src | trait Handler { | Trait Objects |
| src | use std::cell::RefCell; | Reference Cycles |
| src | use std::rc::Rc; | Leak, Drop Order, and the Hazards Nobody Tests |
| src | use std::rc::Rc; | Choosing the Pointer |
Key takeaways
Common mistakes to avoid
7 patternsUsing Rc or RefCell in threaded code and meeting the Send wall at deploy time
Cloning the inner value when Rc::clone sharing was intended
Holding mutex guards across I/O, callbacks, or second locks
Building parent-child graphs with strong edges in both directions
Wrapping everything in Arc<Mutex<T>> by default
Ignoring PoisonError with unwrap in request paths
Leaking per-request allocations to satisfy a 'static API
Interview Questions on This Topic
When do you choose Box versus Rc versus Arc?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Core. Mark it forged?
25 min read · try the examples if you haven't