Rust Macros: Declarative macro_rules and Derive Done Right
Rust macros explained: macro_rules matchers, repetition, hygiene, derive vs attribute macros.
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓Comfortable with Rust ownership, structs, enums, and traits
- ✓Written generic functions and seen trait bounds
- ✓Built Cargo projects and read rustc error output
- macro_rules! matches input token trees against arms like ($x:expr) and expands to code at compile time
- Repetition $(...),* with separators handles variable argument lists cleanly without recursion
- Hygiene keeps macro locals from colliding with caller names; local_inner_macros fixes cross-crate helper paths
- Four kinds exist: declarative (pattern matching), derive (struct impls), attribute (custom annotations), function-like (new syntax)
- Prefer generics and functions first; reach for macros only for syntax, arity tricks, or compile-time code generation
- Debug with cargo expand to see output and trace_macros!(true) to watch arm selection step by step
Imagine you run a bakery and every morning you write the same 40 labels by hand for each bread variant. A macro is a label printer with stencils: you design the stencil once with blanks for the name, weight, and price, then feed it a list and it stamps out all 40 labels identically. Declarative macros are simple fill-in-the-blank stencils, derive macros are automatic stamps that read a recipe card and print the nutrition label for you, and debugging tools let you preview the printed label before it goes on the shelf. The catch is that a bad stencil prints bad labels at scale, so experienced bakers use plain handwriting for one-offs and save the printer for jobs that truly repeat.
You've copied your third nearly identical impl block and you're wondering whether a macro would help. It might. Rust macros generate code at compile time, which means they can stamp out trait impls, build DSLs, and accept syntax that functions can't touch. Used well, they delete hundreds of lines. Used casually, they create error messages nobody can decode.
Here's the honest tradeoff: macros are the most powerful abstraction in Rust and the hardest to debug. A generic function gives you clear signatures and jump-to-definition. A macro gives you flexible syntax and expansion output you can't see without extra tools. That's why senior teams treat macros as a last resort with a clear trigger, not a first instinct.
But don't avoid them entirely. Derive macros like Debug, Clone, and Serde's Serialize remove mountains of boilerplate you'll otherwise hand-write and drift out of sync. A small macro_rules helper can unify 12 call sites that differ only in a name and a type. You'll use macros regularly once you know which kind fits each job.
By the end of this guide you'll read matcher arms like ($x:expr) fluently, write repetitions with separators, predict hygiene behavior, pick correctly among declarative, derive, attribute, and function-like macros, know when plain generics win, and debug expansions with cargo expand and trace_macros. You'll also see the production discipline that keeps macro APIs stable.
macro_rules! Matchers: Reading ($x:expr) Arms Like a Compiler
A declarative macro is a list of arms, each pairing a matcher pattern with a transcriber template. When you invoke the macro, the compiler tries each arm top to bottom, binds fragments such as $x:expr, and pastes the bindings into the template. The fragment specifier is a kind check, not a name: expr accepts one complete expression, ident accepts one identifier, ty accepts a type, pat accepts a pattern, block accepts a braced block, and tt accepts a single token tree. Choosing the wrong kind is the most common beginner error, and the compiler reports it as no rules expected this token.
Consider a tiny assert-like helper with two arms: one matching ($cond:expr) and one matching ($cond:expr, $msg:expr). The first arm expands to a check with a default message, the second to a check with a custom message. Because arms are ordered, the two-expression call matches the second arm while the single-expression call falls to the first. Reversing the order would shadow the specific arm and misroute calls, which is why style guides insist on specific-first ordering.
Fragment kinds also control what the template can do. An expr binding pastes as a value you can call, add, or compare. A ty binding pastes where a type is expected, such as a turbofish or a let annotation. An ident binding can name a function, module, or variable. Mixing them up produces errors at the paste site rather than the match site, so read the error's span: a failure naming the transcriber points at a kind mismatch upstream.
Visibility and import rules are simpler than folklore suggests. A macro_rules! macro lives in its module until exported with #[macro_export], which hoists it to the crate root. Callers then import it by path like any item under the 2018 and later paths. Textual scope, the old order-dependent behavior, still bites when macros call macros in the same file, so define helpers before use or route them through $crate.
Practice reading arms as contracts: the matcher promises which call shapes are legal, the transcriber promises what code each shape becomes. Write three unit tests per macro, one per arm, and the contract stays honest as the macro grows. That habit alone prevents half the expansion mysteries teams blame on the compiler.
Testing matchers deserves a systematic approach because arm interactions surprise even experienced authors. Write one test per arm with minimal input, then add adversarial cases: an integer where a type was expected, a block where an expression suffices, an empty invocation where repetition allows zero items. Each adversarial test documents a deliberate decision about which arm should win. When a future contributor adds a new arm, the suite shows immediately whether existing calls reroute. That protection matters because arm shadowing produces no warning; the compiler simply matches earlier arms first and stays silent about the overlap.
Fragment choice also affects IDE behavior and rust-analyzer support. Expressions and idents complete reasonably inside macro calls, while opaque tt fragments defeat completion entirely. Prefer the narrowest fragment that accepts all legitimate calls: expr over tt for values, ident over tt for names, ty over tt for types. Narrow fragments give better errors at the call site instead of deep inside the transcriber. The precision compounds across large codebases where hundreds of engineers invoke shared macros daily without reading their definitions.
Lifetime and visibility edge cases round out matcher literacy. A macro that pastes $t:ty into a static must respect the type's visibility or downstream crates fail with private-in-public errors. A macro generating functions from $name:ident should document whether the functions are pub or private, since callers cannot override the choice. Generics add another axis: matchers can bind lifetimes and bounds as token groups, but the template must reproduce them faithfully. Cover each axis with one test and the macro's contract stays legible as it evolves across releases.
Repetition With Separators: $(...),* Without the Recursion
Repetition is what makes declarative macros scale from one-offs to lists. The syntax $( PATTERN ) SEP REP describes a repeated group: the pattern binds fragments per iteration, the separator sets the call-site punctuation, and the repeat operator sets cardinality. A comma separator with * accepts zero or more comma-separated items. A semicolon separator accepts statement lists. Plus demands at least one item, star allows empty, and question makes a group optional. The transcriber then repeats its template once per iteration with each binding substituted.
Separators deserve care because call sites vary. Some authors leave trailing commas, others do not. A macro written as $( $x:expr ),* accepts a, b, c but rejects a, b, c, with a trailing comma. Appending an optional $(,)? tail accepts both shapes and ends the debate. Empty invocations need thought too: a sum macro over zero items must expand to a valid identity value such as 0, or the empty call becomes a compile error that surprises the next user.
Nesting repetitions handles pairs and tables. A route table macro might repeat over $( $method:ident $path:expr => $handler:expr ),* and emit one registration per row. Each row binds three fragments, and the template pastes all three into a builder call. Because the repetition is flat, expansion cost stays linear in the number of rows: 200 routes expand to 200 calls, not to nested recursion depth 200.
The classic mistake is hand-rolled recursion where repetition suffices. Recursive munchers peel one item per expansion layer, holding intermediate token buffers at every level; over 2,400 items that pattern produced 180,000 lines and 34 extra CI minutes in the opening incident. Flat repetition over the same list emitted 11,000 lines in linear time. Prefer repetition by default and reserve recursion for genuinely nested grammars such as balanced trees.
Test the shapes explicitly. Every repetition macro deserves four tests: empty, single, multiple, and trailing-comma invocations. Those four cases catch separator bugs, identity-value gaps, and off-by-one template errors before downstream crates depend on the macro. Four small tests protect every future call site.
Identity values and empty expansions are the details that decide whether a repetition macro feels professional. A sum macro must define sum_all!() as 0, a concat macro must define empty as an empty String, and a route table with zero rows must produce an empty table rather than a type error. The transcriber template determines this: 0 $( + $x )* handles empty naturally because the repetition vanishes and leaves the identity. Templates that unconditionally emit separators or commas fail on empty input. Test the empty call first, not last, since library users hit it in generated code paths the author never imagined.
Separator flexibility extends to newlines and comments. Callers format long invocations across lines with trailing commas under rustfmt; a macro that rejects trailing commas fights the formatter on every save. Comments between items are stripped before matching, so they need no special handling, but attributes between items do need explicit $(#[$attr:meta])* groups when the macro stamps annotated items. Document which decorations are accepted: plain items only, or items with doc comments and cfg attributes. The 30 seconds spent documenting saves repeated issues from users whose reasonable call shapes fail.
Performance of flat repetition stays linear and predictable. Expanding 500 routes emits 500 registration calls with no nesting, and rustc processes them like hand-written code. Contrast this with the recursive muncher from the opening incident, where 2,400 items produced 180,000 lines through nested re-scans. When reviewers ask why a macro uses repetition instead of recursion, the answer is arithmetic: linear scaling with visible output versus quadratic scaling with hidden buffers. Show both expansion counts in the pull request and the choice defends itself.
Cover each axis with focused tests and the macro stays reviewable as it grows across many releases.
Hygiene and local_inner_macros: Why Helpers Break Downstream
Hygiene means identifiers introduced inside an expansion do not accidentally capture caller bindings. When your macro declares let tmp, that tmp is scoped to the expansion and will not steal the caller's tmp. This protection is why declarative macros are safer than C preprocessor pasting: the compiler tracks which names came from the definition site and which came from the call site. Most hygiene bugs therefore involve paths, not locals.
Paths resolve at the call site by default. If your macro transcriber calls helper_fn() by bare name, the compiler looks for helper_fn where the macro is invoked, not where it is defined. That works while tests live in the same crate and breaks the day a downstream crate invokes the macro without importing the helper. The error reads as unresolved import, and the macro author insists it works on their machine. Both sides are right about different scopes.
The standard fix has two parts. Reference helpers through $crate, as in $crate::helper_fn(), which anchors the path at the defining crate regardless of call site. Then annotate the macro with #[macro_export(local_inner_macros)] so nested macro calls inside the transcriber also resolve at the definition crate. Together these make the macro self-contained: callers import one name and get correct resolution everywhere.
Local bindings still deserve discipline. Prefix macro-introduced names distinctively, wrap multi-statement templates in braces so bindings cannot leak, and avoid inventing bindings when passing values as arguments suffices. A macro that takes $value:expr and pastes it twice is clearer than one that binds an implicit scratch variable the caller cannot see.
Verify with a consumer test. Add an integration test crate, or at minimum a tests module in a different module path, that invokes the macro without importing helpers. If that test passes, downstream users inherit the same success. Hygiene issues caught by a 10-line consumer test never become release-day support tickets.
The $crate anchor deserves deeper treatment because it solves more than helper functions. Constants, trait imports, and nested macro invocations all resolve correctly through $crate paths. A macro emitting Default::default() needs no anchor since Default is in the prelude, but one emitting mycrate::Config::load() must route through $crate or downstream renames break it. Renamed dependencies via Cargo's package renaming make this concrete: a user depending on your crate as telemetry_v2 still gets correct paths because $crate follows the local rename automatically. Bare crate-name paths would fail.
Nested macro calls inside transcribers are the second beneficiary of local_inner_macros. Without the attribute, an inner vec! or log! call resolves at the final call site, which usually works for std macros but fails for crate-local helpers. With the attribute, inner invocations resolve at the definition site where the helpers live. Test this with a two-crate workspace in the macro's own repository: the macro crate plus a minimal consumer crate that imports nothing but the macro. If the consumer builds, the anchoring is complete.
Name-mangling conventions for macro locals vary by team, but the principle is constant: distinctive prefixes plus minimal scope. Double-underscore prefixes like __retry_count signal generated bindings in expanded output and reduce collision odds to near zero. Wrapping the template body in an extra block confines those bindings to the expansion. Some teams run cargo expand in review specifically to read generated binding names; that five-minute check catches collisions before they become Heisenbugs in downstream test suites.
Hygiene checks with a downstream consumer test catch path bugs before they become release-day support tickets for the whole team. Consumer-crate tests prove the anchoring works for real downstream users, not just local modules in the same package. Distinctive prefixes plus tight blocks keep generated bindings safe at every call site.
flush() by bare name and passed 300 local tests; the first downstream service failed with 47 unresolved-import errors. Anchoring to $crate::flush fixed all 47 in a single patch release.Four Macro Kinds: Declarative, Derive, Attribute, Function-Like
Rust has four macro kinds because four different inputs need generating. Declarative macros match call-site token patterns and transcribe them; they live in the same crate and handle repetition, tiny DSLs, and constructor shorthands. Derive macros inspect a struct or enum definition and emit a trait impl for that type; they run as procedural plugins and power Debug, Clone, Serialize, and friends. Attribute macros wrap whole items such as functions or modules with custom behavior; test harnesses and web frameworks use them to register code. Function-like macros take an arbitrary token stream at a call position and return code; vec! is the canonical example.
Choosing among them starts with the input shape. If the job is stamping similar code over a caller-provided list, declarative fits. If the job is implementing a trait from a type's fields, derive fits. If the job wraps an entire function with setup, timing, or registration, attribute fits. If the job needs brand-new call syntax that patterns cannot describe, function-like fits. Most application code needs only the first two; frameworks reach for the last two.
Costs rise across that list. Declarative macros compile within the current crate with minimal overhead and readable-ish errors. Derive and other procedural macros require a separate macro crate, add dependency weight, and slow clean builds because they run as compiled plugins. Attribute macros also obscure control flow: a reader sees #[handler] above a function and must learn what wrapper was injected. Reserve heavier kinds for jobs where the generated shape genuinely depends on type structure.
A practical map keeps teams consistent. Local repetition over routes, events, or test cases goes declarative. Trait impls for data types go derive, standard first and ecosystem second. Cross-cutting wrappers such as retries or transactions go attribute, but only behind well-documented facades. Novel syntax goes function-like and ships with expansion examples in the docs. Anything else is a function or a generic waiting to be written.
Teach the map with one example per kind in onboarding. Engineers who can name the four inputs and their costs stop proposing attribute macros for list stamping and stop hand-writing the tenth Debug impl. The taxonomy pays for itself the first time a design review picks derive over a 90-line declarative workaround.
Procedural macro crates impose structural costs that declarative macros avoid. A proc-macro crate compiles as a separate dylib with its own dependency tree, links into every downstream build, and runs arbitrary code at compile time. That power enables derive macros that read field types and attribute macros that rewrite functions, but it adds build minutes and supply-chain surface. Audit the tradeoff per dependency: serde_derive earns its place in nearly every service, while a clever one-off attribute for 3 call sites rarely does. Count proc-macro dependencies in CI and question each addition like any other heavyweight import.
Error spans separate beloved macro crates from dreaded ones. A derive that emits errors pointing at the offending field with a plain-English message gets thanked in chat; one that panics with index out of bounds inside generated code gets replaced. Procedural authors should map errors back to input spans with syn::Error and test messages with trybuild snapshots. Declarative authors get spans for free from the transcriber but should still add compile_error! guards for unsupported shapes. Error quality is the highest-leverage investment in any macro crate.
Migration paths deserve planning because syntax sticks. When a declarative macro must grow type-aware features, teams often rewrite it as a function-like proc macro while keeping call syntax compatible. Ship the new implementation under the same name, verify all existing call shapes still match, and document any intentional rejections. Conversely, when a proc macro's job shrinks to simple repetition, rewrite it declaratively and delete the extra crate. Both migrations are routine in maturing codebases; the test suite of call shapes makes them safe.
Deriving Standard Traits: Debug, Clone and Friends Without Drift
The std derive set is the highest-value macro usage in Rust: small annotations that generate correct, boring code you would otherwise hand-write and let drift. Debug gives printable representations for logs and asserts. Clone gives explicit duplication with a visible cost. PartialEq and Eq give equality, Hash gives map keys, Default gives zero-values, and PartialOrd with Ord give ordering. Each derive reads the type's shape and emits an impl that tracks future field changes automatically.
Derive deliberately, not maximally. Debug, Clone, and PartialEq belong on nearly every data type because tests and logs need them. Copy fits small plain-data types under about 32 bytes with no ownership; larger types should move or clone explicitly so costs stay visible. Default makes sense when a sensible empty value exists, and Hash plus Eq make sense when the type keys a map. Deriving Ord on a type with no natural order invites misuse, so leave it off until a sort call demands it.
Customization stays attribute-driven. Skip a secret field from Debug with a manual impl or a wrapper that redacts; never log tokens because a blanket derive printed them. Set struct update syntax with ..Default::default() so new fields with defaults do not break call sites. For enums, derived PartialEq compares discriminants plus payloads, which is exactly the exhaustiveness you want in state-machine tests.
The drift argument is economic. A hand-written PartialEq for a 12-field struct costs 20 lines and rots when field 13 arrives; the derive costs one token and updates itself. Across a 400-type codebase that difference is thousands of lines that never desync. Code review should therefore ask why a manual impl exists whenever a derive could serve: the manual version needs a behavioral reason, not a stylistic one.
Teach derives as public promises. Clone promises duplication cost, Eq promises total equality, Hash promises stable hashing, Debug promises printable output. Each promise constrains future refactors, which is good: the compiler will point at violated assumptions instead of letting them slide. Derive the honest set early and your types stay truthful as they grow.
Redaction and sensitive fields are the production edge of derives. A blanket #[derive(Debug)] on a struct holding tokens, passwords, or PII prints secrets into logs on the first debug! call. The fix is a manual Debug impl that writes [redacted] for sensitive fields, or a wrapper type such as Secret<String> with a redacting Debug impl used consistently. Audit derives on every type crossing trust boundaries: auth tokens, payment instruments, health records. One grep for derive.*Debug near sensitive field names during security review prevents the log-scraping incident nobody wants to disclose.
Copy semantics deserve similar scrutiny. Deriving Copy on a 256-byte struct makes every assignment a memcpy that looks free in source; passing such values through 5 function layers copies kilobytes silently. Reserve Copy for small plain-data types under roughly 32 bytes: coordinates, and small flags. Larger types should move or clone explicitly so costs appear at the call site. When performance profiles show unexpected memcpy hotspots, oversized Copy types are a frequent culprit. Clippy lints flag some cases, but judgment about domain semantics stays human.
Enum derives carry their own subtleties. Derived Clone on an enum with a 2 KB payload variant clones the full payload on every duplication; consider Box on large variants to keep cloning cheap. Derived Default on enums requires #[default] on one variant, which should be the genuinely empty state rather than a convenient first variant. Derived ordering on enums follows declaration order, so reorder variants only with full awareness that persisted discriminants and sort orders shift. Each derive on an enum is a promise about representation; review variant changes accordingly.
Keep one source per trait so manual impls and derives never collide on the same type in future refactors.
Serde Derive in Production: Wire Formats Without Hand Parsers
Serde derive moves serialization from hand-written parsers to annotated types: #[derive(Serialize, Deserialize)] on a struct generates wire conversion that tracks fields automatically. Config files, REST payloads, and queue messages become struct definitions with rename, default, and skip attributes instead of stringly map lookups. A 40-field API response that once needed 200 lines of manual mapping becomes a 45-line struct with five attributes.
Attributes carry the production semantics. Rename aligns Rust snake_case with JSON camelCase without renaming fields. Default plus missing-field tolerance keeps deploys compatible when producers add fields before consumers upgrade. Skip hides secrets and caches from the wire, while skip_serializing_if trims None payloads that would otherwise double message sizes. Deny_unknown_fields on ingestion types turns silent schema drift into loud deploy-time errors, which is the behavior you want for payment and auth payloads.
Versioning discipline matters more than the derive itself. Additive field changes with defaults deploy in either order; renames and type changes need dual-read windows. Gate wire types behind explicit modules so a refactor cannot silently alter the external schema. Snapshot-test serialized output for the 20 most critical message types: a golden JSON file per type fails the build when bytes change, forcing a conscious schema decision.
Error handling stays typed. Serde errors name the missing field and the failure path, which beats a generic parse failed log. Map them into your error enum with context about the source file or endpoint, and include the first 200 bytes of payload in debug logs so on-call can diagnose without replaying traffic. Never unwrap deserialization in request paths; a single malformed client message should return 400, not crash a worker.
Performance is rarely the bottleneck, but know the shape. Derived impls serialize at near-manual speed; buffering with BufReader and reusing allocations matter more than template tweaks. For the 2,400-event registry pattern, one derived impl on a shared envelope plus validation beats 2,400 generated impls by 170,000 lines and 30 CI minutes. Derive once per shape, validate registries with code, and keep wire types boring.
Schema evolution with serde is where teams win or lose quarters. Additive changes, new fields with defaults, new optional variants, are safe in both deploy orders when consumers tolerate missing fields and producers ignore unknown ones. Destructive changes, renames, type swaps, removed variants, need coordinated windows: producers emit both shapes during migration, consumers read both, then the old shape retires. Deny_unknown_fields must be relaxed during dual-read windows and re-tightened after. Document each wire type's compatibility level in its module header so deploy sequencing is explicit rather than tribal.
Payload size discipline matters at scale. A 40-field struct serializing None as null for 30 fields doubles message bytes across millions of queue messages daily. skip_serializing_if trims those to nothing; flatten and untagged enums deserve benchmarks before adoption since they complicate deserialization. Measure serialized bytes per message type in tests with assert size guards that fail when payloads grow 20 percent. Those guards convert silent bloat into reviewable diffs.
Testing wire types needs golden files plus property checks. Golden JSON snapshots pin exact bytes for the 20 critical message types, failing on any drift. Round-trip property tests generate random values, serialize, deserialize, and assert equality across 10,000 cases, catching asymmetric attributes like serialize-only renames. Fuzz the deserializer with malformed inputs to confirm 400 responses instead of panics. Together these layers make wire formats the most tested code in the service, which matches their blast radius.
Derive once per envelope shape and validate registries with code to keep compile times flat as event counts grow.
When Not to Macro: Generics and Functions Win Most Fights
The macro temptation peaks at the third copy-paste. Three impl blocks differ only in a type, and macro_rules! promises to stamp all three from one template. Sometimes that promise is right. More often a generic function with a trait bound covers the variation in fewer lines with better errors, working jump-to-definition, and no expansion overhead. The senior move is writing the function first and promoting to a macro only on concrete evidence.
The decision test is mechanical. If the varying parts are values and types that fit parameters, write a function. If the varying parts are behavior with shared skeleton, write a trait with default methods and implement the hooks per type. If the varying part is arity, syntax, or identifier generation, such as stamping 12 named constructors or accepting trailing-comma tables, then a macro earns its keep. Functions cannot invent identifiers or accept variable-arity syntax; macros can.
Error quality is the tiebreaker newcomers undervalue. A generic function failure names the trait bound, the type, and the call line. A macro failure names an expansion span the user never wrote. Multiply that gap by 30 call sites and onboarding cost dominates any line savings. Documentation follows the same split: rustdoc shows function signatures beautifully and macro templates opaquely unless you write expansion examples by hand.
Compile-time cost reinforces restraint. Each macro expansion adds tokens for rustc to process; generic monomorphization also duplicates code, but with predictable per-type cost instead of template recursion risk. The incident macro that cost 34 CI minutes would have been a 300-line build script plus one generic envelope from the start. Measure both with cargo build timings before committing to cleverness.
Adopt a team rule: third repetition gets a function, fifth repetition with syntax needs gets a macro, and every macro ships with a non-macro alternative documented in its rustdoc. That rule keeps abstractions honest. Most proposed macros dissolve into generics under review, and the survivors arrive with tests, examples, and a clear reason functions could not serve.
Const generics and impl specialization trends keep shrinking macro territory. Patterns that once needed macros, such as fixed-size array helpers for N in 1 to 32, now compile as const-generic functions with real signatures. Trait specialization, where available, replaces macro-stamped per-type impls with conditional default logic. Before writing a macro in 2026, check whether const generics, generic const expressions, or trait defaults cover the case; the language gains yearly ground. Macros should occupy only the ground functions cannot reach: syntax, arity, and identifier generation.
Readability economics favor functions by compounding margins. Every engineer reads function signatures fluently; macro templates require learning matchers, fragments, and repetition per macro. A codebase with 5 macros stays navigable; one with 60 becomes a private dialect where only authors can review call sites. Cap macro count per crate in review guidelines and require a generics-rejected rationale in each macro's rustdoc. The cap forces consolidation: three overlapping helper macros merge into one generic plus one macro, halving the dialect surface.
Staffing risk is the quiet argument. New hires become productive with functions in days and with macro-heavy code in weeks. During incidents, responders debug function call stacks confidently and macro expansions hesitantly. Keeping macros few and well-documented shortens onboarding and incident response alike. When the team grows 3x in a year, the restraint pays out in every sprint: fewer macro mysteries, faster reviews, calmer pages.
Most proposed macros dissolve into generics under review, and the survivors arrive with tests and documented rationale. The restraint compounds across quarters: fewer macro mysteries, faster reviews, and calmer incident pages for everyone. Restraint pays out.
Debugging Expansions: cargo expand and trace_macros in Practice
Macro errors point at code you never wrote, so the debugging workflow makes the invisible visible. Two tools carry the load: trace_macros! shows which arms match during compilation, and cargo expand prints the generated source as ordinary Rust. Used together they answer both questions every macro failure poses: which arm fired, and what code did it produce.
Start with trace_macros!(true) at the top of the failing module. Rebuild with cargo build and read the expansion log: each macro call logs its matched arm and transcribed tokens. A call matching the wrong arm is immediately obvious, as is a repetition binding fewer items than expected. Narrow the log with a small reproduction file containing only the failing invocation; full-crate traces drown the signal. Remember to remove or gate the trace behind cfg debug before merging, since the output is verbose.
Then expand. cargo install cargo-expand once, then run cargo expand for the failing target to see post-expansion source. Search the output for your generated function names and read them as plain Rust: missing commas, wrong fragment pastes, and doubled bindings all read clearly at this stage. For stubborn type errors, copy the expanded function into a scratch file and compile it directly with rustc; the resulting line numbers and E-codes reference real lines instead of macro spans.
Pin the behavior with tests at three levels. Unit tests invoke each arm and assert runtime results, covering empty, single, multi, and trailing-comma shapes. Trybuild UI tests snapshot the compiler errors for invalid invocations so message regressions fail the build. Expansion-size tests assert cargo expand line counts stay under budget, catching the next 2,400-item blowup before CI pays for it.
Build the workflow into onboarding. Every engineer who touches macros should run trace once and expand once in their first week. Those two commands convert macro debugging from folklore into procedure, and procedure is what keeps a 90-line macro from costing 34 CI minutes unnoticed. Visibility is the entire game: generated code you can read is generated code you can fix.
Trace output reading is a skill worth practicing deliberately. Each logged expansion shows the macro name, the matched arm index, and the transcribed tokens with bindings substituted. Learn to spot the arm index first: arm 0 versus arm 2 tells you immediately whether the specific or general shape matched. Then check bindings: a $x bound to an unexpected token range reveals fragment misclassification. Practice on a scratch crate with 3 arms and 6 calls before debugging production failures; the log format becomes familiar in 20 minutes and saves hours later.
Cargo expand options repay familiarity. Expand a single module or item to limit output on huge crates; use --ugly for line counts and default pretty output for reading. Diff expanded output before and after template changes to confirm exactly which generated lines moved. Store canonical expansions for the 5 most critical macros as checked-in reference files, updated intentionally. Reviewers then see expansion diffs alongside template diffs, making code generation reviewable like any other code change.
CI integration turns ad hoc debugging into continuous assurance. Record per-crate expansion line counts on every merge and graph the trend; sudden jumps flag accidental recursion or duplicated bounds. Time clean builds of macro-heavy crates with a budget comment in the pipeline definition. Fail builds when trace-gated debug macros leak into release paths. These gates cost minutes to set up and catch the 34-minute regressions while they are still 2-minute anomalies in a single pull request.
Visibility into generated code turns macro debugging from folklore into repeatable procedure for the entire team. Those two commands convert invisible code generation into readable diffs that any reviewer can verify with confidence. Generated code you can read is generated code you can fix quickly.
Recursive Macros and TT Munching: Power Tools With a Budget
Some grammars nest: balanced trees, chained builders, and s-expression DSLs cannot be described by flat repetition alone. Recursive declarative macros handle them by peeling one token tree per layer, processing the head, and re-invoking themselves on the tail. This tt-munching style is genuinely expressive: a 40-line recursive macro can parse nested configuration that would take 300 lines of builder code. It is also the pattern behind the worst compile-time blowups, so it ships with a budget.
The mechanics are simple and the cost model is not. Each recursion layer holds the remaining token buffer while expanding the current head, so processing N items touches roughly N squared tokens across all layers. At 50 items nobody notices. At 2,400 items the compiler holds gigabytes and clean builds gain half an hour. Depth limits add a second cliff: deeply nested inputs can hit the recursion limit and fail with an inscrutable overflow note. Both cliffs arrive silently because incremental builds cache the pain.
Budget rules keep recursion safe. Cap recursive macros at roughly 200 items with a compile-time assertion or a documented limit, and route larger registries to flat repetition or build scripts. Keep each layer's template minimal: emit one small item per recursion, never duplicate the tail into two branches. Test with 10, 100, and 500 items while timing clean builds; the curve reveals quadratic behavior long before production data does.
Consider the exit ramps early. Flat repetition covers most lists that look recursive at first glance. Build scripts cover registries that change weekly with thousands of entries. Procedural macros cover type-aware generation that patterns cannot express. Recursion remains for truly nested syntax such as tree literals and chained combinators where depth stays under a dozen.
Document the limit in the macro's rustdoc with numbers: supports up to 200 entries, tested to 500 at 40 seconds clean. That sentence forces the author to measure and warns the next team before they register entry 2,401. Power tools with posted limits stay useful; unlimited ones become incidents.
Recursion limits and macro depth settings interact with real inputs in ways worth testing explicitly. The default recursion limit accommodates modest nesting, but deeply nested DSL inputs can exceed it and fail with a limit-reached error that reads like a compiler bug. Raising #![recursion_limit] papers over inputs that should be restructured; better to flatten the input grammar or split large invocations across multiple macro calls. Test with the largest realistic input plus 2x headroom, and document the supported maximum in rustdoc with measured numbers.
Tail patterns and accumulator designs separate clean recursive macros from fragile ones. A muncher that carries an accumulator group, emitting completed items into it while consuming input, keeps each layer's template small and avoids duplicating the tail. One that branches the tail into two recursive calls doubles work per layer and explodes exponentially. Review recursive templates for single-tail-call shape the way you review functions for single responsibility. Two recursive calls in one arm is the code smell that precedes CI blowups.
Know when to graduate from recursion to procedural macros. Grammars needing type information, such as generating impls conditioned on field types, exceed declarative pattern power regardless of recursion budget. Grammars with operator precedence or infix nesting parse more robustly with real parsers like syn. The declarative recursion stays for simple nested shapes under a dozen levels: tree literals, chained combinators, bracketed groups. Beyond that, a proc macro with proper parsing and spanned errors serves users better than heroic pattern gymnastics.
Power tools with posted limits stay useful; unlimited recursion depth becomes an incident under production data volumes. Document the supported maximum with measured clean-build numbers for reviewers.
Shipping Macro APIs: Docs, Errors and Semver That Hold Up
Publishing a macro is publishing a language extension: every arm, fragment kind, and separator becomes a compatibility promise. Callers embed your syntax in thousands of lines, and a minor template change can break them all with errors pointing at their code. Treat macro crates with stricter discipline than ordinary libraries: document expansions, test error messages, and version syntax changes as breaking.
Documentation must show output, not just input. Each arm deserves a compilable example plus a comment sketching the expansion shape in plain Rust. Repetition macros need empty, single, multi, and trailing-comma examples. Derive macros need before-and-after field listings. Attribute macros need the wrapped function shown with its injected behavior named. Without these, users guess at semantics and file issues that docs could have prevented.
Error messages need design because the compiler reports expansion spans by default. Validate early inside the template where possible: static assertions on arity, clear panic messages naming the expected shape, and compile_error! invocations for unsupported combinations. A message reading expected status as ident = expr, got string literal saves hours versus no rules expected this token. UI tests with trybuild pin those messages so refactors cannot silently degrade them.
Semver rules are unforgiving for syntax. Adding a new arm is minor only when no existing call can match it first; arm order changes are breaking. Tightening a fragment from tt to expr is breaking. Changing expansion output that callers depend on, such as generated method names or trait bounds, is breaking even when call syntax is unchanged. When in doubt, release a major version and migration notes with before-and-after call sites.
Operate macros like infrastructure. Record expansion line counts and clean-build minutes in CI, review generated code with cargo expand on every macro pull request, and keep a downstream consumer test that builds a sample service against the new version. Those three habits caught 4 of 5 macro regressions in one surveyed team before release. Macros that ship with docs, errors, and budgets earn trust; macros without them earn reverts.
Changelog discipline for macro crates differs from ordinary libraries because callers cannot adapt gradually to syntax changes. Every release note should list added arms, removed shapes, and altered expansions separately, with before-and-after call examples for each. A minor version adding an arm includes the arm-order proof that existing calls still match identically. A major version changing expansions includes a migration script or sed recipe for mechanical call updates. Callers upgrade macro crates on trust; detailed notes are how that trust compounds.
Deprecation inside macros needs explicit machinery. Functions deprecate with #[deprecated], but macro arms have no equivalent attribute; instead, route deprecated shapes to compile_error! messages naming the replacement, or to deprecated helper functions that emit warnings when expanded code calls them. Give at least one minor version of dual support before removal. Announce upcoming removals in the previous release notes with example rewrites. The 60-service misroute incident cited earlier would have been a deprecation cycle instead of an outage with this discipline.
Ownership and review gates complete the picture. Assign each macro crate a named owner who reviews every template change and expansion diff. Require two reviewers for arm-order modifications, the highest-risk edit class. Run downstream consumer builds in CI before publishing, not after, so breakage surfaces in the macro repository. These gates feel heavy until the first prevented fleet-wide misroute; after that they feel like the cheapest insurance on the release calendar.
Macros that ship with docs, errors, and budgets earn trust; macros without them earn reverts and migration work. Detailed notes are how caller trust compounds across many releases.
The Logging Macro That Added 34 Minutes to Every CI Build for 6 Weeks
- Recursive tt-munching over large lists is quadratic: 180 events feel free while 2,400 events cost 34 minutes, so prefer flat $(...),* repetition and measure expansion with cargo expand counts in CI.
- Clean-build time is the metric that matters for releases: incremental caching hid a 6-week regression from every developer, so gate merges on a timed clean build of macro-heavy crates.
- Generate validation, not code, at scale: one derived impl plus a 300-line build script replaced 2,400 generated impls and kept per-event cost near zero as the registry grows past 3,000 names.
| File | Command / Code | Purpose |
|---|---|---|
| io | macro_rules! check { | macro_rules! Matchers |
| io | macro_rules! sum_all { | Repetition With Separators |
| io | macro_rules! emit_pair { | Hygiene and local_inner_macros |
| io | use std::fmt::Debug; | Four Macro Kinds |
| io | use std::collections::HashMap; | Deriving Standard Traits |
| io | use serde::{Deserialize, Serialize}; | Serde Derive in Production |
| io | use std::fmt::Display; | When Not to Macro |
| io | macro_rules! pick { | Debugging Expansions |
| io | macro_rules! count_tt { | Recursive Macros and TT Munching |
| io | macro_rules! status { | Shipping Macro APIs |
Key takeaways
Common mistakes to avoid
7 patternsReaching for macro_rules! before trying generics or functions
Writing expr arms that must also accept types or patterns
Recursive tt-munching over lists with hundreds of entries
Forgetting trailing-comma and separator shapes in repetitions
Macro helpers that resolve locally but break downstream
Letting macro locals collide with caller names (hygiene bugs)
Deriving and hand-implementing the same trait on one type
Interview Questions on This Topic
How does a macro_rules! matcher arm work, and what is $x:expr?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's Core. Mark it forged?
25 min read · try the examples if you haven't