Rust Testing: cargo test, Clippy Lints, and CI Gates That Work
Put unit tests in-module, integration tests in tests/, doc-test examples, and gate merges on clippy -D warnings, rustfmt, plus CI matrix..
20+ years shipping production backend systems. Drawn from code that ran under real load.
- ✓Comfortable writing basic Rust (functions, structs, Result handling)
- ✓Cargo installed via rustup with a working build (cargo build succeeds)
- ✓A GitHub account and a repo where you can run Actions workflows
- Rust has three native test layers: unit tests inside src/ under #[cfg(test)], integration tests in tests/ that use only your public API, and doc-tests compiled straight from ``` examples in doc comments
- Run them with cargo test; narrow the run with filters like cargo test transfer, split harnesses with --lib --tests --doc, and control parallelism with -- --test-threads=4 or -- --nocapture when you need the println! output
- Mark known-failing paths with #[should_panic(expected = "...")] and park slow or flaky suites behind #[ignore], then run them on purpose with cargo test -- --ignored
- Build fixtures from plain setup functions that return values or guard structs with Drop — no framework needed for 90% of suites
- Gate every merge on cargo clippy --all-targets --all-features -- -D warnings plus cargo fmt --check in a GitHub Actions matrix over stable and beta, and track (never worship) coverage with cargo llvm-cov
Think of a restaurant kitchen before dinner service. The cooks taste each sauce in its own pan (unit tests), then they run a full fake dinner service with real plates and real tickets to see if the whole line works together (integration tests), and the recipes pinned on the wall get cooked exactly as written to prove the instructions are correct (doc-tests). A strict head chef watches every plate and rejects anything sloppy before it leaves the kitchen (Clippy), someone else checks every plate looks identical (rustfmt), and the whole rehearsal repeats automatically every time anyone changes a recipe (CI).
You've shipped your first Rust crate. It compiles, clippy is quiet, and you're feeling good. Then a refactor breaks a helper three modules deep, nothing catches it, and you find out from a user at midnight. That's the moment most Rust developers stop treating tests as homework and start treating them as load-bearing walls.
Rust makes the first step almost unfairly easy. There's no framework to install, no config file to bless. You slap #[test] on a function, run cargo test, and the compiler builds a harness, runs everything in parallel, and reports failures with line numbers. It feels like the language wants you to test, because it does.
But the easy start hides real decisions. Code that lives in src/ can touch private functions, while code in tests/ can only use your public API — and picking the wrong home for a test costs you either coverage or honesty. Doc examples silently become tests, which is wonderful until one rots and blocks a release. Clippy's 700-plus lints can save you or drown you, depending on how you configure them.
Don't learn these lessons from a midnight page. We'll walk through each test layer with compilable examples, show how should_panic and ignore handle the awkward cases, build fixtures from plain functions, and wire the whole thing into a GitHub Actions matrix that runs clippy, fmt, and tests on every push. By the end you'll have a setup that catches regressions while they're still cheap.
Unit Tests Live Next to the Code: #[test] Inside #[cfg(test)]
Unit tests in Rust live in the same file as the code they verify, tucked inside a mod tests block gated by #[cfg(test)]. That placement is the whole trick. Code inside the module can call private functions directly, so you can pin down edge cases in a parser or a rounding helper without making anything public just for testing. The cfg gate means the test module compiles out of release builds entirely — zero binary cost, zero runtime cost, no dead code shipped to users. When you run cargo test, the compiler builds your crate twice: once as a library and once as a test harness with each #[test] function registered as a case.
The mechanics reward small, boring tests. Each #[test] function returns () and panics on failure, which is why plain assert_eq! does most of the heavy lifting. A failure prints the left and right values plus the exact source line, so a message like assertion failed: left: 3, right: 4 at src/cart.rs:42 tells you everything without a custom matcher library. Keep tests focused on one behavior each. A test named applies_tiered_discount_at_boundary is worth five tests named test1 through test5, because the name shows up in failure output and becomes the first line of your debugging session.
Panic messages deserve the same care as the assertions. The format string in assert!(balance >= 0, "balance went negative after refund of {amount}") costs nothing when the test passes and saves ten minutes when it fails at 1 AM. For floating point, never assert exact equality — compare with an epsilon like (got - want).abs() < 1e-9. For Results, prefer assert!(r.is_ok()) or unwrap into a match; calling .unwrap() inside tests is idiomatic and the panic output includes the Err value automatically.
Organization scales with a simple rule: one tests module per source file, nested deeper only when a submodule has real complexity. A tests module at the bottom of src/pricing.rs covering that file's five functions beats a single giant tests module at the crate root that nobody can navigate. Shared helpers used by several test modules belong in a private test_support module inside src/, also cfg-gated, so integration tests in tests/ cannot accidentally depend on them.
Watch the compile-time bill as the suite grows. Every #[cfg(test)] module still compiles its dependencies when you run cargo test, and heavyweight dev-dependencies like full HTTP clients can double harness build time. Prefer tiny fakes over real servers in unit tests — they run in microseconds, never flake, and fail with messages you wrote yourself. Save the real database and the real network for the integration layer, where their cost buys honesty you cannot fake.
Test doubles in unit tests should be boring on purpose. A fake payment gateway that returns canned approvals from a HashMap will outperform a mock framework in readability every time, because the behavior sits in plain code the whole team can read. When a function takes a trait, write a three-line struct implementing it for tests instead of pulling in a mocking crate — the manual fake compiles faster, fails with messages you control, and never breaks because a macro updated. Save mocking libraries for the rare case where a trait has twelve methods and your test cares about call ordering across all of them.
Result-heavy code benefits from a small assertion vocabulary used consistently. Use assert_eq! for values, assert!(matches!(...)) for enum shapes, and assert!(err.to_string().contains(...)) for error surfaces. Consistency matters because failure output becomes predictable: every developer knows where to look and what the message shape will be. Add one custom assertion helper per crate at most — something like assert_cents_eq! — and only when three or more test modules repeat the same comparison logic with the same tolerance.
Snapshot testing earns its place for complex serialized output. When a function's return is a fifty-line config rendering, asserting the full string inline buries the test; writing it to a checked-in snapshot file with a crate like insta keeps the test readable while still catching drift. Review snapshot updates the way you review code — a changed snapshot is a behavior change wearing a diff costume. Reject snapshots for simple values where plain assert_eq! reads better; the tool should reduce noise, not add ceremony.
Integration Tests in tests/: Black-Box Layout That Scales
Everything under tests/ compiles as its own crate that links against your library's public API — and only the public API. That restriction is a feature, not a limitation. Code in tests/checkout.rs cannot call your private helpers even if it wants to, so every integration test exercises the same surface your users touch. When a refactor renames internals but keeps behavior, white-box unit tests may need edits while black-box integration tests pass untouched. That stability is exactly what you want from the suite that guards releases.
Layout follows a convention that scales to dozens of files. Each .rs file directly under tests/ becomes a separate test binary with its own main harness, while subdirectories do not compile unless pulled in with a mod declaration. So tests/checkout.rs and tests/refunds.rs are two binaries, and tests/common/mod.rs holds shared fixtures imported via mod common;. The common directory pattern matters: a file at tests/common.rs would itself compile as a test binary (usually with zero tests and a warning), while tests/common/mod.rs compiles only as a helper module. Get this right once and your tests/ directory stays navigable at 40 files.
Each binary links the library fresh, which has two consequences worth internalizing. First, integration tests cannot see #[cfg(test)] helpers inside src/ — duplicate the tiny fixtures you need in tests/common/ instead of widening visibility. Second, binaries run in parallel with each other, so two files that bind port 8080 or write /tmp/report.csv will collide. Give every test its own temp dir, its own random port, or use a lock for the genuinely shared resources.
Keep integration tests behavior-shaped, not implementation-shaped. Assert on returned values, status codes, and file contents — never on internal call counts or private struct fields. A test that posts an order and asserts the confirmation email file exists will survive three refactors of the mailer. A test that reaches into the mailer's queue length will break on the first cleanup commit and teach the team to distrust the suite.
Speed is the failure mode to plan for. Integration binaries that hit real Postgres or real S3 take seconds each, and seconds times forty files becomes a CI job nobody waits for. Split the layer: fast black-box tests against in-memory fakes run on every save, while the dozen tests needing real services carry #[ignore] locally and run in CI against containers. The directory layout stays identical; only the execution policy differs.
Naming the test binaries deliberately pays off as the directory grows. A file named checkout_refunds.rs reads better in cargo test output than helpers2.rs, and the binary name appears in every failure line CI prints. When two binaries need the same fixture, resist duplicating it — promote the helper to tests/common/ immediately, because duplicated fixtures drift within weeks and then two tests assert subtly different versions of the same setup. The common module is also the right home for deterministic builders like confirmed_order() that five files share.
External services in integration tests deserve containers, not shared staging. A postgres:16 service container started by CI gives each run a private database that can be migrated, seeded, and destroyed without affecting anyone else. Tests connect through an env var like TEST_DATABASE_URL with a localhost default, so developers run the same suite against a local container with one docker run command. The suite stays honest about real I/O while remaining reproducible — the two properties that shared staging environments always trade against each other.
Feature-gated tests keep the matrix honest without multiplying CI jobs. Gate backend-specific suites with #![cfg(feature = "postgres")] so cargo test runs the default path fast while cargo test --all-features exercises everything before merge. Document the mapping between features and test files in the module header so contributors know which invocation covers their change. A test that only runs under a feature nobody enables locally will rot first, so the all-features CI job is the backstop that keeps those files alive.
Doc-Tests: Every Example You Write Is a Test Running in CI
Rust compiles every fenced code block in your doc comments and runs it as a test. That single design decision does more for documentation quality than any style guide. An example that claims client.connect() returns a Session either keeps compiling or fails the build — there is no third option where docs quietly rot for two years. Teams that embrace this get living documentation: every release proves its own examples still work.
The harness wraps your snippet in fn main() automatically, with extern crate wiring handled for the 2024 edition. A block tagged rust-fenced blocks becomes a full program; add use shop::Cart; at the top and write straight-line code ending in assertions. Compilation errors surface exactly like test failures, with the doc path as the test name — shop::Cart::total (line 18) tells you which comment to open. Run cargo test --doc to execute only this layer when you are editing docs, which keeps the loop tight.
Attributes tune the behavior for honest edge cases. Mark snippets that should not compile with compile_fail blocks to lock in a desired error — for example, proving that a negative Money value cannot be constructed. Use no_run blocks for examples with side effects you want compiled but not executed, like a snippet that binds a socket. Add ignore blocks only for examples needing unavailable hardware, and treat each one as debt, because an ignored example is documentation nobody verifies.
Hidden lines keep examples readable without lying. A line starting with # inside the block compiles but stays invisible in rendered docs, so # use shop::Cart; sets up imports without cluttering the page. Readers see three clean lines; the test runs five. This is how you keep docs skimmable and tests complete at the same time.
The failure mode is version drift in prose around the code. The harness checks code, not sentences, so a paragraph claiming milliseconds while the example passes seconds stays green and wrong. Review doc sentences with the same care as the snippet when signatures change. A doc-test proves the code runs; only human review proves the words still describe it.
Choosing which examples become doc-tests is a judgment call worth making explicitly. Public API that users must call correctly — builders, parsers, client constructors — deserves runnable examples, because the doc-test then guards the exact code path newcomers copy. Internal helpers and trivial getters do not need examples; forcing one onto every private function produces noise that slows cargo test --doc without teaching anything. Aim doc examples at the reader who opened the docs page confused, and the tests will cover the paths that matter most.
Longer workflows belong in a tests/ file with prose pointing at them, not in a forty-line doc comment. A doc example should fit on one screen: setup, one call, one assertion. When the realistic scenario needs a running server or three chained calls, write the full version as an integration test and keep the doc example to the smallest honest fragment. Link between them with a sentence naming the test file, so readers can graduate from the taste to the meal without guessing where it lives.
Test the failure examples too, not just the happy path. A compile_fail block proving that negative money cannot be constructed is as valuable as the example showing normal construction, because it locks the type-system guarantee users rely on. Keep compile_fail snippets tiny — one construction, one error — since large failing programs produce brittle stderr assertions across toolchain versions. When the compiler's message wording shifts under a new stable, update the expected error promptly rather than deleting the test and losing the guarantee.
Render checks catch the rest. Periodically read the rendered docs for your most-visited items and confirm the prose still matches the tested code. A screenshot-free habit — open the docs page, run the example mentally, check the surrounding sentences — takes five minutes per item and keeps words and code telling the same story to every newcomer who lands there.
#[should_panic] and #[ignore]: Testing Failure Paths Without Slowing the Suite
Some behaviors only show up when things go wrong, and others take too long to run every time. Rust gives you two attributes for exactly these cases, and using them well separates a suite developers trust from one they work around. #[should_panic] asserts that code fails in the way you designed — a divide-by-zero guard, an index check, an invariant violation. #[ignore] parks a test so the default run skips it, keeping slow or environment-dependent suites out of the inner loop while still runnable on demand.
should_panic works best with an expected substring. Writing #[should_panic(expected = "capacity exceeded")] pins the failure message, so a panic with a different message — a real bug wearing the costume of an expected failure — still fails the test. Without expected, any panic passes, including the ones from an unrelated unwrap three calls deep. Keep the substring specific enough to identify the guard but stable enough to survive refactors; matching on a full sentence with punctuation invites churn.
Unwrap-heavy code needs care around should_panic. A test that calls build("").unwrap() and expects a panic passes for the wrong reason if unwrap panics on None instead of your validation firing. Prefer asserting on the Result directly — assert!(matches!(build(""), Err(BuildError::EmptyName))) — and reserve should_panic for code that genuinely panics by contract, like indexing or explicit panic! guards. The goal is a test that fails when the guard disappears, not one that passes no matter what breaks.
ignore is a scheduling tool, not a trash can. A 40-second end-to-end test behind #[ignore] runs in CI with cargo test -- --ignored while developers iterate in seconds locally. The danger is the quiet pile-up the incident story showed: ignored tests nobody runs become lies about coverage. Every ignore needs either a tracking issue link in a comment or a CI job that runs the ignored set, ideally both.
Pair the two attributes with explicit intent comments. A one-line comment like // Slow: spins up a local SMTP sink; runs in CI ignored-job explains the why to the next reader and resists drive-by deletion. Review ignored tests quarterly the way you review dependencies: re-enable what got fast, fix what got flaky, delete what stopped mattering.
Timeouts deserve the same explicit treatment as panics and ignores. A test that waits on a channel or a spawned thread should wrap the wait in recv_timeout with a named duration, then assert with a message describing what never arrived. Bare recv() in a test is a hang waiting for a CI timeout — forty minutes of billable runner time ending in a useless Canceling log. The timeout converts hangs into failures with names, and named durations make the suite's patience visible during review.
Document the reason beside every attribute, not just the attribute itself. A comment like panics by design: zero capacity would silently admit all traffic explains the contract to the next reader and survives refactors that would otherwise delete a mysterious annotation. Review these comments when the surrounding code changes. An expected message that no longer matches any panic! string in the source is a test guarding a ghost — update the string or remove the test before it teaches the suite to cry wolf.
Property tests complement example tests where input spaces are large. A quickcheck-style test asserting that total_cents never exceeds its input for any u64 runs thousands of cases in milliseconds and finds the overflow your three hand-picked examples missed. Cap the case count for local runs and raise it in nightly CI so developers stay fast while the long tail still gets explored. When a property test finds a failure, shrink the input and freeze the minimal case as a plain regression test so the bug stays caught even if the property harness changes.
Reserve bare should_panic for code that panics by contract, such as explicit invariant guards and indexing paths. Everywhere else, match on the Result so the test names the error variant it requires. That discipline keeps panic tests rare, loud, and meaningful — a failing one always signals a broken promise rather than background noise the team learned to ignore.
Fixtures Without Magic: Plain Setup Functions Over Hidden Frameworks
Newcomers to Rust testing often ask which fixture framework to install. The senior answer is usually none. A plain function that builds a value and returns it covers most needs with zero dependencies, zero macros, and zero hidden lifecycle. fn sample_cart() -> Cart constructs a cart with two items; every test calls it, each gets a fresh value, and ownership rules guarantee tests cannot contaminate each other. Fresh values per test beat shared fixtures in almost every Rust suite because the borrow checker already prevents the aliasing bugs that make shared fixtures scary in other languages.
When setup needs teardown — temp directories, test databases, spawned servers — the Drop trait is your after-each hook. A guard struct like TempDir holds a path, its constructor creates the directory, and its Drop impl removes it. Tests bind let _guard = TempDir::new(); and cleanup runs even if the test panics, because Drop runs during unwinding. No framework, no registration, no ordering puzzles. The same pattern wraps database transactions: begin in the constructor, roll back in Drop, and every test runs in isolation against a pristine snapshot.
Builders handle combinatorial setup without boolean soup. Instead of setup_complex(true, false, true), expose Cart::fixture().with_discount(0.1).with_tax(0.07).build() so each test states only what it varies. Default values live in one place, tests read like specifications, and adding a new option never breaks existing call sites. This pays off fastest in parser and config tests where valid inputs need ten fields but each test cares about one.
Share fixtures through modules, not macros. A tests/common/mod.rs with pub fn funded_account() -> Account beats a fixture macro because go-to-definition works, compile errors point at real code, and new hires can read it without learning a DSL. Reserve external crates like rstest or once_cell for genuine pain — parameterized cases with 30 rows, or one truly expensive immutable resource — and even then isolate them to the files that need them.
Keep fixtures honest about cost. A setup function that sleeps 200 milliseconds for eventual consistency multiplies into minutes across 300 tests. Prefer deterministic construction: fake clocks, in-memory channels, seeded RNGs. When real waiting is unavoidable, centralize it in one helper with a timeout and a clear assertion message so the eventual failure says waited 5s for worker, got 0 jobs instead of a bare timeout panic.
Randomness in tests must be seeded and the seed must be visible. A property test or fuzzed input generator running on a random seed finds different bugs on every run, which sounds good until CI fails on a seed nobody can reproduce. Pass the seed explicitly — SmallRng::seed_from_u64(42) — and print it in the failure message so any developer can replay the exact failing case. Deterministic-by-default with an opt-in nightly random sweep gives you reproducibility on every commit and exploration where it is cheap.
Layer your fixtures by cost so the suite stays fast as it grows. Pure constructors returning owned values cost microseconds and belong everywhere. File-system fixtures cost milliseconds and belong in tests that genuinely need paths. Container-backed fixtures cost seconds and belong behind ignore or a feature flag, running in CI only. When a test uses a heavier layer than its assertions require, demote it during review. Suite speed is a feature, and fixture cost is where that feature is won or lost.
Test-only dev-dependencies deserve the same review bar as production ones. A fixture crate pulled in for one helper adds compile time to every cargo test invocation and a supply-chain entry to every audit. Prefer std-only fixtures, and when a dev-dependency genuinely pays — a snapshot crate used by thirty tests, a fake-clock crate used by fifty — record the reason in a comment near the Cargo.toml entry. Quarterly, check whether the dependency is still pulling its weight or whether three plain functions could replace it.
cargo test Like a Pro: Filters, Threads, and Output You'll Actually Read
cargo test has a split personality: arguments before the -- belong to Cargo (which targets to build), and arguments after belong to the libtest harness (how to run). Internalizing that split turns a blunt command into a precision tool. cargo test checkout --test integration runs only the checkout filter inside the integration binary. cargo test --lib runs unit tests without touching integration binaries. cargo test --doc runs only doc examples. Most developers who complain the suite is slow are running all three layers when they need one.
Filters match substrings of the full test path, so cargo test pricing::discount runs every test under that module while cargo test -- --exact pricing::discount::rounds runs exactly one. Combine with --test target to scope further: cargo test burst --test limiter hits only burst tests in the limiter binary. The -- --list flag prints every test path without running anything, which is the fastest way to check your filter before a long run. Pair it with grep to audit naming: cargo test -- --list | grep ignore finds nothing, but the count line tells you the suite shape.
Output capture is the default that confuses everyone once. Passing tests swallow their println! output; only failures print. When you need to see inside a passing test, cargo test -- --nocapture shows everything, and -- --nocapture --test-threads=1 keeps the lines in order instead of interleaved across 8 threads. For structured logs, Nocapture plus an eprintln! in the code under test beats printf debugging through three layers of harness buffering.
Thread control is both a diagnostic and a policy. The harness defaults to one thread per CPU core, which is right for CPU-bound unit tests and wrong for tests sharing a port or a file. Diagnose flakiness with cargo test -- --test-threads=1: if serial passes and parallel fails, you have shared state, not bad logic. Fix the sharing rather than pinning threads in CI, but know that --test-threads=4 is a legitimate middle ground for suites mixing fast unit tests with a few resource-hungry integration cases.
Reporting flags close the loop. cargo test -- --format=terse trims per-test lines in CI logs, -- --report-time prints durations so you can find the three tests eating half the job, and -- --skip slow skips name patterns without touching attributes. Chain them: nightly CI runs the full suite with --report-time, and the slowest test each week gets either a fake clock or an ignore with a tracking issue. Suites stay fast because slowness is measured, not felt.
Junit output bridges Rust suites into the dashboards managers actually open. Adding -- -Z unstable-options --format json with a conversion tool, or a runner like cargo-nextest with its junit reporter, turns per-test results into the XML your CI dashboard renders as pass-rate graphs. nextest deserves special mention: it runs each test binary in parallel with per-test timeouts, retries flaky tests a configured number of times with quarantine tracking, and archives failing output separately. Teams with suites over five minutes usually recover the migration cost within a month through faster feedback alone.
Environment parity closes the last gap between local green and CI red. Run the exact CI command locally at least once per month — the same --locked flags, the same RUSTFLAGS, the same toolchain — because drift between make test shortcuts and the real pipeline is where phantom failures breed. A just test-ci recipe or an xtask alias that shells out to the identical cargo invocation removes the temptation to approximate. When CI and local runs disagree, suspect the environment first, the lockfile second, and your code last; that ordering resolves most mysteries within the hour.
Teach the team three commands and the suite stays healthy: the filtered loop for editing, the full run before pushing, and the ignored-job command before leaving for the day. Post them in the repo readme so onboarding installs the habit on day one. Developers who know the fast loop run tests ten times a day; developers facing a fourteen-minute all-or-nothing gate run them the night before the deadline. The command list is culture encoded as text, and it costs four lines to write.
Clippy Lints: Your Strictest Reviewer, Plus --fix and -D warnings
Clippy ships with the toolchain and carries over 700 lints across correctness, suspicious patterns, style, complexity, and performance. Running it is one command — cargo clippy --all-targets --all-features — but running it well is a policy. The flags matter: --all-targets includes tests, benches, and examples, which is where needless clones love to hide; --all-features builds every feature combination so a lint cannot sneak in behind a disabled flag. Without both, CI checks a subset and developers learn the gate means nothing.
The -D warnings argument turns every warning into a hard error. That single flag is the difference between a linter and a gate. Without it, new lints arrive with each toolchain release as yellow text nobody reads; with it, the build breaks and someone fixes the code or writes a scoped allow with a reason. Apply -D warnings in CI, not necessarily on every local save — developers iterating on a prototype do not need a denied build for a dbg! they will delete in ten minutes.
cargo clippy --fix applies machine suggestions directly, and --allow-dirty lets it run with uncommitted changes during a cleanup pass. The workflow: commit first, run cargo clippy --fix --all-targets, review the diff hunk by hunk, then run the tests. Most suggestions are safe (collapsing needless returns, removing redundant clones), but a few change semantics subtly — len() == 0 to is_empty() is fine, while a suggested iterator rewrite in hot code deserves a benchmark glance. Never --fix blindly on a dirty tree; the diff is the review.
Configuration lives in Clippy.toml and lint attributes, and restraint is the skill. Crate-level allows like #![allow(clippy::too_many_arguments)] on a seven-parameter constructor may be honest; sprinkling allows to silence pedantic lints you never enabled is noise. Start with the default warn set plus -D warnings, add nursery or pedantic groups only if the team agrees to maintain them, and document every workspace-level allow with the reason it exists. A allow without a comment is a mystery the next hire gets to solve.
Treat new-toolchain lint breakage as scheduled maintenance, not surprise. Every six weeks stable Rust can promote lints, and a green gate turns red with zero code changes. The fix is process: pin CI to a known stable, run a non-blocking beta clippy job to preview incoming lints, and fix them during the week instead of on release day. Teams that do this spend twenty minutes a month on lint upkeep; teams that do not spend a frantic afternoon.
Pedantic and nursery lint groups are worth a deliberate team decision, not a drive-by enable. The pedantic group flags patterns like must_use candidates and needless lifetimes that improve API quality but demand real annotation work across a large codebase. Nursery lints are explicitly experimental and can change meaning between releases, which makes denying them in CI a source of surprise breakage. A pragmatic path: deny the default warn set plus correctness-adjacent groups in CI, run pedantic as warnings during review weeks, and promote individual pedantic lints to denies only after fixing the whole codebase once.
Document lint exceptions where they live, with the reason attached. A #[allow(clippy::too_many_arguments)] on a constructor that genuinely needs seven parameters is honest engineering; the same allow without a comment is a puzzle for the next reader. Crate-level allows in lib.rs should be rare enough to count on one hand — each one disables a check for every line in the crate, so prefer item-level allows that bound the blast radius. During review, treat a new allow the way you treat a new unsafe block: small, justified, and impossible to miss.
Run clippy on the diff during review, not just in CI. A comment pointing at a specific lint — clippy::needless_collect flags the intermediate Vec here — teaches the author while the code is still warm. Over months this builds a team that writes lint-clean code on the first pass, which is the only way the gate stays painless at scale. CI catches what humans miss; humans teach what CI cannot explain. Both directions matter, and the review comment is where the teaching happens.
rustfmt: End Style Arguments Forever With One CI Gate
rustfmt is the rare tool that ends arguments instead of starting them. It owns line width, indentation, import order, and trailing commas, and its output is the definition of correct Rust style. There is nothing to configure for most teams — rustfmt.toml exists, but every option you set is a debate you chose to keep having. The default formatting plus a CI gate resolves thousands of micro-decisions per pull request and gives reviewers their attention back for logic.
The workflow has two commands and one rule. cargo fmt formats the workspace in place; cargo fmt --check verifies formatting and exits nonzero on any diff, printing exactly what would change. CI runs --check and fails the job on drift. Developers run cargo fmt before committing, ideally via a pre-commit hook or editor format-on-save, so the gate never fires on their pull request. When --check fails in CI, the fix is mechanical: pull, run cargo fmt, push. No discussion, no style comments, no seniority spent on brace placement.
Edition awareness matters in 2024. rustfmt formats according to the crate's edition, and edition 2024 changed details like let-chain layout and gen-block spacing. A crate mid-migration can show surprising diffs if the toolchain and the edition flag disagree, so keep rust-toolchain pinned and the edition field in Cargo.toml deliberate. Mixed-edition workspaces are fine — rustfmt reads each crate's own edition — but every crate should declare one explicitly rather than inheriting ambiguity.
The check-everything instinct needs one boundary: generated code. Files produced by build scripts or protobuf compilers should carry #[rustfmt::skip] at the top or live behind a rustfmt.toml ignore, because reformatting generated output creates diff noise and merge conflicts with the generator. Everything humans write stays formatted; everything machines write stays skipped. Reviewers should never see a 2,000-line generated diff because someone ran cargo fmt without the skip.
Treat formatting as hygiene, not virtue. rustfmt will occasionally format a match arm or a builder chain in a way that reads worse to human eyes. Accept it anyway. The value of a formatter is uniformity across 40 contributors and 200,000 lines, not perfection in any single function. Teams that allow hand-tuned exceptions accumulate them; teams that accept the tool's output stop thinking about style entirely, which was the point.
Check-only verification belongs in CI, but local hooks make it painless. A pre-commit hook running cargo fmt --check on staged files catches drift before the push, and editors with rust-analyzer format-on-save remove the failure class entirely. The hook should be fast and quiet: check, reject with the diff, and let the developer run cargo fmt once. Hooks that auto-modify staged files create confusing partial commits, so prefer the check-and-tell pattern over silent rewriting.
Diff hygiene keeps formatting changes reviewable. Never mix cargo fmt fallout with logic changes in one commit — run the formatter first, commit the result alone with a message like fmt: edition 2024 baseline, then build the feature on top. Reviewers can then verify the format commit touches only whitespace and spend their attention on behavior. The same rule applies to the one-time migration when adopting rustfmt on a legacy codebase: one formatting commit, reviewed for diff-only purity, followed by normal work that stays clean forever.
Onboard new hires with the formatter, not against it. The first-day setup guide should enable format-on-save before the first commit, so nobody's opening pull request fails on brace placement. Frame rustfmt as team infrastructure during onboarding — the tool that lets forty people edit the same files without style collisions — and newcomers adopt it gladly. Nobody ever quit over trailing commas they never had to think about.
Measure the win so the habit sticks. Track review comments per pull request before and after enforcement; most teams watch style nits vanish within a quarter while time-to-merge drops by a third. Share the numbers once, then let the quiet speak. A formatter nobody discusses is a formatter doing its job, freeing every review for the logic that actually ships.
GitHub Actions CI Matrix: Stable, Beta, Clippy, and Fmt Gates
A Rust CI pipeline has a standard shape that has emerged across thousands of repos: a matrix over stable and beta toolchains running tests, a dedicated clippy job with denied warnings, a fmt check, and aggressive caching of the target directory and registry. The matrix exists because stable proves your code works today while beta previews what breaks in six weeks. Make beta non-blocking (continue-on-error) so it informs without holding releases hostage, and make stable blocking so nothing merges red.
Caching decides whether CI takes 4 minutes or 19. The Swatinem/rust-cache action keyed on Cargo.lock restores compiled dependencies across runs, and --locked in every cargo invocation guarantees CI builds the exact dependency tree developers tested. Without --locked, a fresh semver-compatible release can change behavior between your local run and the CI run, producing the dreaded passes-here-fails-there that eats afternoons. Commit Cargo.lock for binaries and services; the lockfile is what makes CI reproducible.
Job separation keeps failures legible. One job runs cargo test --locked --all-features on the matrix; a second runs clippy with -D warnings pinned to stable; a third runs fmt --check. When clippy fails, the log shows only clippy — not 400 lines of test output above it. Fail-fast false on the matrix lets you see both stable and beta results even when one breaks, which matters exactly when a toolchain release lands mid-week.
Concurrency control protects both your runners and your sanity. A concurrency group keyed on ref with cancel-in-progress cancels stale runs when a developer pushes three commits in ten minutes, so the queue stays short and the signal stays fresh. Pair it with explicit permissions (contents: read) and pinned action versions — @v4 or a full SHA — because an unpinned action updating under you is an outage you scheduled yourself.
Secrets and environments need the same discipline as code. Database-backed integration tests should spin up services: postgres:16 containers in the job rather than reaching for a shared staging database that another run might wipe. Test secrets live in GitHub encrypted secrets, never in the yaml, and the workflow file itself gets the same review bar as source code. A CI pipeline that touches production credentials without review is a deployment pipeline with extra steps.
Branch protection turns the pipeline from advice into law. Require the stable test job, the clippy job, and the fmt job as status checks before merging, and dismiss stale approvals on new pushes so a green review cannot bless untested code. Dependabot or Renovate should update actions and the toolchain pin on a schedule, with the beta job absorbing the surprise so stable updates land as routine pull requests. Teams that skip protection discover that CI exists but merges happen anyway — the worst of both worlds.
Observability on the pipeline itself prevents slow rot. Track median job duration per week; a test job creeping from 5 to 11 minutes is a bug report about the suite, not background noise. Archive --report-time output or nextest junit files as build artifacts so anyone can name the slowest tests without re-running anything. Review the ignored-test count and the beta-job failures in the same weekly slot as flaky-test triage. A pipeline nobody watches becomes decoration within a quarter; a pipeline with an owner stays fast for years.
Keep the workflow file boring and the tooling pinned. Renovate or Dependabot pull requests bumping actions/checkout or rust-cache get the same test run as code changes, because a broken action version breaks every pipeline at once. Review workflow edits with two eyes on permissions and secrets handling. The CI file is production infrastructure that happens to be written in YAML, and it deserves the senior-review bar that status implies.
Test the pipeline itself when it changes. A workflow edit that adds a job or bumps a pin deserves a draft pull request whose CI run you watch end to end before merging. Broken CI blocks every developer at once, so the pipeline file carries more blast radius per line than almost any source file. Review it with that weight in mind.
Coverage With cargo-llvm-cov: Numbers That Guide, Never Vanity Metrics
Coverage answers one question well — which lines never execute under test — and every other question poorly. Teams that treat it as a treasure map find untested error branches and dead code within days. Teams that treat it as a target hit 90% with assertion-free tests that execute everything and verify nothing. The tool is cargo-llvm-cov, the honest workflow uses it to find gaps, and the number on the badge is the least interesting output.
Setup is two installs: the llvm-tools-preview component plus the cargo-llvm-cov wrapper. From there cargo llvm-cov --all-features --workspace --lcov --output-path lcov.info produces both terminal summaries and files your CI can archive. Add --doctests when your public examples carry contract weight, and --ignore-filename-regex to exclude generated code that would otherwise drag the percentage without meaning. The HTML report (cargo llvm-cov --open) paints uncovered lines red — open it, click the reddest file, and you have next week's test plan.
Enforcement needs a light touch. A CI step that fails when total coverage drops by more than a point catches the pull request that adds 400 untested lines; a hard 85% gate blocks refactors that delete well-tested code and add small clean code, punishing exactly the changes you want. Prefer diff coverage — new lines must be covered — over absolute thresholds, and exempt the genuinely untestable (platform-specific syscalls, defensive unreachable! branches) with explicit annotations rather than gaming the denominator.
Branch coverage matters more than line coverage for Rust's favorite constructs. A match with five arms where tests hit two shows 100% line coverage on the match keyword while three behaviors go unverified. llvm-cov's branch data exposes this; when a match, if-let chain, or combinator pipeline shows partial branches, write the missing cases before chasing any global number. Error paths deserve first priority — the Err arm you never test is the one that pages you.
Combine coverage with mutation thinking for the final mile. If a covered test still passes after you flip < to <= in the source, the test executes the line without verifying it — coverage lied. You do not need a mutation framework to apply this: change one operator, run the focused test, confirm red, revert. Do that for the five most critical functions and your suite earns trust no percentage can buy.
Triage the report on a cadence, not on vibes. A thirty-minute weekly slot where one engineer opens the HTML report, sorts by uncovered lines, and files test tasks for the reddest error-handling paths beats every coverage badge ever shipped. File the tasks with the file and line range attached so they are actionable without re-derivation. Over a quarter, this habit typically lifts meaningful coverage — error arms and partial branches — by double digits while the headline number barely moves, which is exactly the trade you want.
Know what to exclude so the number stays honest. Generated code, build-script output, and platform-specific shims behind cfg(windows) inflate the denominator without representing behavior anyone can test on a Linux runner. Exclude them with --ignore-filename-regex and document the pattern in the coverage job's comments so the next engineer does not re-add them chasing a dip. What remains is code your team owns and can test — the only denominator worth measuring against.
Present coverage to the team as a map, never as a grade. The weekly triage message should read like a treasure note — 212 uncovered lines in error paths, three already-fixed bugs hiding among them — rather than a scoreboard nobody wants to top. Celebrate the deleted untestable code that raises honest coverage and the new assertions that turn red lines green. Teams that hunt gaps together write better tests than teams ranked against a percentage.
Delete tests that coverage proves redundant. When two tests execute identical lines and assert identical behavior, keep the clearer one and remove the other — suite time is a budget and duplication spends it twice. The HTML report shows overlap plainly: identical green blocks across sibling tests are consolidation candidates. Fewer, sharper tests beat a large suite nobody fully understands.
The #[ignore]d Rate-Limiter Test That Cost 41,000 Throttled Checkouts
throttle() example so the contract is verified in two harnesses at once.- Treat ignored tests as production debt with an owner and an expiry date, not as a quiet delete. A CI step that lists ignored tests on every run — or fails when a new one appears — would have surfaced the parked burst test within a day instead of six weeks and 214 green runs later.
- Flaky tests should be made deterministic, not silenced. Replacing wall-clock sleeps with a fake clock cut the burst test from 4.1 seconds to 90 milliseconds and removed the flakiness at its root; the ignore was treating the symptom while the disease stayed in the suite.
- One behavior deserves verification in two harnesses when money flows through it. The throttle contract now has an integration test for burst shape plus a doc-test on the public example, so a future refactor must break two independent checks to reach production silently.
| File | Command / Code | Purpose |
|---|---|---|
| src | pub fn total_cents(subtotal_cents: u64) -> u64 { | Unit Tests Live Next to the Code |
| tests | use shop::{Cart, Money}; | Integration Tests in tests/ |
| src | pub struct Cart { | Doc-Tests |
| src | pub struct Limiter { | #[should_panic] and #[ignore] |
| tests | use std::fs; | Fixtures Without Magic |
| src | mod cli_demo { | cargo test Like a Pro |
| src | pub struct User { | Clippy Lints |
| src | use std::collections::HashMap; | rustfmt |
| .github | name: ci | GitHub Actions CI Matrix |
| Cargo.toml | [package] | Coverage With cargo-llvm-cov |
Key takeaways
Common mistakes to avoid
7 patternsTesting private helpers from tests/ by widening visibility to pub
Writing #[should_panic] without an expected message
Leaving #[ignore] tests without an owner or tracking issue
Running clippy without --all-targets in CI
Pinning --test-threads=1 permanently to hide flakiness
Gating merges on an absolute coverage percentage
Forgetting --locked in CI invocations
Interview Questions on This Topic
Why do Rust unit tests live in src/ under #[cfg(test)] while integration tests live in tests/, and what can each access?
Frequently Asked Questions
20+ years shipping production backend systems. Drawn from code that ran under real load.
That's Testing. Mark it forged?
27 min read · try the examples if you haven't