Home › Rust › Rust Embedded Bare-Metal: no_std, HALs, Embassy on RP2040
Advanced 28 min · September 26, 2026

Rust Embedded Bare-Metal: no_std, HALs, Embassy on RP2040

Build bare-metal Rust firmware with no_std, cortex-m-rt, rp2040-hal and Embassy: memory.x, defmt RTT, probe-rs flashing and heapless patterns explained..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 38 min
  • ✓Comfortable with Rust ownership, traits, and Cargo workspaces
  • ✓Built one std binary crate and read a microcontroller datasheet once
  • ✓Basic C-level grasp of interrupts, clocks, and memory-mapped registers
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Rust bare-metal means #![no_std] plus #![no_main]: you lose std collections, threads, and the filesystem, and you take over the entry point, panic handler, and linker layout yourself
  • Cross-compile with rustup target add thumbv7em-none-eabihf for Cortex-M4F/M7F or riscv32imc-unknown-none-elf for ESP32-C3 class RISC-V, set it in .cargo/config.toml, and describe FLASH and RAM origins in memory.x
  • The cortex-m and cortex-m-rt crates own the vector table and startup code on ARM, rp2040-hal adds safe GPIO/clock drivers on top, and Embassy adds a no-heap async executor with Timer and interrupt-aware sleep
  • Log with defmt over RTT and flash with probe-rs run --chip RP2040 for speed; keep semihosting and QEMU for host-side tests where halting the core to print is acceptable
  • Share state without an OS using heapless fixed-capacity Vec/String and critical-section Mutex guards, never a bare static mut polished over with unsafe
✦ Definition~90s read
What is Rust Embedded Bare Metal?

Rust bare-metal embedded is Rust compiled to a #![no_std] binary that runs directly on a microcontroller with no operating system, no main function in the C sense, and no standard library. The cortex-m-rt crate supplies the vector table and reset handler on ARM Cortex-M, a panic handler crate defines failure behavior, and a linker script named memory.x tells the linker where FLASH and RAM live.

★
Think of a normal Rust program as cooking in a fully equipped restaurant kitchen: running water, gas lines, dishwashers, and staff who clean up after you.

On top of that foundation sit peripheral access crates generated from vendor SVD files, hardware abstraction layers such as rp2040-hal that expose typed GPIO, SPI, I2C, and clock APIs, and optionally an async executor such as Embassy that multiplexes tasks without threads.

What disappears compared with std is larger than newcomers expect: there is no heap unless you add an allocator, no growable Vec or HashMap unless you use fixed-capacity replacements from heapless, no threads or std::sync primitives, no filesystem or network stack, and no println that goes anywhere useful. Logging moves to deferred formatting with defmt over RTT or ITM, flashing moves to probe-rs or picotool, and testing splits into host-side unit tests plus on-target or QEMU-driven integration runs.

The payoff is total control: single static binary, deterministic memory use verified at link time, and zero-cost peripheral access checked by the type system.

Peripheral access follows a layered model that keeps vendors and volunteers collaborating: SVD files describe registers, svd2rust generates peripheral access crates with type-checked registers, HAL crates add clock and driver logic, and embedded-hal traits let sensor drivers port across chips. Teams that learn this stack read any new ARM or RISC-V part quickly, because only the PAC and clock tree change while the patterns repeat.

Choose bare metal when determinism, boot time, and unit cost dominate; reach for Linux instead when you need processes, virtual memory, or rich networking that no static firmware should reimplement.

Plain-English First

Think of a normal Rust program as cooking in a fully equipped restaurant kitchen: running water, gas lines, dishwashers, and staff who clean up after you. Bare-metal Rust is cooking over a campfire with a single pan you brought yourself. There is no tap, no oven timer, no one to wash dishes. You decide exactly where the fire goes, how much wood fits in the pit, and what happens when something burns. The no_std keyword is you telling Rust that the restaurant is gone, the entry attribute is you lighting the fire yourself, the memory.x file is the chalk outline of your campsite showing where the tent and fire can sit, and tools like probe-rs and defmt are the walkie-talkie and headlamp that let you see what is happening in the dark.

You've shipped Rust on servers and you trust the borrow checker with your business logic. Then someone hands you a Raspberry Pi Pico and says the whole program must fit in 264 kilobytes of RAM with no operating system underneath. That's the moment bare-metal stops being a curiosity and becomes a completely different contract with the machine.

On the desktop, main returns, panics print a backtrace, and Vec grows until the allocator says stop. On a Cortex-M4F none of that exists. You'll write #![no_std] and #![no_main], you'll provide the reset handler through cortex-m-rt, and you'll decide what a panic even means when there's no stderr to print to.

The good news is the ecosystem has converged. You've got cortex-m for core registers, rp2040-hal for GPIO and clocks, Embassy for async tasks without an RTOS, defmt plus RTT for logging that doesn't stall your control loop, and probe-rs for flashing with a single cargo run.

But the sharp edges are real and they don't forgive sloppy assumptions. A wrong FLASH origin in memory.x bricks your boot, a blocking semihosting print inside a 1 kHz interrupt wrecks your timing, and a shared static without a critical-section guard corrupts state exactly when the demo starts.

You'll leave here with a working mental model and copy-pasteable patterns: the RP2040 blink that actually compiles, the Embassy task layout that sleeps instead of spins, the heapless buffers that can't fragment, and the host-versus-target test split that keeps your CI fast without lying to you.

no_std vs std: Exactly What You Lose and What Replaces It

The std crate is an operating-system compatibility layer, and #![no_std] removes it in one line. You keep core with its primitives, iterators, and Option and Result, plus compiler builtins, and you lose everything that assumes a syscall interface underneath. That means no heap-allocated Vec, String, HashMap, or Box unless you supply an allocator, no threads, no std::fs or std::net, no time, no environment variables, and no println that reaches anywhere. The embedded book documents this split explicitly: heap and collections are only available with an allocator such as alloc-cortex-m, and HashMap stays unavailable without a secure random source. Your first job is therefore inventory, not code: list every std API your design assumes, then assign each one a bare-metal replacement or delete it.

Collections get replaced by fixed-capacity types, not by willpower. The heapless crate provides Vec, String, IndexMap, LinearMap, BinaryHeap, and spsc and mpmc queues whose capacity is part of the type, so heapless::Vec<u8, 128> can never hold 129 bytes and its push returns a Result you must handle. This changes your error model: insertion failure becomes a normal, testable event instead of an OOM abort at 3 AM. Threads get replaced by interrupts plus a main loop, or by an async executor such as Embassy that multiplexes tasks on one stack. Files get replaced by flash wear-leveling stores such as sequential-storage or littlefs bindings, and networking gets replaced by crates like smoltcp with statically allocated socket buffers you size up front.

The practical consequence is that memory use becomes a compile-time and link-time fact rather than a runtime surprise. A heapless buffer reserves its bytes inline on the stack or in statics, so cargo size tells you the truth before you flash. That determinism is why safety-critical firmware accepts the ergonomic cost: no allocator fragmentation, no unbounded growth, no midnight malloc failure inside an interrupt. You pay with explicit capacities everywhere and with APIs that return Result on push, which forces you to decide what full means for each queue. Teams that fight this by bolting on a global allocator on day one usually reintroduce the exact nondeterminism they chose Rust to escape.

Start every port with a three-column audit: std API, bare-metal substitute, capacity bound. Map Vec to heapless::Vec with a named constant, HashMap to LinearMap or a sorted array, threads to Embassy tasks or RTIC tasks, and println to defmt levels you can filter at compile time. Forbid bare static mut in the same commit, because shared state without a critical section will corrupt the moment interrupts go live. When the audit is done, your firmware links for thumbv6m-none-eabi with no linker complaints about missing _start or start, and that clean link is the first real proof the design fits the chip.

A useful rule of thumb is that anything with implicit growth or implicit blocking has no place in interrupt context. Allocate all buffers at startup, size every queue from a measured worst case plus margin, and keep ISRs to sampling and flagging. The main loop or executor then owns formatting, storage, and transmission where backpressure can be observed. Engineers who respect that split ship firmware whose RAM report barely moves between releases; engineers who smuggle String formatting into ISRs ship field incidents with 6-hour periods.

There is a middle ground worth knowing before you swear off dynamic memory entirely. The alloc crate brings Box, Rc, and Vec behind a global allocator you provide, such as an embedded-alloc bump or linked-list heap carved from a static byte array. Some USB and network stacks expect alloc, and granting them a bounded 8 or 16 KB region can be pragmatic. The cost is nondeterminism: fragmentation and worst-case latency return the moment allocation succeeds conditionally. If you go this route, cap the heap in the linker layout, track high-water marks over defmt, and keep ISRs alloc-free so interrupt timing never depends on heap state.

Two core concepts deserve explicit attention because they change daily coding habits. First, core still gives you iterators, slices, Option, Result, and core::fmt, so parsing, checksums, and state machines port almost untouched; only the OS-backed conveniences vanish. Second, Send and Sync still govern cross-context sharing, and Cortex-M interrupts act like preemptive threads the compiler cannot see. Mark shared types Send only when critical-section guards make that true, wrap mutability in Cell or RefCell inside a Mutex, and run cargo clippy with the target triple set so lints evaluate your real build rather than the host illusion.

src/main.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
#![no_std]
#![no_main]
// core only: no Vec, no String, no threads, no fs.
// heapless replaces growable types with fixed capacity.
use heapless::{String, Vec};
const LOG_CAP: usize = 128;
const MSG_CAP: usize = 64;
fn buffer_demo() -> Result<(), ()> {
    let mut samples: Vec<u16, LOG_CAP> = Vec::new();
    samples.push(ADC_CODE_1).map_err(|_| ())?;
    let mut msg: String<MSG_CAP> = String::new();
    // push_str fails cleanly when full instead of aborting.
    use core::fmt::Write;
    write!(msg, "n={}", samples.len()).map_err(|_| ())?;
    Ok(())
}
const ADC_CODE_1: u16 = 2048;
⚠ No Heap Means No Silent Growth
Every buffer needs an explicit capacity chosen from measurements. A push that returns Err(Full) is your design working, not a bug to unwrap away.
📊 Production Insight
A telemetry team replaced std Vec with heapless::Vec<u16, 256> sized from a week of field captures plus 40 percent margin. Their release RAM report stopped drifting, queue-full events became countable metrics, and the midnight OOM reset that had fired twice a month disappeared for 11 straight months.
🎯 Key Takeaway
Audit every std dependency into a fixed-capacity substitute with a named bound. Full queues must be handled events, not panics, and ISRs must never allocate or format.

The Runtime Contract: #![no_main], #[entry], and #[panic_handler]

A hosted Rust binary starts in crt0, runs init arrays, then calls your main. A bare-metal binary starts at the reset vector with RAM uninitialized, the stack pointer loaded from the first vector word, and no one to call. The cortex-m-rt crate fills that gap: it provides the vector table, copies .data from flash to RAM, zeroes .bss, optionally enables the FPU on thumbv7em targets, then jumps to your entry function. Your part of the contract has three pieces: #![no_main] to drop the hosted entry expectation, #[cortex_m_rt::entry] on a fn main() -> ! that never returns, and a #[panic_handler] fn(&PanicInfo) -> ! linked exactly once. Miss any piece and the link fails with an unmistakable undefined reference rather than a runtime mystery.

The entry macro is stricter than it looks. It must appear once in the final binary, the function must be named main by convention and must diverge, and it runs with interrupts enabled unless you mask them first. Initialization order matters: static constructors do not exist, so peripherals start in reset state and every clock, pin, and baud rate is your explicit decision. The panic handler runs in whatever context panicked, possibly inside an interrupt with limited stack, so it must be short, non-allocating, and non-blocking. The standard choices are panic-halt which spins, panic-probe which emits a defmt frame then faults, and panic-semihosting with its exit feature for QEMU-driven tests that need a pass-fail exit code.

Crate selection encodes intent for reviewers. Use panic-halt for bring-up when you want the debugger to catch the spin, panic-probe with defmt for field builds where every panic must leave an RTT trace, and panic-semihosting only for host-driven QEMU suites where debug::exit maps test results to process codes. Exactly one panic crate may be linked; two produces a duplicate-symbol error that is genuinely helpful. Keep the handler body under a dozen lines: capture the location, log it through the approved channel, then halt or reset. Anything fancier risks a secondary fault inside the fault path, which on Cortex-M escalates to HardFault and then to lockup with no trace at all.

Startup code also owns the exception story. Handlers you do not override default to a tight loop, which is safe but silent, so override HardFault early in development to dump the ExceptionFrame through defmt. Name interrupt handlers exactly as the PAC expects or they bind to the default and your peripheral appears dead. Record the reset cause where the chip offers it, because distinguishing power-on from watchdog from software reset cuts triage time in half. When these pieces are in place, reset behavior becomes boring: vectors valid, RAM initialized, clocks set, panic path proven, and every future bug arrives with a frame instead of a shrug.

Treat the runtime files as product code, not scaffolding. Pin cortex-m-rt and cortex-m versions in Cargo.toml, review their changelogs on upgrade, and keep a reset-cause log in retained RAM or backup registers. Teams that version their runtime deliberately spend release week on features; teams that float on wildcards spend it bisecting a startup regression that only shows on cold boot.

The vector table deserves a closer look because it is the first 256-plus bytes of your image and the chip reads it before a single Rust statement runs. Entry zero holds the initial stack pointer, entry one holds the reset handler from cortex-m-rt, and subsequent entries map exceptions and device interrupts in an order fixed by ARM. On RP2040 the table sits behind the boot2 stage in the XIP window, while on STM32 it lives at 0x08000000 with VTOR available for RAM relocation during RAM-run debugging. Misaligned tables or missing entries escalate to HardFault on first interrupt, so inspect them with cargo objdump when bring-up stalls.

Two more runtime hooks round out the picture. The #[pre_init] attribute runs code before .data and .bss init for clock or watchdog quirks that cannot wait, though most firmware never needs it and should avoid the footgun. Retained-RAM sections marked with linker attributes survive soft resets and carry reset-cause codes plus last-panic frames across reboots, turning the next boot into a postmortem reader. Record whether each reset was power-on, watchdog, or software-requested, because that single discriminant halves triage time on intermittent field lockups.

src/main.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
#![no_std]
#![no_main]
use cortex_m_rt::entry;
use panic_probe as _;
use defmt_rtt as _;
#[entry]
fn main() -> ! {
    // .data/.bss already init, FPU on if thumbv7em-none-eabihf.
    defmt::info!("boot: clocks reset, entering loop");
    loop {
        cortex_m::asm::wfi();
    }
}
// panic-probe supplies #[panic_handler]: logs frame over RTT.
#[cortex_m_rt::exception]
unsafe fn HardFault(ef: &cortex_m_rt::ExceptionFrame) -> ! {
    defmt::error!("hardfault: {:?}", defmt::Debug2Format(ef));
    cortex_m::asm::udf();
}
💡One Entry, One Panic Path, Proven Early
Link exactly one panic crate and override HardFault on day one. A fault that prints its frame beats a silent spin every time.
📊 Production Insight
A motor-driver team shipped with panic-halt and a default HardFault loop, so field lockups left zero evidence across 34 returns. After switching to panic-probe with an RTT trace plus a retained-RAM reset log, the next 6 faults each arrived with register dumps and the root cause was fixed in one sprint.
🎯 Key Takeaway
Own the reset contract explicitly: single diverging entry, single minimal panic handler, named exception overrides. Prove the fault path before you need it.

Cross Targets and Toolchains: thumbv6m, thumbv7em, and riscv32imc

Bare-metal Rust cross-compiles from your laptop to an ARM or RISC-V triple that has no OS, and the triple name encodes the contract. thumbv6m-none-eabi covers Cortex-M0 and M0+ as on the RP2040, thumbv7m-none-eabi covers M3, thumbv7em-none-eabi covers M4/M7 without FPU, and thumbv7em-none-eabihf covers M4F/M7F with hardware float. On RISC-V the pattern is riscv32imc-unknown-none-elf for ESP32-C3 class cores, with imac and imafc variants where atomics or float exist. The eabi versus eabihf suffix decides float calling convention, and mixing them across crates produces link errors about float ABI that no amount of cleaning will fix except aligning every crate on one suffix.

Installation is a single rustup command per triple, but pinning is where teams win. Run rustup target add thumbv6m-none-eabi for the Pico, rustup target add thumbv7em-none-eabihf for an STM32F4 with FPU, and rustup target add riscv32imc-unknown-none-elf for ESP32-C3 bring-up. Record the choice in .cargo/config.toml under [build] target so plain cargo build does the right thing, and add target-specific runners and rustflags in the same file. Verify with rustup target list --installed and rustc --print target-list filtered for your architecture. Reproducible CI images should install the same triples in the Dockerfile rather than relying on a developer laptop that happens to have them.

The RP2040 deserves a special note because newcomers pick the wrong ARM triple. It is a Cortex-M0+ and needs thumbv6m-none-eabi, not the thumbv7em triples that dominate STM32 examples. Building RP2040 code for thumbv7em-none-eabihf may compile yet emit instructions the M0+ cannot execute, failing only on hardware. Conversely, building an M4F project for the soft-float thumbv7em-none-eabi silently drops hardware float and slows control math by an order of magnitude. Match the triple to the core datasheet first, then let the HAL and PAC features follow: rp2040-hal with thumbv6m, stm32f4xx-hal with thumbv7em-none-eabihf, esp32c3-hal with riscv32imc-unknown-none-elf.

Linker behavior follows the triple. ARM Cortex-M builds link through cortex-m-rt link.x plus your memory.x, while RISC-V builds use riscv-rt with its own layout file, and both expect rust-lld as the default linker with no external GCC required for pure Rust. When C objects join, for example a vendor DSP blob, you may need arm-none-eabi-gcc or riscv32-unknown-elf-gcc as linker with matching float flags. Keep build-std off unless you truly need it; the prebuilt core for these triples is current and rebuilding it adds fragility. A clean cargo build --target thumbv6m-none-eabi plus cargo readobj confirming Machine ARM and cargo size confirming section fits is the exit gate for this stage.

Document the triple decision in the README next to the chip part number, because every future contributor, CI agent, and hardware revision will re-ask it. One line stating RP2040 uses thumbv6m-none-eabi via rustup target add prevents a whole class of it works on my STM32 branch failures.

Per-target Cargo configuration keeps multi-board workspaces honest. Use [target.thumbv6m-none-eabi] runner and rustflags blocks so Pico builds flash with --chip RP2040 while STM32 builds use their own chip string, and add aliases such as cargo run-pico to remove guesswork. Linker args belong here too: -Tlink.x for Cortex-M via cortex-m-rt, or the riscv-rt equivalent on RISC-V, so every contributor links identical scripts. Avoid -Zbuild-std unless a vendor target demands it; rebuilding core adds toolchain fragility for benefits most projects never observe, and the prebuilt Triple libraries track stable releases closely.

CI should treat the triple matrix as a first-class test dimension. Build thumbv6m-none-eabi and riscv32imc-unknown-none-elf images on every push, run cargo clippy --target for each, and assert the ELF machine field with cargo readobj so a wrong-triple artifact fails in seconds rather than on a customer bench. Cache the rustup targets in the builder image to keep cold builds under five minutes. When a new board revision changes the core, for example an M0+ to M4 migration, the triple change lands as a reviewed config diff with size reports attached, not as a surprise discovered during factory flashing.

.cargo/config.tomlTOML
1
2
3
4
5
6
7
8
9
10
11
12
13
[build]
# RP2040 is Cortex-M0+: thumbv6m, not thumbv7em.
target = "thumbv6m-none-eabi"
[target.thumbv6m-none-eabi]
# probe-rs flashes and streams RTT on cargo run.
runner = "probe-rs run --chip RP2040"
rustflags = ["-C", "link-arg=-Tlink.x"]
[target.thumbv7em-none-eabihf]
# STM32F4 with FPU keeps hard-float ABI consistent.
runner = "probe-rs run --chip STM32F407VG"
[alias]
b = "build --target thumbv6m-none-eabi"
run-pico = "run --target thumbv6m-none-eabi"
🔥Triple Must Match the Core
RP2040 takes thumbv6m-none-eabi, M4F takes thumbv7em-none-eabihf, ESP32-C3 takes riscv32imc-unknown-none-elf. Pin it in config so every build agrees.
📊 Production Insight
A dual-board project built Pico and STM32F4 firmware from one workspace with a single default triple, so Pico images carried M4 instructions and died at first interrupt on 200 pilot units. Pinning per-target config plus a CI assert on cargo readobj output caught the mismatch in 40 seconds on the next commit.
🎯 Key Takeaway
Install triples with rustup, pin them in .cargo/config.toml, and verify Machine and float ABI before flashing. The triple is a hardware contract, not a preference.

memory.x: Drawing the Map the Linker Must Obey

Every bare-metal link needs a map of physical memory, and memory.x is that map. It declares FLASH with an ORIGIN address and LENGTH, plus RAM with its own ORIGIN and LENGTH, in a MEMORY block the linker script consumes. cortex-m-rt link.x places .text and .rodata into FLASH, reserves .data load addresses in flash with runtime copies to RAM, zeroes .bss in RAM, and anchors the initial stack pointer at the top of RAM via _stack_start. If ORIGIN is wrong the image boots from an alias that mirrors nothing; if LENGTH is optimistic the link overflows only when your buffers grow past the fiction, usually two releases later.

Concrete numbers ground the concept. The RP2040 exposes 264 KB of SRAM at 0x20000000 and up to 16 MB of external QSPI flash windowed at 0x10000000, with the common 2 MB Pico board expressed as FLASH ORIGIN 0x10000000 LENGTH 2048K and RAM ORIGIN 0x20000000 LENGTH 264K. An STM32F303VCT6 instead uses FLASH ORIGIN 0x08000000 LENGTH 256K with RAM ORIGIN 0x20000000 LENGTH 40K. These are not interchangeable décor; the first word of the vector table, the VTOR relocation, and the boot2 second-stage loader on RP2040 all assume the right base. Copying an STM32 memory.x into a Pico project links cleanly and then does nothing on reset, which is a miserable afternoon.

The RP2040 adds a boot wrinkle that STM32 developers miss. It boots through a 256-byte second-stage bootloader in flash that configures the QSPI interface before your reset handler runs, supplied by the rp2040-boot2 crate in Rust projects. Your memory.x must leave room for that header layout and your .cargo/config.toml must pass -Tlink.x so the sections order correctly. Edit memory.x after a first build and run cargo clean before rebuilding, because cargo does not always track the linker script as a dependency and stale artifacts will happily link against your old map. Check the map after every RAM-affecting change with cargo size -- -A and keep 8 to 12 KB of stack margin on small parts.

Treat memory.x as reviewed product code with comments citing the datasheet page. Name the exact part, the flash chip size for external-flash parts, and the stack reserve rationale. Add a CI check that fails when RAM sections exceed 90 percent of LENGTH, because a link that barely fits today breaks the moment someone adds a 4 KB log buffer. When the map is honest, the linker becomes your strictest reviewer: overflow errors arrive at build time with section names attached instead of as HardFaults in a customer rack.

Keep a per-board memory.x rather than a shared fantasy file. Pico, Pico W, and custom RP2040 boards with different flash sizes each deserve their own layout behind a Cargo feature or a board directory, selected explicitly. Shared maps drift toward the largest board and hide overflows on the smallest, which is exactly backward.

Inside link.x the section order tells the boot story in miniature. .vector_table opens FLASH so reset finds it, .text and .rodata follow with code and constants, then load regions stage .data for its RAM copy while .bss reserves zeroed RAM and .uninit reserves uninitialized RAM for retained state. _stack_start anchors at the top of RAM and grows down toward your statics, which is why stack margin math matters: sum worst-case main plus worst-case ISR nesting plus async task needs, then keep 8 to 12 KB clear on small parts. Heapless buffers appear as fixed RAM occupants you can read directly in the size report, unlike allocator heaps that hide behind a single brk pointer.

External-flash parts add one more layout decision: the XIP window mapping. On RP2040 code executes in place from the 0x10000000 window after boot2 configures QSSI timings, so flash wait states and caching directly shape interrupt latency. Keep hot ISRs in RAM with section attributes when jitter budgets demand it, and place logging buffers where DMA or the debugger can reach them without cache maintenance. Document each placement choice beside the MEMORY block with the datasheet section cited, because the engineer debugging a 2 AM HardFault needs the map and its rationale in one glance, not scattered across chat history.

memory.xTEXT
1
2
3
4
5
6
7
8
9
10
11
/* RP2040 on Raspberry Pi Pico: 2MB QSPI flash, 264KB SRAM. */
MEMORY
{
  /* XIP window: boot2 + vector table + .text live here. */
  FLASH : ORIGIN = 0x10000000, LENGTH = 2048K
  /* SRAM: .data/.bss, heapless statics, main + ISR stacks. */
  RAM   : ORIGIN = 0x20000000, LENGTH = 264K
}
/* Stack starts at top of RAM; link.x places vectors at FLASH origin. */
_stack_start = ORIGIN(RAM) + LENGTH(RAM);
/* STM32F303VCT6 contrast: FLASH 0x08000000 x 256K, RAM 0x20000000 x 40K. */
⚠ Wrong Origin Boots Nowhere
Linking a Pico image against an STM32 origin produces a clean build and a dead board. Comment the part number and flash size in memory.x.
📊 Production Insight
A custom RP2040 board with 16 MB flash shipped using the Pico 2 MB memory.x, so OTA images beyond 2 MB truncated silently and 9 percent of units boot-looped after update. A per-board memory.x with LENGTH 16384K plus a CI size gate ended the incident in one release.
🎯 Key Takeaway
memory.x is the hardware truth: correct ORIGIN and LENGTH per board, cargo clean after edits, and cargo size gates so overflow fails the build instead of the field.

The cortex-m crate gives you the core: NVIC, SCB, SysTick, interrupt masking, and delay drivers built on the system timer. The rp2040-hal crate adds the chip: clocks, PLLs, GPIO banks, UART, SPI, I2C, ADC, PWM, and PIO, all typed so a pin configured as input cannot be passed where an output is required. The rp-pico board support crate then maps those chip pins to physical board labels such as the on-board LED on GPIO25. Blink is the correct first program because it exercises the full chain that every later driver reuses: take peripherals once, configure clocks, split GPIO, set pin direction, then toggle with a timed delay.

Peripheral ownership is the pattern newcomers must internalize. pac::Peripherals::take returns Some exactly once after reset; a second take returns None, which prevents two drivers from configuring the same register block. The HAL consumes those raw blocks into safe wrappers: Sio for the single-cycle I/O block, Pins for the bank split gated on RESETS, and a Delay built from the system timer after clocks init. Watchdog setup is mandatory on RP2040 HAL init because the chip boots with it ticking, and the standard template configures a 125 MHz system clock from the 12 MHz crystal before touching GPIO. Skip clock init and your delays run at the wrong rate while UART baud errors corrupt every byte.

GPIO25 blink on the Pico follows a fixed script that is worth memorizing. Take PAC peripherals, create Watchdog, Sio, and Pins with IO_BANK0, PADS_BANK0, and RESETS, convert pins.led into a push-pull output, then loop with set_high, a 500 ms delay, set_low, and another 500 ms delay. The embedded-hal OutputPin trait supplies set_high and set_low returning Results you handle, and cortex-m::asm::wfi between toggles is optional for power. Compile for thumbv6m-none-eabi and flash with probe-rs run --chip RP2040 or picotool over BOOTSEL USB. When the LED blinks at the expected rate, clocks, GPIO, linker, and runner all agree, which is a stronger health signal than any single unit test.

Board crates versus raw HAL is a packaging choice, not a religious one. rp-pico pins the LED and crystal for the Pico board so examples stay short; rp2040-hal alone suits custom boards where GPIO25 means nothing and the crystal differs. Either way, keep pin assignments in one board module so a revision change touches one file. Version-pin rp2040-hal, embedded-hal, and cortex-m together because trait revisions across the HAL boundary break compilation in ways that look like your bug but are really a semver drift.

From blink, grow outward one peripheral at a time: UART transmit next for a second observability path, then ADC sampling into a heapless queue, then PWM for the same LED without busy delays. Each step reuses take, clocks, and pin discipline, so progress compounds instead of restarting. A team that can blink, print, and sample on demand has the skeleton of every sensor product it will ever ship.

Under the HAL sits a layered model that explains both its power and its compile errors. Vendors publish SVD register descriptions, svd2rust generates peripheral access crates with typed registers, and HAL crates like rp2040-hal build clock trees, pin multiplexing, and driver state machines on top. The embedded-hal traits then abstract GPIO, SPI, I2C, and serial across vendors, so a sensor driver written against OutputPin and I2c traits ports from Pico to STM32 with only construction code changing. When versions drift across these layers the errors point at trait bounds rather than your logic; pinning the trio together in Cargo.toml and upgrading them as a unit keeps the abstraction earning its keep.

Timing discipline separates demos from products. Blocking Delay built on SysTick suits init and blink, but control loops want hardware timers, PWM slices, or Embassy Timer awaits that free the core between edges. Verify every rate claim with a scope on a toggled pin before trusting delay_ms arithmetic, because PLL misconfiguration scales all software delays together and only hardware measurement exposes it. From blink, the canonical growth order is UART transmit for a second log path, ADC sampling into a bounded queue, then PWM or PIO for outputs that must not jitter with main-loop load.

src/bin/blinky.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
#![no_std]
#![no_main]
use panic_halt as _;
use rp_pico::entry;
use embedded_hal::digital::v2::OutputPin;
use rp_pico::hal::{pac, sio::Sio, watchdog::Watchdog};
use rp_pico::hal::clocks::init_clocks_and_plls;
use cortex_m::delay::Delay;
#[entry]
fn main() -> ! {
    let mut pac = pac::Peripherals::take().unwrap();
    let mut watchdog = Watchdog::new(pac.WATCHDOG);
    let sio = Sio::new(pac.SIO);
    let clocks = init_clocks_and_plls(
        rp_pico::XOSC_CRYSTAL_FREQ,
        pac.XOSC, pac.CLOCKS, pac.PLL_SYS, pac.PLL_USB,
        &mut pac.RESETS, &mut watchdog,
    ).ok().unwrap();
    let pins = rp_pico::Pins::new(
        pac.IO_BANK0, pac.PADS_BANK0, sio.gpio_bank0, &mut pac.RESETS,
    );
    let mut led = pins.led.into_push_pull_output();
    let mut delay = Delay::new(cortex_m::Peripherals::take().unwrap().SYST, clocks.system_clock.freq().to_Hz());
    loop {
        led.set_high().unwrap();
        delay.delay_ms(500);
        led.set_low().unwrap();
        delay.delay_ms(500);
    }
}
💡Blink Proves the Whole Chain
Take peripherals once, init clocks before GPIO, and flash with the right chip flag. A steady 1 Hz blink means linker, clocks, and runner agree.
📊 Production Insight
A contract manufacturer swapped a 12 MHz crystal for an 8 MHz variant without notice, and blink ran 1.5 times slow on 400 boards, which UART loopback confirmed as baud drift. Centralizing XOSC_CRYSTAL_FREQ in one board module turned the fix into a one-line change verified by blink rate.
🎯 Key Takeaway
Build blink on rp2040-hal through take, clocks, Pins, and OutputPin. One board module owns crystal and pin maps so hardware revisions stay one-line changes.

Embassy Async HAL: Tasks That Sleep Instead of Spin

Embassy is an async executor and HAL family designed for microcontrollers with no heap and no OS threads. Instead of one super-loop polling every device, you write small async tasks that await Timer, GPIO edges, or channel messages, and the executor parks the core with WFE or WFI until an interrupt wakes exactly the task that can progress. Tasks are statically allocated, sized at compile time, with no arena tuning for one or a thousand tasks on current releases. The mental shift is from who runs when, managed by priorities and preemption, to what each task waits on, expressed as await points the compiler checks.

The RP2040 Embassy path centers on embassy-rp with its interrupt and thread executors plus multicore support. A minimal app declares #[embassy_executor::main] on an async main that receives a Spawner, initializes embassy time drivers and chip peripherals, then spawns worker tasks annotated #[embassy_executor::task], often with pool_size for multiple instances. Workers blink LEDs with Timer::after_millis(500).await instead of blocking delays, read sensors without stalling siblings, and share buses through embassy-embedded-hal mutex wrappers. Because waiting tasks consume no CPU, a battery node can sleep between samples and wake on time or pin edge, which polling loops cannot match without hand-rolled state machines.

Choosing between blocking HAL calls and Embassy async drivers is the key design call. Use blocking rp2040-hal calls for bring-up, one-shot init, and tight bit-banged sequences where control is simpler single-threaded. Use Embassy tasks for everything concurrent: LED heartbeat plus button handling plus sensor sampling plus USB logging, each as its own task with clear await points. Shared SPI or I2C buses get a single async mutex owner with device handles checked out per transaction, which removes the chip-select races that plague interrupt-driven sharing. Timer replaces Delay everywhere in async code, because Delay::delay_ms blocks the executor and starves siblings while Timer::after yields the core.

Multicore and priority need deliberate setup on RP2040. Embassy supports spawning executors on core 1, with channels bridging cores through static queues, and interrupt-mode executors for latency-sensitive paths. Keep core 0 for timing and core 1 for communication or DSP, and size channel depths from measured bursts rather than hope. Avoid spawning unbounded task instances; pool_size bounds concurrency at compile time and turns overload into a spawn error you log instead of a stack overflow you debug with a scope. Document which interrupts belong to the executor versus your raw handlers, because double-claiming a priority breaks sleep and burns power.

Adopt Embassy incrementally to control risk. Start with one #[embassy_executor::main] hosting a heartbeat task beside your existing blocking init, prove sleep current drops on a meter, then migrate one driver at a time behind async wrappers. Teams that rewrite everything in week one drown in borrow errors across task boundaries; teams that migrate the LED, then the button, then the sensor ship each Friday with a working image and a power graph that keeps improving.

Communication between tasks runs through static channels rather than shared globals. embassy-sync provides Channel, PubSubChannel, and Watch types with compile-time depths, so a sensor task publishes samples while USB and storage tasks subscribe without locks or copies beyond the queue slots. Size each depth from burst math: publish rate times worst consumer stall plus margin, mirroring the heapless sizing discipline. For single-slot latest-value semantics such as setpoints, Watch overwrites cleanly and wakes all waiters, which beats a mutex-protected global that every reader must poll.

The time driver and sleep behavior close the power story. Embassy installs a timer driver on a hardware alarm that wakes the core exactly when the next task becomes ready, letting WFE sleep dominate the duty cycle on sensor nodes. Measure with a current profiler: a well-built node shows sharp wake spikes separated by microamp floor, while a polling design shows a flat milliamp plateau. Keep interrupt priorities assigned so executor interrupts preempt application ISRs correctly, document the priority map in one module, and never claim the executor's alarm for raw use. When sleep current matches the datasheet floor, the async model has paid for its learning curve in battery invoices avoided.

src/bin/embassy_blink.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
#![no_std]
#![no_main]
use embassy_executor::{main, Spawner, task};
use embassy_rp::gpio::{Level, Output};
use embassy_time::{Duration, Timer};
use panic_probe as _;
use defmt_rtt as _;
#[main]
async fn main(spawner: Spawner) {
    let p = embassy_rp::init(Default::default());
    spawner.spawn(heartbeat(p.PIN_25)).unwrap();
    spawner.spawn(reporter()).unwrap();
    defmt::info!("embassy up: 2 tasks");
}
#[task]
async fn heartbeat(pin: embassy_rp::Peri<'static, embassy_rp::peripherals::PIN_25>) {
    let mut led = Output::new(pin, Level::Low);
    loop {
        led.set_high();
        Timer::after(Duration::from_millis(500)).await;
        led.set_low();
        Timer::after(Duration::from_millis(500)).await;
    }
}
#[task]
async fn reporter() {
    loop {
        defmt::info!("tick");
        Timer::after(Duration::from_secs(5)).await;
    }
}
🔥Await Is Your Scheduler
Timer and channel awaits park the core until work exists. Blocking delays inside async tasks defeat the executor and starve siblings.
📊 Production Insight
A door sensor polling its reed switch at 100 Hz drew 8.2 mA average and killed its coin cell in 19 days. After moving to an Embassy GPIO-edge wait plus Timer heartbeat, average current fell to 310 uA and field life passed 14 months on the same cell.
🎯 Key Takeaway
Model concurrency as await points on Timer, pins, and channels under one static executor. Migrate one task at a time and measure sleep current after each step.

Logging That Fits: defmt and RTT vs Semihosting

Printing on bare metal is a transport problem disguised as a formatting problem. defmt solves formatting by moving string assembly to the host: the target sends tiny numeric frames with interned string indices, and a host decoder rehydrates human text from the ELF. This cuts flash use by an order of magnitude versus core::fmt strings and cuts wire bytes similarly, which matters when logging at 1 kHz from an ISR-adjacent path. RTT solves transport by exposing a small ring buffer in target RAM that the debug probe reads via the debug port without halting the core. Together with probe-rs run streaming decoded frames to your terminal, you get printf-style visibility at near-zero target cost.

Semihosting sits at the opposite end of the trade space. Each hprintln traps through a breakpoint instruction, halts the core, and waits for the debug host to service the request, costing milliseconds per line depending on the probe. Under a debugger on the bench that feels merely slow; in the field with no debugger it hangs or stretches into tens of milliseconds per call, exactly inside the timing path you most wanted to observe. Semihosting needs almost no wiring and works over the same debug connection, which explains its popularity in tutorials. It is a bench-only tool, and every line left in release is a latent stall whose cost depends on probe presence, the worst kind of Heisenbug.

Configure defmt deliberately. Link exactly one transport, either defmt-rtt for hardware or defmt-semihosting for QEMU suites, because linking both or adding rtt-target alongside defmt-rtt collides. Import with use defmt_rtt as _ so the global logger registers, set DEFMT_RTT_BUFFER_SIZE to a power of two such as 1024 when RAM is tight, and gate verbosity with compile-time levels so trace frames vanish from release binaries. Prefer defmt::info, warn, error, and trace over unconditional prints, and format integers and byte slices rather than String to keep frames small. Note that probe-rs forces RTT into blocking mode to avoid losing data, so a disconnected host with a full buffer can still stall; size the buffer from burst measurements and keep logging out of hard-real-time ISRs regardless.

Keep a second observability path for when the probe is absent. A UART TX pin at 115200 baud costs one pin and gives field logs through any USB serial dongle, while retained-RAM panic slots preserve the last fault across resets. For QEMU-driven CI, switch the transport to defmt-semihosting and decode with the qemu-run tool, keeping target code identical apart from the transport import. The discipline is simple: defmt plus RTT for development speed, UART or retained RAM for field truth, semihosting only for automated QEMU tests that need exit codes.

Audit logging cost like any other budget. Measure worst-case frames per second, multiply by frame bytes, confirm RTT bandwidth headroom, and load-test with the debugger detached. Teams that do this keep control loops clean at full log verbosity; teams that log first and measure never discover the 14 ms stall until the cold night it matters.

ITM tracing offers a third transport worth knowing on debug-capable parts. Where RTT needs a debug probe halt-free read, ITM streams through the SWO pin at baud-derived rates with minimal target code, and defmt-itm carries the same interned frames over that wire. It shines for instruction-trace correlation on M3/M4 parts with ETM cells, but the Pico M0+ lacks the full trace macrocell, which is why RTT dominates RP2040 practice. Choose per board: RTT where probes are standard, ITM where trace correlation matters, UART where no probe exists, and semihosting nowhere outside QEMU.

Level hygiene and panic integration complete the setup. Assign trace to per-sample paths, debug to state transitions, info to boot and mode changes, warn to recovered faults like queue-full, and error to unrecoverable paths that precede reset. Wire panic-probe so panics emit defmt error frames with location before faulting, guaranteeing the last RTT kilobyte holds the cause. Size the RTT buffer from burst math: peak frames per second times average frame bytes times the longest host-detached interval you must survive. When that budget is written down and tested with the probe unplugged, logging graduates from debug aid to flight recorder.

src/logging.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
// Hardware build: decoded, low-cost logging over RTT.
use defmt_rtt as _; // exactly one transport; conflicts with rtt-target
use panic_probe as _; // panic frames ride the same RTT channel
pub fn sample_path(code: u16) {
    // interned on host: target sends ~4 bytes, not the string.
    defmt::info!("adc code={=u16} vbat_ok={}", code, code > 900);
    defmt::trace!("raw bytes={=[u8; 2]}", code.to_le_bytes());
}
// QEMU build: swap the single import, keep all call sites identical.
// use defmt_semihosting as _; // plus qemu-run decoder on host
// Bench-only fallback, never in release ISRs:
// cortex_m_semihosting::hprintln!("slow: {}", code).ok();
⚠ Semihosting Stalls, RTT Streams
Semihosting halts the core per line and hangs without a debugger. Ship defmt over RTT and reserve semihosting for QEMU tests.
📊 Production Insight
A vibration monitor logged 40-byte formatted lines over semihosting at 200 Hz, adding 9 ms of stall per sample and aliasing its own FFT on 63 deployed units. Switching to defmt trace frames over RTT cut per-sample cost to 6 us and restored the spectrum without touching the DSP code.
🎯 Key Takeaway
Log with defmt levels over RTT, one transport per build, verbosity filtered at compile time. Keep ISRs log-free and keep a UART path for probeless fields.

Flashing and Live Debug with probe-rs

probe-rs is the modern flashing and debug toolkit for ARM and RISC-V that replaces the OpenOCD plus GDB plus telnet ritual with one Cargo runner. Install it once with cargo install probe-rs-tools, confirm your probe with probe-rs list, then set runner = "probe-rs run --chip RP2040" for the Pico or --chip nRF52840_xxAA and STM32F407VG for those families. From then on cargo run builds for your ARM triple, flashes through SWD, resets the target, and streams RTT and defmt output back into the same terminal as if the firmware printed to stdout. That tight loop is why bring-up that took an afternoon with UF2 drag-and-drop now takes minutes.

Chip names must be exact, and that precision is a feature. probe-rs run --chip RP2040 selects the correct flash algorithm for the external QSPI, while a bare cargo-flash without a chip argument guesses and programs the wrong sectors. Keep the chip string in .cargo/config.toml per target triple so Pico and STM32 builds cannot be crossed, and override on the command line only for one-off boards. For RP2040 USB workflows, picotool remains the alternative: hold BOOTSEL, mount the mass-storage volume, and run picotool load firmware.uf2 followed by picotool reboot. Prefer probe-rs for daily work because it also carries RTT and panics; reserve picotool for factory provisioning where no debug probe exists.

Debugging with probe-rs stays close to the metal without GDB ceremony. probe-rs run prints defmt frames and panic locations inline, probe-rs attach holds a session for memory inspection, and the VS Code extension speaks the same backend for breakpoints and RTT views. When RTT stays silent, check three things in order: the runner chip string, a single RTT transport linked, and physical SWD wiring with adequate power, since brownout during flash mimics every software fault. Measure flash time per board; a healthy RP2040 programs in 2 to 4 seconds, and a sudden jump to 30 seconds usually means a worn cable or a probe firmware mismatch rather than your code.

Production flashing deserves a separate path from development cargo run. Bake release with cargo build --release, record the git SHA into firmware via rp binary-info or vergen, then flash with probe-rs download plus verify, capturing the log per unit. Keep debug builds RTT-verbose and release builds defmt-filtered to warn and above so factory logs stay quiet yet faults still surface. Teams that standardize on one runner config per board eliminate the it flashed on my machine class of support tickets entirely.

Learn the four commands that cover 90 percent of bench life: probe-rs list for probe health, cargo run for flash plus logs, probe-rs download for scripted programming, and cargo size for fit checks before you flash. Everything else is a refinement. When those four are muscle memory, the debug cycle shrinks to edit, run, read, fix, and hardware starts feeling as iterative as the host.

Probe hardware choice shapes daily life more than datasheets suggest. A PicoProbe built from a second Pico costs little and speaks CMSIS-DAP well enough for flash plus RTT, while J-Link and ST-Link probes add speed and GDB-server polish at higher cost. Wire SWDIO, SWCLK, ground, and target power sensing carefully; half of all mysterious flash failures trace to ground bounce or underpowered targets browning out mid-erase. Keep probe firmware current because protocol mismatches present as sudden 30-second flashes or vanished targets rather than clean errors. Label probes per bench so two developers never race one pod.

Factory and field flows diverge from bench habits in one key way: they must be scriptable and attested. Script probe-rs download with explicit chip and ELF path, verify after write, capture per-unit logs keyed by serial, and embed the firmware git SHA plus build profile into binary info so any unit answers what it runs. Keep a recovery path documented: BOOTSEL USB with picotool load for RP2040 units whose SWD pads are unpopulated, and a known-good golden image for fast triage. When flashing is a logged, versioned procedure instead of a developer ritual, manufacturing yield meetings stop discussing firmware as a suspect.

.cargo/config.tomlBASH
1
2
3
4
5
6
7
8
9
10
11
12
# Daily bench loop for RP2040 over SWD.
cargo install probe-rs-tools --locked
probe-rs list
# expect: J-Link / CMSIS-DAP / PicoProbe with RP2040 target
cargo run --target thumbv6m-none-eabi
# flashes, resets, streams defmt/RTT + panics to this console
probe-rs run --chip RP2040 target/thumbv6m-none-eabi/debug/firmware
# explicit flash without Cargo runner indirection
DEFMT_RTT_BUFFER_SIZE=1024 cargo run --target thumbv6m-none-eabi
# larger RTT buffer when bursts exceed 1 KB
cargo size --bin firmware --target thumbv6m-none-eabi --release -- -A
# confirm .text/.data fit before factory programming
💡One Runner Per Board
Pin runner plus chip string per triple in config. Explicit chip names pick the right flash algorithm and end cross-board flash accidents.
📊 Production Insight
A pilot run of 150 Picos was flashed with a stale STM32 runner after a config merge, bricking the morning batch until lunch. Splitting .cargo/config.toml by triple with RP2040 pinned plus a CI grep for the chip string cut misflashes to zero across the next 2,400 units.
🎯 Key Takeaway
Standardize on probe-rs run with exact chip names for dev, scripted download plus verify for factory. Slow flashes mean cables or power before code.

heapless Queues and Critical-Section Sharing Without a Heap or OS

heapless gives you std-like containers with the growth knob removed: Vec, String, IndexMap, IndexSet, LinearMap, BinaryHeap, HistoryBuffer, and single- and multi-producer queues, each parameterized by capacity in the type. heapless::Vec<u8, 128> stores 128 bytes inline, never calls the allocator, and returns Err on the 129th push instead of growing. Modern releases take const-generic capacities directly, so Vec<u8, 8> reads plainly after older typenum U8 spellings. Because storage is inline, instances live on the stack, in statics, or inside drivers with identical semantics, and their sizes show up honestly in cargo size rather than hiding behind a heap high-water mark.

The API forces capacity thinking at every insertion. push, push_str, insert, and extend return Results carrying the rejected value, so callers decide among drop-oldest, drop-newest, overwrite, or signal backpressure. That explicitness is the point: a telemetry queue that is full is a fact about the world, and your code should count it, sample it, or shed it, not panic inside an ISR. Prefer spsc::Queue for interrupt-to-main handoff with its lock-free single-producer single-consumer contract, mpmc::Q for multi-core RP2040 paths, LinearMap for tiny lookup tables under 32 entries, and HistoryBuffer for rolling filters that overwrite oldest automatically. Reach for IndexMap only when hashing earns its code size over linear search.

Sizing is an engineering task with a paper trail. Capture worst-case burst rates from logic-analyzer traces or field logs, multiply by the longest consumer stall, add 50 percent margin, then round up to a power of two where the structure benefits. A 1 kHz ADC feeding a main loop that stalls 40 ms during flash writes needs at least 40 samples plus margin, so a 128-entry queue is prudent and a 16-entry one is a scheduled incident. Record the sizing math in a comment above each declaration with the measured numbers, because the next engineer will otherwise shrink the pretty constant that looked arbitrary. Add counters for full events exposed over defmt so capacity pressure pages you before it drops data silently.

Ownership patterns keep heapless buffers fast and safe. Declare queues as statics behind a critical-section Mutex, split producer and consumer handles at startup, and never pass borrowed halves across interrupt boundaries by value. Keep formatting out of the producer: push raw codes and timestamps, format downstream where allocation-free defmt handles integers directly. For strings, cap aggressively; a 64-byte heapless::String holds most diagnostic lines, and longer text belongs in chunked writes rather than one heroic buffer. When a buffer must cross cores on RP2040, use the multicore-aware queue from embassy-sync or the mpmc variant rather than hand-rolled atomics around a Vec.

Test fullness the way you test emptiness. Unit-test push-past-capacity on the host, assert the count increments, inject consumer stalls in QEMU runs, and soak-test overnight with fault injection on the producer rate. Firmware whose queues have proven full behavior under test degrades gracefully in storms; firmware whose queues were only tested half-empty drops the exact packets you needed for the postmortem.

Sharing those queues between interrupt and main needs mutual exclusion without a kernel, and the critical-section crate is the portable answer. critical_section::with masks interrupts briefly and hands you a token, while critical_section::Mutex<T> reveals data only through borrow(cs) tied to that token lifetime, so the compiler enforces no token, no access. Enable the cortex-m critical-section-single-core provider on single-core builds, and switch to the embassy-rp multicore provider with its hardware spinlock on dual-core RP2040 work where masking one core cannot stop the other. Libraries stay neutral by depending only on critical-section while the final app enables exactly one provider, since two providers duplicate symbols and fail the link in a genuinely helpful way.

Cell choice and guard discipline finish the pattern. Wrap Copy scalars in Mutex<Cell<T>> with get and set, and larger configs in Mutex<RefCell<T>> with borrow scopes kept tiny: copy the flag or push the sample, exit the guard, then format or transmit outside. Forbid static mut with a lint so every shared access flows through the token system, and never hold a guard across await, flash writes, or formatting calls that inflate every interrupt deadline. Measure the worst mask window with a GPIO toggle and a scope; if it exceeds a tenth of your fastest ISR period, split the work or move the hot path to lock-free queues.

src/telemetry.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
use heapless::{String, Vec, spsc::Queue};
use critical_section::Mutex;
use core::cell::RefCell;
// Burst math: 1 kHz ADC x 40 ms flash stall = 40 + 50% margin -> 128.
static SAMPLE_Q: Mutex<RefCell<Queue<u16, 128>>> = Mutex::new(RefCell::new(Queue::new()));
static FULL_COUNT: Mutex<RefCell<u32>> = Mutex::new(RefCell::new(0));
pub fn isr_push(sample: u16) {
    critical_section::with(|cs| {
        let mut q = SAMPLE_Q.borrow_ref_mut(cs);
        if q.enqueue(sample).is_err() {
            *FULL_COUNT.borrow_ref_mut(cs) += 1;
        }
    });
}
pub fn drain_into(buf: &mut Vec<u16, 128>) {
    critical_section::with(|cs| {
        let mut q = SAMPLE_Q.borrow_ref_mut(cs);
        while let Some(s) = q.dequeue() {
            buf.push(s).ok(); // buf sized to queue: cannot overflow here
        }
    });
}
pub fn fmt_line(n: usize, out: &mut String<64>) {
    use core::fmt::Write;
    write!(out, "batch n={}", n).ok();
}
🔥Bounded Buffers, Guarded Sharing
Size every queue from traces and count drops as telemetry. Guard shared state with critical-section Mutex and keep each mask window tiny.
📊 Production Insight
A vibration logger sized its queue at 16 entries and dropped 31 percent of bursts during flash writes across 63 units; resizing to 128 from trace math plus a graphed full counter cut drops to zero. In the same codebase a 3 ms flash write inside a critical section had delayed a fast ISR by 2.8 ms, and snapshotting config outside the guard cut the worst mask window to 900 ns.
🎯 Key Takeaway
Size heapless queues from measured bursts with margin and count every drop as data. Guard sharing with critical-section Mutex matched to core count, keep windows sub-microsecond, and forbid static mut.

Testing Split: Host Unit Tests, QEMU Rigs, and On-Target Proof

Bare-metal testing splits along one axis: does this test touch hardware. Pure logic such as filters, parsers, CRCs, and queue policies runs on the host with plain cargo test, no target flags, full std test harness, and mocks for peripheral traits. Structure code so drivers sit behind small traits, then test the logic against fakes at full speed with coverage. This layer should hold 70 percent of your assertions because it runs in milliseconds, parallelizes, and catches regressions before any flash cycle. Keep hardware types out of these modules with feature gates or separate crates so host builds never see PAC registers.

Target semantics need QEMU before they need silicon. qemu-system-arm with -M lm3s6965evb emulates a Cortex-M3 sufficient for startup, exception, and semihosting-exit flows, while -M mps2-an385 covers M3/M4 variants for/thumbv7m images. Wire .cargo/config.toml or a test runner so cargo run launches QEMU with -semihosting-config enable=on,target=native -kernel against your thumbv7m binary, and use panic-semihosting with its exit feature so test pass maps to EXIT_SUCCESS and failure to EXIT_FAILURE visible as process codes. Swap defmt-rtt for defmt-semihosting in the QEMU profile and decode with the qemu-run host tool, keeping every defmt call site identical. These runs prove vector tables, boot order, and fault paths without booking bench time.

On-target tests close the gap QEMU cannot: real clocks, real flash timing, real probe behavior. The defmt-test harness annotates #[test] functions that run on the chip with results streamed over RTT, while probe-rs embedded-test provides a runner that flashes, executes, and reports pass-fail into cargo output. Reserve this layer for what only hardware proves: ADC codes against known voltages, UART baud against a scope, flash endurance across write cycles, and interrupt latency under load. Tag them #[ignore] for default runs and gate them in CI behind a self-hosted runner with a PicoProbe attached, because cloud agents without probes will skip or fake them.

CI should run all three layers in order of cost. Stage one builds host tests plus clippy plus fmt in under two minutes on every push. Stage two cross-builds thumbv6m and riscv32imc images, runs QEMU semihosting suites, and asserts cargo size budgets so a RAM overflow fails the merge. Stage three runs the ignored on-target suite nightly on real boards in a temperature chamber, logging RTT to artifacts for a week of trend graphs. A team that keeps host tests under 60 seconds iterates fearlessly; a team whose only test needs a human to press BOOTSEL stops testing by the third sprint.

Seed the suite with the faults you have already paid for. Add a QEMU test that boots without a debugger attached to catch semihosting stalls, a host test that pushes 129 items into a 128 queue to lock the full policy, and an on-target soak that samples for 8 hours at temperature extremes while asserting zero HardFaults. Each regression test should cite the incident date it prevents repeating. That ledger turns testing from ritual into institutional memory that survives turnover.

The on-target layer has matured past printf-and-hope into real harnesses. defmt-test annotates test functions that execute on the chip with assertions streamed over RTT, so a failing ADC range check prints its values inline with the run. probe-rs embedded-test goes further as a Cargo runner: it flashes the test ELF, runs each case, and reports pass-fail into familiar cargo test output that CI can parse. Structure suites by cost: fast self-tests on every boot in debug builds, full hardware suites on demand, and destructive flash-endurance cases gated behind explicit features so nobody erases calibration by accident.

Self-hosted CI runners close the loop for teams shipping weekly. A small runner with two Pico boards, a USB relay for power cycling, and a temperature chamber for soak tests executes the ignored hardware suite nightly and archives RTT logs as trend artifacts. Track metrics across runs: boot time, worst ISR latency, queue-full counts, and RTT frame loss, each graphed week over week. When a metric drifts, the team investigates a living regression instead of a customer report. That pipeline turns the testing split from advice into machinery that defends every release while developers sleep.

tests/host_and_target.rsRUST
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
// Host layer: pure logic, std harness, runs with cargo test.
#[cfg(test)]
mod host {
    use heapless::Vec;
    #[test]
    fn queue_policy_drops_newest_and_counts() {
        let mut v: Vec<u16, 4> = Vec::new();
        for i in 0..4 { v.push(i).unwrap(); }
        assert!(v.push(99).is_err()); // full is a testable event
        assert_eq!(v.len(), 4);
    }
}
// Target layer sketch (defmt-test, runs on chip via probe):
// #![no_std] #![no_main]
// #[defmt_test::tests]
// mod t { #[test] fn adc_within_range() { defmt::assert!(sample() < 4096); } }
// QEMU runner wiring (shell):
// qemu-system-arm -M lm3s6965evb -semihosting-config enable=on,target=native -kernel target/thumbv7m-none-eabi/debug/hello
// cargo test --lib  # host only, sub-second
💡Host Fast, QEMU Faithful, Hardware Final
Prove logic on the host, boot and faults in QEMU, timing and analog on silicon. Each layer catches what the others cannot see.
📊 Production Insight
A filter bug that passed 400 host tests aliased at 1 kHz sampling on real ADC clocks and escaped to 200 units, costing a recall of the calibration constants. Adding a nightly on-target soak with known-voltage fixtures caught the next three DSP regressions before they left the branch.
🎯 Key Takeaway
Keep hardware behind traits for host speed, prove startup in QEMU with semihosting exit codes, and reserve silicon for clocks and analog. CI runs all three in cost order.
● Production incidentPOST-MORTEMseverity: high

RP2040 field units froze: semihosting prints stalled control loops for 41 minutes

Symptom
Twelve deployed RP2040 sensor nodes stopped reporting over USB serial in a rolling pattern starting at 02:10. Watchdog LEDs froze solid instead of blinking at 2 Hz. Bench units with a debugger attached stayed alive for days. The failure interval averaged 6.2 hours, and power-cycling restored each node for another 6 hours. CPU load looked normal in the last transmitted packets, free RAM held steady at 41 KB, and no panic frames were captured over RTT because RTT had been disabled in the release profile.
Assumption
The team assumed a memory leak or heap fragmentation because the 6-hour pattern smelled like exhaustion. They had recently enabled an allocator to use String in log formatting, and the prime suspect was a growing telemetry buffer. Two engineers spent a day auditing every push and format call. They also assumed debug and release behaved identically apart from speed, so they kept testing with the debugger attached, which masked the fault because the debugger serviced the semihosting requests quickly on the bench.
Root cause
A debug hprintln from cortex-m-semihosting was left inside the 1 kHz timer interrupt used for sensor sampling. On the bench with OpenOCD attached, each call cost under 1 ms and the loop survived. In the field with no debugger, each semihosting breakpoint trapped the core for 8-14 ms waiting on a debug host that was not there, so interrupts stacked, the sampling queue overran its 128-entry heapless spsc::Queue within 22 seconds of the first backlog, and the main loop deadlocked waiting on a flag the ISR could no longer set. The 6-hour trigger was a temperature-compensation branch that only logged below 4 degrees, which is why nights failed and days passed. Release builds kept the call because it sat behind a custom dbg flag that was never wired to the release profile.
Fix
The semihosting call was deleted and replaced with a defmt::trace behind a compile-time level filter, transported over RTT with a 1024-byte channel so a disconnected host drops rather than stalls. The timer ISR was cut to 14 lines that only push raw ADC codes into the queue, with all formatting moved to the main loop. A CI gate now fails the build if cortex-m-semihosting appears in release dependencies, verified with cargo tree -e normal --target thumbv6m-none-eabi, and an on-target soak test runs 8 hours in a fridge at 2 degrees with RTT logging asserted loss-free.
Key lesson
  • Never ship semihosting in release firmware: it halts the core on a breakpoint and its cost depends on whether a debugger happens to be attached, so bench results lie. Gate it out with a cargo tree check in CI and use defmt over RTT where a missing host degrades to dropped frames instead of a frozen device.
  • Keep interrupt handlers minimal and infallible: sample, push to a fixed-capacity queue, clear the flag, exit. Formatting, logging, and allocation belong in the main loop or an Embassy task where backpressure is visible and testable rather than hidden inside an ISR.
  • Test the exact release artifact under field conditions: same profile, no debugger, realistic temperature and timing. The 6-hour freeze only reproduced in a cold soak without a probe attached, which is precisely the configuration nobody had tried before deployment.
Production debug guideSeven failure shapes that cover most dead-Pico and silent-firmware nights, each with the exact command that separates hardware faults from your code.7 entries
Symptom · 01
cargo build succeeds on host but fails for the Pico target with missing core or unresolved cortex-m-rt
→
Fix
You are building for the wrong triple or the target is not installed. Run rustup target list --installed to confirm thumbv6m-none-eabi is present, add it with rustup target add thumbv6m-none-eabi, then pin it in .cargo/config.toml with [build] target = "thumbv6m-none-eabi". Rebuild with cargo build --target thumbv6m-none-eabi and inspect cargo tree -e normal to confirm cortex-m-rt and rp2040-hal resolve for the ARM triple rather than the host.
Symptom · 02
Linker errors about memory.x, _stack_start, or FLASH/RAM overflow on an RP2040 build
→
Fix
The linker cannot see your memory layout. Run ls memory.x .cargo/config.toml to confirm both files exist, then cat memory.x and check FLASH ORIGIN 0x10000000 LENGTH 2048K and RAM ORIGIN 0x20000000 LENGTH 264K. Run cargo clean followed by cargo build --target thumbv6m-none-eabi because stale artifacts hide memory.x edits, and run cargo size --bin firmware --release -- -A to see section sizes against those 264 KB of RAM before you blame the HAL.
Symptom · 03
Firmware flashes but the LED never blinks and no RTT output appears
→
Fix
Check the runner and chip name first. Run probe-rs list to confirm the probe sees the RP2040, then verify .cargo/config.toml contains runner = "probe-rs run --chip RP2040". Flash explicitly with cargo run --target thumbv6m-none-eabi and watch for RTT channel banners. If flashing works but nothing runs, hold BOOTSEL, re-seat USB, and retry with probe-rs run --chip RP2040 target/thumbv6m-none-eabi/debug/firmware to rule out a Cargo runner misconfiguration.
Symptom · 04
defmt output shows garbled frames or nothing at all while the app clearly runs
→
Fix
Your transport and decoder disagree. Confirm use defmt_rtt as _ appears exactly once in the binary and that rtt-target is not also linked, since the two RTT implementations conflict. Run probe-rs run --chip RP2040 and check for the defmt frame header; if you switched to semihosting for QEMU, swap back with cargo remove defmt-semihosting and cargo add defmt-rtt plus restoring the import in src/main.rs. Set DEFMT_RTT_BUFFER_SIZE to 1024 or 2048 as a power of two when RAM is tight.
Symptom · 05
HardFault or UsageFault immediately after reset on Cortex-M, often after adding an interrupt handler
→
Fix
Capture the exception frame instead of guessing. Run cargo run with panic-probe plus defmt so the fault prints registers over RTT, then check that every #[cortex_m_rt::exception] and interrupt binding matches your PAC features. Run cargo objdump --bin firmware --target thumbv6m-none-eabi -- -d to disassemble the vector table region and confirm the stack pointer word points inside 0x20000000 plus 264K. A single wrong interrupt name silently binds to the default handler that loops forever.
Symptom · 06
Intermittent data corruption between main loop and ISR sharing a counter or buffer
→
Fix
You have an unguarded shared static. Search with grep -rn "static mut" src/ and replace each hit with critical_section::Mutex<core::cell::RefCell<T>> accessed only inside critical_section::with or cortex_m::interrupt::free. Verify no direct accesses remain, then run cargo clippy --target thumbv6m-none-eabi to catch new unsafe shortcuts. On RP2040 with both cores active, confirm you use the multicore-safe critical section from embassy-rp rather than the single-core cortex-m one.
Symptom · 07
Tests pass on the laptop but firmware misbehaves on the Pico around parsing or math
→
Fix
Split the test strategy. Run cargo test --lib to execute pure logic on the host, then run the same logic under qemu-system-arm -M lm3s6965evb -semihosting-config enable=on,target=native -kernel target/thumbv7m-none-eabi/debug/hello for target semantics. Keep hardware-touching code behind traits so host tests use fakes, and gate endianness or alignment assumptions with explicit assertions rather than assuming x86 behavior ports cleanly to ARMv6-M.
Bare-Metal Rust Stacks Compared
ApproachConcurrencyHeap neededBest for
rp2040-hal blocking loopSuper-loop plus ISRsNoBring-up, simple sensors, one-loop products
Embassy async tasksStatic async executorNoConcurrent sensors, low-power sleep, multicore Pico
RTIC prioritized tasksInterrupt prioritiesNoHard deadlines, motor control, zero-crossing ISRs
Heap plus allocatorAny of the aboveYesJSON parsing, USB strings, prototype speed over determinism
defmt plus RTT loggingLock-free channelNoDevelopment visibility without stalling control loops
Semihosting printsHalting trapNoQEMU CI tests only, never field firmware
QEMU plus on-target testHost and siliconMixedFull confidence: fast host tests plus hardware proof
⚙ Quick Reference
10 commands from this guide
FileCommand / CodePurpose
srcmain.rsuse heapless::{String, Vec};no_std vs std
srcmain.rsuse cortex_m_rt::entry;The Runtime Contract
.cargoconfig.toml[build]Cross Targets and Toolchains
memory.x/* RP2040 on Raspberry Pi Pico: 2MB QSPI flash, 264KB SRAM. */memory.x
srcbinblinky.rsuse panic_halt as _;cortex-m and rp2040-hal Blink
srcbinembassy_blink.rsuse embassy_executor::{main, Spawner, task};Embassy Async HAL
srclogging.rsuse defmt_rtt as _; // exactly one transport; conflicts with rtt-targetLogging That Fits
.cargoconfig.tomlcargo install probe-rs-tools --lockedFlashing and Live Debug with probe-rs
srctelemetry.rsuse heapless::{String, Vec, spsc::Queue};heapless Queues and Critical-Section Sharing Without a Heap
testshost_and_target.rsmod host {Testing Split

Key takeaways

1
no_std removes heap, threads, and I/O; replace them with heapless capacities, interrupt or Embassy concurrency, and defmt over RTT, each sized and gated deliberately.
2
Own the reset contract
one diverging entry via cortex-m-rt, one minimal panic handler, named HardFault override, and a proven fault path before deployment.
3
Pin the triple to the core
thumbv6m-none-eabi for RP2040, thumbv7em-none-eabihf for M4F, riscv32imc-unknown-none-elf for ESP32-C3 class RISC-V.
4
memory.x is hardware truth with per-board ORIGIN and LENGTH; clean rebuild after edits and fail CI on size overflow.
5
Blink through take, clocks, Pins, and OutputPin proves linker, clocks, and runner agree; centralize crystal and pin maps in one board module.
6
Model Embassy concurrency as await points on Timer and channels; never block the executor with Delay or long critical sections.
7
Keep ISRs to sample and push; format downstream, count every full queue, and ban semihosting from release builds.
8
Test in cost order
host unit tests for logic, QEMU for boot and faults, on-target soaks for clocks and analog, all wired into CI.

Common mistakes to avoid

7 patterns
×

Building RP2040 firmware for the wrong ARM triple

Symptom
Image links fine for thumbv7em but dies at the first interrupt on the Pico, or float math runs 10x slow under a soft-float triple. Developers blame the HAL for a toolchain mismatch.
Fix
Use thumbv6m-none-eabi for RP2040, thumbv7em-none-eabihf for M4F parts, riscv32imc-unknown-none-elf for ESP32-C3. Pin in .cargo/config.toml and assert with rustup target list --installed plus cargo readobj in CI.
×

Copying an STM32 memory.x into a Pico project

Symptom
Clean link, dead board: vectors at 0x08000000 instead of the RP2040 XIP window at 0x10000000, or RAM sized 40K on a 264K part so stacks collide silently after two features land.
Fix
Keep per-board memory.x with commented part numbers, run cargo clean after edits, and gate cargo size output in CI at 90 percent of LENGTH so overflow fails the merge.
×

Shipping semihosting prints in release firmware

Symptom
Bench runs fine with a probe attached, field units stall 8-14 ms per log line or hang outright. Failures cluster at night when a cold-only branch starts logging.
Fix
Use defmt over RTT for hardware, defmt-semihosting only for QEMU. Fail CI if cortex-m-semihosting appears in release deps via cargo tree, and keep all logging out of ISRs.
×

Sharing state through static mut instead of a critical-section Mutex

Symptom
Counters skip, queues tear, and corruption appears only under interrupt load. The bug vanishes under a debugger that changes interrupt timing.
Fix
Wrap shared data in critical_section::Mutex with RefCell or Cell, access only inside critical_section::with, and lint out static mut. Match the provider to core count on RP2040.
×

Unwrapping heapless push results and panicking when full

Symptom
Telemetry gaps become HardFaults: the exact burst you needed for diagnosis triggers a panic inside the producer, wiping the evidence and resetting the node.
Fix
Handle Err(Full) as design: count drops, shed oldest, or apply backpressure. Size queues from trace math plus margin and graph the full counter.
×

Blocking the Embassy executor with Delay or flash writes inside tasks

Symptom
Sibling tasks starve, heartbeats jitter by hundreds of milliseconds, and sleep current stays high because the core never parks.
Fix
Use Timer::after in async code, move slow work to dedicated tasks with channels, and keep critical sections sub-microsecond. Measure sleep current after each migration.
×

Testing only on the host and skipping QEMU plus on-target suites

Symptom
400 green host tests, broken firmware: boot faults, ADC aliasing, and probe-dependent stalls all escape because no test ever ran the real startup or clocks.
Fix
Run cargo test for logic, QEMU semihosting suites for boot and faults, and nightly on-target soaks for analog and timing. Order CI by cost so speed never excuses gaps.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01SENIOR
What exactly disappears when you add #![no_std], and how do you replace ...
Q02SENIOR
Explain the three parts of the Cortex-M runtime contract: #![no_main], #...
Q03SENIOR
Your RP2040 image links but never boots after you borrowed an STM32 memo...
Q04SENIOR
Why is semihosting dangerous in production firmware, and what replaces i...
Q05SENIOR
How do you share a counter between main and an ISR without an OS, and wh...
Q06SENIOR
Sketch your test pyramid for a sensor node: what runs on host, in QEMU, ...
Q01 of 06SENIOR

What exactly disappears when you add #![no_std], and how do you replace Vec, threads, and println?

ANSWER
You lose the OS-backed standard library: heap containers, threads, filesystem, networking, and console I/O. Replacement pattern: heapless fixed-capacity Vec/String with handled Err(Full), concurrency via interrupts plus main loop or Embassy tasks, and logging via defmt over RTT. A heap returns only if you add an allocator, which reintroduces fragmentation risk you must justify.
FAQ · 8 QUESTIONS

Frequently Asked Questions

01
Which Rust target do I use for the Raspberry Pi Pico?
02
Do I need an allocator for bare-metal Rust?
03
How is Embassy different from an RTOS or plain interrupts?
04
Why does my firmware work with the debugger but freeze without it?
05
How do I share data between main and an interrupt safely?
06
What belongs in memory.x for the RP2040?
07
How do I run tests when the code needs hardware registers?
08
probe-rs or picotool for flashing the Pico?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Lessons pulled from things that broke in production.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Embedded. Mark it forged?

28 min read · try the examples if you haven't

←
Previous
Rust Unsafe FFI C Interop
1 / 1 · Embedded