Overview
Curated: · Written: · Reviewed:
Embedded debugging is a chain of evidence from the exact source revision and toolchain to an ELF with symbols, the programmed image, target silicon and board, debug probe, transport, debug server, architecture description, reset sequence, and debugger UI. GDB requires host symbols that exactly match the executable on the target. A breakpoint reached with a convenient debug build is not proof about a differently optimized production image.
Halting, single-stepping, software breakpoints, register/memory reads, watchpoints, semihosting, and debugger-driven reset can all perturb the system. They can stop timers while peripherals continue, change interrupt timing, clear read-to-clear or other read-sensitive registers, patch flash/RAM, keep clocks or power domains active, trip watchdogs, or hide races. Optimized code legitimately has inlined functions, reordered statements, and variables with no current storage location.
A disciplined workflow begins with reproduction and evidence capture, identifies whether the fault is reset, exception, memory corruption, concurrency, timing, peripheral, power, or tool mismatch, and uses the least intrusive instrument that can falsify a hypothesis. It combines exact ELF/map/disassembly, fault registers and stacked context, retained crash records, logs/counters, hardware trace, logic/power analysis, sanitizers/static analysis on representative builds, and controlled fault injection.
The production invariant is observation-to-binary fidelity: every debugging claim identifies the exact binary, target/configuration, reset and run-control state, observation mechanism, its perturbation, and corroborating evidence. Debug access must also be authenticated or disabled according to product lifecycle; an exposed remote debug server or unlocked field port is privileged control, not harmless telemetry.
A halt is not a camera. Stopping the core on a data watchpoint leaves timers, DMA, and the other core running; a UART FIFO that overruns 80 µs later is an artifact of the halt, not of the bug. ETM/ITM trace at 4-bit TPIU and 84 MHz can export ~42 MB/s and still drop packets when instrumentation exceeds that; dropped packets look like missing function entries. Prefer SWO/ITM timestamps on the release image, or GPIO bits sampled by a logic analyzer at 10 ns, when the question is timing. Single-stepping through an ISR that feeds a watchdog of 8 ms will reset the part because step latency is hundreds of milliseconds.
Symbols must match the bytes in flash. An ELF built at commit abc123 with -O2 and the image flashed from a CI artifact at def456 produces a backtrace that names the wrong function for 40% of PCs after LTO. Hash the programmed region against the ELF’s load segments before trusting a stack. Semihosting BKPT on a board with no debugger attached is an infinite wait; shipping firmware that still links printf to semihosting hangs in the field the first time a log path is hit. RTT is less invasive but still steals RAM and a background loop; budget it.
Production debug policy is a security control. Leaving SWD enabled with no RDP on a 2-layer board is equivalent to UART root. RDP level 1 still allows a mass-erase path; level 2 is irreversible on many STM32 parts and must be a manufacturing step with a recovery story that does not need the debug port. Crash breadcrumbs in a no-init region — stacked PC, CFSR, HFSR, reset cause, 32-byte ring — survive reset if the bootloader does not zero that section. That 64-byte record is worth more than a debugger that cannot be attached in a sealed enclosure.
ITM stimulus ports overflow silently. Port 0 at 2 Mbps SWO with 200 kB/s of printf will drop; the debugger shows a truncated line and the developer "fixes" a bug that was a dropped digit. Count ITM overflow in DWT or throttle. Watchpoints that halt on a DMA-written variable halt the CPU while DMA continues, so the value that triggered is not the value you then read. Use trace watchpoints or a circular buffer in SRAM sampled by DMA to a GPIO-timed log.
JTAG versus SWD pinmux: PA13/PA14 as SWDIO/SWCLK versus GPIO; a board that remuxes those pins after boot to bit-bang a display loses the debug port until reset, which is fine, unless an error path remuxes them in the field and a RDP-locked part can never be diagnosed. Document the pinmux windows. Core dumps: 256 bytes of no-init (r0–r3, r12, lr, pc, xpsr, cfsr, hfsr, mmfar, bfar, reset_cause, sp, 8 stacked extra) plus a 16-bit CRC is enough to classify 80% of field faults. If the bootloader zeroes all SRAM, the dump is theatre. Exclude the dump region from BSS.
Debug verification is a release-image ritual that starts at an ELF hash and ends at a field breadcrumb. Record SYSCLK, wait states, compiler flags, .map sizes, painted stack high-water marks, ISR GPIO timing, logic-analyzer traces of CS/SCK/SDA, current-shunt waveforms at not less than 100 kHz, reset-cause and fault registers, and the boot slot/security counter after every power-loss injection. A pass is a number that can be recomputed from those artifacts: flash LOAD versus FLASH LENGTH, ISR high-water versus period, Stop current versus the schematic budget, confirm window versus the health checks, and disable-to-deny for debug and keys. If the only evidence is a green LED, a UART log, or a debugger session on an -O0 build, the claim is unpublished. Repeat the same measurements at the temperature and voltage corners the datasheet allows, because flash wait states, Stop leakage, crystal error, and brownout thresholds all move, and a 25 °C passing suite is not a 85 °C passing suite.
Production lock is staged: development leaves SWD on, factory burns RDP1 after a dump-fail test, field never offers a static unlock. A stale OpenOCD script that reset-inits clocks differently from the application hides a hang that only happens on power-on reset. Always reproduce with NRST and power-cycle, not only with a debugger reset that runs reset-init. Record probe firmware, server version, and the ELF SHA-256 next to the ticket.
Worked example: debugger reset is not a power cycle
Hang only after POR. Flashed image SHA-256 c0ffee. Host ELF is -O0 from commit abc123.
| reproduction | ELF matches flash | hang appears |
|---|---|---|
| OpenOCD reset-init, -O0 symbols | no | no (script clocks hide it) |
| NRST + power cycle, hashed LOAD vs ELF | yes (-O2 CI artifact) | yes |
A backtrace on the wrong binary is fiction. Hash first, then NRST, then the debugger.
