Top 100 Embedded Systems Engineer Interview Questions and Answers
The questions most likely to actually come up in your Embedded Systems Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1The map file shows .text plus .rodata larger than the 512 KiB flash. The image still flashes. What failed?(show answer)
I pin linker overflow of on-chip flash to a scope channel or a current shunt before I trust any LED or printf about it.
I treat the linker map as the only size that matters. A flasher that wraps or a bootloader that accepts a truncated image will still light an LED while the vector that lives past the last sector is garbage.
Concretely, i add an ASSERT in the linker script that .text + .rodata + .data load is strictly less than FLASH LENGTH, fail the CI build on that ASSERT, and refuse a .bin whose size is not equal to the LOAD region.
The reason for that specificity is a failure I have seen: I shipped a 530,944-byte image onto 524,288 bytes of flash. The programmer wrote 524,288 bytes and returned OK. HardFault_Handler at 0x0807_F800 was the first 2,048 bytes of the next function, so every usage fault jumped into data.
512 KiB part, image that does not fit.
| Region | Bytes |
|---|---|
| FLASH LENGTH | 524,288 |
| .text | 498,176 |
| .rodata | 32,768 |
| LOAD total | 530,944 |
| overflow | 6,656 |
I would not consider it settled without evidence: I would require the map file's LOAD size to be printed next to FLASH LENGTH, and a post-flash CRC of the whole 524,288-byte region to match the built .bin. That check must fail on this image, which is 6,656 bytes over.
A flasher OK is not a flash that holds the image.
Curated: · Written: · Reviewed:
QA-2The application lives at 0x0800_4000. Interrupts still hit bootloader handlers. What did you forget to write?(show answer)
My first question on VTOR after bootloader handoff is which register bit actually flipped, not which breakpoint I hit.
I treat VTOR as part of the handoff contract. Relocating the image without writing SCB->VTOR leaves IRQ vectors at the bootloader's table, so the LED that the reset handler toggles is not evidence that the application's USART ISR will run.
Concretely, i write SCB->VTOR = 0x08004000 after the bootloader jumps, with a DSB, then I fire a software interrupt whose vector exists only in the application table and require that handler to run.
The reason for that specificity is a failure I have seen: I left VTOR at 0x0800_0000. USART1 IRQ 37 still entered the bootloader stub at offset 0xD4. The application printed "up" on UART from main and never received a byte. We spent two days blaming the baud rate.
IRQ 37 vector with VTOR left at the bootloader.
| Table base | IRQ 37 vector address | Who runs |
|---|---|---|
| 0x0800_0000 bootloader | 0x0800_00D4 | bootloader stub |
| 0x0800_4000 application | 0x0800_40D4 | application ISR |
I would not consider it settled without evidence: I would pend a software IRQ whose application vector writes a known SRAM cell, then read that cell on a debug probe with the bootloader table zeroed in a second capture.
Handoff is VTOR plus the stack pointer, not a branch.
Curated: · Written: · Reviewed:
QA-3You write GPIO BSRR then spin 2 cycles and sample the pin on an input. It is still low. Why is volatile not enough?(show answer)
I treat volatile store versus DSB on posted MMIO as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat volatile as a compiler constraint and DSB as a bus-completion barrier. DSB drains the store so later observers on the AHB see BSRR; it does not wait for the pad voltage to finish rising.
Concretely, after a store another bus master must observe, I issue DSB. To prove the pin, I wait the datasheet output delay or I measure the pad with a scope. I do not treat a CPU load of IDR two instructions later as pad proof.
The reason for that specificity is a failure I have seen: At 80 MHz a cycle is 12.5 ns. I stored BSRR, issued DSB, then two NOPs (25 ns) and read IDR. The bus write had completed; the pad still needed about 40 ns to rise with a 22 pF load. IDR stayed 0, the firmware took the "pin stuck" path, and the LED still looked fine because a later loop saw the rise.
Posted GPIO write at 80 MHz.
| Event | Time |
|---|---|
| CPU cycle | 12.5 ns |
| DSB then two NOPs | bus store complete, ~25 ns |
| pad rise with 22 pF | ~40 ns (datasheet / scope) |
| IDR sample before t_OP | still 0 |
I would not consider it settled without evidence: I would quote the GPIO t_OP from the STM32F407 datasheet against a scope of the pad, and I would not call the path done because DSB returned.
DSB completes the bus access. The pad is a datasheet delay or a probe.
Curated: · Written: · Reviewed:
QA-4The USART ISR writes SR = 0 to "clear everything." Overrun never comes back. What did that write do?(show answer)
On write-1-to-clear status bits I start from the silicon effect the datasheet names, then ask how I would observe it.
I read the access type before I write a status register. On STM32F407 USART, RXNE is not W1C: the clear sequence is a read of SR followed by a read of DR. Writing 0 to SR is a no-op on RXNE. Writing 1s into W1C bits such as TC still clears more than I handled.
Concretely, i read SR, then DR, to clear RXNE and ORE. I write 1 only to bits the RM lists as W1C. I never store 0 or 0xFFFFFFFF into SR "to be safe."
The reason for that specificity is a failure I have seen: I wrote USART->SR = 0 thinking I cleared RXNE. RXNE stayed set as a level source, so the NVIC re-entered immediately after the handler returned, not once per 86.8 µs byte time. CPU occupancy was 100 percent. A later "fix" of SR = 0xFFFFFFFF cleared TC on a byte still in the shift register and dropped 1 in 40 frames.
STM32F407 USART RXNE clear styles.
| Sequence | RXNE after | ISR behavior |
|---|---|---|
| SR = 0 | still 1 | re-enters immediately |
| SR = 0xFFFFFFFF | TC cleared, RXNE still 1 | frames lost, still nested |
| read SR then DR | 0 | one entry per pending byte |
I would not consider it settled without evidence: I would GPIO-count ISR entries in 1 ms of idle line (must be 0 after one drain) and single-step the STM32 sequence SR-then-DR, requiring RXNE clear only after the DR read.
STM32 RXNE clears by reading SR then DR, not by writing RXNE.
Curated: · Written: · Reviewed:
QA-5You set RCC enable then write the USART BRR in the next instruction. BRR reads back 0. What happened?(show answer)
I refuse to call posted write before a dependent peripheral enable done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat the RCC enable bit as a posted write into another clock domain. The USART registers are not yet on the bus for a number of AHB cycles, so a store to BRR can vanish.
Concretely, after RCC->APB2ENR |= USART1EN I read the bit back, then I insert the two dummy reads the errata names (or a DSB plus a load of the enable), then I program BRR.
The reason for that specificity is a failure I have seen: On an STM32F407 with APB2 at 84 MHz I wrote USART1 BRR immediately after USART1EN. The store vanished. BRR read back 0x0000, the STM32 reset value, which is not 9,600 baud: the generator is off until BRR is programmed. TX idled. A USB-UART adapter showed 100 percent framing when we later assumed 115,200. The LED still blinked because main never read BRR.
USART1 BRR after a posted RCC enable, APB2 = 84 MHz.
| APB2 | Oversample | BRR | Baud |
|---|---|---|---|
| 84 MHz | 16× | 0x02D9 (45 + 9/16) | 84e6/(16×45.5625)=115,226 |
| 84 MHz | 16× | 0x0000 (store dropped / reset) | generator off, not 9,600 |
| wanted | 16× | 0x02D9 | 115,200 (error +0.023%) |
I would not consider it settled without evidence: I would dump BRR over SWD after the dummy-read sequence and require 0x02D9 for 84 MHz APB2, 16×, 115,200, then capture TX bit time 8.68 µs on a scope.
Enable is not instant on the APB.
Curated: · Written: · Reviewed:
QA-6The process stack is 2,048 bytes. How do you make a 64-byte overrun a MemManage, not silent corruption of the TCB?(show answer)
The firmware question under MPU stack-guard region is which path still races after the path I just stepped through.
I put a no-access MPU region at the bottom of every FreeRTOS task stack (PSP). Thread code and exception entry from a task use PSP. Handler locals and nested exceptions use MSP. Those are two budgets; an ISR nest does not overwrite a TCB that lives under the task stack.
Concretely, i align the PSP guard to 32 bytes, set RASR to XN plus AP=no-access, size MSP for the worst nested handler frame separately, enable MemManage, and fail boot if the MPU cannot cover every task plus the MSP guard.
The reason for that specificity is a failure I have seen: Task stack was 2,048 bytes on PSP, watermark 48 bytes free. A 96-byte local in the task ran through the guard into the TCB. pxTopOfStack became 0xA5A5A5A5. The nested EXTI on MSP was 104 bytes and never touched that TCB. The LED still ran on idle's PSP. First context switch HardFaulted 18 minutes later.
2,048-byte PSP versus MSP handler budget.
| Stack | Role | Budget |
|---|---|---|
| PSP task | 2,048 B, 48 B free | 96 B task local smashes TCB |
| MSP | nested EXTI + FP frame | 104 B, separate |
| TCB under PSP | 48 B overwritten | 18 minutes to HardFault |
I would not consider it settled without evidence: I would force a 96-byte local on the task, require MemManage with CFAR in the PSP guard, confirm the TCB still matches 0xA5, and separately watermark MSP after a forced nest.
PSP guards the TCB. MSP guards the handlers. Do not mix them.
Curated: · Written: · Reviewed:
QA-7The datasheet says bits 31:16 of CR1 are reserved, must be kept at reset. You OR in 0xFFFF0000 "for future use." What breaks?(show answer)
I would budget a measurement for reserved bits in a control register the way I budget flash: if it is not on the bring-up list, it will not happen.
I copy reserved bits from a read of the reset value, never from a mask I invented. Reserved can mean must-be-zero, must-be-one, or will-change-the-clock-tree on the next die.
Concretely, i RMW with a mask that names only documented fields, and I add a static assert that the constant I write equals (reset | intended_bits) & ~reserved_must_be_zero.
The reason for that specificity is a failure I have seen: I ORed 0xFFFF0000 into SPI_CR1. On that die bits 21:20 selected a /256 prescaler. SCLK dropped from 8 MHz to 125 kHz. The IMU's 4 MHz max was "fine" so the LED still strobed. Samples arrived 64× late and the complementary filter diverged.
SPI_CR1 reserved bits selecting a prescaler.
| CR1[21:20] | SCLK | Period |
|---|---|---|
| documented 00 | 8 MHz | 125 ns |
| accidental 11 | 125 kHz | 8 µs |
| sample latency | 64× | filter diverged |
I would not consider it settled without evidence: I would read CR1 after init and require it equal the documented field packing, then scope SCLK period against the intended 125 ns at 8 MHz.
Reserved is a contract with a future errata sheet.
Curated: · Written: · Reviewed:
QA-8M4 writes a 32-bit slot then sets a flag. M7 reads the flag then the slot. You still see torn values. Why?(show answer)
Where I have been burned on dual-core mailbox occupancy versus empty is accepting a known race because "that is how MCUs work."
I treat a mailbox as a protocol, not a pair of SRAM words. A flag store can become visible before the payload on a Cortex-M7 with a store buffer, and the M4 can overwrite the slot while the M7 is still copying.
Concretely, i use a hardware HSEM or a pair of occupancy bits with DMB on both sides, copy into a local, then release. I never share a packed struct across cores without a sequence number I validate.
The reason for that specificity is a failure I have seen: Payload was 8 bytes at 0x3000_0000, flag at 0x3000_0008. M7 saw flag=1 with the first 4 bytes of packet n and the last 4 of packet n+1. CRC failed on 2.1 percent of messages at 1 kHz. The LED on M4 toggled on every send so bring-up called it done.
Unsynchronized 8-byte mailbox at 1 kHz.
| Traffic | Count |
|---|---|
| messages in 10 minutes | 600,000 |
| torn CRC failures | 12,600 |
| failure rate | 2.1% |
I would not consider it settled without evidence: I would run a 1 kHz producer with an incrementing 64-bit counter and require zero torn reads over 10 minutes (600,000 messages) on a logic analyzer of the flag pin plus a dump of mismatches.
A flag is not a memory barrier and not a lock.
Curated: · Written: · Reviewed:
QA-9The errata says USART RX overrun on rev Y when BRR is odd. You apply it on rev X. What did you just break?(show answer)
I answer wrong silicon errata workaround by naming the ISR, bus, power state, and boot slot that the demo never entered.
I key every workaround to DBGMCU_IDCODE revision. Applying a later-rev recipe on an earlier die can disable a feature that was already correct, and skipping it on the bad rev leaves the bug.
Concretely, i read the revision at boot, select a table of workarounds, and test each recipe on the rev it names, including a negative test that the recipe is a no-op on other revs.
The reason for that specificity is a failure I have seen: Rev X IDCODE was 0x1001. I blindly set USART_CR3 DMAT on every part because rev Y needed it. On rev X that bit routed RX to a DMA channel we had not enabled. Bytes sat in DR unread. Overrun after 2 bytes at 115,200. LED still blinked from SysTick.
USART DMA bit applied to the wrong revision.
| Rev | IDCODE | DMAT always-on result |
|---|---|---|
| X | 0x1001 | DR stalls, overrun after 2 bytes |
| Y | 0x2000 | required to avoid the errata |
| test length | 10,000 bytes | at 115,200 baud |
I would not consider it settled without evidence: I would log IDCODE, run a 10,000-byte RX test per rev in the matrix, and require zero overrun only on the recipe that matches that IDCODE.
Errata are per die, not per family name.
Curated: · Written: · Reviewed:
QA-10A naked 10 kHz ISR uses the FPU. After it returns, a task's S16 value is garbage. What did the naked ISR not save?(show answer)
For lazy FPU stacking in a mixed ISR I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat a naked ISR that touches the FPU as responsible for the full FP context. Hardware lazy stacking saves the caller-saved S0–S15 plus FPSCR when the ISR uses FP; it does not save callee-saved S16–S31. Lazy stacking also does not corrupt a non-FP ISR when an FP ISR preempts it.
Concretely, i avoid naked FP ISRs, or I explicitly VSTM S16–S31 on entry and VLDM on exit. I compile the FreeRTOS port with a consistent FPU setting so task switch and ISRs agree on the extended frame.
The reason for that specificity is a failure I have seen: I marked the current-loop ISR naked for cycle count. It used S16 as a local. Lazy stacking left S16 as the task's filter state. After return, 1.00f became a denormal. Torque jumped 40 percent for one tick. Preempting a GPIO ISR that did not use FP was fine; the bug was the naked callee-saved path.
Naked FP ISR versus lazy stacking.
| Context | What hardware saves | After return |
|---|---|---|
| FP ISR preempts non-FP ISR | S0–S15 + FPSCR (lazy) | non-FP ISR intact |
| naked ISR uses S16 | nothing for S16–S31 | task S16 corrupted |
| torque glitch | 1 tick | +40% |
I would not consider it settled without evidence: I would dump S16–S31 at ISR entry and exit across 10,000 turns and require them restored, plus a build that is not naked or that saves S16–S31.
Naked plus FPU means you own S16–S31. Lazy stacking does not.
Curated: · Written: · Reviewed:
QA-11uxTaskGetStackHighWaterMark returns 12 words. Do you ship?(show answer)
I would write the check for stack high-water mark before ship before the driver, because the driver will otherwise certify itself.
I treat the watermark as a remaining-margin measurement under the worst ISR nest I can force, not as a happy-path number from a UART console session.
Concretely, i fill stacks with 0xA5, run the densest IRQ burst plus the deepest call tree, then I require remaining margin of at least a basic-plus-FP exception frame (26 words, 104 bytes) plus 25 percent.
The reason for that specificity is a failure I have seen: Watermark was 12 words (48 bytes) after a quiet console test. Nested EXTI plus FP stacking is 26 words, 104 bytes on that ABI. First ESD hit overflowed. The 0xA5 pattern in the TCB became 0x00000000. Ship current still looked like 18 mA because the core was HardFault-looping in RUN.
2,048-byte task stack, quiet versus nested.
| Condition | Remaining |
|---|---|
| UART console only | 48 bytes (12 words) |
| nested EXTI + FP frame | 26 words = 104 bytes |
| ship rule | at least 104 + 32 bytes remaining |
| ESD hit | TCB smashed |
I would not consider it settled without evidence: I would record watermark after a forced 3-level nest, compare to 104 bytes, and refuse ship below 104 plus 32 bytes remaining on that 2,048-byte stack.
A console watermark is not a worst-case stack.
Curated: · Written: · Reviewed:
QA-12The UART ISR does pvPortMalloc(64) to stash a line. Why is that a field-reset even if it usually works?(show answer)
The trade-off in malloc from a UART ISR is which failure I am willing to ship, and I will not ship an unmeasured one.
I never allocate from an ISR. FreeRTOS heap_4 takes the scheduler lock or a critical section, not a task mutex. FromISR malloc can deadlock a task that already suspended the scheduler, or it can NULL-deref.
Concretely, i use a static pool or a FromISR queue of pointers into a preallocated ring, sized to the burst I measured, and I count drops rather than growing.
The reason for that specificity is a failure I have seen: At 115,200 baud a 64-byte line every 6 ms is 10,667 bytes/s. The ISR called pvPortMalloc. After 40 minutes the heap had 3 free blocks of 32 bytes and malloc returned NULL. The ISR wrote to 0. HardFault. A second crash was a task in vTaskSuspendAll around malloc when the ISR also entered the heap critical section. The LED had been green in the 5-minute demo.
64-byte lines malloc'd in the USART ISR.
| Item | Value |
|---|---|
| baud | 115,200 |
| line every | 6 ms |
| bytes/s | 10,667 |
| time to heap stall | 40 minutes |
| demo that passed | 5 minutes |
I would not consider it settled without evidence: I would run a 1-hour soak with heap tracing disabled in the ISR and a watermark on the static pool, requiring zero ISR allocations in the map file and zero NULL paths.
Heap is a scheduler lock. The ISR is not a task.
Curated: · Written: · Reviewed:
QA-13A file-scope uint32_t filter_state is not 0 after reset. Where did the zeroing go?(show answer)
I pin BSS zeroing versus a load-time assumption to a scope channel or a current shunt before I trust any LED or printf about it.
I treat .bss as a contract with the reset handler. If I skip the zero loop to save 12 ms of boot, or I put the object in .noinit, the C abstract machine I wrote the filter against is a lie.
Concretely, i keep the BSS zero loop, I time it, and I only place objects in .noinit with a comment and a first-boot sentinel. I fail a unit test that reads filter_state before any store.
The reason for that specificity is a failure I have seen: BSS was 48,128 bytes at 48 MHz, about 1 ms to zero with a word store. Someone #if 0'd the loop to hit a 20 ms boot budget. filter_state came up as 0xFFFFFFFF from SRAM remnant. The IIR output pegged. Current draw 42 mA instead of 18 mA because the actuator saturated.
Skipping BSS zero to save boot time.
| Item | Value |
|---|---|
| .bss | 48,128 bytes |
| zero loop at 48 MHz | ~1 ms |
| boot budget | 20 ms |
| filter_state remnant | 0xFFFFFFFF |
| RUN current | 42 mA vs 18 mA |
I would not consider it settled without evidence: I would dump the first 64 bytes of .bss over SWD immediately after Reset_Handler and require zeros, then measure boot time with the loop in and out.
C says zero. SRAM says remnant.
Curated: · Written: · Reviewed:
QA-14You DMA a packed 5-byte frame into a struct with a uint32_t field. The CRC at the end is always wrong. Why?(show answer)
My first question on packed struct versus DMA alignment is which register bit actually flipped, not which breakpoint I hit.
I treat packed as a CPU access tax and a DMA beat-alignment rule. Cortex-M4 LDR of an unaligned uint32_t succeeds unless CCR.UNALIGN_TRP is set. The CRC bug is the DMA engine writing 5 tight bytes while the CPU interprets a padded struct, or a UsageFault when UNALIGN_TRP is on.
Concretely, i DMA into a uint8_t buffer aligned to the DMA beat, then memcpy field-by-field, or I pad the on-wire layout. I do not blame M4 unaligned LDR by default.
The reason for that specificity is a failure I have seen: Frame was 1 header + 4 payload = 5 bytes. DMA to a packed struct at 0x2000_0101. UNALIGN_TRP was 0, so LDR did not fault; it assembled a word that was not the 32-bit payload on the wire. CRC of the CPU view failed 100 percent. The logic analyzer showed a correct 5-byte frame at 1 Mbps. The LED toggled on DMA TC so the driver was "done."
5-byte packed frame DMA'd into a uint32_t field.
| View | Bytes |
|---|---|
| on the wire | 5 |
| DMA beat | 32-bit |
| destination address | 0x2000_0101 |
| CRC fail rate | 100% |
| analyzer CRC | 0% fail |
I would not consider it settled without evidence: I would dump the SRAM destination as bytes versus the struct, require them to match a hand-packed fixture, refuse DMA to an address not aligned to the beat, and document UNALIGN_TRP.
Packed is for the wire. DMA is for aligned SRAM.
Curated: · Written: · Reviewed:
QA-15now is 2,147,483,647 ticks. deadline is now + 200. The wait returns immediately. What rule did you break?(show answer)
I treat signed overflow in a tick timeout as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat C signed overflow as undefined, and I treat a 31-bit tick as wrapping. A deadline computed as now + dt that overflows a signed 32-bit now compares wrong.
Concretely, i use unsigned ticks, subtract with wrap (int32_t)(now - start) >= dt, and I never write now + dt > now as the still-waiting test.
The reason for that specificity is a failure I have seen: Tick was int32_t at 1 kHz. At 2,147,483,647 I added 200 and the sum was -2,147,483,449. The wait saw deadline < now and skipped a 200 ms debounce. A 20 ms contact bounce closed a relay. Coil current 120 mA for 1.2 s instead of a 12 mA hold.
Signed 1 kHz tick wrapping a 200 ms wait.
| Item | Value |
|---|---|
| now | 2,147,483,647 |
| dt | 200 ticks |
| now + dt as int32 | -2,147,483,449 |
| bounce closed relay | 20 ms |
| coil current | 120 mA for 1.2 s |
I would not consider it settled without evidence: I would unit-test the wait at tick values 2,147,483,000 through wrap, and on-target I would set the tick near INT32_MAX with a debugger and require the 200 ms GPIO pulse width on a scope.
now + dt > now is not a timeout. It is overflow.
Curated: · Written: · Reviewed:
QA-16You enable -flto --gc-sections. USART1_IRQHandler is absent and the slot is the weak Default_Handler. What root did the linker drop?(show answer)
On KEEP missing on the vector table I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat the vector table as a real reference to each handler symbol. LTO does not rewrite a live USART1_IRQHandler into Default_Handler. What does is --gc-sections without KEEP(.isr_vector), so the table and the strong handler look unused, or a misspelled name that leaves the weak alias.
Concretely, i KEEP the .isr_vector section, I nm the .elf so the USART1 slot is not Default_Handler, and I fail CI if a used IRQ still aliases the weak default.
The reason for that specificity is a failure I have seen: Without KEEP, gc-sections dropped our USART1_IRQHandler object. The startup weak alias stayed at 0x0800_01A0. RX entered Default_Handler and BKPT. With KEEP, the handler was 0x0800_4C20. The UART printf from main still worked in polling, so the console demo passed.
USART1 handler with and without KEEP(.isr_vector).
| Build | USART1 vector | RX path |
|---|---|---|
| KEEP + LTO | 0x0800_4C20 | live handler |
| gc-sections, no KEEP | 0x0800_01A0 weak Default | BKPT hang |
| misspelled IRQHandler | weak alias remains | same hang |
| main polling printf | works | demo passed |
I would not consider it settled without evidence: I would require nm and the vector dump to agree on every IRQ we enable, and a CI step that fails if a used IRQ's handler is Default_Handler.
KEEP the table. Do not blame LTO for a weak alias.
Curated: · Written: · Reviewed:
QA-17A library throws std::bad_alloc. The image was built -fno-exceptions. What runs?(show answer)
I refuse to call C++ exceptions on a firmware with -fno-exceptions done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat -fno-exceptions as a promise that no throw exists in the link. A throw then calls terminate, which on this runtime is a BKPT, which is a field brick if the debugger is absent.
Concretely, i compile the whole tree -fno-exceptions -fno-rtti, I nm for __cxa_throw, and I replace the one library that throws with a noexcept allocator that returns a static error object.
The reason for that specificity is a failure I have seen: new uint8_t[256] failed after 80 KB of heap was used of 96 KB. throw called terminate. Without a debugger the core spun in a 4-instruction loop at 64 MHz, 18 mA, instead of the 120 µA sleep we certified. The green LED was still on because it is a GPIO left high.
bad_alloc with -fno-exceptions in the field.
| Item | Value |
|---|---|
| heap | 96 KB |
| used at throw | 80 KB |
| request | 256 bytes |
| field current | 18 mA |
| certified STOP | 120 µA |
I would not consider it settled without evidence: I would nm the .elf for throw/terminate, run an OOM test with the probe detached, and require a defined error path plus sleep current ≤ 200 µA.
No exceptions means no throw in the map file.
Curated: · Written: · Reviewed:
QA-18DMA fills a 1,024-byte RX buffer. The CPU CRC is always stale by 32 bytes. What did you not clean?(show answer)
The firmware question under DMA into a cacheable SRAM buffer is which path still races after the path I just stepped through.
I treat D-cache as a second copy of SRAM. DMA writes the memory. The CPU reads the cache. Without an invalidate after DMA complete, I CRC yesterday's bytes.
Concretely, i place DMA buffers in a non-cacheable MPU region or I SCB_InvalidateDCache_by_Addr on the exact range after TC, before any CPU load, and I align the range to 32-byte lines.
The reason for that specificity is a failure I have seen: Buffer was 1,024 bytes in DTCM-off SRAM, cacheable. DMA wrote 1,024. CPU CRC ran on 992 fresh bytes plus 32 cached. Fail rate 100 percent on the last line. The DMA TC LED fired, so the driver returned success.
1,024-byte DMA RX with a 32-byte D-cache line.
| Item | Bytes |
|---|---|
| DMA transfer | 1,024 |
| cache line | 32 |
| stale at CRC | 32 |
| CPU CRC fail | 100% |
| TC LED | success |
I would not consider it settled without evidence: I would dump cache-line 31 before and after invalidate, require the CRC of the SRAM view to match a scope decode of the 1,024 bytes, and fail the driver if the buffer is not line-aligned.
DMA complete is not CPU-visible until the cache agrees.
Curated: · Written: · Reviewed:
QA-19You store a crash PC in .noinit. After IWDG it is 0. Who cleared it?(show answer)
I would budget a measurement for no-init RAM across a watchdog reset the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat .noinit as a section the reset handler must not zero and the bootloader must not use as scratch. A watchdog reset still runs BSS code unless I branch around it for a software-reset cause.
Concretely, i put the breadcrumb in a section with NOLOAD, skip zeroing when RCC_CSR shows IWDG, and I verify the linker places it outside the bootloader's SRAM window.
The reason for that specificity is a failure I have seen: IWDG fired after 8 s. Reset_Handler zeroed all of SRAM including the 16-byte .noinit at 0x2000_7FF0. The field log always showed PC=0. We blamed "random" faults. STOP current still 8 µA because the part did reset.
16-byte crash breadcrumb lost on IWDG.
| Item | Value |
|---|---|
| IWDG timeout | 8 s |
| .noinit | 16 bytes at 0x2000_7FF0 |
| after reset PC log | 0 |
| STOP current after reset | 8 µA |
I would not consider it settled without evidence: I would write 0xA5A5A5A5 into the breadcrumb, trip IWDG, and require the same word back over SWD before any C++ constructor runs.
NOLOAD is a linker flag. The reset handler still has a loop.
Curated: · Written: · Reviewed:
QA-20crc32() calls itself on 512-byte chunks. The task stack is 1,024 bytes. When does it die?(show answer)
Where I have been burned on recursion in a CRC helper on a 1 KB stack is accepting a known race because "that is how MCUs work."
I treat recursion as stack = frame × depth, and I measure the frame. A CRC that splits in half on a 1 KB stack will nest until the frame does not fit.
Concretely, i write the CRC as a loop, I compile with -Wstack-usage=256, and I fail the build if any frame exceeds the budget I set from the watermark.
The reason for that specificity is a failure I have seen: Frame was 44 bytes. Depth for 32 KiB of flash CRC by 512-byte halves is log2(32768/512)+1 = 7, so 7×44=308 bytes, which fit. Someone changed the chunk to 32 bytes: depth log2(32768/32)+1 = 11, 11×44=484 bytes. Then they CRC'd 1 MiB of QSPI: depth log2(1048576/32)+1 = 16, 16×44=704 bytes, plus the 400-byte caller, overflow of a 1,024-byte stack. HardFault at 12 ms into the CRC. LED still on.
Recursive CRC32 on a 1,024-byte stack, 44-byte frames.
| Chunk | Payload | Depth | Stack used |
|---|---|---|---|
| 512 B | 32 KiB | 7 | 7×44=308 B |
| 32 B | 32 KiB | 11 | 11×44=484 B |
| 32 B | 1 MiB QSPI | 16 | 16×44=704 B + 400 B caller |
I would not consider it settled without evidence: I would -fstack-usage and a recursion-depth assert, then CRC the real image size on the target and require the watermark to stay above 64 bytes.
Halving is still depth. Depth is still stack.
Curated: · Written: · Reviewed:
QA-21USART ISR runs 140 µs. Bytes arrive every 86.8 µs at 115,200. What do you measure before you "optimize later"?(show answer)
I answer ISR duration versus UART byte time by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat ISR wall time as a budget against the shortest re-arrival I have enabled. If the handler is longer than the byte time, I lose bytes regardless of the LED that toggles at the end of the handler.
Concretely, i toggle a GPIO around the ISR, scope the high time, and I require it ≤ 50 percent of the byte time so a nest still fits. I move work to a task.
The reason for that specificity is a failure I have seen: Byte time at 115,200 8N1 is 86.8 µs. My handler CRC'd 64 bytes in 140 µs at 72 MHz (about 10,080 cycles). Overrun SR bit set every other byte. Throughput 5,760 bytes/s instead of 11,520. The printf of the last good line still looked fine.
115,200 baud versus a 140 µs USART ISR.
| Item | Value |
|---|---|
| bit time | 8.68 µs |
| byte 8N1 | 86.8 µs |
| ISR high time | 140 µs |
| cycles at 72 MHz | 10,080 |
| goodput | 5,760 B/s of 11,520 |
I would not consider it settled without evidence: I would scope ISR GPIO high time and the UART RX line together, and refuse the handler if high time > 43 µs on that baud.
A finished ISR is not a caught byte.
Curated: · Written: · Reviewed:
QA-22Both IRQs are priority 5. You expected the 20 kHz loop to preempt the 1 kHz log. Why does grouping stop that?(show answer)
For NVIC grouping and preemption I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat NVIC grouping as a split between preemption and subpriority. Same preemption group means tail-chain or wait, never preempt, even if the numbers in the CMSIS call look different.
Concretely, i set grouping once, I encode priorities with NVIC_EncodePriority, and I document which IRQs may preempt which. I test with GPIO latency from the high-rate IRQ while the low-rate IRQ is artificially long.
The reason for that specificity is a failure I have seen: Grouping was 4 bits subpriority, 0 bits preemption. Both "priority 5" were the same group. The 1 kHz log ISR ran 180 µs. The 20 kHz loop jittered to 180 µs. Control error doubled. A debugger halt in the log ISR made it look like the loop still ran, because halt stops both.
20 kHz loop stuck behind a 1 kHz log ISR.
| IRQ | Rate | Handler | Observed start |
|---|---|---|---|
| log | 1 kHz | 180 µs | on time |
| loop | 20 kHz | 12 µs | delayed 180 µs |
| control error | — | — | ×2 |
I would not consider it settled without evidence: I would capture both ISR GPIOs on a logic analyzer and require the 20 kHz high time to start inside the 1 kHz handler when grouping allows preemption.
Subpriority is a tie-break, not a preempt.
Curated: · Written: · Reviewed:
QA-23Two IRQs tail-chain. The second always meets its "entry" timestamp. Who is late?(show answer)
I would write the check for tail-chaining hiding a missed deadline before the driver, because the driver will otherwise certify itself.
I treat tail-chaining as a skipped stack pop/push, not as proof the second ISR met its deadline. The second handler can start 6 cycles after the first returns while the first already ate the whole period.
Concretely, i timestamp against a free-running timer at ISR entry and at work complete, and I budget the sum of chained handlers, not each entry time in isolation.
The reason for that specificity is a failure I have seen: IRQ A ran 38 µs, IRQ B 12 µs, period of B was 40 µs. Tail-chain started B 0.08 µs after A. B's entry looked punctual. B's work completed at 50 µs, 10 µs late. PWM compare was written after the update event. Duty stuck at 0 for one 20 kHz cycle. Motor current spiked 2.4 A.
Tail-chained 12 µs ISR after a 38 µs sibling.
| Event | Time |
|---|---|
| IRQ A | 38 µs |
| tail-chain gap | 0.08 µs |
| IRQ B work | 12 µs |
| B complete | 50 µs |
| B period | 40 µs |
| current spike | 2.4 A |
I would not consider it settled without evidence: I would log TIM CNT at entry and at the PWM write, and require complete ≤ period, not entry ≤ period.
Entry time is not completion time.
Curated: · Written: · Reviewed:
QA-24On STM32F407, TIM1_BRK_TIM9 is one vector. You clear TIM1 break and return. TIM9 still reenters. Why?(show answer)
The trade-off in shared IRQ line for TIM1_BRK and TIM9 is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat a shared vector as a list of sources. Clearing the one I stepped in the debugger leaves the other pending, which looks like a "spurious" storm. EXTI9_5 is a different STM32F407 vector; it is not TIM1_BRK.
Concretely, i read every status bit that can raise TIM1_BRK_TIM9, clear only those I handle, and I have a default path that logs an unexpected source instead of returning.
The reason for that specificity is a failure I have seen: I cleared TIM1_SR BIF. TIM9 UIF stayed set. Vector reentered at the 20 kHz TIM9 rate. CPU 100 percent. UART printf still flushed one line so the console looked alive. Current 28 mA versus 6 mA in the idle budget.
STM32F407 TIM1_BRK_TIM9 shared vector.
| Source | Rate | Cleared |
|---|---|---|
| TIM1_BRK | 100 Hz | yes |
| TIM9 update | 20 kHz | no |
| CPU | 100% | 28 mA vs 6 mA |
I would not consider it settled without evidence: I would dump TIM1 SR and TIM9 SR at ISR entry for 1,000 entries and require a named source each time, plus a GPIO count matching the expected break rate not the TIM9 rate.
One vector is not one peripheral.
Curated: · Written: · Reviewed:
QA-25The ALERT pin stays low until you read the I2C chip. You configured falling-edge EXTI. You miss the next alert. Why?(show answer)
I pin level versus edge EXTI on a held pin to a scope channel or a current shunt before I trust any LED or printf about it.
I treat a level-held line as a level-sensitive interrupt. An edge already happened while I was in the handler. If I return without the pin being high, I will not get another edge.
Concretely, i use level EXTI if the part has it, or I re-read the pin at the end of the handler and loop until it is released, or I sample in a 1 kHz task while the line is low.
The reason for that specificity is a failure I have seen: ALERT stayed low 4.2 ms while I did a 400 kHz I2C read of 6 bytes (~150 µs) plus a printf (3.8 ms). The falling edge was consumed. Next ALERT was a 80 µs dip that arrived while still printing. Missed. IMU ran unserviced for 200 ms. Angle error 12 degrees.
Falling-edge EXTI on a level-held ALERT.
| Event | Time |
|---|---|
| ALERT held low | 4.2 ms |
| I2C 6 bytes at 400 kHz | ~150 µs |
| printf in ISR | 3.8 ms |
| missed next dip | 80 µs |
| unserviced | 200 ms, 12° error |
I would not consider it settled without evidence: I would scope ALERT versus ISR GPIO and require a handler invocation for every low interval, including a held-low of 5 ms.
A held low is not a second falling edge.
Curated: · Written: · Reviewed:
QA-26The current-loop ISR uses VLDR. CPACR left CP10/CP11 disabled. The hard-fault CFSR says NOCP. What mismatched?(show answer)
My first question on FPU instructions in an ISR compiled as soft-float is which register bit actually flipped, not which breakpoint I hit.
I treat the FPU as a coprocessor that must be enabled in CPACR and as an ABI that must match the rest of the image. A translation unit built -mfloat-abi=soft does not emit VLDR just because a hard-float header was included. NOCP is FPU opcodes while CP10/CP11 are disabled, or a hard-float object linked into a CPACR-off image.
Concretely, i enable CP10/CP11 at reset before any FP opcode, I compile every ISR that touches float with the same ABI, and I fail the link on mixed Tag_ABI_VFP_args.
The reason for that specificity is a failure I have seen: CPACR left CP10 disabled to "save 18 extra stacking words." The 10 kHz ISR was hard-float and executed VLDR. NOCP at 9,200 faults/s. A separate soft-float .c file stayed integer and never inherited VLDR from a header. The LED toggled in the HardFault spin so the board looked alive. Input current 22 mA versus 7 mA.
10 kHz ISR with CPACR FPU gated off.
| Item | Value |
|---|---|
| loop rate | 10 kHz |
| NOCP faults | 9,200 /s |
| RUN current | 22 mA |
| expected | 7 mA |
| GPIO period | 100 µs ± 2 µs when healthy |
I would not consider it settled without evidence: I would read CPACR and the .elf's Tag_ABI_VFP_args, then run 10,000 current-loop turns with a GPIO period of 100 µs ± 2 µs.
Soft-float source is not a VLDR. Disabled CPACR is a NOCP.
Curated: · Written: · Reviewed:
QA-27You CPSID I around a 2 µs SRAM copy. A 20 kHz PWM ISR is 12 µs late. What should you have raised instead?(show answer)
I treat BASEPRI versus CPSID for a short critical section as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat CPSID as a global mask. BASEPRI lets me block only IRQs at or below a threshold so a motor ISR can still preempt a UART copy.
Concretely, i set BASEPRI to the UART's group for the copy, I restore it, and I never disable SysTick or the PWM IRQ for a buffer the PWM does not touch.
The reason for that specificity is a failure I have seen: CPSID I for 2 µs at 80 MHz is 160 cycles, which was fine, but a 40 µs flash wait inside the same lock stretched it. PWM period 50 µs. One compare missed. Duty 0 for 50 µs. Phase current 3.1 A. The debugger halt made the PWM look healthy because halt freezes TIM.
CPSID holding off a 20 kHz PWM ISR.
| Item | Value |
|---|---|
| intended lock | 2 µs (160 cycles at 80 MHz) |
| lock with flash wait | 40 µs |
| PWM period | 50 µs |
| missed compare | 1 |
| phase current | 3.1 A |
I would not consider it settled without evidence: I would scope PWM update versus a GPIO around the critical section and require the PWM ISR to still enter while BASEPRI is set to the UART level.
CPSID is a blunt mask. BASEPRI is a split.
Curated: · Written: · Reviewed:
QA-28A 1.2 µs encoder index pulse arrives while you are in a 6 µs GPIO ISR on the same EXTI. You never see the index. Why?(show answer)
On missed edge on a pulse shorter than the ISR I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat an edge detector as a one-bit latch. If a second edge arrives while the line is still pending and I clear PR at the end, I can lose the second edge unless I sample the pin or use a timer input-capture.
Concretely, i capture index on a TIM channel in hardware, or I read the pin before clearing PR and I keep servicing until it is idle.
The reason for that specificity is a failure I have seen: Index pulse 1.2 µs, GPIO ISR 6 µs, encoder 4,096 CPR at 3,000 rpm is 204,800 edges/s (4.88 µs apart). Index coincided with a GPIO ISR. Missed. Commutation offset wrong by one mechanical turn. Current 1.8 A RMS versus 0.4 A.
1.2 µs index overlapping a 6 µs EXTI ISR.
| Item | Value |
|---|---|
| index width | 1.2 µs |
| GPIO ISR | 6 µs |
| 4,096 CPR at 3,000 rpm | 204,800 Hz, 4.88 µs |
| missed index | 1 mechanical turn |
| RMS current | 1.8 A vs 0.4 A |
I would not consider it settled without evidence: I would inject a 1.2 µs pulse during a forced 6 µs ISR on a generator and require the capture counter to increment, not an EXTI count.
Pending is one bit. A pulse is not a queue.
Curated: · Written: · Reviewed:
QA-29USART1 stays pending while its handler prints. USART2 at a lower group never runs. What Cortex-M rule did you miss?(show answer)
I refuse to call tail-chaining starving a lower IRQ done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat a Cortex-M exception as unable to preempt itself. The same IRQ remains pending and tail-chains after return. Nesting requires a different IRQ at a higher preemption group.
Concretely, i clear the USART1 source before long work, I keep print out of the ISR, and I give USART2 a group that can preempt only if I truly need nest, which still will not be USART1 preempting USART1.
The reason for that specificity is a failure I have seen: I left RXNE set through a 400 µs printf. USART1 tail-chained every 86.8 µs byte and never returned to thread. USART2 at a lower group starved. After 22 tail-chains the 400 µs print still ran to completion on the same frame; MSP did not grow 22 times. We had sized MSP for illegal same-IRQ nest of 22×96=2,112 bytes. LED on.
USART1 pending versus USART2 starved.
| Item | Value |
|---|---|
| byte time | 86.8 µs |
| printf in USART1 ISR | 400 µs, tail-chain, depth 1 |
| USART2 | lower group, 0 entries during flood |
| illegal self-preempt | cannot happen on Cortex-M |
I would not consider it settled without evidence: I would GPIO-count ISR depth and require depth=1 on USART1, plus USART2 entries during a 1-second flood at 115,200.
Same IRQ tail-chains. It does not nest on itself.
Curated: · Written: · Reviewed:
QA-30EXTI0 preempts EXTI1. EXTI0 calls pvPortMalloc. The heap looks fine from a task. Why do you still die?(show answer)
The firmware question under pvPortMalloc inside a nested EXTI ISR is which path still races after the path I just stepped through.
I treat the FreeRTOS heap as locked by vTaskSuspendAll or a critical section, not a task mutex. FromISR pvPortMalloc can deadlock with a task that already suspended the scheduler, or it can corrupt the free list.
Concretely, i use a lock-free pool sized at init, I assert configSUPPORT_DYNAMIC_ALLOCATION is not used from ISR files, and I grep the IRQ directory for malloc.
The reason for that specificity is a failure I have seen: A task was inside pvPortMalloc with the scheduler suspended. EXTI0 preempted and malloc'd 32 bytes, re-entering the heap critical section. List corruption. First vTaskSwitchContext HardFault 900 ms later. The EXTI0 LED still fired at 200 Hz.
32-byte malloc from a nested EXTI.
| Item | Value |
|---|---|
| EXTI0 rate | 200 Hz |
| alloc | 32 bytes |
| time to HardFault | 900 ms |
| LED | still 200 Hz |
I would not consider it settled without evidence: I would objdump the ISR objects for bl pvPortMalloc and run a 10-minute preempt soak with heap poisoning, requiring zero ISR heap use.
Nested malloc is two bugs, not a pool.
Curated: · Written: · Reviewed:
QA-31SCL and SDA use 4.7 kΩ to 3.3 V. Cbus is 200 pF. Fast-mode 400 kHz NACKs. What do you change first?(show answer)
I would budget a measurement for I2C RC pull-up rise time the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat I2C rise time as RC, not as a firmware baud. Fast-mode t_r max is 300 ns. tau = R C. The 30–70 percent rise is about 0.847 tau.
Concretely, i measure t_r on a scope, I compute R from Cbus, and I pick 1.0 kΩ when 4.7 kΩ misses 300 ns, watching IOL at 3 mA.
The reason for that specificity is a failure I have seen: R=4.7 kΩ, C=200 pF, tau=940 ns, 0.847 tau=796 ns rise, over the 300 ns limit. At 400 kHz high time is 1,250 ns so the high was still rising at sample. NACK rate 18 percent. The LED on ACK of the first byte still blinked.
I2C Fast-mode rise time versus 4.7 kΩ.
| Rp | tau = RC | 0.847 tau | vs 300 ns |
|---|---|---|---|
| 4.7 kΩ | 940 ns | 796 ns | fail |
| 1.0 kΩ | 200 ns | 169 ns | pass |
| NACK at 4.7 kΩ | 18% | — | — |
I would not consider it settled without evidence: I would scope 30–70 percent on SCL and require t_r ≤ 300 ns at 400 kHz before I change the driver.
Pull-ups are analog. The ISR cannot fix RC.
Curated: · Written: · Reviewed:
QA-32A sensor stretches SCL for 900 µs. Your MCU timeout is 100 µs. Who hangs?(show answer)
Where I have been burned on I2C clock stretching timeout is accepting a known race because "that is how MCUs work."
I treat clock stretching as a slave-legal wait. A master timeout shorter than the slave's worst stretch aborts a transaction that the datasheet allows.
Concretely, i read t_stretch max from the sensor, I set the MCU timeout above that plus bus RC, and I fail bring-up if a scope shows SCL low longer than the timeout without a NACK path.
The reason for that specificity is a failure I have seen: Sensor stretch 900 µs on a 16-bit ADC conversion. MCU timeout 100 µs. Master issued STOP mid-stretch. Sensor held SDA. Bus stuck 42 ms until a bit-bang recovery. Sample rate 10 Hz instead of 200 Hz. The first conversion's LED still strobed.
900 µs stretch versus a 100 µs master timeout.
| Item | Time |
|---|---|
| ADC stretch | 900 µs |
| MCU timeout | 100 µs |
| bus stuck | 42 ms |
| intended sample rate | 200 Hz |
| achieved | 10 Hz |
I would not consider it settled without evidence: I would scope SCL low time over 1,000 conversions and require timeout ≥ max stretch + 20 percent, here ≥ 1.08 ms.
A timeout shorter than stretch is a stuck bus.
Curated: · Written: · Reviewed:
QA-33You set CPHA=0 because "mode 0 is default." The flash samples on the trailing edge. What do you capture?(show answer)
I answer SPI CPHA versus the device datasheet by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat CPOL/CPHA as a pairing with the slave's launch and sample edges, not as a Linux spidev habit. Wrong CPHA reads a bit-shifted stream that still looks like data.
Concretely, i scope MOSI/MISO against SCLK, I match the datasheet's mode number, and I refuse a driver that cannot name sample-on-rising versus falling.
The reason for that specificity is a failure I have seen: Flash wanted mode 3 (CPOL=1, CPHA=1). I used mode 0. JEDEC ID came back 0x00 0x7F 0xFF instead of 0xEF 0x40 0x16. I still got "a ID" so the LED went green. Every read was shifted one bit. CRC of a 256-byte page failed 100 percent.
SPI mode 0 against a mode-3 flash.
| Mode | CPOL | CPHA | JEDEC ID |
|---|---|---|---|
| wanted 3 | 1 | 1 | EF 40 16 |
| used 0 | 0 | 0 | 00 7F FF |
| 256-byte CRC fail | 100% | — | — |
I would not consider it settled without evidence: I would capture SCLK versus MISO for the 0x9F command and require the first bit sampled on the documented edge, ID bytes matching the marking.
A shifted ID is not an ID.
Curated: · Written: · Reviewed:
QA-34APB is 16 MHz. You want 115,200 16× oversample. What BRR do you program, and do you ship the error?(show answer)
For UART baud divisor error at 16 MHz I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat baud error as (actual-wanted)/wanted, and I refuse over about 2 percent with a typical 2 percent crystal and a 2 percent peer. I program STM32 BRR as mantissa plus fraction, not as round(f/(16×baud)) to 9.
Concretely, i compute USARTDIV=f/(16×baud)=8.6806, I pack BRR=0x008B (8 + 11/16), I recompute actual baud, and I pick 8× oversample or a PLL multiple when 16× cannot hit.
The reason for that specificity is a failure I have seen: 16e6/(16×115200)=8.6806. Treating USARTDIV as the integer 9 gives 16e6/(16×9)=111,111 baud (−3.55 percent) and I would not ship that. Packed BRR 0x008B is USARTDIV=8+11/16=8.6875, actual=16e6/139=115,108, error=(115108-115200)/115200=−0.08 percent, which I would ship. At 9,600, USARTDIV=104.1667, BRR=0x0683 (104+3/16), actual=16e6/1667≈9,598, error=−0.02 percent.
16 MHz APB, 16× oversample.
| Wanted | BRR | Actual | Error |
|---|---|---|---|
| 9,600 | 0x0683 (104+3/16) | 16e6/1667≈9,598 | −0.02% |
| 115,200 as integer 9 | mantissa 9, 16e6/(16×9) | 111,111 | −3.55%, do not ship |
| 115,200 packed | 0x008B (8+11/16) | 16e6/139=115,108 | −0.08%, ship |
I would not consider it settled without evidence: I would measure bit time on a scope (wanted 8.68 µs, got 8.687 µs at 115,108 baud) and a BER test of 100,000 bytes against a GPSDO UART.
Rounded integer BRR is not the STM32 fraction field.
Curated: · Written: · Reviewed:
QA-35You need 9-bit addressing. You set M=1 then write 8-bit bytes. Slaves never wake. What bit did you not set?(show answer)
I would write the check for 9-bit UART address mark before the driver, because the driver will otherwise certify itself.
I treat the 9th bit as an address mark in the shift register, not as a software convention on an 8-bit DR write.
Concretely, i set the TXE address bit (or write 9 bits into DR) for the address byte, then I clear it for data, and I verify on a scope that the 9th bit is 1 only on the address frame.
The reason for that specificity is a failure I have seen: I wrote 0x42 as 8 bits with M=1, so the 9th bit was 0. Slaves stayed asleep. Traffic 0 bytes/s of payload. The UART TC LED still fired at 9,600 baud because the master was transmitting.
9-bit address byte versus 8-bit DR writes.
| Frame | 9th bit | Slave |
|---|---|---|
| address 0x42 intended | 1 | wake |
| 8-bit write with M=1 | 0 | sleep |
| master TC LED | on | 9,600 baud |
I would not consider it settled without evidence: I would decode 9-bit frames on a logic analyzer and require mark=1 on the address byte only, then a slave GPIO toggle within 2 byte times.
M=1 without the mark bit is just 9 zeros of data.
Curated: · Written: · Reviewed:
QA-36The sensor is 10-bit address 0x2A5. Your scanner finds nothing at 0x55. Where is the device?(show answer)
The trade-off in 10-bit I2C address versus a 7-bit scan is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat 10-bit addressing as a two-byte header: 11110 | AA | W then the low 8 bits. 0x2A5 write is 0xF4 then 0xA5. A read is 0xF4, 0xA5, repeated start, then 0xF5. It is not 0xF5 then 0xA5 as the first two bytes.
Concretely, i send the 10-bit write header the spec names, I do a repeated-start 0xF5 only when reading, I do not trust a 7-bit ping scan, and I document both the 10-bit value and the first-byte pattern.
The reason for that specificity is a failure I have seen: 0x2A5 >> 3 = 0x54, so a naive scan of 0x55 never ACKed. We declared the part DNM. Scope showed the part ACKing 0xF4 then 0xA5 on a write. A read needed F4, A5, Sr, F5. Someone tried 0xF5 then 0xA5 as the first bytes and NACKed. NPI delay 9 days. LED on the 7-bit NACK path blinked "no device."
10-bit 0x2A5 missed by a 7-bit ping.
| Scan | Bytes on the wire | ACK |
|---|---|---|
| 7-bit 0x55 | one address byte | NACK |
| 10-bit write 0x2A5 | 0xF4 then 0xA5 | ACK ACK |
| 10-bit read 0x2A5 | 0xF4, 0xA5, Sr, 0xF5 | ACK ACK then read |
| NPI delay | 9 days | — |
I would not consider it settled without evidence: I would capture the first two bytes of a 10-bit write as 0xF4 then 0xA5 with ACK on both, plus a 7-bit scan that is documented as inconclusive.
A 7-bit scan is not a 10-bit search.
Curated: · Written: · Reviewed:
QA-37You issue 0x3B dual-output read with 8 dummy clocks. The flash wants 4. The payload is shifted. Who is right?(show answer)
I pin SPI dummy clocks for a dual-output flash read to a scope channel or a current shunt before I trust any LED or printf about it.
I treat dummy clocks as part of the command, in SCLK edges, not as a delay in microseconds. Wrong dummy count bit-shifts the entire payload.
Concretely, i take dummy cycles from the command table at that frequency and dummy-cycle register, I count SCLK edges on a scope, and I CRC a known page.
The reason for that specificity is a failure I have seen: Command 0x3B at 48 MHz wanted 4 dummy clocks. I ran 8, four extra. Dual I/O moves 2 bits per clock, so 4 extra clocks shifted 8 bits, not 4. 256-byte page CRC failed 100 percent. DMA TC LED still fired after 256+overhead bytes. I called the flash "bad silicon."
0x3B dual-output dummy mismatch at 48 MHz.
| Dummy clocks | Dual-I/O shift | 256-byte CRC |
|---|---|---|
| 4 (wanted) | 0 bits | pass |
| 8 (used), 4 extra | 4 clocks × 2 bits = 8 bits | 100% fail |
| DMA TC | fired | — |
I would not consider it settled without evidence: I would count SCLK edges between command and first MISO nibble and require 4, then CRC page 0 against a programmer dump.
Dummy clocks are edges, not a sleep.
Curated: · Written: · Reviewed:
QA-38A task takes a SPI mutex. The ISR also bit-bangs CS on that bus. You see torn 16-bit ADC words. Why?(show answer)
My first question on bus mutex shared by a task and an ISR is which register bit actually flipped, not which breakpoint I hit.
I treat a mutex as a task protocol. An ISR cannot take it. Sharing a bus with an ISR requires a FromISR lock or moving the ISR work to a deferred task.
Concretely, i do all SPI from one priority or I use a disable of that IRQ around the task transaction, and I never assert CS from two contexts.
The reason for that specificity is a failure I have seen: Task held mutex 42 µs for a 16-bit read at 8 MHz (2 µs clocks plus setup). ISR stole CS at 10 kHz for a 2-byte status. ADC word 0x03E8 became 0x03 then 0x11 from the status. 12 percent of samples. LED on mutex take still toggled.
ISR stealing CS during a 16-bit SPI read.
| Owner | Time | Result |
|---|---|---|
| task 16-bit at 8 MHz | 42 µs | 0x03E8 |
| ISR 2-byte at 10 kHz | steal | 0x0311 torn |
| torn sample rate | 12% | — |
I would not consider it settled without evidence: I would logic-analyze CS versus SCLK and require a single owner, plus zero mixed-length transactions in 10,000 samples.
A mutex does not bind an ISR.
Curated: · Written: · Reviewed:
QA-39The slave NACKs when busy. You retry immediately in a 2 µs loop. The bus never recovers. What backoff did you skip?(show answer)
I treat I2C NACK retry storm as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat a NACK as either address-not-present or slave-busy. Immediate retry on a busy slave starves the stretch/release it needed and can look like a 400 kHz flood.
Concretely, i cap retries, I wait the datasheet busy time (here 1.5 ms), and I run a recovery clock if SDA sticks, not a tight loop.
The reason for that specificity is a failure I have seen: Busy NACK for 1.5 ms conversion. I retried every 2 µs, 750 tries. SCL 400 kHz continuous. Slave never started the conversion. Success 0 percent. After a 5 ms backoff, success 100 percent in 1.5 ms + 80 µs. The NACK LED looked like "activity."
Immediate I2C retry versus a 1.5 ms conversion.
| Policy | Retries | Success |
|---|---|---|
| every 2 µs for 1.5 ms | 750 | 0% |
| wait 1.5 ms then one try | 1 | 100% in 1.58 ms |
I would not consider it settled without evidence: I would count retries per transaction on a GPIO and require ≤ 3, plus a 1.5 ms wait, then a 1,000-sample success rate ≥ 99.9 percent.
A NACK is not a request to spin.
Curated: · Written: · Reviewed:
QA-40You drain a 32-byte SPI FIFO in a 1 kHz task. The fill is 4 Mbps. How many bytes overflow per millisecond?(show answer)
On SPI FIFO versus DMA at 4 Mbps I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat FIFO depth against arrival rate. 4 Mbps is 500 kB/s = 500 bytes/ms. A 32-byte FIFO without DMA overflows in 64 µs.
Concretely, i enable DMA circular at least 2× the millisecond burst, or I raise a 50 µs ISR, and I count the overflow flag.
The reason for that specificity is a failure I have seen: 4 Mbps = 4e6/8 = 500,000 bytes/s. FIFO 32 bytes lasts 32/500000 = 64 µs. 1 kHz task period 1,000 µs. Overflow 1,000-64=936 µs × 500 bytes/ms = 468 bytes lost per ms. 93.6 percent loss. The FIFO-not-empty LED still blinked.
32-byte SPI FIFO versus 4 Mbps fill.
| Item | Value |
|---|---|
| bit rate | 4 Mbps |
| byte rate | 500,000 B/s |
| FIFO | 32 B = 64 µs |
| task period | 1,000 µs |
| lost per ms | 468 B (93.6%) |
I would not consider it settled without evidence: I would scope SCLK for 1 ms and compare received count to 500, requiring DMA or a ≤ 50 µs drain.
A 1 kHz drain is not a 4 Mbps FIFO.
Curated: · Written: · Reviewed:
QA-41A low-priority log task holds SPI. A high-priority control task blocks 8 ms. What RTOS feature did you not enable?(show answer)
I refuse to call priority inversion on a SPI mutex done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat a mutex as needing priority inheritance. A counting semaphore does not boost the holder. The high-priority task waits the full low-priority critical section plus any medium-priority CPU hog.
Concretely, i use a mutex with inheritance, I keep the hold to the transaction (42 µs), and I never print while holding the bus.
The reason for that specificity is a failure I have seen: Log task priority 1 held SPI for an 8 ms printf. Control priority 8 blocked. Medium GUI priority 4 ran 7.9 ms. Control jitter 8 ms on a 1 ms loop. Position error 3.2 mm. The SPI mutex-taken LED was high the whole time so it "worked."
8 ms printf while holding SPI.
| Task | Priority | Hold or wait |
|---|---|---|
| log | 1 | 8 ms printf on SPI |
| GUI | 4 | 7.9 ms CPU |
| control | 8 | blocked 8 ms |
| position error | — | 3.2 mm |
I would not consider it settled without evidence: I would trace task ready times and require control start ≤ 50 µs after it posts, with inheritance on and printf outside the lock.
A semaphore is not inheritance.
Curated: · Written: · Reviewed:
QA-42The loop does work then vTaskDelay(1). The phase wanders versus the tick. Which API did you want?(show answer)
The firmware question under vTaskDelay versus vTaskDelayUntil on a 1 kHz loop is which path still races after the path I just stepped through.
I treat vTaskDelay as relative to the current tick count. vTaskDelay(1) unblocks on the next tick, about 1 ms at 1 kHz from the tick boundary, not work-time plus 1 ms. vTaskDelayUntil is absolute from a stored wake tick so the loop keeps phase.
Concretely, i capture last_wake, I DelayUntil 1 tick at a 1 kHz tick, and I measure period and phase on a GPIO versus SysTick.
The reason for that specificity is a failure I have seen: Work was 400 µs. Delay(1) fired on the next tick (~1.0 ms from the previous tick, ~600 µs after the call), so period stayed near 1 ms but started 400 µs late relative to the plant's sample instant. Integrator of a 1 kHz plant saw a moving phase. Overshoot 22 percent. DelayUntil kept the 1.000 ms grid. A debugger halt made both look like 1 ms because halt stops the tick.
1 kHz loop with 400 µs work.
| API | Work | Period | Phase vs tick |
|---|---|---|---|
| vTaskDelay(1) | 400 µs | ~1.0 ms (next tick) | +400 µs slip |
| vTaskDelayUntil | 400 µs | 1.0 ms | locked |
| overshoot | — | — | 22% with Delay |
I would not consider it settled without evidence: I would scope the loop GPIO versus the tick for 10 s and require mean period 1.000 ms ± 50 µs and phase locked with DelayUntil, not Delay.
Delay waits until the next tick. DelayUntil keeps phase.
Curated: · Written: · Reviewed:
QA-43Two tasks "take" a counting semaphore around a UART. Bytes interleave. Why was a mutex the right object?(show answer)
I would budget a measurement for mutex versus counting semaphore for a UART the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat a counting semaphore as a counter, not as ownership. Two takes can succeed if the count was 2, and neither inherits priority.
Concretely, i use a mutex with count 1 semantics, I hold only for the frame, and I never give from an ISR unless it is a binary semaphore dedicated to deferred work.
The reason for that specificity is a failure I have seen: uxSemaphoreCreateCounting(2,2). Both log and CLI took. TX 0x41 0x42 from one and 0x31 0x32 from the other became 0x41 0x31 0x42 0x32. Parser error 40 percent. The take-success LED blinked twice so both tasks were "correct."
Counting semaphore(2,2) on a UART.
| Sender | Intended | On the wire |
|---|---|---|
| log | 41 42 | interleaved |
| CLI | 31 32 | 41 31 42 32 |
| parser errors | 40% | — |
I would not consider it settled without evidence: I would record TX bytes on an analyzer during dual senders and require atomic frames, plus a mutex owner dump.
A count of two is two owners.
Curated: · Written: · Reviewed:
QA-44configCHECK_FOR_STACK_OVERFLOW is 0. A task returns. Later the idle task HardFaults. What did you skip?(show answer)
Where I have been burned on stack overflow hook versus silent TCB smash is accepting a known race because "that is how MCUs work."
I treat overflow checking as a ship gate. Mode 2 canary is cheap. Without it a smash is a delayed fault in another task.
Concretely, i set overflow check to 2, I implement the hook as a breadcrumb plus reset, and I fail CI if the hook is a no-op.
The reason for that specificity is a failure I have seen: Overflow of 80 bytes into the next TCB. Idle crashed 3.4 s later. Hook was empty. Field units 12 per week. With mode 2 the canary of 16 bytes would have fired at the smash, 3.4 s of mis-control avoided. LED still ran on the overflowing task.
80-byte overflow with the hook compiled out.
| Item | Value |
|---|---|
| smash | 80 bytes into TCB |
| idle HardFault | 3.4 s later |
| field returns | 12 / week |
| canary | 16 bytes in mode 2 |
I would not consider it settled without evidence: I would force a 128-byte local on a 256-byte stack and require the hook to run before any other TCB is used.
The victim of a smash is not the smasher.
Curated: · Written: · Reviewed:
QA-45eExpectedIdleTime is 0xFFFFFFFF. The 32-bit tick is 0xFFFF0000. You wanted an 18-hour STOP. What wrap did you miss?(show answer)
I answer tickless idle wrap of the expected-wake tick by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat tickless sleep as a 32-bit subtraction that must use unsigned wrap, and I cap the sleep to the timer's 24-bit ARR if that is the hardware. 0xFFFFFFFF ticks at 1 kHz is 49.7 days, not 18 hours. 18 hours is 64,800,000 ticks.
Concretely, i compute wake = now + idle with uint32_t, I min() with the SysTick reload max and with 64,800,000 when the product wants 18 h, and I test near 0xFFFF0000.
The reason for that specificity is a failure I have seen: Tick 0xFFFFF000, idle 0x20000 ticks at 1 kHz = 131,072 ms. Sum wrapped to 0x0001F000. The compare was already in the past. Immediate wake. STOP never entered. Average current 6.2 mA versus 18 µA. A separate 18 h cap is 64,800,000 ticks, not 0xFFFFFFFF (4,294,967,295 ms ≈ 49.7 days). Battery 4 days instead of 14 months.
uint32 tickless compare already in the past.
| Item | Value |
|---|---|
| now | 0xFFFFF000 |
| idle | 0x20000 ticks (131.072 s) |
| sum wrap | 0x0001F000 |
| 0xFFFFFFFF at 1 kHz | 49.7 days |
| 18 h separately | 64,800,000 ticks |
| RUN current | 6.2 mA |
| STOP target | 18 µA |
I would not consider it settled without evidence: I would unit-test wrap and on-target set the tick high, then require a current trace of STOP for the capped ARR time, with 18 h requested as 64,800,000 ticks.
Tickless plus wrap is a past compare.
Curated: · Written: · Reviewed:
QA-46The ISR queues a sample and returns. The consumer task wakes 1 ms later. What flag did you not set?(show answer)
For xQueueSendFromISR without a yield I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat FromISR as needing pxHigherPriorityTaskWoken. Without it the task waits for the next tick, which at 1 kHz is up to 1 ms plus the remainder of the current tick.
Concretely, i pass the woken flag, I portYIELD_FROM_ISR, and I measure task GPIO versus ISR GPIO.
The reason for that specificity is a failure I have seen: ADC ISR at 20 kHz queued a sample. Consumer woke on the 1 kHz tick. 19 samples piled up. Queue length 8, so 11 overflowed per millisecond. Loss 11/20=55 percent. The queue-send-OK path in the ISR still toggled a LED.
20 kHz ADC queue without ISR yield.
| Item | Value |
|---|---|
| ADC ISR | 20 kHz |
| tick | 1 kHz |
| queue depth | 8 |
| overflow per ms | 11 |
| loss | 55% |
I would not consider it settled without evidence: I would scope ISR to task latency and require ≤ 15 µs at 80 MHz with yield, not 1 ms.
FromISR without yield is a tick later.
Curated: · Written: · Reviewed:
QA-47CAN FIFO is 3. You queue to a length-4 RTOS queue from ISR. You still drop. Where?(show answer)
I would write the check for queue overflow dropping a CAN frame before the driver, because the driver will otherwise certify itself.
I treat two queues in series as the min of their depths minus in-flight. A 3-frame hardware FIFO plus a 4-slot software queue still drops if the task is late by more than 7 frames.
Concretely, i size the software queue to burst × worst task latency, I count hw overflow and sw overflow separately, and I do not claim "queue 4 is enough" from the CAN FIFO number.
The reason for that specificity is a failure I have seen: Bus 1 Mbps, standard 8-byte frames stuffed about 120 µs (not 11 bytes × 8 = 88 µs). Peak 1/120 µs ≈ 8.3 kHz. Task latency 2.1 ms = 2.1e-3/120e-6 = 17.5 frames. HW FIFO 3 + SW 4 = 7. Dropped 10.5 frames per burst, 60 percent of a 17.5-frame burst. The SW send-from-ISR success LED fired for the 4 that fit.
1 Mbps CAN burst versus a 2.1 ms task stall.
| Item | Value |
|---|---|
| stuffed standard frame | ~120 µs at 1 Mbps |
| peak | ~8.3 kHz |
| stall | 2.1 ms = 17.5 frames |
| HW+SW depth | 3+4=7 |
| dropped | 10.5 (60% of burst) |
I would not consider it settled without evidence: I would log both overflow counters during a 2.1 ms stall and require zero, then resize to ≥ 24.
Hardware FIFO plus RTOS queue is still one bucket.
Curated: · Written: · Reviewed:
QA-48Idle calls WFI. DMA to SRAM stalls. Who gated the clock?(show answer)
The trade-off in idle-task sleep while a DMA needs SRAM clocks is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat STM32F407 Sleep as a chance for RCC_AHB1LPENR to gate SRAM1. WFI by itself does not turn off SRAM. If DMA runs in idle, SRAM1LPEN and DMA1LPEN must stay set, or I use a sleep mode the RM lists as keeping that master.
Concretely, i leave SRAM1LPEN and DMA1LPEN set during Sleep, I test a 1,024-byte DMA during idle, and I refuse a current number taken with DMA off.
The reason for that specificity is a failure I have seen: Idle entered Sleep with SRAM1LPEN cleared. WFI did not gate SRAM; the LPEN bit did. DMA TC never came. 1,024-byte SPI RX froze. Overrun after 64 µs of FIFO. I still measured 1.8 mA idle with DMA disabled and published that number. With DMA, 6.1 mA and a hung bus.
STM32F407 Sleep gating SRAM1 under a 1,024-byte DMA.
| Item | Value |
|---|---|
| DMA | 1,024 B at 4 Mbps |
| expected TC | 2.05 ms |
| Sleep, SRAM1LPEN=0 | TC never |
| advertised idle | 1.8 mA |
| actual with DMA | 6.1 mA hung |
I would not consider it settled without evidence: I would run DMA during Sleep, require TC in 1,024 / 500000 = 2.05 ms at 4 Mbps, and publish the current of that run with LPEN bits dumped.
Idle current with DMA off is a different product.
Curated: · Written: · Reviewed:
QA-49A 1 ms software timer runs a 400 µs SPI burst. The daemon is idle priority. The control task hogs 82 ms then delays 1 tick. What did you mis-prioritize?(show answer)
I pin software timer daemon priority to a scope channel or a current shunt before I trust any LED or printf about it.
I treat timer callbacks as running in the daemon task. If the daemon is below an always-ready control task, the timer only runs in the leftover delay. That leftover is what sets the observed rate.
Concretely, i keep callbacks short, I set daemon priority above background and below control, or I move the SPI to a dedicated task signaled by the timer.
The reason for that specificity is a failure I have seen: Daemon priority 0. Control priority 5 ran 82 ms of compute then vTaskDelay(1) at 1 kHz, a 83 ms cycle. Daemon ran in that 1 ms window. One 400 µs SPI callback per 83 ms is 1/0.083 ≈ 12 Hz. SPI burst 12×400 µs = 4.8 ms/s instead of 1 kHz × 400 µs = 400 ms/s. Sensor 12 Hz versus 1 kHz. The timer-create LED still went on at init.
1 ms software timer under an 82 ms control hog.
| Item | Value |
|---|---|
| intended callback | 1 kHz, 400 µs SPI |
| daemon priority | 0 |
| control | 82 ms work + 1 ms delay = 83 ms |
| achieved callback | 1/0.083 ≈ 12 Hz |
| SPI duty | 4.8 ms/s vs 400 ms/s |
I would not consider it settled without evidence: I would GPIO the callback and require 1,000 Hz ± 1 percent, then show the daemon priority and the 82 ms hog in the report.
A timer is only as real-time as its daemon.
Curated: · Written: · Reviewed:
QA-50Two tasks at priority 3. One hogs 3 ms of CPU. The other needs 200 µs every 1 ms. Time slicing is off. Who never runs?(show answer)
My first question on equal-priority round-robin time slice is which register bit actually flipped, not which breakpoint I hit.
I treat equal priority without slicing as run-to-block. A CPU hog at the same priority starves a peer until it delays or waits.
Concretely, i enable slicing at 1 tick, I split priorities, or I insert a wait in the hog, and I measure both GPIOs.
The reason for that specificity is a failure I have seen: Slicing off. Hog ran 3 ms compute without a block. Peer needed 200 µs each 1 ms and never ran. UART RX queue of 8 overflowed at 115,200 (86.8 µs × 8 = 694 µs) because the peer never drained. LED on the hog still 100 Hz of "progress."
Same-priority hog versus a 200 µs peer.
| Task | Need | Observed |
|---|---|---|
| hog | 3 ms compute, no block | runs forever |
| peer | 200 µs / 1 ms | never |
| UART queue | 8 × 86.8 µs = 694 µs | overflow |
I would not consider it settled without evidence: I would capture both task GPIOs and require the peer's 200 µs inside every 1 ms window, which means a slice, a higher peer priority, or a block in the hog.
Equal priority without a slice is a single task.
Curated: · Written: · Reviewed:
QA-51The bootloader CRCs 384 KB and jumps. An attacker swaps a 384 KB image with the same CRC. What check was missing?(show answer)
I treat image CRC versus a signature as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat CRC as integrity against random flash errors, not as authenticity. A signature over the hash is what binds the image to a key I hold in OTP.
Concretely, i SHA-256 the image, I verify a 256-bit ECDSA P-256 signature with a public key in OTP, and I still CRC to catch bit rot before I spend the verify time.
The reason for that specificity is a failure I have seen: CRC-32 of 393,216 bytes passed in 18 ms at 48 MHz. A swapped image with a patched CRC in the header also passed. We shipped 4,000 units. Signing is 42 ms extra. The green LED after CRC still meant boot.
384 KiB image, CRC-only boot.
| Check | Time at 48 MHz | Stops a swapped image |
|---|---|---|
| CRC-32 | 18 ms | no |
| SHA-256 + ECDSA P-256 | 42 ms | yes |
| total budget | 80 ms | — |
| units shipped | 4,000 | CRC LED green |
I would not consider it settled without evidence: I would load a CRC-valid but unsigned image and require a refuse, plus a signed image that boots in at most 80 ms including CRC.
CRC finds rot. It does not find a foe.
Curated: · Written: · Reviewed:
QA-52You copy slot B then reset. Slot B is now active without a confirm. A boot loop bricks both. What state did you skip?(show answer)
On A/B slot confirm after self-test I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat an unconfirmed image as a trial boot. Confirm is a write after self-test, not after flash program-complete.
Concretely, i boot the new slot in trial, I run a 2-second self-test, I write CONFIRMED, and I revert to A if I reset before that write.
The reason for that specificity is a failure I have seen: I marked B active at flash TC. B HardFaulted in 120 ms. Bootloader saw B active, retried B 8 times in 8 times 120 ms equals 960 ms, then still B. A was intact. Field brick rate 2.4 percent of a 500-unit wave (12 units). The program-complete LED was green.
Unconfirmed slot B in a boot loop.
| Event | Time |
|---|---|
| B HardFault | 120 ms |
| retries | 8 times 120 ms = 960 ms |
| A used | never |
| wave | 500 units, 12 bricked (2.4%) |
I would not consider it settled without evidence: I would flash a trial image that resets at 100 ms and require A to run after the retry budget, with B still unconfirmed.
Program complete is not confirmed.
Curated: · Written: · Reviewed:
QA-53Slot B is signed version 6. OTP monotonic is 8. Do you boot B because the signature is valid?(show answer)
I refuse to call anti-rollback version monotonicity done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat anti-rollback as a second gate after signature. A valid old image is still an attack if it has a known hole.
Concretely, i compare version to OTP monotonic, I refuse older, I permit equal, I increment OTP only after confirm of a newer image, and I still require ECDSA.
The reason for that specificity is a failure I have seen: I booted signed v6 because ECDSA passed. v6 had the UART bootrom hole we patched in v8. An attacker replayed v6. 900 devices. OTP was 8. Verify time 42 ms well spent on v8, wasted on v6. The signature-OK LED lied.
Signed v6 versus OTP monotonic 8.
| Image | ECDSA | OTP | Boot |
|---|---|---|---|
| v6 replay | pass | 8 | refuse, older |
| v8 current | pass | 8 | permit equal |
| v9 trial | pass | 8 until confirm | then OTP 9 |
| devices replayed | 900 | — | hole in v6 |
I would not consider it settled without evidence: I would present v6 signed against monotonic 8 and require a refuse, present v8 equal to OTP 8 and require boot without bumping OTP, plus v9 that increments OTP only after confirm.
A signature on yesterday is not a license to roll back.
Curated: · Written: · Reviewed:
QA-54Erase of slot B is 40 kB in 220 ms. Mains dies at 80 ms. What must still boot?(show answer)
The firmware question under power-loss during a flash erase is which path still races after the path I just stepped through.
I treat erase as destroying the slot. Power loss mid-erase leaves B unreadable. A stays the active slot until B is fully programmed, authenticated, and then marked trial.
Concretely, i erase B, program B, CRC and verify signature, then mark B trial. A remains active through that sequence. I never mark B active before program plus CRC, and I test a supply drop at 80 ms of erase.
The reason for that specificity is a failure I have seen: I set B active, then erased. Drop at 80 ms of 220 ms. B unreadable, active equals B. Brick. A still had the old image. 7 of 40 lab drops. LED on erase-started was our only breadcrumb.
220 ms erase, 80 ms power loss.
| Item | Value |
|---|---|
| erase | 40,960 bytes, 220 ms |
| drop | 80 ms |
| B | unreadable |
| lab bricks if active=B first | 7 / 40 |
| required | A boots 40 / 40 |
I would not consider it settled without evidence: I would brownout the 3.3 V rail at 80 ms into erase on 40 tries and require A to boot 40 of 40.
Active must point at a slot that still exists.
Curated: · Written: · Reviewed:
QA-55A and B both CRC-fail. The product has no SWD in the field. What boots?(show answer)
I would budget a measurement for recovery image when both A and B fail CRC the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat a recovery slot as a third image, write-protected, that only talks a signed updater. Without it a dual CRC fail is a brick.
Concretely, i lock recovery in OTP or RDP, I keep it under 32 KB, I test a double-fail fixture, and I never erase recovery from the application.
The reason for that specificity is a failure I have seen: A and B both failed after a botched dual write. No recovery. 3,200 units needed JTAG. Recovery would have been 28 KB, 9 ms CRC. UART at 115,200 8N1 is about 11,520 B/s; a 384 KiB = 393,216-byte image is 393216/11520 = 34.13 s, not 33 s. The dual-fail LED did not exist.
Dual CRC fail without a recovery slot.
| Slot | Size | CRC |
|---|---|---|
| A | 384 KB | fail |
| B | 384 KB | fail |
| recovery (missing) | 28 KB | 9 ms |
| UART 384 KiB (393,216 B) at 11,520 B/s | 34.13 s | — |
| field units | 3,200 | JTAG |
I would not consider it settled without evidence: I would corrupt A and B CRCs on a fixture and require recovery to enumerate a signed update within 2 s of reset.
Two slots can still be zero slots.
Curated: · Written: · Reviewed:
QA-56You rotate the public key in OTP. Slot A was signed with key 0. Do you still boot A?(show answer)
Where I have been burned on key rotation leaving old slots signed is accepting a known race because "that is how MCUs work."
I treat a rotated key as invalidating old signatures unless I keep a window of N keys and resign. Booting A with key 0 after OTP holds only key 1 is a brick.
Concretely, i resign both slots before burning the new OTP key, I keep a 2-key window during rollout, then I disable key 0.
The reason for that specificity is a failure I have seen: OTP switched to key 1. A still signed with key 0. Next reset refused A. B was empty. 180 units in a warehouse needed a jig. Resign is 42 ms times 2 slots. The OTP-burn LED was the last thing that worked.
OTP key 1 versus slot A still on key 0.
| Item | Value |
|---|---|
| resign per slot | 42 ms |
| both slots | 84 ms |
| warehouse bricks | 180 |
| window | 2 keys during rollout |
I would not consider it settled without evidence: I would run a rotation drill: resign, verify both slots with key 1, then burn, on 10 boards 10 of 10.
A new key does not resign flash by itself.
Curated: · Written: · Reviewed:
QA-57The trailer is 16 bytes: magic, length, CRC. A reset hits after 8 bytes. The bootloader reads length 0xFFFF. What protocol did you skip?(show answer)
I answer torn trailer in slot metadata by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat metadata as a two-copy or copy-on-confirm structure. A torn 16-byte write is a race I must detect, not how flash works.
Concretely, i write trailer B, CRC it, then flip an atomic bank-select word that is a single 32-bit program, and I reject magic mismatch.
The reason for that specificity is a failure I have seen: Reset at 8 bytes. Length 0xFFFF = 65,535. Bootloader copied 65,535 bytes, which is in-range for a 384 KiB slot and is not a wrap of 393,216. It was an out-of-range metadata length: 180 ms of garbage used as a jump target. Brick. Two-copy trailer is 32 bytes, 2 ms extra. The first-8-bytes LED in the programmer was green.
16-byte trailer torn at 8 bytes.
| Item | Value |
|---|---|
| trailer | 16 bytes |
| torn after | 8 bytes |
| length read | 0xFFFF |
| copy attempt | 65,535 bytes |
| two-copy extra | 32 bytes, 2 ms |
I would not consider it settled without evidence: I would reset during trailer program 50 times and require either old trailer or new, never mixed length 0xFFFF.
A torn trailer is a known race, not a platform tax.
Curated: · Written: · Reviewed:
QA-58Factory sets CONFIRMED in the same script as PROGRAM. First power in the truck HardFaults. Why can you not revert?(show answer)
For confirmed bit on an image that never ran I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat CONFIRMED as a statement that this silicon ran this image. Factory confirm without a run is a lie that disables anti-brick.
Concretely, i confirm only after a power cycle and a 2 s self-test on the unit, not in the programmer, and I keep trial boots for first factory power.
The reason for that specificity is a failure I have seen: 10,000 units confirmed in the fixture with NRST held. Image crashed on a GPIO that the fixture pulled. No revert. 10,000 trucks. Self-test is 2 s. Fixture time was 1.1 s program plus 0 s run. The confirm LED on the jig was the metric.
Factory CONFIRMED with NRST held.
| Step | Time |
|---|---|
| program | 1.1 s |
| run on unit | 0 s |
| self-test wanted | 2 s |
| units | 10,000 |
| revert | impossible |
I would not consider it settled without evidence: I would require a unit-side confirm token after 2 s of RUN current within 2 mA of golden, then set the bit.
Confirmed means it ran, not that it programmed.
Curated: · Written: · Reviewed:
QA-59On-the-fly AES-XIP hits a 12 µs QSPI stall. The core fetches a 6-cycle instruction. What do you budget?(show answer)
I would write the check for encrypted XIP with a stalled QSPI before the driver, because the driver will otherwise certify itself.
I treat encrypted XIP as adding decrypt latency plus flash stalls to every miss. A 6-cycle Thumb fetch is not 6 cycles when the line comes from QSPI.
Concretely, i measure miss latency with DWT CYCCNT, I place ISRs in SRAM, and I refuse a WCET that used SRAM-only numbers.
The reason for that specificity is a failure I have seen: QSPI stall 12 µs at 80 MHz is 960 cycles. ISR in XIP took 12 µs extra per miss. Three misses: 36 µs. PWM period 50 µs. Missed update. Current 2.8 A. SRAM placement is 0 extra. The XIP-enable LED was on so encryption looked done.
12 µs QSPI stall inside a 50 µs PWM ISR.
| Item | Value |
|---|---|
| stall | 12 µs = 960 cycles at 80 MHz |
| misses in ISR | 3 then 36 µs |
| PWM period | 50 µs |
| phase current | 2.8 A |
| SRAM ISR extra | 0 µs |
I would not consider it settled without evidence: I would CYCCNT an ISR in XIP versus SRAM and require the PWM ISR in SRAM with miss latency published.
Encrypted XIP is a cache miss tax.
Curated: · Written: · Reviewed:
QA-60A TLV length field is 0xFFF0. The header is 256 bytes. Do you memcpy?(show answer)
The trade-off in TLV bounds in a boot header parser is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat TLV length as hostile. I check remaining bytes before every copy, and I cap the total walk at the header size.
Concretely, i parse with remaining minus size, I refuse length greater than remaining, and I fuzz the header with 0xFFFF lengths.
The reason for that specificity is a failure I have seen: Length 0xFFF0, remaining 20. memcpy 65,520 bytes from flash into a 256-byte stack. Overflow 65,264 bytes. Bootloader MSP smashed. Brick. A 2-byte remaining check is 1 cycle. The parse-OK LED fired after the first tag.
0xFFF0 TLV length in a 256-byte header.
| Item | Bytes |
|---|---|
| header | 256 |
| remaining | 20 |
| claimed length | 65,520 |
| stack smash | 65,264 |
| bound check | 1 cycle |
I would not consider it settled without evidence: I would fuzz 10,000 headers including 0xFFFF lengths and require zero stack movement beyond the 256-byte header buffer.
TLV length is not a memcpy length.
Curated: · Written: · Reviewed:
QA-61RUN is 12 mA. STOP is 8 µA. A wake costs 40 µJ at 3.3 V. How long must you stay in STOP to win?(show answer)
I pin sleep break-even versus wake energy to a scope channel or a current shunt before I trust any LED or printf about it.
I treat break-even as wake energy divided by the current delta times voltage. Shorter sleeps cost more than staying in RUN.
Concretely, i measure run current, stop current, and wake energy on a shunt, I compute break-even time, and I refuse STOP for periods below that.
The reason for that specificity is a failure I have seen: Break-even is 40e-6 / ((0.012 - 0.000008) times 3.3) = 40e-6 / 0.0395736 = 1.011 ms. I slept 200 µs once per 1 ms of a 1 kHz loop: one wake, not five. That wake is 40 µJ versus 12 mA × 3.3 V × 1 ms = 39.6 µJ to stay in RUN. Battery 8 days versus 40. The STOP-entry LED was busy.
Break-even at 3.3 V, 40 µJ wake.
| Item | Value |
|---|---|
| I_run | 12 mA |
| I_stop | 8 µA |
| E_wake | 40 µJ |
| t_be | 1.011 ms |
| 200 µs STOP once per 1 ms | 1 wake, 40 µJ vs 39.6 µJ RUN |
I would not consider it settled without evidence: I would current-trace a 1 kHz loop in RUN versus one 200 µs STOP per millisecond and require the lower integral, here RUN.
A 200 µs STOP can cost more than RUN.
Curated: · Written: · Reviewed:
QA-62You enter STOP then a 4 µs pulse arrives on EXTI. You never wake. The pulse is legal. What race did you accept?(show answer)
My first question on lost wakeup from a GPIO in STOP is which register bit actually flipped, not which breakpoint I hit.
I treat STOP entry and EXTI as a protocol. An EXTI already pending before WFI wakes immediately; the pulse is not lost. Loss is a mask/clear race (pending cleared after the edge, then WFI) or a pin that is not in the STOP wakeup mux.
Concretely, i sample the wake pin after configuring EXTI, I do not clear PR after the last edge without re-reading, I confirm the pin is a STOP wakeup source in the RM, then WFI.
The reason for that specificity is a failure I have seen: Pulse 4 µs set EXTI PR. The STOP path cleared PR, then WFI. STOP held. Lost wakeup 6 of 1,000 lab pulses (0.6 percent). A second board used PE2, which is not an STM32F407 STOP wakeup pin; pending did nothing in STOP. Field: 18-hour freeze. Current 8 µA looking like success. A debugger halt prevented STOP so the race vanished on the bench with SWD.
4 µs EXTI pulse overlapping STOP entry.
| Item | Value |
|---|---|
| pulse | 4 µs, EXTI pending |
| pending then WFI | wakes immediately |
| clear PR then WFI | lost, 6 / 1,000 (0.6%) |
| PE2 in STOP | not a wakeup pin |
| STOP current | 8 µA looks healthy |
I would not consider it settled without evidence: I would inject 4 µs pulses on a generator during STOP entry 10,000 times and require 10,000 wakes, plus a re-read before WFI and a wakeup-capable pin.
Pending-before-WFI wakes. A clear race or a non-wakeup pin does not.
Curated: · Written: · Reviewed:
QA-63IWDG timeout is 8 s. You STOP for 30 s on RTC. You reset at 8 s. Which bit did you not set?(show answer)
I treat IWDG still counting in STOP as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat IWDG as optionally running in STOP. If I need 30 s STOP, I either refresh from RTC wake every less than 8 s or I set the bit that stops IWDG in STOP if the product allows.
Concretely, i read the reference manual STOP and IWDG table, I pick RTC-wake refresh at 4 s, and I current-trace to prove IWDG is not resetting me.
The reason for that specificity is a failure I have seen: STOP 30 s, IWDG 8 s still clocked from LSI 32 kHz. Reset at 8 s. Duty became 8 s STOP plus 200 ms RUN. Average current (8e-6 times 8 plus 0.012 times 0.2) / 8.2 is about 301 µA versus 8 µA. On 250 mAh that is 0.250/0.000301 ≈ 830 h ≈ 34.6 days, not 11 months. Target 8 µA is 0.250/0.000008 = 31,250 h ≈ 3.6 years. The STOP LED fired, then reset.
8 s IWDG during a 30 s STOP.
| Item | Value |
|---|---|
| intended STOP | 30 s |
| IWDG | 8 s, LSI 32 kHz |
| average current | ~301 µA |
| 250 mAh at 301 µA | 830 h ≈ 34.6 days |
| target 8 µA | 31,250 h ≈ 3.6 years |
| refresh plan | RTC every 4 s |
I would not consider it settled without evidence: I would log reset cause for 10 STOP cycles of 30 s and require IWDG reset cause clear, plus average current at most 15 µA.
STOP is not a paused watchdog unless you said so.
Curated: · Written: · Reviewed:
QA-64STOP is specified 8 µA. You measure 1.01 mA. 20 pins are floating inputs. What do you do before blaming the regulator?(show answer)
On GPIO leakage in STOP I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat floating CMOS inputs as shoot-through. 20 pins times 50 µA is 1 mA, which swallows the STOP spec.
Concretely, i analog-config or pull every unused pin, I measure current pin-group by pin-group, and I refuse a STOP number with SWD connected.
The reason for that specificity is a failure I have seen: 20 floating pins times 50 µA = 1,000 µA, plus 8 µA STOP = 1,008 µA. Spec 8 µA is 126 times off. Battery 250 mAh / 1.008 mA = 248 h = 10.3 days, not 14. At 9 µA after pulls, 250/0.009 ≈ 27,778 h ≈ 3.2 years. The regulator was innocent. The green LED on a pushed-pull pin was our only configured pad.
20 floating GPIOs versus 8 µA STOP.
| Item | Current |
|---|---|
| STOP spec | 8 µA |
| 20 times 50 µA leak | 1,000 µA |
| measured | 1,008 µA |
| 250 mAh life | 248 h = 10.3 days |
| after analog config | 9 µA |
I would not consider it settled without evidence: I would current-trace with all unused pins analog, SWD off, and require at most 15 µA at 3.3 V 25 C.
Floating inputs are milliamps, not STOP.
Curated: · Written: · Reviewed:
QA-65You expect a 1 kHz SysTick in STOP. You wake every 30 s. What clock did you actually keep?(show answer)
I refuse to call RTC wake versus SysTick in STOP done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat SysTick as HCLK-derived. STOP gates HCLK. A 1 kHz RTOS tick does not run. RTC or LPTIM is the wake source.
Concretely, i use RTC wake for long STOP, I reconstruct tick on wake, and I never call vTaskDelay during STOP expecting 1 ms granularity.
The reason for that specificity is a failure I have seen: I WFI-STOP expecting SysTick 1 ms. HCLK off. Woke on IWDG 8 s. Delay of 10 ticks in a task before STOP was a lie: 10 ms became 8 s. User-visible 8 s button latency. Current 8 µA, so the sleep looked better than spec. Debugger kept HCLK so the 1 kHz tick appeared to work.
SysTick assumed during STOP.
| Clock | In STOP | Wake |
|---|---|---|
| SysTick / HCLK | gated | none |
| IWDG LSI | on | 8 s |
| RTC 32,768 Hz | on | 30.5 µs tick |
| SWD attached | HCLK on | false 1 kHz |
I would not consider it settled without evidence: I would detach SWD, STOP, and require RTC EXTI as the only wake, with tick reconstruction error at most 1 RTC tick (30.5 µs at 32,768 Hz).
SysTick dies in STOP. RTC does not.
Curated: · Written: · Reviewed:
QA-66Vdd sags to 2.4 V for 80 µs while you program a 32-bit word. The part's PVD is 2.7 V but disabled. What is in flash?(show answer)
The firmware question under brownout during a flash write is which path still races after the path I just stepped through.
I treat a flash write below the programming spec as a possible torn word. PVD must reset or NMI before the write, not after.
Concretely, i enable PVD at 2.7 V, I hold writes if Vdd is under 2.8 V on an ADC, and I ECC-check the word after program.
The reason for that specificity is a failure I have seen: Sag 2.4 V, 80 µs, PVD off. Programmed 0x0000FFFF instead of 0x00000000. Trailer magic broken. Next boot CRC fail. 1.8 percent of a 2,000-unit drop test (36 units). Write time 40 µs, sag overlapped. The program-IRQ LED still fired.
80 µs sag to 2.4 V during a 32-bit program.
| Item | Value |
|---|---|
| PVD threshold | 2.7 V (disabled) |
| sag | 2.4 V, 80 µs |
| program | 40 µs |
| torn word | 0x0000FFFF vs 0x00000000 |
| drop-test fails | 36 / 2,000 (1.8%) |
I would not consider it settled without evidence: I would sag 2.4 V on a MOSFET during program 200 times with PVD on and require 200 resets before write, zero torn magics.
A program IRQ is not a verified word.
Curated: · Written: · Reviewed:
QA-67Internal LDO is in high-power mode. STOP current is 180 µA not 8 µA. Which PWR bit is still RUN?(show answer)
I would budget a measurement for regulator mode in RUN versus STOP the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat regulator scale and LDO mode as part of the sleep entry. Leaving LPDS clear keeps the LDO in RUN in STOP.
Concretely, i set LPDS per the manual, I measure current, and I do not copy a RUN init into the STOP path.
The reason for that specificity is a failure I have seen: LPDS was 0. STOP 180 µA versus 8 µA, 22.5 times. 250 mAh / 180 µA = 1,389 h about 58 days versus 3.5 years at 8 µA. The STOP WFI LED toggled. A debugger 180 µA looked about right because SWD adds about 100 µA.
STOP with LDO left in high-power.
| Mode | Current | 250 mAh life |
|---|---|---|
| LPDS=0 | 180 µA | ~58 days |
| LPDS=1 | 8 µA | ~3.5 years |
| SWD extra | ~100 µA | hides the bit |
I would not consider it settled without evidence: I would read PWR_CR after STOP entry via a breadcrumb and require LPDS set, current at most 15 µA with SWD off.
WFI is not STOP if the LDO is still in RUN.
Curated: · Written: · Reviewed:
QA-68You retain 128 KB in STOP for 8 µA extra versus 2 KB. Do you need the 128 KB?(show answer)
Where I have been burned on RAM bank retention cost is accepting a known race because "that is how MCUs work."
I treat retained RAM as a current menu. Each bank has a µA price. Keeping BSS retained because it is convenient is a battery decision.
Concretely, i place overnight state in a 2 KB retained section, I power-down the rest, and I measure the delta.
The reason for that specificity is a failure I have seen: 128 KB retain added 8 µA (datasheet 1 µA per 16 KB). I needed 48 bytes of RTC calendar. Extra 8 µA times 3.3 V = 26.4 µW. Over a year 232 mWh, about 70 mAh at 3.3 V. On 250 mAh that is 28 percent of the budget. 2 KB retain is 0.125 µA. The retain-all LED in the HAL example was on.
128 KB versus 2 KB retain in STOP.
| Retain | Extra current | Year at 3.3 V |
|---|---|---|
| 128 KB | 8 µA | ~70 mAh |
| 2 KB | 0.125 µA | ~1.1 mAh |
| state needed | 48 bytes | — |
| pack | 250 mAh | 28% wasted |
I would not consider it settled without evidence: I would current-trace 2 KB versus 128 KB retain and require the smaller bank if state is at most 2 KB.
Retained RAM is a bill, not a default.
Curated: · Written: · Reviewed:
QA-69A button bounces 3.2 ms. You wake STOP on EXTI and run 12 mA for 20 ms of debounce in software. What hardware filter did you skip?(show answer)
I answer wake filter on a bouncing pin by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat bounce as a train of wakes. A 3.2 ms bounce of 12 edges is 12 STOP exits if I do not ignore further edges. A 32 kHz 4-cycle EXTI filter is 122 µs, which is shorter than 3.2 ms, so it still passes the bounce.
Concretely, i use a 10 ms RTC or software lockout after the first wake, I do not RUN-debounce 20 ms at 12 mA for each of 12 edges, and I do not call a 122 µs filter a 3.2 ms debounce.
The reason for that specificity is a failure I have seen: 12 edges times 20 ms RUN = 240 ms at 12 mA per press. 1,000 presses/day = 240 s times 12 mA = 0.80 mAh/day. Hardware 4-cycle filter at 32 kHz is 122 µs and still saw all 12 edges. A 10 ms lockout is one STOP exit per press. The first-edge LED counted 12 per press and we called it alive.
3.2 ms bounce waking STOP 12 times.
| Item | Value |
|---|---|
| bounce | 3.2 ms, 12 edges |
| RUN debounce | 20 ms times 12 = 240 ms at 12 mA |
| 1,000 presses/day | 0.80 mAh/day |
| 32 kHz 4-cycle filter | 122 µs, still 12 wakes |
| 10 ms lockout | 1 wake per press |
I would not consider it settled without evidence: I would scope the button and the RUN current for one press and require a single STOP exit with a 10 ms lockout, not 12.
Bounce is extra wakes, not extra user input.
Curated: · Written: · Reviewed:
QA-70You STOP the core. ADC DMA to SRAM should fill 256 samples. The buffer stays 0xA5. Who stopped the DMA clock?(show answer)
For DMA running while the core is in STOP I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat STOP as gating the CPU and possibly AHB masters. ADC plus DMA in STOP needs the mode that keeps ADC and DMA clocks, or I use STOP0 not STOP2.
Concretely, i pick the sleep mode that lists ADC and DMA, I start DMA, then STOP, I wake on DMA TC, and I current-trace that path.
The reason for that specificity is a failure I have seen: STOP2 gated DMA. 256 samples times 2 bytes = 512 bytes stayed 0xA5. We averaged zeros. Control offset 0.4 V. STOP2 current 2 µA looked excellent. STOP0 with DMA is 180 µA for 256 / 200e3 = 1.28 ms at 200 kS/s, energy 180e-6 times 3.3 times 1.28e-3 = 0.76 µJ.
256-sample ADC DMA in STOP2 versus STOP0.
| Mode | DMA | Current | Fill time at 200 kS/s |
|---|---|---|---|
| STOP2 | gated | 2 µA | never, 0xA5 |
| STOP0 | runs | 180 µA | 1.28 ms, 0.76 µJ |
| samples | 256 times 2 B | 512 B | — |
I would not consider it settled without evidence: I would require a buffer that is not 0xA5 after STOP plus TC, and publish 180 µA not 2 µA for that mode.
The core sleeping is not the ADC sleeping.
Curated: · Written: · Reviewed:
QA-71RDP is 0. SWD pins are the default AF. A competitor dumps 384 KB in 12 s. What did factory skip?(show answer)
I would write the check for SWD left enabled in production before the driver, because the driver will otherwise certify itself.
I treat SWD as a debug aperture. Production sets RDP and remuxes the pins unless a licensed fixture needs them.
Concretely, i set RDP level 1 after the last fixture test, I analog the SWD pins, and I sample 10 units for dump-fail.
The reason for that specificity is a failure I have seen: RDP 0, SWD on. 384 KB at 32 kB/s dump about 12 s. 50,000 units. Fixture time saved 200 ms by skipping RDP. The SWD-OK LED on the programmer was the factory KPI.
Open SWD on a 384 KB image.
| Item | Value |
|---|---|
| image | 384 KB |
| dump rate | 32 kB/s |
| dump time | 12 s |
| units | 50,000 |
| RDP skip saved | 200 ms |
I would not consider it settled without evidence: I would attempt an openocd dump on 10 sealed units and require 10 fails, plus pin analog leakage at most 50 nA.
A production SWD port is a download cable.
Curated: · Written: · Reviewed:
QA-72The race vanishes when you halt. You conclude there is no race. What trace did you skip?(show answer)
The trade-off in halt-mode debug versus ETM of a race is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat halt as freezing timers, DMA, and the other core. A race that needs time to exist will not exist under halt. ETM or GPIO trace keeps time.
Concretely, i use ETM or 4 GPIOs at 10 ns, I never single-step the mailbox, and I reproduce at full 80 MHz.
The reason for that specificity is a failure I have seen: Mailbox torn 2.1 percent at 1 kHz. Halt-step showed flag then payload always ordered. We shipped. 2.1 percent times 600,000 messages per 10 min = 12,600 errors. ETM would have shown the store buffer. The halt LED in the IDE was our evidence.
Mailbox race that halt cannot see.
| Method | Tear rate |
|---|---|
| halt step | 0% |
| full-speed 1 kHz | 2.1% |
| 10 min | 12,600 tears |
| ETM | shows store buffer |
I would not consider it settled without evidence: I would capture ETM around the flag store and require payload visibility before the flag, 0 tears in 600,000.
Halt is not the product. Trace is.
Curated: · Written: · Reviewed:
QA-73GDB says the PC is in HAL_Delay. The part is looping in your HardFault. What file did you load?(show answer)
I pin mismatched ELF versus the flashed image to a scope channel or a current shunt before I trust any LED or printf about it.
I treat the ELF as matching the CRC of flash. A stale ELF maps the PC into the wrong function and you fix a delay that is not running.
Concretely, i CRC flash against the ELF load segments before I debug, I embed a build-id in both, and I refuse a session on mismatch.
The reason for that specificity is a failure I have seen: ELF was yesterday's 0x08004A10 HAL_Delay. Flash was today's HardFault at 0x08004A10. We increased the delay 3 times. Fault was a null at that address. 6 hours. Build-id compare is 16 bytes. The GDB-connected LED was green.
Stale ELF, same address, different function.
| Item | Address | Symbol |
|---|---|---|
| yesterday ELF | 0x08004A10 | HAL_Delay |
| today's flash | 0x08004A10 | HardFault |
| debug time | 6 hours | — |
| build-id | 16 bytes | — |
I would not consider it settled without evidence: I would require matching build-id in flash and ELF, then a PC that lands in the HardFault file and line.
GDB without a CRC is fiction.
Curated: · Written: · Reviewed:
QA-74SEGGER_RTT_Write blocks for 40 ms. There is no debugger. Your 50 µs control loop is dead. What config did you ship?(show answer)
My first question on RTT blocking when the host is absent is which register bit actually flipped, not which breakpoint I hit.
I treat RTT as a host-present buffer. Blocking mode waits for a debugger that production does not have.
Concretely, i set MODE_NO_BLOCK_SKIP, I compile RTT out of production, and I never log from the 50 µs path.
The reason for that specificity is a failure I have seen: BLOCK_IF_FIFO_FULL, 1 KB buffer, no host. Write of 80 bytes stalled 40 ms waiting. PWM 20 kHz missed 800 updates. Current 4.2 A. With SKIP, 0 wait, 80 bytes dropped. The RTT-connected LED is only on the host.
Blocking RTT in a 50 µs loop, no host.
| Mode | Wait | PWM misses at 20 kHz |
|---|---|---|
| BLOCK_IF_FIFO_FULL | 40 ms | 800 |
| NO_BLOCK_SKIP | 0 | 0, 80 B dropped |
| loop | 50 µs | — |
| current | 4.2 A | — |
I would not consider it settled without evidence: I would run with SWD unplugged and require the 50 µs GPIO period to hold, plus production builds without RTT.
RTT wait is a debugger, not a log.
Curated: · Written: · Reviewed:
QA-75You set a data watchpoint on the DMA RX byte. The overrun disappears. Why?(show answer)
I treat watchpoint Heisenbug on a DMA buffer as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat a DWT data watchpoint as a comparator that can halt on match. Halt freezes timers and DMA and hides an overrun. DWT is not an extra AHB master that adds 12 cycles per access.
Concretely, i debug DMA with GPIOs and a logic analyzer, I do not halt-on-match the buffer, and I reproduce without the probe when I can.
The reason for that specificity is a failure I have seen: Watchpoint halt-on-match stopped the core on the first RX byte. DMA versus CPU timing changed; the 32-byte FIFO no longer overran in the lab. Without the watchpoint, overrun 40 per second. We shipped the so-called fixed build. Field overrun 40 per second. The watchpoint-hit LED was the lab's success.
Watchpoint halt hiding a 32-byte FIFO overrun.
| Setup | Timing | Overrun |
|---|---|---|
| halt-on-match | core stopped | 0 /s |
| GPIO only | full speed | 40 /s |
| DWT comparator | no extra 12 AHB cycles | — |
I would not consider it settled without evidence: I would reproduce overrun with watchpoints off, GPIO-only, 10-second count matching 40 per second plus or minus 10 percent.
Halt-on-match is the Heisenbug. DWT is not a 12-cycle master.
Curated: · Written: · Reviewed:
QA-76You set RDP 1 in the first factory script. Stations 2 to 4 still need SWD calibration. What did you skip?(show answer)
On RDP level that bricks the fixture I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat STM32 RDP 1 as blocking debug. After RDP 1 the fixture cannot attach SWD to write calibration without going to RDP 0, which erases flash. RDP 2 is a later, irreversible door. Calibrate under RDP 0, then set RDP 1.
Concretely, i sequence RDP 0 for program and calibration, RDP 1 after the last SWD write, RDP 2 only on SKUs that forbid any later dump, and I test the sequence on 5 boards.
The reason for that specificity is a failure I have seen: RDP 1 at station 1. Stations 2 to 4 could not write calibration over SWD. 2,400 boards scrap. Calibration before RDP 1 would have stuck. Time to RDP 1 was 40 ms we spent 9 days of scrap. The RDP-1 LED was on the first jig.
RDP 1 burned before calibration writes.
| Station | Need | RDP 1 too early |
|---|---|---|
| 1 program | write | burned RDP 1 |
| 2 to 4 cal | SWD write | blocked |
| scrap | 2,400 | — |
| RDP 1 time | 40 ms | 9 days of boards |
I would not consider it settled without evidence: I would run the station sequence on 5 boards and require calibration write under RDP 0, dump fail after RDP 1, then RDP 2 only at the last cell.
Calibrate before RDP 1. RDP 1 already ends SWD.
Curated: · Written: · Reviewed:
QA-77printf goes through SYS_WRITE. There is no debugger. The core sits on BKPT. What library did you link?(show answer)
I refuse to call semihosting BKPT in a field image done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat semihosting as a BKPT 0xAB. Without a debugger it is a lockup at 18 mA, not a log.
Concretely, i link nosys or a UART retarget, I nm for BKPT 0xAB, and I fail CI if semihosting is in the map.
The reason for that specificity is a failure I have seen: First printf at 80 ms after reset. BKPT. Field current 18 mA. Certified 120 µA. 8,000 units returned in a week. UART retarget is 80 bytes. The first printf boot string never left a pin.
SYS_WRITE BKPT 80 ms after reset.
| Item | Value |
|---|---|
| first printf | 80 ms |
| field current | 18 mA |
| certified STOP | 120 µA |
| returns | 8,000 / week |
| UART retarget | 80 bytes |
I would not consider it settled without evidence: I would nm for semihosting syscalls and run 60 s detached, requiring no BKPT and current at most 200 µA after boot.
Semihosting is a debugger trap, not stdio.
Curated: · Written: · Reviewed:
QA-78You ITM-printf 80-byte lines at 1 kHz. SWO is 2 MHz. You drop 40 percent of lines. What rate did you not compute?(show answer)
The firmware question under ITM stimulus overflow is which path still races after the path I just stepped through.
I treat ITM as a FIFO with a SWO UART bit rate. 80 bytes × 10 bits × 1 kHz = 800 kbit/s of payload does not imply 40 percent loss on 2 Mbit/s SWO. Overflow is encoding, extra ports, or a measured FIFO drop.
Concretely, i budget SWO from the on-wire encoding I actually use, I count the overflow bit, I use binary traces, and I refuse a 40 percent claim that is only 800 versus 2,000.
The reason for that specificity is a failure I have seen: ITM_SendChar with a 1-byte stimulus header makes 160 kB/s of UART-8N1 on SWO for 80-byte lines at 1 kHz. Three ports plus timestamp packets burst the 6-byte FIFO. In 10 s we counted 400 of 1,000 lines with the overflow flag (40 percent measured). 800 kbit/s versus 2 Mbit/s would have looked like headroom. We missed the HardFault line. GPIO trace at 4 pins would have kept it. The ITM-enable LED was on.
80-byte ITM lines at 1 kHz on 2 MHz SWO.
| Item | Value |
|---|---|
| payload bits | 80 B × 1 kHz × 10 = 800 kbit/s |
| SWO UART | 2 Mbit/s, not a 40% proof |
| overflow flag in 10 s | 400 / 1,000 lines (40% measured) |
| HardFault lines kept | 60% |
I would not consider it settled without evidence: I would count ITM overflow for 10 s at 1 kHz 80-byte lines and require 0, or drop to 20-byte binary at 100 Hz.
SWO MHz is not printf bandwidth.
Curated: · Written: · Reviewed:
QA-79HardFault writes PC to 0x20007FF0 then spins. After reset the word is 0. Where should it have lived?(show answer)
I would budget a measurement for crash breadcrumb in retained RAM the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat a breadcrumb as needing retained or backup SRAM and a reset handler that skips zeroing it. A spin without a stored PC is a blink.
Concretely, i write PC, LR, CFSR, then IWDG-reset into no-init, I skip BSS for that section, and I dump it on the next boot UART.
The reason for that specificity is a failure I have seen: Spin at 18 mA with PC only in a register. IWDG 8 s later. BSS zeroed the SRAM. 400 RMAs with PC=0. 16-byte no-init plus skip is 20 lines of reset. The HardFault LED blinked at 2 Hz, which we called a breadcrumb.
HardFault spin without retained PC.
| Item | Value |
|---|---|
| spin current | 18 mA |
| IWDG | 8 s |
| PC after BSS | 0 |
| RMAs | 400 |
| no-init | 16 bytes |
I would not consider it settled without evidence: I would force HardFault, IWDG, and require PC in the 0x08000000 flash window on the UART of the next boot, 10 of 10.
A blinking LED is not a PC.
Curated: · Written: · Reviewed:
QA-80PA13 and PA14 are SWD. At 50 ms you remux them to PWM. The fixture cannot attach. When do you remux?(show answer)
Where I have been burned on JTAG pins remuxed as GPIO after boot is accepting a known race because "that is how MCUs work."
I treat SWD pins as a time window. Factory needs them until the last test. Field may remux. Remux at 50 ms races the fixture attach.
Concretely, i wait a 3 s window or an OTP fixture-done bit, then remux, and I never remux on engineering builds.
The reason for that specificity is a failure I have seen: Remux at 50 ms. Fixture attach 80 ms. 100 percent fail at ICT. Window of 3 s costs 3 s times 12 mA = 36 mC about 0.01 mAh, nothing. We delayed NPI 4 days. The PWM-on-PA13 LED was the app team's win.
SWD remux at 50 ms versus fixture at 80 ms.
| Event | Time |
|---|---|
| remux | 50 ms |
| fixture attach | 80 ms |
| ICT | 100% fail |
| 3 s window energy | 0.01 mAh |
| NPI delay | 4 days |
I would not consider it settled without evidence: I would attach SWD at 2.9 s and 3.1 s on 10 boards and require 10 of 10 attach before, 0 of 10 after, with OTP override.
PWM on SWD pins is a factory race.
Curated: · Written: · Reviewed:
QA-81RDP 1 is locked with a 64-bit password stored in firmware .rodata. Who can unlock?(show answer)
I answer debug unlock password in OTP by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat an unlock secret in the image as public. The password belongs in OTP written at a controlled station, not in Git.
Concretely, i generate per-unit passwords, I write OTP, I never put them in .rodata, and I audit the image for the constant.
The reason for that specificity is a failure I have seen: Password 0x0123456789ABCDEF as eight raw bytes in .rodata at 0x08002200. strings of the ELF found nothing; the constant is not ASCII. A byte-pattern scan of the .bin found it in 2 s. 12,000 units unlockable. OTP write is 40 ms. The unlock-success LED on our own jig used the same constant.
64-bit debug password in .rodata.
| Item | Value |
|---|---|
| password location | 0x08002200 .rodata |
| strings | 0 hits (not ASCII) |
| byte-pattern scan | 2 s |
| units | 12,000 |
| OTP write | 40 ms |
I would not consider it settled without evidence: I would scan the .bin for the 8-byte pattern and require zero hits, plus a per-unit OTP verify on 10 samples. I would not trust strings.
A password in flash is a comment.
Curated: · Written: · Reviewed:
QA-82You set RDP 1. After NRST, SWD still dumps RAM. Is that a fail?(show answer)
For readout protection versus SWD after reset I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat RDP 1 as blocking flash dump and some SRAM, with a defined window. I read the RM for what remains visible. Surprised RAM dumps are not how RDP works if the RM says SRAM is protected.
Concretely, i dump-test flash and SRAM after RDP 1, I compare to the RM table, and I add an MPU plus SRAM erase on RDP if the part leaks.
The reason for that specificity is a failure I have seen: RDP 1 blocked flash. SRAM still dumped 96 KB of keys in 3 s. RM said SRAM protected on this rev. We were on rev X, errata Y. 6,000 units. SRAM erase on lock is 1 ms. The RDP-1 option-byte LED was green.
RDP 1 still dumping 96 KB SRAM.
| Region | After RDP 1 | Time |
|---|---|---|
| flash | blocked | — |
| SRAM keys | 96 KB dumped | 3 s |
| units | 6,000 | — |
| SRAM erase | 1 ms | wanted |
I would not consider it settled without evidence: I would dump SRAM after RDP 1 on the shipping rev and require 96 KB of zeros or fail, not keys.
RDP is per rev, and RAM is not flash.
Curated: · Written: · Reviewed:
QA-83You use the same 48-bit MAC from a header file on every board. Provisioning thinks there is one device. What ID did you skip?(show answer)
I would write the check for unique device identity versus a cloned MAC before the driver, because the driver will otherwise certify itself.
I treat silicon UID as the identity, optionally hashed with a factory secret. A Git MAC is a clone army.
Concretely, i read the 96-bit UID, I derive a MAC in a documented LAA range, I store it once, and I fail factory if two jigs read the same UID.
The reason for that specificity is a failure I have seen: One MAC 02:00:00:00:00:01 on 20,000 boards. Cloud saw one device. UID is 96 bits, 12 bytes, 1 µs to read. Collision of UID on one lot: 0. The provision-OK LED used the header MAC.
Shared 48-bit MAC versus 96-bit UID.
| ID | Width | Unique on 20,000 |
|---|---|---|
| header MAC | 48 bits | 1 |
| silicon UID | 96 bits | 20,000 |
| UID read | 1 µs | 12 bytes |
I would not consider it settled without evidence: I would log UID and MAC on 100 boards and require 100 unique UIDs and 100 unique MACs.
A header MAC is not an identity.
Curated: · Written: · Reviewed:
QA-84You encrypt the 384 KB image with AES-CTR. An attacker flips one ciphertext bit. What happens to a jump opcode?(show answer)
The trade-off in AES-CTR without authentication is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat CTR as malleable. A bit flip in ciphertext is a bit flip in plaintext. Without GCM or a signature, the attacker patches a jump.
Concretely, i use AES-GCM or encrypt-then-ECDSA, I refuse CTR-only, and I verify before I jump.
The reason for that specificity is a failure I have seen: Flipped 1 bit in a B instruction. Destination moved 2 bytes into a gadget. 384 KB CTR encrypt 28 ms. GCM tag 16 bytes extra, 31 ms. We shipped CTR because the encrypt-done LED fired. 0 of 384 KB authenticated.
AES-CTR image with a 1-bit patch.
| Mode | Time | 1-bit flip |
|---|---|---|
| AES-CTR | 28 ms | jump patched |
| AES-GCM | 31 ms | tag fail |
| image | 384 KB | 0 bytes authenticated in CTR |
I would not consider it settled without evidence: I would flip one bit of a CTR image and require a verify fail, plus a GCM image that still boots.
CTR without a tag is editable flash.
Curated: · Written: · Reviewed:
QA-85The TRNG returns 0xFFFFFFFF forever. You still make keys. Which health test did you skip?(show answer)
I pin TRNG health test versus a stuck bit to a scope channel or a current shunt before I trust any LED or printf about it.
I treat a TRNG as failing unless Adaptive Proportion and Repetition Count tests run. A stuck-at-one is not entropy.
Concretely, i run NIST SP 800-90B startup tests, I refuse keys if they fail, and I fall back to a locked boot not to a counter.
The reason for that specificity is a failure I have seen: Stuck 32-bit ones. 256-bit key was 32 copies of 0xFF. We made 5,000 device keys. Health test is 1,024 samples, 2 ms. The TRNG-ready bit was 1 because the FIFO was full of ones. LED green.
Stuck TRNG filling a 256-bit key.
| Item | Value |
|---|---|
| TRNG output | 0xFFFFFFFF |
| 256-bit key | 32 times 0xFF |
| devices | 5,000 |
| health samples | 1,024 in 2 ms |
I would not consider it settled without evidence: I would force a stuck TRNG in a test harness and require boot refuse, plus 1,024-sample health pass on good silicon.
A full FIFO of ones is not entropy.
Curated: · Written: · Reviewed:
QA-86You store a 32-byte key at flash address 0x0807F000. PSA ITS exists on the part. Why is the raw address a defect?(show answer)
My first question on PSA ITS versus a raw flash key-value is which register bit actually flipped, not which breakpoint I hit.
I treat a raw flash address as dumpable, torn on power loss, and reachable by any bug with a pointer. PSA Internal Trusted Storage is not universally authenticated or wear-leveled; those properties are implementation-defined. On this part I use Protected Storage with authentication and anti-rollback, or I name the ITS driver's actual guarantees.
Concretely, i psa_ps_set with a UID and PS_CRYPTO_MODE, I enable rollback protection, and I never pass the flash address to the application.
The reason for that specificity is a failure I have seen: Power loss during 32-byte write tore the key to 16 valid bytes. Next boot used 16 bytes plus 0xFF. Session fail 100 percent of 40 drop tests. Protected Storage write is 12 ms with a 2-copy on this TF-M build. Raw write 0.4 ms. The flash-write LED fired at 0.4 ms.
32-byte raw flash key torn at 16 bytes.
| Store | Write time | Power-loss result |
|---|---|---|
| raw 0x0807F000 | 0.4 ms | 16 B plus 0xFF, 40 of 40 bad |
| Protected Storage (auth, 2-copy) | 12 ms | 0 of 40 mixed |
| generic PSA ITS | not assumed auth/wear-level | name the driver |
| key | 32 bytes | — |
I would not consider it settled without evidence: I would brownout during psa_ps_set 40 times and require 40 good keys or a defined empty, never a 16-byte mix.
A flash address is not a key store.
Curated: · Written: · Reviewed:
QA-87UNLOCK with a 32-byte HMAC is sniffed. The attacker plays it back 2 s later. Why does it still work?(show answer)
I treat UART command replay without a nonce as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat HMAC of a static command as replayable. I bind a nonce or a monotonic counter that the device emits.
Concretely, i send a 12-byte nonce, I HMAC the command with nonce and counter, I increment OTP or SRAM counter, and I refuse old counters.
The reason for that specificity is a failure I have seen: Sniff at 115,200 of 32 plus overhead about 3.5 ms. Replay at 2 s unlocked 100 percent of 30 tries. Counter is 4 bytes. We shipped static HMAC because the unlock LED matched the test script.
Replay of a static UART HMAC unlock.
| Item | Value |
|---|---|
| capture | ~3.5 ms at 115,200 |
| replay delay | 2 s |
| unlocks | 30 / 30 |
| nonce | 12 bytes, window 1 |
I would not consider it settled without evidence: I would replay a captured frame 30 times and require 0 unlocks, plus a nonce window of 1.
A static HMAC is a password with extra steps.
Curated: · Written: · Reviewed:
QA-88The device signs a quote of its UID. The cloud accepts the quote as a login. What binding is missing?(show answer)
On attestation quote versus authentication I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat attestation as this silicon ran this measurement. Authentication is this session is allowed. A quote without a challenge is a cloned login.
Concretely, i include a cloud nonce in the signed quote, I expire it in 5 s, and I authorize from a separate credential.
The reason for that specificity is a failure I have seen: Quote of UID only, no nonce. Attacker replayed the 64-byte signature. 800 sessions. Challenge is 32 bytes, ECDSA 42 ms. The quote-OK LED on first boot was reused as auth.
UID quote replayed as a login.
| Item | Value |
|---|---|
| quote | UID, no nonce, 64 B sig |
| replayed sessions | 800 |
| ECDSA | 42 ms |
| challenge | 32 bytes, 5 s expiry |
I would not consider it settled without evidence: I would replay a quote without a fresh nonce and require refuse, plus a 5 s challenge success.
A quote is not a session.
Curated: · Written: · Reviewed:
QA-89UART command 0xFE dumps flash in the fixture. The same firmware ships. How do you close it?(show answer)
I refuse to call factory test backdoor left enabled done from a green LED, a halt in the debugger, or a UART line that printed OK.
I treat factory commands as a compile-time or OTP gate, not as a promise we will remember to disable it.
Concretely, i compile them out of production, I blow an OTP factory-done bit that the parser checks first, and I fuzz 0xFE on a production image.
The reason for that specificity is a failure I have seen: 0xFE dumped 384 KB in 33 s at 115,200. 50,000 images. OTP gate is 1 bit, 40 ms. The fixture still used 0xFE because it was faster than SWD. Dump LED on the fixture was also in the field parser.
0xFE flash dump left in production.
| Item | Value |
|---|---|
| dump | 384 KB in 33 s at 115,200 |
| units | 50,000 |
| OTP gate | 1 bit, 40 ms |
| production 0xFE | must NACK |
I would not consider it settled without evidence: I would send 0xFE to 10 production units and require 10 NACKs, plus OTP bit set.
A fixture command is a product command if it is in the image.
Curated: · Written: · Reviewed:
QA-90You AES-128 in 12 µs in software on a Cortex-M0. A 3.3 V shunt at 1 MS/s recovers the key. What did you not budget?(show answer)
The firmware question under DPA on a naive AES round is which path still races after the path I just stepped through.
I treat software AES as leaking in amplitude. DPA needs traces, not a clock. Random delays are not a DPA countermeasure: they stretch traces that still align on the round. Hardware AES is not necessarily masked.
Concretely, i use a masked hardware AES if the part has one, I run a leakage assessment, and I never claim 12 µs AES or a jittered delay as security.
The reason for that specificity is a failure I have seen: 10,000 traces times 12 µs = 120 ms of encrypt, plus 10 ms each setup about 100 s. Key recovered. Random delays of 0–8 µs still aligned after integration. Unmasked HW AES 1.2 µs also leaked on the same shunt. Masked HW plus a passing TVLA is what we required. We had shipped software because 12 µs met the 50 µs loop. The AES-done GPIO was a perfect trigger.
Software AES-128 versus a 1 MS/s shunt.
| AES | Time | DPA |
|---|---|---|
| software M0 | 12 µs | key in ~100 s, 10,000 traces |
| random delays | 0–8 µs extra | still recovered |
| HW unmasked | 1.2 µs | still leakable |
| HW masked + TVLA | 1.2 µs | pass required |
| loop budget | 50 µs | GPIO trigger |
I would not consider it settled without evidence: I would not ship software AES on a product with a shunt-accessible 3.3 V; I would require masked HW and a leakage test, and no GPIO trigger on the round.
A fast AES is a loud AES.
Curated: · Written: · Reviewed:
QA-91You reuse the TX buffer 2 µs after you start DMA. TC is 256 µs later. What is on the wire?(show answer)
I would budget a measurement for DMA ownership until transfer complete the way I budget flash: if it is not on the bring-up list, it will not happen.
I treat the buffer as owned by DMA until TC or abort. CPU writes in between are a race, not a speedup.
Concretely, i double-buffer, I wait TC or I use a memory barrier plus TC IRQ before reuse, and I never stack-allocate a DMA buffer.
The reason for that specificity is a failure I have seen: Start DMA 256 bytes at 8 MHz SPI, 256 µs. Reused at 2 µs. First 4 bytes 0xA1, rest 0x00 from the next stack frame. CRC fail 100 percent. TC LED at 256 µs was success.
256-byte SPI DMA reused at 2 µs.
| Item | Time |
|---|---|
| SPI 8 MHz, 256 B | 256 µs |
| CPU reuse | 2 µs |
| wire | 4 B good, 252 B stack |
| CRC | 100% fail |
I would not consider it settled without evidence: I would logic-analyze the 256 bytes and require they match the snapshot at start, plus no CPU write to that range until TC.
Start is not complete. Complete is TC.
Curated: · Written: · Reviewed:
QA-92Rev Y DMA FIFO drops the last beat. You workaround on the app core. The bootloader still uses DMA. Who still drops?(show answer)
Where I have been burned on DMA FIFO errata on one silicon rev is accepting a known race because "that is how MCUs work."
I treat an errata workaround as required on every DMA client, including bootloaders and the other core, keyed by IDCODE.
Concretely, i put the FIFO flush in a shared HAL, I test bootloader XIP load and app, both revs.
The reason for that specificity is a failure I have seen: App flushed FIFO on rev Y. Bootloader did not. Last 4 bytes of a 384 KB image dropped. CRC failed 1 of 96 pages (4,096-byte pages, last page). Brick 1.0 percent. IDCODE check is 8 cycles. The app-workaround LED never ran in the bootloader.
Rev Y DMA last-beat drop in the bootloader.
| Path | Workaround | Last 4 B |
|---|---|---|
| app | yes | kept |
| bootloader | no | dropped |
| image | 384 KB, 96 times 4,096 | 1 page CRC fail |
| brick | 1.0% | — |
I would not consider it settled without evidence: I would load 384 KB on rev Y through the bootloader DMA path and require CRC pass 50 of 50, same as the app.
Errata on one core is not errata on the boot path.
Curated: · Written: · Reviewed:
QA-93You disable USART clock then write CR1. The write is ignored. You re-enable and the old CR1 is back. What did you read?(show answer)
I answer register access after the peripheral clock is gated by naming the ISR, bus, power state, and boot slot that the demo never entered.
I treat a gated APB as a black hole. Writes vanish. Reads return 0. I do not save power by gating then configuring.
Concretely, i enable clock, DSB, configure, operate, then gate. I never configure while gated.
The reason for that specificity is a failure I have seen: Gated write of UE=1 vanished. USART off. We spun 50 ms. Then ungated: CR1 still 0. 50 ms times 12 mA = 0.60 mC per boot. 1,000 boots/day = 0.17 mAh. The clock-gate LED was our power win. UART never started.
USART CR1 write with APB clock off.
| Step | CR1 UE |
|---|---|
| write while gated | vanished |
| after ungate | 0 |
| spin | 50 ms at 12 mA |
| 1,000 boots/day | 0.17 mAh |
I would not consider it settled without evidence: I would read CR1 after the sequence and require UE=1, plus a 8.68 µs bit on TX.
A gated write is a discarded store.
Curated: · Written: · Reviewed:
QA-94PB3 is SWO in AF0 and UART TX in AF7. You enable both. The console is garbage. Who owns the pin?(show answer)
For pinmux conflict between UART TX and SWO I want a number I can recompute from a capture, not a story about a board that seemed alive.
I treat a pin as one AF. SWO and UART cannot share PB3. The last AFR write wins, and the other peripheral's LED still lights.
Concretely, i allocate pins in a single table, I fail build on double-assign, and I scope PB3 for UART bit time versus SWO edges.
The reason for that specificity is a failure I have seen: AFR AF7 then debug init set AF0. Console 115,200 became SWO 2 MHz edges. BER 100 percent. UART TC still interrupted at 86.8 µs. We blamed baud. 2 days. A pin table is 1 page.
PB3 as UART TX then remuxed to SWO.
| AFR | Pin function | Console |
|---|---|---|
| AF7 | UART TX 115,200 | 8.68 µs/bit |
| AF0 | SWO 2 MHz | garbage, BER 100% |
| UART TC | still 86.8 µs | — |
I would not consider it settled without evidence: I would read AFR for PB3 and require one AF, then a UART U at 8.68 µs per bit on the scope.
Two AFs is one pin lying to both.
Curated: · Written: · Reviewed:
QA-95You DMA-circular 256 bytes. A 9-byte NMEA sentence waits 200 ms for the FIFO to fill. What interrupt did you not enable?(show answer)
I would write the check for UART FIFO timeout versus DMA circular RX before the driver, because the driver will otherwise certify itself.
I treat DMA TC as a full-buffer event. Short frames need idle-line or FIFO timeout, not a 256-byte wait.
Concretely, i enable USART idle IRQ or a 1 ms timeout, I process the NDTR remainder, and I never wait TC for a 9-byte line.
The reason for that specificity is a failure I have seen: 9-byte sentence, 256-byte DMA. At 1 sentence per 200 ms the buffer never fills. Latency until 256/9 about 28 sentences is about 5.6 s. Navigation 5.6 s stale. DMA HT LED never fired. Idle IRQ is 1 bit.
9-byte NMEA in a 256-byte DMA circular buffer.
| Item | Value |
|---|---|
| sentence | 9 B every 200 ms |
| DMA buffer | 256 B |
| sentences to TC | 28.4 |
| stale | ~5.6 s |
| idle IRQ latency | at most 2 ms |
I would not consider it settled without evidence: I would send one 9-byte line and require a task GPIO within 2 ms, not 5.6 s.
Circular DMA is not a line discipline.
Curated: · Written: · Reviewed:
QA-96SPI TXE fires on an STM32F407. You deassert CS 0 ns later. The last bit is still in the shift register for 125 ns. The slave drops it. Why?(show answer)
The trade-off in transfer-complete versus the register effect is which failure I am willing to ship, and I will not ship an unmeasured one.
I treat STM32F407 SPI TXE as "FIFO can take another byte," not as the last edge on the wire. TC can still be early relative to BSY. I wait BSY (F4) or EOT (H7), then I hold CS for the datasheet t_CSH.
Concretely, i wait SPI_SR BSY clear, I then wait the slave's CS hold from the datasheet, then I raise CS, and I scope CS versus the last SCLK edge.
The reason for that specificity is a failure I have seen: 8 MHz SCLK, period 125 ns. CS high at TXE, 40 ns before the last falling edge. Slave sampled 0. CRC 12 percent fail on 32-byte frames. TXE LED was our CS timing. BSY wait plus 125 ns t_CSH from the ADC datasheet fixed it.
STM32F407 CS raised at SPI TXE, 8 MHz last bit still shifting.
| Item | Time |
|---|---|
| SCLK | 8 MHz, 125 ns |
| CS at TXE | 40 ns before last edge |
| frame | 32 B |
| CRC fail | 12% |
| BSY plus datasheet t_CSH | 0% fail |
I would not consider it settled without evidence: I would capture CS versus SCLK and require CS high at least 1 period after the last edge, CRC 0 fail in 10,000 frames.
TXE is not the last edge. BSY or EOT plus t_CSH is.
Curated: · Written: · Reviewed:
QA-97You init USART before GPIO AF and before PLL. Baud is 9,600 on a 16 MHz HSI, then PLL goes to 80 MHz. What is the on-wire baud?(show answer)
I pin init order of clocks, pins, then peripherals to a scope channel or a current shunt before I trust any LED or printf about it.
I treat init as clocks, then GPIO, then peripheral divisors that depend on those clocks. BRR computed against HSI is wrong after PLL.
Concretely, i set PLL, I update SystemCoreClock, I then GPIO AF, then USART BRR from the real APB, and I fail a test if SystemCoreClock is not 80e6.
The reason for that specificity is a failure I have seen: BRR packed for 115,200 at 16 MHz is 0x008B, USARTDIV=8.6875, actual 16e6/139=115,108. Integer USARTDIV=9 would have been 111,111. PLL then 80 MHz, same 0x008B, actual 80e6/139=575,540 (not 5×111,111=555,555). Scope bit 1.74 µs not 8.68 µs. Framing 100 percent. LED blinked from SysTick on the new clock so boot looked fine.
USART BRR set on 16 MHz, PLL then 80 MHz.
| Stage | f_APB | Baud on the wire |
|---|---|---|
| BRR=0x008B at 16 MHz | 16 MHz | 115,108 |
| same BRR after PLL | 80 MHz | 575,540 |
| integer 9 myth | 80 MHz | 555,555 (wrong) |
| wanted | 80 MHz | 115,200, 8.68 µs |
I would not consider it settled without evidence: I would measure TX bit time after init and require 8.68 µs plus or minus 2 percent at 115,200, with SystemCoreClock 80 MHz.
BRR is only right for the clock you had.
Curated: · Written: · Reviewed:
QA-98CPU fills a 512-byte TX buffer. DMA reads 512 zeros. You invalidated after. What did you need before?(show answer)
My first question on D-cache clean before DMA from SRAM is which register bit actually flipped, not which breakpoint I hit.
I treat CPU stores as landing in D-cache. DMA reads SRAM. I must clean (push) before DMA-from-memory, and invalidate after DMA-to-memory.
Concretely, i SCB_CleanDCache_by_Addr on the TX range before DMA start, 32-byte aligned, and I never swap clean and invalidate.
The reason for that specificity is a failure I have seen: 512 bytes in cache. DMA read 512 zeros from SRAM. SPI sent zeros. CRC fail 100 percent. Invalidate-after did nothing useful on TX. Clean is 512/32=16 lines. TC LED fired. Analyzer showed 512 times 0x00.
512-byte DMA TX without D-cache clean.
| Step | SRAM view | Wire |
|---|---|---|
| CPU fill, no clean | 0x00 | 512 times 0x00 |
| clean 16 lines | 512 B | match |
| CRC | 100% fail | 0% after clean |
I would not consider it settled without evidence: I would dump SRAM with cache off after fill plus clean and require the 512 bytes, then the SPI capture to match.
Invalidate is for RX. Clean is for TX.
Curated: · Written: · Reviewed:
QA-99M4 and M7 both write a 64-byte log slot. You see mixed ASCII. What did you call how dual-core works?(show answer)
I treat dual-core shared SRAM without a protocol as unfinished until a probe, a WCET bound, or a current trace can contradict me.
I treat shared SRAM as needing a lock or a single-producer ring. Two cores writing bytes is a race I can measure, not a silicon tax.
Concretely, i use HSEM, I give each core a slot, or I DMA through a mailbox with occupancy bits, and I fail a test that finds mixed lines.
The reason for that specificity is a failure I have seen: 64-byte slot, both cores 1 kHz. Mixed lines 7.4 percent. 10 min times 2,000 writes per second is 1,200,000 writes, 88,800 mixed. HSEM is 12 cycles. We wrote MCU artifact in the ticket. The log LED on each core toggled.
Two cores, one 64-byte log slot, 1 kHz each.
| Item | Value |
|---|---|
| writers | 2 times 1 kHz |
| 10 min writes | 1,200,000 |
| mixed | 7.4% = 88,800 |
| HSEM | 12 cycles |
I would not consider it settled without evidence: I would grep mixed signatures in 10 minutes of logs and require 0, with HSEM held around the 64 bytes.
Shared SRAM without a protocol is a torn log.
Curated: · Written: · Reviewed:
QA-100You apply the rev-Y FLASH prefetch workaround on every die. Rev W hangs in the workaround. What do you read first?(show answer)
On silicon revision in the errata table at boot I start from the silicon effect the datasheet names, then ask how I would observe it.
I treat DBGMCU_IDCODE as a boot input to a table of workarounds. A recipe is not a family default.
Concretely, i switch on rev, I skip unknown revs to a safe halt with a breadcrumb, and I test each recipe on the die it names.
The reason for that specificity is a failure I have seen: Rev W ID 0x1000 hung in a prefetch-off sequence that is illegal on W, 14 ms after reset, 18 mA. Rev Y 0x2001 needs it. 15 percent of a mixed reel (1,200 of 8,000). IDCODE read is 1 load. The workaround-applied LED was in an always-on compile flag.
FLASH prefetch workaround on mixed revs.
| Rev | IDCODE | Prefetch-off recipe |
|---|---|---|
| W | 0x1000 | hang at 14 ms, 18 mA |
| Y | 0x2001 | required |
| mixed reel | 1,200 / 8,000 (15%) | hung |
| IDCODE | 1 load | — |
I would not consider it settled without evidence: I would boot 10 of each rev and require 10 of 10 RUN with the matching recipe, 0 hangs, IDCODE logged.
A workaround without IDCODE is a hang on the other rev.
Curated: · Written: · Reviewed:
