Overview
Curated: · Written: · Reviewed:
A bootloader establishes which firmware is allowed to execute and how a device returns to service after interrupted or faulty updates. Its trust boundary includes immutable or protected root code, keys and policy, image storage, update transport, manifest parsing, signature verification, version and dependency checks, selection metadata, reset cause, recovery entry, and the handoff to an application. A checksum detects accidental corruption; it does not authenticate an image. A valid signature is insufficient when hardware identity, compatibility, dependency, or rollback policy fails.
A resilient design separates download from authorization and activation. It validates bounds before reads or copies, authenticates metadata and payload, enforces hardware/product identity and anti-rollback policy, writes power-fail-safe state transitions, boots a candidate provisionally, and commits only after an application health decision. Power loss must be tested at every program/erase boundary. A/B, swap, overwrite, execute-in-place, and recovery designs have different atomicity, scratch-space, wear, and availability properties.
The handoff must establish the application's expected clock, memory, privilege/security, stack, vector, interrupt, cache, peripheral, and watchdog state. Update keys require rotation and revocation plans; factory and field recovery must be authenticated and physically/logically constrained. Image encryption does not replace signature verification and often leaves plaintext in execution storage.
The production invariant is authenticated recoverable boot: after any reset or credible interruption, the device selects only policy-authorized compatible firmware or enters a bounded authenticated recovery path, while retained evidence explains selection, rejection, rollback, and commit.
Power-loss during flash is the test that separates a bootloader from a coin flip. A 2 KiB erase on a 128 KiB sector that also holds the 32-byte swap trailer, interrupted at 1.1 kV brownout, leaves 0xFF in the magic and 0x00 in the copy-done flag. The next boot reads “swap in progress”, tries to resume, erases the other slot, and both slots are gone. A/B designs survive this only if metadata lives in a separately erased, journaled, or dual-copied region and every reset point is named: downloading, candidate-unconfirmed, confirmed, reverting. Confirm must be an application health decision after 30 s of sensor and comms checks, not “main() was reached”: a firmware that boots and then hard-faults in the 31st second looping on an unconfirmed slot is the point of the confirm window.
CRC32 of a 384 KiB image is 2^{-32} collision for accidental flash errors and zero protection against an attacker who recomputes the CRC. ECDSA P-256 over SHA-256 on the manifest, with the hardware identity and security counter inside the signed object, is the authorization step. Anti-rollback that lives in RAM is not anti-rollback: a glitch reboot restores counter 3 and accepts the signed-but-vulnerable 2.1.0 after 2.2.0 was installed. Put the monotonic counter in OTP, eFuse, or a wear-leveled protected region updated only after a successful confirmed boot. Key rotation needs two slots in the root: devices still verifying key A must accept images that add key B, then a later image that revokes A; skipping the overlap bricks the 12% of the fleet that missed the first update.
Encrypted execute-in-place without a signature still boots attacker ciphertext that decrypts to whatever the attacker chose if they also stole the device key; encryption is confidentiality of the image at rest, not authorization. Bounds-check every TLV length before pointer arithmetic: a 0xFFFF length on a 48-byte trailer wraps and copies 64 KiB of RAM over the vector table. Test reset at every program/erase boundary on a sample of parts, including the exact brownout threshold, and keep a recovery image that is itself signed, anti-rollback constrained, and physically or cryptographically gated.Handoff state is a checklist that bricks when skipped. The application expects VTOR at its vector table, MSP at the top of its stack, interrupts disabled until it enables them, the PLL at the documented SYSCLK, and the watchdog either fed or explicitly stopped for the first 100 ms. A bootloader that leaves Systick enabled with the bootloader handler will fire 1 ms after jump into an invalid vector. Cache and MPU must be left in a defined state; an enabled I-cache with the old mapping fetches stale vectors.
Recovery images must themselves be updatable under policy or they become a second unpatched OS. Size the recovery to a UART XMODEM or USB DFU path that works with the product's only remaining interface after a bad app, and authenticate it. Trailer TLVs: a length of 0x0000 is end; a length of 0xFF00 on a 40-byte remainder is an overflow. Parse with remaining-bytes checks before memcpy. Version compatibility: app 3.4 requiring bootloader ≥2.1 must be in the signed manifest; installing 3.4 on bootloader 2.0 that does not understand the new TLV layout is a brick even if the signature is valid. Test the matrix of bootloader N / app N-1 / N / N+1 on hardware, including power loss, not only the happy pair.
Boot verification is a release-image ritual that starts at a brownout injector and ends at a slot counter. Record SYSCLK, wait states, compiler flags, .map sizes, painted stack high-water marks, ISR GPIO timing, logic-analyzer traces of CS/SCK/SDA, current-shunt waveforms at not less than 100 kHz, reset-cause and fault registers, and the boot slot/security counter after every power-loss injection. A pass is a number that can be recomputed from those artifacts: flash LOAD versus FLASH LENGTH, ISR high-water versus period, Stop current versus the schematic budget, confirm window versus the health checks, and disable-to-deny for debug and keys. If the only evidence is a green LED, a UART log, or a debugger session on an -O0 build, the claim is unpublished. Repeat the same measurements at the temperature and voltage corners the datasheet allows, because flash wait states, Stop leakage, crystal error, and brownout thresholds all move, and a 25 °C passing suite is not a 85 °C passing suite.
Worked example: confirm at main() bricks the 31st second
384 KiB candidate. Application hard-faults 31 s after jump.
| confirm | slot after reset | next boot |
|---|---|---|
| confirm when main() is reached | confirmed bad image | loops the fault |
| 30 s sensor/comms health, then confirm | unconfirmed trial | revert to last good |
Confirmation is a health decision after the window, not a successful C jump.
