Skip to content
Tech Interview Prep home
Technical interview guide

Embedded Security

How a device establishes that its firmware hasn't been tampered with, and the physical-access attack classes that don't exist in typical server-side security.

Read
65 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: PSA Crypto 1.5, Storage 1.0, Attestation 2.0; current TF-M and MCUboot; NIST IR 8259A/SP 800-193; OWASP ISVS 1.0 reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Embedded Security

Review status: rewritten from reviewer feedback; coverage and framing checks previously failed. Quality score: pending re-review.

The mental model: a chain of trust enforced at every lifecycle state

An embedded device is a small computer an adversary can hold in their hand, flash, reset, and probe with a $20 debugger. Security is not a feature you add; it is an invariant you enforce across the whole lifecycle — development, manufacturing, provisioning, operation, update, service, and decommissioning. The invariant: at every lifecycle state, every sensitive asset and privileged transition has an identified authority, authenticated policy, a least-privilege enforcement point, defined failure/recovery behavior, and auditable evidence.

"The device is behind a firewall" is not a hardware trust boundary. Neither is "the case is screwed shut." The threat model has to name assets, attackers, the physical/network/debug interfaces they can reach, trust anchors, privilege and isolation boundaries, supply-chain assumptions, recovery authority, and consequences. A device that boots successfully has not demonstrated security; security status must survive adversarial inputs, resets, rollback attempts, credential events, and operational updates.

The chain begins with an immutable or protected root of trust, authenticated boot and update policy, unique device identity, protected keys and counters, trustworthy entropy, least-privilege interfaces, memory/domain isolation, and robust untrusted-input parsers. Encryption without authenticity authorizes nothing: firmware encrypted but not signed still lets an attacker who can write flash install ciphertext that decrypts under the device key to attacker code.

Chain of trust mechanics: what runs before any check is possible

Boot verification is a chain because something must run first, unchecked. That something is the root of trust: mask ROM burned at silicon fabrication, or OTP fuses/eFuse programmed at provision time. It is immutable in the strong sense — the attacker cannot rewrite it without decapsulating the die, which is the expensive end of the attack-cost ladder.

The chain then works stage by stage:

  1. The immutable root (ROM or a first-stage loader whose public-key hash is fused) verifies the signature of the next stage — typically the bootloader — before jumping to it.
  2. The bootloader verifies the application image (and, in an A/B scheme, the inactive slot) before booting or swapping into it.
  3. Verification terminates at the root. Everything below the root is verified; the root itself is trusted by construction, not by signature. The design question interviewers probe: what code executes before any verification can happen, and why do you trust it? The answer is that it is minimal, immutable, and reviewed — not that it is magically safe.

Each signed object should cover not just the payload but the security-relevant metadata used to select and interpret it: image, hardware identity/version compatibility, dependencies, memory layout, and entry policy. Signing only the payload CRC, or leaving the load address unsigned, means an attacker can reposition a valid image or pair it with hostile metadata.

Anti-rollback: why a valid signature is not authorization

A signature proves an image was signed by the right key at some point in the past. It says nothing about now. An attacker who captures a legitimately signed v7 image — one with a known, patched vulnerability — can flash it back and re-open the hole. Signature validity alone is insufficient; you need monotonic version state.

The mechanism is a rollback index / monotonic security counter stored in a protected, replay-resistant location: OTP fuses (burn one bit per version step — irreversible, which is the point), a secure element's monotonic counter, or replay-protected flash (e.g., PSA Internal Trusted Storage with a hardware-backed counter). The boot chain refuses any image whose version is below the stored counter, and bumps the counter only after a new image is confirmed good.

The same store serves the command layer. A signed "open relay" CAN frame without a counter or challenge is valid forever; capturing it once opens the relay for the life of the key. Bind commands to a monotonic counter or a device-issued challenge, and reject stale counters after reboot using the same protected store. Counter overflow at 2^32 must trigger key rotation, not a wrap to zero. And the counter must be tested: counter wear or a counter that resets on power loss is a rollback vulnerability with extra steps.

Key provisioning and storage: off-device, one-way, and never in plaintext flash

The rule that decides most design questions: signing keys live off the device; verification keys on the device are public and one-way; secret keys never live in plaintext flash.

  • A fleet-wide AES key in flash is one theft away from universal impersonation. Unique device identity provisioned in the factory — private key in a locked key-management unit or OTP, public key enrolled in the backend — bounds compromise to that serial.
  • Firmware update signing keys belong in an HSM on the build/signing side. If the update key is compromised, every device in the fleet accepts attacker firmware; that is a fleet-wide event requiring a key-rotation path in the update protocol. A per-device key compromise is one serial number.
  • On-device secret keys go in a secure element, a hardware KMU, or a TrustZone-class TEE with locked storage — not a #define KEY in a header, and not a hidden flash address.
  • The provision-vs-production lifecycle matters: debug interfaces are enabled in development, and locked (fused off, or gated behind signed, nonce-fresh, per-device unlock tokens with backend audit) at provisioning. A static debug password printed in the service manual is a master key. Factory test mode that raises SWD and loads an unsigned RAM image is required on the line and forbidden in the field — fuse it off at provision or require a signed, time-bound fixture identity, with a separate factory key hierarchy revoked at provision.

Entropy is part of this story: a TRNG that is actually an LFSR seeded from ADC LSBs fails entropy tests and gives every device correlated keys. Use the platform crypto API (e.g., PSA Crypto with a health-tested TRNG, or a certified secure element) and test startup, low-entropy, reseed, and fault states. Identical boot seeds across units is a production failure, not a curiosity.

Physical attacks, in attacker-cost order

Name the classes by what they cost, because that is how you justify countermeasures:

AttackWhat it doesCostCountermeasure
JTAG/SWD readoutDump flash/RAM over debug port$20 probeRDP/lock bits, fused off at provision
Bus sniffing (SPI/UART/I2C)Capture traffic, extract keys from logs or plaintext protocolsLogic analyzer, ~$50Authenticate and encrypt on the wire; no secrets on debug UART
Voltage/clock glitching, fault injectionSkip a signature-check branch or a counter compareBench gear, skill-dependentRedundant checks, control-flow hardening, sensors
Power/EM side channels (DPA/CPA, SPA)Recover AES keys from power tracesOscilloscope + shunt, free software (e.g., ChipWhisperer-class setups)Hardware AES with masking; limit encryptions per key
Decapsulation, microprobingRead fuses, patch ROMLab, thousands of dollarsUsually out of scope — say so explicitly

The cheap end of that list is the realistic threat. A 3.3 V software AES-128 at 48 MHz leaks the key in a few thousand traces through a 1 Ω shunt; if you must run AES in software, limit encryptions under one key and never encrypt attacker-controlled plaintext at a triggerable point. But the first question is always the debug port: readout protection that is not tested at the end of the line is not on. A fixture that dumps 1 KiB from 0x08000000 after RDP1 should fail the unit; if it succeeds, the fuse did not burn. Pair that with a unique serial already enrolled so a stolen dump cannot be replayed as another device.

Signed updates and OTA: the recovery path is the security boundary

A secure update is a state machine, not a flash write:

  1. Download — integrity-checked (hash) but not yet trusted.
  2. Authorize — signature verification against the fused verification key, plus version ≥ rollback counter, plus hardware compatibility.
  3. Install — into the inactive slot (A/B) or a staged image, never over the running image in place.
  4. Trial and confirm — boot the new image, confirm health, then burn the anti-rollback counter. If the counter is bumped before confirmation, a failed update bricks rollback.
  5. Recovery — power loss at any point must land in a verified, bootable state. A/B swap with a verified fallback is the standard answer; a single image with no recovery path turns every failed update into an RMA.

Unsigned metadata (version numbers, compatibility flags) alongside a signed payload reopens rollback and misbinding attacks. And "newest timestamp wins" is not a rollback policy — timestamps reset with the real-time clock battery.

Update key compromise differs from device-side key compromise in blast radius: the update key is fleet-wide and lives in your signing infrastructure (HSM, access-controlled); a device key compromise is one unit. That asymmetry is why the update protocol needs a key-rotation mechanism designed in from day one, and why vulnerability intake needs a path to a signed update — a device that cannot update is a liability with a serial number.

Worked example: signed ≠ currently authorized

Relay command and firmware slot on a Cortex-M33 with PSA ITS:

controlattacker actiondevice outcome
AES wrap of firmware, no signaturewrite attacker ciphertext to slot Bdecrypts under the device key → attacker code
signed image, security counter 12replay still-valid image v7boot rejects counter 7 < 12
HMAC over "open relay", no noncecapture one CAN framereplay forever
HMAC + ITS monotonic countersame capture after rebootreject ≤ last-seen

Authenticity is the signature. Authorization is the counter, identity, and lifecycle state. Attestation fits here too: a signed "I booted image X" quote is evidence for a verifier, tied to a challenge and device identity — it is not authorization to open a lock, and a quote without a challenge can be relayed.

Misconceptions, failure modes, and what interviewers probe

Common misconceptions that sink answers:

  • "Encrypted means authenticated." No — encryption without a signature/MAC lets an attacker substitute ciphertext that decrypts to attacker code.
  • "A MAC includes time." No — a MAC covers exactly the bytes you feed it. Freshness is a counter or challenge you add and enforce.
  • "Physical access is not adversarial" / "the case is sealed." A $20 SWD probe defeats both.
  • "Attestation authorizes." It appraises; the verifier decides, and only for the moment of the quote.
  • "Security testing stops at release." The vulnerability lifecycle — SBOM, intake, triage, disclosure, signed fixes, deployment measurement — is part of the design.

Failure modes to name unprompted: counter wear or counter reset on power loss; debug unlock left enabled after provision; factory test mode reachable in the field; log flooding via security telemetry (telemetry must be bounded and itself not a DoS vector); ignored crypto return values; integer wrap in a length field leading to an undersized allocation.

What interviewers probe, and what a weak answer sounds like:

  • "Walk me through secure boot on your part." Weak: "the ROM checks the bootloader signature." Strong: names the root (mask ROM / fused key hash), what the signed object covers (image, version, layout, hardware compatibility), where verification terminates, and what runs before any check.
  • "How do you stop rollback?" Weak: "we check the version number." Strong: names the protected monotonic store, when the counter is bumped (after confirmation, not install), and the overflow/recovery policy.
  • "What happens when your update signing key leaks?" Weak: "we'd rotate it." Strong: names the blast radius (fleet-wide vs per-device), the rotation mechanism in the update protocol, and why the key lives in an HSM.
  • "Someone steals a unit off a truck. What do they get?" Weak: "nothing, it's encrypted." Strong: walks the attack-cost ladder from debug readout to decapsulation and says which class is in scope.

Likely follow-ups: how you provision keys at the factory without the fixture becoming a fleet key; how RMA/debug unlock works without a master bypass; how you verify the fuses actually burned (production test, not a design slide); what telemetry you keep and how an attacker can't flood it; what your decommissioning story is (secure erasure, credential revocation).

What changes at scale, and when not to over-engineer

At fleet scale, per-device identity and enrollment become the operational core: compromise scoping, credential rotation/revocation, and update reachability are all per-serial. A fleet-wide secret turns one physical theft into a fleet event. Telemetry, attestation, and update infrastructure must themselves be hardened — a compromised update server is a fleet-wide signing-key event by proxy.

Not every device needs a secure element. The decision criterion is asset value against attacker cost: a $3 toy with no secrets needs a signed-update check and locked debug, not a TPM-class subsystem; a payment terminal or a vehicle relay controller justifies hardware crypto, masking, and lab evaluation against fault injection. Say the trade-off out loud — over-engineering burns BOM cost and battery, under-engineering burns the fleet. The weak answer treats security as a checklist maximum; the strong one allocates it by risk.

Verifying the claim: evidence, not green LEDs

Security verification is a release-image ritual that starts at a readout attempt and ends at a replayed command. Measure disable-debug, readout protection, and update authenticity as production tests on every unit. Record reset-cause and fault registers, the boot slot and security counter after every power-loss injection, and current-shunt waveforms for side-channel work. A pass is a number recomputable from artifacts: flash LOAD versus FLASH LENGTH, the confirm window versus health checks, disable-to-deny for debug and keys. If the only evidence is a green LED, a UART log, or a debugger session on an -O0 build, the claim is unpublished. Repeat at the temperature and voltage corners the datasheet allows — flash wait states, brownout thresholds, and Stop leakage all move, and a 25 °C passing suite is not an 85 °C passing suite.