Skip to content
Tech Interview Prep home
Technical interview guide

Microcontroller Architecture & Memory-Mapped I/O

How a microcontroller's CPU, memory, and peripherals are organized around a single address space, and the architectural tradeoffs between common MCU families.

Read
54 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Arm CMSIS 6 and Cortex-M programmer guidance; Linux MMIO ordering/accessor guidance; STM32 Cortex-M4 programming manual reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Why this matters in the interview

Memory-mapped I/O is the question senior embedded interviews use to separate people who have shipped firmware from people who have read about firmware. The probe chain is predictable: "What does volatile actually guarantee?" → "So why did your write not take effect?" → "How do you know the hardware did what your C code said?" A weak answer stops at "volatile prevents the compiler from optimizing the access away." A strong answer names the three separate layers — compiler reordering, CPU/store-buffer reordering, and bus/posted-write completion — and says which mechanism addresses which layer, and which layer volatile does not touch at all.

Expect follow-ups on read-modify-write races with ISRs, write-one-to-clear semantics, why a debugger read can destroy a FIFO, and what evidence you would accept that a register write worked. If you can trace one real bug — a posted write, a wrong VTOR, a torn mailbox word — with numbers, you pass this topic.

The organization: one address space, a decoder, and no separate I/O instructions

A microcontroller integrates a processor core, flash, SRAM, clocks, reset logic, buses, an interrupt controller, debug block, DMA, and peripherals into one device. The CPU sees an address map, not C objects. When the core issues a load or store, the address travels on the bus's address lines alongside data and control lines, and a bus decoder — a piece of combinational and clocked logic in the fabric — compares the address against region base/limit ranges and routes the access to flash, SRAM, a core peripheral, a device register block, or an external bus. That is why peripherals appear as addresses rather than as separate devices: the same load/store instructions and the same interconnect reach everything, and the decoder is the only thing that decides what physically answers.

This has consequences interviewers probe. Different regions can sit on different buses (I-code/D-code, AHB, APB), behind bridges with different widths and rates, in different clock and power domains, with different synchronization and timeout behavior. An access is not "a CPU register write"; it is a transaction that crosses a specific path with specific prerequisites — clock enabled, reset released, pin muxed — and the fabric may post it, merge it, or stall it.

The memory map and what it constrains

A typical Cortex-M map: flash at 0x08000000, SRAM at 0x20000000, core peripherals including the vector table at 0xE0000000, vendor peripheral blocks at vendor-defined bases, plus aliases and external memory regions. Region size, width, and placement constrain the program directly:

  • A 512 KiB flash part with a 32 KiB bootloader at 0x08000000 leaves 480 KiB for the application at 0x08008000. If VTOR is left on the bootloader table, every interrupt vectors through the old image; a 1 kHz SysTick runs the bootloader's Default_Handler loop and the application looks "hung" with a green main() that already returned. VTOR relocation on Cortex-M4 requires alignment (the architecture requires the vector-table offset to be aligned to a minimum of 128 bytes on implementations with 128 or fewer vectors — check your core); get it wrong and you fault immediately.
  • Stack collision is the other map failure: a 2 KiB MSP at the top of 64 KiB SRAM with a 48 KiB .bss and 16 KiB heap has zero slack; one 1.2 KiB ISR nesting on a 900-byte call chain overflows into .bss and corrupts a DMA descriptor 12 bytes below the stack limit. The map file plus a painted-stack high-water mark on the release image is the check; average "free heap 18 KiB" is not.
  • Running from flash vs RAM is a map decision: flash wait states, prefetch, and caches determine ISR latency. A PLL multiplying 8 MHz HSE by 21 to 168 MHz must have flash wait states set before the switch; ST's table for the F4 family wants 5 wait states above 150 MHz at 2.7–3.6 V. Teams that switch the PLL first and patch wait states later see random HardFaults in flash-resident ISRs, never in SRAM-resident handlers.

A linker script places sections into those ranges; startup code establishes stack and vector state, initializes data and BSS, configures the system, and enters the language runtime. Assert sizes, origins, alignment, and load/run addresses in CI — a hidden flash overflow is a production miss, not a build warning.

Register flavors and their side-effect contracts

A memory-mapped register is a hardware latch or logic block wired so that a bus write changes device state and a bus read returns device state. Three flavors, each with different read/write direction and side effects:

  • Control registers — you write configuration; reads may return the written value, a shadow, or garbage on genuinely write-only fields. Never assume a readback of a write-only register is meaningful.
  • Status registers — hardware writes, software reads. Flags may be clear-on-read (a read is an acknowledge and pops a FIFO entry — which is why a debugger watch window silently destroys state), or write-one-to-clear, where you write 1 to the bit you want to clear without read-modify-writing unrelated asserted flags.
  • Data registers — FIFO push/pop, self-clearing command bits, set/clear aliases that update selected bits without a read-modify-write race.

Read-modify-write is unsafe when a register has side effects, concurrent writers (an ISR touching the same register), or mixed write semantics — a plain RMW on a W1C status register can clear an event that arrived between your read and your write. Classify register semantics and ownership before choosing the update operation. Vendor headers and SVD-generated definitions reduce address mistakes but do not replace the reference manual and errata.

MMIO vs port-mapped I/O

Port-mapped (isolated) I/O gives devices a separate address space reached by dedicated instructions — x86's IN/OUT with a 64 KiB port space is the surviving example. The tradeoffs: port I/O keeps the memory address space free and can block speculative/wrong access types by construction, but it costs instruction-set real estate, needs separate compiler/toolchain support, and gives you a tiny space. MMIO uses ordinary loads and stores, so the compiler, debugger, DMA, and C pointers all work uniformly, at the price of device accesses being indistinguishable from memory accesses to the compiler — which is exactly why volatile and memory-type attributes exist. Virtually all modern microcontrollers use MMIO; port I/O survives on x86 and legacy parts.

volatile, the compiler, and the three layers of reordering

The C volatile qualifier forces observable accesses in the abstract machine: the compiler may not merge, elide, or cache volatile accesses in a register, and it must preserve their ordering among other volatile accesses. What it does not do: make an operation atomic, order it against all memory or devices, flush posted writes, protect a shared invariant, or provide a cache/DMA coherency protocol.

Separate the layers:

  1. Compiler reordering/caching — solved by volatile (and by language-level ordering rules).
  2. CPU and store-buffer reordering — volatile does not touch this; use DMB/DSB/ISB or platform primitives at protocol boundaries.
  3. Bus and device observation — a posted write makes a C store a request, not a completion. On an AHB-to-APB bridge, a store into a 48 MHz APB peripheral from a 168 MHz Cortex-M4 can sit in a write buffer for several CPU cycles; the next instruction reading a different SRAM address does not wait. If the sequence is enable-clock → write-peripheral → start-DMA, the second write can hit a clock-gated bus and be dropped or applied to a still-reset block. The cheap confirmation is a readback of a side-effect-free status bit in the same peripheral after a DSB, timed on a logic analyzer: at 168 MHz one missed cycle is 6 ns, but a 12-cycle posted-write stall is 71 ns — enough for a 10 MHz SPI clock to emit a bit before chip-select is actually asserted. Teams that skip the readback pass every unit test that mocks registers as ordinary RAM.

Use the architecture's device memory attributes, barriers, atomic/set-clear registers, critical sections, and platform accessors as required by the exact bus and peripheral. Dual-core mailboxes make this concrete: on an M4+M0 part, SRAM seen as Normal by both cores lets the M4 store buffer merge a 32-bit sequence number; mark the 64-byte channel Device-nGnRnE or insert DMB on both sides and write a single producer word last. A 50,000-message soak that still shows a torn 0x0001FFFF is the test; a code review that says "volatile uint32_t" is not.

Binding to silicon, and the evidence standard

Production firmware binds to exact silicon, package, revision, memory map, clock/reset tree, boot mode, and errata. Reserved bits and errata close the contract: writing 1 into a reserved field the reference manual says must remain 0 can enable undocumented behavior; treat silicon revision, errata worksheets, and the exact access width as part of the source, not as comments. Use named masks and an approved preserve policy rather than magic full-word constants, and gate workarounds on revision identifiers so they run only on affected silicon.

The production invariant is address-to-effect traceability: every low-level access names the authoritative register contract, establishes required ordering and ownership, and has evidence that the intended hardware effect — not merely a successful C assignment — occurred. Verification starts at the map file and ends at a probe on the pin: record SYSCLK, wait states, .map sizes, painted stack high-water marks, logic-analyzer traces of CS/SCK/SDA, current-shunt waveforms, reset-cause and fault registers after power-loss injection. A pass is a number that can be recomputed from those artifacts. If the only evidence is a green LED, a UART log, or a debugger session on an -O0 build, the claim is unpublished. Repeat the measurements at the temperature and voltage corners the datasheet allows — flash wait states, Stop leakage, crystal error, and brownout thresholds all move, and a 25 °C passing suite is not an 85 °C passing suite.

Worked example: VTOR left on the bootloader table

512 KiB flash, 32 KiB bootloader at 0x08000000, application at 0x08008000. SysTick at 1 kHz.

VTORSysTick vectors tomain()
0x08000000 (bootloader)Default_Handler loopreturned, looks hung
0x08008000, aligned per core requirementsapplication handlerruns

The C store that "enabled SysTick" succeeded. The table it indexed was the wrong image. This is the shape of most MMIO bugs in interviews and in production: the access completed, the address was right, and the contract — which table, which clock, which ordering — was wrong.

What interviewers probe, and what a weak answer sounds like

  • "Why is the register pointer volatile?" Weak: "so the compiler doesn't optimize it away." Strong: name what volatile guarantees (observable accesses, ordering among volatile accesses) and what it doesn't (atomicity, CPU reordering, bus completion), then say which mechanism covers each gap.
  • "Your write to enable the peripheral didn't take effect. Debug it." Weak: repeated the write. Strong: clock/reset gating first, then posted-write completion — readback of a side-effect-free register after a DSB, confirmed on a logic analyzer.
  • "An ISR and main both touch a register. What breaks?" Weak: "make it volatile." Strong: the RMW race on a W1C or mixed-semantics register, and the fix — set/clear aliases, a critical section, or single-owner access.
  • "How do you know the DMA descriptor wasn't corrupted?" Weak: "it works." Strong: stack high-water mark versus ISR nesting depth, from the release image, with numbers.
  • "What's the difference between a debugger read and a bus read?" Weak: nothing. Strong: a debugger read is a bus access with all the side effects of one — it can clear a read-to-clear flag or pop a FIFO.

Likely follow-ups: cache and MPU attributes for shared buffers, DMA coherency, errata-driven workarounds, and what evidence you would accept at temperature corners. Bring one traced bug with figures and you have the answer to all of them.