Overview
Curated: · Written: · Reviewed:
Hardware/software co-design turns product requirements into a jointly feasible allocation of sensing, actuation, compute, memory, buses, clocks, power, safety, security, test, cost, and update behavior. Partitioning is iterative: dedicated hardware can reduce latency and energy but adds silicon/board cost and fixed behavior; software adds flexibility but consumes execution, memory, bandwidth, and scheduling margin. The interface contract must define units, ranges, timing, ownership, reset values, state machines, errors, and lifecycle—not merely register addresses.
A peripheral driver binds one or more hardware instances to a stable subsystem API. Immutable instance configuration (MMIO regions, clocks, resets, pins, interrupts, DMA requests and properties) should be separated from mutable runtime state. Devicetree/SVD/generated descriptions reduce duplication but do not replace the silicon reference manual, board schematic, errata, or validation. Initialization ordering must follow dependencies; APIs must specify calling context, blocking/asynchronous completion, buffer lifetime, concurrency, cancellation, timeout, and power-state behavior.
Correct register access honors width, alignment, reserved bits, write-one-to-clear/read-to-clear effects, posted writes, barriers, and hardware state. Interrupts and DMA are independent execution agents; volatile is not a synchronization or cache-coherency protocol. Driver recovery must restore a known device state without inventing success, duplicating side effects, leaking ownership, or destroying the evidence needed to diagnose the original failure.
The production invariant is requirement-to-effect fidelity: every requested operation maps through a versioned API and exact hardware description to observable register/bus/physical effects, bounded timing, explicit ownership, and a classified outcome/recovery. Unit models, emulation, hardware-in-loop tests, logic/power traces, fault injection, release-build timing, and field counters must agree at their declared abstraction boundaries.
DMA is a second bus master with its own idea of completion. A driver that returns after writing the enable bit, before the last beat has landed in SRAM, hands a buffer to the task 12 cycles early; at 84 MHz AHB that is 143 ns, enough for the CPU to parse a length field that DMA is still updating. Ownership must move on the transfer-complete interrupt or a polled TC flag after a DSB, and the buffer must be in a region the DMA can see — not Tightly Coupled Memory on a core the DMA cannot address. Cache: on Cortex-M7, clean the TX buffer to point of unification before the transfer and invalidate the RX buffer after, by address, or map the DMA region as non-cacheable.
Clock gating is an access-time bug. Writing a UART while its APB clock is off can bus-fault or silently drop; the init order is clock, reset release, pinmux, then registers. Pinmux conflicts are schematic bugs that firmware will “fix” by overwriting another driver’s AFR bits: PA9 as USART1_TX and as TIM1_CH2 cannot be both; the last writer wins and the motor PWM stops. FIFO versus DMA is a rate decision: a 12.5 Mbps UART with a 16-byte FIFO interrupt at ≥11.5 µs; if the ISR worst case is 18 µs, bytes are lost and DMA is not optional.
Errata and silicon revision belong in the driver. Rev Y needs a dummy read after a specific ADC sequence; skipping it yields a stuck 0xFFF. Shared SRAM between two cores needs a memory-ordering protocol and a single owner per cache line; two cores writing adjacent bytes on a 32-bit bus tear. Init levels in Zephyr that call k_mutex_lock before the kernel is up hang forever. The contract is the oscilloscope and the errata sheet together: if the driver says “transfer complete” the CS line, the last clock, and the DMA TC flag must agree on that millisecond.Completion is not effect. A PWM driver that writes CCR and returns has not produced a pulse until the timer is running, the output compare is enabled, the pin is in the AF mode, and the dead-time register matches the FET spec (e.g. 400 ns at 20 kHz). Probe the gate with an oscilloscope; if the software log says "PWM start" 80 µs before the first edge, the driver's completion is a register write. Init order: clocks, then reset pulse of at least one APB cycle, then pinmux, then registers, then enable. Reversing pinmux and clocks glitches the pin.
Silicon revision in the driver: read DBGMCU or the equivalent ID, branch to the errata sequence, and fail closed on unknown rev rather than guessing. Shared SRAM 4 KiB at 0x3000_0000 between CM4 and CM7 with D-cache on the M7 requires the M7 to clean by address after writing a command and the M4 to invalidate before reading; skipping either side is a stale-command bug that reproduces after 30 minutes when a cache line finally evicts. FIFO-only drivers at 25 Mbps will interrupt more often than the CPU can run; the allocation to DMA is a requirement, not an optimization. Record the bit rate, ISR budget, and DMA decision in the interface contract.
Driver verification is a release-image ritual that starts at an oscilloscope and ends at an errata branch. Record SYSCLK, wait states, compiler flags, .map sizes, painted stack high-water marks, ISR GPIO timing, logic-analyzer traces of CS/SCK/SDA, current-shunt waveforms at not less than 100 kHz, reset-cause and fault registers, and the boot slot/security counter after every power-loss injection. A pass is a number that can be recomputed from those artifacts: flash LOAD versus FLASH LENGTH, ISR high-water versus period, Stop current versus the schematic budget, confirm window versus the health checks, and disable-to-deny for debug and keys. If the only evidence is a green LED, a UART log, or a debugger session on an -O0 build, the claim is unpublished. Repeat the same measurements at the temperature and voltage corners the datasheet allows, because flash wait states, Stop leakage, crystal error, and brownout thresholds all move, and a 25 °C passing suite is not a 85 °C passing suite.
Worked example: enable bit is not transfer complete
SPI RX into SRAM at 84 MHz AHB. Driver returns after writing DMA enable. CPU reads length 12 cycles later.
| completion signal | CPU sees length | DMA still writing |
|---|---|---|
| enable bit | 0xFFFF (stale) | last 4 bytes |
| TC IRQ, then DSB, then invalidate RX | 64 | 0 |
Ownership moves on documented completion, not on the MMIO write that started the engine.
