Skip to content
Tech Interview Prep home
Technical interview guide

ROS (Robot Operating System) Fundamentals

The de facto standard middleware for robotics software — nodes, topics, and services that let components communicate.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Robotics Engineer

Scope: ROS 2 Rolling node, topic, service, action, QoS, executor, callback-group, lifecycle, tf2, parameter, security, rosbag2, and composition documentation reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Design the graph as a distributed system with physical consequences

ROS 2 is middleware and a software architecture toolkit for robotic systems. Nodes expose typed publishers, subscriptions, services, actions, parameters, and transforms over a graph implemented on DDS. Interviewers at senior level are not checking whether you know the API names — they are checking whether you treat the graph as a distributed system with physical consequences: asynchronous discovery, delayed or lost messages, disagreeing clocks, restarting peers, evolving schemas, overflowing queues, blocking callbacks, and network identities that need governance. The weak answer in a system-design round is the one that talks only about nodes and topics; the strong answer reaches for QoS, executors, time, and failure behavior unprompted.

Nodes, topics, services, actions: the decision rule

A node represents a coherent responsibility and lifecycle, not every function in your program. Its public contract includes names and namespaces, message or service types, units and frames, QoS, timing, parameters, diagnostics, ownership, failure behavior, and version compatibility. Hidden coupling through graph discovery or implicit remaps is a defect, not a convenience. Duplicate node names, unexpected publishers, and multiple transform authorities are configuration faults that must be detectable — interviewers probe this directly: "what happens if two nodes claim the same name?" The weak answer shrugs; the strong answer says it's a fault you detect and alarm on.

Choosing the interface is the first design decision, and it follows one rule:

  • Topic — one-to-many, asynchronous, fire-and-forget or QoS-managed delivery. Use it for continuous streams: sensor samples, state estimates, events. No notion of a requester, so no way to know anyone consumed it.
  • Service — request/response, bounded. The caller blocks (or times out) waiting for a reply. Poor fit for long work or continuous data: one slow service call stalls the caller, and there is no feedback or cancellation.
  • Action — a goal with acceptance, feedback, cancellation, and a result. Use it for long-running operations: navigate to pose, execute a grasp. Clients must handle rejection, timeout, cancellation races, server restart, duplicate intent, and uncertain outcome.

A received request is not automatically authorized or physically applied. That sentence is the one interviewers want to hear when they ask about services and safety.

Why ROS 2 exists: the ROS 1 comparison

Expect "why did your team move to ROS 2?" or "what's actually different?" The honest answer is architectural, not syntactic:

  • No roscore. ROS 1's master handled name registration and lookup: nodes registered with it and asked it how to reach each other. That made it a single point of failure for discovery — kill the master and no new connections can be established — even though message transport itself was direct peer-to-peer (TCPROS/UDPROS) between nodes. ROS 2 discovery is peer-to-peer over DDS, so no process holds the graph's naming hostage.
  • DDS underneath. ROS 1 had custom TCP transport with essentially one QoS policy. DDS brings negotiated per-topic QoS, discovery, and security hooks.
  • Security. ROS 1 had none. ROS 2 can provide participant identity, authentication, encryption, integrity, and access control through DDS security.
  • Real-time and production targets. ROS 1's tooling and master process made deterministic deployment painful; ROS 2 supports lifecycle management, composition, and embedded/RTOS targets.
  • Multi-robot and multi-machine. Peer discovery and configurable QoS make lossy, multi-machine networks workable where ROS 1's assumptions broke.

The API and build system changed too: a client library layer (rcl) under language bindings (rclcpp, rclpy), ament/colcon instead of catkin, and executors you configure rather than a hidden spin loop. If you used ROS 1, name the differences you personally hit — the master as a single point of failure for discovery and the inability to express "best-effort, keep only the latest" are the two that come up most.

The DDS layer underneath

Under every topic is a DDS publisher and subscriber pair, matched per-topic after discovery. Peers find each other through multicast or unicast discovery traffic, exchange endpoint information, and then negotiate whether the offered and requested QoS are compatible. Two consequences interviewers care about:

  • Discovery is asynchronous. A node that publishes before its subscriber exists may lose those messages entirely unless durability covers the gap. Never assume a connection exists at time zero.
  • Lossy and multi-machine networks change behavior. Reliable QoS retransmits; best-effort drops. On Wi-Fi between robots, a reliable stream with a small depth can fall behind and then burst stale data, while best-effort keeps only fresh samples. "It worked on the bench with one machine" is the failure story to tell here.

QoS: part of every topic contract

Reliability, history, depth, durability, deadline, lifespan, liveliness, and lease duration affect both compatibility and runtime semantics. Reliable does not mean unbounded or timely — it means retransmission, which adds delay and can still fail under disconnection. Transient-local durability replays state to late joiners (the ROS 1 "latched" behavior), but a retained stale command must never become valid merely because it was retained. Sensor streams usually want best-effort with shallow keep-last depth — the latest sample is the only one that matters; state and configuration topics may want reliable or transient-local.

The trap: mismatched QoS silently prevents two nodes from connecting. A reliable publisher and a best-effort subscriber match (offered satisfies requested); a best-effort publisher and a reliable subscriber do not, and the failure is silent — no data, no error, only a warning in the logs if you look. Interviewers ask "your subscriber receives nothing, the topic echoes fine — what do you check?" QoS compatibility is the first answer. Test offered/requested compatibility, late joins, loss, reordering, partitions, and reconnect as part of integration, not as an afterthought.

Executors, time, and transforms

Executors decide when ready callbacks run. A single-threaded executor gives simple serialization, but one blocking callback delays all others — the classic "my control loop stutters when the camera callback runs" bug. A multithreaded executor creates concurrency only where callback groups and application locks allow it: mutually exclusive groups serialize, reentrant groups permit overlap and require thread-safe state. Synchronous waiting inside a callback can deadlock when the completion callback needs the same executor or group. Budget callback runtime, queue age, timer jitter, and blocking; isolate physical control from slow I/O and unbounded work.

ROS time, system time, and steady time serve different purposes. Simulation playback may pause or jump ROS time; deadlines and elapsed intervals need a monotonic source. Every message timestamp must identify when and on which clock the measurement occurred — "the timestamp is wrong" is usually "the timestamp is on a different clock."

tf2 stores a time-indexed graph of frame transforms. Request the target-source relationship at the observation time, avoid excessive extrapolation, ensure one authority per dynamic edge, and keep the map/odom/base/sensor frame responsibilities distinct. "Who publishes odom and who publishes map→odom?" is a standard probe; the weak answer doesn't know that the map→odom edge is where localization corrections land.

Parameters, lifecycle, and composition

Parameters are node-owned runtime configuration, not an untyped global database. Declare type, range, units, defaults, mutability, validation, and whether a change can be applied atomically while active. A callback that accepts an invalid partial parameter set leaves a node in mixed configuration. For safety- or timing-relevant changes: stage, validate dependencies, transition lifecycle state, apply atomically, record provenance, support rollback.

Managed lifecycle nodes make configuration, activation, deactivation, cleanup, shutdown, and error processing explicit. Activation is an authority boundary: publishers, timers, actuators, and subscriptions should not energize behavior before configuration, calibration, transforms, interlocks, and dependencies are valid. Restart must not replay stale goals or duplicate physical effects. Supervisors order dependent transitions, enforce readiness and timeouts, and choose bounded recovery.

Composition loads multiple node components into one process, reducing serialization and copies but sharing fate, address space, executor resources, and security context. Separate processes buy fault and resource isolation at communication cost. Choose boundaries from latency, copying, crash containment, permissions, deployment, observability, and real-time requirements — then test both callback interaction and process failure.

Launch files, remapping, and namespacing are the composition layer that lets you reuse a node without editing code: the same driver node, remapped to a different topic and placed in a robot-specific namespace, becomes two instances. Interviewers probe "how do you run the same node twice against different hardware?" — remap and namespace, and then check that parameters and frames follow.

Security, evidence, and operating the graph

ROS 2 security provides participant identity, authentication, encryption, integrity, and access control through DDS security. It requires keystore and certificate lifecycle, least-privilege policy for exact domains/topics/services/actions, secure time and provisioning, rotation and revocation, defined failure behavior, and protected debug/update paths. A namespace is not an authorization boundary. Local safety interlocks must remain authoritative when a remote peer is compromised or disconnected.

rosbag2 is evidence, but recording is selective and can alter timing or drop data. Preserve topic types, QoS metadata, clocks, transforms, parameters, calibration, code/config versions, storage and compression settings, recorder drop/error counters, and raw-data identity. Playback does not recreate hardware timing, discovery, nondeterministic executor scheduling, external state, or actuator physics — use deterministic simulation and hardware-in-the-loop to close those boundaries.

Operate the graph with metrics for discovery and endpoint counts, QoS incompatibility, publish-to-callback and sensor-to-actuator latency, age and drop, deadline and liveliness events, executor queue and callback duration, timer jitter, transform age/extrapolation, parameter and lifecycle changes, action acceptance/cancel/result, service timeout, CPU/memory/network, security denial, restart and recovery, and safe-stop time. Test cold start, reordered startup, duplicate names, dependency loss, slow callback, QoS mismatch, packet loss, partition, clock jump, stale transform, parameter race, node/container crash, malformed payload, credential expiry, bag saturation, and repeated restart.

What interviewers probe, and the weak answers

  • "Design the software for this robot." Weak: a list of nodes and topics. Strong: responsibilities and lifecycles, interface choices justified by delivery semantics, QoS per contract, executor and callback budget, time and frame authorities, and what happens on failure.
  • "Your subscriber gets nothing but ros2 topic echo works." Weak: "check the code." Strong: QoS incompatibility first, then discovery timing, then namespace/remap mistakes.
  • "The robot jerks when the network is busy." Weak: "add bandwidth." Strong: best-effort shallow-depth QoS for sensor streams, executor isolation of control from I/O, and measured sensor-to-actuator latency.
  • "Two nodes publish the same transform." Weak: "it usually works." Strong: one authority per dynamic edge, detection of multiple authorities as a fault.
  • "A service call takes ten seconds and the caller hangs." Weak: "increase the timeout." Strong: this work is an action, not a service; the caller needs feedback and cancellation.

Likely follow-ups: how you'd test QoS mismatch before deployment, how you'd make a restart idempotent for physical effects, and how you'd secure a multi-robot deployment. Have a concrete story with numbers for each.