Skip to content
Tech Interview Prep home
Technical interview guide

Perception & SLAM

How a robot builds a map of an unknown environment while simultaneously figuring out its own location within it.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed
Relevant for
Robotics Engineer

Scope: ORB-SLAM3, Cartographer, GTSAM, Ceres, OpenCV 4.x, ROS 2 Rolling, Kalibr, PCL, Nav2, TUM RGB-D, KITTI, and evo references reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

Estimate state with explicit uncertainty, provenance, and failure boundaries

Perception converts sensor measurements into useful estimates about a robot and its environment. Localization estimates pose in a known map; mapping estimates the environment from poses and observations; SLAM estimates both together. The coupling is fundamental: mapping errors corrupt localization, while pose errors distort the map. A production system therefore manages probabilistic state, calibration, timing, coordinate frames, data association, optimization, map lifecycle, and safe degradation—not just a point-cloud or image algorithm.

Begin with sensor physics and task requirements. Cameras have projection, distortion, exposure, rolling-shutter, lighting, and texture constraints. LiDAR has range, angular sampling, reflectivity, motion distortion, weather, and return artifacts. IMUs measure specific force and angular rate with bias, noise, scale, saturation, vibration, and temperature effects. Wheel odometry depends on geometry and slip. Characterize rate, latency, jitter, timestamp origin, units, covariance, field of view, blind zones, failure modes, and environmental envelope.

Spatial and temporal calibration are part of the estimator. Intrinsics, sensor-to-body extrinsics, encoder geometry, IMU parameters, and clock offsets need versioned provenance and uncertainty. The freshest sample from each device may not describe the same instant. Transform observations using a connected frame graph at measurement time, compensate motion where the model supports it, and reject excessive interpolation or extrapolation. A centimetre extrinsic or millisecond timing error can become a systematic trajectory error during motion.

The front end extracts and associates observations. Visual features, direct photometric residuals, scan descriptors, ICP, or learned representations each have assumptions. Data association is the core ambiguity: deciding whether two observations correspond to the same physical place or landmark. Outliers must be rejected using geometry, robust estimation, sensor quality, and temporal consistency. Dynamic people, vehicles, reflections, repeated corridors, and seasonal or lighting change can violate static-world or appearance assumptions.

Odometry estimates local incremental motion and normally drifts. Visual odometry may lose scale with a monocular camera and becomes weak with blur, low texture, pure rotation, or poor parallax. LiDAR odometry degenerates in geometrically repetitive or underconstrained scenes. IMU integration supplies fast motion information but bias makes unchecked error grow. Fusion is valuable only with correct frames, time, measurement models, covariance, correlations, and observability. Adding a noisy or miscalibrated sensor can make the estimate more confidently wrong.

The back end represents poses, landmarks, biases, or map variables in a filter or factor graph. Factors encode measurement residuals and uncertainty. Nonlinear least squares needs a valid manifold representation, initialization, scaling, robust loss, sparsity-aware solving, termination criteria, and marginalization policy. Global position and yaw may have gauge freedom without an anchor. Covariance and residual statistics require interpretation; low reported uncertainty is not proof of accuracy when the model is wrong.

Loop closure recognizes a previously visited place and adds a long-range constraint that can correct accumulated drift. It can also catastrophically deform a map when perceptual aliasing produces a false match. Candidate retrieval must be followed by independent geometric verification, consistency with current graph and sensor context, robust optimization, and a policy for rejecting or rolling back damaging constraints. Observe corrections and preserve the evidence needed to explain them.

The map is a deployed artifact. Define representation—landmarks, occupancy, voxels, surfels, meshes, semantics—resolution, coordinate origin, uncertainty, freshness, ownership, and update policy. Separate locally continuous odom from globally corrected map frames so loop closure does not jerk the controller. Version maps with calibration, sensor configuration, environment assumptions, and compatibility metadata. Detect environmental change and decide when to localize, update, fork, rebuild, or retire a map.

Relocalization and recovery are explicit modes. On low tracking quality, bound reliance on stale pose, slow or stop according to hazard, search for a verified relocalization, and avoid silently jumping between hypotheses. The kidnapped-robot case, map mismatch, changed layout, sensor obstruction, and compute overload need deterministic responses. Localization confidence must influence planning and control, but normal confidence scores do not replace independent collision sensing or safety-rated protection.

Evaluate with independent ground truth and exact protocol. Absolute trajectory error captures global disagreement after a declared alignment; relative pose error captures local drift over intervals. Alignment can hide scale, yaw, or origin defects, so report the allowed transform. Segment by speed, distance, rotation, scene, lighting, weather, texture, geometry, sensor mode, and failure/recovery event. Include latency, CPU/GPU/memory, dropped frames, map size, loop precision, time to relocalize, pose jumps, uncertainty calibration, and downstream planning safety—not only a mean trajectory score.

Test calibration perturbation, clock offset and drift, packet loss and reordering, blur, darkness, glare, textureless surfaces, repetitive structure, glass, rain or dust, dynamic crowds, rapid rotation, vibration, sensor saturation, map change, false loop candidates, no overlap, feature scarcity, geometric degeneracy, restart, corrupted map, and resource pressure. Record raw-data identity, code/config/model/calibration/map versions, frames and clocks, random seeds, hardware, solver diagnostics, residuals, covariance, loop decisions, pose corrections, and failure transitions so results are reproducible and operationally actionable.