Top 100 Robotics Engineer Interview Questions and Answers
The questions most likely to actually come up in your Robotics Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A colleague writes `T_base_camera` and uses it to convert a point measured by the camera into base coordinates. When is that right, and how would you make the mistake impossible to repeat?(show answer)
The first thing I would pin down about transform direction and naming is what the machine physically does when the assumption is wrong.
A transform is directional. The convention that reads "the pose of the camera expressed in base" is exactly the matrix that maps a point in camera coordinates into base coordinates, and its inverse does the opposite. Both are valid conventions; only one can be in force in a codebase.
Concretely, write the convention down once, name every transform with both frames in a fixed order, and make composition check that the adjacent frames match rather than trusting the caller. A helper that takes source and target frame names and refuses an unmatched pair costs a few lines and removes the whole class of error.
The reason for that specificity is a failure I have seen: A pick cell composed a camera-to-tool transform in the wrong order and placed parts 214 mm from the target, which reads as a plausible calibration error rather than as an inverted matrix, so the team spent two days re-running calibration.
The same 214 mm error, two explanations.
| Hypothesis | Predicted residual | Varies with pose? | Matches log |
|---|---|---|---|
| calibration drift | 0.5-2 mm | slightly | no |
| inverted transform | 100-400 mm | strongly | yes |
| observed | 214 mm | strongly | inverted |
I would not consider it settled without evidence: Assert on a known fixture: a point at a measured location in the camera frame must map to its surveyed location in the base frame within the stated tolerance, and the test must fail if the transform is inverted.
A transform without a stated direction is not data, it is a guess.
Curated: · Written: · Reviewed:
QA-2Your transform lookup always uses the latest available value. What does that cost on a moving robot, and what would you use instead?(show answer)
I would start transform timestamps from the measured envelope, not from the number the datasheet promises.
A transform tree is a time series, not a set of constants. Asking for the newest value from each link gives a chain that describes several different instants, which is a pose the robot never held.
Concretely, look the whole chain up at one stamp — the stamp of the measurement being transformed — and accept a bounded wait or an interpolation rather than the newest sample. Declare a maximum age and fail loudly past it instead of silently using stale geometry.
The reason for that specificity is a failure I have seen: A mobile base transformed lidar returns using the newest odometry rather than odometry at the scan time; at 1.2 m/s and 40 ms of skew the map smeared walls by about 48 mm and the localiser drifted rather than converged.
Skew cost at 1.2 m/s.
| Skew | Position error | Effect on map |
|---|---|---|
| 5 ms | 6 mm | within noise |
| 40 ms | 48 mm | walls smear |
| 120 ms | 144 mm | localiser diverges |
I would not consider it settled without evidence: Replay a recorded run with deliberately delayed odometry and confirm the transform layer refuses the lookup rather than returning a mixed-time chain.
Geometry is only correct at an instant, so the instant is part of the query.
Curated: · Written: · Reviewed:
QA-3When would you store orientation as Euler angles rather than a quaternion, and what breaks if you average quaternions componentwise?(show answer)
This is an area where a clean simulation and a correct handling of rotation representation choice are not the same event.
Euler angles are a human interface with axis-sequence dependence and coordinate singularities; unit quaternions compose and interpolate without gimbal lock but represent each rotation twice, since q and -q are the same orientation. Neither is a vector space, so componentwise averaging is not an average of rotations.
Concretely, store and compute in quaternions or matrices, convert to Euler only for display with the sequence stated, normalise after every accumulation, and canonicalise sign before comparing or blending. Average rotations with a proper method rather than by adding components.
The reason for that specificity is a failure I have seen: A pose filter averaged quaternions componentwise across four estimates that straddled the sign convention; the result had a norm of 0.31 and, once normalised, pointed 68 degrees away from every input.
Four estimates, two conventions.
| Input | w | Sign canonicalised | Naive mean |
|---|---|---|---|
| a | 0.71 | 0.71 | included |
| b | -0.70 | 0.70 | cancels a |
| result norm | 1.00 | 0.31 |
I would not consider it settled without evidence: Unit-test that averaging q and -q of the same orientation returns that orientation, and that a 179-degree pair does not produce an arbitrary result.
The representation is part of the algorithm, not a storage detail.
Curated: · Written: · Reviewed:
QA-4The URDF parses, rviz looks right, and the arm still misses by 4 mm at reach. Where do you look?(show answer)
My answer to URDF as a claim about hardware begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
A model file describes the machine somebody intended to build. Joint axis sign, zero offset, link length, tool frame, and base mounting are each an independent claim that can be wrong while the file remains perfectly valid.
Concretely, measure the tool pose at a spatially diverse set of configurations with an independent instrument, fit the residual to a parameter error rather than adjusting a single offset until one point matches, and version the model with the calibration that belongs to it.
The reason for that specificity is a failure I have seen: A cell carried a 4.1 mm systematic tool error because the URDF listed a link at 300 mm where the machined part was 299.6 mm and the tool plate added 3.7 mm that nobody had modelled; a single-pose touch-up hid it at the taught point and doubled it elsewhere.
Single-pose touch-up versus envelope fit.
| Pose | Before | After touch-up | After parameter fit |
|---|---|---|---|
| taught | 4.1 mm | 0.1 mm | 0.3 mm |
| far corner | 4.4 mm | 8.2 mm | 0.5 mm |
| holdout | 3.9 mm | 7.6 mm | 0.4 mm |
I would not consider it settled without evidence: Report positional residual across at least twenty poses spanning the working envelope, with a holdout set the fit never saw.
A file that parses has been validated as XML, not as geometry.
Curated: · Written: · Reviewed:
QA-5Your IK returns a valid solution and the arm swings through the fixture on the way. What was missing?(show answer)
I would treat inverse kinematics branch selection as a claim about the physical world that has to survive a bench test.
A six-axis arm typically offers up to eight closed-form solutions for a reachable pose. Every one places the tool identically and sweeps a different volume, so choosing among them is a policy decision, not a numerical one.
Concretely, state the rule — nearest to current configuration, or a fixed branch — log the branch actually taken with each solve, and treat a branch change inside a taught path as an event that requires a fresh collision check rather than a routine solver result.
The reason for that specificity is a failure I have seen: A palletising program flipped from elbow-up to elbow-down at rank 34 of a 60-point path because the nearest-solution rule crossed a boundary; the elbow swept through a fixture that the taught path had never approached.
Same tool pose, different swept volume.
| Branch | Elbow height | Swept volume | Clears fixture |
|---|---|---|---|
| elbow-up | 1180 mm | 0.42 m³ | yes |
| elbow-down | 610 mm | 0.51 m³ | no |
| chosen at rank 34 | 610 mm | collision |
I would not consider it settled without evidence: Log the solution branch per waypoint and assert that a branch change never occurs inside a segment that was collision-checked as a unit.
The tool pose does not determine the arm, so the arm has to be chosen.
Curated: · Written: · Reviewed:
QA-6A Cartesian move slows to a crawl and then the wrist snaps around. What is happening and what would you change?(show answer)
The useful question for Jacobian conditioning near singularity is what still holds at the edge of the workspace, cold, and under full payload.
The Jacobian is configuration dependent. Approaching a singularity, one Cartesian direction costs increasingly large joint rates, and a plain pseudoinverse will happily command them because the mathematics does not know about the motor.
Concretely, monitor the smallest singular value or the condition number along the path, apply damped least squares with a damping term that grows as conditioning degrades, cap joint rate independently of the solver, and refuse the segment rather than degrading it silently.
The reason for that specificity is a failure I have seen: A wrist-singular waypoint drove joint 4 to a commanded 480 deg/s against a 200 deg/s limit; the controller clipped it, the tool lagged the path by 9 mm, and the recovery motion looked to the operator like a fault rather than a limit.
Conditioning along one path.
| Waypoint | σ_min | Commanded J4 rate | Result |
|---|---|---|---|
| 10 | 0.21 | 34 deg/s | fine |
| 22 | 0.04 | 210 deg/s | at limit |
| 24 | 0.008 | 480 deg/s | clipped |
I would not consider it settled without evidence: Record the minimum singular value along every commissioned path and gate release on a declared floor rather than on the absence of complaints.
A singularity is a property of the pose, so it has to be detected before the move, not after.
Curated: · Written: · Reviewed:
QA-7A seven-axis arm keeps drifting into a joint limit over a long task even though every Cartesian point is reached. Why, and what would you add?(show answer)
I would settle redundancy and null-space objectives against an instrumented run on real hardware before trusting the model.
A redundant arm has a null space: joint motion that changes the configuration without moving the tool. Left unmanaged it is not zero, it is whatever the solver's least-norm step happens to produce, and it accumulates.
Concretely, add an explicit secondary objective projected into the null space — joint-limit avoidance, manipulability, or a posture preference — with stated priority and weights, and bound the drift per cycle so the configuration cannot wander across a long path.
The reason for that specificity is a failure I have seen: A seven-axis inspection robot reached every one of 900 waypoints and arrived at waypoint 780 with joint 3 at 2 degrees from its limit; the next segment was unreachable and the cycle aborted mid-part.
Joint 3 margin over 900 waypoints.
| Waypoint | No null-space term | With limit avoidance |
|---|---|---|
| 100 | 41° | 44° |
| 500 | 17° | 39° |
| 780 | 2° | 36° |
I would not consider it settled without evidence: Plot each joint's distance to its limit across the full task and require a declared margin at every waypoint, not only at the first.
Unmanaged redundancy is not freedom, it is uncontrolled state.
Curated: · Written: · Reviewed:
QA-8Your velocity-level IK tracks a Cartesian path and the tool ends 3 mm off after thirty seconds. Where did the error come from?(show answer)
The judgement in differential IK drift is which timestamps and frames are pinned, not which library is fashionable.
Differential IK integrates joint rates. Integration accumulates the small residual between the twist you asked for and the twist the linearisation delivered, so position error grows even when every individual step is correct to first order.
Concretely, close the loop on pose rather than on velocity alone: recompute the pose error each cycle and feed a proportional correction into the commanded twist, so the integration is corrected rather than trusted. Bound the correction so it cannot fight the path.
The reason for that specificity is a failure I have seen: A welding pass ran 28 seconds of pure velocity IK at 250 Hz; each step was accurate to 0.4 µm and the accumulated tool offset at the end of the seam was 3.1 mm, enough to miss the joint.
Accumulated offset over one seam.
| Elapsed | Open-loop drift | With pose feedback |
|---|---|---|
| 5 s | 0.6 mm | 0.02 mm |
| 15 s | 1.7 mm | 0.02 mm |
| 28 s | 3.1 mm | 0.03 mm |
I would not consider it settled without evidence: Compare commanded and forward-kinematic pose at every cycle and alarm on a residual trend rather than on an instantaneous threshold.
An integrator without feedback is a drift generator with good intentions.
Curated: · Written: · Reviewed:
QA-9A planner reports success and the executed path violates a joint limit at run time. How do you make those two agree?(show answer)
Where candidates lose the interview on joint limits versus Cartesian tolerance is treating a replayed bag as evidence about the robot.
A plan is feasible only against the constraint set the planner was given. If the planner checks Cartesian collision but the controller enforces joint limits, velocity limits, and torque limits, the two are solving different problems and will disagree.
Concretely, give the planner the same limit set the controller enforces, including rate and acceleration bounds and any tool-dependent restriction, and validate the time-parameterised trajectory rather than the geometric path, since time parameterisation is where rate limits are actually violated.
The reason for that specificity is a failure I have seen: A path that was collision-free and within position limits demanded 6.1 rad/s on joint 6 after time parameterisation against a 4.4 rad/s limit; the controller scaled the whole segment and the tool arrived 180 ms late into a moving conveyor pick.
Where the plan and the controller disagree.
| Check | Planner | Controller | Agrees |
|---|---|---|---|
| collision | yes | no | — |
| joint position | yes | yes | yes |
| joint rate | no | yes | no |
I would not consider it settled without evidence: Run the parameterised trajectory through the controller's own limit checker before release and reject on any violation, not only on collision.
Feasibility is a property of the constraint set, so both stages need the same one.
Curated: · Written: · Reviewed:
QA-10How do you establish a tool centre point, and why is a four-point pivot not enough on its own?(show answer)
I would answer tool centre point calibration by separating what the sensor measured from what the visualiser drew.
A pivot procedure constrains the tool position because several orientations must agree on one point, but it says nothing about tool orientation, and a poorly spread set of orientations leaves the fit badly conditioned even for position.
Concretely, use well-separated orientations for the position fit, add an explicit orientation procedure against a known direction, report the residual of the fit rather than only the resulting numbers, and re-verify after any tool change or crash.
The reason for that specificity is a failure I have seen: A gripper TCP was taught with four poses spanning 22 degrees of wrist rotation; the position residual was 0.8 mm but the tool Z axis was 3.4 degrees off, which put a 180 mm long tool 10.7 mm out when the wrist flipped.
Orientation spread and fit quality.
| Spread | Position residual | Orientation error | Tip error at 180 mm |
|---|---|---|---|
| 22° | 0.8 mm | 3.4° | 10.7 mm |
| 90° | 0.3 mm | 0.4° | 1.3 mm |
| 150° | 0.2 mm | 0.2° | 0.6 mm |
I would not consider it settled without evidence: Report the pivot residual and the orientation check separately, and reject a fit whose orientation spread is below a declared threshold.
A tool has six numbers, and a pivot measures three of them.
Curated: · Written: · Reviewed:
QA-11Explain eye-in-hand versus eye-to-hand calibration and what each one leaves unconstrained.(show answer)
The engineering content of hand-eye calibration is the tolerance budget and the stop condition, not the algorithm name.
Hand-eye calibration solves for a rigid transform from a set of paired robot motions and camera observations. The solution is constrained only by the rotations present in the motion set: pure translations leave the rotational part unidentifiable.
Concretely, collect poses with rotation about at least two independent axes, keep the target well inside the depth of field across the set, solve with a method appropriate to the noise, and validate on a holdout pose that the fit never saw rather than on the fit residual.
The reason for that specificity is a failure I have seen: An eye-in-hand set was collected by translating the camera across a table with under 5 degrees of rotation; the solver returned a confident answer whose rotational part was wrong by 6 degrees, which appeared as a depth-dependent pick error nobody could reproduce at the teach distance.
Rotational coverage and identifiability.
| Set | Max rotation | Residual | Holdout error |
|---|---|---|---|
| translation only | 5° | 0.4 px | 6.0° |
| two axes | 45° | 0.6 px | 0.5° |
| three axes | 70° | 0.5 px | 0.3° |
I would not consider it settled without evidence: Report the rotational spread of the calibration set alongside the residual, and reject a set below a declared angular coverage.
A calibration is only as identifiable as the motion that produced it.
Curated: · Written: · Reviewed:
QA-12A datasheet quotes 0.02 mm repeatability. What does that let you promise a customer, and what does it not?(show answer)
Before letting it move at speed I would write down what a wrong result for repeatability versus accuracy looks like on the log.
Repeatability is the spread on returning to the same commanded configuration. Accuracy is the distance between the commanded pose and the achieved pose in world coordinates. A machine can be extremely repeatable and consistently wrong.
Concretely, design taught applications around repeatability and calibrated applications — vision-guided placement, offline programming, tool exchange — around accuracy, which is typically an order of magnitude worse and improves only with kinematic calibration.
The reason for that specificity is a failure I have seen: An offline-programmed cell was scoped on a 0.02 mm repeatability figure; measured absolute accuracy at reach was 0.9 mm, and every program generated from CAD needed manual touch-up, which erased the reason for programming offline.
One arm, two figures.
| Measure | Value | Applies to |
|---|---|---|
| repeatability | 0.02 mm | taught points |
| accuracy, calibrated | 0.25 mm | CAD programs |
| accuracy, uncalibrated | 0.9 mm | CAD programs |
I would not consider it settled without evidence: Measure absolute accuracy over the working envelope with an independent instrument before promising any CAD-driven placement tolerance.
Repeatability is about coming back, accuracy is about being right.
Curated: · Written: · Reviewed:
QA-13Position feedback says the joint is at target and the tool is 0.4 mm off under load. Explain.(show answer)
The first thing I would pin down about backlash and compliance in the loop is what the machine physically does when the assumption is wrong.
Motor-side feedback measures the motor, not the tool. Gearbox backlash, belt stretch, and structural compliance sit between them, so the encoder is telling the truth about a quantity that is not the one you care about.
Concretely, characterise the deflection against load rather than assuming it, approach critical points from a consistent direction so backlash contributes a repeatable offset, and add load-side sensing where the tolerance demands it rather than tightening the control gains.
The reason for that specificity is a failure I have seen: A press-fit station held 0.02 mm on the encoders and produced a 0.41 mm tool deflection at 60 N of process force; raising the position gain made the loop oscillate without moving the tool any closer.
Deflection against process force.
| Force | Encoder error | Tool deflection |
|---|---|---|
| 0 N | 0.02 mm | 0.02 mm |
| 30 N | 0.02 mm | 0.21 mm |
| 60 N | 0.02 mm | 0.41 mm |
I would not consider it settled without evidence: Measure tool deflection against applied load across the envelope and publish the curve, rather than quoting the encoder resolution.
The encoder is honest about the motor and silent about the tool.
Curated: · Written: · Reviewed:
QA-14A newly fitted tool causes sag on some joints and not others, and the numbers in the payload dialog came from CAD. What would you do?(show answer)
I would start gravity compensation and payload identification from the measured envelope, not from the number the datasheet promises.
Gravity compensation needs mass, centre of mass, and inertia. CAD gives the first accurately and the second only if every cable, fastener, and workpiece is modelled, which is rarely true for a tool as built.
Concretely, run the controller's payload identification routine with the tool as fitted, including cabling, and repeat it for the loaded and unloaded cases if the part mass is a significant fraction of the tool mass. Record the identified values with the program.
The reason for that specificity is a failure I have seen: A CAD payload listed 3.2 kg at 60 mm from the flange; the tool as built with cabling and a workpiece was 4.1 kg at 104 mm, and joint 5 sagged 0.6 mm at full extension while the near-flange poses looked perfect.
CAD versus identified payload.
| Quantity | CAD | Identified | Sag at reach |
|---|---|---|---|
| mass | 3.2 kg | 4.1 kg | — |
| CoM offset | 60 mm | 104 mm | — |
| joint 5 error | — | — | 0.6 mm |
I would not consider it settled without evidence: Compare the identified payload against the CAD figure and require an explicit sign-off when they differ by more than a declared margin.
Compensation corrects the payload you declared, not the one you bolted on.
Curated: · Written: · Reviewed:
QA-15What do you do before the first automatic-mode enable on a new axis?(show answer)
This is an area where a clean simulation and a correct handling of feedback sign and direction checks are not the same event.
Closed-loop control assumes that a positive error produces actuation that reduces it. A reversed sensor, motor, or configuration turns the same loop into positive feedback, and it runs away at whatever rate the actuator can manage.
Concretely, verify sign at the lowest energy the system permits: move the axis by hand or at minimum torque, confirm the measured value moves the expected way, then confirm the actuator drives the expected way, with limits and protective stop independently armed throughout.
The reason for that specificity is a failure I have seen: A rebuilt axis had its encoder cable swapped during maintenance; the first automatic enable drove it to the hard stop in under 200 ms and bent the mechanism, with the controller reporting a following error only after the impact.
Time to hard stop at first enable.
| Condition | Loop behaviour | Time to limit |
|---|---|---|
| correct sign | converges | n/a |
| reversed encoder | diverges | 0.2 s |
| reversed motor | diverges | 0.2 s |
I would not consider it settled without evidence: Make the sign check an explicit, recorded commissioning step with a pass criterion, not a step somebody remembers to do.
Positive feedback is the fastest thing a servo does.
Curated: · Written: · Reviewed:
QA-16A temperature loop overshoots badly after a long cold start but behaves near setpoint. What is wrong and how would you fix it?(show answer)
My answer to integral windup under saturation begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
While the actuator is saturated the loop is open, but the integrator keeps accumulating error as though its output still mattered. The stored term then has to be unwound before the actuator leaves saturation, which shows up as overshoot proportional to how long the saturation lasted.
Concretely, stop or unwind the integrator during saturation by clamping, conditional integration, or back-calculation from the difference between the requested and delivered output, and make the software limit match the real actuator limit so saturation is visible rather than hidden.
The reason for that specificity is a failure I have seen: An oven loop sat saturated for 380 seconds during warm-up; the integral term reached the equivalent of 140 percent output and drove a 31 degree overshoot past a 200 degree setpoint, which was outside the process window for the parts already inside.
Overshoot against saturation time.
| Saturated for | No anti-windup | With back-calculation |
|---|---|---|
| 30 s | 4°C | 1°C |
| 180 s | 18°C | 1°C |
| 380 s | 31°C | 2°C |
I would not consider it settled without evidence: Log the output before and after limiting alongside the integral term, and test the loop from a cold start rather than only from near setpoint.
An integrator that cannot tell it is saturated is storing a debt.
Curated: · Written: · Reviewed:
QA-17Adding derivative gain makes the actuator chatter. What are your options and what does each cost?(show answer)
I would treat derivative action and measurement noise as a claim about the physical world that has to survive a bench test.
Derivative action responds to rate of change, and measurement noise is mostly rate of change. An unfiltered derivative multiplies the noise by the derivative gain, so the term that was meant to add damping instead adds actuator wear.
Concretely, filter the derivative with a first-order filter whose coefficient bounds the noise gain, take the derivative of the measurement rather than the error so a setpoint step does not produce a spike, and record the filter coefficient with the gains since two otherwise identical tunings behave differently without it.
The reason for that specificity is a failure I have seen: A servo with an unfiltered derivative and 12 bits of encoder noise commanded ±8 percent output at 90 Hz while standing still; the valve logged 2.1 million reversals in a week against a rated 5 million lifetime.
Noise gain against filter coefficient.
| Filter N | Output noise | Reversals per hour |
|---|---|---|
| none | ±8% | 12500 |
| 20 | ±1.6% | 900 |
| 8 | ±0.7% | 210 |
I would not consider it settled without evidence: Measure output noise amplitude and reversal count at steady state, not only step response, before accepting a derivative term.
Derivative gain amplifies whatever the sensor is doing, including nothing.
Curated: · Written: · Reviewed:
QA-18A loop has 400 ms of transport delay and a 600 ms dominant time constant. What tuning do you expect to achieve, and when do you change structure?(show answer)
The useful question for dead time and achievable bandwidth is what still holds at the edge of the workspace, cold, and under full payload.
Dead time bounds what feedback can do, because every correction arrives after the plant has already moved. Once the delay approaches the dominant time constant, raising gain buys oscillation rather than speed.
Concretely, measure the delay rather than inferring it from the tune, since a network hop, a filter, and a slow front end all present as dead time and only some are removable. Beyond a ratio of roughly one, move to a Smith predictor or model predictive control with an explicit delay model, or remove the delay at its source.
The reason for that specificity is a failure I have seen: A blending loop with a 0.67 delay-to-constant ratio was tuned aggressively to hit a settling target; it oscillated with a 2.4 second period and the operators put it in manual, which is where it stayed for a year.
Delay ratio and what is achievable.
| θ/τ | Practical settling | Structure |
|---|---|---|
| 0.1 | 3τ | PID |
| 0.67 | 12τ | PID marginal |
| 1.5 | oscillatory | predictor needed |
I would not consider it settled without evidence: Report the measured delay-to-time-constant ratio in the tuning record, and require a structure decision above a declared threshold rather than a gain change.
You cannot tune your way out of the speed of the process.
Curated: · Written: · Reviewed:
QA-19Switching from manual to automatic steps the output by 20 percent. What was not initialised?(show answer)
I would settle bumpless transfer between modes against an instrumented run on real hardware before trusting the model.
A controller carries state. Entering automatic with an integral term that does not correspond to the output currently being delivered produces a step the size of the mismatch, which the process feels immediately.
Concretely, track the actual delivered output while in manual and initialise the integral so that the first automatic output equals it, then let the loop move from there. Do the same for cascade, override, and restart, and test each transition rather than assuming symmetry.
The reason for that specificity is a failure I have seen: A cascade inner loop initialised its integral to zero on transfer; the output stepped from 62 percent to 41 percent in one cycle, the flow dropped, and the outer loop responded to a disturbance that the transfer had created.
Output step at transfer.
| Transition | Uninitialised | Tracked |
|---|---|---|
| manual to auto | 21% | 0.2% |
| auto to cascade | 14% | 0.1% |
| restart | 62% | 0.3% |
I would not consider it settled without evidence: Record output either side of every mode change and assert the step is below a declared threshold for each transition in both directions.
A transfer that the process can feel is not a transfer, it is a disturbance.
Curated: · Written: · Reviewed:
QA-20A conveyor-tracking robot lags a moving part by a constant distance. Would you raise the gain?(show answer)
The judgement in feedforward versus feedback roles is which timestamps and frames are pinned, not which library is fashionable.
Feedback acts on error that has already appeared, so tracking a ramp with proportional-plus-integral action leaves a lag set by the loop gain. Feedforward acts on the known command or measured disturbance before the error develops, which is the only way to remove that lag without pushing the loop toward instability.
Concretely, add a velocity feedforward term derived from the commanded or measured conveyor speed, keep feedback for the model mismatch it is good at, and tune the two separately so an error in the feedforward model does not masquerade as a gain problem.
The reason for that specificity is a failure I have seen: A tracking application raised proportional gain from 12 to 40 to close a 6 mm lag; the lag fell to 2 mm and the axis began oscillating at 14 Hz whenever the conveyor jogged, which produced worse placement than the original lag.
Lag against conveyor speed.
| Speed | Feedback only | With feedforward |
|---|---|---|
| 0.1 m/s | 2 mm | 0.3 mm |
| 0.3 m/s | 6 mm | 0.4 mm |
| 0.5 m/s | 10 mm | 0.4 mm |
I would not consider it settled without evidence: Measure tracking error against conveyor speed with and without the feedforward term and confirm the residual is speed-independent.
Feedback corrects what has gone wrong; feedforward prevents what is about to.
Curated: · Written: · Reviewed:
QA-21When does a cascade loop help, and what makes one actively worse than a single loop?(show answer)
Where candidates lose the interview on cascade control structure is treating a replayed bag as evidence about the robot.
Cascade puts a fast inner loop inside a slow outer one so that disturbances entering the inner process are corrected before the outer variable moves. The benefit exists only if the inner loop is substantially faster than the outer and is itself well tuned.
Concretely, require the inner loop to settle several times faster than the outer, tune inner first and then outer, and give the inner loop its own limits and anti-windup so the outer cannot drive it into a saturated state it cannot report.
The reason for that specificity is a failure I have seen: A cascade with an inner loop only 1.4 times faster than the outer produced a 0.8 Hz interaction that neither loop caused alone; operators tuned the outer down until the cascade was slower than the original single loop.
Speed ratio and outcome.
| Inner/outer speed | Disturbance rejection | Stability |
|---|---|---|
| 1.4x | worse than single | oscillates |
| 5x | 3x better | stable |
| 12x | 4x better | stable |
I would not consider it settled without evidence: Measure both loop bandwidths and require a declared separation ratio before commissioning the cascade.
A cascade is a bet on separation of timescales, so measure the separation.
Curated: · Written: · Reviewed:
QA-22A loop is well behaved at low flow and unstable at high flow. What would you do beyond retuning?(show answer)
I would answer gain scheduling across operating regions by separating what the sensor measured from what the visualiser drew.
A single set of gains assumes one linear plant. When the process gain varies substantially across the operating range, a tune that is correct in one region is either sluggish or unstable in another, and no single compromise fixes both.
Concretely, define regions with measured plant gain, schedule gains against a reliable scheduling variable, and interpolate between regions so the transition is bumpless rather than a step. Validate at the boundaries, which is where scheduling errors appear.
The reason for that specificity is a failure I have seen: A valve loop had 4.2 times the process gain at 80 percent open than at 20 percent; gains tuned at low flow oscillated at high flow, and gains tuned at high flow took 40 seconds to respond at low flow.
Process gain across the range.
| Valve position | Process gain | Single tune result |
|---|---|---|
| 20% | 0.35 | sluggish |
| 50% | 0.9 | good |
| 80% | 1.47 | oscillates |
I would not consider it settled without evidence: Measure open-loop process gain at several operating points and publish the curve before choosing between one tune and a schedule.
One set of gains is a claim that the plant is the same everywhere.
Curated: · Written: · Reviewed:
QA-23A loop tuned on a 100 Hz controller behaves differently after a move to 500 Hz hardware with the same gains. Why?(show answer)
The engineering content of sample rate and controller portability is the tolerance budget and the stop condition, not the algorithm name.
Integral and derivative terms are defined against time, and an implementation that accumulates per cycle rather than per elapsed second silently rescales them when the rate changes. The same numbers then mean five different things.
Concretely, scale integral and derivative by the measured elapsed interval rather than assuming a nominal one, record the design sample rate with the gains, and detect a missed deadline rather than integrating the same stale sample as though it were new.
The reason for that specificity is a failure I have seen: A port from 100 Hz to 500 Hz kept per-cycle accumulation; the effective integral gain rose fivefold and the axis overshot 22 percent on a move that had previously overshot 3 percent.
Same gains, two rates.
| Rate | Effective Ki | Overshoot |
|---|---|---|
| 100 Hz | 1.0x | 3% |
| 500 Hz, per cycle | 5.0x | 22% |
| 500 Hz, per second | 1.0x | 3% |
I would not consider it settled without evidence: Run the same step at two sample rates and require the response to match within a declared tolerance before accepting the port.
A gain without a time base is a number, not a tuning.
Curated: · Written: · Reviewed:
QA-24How much does 10 ms of unaccounted camera latency cost a robot moving at 2 m/s, and what do you do about it?(show answer)
Before letting it move at speed I would write down what a wrong result for sensor latency in a moving frame looks like on the log.
Latency between the instant a sensor sampled and the instant its value is used converts directly into position error through the platform's own motion. It is systematic rather than random, so filtering does not remove it.
Concretely, stamp at the earliest honest point — a capture interrupt or a device-side clock — carry the uncertainty of that stamp, and transform measurements at their own sample time rather than at arrival time. Where the delay is fixed, compensate it explicitly.
The reason for that specificity is a failure I have seen: A mobile manipulator used arrival time for a camera running 34 ms behind capture; at 2 m/s the detected pallet edge sat 68 mm ahead of where the robot believed it was, and the fork entered 68 mm high.
Latency as distance at 2 m/s.
| Latency | Position error | Consequence |
|---|---|---|
| 5 ms | 10 mm | within tolerance |
| 34 ms | 68 mm | mis-pick |
| 100 ms | 200 mm | collision risk |
I would not consider it settled without evidence: Measure end-to-end latency with a physical event visible to both the sensor and an independent trigger, rather than trusting the driver's stated figure.
Latency in a moving frame is distance, and distance is the thing you were measuring.
Curated: · Written: · Reviewed:
QA-25Two sensors both report good data and the fused result is wrong. How do you establish whether time is the problem?(show answer)
The first thing I would pin down about time synchronisation across devices is what the machine physically does when the assumption is wrong.
Each device has its own clock, its own sample instant, and its own transport delay. Fusing the newest sample from each produces an estimate of a moment that never existed, and the error scales with relative motion rather than with sensor noise.
Concretely, discipline clocks to a common source where the hardware allows it, resample deliberately onto a shared time base instead of pairing whatever is newest, and log per-sensor stamp and arrival so skew is measurable after the fact rather than a hypothesis.
The reason for that specificity is a failure I have seen: A lidar and an IMU differed by 18 ms of clock offset that nobody had measured; the fused attitude showed a 1.1 degree oscillation at the turn rate of the vehicle, which the team spent three weeks attributing to IMU bias.
Clock offset and attitude error.
| Offset | Attitude oscillation | Attributed to |
|---|---|---|
| 0 ms | 0.05° | noise |
| 18 ms | 1.1° | wrongly, IMU bias |
| 40 ms | 2.4° | divergence |
I would not consider it settled without evidence: Inject a physical event both sensors observe — a sharp motion or a flash — and measure the offset directly rather than assuming the network synchronised them.
Two correct sensors on two clocks make one wrong estimate.
Curated: · Written: · Reviewed:
QA-26Which sensor failures does a range check catch, and which does it miss?(show answer)
I would start sensor fault detection from the measured envelope, not from the number the datasheet promises.
The dangerous sensor faults produce plausible values. A disconnected thermocouple reads ambient, a transducer that has lost excitation reads mid-scale, and a stuck channel repeats its last value, all of which sit comfortably inside any range check.
Concretely, place the operating band inside the sensing band so open and short circuit fall outside it and are distinguishable from real values, check rate of change against what physics permits, watch for a signal that has gone implausibly quiet since real measurements carry noise, and cross-check against an independent device where consequence justifies it.
The reason for that specificity is a failure I have seen: A stuck ADC channel held 42.6 degrees for 90 minutes while the vessel reached 71 degrees; every range, rate, and alarm check passed because the value never moved and never left the band.
What each check catches.
| Fault | Range check | Rate check | Variance check |
|---|---|---|---|
| open circuit | yes | no | yes |
| stuck value | no | no | yes |
| slow drift | no | no | no |
I would not consider it settled without evidence: Test each declared fault by inducing it physically — unplug, short, freeze — and confirm the diagnostic fires rather than reasoning that it would.
A plausible number from a dead sensor is the worst reading you can get.
Curated: · Written: · Reviewed:
QA-27A vibration signal shows a 3 Hz component that no mechanism can produce. What happened?(show answer)
This is an area where a clean simulation and a correct handling of anti-alias filtering before sampling are not the same event.
Content above half the sample rate does not disappear when you sample; it folds into the band you are looking at and becomes indistinguishable from a real low-frequency signal. Digital filtering after conversion cannot separate them because the information is already gone.
Concretely, put an analogue filter ahead of the converter sized for the actual sample rate, sample fast enough for the genuine signal bandwidth, and verify with a swept input that out-of-band content is attenuated rather than folded.
The reason for that specificity is a failure I have seen: A 100 Hz accelerometer channel with no analogue filter aliased a 103 Hz bearing tone down to 3 Hz; a condition-monitoring alarm fired on a mechanism whose slowest natural frequency was 40 Hz.
Where 103 Hz lands at 100 Hz sampling.
| Input | Nyquist | Apparent frequency |
|---|---|---|
| 40 Hz | 50 Hz | 40 Hz |
| 103 Hz | 50 Hz | 3 Hz |
| 197 Hz | 50 Hz | 3 Hz |
I would not consider it settled without evidence: Sweep a known input across the band and above it, and confirm the recorded spectrum shows attenuation rather than mirrored content.
Aliasing is not noise, it is a wrong answer that looks like a signal.
Curated: · Written: · Reviewed:
QA-28Would moving from a 17-bit to a 23-bit encoder improve your placement accuracy?(show answer)
My answer to encoder resolution versus achievable precision begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
Resolution is the smallest change the device can represent. Accuracy is limited by whatever dominates the error budget, which on most machines is thermal growth, compliance, and kinematic parameter error rather than quantisation.
Concretely, build the error budget before buying resolution: list quantisation, bearing runout, thermal expansion, deflection under load, and calibration residual with measured magnitudes, and spend where the largest term is.
The reason for that specificity is a failure I have seen: An upgrade from 17-bit to 23-bit encoders cost a shift of downtime and moved placement error from 0.31 mm to 0.30 mm, because thermal growth across the 1.4 m frame contributed 0.24 mm and nobody had measured it.
Error budget at 1.4 m reach.
| Term | Contribution |
|---|---|
| thermal growth | 0.24 mm |
| compliance | 0.14 mm |
| kinematic residual | 0.11 mm |
| 17-bit quantisation | 0.01 mm |
I would not consider it settled without evidence: Publish the measured error budget with each term's magnitude before approving a resolution upgrade.
Resolution is cheap to quote and rarely the binding constraint.
Curated: · Written: · Reviewed:
QA-29Why does dead reckoning from an IMU alone fail, and how fast?(show answer)
I would treat IMU bias and integration drift as a claim about the physical world that has to survive a bench test.
Accelerometer bias integrates twice into position, so a constant offset becomes an error growing with the square of time. Gyro bias integrates once into attitude, which then tilts the gravity projection and feeds the accelerometer error, so the two compound.
Concretely, estimate bias online with an aiding source — wheel odometry, GNSS, a visual fix, or a zero-velocity update when the platform is known to be stationary — and treat unaided integration as usable only across the interval where the growing error stays inside the tolerance.
The reason for that specificity is a failure I have seen: A 0.02 m/s² residual accelerometer bias produced 9.0 m of position error after 30 seconds of tunnel travel without aiding, which put the vehicle in the wrong lane on re-acquisition.
Unaided drift from 0.02 m/s² bias.
| Elapsed | Position error |
|---|---|
| 5 s | 0.25 m |
| 15 s | 2.3 m |
| 30 s | 9.0 m |
I would not consider it settled without evidence: Measure position error against unaided duration on the actual device and publish the curve that says how long the estimate is usable.
An IMU tells you about change, and change integrates.
Curated: · Written: · Reviewed:
QA-30Wheel odometry is accurate on the test floor and drifts badly in the warehouse. What changed?(show answer)
The useful question for wheel odometry error sources is what still holds at the edge of the workspace, cold, and under full payload.
Odometry converts wheel rotation into displacement through an assumed rolling radius and track width. Load, tyre pressure, surface, and slip all change the effective radius, so the calibration is a property of the conditions it was measured in.
Concretely, calibrate over the surfaces and loads the robot actually sees, estimate the systematic scale and heading error separately with a bidirectional square-path test, and treat slip as a detectable event rather than as noise by comparing against an independent heading source.
The reason for that specificity is a failure I have seen: A calibration taken unloaded on epoxy floor gave a 1.4 percent scale error on painted concrete under a 400 kg load; over a 60 m aisle that was 840 mm of along-track error, enough to miss the rack entirely.
Scale error by condition.
| Condition | Scale error | Error over 60 m |
|---|---|---|
| unloaded, epoxy | 0.1% | 60 mm |
| loaded, epoxy | 0.6% | 360 mm |
| loaded, concrete | 1.4% | 840 mm |
I would not consider it settled without evidence: Run a bidirectional square-path test under each load and surface combination and record the scale and heading errors separately.
Odometry measures wheels, and the floor is not part of the calibration unless you put it there.
Curated: · Written: · Reviewed:
QA-31A lidar-based safety function misses a dark object at a range where it detects a white one. Is that a fault?(show answer)
I would settle lidar intensity and material response against an instrumented run on real hardware before trusting the model.
A time-of-flight sensor detects returned energy, and returned energy depends on target reflectivity. Detection range is therefore a function of the target, and a single range figure without a reflectivity qualifier describes one material.
Concretely, specify detection range against the lowest reflectivity that matters — safety standards commonly reference a low-reflectivity test body for exactly this reason — verify with physical targets across the reflectivity range, and never derive a safety distance from a figure measured on a cooperative surface.
The reason for that specificity is a failure I have seen: A safety scanner rated to 4 m detected a 90 percent reflective panel at 4.1 m and a 1.8 percent reflective matt black test body at 2.3 m; the protective field had been laid out on the 4 m number.
Detection range against reflectivity.
| Target reflectivity | Detection range |
|---|---|
| 90% | 4.1 m |
| 18% | 3.4 m |
| 1.8% | 2.3 m |
I would not consider it settled without evidence: Measure detection range with the standard low-reflectivity test body across the field, not with whatever object was to hand.
Range is a property of the pair, not of the sensor.
Curated: · Written: · Reviewed:
QA-32Why does a stereo depth camera get worse with distance faster than a time-of-flight one, and what does that mean for bin picking?(show answer)
The judgement in depth camera error models is which timestamps and frames are pinned, not which library is fashionable.
Stereo depth is recovered from disparity, and the same pixel of disparity error corresponds to a depth error that grows with the square of range. Time-of-flight error is dominated by timing and signal-to-noise, which degrades more gently with range but suffers from multipath on concave geometry.
Concretely, pick the sensor against the geometry and range of the task, state depth uncertainty as a function of range rather than as a single figure, and validate on the actual parts including the shiny and concave ones that break each technology differently.
The reason for that specificity is a failure I have seen: A bin-picking cell specified on a 2 mm depth accuracy figure quoted at 0.5 m was mounted at 1.4 m; the stereo error at that range was 15 mm, which exceeded the grasp tolerance and produced a 12 percent pick failure rate.
Stereo depth error against range.
| Range | Depth error | Grasp tolerance |
|---|---|---|
| 0.5 m | 2 mm | 5 mm |
| 1.0 m | 8 mm | 5 mm |
| 1.4 m | 15 mm | 5 mm |
I would not consider it settled without evidence: Measure depth error against range on the actual parts and publish the curve alongside the mounting height decision.
A depth accuracy figure without a range is half a specification.
Curated: · Written: · Reviewed:
QA-33A checkerboard calibration reports 0.2 pixel reprojection error and vision-guided placement is still poor at the edges. What went wrong?(show answer)
Where candidates lose the interview on camera intrinsic calibration quality is treating a replayed bag as evidence about the robot.
Reprojection error measures how well the model fits the images it was fitted to. A set collected near the image centre and at one depth constrains the distortion parameters weakly, so a low residual can coexist with large error where the set had no data.
Concretely, cover the whole image area including corners, vary target distance and tilt across the working range, hold out images the fit never saw, and report error as a map across the frame rather than as one number.
The reason for that specificity is a failure I have seen: A calibration set that filled the central third of the frame produced 0.2 pixel residual and 3.1 pixels of undistortion error at the corners, which at the working distance was 2.4 mm of placement error on parts presented near the edge of the field.
Error by image region.
| Region | Fitted residual | Holdout error |
|---|---|---|
| centre | 0.18 px | 0.2 px |
| mid | 0.21 px | 0.9 px |
| corner | 0.24 px | 3.1 px |
I would not consider it settled without evidence: Publish reprojection error as a heat map over the image with the holdout set marked, rather than a single scalar.
A residual measures the fit, and the fit only knows where you pointed the target.
Curated: · Written: · Reviewed:
QA-34Detection accuracy falls when line speed rises even though the model is unchanged. What is the first thing to check?(show answer)
I would answer exposure and motion blur in inspection by separating what the sensor measured from what the visualiser drew.
Exposure time and platform speed set the blur length in pixels. A model trained on sharp images sees a different distribution once blur exceeds a small fraction of the feature size, and no amount of retraining on the sharp set recovers it.
Concretely, compute blur as speed times exposure divided by the pixel footprint, size lighting so exposure can be short enough to keep blur below a declared fraction of the smallest feature, and validate the detector on images captured at line speed rather than on a stationary rig.
The reason for that specificity is a failure I have seen: A line moved from 0.4 to 1.1 m/s while exposure stayed at 4 ms; blur grew from 3.2 to 8.8 pixels against a 12 pixel defect, and detection recall fell from 0.97 to 0.71 with no change to the model.
Blur and recall against line speed.
| Speed | Blur | Recall |
|---|---|---|
| 0.4 m/s | 3.2 px | 0.97 |
| 0.7 m/s | 5.6 px | 0.88 |
| 1.1 m/s | 8.8 px | 0.71 |
I would not consider it settled without evidence: Capture the validation set at production line speed and report blur in pixels alongside recall.
The camera and the conveyor are one system, so tune them together.
Curated: · Written: · Reviewed:
QA-35A SLAM map looks locally clean and has a duplicated corridor. What does that tell you?(show answer)
The engineering content of SLAM loop closure and map consistency is the tolerance budget and the stop condition, not the algorithm name.
Local odometry is accurate over short intervals and accumulates error over long ones. Loop closure is what converts a locally consistent trajectory into a globally consistent map, and without an accepted closure the same place appears twice.
Concretely, instrument closure rate and rejection reason, tune place recognition against the environment's actual self-similarity, and treat a map with duplicated structure as unusable rather than as cosmetically imperfect, because the planner will route through the duplicate.
The reason for that specificity is a failure I have seen: A warehouse map closed no loops in a 140 m aisle because every bay looked alike to the descriptor; the corridor appeared twice, 2.1 m apart, and the planner routed a robot into a rack it believed was free space.
Closure outcome in a self-similar aisle.
| Descriptor | Closures found | Duplicate offset |
|---|---|---|
| geometric only | 0 | 2.1 m |
| geometric + intensity | 4 | 0.1 m |
| with fiducials | 9 | 0.02 m |
I would not consider it settled without evidence: Report loop-closure count and rejection reasons per run, and check the map against surveyed control points rather than against how it looks.
Locally clean and globally wrong is the normal failure of odometry.
Curated: · Written: · Reviewed:
QA-36Should a robot keep driving when its localisation covariance grows? What would you do with the number?(show answer)
Before letting it move at speed I would write down what a wrong result for localisation covariance as a gate looks like on the log.
A pose estimate without an uncertainty is not usable for a safety or speed decision. Covariance is the estimator's own statement about how much it should be trusted, and it is the natural input to a behaviour gate.
Concretely, define speed and task limits as a function of covariance, degrade rather than stop where that is safe, require re-localisation past a declared bound, and validate the covariance itself against measured error rather than assuming the filter is calibrated.
The reason for that specificity is a failure I have seen: A robot drove at full speed with a covariance corresponding to 0.8 m of one-sigma uncertainty because nothing consumed the number; it clipped a rack leg that the map said was 0.6 m away.
Speed gate against localisation uncertainty.
| 1σ uncertainty | Permitted speed |
|---|---|
| < 0.10 m | 1.5 m/s |
| 0.10-0.35 m | 0.6 m/s |
| > 0.35 m | re-localise |
I would not consider it settled without evidence: Compare reported covariance against measured error over many runs and confirm the filter is neither optimistic nor useless before gating on it.
An uncertainty nobody reads is a comment, not a safeguard.
Curated: · Written: · Reviewed:
QA-37A map that worked for six months starts producing localisation failures. Nothing on the robot changed.(show answer)
The first thing I would pin down about map lifecycle and environment change is what the machine physically does when the assumption is wrong.
A map is a snapshot of an environment that keeps moving. Racking changes, seasonal light, new fixtures, and stacked pallets all alter the features localisation depends on, so map validity has a lifetime.
Concretely, monitor match rate and inlier count against the map as a leading indicator, define an update process with a review step rather than continuous silent adaptation, and keep a versioned map with the ability to roll back when an update makes things worse.
The reason for that specificity is a failure I have seen: A racking reconfiguration moved 40 percent of the observable structure in one aisle; inlier count fell from 780 to 210 over two weeks while localisation still nominally succeeded, then failed outright and stopped the shift.
Inlier count before the failure.
| Week | Inliers | Localisation |
|---|---|---|
| 0 | 780 | fine |
| 1 | 460 | fine |
| 2 | 210 | fails |
I would not consider it settled without evidence: Trend inlier count and match residual per zone and alarm on the trend rather than waiting for the failure.
The map ages even when the software does not.
Curated: · Written: · Reviewed:
QA-38A safety-rated scanner covers the approach and a person is still struck by the load. How is that possible?(show answer)
I would start sensor field of view and occlusion from the measured envelope, not from the number the datasheet promises.
A protective field covers the volume the sensor can see. A load carried ahead of the sensor, a low object below the scan plane, or a person behind a pallet occupies space the field never observed, and the sensor reports clear because it is clear where it looked.
Concretely, analyse the swept volume of the robot including its load against the observed volume, add sensing or restrict motion where they differ, and treat a single scan plane as covering a plane rather than a space.
The reason for that specificity is a failure I have seen: A tugger scanned at 200 mm above floor level and carried a 1.2 m overhanging load; a person seated on a low pallet at 400 mm was outside the scan plane and inside the swept volume of the load.
Covered versus swept volume.
| Height | Scanned | Swept by load |
|---|---|---|
| 200 mm | yes | yes |
| 400 mm | no | yes |
| 1200 mm | no | yes |
I would not consider it settled without evidence: Overlay the measured protective field on the swept volume of the robot with its largest load and show the uncovered region explicitly.
A clear field means the sensor saw nothing, not that there is nothing.
Curated: · Written: · Reviewed:
QA-39Lidar says the aisle is clear and the camera says it is blocked. What should the robot do?(show answer)
This is an area where a clean simulation and a correct handling of multi-sensor disagreement policy are not the same event.
Fusion has to define what happens when sources disagree, because disagreement is information rather than a fault. Silently preferring one sensor makes the other decorative and hides the very cases fusion was added for.
Concretely, write the arbitration rule explicitly against consequence: take the conservative reading for safety-relevant decisions, log every disagreement with both raw observations, and treat a rising disagreement rate as a maintenance signal rather than as noise.
The reason for that specificity is a failure I have seen: A stack preferred lidar whenever the two disagreed; a fogged camera lens went unnoticed for eleven days because its disagreements were discarded rather than counted, and the redundancy the design claimed had not existed for over a week.
Disagreement rate as a fault signal.
| Day | Disagreements/hour | Camera state |
|---|---|---|
| 0 | 2 | clean |
| 5 | 31 | filming |
| 11 | 140 | fogged |
I would not consider it settled without evidence: Count and trend disagreements per sensor pair, and alarm when the rate departs from its established baseline.
Redundancy that resolves silently is not redundancy.
Curated: · Written: · Reviewed:
QA-40A control loop averages 400 µs on a general-purpose kernel and occasionally misses its 1 ms deadline. Why is the average not the point?(show answer)
My answer to real-time determinism versus average speed begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
A real-time system is judged on its worst case, not its mean. A loop that meets its deadline 99.9 percent of the time misses it roughly once a second at 1 kHz, and the controller has no way to make up the lost interval.
Concretely, measure the distribution including the tail, isolate the control thread from the sources of jitter — scheduling, interrupts, power management, shared cache — and detect a missed deadline explicitly so the loop can respond rather than integrating a stale sample as new.
The reason for that specificity is a failure I have seen: A 1 kHz loop with a 400 µs mean showed a 2.3 ms 99.99th percentile caused by a power-management transition; the axis logged a following-error spike every few minutes that the mean execution time never revealed.
Execution time distribution at 1 kHz.
| Percentile | Duration | Deadline |
|---|---|---|
| 50th | 0.40 ms | met |
| 99.9th | 0.95 ms | met |
| 99.99th | 2.30 ms | missed |
I would not consider it settled without evidence: Publish the execution-time distribution to the 99.99th percentile under production load, not the average under an idle system.
Real-time is a promise about the worst cycle.
Curated: · Written: · Reviewed:
QA-41A high-priority control task misses deadlines only when a logging task runs. Explain the mechanism.(show answer)
I would treat priority inversion in a control stack as a claim about the physical world that has to survive a bench test.
When a high-priority task waits on a lock held by a low-priority one, it runs at the low task's effective priority. A medium-priority task can then preempt the lock holder and delay the high-priority task indefinitely, without either of them appearing to be at fault.
Concretely, use priority inheritance or a ceiling protocol on any lock a real-time task can take, keep shared state out of the control path where possible with lock-free handoff, and bound the critical section rather than assuming it is short.
The reason for that specificity is a failure I have seen: A control thread blocked 3.1 ms on a mutex held by a logging thread that had been preempted by a network handler; the 1 ms loop missed three consecutive deadlines and the axis lurched.
Blocking time by lock protocol.
| Protocol | Worst blocking | Deadlines missed |
|---|---|---|
| plain mutex | 3.1 ms | 3 |
| priority inheritance | 0.2 ms | 0 |
| lock-free handoff | 0.01 ms | 0 |
I would not consider it settled without evidence: Trace lock acquisition with the holder's priority recorded and assert a bound on blocking time under load.
Priority describes what runs, and a lock quietly overrides it.
Curated: · Written: · Reviewed:
QA-42A watchdog fires and the robot restarts into the same fault. What is missing from the design?(show answer)
The useful question for watchdog design is what still holds at the edge of the workspace, cold, and under full payload.
A watchdog proves that something is still executing. It does not prove the system is doing the right thing, and a restart that returns to the same state converts a fault into a loop rather than into recovery.
Concretely, kick the watchdog from a point that only executes when the real work completed, not from a timer, record the state that preceded the reset in non-volatile storage, and define what a repeated reset does — safe state and hold, rather than an eleventh attempt.
The reason for that specificity is a failure I have seen: A watchdog kicked from a timer callback kept a hung control loop alive for 40 minutes because the timer was fine; the axis held its last command throughout.
Kick source and what it detects.
| Kick from | Detects hang | Detects wrong output |
|---|---|---|
| timer callback | no | no |
| loop completion | yes | no |
| completion + plausibility | yes | partly |
I would not consider it settled without evidence: Test the watchdog by hanging the actual work, not by stopping the process, and confirm both the reset and the recorded cause.
A watchdog that a hung system can still feed is decoration.
Curated: · Written: · Reviewed:
QA-43An emergency stop cuts power to everything. When is that the wrong design?(show answer)
I would settle safe state is not always de-energised against an instrumented run on real hardware before trusting the model.
Safe state is determined by hazard analysis, not by a convention. Removing power releases a brake on a vertical axis, stops cooling on a hot process, and drops a magnetic gripper, each of which creates the hazard the stop was meant to prevent.
Concretely, identify the safe state per hazard, keep the energy needed to hold it available through the stop, and verify by inducing the stop under the worst case — loaded, at height, at temperature — rather than at rest.
The reason for that specificity is a failure I have seen: A vertical axis holding a 60 kg fixture used a power-off brake but the stop also cut the servo before the brake engaged; the fixture fell 40 mm during the 120 ms gap and struck the fixture plate.
Stop sequence timing under load.
| Event | Time | Vertical drop |
|---|---|---|
| stop asserted | 0 ms | 0 mm |
| servo disabled | 15 ms | 2 mm |
| brake engaged | 135 ms | 40 mm |
I would not consider it settled without evidence: Measure the actual sequence and timing of the stop under full load, including brake engagement delay, rather than reading the schematic.
De-energised is a state, and it is not automatically the safe one.
Curated: · Written: · Reviewed:
QA-44The control software already limits speed near a person. Why add a separate safety function?(show answer)
The judgement in separating control from protection is which timestamps and frames are pinned, not which library is fashionable.
A protective function must be independent of the thing it protects against. A limit implemented in the same software that computes the motion shares its faults, its update path, and its bugs, so a single defect removes both the motion plan and the protection.
Concretely, implement the protective function in a separate, rated channel with its own sensing and its own logic, keep the normal control limit as well because it prevents nuisance trips, and prove independence rather than asserting it.
The reason for that specificity is a failure I have seen: A speed limit and the trajectory generator shared a units conversion; a change that introduced a degrees-for-radians error raised both the commanded speed and the limit, and the cell ran at 3.2 times the intended speed with no trip.
Common-cause exposure.
| Shared element | Control | Protection | Independent |
|---|---|---|---|
| units conversion | yes | yes | no |
| encoder channel | yes | yes | no |
| rated safe encoder | no | yes | yes |
I would not consider it settled without evidence: Demonstrate the protective function with the control software deliberately faulted, not with it behaving correctly.
Protection that shares a fault with control is not a second layer.
Curated: · Written: · Reviewed:
QA-45A robot is described as collaborative. What does that entitle you to assume about the installation?(show answer)
Where candidates lose the interview on collaborative operation and force limits is treating a replayed bag as evidence about the robot.
Collaborative capability is a property of the application, not of the arm. The permitted force and pressure depend on the body region that can be contacted, the tool geometry, the payload, and the speed, so the same arm is collaborative in one cell and not in another.
Concretely, perform the risk assessment against the actual tool, payload, and reachable body regions, measure contact force and pressure with an instrumented body model rather than computing it, and re-do the assessment when the tool changes.
The reason for that specificity is a failure I have seen: An arm rated for collaborative operation was fitted with a probe whose 12 mm tip contacts over about 113 mm²; at 110 N that is 0.97 N/mm² against a 0.28 N/mm² hand limit, so contact pressure exceeded it 3.5-fold while the assessment had used the 900 mm² bare flange at 0.44 of the limit. Pressure is force divided by contact area, so shrinking the contact multiplies it.
Same force, different tool.
| Tool | Contact area | Pressure at 110 N | vs 0.28 N/mm² limit |
|---|---|---|---|
| bare flange | 900 mm² | 0.12 N/mm² | 0.44x |
| 12 mm probe tip | 113 mm² | 0.97 N/mm² | 3.5x |
| padded plate | 2400 mm² | 0.05 N/mm² | 0.16x |
I would not consider it settled without evidence: Measure force and pressure with the fitted tool on an instrumented body model at the maximum speed the application permits.
Collaborative describes an installation, and the tool is part of it.
Curated: · Written: · Reviewed:
QA-46How does a speed-and-separation scheme decide how fast the robot may move?(show answer)
I would answer speed and separation monitoring by separating what the sensor measured from what the visualiser drew.
The permitted speed follows from the protective separation distance, which is the sum of how far the human can travel while the system reacts, how far the robot travels in its own reaction and stopping time, and allowances for sensor and position uncertainty. It is a distance budget, not a speed setting.
Concretely, measure each term rather than taking it from a manual: human approach speed by standard, system reaction time end to end, robot stopping distance at the relevant speed and payload, and sensor uncertainty at the relevant range. Recompute when payload, speed, or tooling changes.
The reason for that specificity is a failure I have seen: A cell used a stopping distance measured unloaded; with a 12 kg payload the arm needed 310 mm rather than 180 mm to stop, and the separation distance was 130 mm short at every point of the cycle.
Separation budget at 1.2 m/s tool speed.
| Term | Unloaded | 12 kg payload |
|---|---|---|
| human travel | 480 mm | 480 mm |
| robot stopping | 180 mm | 310 mm |
| uncertainty | 120 mm | 120 mm |
| required separation | 780 mm | 910 mm |
I would not consider it settled without evidence: Measure stopping distance physically at the maximum speed and payload combination and use that figure in the budget.
The speed is an output of the distance budget, not an input.
Curated: · Written: · Reviewed:
QA-47A control node subscribes to a sensor topic and occasionally acts on old data. What in the middleware configuration would you look at?(show answer)
The engineering content of publish-subscribe delivery guarantees is the tolerance budget and the stop condition, not the algorithm name.
A publish-subscribe transport buffers. Queue depth, reliability setting, and history policy together decide whether a slow subscriber receives the newest sample or works through a backlog, and the default is usually not what a control loop wants.
Concretely, use a depth-one, keep-last policy on anything a controller acts on so a slow consumer drops old data rather than queuing it, separate high-rate telemetry from control topics, and measure queue occupancy rather than assuming it is zero.
The reason for that specificity is a failure I have seen: A grasp node with a queue depth of 50 fell behind by 0.4 s during a garbage-collection pause and executed a grasp against a pose from before the part moved, closing the gripper on empty air.
Queue depth and message age at use.
| Depth | Age at use | Behaviour under stall |
|---|---|---|
| 1 | 12 ms | drops old |
| 10 | 95 ms | lags |
| 50 | 400 ms | acts on stale |
I would not consider it settled without evidence: Instrument the age of every message at the point of use and alarm on age rather than on message rate.
A deep queue turns a dropped message into a wrong action.
Curated: · Written: · Reviewed:
QA-48Monitoring shows the sensor topic publishing at its full rate and the robot still behaves as though the data is late. What would you measure instead?(show answer)
Before letting it move at speed I would write down what a wrong result for message age versus message rate looks like on the log.
Rate says how often messages arrive, not how old their content is. A pipeline that buffers, batches, or re-publishes can maintain a perfect rate while the content it carries is arbitrarily stale.
Concretely, stamp at capture, measure the difference between that stamp and the moment of use, and monitor that distribution. Keep the rate metric as a liveness check, and make the age metric the one that gates behaviour.
The reason for that specificity is a failure I have seen: A vision pipeline republished at a steady 30 Hz while an intermediate node held a 6-frame buffer; content age at use was 200 ms and the rate dashboard showed nothing unusual for the eleven days it took to find.
Steady rate, growing age.
| Stage | Rate | Content age |
|---|---|---|
| camera | 30 Hz | 0 ms |
| after buffer | 30 Hz | 200 ms |
| at consumer | 30 Hz | 210 ms |
I would not consider it settled without evidence: Plot the age distribution at the consumer alongside the rate, and gate on age.
Arriving on time and being current are different properties.
Curated: · Written: · Reviewed:
QA-49The system works when started by hand and fails intermittently on boot. What is the usual cause?(show answer)
The first thing I would pin down about node lifecycle and startup ordering is what the machine physically does when the assumption is wrong.
Independent processes come up in an order the launcher does not control. A node that reads a parameter, subscribes to a transform, or opens a device before the provider exists will sometimes succeed and sometimes fail, which is exactly the profile of a boot-order bug.
Concretely, give nodes an explicit lifecycle with a configure step that can fail and be retried, declare dependencies rather than sleeping, and make every startup read either wait with a bound or fail loudly rather than proceeding with a default.
The reason for that specificity is a failure I have seen: A calibration node read its extrinsics parameter 40 ms before the parameter server had loaded them and fell back to an identity transform; the cell ran a full shift placing parts against an unrotated camera frame roughly one boot in seven.
Boot outcome by ordering.
| Order | Extrinsics loaded | Result |
|---|---|---|
| params first | yes | correct |
| node first, waits | yes | correct |
| node first, defaults | no | identity frame |
I would not consider it settled without evidence: Force adverse startup ordering in test — start dependencies late deliberately — and require the system to refuse rather than default.
A silent default at startup is a fault that waits for a reboot.
Curated: · Written: · Reviewed:
QA-50Two identical cells behave differently and both are running the same software version. Where do you look?(show answer)
I would start parameter provenance and configuration drift from the measured envelope, not from the number the datasheet promises.
Behaviour is set by code and configuration together. A version number that covers only the code leaves gains, limits, calibration, and tool data free to differ, and those are usually the numbers that were touched on site.
Concretely, version configuration with the code, record the effective values at startup rather than the file contents, and provide a diff between the running configuration and the released baseline so a site change is visible rather than archaeological.
The reason for that specificity is a failure I have seen: Two cells on the same release differed by a 0.6 factor on one axis gain applied by a commissioning engineer eighteen months earlier; the slower cell had been accepted as normal and the difference was found only during a cycle-time study.
Same release, different behaviour.
| Cell | Software | Config hash | Cycle time |
|---|---|---|---|
| A | 4.2.1 | 9c31 | 8.4 s |
| B | 4.2.1 | 7ae0 | 11.9 s |
I would not consider it settled without evidence: Emit the effective configuration hash at startup and compare it across cells rather than comparing software versions.
The running configuration is what determines behaviour, so record that.
Curated: · Written: · Reviewed:
QA-51A robot stopped unexpectedly last week and the logs do not explain it. What should have been recorded?(show answer)
This is an area where a clean simulation and a correct handling of recording for post-incident analysis are not the same event.
A motion incident is reconstructed from what was commanded, what was measured, and what the safety layer saw, at the resolution of the control loop. Application-level logging at one hertz cannot resolve an event that lasted 40 ms.
Concretely, keep a rolling high-rate buffer of commanded and measured joint state, mode, limits, safety inputs, and fault codes, and freeze it on any protective stop. Store the frozen window rather than trying to keep everything.
The reason for that specificity is a failure I have seen: An unexplained protective stop had 1 Hz application logs; the following-error spike that triggered it lasted 40 ms and was invisible in the record, so the cell ran for another three weeks before the loose connector was found.
What each log rate can resolve.
| Rate | Resolves 40 ms spike | Storage per hour |
|---|---|---|
| 1 Hz | no | 0.2 MB |
| 100 Hz | partly | 20 MB |
| 1 kHz ring, frozen | yes | 20 MB retained |
I would not consider it settled without evidence: Trigger a test stop and confirm the frozen buffer contains the preceding window at control-loop resolution.
The evidence has to exist before the incident, not after it.
Curated: · Written: · Reviewed:
QA-52A grasp policy succeeds 98 percent of the time in simulation and 62 percent on hardware. What would you do first?(show answer)
My answer to simulation fidelity and the reality gap begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
A simulator reproduces the physics somebody modelled. Contact, friction, compliance, sensor noise, and latency are the terms usually simplified, and they are exactly the terms a grasp depends on, so the gap is structural rather than a tuning issue.
Concretely, measure the gap per factor rather than in aggregate: replay the same commands on both, compare the divergence, and identify which modelled quantity explains it. Randomise the factors that matter and validate on hardware trials, not on a better simulator score.
The reason for that specificity is a failure I have seen: A policy trained with a friction coefficient fixed at 0.8 met parts whose measured coefficient ranged from 0.35 to 0.9; the failures clustered entirely on the low-friction parts, which the aggregate success rate had hidden.
Success by measured friction.
| Friction | Sim success | Hardware success |
|---|---|---|
| 0.35 | 98% | 21% |
| 0.60 | 98% | 74% |
| 0.90 | 98% | 96% |
I would not consider it settled without evidence: Report hardware success rate broken down by the physical factor suspected, not as a single number.
The gap is a list of specific unmodelled quantities, so find out which one.
Curated: · Written: · Reviewed:
QA-53Would randomising the simulator harder have fixed the grasp policy?(show answer)
I would treat domain randomisation limits as a claim about the physical world that has to survive a bench test.
Randomisation broadens the training distribution over the parameters you chose to vary. It cannot cover a factor that is absent from the model, and widening the wrong parameters buys a more conservative policy rather than a more capable one.
Concretely, choose randomisation ranges from measured hardware variation rather than from intuition, verify that the real system falls inside the sampled range, and add the missing physics where the gap is structural rather than widening what is already modelled.
The reason for that specificity is a failure I have seen: A team widened lighting and texture randomisation by a factor of three while the actual failure was a 30 ms control latency that the simulator did not model at all; success on hardware moved from 62 to 64 percent after two weeks of training.
Where the effort went.
| Factor | Randomised | Present in sim | Explains gap |
|---|---|---|---|
| lighting | wide | yes | no |
| texture | wide | yes | no |
| control latency | no | no | yes |
I would not consider it settled without evidence: Show that measured hardware parameters fall inside the randomised range, per parameter, before attributing a gap to insufficient randomisation.
You cannot randomise over a variable the model does not have.
Curated: · Written: · Reviewed:
QA-54A geometric path is collision-free. What still has to be decided before it can be executed?(show answer)
The useful question for trajectory time parameterisation is what still holds at the edge of the workspace, cold, and under full payload.
A path is a sequence of configurations; a trajectory adds when each is reached. Velocity, acceleration, and jerk limits are applied at that stage, which is where a feasible path becomes an infeasible motion or an unnecessarily slow one.
Concretely, parameterise against the real per-joint limits including any payload dependence, keep jerk bounded so the mechanism is not excited, and validate the parameterised result against the controller's own checker rather than the planner's.
The reason for that specificity is a failure I have seen: A path parameterised without a jerk bound excited a 22 Hz structural mode; the tool oscillated 0.8 mm for 300 ms after each stop and the vision system measured the part before it settled.
Jerk limit against settling.
| Jerk limit | Cycle time | Settling |
|---|---|---|
| none | 6.8 s | 300 ms |
| 8000 °/s³ | 7.1 s | 90 ms |
| 3000 °/s³ | 7.9 s | 25 ms |
I would not consider it settled without evidence: Measure settling time at the end of the motion and require it inside the cycle budget, rather than checking only that the motion completed.
Timing is where a path becomes something the machine can actually do.
Curated: · Written: · Reviewed:
QA-55The planner reports a collision-free path and the robot clips a fixture between waypoints. Explain.(show answer)
I would settle collision checking resolution against an instrumented run on real hardware before trusting the model.
Discrete collision checking samples the path. Between samples the robot is unchecked, and the distance travelled between samples grows with joint speed, so a resolution that is adequate for a slow move is not adequate for a fast one.
Concretely, choose the check resolution from the maximum swept distance rather than from a fixed number of samples, use continuous collision checking where the tolerance demands it, and inflate the collision model by the residual uncertainty rather than checking the nominal geometry.
The reason for that specificity is a failure I have seen: A path checked at 20 samples per segment moved 62 mm of tool travel between checks near the fast portion; a 40 mm fixture sat entirely inside one gap and the check passed.
Unchecked gap against sampling.
| Samples/segment | Max swept gap | 40 mm fixture |
|---|---|---|
| 20 | 62 mm | missed |
| 100 | 12 mm | caught |
| continuous | 0 mm | caught |
I would not consider it settled without evidence: Report the maximum unchecked swept distance for every planned path and gate on it rather than on the sample count.
A discrete check is a statement about the samples, not about the path.
Curated: · Written: · Reviewed:
QA-56The same planning request produces a different path each run. Is that acceptable in production?(show answer)
The judgement in sampling-based planner determinism is which timestamps and frames are pinned, not which library is fashionable.
Sampling-based planners are randomised, so path, length, and planning time all vary between runs. That is acceptable for exploration and awkward for a production cycle whose time budget and swept volume were validated on one particular path.
Concretely, fix the seed and cache the resulting path for repeated production motions, validate the cached path once, and re-plan only when the scene changes. Keep the randomised planner for the cases that genuinely vary.
The reason for that specificity is a failure I have seen: A cell re-planned every cycle; planning time ranged from 40 ms to 1.9 s and one path in roughly two hundred swept close enough to a fixture to trip the protective field, stopping the line for reasons that never reproduced.
Planning variability over 1000 runs.
| Percentile | Planning time | Path length |
|---|---|---|
| 50th | 120 ms | 2.10 m |
| 99th | 940 ms | 2.80 m |
| max | 1.9 s | 3.40 m |
I would not consider it settled without evidence: Record planning time and path length distributions across at least a thousand runs before accepting a randomised planner in a timed cycle.
A path that varies has to be validated every time it varies.
Curated: · Written: · Reviewed:
QA-57What makes a planning scene wrong, and how would you catch it before the robot moves?(show answer)
Where candidates lose the interview on planning scene freshness is treating a replayed bag as evidence about the robot.
The planner reasons about the world it was told about. Objects that have moved, been removed, or never been added are invisible to it, and a stale scene produces a confidently collision-free path through a real obstacle.
Concretely, stamp the scene, refuse to plan against one older than a declared bound, reconcile the scene against live perception before executing rather than only before planning, and treat an object whose observation has lapsed as present rather than absent.
The reason for that specificity is a failure I have seen: A scene retained a pallet that had been removed 90 seconds earlier and omitted a tote placed 8 seconds earlier; the plan avoided empty space and drove through the tote.
Scene versus reality at execution.
| Object | In scene | Present | Planner behaviour |
|---|---|---|---|
| pallet | yes | no | avoids nothing |
| tote | no | yes | drives through |
| fixture | yes | yes | avoids |
I would not consider it settled without evidence: Check scene age and unobserved-object count at execution time and refuse motion rather than logging a warning.
Absence in the scene is not evidence of absence in the cell.
Curated: · Written: · Reviewed:
QA-58A grasp detector returns a high-scoring pose that the arm cannot execute. Whose problem is that?(show answer)
I would answer grasp pose ranking and reachability by separating what the sensor measured from what the visualiser drew.
A grasp score measures the quality of the contact, not whether the arm can reach it without collision, singularity, or a joint-limit violation. Ranking on score alone systematically prefers poses the robot cannot use.
Concretely, filter candidates through reachability, collision, and approach-path feasibility before ranking, and rank on a combined score so a marginally worse but comfortably reachable grasp wins. Log why each rejected candidate failed.
The reason for that specificity is a failure I have seen: A detector's top grasp was reachable in 34 percent of bin positions; the cell retried up to five times per part and cycle time rose from 9 s to 21 s, which read as a perception problem rather than a ranking one.
Score-only versus feasibility-filtered ranking.
| Ranking | Top grasp executable | Mean retries | Cycle |
|---|---|---|---|
| score only | 34% | 2.9 | 21 s |
| feasibility filtered | 91% | 0.3 | 10 s |
I would not consider it settled without evidence: Report the fraction of top-ranked grasps that are executable, per bin region, rather than the detector's own score distribution.
An unreachable grasp scores zero however good the contact is.
Curated: · Written: · Reviewed:
QA-59An assembly task chatters when the part contacts the fixture. What is happening?(show answer)
The engineering content of force control and contact stability is the tolerance budget and the stop condition, not the algorithm name.
Contact couples the robot's controller to the stiffness of the environment. A gain that is stable in free space can be unstable against a stiff surface, because the effective loop gain rises with contact stiffness.
Concretely, reduce gain or add compliance on contact, use an explicit impedance or admittance formulation with a stated target stiffness rather than a position loop pushed into a surface, and validate against the stiffest surface the task can meet.
The reason for that specificity is a failure I have seen: A position-controlled insert against a steel fixture chattered at 60 Hz and marked 400 parts before the loop was switched to admittance control with a 200 N/m target stiffness.
Stability against environment stiffness.
| Surface | Stiffness | Position control | Admittance |
|---|---|---|---|
| foam | 2 kN/m | stable | stable |
| plastic | 40 kN/m | marginal | stable |
| steel | 900 kN/m | chatters | stable |
I would not consider it settled without evidence: Test against the stiffest and softest surfaces in the task and record contact force over time, not just success or failure.
Contact changes the plant, so the controller has to change with it.
Curated: · Written: · Reviewed:
QA-60A wrist force sensor reads 14 N with nothing touching the tool. What do you do before using the number?(show answer)
Before letting it move at speed I would write down what a wrong result for force-torque sensor bias and tool weight looks like on the log.
A wrist sensor measures everything distal to it, which includes the tool's own weight projected through the current orientation, plus sensor bias and temperature drift. The contact force is what remains after those are removed.
Concretely, compensate tool weight using the identified mass and centre of mass at the current orientation, re-zero bias at a known no-contact pose rather than at power-on only, and monitor drift so a slow thermal shift does not become a phantom contact.
The reason for that specificity is a failure I have seen: An uncompensated 1.4 kg tool produced up to 14 N of apparent force as the wrist rotated; a 10 N contact threshold triggered on orientation alone and the robot reported contact in mid-air on every fourth approach.
Apparent force by wrist orientation, no contact.
| Orientation | Raw | Compensated |
|---|---|---|
| tool down | 14.0 N | 0.3 N |
| tool horizontal | 6.8 N | 0.2 N |
| tool up | -13.6 N | 0.4 N |
I would not consider it settled without evidence: Sweep the wrist through its orientation range with no contact and require the compensated reading to stay inside a declared band.
A force reading is a sum, and the contact is only one term.
Curated: · Written: · Reviewed:
QA-61A gripper works on the sample parts and fails on production ones. What did the sample set not represent?(show answer)
The first thing I would pin down about gripper design and part tolerance is what the machine physically does when the assumption is wrong.
A grasp succeeds over a range of part geometry, surface, and presentation. Sample parts are usually from one lot, clean, and presented consistently, which is a much narrower distribution than production.
Concretely, characterise the actual variation — dimensional tolerance, surface finish, contamination, and presentation spread — and design the grasp to span it, with compliance or self-centring geometry rather than tighter positioning.
The reason for that specificity is a failure I have seen: A gripper tuned on parts from a single mould cavity met a second cavity 0.35 mm larger; the fingers bottomed out before closing and the pick rate on that cavity's parts was 41 percent while the overall figure looked acceptable.
Pick rate by mould cavity.
| Cavity | Part width | Pick rate |
|---|---|---|
| 1 | 24.10 mm | 99% |
| 2 | 24.45 mm | 41% |
| pooled | 78% |
I would not consider it settled without evidence: Test across lots, cavities, and contamination states, reporting success per subgroup rather than pooled.
A pooled success rate hides the subgroup that is failing.
Curated: · Written: · Reviewed:
QA-62A cell must hit 8 seconds and runs at 11. How do you decide what to change?(show answer)
I would start cycle time budgeting from the measured envelope, not from the number the datasheet promises.
Cycle time is a sum of specific intervals, several of which overlap. Optimising the one that is easiest to change rather than the one on the critical path buys nothing, and the critical path is rarely the motion everybody watches.
Concretely, instrument each interval, identify what is genuinely serial, and look for overlap first — moving while the vision system processes, pre-positioning during the fixture cycle — before increasing speed, which costs settling time and wear.
The reason for that specificity is a failure I have seen: A team raised joint speeds 20 percent to recover 3 seconds and gained 0.4 s, because 2.6 s of the gap was a serial vision exposure and inference that could have run during the return move.
Where the 11 seconds goes.
| Interval | Duration | Overlappable |
|---|---|---|
| approach | 1.9 s | no |
| vision | 2.6 s | yes |
| grasp | 1.2 s | no |
| transfer | 3.1 s | no |
| place + return | 2.2 s | partly |
I would not consider it settled without evidence: Publish the measured interval breakdown with serial and overlapped portions marked before approving any speed increase.
Speed is the expensive lever and usually not the binding one.
Curated: · Written: · Reviewed:
QA-63Placement accuracy degrades over the first two hours of a shift and then stabilises. What is going on?(show answer)
This is an area where a clean simulation and a correct handling of thermal drift over a shift are not the same event.
Motors, gearboxes, and structure warm during operation and the machine grows. The growth is systematic, roughly proportional to temperature rise and reach, and it stops when the thermal state settles rather than when the software changes.
Concretely, measure drift against temperature rather than against time, warm up before critical work or compensate from a temperature model, and re-reference against a fixed fiducial periodically rather than trusting the morning calibration all day.
The reason for that specificity is a failure I have seen: A cell drifted 0.28 mm over 110 minutes of a shift; the first-off inspection passed and parts made between 40 and 110 minutes drifted outside tolerance, which the daily calibration at start-up guaranteed would recur.
Drift against temperature rise.
| Elapsed | ΔT | Placement error |
|---|---|---|
| 10 min | 2 K | 0.03 mm |
| 60 min | 11 K | 0.19 mm |
| 110 min | 16 K | 0.28 mm |
I would not consider it settled without evidence: Log structure temperature alongside placement error and show the relationship, rather than describing the drift as time-dependent.
The machine is a different size warm than cold.
Curated: · Written: · Reviewed:
QA-64How would you detect a failing gearbox before it stops the line?(show answer)
My answer to wear as a measurable trend begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
Mechanical degradation shows up in quantities the controller already computes long before it shows up as a fault. Motor current for the same commanded motion, following error, and backlash measured at a reference move all trend before the failure.
Concretely, run a short reference motion at fixed speed and payload on a schedule, record current, following error, and measured lost motion, and trend them per axis against that axis's own baseline rather than against a fleet number.
The reason for that specificity is a failure I have seen: A gearbox failed without warning to the operator while its reference-move current had risen 34 percent over six weeks in data nobody trended; the unplanned stop cost eleven hours.
Reference-move signature before failure.
| Week | Peak current | Lost motion |
|---|---|---|
| 0 | 4.2 A | 0.04° |
| 3 | 4.9 A | 0.09° |
| 6 | 5.6 A | 0.21° |
I would not consider it settled without evidence: Trend the reference-move signature per axis and alarm on the slope, not on an absolute threshold.
The controller already knows, if somebody keeps the record.
Curated: · Written: · Reviewed:
QA-65Why run a new program at reduced speed first, and what does that not tell you?(show answer)
I would treat commissioning at reduced speed as a claim about the physical world that has to survive a bench test.
Reduced speed limits the energy of a mistake, which is why it is the right first step. It also changes the dynamics, so path deviation from inertia, settling, and any speed-dependent limit is not exercised at all.
Concretely, step speed up in stages with a check at each, watch following error and path deviation as speed rises rather than only at the end, and re-run collision-critical segments at full speed before release since the swept path differs.
The reason for that specificity is a failure I have seen: A program verified at 20 percent speed deviated 12 mm from the taught path at 100 percent through a fast corner and struck a fixture that had 8 mm of clearance at low speed.
Path deviation against speed.
| Speed | Deviation | Clearance |
|---|---|---|
| 20% | 0.4 mm | 8 mm |
| 60% | 4.1 mm | 4 mm |
| 100% | 12.0 mm | collision |
I would not consider it settled without evidence: Measure path deviation at each speed step and require the full-speed swept path to clear obstacles, not the taught path.
Slow proves the geometry, not the dynamics.
Curated: · Written: · Reviewed:
QA-66A Cartesian move between two reachable poses fails midway. Both endpoints are fine.(show answer)
The useful question for singularity-free path planning in Cartesian space is what still holds at the edge of the workspace, cold, and under full payload.
Reachability is a property of each pose; feasibility is a property of the whole path. A straight line in Cartesian space can pass through a singularity or leave the reachable set even when both endpoints are comfortable.
Concretely, check the path rather than the endpoints: sample conditioning and joint limits along it, and where the direct route fails, insert a via point or plan in joint space, which trades path shape for feasibility.
The reason for that specificity is a failure I have seen: A straight-line move between two comfortable poses passed within 12 mm of a wrist alignment singularity at its midpoint; the controller aborted with a velocity limit error that named a joint neither endpoint stressed.
Conditioning along a straight-line move.
| Fraction of path | σ_min | Status |
|---|---|---|
| 0.0 | 0.28 | fine |
| 0.5 | 0.006 | aborts |
| 1.0 | 0.31 | fine |
I would not consider it settled without evidence: Report minimum conditioning along the path, not at the endpoints, as a release gate for Cartesian moves.
Both ends being fine says nothing about the middle.
Curated: · Written: · Reviewed:
QA-67Why is a state machine preferable to a linear script for a robot task, and what do people get wrong about it?(show answer)
I would settle state machines for task sequencing against an instrumented run on real hardware before trusting the model.
A robot task has to handle interruption, recovery, and resumption from wherever it stopped. A linear script encodes only the successful path, so every error handler becomes a special case rather than a transition.
Concretely, enumerate states including the recovery and abort states, define transitions and their guards explicitly, and make the current state observable and persistent so a restart resumes deliberately rather than beginning again from the top with a part already in the gripper.
The reason for that specificity is a failure I have seen: A script-based cell restarted after a stop and ran its home sequence while still holding a part; the part struck the fixture on the way to home, and the recovery then had to be written as a special case at each of the 11 stop points, 3 of which were missed.
Restart behaviour by design.
| Stopped at | Script restart | State machine | Special cases |
|---|---|---|---|
| approach | homes safely | resumes | 0 |
| holding part | collides | places then homes | 4 |
| in fixture | collides | retracts first | 7 |
I would not consider it settled without evidence: Test recovery from a stop injected at every state, not only at the convenient ones.
Recovery is a set of transitions, so give it states to transition between.
Curated: · Written: · Reviewed:
QA-68A teleoperated arm with force feedback becomes unstable over a network link. Why?(show answer)
The judgement in teleoperation latency and stability is which timestamps and frames are pinned, not which library is fashionable.
Force feedback closes a loop through the operator, the link, and the robot. Delay in that loop reduces phase margin exactly as it does in any feedback system, so a link that is merely slow makes a stable system oscillate.
Concretely, bound the round-trip delay, reduce feedback gain as measured delay rises, and use a formulation designed for delay — wave variables or a passivity observer — rather than tuning against a delay that varies with the network.
The reason for that specificity is a failure I have seen: A haptic link stable at 20 ms round trip oscillated at 4 Hz when the path shifted to 95 ms; the operator read it as a stiff surface and pushed harder, which increased the amplitude.
Stability against round-trip delay.
| Round trip | Feedback gain | Behaviour |
|---|---|---|
| 20 ms | 1.0 | stable |
| 95 ms | 1.0 | 4 Hz oscillation |
| 95 ms | 0.3 | stable, soft |
I would not consider it settled without evidence: Measure round-trip delay continuously and gate feedback gain on it rather than on a nominal figure from commissioning.
Delay in a haptic loop is a stability parameter, not a latency metric.
Curated: · Written: · Reviewed:
QA-69Two robots in a narrow aisle both stop and wait. What design avoids it?(show answer)
Where candidates lose the interview on fleet coordination and deadlock is treating a replayed bag as evidence about the robot.
Independent local avoidance produces deadlock whenever two agents each need space the other occupies. No amount of better local planning resolves it, because the resource conflict is global.
Concretely, allocate contended space as a resource with an explicit owner and an ordering, so one robot proceeds and the other waits deliberately, and detect the wait state rather than letting two robots sit indefinitely with valid local plans.
The reason for that specificity is a failure I have seen: Two units met in a 1.6 m aisle with 1.1 m footprints; each yielded to the other, both stopped, and the aisle was blocked for 26 minutes until an operator intervened, with neither robot reporting a fault.
Aisle width and outcome.
| Aisle | Footprint | Local avoidance | With allocation |
|---|---|---|---|
| 2.6 m | 1.1 m | passes | passes |
| 1.6 m | 1.1 m | deadlock | one waits |
I would not consider it settled without evidence: Detect and alarm on mutual wait rather than relying on either robot to notice, and test the narrow-aisle case explicitly.
Deadlock is a resource problem wearing a navigation costume.
Curated: · Written: · Reviewed:
QA-70A mobile robot reports 30 percent charge and shuts down twenty minutes later than its estimate, or five minutes earlier. Which is worse and why?(show answer)
I would answer battery state of charge under load by separating what the sensor measured from what the visualiser drew.
State of charge inferred from voltage is load-dependent, because internal resistance drops the terminal voltage under current. The same cell reads lower while accelerating and recovers at rest, so the estimate moves with duty rather than with energy.
Concretely, estimate with coulomb counting corrected by a rested-voltage reference and a temperature model, and reserve margin against the worst-case remaining task rather than against average consumption.
The reason for that specificity is a failure I have seen: A robot estimating from loaded voltage reported 30 percent on a ramp and 44 percent at rest; it was dispatched on a task needing 35 percent, stopped mid-aisle, and blocked the route for forty minutes.
Reported charge by condition.
| Condition | Terminal voltage | Reported | Actual |
|---|---|---|---|
| at rest | 26.1 V | 44% | 38% |
| on ramp | 24.6 V | 30% | 37% |
| coulomb counted | — | 37% | 38% |
I would not consider it settled without evidence: Compare the estimate against measured energy to shutdown across duty cycles and report the error distribution rather than a single accuracy figure.
A charge estimate that moves with load is measuring the load.
Curated: · Written: · Reviewed:
QA-71An analogue sensor reads correctly on the bench and shows 50 Hz noise in the cell. What would you change?(show answer)
The engineering content of electromagnetic interference on sensor lines is the tolerance budget and the stop condition, not the algorithm name.
A long analogue run in a cabinet with servo drives picks up conducted and radiated interference. The sensor is fine; the measurement chain between it and the converter is the part that changed when the system moved.
Concretely, move the conversion close to the sensor and transmit digitally where possible, use differential signalling and correct shield termination at one end, separate signal from power routing, and verify with the drives running rather than with the cell idle.
The reason for that specificity is a failure I have seen: A 6 m single-ended thermocouple run alongside a servo cable picked up 40 mV of noise against a 41 µV/K sensitivity, which is roughly 976 K of apparent variation; the bench test used a 300 mm lead.
Noise by wiring approach.
| Wiring | Length | Noise | Apparent ΔT |
|---|---|---|---|
| single-ended | 6 m | 40 mV | 976 K |
| differential | 6 m | 1.2 mV | 29 K |
| digital at sensor | 6 m | 0 | 0 K |
I would not consider it settled without evidence: Measure noise with the drives energised and moving, at the production cable length and routing.
The bench test measured a different cable.
Curated: · Written: · Reviewed:
QA-72An IP65 robot fails in a washdown food plant. Was the rating wrong?(show answer)
Before letting it move at speed I would write down what a wrong result for ingress protection and duty environment looks like on the log.
An ingress rating describes performance against a specified test, not against an arbitrary environment. Water jets at defined pressure and distance are not the same as heated caustic washdown at close range, and cable glands, connectors, and tooling each carry their own rating.
Concretely, match the rating to the actual cleaning procedure including temperature and chemistry, check the rating of every element in the assembly rather than only the arm, and inspect seals as a scheduled task since they degrade.
The reason for that specificity is a failure I have seen: An IP65 arm with an IP54 tool connector was washed with 60 °C caustic at 300 mm; ingress reached the connector within three weeks and corroded the pins, and the arm itself was undamaged.
Ratings along the washdown path.
| Element | Rating | Survives procedure |
|---|---|---|
| arm | IP65 | yes |
| tool connector | IP54 | no |
| cable gland | IP67 | yes |
I would not consider it settled without evidence: List the rating of every component in the washdown path and test the actual cleaning procedure rather than the standard's test.
The assembly is rated at its weakest element.
Curated: · Written: · Reviewed:
QA-73How do you update firmware on a robot fleet without creating a fleet-wide outage?(show answer)
The first thing I would pin down about firmware update on a moving machine is what the machine physically does when the assumption is wrong.
A firmware update changes the layer that has physical authority. It cannot be rolled back by restarting a process, and a bad image on a machine that will not boot is a site visit rather than a redeploy.
Concretely, require an A/B image with automatic rollback on failed health check, update only machines that are idle and safe, stage across the fleet rather than all at once, and verify the rolled-back path works before relying on it.
The reason for that specificity is a failure I have seen: A fleet-wide push bricked 14 of 60 units because a bootloader assumption held only on the newer hardware revision; the rollback path had never been exercised on that revision and the units needed physical recovery.
Staged rollout versus fleet-wide push.
| Strategy | Units at risk | Recovery |
|---|---|---|
| fleet-wide | 60 | 14 site visits |
| 5% canary | 3 | 0, caught early |
| canary + A/B | 3 | automatic |
I would not consider it settled without evidence: Prove rollback on every hardware revision in the fleet before the first staged update, not after.
An update you cannot undo remotely is a site visit waiting to be scheduled.
Curated: · Written: · Reviewed:
QA-74A cloud service sends motion commands to a robot. What must the robot enforce locally?(show answer)
I would start remote commands and local authority from the measured envelope, not from the number the datasheet promises.
A remote command arrives across a link that can be delayed, replayed, or reordered, from a system whose state may be stale. Local protection has to remain the final authority because it is the only layer that knows what the machine is currently doing.
Concretely, authenticate and bound every command, reject stale or out-of-order ones by sequence and freshness, clamp against local limits regardless of what was requested, and keep protective functions outside the path that the remote system can influence.
The reason for that specificity is a failure I have seen: A queued command replayed after a 40 second network stall told a robot to move to a pose that had been safe when issued; the fixture had since been loaded and the arm drove into it, with the command perfectly valid on arrival.
Command validity checks.
| Check | Catches replay | Catches stale | Catches over-limit |
|---|---|---|---|
| authentication | no | no | no |
| sequence number | yes | no | no |
| freshness window | yes | yes | no |
| local clamp | no | no | yes |
I would not consider it settled without evidence: Test with deliberately delayed, replayed, and reordered commands, and confirm the robot rejects rather than executes them.
The robot is the only party that knows where it is now.
Curated: · Written: · Reviewed:
QA-75Debug logging is enabled on a robot controller and the control loop starts missing deadlines. What is the mechanism?(show answer)
This is an area where a clean simulation and a correct handling of logging volume on an embedded target are not the same event.
Logging from a real-time path costs whatever the slowest part of the write costs: allocation, formatting, a lock, and eventually a blocking flush to storage. The cost is unbounded and lands on the thread that can least afford it.
Concretely, format nothing in the real-time path — push binary records into a lock-free ring and let a lower-priority thread format and write — bound the ring so a slow consumer drops rather than blocks, and count the drops.
The reason for that specificity is a failure I have seen: A debug build formatted strings inside the 1 kHz loop; the p99.9 iteration cost rose from 0.4 ms to 3.2 ms whenever the log file rotated, and the axis logged following errors that only ever appeared with logging on.
Loop cost against logging design.
| Design | p50 | p99.9 | Deadline 1 ms |
|---|---|---|---|
| format in loop | 0.6 ms | 3.2 ms | missed |
| ring + writer | 0.41 ms | 0.52 ms | met |
| logging off | 0.40 ms | 0.48 ms | met |
I would not consider it settled without evidence: Measure loop timing with logging at production verbosity, and confirm the real-time path performs no allocation.
A log line in a control loop is a syscall with a deadline attached.
Curated: · Written: · Reviewed:
QA-76Why is dynamic allocation discouraged in a real-time control loop when the allocator is fast on average?(show answer)
My answer to memory allocation in the control path begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
The allocator's average is irrelevant to a deadline. Fragmentation, a coalescing pass, or a page fault produces a tail that arrives rarely and lands unpredictably, and a loop is judged on that tail.
Concretely, allocate at initialisation, use fixed-size pools and preallocated buffers in the loop, lock pages into memory where the platform supports it, and assert on allocation from the real-time thread so a regression is caught at the point it is introduced.
The reason for that specificity is a failure I have seen: A trajectory buffer that grew on demand triggered a 2.8 ms allocation once every few thousand cycles; the loop deadline was 1 ms and the resulting jerk was audible before it was measurable.
Allocation cost distribution.
| Percentile | Allocation | Loop budget |
|---|---|---|
| 50th | 0.9 µs | 1000 µs |
| 99th | 41 µs | 1000 µs |
| 99.99th | 2800 µs | exceeded |
I would not consider it settled without evidence: Instrument the real-time thread to fail on any allocation, rather than reasoning that the allocation is small.
Deadlines care about the worst allocation, not the average one.
Curated: · Written: · Reviewed:
QA-77When would you implement a control law in fixed point, and what is the risk?(show answer)
I would treat fixed-point versus floating-point on a microcontroller as a claim about the physical world that has to survive a bench test.
Fixed point gives deterministic cost and adequate precision on hardware without a floating-point unit, at the price of range: every intermediate result has to fit, and an overflow wraps silently rather than saturating or signalling.
Concretely, analyse the range of every intermediate against the worst-case input, choose the scaling deliberately and record it, use saturating arithmetic where wrap would be dangerous, and test at the extremes of the input range rather than around nominal.
The reason for that specificity is a failure I have seen: A Q15 integral term wrapped from positive to negative maximum during a long saturation; the actuator reversed to full output in one cycle, and the fault reproduced only after a stall lasting more than 90 seconds.
Q15 integral during long saturation.
| Elapsed saturated | Accumulator | State |
|---|---|---|
| 30 s | 18400 | fine |
| 80 s | 32600 | near limit |
| 95 s | -32700 | wrapped |
I would not consider it settled without evidence: Test at input extremes and during extended saturation, asserting that no intermediate reaches its representable limit.
Fixed point trades a hardware requirement for an analysis requirement.
Curated: · Written: · Reviewed:
QA-78What can you meaningfully test for a robot without a robot?(show answer)
The useful question for unit testing motion code without hardware is what still holds at the edge of the workspace, cold, and under full payload.
Geometry, limit logic, state transitions, and message handling are pure computation and can be tested exhaustively. Contact, friction, timing under load, and anything involving the physical plant cannot, and pretending otherwise produces a green suite and a broken machine.
Concretely, test the computable parts hard — transform round trips, IK against forward kinematics, limit clamps at their boundaries, state recovery from every state — and put the physical claims behind hardware-in-the-loop tests that are explicitly separate rather than mocked.
The reason for that specificity is a failure I have seen: A suite of 340 tests mocked the motion interface and asserted the commands sent; all 340 passed while the real controller rejected every command because the joint values were in degrees against a radians interface, and the mock had accepted whatever it was given.
What each layer can prove.
| Property | Unit test | Sim | Hardware | Escaped defects |
|---|---|---|---|---|
| IK correctness | yes | yes | yes | 0 |
| limit clamping | yes | yes | yes | 0 |
| contact stability | no | partly | yes | 6 |
| loop timing | no | no | yes | 3 |
I would not consider it settled without evidence: Round-trip against the real interface's validation, and keep a hardware suite that runs before release rather than mocking the part that carries the risk.
A mock validates the caller and asserts nothing about the machine.
Curated: · Written: · Reviewed:
QA-79What belongs in a hardware-in-the-loop rig and what does putting it there cost you?(show answer)
I would settle hardware-in-the-loop test design against an instrumented run on real hardware before trusting the model.
A hardware-in-the-loop rig runs the real controller against a simulated plant, which exercises the timing, drivers, and interfaces that pure simulation cannot while keeping the plant reproducible. It is a compromise, and knowing which half is real is the whole point.
Concretely, keep the controller, its scheduler, and its I/O real; simulate the plant with a model whose fidelity is stated per quantity; and record which failures the rig can and cannot reproduce so a green rig is not read as a green machine.
The reason for that specificity is a failure I have seen: A rig with an idealised encoder interface passed every test while the production encoder's 40 µs update jitter caused a following-error trip on hardware; the rig had been treated as equivalent to the machine.
Fidelity by quantity.
| Quantity | Rig fidelity | Covered |
|---|---|---|
| loop scheduling | real | yes |
| bus timing | real | yes |
| encoder jitter | idealised | no |
| contact force | modelled | partly |
I would not consider it settled without evidence: List the quantities the rig models and their fidelity, and keep a hardware acceptance test for the ones it does not.
The rig is real on one side, and the release depends on which side.
Curated: · Written: · Reviewed:
QA-80You need to test candidate poses against a 40 m by 40 m workspace at 5 cm resolution. What do you store?(show answer)
The judgement in choosing a data structure for occupancy is which timestamps and frames are pinned, not which library is fashionable.
A dense grid at that size is 640,000 cells, which is small enough to hold and gives constant-time lookup. The trade turns as resolution or dimensionality rises: at 1 cm in three dimensions the same volume is far beyond dense storage and sparsity becomes the deciding property.
Concretely, pick against the measured occupancy fraction and the query pattern: dense arrays for small, mostly-full spaces with random access; hierarchical or hash-backed sparse structures when occupancy is a few percent; and measure the constant factor rather than reasoning from complexity alone.
The reason for that specificity is a failure I have seen: A team moved from a dense 2-D grid to an octree for a workspace that was 60 percent occupied; lookups went from 40 ns to 900 ns and the planner's checking budget tripled with no memory saved worth having.
Representation against occupancy.
| Occupancy | Dense memory | Octree memory | Query |
|---|---|---|---|
| 60% | 640 kB | 2.1 MB | 900 ns |
| 5% | 640 kB | 90 kB | 700 ns |
| 0.5% | 640 kB | 12 kB | 650 ns |
I would not consider it settled without evidence: Measure occupancy fraction and per-query cost on real maps before changing the representation.
Sparsity is the property that decides this, so measure it.
Curated: · Written: · Reviewed:
QA-81Why is A-star usually preferred to Dijkstra for a navigation graph, and when does the heuristic hurt?(show answer)
Where candidates lose the interview on graph search for navigation is treating a replayed bag as evidence about the robot.
A-star is Dijkstra with a heuristic that orders the frontier toward the goal. It returns an optimal path when the heuristic never overestimates the true remaining cost, and an inflated heuristic returns a path faster and gives up that guarantee.
Concretely, use straight-line distance scaled by the minimum traversal cost per metre so the heuristic stays admissible, and if you inflate it for speed, state the resulting bound on suboptimality rather than describing the result as optimal.
The reason for that specificity is a failure I have seen: A heuristic used straight-line distance while the cost map charged 3.5 times for carpet; the heuristic was admissible but so weak that the search expanded 92 percent of Dijkstra's nodes, and the team concluded A-star did not help.
Heuristic quality on one map.
| Heuristic | Nodes expanded | Path cost | Optimal |
|---|---|---|---|
| none (Dijkstra) | 41200 | 38.2 | yes |
| unscaled straight line | 37900 | 38.2 | yes |
| cost-scaled | 6100 | 38.2 | yes |
| inflated 1.5x | 2300 | 41.0 | no |
I would not consider it settled without evidence: Report nodes expanded and path cost against plain Dijkstra on real maps, rather than assuming the heuristic is doing work.
A heuristic is a promise about remaining cost, so it has to know the cost model.
Curated: · Written: · Reviewed:
QA-82A planner spends 40 percent of its time in the open list. What would you change?(show answer)
I would answer priority queues in a planner's inner loop by separating what the sensor measured from what the visualiser drew.
A binary heap gives logarithmic push and pop, which is usually right. The cost that surprises people is the decrease-key operation, which a plain heap cannot do, so implementations push duplicates and discard stale entries on pop instead.
Concretely, measure where the time actually goes before changing structure: duplicate entries inflate the heap and the pop count, so tracking the best known cost per node and skipping stale pops often recovers more than switching to an exotic queue.
The reason for that specificity is a failure I have seen: A planner pushed a duplicate on every improvement and never skipped stale pops; the open list reached 340,000 entries for a graph with 41,000 nodes, and 88 percent of pops were discarded work.
Open list behaviour on a 41k-node graph.
| Policy | Peak entries | Discarded pops |
|---|---|---|
| duplicates, no skip | 340000 | 88% |
| duplicates, skip stale | 340000 | 88% but cheap |
| tracked best cost | 52000 | 6% |
I would not consider it settled without evidence: Count pushes, pops, and discarded pops against node count, rather than profiling the heap in isolation.
The queue is usually fine and the duplicate policy usually is not.
Curated: · Written: · Reviewed:
QA-83A long chain of transforms accumulates error even in double precision. What would you do?(show answer)
The engineering content of numerical stability in pose composition is the tolerance budget and the stop condition, not the algorithm name.
Repeated matrix multiplication accumulates rounding, and a rotation matrix slowly loses orthonormality. The result stops being a rotation, which shows up as scaling or shear rather than as an obvious error.
Concretely, re-orthonormalise periodically, keep chains short by composing from a canonical reference rather than incrementally, and check the determinant and orthonormality residual rather than assuming double precision is sufficient.
The reason for that specificity is a failure I have seen: An incremental pose accumulated over 400,000 updates drifted to a determinant of 1.0004; the apparent 0.04 percent scale put a 2 m reach 0.8 mm out and nobody suspected the arithmetic.
Drift over accumulated updates.
| Updates | Determinant | Error at 2 m |
|---|---|---|
| 1000 | 1.0000001 | 0.0002 mm |
| 100000 | 1.0001 | 0.2 mm |
| 400000 | 1.0004 | 0.8 mm |
I would not consider it settled without evidence: Assert orthonormality residual and determinant on every accumulated rotation, with a threshold rather than a spot check.
A matrix that is nearly a rotation is not one.
Curated: · Written: · Reviewed:
QA-84Where does Python belong in a robot system, and where does it not?(show answer)
Before letting it move at speed I would write down what a wrong result for Python in the robotics stack looks like on the log.
Python is excellent for orchestration, configuration, offline analysis, and anything whose deadline is measured in tenths of a second. It is a poor fit for a hard real-time loop because garbage collection and the interpreter give an unbounded tail rather than a slow average.
Concretely, keep the control loop in a compiled, deterministic layer and Python above it for sequencing and tooling, pass data across the boundary in preallocated buffers rather than per-cycle objects, and never let a Python thread hold anything the real-time path waits on.
The reason for that specificity is a failure I have seen: A 200 Hz behaviour node in Python showed a 34 ms pause during a collection cycle; the arm continued its last command through the pause and overshot the handover point by 27 mm.
Pause distribution by layer.
| Layer | p50 | p99.9 | Suitable for 5 ms |
|---|---|---|---|
| Python behaviour | 0.9 ms | 34 ms | no |
| C++ control | 0.2 ms | 0.4 ms | yes |
| Python tooling | 12 ms | 90 ms | not in loop |
I would not consider it settled without evidence: Measure the pause distribution of every Python component in the loop and gate on the tail rather than the mean rate.
Python's tail is the property that decides where it can sit.
Curated: · Written: · Reviewed:
QA-85A function modifies a pose array and the caller's data changes too. What happened and how do you prevent a class of these?(show answer)
The first thing I would pin down about NumPy array semantics in pose math is what the machine physically does when the assumption is wrong.
Slicing a NumPy array returns a view sharing the same buffer, so writing through it modifies the original. That is a deliberate performance choice and it makes aliasing the default rather than the exception.
Concretely, copy explicitly at API boundaries where ownership is not obvious, mark arrays read-only when they must not change, and prefer returning new arrays over in-place modification in code that other people call.
The reason for that specificity is a failure I have seen: A trajectory smoother took a slice of the waypoint array and normalised it in place; the caller's original waypoints were silently rewritten, and the second execution of the same trajectory followed a different path from the first.
View versus copy.
| Operation | Shares buffer | Caller affected |
|---|---|---|
| a[2:5] | yes | yes |
| a[2:5].copy() | no | no |
| a[[2,3,4]] | no | no |
I would not consider it settled without evidence: Assert that the caller's input is unchanged after the call in tests covering every function that takes an array.
A view is the same memory with a different name.
Curated: · Written: · Reviewed:
QA-86A collision check occasionally reports contact between objects 0.0000001 mm apart and occasionally misses touching ones. What is the fix?(show answer)
I would start floating-point comparison in geometric predicates from the measured envelope, not from the number the datasheet promises.
Exact equality on floating-point results of geometric computation is meaningless, because the last bits carry accumulated rounding rather than information. A predicate needs a tolerance chosen from the physical problem rather than from the type.
Concretely, choose the tolerance from the manufacturing and calibration uncertainty of the objects being tested, apply it consistently to both sides of the comparison, and inflate the collision geometry by the uncertainty rather than testing the nominal shapes exactly.
The reason for that specificity is a failure I have seen: A predicate tested for exact zero clearance; with 0.4 mm of real positional uncertainty it reported contact on 3 percent of clearly separated pairs and passed pairs that were physically touching.
Predicate behaviour against tolerance.
| Tolerance | False contacts | Missed contacts |
|---|---|---|
| exact zero | 3.0% | 2.1% |
| 0.1 mm | 0.4% | 0.6% |
| 0.4 mm (measured) | 0.0% | 0.0% |
I would not consider it settled without evidence: Set the tolerance from the measured uncertainty budget and show that the predicate's decision is stable across that band.
The tolerance comes from the physics, not from the floating-point type.
Curated: · Written: · Reviewed:
QA-87Perception runs at 10 Hz and control at 1 kHz. How do they share data?(show answer)
This is an area where a clean simulation and a correct handling of concurrency between perception and control are not the same event.
A slow producer and a fast consumer must not be coupled by a lock, because the fast side would then wait for the slow one. The consumer needs the most recent complete result, never a partially written one, and never a wait.
Concretely, publish complete results by atomic pointer swap into a double or triple buffer so the consumer always reads a consistent snapshot without blocking, and carry the result's timestamp so the consumer can decide whether it is still usable.
The reason for that specificity is a failure I have seen: A shared struct guarded by a mutex blocked the 1 kHz loop for up to 1.4 ms while perception wrote; the loop missed deadlines every perception cycle, which appeared as a 10 Hz vibration in the arm.
Handoff design and loop blocking.
| Mechanism | Blocking | Torn reads |
|---|---|---|
| mutex | 1.4 ms | no |
| unguarded struct | 0 | yes |
| pointer swap | 0.001 ms | no |
I would not consider it settled without evidence: Measure control-loop blocking time attributable to the perception handoff and require it to be effectively zero.
The fast side should never wait for the slow one.
Curated: · Written: · Reviewed:
QA-88A camera vendor's SDK and the robot controller disagree about which axis is up. How do you handle it once?(show answer)
My answer to coordinate conventions across vendors begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
Axis convention, handedness, rotation order, and angle units differ between vendors and are all equally defensible. The cost is not in any single conversion but in the conversions being scattered so that some paths get them and others do not.
Concretely, convert once at the boundary into a single internal convention, keep the vendor convention out of everything downstream, and make the boundary the only place where a conversion appears so a missed one is a missing adapter rather than a silent bug.
The reason for that specificity is a failure I have seen: A codebase converted camera coordinates in the pick path and not in the calibration-verification path; the verification reported a clean result against the unconverted data while production picks were 90 degrees out about one axis.
Conversion coverage by path.
| Path | Converts | Result |
|---|---|---|
| pick | yes | correct |
| verification | no | falsely clean |
| diagnostics | no | misleading |
I would not consider it settled without evidence: Assert on a fixture that every path — production, verification, and diagnostics — produces the same value from the same raw input.
Convert once at the edge, and there is only one place to be wrong.
Curated: · Written: · Reviewed:
QA-89Adding a second camera did not improve pose estimation. Why might that be?(show answer)
I would treat sensor placement and observability as a claim about the physical world that has to survive a bench test.
Estimation quality depends on geometry, not on sensor count. Two cameras with nearly the same viewpoint contribute nearly the same constraint, so the estimate gains redundancy against failure without gaining observability along the weak direction.
Concretely, analyse which directions are poorly constrained before adding hardware, place the second sensor to observe those directions, and validate with the covariance along the weak axis rather than with an overall residual.
The reason for that specificity is a failure I have seen: A second camera mounted 120 mm from the first improved overall reprojection residual by 4 percent and left depth uncertainty at 11 mm, because both viewpoints constrained the lateral directions and neither constrained range.
Per-axis uncertainty after adding a camera.
| Axis | One camera | Two, 120 mm apart | Two, 900 mm apart |
|---|---|---|---|
| lateral | 0.9 mm | 0.7 mm | 0.6 mm |
| vertical | 1.0 mm | 0.8 mm | 0.7 mm |
| depth | 12 mm | 11 mm | 1.4 mm |
I would not consider it settled without evidence: Report per-axis uncertainty before and after, not a pooled residual.
Two sensors looking the same way answer the same question twice.
Curated: · Written: · Reviewed:
QA-90A square fiducial's estimated pose flips between two orientations at long range. What causes it?(show answer)
The useful question for fiducial markers and pose ambiguity is what still holds at the edge of the workspace, cold, and under full payload.
A planar marker viewed nearly head-on has two poses that project to almost the same image, and the ambiguity resolves only through the small perspective difference between them. At long range or low resolution that difference falls below the noise and the estimate flips.
Concretely, increase apparent marker size or resolution so the perspective cue is measurable, use multiple markers in a rigid arrangement so the ambiguity of one is resolved by the others, and reject a pose whose two candidate solutions have similar residuals rather than taking the marginally better one.
The reason for that specificity is a failure I have seen: A single 60 mm marker at 2.4 m produced pose flips on roughly one frame in eight; the robot's approach direction reversed on those frames and the resulting motion looked to operators like a random jerk.
Ambiguity against apparent marker size.
| Marker | Range | Apparent size | Flip rate |
|---|---|---|---|
| 60 mm | 2.4 m | 22 px | 12% |
| 150 mm | 2.4 m | 55 px | 0.4% |
| 60 mm x4 rigid | 2.4 m | 22 px | 0.0% |
I would not consider it settled without evidence: Report the residual ratio between the two candidate poses per detection and reject below a declared margin.
Two solutions that fit equally well is a measurement, not a tie to break.
Curated: · Written: · Reviewed:
QA-91A robot picks from a moving conveyor and placement error grows with belt speed. What is the likely cause?(show answer)
I would settle conveyor tracking and encoder coupling against an instrumented run on real hardware before trusting the model.
Tracking requires the robot to know belt position at the same instants it knows its own. An encoder read on a different clock, at a different rate, or through a slow path introduces a phase error that scales directly with belt speed.
Concretely, latch belt position in the same time base as the robot's own state, ideally in the drive itself, and validate the coupling by picking at several speeds and checking that the residual does not trend.
The reason for that specificity is a failure I have seen: A belt encoder polled over a 20 ms application cycle gave a phase error that put placement 6 mm out at 0.3 m/s and 20 mm out at 1.0 m/s; the fix was hardware latching, not a better vision model.
Placement error against belt speed.
| Speed | Polled encoder | Latched encoder |
|---|---|---|
| 0.3 m/s | 6 mm | 0.4 mm |
| 0.6 m/s | 12 mm | 0.4 mm |
| 1.0 m/s | 20 mm | 0.5 mm |
I would not consider it settled without evidence: Plot placement error against belt speed and require a flat relationship rather than a low value at one speed.
A residual that grows with speed is a timing problem, not an accuracy one.
Curated: · Written: · Reviewed:
QA-92A gantry rings after every move. What determines whether you fix it in mechanics or in control?(show answer)
The judgement in vibration and structural modes is which timestamps and frames are pinned, not which library is fashionable.
A structure has natural frequencies set by its stiffness and mass. Control can avoid exciting a mode by shaping the command, but it cannot raise the frequency, and a mode close to the desired motion bandwidth constrains what any controller can achieve.
Concretely, measure the mode frequencies and damping first, then choose: input shaping or a notch where the mode is well separated and stable, and mechanical stiffening where the mode sits inside the bandwidth the application needs.
The reason for that specificity is a failure I have seen: A team notched a 14 Hz mode and the required bandwidth was 9 Hz; the notch removed the ringing and the axis could no longer track the required profile, so cycle time rose 18 percent.
Mode frequency against required bandwidth.
| Mode | Required bandwidth | Remedy |
|---|---|---|
| 60 Hz | 9 Hz | notch |
| 14 Hz | 9 Hz | stiffen |
| 14 Hz | 3 Hz | input shaping |
I would not consider it settled without evidence: Measure the frequency response of the axis and show the mode against the required bandwidth before choosing the remedy.
Control can dodge a mode, and only mechanics can move it.
Curated: · Written: · Reviewed:
QA-93What does a meaningful acceptance test for a cell contain beyond running the cycle?(show answer)
Where candidates lose the interview on acceptance testing a robotic cell is treating a replayed bag as evidence about the robot.
A cycle that runs proves the happy path once. Acceptance has to establish capability over the variation the cell will see and behaviour at the boundaries, because those are what determine whether it still works in month three.
Concretely, run a defined number of consecutive cycles at production rate with production parts spanning the tolerance range, measure capability rather than pass or fail, exercise each fault and recovery path deliberately, and record the baseline that later performance is compared against.
The reason for that specificity is a failure I have seen: A cell was accepted on 20 consecutive good cycles with parts from one lot; production capability against the placement tolerance was well below what the specification implied and the scrap rate settled at 4 percent.
What each acceptance approach establishes.
| Test | Cycles | Establishes |
|---|---|---|
| demo run | 20 | it works once |
| 500-cycle run | 500 | rate and capability |
| plus fault injection | 500 | recovery behaviour |
I would not consider it settled without evidence: Report a capability index against the stated tolerance from a run long enough to estimate it, not a count of successful cycles.
Twenty good cycles is a demonstration, not a capability.
Curated: · Written: · Reviewed:
QA-94An operator jogs an axis expecting tool-frame motion and the arm moves in base frame. Whose fault is that?(show answer)
I would answer operator interface and mode clarity by separating what the sensor measured from what the visualiser drew.
A jog command is meaningless without a frame, and the interface is where that gets decided. If the active frame is not visible at the moment of the command, the operator is guessing, and the machine will do exactly what it was told.
Concretely, show the active frame, speed limit, and mode where the operator's attention already is, require an explicit action to change frame rather than inheriting the last one, and default to the most conservative interpretation after any interruption.
The reason for that specificity is a failure I have seen: A pendant retained tool frame from a previous session; an operator jogged what they believed was straight up and moved 40 mm laterally into a fixture, and the interface had shown the active frame two screens away.
Same jog, two active frames.
| Active frame | Operator intent | Actual motion |
|---|---|---|
| base | up | up |
| tool, wrist level | up | up |
| tool, wrist rolled 90° | up | 40 mm lateral |
I would not consider it settled without evidence: Test the interface with operators who did not build it, at the moment of the command, rather than reviewing the screen design.
The frame is part of the command, so it belongs where the command is given.
Curated: · Written: · Reviewed:
QA-95One of three cameras fails mid-shift. Should the cell stop?(show answer)
The engineering content of handling degraded modes deliberately is the tolerance budget and the stop condition, not the algorithm name.
A degraded mode is a designed state or an accident. If the remaining sensing genuinely supports a reduced task at a reduced rate, continuing is a decision; if nobody decided, the system is running outside its validated envelope while reporting normal operation.
Concretely, enumerate degraded modes with what each still guarantees, make entry to one explicit and visible, and constrain speed or task scope to what the remaining sensing supports rather than continuing unchanged.
The reason for that specificity is a failure I have seen: A cell lost one of three cameras and continued at full rate; the remaining pair could not resolve part orientation for one variant, and 1,400 parts of that variant were placed rotated before the shift ended.
Capability by sensor availability.
| Cameras | Position | Orientation | Permitted rate |
|---|---|---|---|
| 3 | yes | yes | 100% |
| 2 | yes | one variant only | 60% |
| 1 | coarse | no | stop |
I would not consider it settled without evidence: Test each single-sensor failure explicitly and record what the cell can still do, rather than assuming graceful degradation.
Degraded operation is a mode, so it needs a definition and a limit.
Curated: · Written: · Reviewed:
QA-96Two designs have the same reliability figure. What else determines line availability?(show answer)
Before letting it move at speed I would write down what a wrong result for spare parts and mean time to repair looks like on the log.
Availability depends on how long repair takes as much as on how often failure happens. A component that fails rarely but takes two days to source and four hours to align costs more downtime than one that fails more often and swaps in ten minutes.
Concretely, design for replacement — keyed mounting, stored calibration, no realignment — hold spares for the long-lead items, and measure repair time in practice rather than estimating it from the assembly drawing.
The reason for that specificity is a failure I have seen: A camera mount requiring full recalibration took 3.5 hours to replace against a 6-month failure interval; a keyed mount with stored extrinsics brought it to 12 minutes and moved annual downtime from 7 hours to 24 minutes.
Annual downtime from one component.
| Design | Interval | Repair time | Annual downtime |
|---|---|---|---|
| recalibrate | 6 months | 3.5 h | 7.0 h |
| keyed mount | 6 months | 12 min | 24 min |
I would not consider it settled without evidence: Time an actual replacement during commissioning rather than estimating it, and record it as an acceptance figure.
Availability is a ratio, and the denominator is repair time.
Curated: · Written: · Reviewed:
QA-97Which datasheet numbers should you distrust for your application, and what do you measure instead?(show answer)
The first thing I would pin down about reading a robot datasheet critically is what the machine physically does when the assumption is wrong.
Datasheet figures are measured under the manufacturer's stated conditions, which are usually favourable: rated payload at the flange, repeatability at moderate reach and temperature, speed per joint rather than at the tool. Every one of them can be true and inapplicable.
Concretely, restate each figure with your own conditions — payload at your centre of mass, accuracy across your envelope, tool speed on your path — and measure the ones your tolerance depends on rather than taking the number into the specification.
The reason for that specificity is a failure I have seen: A cell was specified on a 10 kg rated payload; with the tool's centre of mass at 180 mm the permitted payload was 6.4 kg, and the arm ran into torque limits on every acceleration once the part was added.
Rated payload against centre of mass.
| CoM offset | Permitted payload |
|---|---|
| 50 mm | 10.0 kg |
| 120 mm | 7.8 kg |
| 180 mm | 6.4 kg |
I would not consider it settled without evidence: Recompute the payload against the actual centre of mass on the manufacturer's own load chart before committing to a tool design.
The number is true under conditions that are not yours.
Curated: · Written: · Reviewed:
QA-98A cell handles four product variants and errors cluster in the first ten minutes after a changeover. Why?(show answer)
I would start changeover between product variants from the measured envelope, not from the number the datasheet promises.
A changeover swaps several things at once — tooling, program, fixture, and part presentation — and any one of them can be left in the previous variant's state. The clustering is the signature of a manual step that has no verification.
Concretely, make the variant a single selected value that drives every dependent setting, verify each physical change with a sensor or a scan rather than a checklist, and refuse to start until the verified state matches the selection.
The reason for that specificity is a failure I have seen: A changeover left the previous variant's gripper fingers fitted while the program had switched; the first eight parts were crushed before an operator noticed, and the same failure recurred at roughly one changeover in twenty.
Changeover errors by verification.
| Method | Errors per 100 changeovers |
|---|---|
| operator checklist | 5 |
| scanned tooling id | 0.2 |
| interlocked detection | 0 |
I would not consider it settled without evidence: Detect the fitted tooling rather than asking the operator to confirm it, and log any mismatch between selection and detection.
A changeover is a state transition, so verify the state.
Curated: · Written: · Reviewed:
QA-99What makes a robotics incident report useful six months later?(show answer)
This is an area where a clean simulation and a correct handling of writing a robotics postmortem are not the same event.
An incident involves a physical machine, so the report has to reconstruct what the machine did, not only what the software decided. Without the recorded motion, the conclusion is a hypothesis that later readers cannot check.
Concretely, include the frozen high-rate record, the configuration and calibration in force, the physical state of tooling and parts, the sequence with timings, and the specific change that would have prevented it. Name the detection gap separately from the cause.
The reason for that specificity is a failure I have seen: A report attributed a collision to operator error with no motion record; 8 months later the same collision recurred on another of the 12 cells, and the actual cause — a scene that retained a fixture removed 90 seconds earlier — had to be found from scratch over 3 days.
What the report needs to answer.
| Question | Source | Present | Days to re-derive |
|---|---|---|---|
| what did it command | frozen log | no | 1 |
| what did it measure | frozen log | no | 1 |
| what config was live | config hash | no | 1 |
| what would prevent it | analysis | partly | 0 |
I would not consider it settled without evidence: Require the motion record and the effective configuration as attachments before a report is accepted as complete.
A conclusion nobody can re-derive is not evidence.
Curated: · Written: · Reviewed:
QA-100When would you advise against automating a task?(show answer)
My answer to deciding when a robot is the wrong answer begins at the failure mode: if I cannot name how it hurts someone, I do not have a design.
Automation pays where the task is repeatable, the presentation is controlled, and the volume amortises the engineering. Where variation is high and volume is low, the cost lands in fixturing and exception handling rather than in the robot, and the payback disappears.
Concretely, cost the whole system including fixturing, presentation, exception handling, and the labour that remains, estimate the exception rate from observation rather than from the ideal case, and compare against the actual current cost rather than an assumed one.
The reason for that specificity is a failure I have seen: A cell justified on a 4 percent assumed exception rate met 19 percent in production because incoming parts arrived in eleven presentations rather than the two that were specified; an operator was needed full time and the payback moved from 14 months to never.
Payback against exception rate.
| Exception rate | Operator load | Payback |
|---|---|---|
| 4% | 0.1 FTE | 14 months |
| 10% | 0.4 FTE | 38 months |
| 19% | 1.0 FTE | none |
I would not consider it settled without evidence: Measure the actual variation in incoming presentation over a representative period before committing to the design.
The robot is usually the cheap part of the answer.
Curated: · Written: · Reviewed:
