Friday, October 9, 2026

Continued from previous post

 

This architecture combines elements commonly associated with stereo-inertial and monocular-inertial SLAM. It can use stereo where overlap exists without requiring every camera pair to overlap, and it can retain useful constraints from divergent or lateral views. The authors state that the same binaries run across handheld fisheye rigs, aerial stereo platforms, virtual-reality headsets, and research glasses, with rig-dependent choices derived from each rig rather than selected by dataset. The report does not disclose the complete factor graph, numerical weights, feature and place-recognition implementations, thresholds, software interface, licensing terms, or aerial-platform test results, so those portions should be treated as vendor implementation details requiring evaluation rather than assumed capabilities.

Calibration is a major part of the technique. Each RoboCap unit is factory-calibrated by moving it around a stationary AprilGrid while all camera groups and both IMUs observe the sequence. Continuous-time batch optimization estimates camera intrinsics, camera-to-IMU extrinsics, and timing offsets; the camera model uses Kannala–Brandt fisheye distortion, and IMU noise terms are characterized by Allan-variance analysis. The report also describes a per-session self-calibration step because thermal and mechanical stress can alter relative camera orientation after factory calibration. Natural feature correspondences are used to estimate small rotational corrections against Sampson epipolar distance, with one camera fixed as the anchor; a correction is accepted only when it improves a held-out residual and remains within approximately one degree of that unit’s factory calibration.

The measured calibration drift in the reported RoboCap fleet was 0.2 to 1.1 degrees of relative rotation, with a median near 0.5 degrees, corresponding to 1 to 5 pixels of epipolar residual and 1 to 2 cm of depth bias at manipulation range. Across 150 devices, the paper reports a reduction in median front-pair epipolar residual from 3.8 pixels under factory calibration to 0.56 pixels after session calibration. These measurements concern RoboCap geometry at close range and do not establish equivalent drone performance, but they illustrate why a nominal camera model is insufficient when metric outputs are expected from inexpensive or mechanically exposed sensor assemblies.

For a drone, the analogous calibration service would need a different parameterization and acceptance policy. A rigid gimbal-mounted stereo pair might require estimation of residual camera rotation and camera-to-IMU or camera-to-airframe alignment, while a vibration-isolated payload might also require monitoring slowly varying translation, gimbal angles, and timestamp offsets. Rolling-shutter cameras would require line-time and readout-direction modeling, or their measurements would need to be excluded during motion regimes where a global-shutter approximation is invalid. Zoom, autofocus, digital stabilization, and electronic image cropping can change the effective intrinsics, so the calibration state should be versioned by camera mode and tied to every produced pose and depth map. Unlike RoboCap’s once-per-session rotational correction, drone calibration may need segmented validity intervals because temperature, vibration, gimbal motion, payload replacement, and hard landings can change geometry during an operational day.

No comments:

Post a Comment