This
architecture combines elements commonly associated with stereo-inertial and
monocular-inertial SLAM. It can use stereo where overlap exists without
requiring every camera pair to overlap, and it can retain useful constraints
from divergent or lateral views. The authors state that the same binaries run
across handheld fisheye rigs, aerial stereo platforms, virtual-reality
headsets, and research glasses, with rig-dependent choices derived from each
rig rather than selected by dataset. The report does not disclose the complete
factor graph, numerical weights, feature and place-recognition implementations,
thresholds, software interface, licensing terms, or aerial-platform test
results, so those portions should be treated as vendor implementation details requiring
evaluation rather than assumed capabilities.
Calibration
is a major part of the technique. Each RoboCap unit is factory-calibrated by
moving it around a stationary AprilGrid while all camera groups and both IMUs
observe the sequence. Continuous-time batch optimization estimates camera
intrinsics, camera-to-IMU extrinsics, and timing offsets; the camera model uses
Kannala–Brandt fisheye distortion, and IMU noise terms are characterized by
Allan-variance analysis. The report also describes a per-session
self-calibration step because thermal and mechanical stress can alter relative
camera orientation after factory calibration. Natural feature correspondences
are used to estimate small rotational corrections against Sampson epipolar
distance, with one camera fixed as the anchor; a correction is accepted only
when it improves a held-out residual and remains within approximately one
degree of that unit’s factory calibration.
The
measured calibration drift in the reported RoboCap fleet was 0.2 to 1.1 degrees
of relative rotation, with a median near 0.5 degrees, corresponding to 1 to 5
pixels of epipolar residual and 1 to 2 cm of depth bias at manipulation range.
Across 150 devices, the paper reports a reduction in median front-pair epipolar
residual from 3.8 pixels under factory calibration to 0.56 pixels after session
calibration. These measurements concern RoboCap geometry at close range and do
not establish equivalent drone performance, but they illustrate why a nominal
camera model is insufficient when metric outputs are expected from inexpensive
or mechanically exposed sensor assemblies.
For a drone, the analogous
calibration service would need a different parameterization and acceptance
policy. A rigid gimbal-mounted stereo pair might require estimation of residual
camera rotation and camera-to-IMU or camera-to-airframe alignment, while a
vibration-isolated payload might also require monitoring slowly varying
translation, gimbal angles, and timestamp offsets. Rolling-shutter cameras
would require line-time and readout-direction modeling, or their measurements
would need to be excluded during motion regimes where a global-shutter
approximation is invalid. Zoom, autofocus, digital stabilization, and
electronic image cropping can change the effective intrinsics, so the
calibration state should be versioned by camera mode and tied to every produced
pose and depth map. Unlike RoboCap’s once-per-session rotational correction,
drone calibration may need segmented validity intervals because temperature,
vibration, gimbal motion, payload replacement, and hard landings can change
geometry during an operational day.
No comments:
Post a Comment