Cross-camera Monocular 3D Detection for Autonomous Racing
Abstract
A robust perception stack is critical, and relying on multiple sensors, such as LiDAR, radar, and cameras, ensures redundancy and accuracy. Having a reliable monocular 3D detection network acts as a safety net when other sensors fail due to harsh conditions, hardware faults, or unexpected inferences.We propose a methodology for cross-camera monocular 3D detection in the Autonomous Racing scenario.
Full text
Cross-camera Monocular 3D Detection for Autonomous Racing Elena Govi*, Davide Malvezzi*, Davide Sapienza*, Micaela Verucchia, Marko Bertogna* *Università di Modena e Reggio Emilia, Italy a Hipert srl, Modena, Italy Introduction Qualitative Results References Quantitative Results Motivation Methodology Dataset This study was carried out within the MOST – Sustainable Mobility National Research Center and received funding from the European Union Next-GenerationEU (PIANO NAZIONALE DI RIPRESA E RESILIENZA (PNRR) – MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1.4 – D.D. 1033 17/06/2022, CN00000023). This manuscript reflects only the authors’ views and opinions, neither the European Union nor the European Commission can be considered responsible for them. A robust perception stack is critical, and relying on multiple sensors, such as LiDAR, radar, and cameras, ensures redundancy and accuracy. Having a reliable monocular 3D detection network acts as a safety net when other sensors fail due to harsh conditions, hardware faults, or unexpected inferences. Scope We propose a methodology for cross-camera monocular 3D detection in the Autonomous Racing scenario MonoDETR [1] was chosen as 3D detector; Inspired by [2], we used an adaptive strategy for better inter-cameras generalization, called virtual depth. We further extend it to dimensions. The dataset is composed by acquisitions from five different logs, acquired in different race and countries; 4490 samples for training; KITTI format with RGB images, labels and calib. files; Resolution is 1032 x 772; Inputs taken from three different frontal cameras; Depth range is between 3.8 and 135.0 meters, depth mean equal to 44.7808 m; One log is excluded from the dataset and used for test only. [1] Zhang, Renrui, et al. "MonoDETR: Depth-guided transformer for monocular 3D object detection." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. [2]Brazil, Garrick, et al. "Omni3d: A large benchmark and model for 3d object detection in the wild." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023.