scieee AI-readable full text Open interactive document viewer

A DVS-RADAR dataset to bridge the gap towards conventional RGB-RADAR sensing

Pieter Meiresone; David Van Hamme; Jan Aelterman; Wilfried Philips

Abstract

Dynamic Vision Sensors (DVS), also known as event cameras, are bioinspired imaging sensors that are emerging as a potential valuable tool in the automotive industry. Their extremely low latency and high dynamic range make them well suited to enhance perception in autonomous driving systems. However, despite these advantages, event cameras remain challenging to work with, and their detection performance still lags behind that of conventional RGB sensor based systems. There is a broad consensus that a robust autonomous driving system will consist of multiple sensors including RGB and lidar/radar. According to the authors knowledge, no automotive dataset exist containing event and radar data. This gap limits innovation with DVS cameras in the automotive industry. We created our own measurement setup and captured our own dataset to encourage more researches to work on DVS and RADAR based sensing systems to bridge the gap towards conventional RGB and RADAR based sensing systems.

Full text

A DVS-RADAR dataset to bridge the gap towards conventional RGB-RADAR sensing IPI – IMAGE PROCESSING AND INTERPRETATION Pieter Meiresone, David Van Hamme, Jan Aelterman, Wilfried Philips Capture platform A car was equipped with a DAVIS 346 event camera and TI IWR6843 radar. An additional Intel Realsense camera was used as reference camera to obtain high quality RGB images. An overview of the radar configuration is show in the table on the right. An important design choice was made to have a Doppler resolution of 0.059m/s with a maximum unambiguous velocity of 14km/h. While this creates aliasing, the high doppler velocity resolution should capture clearly the micro doppler patterns from VRUs [4]. An Intel Nuc is used to connect the different sensors and to timestamp the data streams. We briefly outline how the annotation process works. A state of the art detector on the reference camera is used to generate high quality candidate labels, each label is transformed to the image plane of the DVS camera and manually verified. This process is visualized on the left. Contact [email protected] Universiteit Gent @ugent Ghent University Introduction Dynamic Vision Sensors (DVS), also known as event cameras, are bioinspired imaging sensors that are emerging as a potential valuable tool in the automotive industry. Their extremely low latency and high dynamic range make them well suited to enhance perception in autonomous driving systems. However, despite these advantages, event cameras remain challenging to work with, and their detection performance still lags behind that of conventional RGB sensor based systems. There is a broad consensus that a robust autonomous driving system will consist of multiple sensors including RGB and lidar/radar. According to the authors knowledge, no automotive dataset exist containing event and radar data. This gap limits innovation with DVS cameras in the automotive industry. We created our own measurement setup and captured our own dataset to encourage more researches to work on DVS and RADAR based sensing systems to bridge the gap towards conventional RGB and RADAR based sensing systems. Proposed dataset In total we have driven around 100km which is 3 hours of recorded data around the city of Gent in Belgium as show in Figure 7. From these recordings we create an annotated subset with our semi-automatic labelling process, in total we annotated 55k frames with around 130k bounding boxes. Of these bounding boxes there are approximately 115k bounding boxes for cars and 15k bounding boxes for VRUs. We provide ground truth for cars and Vulnerable Road User (pedestrians and cyclists). The ground truth consist of a timestamp and position in the image plane of the DVS camera. Although the DVS camera produces an almost continuous event stream, we propose to evaluate the detector only at the discrete timestamps specified in the ground truth. We propose to use the same evaluation protocol as in [5] and omit bounding boxes in the dataset with a side length of less than 10 pixels and a diagonal of less than 30 pixels. Conclusion With the release of this dataset, we hope that we fill a gap in automative datasets to support innovation with bioinspired imaging sensors to close the gap towards conventional RGB + RADAR sensing. Furthermore, it has the potential benefits of •Neuromorphic processing •Extreme low latency DVS sensor complementary to high latency radar sensor •High dynamic range which can outperform RGB sensor in corner cases •Optical flow field gives speed estimate which can potentially improve tracking In summary it holds potential to create new sensing algorithms resilience to motion blur, lowlight conditions, low-latency and high dynamic range challenges. Difficult settings with motion blur and lens flare [1] P. de Tournemire, D. Nitti, E. Perot, D. Migliore, and A. Sironi, “A large scale event-based detection dataset for automotive.” [Online]. Available: http://arxiv.org/abs/2001.08499 [2] E. Perot, P. de Tournemire, D. Nitti, J. Masci, and A. Sironi, “Learning to detect objects with a 1 megapixel event camera,” in Advances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., pp. 16 639–16 652. [3] Gehrig, M., Aarents, W., Gehrig, D. & Scaramuzza,D. DSEC: a stereo event camera dataset for driving scenarios. IEEE Robot. Automat.Lett. 6,4947–4954 (2021). [4] M. Dimitrievski, I. Shopovska, D. V. Hamme, P. Veelaert, and W. Philips, “Weakly supervised deep learning method for vulnerable road user detection in FMCW radar,” in 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), pp. 1–8. [5] M. Gehrig and D. Scaramuzza, “Recurrent vision transformers for object detection with event cameras,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13 884–13 893, ISSN: 2575-7075. [Online]. Available: https://ieeexplore.ieee.org/document/10204090 Existing datasets As of the time of writing, only few big annotated automotive event datasets are available. Gen1 [1] is an automotive dataset containing 39 hours of event camera recordings at a resolution of 304 ×240. In total bounding boxes for 228k cars and 28k pedestrian were manually annotated. 1 Mpx [2] contains 15 hours of recordings in HD resolution (1280 ×720) with a total of 25M bounding boxes for cars, pedestrians and two-wheelers. DSEC [3] presents the first stereo dataset for driving scenarios. Overview of recorded locations, covering a mix of urban traffic scenes, countryside roads, and national roads. Gen1 1Mpx DSEC Duration 39 hours 15 hours 70k frames Bounding boxes 256k 25M 390k Other sensors No No 2xDVS, 2xRGB, Lidar Left: Events aggregated in an event frame (green are positive and red are negative events) annotated cars and VRU. Right: Range-azimuth and range-doppler projections. Top left: capturing vehicle with a roof box converted to measurement box including batteries to power the setup. This setup is flexible in adding or removing sensors. Bottom left: Transfer of labels from reference RGB to DVS Left: radar cube with range, doppler and azimuth information. Right: specification of the radar sensor. Bounding box heatmap in the image plane, with near-uniform distribution on the ground plane.