scieee AI-readable full text Open interactive document viewer

Large Scale Asset Detection Within Railway Scene Point Cloud Data From Mobile Laser Scanning

Ton, Bram; Akster, Rick

Abstract

To reduce greenhouse gas emissions from the transport sector, shifting to rail transport is crucial. This transition will increase the demand on existing rail infrastructure, necessitating large-scale monitoring to maintain its resilience. Point cloud data are an ideal candidate for this purpose, as they provide immediate, precise 3D geometric information independent of illumination conditions. This study investigates two object detection models, the PointPillar and the CenterPoint model, to automatically create a digital representation of the rail environment. Using a custom open dataset, these two models are evaluated to detect masts, tension rods, signals, and relay cabinets. A mean Average Precision ([email protected]) of 70.6% is achieved. A unique contribution of this study is an in-depth analysis of the locational error in terms of the x and y components of the detected positions. This analysis reveals that location accuracy is not yet sufficient for engineering applications. The analysis indicates that the largest contribution to this error originates from the random error. Additionally, this study demonstrates that transfer learning effectively reduces the labeling burden. For instance, when using 25% of the training data, the average Precision (AP) for the tension rod class improves from 9.5% without transfer learning to 70.8% with transfer learning.

Full text

Date of publication xxxx 00, 0000, date of current version xxxx 00, 0000. Digital Object Identifier 10.1109/ACCESS.2024.0429000 Large Scale Asset Detection Within Railway Scene Point Cloud Data From Mobile Laser Scanning B. TON1, and R. AKSTER2 1Saxion University of Applied Sciences, Ambient Intelligence, M.H. Tromplaan 28, 7513 AB Enschede, Netherlands (e-mail: b[email protected]) 2Strukton Rail, Westkanaaldijk 2, 3542 DA Utrecht, Netherlands (e-mail: [email protected]) Corresponding author: B. Ton (e-mail: b[email protected]). This work was supported by the TechForFuture Grant 2207 ‘‘Digital Twinning voor Spoorontwerp" and by the Nederlandse Organisatie voor Wetenschappelijk Onderzoek (NWO) grant number NWA.1160.18.238. Article Processing Charges (APC) were funded by the Saxion University of Applied Sciences. Data are available on-line at: https://doi.org/10.4121/fa259c52-a585-420c-8a0c-af5e91518e29 ABSTRACT To reduce greenhouse gas emissions from the transport sector, shifting to rail transport is crucial. This transition will increase the demand on existing rail infrastructure, necessitating large-scale monitoring to maintain its resilience. Point cloud data are an ideal candidate for this purpose, as they provide immediate, precise 3D geometric information independent of illumination conditions. This study investigates two object detection models, the PointPillar and the CenterPoint model, to automatically create a digital representation of the rail environment. Using a custom open dataset, these two models are evaluated to detect masts, tension rods, signals, and relay cabinets. A mean Average Precision ([email protected]) of 70.6% is achieved. A unique contribution of this study is an in-depth analysis of the locational error in terms of the x and ycomponents of the detected positions. This analysis reveals that location accuracy is not yet sufficient for engineering applications. The analysis indicates that the largest contribution to this error originates from the random error. Additionally, this study demonstrates that transfer learning effectively reduces the labeling burden. For instance, when using 25% of the training data, the average Precision (AP) for the tension rod class improves from 9.5% without transfer learning to 70.8% with transfer learning. INDEX TERMS Deep learning, LiDaR, mobile laser scanning, object detection, point cloud, PointPillar, railway I. INTRODUCTION Climate change, exacerbated by increased greenhouse gas emissions, has a serious impact on the world [1]. To mitigate these emissions, particularly from the transport sector, rail transport offers a viable solution [2]. The European Union advocates for a shift from road to rail transport as a means to significantly reduce greenhouse gas emissions [3]. The anticipated rise in passenger and freight rail traffic in the near future will inevitably increase the strain on railway infrastructure. Ensuring the continued reliability, availability, maintainability and safety of the railway network in an efficient way, requires accurate and up-to-date digital representations of the railway environment. These as-is models are essential for various applications, including work planning, immersive visualizations, infrastructure health monitoring, automated inventory assessment, and predictive maintenance. However, the automated creation of these as-is models presents significant challenges due to the complex and expansive nature of railway environments. These environments encompass both discrete objects, such as signals, catenary arches, and relay cabinets, and continuous structures, such as overhead catenary wires and rail tracks. The diversity and number of distinct object classes, especially in countries with dense and historically rich rail networks like the Netherlands, add to the complexity. Point clouds serve as an appropriate data source for creating these as-is models. Point clouds are collected by using a mobile laser scanner, which uses the time-of-flight of light pulses to determine the distance between sensor and object. Point clouds can be captured regardless of external illumination variations and provide immediate, precise 3D geometric information. This in contrast with image data, where the quality is influenced by external illumination conditions and where advanced photogrammetric computations are required to obtain 3D information from images. A research area which is well established with regard to scene understanding from point cloud data is the area of autonomous driving vehicles. Typically autonomous driving VOLUME 11, 2023 1 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS vehicles are equipped, among a plethora of other sensors, with laser scanners. The data from these laser scanners are used to detect other road users such as other vehicles, pedestrians, or cyclists. This study aims to leverage the knowledge and techniques developed for this domain and evaluate its applicability to object detection within railway environments. Recent surveys on monitoring critical infrastructure using point cloud data indicate that most approaches do not leverage deep learning techniques. Instead, they predominantly rely on heuristic methods [4], [5]. A drawback of heuristic approaches is their dependence on numerous hand-crafted features and manually tuned parameters. To overcome this drawback, deep learning methods are able to automatically learn by providing them with ample examples. This also means that deep learning models are partially domain agnostic. For instance, this research uses models which have been developed for the road domain, but are now applied to the rail domain. This research aims to evaluate the applicability of deep learning-based object detection in the rail domain. Additionally, it assesses the locational accuracy of deep learning-based object detection models, a factor often overlooked or ignored in other rail domain publications. The difficulty in obtaining ground truth data for object locations, due to stringent safety protocols for track access, likely contribute to this oversight. Nevertheless, accurate locational data is crucial for planning construction or deploying automated construction robots. It is estimated that the required locational accuracy for creating a digital representation of the railway environment in the Netherlands is around ±5 cm 1. This digital representation should then be accurate enough for planning construction works without doing a manual inventory of the site. The remainder of this article is organized as follows: First, Section II presents the related work. Thereafter the custom dataset is presented in Section III together with the preprocessing steps required to prepare the data for ingestion by the machine learning models. This is followed by the methodology in Section IV. This section describes how the models are trained and evaluation, how the effectiveness of transfer learning is evaluated, and how the locational error is analyzed. Section V presents the findings, including two approaches that yielded negative results. Finally, Section VI concludes the article with recommendations. An earlier version of this work was previously published as part of B. Ton his PhD thesis [6]. Interested readers are encouraged to refer to this thesis to understand how the work presented here contributes to the broader goal of digitizing the rail environment. II. RELATED WORK This section reviews automated methods for deriving information from point cloud data to construct an as-is model of railway infrastructure. Most related works focus on specific 1Based on personal communications with Strukton Rail engineers tasks, such as detecting tracks or components of the overhead line equipment, like catenary wires. Zhu and Hyyppa [7] utilize both Airborne Laser Scanning (ALS) and Mobile Laser Scanning (MLS) data to model the railway environment in 3D. Their non-supervised approach employs a variety of parametric models fine-tuned for each element of interest, including ground, trees, buildings, catenary arches, and overhead power lines. Notably, their method extracts building facades from MLS data and roofs from ALS data. Pastucha [8] and Arastounia [9] present similar results, with Arastounia’s work offering a higher level of detail. Both studies use parametric models to detect tracks and overhead components. Arastounia’s research focuses on a 550 m stretch of Austrian railway track, achieving high accuracy metrics specific to that section. Pastucha’s work evaluates around 90 km of Polish railway track, detecting over 97% of support structures. Cheng et al. focus on creating an as-is Building Information Model (BIM) of single-track railway tunnels using Terrestrial Laser Scanner (TLS) data [10]. They heavily rely on parametric models to achieve a high-accuracy BIM in the mm–cm range. Wolf et al. also aim at automatically detecting railway assets from point cloud data [11]. Their approach involves rendering grayscale images from point cloud slices, with pixel values representing intensity values from the original point cloud. These slices are taken perpendicular to the rail track. They present results of both object detection, based on the YOLOv3 model [12], and on semantic segmentation, based on the U-Net [13] model. An image based approach has two major benefits; the field of image processing has advanced much further than point based methods and the processing of raster data can be done much more efficient compared to point data. It should be noted that their work is still in a very premature state. Pre-processing is a common step before modeling. In general the goal of this step is to cull the number of points, which is done to lessen the computation load. Often, this step also partitions a large scene into smaller tractable pieces. When collecting data from the rail environment using a mobile laser scanner, the trajectory of the measurement train is typically logged. This trajectory is valuable for partitioning the large point cloud scene into smaller tractable pieces. Soilán et al. [14] describe a trajectory log based dissection method using the roll, pitch and heading information of the measurement train, with each piece set to 3 m. Their work focuses on automated rail extraction. Lamas et al. also partition the data into smaller pieces, each with a length of 200 m and a width of 20 m, covering 90 km of track in total [15]. Each of the pieces is segmented with a high level of detail by using a heuristic-based workflow. Interestingly, Grandio et al. used the results of this work to train a fully supervised semantic segmentation model [16]. They also subjectively evaluate the generalisability of the trained model by applying it to different scenarios captured 2VOLUME 11, 2023 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS with various types of laser scanners. The model shows some generalisability but is heavily dependent on location and data quality. Intriguingly, to overcome the issue of poor generalisability of models, Wang et al. have added an extra module to the network which can recognize the type of railway scene before segmentation takes place. An exploratory ablation study indicates that by utilizing this prior knowledge the mIoU is increased significantly [17]. A proper data format for exchanging information between different stakeholders is essential. This aspect is often overlooked, but Soilán et al. suggest using the open Industry Foundation Classes (IFC) to ensure interoperability [14]. IFC is also used by Ariyachandra and Brilakis [18] to model overhead line equipment. Most related works make use of professional-grade laser scanners, such as the Riegl-VMX250. An exception is the work of Zou et al. which use a Velodyne VLP-16 laser scanner [19], the same sensor used in this study. The investment required for a professional-grade laser scanner is extremely high, which can be a bottleneck for railway contractors and railway operators. Using lower-priced laser scanners also allows for multiple scanners to be operational simultaneously, which is crucial for moving to a predictive maintenance paradigm that requires monitoring the infrastructural health over time. In summary, most works rely on professional highresolution laser scanners and employ semantic segmentation or parametric models to extract information from point cloud data. This work addresses this gap by using a consumer-grade laser scanner and, in contrast to the commonly used semantic segmentation task, explores the use of object detection models to detect railway assets. A major benefit of utilizing object detection models is that no ambiguous ‘background’ class needs to be defined, labeled, and learned. III. DATA As public point cloud datasets related to railway infrastructure were virtually nonexistent at the initiation of the project, a custom dataset has been collected. This custom dataset has been made available to the public [20]. This dataset was collected by a dedicated measurement train from Strukton Rail, the ‘Leonardo’. This measurement train is equipped with a variety of sensors, including a Velodyne VLP-16 laser scanner to capture point cloud data of the scene. Fig. 1 shows the sensor mounted center front of the train’s top. This sensor has a horizontal field of view of 360 degrees and a vertical field of view of 30 degrees, with a stated typical range accuracy of ±3 cm. Simultaneously with the laser scanner data, the location and orientation of the measurement train are recorded with an Applanix POS LVX sensor. The Applanix POSPac Mobile Mapping Suite was used to post-process the captured data. This post-processing step includes removing redundant data when the train was stationary and registering individual captured scenes into a larger scene. FIGURE 1: Leonardo measurement train. Red circle indicates the position of the LiDAR sensor. Image by Strukton Rail. As an alternative to using a mobile laser scanner, an aerial laser scanner was considered, but had several disadvantages. First, the operating costs of ALS are higher compared to MLS. Second, because the distance between scanner and scene is larger, the density of the point cloud is lower [21]. Third, the vantage point of ALS leads to a certain type of selfshadowing, which typically distorts the base of the object. The first dataset was collected on 14 June 2021 around the Deventer-Twello region, The Netherlands. The measurement train drove four times back and forth on the same piece of track. Unfortunately, the measurement train did not turn at the end of the section, but instead drove backward. This implies that objects were not captured from both directions. Each trip of the measurement train will be referred to as a run. In addition to the captured point cloud, the GPS location, heading, acceleration, and odometer of the train were also logged. An additional smaller dataset was collected around the city of Dronten, The Netherlands. This trajectory was captured on 16 November 2021 and consists of a single run. For both areas, the EPSG:32631 coordinate reference system (CRS) was used for the captured point clouds, and Coordinated Universal Time (UTC) was used for the timestamps. The corresponding trajectory logs of the measurement train were recorded using the EPSG:4258 CRS and made use of the Central European Summer Time (CEST, UTC+02:00) and the Central European Time (CET, UTC+01:00) respectively. Fig. 2 shows the trajectory of both measurement sessions. The average distance of a single run within the DeventerTwello region was ≈6.5 km, and the length of the Dronten trajectory was ≈2.9 km. The key figures of the different runs are summarized in Table 1. The distance value is based on the odometer data. The table also clearly shows the correlation between speed and the number of points collected. It can be seen that run 1 had the lowest average speed of 62.4 km/h and the highest point count, While run 4 had the highest average speed of 72.1 km/h and the lowest point count. The number of points processed will be significantly reduced during the preprocessing steps. These steps are elaborated in Section III-C. To prevent repetitive labeling, all four runs were merged VOLUME 11, 2023 3 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS (a) Deventer-Twello area. (b) Dronten area. FIGURE 2: Trajectories of measurement train indicated in red. Map data from OpenStreetMap. TABLE 1: Key figures of the different runs and the Dronten test set Duration [s] Distance [m] Speed [km/h] Points [-] run 1 373.6 6476.6 62.4 59,607,089 run 2 325.8 6479.3 71.6 52,834,759 run 3 346.4 6476.5 67.3 55,257,052 run 4 323.4 6479.2 72.1 52,423,814 Dronten 119.8 2867.3 86.1 16,532,528 before the labeling process. To separate the individual runs again after labeling, each point has an additional attribute ‘run’. A. LABELING Prior to labeling, the large point cloud scene is divided into sections of 250 m long using the same approach described in Section III-C1. This ensures that the sections are tractable and can be divided among multiple people performing the labeling task. As the point density of the scenes is not very high, labeling is very time-consuming, and given the explorative nature of this research, the following limited set of four object classes are defined: masts, tension rods, signals, and relay cabinets. Fig. 3 provides examples of each object class. To provide a scale perspective, an average point cloud human being with a height of 174 cm is plotted alongside the examples. Labeling the objects within the point cloud data was done using the open-source application CloudCompare. The segment tool was used for drawing polygons around the objects of interest. In general, an unobstructed top view of the object is possible, making it easy to segment the object by drawing a polygon around it from this viewpoint. After cutting out the object, a label and a unique identifier are added as a ‘Scalar Field’. Finally, all the individual pieces of the dissected scene were merged together again. The unique identifier has proven to be very useful when iterating over each of the individual objects during subsequent processing steps. The benefit of labeling points instead of just bounding boxes is that the dataset can serve both segmentation and object detection tasks. FIGURE 3: Samples of object classes, From top to bottom: relay cabinet, signal, tension rod, mast. Last three rows are scaled by 0.3. A human being with an average height of 174 cm is added for scale comparison. B. KEY FIGURES This section provides some key figures of the labeled datasets. Table 2 shows the class distribution and the mean and average number of points per object class. It can be seen that the standard deviation of number of points per object is large. This large variation can have several causes, such as the distance between sensor and object, or obstructions between sensor and object. TABLE 2: Occurrence (Occ.) of objects and their mean (M) and standard deviation (SD) of point count per object type First run Fourth run Dronten Occ. M (SD) Occ. M (SD) Occ. M (SD) Mast 209 992 (516) 203 892 (482) 124 983 (412) Tension rod 29 710 (395) 30 750 (589) 24 494 (152) Signal 20 930 (507) 22 810 (400) 9 854 (232) Relay cabinet 35 392 (175) 34 345 (184) 2 216 (14) The Dronten dataset shows a very low count of relay cabinets. This is attributed to presence of a sound barrier adjacent to large parts of the track. Relay cabinets usually reside on the other side of the sound barrier. This sound barrier obstructs the line-of-sight from the laser scanner, resulting in an extremely low count of relay cabinets. 4VOLUME 11, 2023 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS C. PRE-PROCESSING Before a model can be trained, an elaborate set of preprocessing steps is required to get the data in the right format so that it can be ingested by the model. These steps are described in the following subsections. To ensure that a local object location can be mapped to a real-world location again, it is necessary to keep track of the net transformation during each of the subsequent processing steps. 1) Partitioning The first step in the pre-processing chain is to partition the data into equal-length pieces. After the data are collected by the measurement train, they are post-processed. During this post-processing step, the non-stationary ‘frames’ are merged to form one large point cloud. The frames collected when the measurement train is stationary are omitted. To avoid large time gaps in the timestamps of the individual points, the time is shifted to compensate for the omitted frames. Unfortunately, this time shift is not reflected in the trajectory log, implying that the timestamps of the point cloud data and the trajectory log are not in sync. This has as consequence that the partitioning of the point cloud cannot be based on temporal information but has to be based on spatial information. The trajectory log plays an important role to define simple spatial polygons, which are then used as a cookie cutter to partition the large point cloud scene into smaller, tractable pieces. On several occasions, the measurement train drove multiple times on the same track or on an adjacent track. To distinguish individual trips, a region of interest can be defined and based on the ∆tof the timestamps, making it possible to distinguish them. A large ∆tmeans that the measurement left the region of interest and entered it again at a later time. A caveat with this approach is that the measurement train cannot change direction within the region of interest. Based on the timestamps from the trajectory log relating to a single trip, the cookie cutter polygons are defined. This is done by first calculating the cumulative distance based on the odometer data present in the trajectory log. Together with the desired length of the pieces, it is possible to determine the locations belonging to the start and end of a piece. These locations will be referred to as p1and p2. To define a cookie cutter polygon, the slope of the line between these two points is determined. Based on this slope, it is possible to determine the slope of a line perpendicular to the line passing through the two points (p1and p2). Given the desired extent of a piece and the equations of the perpendicular lines, it is possible to determine the four vertices of the polygon (vx,1,vx,2were x∈ {1,2}). Fig. 5 clarifies the procedure; they gray shaded area indicates the current polygon. The vertices v2,1and v2,2are shared with the next polygon. 2) Alignment The two points, p1and p2, from the previous step, which indicate the start and end position of a piece, are used for centering the scene and aligning the track along the x-axis. First, the intermediate point between p1and p2is determined and is used to set the translational part of an affine transformation. The rotational part of the affine transformation, which aligns the scene along the x-axis, is based on the angle between the x-axis and the point p1. After alignment, the point cloud is clipped in the ydirection to [−15 m,15 m]. Furthermore, a voxel based subsampling method is applied to homogenize the point density. In this scenario, a voxel size of 5 cm was used, within each voxel the centroid of the points is calculated. The calculated centroid is used to query the nearest neighbor within the voxel, which is then returned. A grid based sub-sampling strategy is also used by Grandio et al. [16]. Density variations occur naturally due to the fact that objects close to the sensor are captured with a higher spatial density compared to objects that are further away. 3) Ground Classification The next step in the pre-processing pipeline is to classify the scene into ground and non-ground points. This classification step serves two purposes: the first is to use the classified ground points to determine a transformation that aligns the scene to the xy-plane. The second is to reduce the amount of data being worked with by removing the ground points from the scene during training and inference. The ground classification algorithm used is based on cloth simulation [22], and the result of this classification is stored in the Classification field of the Point Data Record. The removal of the ground plane is also done by Ariyachandra and Brilakis to reduce computation time and false positives [18]. 4) Leveling The least-squares method is used to fit a plane to the classified ground points. The normal vector of this plane is used to determine a rotation matrix that levels the scene and makes it flush with the xy-plane. Special care is taken to ensure the normal of the plane is pointing towards the positive zdirection. If this is not the case, the scene would be flipped upside down. Level normalization is an important prerequisite for the next pre-processing step, determining bounding boxes. 5) Bounding Boxes In general, the task of object detection is to predict (oriented) bounding boxes of objects. This section describes how an oriented bounding box is obtained from a cluster of labeled points. A constraint imposed by the used models is that the top and bottom of the bounding box are flush with the xy-plane. An oriented bounding box is represented by a vector with seven elements. These elements are: the center point of the bounding box (x,y,z), the extent in all directions (ex,ey,ez) and a yaw angle (θ). The yaw angle is defined as the counterclockwise rotation around the positive z-axis. The processing steps to determine an oriented bounding box based on point clusters is detailed below: •Determine convex hull of the xy-coordinates of the cluster. VOLUME 11, 2023 5 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS FIGURE 4: Pre-processing steps. p1 p2 v1,1 v1,2 v2,1 v2,2 x-coordinate y-coordinate Train trajectory FIGURE 5: Partitioning of the point cloud scene. •Use a Singular Value Decomposition (SVD) of the xycoordinates of the hull to obtain the orientation of the object. •Rotate the cluster of points such that it aligns with the x-axis. •Determine the extent (ex,ey,ez) of the object in x,y,zdirections. •If ey >ex, swap ex and ey and add π/2to the yaw angle. The convex hull of the cluster of points is used to avoid point density variations altering the true orientation of the bounding box when the SVD is applied in the next step. The final step ensures that the longest side of the bounding box is always aligned with the x-axis. This is important when calculating the dimensions of the anchor boxes; otherwise, the dimensions of the anchor boxes would be distorted. Fig. 6 shows a sample scene. It shows the labeled points and the calculated bounding boxes. Anchor box statistics The PointPillar model requires an initial estimate of the bounding box sizes of the objects of interests. These estimates are referred to as anchor boxes. The anchor box dimensions are summarized in Table 3. TABLE 3: Dimensions (m) of the anchor boxes of the first run per object class width length height Mast 1.4 1.1 8.9 Tension rod 10.6 1.2 7.1 Signal 2.4 1.3 6.2 Relay cabinet 1.5 0.8 1.6 IV. METHODOLOGY Two successful deep learning-based 3D object detection models, the PointPillar model (2019) [23] and the CenterPoint model (2021) [24], from the domain of autonomous driving will be evaluated for the task of detecting objects within the railway environment. The PointPillar and CenterPoint object detection models are chosen for this application because they offer good performance in terms of average precision and have a low inference time, making them suitable for practical applications [25]. The CenterPoint model supports both a pillar-based discretization and a voxel-based discretisation of the scene. To ensure a fair comparison between both models, the pillarbased discretization is used. Both models utilize the SECOND backbone [26], which is an improved version of the VoxelNet backbone [27]. The main difference between the two models is the type of head. The CenterPoint model uses a two-stage approach: first, the center of the object is located (hence the name); second, a regression head is used to predict the bounding box and orientation. This implies that no prior information about the box size in the form of anchor boxes is required. The PointPillar model, on the other hand, uses a Single Shot Detector (SSD) to estimate bounding boxes in the 2D Bird’s Eye View (BEV). The 3D bounding box height, and elevation are additional regression targets. The original CenterPoint model includes velocity as a regression output; however, for this study, the velocity regression head has been removed completely from the model. The grid size used by the original CenterPoint model is (0.2×0.2m)or (0.32 ×0.32 m), depending on the dataset being evaluated. The grid size used by the original PointPillar mode is (0.16 ×0.16 m). In an ablation study, the authors of 6VOLUME 11, 2023 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS FIGURE 6: Sample scene. Red: masts, blue: tension rods, magenta: signals, yellow: relay cabinets, and black: background. Labeled points and calculated bounding boxes are visible. Best viewed in color. the PointPillar model show that the performance diminishes as the grid size increases. Therefore, for this study, a grid size of ≈(0.1×0.1m)has been used. This grid size has been chosen as a good balance between locational accuracy and memory consumption of the model. The resulting grid size was 736 ×304. A. EVALUATION METHOD All possible permutations for training, validation and testing are evaluated on the four runs of the Deventer-Twello dataset. This means that a of total 4! = 24 experiments are conducted. The model that performs best on the validation set is used for testing. The mean average precision (mAP) specified at an Intersection over Union (IoU) threshold of 0.5 and 0.75 are used as final metric. The results of all permutations are averaged and presented with the standard deviation. The software framework used for evaluating the model is non-deterministic. Hence, to reduce the variability of the validation and test results, the validation and testing datasets are iterated ten times when determining the AP values. As there is no standardized definition of calculating the mAP, the details of the calculation are provided here. The metric used in this work could be considered a vanilla type of mAP. When determining the precision and recall values, the confidence threshold is varied. In our work all encountered prediction certainty scores are used as confidence threshold values. The IoU values are volume-based. In addition to the overall average precision scores, the impact of point count for the mast class is also evaluated. This class was chosen as it is the most common class in the dataset. To evaluate the impact, the masts have been divided into four quantiles based on point count. For each quantile the recall curve is determined; the precision-recall is not considered, as it would be impossible to determine to which quantile a False Positive will contribute. For each confidence threshold, the recall (Equation 1) is calculated, where TP refers to the number of True Positives and FN refers to the number of False Negatives. The IoU threshold is set to 0.5. Recall =TP TP +FN (1) The final metric reported will be the Area Under Recall Curve (AUReC). B. TRANSFER LEARNING Labeling data is an expensive endeavor. Therefore, it is valuable to know how much effort is required to get good results from an unseen, new piece of track. Transfer learning is a successful technique to make use of an existing trained model and fine-tune it for new data [28], [29]. The Dronten dataset is used to get insights into how much new training is required to obtain good results. The dataset is manually split into four distinct subsets, ensuring that the distribution of classes is roughly equal among these subsets. The amount of new training data are varied, by taking 0%, 25%, 50%, and 75% of the Dronten dataset. The remaining data are used for testing. As there are very few examples of relay cabinets present in this dataset, this class if omitted for this experiment. As the dataset is too small to also have a proper validation set, the number of training epochs is fixed to 600. The learning rate is set to 1×10−4. The reported metric will be the mAP of the test set based on the weights of epoch 600. VOLUME 11, 2023 7 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS The number of possible permutations of training and testing sets for the case of 25% and 75% is four. For the case of 50% training data, the number of permutations is twelve. All permutations are trained and evaluated, and the average mAP is reported. An initial single model is trained on the Deventer-Twello dataset with the relay cabinet class ignored. This initial model is then fine-tuned using varying amount of training data from the Dronten dataset. C. LOCATIONAL ERROR The locational error between the ground truth location and the location predicted by the machine learning models can be decomposed into two components: the random error and the systematic error [30]. We will adhere to the terminology set out by the ISO 5725-1 standard [31]. This standard introduces the concept of trueness, which captures the systematical error, and precision, which captures the random error. Both error components combined constitute the accuracy. The errors are analyzed for the masts and cabinets together, along with an overall analysis. The other two asset classes, tension rods and signals, are not considered as it is not possible to define a precise origin for these two classes. For instance, in our case, the tension rods have been labeled together with their concrete foundation, while the reference data only labeled the rods themselves. A similar case holds for the signals; in our case, signals have been labeled together with the attached staircase, while the reference data did not. Evaluation of the locational accuracy of the assets detected by the machine learning model requires accurate ground truth data. Often, these data are not available as it requires considerable resources to obtain them, especially for longer stretches of track. As an alternative, the Dutch railway manager ProRail provides the absolute location of a large number of railway assets situated in the Netherlands. The locations of selected assets are publicly made available using a Web Feature Service (WFS) 2which can be queried. It is not known how these locations are obtained, but it is most likely done using photogrammetry, and the expected accuracy of these locations is unknown. Hence, the first step will be to evaluate the locational accuracy of the data provided by ProRail. To do so, a geodetic survey was carried out using a Leica GS18 GNSS RTK rover during the night of 28–29 January 2024. To get foot access to the tracks a special permit was required and the mandated personal protective equipment was used. The data sheet of this rover lists a mean horizontal accuracy of 7 mm and a mean vertical accuracy of 25 mm. Seventy-four objects (4 signals, 12 catenary rods and 58 masts) were manually surveyed and compared to the dataset provided by ProRail. We would like to emphasize that the rover does not collect point cloud data, but instead a single accurate location for each object is measured manually. The masts were all steel H-profile style masts, the center of the H on the ground plate has been used as measurement point. Based on mast locations, it was concluded that all objects 2https://mapservices.prorail.nl surveyed fall within a locational accuracy of 10 centimeters horizontal (µ8.7 cm, σ4.5 cm with a maximum error of 19.6 cm). Therefore, we consider the locations provided by ProRail a valid proxy source for the ground truth. The ProRail data are then used to evaluate the locational accuracy of the detected objects. For an object to be included in the evaluation, it must be within 3 meters of the location specified by ProRail. In the case of relay cabinets, which are usually clustered and reside within 3 meters of each other, a manual inspection was conducted to correctly match object detections to the ProRail-specified locations. Additionally, the detected object must have the correct label; otherwise, the detection was excluded. To get meaningful results when comparing locations of objects, both datasets must use the same origin. For the catenary masts and cabinets, the horizontal position of the origin is set to the center of these objects. The ProRail data do not include the vertical position (z-component) or the orientation of objects. Therefore, the decision was made to include only the horizontal components in the evaluation of the locational accuracy. The locational error is analyzed for the mast and cabinet asset classes. For each asset class, there are ngt ground truth locations, and for each ground truth location, there are nd,i detections from the models. In total, for a certain asset class, there will be nddetections. Let pibe a vector denoting a specific ground truth location, where iis an index over the number of ground truth locations for this asset class. Let ˜ pi,jbe another vector, specifying a certain detection corresponding to a ground truth location i, and jis an index over the number of detections for this ground truth location. 1) Trueness To determine the trueness, the systematic error vector δcan be determined based on the difference between the ground truth locations and the detections from the machine learning models. Equation 2 shows how this systematic error vector is formally defined. δ=1 nd ngt X i=1 nd,i X j=1 ˜ pi,j−pi(2) The trueness for an asset class is defined as the norm of the systematic error vector, that is ||δ||. 2) Precision The precision of the detections will be represented by the standard deviation of the measurements. The standard deviation will be represented by a vector σ, consisting of the elements σxand σy. As an example, the calculations for determining the standard deviation σxare provided, σyis derived in the same way. To determine the standard deviation σxof the detections corresponding to a ground truth location pi, first the arithmetic mean ˜µiof the detections ˜ pi,jis determined. Thereafter, Equation 3 shows how this mean value is used to determine the precision. 8VOLUME 11, 2023 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS σx=1 ngt ngt X i=1 sPnd,i j=1(˜ pi,jx−˜µix)2 nd,i (3) The overall precision is defined as ||σ||. 3) Accuracy The overall locational accuracy for a certain asset class will be measured using the Root Mean Square Error (RMSE) vector. This metric encompasses both trueness and precision. The RMSE vector for a single ground truth location iis composed of the individual mixand miycomponents. Equation 4 shows how the mixcomponent is calculated. The other component is calculated in similar fashion. mix=v u u t 1 nd,i nd,i X j=1 (˜ pi,jx−pix)2(4) Finally, the RMSE can be calculated using Equation 5. RMSEi=qm2 ix+m2 iy(5) The reported RMSE for a certain asset class will be the average of all RMSE values belonging to this asset class. D. TRAINING SETUP Due to memory consumption of the model during training, a batch size of one is used. Initially the model had difficulty converging; this is why for training a Stochastic Gradient Descent (SGD) optimizer is used instead of the default Adam optimizer. SGD provides a better generalization, with the penalty of having a slower training time [32]. A total of 3000 epochs are trained, with an initial learning rate of 1×10−3, which is reduced to 1×10−4after 1500 epochs. Every 50 epochs, the model is validated, and the weights corresponding to the highest mAP of the validation set are retained. These model weights are later used for testing. The models are trained on a desktop computer with 64 GB of RAM, an AMD Ryzen Threadripper 2950X CPU, and an NVIDIA TITAN V GPU with 12 GB of memory. V. RESULTS First, the object detection results are presented in terms of average precision. Additionally, the impact of point density on the recall rate of the catenary masts is presented. Thereafter, the effectiveness of transfer learning is analyzed. Because accurate locational information is essential when creating ‘asis’ models, an in depth analysis of the locational error is provided. Finally, two rejected hypotheses are presented for completeness. A. OBJECT DETECTION The aggregated object detection results are presented in Table 4. The table shows the average precision results averaged across the different experiments, along with the standard deviation. TABLE 4: Mean [email protected] and [email protected] results for all combinations of training, validation, and testing. Standard deviation in brackets IoU ≥0.5IoU ≥0.75 Asset PointPillar CenterPoint PointPillar CenterPoint Mast 78.5 (11.1) 72.8 ( 8.5) 16.3 (9.6) 7.5 (4.1) Tension rod 71.7 (11.9) 64.3 (14.1) 12.6 (8.7) 9.7 (4.8) Signal 83.5 ( 7.0) 82.1 ( 7.0) 16.3 (11.3) 14.0 (6.5) Relay cabinet 48.6 (11.5) 36.3 (13.1) 2.4 (2.5) 0.5 (1.2) mAP 70.6 (8.0) 63.9 ( 6.8) 14.6 (15.1) 10.3 (11.8) From the table, it can be seen that the performance quickly degrades when an IoU threshold of 0.75 is used. This is already an indication of poor locational accuracy, as the IoU directly depends on how well bounding boxes overlap. An in-depth analysis of the locational accuracy is provided in Section IV-C. The table also shows that the PointPillar model has the best performance for all classes, albeit with a slightly higher standard deviation. The results of determining the impact of point count on the recall rate are provided in Table 5. The first and last quantile show the largest spread in point count. The samples column specifies the number of samples falling within the listed interval. TABLE 5: Area Under Recall Curve (AUReC) for each quantile for the masts using an IoU threshold of 0.5 Quantile Interval Samples AUReC 1 [ 159, 629) 212 58.9 (11.2) 2 [ 629, 783) 210 67.5 (11.3) 3 [ 783, 1076) 214 62.8 ( 9.3) 4 [1076, 4486] 212 63.8 ( 7.4) Surprisingly, the data does not show a positive correlation between point count and the AUReC. This was expected, as a higher point count would also imply more information for the machine learning model to learn from. B. TRANSFER LEARNING The results of the transfer learning experiments are presented in Table 6. It should be noted that the results mentioned for the signal class are very imprecise, as the number of signals per split was very limited (1,2,2,4). TABLE 6: With and without Transfer Learning (TL) for varying amounts of training data. AP@IoU=0.5 No TL With TL 0% 25% 50% 75% 0% 25% 50% 75% Mast – 70.3 84.3 82.5 57.0 86.3 90.0 91.5 Tension rod – 9.5 42.4 60.0 21.0 70.8 71.8 74.3 Signal – 8.5 6.8 12.5 6.0 6.8 19.3 12.5 The table shows a clear benefit when using transfer learning. For instance, when using 25% of the training data, the AP for the tension rod class jumps from 9.5% to 70.8%. VOLUME 11, 2023 9 This article has been accepted for publication in IEEE Access. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/ACCESS.2025.3590779 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/