Full text
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 1 Thermal Infrared Video Stabilization for Aerial Monitoring of Active Wildfires Mario M. Valero, Steven Verstockt, Bret Butler, Dan Jimenez, Oriol Rios, Christian Mata, LLoyd Queen, Elsa Pastor and Eul` alia Planas Abstract—Measuring wildland fire behaviour is essential for fire science and fire management. Aerial thermal infrared (TIR) imaging provides outstanding opportunities to acquire such information remotely. Variables such as fire rate of spread (ROS), fire radiative power (FRP) and fire line intensity may be measured explicitly both in time and space, providing the necessary data to study the response of fire behaviour to weather, vegetation, topography and firefighting efforts. However, raw TIR imagery acquired by Unmanned Aerial Vehicles (UAVs) requires stabilization and georeferencing before any other processing can be performed. Aerial video usually suffers from instabilities produced by sensor movement. This problem is especially acute near an active wildfire due to fire-generated turbulence. Furthermore, the nature of fire TIR video presents some specific challenges that hinder robust inter-frame registration. Therefore, this paper presents a software-based video stabilization algorithm specifically designed for thermal infrared imagery of forest fires. After a comparative analysis of existing image registration algorithms, the KAZE feature-matching method was selected and accompanied by preand post-processing modules. These included foreground histogram equalization and a multireference framework designed to increase the algorithm’s robustness in the presence of missing or faulty frames. Performance of the proposed algorithm was validated in a total of nine video sequences acquired during field fire experiments. The proposed algorithm yielded a registration accuracy between 10 and 1000 times higher than other tested methods, returned 10x more meaningful feature matches and proved robust in the presence of faulty video frames. The ability to automatically cancel camera movement for every frame in a video sequence solves a key limitation in data processing pipelines and opens the door to a number of systematic fire behaviour experimental analyses. Moreover, a completely automated process supports the development of decision support tools that can operate in real time during an emergency. Index Terms—Fire behaviour, image registration, KAZE, remote sensing, UAS, video stabilization, wildland fire M.M. Valero, O. Rios, C. Mata, E. Pastor and E. Planas are with the Centre for Technological Risk Studies, Universitat Polit` ecnica de Catalunya, 08019 Barcelona, Spain (e-mail: mm.v[email protected]; chris- [email protected]; [email protected]; [email protected]) S. Verstockt is with IDLab, Ghent University – imec, 9502 Ghent, Belgium; (e-mail: ste[email protected]) B. Butler and D. Jimenez are with the Missoula Fire Sciences Lab, US Forest Service Rocky Mountain Research Station, Missoula, MT 59808, USA (e-mail: [email protected]v; [email protected]) L. Queen is with the National Center for Landscape Fire Analysis, University of Montana, Missoula, MT 59812, USA (e-mail: [email protected]) Manuscript received XXXX; revised YYYY. This research was supported by the Spanish Ministry of Education, Culture and Sport (PhD Grant FPU13/05876), the Spanish Ministry of Economy and Competitiveness (projects CTM201457448-R and CTQ2017-85990-R, funded with FEDER funds) and the Autonomous Government of Catalonia (2017-SGR-392). Furthermore, two research stays, completed by M.M.V., were funded by the Erasmus+ Traineeship Program and Obra Social La Caixa research mobility grants. I. INTRODUCTION AERIAL Thermal Infrared (TIR) imaging is widely used to acquire detailed spatial information about active wildfires. Data such as the location of the fire perimeter, its rate of spread (ROS), fire line intensity and fire radiative power (FRP) can be computed from geo-referenced thermal infrared footage [1]–[9]. This information has subsequently been used for additional fire behaviour analysis and evaluation of suppression activities [10]–[12]. Furthermore, there is a growing trend to incorporate observed fire perimeter evolution into operational fire spread simulators in order to improve forecasts using data assimilation techniques [13]–[17]. However, there are a number of unsolved issues in the data processing pipeline that prevent full exploitation of UAV potential for wildfire remote sensing. Limitations begin at the very first processing step with image registration. For aerial imagery to be useful, it must be geo-referenced so that image information can be projected onto geographic coordinates. In theory, there are a few alternatives to achieve this goal, but none has proved to be completely successful thus far. A few modern dedicated airborne monitoring systems, such as the NASA Ikhana UAV [18], [19], are able to geo-correct acquired imagery on board, automatically and even in real time, using high-accuracy positioning systems and powerful computing units. However, completely autonomous onboard processing is only possible if very accurate information about camera position and orientation is provided by highperformance Global Positioning Systems (GPS) and Inertial Measurement Units (IMUs). Although achieving sufficient location and orientation accuracy is technically feasible at present, the equipment capable of doing this is usually tailormade, heavy and expensive. For example, the Autonomous Modular Sensor (AMS), installed aboard the NASA Ikhana UAV, weighs 109kg [20]. Most frequently, data is acquired by less sophisticated platforms, usually consisting in commercial off-the-shelf (COTS) small unmanned aerial systems (UAS) or hand held cameras operated manually from a helicopter –see for instance [3], [5], [10], [21], [22]. In none of these studies could aerial footage be automatically georeferenced – i.e. registered onto a Digital Elevation Model (DEM). To date, thermal image geo-correction for wildfire research has required manual identification of Ground Control Points (GCPs) [2], [3], [6], [7], [9], [23], [24]. This approach is easy to implement and allows achieving high accuracy results under certain conditions, but it has very important limitations in practice. First, identifying and annotating GCPs in every video frame
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2 is extremely time consuming. A minimum of 4 GCP pairs are needed to estimate the projective transformation between two planes, and higher amounts of GCPs are required when remote sensing information must be projected onto non-flat terrain. Moreover, manual identification and annotation of GCPs is prone to errors in feature detection, localisation and matching. Resulting inaccuracies are hard to quantify and minimize due to the small amounts of feature matches typically available, which prevents the application of statistical methods. While the existence of these errors encourage the annotation of additional GCPs, reaching a sufficient amount of point matches is usually unfeasible. Furthermore, the user time required for point annotation imposes a limit to the maximum temporal resolution at which fire behaviour can be studied. Modern TIR imaging sensors are capable of recording high-frequency video with high spatial resolution and low hardware requirements [25]. There are a number of fire dynamics aspects that could be studied if highfrequency infrared (IR) video could be correctly geo-registered [26]. However, due to the cost of GCP annotation, video sampling frequency is usually reduced to some value between 1Hz and 0.1Hz, which is around 30 to 300 times lower than the technical limit. Finally, limiting the information used for registration to a few points per image may result in further unexpected data losses. Several GCPs may stop being identifiable for various reasons, and this may prevent image registration. For example, GCPs may disappear from the field of view or they may become occluded by trees, buildings or other aircraft. Similarly, a common practice in experimental controlled burns consists in using fire beacons as GCPs, but this approach may fail if beacons burn or blow out. Due to these limitations, image registration has recently been identified as one of the most significant sources of uncertainty when computing fire behaviour properties from aerial IR imagery [27], and several authors have reported significant challenges when registering, stabilizing and georeferencing aerial TIR video of active fires [6], [7], [20]. A promising strategy to solve existing limitations consists in automated video stabilization. Most of the times that remote sensing equipment is deployed during an active wildfire, the fire is recorded from a quasi-static overhead position. A nominal vantage point is selected and a camera is set in that position either in a tower or on hovering rotary-wing aircraft. With this configuration, and if the camera was totally still, GCP annotation in one single reference frame would suffice to georeference the complete video sequence. However, the wildfire environment is characterised by highly turbulent winds fostered by the interaction between atmospheric wind, topography and buoyancy originated from the fire [28]–[33]. This environment generates undesired displacements in any rotary-wing aircraft hovering near the fire [7], with small UAS being especially affected by aerodynamic turbulence. Mechanical stabilization systems can reduce camera movement, but they do not cancel it completely [25]. Video jitter has been observed even when installing IR cameras on ground boom lifts and towers [24]. On the contrary, if camera movement can be estimated and cancelled using software-based solutions, resulting still video becomes significantly easier to analyse both qualitatively and through computer vision algorithms. Stable footage can be georeferenced using a single geometric transformation, and the optimum transformation can be robustly estimated combining GCPs visible at different times. Video stabilization is a rather mature field in other remote sensing areas, and it has successfully been used in combination with GCPs to improve geo-correction accuracy of visible imagery [34]. However, to the best of our knowledge, this approach has not yet been extended to thermal IR imagery of fire. Stabilization of fire imagery acquired in the TIR range entails additional challenges due to limitations in image resolution and a lower level of detail. Because of the hightemperature measurement ranges required to observe fire, cold details in the image background are rarely visible in aerial TIR imagery. Furthermore, the fire itself, which is sometimes the only visible item, is highly dynamic, and this hinders the identification of persistent features in consecutive frames. Some authors have proposed the use of phase-correlation in the Fourier-Mellin space for registration of non-fire aerial TIR imagery [35]. However, we are not aware of any system capable of solving this issue in a wildfire scenario. In fact, recent review articles have concluded that, in the field of forest fire monitoring, a simple and low-cost image processing approach that can be used for image motion calculation and elimination is still in high demand [23]. Therefore, this paper presents a video stabilization algorithm specifically designed for aerial TIR imagery of wildfire. First, Section II describes the type of data to be processed. Afterwards, Section III details the study design we followed, the registration methods that were considered and the proposed video stabilization algorithm. Section IV presents and discusses the results of this study, which include data about registration accuracy, method robustness and video stabilization performance, plus a demonstration of algorithm deployment in three real study cases. Finally, Section V summarizes the contributions of this paper and future work. II. DATA The thermal infrared video used in this study was recorded during nine field experiments conducted between 2008 and 2018. Remote sensing data was acquired by different teams, using dissimilar setups and equipment. These nine fires were selected as a representative sample of the varied fire behaviour and imaging conditions present in wildfire research. In all scenarios, experimental fires propagated on flat terrain and fire evolution was recorded from vantage points using thermal infrared cameras. Still, the scenarios showed significant differences in experimental setup, fuels and spatial scale. This variability affected the observed fire behaviour. Appreciable differences occurred in flaming activity and rate of spread, both of which have an important impact on image content as well as its evolution with time. Moreover, the distance between camera and fire ranged from a few to more than a hundred metres. Different camera models, provided by various manufacturers, were used to monitor fire behaviour from different perspectives. Disparity in camera
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 3 resolution, together with variations in their distance to fire, contributed to inhomogeneous pixel sizes. Moreover, although all cameras worked in the same spectral range (long-wave infrared), radiance measurement ranges also differed among tests. The radiance measurement range selected to record a TIR video sequence has a critical impact on the amount of image detail. Background objects with brightness temperature below the selected measurement range are invisible, whereas zones with brightness temperature above the selected range are likely to cause sensor saturation. Both situations are highly detrimental to image processing techniques. Finally, video recording frequency varied between 1 and 32 frames per second. The experiment in Scenario 1 was conducted in 2018 at the Centre for Technological Risk Studies (Universitat Polit` ecnica de Catalunya - BarcelonaTech). A homogeneous bed of straw was burned over a 1.5mx3mcombustion table to reproduce fire spread on a flat horizontal surface with no wind. Scenarios 2and 3were recorded at the Tall Timbers Research Station in Tallahassee, FL, USA, in April 2017. These video sequences were acquired during a set of small-scale experimental burns on mixed rough/long leaf pine fuels. Scenarios 4,5and 6 correspond to tests S3, S4 and S5, respectively, pertaining to the Prescribed Fire Combustion and Atmospheric Dynamics Research Experiment (RxCADRE, FL, USA, 2012) [24], [36]. Burned vegetation was a mix of grass and shrubs, predominantly turkey oak. Finally, Scenarios 7,8and 9were recorded during another set of large-scale field experiments conducted in the Ngarkat Conservation Park, South Australia [11], [37]. These three video sequences correspond to a series of controlled burns in horizontal mallee-heath shrub plots with areas ranging from 4 to 25 ha. Despite the significant variability among test scenarios, video sequences 1-6 have an important property in common: they were all acquired from stable vantage positions with a clear overhead view of the fire. This property was exploited here for algorithm development and validation: the application of known synthetic vibrations to the stable video facilitated the quantitative measurement of image registration performance and therefore the comparison of different registration methods. Conversely, video sequences 7-9 were recorded from a hovering helicopter under circumstances similar to an actual wildfire. These datasets were used for the demonstration of the proposed video stabilization algorithm in independent footage with real jitter. Sample frames from each video sequence are displayed in Fig. 1, while Table I summarises the most relevant technical information about the deployed thermal cameras. III. METHODOLOGY Video stabilization software is primarily based on image registration. Owing to the wide range of applications of image registration and the variety of image types available, a good number of registration methods have been developed [38], [39], some of them specifically designed for remote sensing applications (e.g. [40]–[44]). However, to date there has been no detailed analysis of image registration in wildfire monitoring scenarios. As the first step to build our video stabilization algorithm, we conducted a comparative analysis of existing image registration techniques in order to assess their performance on wildfire TIR imagery. Subsequently, we selected the best-performing method and designed additional preand post-processing operations to optimize the accuracy and robustness of the algorithm. Section III-A summarizes the tested registration methods, Section III-B describes the study design and Section III-C details the implementation of our complete video stabilization algorithm. A. Comparative analysis of image registration methods Despite the wide variety of existing image registration methods, the majority of the algorithms follow one of three approaches: image similarity maximisation, phase correlation or feature matching. In this study, we included some of the most widespread methods within each of these categories. 1) Image similarity metrics and the optimisation problem: Assuming that two images are correctly aligned when their similarity is maximum, registration may be understood as an optimisation problem with a certain similarity metric as cost function. This constitutes the most generic approach with barely any need for a priori information. Image similarity can be estimated using various metrics, and how image similarity is measured has implications for registration algorithms. After a comparative analysis of various popular alternatives, our previous results encouraged the use of Mutual Information (MI) as similarity metric for TIR image registration in wildfire contexts [45]. MI is a widely adopted metric to measure image similarity during registration problems in remote sensing applications [41], [43]. It has favourable properties for image similarity measurement because it does not rely on any assumption about the nature of the images or their relation. Therefore, MI provides precise and reliable information about how much information two images share even if they were acquired by different sensors or under dissimilar lighting conditions. Furthermore, MI has showed sharper peaks than cross correlation at the point of correct alignment [46] and higher robustness in front of Gaussian noise and non-linear intensity relationships between images [43]. Nevertheless, Mutual Information itself is a similarity measure, not a registration technique. Similarity information is not accompanied by an estimation of the optimal transformation that should be applied to register two images. This limitation has traditionally made it challenging to find global optima [41]. Although some authors have addressed this issue [41], [47] and proposed mathematical and computational models that facilitate and accelerate convergence [43], poor convergence and high computation cost remain two critical limitations of MI-based registration algorithms. At present, there is active research aimed at improving the MI maximisation scheme [48]. 2) Phase correlation and the Fourier-Mellin Transform: Phase correlation was firstly proposed in [49] as an algorithm to align two images that were shifted with respect to each other. Like similarity maximisation algorithms, it uses
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 4 1.5 m (a) Scenario 1 30 m (b) Scenario 2 40 m (c) Scenario 3 100 m (d) Scenario 4 100 m (e) Scenario 5 100 m (f) Scenario 6 200 m (g) Scenario 7 300 m (h) Scenario 8 500 m (i) Scenario 9 Figure 1: Sample frames of the nine video sequences used in this study. Length values indicate approximate ground distances. Table I: Camera properties and parameters used to record the analysed footage. Scenario Camera commercial name Spectral range (wavelength, µm) Brightness temperature range (◦C) Image resolution (pixels) Field of view (◦) Recording frequency (Hz) 1 Optris PI 640 [7.5, 13] [20, 900] 640 x 480 60 x 45 32 2 Optris PI 400 [7.5, 13] [200, 1500] 382 x 288 60 x 45 27 3 Optris PI 400 [7.5, 13] [200, 1500] 382 x 288 60 x 45 27 4 FLIR SC660 [7.5, 13] [300, 1500] 640 x 480 45 x 30 1 5 FLIR SC660 [7.5, 13] [300, 1500] 640 x 480 45 x 30 1 6 FLIR SC660 [7.5, 13] [0, 550] 640 x 480 45 x 30 1 7FLIR AGEMA Thermovison 570-Pro [7.5, 13] [30 800] 240 x 320 24 x 18 4 8FLIR AGEMA Thermovison 570-Pro [7.5, 13] [30 800] 240 x 320 24 x 18 4 9FLIR AGEMA Thermovison 570-Pro [7.5, 13] [30 800] 240 x 320 24 x 18 4 all intensity information contained in the image. The main difference between phase correlation and other intensity-based methods resides in the mathematical space used for image comparison. By studying images in the frequency domain, phase correlation takes advantage of the Fourier shift theorem, which states that a displacement in the space domain is transformed into a change in phase in the frequency domain. Based on this property, relative translation between two identical images can be found uniquely without the need for an optimization algorithm. However, the phase correlation method is valid only for pure translations. In order to accommodate rotations and changes in scale, the authors in [50] proposed a revised algorithm which combines phase correlation with a log-polar transformation. The use of polar coordinates allows handling rotations as translations of the independent variable, while expressing the signal modulus in logarithmic scale also converts scaling to a translational movement. These properties allow applying the Fourier shift theorem in the different spaces in order to sequentially retrieve rotation, scaling and translation. Complete details about the mathematical algorithm and its original implementation can be found in [50].
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 5 The log-polar version of the Fourier transform, also known as Fourier-Mellin transform, has been extensively used in image registration, video stabilization and global motion estimation problems [51]–[55]. Moreover, its original formulation has been revised and coupled with additional algorithms to increase its robustness and reduce its computational requirements [53]. Therefore, registration based on phase correlation and the Fourier-Mellin transform was included in our comparative study. 3) Feature-based methods: In order to reduce the cost of a full 2D intensity comparison between two images, it is a common approach to limit the comparison to a number of sampled features. Usually, small local features that are invariant to scale and rotation are sought and characterized by a combination of descriptors that allow their comparison. Because feature description is highly dependent on the type of identified features, detectors and descriptors are usually coupled together. Widely used feature descriptors include the Scale Invariant Feature Transform (SIFT) [59], SpeedUp Robust Features (SURF) [63], Maximally Stable Extremal Regions (MSER) [64], [65] and KAZE [66]. Features detected in the images to be registered are then compared and matched. However, estimation of the correct geometric transformation from a set of feature matches might not be trivial. Due to the nature of feature detection methods, the amount of mismatches is usually sufficient to mislead leastsquares estimators [67]. A more robust approach that has been widely employed with success relies on random sampling, with the most famous algorithm of this type being the random sample consensus (RANSAC) [68]. RANSAC is an iterative algorithm in which certain hypotheses are generated from reduced sample subsets and subsequently verified with the complete data set. After its original publication in 1981, several revisions and computational optimisations have been proposed [67], [69], [70]. Feature matching has been successfully used in mediumaltitude UAS monitoring systems. The authors in [71] used edge detectors for precise image registration after a previous coarse motion estimation step. However, feature selection, description and matching vary greatly depending on image nature, which prevents the design of a universal feature-based algorithm. Fire TIR imagery poses important challenges for feature-based registration as points, edges and corners are not easily identifiable in TIR fire imagery. When they are, they usually correspond to flames, which are highly dynamic. Therefore, SURF, KAZE and MSER feature detectors and descriptors were included in this study because each of them is built upon different principles and low-level image characteristics. B. Study design In order to assess the performance of different image registration approaches in a wildfire scenario, we performed a systematic comparative study in a controlled environment. Registration accuracy was analysed through the application of synthetic movement to still TIR video. One hundred frames were randomly sampled from each of the video sequences 16. Afterwards, random similarity transformations were applied to the selected frames. Maximum perturbation ranges were set to ±20% of frame width and height for horizontal and vertical translation, respectively, ±25deg for rotation and ±20% for scaling. After perturbation, the transformed frames were registered back to their original position using each of the tested registration algorithms. Under normal working conditions, video stabilization is to be performed by registering each video frame with a previous –i.e. different– frame. However, it is difficult to conduct a generic analysis of registration performance under such realistic operation conditions as method response to the recording frequency may vary widely among registration strategies. In order to account for this, we compared the performance of selected registration algorithms under two different sets of conditions. First, we registered each perturbed frame back to its original position using the same original frame as reference for registration. Registering each frame with itself instead of a previous frame allowed decoupling the algorithm response to camera movement and recording frequency. We refer to this approach as idealised working conditions. Afterwards, we tested the same registration methods under realistic working conditions by sampling video frames at 1Hz and registering each perturbed frame with a previous stable frame, sampled 1 second before the perturbed frame. Registration accuracy was evaluated twofold. On the one hand, by comparing retrieved translations, rotations and scaling with the originally applied transformations. On the other, by computing similarity between the output registered frame and the original –target– frame. Image similarity was measured using 2D correlation as recommended in [45]. Each registered frame was evaluated using as reference the stable version of the same frame, both under idealised and realistic working conditions. Therefore, the maximum achievable correlation was always exactly 1. Furthermore, computation times required by each registration algorithm were measured for every frame and averaged over the complete study sample. These times were measured on a laptop computer equipped with an Intel i7-4700MQ CPU and 8.0 GB of RAM. In addition to registration accuracy, robustness was assessed for feature-based methods. Feature matching registration algorithms are based on the detection of relevant features in every frame of a video sequence and the pairing of these features in consecutive frames. Having a sufficient number of feature matches that conform to a single transformation is essential to obtain a reliable estimation of the relative movement between two images. Therefore, it is generally desirable to use a combination of feature detectors and descriptors that provides large amounts of point matches and a high inlier/outlier ratio. C. Video stabilization algorithm Based on the results obtained during the comparative analysis (see Sections IV-A and IV-B), the KAZE feature detection and description algorithms [66] were selected for inter-frame image registration. This section describes how KAZE feature matching was incorporated into a broader algorithm for video stabilization. First, we added an image preprocessing step to facilitate feature identification. Our approach is based on histogram
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 6 equalization, whose original implementation was adapted to account for the specificities of fire TIR imagery. In wildfire aerial remote sensing, it is frequent to use high temperature measurement ranges in order to avoid sensor saturation. This requirement limits the amount of cold details available in the image. Specifically, if the lower end of the temperature measurement range falls above the brightness temperature of the image background, the intensity value of all background pixels is set to the lower limit of the camera measurement range. Consequently, any background detail is lost. This fact has serious implications not only for image registration and georeferencing, but also for contrast enhancement. If the amount of foreground pixels is too small in comparison with the size of the background, histogram equalization may fail completely. Such failure occurred in scenarios 4 and 5 of our study. We worked around this limitation by applying the histogram equalization methodology only to the bounding box of the image foreground, which usually includes a balanced combination of foreground and background pixels. This approach was implemented following three steps: firstly, the image foreground was identified through pixel intensity segmentation; secondly, histogram equalization was applied to the minimum rectangular bounding box containing the selected foreground; finally, the processed foreground was merged back into the original background in order to maintain frame size. In addition to preprocessing, we propose the introduction of a filtering postprocessing step that increases the robustness of inter-frame registration. Regardless of how reliably image features can be detected and matched, sporadic registration failure is unavoidable. Extreme camera movements, vision occlusion by smoke or objects and hardware malfunction are some examples of the events that may produce faulty frames. Invalid frames should neither be registered, nor used as reference to register other frames. In order to account for faulty frames and prevent them from affecting the overall stabilization performance, we propose the use of a sliding multi-reference framework (Fig. 2). A frequent approach for video stabilization consists in registering each new frame with the latest that has already been stabilized. This approach provides high registration accuracy and is adaptable to changes in image content. However, it is not robust against the presence of faulty frames, especially if these are frequent. Instead, we suggest registering every new frame to the latest 5 stabilized frames, obtaining the 5 corresponding transformations that would map the new frame to the reference and using the median of these transformations for registration. The mathematical representation of 2D similarity transformations are 3x3 matrices. Therefore, the direct implementation of the proposed approach would consist in computing median values for every element in these matrices. However, this would result in more general transformation types not complying with similarity specifications. In order to keep a meaningful description of the applied registration, we suggest applying median filtering individually to the different movement components. Therefore, our algorithm retrieves the estimated translation, rotation and scale for all 5 reference frames, computes their median values and builds a new registration matrix with the resulting coefficients. IV. RESULTS AND DISCUSSION A. Registration accuracy Figures 3 and 4 summarise the differences between actual and estimated image misalignment in the form of pseudoBland-Altman plots. Bland-Altman plots are a common tool used to compare the accuracy of two methods that were designed to measure the same magnitude. This is accomplished by graphically displaying measurement differences between both methods along the complete range of measured samples. When ground truth sample values are unknown, their value is estimated as the average of measurements provided by both methods. Additionally, bias and limits of agreement are superimposed to the scatter plot. Bias is computed as the average difference, whereas limits of agreement are estimated as bias plus and minus 1.96 times the difference standard deviation [72]. Both bias and limits of agreement are accompanied by their respective 95% confidence intervals, which were computed here using the approximated estimations proposed in [73]. Confidence intervals are not always computed in the literature when using Bland-Altman plots, although they have been considered essential by some authors [74]. Because in our case ground truth data is known, Figs. 3 and 4 display pseudoBland-Altman plots where horizontal axes contain ground truth values instead of method output average. Tables II and III summarise mean squared errors in the estimation of individual motion components together with the global registration quality achieved by each method. These results are in agreement with Figs. 3 and 4 and show that the best accuracy was achieved by feature matching algorithms, especially when using the KAZE detector and descriptor together. Phase correlation systematically over-estimated translations (positive bias in pseudo-Bland-Altman plots, Figs. 3 and 4) and reached errors exceeding 100% of frame dimensions. MIbased registration showed better accuracy, especially for translations, although it was seriously affected by scale and errors remained high in general. Furthermore, both MI optimisation and phase-correlation tended to fail completely at moderately high rotation angles. In addition to KAZE, other feature-based algorithms achieved satisfactory performance. SURF feature matching proved especially capable of estimating rotation with high accuracy under idealised working conditions (MSErot = 0.0040). Nonetheless, Fig. 3 suggests that SURF successful performance was limited to light rotations. The accuracy of SURF rotation estimation diminished considerably when rotation angles approached ±20deg. Similar error behaviour was observed for MSER+SURF and MSER+KAZE combinations. Furthermore, these errors followed linear tendencies as observed in rotation plots of Figs. 3 and 4. This fact indicates that absolute errors in rotation estimation coincided with the real rotation value, which suggests that such errors were caused by complete registration failure. Although more subtle, similar failure tendencies were observed in translation and scaling estimation for SURF+SURF, MSER+SURF and
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 7 Frame i 5 last correctly stabilized frames 𝑀!= 𝑎!! 𝑎!" 𝑎!# 𝑎"! 𝑎"" 𝑎"# 𝑎#! 𝑎#" 𝑎## 𝑀" 𝑀#𝑀$𝑀% 𝑀!𝑀"𝑀#𝑀$ 𝑀% Registration matrices 𝑇&,! 𝑇&," 𝑇&,# 𝑇&,$ 𝑇&,% 𝑇(,! 𝑇(," 𝑇(,# 𝑇(,$ 𝑇(,% 𝜃!𝜃"𝜃#𝜃$ 𝜃% 𝑆!𝑆"𝑆#𝑆$ 𝑆% Individual movement components 𝑇&=𝑚𝑒𝑑𝑖𝑎𝑛(𝑇&,%, 𝑇&,!, 𝑇&,", 𝑇&,#, 𝑇&,$) 𝑇(=𝑚𝑒𝑑𝑖𝑎𝑛(𝑇(,%, 𝑇(,!, 𝑇(,", 𝑇(,#, 𝑇(,$) 𝜃 = 𝑚𝑒𝑑𝑖𝑎𝑛(θ%, 𝜃!, 𝜃", 𝜃#, 𝜃$) 𝑆 = 𝑚𝑒𝑑𝑖𝑎𝑛(𝑆%, 𝑆!, 𝑆", 𝑆#, 𝑆$) Filtered stabilization components Estimated registration transformation 𝑀) 𝑀) Faulty or missing frames Samples of the sought registration transformation Figure 2: Workflow of the proposed multi-reference stabilization algorithm. MSER+KAZE combinations. Conversely, KAZE algorithms proved robust in front of all applied perturbations. Finally, SURF suffered a larger decrease in performance under realistic operation conditions than KAZE (Tables II and III). Overall, feature matching approaches outperformed phasecorrelation and mutual information optimisation. Not only did they achieve better registration accuracy but they also ran significantly faster. Among feature-based methods, preliminary results suggested that KAZE was the most accurate and robust. However, the slight performance differences observed using random sampling encouraged a more detailed analysis, which is presented in section IV-B. Because registration failure was likely due to the lack of feature matches, special attention was paid to method robustness. B. Feature matching robustness There are two dominant factors that affect the amount of inlier image features effectively used by registration algorithms: image content and the relative position of compared images. We studied both by applying systematic transformations along the complete video duration in sequences 1-6.
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 8 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 Figure 3: Pseudo-Bland-Altman plots for motion estimation errors committed by the studied registration methods in idealised working conditions. MI = Mutual Information optimization; FM = Fourier-Mellin-based phase correlation. Motion estimation errors are plotted against ground-truth –i.e. actual– movement. Camera movement components are: translation in X direction (Tx), translation in Y direction (Ty), rotation (θ) and scaling. Txand Tywere normalised using frame width and height, respectively. Black dots correspond to individual frames randomly sampled along all studied scenarios and randomly perturbed. Red solid lines indicate mean bias. Red dashed lines account for Limits of Agreement (LoA), whereas red dotted lines represent 95% confidence intervals for estimated bias and LoA. Instead of limiting the study to a reduced set of randomly selected frames and perturbations, frames were evenly sampled at 1Hz for the complete duration of each sequence, which allowed assessing the response of feature detection to image content. In addition, synthetic translations, rotations and scaling were applied to each sampled frame sequentially and independently to observe the effect of each type of movement. Each distorted frame was then registered back to its original position. In this case, the previous stable frame was used as registration reference in order to reproduce normal working conditions. Fig. 5 shows the variation in feature matches with image content, whereas Fig. 6a displays its dependence on image alignment. The data shown in both of these figures correspond to the number of matches effectively used to estimate registration transformations, i.e. inliers accepted after RANSAC optimisation.
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 9 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -100 0 100 -20 0 20 -90 0 90 80 100 120 -100 0 100 Figure 4: Pseudo-Bland-Altman plots for motion estimation errors committed by the studied registration methods in realistic working conditions. MI = Mutual Information optimization; FM = Fourier-Mellin-based phase correlation. Motion estimation errors are plotted against ground-truth –i.e. actual– movement. Camera movement components are: translation in X direction (Tx), translation in Y direction (Ty), rotation (θ) and scaling. Txand Tywere normalised using frame width and height, respectively. Black dots correspond to individual frames randomly sampled along all studied scenarios and randomly perturbed. Red solid lines indicate mean bias. Red dashed lines account for Limits of Agreement (LoA), whereas red dotted lines represent 95% confidence intervals for estimated bias and LoA. In addition to the number of feature matches, there is a second variable representative of registration robustness. If the selected feature detectors and descriptors are not suitable for the type of images at hand, there are occasions when not enough features can be found, they cannot be matched across images or those connections do not follow a unique transformation. In such cases, the algorithm cannot suggest any registration at all. This fact can affect robustness of the complete video stabilization system, which may be able to handle the existence of missing frames only if they are scarce. We assessed this aspect using the percentage of frames for which each method was able to estimate a registration transformation (Fig. 6b). Note that Fig. 6b only evaluates the capability of providing a transformation estimation and it does not provide information about registration accuracy. Fig. 5 shows a strong dependence of the number of matches on the portion of the image filled with fire. Footage sections with a higher amount of successful feature matches correspond
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/JSTARS.2021.3059054, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 16 [38] L. G. Brown, “A survey of image registration techniques,” ACM Computing Surveys, vol. 24, no. 4, pp. 325–376, 1992. [39] B. Zitov´ a and J. Flusser, “Image registration methods: A survey,” Image and Vision Computing, vol. 21, no. 11, pp. 977–1000, 2003. [40] L. M. G. Fonseca and B. S. Manjunath, “Registration techniques for multisensor remotely sensed imagery,” Photogrammetric Engineering & Remote Sensing, vol. 62, no. 9, pp. 1049–1056, 1996. [41] H. M. Chen, P. K. Varshney, and M. K. Arora, “Performance of Mutual Information Similarity Measure for Registration of Multitemporal Remote Sensing Images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 41, no. 11 PART I, pp. 2445–2454, 2003. [42] Y. Bentoutou, N. Taleb, K. Kpalma, and J. Ronsin, “An automatic image registration for applications in remote sensing,” IEEE Transactions on Geoscience and Remote Sensing, vol. 43, no. 9, pp. 2127–2137, 2005. [43] J. P. Kern and M. S. Pattichis, “Robust Multispectral Image Registration Using Mutual-Information Models,” IEEE Transactions on Geoscience and Remote Sensing, vol. 45, no. 5, pp. 1494–1505, 2007. [44] A. Wong and D. A. Clausi, “ARRSI: Automatic registration of remotesensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 45, no. 5, pp. 1483–1492, 2007. [45] M. M. Valero, S. Verstockt, C. Mata, D. Jimenez, L. Queen, O. Rios, E. Pastor, and E. Planas, “Image Similarity Metrics suitable for Infrared Video Stabilization During Active Wildfire Monitoring: a Comparative Analysis,” Remote Sensing, vol. 12, no. 3, p. 540, 2020. [46] K. Johnson, A. Cole-Rhodes, I. Zavorin, and J. Le Moigne, “Mutual information as a similarity measure for remote sensing image registration,” in Proc. SPIE 4383, Geo-Spatial Image and Data Exploitation II, 2001. [47] J. P. W. Pluim, J. B. A. Maintz, and M. A. Viergever, “Image registration by maximization of combined mutual information and gradiant information,” IEEE Transactions on Medical Imaging, vol. 19, no. 8, pp. 809–814, 2000. [48] Y. Zhuang, K. Gao, X. Miu, L. Han, and X. Gong, “Infrared and visual image registration based on mutual information with a combined particle swarm optimization - Powell search algorithm,” Optik, vol. 127, no. 1, pp. 188–191, 2016. [49] C. Kuglin and D. Hines, “The Phase Correlation Image Alignment Method,” in Proceedings of the IEEE International Conference on Cybernetics and Society, 1975, pp. 163–165. [50] B. R. Reddy and B. N. Chatterji, “An FFT-based technique for translation, rotation, and scale-invariant image registration,” IEEE Transactions on Image Processing, vol. 5, no. 8, pp. 1266–1271, 1996. [51] J. Martinez-de Dios and A. Ollero, “A real-time image stabilization system based on fourier-mellin transform,” in Image Analysis and Recognition, 2004, no. October, pp. 376–383. [52] R. Goecke, A. Asthana, N. Pettersson, and L. Petersson, “Visual Vehicle Egomotion Estimation using the Fourier-Mellin Transform,” in 2007 IEEE Intelligent Vehicles Symposium, 2007, pp. 450–455. [53] S. Kumar, H. Azartash, M. Biswas, and T. Nguyen, “Real-time affine global motion estimation using phase correlation and its application for digital image stabilization.” IEEE transactions on image processing, vol. 20, no. 12, pp. 3406–18, 2011. [54] L. Lai and Z. Xu, “Global Motion Estimation Based on Fourier Mellin and Phase Correlation,” in International Conference on Civil, Materials and Environmental Sciences, 2015, pp. 636–639. [55] G. Rabatel and S. Labb´ e, “Registration of visible and near infrared unmanned aerial vehicle images based on Fourier-Mellin transform,” Precision Agriculture, vol. 17, no. 5, pp. 564–587, 2016. [56] C. Harris and M. Stephens, “A Combined Corner and Edge Detector,” in Procedings of Forth Alvey Vision Conference, 1988, pp. 147–151. [57] E. Rosten and T. Drummond, “Fusing Points and Lines for High Performance Tracking,” in Proceedings of the Tenth IEEE International Conference on Computer Vision, 2005. [58] ——, “Machine learning for high-speed corner detection,” in European Conference on Computer Vision, Leonardis A., H. Bischofm, and A. Pinz, Eds. Berlin, Heidelberg: Springer, 2006, pp. 430–443. [59] D. G. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004. [60] T. Kadir and M. B. Scale, “Saliency and Image Description,” International Journal of Computer Vision, vol. 45, no. 2, pp. 83–105, 2001. [61] F. Jurie and C. Schmid, “Scale-invariant shape features for recognition of object categories,” in Conference on Computer Vision and Pattern Recognition, vol. 2, 2004, pp. 90–96. [62] W. Ma, Z. Wen, Y. Wu, L. Jiao, M. Gong, Y. Zheng, and L. Liu, “Remote sensing image registration with modified sift and enhanced feature matching,” IEEE Geoscience and Remote Sensing Letters, vol. 14, no. 1, pp. 3–7, 2017. [63] H. Bay, A. Ess, T. Tuytelaars, and L. Van Gool, “Speeded-Up Robust Features (SURF),” Computer Vision and Image Understanding, vol. 110, no. 3, pp. 346–359, 2008. [64] J. Matas, O. Chum, M. Urban, and T. Pajdla, “Robust Wide Baseline Stereo from Maximally Stable Extremal Regions,” in British Machine Vision Conference, 2002, pp. 384–393. [65] D. Nist´ er and H. Stew´ enius, “Linear Time Maximally Stable Extremal Regions,” in European Conference on Computer Vision, 2008, pp. 183– 196. [66] P. F. Alcantarilla, A. Bartoli, and A. J. Davison, “KAZE Features,” in European Conference on Computer Vision, 2012, pp. 214–227. [67] P. H. S. Torr and A. Zisserman, “MLESAC: A New Robust Estimator with Application to Estimating Image Geometry,” Computer Vision and Image Understanding, vol. 78, no. 1, pp. 138–156, 2000. [68] M. A. Fischler and R. C. Bolles, “Random Sample Paradigm for Model Consensus: A Apphcatlons to Image Fitting with Analysis and Automated Cartography,” Communications of the ACM, vol. 24, no. 6, pp. 381–395, 1981. [69] H. Wang and D. Suter, “Robust adaptive-scale parametric model estimation for computer vision,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 26, no. 11, pp. 1459–1474, 2004. [70] B. J. Tordoff and D. W. Murray, “Guided-MLESAC: Faster image transform estimation by using matching priors,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, no. 10, pp. 1523– 1535, 2005. [71] H. Li, W. Ding, X. Cao, and C. Liu, “Image registration and fusion of visible and infrared integrated camera for medium-altitude unmanned aerial vehicle remote sensing,” Remote Sensing, vol. 9, no. 5, 2017. [72] J. M. Bland and D. G. Altman, “Statistical methods for assessing agreement between two methods of clinical measurement,” The Lancet, vol. 327, no. 8476, pp. 307–310, 1986. [73] ——, “Measuring agreement in method comparison studies,” Statistical Methods in Medical Research, vol. 8, no. 2, pp. 135–160, 1999. [74] A. Carkeet, “Exact Parametric Confidence Intervals for Bland-Altman Limits of Agreement,” Optometry & Vision Science, vol. 92, no. 3, pp. 71–80, 2015.