scieee AI-readable full text Open interactive document viewer

Deep Learning-Based Algorithms for Real-Time Lung Ultrasound Assisted Diagnosis

Muñoz, Mario,Rubio, Adrián,Cosarinsky, Guillermo,Cruza, Jorge F.,Camacho, Jorge

Abstract

This research was partially supported by the project PID2022-143271OB-I00, founded MCIN/AEI/10.13039/501100011033/FEDER, UE and by the fellowship PRE2019-088602 founded by MCIU (Spain).

Full text

Citation: Muñoz, M.; Rubio, A.; Cosarinsky, G.; Cruza, J.F.; Camacho, J. Deep Learning-Based Algorithms for Real-Time Lung Ultrasound Assisted Diagnosis. Appl. Sci. 2024,14, 11930. https://doi.org/10.3390/ app142411930 Academic Editors: Giacomo Cappon and Martina Vettoretti Received: 13 November 2024 Revised: 11 December 2024 Accepted: 14 December 2024 Published: 20 December 2024 Copyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). Article Deep Learning-Based Algorithms for Real-Time Lung Ultrasound Assisted Diagnosis Mario Muñoz 1,2,* , Adrián Rubio 1,2 , Guillermo Cosarinsky 1, Jorge F. Cruza 1and Jorge Camacho 1 1 Institute for Physical and Information Technologies, Spanish National Research Council, 28006 Madrid, Spain; [email protected] (A.R.); [email protected] (G.C.); jor[email protected] (J.F.C.); [email protected] (J.C.) 2Electronic Department, Universidad de Alcalá, 28805 Alcaláde Henares, Spain *Correspondence: [email protected] Abstract: Lung ultrasound is an increasingly utilized non-invasive imaging modality for assessing lung condition but interpreting it can be challenging and depends on the operator’s experience. To address these challenges, this work proposes an approach that combines artificial intelligence (AI) with feature-based signal processing algorithms. We introduce a specialized deep learning model designed and trained to facilitate the analysis and interpretation of lung ultrasound images by automating the detection and location of pulmonary features, including the pleura, A-lines, B-lines, and consolidations. Employing Convolutional Neural Networks (CNNs) trained on a semiautomatically annotated dataset, the model delineates these pulmonary patterns with the objective of enhancing diagnostic precision. Real-time post-processing algorithms further refine prediction accuracy by reducing false-positives and false-negatives, augmenting interpretational clarity and obtaining a final processing rate of up to 20 frames per second with accuracy levels of 89% for consolidation, 92% for B-lines, 66% for A-lines, and 92% for detecting normal lungs compared with an expert opinion. Keywords: lung ultrasound (LUS); artificial intelligence (AI); convolutional neural network; deep learning; pleura; B-line; A-line; consolidations; assisted diagnosis; real-time 1. Introduction Within the field of diagnostic imaging, lung ultrasound (LUS) has become significantly more prominent in recent years, offering a non-invasive, radiation-free approach for dynamic assessment of lung condition in various respiratory diseases [ 1 ]. Yet, despite its growing importance, the interpretation challenges inherent to pulmonary ultrasound images persist, which, combined with a shortage of skilled professionals trained in this technique, limits its application in clinical practice [2]. LUS images do not provide an anatomical view of the lung; this is because ultrasound waves hardly penetrate into the lung parenchyma due to the air present in the alveoli. Instead, most of the acoustic energy is reflected by the pleura, which is the most easily identifiable structure, appearing like a bright horizontal line. Furthermore, replicas of the pleura appear at regular intervals below this line, generated by the reverberation of the incident wave between the transducer and the pleura. This artifact, named A-Line, is indicative of good lung condition. In the presence of pneumonia and other interstitial syndrome diseases, the air inside the alveoli is progressively substituted by liquid, which modifies the acoustic impedance of the lung parenchyma, making it more similar to the impedance of muscle and fat above the pleura. As the acoustic impedance difference between both tissues reduces, more energy of the incident wave passes through the pleura to the lung parenchyma. But if air remains in some alveoli, a local reverberation phenomenon occurs (acoustic trap), and a bright vertical line appears in the image because of the multiple echoes backscattered inside the Appl. Sci. 2024,14, 11930. https://doi.org/10.3390/app142411930 https://www.mdpi.com/journal/applsci Appl. Sci. 2024,14, 11930 2 of 23 lesion [ 3 ]. This vertical artifact is named B-Line, and it is indicative of the presence of pneumonia. Figure 1shows some examples to better understand the morphology of the artifacts explained. Appl. Sci. 2024, 14, x FOR PEER REVIEW 2 of 25 bright vertical line appears in the image because of the multiple echoes backscattered inside the lesion [3]. This vertical artifact is named B-Line, and it is indicative of the presence of pneumonia. Figure 1 shows some examples to better understand the morphology of the artifacts explained. Figure 1. Typical LUS artifacts: Pleura (blue), A-lines (green), B-lines (orange), Consolidation (red). As the disease advances, more air is substituted by liquid and then consolidates into solid material. These consolidations appear in the image as hypoechoic regions and constitute the third artifact usually looked for when imaging the lung. In sum, artifacts like A-lines, B-lines, and consolidations are indicative of lung condition and can be used to diagnose a pulmonary pathology [4]. Detecting and interpreting these artifacts is the key for a correct evaluation of lung images, which has been the principal bottleneck in the dissemination of this technique. Moreover, the shock caused by the COVID-19 pandemic amplified the relevance of pulmonary ultrasound as a vital tool for evaluating pulmonary conditions. As the pandemic disseminated worldwide, LUS emerged as a frontline modality for assessing lung involvement and monitoring pulmonary condition, particularly in patients affected with COVID-19 pneumonia. Its real-time, bedside capabilities offered rapid insights into lung pathology, especially in emergency, critical care units, and domiciliary attention, where other imaging modalities like TAC and magnetic resonance encountered limitations or were less accessible [1,5–8]. In this context, the integration of diagnostic aid algorithms into ultrasound scanners could help reduce the learning curve of the technique, as well as reducing the evaluation time and possibly increasing the diagnosis success. Artificial intelligence (AI) has brought about significant advancements across diverse scientific and medical fields, showcasing its potential in reshaping the landscape of diagnosis and patient welfare. Within the realm of healthcare, AI is revolutionizing the field, showcasing promising outcomes through the provision of sophisticated tools that aid healthcare providers in making clinical decisions. This integration has led to improvements in precision and effectiveness, enhancing the process of diagnosing and treating a wide range of medical conditions [9]. The machine learning approach for detecting and classifying lesions is not exclusive to ultrasound imaging. In [10], the authors propose a Vision Transformer (ViT) based architecture able to classify melanoma versus non-cancerous lesions, which was tested on public skin cancer data with good results. In the X-Ray field, several machine learning (ML) approaches are proposed in [11] to distinguish between benign and malign tumors in mammography images, where the best results were obtained with a Naive Bayes algorithm. In [12], a review is provided of ML methods for lung cancer detection and classification with different imaging modalities (X-rays, computer tomography, and magnetic resonance), reporting a sensitivity between 0.81 and 0.99, a specificity between 0.46 and 1.00, and an accuracy from 77.8% to 100%, which reveals the potential of these techniques. Figure 1. Typical LUS artifacts: Pleura (blue), A-lines (green), B-lines (orange), Consolidation (red). As the disease advances, more air is substituted by liquid and then consolidates into solid material. These consolidations appear in the image as hypoechoic regions and constitute the third artifact usually looked for when imaging the lung. In sum, artifacts like A-lines, B-lines, and consolidations are indicative of lung condition and can be used to diagnose a pulmonary pathology [ 4 ]. Detecting and interpreting these artifacts is the key for a correct evaluation of lung images, which has been the principal bottleneck in the dissemination of this technique. Moreover, the shock caused by the COVID-19 pandemic amplified the relevance of pulmonary ultrasound as a vital tool for evaluating pulmonary conditions. As the pandemic disseminated worldwide, LUS emerged as a frontline modality for assessing lung involvement and monitoring pulmonary condition, particularly in patients affected with COVID-19 pneumonia. Its real-time, bedside capabilities offered rapid insights into lung pathology, especially in emergency, critical care units, and domiciliary attention, where other imaging modalities like TAC and magnetic resonance encountered limitations or were less accessible [1,5–8]. In this context, the integration of diagnostic aid algorithms into ultrasound scanners could help reduce the learning curve of the technique, as well as reducing the evaluation time and possibly increasing the diagnosis success. Artificial intelligence (AI) has brought about significant advancements across diverse scientific and medical fields, showcasing its potential in reshaping the landscape of diagnosis and patient welfare. Within the realm of healthcare, AI is revolutionizing the field, showcasing promising outcomes through the provision of sophisticated tools that aid healthcare providers in making clinical decisions. This integration has led to improvements in precision and effectiveness, enhancing the process of diagnosing and treating a wide range of medical conditions [9]. The machine learning approach for detecting and classifying lesions is not exclusive to ultrasound imaging. In [ 10 ], the authors propose a Vision Transformer (ViT) based architecture able to classify melanoma versus non-cancerous lesions, which was tested on public skin cancer data with good results. In the X-Ray field, several machine learning (ML) approaches are proposed in [ 11 ] to distinguish between benign and malign tumors in mammography images, where the best results were obtained with a Naive Bayes algorithm. In [ 12 ], a review is provided of ML methods for lung cancer detection and classification with different imaging modalities (X-rays, computer tomography, and magnetic resonance), reporting a sensitivity between 0.81 and 0.99, a specificity between 0.46 and 1.00, and an accuracy from 77.8% to 100%, which reveals the potential of these techniques. Several authors propose the application of AI-based algorithms with ultrasound lung images to address the challenges related to artifacts identification and quantification. There are usually two different approaches to this problem: processing a whole ultrasound video Appl. Sci. 2024,14, 11930 3 of 23 and classifying it according to the disease severity, or processing the video frame-by-frame, highlighting the artifacts that are found by labeling the image. The first approach means a classification problem, where the input is a buffer of images, and the output is a severity score. The second one is a segmentation problem, where visual feedback is given to the clinician about the type, position, and extent of the imaging artifact. Segmentation means to detect and isolate a region of interest inside an image, in this case, the A-Line, B-Line and consolidation artifacts, usually highlighting them with a color overlay on the conventional gray-scale image. This feedback would be useful for evaluating the image, but also improves the interpretability of the model, as the physician can confirm that the deep learning model is paying attention to the correct zone of the image. In this work, we propose to combine both approaches, providing frame-by-frame real-time identification and enhancement of artifacts, along with a score of the whole video based on the artifacts found. In [ 13 ], the authors propose a deep learning solution capable of assigning a severity score to a lung ultrasound video. They pose the problem as a classification problem where the network learns to assign a score based on the artifacts found in the input to the network. In [ 14 ], the authors design an algorithm to extract several images that serve as a summary of the ultrasound examination from a lung ultrasound video. These images are subsequently introduced to a neural classification network to assign a score. The network goes on to evaluate the classification problem on an image-by-image basis. In [ 15 ], the authors combine CNN with Long Short-Term Memory (LSTM) networks to extract special and temporal features to classify between healthy and non-healthy at video level. If we talk about image-by-image processing, in [ 16 ], the authors propose a deep learning model for B-line detection trained on images from Dengue patients as a classification problem. In [ 17 ], different solutions for B-line detection and localization are proposed: as a video-level and frame-level classification problem and as a frame-level segmentation problem. The authors conclude that frame-by-frame segmentation has an advantage over the other solutions because it results in less variability in image interpretability by different physicians. Other works, such as [ 18 ], focus on the detection of B lines by visualizing the activation map on the output of a binary classifier convolutional network. Here they contemplate the possibility of using such solutions in real time. The above-mentioned works focus on the detection of a single artifact, the B line, which means that other solutions are needed for the other artifacts typical of pneumonia and lung ultrasound. In [ 19 ], a multiclass segmentation model that distinguishes multiple artifacts is proposed instead, but unlike our work, the model is trained with images obtained from a lung phantom. In this work, we propose a novel complete workflow for computed assisted diagnosis in lung ultrasound imaging, aimed to be implemented in real-time. A specialized deep learning model was designed and trained to detect pulmonary features, including the pleura, A-lines, B-lines, and consolidations. Furthermore, a set of preand post-processing algorithms are proposed to normalize images to a common representation space and take into account a priori knowledge about the problem to enhance robustness. Another contribution of the work is a semi-automated labeling tool that could possibly contribute to extending datasets for further training of new deep learning models. Altogether, this work aims to take a further step towards the implementation of diagnostic aids in lung imaging on clinical scanners. 2. Methods 2.1. Neural Network Architecture Convolutional Neural Networks (CNNs) find extensive usage in real-time segmentation tasks due to their capability of recognizing local patterns and characteristics within images. They provide processing efficiency and a hierarchical feature structure enabling the identification of both intricate details and broader patterns. Specifically, the U-Net architecture stands out for its specialization in semantic and medical segmentation [ 20 ]. Its skip connections facilitate the fusion of information across various spatial scales, resulting Appl. Sci. 2024,14, 11930 4 of 23 in the creation of highly detailed segmented masks. This is particularly advantageous in medical imaging, like lung consolidation segmentation. With the aim of improving the behavior and response of our model, various U-Net network modifications were studied, including the standard U-Net architecture. A study was performed using the classical U-Net; the results can be found in [ 21 ], where it was shown that the performance was inferior compared to the architecture with attention blocks presented in this paper. For the modifications, we used the Python library Keras-unet, developed by K. Zak and published on GitHub [ 22 ]. Finally, the chosen architecture is Attention U-Net, developed by O. Oktay in [ 23 ], where attention gates filters are added in the decoder part, which automatically learns to focus on main patterns with different shapes and sizes, improving the result of the prediction. To find a compromise between computational cost, performance, and network complexity, we chose to use 16, 32, 64, and 128 filters on the encoding layers and the inverse for the decoding layers. Higher filter counts increase the model’s complexity and resource needs significantly, which may not translate into better performance for our task. Two implementation options were explored with the aim of seeing what is more suitable for the described problem: A different model for each pneumonia pattern, and a single model capable of predicting and segmenting all main lung ultrasound patterns (Pleura, A-line, B-line, Consolidation). Both solutions present advantages and disadvantages that were taken into account. The use of several models offers a more specific response and the possibility to tune the network according to each pattern to obtain the best behavior, but, on the other hand, when using several models, the inference time increases, making the real-time implementation more difficult. While using one global network with multiple outputs, the response could be less precise if we compare each output with a network specifically trained for that pattern, but the main advantage is that the inference time is reduced, thereby allowing real-time implementation. 2.2. Dataset The data used to perform this study were obtained in the clinical trial ULTRACOV [ 24 ]. They were generated with a 128-channel ultrasound electronic equipment developed in-house and a medical grade 3.5 MHz convex probe. A total of 689 videos were acquired, corresponding to 30 patients and following a standardized scanning procedure of 12 thoracic regions [25]. The videos were manually annotated by an expert physician, indicating the presence or absence of each artifact. The labeling process was performed in two steps: first, the physician marked the presence or absence of each artifact in the scanner software during the examination, reducing the bias that could be introduced when reviewing the videos of all the patients in a single session. Then, a review process was performed off-line to discard errors and finally validate the dataset labeling. Then, a score for the patient was calculated, taking into account the videos of the 12 scanned regions. Because the videos were not labeled frame-by-frame by the expert, a semi-automated annotation algorithm was developed (Section 2.2.2). 2.2.1. Neural Network Input Data The format of the input data has a great impact on the deep learning model definition and performance. A typical ultrasound image with a curved array (like that used in this study) is a circle sector, defined by an aperture angle α , an initial range r 1 , and a final range r 2 , which is usually framed into a rectangular image with size W × H pixels (Figure 2Left). But, in fact, a sector ultrasound image is originally formed by N scan lines usually equally distributed inside the sector area. These lines, containing M samples each, are the output of the beamforming algorithm, and could be interpreted themselves as a rectangular image (Figure 2Right). The algorithm to obtain the sector image from the B-Scan data is usually called scanconverter, and it is typically implemented by bi-linear interpolation of the acquired samples Appl. Sci. 2024,14, 11930 5 of 23 over the pixel grid. This process is carried out by the scanner, which gives the user the sector image in W × H format. Therefore, a question arises about which image format is more appropriate for implementing the deep learning algorithm aimed at in this work. Appl. Sci. 2024, 14, x FOR PEER REVIEW 5 of 25 2.2.1. Neural Network Input Data The format of the input data has a great impact on the deep learning model definition and performance. A typical ultrasound image with a curved array (like that used in this study) is a circle sector, defined by an aperture angle α, an initial range r 1 , and a final range r 2 , which is usually framed into a rectangular image with size W × H pixels (Figure 2 Left). But, in fact, a sector ultrasound image is originally formed by N scan lines usually equally distributed inside the sector area. These lines, containing M samples each, are the output of the beamforming algorithm, and could be interpreted themselves as a rectangular image (Figure 2 Right). Figure 2. Possible representations of an ultrasound sector image: (left) conventional pixel-based image given by ultrasound scanners; (right) B-Scan rectangular image formed by ultrasound samples only, without geometrical information of the probe. The algorithm to obtain the sector image from the B-Scan data is usually called scanconverter, and it is typically implemented by bi-linear interpolation of the acquired samples over the pixel grid. This process is carried out by the scanner, which gives the user the sector image in W × H format. Therefore, a question arises about which image format is more appropriate for implementing the deep learning algorithm aimed at in this work. Sector image has the advantage of being more easily accessible, because it is the typical output format for most ultrasound equipment. Therefore, it maintains the aspect ratio of the structures to be imaged, which eases interpretation by medical professionals. However, accommodating a sector image inside a rectangular grid generates black margins around it, which, besides adding pixels with no information, could potentially introduce a bias in the automated image analysis. Furthermore, the shape and extension of these zones depends on the scanner model and the configuration, hindering the translation of the resultant model between different scanners. On the other hand, rectangular B-scan images have the advantage of providing only useful information (no black margins), while its rectangular format is highly suitable as input for segmentation models. Another important advantage is that each vertical line represents a physical propagation direction of the beam inside the tissue, which is particularly relevant for artifacts like B-Lines, that appear precisely on those directions. Therefore, a B-Line will be always seen in the B-Scan as a vertical artifact, independent of its Figure 2. Possible representations of an ultrasound sector image: (left) conventional pixel-based image given by ultrasound scanners; (right) B-Scan rectangular image formed by ultrasound samples only, without geometrical information of the probe. Sector image has the advantage of being more easily accessible, because it is the typical output format for most ultrasound equipment. Therefore, it maintains the aspect ratio of the structures to be imaged, which eases interpretation by medical professionals. However, accommodating a sector image inside a rectangular grid generates black margins around it, which, besides adding pixels with no information, could potentially introduce a bias in the automated image analysis. Furthermore, the shape and extension of these zones depends on the scanner model and the configuration, hindering the translation of the resultant model between different scanners. On the other hand, rectangular B-scan images have the advantage of providing only useful information (no black margins), while its rectangular format is highly suitable as input for segmentation models. Another important advantage is that each vertical line represents a physical propagation direction of the beam inside the tissue, which is particularly relevant for artifacts like B-Lines, that appear precisely on those directions. Therefore, a B-Line will be always seen in the B-Scan as a vertical artifact, independent of its position on the pleura and of the scanner configuration. Furthermore, scanner configuration parameters like the aperture angle α or the initial and final range r 1 and r 2 only affect the size M × N of the image, which simplifies adapting images acquired with different configurations to the same network, only by vertical and horizontal scaling. On the other hand, these images are not suitable for visual interpretation, as they present a distorted view of the tissue anatomy. Because not all ultrasound equipment provides direct access to B-Scan data, a Sector-Image-to-B-Scan conversion algorithm would be needed. With a similar approach than scan-converter algorithms from B-Scan to sector image, it could be based on a simple bilinear interpolation algorithm after defining a set of beam lines that cover the useful area of the sector image (green lines in Figure 2Left). In this work, we had access to the B-Scan raw data generated by our system, so no SectorImage-to-B-Scan process was needed. Appl. Sci. 2024,14, 11930 6 of 23 Furthermore, for possible future hardware implementations of the proposed models, using B-scan data would be optimal in the sense that it reduces the amount of information to be processed, and uses data in a raw format available at low hardware level. For defining the size of the B-Scan images, the typical number of beams in an ultrasound scanner must be taken into account. For convex probes like the one used in this work, an active sub-aperture is used for generating each beam, typically with between 32 and 64 elements. That sub-aperture is moved along the array in one or several elements step, forming the B-Scan. For example, a 128 elements array with a 32 element aperture can generate up to 96 scan lines, while a 196 elements array with an aperture of 64 elements generates 132 scan lines. Based on these typical numbers, an input width of 128 lines was selected, being power of 2 for operations optimization. The size of the vertical direction is related to the frequency content of the signal and the sampling rate. For an ideal 100% bandwidth array, the maximum frequency content of the signal envelope is equal to half the array center frequency. To reconstruct the envelope without aliasing, the pixel density in the vertical direction should be, at least, able to sample the signal at double of that frequency (Nyquist criteria). For the array used in this work with 3.5 MHz center frequency and 70% bandwidth, the number of pixels required for sampling up to 70 mm and 90 mm is 210 and 294, respectively. Based on these numbers, a height of 256 was selected for the network, which imposes a trade-off between training and inference cost and image quality. In cases when larger images are used, they should be scaled down using compression algorithms that preserve the artifacts’ information. For example, in the scanner used in this work, a data reduction algorithm without peak information losses is available [ 26 ] and was used to accommodate the B-Scan height to the network size. In Figure 3, a comparison example between sectorial and B-scan images is shown. Appl. Sci. 2024, 14, x FOR PEER REVIEW 6 of 25 position on the pleura and of the scanner configuration. Furthermore, scanner configuration parameters like the aperture angle α or the initial and final range r 1 and r 2 only affect the size M × N of the image, which simplifies adapting images acquired with different configurations to the same network, only by vertical and horizontal scaling. On the other hand, these images are not suitable for visual interpretation, as they present a distorted view of the tissue anatomy. Because not all ultrasound equipment provides direct access to B-Scan data, a Sector-Image-to-B-Scan conversion algorithm would be needed. With a similar approach than scan-converter algorithms from B-Scan to sector image, it could be based on a simple bilinear interpolation algorithm after defining a set of beam lines that cover the useful area of the sector image (green lines in Figure 2 Left). In this work, we had access to the B-Scan raw data generated by our system, so no SectorImage-to-B-Scan process was needed. Furthermore, for possible future hardware implementations of the proposed models, using B-scan data would be optimal in the sense that it reduces the amount of information to be processed, and uses data in a raw format available at low hardware level. For defining the size of the B-Scan images, the typical number of beams in an ultrasound scanner must be taken into account. For convex probes like the one used in this work, an active sub-aperture is used for generating each beam, typically with between 32 and 64 elements. That sub-aperture is moved along the array in one or several elements step, forming the B-Scan. For example, a 128 elements array with a 32 element aperture can generate up to 96 scan lines, while a 196 elements array with an aperture of 64 elements generates 132 scan lines. Based on these typical numbers, an input width of 128 lines was selected, being power of 2 for operations optimization. The size of the vertical direction is related to the frequency content of the signal and the sampling rate. For an ideal 100% bandwidth array, the maximum frequency content of the signal envelope is equal to half the array center frequency. To reconstruct the envelope without aliasing, the pixel density in the vertical direction should be, at least, able to sample the signal at double of that frequency (Nyquist criteria). For the array used in this work with 3.5 MHz center frequency and 70% bandwidth, the number of pixels required for sampling up to 70 mm and 90 mm is 210 and 294, respectively. Based on these numbers, a height of 256 was selected for the network, which imposes a trade-off between training and inference cost and image quality. In cases when larger images are used, they should be scaled down using compression algorithms that preserve the artifacts’ information. For example, in the scanner used in this work, a data reduction algorithm without peak information losses is available [26] and was used to accommodate the B-Scan height to the network size. In Figure 3, a comparison example between sectorial and B-scan images is shown. (a) (b) Figure 3. Comparison between sectorial image (a) and B-Scan (b). Figure 3. Comparison between sectorial image (a) and B-Scan (b). 2.2.2. Neural Network Output Data Labelling Tools One of the limitations of the used dataset [ 24 ] is that it was labeled by an expert physician at video level, while for training a segmentation network, it must be labeled at frame level. While asking a physician to label such a large dataset frame-by-frame is usually impractical, a semi-automated algorithm was developed for this task, based on the initial label of the presence of each artifact in the video. Depending on the type of artifact, this process is fully automated or requires to identify a key frame where the artifact is visible and manually segment it, being then automatically followed in the subsequent frames in the video. These algorithms are explained in the following subsections. Pleural Line Labelling One of the most important indications to identify in a lung ultrasound image is the pleura, because it defines the zone where the artifacts must be looked for. Correctly identifying the pleura helps to discard zones that do not have to be analyzed, like the Appl. Sci. 2024,14, 11930 7 of 23 ribs and their shadows and the muscle and fat area above it. Giving this information to the neural network could improve the results, as it will learn that the B-Lines and consolidations are located bellow the pleura, and not in other areas of the image. The algorithm for automatically segmenting the pleura is based on the fact that when the ultrasound probe remains stationary, the image above the pleura (fat and muscle) remains basically unchanged between frames, while variations in the image occur below the pleura due to the respiratory cycle. Then, by subtracting consecutive frames, an auxiliary image is formed, where the upper region is almost black and the lower part is bright. Furthermore, if several of these images are averaged, the random nature of the speckle in the lower part generates a quite homogeneous region that can be more easily distinguished from the upper region. The frontier between these regions can be considered a first approximation of the pleura line. Given a video V, formed by Kframes with Bdefined by: V={B1,B2, . . . , BK}(1) where Bi=[bi(m,n)] ∈RMxN f or 1≤i≤K, 1 ≤m≤M and 1≤n≤N(2) and b i (m,n) is the pixel intensity of image iat position (m,n), the output of the proposed algorithm is: e V=ne B1,e B2, . . . , e BK−Lo(3) where e Bi=1 L k=i+L−1 ∑ k=i |Bi+1−Bi|(4) with Lbeing the length of the averaging filter. Figure 4shows an example output of the algorithm. On the left is the original B i image of a frame, and on the right is its filtered version V i , showing the two formerly mentioned dark and bright zones. For detecting their boundary, a dynamic threshold was applied to each vertical line of the filtered image, calculated from the average amplitude of the upper part of the image: e Pj=mini|Aj(i)>thj(5) where Ajis the vertical line of the filtered image at column j,e Pjis the vertical index to the pleura guess for that line, and th j the applied threshold. Figure 4b shows, with green dots, an example of this set of initial pleura raw guess points e Pn , which are then used to fit a second order polynomial to obtain a smooth representation Pn of the initial guess (red line in Figure 4b). Appl. Sci. 2024, 14, x FOR PEER REVIEW 8 of 25 fit a second order polynomial to obtain a smooth representation 𝑃 of the initial guess (red line in Figure 4b). It is worth mentioning that directly applying a threshold to the original image (Figure 4a) is very prone to misdetections, as other horizontal bright lines in the upper region of the image are present, with pixel intensity values even larger than those of the pleura. On the other hand, the auxiliary image filtered by the proposed method is very suitable for a simple threshold detection because of its stepped nature. (a) (b) Figure 4. First pleura approximation example. (a) original B-scan image; (b) filtered image. This first approximation does not take into account that a bright line has to be seen in the image to be considered part of the pleura line. Then, the P n indexes are used as the center of a vertical window with W samples, where to look for the maximum amplitude of the original image. If this amplitude is larger than a global threshold T, it is considered that the pleura is present in the window. Then, a −3 dB signal drop criteria is used to find the high amplitude zone around the maximum, and those pixels are labeled as belonging to the pleura in a binary mask called M. The mathematical formulation of this part of the algorithm is as follows: 𝑀󰇛𝑚,𝑛󰇜=󰇫1, 𝑖𝑓 𝑏󰇛𝑚,𝑛󰇜 𝐴 󰇛𝑚󰇜⋅10   ; 𝑓 𝑜𝑟 𝑛 ∈󰇟𝑃 −𝑊, 𝑃 𝑊󰇠 0, 𝑜𝑡ℎ𝑒𝑟 (6) Keeping the previous nomenclature, this condition checks if the amplitude at position (m, n) remains within 3 dB of the maximum amplitude at 𝑃  (𝐴󰇜, which corresponds to an approximate 29% drop in amplitude. If this condition is met, the pixel is considered part of the pleura and is marked as 1 in the binary mask. Otherwise, it is marked as 0. For this particular dataset, the approximation is applied in a searching window W = 20 samples around the pleura maximum amplitude identified in each scan line, but W should be adjusted according to the image resolution to include a region of about 6 mm around the pleura. Finally, morphological opening and closing filters sized at 3 × 3 are applied to refine and smooth the pleura labeling mask (Figure 5c), eliminating outliers and giving a more precise representation of the higher amplitude line. For each video, a set of K-L pleura binary mask images are generated, to be used during the training process of the network. Figure 4. First pleura approximation example. (a) original B-scan image; (b) filtered image. Appl. Sci. 2024,14, 11930 8 of 23 It is worth mentioning that directly applying a threshold to the original image ( Figure 4a ) is very prone to misdetections, as other horizontal bright lines in the upper region of the image are present, with pixel intensity values even larger than those of the pleura. On the other hand, the auxiliary image filtered by the proposed method is very suitable for a simple threshold detection because of its stepped nature. This first approximation does not take into account that a bright line has to be seen in the image to be considered part of the pleura line. Then, the P n indexes are used as the center of a vertical window with Wsamples, where to look for the maximum amplitude of the original image. If this amplitude is larger than a global threshold T, it is considered that the pleura is present in the window. Then, a − 3 dB signal drop criteria is used to find the high amplitude zone around the maximum, and those pixels are labeled as belonging to the pleura in a binary mask called M. The mathematical formulation of this part of the algorithm is as follows: Mi(m,n)=(1, i f bi(m,n)≥Amax(m)·10−3 20 ;f or n ∈he Pj−W,e Pj+Wi 0, other (6) Keeping the previous nomenclature, this condition checks if the amplitude at position (m,n) remains within 3 dB of the maximum amplitude at e Pj(Amax) , which corresponds to an approximate 29% drop in amplitude. If this condition is met, the pixel is considered part of the pleura and is marked as 1 in the binary mask. Otherwise, it is marked as 0. For this particular dataset, the approximation is applied in a searching window W = 20 samples around the pleura maximum amplitude identified in each scan line, but W should be adjusted according to the image resolution to include a region of about 6 mm around the pleura. Finally, morphological opening and closing filters sized at 3 × 3 are applied to refine and smooth the pleura labeling mask (Figure 5c), eliminating outliers and giving a more precise representation of the higher amplitude line. For each video, a set of K-L pleura binary mask images are generated, to be used during the training process of the network. Appl. Sci. 2024, 14, x FOR PEER REVIEW 9 of 25 (a) (b) (c) Figure 5. Post-processing adjustment of automatic pleural annotation. (a) B-scan image; (b) initial pleural annotation approximation; (c) final pleura annotation. Figure 6 summarizes the workflow for the pleural line automatic annotation algorithm. Figure 6. Automatic pleural line annotation flowchart. A-Line Labelling In case of A-lines, the developed algorithm detects these patterns as echoes of the pleura occurring at depths that are multiples of the probe-to-pleura distance. A search window W 󰇟𝑤,𝑤󰇠 is applied at those depths to find signal peaks that cross a threshold, and then verify their validity by comparing with the average amplitude in consecutive lines: 𝑎=arg max ∈󰇟  ,  󰇠 𝐴 󰇛𝑥󰇜 (7) where 𝑎 corresponds to the index in the A-scan signal 𝐴 where the maximum amplitude is located within the window 󰇟𝑤,𝑤󰇠, which corresponds with an initial guess of the A-Line position. As in the case of the pleura line, to define the ground truth mask, a −3 db signal drop criteria equivalent to the Formula (6) is used, followed by the application of opening and closing morphological filters sized at 3 × 3 to refine the mask and eliminating outliers. Figure 7 summarizes the workflow for the A-lines automatic annotation algorithm. Figure 7. Automatic A-line annotation flowchart. Figure 5. Post-processing adjustment of automatic pleural annotation. (a) B-scan image; (b) initial pleural annotation approximation; (c) final pleura annotation. Figure 6summarizes the workflow for the pleural line automatic annotation algorithm. Appl. Sci. 2024, 14, x FOR PEER REVIEW 9 of 25 (a) (b) (c) Figure 5. Post-processing adjustment of automatic pleural annotation. (a) B-scan image; (b) initial pleural annotation approximation; (c) final pleura annotation. Figure 6 summarizes the workflow for the pleural line automatic annotation algorithm. Figure 6. Automatic pleural line annotation flowchart. A-Line Labelling In case of A-lines, the developed algorithm detects these patterns as echoes of the pleura occurring at depths that are multiples of the probe-to-pleura distance. A search window W 󰇟𝑤,𝑤󰇠 is applied at those depths to find signal peaks that cross a threshold, and then verify their validity by comparing with the average amplitude in consecutive lines: 𝑎=arg max ∈󰇟  ,  󰇠 𝐴 󰇛𝑥󰇜 (7) where 𝑎 corresponds to the index in the A-scan signal 𝐴 where the maximum amplitude is located within the window 󰇟𝑤,𝑤󰇠, which corresponds with an initial guess of the A-Line position. As in the case of the pleura line, to define the ground truth mask, a −3 db signal drop criteria equivalent to the Formula (6) is used, followed by the application of opening and closing morphological filters sized at 3 × 3 to refine the mask and eliminating outliers. Figure 7 summarizes the workflow for the A-lines automatic annotation algorithm. Figure 7. Automatic A-line annotation flowchart. Figure 6. Automatic pleural line annotation flowchart. Appl. Sci. 2024,14, 11930 9 of 23 A-Line Labelling In case of A-lines, the developed algorithm detects these patterns as echoes of the pleura occurring at depths that are multiples of the probe-to-pleura distance. A search window W [w1,w2] is applied at those depths to find signal peaks that cross a threshold, and then verify their validity by comparing with the average amplitude in consecutive lines: aj=arg max x∈[w1,w2]Aj(x)(7) where aj corresponds to the index in the A-scan signal Aj where the maximum amplitude is located within the window [w1,w2] , which corresponds with an initial guess of the A-Line position. As in the case of the pleura line, to define the ground truth mask, a − 3 db signal drop criteria equivalent to the Formula (6) is used, followed by the application of opening and closing morphological filters sized at 3 ×3 to refine the mask and eliminating outliers. Figure 7summarizes the workflow for the A-lines automatic annotation algorithm. Appl. Sci. 2024, 14, x FOR PEER REVIEW 9 of 25 (a) (b) (c) Figure 5. Post-processing adjustment of automatic pleural annotation. (a) B-scan image; (b) initial pleural annotation approximation; (c) final pleura annotation. Figure 6 summarizes the workflow for the pleural line automatic annotation algorithm. Figure 6. Automatic pleural line annotation flowchart. A-Line Labelling In case of A-lines, the developed algorithm detects these patterns as echoes of the pleura occurring at depths that are multiples of the probe-to-pleura distance. A search window W 󰇟𝑤,𝑤󰇠 is applied at those depths to find signal peaks that cross a threshold, and then verify their validity by comparing with the average amplitude in consecutive lines: 𝑎=arg max ∈󰇟  ,  󰇠 𝐴 󰇛𝑥󰇜 (7) where 𝑎 corresponds to the index in the A-scan signal 𝐴 where the maximum amplitude is located within the window 󰇟𝑤,𝑤󰇠, which corresponds with an initial guess of the A-Line position. As in the case of the pleura line, to define the ground truth mask, a −3 db signal drop criteria equivalent to the Formula (6) is used, followed by the application of opening and closing morphological filters sized at 3 × 3 to refine the mask and eliminating outliers. Figure 7 summarizes the workflow for the A-lines automatic annotation algorithm. Figure 7. Automatic A-line annotation flowchart. Figure 7. Automatic A-line annotation flowchart. B-Lines Labelling The approach employed to recognize these pneumonia indicators involves fitting each A-scan line Aj within the image to a linear function y(x) that initiates from the identified pleural line representing the amplitude of the signal at a depth x: y(x)=m·x+d(8) where m is the slope and d is the intercept. Then, several thresholds are established to confirm the detection of the B-line based on a maximum slope Sth , standard deviation σth , and average amplitude of the signal Aavg th after the pleural line detected. Therefore, the lines considered as B-line Blj are calculated as follows, as a boolean vector with length n (number of lines): Blj=True,i f m >Sth and σ<σth and Aavg >Aavg th False,others (9) Once the A-scans with B-lines are known, the annotated masks are generated marking B-lines region below the pleura line frame by frame. Figure 8summarizes the workflow for the B-lines automatic annotation algorithm. Appl. Sci. 2024, 14, x FOR PEER REVIEW 10 of 25 B-Lines Labelling The approach employed to recognize these pneumonia indicators involves fitting each A-scan line 𝐴 within the image to a linear function 𝑦󰇛𝑥󰇜 that initiates from the identified pleural line representing the amplitude of the signal at a depth 𝑥: 𝑦󰇛𝑥󰇜=𝑚⋅𝑥𝑑 (8) where 𝑚 is the slope and 𝑑 is the intercept. Then, several thresholds are established to confirm the detection of the B-line based on a maximum slope 𝑆, standard deviation 𝜎, and average amplitude of the signal 𝐴  after the pleural line detected. Therefore, the lines considered as B-line 𝐵𝑙 are calculated as follows, as a boolean vector with length n (number of lines): 𝐵𝑙=𝑇𝑟𝑢𝑒, 𝑖𝑓 𝑚>𝑆 𝑎𝑛𝑑 𝜎𝜎 𝑎𝑛𝑑 𝐴  > 𝐴   𝐹𝑎𝑙𝑠𝑒, 𝑜𝑡ℎ𝑒𝑟𝑠 (9) Once the A-scans with B-lines are known, the annotated masks are generated marking B-lines region below the pleura line frame by frame. Figure 8 summarizes the workflow for the B-lines automatic annotation algorithm. Figure 8. Automatic B-line annotation flowchart. Consolidation Labelling While pleura, A-lines, and B-lines can be robustly detected by quite simple algorithms applied along each scan line, the consolidations cannot be tackled with the same approach. As they are intrinsically two-dimensional structures, it is not possible to detect them on a line-by-line basis, and image algorithms are needed. The approach followed in this work was that, given a video labeled by the expert as containing a consolidation, a key frame where this consolidation is seen is first identified. Then, it is manually delineated using a custom developed interactive tool (Figure 9). Finally, an optical flow algorithm [27,28] tracks the movement of the consolidation in the subsequent frames, automatically generating the ground truth masks for the entire video. The optical flow algorithm is based on the principle of selecting a set of reference points and tracking them through the video. This algorithm is particularly suited for ultrasound images because the texture generated by the image speckle can be used for this purpose. The implemented application requests frame-by-frame validation from the user to ensure correct labelling. In case the algorithm fails in the detection, the points of interest will be re-selected again. These semi-automatic algorithms aim to decrease the time required for labelling videos at frame level, while they only require manual intervention in the case of consolidations and for segmenting a unique frame. Nevertheless, because they can also fail in detecting the indications, supervision of the whole video after it is labeled is required, with the possibility of eliminating those frames where the labeling is considered wrong. Figure 8. Automatic B-line annotation flowchart. Consolidation Labelling While pleura, A-lines, and B-lines can be robustly detected by quite simple algorithms applied along each scan line, the consolidations cannot be tackled with the same approach. As they are intrinsically two-dimensional structures, it is not possible to detect them on a line-by-line basis, and image algorithms are needed. Appl. Sci. 2024,14, 11930 16 of 23 3.1. Model Results As explained in the validation section, a study was conducted to validate the model’s response frame by frame. The results are presented in Tables 1–4, where the Dice coefficients, Intersection over Union (IoU), F1-score, recall, and precision values are calculated. Figure 14 shows the behaviour of the model applying different thresholds to the network output showing the mean values of the metrics in the test dataset. Given the stability of these results, it makes sense to apply 0.5 as a generic threshold to the network output. Appl. Sci. 2024, 14, x FOR PEER REVIEW 18 of 25 Precision 0.81 0.11 0.97 0.14 0.88 0.28 0.69 0.39 Table 2. Table of F1-score of the model at a frame level in test dataset. Artifact Pleura Consolidation B-Line A-Line F1-score 0.83 0.97 0.91 0.73 Table 3. Table of statistic metrics of model at a frame level in out of training patients. Metric Artifact Pleura Consolidation B-Line A-Line Mean Std Mean Std Mean Std Mean Std Dice coefficient 0.77 0.13 0.85 0.34 0.62 0.41 0.48 0.41 IoU 0.64 0.15 0.84 0.35 0.58 0.42 0.42 0.40 Recall 0.78 0.17 0.88 0.31 0.81 0.31 0.59 0.41 Precision 0.79 0.15 0.95 0.20 0.70 0.40 0.64 0.40 Table 4. Table of F1-score of the model at a frame level in out of training patients. Artifact Pleura Consolidation B-Line A-Line F1-score 0.78 0.91 0.75 0.61 Figure 14. Threshold comparison in CNN output applied for test dataset. Figure 15 shows different network output examples modifying the saturation of the image at the input of the network, to simulate the gain variation usually applied on an ultrasound scanner. Segmentation and detection of artifacts remains stable, guaranteeing correct operation in spite of signal saturation. Figure 15a shows the learning and generalization capacity of the trained model. The labelling algorithms of the A-line only considers a single A-line at double the distance Figure 14. Threshold comparison in CNN output applied for test dataset. Table 1. Table of statistic metrics of model at a frame level in test dataset. Metric Artifact Pleura Consolidation B-Line A-Line Mean Std Mean Std Mean Std Mean Std Dice coefficient 0.83 0.08 0.96 0.17 0.87 0.30 0.62 0.40 IoU 0.72 0.11 0.95 0.18 0.84 0.30 0.56 0.40 Recall 0.86 0.10 0.97 0.12 0.94 0.17 0.78 0.33 Precision 0.81 0.11 0.97 0.14 0.88 0.28 0.69 0.39 Table 2. Table of F1-score of the model at a frame level in test dataset. Artifact Pleura Consolidation B-Line A-Line F1-score 0.83 0.97 0.91 0.73 It is interesting to note in Tables 1–4how the results, between patients included in the training and not, present a similar behaviour. The network, despite the fact that in both cases are images that have never been seen, is able to match more accurately in data from known patients (Tables 1and 2). However, the results shown in Tables 3and 4suggest that the neural network has been able to generalize the problem and maintain high success rates. Appl. Sci. 2024,14, 11930 17 of 23 Table 3. Table of statistic metrics of model at a frame level in out of training patients. Metric Artifact Pleura Consolidation B-Line A-Line Mean Std Mean Std Mean Std Mean Std Dice coefficient 0.77 0.13 0.85 0.34 0.62 0.41 0.48 0.41 IoU 0.64 0.15 0.84 0.35 0.58 0.42 0.42 0.40 Recall 0.78 0.17 0.88 0.31 0.81 0.31 0.59 0.41 Precision 0.79 0.15 0.95 0.20 0.70 0.40 0.64 0.40 Table 4. Table of F1-score of the model at a frame level in out of training patients. Artifact Pleura Consolidation B-Line A-Line F1-score 0.78 0.91 0.75 0.61 Figure 15 shows different network output examples modifying the saturation of the image at the input of the network, to simulate the gain variation usually applied on an ultrasound scanner. Segmentation and detection of artifacts remains stable, guaranteeing correct operation in spite of signal saturation. Appl. Sci. 2024, 14, x FOR PEER REVIEW 19 of 25 between the probe and the pleura; however, the model can detect the second and even the third echo in some particular cases, demonstrating the capacity of the model in generalizing and learning the problem. It is worth mentioning that this fact negatively affects the validation metrics calculated in Tables 2 and 4, where the false positive values for A-lines are higher. This is because the second and third occurrences of the A-lines are not labelled in the input dataset, which is a limitation of the labelling algorithm rather than of the model itself. (a) (b) Figure 15. Cont. Appl. Sci. 2024,14, 11930 18 of 23 Appl. Sci. 2024, 14, x FOR PEER REVIEW 20 of 25 (c) Figure 15. Results comparative with different image gain: (a) healthy lung; (b) pathological lung; (c) grave lung. 3.2. Real-Time Implementation Results As already mentioned, the implemented software (version 1.0) is capable of processing and displaying in real-time the pleural line, B-lines, consolidations, and A-lines. The algorithm can calculate and obtain the percentage of pleura affected by B-lines, which could be useful to physicians when classifying and determining the severity of a patient’s condition. The implementation proposed in this study achieves a processing rate of up to 20 predictions and image update per second using an octa-core i5 CPU processor and a NVIDIA GeForce RTX 2060 GPU, which is a medium-range hardware set-up. This figure coincides with the framerate given by the ultrasound scanner, so we can state that the solution operates in strict real-time. In terms of computational cost in the different processes, the main process takes about 10 ms, which gives a fluid and latency-free feeling during the scan, while the real-time computation process takes about 50 ms of which 20– 25 ms is for model inference on the GPU. It is worth mentioning that, in the current implementation, the final image refresh rate depends on the number of artifacts detected, due to the respective post-processing and painting algorithms that are not yet optimized. Therefore, the effective frame rate obtained was between 17 and 20 FPS, which should be improved by parallelizing the last stages of the algorithm. Regarding prediction capability at the video level, the results given by the model are compared with the opinion of an expert physician. Two different calculations are performed: first, its performance with the whole dataset is assessed (Table 5), and second, a specific comparison is made for the set of patients excluded from training (Table 6). With these results, it is worth noting that the number of false positives and false negatives are balanced in the case of detection of consolidations and B-lines, which is positive since it is a sign that there is no bias in the training. Additionally, to ensure that the model performs well for normal cases, specific metrics such as accuracy, false positives, and false negatives were calculated for videos showing only A-lines or no artifacts at all. These metrics are included in Tables 5 and 6 under the “Normal lung” column. The results demonstrate that the model achieves balanced performance, even for healthy lungs, despite the dataset focusing more on pathological findings. Figure 15. Results comparative with different image gain: (a) healthy lung; (b) pathological lung; (c) grave lung. Figure 15a shows the learning and generalization capacity of the trained model. The labelling algorithms of the A-line only considers a single A-line at double the distance between the probe and the pleura; however, the model can detect the second and even the third echo in some particular cases, demonstrating the capacity of the model in generalizing and learning the problem. It is worth mentioning that this fact negatively affects the validation metrics calculated in Tables 2and 4, where the false positive values for A-lines are higher. This is because the second and third occurrences of the A-lines are not labelled in the input dataset, which is a limitation of the labelling algorithm rather than of the model itself. 3.2. Real-Time Implementation Results As already mentioned, the implemented software (version 1.0) is capable of processing and displaying in real-time the pleural line, B-lines, consolidations, and A-lines. The algorithm can calculate and obtain the percentage of pleura affected by B-lines, which could be useful to physicians when classifying and determining the severity of a patient’s condition. The implementation proposed in this study achieves a processing rate of up to 20 predictions and image update per second using an octa-core i5 CPU processor and a NVIDIA GeForce RTX 2060 GPU, which is a medium-range hardware set-up. This figure coincides with the framerate given by the ultrasound scanner, so we can state that the solution operates in strict real-time. In terms of computational cost in the different processes, the main process takes about 10 ms, which gives a fluid and latency-free feeling during the scan, while the real-time computation process takes about 50 ms of which 20–25 ms is for model inference on the GPU. It is worth mentioning that, in the current implementation, the final image refresh rate depends on the number of artifacts detected, due to the respective post-processing and painting algorithms that are not yet optimized. Therefore, the effective frame rate obtained was between 17 and 20 FPS, which should be improved by parallelizing the last stages of the algorithm. Regarding prediction capability at the video level, the results given by the model are compared with the opinion of an expert physician. Two different calculations are performed: first, its performance with the whole dataset is assessed (Table 5), and second, a specific comparison is made for the set of patients excluded from training (Table 6). With these results, it is worth noting that the number of false positives and false negatives are balanced in the case of detection of consolidations and B-lines, which is positive since it is a sign that there is no bias in the training. Additionally, to ensure that the model performs well for normal cases, specific metrics such as accuracy, false positives, and false negatives Appl. Sci. 2024,14, 11930 19 of 23 were calculated for videos showing only A-lines or no artifacts at all. These metrics are included in Tables 5and 6under the “Normal lung” column. The results demonstrate that the model achieves balanced performance, even for healthy lungs, despite the dataset focusing more on pathological findings. Table 5. Table of accuracy at video level in whole dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 97.81 88.74 65.79 88.74 False positives (%) 0.44 4.09 23.54 7.16 False negatives (%) 1.75 7.16 10.67 4.09 Table 6. Table of accuracy at video level in out of training dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 89.29 92.86 66.07 92.86 False positives (%) 3.57 1.79 26.79 5.36 False negatives (%) 7.14 5.36 7.14 1.79 It can be observed that the accuracy of the proposed solution at the video level is quite high, and it behaves stably even with patients out of training. In the case of the detection of A lines, the results are notably worse than for consolidations and B-Lines, and this because, in the labelling process carried out by the physician in [ 24 ], labelling the A lines was not mandatory, given that they are indicative of healthy lungs. Therefore, the dataset labelling contains false negatives in those videos where the A-Line is present, but it was not marked by the physician. This limitation might partially explain the lower performance in detecting A-lines compared to other artifacts. However, the inclusion of normal lung metrics demonstrates that the solution remains robust and accurate when distinguishing healthy from pathological cases. This is a limitation of this study, but it does not imply that the performance of the solution is nevertheless promising. Figure 16 shows several examples of the screen of the implemented software in operation explained in Section 2.5.4. Appl. Sci. 2024, 14, x FOR PEER REVIEW 21 of 25 It can be observed that the accuracy of the proposed solution at the video level is quite high, and it behaves stably even with patients out of training. In the case of the detection of A lines, the results are notably worse than for consolidations and B-Lines, and this because, in the labelling process carried out by the physician in [24], labelling the A lines was not mandatory, given that they are indicative of healthy lungs. Therefore, the dataset labelling contains false negatives in those videos where the A-Line is present, but it was not marked by the physician. This limitation might partially explain the lower performance in detecting A-lines compared to other artifacts. However, the inclusion of normal lung metrics demonstrates that the solution remains robust and accurate when distinguishing healthy from pathological cases. This is a limitation of this study, but it does not imply that the performance of the solution is nevertheless promising. Table 5. Table of accuracy at video level in whole dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 97.81 88.74 65.79 88.74 False positives (%) 0.44 4.09 23.54 7.16 False negatives (%) 1.75 7.16 10.67 4.09 Table 6. Table of accuracy at video level in out of training dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 89.29 92.86 66.07 92.86 False positives (%) 3.57 1.79 26.79 5.36 False negatives (%) 7.14 5.36 7.14 1.79 Figure 16 shows several examples of the screen of the implemented software in operation explained in Section 2.5.4. (a) (b) (c) (d) Figure 16. Cont. Appl. Sci. 2024,14, 11930 20 of 23 Appl. Sci. 2024, 14, x FOR PEER REVIEW 21 of 25 It can be observed that the accuracy of the proposed solution at the video level is quite high, and it behaves stably even with patients out of training. In the case of the detection of A lines, the results are notably worse than for consolidations and B-Lines, and this because, in the labelling process carried out by the physician in [24], labelling the A lines was not mandatory, given that they are indicative of healthy lungs. Therefore, the dataset labelling contains false negatives in those videos where the A-Line is present, but it was not marked by the physician. This limitation might partially explain the lower performance in detecting A-lines compared to other artifacts. However, the inclusion of normal lung metrics demonstrates that the solution remains robust and accurate when distinguishing healthy from pathological cases. This is a limitation of this study, but it does not imply that the performance of the solution is nevertheless promising. Table 5. Table of accuracy at video level in whole dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 97.81 88.74 65.79 88.74 False positives (%) 0.44 4.09 23.54 7.16 False negatives (%) 1.75 7.16 10.67 4.09 Table 6. Table of accuracy at video level in out of training dataset. Consolidations B-Lines A-Lines Normal Lung Accuracy (%) 89.29 92.86 66.07 92.86 False positives (%) 3.57 1.79 26.79 5.36 False negatives (%) 7.14 5.36 7.14 1.79 Figure 16 shows several examples of the screen of the implemented software in operation explained in Section 2.5.4. (a) (b) (c) (d) Figure 16. Application visualization sample: (a) B-lines (orange) and consolidation (red) detection; (b) normal lung with A-lines (green); (c) probe movement detected; (d) B line (orange) and A-line (green) deteccion on a Lung phantom. On the right of each image the C-scan image is shown. 4. Discussion The implementation of software solutions using deep learning algorithms is a step towards assisted diagnosis in lung ultrasound. Both the model and the complete solution demonstrate solid performance, validated by the results shown in the previous section. Based on studies such as [ 34 ], which demonstrate the variability in physicians’ opinions, and, taking into account that the proposed method works in real-time (up to 20 fps), it could be a useful tool for clinical practice, helping physicians to quickly address lung conditions and reducing the learning curve for less experienced healthcare personnel in the field. Despite the promising results, there are some limitations in this work that require further research. One of the main ones is the limited extent of the database used to train the model. Only 689 videos from 30 patients are available, which are insufficient to obtain a fully trained model capable of generalizing the entire problem across the four artifacts sought. For example, in the case of videos with consolidations, only 58 of the videos exhibit that artifact, which also explains the high values shown in Table 1. Another limitation of the dataset is that A-Lines were not consistently labeled by the physician, as they are not pathological findings. Even though A-Lines, B-Lines, and consolidations are common artifacts found in patients with pulmonary affectation of different origins, a limitation of the dataset used is that only patients with COVID-19 are present. This would impose a bias when evaluating patients with other pulmonary diseases, which should be further tested in future studies. It should also be highlighted that the images used in this study belong to a single scanner, which is likely to be a limitation, given the unique image characteristics of each equipment in terms of noise, sampling frequency, etc., which may introduce bias to the network training and limit the cross-platform applicability of the solution. Another limitation lies in the semi-automated labelling approach used for consolidations. While the automatic labelling of pleura, B-lines, and A-lines is supported by a validated algorithm from a previous study [ 24 ], the labelling of consolidations required manual intervention guided by physician annotations. This introduces some variability to the dataset, as the identification of consolidations is inherently subjective and dependent on the physician’s experience. Although this method facilitated the creation of a clinically meaningful dataset, it represents a limitation in terms of ensuring complete consistency across all annotations and highlights the need for improved automated labelling approaches in future work. Additionally, the implemented model shows an inability to always segment the complete B-line. In some cases, only part of the B-line is segmented, as is shown in Figure 17. This behaviour is corrected by the post-processing algorithm, which marks the whole scan line as affected by B-line regardless the vertical extension of the output mask. Therefore, the results shown in Tables 5and 6are not affected by this phenomenon, but it indicates Appl. Sci. 2024,14, 11930 21 of 23 that there is already some improvement margin in the network definition and/or in the training process. Appl. Sci. 2024, 14, x FOR PEER REVIEW 23 of 25 the segmentation result, in real time, in a secondary screen. While it is a flexible approach that could take advantage of currently available scanners, it would probably have limitations related to the limited quality of the received image, the lack of configuration information, and the need to convert the sector-scan image to a B-Mode image before processing it. Alternatively, the solution could be deployed on the scanner hardware, which would be ideal in terms of image quality, accessibility to raw data (B-Scan), ease of usage, etc., but depends on the interest of the scanners manufacturers for introducing these features in their systems. Figure 17. B-line segmentation error examples. Regarding regulatory aspects, a solution like this will likely need to comply with “Software as a Medical Device” standards that define the methodological and implementation aspects to warranty patient safety and diagnosis quality. Complying with these standards should be part of the industrialization process of the solution. 5. Conclusions In this study, we have shown the effectiveness of employing deep learning models along with signal processing algorithms for accurate detection of pleura, A-lines, B-lines, and consolidations in lung ultrasound images. The proposed model has been developed as a help for assisted diagnosis, demonstrating promising capabilities in accurately segmenting and detecting key features in the images. The proposed implementation exhibits efficient real-time processing capabilities with moderate hardware resources, achieving a rate of up to 20 predictions per second in a mid-range computer. This makes it suitable for application in clinical settings, where timely diagnosis is crucial for patient care. While the results are promising, this solution presents opportunities for future research and improvements in model and architecture design. The use of semi-automatic labelling tools could be helpful when working with large amounts of data. Frame-by-frame labelling is a time-consuming task, which cannot be always performed by expert physicians in a fully manual way. In terms of segmentation accuracy, the trained model obtained DICE values of 83% for pleura, 96% for consolidations, 87% for B lines, and 62% for A lines. Additionally, the proposed solution achieved 92% of accuracy for normal lung detection. These results suggest that solutions of this type could be used in clinical environments, to reduce subjectivity in the interpretation of images, and even as a tool to help less experienced physicians to perform lung echography and reduce the learning curve of the technique. Author Contributions: Conceptualization, M.M., G.C., J.F.C. and J.C.; methodology, M.M. and J.C.; software, M.M. and J.C.; validation, M.M., A.R., G.C., J.F.C. and J.C.; formal analysis, M.M.; investigation, M.M. and J.C.; resources, J.F.C. and J.C.; data curation, M.M.; writing—original draft preparation, M.M. and J.C.; writing—review and editing, M.M., A.R., G.C., J.F.C. and J.C.; visualization, Figure 17. B-line segmentation error examples. For future work, these limitations could be addressed by including new data into the training or applying transfer learning technics to adapt the result to other scanners to get a more scalable solution. New modifications to the network architecture could be also studied to improve the behaviour of the model, as well as implementing in FPGAs using proper engines adapted for that hardware architecture [ 35 ]. This can enhance its adaptability to broader clinical scenarios [36]. With regard to the implementation of this method in clinical practice, we foresee two approaches. One is to deploy the algorithm in a dedicated hardware that is connected to a conventional scanner through an available output port (Ethernet, HDMI, etc), showing the segmentation result, in real time, in a secondary screen. While it is a flexible approach that could take advantage of currently available scanners, it would probably have limitations related to the limited quality of the received image, the lack of configuration information, and the need to convert the sector-scan image to a B-Mode image before processing it. Alternatively, the solution could be deployed on the scanner hardware, which would be ideal in terms of image quality, accessibility to raw data (B-Scan), ease of usage, etc., but depends on the interest of the scanners manufacturers for introducing these features in their systems. Regarding regulatory aspects, a solution like this will likely need to comply with “Software as a Medical Device” standards that define the methodological and implementation aspects to warranty patient safety and diagnosis quality. Complying with these standards should be part of the industrialization process of the solution. 5. Conclusions In this study, we have shown the effectiveness of employing deep learning models along with signal processing algorithms for accurate detection of pleura, A-lines, B-lines, and consolidations in lung ultrasound images. The proposed model has been developed as a help for assisted diagnosis, demonstrating promising capabilities in accurately segmenting and detecting key features in the images. The proposed implementation exhibits efficient real-time processing capabilities with moderate hardware resources, achieving a rate of up to 20 predictions per second in a mid-range computer. This makes it suitable for application in clinical settings, where timely diagnosis is crucial for patient care. While the results are promising, this solution presents opportunities for future research and improvements in model and architecture design. The use of semi-automatic labelling tools could be helpful when working with large amounts of data. Frame-by-frame labelling is a time-consuming task, which cannot be always performed by expert physicians in a fully manual way. In terms of segmentation accuracy, the trained model obtained DICE values of 83% for pleura, 96% for consolidations, 87% for B lines, and 62% for A lines. Additionally, the proposed solution achieved 92% of accuracy for normal lung detection. These results suggest that solutions of this type could be used in clinical environments, to reduce subjectivity in the interpretation of images, and Appl. Sci. 2024,14, 11930 22 of 23 even as a tool to help less experienced physicians to perform lung echography and reduce the learning curve of the technique. Author Contributions: Conceptualization, M.M., G.C., J.F.C. and J.C.; methodology, M.M. and J.C.; software, M.M. and J.C.; validation, M.M., A.R., G.C., J.F.C. and J.C.; formal analysis, M.M.; investigation, M.M. and J.C.; resources, J.F.C. and J.C.; data curation, M.M.; writing—original draft preparation, M.M. and J.C.; writing—review and editing, M.M., A.R., G.C., J.F.C. and J.C.; visualization, M.M.; supervision, J.F.C. and J.C.; project administration, J.C.; funding acquisition, J.C. All authors have read and agreed to the published version of the manuscript. Funding: This research was partially supported by the project PID2022-143271OB-I00, founded MCIN/AEI/10.13039/501100011033/FEDER, UE and by the fellowship PRE2019-088602 founded by MCIU (Spain). Institutional Review Board Statement: The study was conducted according to the guidelines of the Declaration of Helsinki and approved by the Institutional Review Board of Hospital Universitario Puerta de Hierro (approval code PI47-21, procol version 3.0 and date of approval 5 April 2021). Informed Consent Statement: Informed consent was obtained from all subjects involved in the study. Data Availability Statement: The data presented in this study are available on request from the corresponding author. Acknowledgments: We would like to thank Yale Tung and Angela Trueba-Vicente for their support and advice in the interpretation of images and labelling process. Conflicts of Interest: The authors declare no conflict of interest. References 1. Ferreiro Gómez, M.; Dominguez Pazos, S.J. The Role of Lung Ultrasound in the Management of Respiratory Emergencies. Open Respir. Arch. 2022,4, 100206. [CrossRef] [PubMed] [PubMed Central] 2. Hansell, L.; Milross, M.; Delaney, A.; Tian, D.H.; Rajamani, A.; Ntoumenopoulos, G. Barriers and facilitators to achieving competence in lung ultrasound: A survey of physiotherapists following a lung ultrasound training course. Aust. Crit. Care 2023, 36, 573–578. [CrossRef] [PubMed] 3. Soldati, G.; Smargiassi, A.; Demi, L.; Inchingolo, R. Artifactual lung ultrasonography: It is a matter of traps, order, and disorder. Appl. Sci. 2020,10, 1570. [CrossRef] 4. Bhoil, R.; Ahluwalia, A.; Chopra, R.; Surya, M.; Bhoil, S. Signs and lines in lung ultrasound. J. Ultrason. 2021,21, e225–e233. [CrossRef] [PubMed] [PubMed Central] 5. Tung-Chen, Y.; Martíde Gracia, M.; Díez-Tascón, A.; Alonso-González, R.; Agudo-Fernández, S.; Parra-Gordo, M.L.; Ossaba-Vélez, S.; Rodríguez-Fuertes, P.; Llamas-Fuentes, R. Correlation between chest computed tomography and lung ultrasonography in patients with coronavirus disease 2019 (COVID-19). Ultrasound. Med. Biol. 2020,46, 2918–2926. [CrossRef] 6. Vieillard-Baron, A.; Goffi, A.; Mayo, P. Lung ultrasonography as an alternative to chest computed tomography in COVID-19 pneumonia? Intensive Care Med. 2020,46, 1908–1910. [CrossRef] [PubMed] [PubMed Central] 7. Demi, L.; Wolfram, F.; Klersy, C.; De Silvestri, A.; Ferretti, V.V.; Muller, M.; Miller, D.; Feletti, F.; Wełnicki, M.; Buda, N.; et al. New International Guidelines and Consensus on the Use of Lung Ultrasound. J. Ultrasound Med. 2023,42, 309–344. [CrossRef] 8. Gil-Rodríguez, J.; Pérez de Rojas, J.; Aranda-Laserna, P.; Benavente-Fernández, A.; Martos-Ruiz, M.; Peregrina-Rivas, J.A.; Guirao-Arrabal, E. Ultrasound findings of lung ultrasonography in COVID-19: A systematic review. Eur. J. Radiol. 2022,148, 110156. [CrossRef] [PubMed] [PubMed Central] 9. Amisha; Malik, P.; Pathania, M.; Rathaur, V.K. Overview of artificial intelligence in medicine. J. Fam. Med. Prim. Care 2019,8, 2328–2331. [CrossRef] [PubMed] [PubMed Central] 10. Cirrincione, G.; Cannata, S.; Cicceri, G.; Prinzi, F.; Currieri, T.; Lovino, M.; Militello, C.; Pasero, E.; Vitabile, S. Transformer-Based Approach to Melanoma Detection. Sensors 2023,23, 5677. [CrossRef] 11. Bekta¸s, B.; Emre, ˙ I.E.; Kartal, E.; Gulsecen, S. Classification of Mammography Images by Machine Learning Techniques. In Proceedings of the 2018 3rd International Conference on Computer Science and Engineering (UBMK), Sarajevo, Bosnia and Herzegovina, 20–23 September 2018; pp. 580–585. [CrossRef] 12. Pacurari, A.C.; Bhattarai, S.; Muhammad, A.; Avram, C.; Mederle, A.O.; Rosca, O.; Bratosin, F.; Bogdan, I.; Fericean, R.M.; Biris, M.; et al. Diagnostic Accuracy of Machine Learning AI Architectures in Detection and Classification of Lung Cancer: A Systematic Review. Diagnostics 2023,13, 2145. [CrossRef] [PubMed] 13. Mento, F.; Perrone, T.; Fiengo, A.; Smargiassi, A.; Inchingolo, R.; Soldati, G.; Demi, L. Deep learning applied to lung ultrasoundvideos for scoring COVID-19 patients: A multicenter study. J. Acoust. Soc. Am. 2021,149, 3626–3634. [CrossRef] Appl. Sci. 2024,14, 11930 23 of 23 14. Khan, U.; Afrakhteh, S.; Mento, F.; Mert, G.; Smargiassi, A.; Inchingolo, R.; Tursi, F.; Macioce, V.; Perrone, T.; Iacca, G.; et al. Low-complexity lung ultrasound video scoring by means of intensity projection-based video compression. Comput. Biol. Med. 2023,169, 107885. [CrossRef] [PubMed] 15. Barros, B.; Lacerda, P.; Albuquerque, C.; Conci, A. Pulmonary COVID-19: Learning Spatiotemporal Features Combining CNN and LSTM Networks for Lung Ultrasound Video Classification. Sensors 2021,21, 5486. [CrossRef] [PubMed] 16. Kerdegari, H.; Phung, N.T.H.; McBride, A.; Pisani, L.; Nguyen, H.V.; Duong, T.B.; Razavi, R.; Thwaites, L.; Yacoub, S.; Gomez, A.; et al. B-Line Detection and Localization in Lung Ultrasound Videos Using Spatiotemporal Attention. Appl. Sci. 2021,11, 11697. [CrossRef] 17. Lucassen, R.T.; Jafari, M.H.; Duggan, N.M.; Jowkar, N.; Mehrtash, A.; Fischetti, C.; Bernier, D.; Prentice, K.; Duhaime, E.P.; Jin, M.; et al. Deep Learning for Detection and Localization of B-Lines in Lung Ultrasound. IEEE J. Biomed. Health Inform. 2023,27, 4352–4361. [CrossRef] [PubMed] 18. van Sloun, R.J.G.; Demi, L. Localizing B-Lines in Lung Ultrasonography by Weakly Supervised Deep Learning, In-Vivo Results. IEEE J. Biomed. Health Inform. 2020,24, 957–964. [CrossRef] 19. Howell, L.; Ingram, N.; Lapham, R.; Morrell, A.; McLaughlan, J.R. Deep learning for real-time multi-class segmentation of artefacts in lung ultrasound. Ultrasonics 2024,104, 107251. [CrossRef] [PubMed] 20. Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015, Proceedings of the 18th International Conference, Munich, Germany, 5–9 October 2015; Proceedings, Part III 18; Springer International Publishing: Berlin/Heidelberg, Germany, 2015; pp. 234–241. 21. Muñoz, M.; Cosarinsky, G.; Cruza, J.F.; Camacho, J. Deep Learning-Based Lung Ultrasound image segmentation for real-time analysis. In Proceedings of the 2023 IEEE International Ultrasonics Symposium (IUS), Montreal, QC, Canada, 3–8 September 2023; pp. 1–4. [CrossRef] 22. Keras-Unet. Available online: https://github.com/karolzak/keras-unet/tree/master (accessed on 1 March 2023). 23. Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention u-net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. 24. Camacho, J.; Muñoz, M.; Genovés, V.; Herraiz, J.L.; Ortega, I.; Belarra, A.; González, R.; Sánchez, D.; Giacchetta, R.C.; TruebaVicente, Á.; et al. Artificial Intelligence and Democratization of the Use of Lung Ultrasound in COVID-19: On the Feasibility of Automatic Calculation of Lung Ultrasound Score. Int. J. Transl. Med. 2022,2, 17–25. [CrossRef] 25. Tung-Chen, Y.; Ossaba-Vélez, S.; Acosta Velásquez, K.S.; Parra-Gordo, M.L.; Díez-Tascón, A.; Villén-Villegas, T.; MonteroHernández, E.; Gutiérrez-Villanueva, A.; Trueba-Vicente, Á.; Arenas-Berenguer, I.; et al. The Impact of Different Lung Ultrasound Protocols in the Assessment of Lung Lesions in COVID-19 Patients: Is There an Ideal Lung Ultrasound Protocol? J. Ultrasound 2022,25, 483–491. [CrossRef] [PubMed] [PubMed Central] 26. Fritsch, C.; Camacho, J.; Ibañez, A.; Brizuela, J.; Giacchetta, R.; González, R. A Full Featured Ultrasound NDE System in a Standard FPGA. In Proceedings of the European Congress on Non-Destructive Testing (ECNDT), Berlin, Germany, 25–29 September 2006; ISBN 3-931381-86-2. 27. Lucas, B.D.; Kanade, T. An Image Registration Technique with an Application to Stereo Vision. In Proceedings of the Image Understanding Workshop, Washington, DC, USA, 23 April 1981; pp. 121–130. 28. Patel, D.; Upadhyay, S. Optical flow measurement using Lucas-Kanade method. Int. J. Comput. Appl. 2013,61, 6–10. [CrossRef] 29. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011,12, 2825–2830. 30. Ikotun, A.M.; Ezugwu, A.E.; Abualigah, L.; Abuhaija, B.; Heming, J. K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Inf. Sci. 2023,622, 178–210. [CrossRef] 31. PyQtGraph. Available online: https://pyqtgraph.readthedocs.io/en/latest/getting_started/introduction.html (accessed on 1 February 2022). 32. Multiprocessing—Process-Based Parallelism. Available online: https://docs.python.org/3/library/multiprocessing.html (accessed on 1 October 2022). 33. Gordon, G.A.; Canumalla, S.; Tittmann, B.R. Ultrasonic C-scan imaging for material characterization. Ultrasonics 1993,31, 373–380. [CrossRef] 34. Herraiz, J.L.; Freijo, C.; Camacho, J.; Muñoz, M.; González, R.; Alonso-Roca, R.; Álvarez-Troncoso, J.; Beltrán-Romero, L.M.; Bernabeu-Wittel, M.; Blancas, R.; et al. Inter-Rater Variability in the Evaluation of Lung Ultrasound in Videos Acquired from COVID-19 Patients. Appl. Sci. 2023,13, 1321. [CrossRef] 35. Reis, M.; Véstias, M.; Neto, H. Designing Deep Learning Models on FPGA with Multiple Heterogeneous Engines. ACM Trans. Reconfigurable Technol. Syst. 2024,17, 1–30. [CrossRef] 36. Xu, Y.; Wang, Y.; Chen, Q.; Hu, H.; Huang, H.; Lin, L.; Chen, Y.-W.; Li, J.; Lin, H. FPGA Oriented Lightweight Deep Learning Inference for Liver Cancer Segmentation. In Proceedings of the 2024 IEEE International Symposium on Biomedical Imaging (ISBI), Athens, Greece, 27–30 May 2024; pp. 1–5. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.