scieee AI-readable full text Open interactive document viewer

Image models for the detection and characterization of a rare liver disease (Porto-Sinusoidal Vascular Disorder)

Suárez Fernández, Martín

Abstract

Porto-Sinusoidal Vascular Disorder (PSVD) is a rare and often underdiagnosed liver condition with significant clinical challenges, frequently requiring invasive procedures like biopsies for diagnosis. This thesis investigates the use of advanced artificial intelligence (AI) models to improve the detection and characterization of PSVD using medical imaging data. The research addresses key challenges, including noisy datasets, data imbalance (through metadata), and the subtle visual markers of PSVD that overlap with conditions such as cirrhosis. Using state-of-the-art deep learning techniques, including 2D and 3D convolutional neural networks, this study evaluates different architectures, transfer learning strategies, and preprocessing methods. A detailed pipeline for noise reduction and image segmentation was developed to help models focus on relevant anatomical features while reducing interference from irrelevant areas. The results show that AI has the potential to improve PSVD detection, with advancements in model robustness and interpretability. Grad-CAM was used to create explainable visualizations, offering insights into model decision-making and supporting clinical validation. Despite these advancements, the research emphasizes the need for larger and more diverse datasets, along with further refinement of AI methods, to achieve higher diagnostic accuracy. This work contributes to the growing field of AI-driven medical imaging by providing a foundation for future innovations in diagnosing rare diseases. It also highlights the importance of sustainability and ethical considerations in healthcare technology.

Full text

id195180   IMAGE MODELS FOR THE DETECTION AND CHARACTERIZATION OF A RARE LIVER DISEASE (PORTO-SINUSOIDAL VASCULAR DISORDER) MARTÍN SUÁREZ FERNÁNDEZ Thesis supervisor JUANCARLOSGARCIAPAGAN(FundacióRecercaClínicBarcelona-IDIBAPS) Thesis co-supervisor DARIOGARCÍAGASULLA(DepartmentofComputerScience) Tutor:JAVIERBÉJARALONSO(DepartmentofComputerScience) Degree Master'sDegreeinArtificialIntelligence Master's thesis School of Engineering Universitat Rovira i Virgili (URV) Faculty of Mathematics Universitat de Barcelona (UB) Barcelona School of Informatics (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech  Abstract Porto-Sinusoidal Vascular Disorder (PSVD) is a rare and often underdiagnosed liver condition with significant clinical challenges, frequently requiring invasive procedures like biopsies for diagnosis. This thesis investigates the use of advanced artificial intelligence (AI) models to improve the detection and characterization of PSVD using medical imaging data. The research addresses key challenges, including noisy datasets, data imbalance (through metadata), and the subtle visual markers of PSVD that overlap with conditions such as cirrhosis. Using state-of-the-art deep learning techniques, including 2D and 3D convolutional neural networks, this study evaluates different architectures, transfer learning strategies, and preprocessing methods. A detailed pipeline for noise reduction and image segmentation was developed to help models focus on relevant anatomical features while reducing interference from irrelevant areas. The results show that AI has the potential to improve PSVD detection, with advancements in model robustness and interpretability. Grad-CAM was used to create explainable visualizations, offering insights into model decision-making and supporting clinical validation. Despite these advancements, the research emphasizes the need for larger and more diverse datasets, along with further refinement of AI methods, to achieve higher diagnostic accuracy. This work contributes to the growing field of AI-driven medical imaging by providing a foundation for future innovations in diagnosing rare diseases. It also highlights the importance of sustainability and ethical considerations in healthcare technology. i Acknowledgments I would like to express my sincere gratitude to my supervisors Darío and Javier for his invaluable guidance and support throughout this study. Special thanks are extended to Juan Carlos and other physicians from Hospital Clinic and IDIBAPS for providing expert insights and facilitating access to the medical imaging data essential for this research. Additionally, I deeply appreciate the contributions of all HPAI team members who assisted in the development of the methodologies described in this work, especially Jaume, who, on a weekly basis helped me discuss aspects of the project. Their dedication and effort were essential in the successful completion of this thesis. And last but not least, thank all my family and close friends for their support during the development of the project, especially during the final stages which supposed the biggest amount of work. iii Table of Contents 1 Introduction 1 1.1 Motivation ......................................... 1 1.2 ResearchQuestions..................................... 2 2 Related work 3 2.1 Clinicalbackground .................................... 3 2.2 Difficulties in Differentiating PSVD from Cirrhosis . . . . . . . . . . . . . . . . . . . 4 2.3 AI Transforming Medical Imaging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.4 AIforPSVDdiagnosis................................... 5 3 Dataset 7 3.1 Securityprotocols ..................................... 7 3.2 Datasetsanalysis...................................... 8 3.3 Datapreprocessing..................................... 16 4 Methodology 27 4.1 InputConfigurations.................................... 27 4.2 Architectures ........................................ 28 4.3 TransferLearning...................................... 31 4.4 ExperimentalSetup .................................... 33 4.5 Evaluation.......................................... 33 4.6 ExpectedOutcomes .................................... 36 5 Experiments 37 5.1 Fullimage.......................................... 38 5.2 UsingBoundingBoxes................................... 40 5.3 UsingSegmentationMasks ................................ 42 5.4 Summary .......................................... 44 6 Problem simplification 47 6.1 Healthy vs PSVD vs Cirrhosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 6.2 HealthyvsPSVD...................................... 49 6.3 PSVDvsCirrhosis..................................... 51 6.4 Summary .......................................... 52 7 Sustainability and Ethical Implications 53 v 8 Conclusions 57 8.1 Research Questions Addressed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 8.2 KeyContributions ..................................... 57 8.3 Limitations and Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 8.4 FinalRemarks ....................................... 58 vi List of Figures 3.1 Data safety treatment measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2 Different types CT scan views [14] ............................ 9 3.3 Class distribution across datasets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.4 Overall datsets sex distributions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.5 Condition prevalence by sex. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.6 Age distributions of patients across datasets. . . . . . . . . . . . . . . . . . . . . . . 12 3.7 Age distribution by condition. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.8 Distribution of slice thickness among datasets volumes. . . . . . . . . . . . . . . . . . 13 3.9 Distribution of number of slices per volume . . . . . . . . . . . . . . . . . . . . . . . 13 3.10 Distribution of slice counts per volume by condition. . . . . . . . . . . . . . . . . . . 14 3.11 Distribution of datasets samples’ study years. . . . . . . . . . . . . . . . . . . . . . . 14 3.12 Study year distribution by condition. . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 3.13 Distribution of scanner models used for imaging. . . . . . . . . . . . . . . . . . . . . 15 3.14 Phases of developed automatic cropping method . . . . . . . . . . . . . . . . . . . . 18 3.15 Closeness of representative slices in absolute and relative difference of indices between the best slice and subsequent best slices . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.16 SAM example segmentations of liver (orange) and spleen (green) . . . . . . . . . . . 19 3.17 GennUNet example segmentations of liver (orange) and spleen (green) . . . . . . . . 20 3.18 MONAI example segmentations of spleen (green) . . . . . . . . . . . . . . . . . . . . 20 3.19 Interface of Slicer3D program for segmentation mask creation . . . . . . . . . . . . . 21 3.20 Segmentation mask creation. Drawing (Left) and corresponding mask (Right) . . . . 21 3.21 Difference between raw and clipped image . . . . . . . . . . . . . . . . . . . . . . . . 22 3.22 Example of slice after data augmentation pipeline . . . . . . . . . . . . . . . . . . . . 25 4.1 Input type 1 (Full) image example slice. . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.2 Input type 2 (Cropped) image example slice. . . . . . . . . . . . . . . . . . . . . . . 28 4.3 Input type 3 (Masked) image example slice. . . . . . . . . . . . . . . . . . . . . . . . 28 4.4 Residual connection [30].................................. 29 4.5 Multiclass confusion matrix [41]. ............................. 34 4.6 Example saliency maps generated using Grad-CAM. . . . . . . . . . . . . . . . . . . 35 5.1 Confusion matrices from best models trained on Full images . . . . . . . . . . . . . . 38 5.2 Saliency maps obtained using Grad-CAM from best 2D models trained on full images 39 5.3 Saliency maps from 3D model Medical Net trained on Full images . . . . . . . . . . 40 5.4 Confusion matrices from best models trained on Cropped images . . . . . . . . . . . 41 5.5 Saliency maps from best 2D models trained on cropped images . . . . . . . . . . . . 42 5.6 Confusion matrices from best models trained on Masked images . . . . . . . . . . . . 43 vii Chapter 2. Related work Early detection of PSVD is crucial to prevent disease progression and reduce the likelihood of severe complications. PSVD’s variability and overlap with other liver conditions make it difficult to diagnose and treat, but they also highlight opportunities for innovation. From a computer science perspective, the challenges posed by PSVD, such as interpreting medical imaging, identifying subtle patterns, and predicting disease progression, are well-suited for advanced computational techniques. Machine learning and image analysis, in particular, have the potential to enhance diagnostic accuracy, reduce reliance on invasive tests, and provide new insights into the condition. In conclusion, PSVD is a rare but important liver disorder that presents unique challenges in both medicine and technology. Its distinct features and unpredictable clinical course require careful evaluation and new approaches to diagnosis and treatment. By integrating computational models into the study and management of PSVD, there is hope for more accurate diagnoses, better patient outcomes, and a deeper understanding of the disease’s progression. 2.2 Difficulties in Differentiating PSVD from Cirrhosis Porto-sinusoidal vascular disease (PSVD) and cirrhosis share many clinical similarities, making it difficult to diagnose and manage these conditions correctly. Both can present with complications related to portal hypertension (PH), such as an enlarged spleen, variceal bleeding, fluid buildup in the abdomen (ascites), and portal vein blockages (PVT) [4,5]. However, PSVD is fundamentally different from cirrhosis. While cirrhosis involves advanced scarring and structural damage to the liver, PSVD is defined by specific vascular changes without significant fibrosis or distortion of liver structure [6]. Differentiating between the two is crucial since their underlying causes and treatments are not the same. One of the key tools for distinguishing PSVD from cirrhosis is analyzing liver tissue under a microscope. PSVD typically shows vascular changes such as blocked small veins (obliterative portal venopathy), partial fibrosis, and nodular regenerative changes, but it does not display the advanced scarring and regenerative nodules seen in cirrhosis [6,7]. However, even liver biopsies are not always definitive, especially in the early stages of PSVD, where the signs can be subtle or missed due to limited sampling. Patients with PSVD often have relatively well-preserved liver function. Key indicators, such as albumin, bilirubin, and blood clotting factors, may appear normal or close to normal. On the other hand, cirrhosis frequently causes significant liver dysfunction due to widespread tissue damage [5]. That said, advanced cases of PSVD can result in complications that resemble decompensated cirrhosis, such as severe fluid retention or gastrointestinal bleeding, making diagnosis even more complicated [8]. Imaging methods like ultrasound or CT scans may also struggle to differentiate the two conditions, as both can show similar features, such as spleen enlargement or abnormalities in the portal vein. The difficulty in distinguishing PSVD from cirrhosis has a major impact on treatment. Treatments for cirrhosis often focus on reducing fibrosis and managing related complications, while PSVD requires a different strategy that targets vascular issues. For instance, anticoagulants to treat PVT in PSVD are used cautiously due to the risk of bleeding, whereas they might be more commonly prescribed in cirrhosis [7]. Misdiagnosing PSVD as cirrhosis can result in inappropriate treatments, and in extreme cases, unnecessary liver transplants [9]. 4 Chapter 2. Related work To tackle these diagnostic challenges, researchers are developing better tools and methods. These include advanced imaging techniques, non-invasive biomarkers, and improved criteria for interpreting liver biopsies [9]. Accurate diagnosis early on is essential to ensure that patients receive the right treatments and to avoid complications caused by mismanagement. 2.3 AI Transforming Medical Imaging Artificial intelligence (AI) has made significant strides in advancing medical imaging and improving diagnostic accuracy and efficiency across various medical fields. In hepatology, AI has been particularly impactful in enhancing the detection and management of liver diseases, such as cirrhosis and its complications. Machine learning (ML) models have shown superior predictive capabilities to traditional diagnostic tools, particularly in assessing liver fibrosis. For example, Chang et al. evaluated various ML models, including logistic regression, random forests, and artificial neural networks, for predicting fibrosis stages in patients with metabolic-associated fatty liver disease (MASLD). These models consistently outperformed conventional methods in accuracy and reliability [10]. Deep learning (DL) techniques, particularly convolutional neural networks (CNNs), have further expanded the capabilities of medical imaging. These models have been successfully applied to imaging modalities such as CT and MRI to detect liver conditions. For instance, Mazumder et al. developed an automated liver segmentation model combining 3D-U-Net and DeepLabv3+ algorithms, accurately identifying cirrhosis from CT scans [11]. Beyond liver diseases, AI has demonstrated its versatility in diagnosing a variety of conditions. Innovative AI models have been designed to analyze non-traditional diagnostic markers, such as tongue color, to detect diseases like diabetes and cancer with high accuracy. Similarly, in oncology, AI-assisted mammography has gained traction for improving breast cancer detection. By leveraging deep learning on extensive mammogram datasets, AI tools have achieved faster and potentially more accurate identification of cancerous patterns compared to human radiologists. Despite these advancements, challenges remain in integrating AI into everyday clinical practice. Key concerns include data privacy, model interpretability, and the lack of standardized validation protocols. Ensuring compliance with medical regulations and safeguarding patient information is critical for successful adoption [12]. In summary, AI continues to revolutionize medical imaging, offering transformative potential for diagnosing liver diseases and beyond. Ongoing research is essential to address current challenges and maximize the utility of AI-driven tools in clinical settings. 2.4 AI for PSVD diagnosis However, despite all the advances made in the medical field regarding AI, the application of this kind of systems for the specific case of PSVD has not been developed that much, being one of the most advanced cases the previous work done with reduced versions of the data used in this study [13]. The mentioned work carefully analyzed the data and successfully created a deep learning model able to predict with high accuracy the diagnosis of the disease by following a classification approach between 4 classes: Cirrhosis, PSVD, Healthy, and Miscellaneous (Group of diverse liver diseases other than the ones already stated). However, this model, when presented with a new dataset 5 Chapter 2. Related work was not able to perform as good as expected, providing almost no predictions for the PSVD class. Additionally, the explainability methods used on the mentioned model showed that it was hugely biased by high-value zones such as ribs or the backbone of the patient, leading the physicians to not trust the model. Therefore, as stated before in this document, one of the goals of this study is to develop a model that consistently performs good and reliable predictions. 6 Chapter 3 Dataset This chapter captures the definition and analysis of the data used for the project, detailing aspects of both datasets used as well as different types of processing applied to such datasets in order to obtain a curated version with reduced noise. 3.1 Security protocols First and foremost, due to the necessary privacy measures that have to be taken for the personal details of the patients not to be publicly revealed, a protocol had to be followed in order to be able to work with the data. As explained in the previous work done by Rubén [13], the data was safely transported from Hospital Clinic to the Barcelona Supercomputing Center using an encrypted disk and were safely stored. However, access to the data had to be made using an HVAC client to a server in BSC, not allowing access to the data from MareNostrum, only allowing the training in machines with fewer computing capabilities. In order to cope with this problem, a small research was performed to find a secure way of storing the data in the high-performance cluster for later access. Such research led to the following results: • Encryption: Data is encrypted in the cluster using Advanced Encryption Standard (AES) in Cipher Block Chaining (CBC) mode. • Encryption key: The key used to access encrypted data is saved in a different container only accessible by authorized users. It is passed as a variable in memory each time a job is run so that there is never a persistent copy of such key. • Re-encryption: In order to increase the security of the proposed method, encrypted data may be re-encrypted with a newly generated key at any desired moment. Figure 3.1 provides a simple representation of the data treatment necessary to access data for model training. 7 Chapter 3. Dataset Figure 3.1: Data safety treatment measures 3.2 Datasets analysis In the study conducted, two distinct datasets were provided by Instituto de Investigaciones Biomédicas August Pi i Sunyer (IDIBAPS), comprised of 201 (Training dataset) and 208 (Test dataset) volumes each. The Training dataset is the one used in [13] and the Test dataset is the one used to measure the performance of the different models developed. In [13] the model developed, which performed really well on the validation partition of the Training dataset was not able to correctly predict a single instance of the PSVD set of patients, which lead us to perform a careful study of both datasets and analyze their differences. The following exploratory data analysis (EDA) focuses on identifying patterns and characteristics within the datasets that are relevant to diagnosing Porto Sinusoidal Vascular Disease (PSVD) and also to reveal any possible key differences between them. Key insights related to gender, age, imaging slice distribution, and scanning protocols are outlined below. These findings serve as a foundation for preprocessing, and model development aimed at PSVD diagnosis. 3.2.1 Definition Both datasets are comprised of CT (Computed Tomography) scans which use X-Ray radiation and information about the attenuation of such radiation to create what we see as images. The main properties of the files comprising both datasets are described below. • Axial view: Each and every one of the CT scans of the Training dataset provides images in axial view (view from the bottom of the patient. See figure 3.2 to understand the different types of views available for CT scans) whereas some of the scans from the Testing set (4 volumes) are presented in coronal view and, given the difference in format, which made them unusable for models trained on axial views, were therefore discarded. • DICOM format: The file format used for the scans is called DICOM (Digital Imaging and Communications in Medicine) which is the standard format for this type of scans in the 8 Chapter 3. Dataset medical field. It provides not only the images that will be used for training but also some additional metadata that will be explained in detail later in this document. Note that there is no difference in the format file between both datasets. • Region of interest: All of the volumes that comprise both datasets should theoretically be composed of images that contain the liver or at least the abdomen. The majority of CT scans of the Training dataset fill out such requisite, with one sample providing images from the chest down to the pelvis. However, the DICOM files presented in the Test dataset present slices belonging to other areas apart from the abdomen, including the chest, waist, or even legs and head. Figure 3.2: Different types CT scan views [14] 3.2.2 Class distribution As previously mentioned, the whole dataset is divided into four subgroups or classes: 1. Healthy: Group of CT scans made on people in a healthy state regarding the liver. 2. PSVD: Set of patients with a confirmed diagnosis for PSVD. 3. Cirrhosis: This group consists of the CT scans made to those patients with liver cirrhosis. Such partition is included with the objective of making the model more robust towards the differentiation between PSVD and Cirrhosis, which is one the main difficulties for experts in the field. 4. Miscellaneous: Finally, a group with other kinds of liver diseases is included with the purpose of not limiting the possibilities of the diagnosis, especially since a patient might not present PSVD or Cirrhosis but still not be healthy. Both datasets are relatively balanced, with the Training dataset presenting almost a perfect distribution of patients across all classes and the Test dataset showing a slight imbalance towards the Cirrhosis (65 samples) and PSVD (42 volumes) groups. It is worth noting that the most important class (PSVD) is the one with the lowest amount of samples, meaning the is a slight underrepresentation of the disease. Figure 3.3 provides a graphical representation of the class distribution for both groups. 9 Chapter 3. Dataset (a) Train dataset class distribution (b) Test dataset class distribution Figure 3.3: Class distribution across datasets 3.2.3 Sex One of the most relevant aspects for many diseases towards the probability of the patient having it or not is the sex of such patient. Biological differences might make males or females more prone to having a specific disease, and thus, a detailed analysis of the sex distribution across datasets is necessary to better understand the disorder. The datasets show a higher representation of male participants (58.2% in the Training set and 62% in the Test set) 3.4, with over 80% of PSVD cases occurring in males in the Training set and a bit less (73%) in the Test set. This aligns with known disparities in disease prevalence across sexes, influenced by factors such as hormonal differences and body composition. Recognizing this imbalance is crucial for developing diagnostic models that are fair and inclusive, ensuring accurate predictions for all patients (Figure 3.5). At first sight, it appears that PSVD and Cirrhosis are significantly more frequent in males compared to females, whereas females dominate the Miscellaneous class. This might be one of the factors making it difficult to differentiate between patients presenting PSVD or Cirrhosis. The overall gender distribution is relatively balanced across other conditions on both datasets with the Test set being the most irregularly distributed set. However, there is enough balance to ensure model generalization and reduce the likelihood of gender bias during training. Additionally, both datasets present the same type of imbalance in all aspects except for the Healthy class which is balanced in both cases. This leads us to believe that there will be no considerable bias regarding the sex of the patients. 10 Chapter 3. Dataset (a) Training dataset (b) Test dataset Figure 3.4: Overall datsets sex distributions. (a) Training dataset (b) Test dataset Figure 3.5: Condition prevalence by sex. 3.2.4 Age Age is one of the most important factors regarding almost any disease, since the older the human body the weaker it becomes. Specially in this case, where the main symptom is the portal hypertension, age becomes a really relevant aspect to take into account. Additionally, regarding the visual aspect of the problem, which is the one concerning this study, it is important to take into account that the liver might experience changes in its morphology as the body ages [15], which might lead to considerable differences regarding the visual examination of the CT scan. It can be seen in Figure 3.6 how for the Training set the age is equally distributed along a wide range of ages concentrated around 30 and 40 years of age, consistent with the typical demographic for PSVD diagnosis. However, the age distribution notably varies for the Test dataset which presents a concentration of patients around the ages of 50 and 60. Such differences are to be taken into account when analyzing the results since biases might be produced by the age factor due to the model learning on a younger set of patients and most likely with slight differences in the base morphology of the liver which may make it struggle to effectively generalize to older populations leading then to reduced predictive accuracy 11 Chapter 3. Dataset (a) Training dataset (b) Test dataset Figure 3.6: Age distributions of patients across datasets. When age is stratified by condition (Figure 3.7), several differences can be observed between the Training and the Test datasets. On the Training set, the class with the youngest patients is Healthy and the one with the oldest ones is Miscellaneous, contrary to the Test dataset. Additionally, Cirrhosis and PSVD show very similar ranges of age in the Test set, whereas for the Training set there is a clear difference. Again, as mentioned previously, differences in the demographics of the datasets’ patients might produce undesired results, especially given the low amount of data available which may not be enough to learn the natural differences produced in the liver as it ages. (a) Training dataset (b) Test dataset Figure 3.7: Age distribution by condition. 3.2.5 Slices per Volume The overall distribution of slice thickness (Figure 3.8) shows for both datasets two main groups of volumes with 4 and 5mm slice thickness with an increased size of the 4mm group for the Test set. Such similarities suggest that the slice thickness will most likely not skew the model’s results on the Test set. However, as for the slice counts (Figure 3.9), it can be seen how the Test set presents a much higher count of slices per volume when compared to the Training set. This is due to the previously mentioned fact that most if not all volumes in the Test set contain slices not just of the liver area but rather of the whole body from the chest down to the beginning of the femur and in some extreme cases, slices where the brain is shown can be seen. 12 Chapter 3. Dataset (a) Training dataset (b) Test dataset Figure 3.8: Distribution of slice thickness among datasets volumes. (a) Training dataset (b) Test dataset Figure 3.9: Distribution of number of slices per volume For both datasets, the distribution of imaging slices by condition (Figure 3.10) highlights a slight imbalance regarding the Miscellaneous class, which presents a higher slice count. However, the dataset presents an overall balance regarding the number of slices per patient, meaning in theory that there is no potential bias regarding this characteristic of the dataset. It must be noted that even though the distribution is balanced, the values for the Test dataset are much higher than for the Training set, caused by the fact that test volumes present slices that do not show the liver. Therefore, even if at first sight it looks like there will be no bias, the true amount of slices per volume that present the liver in the Test set is unknown. Additionally, even if few, there are outliers in slice counts pointing to variability in imaging practices, likely influenced by differences in equipment or institutional protocols. Preprocessing techniques will be critical for handling these inconsistencies. 13 Chapter 3. Dataset Figure 3.17: GennUNet example segmentations of liver (orange) and spleen (green) Figure 3.18: MONAI example segmentations of spleen (green) Several other models such as Aladdin5 [18], Blackbean [19] or others from competitions about liver segmentation such as SLIVER07 [20] or FLARE22 [21] were considered, but due to limitations regarding how the data is accessed, such options could not be tested in the end. After carefully exploring the results it was deduced that, due to the fact that the dataset contains not only healthy livers but also livers of patients with specific and rare diseases, this approach did not succeed. The publicly available liver segmentation models are either trained to work on healthy livers or livers with cancer (tumors), making them unsuitable for the used dataset. Manual segmentation Finally, after a discussion with the physicians from Hospital Clinic, it was decided that a manual approach for this task could be suitable. The main reason behind this is that the rare nature of the disease and the needs of the physicians call more for precision rather than speed, meaning that even if the process takes longer it is better if the result is more accurate. Therefore, the whole dataset was manually segmented following the steps below: • Tool: Since the data is stored as DICOM files, the decision of which tool to use was mainly based on that, making Slicer3D the most suitable as it provides an easy-to-use interface and makes the segmentation process much easier. • Data loading: Despite Slicer3D being one of the most if not the most appropriate program for this task, the fact that data cannot be stored persistently without being encrypted makes it at first impossible to use. However, this program provides a Python console that allows to run custom code, so a script was created to load the data from the safe container and load it directly into the program so that there is no persistent and decrypted copy of the data at any moment, thus, complying with the security measures. 20 Chapter 3. Dataset • Segmentation: The program allows the segmentation of 3D volumes by an interpolation approach. This is, for a given volume, manual segmentations can be made every N slices and then, using interpolation, segmentations for intermediate slices are created. This mechanism allows to segment a patient’s liver in a reasonable amount of time and also with high precision since the tool allows the edition of the automatically created segmentations. Figure 3.19: Interface of Slicer3D program for segmentation mask creation (a) Drawing (b) Mask Figure 3.20: Segmentation mask creation. Drawing (Left) and corresponding mask (Right) 3.3.2 Cropping or Masking As mentioned in section 3.3.1, images are either cropped or masked using the corresponding segmentation, thus, reducing to a large extent the noise present in each image. Benefits of Clipping and Masking: These preprocessing techniques offer several advantages: • Enhanced Model Training: By reducing the input size and noise, these techniques allow the model to learn more efficiently and focus on patterns that are truly relevant to the task. • Improved Generalization: By eliminating irrelevant regions, the model becomes less likely to overfit to extraneous features, improving its performance on unseen data. 21 Chapter 3. Dataset • Reduced Computational Load: Cropped or masked images are smaller in size, reducing memory and computational requirements during training and inference. 3.3.3 Clipping One of the best and most common practices in medical image analysis is a technique called windowing also known as intensity clipping. This technique adjusts the range of pixel intensities to enhance the visualization of specific tissues or structures in an image [22,23,24]. CT scans capture a wide range of intensities corresponding to various tissue densities, from air to dense bone. This technique is especially useful in our case because the volumes in the dataset show pixel value ranges of either [0, 1024] or [0, 8192], therefore, by clipping the image to a certain range of values, we ensure that after normalization, liver values correspond between different samples. Tissues often have overlapping intensity ranges; for instance, soft tissues and fluids can have similar intensities. By narrowing the window level (center) and width (range), specific tissues, such as lungs, bones, or soft tissues, can be emphasized, improving diagnostic accuracy. Additionally, clipping intensities outside the selected range minimizes noise from irrelevant regions, such as background air or very dense materials. This enhances the signal-to-noise ratio in the area of interest. (a) Sample raw slice (b) Result of slice clipped to (-100, 300) Figure 3.21: Difference between raw and clipped image 3.3.4 Normalization Normalization is a crucial preprocessing step for ensuring consistency across the dataset and improving the stability and efficiency of model training. In medical imaging, CT scans typically have pixel intensity values that correspond to tissue densities. These values can vary widely depending on the imaging protocol, equipment, and patient-specific factors. Normalization helps standardize these intensities, making the data more suitable for machine learning models. 22 Chapter 3. Dataset For this dataset, pixel intensity values were scaled to a fixed range between 0 and 255. This approach ensured that all images had comparable intensity distributions, regardless of the original acquisition parameters. Normalizing to this range is particularly useful for models that require standardized input ranges [25], such as convolutional neural networks (CNNs). Benefits of Normalization: • Improved Convergence: By standardizing the input data, normalization reduces the risk of large gradients during training, leading to faster and more stable convergence of the model. • Reduced Sensitivity to Variations: Normalization minimizes the impact of variations in pixel intensity due to differences in equipment or imaging protocols, making the model more robust to real-world variability. • Enhanced Comparability: Ensuring consistent intensity ranges across all images allows for fair comparisons and ensures that the model focuses on meaningful patterns rather than being influenced by intensity scale differences. In summary, normalization is a foundational step in preprocessing that ensures the dataset is consistent and ready for machine learning tasks. It reduces variability, enhances model performance, and contributes to reliable and reproducible results. 3.3.5 Resizing The dataset included images of varying sizes and resolutions, so all images were resized to consistent dimensions (512x512) to match the model’s input requirements. Also, note that even if information might be lost, the differences in sizes between the different volumes of the datasets are not big enough to cause important losses. 3.3.6 Data Augmentation To improve model robustness and reduce overfitting, data augmentation was applied. This step artificially expanded the dataset by applying various transformations to the images. Several types of transformations were applied but only a few of them worked properly: Tested transformations 1. Brightness Adjustment: Modifies the intensity of the image to simulate different lighting conditions, making it brighter or darker. 2. Contrast Adjustment: Alters the difference between light and dark areas to enhance or suppress specific features. 3. Gamma Adjustment: Applies a nonlinear transformation to adjust the intensity of darker or lighter regions, emphasizing subtle details. 4. Rotation: Rotates the image up to a certain random degree. 5. Flip: Randomly flips vertically or horizontally the image. 6. Scaling: Reescales the data to a certain range of so that values of similar regions in two different images present similar values. For example, liver should commonly present values around 60 but due to differences in the amount of contrast provided to the patient, these might vary. 23 Chapter 3. Dataset 7. Erasing: Randomly masks out patches of the image to simulate occlusions and improve the model’s robustness. 8. Zooming: Changes the scale of the image, simulating variations in object size or distance by zooming in or out. 9. Gaussian noise: Applies a Gaussian filter to the image in order to blur it. 10. Elastic Transform: Applies random, localized distortions to the image by deforming it elastically, simulating natural deformations and enhancing the model’s ability to handle geometric variability. Augmentations purposes • Rotations, Flips, and Scaling: Geometric transformations such as rotations, horizontal flips, and scaling were applied to simulate differences in orientation, perspective, and size. Rotations helped the model become invariant to the angle at which scans were taken, while flips accounted for anatomical symmetry. Scaling adjusted the size of image features, reflecting the variability in patient anatomy and scanning protocols. • Brightness, Contrast, and Gamma Adjustments: Adjustments to brightness, contrast, and gamma values were used to replicate variations in image acquisition conditions, such as changes in scanner settings or the application of contrast agents. These augmentations were implemented in order to improve the model’s robustness to intensity differences, enhancing its ability to generalize across datasets with diverse lighting and contrast properties. Such transformations are particularly effective in medical imaging, where small pixel intensity differences can highlight important diagnostic features [26]. • Noise and Blur: Gaussian noise and blur were added to simulate the noise and artifacts that occur during CT scans. These augmentations were meant to help the model handle noisy data, making it less sensitive to image imperfections. Noise augmentation has been shown to reduce overfitting and improve performance on datasets with inherent variability [27]. • Elastic Transformations: Elastic transformations introduced non-linear deformations to the images, simulating distortions in soft tissues or organs caused by patient movement or scanning variability. These transformations stretched and compressed localized regions of the image, maintaining the overall structure while introducing realistic distortions. Elastic transformations are particularly helpful in medical imaging, as they enhance the model’s robustness to anatomical variability and positional shifts. This method is often used in tasks like segmentation and classification to improve performance on deformed or irregular data [28]. However, maybe due to the limitations in the available data, only a subset of these augmentations worked, being these: Random flips, Random rotations, Scaling, and the Elastic transform. Figure 3.22 shows the difference between a sample slice and the same slice after data augmentation. 24 Chapter 3. Dataset (a) Preprocessed unaugmented slice (b) Preprocessed slice after data augmentation is applied Figure 3.22: Example of slice after data augmentation pipeline 3.3.7 Label Encoding and Smoothing Labels were prepared for supervised learning by encoding them in a machine-readable format. For the current task, these were transformed into a one-hot vector and additionally, due to the fact that it is not sure that every slice of a volume presents the disease of the patient, labels were smoothed. This is, using Gaussian noise, the one-hot vector representing the label of a volume had its values redistributed in order to add some uncertainty (see Algorithm 1). The use of noisy or fuzzy labels helps the model avoid overfitting and therefore generalize better [29]. Algorithm 1 Gaussian Fuzzy Labels Require: sparse_label,σ 1: Initialize one-hot vector: one_hot ←zeros(num_classes) 2: Set one_hot[sparse_label]←1 3: Add Gaussian noise: noise ∼ N (0, σ) 4: Compute fuzzy label: fuzzy_label ←one_hot +noise 5: Clip fuzzy label to ensure non-negative values: fuzzy_label ←max(fuzzy_label, 0) 6: Normalize fuzzy label to ensure probabilities sum to 1: fuzzy_label ←fuzzy_label P(fuzzy_label) 7: return fuzzy_label 25 Chapter 3. Dataset 3.3.8 Splitting the Dataset The Training dataset was split into two subsets: training and validation. A standard split ratio of 80%-20% was used, ensuring each subset represented the overall data distribution. Additionally, stratified sampling was applied to maintain class balance across splits. Finally, the Test dataset was left untouched to be used as a blind test partition to properly measure the generalization capabilities of the trained models. 3.3.9 Summary The preprocessing pipeline ensured that the data was clean, well-structured, and optimized for experimentation. These steps played a vital role in improving the quality of the inputs and, ultimately, the performance of the models. 26 Chapter 4 Methodology This chapter outlines the methodology used to evaluate the effectiveness of different neural network architectures, input configurations, and transfer learning strategies for the proposed classification problem. The goal is to identify the optimal combination of these factors to maximize classification accuracy while balancing computational efficiency and maintaining a reasonable model explainability level. Each experiment was designed to systematically analyze the impact of these variables on model performance and generalization. 4.1 Input Configurations Three distinct input configurations were used to explore the role of preprocessing in improving model performance: Full Image (Figure 4.1): This configuration utilized the entire image slice, providing complete anatomical context in addition to some external noise. This approach allows the model to consider all available spatial information, which is particularly useful when relevant features appear across the entire image. However, the amount of noise included in the final image is more than the relevant information gained, making it a weak approach. Figure 4.1: Input type 1 (Full) image example slice. 27 Chapter 4. Methodology Cropped Image (Figure 4.2): In this setup, the images were cropped using the previously explained method to try to reduce the noise present in each slice as much as possible in an automated way. Cropping reduces irrelevant background noise, focuses the model on the most informative areas, and decreases computational complexity. Figure 4.2: Input type 2 (Cropped) image example slice. Masked Image (Figure 4.3): Again, as explained in previous sections, segmentation masks were produced manually, which were then applied to the corresponding slice for every image in the dataset. Additionally, slices where the corresponding mask does not provide any value, are discarded. This method helps us ensure that the model does not get affected by noise in the image and that the final prediction is solely based on information gathered from the liver. Figure 4.3: Input type 3 (Masked) image example slice. 4.2 Architectures Four types of neural network architectures were tested, each offering different levels of complexity and information processing. All of these are based on the ResNet architecture, ResNet 50 to be specific, which is widely used in the medical field. It uses residual connections to help the model 28 Chapter 4. Methodology learn better and avoid problems like vanishing gradients. The first method, a 2D ResNet pretrained on RadImageNet, is the one used in prior research [13], while the other three architectures introduce novel approaches designed to address specific challenges in processing 3D medical imaging data. Below, each architecture is described in detail. 2D ResNet ResNet, short for Residual Network, was introduced by [30] in 2015 to solve the vanishing gradient problem that made it hard to train very deep neural networks. ResNet uses residual connections 4.4, or "shortcut connections" which let gradients pass through the network more easily by skipping one or more layers. This design makes it possible to train networks with hundreds or even thousands of layers, improving performance in image recognition tasks. Figure 4.4: Residual connection [30] This 2D convolutional neural network processes one image slice at a time, providing one prediction for each slice of a volume. However, since each volume has to be given a single prediction, the outputs of the network for each slice have to be aggregated and in this case, the majority voting strategy has been used, the idea is to give the volume the prediction that is repeated most along the whole volume. This method is an exact replica of what was developed in [13] and will be used for two reasons, first of all, because it was a working method on the Training set, and second, to show can data processing and noise reduction strategies can improve a model’s performance, answering RQ1. Moreover, these types of models provide highly detailed saliency maps when Grad-CAM is applied, which favors its usage for the proposed task, that includes concerns about the reliability of the models. CNN + LSTM Nowadays, most 3D approaches for medical imaging processing are aimed at segmentation and these are most commonly based on the use of U-Net. However, for the small subset of 3D approaches whose task is classification, a widespread approach is to combine a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) [31] network, which is a type of recurrent neural network [32,33,34]. This hybrid model combines the feature extraction capabilities of 2D CNNs for extracting features from single slices with LSTMs for analyzing patterns across multiple slices. This setup works well when context from neighboring slices is important. The specific implementation of this architecture consists of a combination of ResNet 50 pretrained 29 Chapter 4. Methodology Together, these methods ensured that the evaluation was rigorous, validating the models not only in terms of numerical performance but also in terms of clinical relevance and reliability. This combination of metrics and interpretability tools supports the development of trustworthy AI systems for medical imaging, ensuring that they align with expert understanding and deliver meaningful predictions. Visualization: Performance metrics were visualized using: •Line graphs to track performance over training epochs. •Heatmaps to show how the models performed across different classes. 4.6 Expected Outcomes The experiments aim to: • Find the best combination of input type, architecture, and transfer learning strategy for classification. • Show the advantages of using RadImageNet weights compared to ImageNet weights and training from scratch. • Study how noise reduction methods (Full, Cropped, Masked) affect model performance and generalization in order to answer RQ1 and RQ2. • Determine if the training set is enough to create a model able to correctly diagnose PSVD robustly. • Provide useful heatmaps that help identify the flaws in each model and that aid physicians in the diagnostic process. • Study the effects of reducing the task’s complexity by removing classes, answering therefore RQ3. 36 Chapter 5 Experiments This chapter details the experiments conducted to evaluate the performance of different approaches for the task. Three main strategies were explored: using full images, using bounding boxes, and using segmentation masks. For each approach, the results are analyzed, and explainable AI (XAI) techniques are applied to better understand how the model makes decisions. For each section, both quantitative and qualitative results are provided to analyze in detail the impact of different approaches for the suggested problem. In the quantitative sections, tables providing macro F1 scores averaged through the five runs of each model on the three partitions are shown. These present seven different rows, each corresponding to a different approach regarding the model used: • Resnet Scratch: Refers to the 2D model using the ResNet architecture with random weight initialization. • Resnet Imagenet: Uses the same architecture as the previous one but is initialized with ImageNet weights. • Resnet Rad: Another ResNet model but initialized with weights of a model pretrained on RAD ImageNet. • CNN + LSTM: Refers to the model that combines a CNN to extract features and an LSTM to combine such features in a sequence-like way. • CNN + Pool: Similar approach as CNN + LSTM but replaces LSTM with an average pooling layer to simplify the aggregation process. • Medical Net 1 GPU: Corresponds to the MedicalNet-based model presented in section 4.2 but using only 1 GPU, which limits the batch size used during the training process to 2 samples per step. • Medical Net 4 GPU: Same model as before but using 4 GPUs during the training process which allows to raise the batch size up to 8, allowing the model to learn from different classes on each step, and therefore having a more regularized learning process. 37 Chapter 5. Experiments 5.1 Full image In this method, the entire image was used as input to the model without any noise reduction methods applied like cropping or segmentation. The goal was to see how well the model could generalize directly from the raw image data. 5.1.1 Quantitative results The results in Table 5.1 show that models pretrained on radiology-specific data (Resnet Rad) performed best across all metrics, achieving the highest validation and test F1 scores. The Resnet Scratch model, trained from scratch on this dataset, had a very high training F1 score (99.88 ±0.3) but struggled with validation and test performance, suggesting overfitting. Sequential architectures like CNN + LSTM and pooling-based models showed weaker performance, possibly because they failed to capture the spatial features essential for 3D medical imaging tasks due to the huge amount of noise present in the inputs. On the other hand, the 3D convolution-based model Medical Net managed to achieve F1 scores similar to the 2D approaches showing that 3D approaches are feasible. Model Avg Train F1 Avg Validation F1 Avg Test F1 Resnet Scratch 99.88 ±0.3 50.24 ±11.0 44.34 ±4.2 Resnet Imagenet 100.0 ±0.0 62.44 ±7.4 47.68 ±5.1 Resnet Rad 100.0 ±0.0 69.27 ±3.7 48.28 ±2.7 CNN + LSTM 49.25 ±3.6 42.44 ±4.4 23.84 ±2.7 CNN + Pooling 37.50 ±4.6 37.56 ±6.8 31.43 ±6.5 Medical Net 1 GPU 31.50 ±4.5 61.82 ±17.5 24.31 ±0.9 Medical Net 4 GPU 84.62 ±7.7 66.36 ±5.4 41.18 ±1.0 Table 5.1: Avg Macro F1 scores for models trained on Full images (a) Resnet Imagenet (b) Resnet Rad Figure 5.1: Confusion matrices from best models trained on Full images 38 Chapter 5. Experiments However, even if the macro average F1 for the best models reaches almost 50%, it can be seen in the confusion matrices of the best-performing models out of all the runs conducted that such metric values are raised by the abilities of the model to identify cirrhotic and healthy patients, rather than those presenting the PSVD disorder, a class for which the best F1 score is 34.5%, revealing the poor performance on the class we are interested in. This limitation shows the need to find ways to help the models focus on important parts of the images while ignoring noise. The next section looks at methods to improve the input data by picking out the key areas of interest. By cutting down noise and guiding the models to focus on the most important parts, these methods aim to make it easier to find small features linked to PSVD, improving how well the models work with this hard-to-detect class. 5.1.2 Qualitative results In addition to the previous results, saliency maps for the best-performing models were obtained using Grad-CAM from PSVD patients from the Test set to see if the predictions made by such models are to be trusted or not. The saliency maps from the 2D models (Figure 5.2 show that they do not focus on specific important areas in the images. Instead, the models have a scattered activation pattern, which suggests they are not finding the key features linked to the target conditions. This could mean the models are relying on unimportant or random parts of the image, rather than learning useful patterns that separate the classes. (a) Trained from scratch (b) Pretrained on RAD Imagenet Figure 5.2: Saliency maps obtained using Grad-CAM from best 2D models trained on full images On the other hand, Medical Net shows more detailed and specific patterns, meaning that it does focus on certain parts of the image such as the liver and spleen (Figure 5.3. However, such saliency maps are homogeneous throughout the respective organs and do not reflect any specific detail. Moreover, heatmaps show high-value activations on the table or the outer part of the abdomen. This problem shows the need for preprocessing methods that help the models focus on the most important parts of the image. Methods like bounding boxes and segmentation masks, which highlight the key areas, are discussed in the next sections to solve this issue and improve both the accuracy and interpretability of the models. 39 Chapter 5. Experiments Figure 5.3: Saliency maps from 3D model Medical Net trained on Full images The saliency maps indicate that models trained on full images often focus on irrelevant regions, suggesting as initially stated in RQ1 that noise impacts the models’ ability to focus on diagnostic features. 5.2 Using Bounding Boxes In this approach, bounding boxes were applied using the cropping method mentioned in Section 3.3.1 to isolate regions of interest in the images. The idea was to reduce noise from irrelevant parts of the image and make the model focus on the critical areas. 5.2.1 Quantitative results Bounding box-based preprocessing improved performance for some models by focusing on regions of interest and removing irrelevant data. As shown in Table 5.2, Resnet Scratch achieved the highest test F1 score (49.95 ±2.8), while Resnet Rad and Resnet Imagenet showed slightly lower scores. This suggests that bounding boxes help reduce noise but not as much as necessary, indicating that most of the noise is produced by elements inside the abdomen and those slices where the liver is not present. Therefore, such noise has to be further reduced to increase the generalization capabilities of the models. Again, as in the previous section, confusion matrices for the best-performing models were retrieved 5.4 to assess the performance of these in the PSVD group. While the overall macro F1 scores slightly improved with noise-reduction techniques, the specific F1 score for PSVD on the best model only reached 38.9%. This is a clear improvement compared to the previous section, showing that using bounding boxes to reduce the amount of noise in the image helps the models focus on features that are more useful for detecting PSVD. However, the performance is still far from ideal, likely because PSVD has subtle features that overlap with other classes, making it hard to distinguish. 40 Chapter 5. Experiments Model Avg Train F1 Avg Validation F1 Avg Test F1 Resnet Scratch 94.1 ±1.4 60.49 ±7.8 49.95 ±2.8 Resnet Imagenet 100.0 ±0.0 69.27 ±7.4 46.40 ±4.5 Resnet Rad 100.0 ±0.0 69.27 ±2.2 46.80 ±3.9 CNN + LSTM 52.12 ±5.3 33.17 ±5.6 24.03 ±3.4 CNN + Pooling 44.25 ±1.2 39.51 ±6.5 28.76 ±3.6 Medical Net 1 GPU 38.50 ±6.7 70.90 ±11.8 27.84 ±5.7 Medical Net 4 GPU 85.99 ±6.0 68.18 ±4.5 32.64 ±7.4 Table 5.2: Avg Macro F1 scores for models trained on Cropped images (a) Resnet Scratch (b) Resnet Rad Figure 5.4: Confusion matrices from best models trained on Cropped images Despite this progress, the confusion matrices show that PSVD cases are still often misclassified, usually as either healthy or cirrhotic. While bounding boxes reduced noise and improved models’ performances, indicating that image noise reduces the quality of the results; slices without distinctive liver features still led to frequent misclassifications, indicating that further refinement in preprocessing or model training is required. 5.2.2 Qualitative results The saliency maps for models using bounding boxes as a noise-reduction strategy 5.5 show a change in activation patterns, with attention now more focused on certain areas of the images. However, these focused areas are still far from the liver, which is the key region for diagnosing PSVD and related conditions. This misplaced focus suggests that, although the models have improved in narrowing their attention, they still fail to identify the important anatomical features needed for accurate diagnosis. 41 Chapter 5. Experiments This issue shows that the preprocessing techniques tested so far may not be enough to guide the models to clinically relevant regions. It highlights the need for additional methods, such as the use of segmentation masks to direct the models explicitly to the important areas. The results obtained for such an approach will be provided in the next section, (a) Trained from scratch (b) Pretrained on RAD Imagenet Figure 5.5: Saliency maps from best 2D models trained on cropped images 5.3 Using Segmentation Masks This approach used segmentation masks to isolate the liver from the rest of the image, ensuring that the model was provided with only information relevant to the diagnosis of the disease. 5.3.1 Quantitative results Using segmentation masks produced the best overall performance, with Resnet Rad achieving the highest test F1 score (53.36 ±2.8) as shown in Table 5.3. Segmentation effectively removed irrelevant information, allowing models to focus on diagnostic features. Resnet Imagenet and MedicalNet also performed well but fell slightly short of Resnet Rad. CNN + LSTM showed improved performance compared to the full-image approach, indicating that segmentation effectively removes noise from each volume, as not only images are limited to the liver but also, slices not containing the organ are discarded. Model Avg Train F1 Avg Validation F1 Avg Test F1 Resnet Scratch 79.12 ±10.3 55.6 ±7.6 50.50 ±2.8 Resnet Imagenet 98.38 ±1.3 61.46 ±18.3 48.71 ±1.6 Resnet Rad 100.0 ±0.0 67.80 ±5.3 53.36 ±2.8 CNN + LSTM 43.25 ±5.4 47.31 ±2.1 48.11 ±2.7 CNN + Pooling 26.12 ±1.6 35.12 ±8.3 26.03 ±2.4 Medical Net 1 GPU 55.25 ±10.9 60.00 ±6.1 48.01 ±5.2 Medical Net 4 GPU 85.25 ±4.9 65.45 ±5.8 46.17 ±3.9 Table 5.3: Avg Macro F1 scores for models trained on Masked images 42 Chapter 5. Experiments The results show that while segmentation masks improved the overall performance for some classes, the F1 score for the PSVD class stayed low at 36.1%. This is slightly lower than the score achieved with the bounding box approach, suggesting that even with isolated regions, the models still struggle to correctly identify PSVD. The confusion matrices highlight this issue, showing that PSVD cases are often misclassified as either healthy or cirrhotic. The low F1 score for PSVD, even with focused preprocessing, shows how difficult it is to detect this condition. It suggests that the visual markers for PSVD may not be clear enough in the isolated liver region, or that the models are not properly learning the patterns needed to identify these markers. (a) Resnet Scratch (b) Resnet Rad Figure 5.6: Confusion matrices from best models trained on Masked images 5.3.2 Qualitative results The saliency maps for the segmentation mask models (Figures 5.7 and 5.8 show inevitably improved focus within the liver region, which aligns with the purpose of the targeted preprocessing. However, the activations remain scattered and do not concentrate on specific features that are likely related to PSVD. This scattered attention suggests that while segmentation masks help reduce noise, they do not automatically guide the models to learn which features are important for detecting PSVD. (a) Trained from scratch (b) Pretrained on RAD Imagenet Figure 5.7: Saliency maps from best 2D models on Masked images Additionally, Figure 5.8 shows low activation values forming a silhouette of the liver as it would appear in slices close to the one being depicted. Such result in the saliency maps is due to the 43 Chapter 5. Experiments fact that the Medical Net-based approach compresses the spatial dimensions too much to provide detailed information when the input volumes consist of a low amount of slices, as is the case on the masked sets. A solution for more detailed information would be to capture information from earlier layers in the network but that would result in less class-oriented heatmaps. Figure 5.8: Saliency maps from 3D model Medical Net trained on Masked images 5.4 Summary The experiments conducted in this study provide answers to the research questions (RQs) stated in 1.2 by analyzing the performance of imaging models for detecting and characterizing PSVD. RQ1: Impact of Noise in Imaging Data: Noise in imaging data, such as non-body elements, irrelevant organs, and irrelevant slices, had a significant negative impact on model performance. The full-image approach struggled with overfitting and poor generalization due to the high level of noise. Using bounding boxes and segmentation masks helped reduce noise and improved the model’s focus on important areas, leading to better results. However, even with segmentation masks, models found it difficult to identify PSVD-specific features. This suggests that even though noise reduction has been proven to improve the models’ performances such preprocessing alone is not enough to address the challenges of detecting PSVD. RQ2: Handling Slices Without Visual Cues: Given the results, it is hard to confirm whether or not the amount of slices without visual cues is big enough or even if these can greatly affect the outcome. One way to measure the impact of such slices is to compare the methods used to aggregate slices’ features within a volume. Since the best-performing models in general are the ones that only used a CNN and aggregate through max voting it would be logical to think that the disease can be identified in all or most of the slices of a volume. On the other hand, the pooling method performed considerably worse than the LSTM approach, meaning that the disease is most likely not present in all of the slices and is not shown with the same degree of intensity in every slice. RQ3: Effect of Other Classes (Hypothesis: Removing one class from the classification task can help improve the F1 score for PSVD by making the task simpler and reducing the complexity the model needs to handle. In a multi-class setup, the model has to differentiate between all classes, which increases the chances of confusion, especially when some classes have overlapping features. By removing one class, the model has fewer categories to learn, allowing it to focus more on the 44 Chapter 5. Experiments remaining classes, including PSVD. For instance, if one class has non-specific features such as Miscellaneous, which presents several diseases and therefore different characteristics, such class can contribute heavily to misclassification errors, and removing it can help the model to focus on the rest of the classes and better recognize the unique features of PSVD. This approach works well when the removed class is not important to the study’s goals or when its removal does not significantly affect the practical use of the model. By simplifying the classification task, the model can focus on the subtle patterns linked to PSVD, which may lead to better detection and higher F1 scores for this condition. Chapter 8 will focus on the study of such hypothesis and will confirm whether or not the removal of one or several classes can increase the performance of a model on the PSVD class. 45 Chapter 6. Problem simplification As for every other experiment, models struggle to effectively identify the characteristics that differentiate PSVD from Cirrhosis, and as for the previous experiment, it looks like there is a subset of patients that is particularly difficult to classify. (a) Resnet Rad (b) CNN + LSTM Figure 6.5: Confusion matrices of best-performing models trained in PSVD vs Cirrhosis experiment 6.3.2 Qualitative results Figure 6.6 shows saliency maps for this binary classification task. Resnet Rad and Medical Net showed minor improvements in attention, with activations slightly more scattered within the liver region. Additionally, the same problem regarding the spatial dimensions compression of the Medical Net model remained present which might make it less reliable for physicians given the uncertainty in the explicability. (a) Pretrained on RAD Imagenet (b) Medical Net Figure 6.6: Saliency maps from best models trained in PSVD vs Cirrhosis experiment 6.4 Summary Answering research question 3, the experiments in this chapter demonstrate that reducing class complexity can improve model performance, both for distinguishing PSVD from healthy cases and PSVD from cirrhotic patients. However, even with binary classification, the models struggle to isolate PSVD-specific features. These findings highlight the need for further refinement in dataset creation and model development. 52 Chapter 7 Sustainability and Ethical Implications Environmental Sustainability Development of the Project • Energy Consumption: The project involves processing a low amount of data but many times due to the number of experiments (6 distinct experiments) and models tested (7 models each ran 5 times), leading to an estimated final consumption of 84 KWh during development over 210 runs or 0.4 KWh on average per training. • Materials Used: The development of the project was carried out using the facilities of the Barcelona Supercomputing Center, which makes use of Nvidia H100 graphics cards. • Reduction Measures: Measures, such as optimizing computational efficiency, have been implemented to minimize energy consumption. Additionally, measures applied during the project development to reduce the noise in the dataset resulted in the reduction of the dataset size, leading to reduced training times and reduced power consumption. For instance, the power consumption for one run on experiment one was on average 0.5 KWh whereas the consumption for one training on the last experiment (binarized) was on average 0.25 KWh. Execution of the Project • Resource Utilization: All of the models developed can be executed on computers that do not even have a GPU (not using one leads to longer execution times, but still these make its use feasible). Economic Sustainability Development of the Project • Costs: Primary costs involve computational resources and personnel hours which sum up to more than 11k €. • Efficiency Measures: As mentioned before, the optimization of the data helps reduce the computing power needed and therefore its cost. 53 Chapter 7. Sustainability and Ethical Implications Execution of the Project • Viability: Any available decent computer in a hospital could execute these models, so no cost would be assumed in regards to new hardware. Social Sustainability Development of the Project • Skill Development: Development of the project in a professional environment improved the technical and analytical skills of the student. Execution of the Project • Impact on Stakeholders: Potential beneficiaries include researchers and clinicians. Additionally, if finally applied in a real-life scenario, all of the patients whose diagnosis is influenced by results shown by the developed model will be affected too. Risks and Limitations • Ethical Concerns: Patients’ data privacy and potential biases in AI algorithms need continuous assessment to make sure that none of the data is unsafely released or leaked and that the models maintain nondiscriminatory behavior. Ethical Implications • Responsiveness to Needs: The project addresses a specific gap in medical imaging analysis, aiming to improve healthcare outcomes by both providing insightful advances in the research of such novel disease and improving the decision process of the diagnosis. • Anticipated Consequences: Risks such as data misuse or biased outputs are acknowledged and monitored. • Model’s role: The project aims to provide a tool for physicians to better understand a rare disease and a tool to aid in the diagnosis process rather than doing such a diagnosis alone. In the short term, it is a decision support system. Alignment with Sustainable Development Goals (SDGs) Relevant SDG Contribution SDG 3 (Good Health and Wellbeing) Enhancing diagnostic tools for better healthcare outcomes. SDG 9 (Industry, Innovation, and Infrastructure) Promoting sustainable innovation in medical imaging. Table 7.1: Suistainable Development Goals 54 Chapter 7. Sustainability and Ethical Implications Conclusion The project shows a clear effort to evaluate its environmental, economic, and social impacts while considering ethical issues. Ongoing improvements and involvement of stakeholders are needed to ensure long-term sustainability and alignment with ethical standards. 55 Chapter 8 Conclusions This thesis explored the use of advanced imaging models for detecting and characterizing PortoSinusoidal Vascular Disorder (PSVD), addressing the challenges of diagnosing this rare liver disease. By applying deep learning techniques to medical imaging data, the research aimed to improve diagnostic accuracy while tackling issues such as data noise, overlapping class features, and the complexity of features characterizing PSVD. 8.1 Research Questions Addressed RQ1: Impact of Noise in Imaging Data: The research showed that noise, such as non-liver regions and irrelevant anatomical structures, significantly affected model performance. Preprocessing methods like cropping and segmentation masks reduced noise and helped the models focus on relevant regions. However, even with these methods, the models struggled to isolate PSVD-specific features, highlighting the subtle and complex nature of the disorder’s visual markers. RQ2: Handling Slices Without Visual Cues: The analysis revealed that slices without clear visual indicators of PSVD still contributed to predictions when processed together. Aggregation strategies, such as majority voting for slice-level predictions, improved overall accuracy and demonstrated that the disease is visible in most of the slices at least. Other strategies such as pooling, showed by failing in comparison to others that the disease is present clearer in some slices than others. However, these findings show the need for the development of better strategies to make full use of the data. RQ3: Effect of Other Classes: Reducing the number of classes in the classification task improved performance. Binary classifications, such as PSVD vs. Healthy or PSVD vs. Cirrhosis, helped the models focus on subtle features unique to PSVD. However, the experiments also showed the difficulty of distinguishing PSVD from conditions with overlapping imaging characteristics, like cirrhosis. Additionally, as agreed with physicians from IDIBAPS, binary classification problems would be more appropriate in future work, especially tasks were PSVD is to be discerned from Cirrhosis. 8.2 Key Contributions • Developed a preprocessing pipeline that includes noise reduction, data augmentation, and normalization to improve dataset quality and consistency. 57 Chapter 8. Conclusions • Evaluated multiple deep learning architectures (2D, 3D, and hybrid models) to identify optimal configurations for medical imaging tasks. • Demonstrated the value of domain-specific pretrained models (e.g., RadImageNet) for improving diagnostic accuracy and robustness. • Highlighted the importance of explainable AI tools, like Grad-CAM, to provide insights into model decisions, building trust and utility in clinical practice. 8.3 Limitations and Future Work Despite the progress made, the research faced several limitations: • Dataset Size and Diversity: The small and imbalanced (in terms of metadata, not classes) dataset limited the models’ ability to generalize across different populations and imaging conditions. • Subtle Features of PSVD: The lack of clear visual markers for PSVD reduced model performance, requiring further research into feature engineering and interpretability. • Clinical Validation: Additional validation with external datasets and clinical trials is needed to ensure the models are reliable in real-world applications. Future work should focus on: •Increasing dataset size and diversity to improve generalization. • Further improving the preprocessing pipeline to ensure all images present the same conditions, like for example, standardized liver values, invariant to the amount or type of contrast supplied to the patient. • Exploring new architectures better suited for rare disease detection that makes better use of the third dimension of samples. • Discuss with medical professionals to refine the models’ utility and interoperability, ensuring also ethical and responsible development. 8.4 Final Remarks Previous work in the field has been improved, especially since the work presented by [13] was not able to generalize to the Test dataset concerning the PSVD group of patients, therefore showing that improvements in the PSVD research line can be obtained by using AI. This research demonstrates the potential of artificial intelligence in diagnosing rare diseases like PSVD. By addressing key challenges and proposing clinically relevant solutions, this thesis lays a strong foundation for developing more robust, interpretable, and impactful diagnostic tools. The findings not only improve the understanding of medical datasets referred to PSVD but can also be extrapolated to the broader field of AI-based medical image analysis. 58 Bibliography [1] De Gottardi A. et al. “Porto-sinusoidal vascular disease: proposal and description of a novel entity.” en. In: Lancet Gastroenterol. Hepatol. 4.5 (May 2019), pp. 399–411. [2] Andrea De Gottardi, Christine Sempoux, and Annalisa Berzigotti. “Porto-sinusoidal vascular disorder”. en. In: J. Hepatol. 77.4 (Oct. 2022), pp. 1124–1135. [3] Mislav Barisic-Jaman et al. “Porto-sinusoidal vascular disease: a new definition of an old clinical entity”. en. In: Clin. Exp. Hepatol. 9.4 (Dec. 2023), pp. 297–306. [4] Elkrief L. et al. “Porto-sinusoidal vascular disorder: Diagnostic challenges and management”. In: Journal of Hepatology (2023). url:https://pubmed.ncbi.nlm.nih.gov/30957754/. [5] D. et al. Cazals-Hatem. “Histopathology of PSVD”. In: Hepatology International (2019). url: https://pubmed.ncbi.nlm.nih.gov/35690264/. [6] J. N. Schouten et al. “Idiopathic non-cirrhotic portal hypertension: A review of its pathophysiological mechanisms”. In: Hepatology (2011). doi: 10.1002/hep.24445 .url: https: //aasldpubs.onlinelibrary.wiley.com/doi/full/10.1002/hep.24445. [7] Sundaram V. et al. “Emerging diagnostic tools in PSVD”. In: Liver International (2020). url: https://pmc.ncbi.nlm.nih.gov/articles/PMC11103802/. [8] D. Tripathi et al. “Management of portal hypertension: UK consensus guidelines”. In: Gut (2020). doi: 10.1136/gutjnl-2020-320402 .url: https://gut.bmj.com/content/69/11/ 2116. [9] Verheij J. et al. “Advances in noninvasive diagnostics for PSVD”. In: Liver Research (2021). url: https://www.elsevier.es/en-revista-medicina-clinica-english-edition-- 462-articulo-portosinusoidal-vascular-disorder-a-paradigm-S2387020624001475. [10] Fahad Muflih Alshagathrh and Mowafa Said Househ. “Artificial Intelligence for Detecting and Quantifying Fatty Liver in Ultrasound Images: A Systematic Review”. In: Bioengineering (2022). url:https://www.mdpi.com/2306-5354/9/12/748. [11] Zhou LQ et al. “Artificial intelligence in medical imaging of the liver”. In: World Journal of Gastroenterology (2019). 59 Bibliography [12] Naoshi Nishida. “Advancements in Artificial Intelligence-Enhanced Imaging Diagnostics for the Management of Liver Disease—Applications and Challenges in Personalized Care”. In: Bioengineering (2024). url:https://www.mdpi.com/2306-5354/11/12/1243. [13] Ruben Cuervo Noguera. “Deep learning approaches for detecting and explaining hepatic disorders from CT scans”. PhD thesis. UPC, Facultat d’Informàtica de Barcelona, Departament de Ciències de la Computació, Jan. 2024. url:http://hdl.handle.net/2117/406465. [14] Daria Kern and Andre Mastmeyer. “3D bounding box detection in volumetric medical image data: A systematic literature review”. In: J. Image Graph. 10.1 (2022). [15] Nirupama Chatterjee et al. “Is the liver resilient to the process of ageing?” en. In: Ann. Hepatol. 30.2 (Sept. 2024), p. 101580. [16] Nicolás Álvarez Llopis. “From diverse CT scans to generalization: towards robust abdominal organ segmentation”. PhD thesis. UPC, Facultat d’Informàtica de Barcelona, Departament de Ciències de la Computació, June 2024. url:http://hdl.handle.net/2117/413967. [17] M. Jorge Cardoso and Wenqi Li et al. MONAI: An open-source framework for deep learning in healthcare. 2022. arXiv: 2211.02701 [cs.LG].url:https://arxiv.org/abs/2211.02701. [18] Ershuai Wang, Yaliang Zhao, and Yajun Wu. “Cascade Dual-decoders Network for Abdominal Organs Segmentation”. In: Fast and Low-Resource Semi-supervised Abdominal Organ Segmentation. Ed. by Jun Ma and Bo Wang. Cham: Springer Nature Switzerland, 2022, pp. 202–213. isbn: 978-3-031-23911-3. [19] Ziyan Huang et al. “Revisiting nnU-Net for Iterative Pseudo Labeling and Efficient Sliding Window Inference”. In: Fast and Low-Resource Semi-supervised Abdominal Organ Segmentation. Ed. by Jun Ma and Bo Wang. Cham: Springer Nature Switzerland, 2022, pp. 178–189. isbn: 978-3-031-23911-3. [20] Tobias et al Heimann. “Comparison and Evaluation of Methods for Liver Segmentation From CT Datasets”. In: IEEE Transactions on Medical Imaging 28.8 (2009), pp. 1251–1265. doi: 10.1109/TMI.2009.2013851. [21] Jun Ma et al. “Unleashing the strengths of unlabelled data in deep learning-assisted pan-cancer abdominal organ quantification: the FLARE22 challenge”. In: The Lancet Digital Health 6.11 (2024), e815–e826. doi:https://doi.org/10.1016/S2589-7500(24)00154-7. [22] Sahar Sabouri et al. “Adding Liver Window Setting to the Standard Abdominal CT Scan Protocol: Is It Useful?” In: Iranian Journal of Radiology (June 2008). [23] Muhammad Waheed Sabir et al. “Segmentation of Liver Tumor in CT Scan Using ResUNet”. In: Applied Sciences 12.17 (2022). issn: 2076-3417. doi: 10.3390/app12178650 .url: https://www.mdpi.com/2076-3417/12/17/8650. 60 Bibliography [24] Muhammad Islam, Kaleem Nawaz Khan, and Muhammad Salman Khan. Evaluation of Preprocessing Techniques for U-Net Based Automated Liver Segmentation. 2021. arXiv: 2103. 14301 [eess.IV].url:https://arxiv.org/abs/2103.14301. [25] Yann A. LeCun et al. “Efficient BackProp”. In: Neural Networks: Tricks of the Trade: Second Edition. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 9–48. isbn: 978-3-642-35289-8. doi: 10.1007/978-3-642-35289-8_3 .url: https://doi.org/10.1007/978-3-642-352898_3. [26] C. Shorten and T. M. Khoshgoftaar. “A survey on image data augmentation for deep learning”. In: Journal of Big Data 6.1 (2019), pp. 1–48. doi: 10.1186/s4053701901970 .url: https://doi.org/10.1186/s40537-019-0197-0. [27] S. C. Wong et al. “Understanding data augmentation for classification: When to warp?” In: 2016 International Conference on Digital Image Computing: Techniques and Applications (DICTA). 2016, pp. 1–6. doi: 10.1109/DICTA.2016.7797091 .url: https://doi.org/10. 1109/DICTA.2016.7797091. [28] P. Y. Simard, D. Steinkraus, and J. C. Platt. “Best practices for convolutional neural networks applied to visual document analysis”. In: Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings. 2003, pp. 958–963. doi: 10.1109/ICDAR.2003. 1227801.url:https://doi.org/10.1109/ICDAR.2003.1227801. [29] Scott Reed et al. Training Deep Neural Networks on Noisy Labels with Bootstrapping. 2015. arXiv: 1412.6596 [cs.CV].url:https://arxiv.org/abs/1412.6596. [30] Kaiming He et al. “Deep Residual Learning for Image Recognition”. In: CoRR abs/1512.03385 (2015). arXiv: 1512.03385.url:http://arxiv.org/abs/1512.03385. [31] Sepp Hochreiter and Jürgen Schmidhuber. “Long Short-Term Memory”. In: Neural Computation 9.8 (1997), pp. 1735–1780. doi:10.1162/neco.1997.9.8.1735. [32] Shreyasi Roy Chowdhury, Yash Khare, and Susmita Mazumdar. “Chapter 12 - Classification of diseases from CT images using LSTM-based CNN”. In: Diagnostic Biomedical Signal and Image Processing Applications with Deep Learning Methods. Ed. by Kemal Polat and Saban Öztürk. Intelligent Data-Centric Systems. Academic Press, 2023, pp. 235–249. isbn: 978-0-323-96129-5. doi: https://doi.org/10.1016/B978-0-323-96129-5.00008-1 .url: https://www.sciencedirect.com/science/article/pii/B9780323961295000081. [33] Sunil Kumar et al. “CNN-BO-LSTM: an ensemble framework for prognosis of liver cancer”. en. In: Int. J. Inf. Technol. (Sept. 2024). doi: https://doi.org/10.1007/s41870-024-02190-5 . [34] Dong Liang et al. “Combining Convolutional and Recurrent Neural Networks for Classification of Focal Liver Lesions in Multi-phase CT Images”. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2018. Ed. by Alejandro F. Frangi et al. Cham: Springer International Publishing, 2018, pp. 666–675. isbn: 978-3-030-00934-2. doi: 10.1007/978-3030-00934-2_74. 61