Full text
PhD Program in Artificial Intelligence ADVANCES TOWARDS A ROBUST AI-BASED GLIOMA GRADING SYSTEM Carla Pitarch Abaigar Thesis Supervisors: Prof. Alfredo Vellido Alcacena Dr. Vicent Ribas Ripoll September 2024
ABSTRACT Gliomas represent the most common and aggressive form of primary brain tumors in adults, posing significant diagnostic challenges due to their heterogeneous molecular, histological, and radiological characteristics. The World Health Organization (WHO) classifies gliomas on a four-stage scale, ranging from the most benign to the most malignant. The latest WHO Central Nervous System (CNS) diagnostic criteria releases incorporate molecular and histological features for subtyping and grading. However, this process is complex, time-consuming, and prone to inter-observer variability. While histological assessment remains the standard method for glioma characterization, it is invasive and limited in cases where tissue extraction is not feasible, highlighting the need for non-invasive alternatives. This dissertation investigates the application of deep learning algorithms to multi-sequence Magnetic Resonance Imaging (MRI) for the automatic grading of gliomas. First, we integrated and homogenized MRI data from three well-known publicly available datasets resulting in a collective database exceeding one thousand cases. Initial efforts focused on understanding the multi-sequence MRI data, assessing artifacts, and evaluating techniques such as data normalization, data augmentation, transfer learning, and tumor patching, to enhance the classification between lowerand high-grade gliomas, benchmarking with the state-of-the-art. In the second phase, we extend the classification to glioma grades as defined by the WHO CNS 2016 criteria. We developed and compared different Convolutional Neural Networks (CNNs) architectures based on 2D slices from single anatomical planes, multi-planar integrations, and entire 3D volumes. Recognizing the strong association between glioma aggressiveness and genetic features like IDH mutations and 1p/19q co-deletions, we developed a multi-task framework capable of classifying the glioma grade and these molecular features simultaneously, allowing the network to learn contextual relationships among variables. To prevent healthcare disparities and foster the reliability and transparency of our findings, we evaluated model performance across different demographic and clinical populations, including sex, age, and molecular subgroups. Our results reveal consistent challenges in accurately classifying grade 3 gliomas, underscoring the need for further research, as well as more diverse and up-to-date data that aligns with current guidelines. While Artificial Intelligence (AI) holds immense potential to assist in medical diagnosis, prognosis, and therapy planning, its full adoption in the medical field remains limited largely due to a lack of trust and widespread acceptance among healthcare professionals and patients. The ultimate goal of this thesis is to explore the implications and limitations of integrating AI algorithms in healthcare and foster collaboration among clinicians and technologists to achieve patient-centric, reliable AI within clinical practice. iii
RESUM Els gliomes representen la forma m´ es comuna i agressiva de tumors cerebrals primaris en adults, presentant un desafiament diagn` ostic significatiu a causa de la seva heterogene¨ ıtat molecular, histol ` ogica i radiol` ogica. L’Organitzaci´ o Mundial de la Salut (OMS) classifica els gliomes en una escala de quatre estadis, des de benignes fins a malignes. Les ´ ultimes edicions dels criteris diagn ` ostics del Sistema Nervi´ os Central (SNC) de l’OMS incorporen caracter´ ıstiques moleculars i histol ` ogiques per a la subtipificaci ´ o i gradaci´ o. Tanmateix, aquest proc´ es ´ es complex, requereix molt de temps i ´ es propens a la variabilitat entre observadors. Encara que l’avaluaci ´ o histol ` ogica continua sent el m` etode est` andard per a la caracteritzaci ´ o dels gliomes, ´ es invasiva i est` a limitada en casos on no ´ es factible l’extracci ´ o de teixit, cosa que ressalta la necessitat d’alternatives no invasives. Aquest treball investiga l’aplicaci ´ o d’algoritmes d’aprenentatge profund en imatges de resson` ancia magn` etica (IRM) multisecuencia per a la classificaci ´ o autom` atica dels gliomes. Primer, vam integrar i homogene¨ ıtzar les IRM de tres bases de dades p ´ ubliques reconegudes, resultant en una base de dades col·lectiva que supera els mil casos. Els esforc¸os inicials es van centrar en comprendre les dades d’IRM multisecuencia, avaluar artefactes i valorar t` ecniques com la normalitzaci´ o de dades, l’augment de dades, l’aprenentatge per transfer` encia i l’extracci´ o de tumors, per millorar la classificaci´ o entre gliomes de grau baix i alt, comparant-los amb l’estat de l’art. En la segona fase, vam ampliar la classificaci´ o als graus de gliomes segons els criteris del SNC de l’OMS establerts el 2016. Vam desenvolupar i comparar diferents arquitectures de xarxes neuronals convolucionals basades en talls 2D de plans anat ` omics individuals, integracions multiplanares i volums 3D. Reconeguent l’associaci´ o forta entre l’agressivitat dels gliomes i caracter´ ıstiques gen` etiques com les mutacions d’IDH i la codeleci´ o 1p/19q, vam desenvolupar una arquitectura multitasca capac¸ de classificar simult` aniament el grau del glioma i aquestes caracter´ ıstiques moleculars, permetent que la xarxa aprengui relacions contextuals entre les variables. Per prevenir disparitats en el context cl´ ınic i fomentar la fiabilitat i transpar` encia dels nostres resultats, vam avaluar el rendiment del model en diferents poblacions demogr` afiques i cl´ ıniques, incloent sexe, edat i subgrups moleculars. Els nostres resultats revelen reptes consistents en la classificaci ´ o precisa dels gliomes de grau 3, cosa que ressalta la necessitat de m´ es investigaci ´ o, aix´ ı com de dades m´ es diverses i actualitzades que s’aline¨ ın amb les guies actuals. Encara que la Intel.lig` encia Artificial (IA) posseeix un immens potencial per assistir en el diagn` ostic m` edic, el pron` ostic i la planificaci´ o terap` eutica, la seva adopci´ o completa en el camp m` edic continua sent limitada en gran mesura a causa de la manca de confianc¸a i d’acceptaci ´ o generalitzada entre els professionals de la salut i els pacients. L’objectiufinald’aquestatesi ´ esexplorarlesimplicacions ilimitacionsde laintegraci ´ o d’algoritmes de IA en el context m` edic i fomentar la col·laboraci ´ o entre cl´ ınics i tecn ` olegs per aconseguir una IA fiable i centrada en el pacient dins de la pr` actica cl´ ınica. v
RESUMEN Los gliomas representan la forma m´ as com ´ un y agresiva de tumores cerebrales primarios en adultos, presentando un significativo desaf´ ıo diagn´ ostico por su heterogeneidad molecular, histol ´ ogica y radiol´ ogica. La Organizaci ´ on Mundial de la Salud (OMS) clasifica los gliomas en una escala de cuatro estadios, de benignos a malignos. Las ´ ultimas ediciones de los criterios diagn´ osticos del Sistema Nervioso Central (SNC) de la OMS incorporan caracter´ ısticas moleculares e histol´ ogicas para la subtipificaci ´ on y gradaci´ on. Sin embargo, este proceso es complejo, requiere mucho tiempo y es propenso a variabilidad entre observadores. Aunque la evaluaci ´ on histol ´ ogica sigue siendo el m´ etodo est´ andar para la caracterizaci ´ on de gliomas, es invasiva y est´ a limitada en casos donde no es factible la extracci´ on de tejido, lo que resalta la necesidad de alternativas no invasivas. Este trabajo investiga la aplicaci ´ on de algoritmos de aprendizaje profundo en im´ agenes de resonancia magn´ etica (IRM) multi-secuencia para la clasificaci´ on autom´ atica de gliomas. Primero, integramos y homogeneizamos IRM de tres reconocidas bases de datos p ´ ublicas, resultando en una base de datos colectiva que supera los mil casos. Esfuerzos iniciales se centraron en comprender los datos de IRM multisecuencia, evaluar artefactos y evaluar t´ ecnicas como la normalizaci´ on de datos, aumento de datos, el aprendizaje por transferencia y la extracci´ on de tumores, para mejorar la clasificaci´ on entre gliomas de grado bajo y alto, compar´ andolos con el estado del arte. En la segunda fase, ampliamos la clasificaci´ on a los grados de gliomas seg ´ un los criterios del SNC de la OMS establecidos en 2016. Desarrollamos y comparamos diferentes arquitecturas de redes neuronales convolucionales basadas en cortes 2D de planos anat´ omicos individuales, integraciones multiplano y vol ´ umenes 3D. Reconociendo la fuerte asociaci ´ on entre la agresividad de los gliomas y caracter´ ısticas gen´ eticas como las mutaciones de IDH y la codeleci´ on 1p/19q, desarrollamos una arquitectura multitarea capaz de clasificar simult´ aneamente el grado del glioma y estas caracter´ ısticas moleculares, permitiendo que la red aprenda relaciones contextuales entre las variables. Para prevenir disparidades en el contexto cl´ ınico y fomentar la fiabilidad y transparencia de nuestros hallazgos, evaluamos el rendimiento del modelo en diferentes poblaciones demogr´ aficas y cl´ ınicas, incluyendo sexo, edad y subgrupos moleculares. Los resultados revelan desaf´ ıos consistentes en la clasificaci´ on precisa de los gliomas de grado 3 que resaltan la necesidad de m´ as investigaci ´ on, as´ ı como de datos m´ as diversos y actualizados que se alineen con las gu´ ıas actuales. Aunque la Inteligencia Artificial (IA) posee un inmenso potencial para asistir en el diagn´ ostico m´ edico, el pron ´ ostico y la planificaci ´ on terap´ eutica, su adopci´ on completa en el campo m´ edico sigue siendo limitada en gran medida debido a la falta de confianza y aceptaci ´ on generalizada entre los profesionales de la salud y los pacientes. El objetivo final de esta tesis es explorar las implicaciones y limitaciones de la integraci´ on de algoritmos de IA en el contexto m´ edico y fomentar la colaboraci ´ on entre cl´ ınicos y tecn ´ ologos para lograr una IA confiable y centrada en el paciente dentro de la pr´ actica cl´ ınica. vii
PRACTICE 075 4 THE DATA 081 4.1 Description ............................083 4.2 Exploratory data analysis ....................087 4.2.1 MRI data ........................087 4.2.2 Demographic & clinical data .............092 Key Takeaways ..............................096 5 PRELUDE: A CLASSIFICATION OF LOWER AND HIGH GRADES 101 5.1 Preamble .............................103 5.2 Data ................................104 5.3 Methods ..............................106 5.3.1 CNN architectures ...................106 5.3.2 Model training .....................107 5.4 Results ...............................108
5.4.1 Anatomical plane selection .............108 5.4.2 Pre-processing assessment .............109 5.4.3 Evaluating network architectures ..........116 5.4.4 Evaluating the sample size ..............117 Key Takeaways ..............................118 6 ON THE WHO GLIOMA GRADE CLASSIFICATION 123 6.1 Preamble .............................125 6.2 Data ................................127 6.3 Methods ..............................128 6.3.1 The 2D single-planar classifier ...........128 6.3.2 The 2.5D multi-planar classifier ...........128 6.3.3 The 3D volumetric classifier .............129 6.3.4 The multi-task learning framework .........130 6.4 Model training ...........................131 6.5 Results ...............................131 6.5.1 Glioma grade classification performance .....131 6.5.2 Combining grade and molecular features .....137 6.5.3 Evaluating the classifier disparities .........139 Key Takeaways ..............................142
CLOSING 147 DISCUSSION 151 FUTURE DIRECTIONS 156 LIST OF PUBLICATIONS 158 BIBLIOGRAPHY 163 APPENDICES 203 A. Literature review .........................203 B. Supplementary results .....................210 C. Models’ hyperparameters ....................226
LIST OF FIGURES Fig. 0.1 General thesis outline, including its parts and chapters. 009 Fig. 1.1 Flow diagram that illustrates the evolution of the WHO 2016 classification into the WHO 2021 classification of gliomas .................. 025 Fig. 1.2 Patches extracted from glioblastoma (first row), astrocytoma (second row), and oligodendroglioma (third row) WSI images. ................... 027 Fig. 1.3 MRI scanner (left) and a radiologist interpreting a brain MRI (right) ....................... 028 Fig. 1.4 Sagittal, coronal, and axial views of an MRI scan. ... 029 Fig. 1.5 Conventional sequences of an MRI scan in axial view. 030 Fig. 2.1 Illustration of interand intra-subjects MRI registration. 036 Fig. 2.2 Raw (top) and bias field corrected (bottom) images. A color map has been applied to facilitate the visualization of tissue differences after BFC. ...... 038 Fig. 2.3 Raw (top) and skull-stripped (bottom) images. ..... 039 Fig. 2.4 Raw (top) and equalized (bottom) images. ....... 040 Fig. 2.5 Visualization of 3D brain MRI volume. .......... 042 xvii
Fig. 2.6 Visualization of ten (a) sagittal, (b) coronal, and (c) axial 2D slices. ........................ 042 Fig. 2.7 Architecture of a 2D vanilla CNN with two convolutional layers, two pooling layers, one flattened layer, and one fully connected layer. ..... 048 Fig. 2.8 K-fold Cross-Validation scheme. ............. 050 Fig. 2.9 Data integration strategies. ................ 051 Fig. 2.10 Multi-task learning. ..................... 052 Fig. 3.1 Publicly available datasets usage frequency from 2018 to 2024. ........................ 064 Fig. 4.1 Flow diagram of participant numbers initially and post-quality assessment of MRI modalities, alongside the implemented data split strategy. ..... 084 Fig. 4.2 Pixel distribution of EGD flair MRIs across imaging acquisition scans. ...................... 087 Fig. 4.3 Histograms of pixel values extracted from sagittal 2D MRIs containing the largest area of tumor. ....... 089 Fig. 4.4 Histograms of pixel values extracted from coronal 2D MRIs with largest area of tumor. ............. 090 Fig. 4.5 Histograms of pixel values extracted from axial 2D MRIs containing the largest area of tumor. ....... 091 Fig. 4.6 Histograms of 2D axial slices using (a) min-max normalization, (b) standardization on complete image region, and (c) standardization on the brain region. ............................. 092 xviii
Fig. 4.7 Distribution of two-class tumor grade by (a) sex, (b) age, (c) IDH, and (d) 1p/19q co-deletion status. .... 093 Fig. 4.8 Distribution of tumor grade by (a) sex, (b) age, (c) IDH, and (d) 1p/19q co-deletion status. ......... 094 Fig. 4.9 Distribution of IDH and 1p/19q co-deletion status by glioma grade. ......................... 095 Fig. 5.1 2D MRI slices containing the largest tumor area from (a) sagittal, (b) coronal, and (c) axial views. ....... 104 Fig. 5.2 Example of the selected 10 contiguous slices both before and after the slice containing the maximum tumor area with a skip of 2 for (a) LGG and a skip of 5 slices for (b) HGG patients. ............... 105 Fig. 5.3 Proposed ResNet architecture based on 18 and 34 layers. ............................. 106 Fig. 5.4 Proposed VGGNet architecture based on 11 and 16 layers. ............................. 107 Fig. 5.5 Classification performance of LGG and HGG across anatomical planes. ..................... 108 Fig. 5.6 (a) Overall AUC, (b) LGG recall, and (c) HGG recall, using individual MRI sequences and their fusion across various pre-processing techniques. ....... 111 Fig. 5.7 GradCAM attention maps across different pre-processing and sequences. .............. 113 Fig. 5.8 Model’s output probability distribution across independent sequences and their fusion, using mean-std standardized images within the brain region. 114 Fig. 5.9 Multi-modal model’s output probability distribution incorporating data augmentation. ............. 115 xix
Fig. 6.1 2D MRI slices containing the largest tumor area cropped to the tumor ROI from (a) sagittal, (b) coronal, and (c) axial views. ................ 127 Fig. 6.2 Proposed multi-planar architecture using intermediate aggregation to integrate the output features of each plane-specific model. .......... 129 Fig. 6.3 Proposed 3D model for volumetric data. ......... 130 Fig. 6.4 Recall for glioma grades 2, 3, and 4 across 2D, 2.5D, and 3D model strategies. .................. 134 Fig. 6.5 Confusion matrices obtained using the multi-planar model with intermediate feature concatenation. .... 135 Fig. 6.6 Output probability distribution across grades using the multi-planar model with intermediate feature concatenation. ........................ 135 Fig. 6.7 Mean 3-fold CV recall on test set across demographic and clinical populations. .......... 140 Fig. A.1 Confusion matrices using the multi-planar model for the female group. ...................... 223 Fig. A.2 Confusion matrices using the multi-planar model for the male group. ....................... 223 Fig. A.3 Confusion matrices using the multi-planar model for the age <40 group. ..................... 223 Fig. A.4 Confusion matrices using the multi-planar model for the age 40 −59 group. ................... 224 Fig. A.5 Confusion matrices using the multi-planar model for the age ≥60 group. ..................... 224 xx
Fig. A.6 Confusion matrices using the multi-planar model for the IDH wild-type group. .................. 224 Fig. A.7 Confusion matrices using the multiplanar model for the IDH mutated group. ................... 225 Fig. A.8 Confusion matrices using the multi-planar model for the 1p/19q intact group. .................. 225 Fig. A.9 Confusion matrices using the multi-planar model for the 1p/19q co-deleted group. ............... 225 xxi
LIST OF TABLES Table 3.1 An overview of publicly available MRI datasets for brain tumor classification benchmarking. .........062 Table 4.1 Grade distribution across datasets. .............084 Table 4.2 Demographic characteristics over the different data partitions. ............................086 Table 4.3 Frequency of imaging acquisition scanner manufacturers in EGD dataset. ...............087 Table 4.4 Statistics computed from sagittal 2D MRIs containing the largest area of tumor. ............089 Table 4.5 Statistics computed from coronal 2D MRIs containing the largest area of tumor. ............090 Table 4.6 Statistics computed from axial 2D MRIs containing the largest area of tumor. ...................091 Table 4.7 Distribution of LGG and HGG by clinical features. ....093 Table 4.8 Distribution of tumor grade by clinical features. ......094 Table 5.1 Mean 3-fold CV 2D classification performance on the test set using ResNet-18 to distinguish between LGG and HGG, evaluated across the three anatomical planes using the largest tumorous slice. ...109 xxii
Table 5.2 Mean 3-fold CV performance on the test set for the classification of LGG vs. HGG obtained using the tumor ROI, multi-slice oversampling, and fine-tuning from ImageNet. ........................115 Table 5.3 Mean 3-fold CV performance on the test for the classification of LGG vs. HGG using ResNet-18, ResNet-34, VGGNet-11, and VGGNet-16. ........116 Table 5.4 Comparison of model parameters and training time for ResNet-18, ResNet-34, VGGNet-11, and VGGNet-16. ..........................117 Table 5.5 Mean 3-fold CV performance on the test set for the classification of LGG vs. HGG considering different sample sizes. ..........................117 Table 6.1 Mean 3-fold CV performance on the test set to classify the WHO glioma grade using 2D single-planar, 2.5D multi-planar, and 3D volume models. 133 Table 6.2 Mean 3-fold CV performance on the test set for pairwise grade comparisons using the intermediate fusion multi-planar architecture. ...............136 Table 6.3 Mean 3-fold CV performance on the test set for the classification of IDH mutation status and 1p/19q codeletion status using the multi-planar architecture. ....137 Table 6.4 Mean 3-fold CV performance on the test set for the classification of the WHO glioma grades, the IDH mutation status, and the 1p/19q co-deletion status using the multi-task framework. ...............138 Table 6.5 Mean 3-fold CV performance on the test set for the classification of the WHO glioma grades across demographic groups using the multi-planar architecture. 141 xxiii
INTRODUCTION CLINICAL CONTEXT In recent decades, the prevalence of brain cancer in the adult population has increased by approximately 40% [1]. The latest data available from Global Cancer Observatory [2] reveals that in 2022, brain tumors were diagnosed in 321 731 individuals globally. This represents an incidence rate of 4.1 per 100 000 people, accounting for 1.6% of all new cancer cases diagnosed. Furthermore, brain cancer was responsible for 248 500 deaths, corresponding to a crude mortality rate of 3.2 per 100 000 person-years. Gliomas represent the most prevalent form of malignant primary brain tumors, constituting the 75% [3]. Managing these tumors is particularly challenging due to their heterogeneity, infiltrative nature, and variable response to treatment [4]. Notably, Glioblastoma Multiforme (GBM), recognized as one of the most aggressive forms of cancer, accounts for 55% of all gliomas and is often accompanied by a poor prognosis, with only 5.5% of the patients estimated to achieve a 5-year relative survival [5]. Glioblastoma, occurying at a rate of 3 cases per 100 000 person-years [6], is a rare disease, making the collection of large and diverse datasets for large-scale population research challenging. According to the World Health Organization (WHO) system, gliomas are graded on a scale ranging from 1 (most benign) to 4 (most malignant). Grade 1 tumors are considered non-cancerous and are rare in adults. In the fourth edition of the WHO Classification of Tumors of the Central Nervous System (CNS) published in 2016, molecular parameters were included along with histological features as decisive markers for glioma classification, mostly based on the presence of Isocitrate Dehydrogenase (IDH) mutation and co-deletion of chromosome arms 1p/19q [7]. The 2021 fifth edition [8] introduced significant changes 003
that advance the role of molecular diagnostics in CNS tumor classification. Inconsistencies arise in how certain criteria are interpreted, which can differ among clinicians, consequently resulting in a misgrading rate of up to 30% [9]. Although imaging features are not included in the WHO classification, they still can serve as a powerful tool to impact clinical practice. Before tissue confirmation and also in rare cases when tissue confirmation is not possible, imaging features are the key information that drives the clinical decision [10]. The treatment for a glioma depends on its grade and may include observation, surgery, radiation therapy, and chemotherapy [11]. The grade of gliomas is significant in determining their treatment since high grades are treated in more aggressive ways, often with radiation or chemotherapy. Accurate pathological classification is essential for determining treatment and prognosis prediction. Several studies highlight significant variations among observers in the histological analysis of gliomas, affecting tumor typing and grading [12]. Such discrepancies can result in inappropriate treatment decisions, risking both overand under-treatment, especially in determining radiotherapy and chemotherapy interventions. The differentiation between grades 2 and 3 is particularly challenging due to the subjective nature of the WHO glioma classification standards, which rely on morphological descriptions with interpretive boundaries between different tumor grades, with terms like “moderately” or “increased” cellularity. Imaging plays a fundamental role in the management of gliomas from the diagnosis to the post-treatment follow-up, which varies according to the histological grade, location, resectability of the tumor, and performance status of the patient [13]. Regarding the post-treatment follow-up, the gold standard for assessing treatment response is provided by the Response Assessment in Neuro-Oncology Working Group (RANO) [14] criteria, which integrates both imaging features and clinical factors. Recently, the introduction of RANO 2.0 criteria [15], represents a notable refinement distinguishing between low-grade and high-grade gliomas. Notably, RANO 2.0 also considers the IDH mutation status to determine the assessment of the surrounding non-enhancing region. The current gold-standard procedure for glioma grading involves surgery and histological evaluation of the tumor, which play a pivotal role in both diagnosis and prognosis. Nonetheless, due to the invasive and time-consuming nature of this method, there is a growing interest in exploring non-invasive and pre-operative procedures to characterize the tumor, accelerate the diagnosis, and plan personalized treatments. 004
METHODOLOGICAL CONTEXT AI in Healthcare In recent decades, significant strides have been made in developing robust Artificial Intelligence (AI) based models capable of addressing intricate challenges such as image classification [16–19], object detection [20–24], and image segmentation [25–27]. With the maturation of Machine Learning (ML) techniques in recent decades, there has been a notable surge in interest regarding the development of automated systems for medical diagnosis. The emergence of Deep Learning (DL) has brought about new possibilities in the area of medical imaging analysis and diagnosis. One of its arguably most successful models is Convolutional Neural Network (CNN), a widely used type of DL algorithm, well known for its ability to capture spatial correlations within image pixel data hierarchically. DL models are data-hungry, which means that the full potential of a model is often inhibited by a lack of data. This issue is particularly pronounced in the medical context, where data acquisition is costly and dependent on medical expertise. The availability of good-quality datasets that are clinically relevant is one of the key challenges that developers face. The lack of good-quality datasets for use in the development of AI systems may hinder their effectiveness and potential benefits. Data that are not of sufficient quality for the intended purpose can also lead to many problems, such as bias and errors [28]. Even though AI-based models showed promising results, one of the limitations in the current integration of AI into clinical practice is the difficulty in understanding their internal mechanisms and explaining the rationale behind their decisions. For AI-based applications to be effectively incorporated into the clinical routine, it is essential to comprehend and interpret their outputs. AI systems should ideally be used to complement human intelligence capabilities and decision-making, in order to reduce errors in diagnostic classification or prognosis. They have shown promise in medical imaging tasks [29–31], enabling improved tumor detection, classification, and prognosis assessment. In recent years, the application of these cutting-edge methods to brain tumor assessment has also seen a marked increase in interest [32–36], further showcasing the potential of DL in revolutionizing the field of neuro-oncology. Ethics of AI in Healthcare In recent years, Europe has seen the emergence of regulations aimed at addressing the complexities of the digital age. The General Data Protection Regulation (GDPR) set the stage in 2018 by establishing strict guidelines for data privacy and protection. In 2021, the United Nations Educational, Scientific and Cultural Organization (UNESCO) elaborated the Recommendation on the Ethics of AI, which calls for legal frameworks to govern these technologies and ensure they contribute to the public good, emphasizing transparent, unbiased, inclusive, and fair solutions. Similarly, the WHO released the guidance on Ethics and Governance of AI for Health [37], which emphasizes the need for AI ethical regulations, especially in the healthcare domain. This guidance endorses a set of key ethical principles: (1) Protecting human autonomy in medical decisions; (2) Promoting human well-being and safety and the public interest; (3) Ensuring transparency, explainability and intelligibility by developers, medical professionals, patients, 005
users, and regulators; (4) Fostering responsibility and accountability; (5) Ensuring inclusiveness and equity irrespective of sex, gender, income, race, ethnicity, sexual orientation, among others; (6) Promoting responsive and sustainable AI. The WHO guidance has been further extended recently in 2024 to include large multimodal models. In response to the growing need for ethical and secure AI, the European Commission has recently released the AI Act [38,39] as the first legal framework dedicated to AI that aims to mitigate the risks associated with AI technologies while establishing measures to promote the development of trustworthy AI. Multimodal Data Integration The integration of diverse data types and modalities presents opportunities to enhance the robustness and precision of diagnostic and prognostic models, thereby bridging the gap between AI advancements and clinical application. In recent years, AI models have exhibited impressive efficacy across numerous clinically relevant tasks, even those that may pose challenges for human interpretation. These models can draw upon a range of data sources, such as radiology, pathology, genomics, and electronic health records, to deliver precise patient predictions. In order to get closer to the clinical practice routine, models that can handle multiple types of data are gaining popularity. Constraining to a single modality can significantly reduce the clinical potential. According to [40], multimodal models could extend AI-based systems throughout diagnosis and clinical care with a long-term vision of developing a “generalist medical AI”. Model-centric vs. Data-centric AI In the realm of AI research, the primary focus lies in developing a model good enough to handle the noise in the data, with a minority focusing on understanding and improving the data. As stated by Andrew NG [41], more than 90% of the research papers are model-centric. The model-centric approach focuses on developing experimental research to improve the ML model performance while keeping the data the same and improving the code or model architecture. On the contrary, in a data-centric view, the consistency of the data is paramount. This approach involves systematically improving the datasets to increase the performance of ML applications. Despite living in an era where data is at the core of every decision-making process, AI initiatives often neglect and mishandle data. Consequently, countless hours are wasted fine-tuning models based on faulty data, potentially compromising accuracy despite optimization efforts [42]. As data serves as the foundational material for data-driven approaches, it is essential to prioritize high-quality, adequately sized, and consistently labeled datasets. While singular viewpoints dominate, adopting a hybrid approach that incorporates both model and data considerations may offer the most effective strategy. According to the WHO Regulatory Considerations on AI for health [28], data quality greatly influences the success of AI solutions’ safety and effectiveness. 006
MOTIVATION & OBJECTIVES Glioma tumors, with their complex and heterogeneous nature, pose significant challenges within the medical field, demanding the development of innovative and robust diagnostic and prognostic solutions. Automatic reliable identification and characterization of brain tumors for their later removal is a growing public health concern worldwide. The current practice for glioma diagnosis has some associated limitations, such as clinicians’ skills, inter-expert variability, and large waiting times to obtain results. These issues highlight the importance of developing computer-aided diagnosis systems to help radiologists interpret and quantify abnormalities from brain images through the automation and standardization of tedious and time-consuming tasks involving identification, classification, or segmentation. Algorithms that could at least partially automate the process of tumor diagnosis, malignancy quantification, and prognosis assessment would be very valuable. AI-driven approaches have emerged as promising tools in this regard. By extracting quantitative features from medical imaging data, they can contribute to the development of non-invasive and pre-operative decision-making systems. This thesis was originally conceived to develop an accurate automatic system to support clinicians in the complex realm of glioma diagnosis and grading. However, this research faced unexpected hurdles. Despite genuine intentions to contribute meaningfully to the field, efforts to procure essential data from hospitals and secure the cooperation of clinicians unfortunately resulted in setbacks. These challenges unveiled a critical aspect of the broader landscape - the indispensable need for collaboration among diverse stakeholders. The integration of AI in healthcare holds immense promise, yet its successful implementation relies heavily on multidisciplinary and interdisciplinary collaboration involving domain and AI experts. Although data is at the cornerstone of every decision-making process, the focus is often primarily on the models. However, despite the potential for transformative change of AI in healthcare applications, significant challenges remain. One major concern is the opacity of AI systems, which can hinder trust and adoption among healthcare professionals and patients alike. This thesis emphasizes the critical importance of large-scale, robust datasets and the establishment of reliable processes as foundational elements for developing solutions that can be seamlessly integrated into healthcare routines. By prioritizing transparency, data quality, and process reliability, we aim to build AI systems that not only perform well but are also trusted and widely adopted in clinical settings. 007
This thesis is structured around the following key objectives: •Review the openly available databases for brain tumor classification research. Thisreviewaimstoassessthequalityoftheopendatabasesforsupportingdeeplearning research in brain tumor classification, particularly for glioma grading, and to identify potential gaps or the need for new datasets. •Review the current state-of-the-art deep learning-based methodologies for brain tumor classification from MRI, with a focus on glioma grading. This review aims to identify gaps in the current methods and highlight areas for further research. •Evaluate various MRI pre-processing methods to enhance the distinction of glioma grades using deep learning. This evaluation aims to identify the most effective strategies for an accurate and robust glioma classification, considering techniques utilized in the current state of the art. •Develop an accurate and robust deep learning-based model to classify the WHO glioma grade from MRI. This involves comparing various data preparation and architecture strategies to advance toward a clear understanding of how these factors impact the performance of deep learning-driven glioma grading systems. •Compare the efficacy of 2D-based and 3D-based methodologies to extract meaningful patterns in MRI using deep learning. The goal is to develop a feasible, comprehensive, and robust model suitable for clinical implementation, balancing computational efficiency, performance, and interpretability. •Develop and evaluate a multi-task framework design to differentiate glioma grades and the presence of molecular features simultaneously. This goal seeks to explore the contribution of molecular characteristics to glioma grade classification aiming to align more closely with clinical practice. •Evaluate possible demographic disparities in the proposed models. The aim is to promote transparency and fairness in performance evaluation and result reporting, ensuring the responsible utilization and distribution of AI technologies and mitigating biases present in the algorithms. •Foster collaborative multidisciplinary and interdisciplinary efforts. The goal is to raise awareness of the need for further research and to promote the creation of meaningful, up-to-date, large-scale, and robust datasets. 008
THESIS OUTLINE & CONTRIBUTIONS This dissertation is structured into three parts. The first, background, provides essential clinical and methodological concepts necessary for understanding the thesis, accompanied by an extensive review of the current cutting-edge research in the field. The second part, practice, delves into the specific methods employed and the corresponding results obtained. Lastly, the closing part, provides an in-depth discussion of the research conclusions and potential future directions. The general outline is illustrated in Figure 0.1. Figure 0.1. General thesis outline, including its parts and chapters. 009
This part is organized into three chapters, each laying the foundational groundwork essential to this thesis. Chapter 1provides a thorough overview of the clinical foundation of this study. It begins with the definition of brain cancer and traces the evolution of the World Health Organization glioma classification to its current form. It subsequently delves into MRI, highlighting its role as the principal non-invasive diagnostic tool in neuro-oncology. Chapter 2explores the methodological foundation of the thesis. Firstly, we delve into the fundamentals of deep learning for image classification, covering essential methods and the most successful network architectures. Secondly, we discuss various strategies employed in deep learning for medical imaging classification assessment. Next, we explore approaches aimed at addressing the challenge of data scarcity in deep learning. Lastly, we examine interpretability mechanisms. Chapter 3reviews the latest advancements in MRIand DL-based methodologies for brain tumor classification, with a particular focus on glioma grading. This chapter provides an in-depth analysis of the state-of-the-art techniques, the intricacies of the methodologies employed, and the datasets utilized in these studies. 017
1 CLINICAL BACKGROUND
1.1 THE BRAIN CANCER The brain is comprised of two distinct categories of specialized cells: neurons and glial cells. Neurons are highly specialized in function and interconnected, while glial cells provide structural support. Healthy cells exhibit a regulated growth pattern, halting proliferation when a sufficient number of cells are present. This ensures the replacement of old or damaged cells. In contrast, tumor cells proliferate uncontrollably and fail to adhere to the typical interactions observed between normal cells. A primary brain tumor arises from an uncontrollable proliferation of cells originating within the brain itself, which is different from brain metastases, as the latter originate from cells elsewhere in the body. Most brain tumors are not linked with any specific cause, but some factors can raise the risk, including age, genetics, lifestyle, and environmental factors. Radiation exposure is the best-known risk factor for brain tumors and other common factors include age, family history, or bad habits (alcohol, smoking, or poor nutrition) [1,43]. Brain tumors cause different symptoms by pressing on the brain or the spinal cord, taking up space inside the skull when they grow. Headaches, seizures, vision loss, weakness, and speech impairment are among the most common symptoms [1,43,44]. There are hundreds of types of brain tumors. Among the benign varieties, meningioma originates in the meninges and accounts for more than 30% of all brain tumors. Approximately 85% of meningiomas are non-cancerous and exhibit slow growth rates. Pituitary adenomas, deriving from the pituitary glands, constitute around 10% of brain tumors, characterized by their gradual development [45,46]. Gliomas originate from glial cells and account for over 30% of all brain tumors, encompassing 75% of malignant primary tumors [3] and are often accompanied by a poor clinical prognosis. Unlike meningiomas and pituitary adenomas, gliomas possess a diffuse nature, spreading widely throughout the brain and extensively infiltrating surrounding brain tissue. Gliomas can occur everywhere, with the densest occurrence found in the frontal lobe [47]. The male-to-female ratio among affected patients is about 3:2. The peak age at onset for anaplastic astrocytomas is in the fourth or fifth decade, whereas glioblastomas usually present in the sixth or seventh decade [10,48]. 023
1.2 THE WHO CNS CRITERIA According to the initial WHO CNS classification, gliomas were primarily diagnosed histopathologically conceptually based on the presumed cell of origin [49]. Brain tumors exhibiting cell characteristics similar to astrocytes were classified as astrocytomas, while those resembling oligodendrocytes were categorized as oligodendrogliomas. WHO CNS classification has been evolving since 1979, where only miotic activity, necrosis, and tumor infiltration were determinant elements. In 1993, the immune histochemistry status was added to the diagnosis criteria. Later in 2000, the genetic profile started to play a role in determining the grade of the CNS tumors, followed by the addition of the histological variation in 2007, and further extended to molecular features in 2016. Until the 2016 WHO classification, diffuse gliomas were classified based on their morphologic features, with molecular testing merely playing an ancillary role [10]. The WHO 2021 CNS fifth edition, featured substantial changes by moving further to advance the role of molecular diagnostics [8]. For the first time in the WHO classification, adult-type gliomas were separated from the pediatric-type. The molecular parameters that have now been added as biomarkers of grading and for further estimating prognosis within multiple tumor types include IDH mutation, 1p/19q co-deletion, O-6-methylguanine-DNA methyltransferase (MGMT) methylation, telomerase reverse transcriptase (TERT) promoter mutation, epidermal growth factor receptor (EGFR) amplification, and tumor protein TP53 mutation. IDH mutation and 1p/19q co-deletion are linked to improved prognosis and treatment response. On the other hand, the MGMT promoter has been shown increased sensitivity to alkylating agents such as temozolomide, whereas alterations in the TERT promoter, EGFR amplification, and TP53 mutation are associated with more aggressive phenotype and poorer treatment response [50,51]. Previous WHO CNS classification releases have considered a four-stage grading system across different tumor types, from the most benign to the most malignant. For example, an anaplastic1astrocytoma and an anaplastic meningioma would be assigned grade 3 even though they are biologically unrelated and have different clinical courses. Conversely, the WHO CNS fifth edition (2021) moved grades within types aligning with tumors in other organ systems, where the tumor is graded according to its particular grading system [8,10,49]. The scheme presented in Figure 1.1 shows how the entities from WHO CNS 2016 evolved in WHO CNS 2021 criteria. 1Anaplastic cells proliferate rapidly and have little or no resemblance to normal cells 024
Figure 1.1. Flow diagram that illustrates the evolution of the WHO 2016 classification into the WHO 2021 classification of gliomas. Solid lines denote strong correlations between the two classifications, while dotted lines denote how a WHO 2016 disease entity would likely, but not definitively, be defined. NOS: Not Otherwise Specified. Adapted from Whitfield and Huse [49]. 025
2 METHODOLOGICAL BACKGROUND
2.1 MRI PRE-PROCESSING Image pre-processing plays a critical role in enhancing the quality, consistency, and comparability of MRI data for further analysis. In this section we present a series of essential steps designed to mitigate various sources of noise inherent in MRI acquisition. These steps involve image alignment or registration, image resampling to a common resolution, bias field correction, skull removal, contrast enhancement, and intensity normalization. 2.1.1 REGISTRATION Image registration is a crucial step in medical imaging, that involves aligning and combining data from multiple sources into a unified coordinate system. This process geometrically aligns multiple images captured at saying times, viewpoints, or by different sensors [62,63]. Image registration involves identifying a spatial transformation that maps multiple images to a common coordinate frame [64]. There are three different coordinate systems and spaces. The world coordinates refer to a Cartesian coordinate system in which the MRI scanner or the patient is positioned. The anatomical space consists of the three planes that describe the standard anatomical position of a human: axial, coronal, and sagittal. The image coordinate system (i, j, k)describes how an image (voxels) was acquired concerning the anatomy. In essence, registration involves modifying an image so that it aligns with the physical space and shape of a reference image. The alignment of images between different coordinate spaces can occur through two primary approaches: inter-subject and intra-subject registration. Inter-subject registration, commonly known as spatial normalization, involves mapping MRI data from multiple subjects into a standardized anatomical template. Frequently used templates include the Talairach and MNI atlases, providing a common reference for analysis across subjects. Conversely, intra-subject registration, often termed co-registration, focuses on aligning data from a single subject’s distinct anatomic imaging studies, such as different modalities or time points, to ensure internal consistency within that individual’s data. The idea of interand intra-subject registration is illustrated in Figure 2.1. Anatomical templates refer to an average MR volume that contains an average 035
probability of different tissues at a particular spatial location in their voxels. Since template generation is computationally expensive, it is common to use publicly available templates such as MNI305[65] in which 305 T1-weighted MRI brains were linearly co-registered to 241 brains that had been co-registered to the Talairach atlas. The process of image registration is typically iterative, adjusting the transformation parameters to minimize a cost function that quantifies the quality of alignment between two images [66]. Commonly used cost functions include the Sum of Square Differences (SSD) and Normalized Cross-Correlation (NCC) for intra-modality registration, and Mutual Information (MI) and Normalized Mutual Information (NMI) for inter-modality registration. The spatial relationships between the images can be modeled as rigid transformations (involving rotations and translations), affine transformations (including scaling and shearing), or deformable transformations (allowing non-linear adjustements). The parameters optimized during registration correspond to the Degrees of Freedom (DoF), which define the number of independent ways that the transformation can vary [67]. For rigid registration, typically used for intra-subject alignment, the transformation involves 6 DoF: 3 for rotations and 3 for translations. For affine registration, commonly applied in inter-subject registration, the transformation expands to 12 DoF, adding 3 for scaling and 3 for shearing to the rigid parameters [68]. Deformable transformations are used for complex alignments and have variable and often in the order of hundreds or thousands DoF. Figure 2.1. Illustration of interand intra-subjects MRI registration. 036
2.1.2 RESAMPLING MRI scans can vary in resolution and voxel sizes depending on the acquisition system. Voxel size, analogous to a pixel in three dimensions, is crucial for image quality. Voxel size is determined by both the pixel size and slice thickness. Pixel size, dependent on both the field of view and the image matrix, directly impacts in-plane spatial resolution. Typically expressed in millimeters, pixel size commonly falls within the range of 0.5 to 1.5 mm. The smaller the pixel size, the greater the image spatial resolution [69]. In MR image processing, variations in voxel size and resolution across scans can pose challenges for subsequent analysis and comparison, especially when using analytical techniques that require uniform input dimensions. To mitigate these disparities, resampling involves the standardization of the spatial resolution across the MRI images to ensure uniform dimensions. The process involves interpolating the intensity values of the original image to generate a new image with desired spatial characteristics, such as pixel size, slice thickness, or overall dimensions. 2.1.3 BIAS FIELD CORRECTION MRI images might exhibit variations in intensity, which are commonly referred to as inhomogeneity artifacts or bias fields, along with noise. The image formation model can be represented by the equation: ν(x) = u(x)f(x) + n(x)(2.1) where the image ν(x)is a combination of the original uncorrupted image u(x), the bias field f(x), and the noise n(x), assumed to be independent and Gaussian. Bias Field Correction (BFC) consists of an iterative algorithm that gradually estimates and rectifies intensity variations across the image. This method relies on a multi-resolution approach, where the image is successively smoothed at various resolutions to estimate the bias field. N4ITK’s (N4 Bias Field Correction) [70], refinement over N3 (non-parametric non-uniform normalization) [71] includes a more robust B-spline fitting strategy for bias field estimation, providing enhanced accuracy and efficiency in correcting intensity inhomogeneities. By addressing these variations, BFC not only improves the overall image quality but also facilitates more accurate and consistent quantitative analysis in MRI-based studies. 037
Figure 2.2. Raw (top) and bias field corrected (bottom) images. A color map has been applied to facilitate the visualization of tissue differences after BFC. 2.1.4 SKULL-STRIPPING The main objective of the skull-stripping step is to efficiently isolate the cerebral region of interest from non-cerebral tissues, enabling DL models to focus exclusively on those brain tissues. Skull-stripping is not only an essential pre-processing step but also crucial for privacy preservation. Traditionally, this step was performed through manual segmentation, but in the past decade, significant efforts have been made to automate this process with considerable success. For instance, BET (Brain Extraction Tool) [72] employs a deformable model with adaptive intensity thresholds to isolate the brain from non-brain tissues. More recently, DL-based methods have been proposed in the literature. HD-BET, implemented by Isensee et al. [73], is an ANN inspired in the UNet [25] capable of extracting the brain from the four conventional MRI modalities (flair, t1ce, t1, and t2). This model was trained using both healthy and pathological 3D volumes, demonstrating its robustness and versatility in various clinical scenarios. Another important contribution was made by [74], who compared multiple DL-based models trained on a large brain tumor cohort and proposed a “modality-agnostic” approach, that is robust to the presence of missing modalities. 038
Figure 2.3. Raw (top) and skull-stripped (bottom) images. 2.1.5 CONTRAST ENHANCEMENT Contrast enhancement methods play a pivotal role in highlighting intricate details and enabling improved analysis of images. One widely used technique in image processing for this purpose is histogram equalization. This method aims to stretch the histogram of intensity values within an image, effectively distributing them to achieve a more uniform distribution. This redistribution of intensities helps unveil subtle details, enhancing the overall clarity of the image. The implementation of histogram equalization is as follows. Consider I(x, y)as an image with intensity values at pixel (x, y). The histogram h(i) represents the count of occurrences of each intensity level iin the image I(x, y). hist(i) = X pixels(x,y) p(I(x, y) = i)(2.2) where p(I(x, y) = 0) is the probability of an occurrence of a pixel of level iin the image I(x, y). The Cumulative Distribution Function CDF(i)accumulates the histogram values up to a particular intensity level i: CDF(i) = X j≤i h(j≤i)(2.3) 039
The normalization of the CDF ensures that its range remains between 0 and 1. CDF′(i) = CDF(i) CDF(max(i)) (2.4) The original intensity values I(x, y)are mapped to new intensity values I′(x, y)using the inverse of the normalized CDF function, redistributing the intensity values and thus enhancing the contrast. I′(x, y) = CDF ′−1(CDF ′(I(x, y))) (2.5) which maps original intensity values I(x, y)to new intensity values I′(x, y). Figure 2.4. Raw (top) and equalized (bottom) images. 2.1.6 INTENSITY NORMALIZATION A technique adopted to rescale intensity values of MRI scans to a numeric range, rendering them consistent across the dataset. This process mitigates scale-related disparities. Two prominent approaches commonly applied to MRI data as input for DL models are min-max 040
normalization and z-score normalization. Min-max achieves its goal by rescaling intensity values within MRI scans, spanning their range between 0 and 1, through the formula Xnorm =X−Xmin Xmax −Xmin (2.6) In contrast, z-score normalization or standardization, transforms the distribution of intensity values by centering it around a zero mean and standard deviation of value 1, through the formula Xstd =X−µ σ(2.7) Normalizing intensities is a widely acknowledged process to enhance the efficiency of DL networks. By adjusting inputs to a range around zero, this practice enables more consistent weight updates within the model. Consequently, it mitigates the issue of fluctuating step sizes, thereby alleviating problems associated with exploding gradients, a phenomenon that occurs when the gradients of the network loss become excessively large concerning the weights. Since the network output is a linear combination of each feature vector, this means that the network learns weights for each feature that are also on different scales. Otherwise, the large feature will simply drown out the small feature. Then during gradient descent, to “move the needle” for the Loss, the network would have to make a large update to one weight compared to the other weight. This can cause the gradient descent trajectory to oscillate back and forth along one dimension, thus taking more steps to reach the minimum. Data normalization is not only important between samples but also within. When images have multiple channels, the intensities in one channel can be high compared to the other channels and the values of this channel would become dominant, diminishing the others [75]. 2.2 MRI PROCESSING MRI data constitutes three-dimensional (3D) volumes that enable visualization across axial, coronal, and sagittal anatomical planes. These volumes are composed of 3D cubes known as voxels, in contrast to the standard two-dimensional (2D) images which are made up of 2D squares called pixels. While 3D volumes offer comprehensive data, they can also be decomposed into 2D slices. MRI images are commonly provided in NIfTI (Neuroimaging Informatics Technology Initiative) format or DICOM (Digital Imaging and Communications in Medicine). NIfTI is a standard format for storing and exchanging neuroimaging data, while DICOM is a widely used standard for medical imaging. 041
Figure 2.7. Architecture of a 2D vanilla CNN with two convolutional layers, two pooling layers, one flattened layer, and one fully connected layer. Several networks have made significant contributions to the world of DL. The first CNN, LeNet5, was presented by LeCun et al. [83] in 1989 and it was called LeNet5. It consists of multiple convolutional and pooling layers, followed by a fully-connected layer. Further to LeNet’s emergence, several CNNs made significant contributions to the world of DL. AlexNet [16], was the first architecture to employ max-pooling layers, ReLU activation functions, and dropout regularization method. It was the first deep CNN trained on ImageNet and won the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2012. In the ILSVRC 2014, Szegedy2014 won with a CNN called GoogleNet, achieving a significantly low error rate. Its main contribution was the implementation of an inception module that allows for more efficient computation and a substantial reduction in the number of parameters in the model. The evolved version of GoogLeNet is calledInceptionNet. Another relevant network, VGGNet, was introduced by Simonyan and Zisserman [18] with the particularity of using small-size convolution filters that allow increasing the depth of the network, while significantly reducing the number of necessary parameters. Among the most popular CNNs, ResNet was developed by He et al. [19] to win the ILSVRC challenge in 2015. This work proposed the use of residual blocks, which help to address the problem of vanishing gradients when increasing the depth of the network. In 2017, the concept of dense blocks was introduced with DenseNet [84], in which the feature maps obtained from each layer are concatenated are feed the subsequent layer. In 2019, Tan and Le [85] introduced EfficientNet, which introduces the idea of identifying optimal depth, width, and resolution. More recently, Transformers, initially with the ViT [27], gained popularity in the field of computer vision through the addition of self-attention layers. DL models are considered data-hungry since they require substantial amounts of data for effective training. In the realm of medical data analysis, a primary challenge, as previously mentioned, is the inherent data scarcity. To address this challenge, common solutions include the application of Data Augmentation methods and Transfer Learning (TL) 048
techniques. Data Augmentation techniques are a crucial strategy to mitigate the challenge of limited annotated data in medical image analysis. These methods encompass a range of transformations applied to existing images, effectively expanding the dataset in terms of both size and diversity. Former approaches involve a wide range of geometric modifications such as rotation, scaling, flipping, cropping, zooming, or color changes. Beyond traditional augmentations, advanced methods like Generative Adversarial Networks (GANs) [86] are used to generate new synthetic and realistic examples. The idea behind TL is to transfer the knowledge from a source task to a new domain. TL involves using a pre-trained model, typically trained on large and diverse datasets, as a starting point and adapting it for a new task, especially when the available data may lack diversity and representation. Different strategies are available depending on the similarity between the two domains. When the new task closely resembles the original task, a practical approach is to initiate with the pre-trained model, freezing certain layers that capture low-level features, while fine-tuning high-level layers that adapt to the new domain and capture more complex features. Conversely, when the new task diverges significantly from the original task, the approach involves unfreezing all the layers of the pre-trained model, allowing them to adapt and learn specific features from the new dataset, constituting a weight initialization strategy. Widely used pre-trained CNNs, such as ImageNet [87] or MS-COCO [88], have been originally developed from 2D large-scale datasets. However, a notable challenge when dealing with medical image data is the limited availability of large and diverse 3D datasets for universal pretraining [89]. Transferring the knowledge acquired from the 2D to the 3D domain proves to be a non-trivial task, primarily due to the fundamental differences in data structure and representation between these two contexts. To tackle this challenge and address the limitation of limited data, a broadly used strategy is to decompose 3D volumes into individual 2D slices within a determined anatomical plane. However, the decomposition of 3D volumes into individual 2D slices introduces a potential data leakage concern. This issue arises when 2D slices from the same individual inadvertently end up in both the training and testing datasets in an analytical pipeline. Such data leakage can lead to overestimations of model performance and affect the validity of experimental results. In addition, it is important to note that this approach comes with the trade-off of losing the 3D context present in the original data. 2.5 MODEL EVALUATION In order to evaluate the generalization ability in ML, various methodologies are available. The hold-out method involves dividing the dataset into two subsets, a training set used to train the model and a testing set to evaluate the performance. It is also common to add a third set, known as validation, for monitoring and tuning hyperparameters to optimize 049
model performance. However, a primary drawback of this method is its susceptibility to selection bias, particularly when the sample size is small. To address this limitation, K-fold Cross-Validation (CV) emerges as a powerful alternative. In K-fold CV, the dataset is partitioned into K subsets (or folds) and the model is trained and evaluated K times, each time utilizing a different fold as the validation set and the remaining folds for training. Performance metrics are computed for each validation set and, typically, the average performance is reported. This approach provides a more robust estimate of model performance compared to the hold-out method, especially in scenarios with limited data. It is also common to have an independent test set to ensure that the final model evaluation is unbiased and representative of its true generalization performance. Figure 2.8. K-fold Cross-Validation scheme. 2.6 DATA INTEGRATION Data integration involves combining information from various modalities, sources, or perspectives. By leveraging multiple data sources, models can capture a richer set of features and complementary contextual information, enhancing the robustness and accuracy of the predictions [90,91]. Data integration strategies are divided in terms of the stage at which the information fusion occurs. Early Fusion 1integrates raw data from various modalities at the initial stage of the processing pipeline. The fused representation is then fed into a single model. An advantage of early fusion is its computational efficiency due to a single model. However, a drawback is that it might struggle to effectively capture complex interactions between different modalities, potentially leading to information dilution. Intermediate Fusion 2Occurs after some initial processing has been performed independently on each modality. The extracted features are then combined, and the fused 050
representation undergoes further processing to produce the output scores. The advantage of intermediate fusion is that the model can learn rich interactions across each modality. Late Fusion 3Involves training separate models are trained for each modality independently, and the output scores are combined at the final stage. This approach is akin to an ensemble method. While late fusion can increase the robustness of the model by mitigating the impact of errors from a single model, it does not enable the model to learn complex interactions across modalities. Additionally, it can be computationally intensive, as it requires training multiple models. Figure 2.9. Data integration strategies. 1Also known as input-level fusion. 2Also known as feature-level fusion. 3Also known as decision-level fusion. 051
2.7 MULTI-TASK LEARNING In traditional DL, a separate model is typically trained for each specific task. Multi-task Learning (MTL) was first introduced by Caruana [92] and consists of training a single model to perform multiple tasks simultaneously. In the context of DL, this involves training a neural network to learn representations that are useful for multiple related tasks. The concept is inspired by human learning, for new tasks we apply the knowledge we have acquired by learning related tasks. The main advantage is that the model can learn the context between different features. By sharing representations between related tasks, we can enable our model to generalize better on our original task [93]. Multi-task learning improves generalization by leveraging the domain-specific information contained in the training signals of related tasks [94]. By jointly learning multiple tasks the model facilitates TL, in the sense that the knowledge gained from learning one task can be transferred to improve performance on related tasks, but it can also implicitly learn the relationship between tasks. As different tasks have different noise patterns, a model that learns multiple tasks simultaneously can learn a more general representation reducing the risk of overfitting [93]. Figure 2.10. Multi-task learning. In multi-task DL we use one network that performs multiple tasks simultaneously which involves hard or soft parameter sharing of hidden layers. Hard parameter sharing involves the parameters of hidden layers being shared across all tasks while keeping several task-specific output layers. In soft parameter sharing, each task has its model with its parameters, but the parameters are encouraged to be similar across tasks through regularization techniques. This last approach requires high computational cost. Multi-task learning emerges as a useful approach when there is an abundance of labeled data for one task that can be shared with another task with less labeled data [95]. 052
2.8 EXPLAINABLE DEEP LEARNING Even though sophisticated algorithms can produce fancy outcomes, they frequently come at a cost – opacity. Artificial Neural Network models, and DL models in particular, are often referred to as “black boxes” due to the opacity of their internal mechanisms. This lack of transparency poses a barrier to their acceptance and approval for clinical use [96,97]. Clinical decisions can be considered high-risk tasks since they have a strong impact and the model behavior comprehensibility is desired to be high. The more complex the architectures, the more difficult the explanation of how and why a particular network prediction is obtained, or the elucidation of which components of the complex system contributed essentially to the obtained decision [98]. In recent years, there has been an increasing emphasis on the necessity for these models to be both transparent and interpretable. Interpretable ML refers to methods and models that make the behavior and predictions of machine learning systems understandable to humans. Interpretability can be defined as the degree to which a human can understand the cause of a decision [99]. An ML model is considered interpretable due to its simple structure, such as linear models or decision trees. Explainability is closely tied to interpretability but involves the useofadditional post-hoc methodstocomprehendthemodel’sdecision. Ittacklesthe problem whereinhumanusersfacedifficultyindirectlycomprehendingtheintricatebehaviorsofDeep Neural Network (DNN) or elucidating their underlying decision-making mechanisms. Explainability in CNNs is frequently addressed through feature visualization. Pixel Attribution methods, also known as Saliency Maps, highlight the pixels that are most relevant to a specific classification task. For enhancing image classification transparency and understanding, several techniques have emerged, broadly categorized into two main types: those that rely on perturbing features to observe changes in the model output and those that utilize gradient information to ascertain the influence of each pixel on the decision-making process. A popular perturbation-based method is Local Interpretable Model-Agnostic Explanation (LIME), as introduced by Ribeiro et al. [100] in 2016. The key intuition behind LIME is to generate a new dataset containing perturbed samples by turning some of the components “off” and the corresponding predictions of the model. For each perturbed sample, LIME computes the class probability and then learns an interpretable model (e.g. linear regression) which is weighted by the distance between the perturbed image and the original image [101]. The bigger the weight, the bigger the importance of the feature. Mathematically, ξ(x) = arg min g∈G L(f, g, πx) + Ω(g)(2.24) where Lis a measure of how unfaithful gis in approximating fin the locality defined by πx. Ω(g)is a measure of the complexity of the explanation g∈G. 053
Gradient-based methods utilize the gradients of the output with respect to the extracted features calculated through backpropagation to reveal which parts of the input contribute the most to the final decision. A widely used gradient-based method is Gradient-weighted Class Activation Mapping (GradCAM) [102]. An intuitive idea of GradCAM is to understand which parts of the image the convolutional layer gives more attention for a certain decision by analyzing which regions are activated in the feature maps. In order to decide the importance of each feature map to the class of interest, GradCAM weights each pixel of each feature map with the gradient averaging over the feature maps [101]. Then, for each input feature (pixel) it produces a value between 0 and 1 representing a low or high contribution in the class prediction, which results in a heatmap that highlights the regions that contribute either negatively or positively to the decision. The objective is to produce a localization map, represented mathematically as: Lc GradCAM ∈Ruxv =ReLU |{z} Pick positive values X k αc kAk!(2.25) Here, uand vrepresent the width and the height, respectively, of the explanation map for a given classc. The termαkdenotestheimportanceofthek-thfeature mapforthe targetclassc, and Akrefers to the activations of a convolutional layer’s feature map. The application of the ReLU function ensures that only positive contributions to the classof interest are considered. To calculate αc kthe equation is as follows: αc k= global average pooling z }| { 1 ZX iX j δyc δAk ij |{z} gradients via backprop (2.26) This represents the global average pooling of the gradients computed via backpropagation across the feature map Ak.Ak ij refers to the activation at location (i, j)of the feature map Ak and Ycrefers to the score for each class c. The variable Zaccounts for the total number of pixels in the feature map. The class score yccan be expressed as: Yc=X k wc k |{z} class feature weights GAP z }| { 1 ZX iX j Ak ij |{z} feature map (2.27) where wc kare weights indicating the relevance of features within the map Akto the class c. 054
2.9 BIAS IN MACHINE LEARNING MODELS From a legal perspective, bias refers to an inclination or prejudice towards a particular individual or entity. In statistics, bias refers to the systematic deviation of an estimator from the true value of the parameter being estimated. In the realm of ML, bias translates to the systematic error in the model’s predictions from the actual outcomes. This implies the possibility of having an unbiased model in terms of ML while still possessing bias in legal or statistical terms. Algorithmic decisions are often perceived as fair or unbiased due to minimal human intervention, which theoretically eliminates human biases from the decision-making process. When we observe disparities, it does not imply that the designer of the system intended for such inequalities to arise, however, algorithmic decision-making may unintentionally reflect existing inequalities present in the society, which are inherently present in the data [103]. Machines are not inherently unethical [104], however, an algorithm is only as good as the data it works with [105]. If the dataset is skewed or imbalanced, the model may develop a propensity towards certain patterns or outcomes, potentially yielding unfair or discriminatory results for specific populations. This bias poses a significant risk, as it can lead to outcomes that unfairly impact certain groups or individuals, which in healthcare would be translated to inequitable healthcare outcomes. Recent works have demonstrated that ML models can exhibit bias based on certain features such as sex [106–108] or race [109,110]. As ML increasingly permeates decision-making processes, it becomes crucial to develop models that are not only accurate but also objective and fair. Three common notions of (un)fairness can be defined as follows [111,112]: Disparate treatment1A decision-making process suffers from disparate treatment if its outcomes change based on a change in the sensitive feature value with all other features being the same. Disparate impact2A decision-making process suffers from disparate impact if it grants a disproportionately large fraction of beneficial (or positive classification) outcomes to certain sensitive feature groups. Disparate mistreatment3A decision-making process suffers from disparate mistreatment if its misclassification rates are different for different sensitive groups. While disparate treatment accounts for direct unfairness where a decision-making process intentionally uses the sensitive information for the classification outcome, disparate impact 1Disparate treatment has been also referred to as direct discrimination [113]. 2Disparate impact has been also referred to as statistical parity [114] or demographic parity [115]. 3Disparate mistreatment has been also referred to as equality of opportunity [116] and predictive equality [114]. 055
and disparate mistreatment account for indirect unfairness, where the model unintentionally uses the correlation between sensitive features and class labels in a way that places a sensitive feature group at a relative disadvantage. In the context of fairness in binary classification, we can express the absence of unfairness concepts as follows [111]: No disparate treatment. A binary classifier does not suffer from disparate treatment if the probability that the classifier outputs a specific value of ˆygiven a feature vector xdoes not change after observing the sensitive feature z, P(ˆy|x, z) = P(ˆy|x)(2.28) No disparate impact. A binary classifier does not suffer from disparate impact if the probability that a classifier assigns a user to the positive class is the same for both values of the sensitive feature z, P(ˆy= 1|z= 0) = P(ˆy= 1|z= 1) (2.29) Notice that neither disparate treatment nor disparate impact depends on the subjects’ ground truth label (y). No disparate mistreatment. A binary classifier does not suffer from disparate mistreatment if the misclassification rates for different groups of people having different values of the sensitive feature zare the same. Approaches to mitigate bias can be classified into three categories: pre-processing, in-processing, and post-processing. Pre-processing involves removing inequities prior to model training, for example by incorporating new data. In-processing methods are applied during the learning process, by for example imposing constraints or regularization. Finally, post-processing involves adjusting the model’s predictions to correct any remaining biases [117]. 056
Figure 3.1. Publicly available datasets usage frequency from 2018 to 2024. For a detailed review, the author refers to Pitarch et al. [257], where Table A1 offers a detailed overview of thedatasets, focusing onessentialaspects suchasthedimensionalityof theimages, sample size, MRI details, and pre-processing methods used. Table A2 in the referred article, delves into the specifications of the employed DL models, highlighting the brain tumor classification task, data partitioning, architecture, and the reported performance metrics. 3.2.1 BRAIN TUMOR CLASSIFICATION OVERVIEW The performance, generalizability, and robustness of machine learning models are significantly impact by the size and diversity of the training data. Several studies have explored the impact of varying the size of the training data set on model performance [130,135,143,157, 162,164,218,235,248]. Their findings highlight the value of ensuring that a substantial volume of data is available for training, as it significantly contributes to the model’s ability to make more accurate and reliable predictions. Recent efforts have aimed at overcoming data scarcity challenges. Approximately 60% of the examined studies employed data augmentation techniques, and 40% incorporated transfer learning in a 2D domain as a viable solution. A number of these investigations [139,144,153,159,170,216,220,236] have demonstrated the advantages of increasing both the quantity and variability of the samples through the 064
inclusion of augmented images. Applying traditional data augmentation techniques, such as geometric variations from original images, was the most widely used strategy, while only a few studies opted for the use of DL generative models [147,159,168,228]. Several studies[183,220,239,258–260]haveintegrateddataaugmentationasan oversamplingtechnique to address the problem of imbalanced data in the context of brain tumor classification. Other works have explored the inclusion of multi-view 2D slices from sagittal, coronal, and axial planes, in addition to employing image flipping and rotations to augment the dataset [221]. Pre-trained models have demonstrated performance enhancements in the classification of glioma grades in several studies [223,224,235]. However, not all investigations have reported equivalent advantages when employing pre-trained models to discriminate between healthy andtumorous samples[187] or to differentiate tumor types[141]. Thesevariations infindings underscore the complexity of the observed performance disparities, which may not be solely ascribed to the classification task itself, but also be influenced by intrinsic dataset variations. The predominant approach for data partitioning involves the use of hold-out validation, with training and validation sets. This was followed by the adoption of K-fold cross-validation, which enhances the robustness of model evaluation. A less frequently employed method was the three-way split, which includes training, validation, and testing sets. In total, only 36% of the studies assessed their final results using an independent test set. Alanazi et al. [160], Decuyper et al. [233], Gilanie et al. [234] took a step further by assessing the generalizability and robustness of their models using external testing sets. While the authors of the Figshare dataset thoughtfully included a 5-fold CV setup alongside the data to promote comparability and reproducibility, it is still important to remark that a substantial majority of studies using this database continue to prefer custom data partitioning methods. Pre-processing techniques are pivotal in medical image analysis. A substantial 80% of the considered papers, provide insights into the specific pre-processing methodologies that were employed. Withinthissubset,35%employedregistrationtechniques,involvingregistrationto a common anatomical template and co-registration to the same MRI modality. Furthermore, 40% employed segmentation as a critical step to isolate the brain from the surrounding skull structures. Nearly half of the papers embraced normalization techniques to standardize the intensity of the image data before it was fed to the models. Attention has also been directed toward identifying the optimal MR image intensity normalization approach, which plays a critical role in enhancing model performance in DL models. Many studies prefer mean-std standardization [54,57,143,146,160,176,232,236,237,240,243,253,254,260–267], while others lean towards min-max normalization [138,155,162,173,177,182,183,212,215, 218,247,268]. Additionally, some studies incorporate contrast enhancement methods to improve the contrast and visibility of crucial anatomical structures. [143,149,151,154,156, 158,176,181,214,229,231,251]. Additionally, 30% of the papers extract the brain tumor region through methods such as bounding box delineation or tumor segmentation. In several studies [222,258,269], researchers investigated the advantages of utilizing the tumor area as opposed to the entire image, highlighting the significant benefits of concentrating on the tumor region rather than the entire image. 065
The ability of CNNs to automatically extract meaningful features from brain MRI images, as opposed to the conventional need for manual feature engineering in certain ML algorithms like RF, GrB, and SVM, has been emphasized by numerous studies. These studies underscore the potential of CNNs in revolutionizing the landscape of MRI feature extraction for enhanced accuracy and efficiency in brain tumor classification [61,172,193,206,210,217,270]. A majority of the reviewed papers (approximately 60%) utilized established state-of-the-art CNN architectures to obtain brain tumor classification. Among these, ResNet and VGGNet backbones were the most prevalent choices, closely followed by AlexNet, GoogLeNet, and Inception. In contrast, the remaining 40% of the papers concentrated on enhancing brain tumor classification by introducing novel model architectures. The inherent black-box nature of CNNs highlights the importance of delving into the comprehension of their predictions, especially in a medical context. Several studies [176,195,222,243,253,263,266] have applied post-processing explainability tools to validate that the network’s decision-making process aligns with the intended diagnostic criteria, thereby enhancing the reliability of CNN-based medical applications. Merely 70% of the reviewed studies disclosed the MRI modalities utilized for the analysis. Among these, close to 50% exclusively employed the T1ce sequence, while 26% used a combination of T1ce, T1, T2, and FLAIR sequences, 12% used three sequences, and the rest chose one sequence. Various strategies were employed to integrate information from multiple modalities. The prevalent method involved fusing them as input channels, comparable to the treatment of channels in RGB images. In their study, Ge et al. [221] evaluated the sensitivity of T1ce, T2, and Flair modalities in glioma grade classification. Their investigation highlighted the T1ce sequence as the most informative among these modalities. To further enhance the classification performance, they incorporated information from each source using an aggregation layer within the network architecture. Subsequently, similar ensemble learning approaches were adopted by Gutta et al. [61], Hussain et al. [250], Rui et al. [252]. Notably, Guo et al. [271] directly compared the performance of a modality-fusion approach, where the four MRI modalities were concatenated as a four-channel input, with a decision-fusion approach, where final predictions were derived through a linear weighted sum from the probabilities obtained through four independent pre-trained unimodal models. This study reinforced the notion of the T1ce modality’s significance in glioma subtype classification. Moreover, it revealed that any multimodal approach consistently outperformed unimodal models. While brain MRIs inherently capture 3D data, a notable observation is that over 80% of the studies conducted their analyses within a 2D domain, focusing on 2D MRI slices. Nonetheless, some investigations [54,57,212,221,222,227,229,232,233,238–240,246,249, 250,253,260–264,272,273] have actively explored the significance of incorporating 3D volumetric information into the realm of brain tumor classification. While 3D volumes inherently capture information from the three anatomical planes, 2D slices are restricted to a specific view. Among the studies that adopted a 2D approach, only 44% provided details 066
about the chosen anatomical plane. Among this subset, more than 50% utilized the three anatomicalviews, whileover40%exclusivelyemployedaxialviews. Decomposing3Dvolumes into individual 2D slices may introduce the potential for data leakage. Maintaining the reliability of the analysis is crucial for obtaining robust and trustworthy findings. It is worth noting, however, that only a limited number of studies that use multiple 2D slices [61,130, 135,146,149,158,182,207,210,212,221,223,227,228,234,235,237,252,254,266,267], explicitly detailed their approach to data splitting at the patient level, addressing this critical concern. An insightful comparison was carried out in the work of Badˇ za and Barjaktarovi´ c [137] between data splitting strategies at the patient and image levels. The findings elucidate that an image-wise approach yields accuracy results as high as 96% for brain tumor type classification, while a patient-level split demonstrates a higher degree of reliability with an accuracy of 88%. These results underscore the critical importance of utilizing a patient-wise training approach to assess the model’s generalization capacity. Similarly, Ghassemi et al. [139], Ismael et al. [140] also provided evidence of superior performance when using an image-wise split, further reinforcing the importance of thoughtful data splitting and reliable and transparent performance reporting. It is also important to note that 3D models operate on complete 3D volumes and are inherently structured at the patient level. This approach substantially reduces the likelihood of data leakage, thereby enhancing the reliability of the analysis and ensuring that the results faithfully represent the model’s performance. This aspect may provide a valuable perspective when interpreting differences in accuracy between 3D and 2D models. The integration of information from various data sources has garnered growing interest in the medical field. Brain tumors, due to their distinct features both at the histopathological and radiological level, have motivated numerous studies to explore the synergy between WSI and MRI data [54,57,260,264,273]. These investigations consistently highlight the richer information content in WSI as compared to MRI. However, they also reveal that combining data from both sources leads to improved overall performance in brain tumor characterization. Ensemble learning methods have also shown promise in combining predictions from multiple DL models on MRI to improve overall performance [36,152,155,164,172,173,184,189,192,195,210,212,216,269] and the outputs of radiomics and DL models [219,241,267]. 3.2.2 TARGETING GLIOMA GRADING In this subsection, we delve deeper into studies that have contributed to advancing our understanding of the classification of glioma grades. Earlier mentioned, the majority of the articles have assessed this question from the distinction of LGG versus HGG [36,144,145,176,206, 210,221–223,227–229,231–233,235,236,238–246,248–256,258,267–269,272,274] while several studies assessed the exact grade according to the WHO CNS criteria [61,129, 132–134,144,150,172,180,224–226,230,234,237,240,247,253]. Tables A.1 and A.2 provide a comprehensive summary of the data, methods, and results obtained in each study 067
for the differentiation of lower-grade gliomas versus glioblastoma and the exact WHO grade, respectively. The median accuracy for differentiating LGG and HGG is 95% when using a 2D model and91% witha 3Dmodel. In contrast, when differentiating WHO grades, theaccuracy is 92% and 71%, respectively. Interestingly, 3D models generally result in lower accuracy, especially for WHO grade differentiation. This could be due to several factors, including the fact that 3D models ensure no slice leakage is present. Moreover, 3D models require more data to train effectively due to its complexity and more computational power, what can lead to limitations such as small batch size or fewer training epochs. Due to these computational requirements, it is also common to decrease the resolution of the images, which can lead to a loss of details and impact the model’s ability to learn useful discriminative patterns. Yang et al. [223] proposed the usage of AlexNet and GoogLeNet for classifying LGG and HGG on a private database composed of 113 diagnosed glioma patients. In this study, the 2D T1ce axial images that contain at least 80% of the tumor visible were selected. Then, the image was cropped at the tumor bounding box. The data were split into the train, validation, and test subsets at the patient level. The best performance (96.80 ROC-AUC) was achieved with a pre-trained GoogLeNet. Despite the good overall performance, only accuracy and ROCAUC were reported which, along with the small sample size, are limiting factors in this work. Ge et al. [258] proposed a 3D CNN combined with a saliency-aware strategy to enhance tumor regions by reducing the pixel intensities outside the brain mask to one-third their original values. This model achieved a sensitivity of 90.48% for HGG and 86.67% for LGG, with an overall accuracy of 89.47%, compared to 84.21% without mask enhancement. In a related study, Pereira et al. [222] examined the effect of standardizing either the entire image or solely the brain region for the classification of 3D LGG and GBM images. Their findings indicated that standardization, applied specifically to brain masks, significantly improved performance when the model was trained on the tumor region, increasing accuracy from 87.70% to 95.24%. Interestingly, GradCAM maps revealed that when intensity standardization was applied across the entire image, the CNN considered the border of the brain as discriminative. This emphasizes the importance of using explainability methods for validation purposes. Tandel et al. [36] conducted an evaluation of various 2D DL-based models for classifying LGG and GBM utilizing Flair, T1, and T2 sequences independently. Their findings underscored Flair as the most pivotal sequence, followed by T2, and T1 as the least significant. In the literature, strategies involving aggregation layers, or combining sequences as RGB image channels have been used to leverage information from multiple modalities. In Ge et al. [221], a 2D classifier processing T1ce, T2, and Flair sequences independently was developed, with subsequent feature combinations through a linear aggregation layer. The reported results indicated T2 as the least informative sequence for glioma grading (69.84% accuracy), followed by Flair (75.40%), and T1ce (83.73%), with the aggregated model yielding the highest accuracy (90.87%). Similarly, Gutta et al. [61] reported consistent findings and introduced T1, which emerged as the least relevant sequence to grade gliomas. 068
In Rui et al. [252], a combination of T1, Flair, and T1ce was leveraged as input image RGB channels, achieving superior performance (86%) when compared to independent or dual-sequences combinations across various classification tasks (LGG versus GBM, grade 2 versus higher grades, and molecular subtypes). This approach was further reinforced in Coupet et al. [212] by incorporating T2, and it was concluded that the combination of sequences yielded better results (86.38%) than individual sequences in the problem of distinguishing healthy from glioma images. The study reported in Guo et al. [271] compared the modality-ensemble (87.80%) and modality-fusion (84.60%) methodologies for classifying glioma types into Astrocytoma, Oligodendroglioma, and Glioblastoma, both outperforming independent sequences with 71.90% accuracy using T1, 73.30% with T2, 74.20% with Flair, and 83.30% with T1ce. Overall, these studies provide evidence of the information provided by different MRI sequences for brain tumor classification, highlighting that a joint representation significantly enhances model performance. Banerjee et al. [227] classified LGG and HGG multi-sequence brain MRIs from TCGA and BraTS2017 data using multiple slice-based approaches. In their work, they provided a comparison of the performance obtained with CNNs trained from scratch on 2D image patches (PatchNet), entire 2D slices (SliceNet), and multi-planar slices through a final ensemble method that averages the classification obtained from each anatomical view (VolumeNet). The classification obtained with these models is also compared with pre-trained VGGNet and ResNet on ImageNet. The multi-planar method outperformed the rest of the approaches with an accuracy of 94.74%, and the lowest accuracy (68.07%) was obtained with pre-trained VGGNet. Unfortunately, TCGA and BraTS data share some patient data, which could involve an overlap between training and testing samples and hence be prone to data leakage. Ding et al. [241] combined radiomics and DL features using 2D pre-trained CNNs using single-plane images and performing a subsequent multi-planar fusion. VGG16, in combination with radiomics and RF, achieved the highest accuracy of 80% when combining the information obtained from the three views. Even though the multi-planar approach processes the information gathered from the axial, coronal, and sagittal views, it is still essentially a 2.5D approach, weak at fully capturing 3D contexts. Zhuge et al. [232] presented a properly native 3D CNN for tumor segmentation and subsequent binary glioma grade classification and compared it with a pre-trained 2D ResNet50 on ImageNet with previous tumor detection, employing a Mask R-CNN. The results of the 3D approach were slightly higher than the 2D ones, reporting 97.10% and 96.30% accuracy, respectively. In their study, Chatterjee et al. [272] explored the role of (2+1)D, mixed 2D-3D, and native 3D convolutions based on ResNet. This study highlights the effectiveness of mixed 2D-3D convolutions achieving an accuracy of 96.98%, surpassing both the (2+1)D and the pure 3D approaches. Furthermore, the use of pre-trained networks demonstrated enhanced performance in the spatial models, yet, intriguingly, the pure 3D model performed better when trained from scratch. Glioma diagnosis and prognosis are significantly influenced by molecular factors, prompting numerous studies to investigate the potential of DL in extracting meaningful features from 069
MRI scans for classifying these genetic frameworks [184,221,227,228,233,248,249,252, 253]. Ge et al. [221] and Banerjee et al. [227] utilized CNNs to distinguish lowand highgrade gliomas and further differentiate low-grade gliomas with and without 1p/19q codeletion. Tripathi and Bag [248] proposed a two-level classification system, first categorizing the tumor grade and then predicting 1p/19q status for low-grade tumors. Building upon this, Tripathi and Bag [249] expanded their framework to a multi-task learning approach, simultaneously addressing low and high tumor grading, as well as tumor characterization based on both IDH and 1p/19q status. This approach enables joint feature learning, facilitating the network to discern patterns such as the absence of 1p/19q codeletion in IDH-wildtype tumors. Multi-task learning has garnered attention [233,253,275]. van der Voort et al. [253] further extended the two-class grade classification to predict the WHO tumor grade (2/3/4) simultaneously with IDH mutation and 1p/19q codeletion. Their results demonstrated a 90% ROC-AUC for IDH mutation status, albeit with a sensitivity of 40% for grade 4 glioma patients and 73% for grade 2-3 glioma patients. Additionally, they achieved an 85% ROCAUC for 1p/19q codeletion prediction, although with a sensitivity of 39% specifically for grades 2-3. A recent study carried by Jin et al. [256] compared the performance of 35 neurosurgeons with and without AI assistance on glioma grading. The 3D AI-based with a VGG backbone system was trained on BraTS2020 with an accuracy of 88%. The AI-based tool showed the physicians the prediction along with a color map on the MRI to highlight the important regions for the AI decision. The average accuracy improved from 82.50% up to 87.70% with the AI assistance, and there was no additional boost with the assistance of post-hoc explanations. Instead, the feature maps help the physicians verify AI decisions but failed to give explicit reasons. This indicates that the existing AI post-hoc explanation failed to indicate for physicians when to rely on AI recommendations. With AI assistance physicians’ performance significantly increased but did not exceed AI alone. According to the results of this study, a physician has a probability of 67.20% of having higher accuracy when assisted by AI prediction and 71.00% when the prediction is accompanied by a post-hoc explanation than performing the task alone. In the experiments, doctors and AI decisions agreed with each other in 81% decisions. The decision agreement increased to 86.80% when physicians were assisted by AI prediction and further increased to 87.20% when incorporated the heatmap explanation. Doctors rated a statistically higher trust and willingness to use the AI system after viewing the AI-driven model performance metrics, however, there was no significant difference after using AI with post-hoc explanations. The study also shows that when there was a case of disagreement between the clinician and the AI decisions, physicians were more likely to check the feature attribution AI explanation. 070
KEY TAKEAWAYS •Open datasets have played a crucial role in advancing AI-driven technologies for brain tumor diagnosis. Over 85% of the studies reviewed between 2018 and 2024 utilize public data for brain tumor classification tasks, emphasizing the role of accessible data in supporting the evolution of AI technologies in healthcare and enabling their benchmarking. However, the size of these datasets typically ranges in the hundreds, which may limit the robustness and reliability of the conclusions. A larger volume of data is necessary to enhance the generalizability and accuracy of AI models in this field. •Current publicly available databases containing glioma-grade labels are labeled according to the WHO CNS criteria released in 2016. However, these criteria have since been updated in 2021, resulting in re-defined categories. Consequently, the existing publicly availablelabelsareoutdated,highlighting the need for updated data. •A complete disclosure of the data utilized, the methodologies employed, and the results obtained are not consistently observed across all papers, highlighting a lack of standardized reporting procedures. Such transparency is crucial as it enhances trust in research outcomes and significantly improves the reproducibility of the findings in research. •The vast majority of papers utilize 2D approaches rather than 3D methods for analyzing brain tumor data. The most used database, Figshare, provides 2D slices instead of full 3D volumes. However, even among studies that use other databases offering 3D volumes, there is a notable preference for 2D approaches. This may be attributed to the computational efficiency and relative simplicity of interpreting 2D data compared to 3D. •The classification performance of 2D approaches is generally higher compared to 3D methods, despite the additional spatial context provided by 3D volumes. However, the use of 2D slice data poses a risk of slice leakage during data splitting at the image level, potentially compromising the validity of the results. •Numerous studies have demonstrated the potential of MRI data pre-processing techniques, such as skull-stripping, intensity normalization, and tumor detection, in producing more accurate results. This underscores the significant role of data-centric approaches in enhancing the quality of the analysis. However, the literature often prioritizes model-centric perspectives, focusing on architectural improvements rather than on data enhancement. 071
4 THE DATA
4.1 DESCRIPTION To conduct this research, we integrated data from three publicly available datasets: BraTS [118] (n=369), EGD [123] (n=716), and a subset of the Rembrandt dataset [276] provided with tumor segmentation masks (n=59). Data from BraTS included grade tumor labels categorized in a binary manner as lower-grade gliomas (grades 2 and 3) and high-grade gliomas (grade 4 or glioblastoma), provided alongside age and survival data. The BraTS dataset contains samples from TCGA-LGG and TCGA-GBM collections from the TCIA [119,120] from which the grade label (2, 3, and 4) under the WHO CNS criteria is available. Demographic information for BraTS samples, including sex, age, and genetic labels, was collected from the data published by Ceccarelli et al. [277]. The EGD images also included clinical labels for sex, age, genetic information, and detailed scan characteristics. The Rembrandt database includes sex, age, race, and survival data. All data labeling from the three datasets was conducted under the WHO CNS criteria established in 2016 [7]. All the scans from the three datasets are distributed after some pre-processing steps, involving co-registered to the same anatomical template (RSI24 in BraTS and Rembrandt and MNI152 in EGD), interpolated to a consistent resolution of 1mm3, and additionally, BraTS images are provided after skull removal. We further applied some pre-processing steps to harmonize the MRI scans. Images were spatially aligned following the standard of the RSI24 template. Therefore, all the scans were resampled to a common resolution of [240,240,155] with an isotropic voxel size (1mm3). Skull-stripping was performed using the HD-BET algorithm [73], which is an automatic brain extraction tool based on artificial neural networks that yielded median Dice coefficients above 96% for the four MRI modalities. BFC was conducted using N4BiasFieldCorrectionImageFilter() from SimpleITK. 083
BraTS EGD Rembrandt Total LGG 76 214 40 330 HGG 293 502 19 814 Total 369 716 59 1144 G.2 28 135 24 187 G.3 37 79 13 129 G.4 293 502 19 814 Total 358 716 59 1130 Table 4.1. Grade distribution across datasets. Figure 4.1. Flow diagram of participant numbers initially and post-quality assessment of MRI modalities, alongside the implemented data split strategy. 084
The initial integrated dataset was composed of 1144 samples. A thorough data quality analysis was undertaken and images with compromised quality – whether due to flawed sequences, truncation, or inadequate depiction of the tumor region – were meticulously identified and subsequently removed from the dataset. This curation process was employed to ensure the integrity and reliability of the remaining samples for subsequent analyses and model training. The resultant dataset comprised a total of 1125 samples (see Figure 4.1). To maintain an independent evaluation standard, a distinct test set was curated from the initial 25% of the entire dataset. This separate test set was carefully set aside at the outset of the experiment and remained untouched during the model training and validation phases, thereby providing an unbiased measure of the model’s performance on entirely new data. Given the constraints of a limited dataset size, a 3-fold cross-validation (CV) methodology was adopted to enhance the robustness of the results. Each iteration of the 3-fold CV utilized a ensured 2/3 for training to facilitate a comprehensive understanding of the underlying features. The remaining 1/3 served as a validation set to assess generalization performance. Stratified sampling preserved the original distribution. The partition was performed by mantaining the same grade and database ratio in each set. Table 4.2 provides a summary of the patient demographics and class distribution for the complete sample of study as well as for the different data partition sets, including the three folds for CV and the independent test set. Age was segmented into 5-year ranges to align with the Rembrandt dataset age variable format. The statistical differences in patient demographics (sex and age) between data splits were assessed using the χ2test. No significant statistical differences were found among the different sets with regard to sex or age, indicating that the data partitioning process effectively maintained demographic balance across the experimental subsets. This ensures that any observed differences in model performance are not confounded by variations in patient characteristics. 085
Parameter Total Fold 1 Fold 2 Fold 3 Test No. of Patients 1125 (100.00) 281 (24.98) 281 (24.98) 281 (24.98) 282 (25.06) Sex 1 Female 354 (31.46) 82 (29.18) 93 (33.10) 88 (31.32) 91 (32.27) Male 552 (49.06) 140 (49.82) 139 (49.47) 137 (48.75) 136 (48.23) Unknown 219 (19.48) 59 (21.00) 49 (17.43) 56 (19.93) 55 (19.50) Age 2 [10–14] 1 (0.09) 1 (0.36) 0 (0.00) 0 (0.00) 0 (0.00) [15–19] 4 (0.36) 0 (0.00) 1 (0.36) 2 (0.71) 1 (0.35) [20–24] 19 (1.69) 5 (1.78) 4 (1.42) 4 (1.42) 6 (2.13) [25–29] 35 (3.11) 10 (3.56) 11 (3.91) 10 (3.56) 4 (1.42) [30–34] 39 (3.47) 8 (2.85) 6 (2.14) 16 (5.69) 9 (3.19) [35–39] 57 (5.07) 19 (6.76) 11 (3.91) 12 (4.27) 15 (5.32) [40–44] 56 (4.98) 17 (6.05) 12 (4.27) 15 (5.34) 12 (4.26) [45–49] 100 (8.89) 32 (11.39) 28 (9.96) 14 (4.98) 26 (9.22) [50–54] 129 (11.47) 28 (9.96) 38 (13.52) 35 (12.46) 28 (9.93) [55–59] 116 (10.31) 27 (9.61) 29 (10.32) 29 (10.32) 31 (10.99) [60–64] 145 (12.89) 32 (11.39) 39 (13.88) 37 (13.17) 37 (13.12) [65–69] 128 (11.38) 40 (14.23) 29 (10.32) 28 (9.96) 31 (10.99) [70–74] 125 (11.11) 26 (9.25) 36 (12.81) 29 (10.32) 34 (12.06) [75–79] 70 (6.22) 13 (4.63) 18 (6.41) 18 (6.41) 21 (7.45) [80–84] 23 (2.04) 8 (2.85) 1 (0.36) 8 (2.85) 6 (2.13) [85–89] 8 (0.71) 0 (0.00) 3 (1.07) 4 (1.42) 1 (0.35) Unknown 70 (6.22) 15 (5.34) 15 (5.34) 20 (7.12) 20 (7.09) Grade Label 3 LGG 320 (28.44) 80 (28.47) 80 (28.47) 80 (28.47) 80 (28.37) HGG 805 (71.56) 201 (71.53) 201 (71.53) 201 (71.53) 202 (71.63) Data is shown as count (percentage) 1P-value: .89 2P-value: .40 3P-value: 1.0 Table 4.2. Demographic characteristics over the different data partitions. 086
4.2 EXPLORATORY DATA ANALYSIS 4.2.1 MRI DATA Absolute MRI intensity values lack precise physical meaning, as voxel intensities can vary significantly based on the scanner settings and acquisition parameters. Instead, intensity values in MRI are relative, meaning their interpretation relies on comparisons within the same image across different regions within the image. Hence, MRIs are not comparable across scanners, subjects, or even studies. Within the EGD database, information regarding scanner vendors is available. The number of scans acquired with systems from each vendor is presented in Table 4.3. Figure 4.2 illustrates the pixel distributions for Siemens, Philips, and GE. Toshiba was removed from this analysis since only one image was acquired using this system manufacturer. System Manufacturer Count Siemens 306 Philips 231 GE Medical Systems 163 Toshiba 1 Table 4.3. Frequency of imaging acquisition scanner manufacturers in EGD dataset. Figure 4.2. Pixel distribution of EGD flair MRIs across imaging acquisition scans. 087
This relative nature emphasizes the importance of assessing intensity values within the context of the specific scan or study, rather than having a standardized, universally applicable scale across different MRI machines or datasets. Since these intensities lack standardized physical meaning due to variations in scanner settings and image acquisition parameters, their interpretation relies heavily on comparisons within the same image and across different regions within that image. Understanding and assessing intensity values within their particular context becomes crucial for accurate analysis and interpretation in MRI studies. To provide a comprehensive understanding of the MRI data, this section explores the pixel distribution of the maximum tumor 2D slices in each antomical view and MRI modality. Tables 4.4,4.5,4.6 present the pixel distribution statistics for the sagittal, coronal, and axial planes, respectively, encompassing both the entire image and the brain region across our complete set of images. These tables are complemented by the histograms of all the images shown in Figures 4.3,4.4, and 4.5. Background pixels were consistently set to 0.0. From Tables 4.4,4.5,4.6 we can observe that 50% of the image pixels are background pixels and, consequently, omitted from the image intensity histograms to enhance the visualization of the brain pixel distribution. As previously mentioned, MRI intensity values are relative and the maximum value is unbounded. From the statistics, we can observe a considerable distance between the 99th percentile and the maximum pixel value, which indicates the existence of outliers. For example, 99% of the sagittal Flair 2D slices have a maximum pixel value inside the brain mask lower than 2 834.59, while the remaining 1% have a maximum pixel value under 25 502.74. The results for the coronal, axial planes, and the rest of the sequences are analogous. Notably, pixel values in T1ce and T2 sequences are higher, yielding brighter areas. Furthermore, intensity values vary among MRI sequences and planes, which difficulties the direct image comparison or analysis without previous pixel value standardization. Intensity normalization aligns histograms within a comparable intensity range, which helps the network to converge in a more stable manner and has become a standard pre-processing step for the development of image-based DL models. Identifying the optimal normalization technique can significantly enhance the comparability and consistency of the data. In this chapter, we provide a comprehensive overview of the impact of different data normalization on our data distribution applied in an imageand sequence-wise fashion, meaning that each image was normalized independently. Figure 4.6 illustrates the histograms of 2D axial slices using three normalization techniques: min-max normalization, standardization using the whole image, and standardization inside the brain mask. We can observe that each normalization method presents a distinct distribution pattern. Min-max normalization scales the data to a fixed range between 0 and 1. Standardization shifts the data to have zero min and unit variance, creating a standard normal distribution. When the background pixels hold no significance for the task at hand, such as in our scenario where background pixels are consistently set to 0.0, standardizing solely on the brain mask emerges as an efficient approach, minimizing the impact of insignificant features. 088
Statistic Min Mean Std p50 p75 p90 p95 p99 Max Image Flair 0.00 244.20 634.85 0.00 306.31 752.72 1094.77 1963.24 25502.74 T1ce 0.00 311.65 759.10 0.00 364.10 764.38 1431.91 3719.42 27866.77 T1 0.00 261.29 584.49 0.00 347.72 720.52 1109.04 2429.50 25287.51 T2 0.00 291.99 706.25 0.00 387.48 821.47 1259.92 2580.09 29774.06 Brain Mask Flair 1.36×10-30 592.97 878.53 372.35 738.40 1210.33 1578.26 2834.59 25502.74 T1ce 5.29×10-16 763.93 1032.95 429.64 755.13 1749.07 2809.08 4988.68 27866.77 T1 8.21×10-16 630.49 769.13 400.35 703.92 1236.54 1866.97 3119.36 25287.51 T2 9.11×10-42 710.88 957.36 468.04 807.27 1397.13 1879.82 3719.51 29774.06 Table 4.4. Statistics computed from sagittal 2D MRIs containing the largest area of tumor. Figure 4.3. Histograms of pixel values extracted from sagittal 2D MRIs containing the largest area of tumor. 089
KEY TAKEAWAYS •MRI intensity values are inherently relative and not directly comparable across images. Variations in intensity can occur due to differences in scanning protocols or MRI systems. Consequently, applying image intensity normalization or standardization is crucial to ensure that these values are comparable across scans. •Grade 4 gliomas dominate the dataset, representing over 70% of the total cases, followed by nearly 17% of grade 2 cases, and 11% grade 3. This significant imbalance might imply potential challenges for model training, particularly in classifying minority classes. •The incidence of gliomas is higher in males compared to females; however, there are no significant sex-based differences in tumor grade distribution. Both male and female classes exhibit comparable proportions across different grades. •Gliomaincidenceandtumoraggressiveness both increase with advancing age. Glioblastomas, in particular, become more prevalent in older age groups, with a significant rise observed in individuals aged 60 to 69. •The presence or absence of IDH mutations show a significant association with tumor grade. Grade 2 and grade 3 gliomas predominantly exhibit IDH mutations, while the absence of the mutation characterizes grade 4 gliomas. •The co-deletion of chromosome arms 1p and 19q is frequently observed in IDHmutant lower-grade tumors, with the highest prevalence in grade 2 tumors. •Molecular features, such as IDH mutations and 1p/19q co-deletion, are strongly associated with lower-grade gliomas. This observed pattern aligns with the current established molecular characteristics of gliomas. 096
5 PRELUDE: A CLASSIFICATION OF LOWER AND HIGH GRADES
5.1 PREAMBLE Numerous studies in the literature aim to differentiate glioblastomas from lower-grade gliomas. Several works have addressed the benefits of pre-processing methods before feeding the brain MRI data into the neural network for glioma grade classification [36,222,258]. Others have evaluated the impact of augmenting the dataset’s size [133,144,228,236], using pre-trained models [223,224,235], adding more 2D slices [227], or combining MRI modalities [61,212,221,250,252]. Other works have centered on evaluating the effectiveness of different network configurations in the glioma grading task [172,210,226,227,237,242,254,268]. In this chapter, we focus on evaluating different strategies to enhance MRI data and improve the accuracy and robustness of distinguishing between LGG and HGG. We will assess the informativeness of each anatomical plane by comparing the model’s performance using different views as input data. Next, we will explore different data normalization procedures and examine the significance of the multiple MRI sequences, including their fusion as a strategy to combine the useful information each sequence provides. We will thoroughly evaluate the critical role of sample quantity and diversity in model generalizability through data augmentation and transfer learning. Additionally, we will examine the impact of using a single slice containing the maximum tumor area versus incorporating multiple slices, as well as comparing the use of the entire brain MRI image to focusing solely on the tumor region. Given that numerous studies have achieved high classification performance using SoA networks, we will compare the performance of ResNet and VGGNet backbones, as they were among the most popular choices. 103
5.2 DATA For each patient, the 2D slices with the largest tumor area have been selected from the sagittal, coronal, and axial planes. This is achieved by identifying the slices that contain the maximum number of pixels within the tumor segmentation mask, which is provided with the MRI images for each of the three datasets. Figure 5.1 illustrates the MRI modalities and tumor masks from the maximum tumor 2D slices of a specific patient for the three anatomical views. Figure 5.1. 2D MRI slices containing the largest tumor area from (a) sagittal, (b) coronal, and (c) axial views. To assess if additional information is provided to the model by incorporating more slices, we implemented an oversampling technique inspired by the strategy proposed by Banerjee et al. [227]. This technique also helps mitigate the challenges posed by imbalanced data. This approach entailed the selection of 10 contiguous slices both before and after the slice containing the largest tumor area, with a skip of 5 slices for high-grade tumors and a skip of 104
2 slices for lower-grade tumors. This resulted in 5 slices for high-grade gliomas and 11 slices for lower grades. An example of the slice selection approach for a LGG and an HGG patient is illustrated in Figure 5.2. The rationale behind this strategy was to balance the data distribution across different tumor grades while adding more views of the tumor. Notice that the representativity of the lower-grade gliomas is enhanced. After the training process, the ultimate patient predictions are determined through a majority voting mechanism. It is worth clarifying that the split of the images in train, validation, and test sets following a 3-fold CV approach was done on patient ID, preventing slices from the same patients from falling into train and validation/test subsets. Figure 5.2. Example of the selected 10 contiguous slices both before and after the slice containing the maximum tumor area with a skip of 2 for (a) LGG and a skip of 5 slices for (b) HGG patients. 105
Figure 5.8 illustrates the probability distributions of models (which is a direct quantitative indication of the model’s certainty about its predictions) trained on images standardized using the mean and standard deviation focusing on the brain region. The distribution for models utilizing Flair, T1, and T2 sequences are predominantly centered around 0.5, suggesting a degree of uncertainty in the predictive outcomes. Interestingly, the probability distribution derived from the model incorporating the T1ce sequence, as well as considering the fusion of all the sequences, presents a more nuanced and spread-out distribution across the entire range of 0 to 1, which is indicative of enhanced confidence in their ability to discern between classes accurately. This observation aligns logically with our numerical results, as the utilization of the T1ce sequence has consistently demonstrated superior performance when applied to independent sequences, followed by the combination of all modalities. Figure 5.9 shows the behavior of the distribution after adding data augmentation transforms to the training set as part of the pre-processing pipeline to the model using the fusion of the four sequences. The results show a notable shift in the probability distribution towards the extremities. This shift suggests increased confidence in the model’s predictions, as the augmented dataset introduces a richer variety of scenarios for the model to learn from. The broader and more pronounced distribution reflects the enhanced predictive certainty. Numerical results are presented in Table A.9. Overall, these findings underscore the significance of sequence selection and data augmentation in influencing the performance and confidence level of the model. 112
Figure 5.7. GradCAM attention maps across different pre-processing and sequences. 113
Figure 5.8. Model’s output probability distribution across independent sequences and their fusion, using mean-std standardized images within the brain region. 114
Figure 5.9. Multi-modal model’s output probability distribution incorporating data augmentation. Additional pre-processing strategies such as cropping the image to the tumor ROI, adding multiple slices from the same patient, and employing fine-tuning from pre-trained models on ImageNet have also been considered to enhance the distinction between LGG and HGG cases. Table 5.2 shows the results obtained. As per our experiments, utilizing only the tumor region versus the entire brain image enhances the true HGG positive rate in 5% while no significant differences are observed in LGG sensitivity. On the other hand, adding multiple slices or using ImageNet knowledge do not seem to contribute substantial relevant features for the model to better discern between LGG and HGG patients. These results imply that focusing on the immediate vicinity of the largest tumor region is sufficient for the model to make accurate grade classifications. Including more slices may not yield discernible improvements in the classification task. Accuracy ROC-AUC Precision Recall F1 Tumor ROI 87.70 ±0.008 92.10 ±0.002 LGG - - 75.70 83.88 73.87 HGG - - 93.30 89.27 87.83 Multi-slice 82.97 ±0.011 88.03 ±0.005 LGG - - 65.77 83.73 73.63 HGG - - 92.80 82.70 87.43 Fine-tuning 82.27 ±0.013 86.77 ±0.005 LGG - - 66.43 77.07 69.50 HGG - - 90.33 84.33 85.30 Table 5.2. Mean 3-fold CV performance on the test set for the classification of LGG vs. HGG obtained using the tumor ROI, multi-slice oversampling, and fine-tuning from ImageNet. 115
5.4.3 EVALUATING NETWORK ARCHITECTURES Parallel experiments were conducted employing various architectures, such as ResNet-34, VGGNet-11, and VGGNet-16 to evaluate and compare their performance in terms of accuracy metrics and computational efficiency. Tables A.10 to A.12 present the performance across these architectures using the previously assessed pre-processing methods. In essence, all the architectures performed better in classifying LGG versus HGG when employing mean-std standardization compared to min-max normalization. The evaluation results using the images with mean-std standardization on the brain mask, extracted from the test confusion matrices and reported as 3-fold CV averages, are displayed in Table 5.3. The accuracy and ROC-AUC values are 79.43 and 88.20, respectively, using the ResNet-18 backbone, which improves to 83.00 and 89.20 with ResNet-34. Similarly, VGGNet-11 and VGGNet-16 achieved accuracy and ROC-AUC values of 84.73 and 89.87, and 83.00 and 89.53, respectively. Although VGGNet-11 demonstrates a slightly higher accuracy, VGGNet-16 provides a more balanced classification of both lower and high-grade classes. Specifically, VGGNet-16 correctly classifies 82.50% of LGG cases, showing a 6.27% improvement over VGGNet-11, and 84% of HGG cases, a decrease of 3.97 compared to VGGNet-11. In the subsequent analyses, we will utilize VGGNet-16 as the backbone model. Regarding the computational efficiency, Table 5.4 shows the number of parameters and the training time for each network architecture. Although ResNet-34 has fewer parameters than VGGNets, it requires more training time. This suggests that the architectural complexity and the presence of residual connections in ResNet models can impact the training duration. Accuracy ROC-AUC Precision Recall F1 ResNet-18 79.43 ±0.013 88.20 ±0.005 LGG 59.90 83.73 69.10 HGG 92.37 77.73 82.70 ResNet-34 83.00 ±0.006 89.20 ±0.010 LGG 67.17 78.33 72.43 HGG 90.93 84.80 86.73 VGGNet-11 84.73 ±0.008 89.87 ±0.003 LGG 71.87 76.23 74.03 HGG 90.40 88.10 88.30 VGGNet-16 83.70 ±0.013 89.93 ±0.004 LGG 67.57 82.50 72.10 HGG 92.40 84.13 86.00 Table 5.3. Mean 3-fold CV performance on the test for the classification of LGG vs. HGG using ResNet-18, ResNet-34, VGGNet-11, and VGGNet-16. 116
Parameter count Training time ResNet-18 11 177 602 1h 18 min ResNet-34 21 285 762 1h 57 min VGGNet-11 128 780 034 1h 27 min VGGNet-16 134 277 186 1h 46 min Table 5.4. Comparison of model parameters and training time for ResNet-18, ResNet34, VGGNet-11, and VGGNet-16. 5.4.4 EVALUATING THE SAMPLE SIZE Table 5.5 shows the results on the test set across different percentages of the sample size. The ROC-AUC improved significantly from 70% with 25% of the data to 87% with the entire data set. Similarly, the sensitivity for lower-grade gliomas showed an initial value of 38% experiencing a substantial rise to 77%. As the dataset size decreases, there is an observable bias in the classification ability towards the majority class. Considering the results obtained in the previous section, even enhanced with data augmentation boosting the recall of lower-grade class 82.50%. Given the imbalanced nature of our dataset (70%–30%), it is essential to exercise caution when interpreting metrics, as the minimum accuracy of 70% can be achieved when the model is entirely biased towards the majority class and fails to classify the minority group. Accuracy ROC-AUC Precision Recall F1 25% 70.07 ±0.021 70.70 ±0.067 LGG 46.00 38.73 54.00 HGG 77.40 82.50 84.00 50% 80.13 ±0.003 83.83 ±0.004 LGG 67.33 58.70 71.57 HGG 84.47 88.60 88.43 75% 82.73 ±0.013 86.90 ±0.002 LGG 68.73 72.13 70.60 HGG 88.70 86.93 86.60 100% 83.00 ±0.012 87.87 ±0.001 LGG 67.43 77.90 72.67 HGG 90.70 84.97 86.40 Table 5.5. Mean 3-fold CV performance on the test set for the classification of LGG vs. HGG considering different sample sizes. 117
KEY TAKEAWAYS •The axial plane proved to be more effective than the sagittal and coronal planes in highlighting relevant features for distinguishing between LGG and HGG cases in a 2D context. •Image mean-std standardization outperforms min-max normalization for intensity homogenization across scans. This enhancement is further achieved by removing background pixels from the computation and concentrating solely on the brain region. This improvement is reflected not only in quantitative metrics but also in the network’s increased attention to the tumor area. Additionally, contrast enhancement does not offer substantial benefits for the classification. •T1ce emerged as the most effective individual MRI modality, followed by Flair, T1, and, lastly T2, in highlighting relevant imaging features for the classification of LGG and HGG cases. The multi-sequence approach, significantly enhances the model’s performance, particularly in distinguishing LGG cases. This improvement is evident in the quantitative results and the increased confidence in the model’s probability outputs. Additionally, the network’s attention appears to be more concentrated on the tumor area when using T1ce and FLAIR sequences. •Focusing on the 2D slice containing the largest tumor area is sufficient for accurate differentiation between lower and high-grade gliomas. The addition of additional slices does not show further classification improvements. •Using Kaiming weight initialization proved to be a more effective strategy than fine-tuning the weights from pre-trained models on ImageNet data. This strategy provided a better starting point for training, yielding improved convergence and more accurate classification. •Focusing on the tumor region in brain MRI, rather than using the entire image, enhanced the model’s ability to differentiate between lower and highgrade gliomas by emphasizing the most relevant features, thereby reducing the noise from non-tumorous areas. •Amongthevariousnetworkarchitectures evaluated for classifying LGG and HGG groups, ResNet-34 demonstrated better performance than ResNet-18, which was further improved using VGGNet11 and VGGNet-16, this achieving the best classification for LGG cases. •Increasing the dataset’s size significantly enhances the model’s classification performance, particularly in enhancing the sensitivity of lower-grade gliomas, which constitute the minority class. Smaller datasets may lead to a bias towards the majority class, highlighting the importance of large and diverse datasets for robust classification. 118
Figure A.4. Confusion matrices using the multi-planar model for the age 40−59 group. Figure A.5. Confusion matrices using the multi-planar model for the age ≥60 group. Figure A.6. Confusion matrices using the multi-planar model for the IDH wild-type group. 224
Figure A.7. Confusion matrices using the multiplanar model for the IDH mutated group. Figure A.8. Confusion matrices using the multi-planar model for the 1p/19q intact group. Figure A.9. Confusion matrices using the multi-planar model for the 1p/19q codeleted group. 225
C. MODELS’ HYPERPARAMETERS The hyperparameters involved in training the models needed to be tuned to achieve the best performance. In this section, we present the hyperparameters tuned in each classification problem and the values that lead to the best performance. 226
C.1 TESTED HYPERPARAMETERS Parameter Values Learning rate initial learning rate 1×10−3,1×10−4,5×10−5,1×10−5 scheduler patience 5,10,15 factor 0.25,0.5 Batch size 8,16,32,64 Augmentation factor 1,2,3 Dropout rate 0.10,0.20,0.25,0.5 Optimizer Adam, SGD Weight decay 1×10−3,1×10−4,5×10−5,1×10−5 Momentum 0.50,0.80,0.90 Weight initialization Xavier, Kaiming, ImageNet TL Table A.21. Hyperparameters tuned along with the tested values. 227
C.2 OPTIMAL HYPERPARAMETERS SLICE-BASED MODELS Parameter LGG vs. HGG G.2 vs. G.3 G.3 vs. G.4 G.2 vs. G.4 WHO Grade Learning rate initial value 1×10−45×10−51×10−51×10−51×10−4 patience 5 5 5 5 5 factor 0.5 0.5 0.5 0.5 0.5 Batch size 16 16 16 16 16 Augmentation factor 2 2 2 2 2 Dropout rate 0.5 0.5 0.5 0.5 0.5 Weight decay 1×10−31×10−31×10−31×10−31×10−3 Momentum 0.9 0.9 0.9 0.9 0.9 MULTI-PLANAR MODELS Parameter LGG vs. HGG G.2 vs. G.3 G.3 vs. G.4 G.2 vs. G.4 WHO Grade Learning rate initial value 1×10−45×10−55×10−55×10−51×10−4 patience 5 5 5 5 5 factor 0.5 0.5 0.5 0.5 0.5 Batch size 8 8 8 8 8 Augmentation factor 2 2 2 2 2 Dropout rate 0.1 0.1 0.1 0.1 0.1 Weight decay 1×10−31×10−31×10−31×10−31×10−3 Momentum 0.9 0.9 0.9 0.9 0.9 Linear units 4096 1024 1024 1024 4096 228
VOLUME-BASED MODELS Parameter LGG vs. HGG G.2 vs. G.3 G.3 vs. G.4 G.2 vs. G.4 WHO Grade Learning rate initial value 5×10−45×10−55×10−55×10−55×10−4 patience 5 5 5 5 5 factor 0.5 0.5 0.5 0.5 0.5 Batch size 4 4 4 4 4 Augmentation factor 1 1 1 1 1 Dropout rate 0.5 0.5 0.5 0.5 0.5 Weight decay 1×10−31×10−31×10−31×10−31×10−3 Momentum 0.9 0.9 0.9 0.9 0.9 229