scieee AI-readable full text Open interactive document viewer

Automation of ophthalmic diagnosis: glaucoma as a case study

Faria, Marco André Lopes

Abstract

With ocular diseases increasing at an alarming rate due to population aging, it is relevant to explore new technologies to meet patients' needs. As we witness the exponential evolution of the huge field that is Artificial Intelligence (AI), new methodologies and cognitive models have been developed for sustainability and optimization in several sectors of our society, namely in healthcare data analysis. By incorporating intelligent systems in clinical diagnostics, it is possible to achieve a high degree of information integration, data processing, and a theoretical improvement in the speed of diagnosis. With the presented project we had the goal of training and testing supervised models with ophthalmic data, to predict and diagnose eye diseases, such as Glaucoma. To this effect, it was carried out a Machine Learning (ML) project following a typical ML methodology. The methodology is composed of stages such as Business Understanding, Data Acquisition and Transformation, Modeling and Deployment. The document is structured according to these stages and in each of them is presented all the work that was done. The main objective of this research was to test several features and algorithms (LGBM, Random Forest, KNN, SVM and Logistic Regression) to see which set would yield the best results. For that, activities like data collection and transformation, model training and testing, and evaluation of the final outcomes were performed. Random Forest was the best ranked algorithm according to the chosen metric, F1 Score. LGBM was also considered due to its good results in Recall, one of the most important metrics in the healthcare area. This dissertation was carried out during an academic internship at Altice Labs, which provided all the necessary material, from access to technological tools and data. In turn, Altice Labs is working with the Centro Cirúrgico de Coimbra (CCC) that provided the necessary data for the study of automation of ophthalmologic diagnosis.

Full text

Universidade do Minho Escola de Engenharia Marco André Lopes Faria Automation of Ophthalmic Diagnosis: Glaucoma as a Case Study October 2022 Automation of Ophthalmic Diagnosis: Glaucoma as a Case Study Marco André Lopes Faria UMinho | 2022 Marco André Lopes Faria A80719 Automation of Ophthalmic Diagnosis: Glaucoma as a Case Study October 2022 MSc Dissertation Integrated Master’s in Engineering and Management of Information Systems Thesis performed under supervision of Professor José Luís Mota Pereira COPYRIGHT Third parties can use this academic work as long as the internationally accepted rules and good practices are respected, with regard to copyright and related rights. Thus, the present work may be used under the terms set out in the license below. If the user needs permission to be able to use the work under conditions not foreseen in the above-mentioned licensing, he/she should contact the author, through the RepositóriUM of the University of Minho. Licença concedida aos utilizadores deste trabalho Atribuição CC BY https://creativecommons.org/licenses/by/4.0/ v ACKNOWLEDGEMENTS First, I would like to thank my family for the support and motivation that was given to me throughout all this time that made it possible for me to finish my education and achieve all my goals. I would like to thank Luís Cortesão and all the Altice Labs team for the tireless support that was given to me throughout these 11 months of internship. Finally, I would like to thank Professor José Luís Mota Pereira for his availability and support whenever it was needed. vi DECLARATION OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. vii ABSTRACT Automation of Ophthalmic Diagnosis: Glaucoma as a Case Study With ocular diseases increasing at an alarming rate due to population aging, it is relevant to explore new technologies to meet patients' needs. As we witness the exponential evolution of the huge field that is Artificial Intelligence (AI), new methodologies and cognitive models have been developed for sustainability and optimization in several sectors of our society, namely in healthcare data analysis. By incorporating intelligent systems in clinical diagnostics, it is possible to achieve a high degree of information integration, data processing, and a theoretical improvement in the speed of diagnosis. With the presented project we had the goal of training and testing supervised models with ophthalmic data, to predict and diagnose eye diseases, such as Glaucoma. To this effect, it was carried out a Machine Learning (ML) project following a typical ML methodology. The methodology is composed of stages such as Business Understanding, Data Acquisition and Transformation, Modeling and Deployment. The document is structured according to these stages and in each of them is presented all the work that was done. The main objective of this research was to test several features and algorithms (LGBM, Random Forest, KNN, SVM and Logistic Regression) to see which set would yield the best results. For that, activities like data collection and transformation, model training and testing, and evaluation of the final outcomes were performed. Random Forest was the best ranked algorithm according to the chosen metric, F1 Score. LGBM was also considered due to its good results in Recall, one of the most important metrics in the healthcare area. This dissertation was carried out during an academic internship at Altice Labs, which provided all the necessary material, from access to technological tools and data. In turn, Altice Labs is working with the Centro Cirúrgico de Coimbra (CCC) that provided the necessary data for the study of automation of ophthalmologic diagnosis. Keywords: Artificial Intelligence, Glaucoma, Machine Learning, Ophthalmology viii RESUMO Automação de Diagnósticos Oftalmológicos: Glaucoma como Caso de Estudo Com o número de doenças oftalmológicas a aumentar a um ritmo alarmante devido ao envelhecimento da população, é importante explorar novas tecnologias com o objetivo de satisfazer as necessidades dos pacientes. Com a evolução exponencial da Inteligência Artificial, novas metodologias e modelos cognitivos têm sido desenvolvidos de forma a otimizar vários setores na nossa sociedade, entre eles a análise de dados no ramo da saúde. Ao incorporar sistemas inteligentes nos diagnósticos clínicos, é possível atingir um alto nível de integração da informação, processamento de dados e uma melhoria teórica na rapidez do diagnóstico. O objetivo do projeto apresentado é treinar e testar modelos supervisionados com dados oftalmológicos, de modo a prever e diagnosticar doenças oculares como o Glaucoma. Para esse efeito, foi realizado um projeto de Machine Learning , seguindo uma metodologia típica. A metodologia é composta por fases como Business Understanding, Data Acquisition and Transformation, Modeling e Deployment. O documento está estruturado de acordo com essas fases e em cada uma delas será apresentado todo o trabalho realizado. O principal objetivo desta investigação foi testar vários conjuntos de características e algoritmos (LGBM, Random Forest, KNN, SVM e Logistic Regression) para ver quais produziriam os melhores resultados. Para isso, foram realizadas atividades como a recolha e transformação de dados, treino e teste de modelos e avaliação dos resultados. Random Forest foi o algoritmo mais bem classificado de acordo com a métrica escolhida, F1 Score. LGBM foi também considerado devido aos seus bons resultados em Recall, uma das métricas mais importantes na área da saúde. Esta dissertação foi realizada no contexto de um estágio curricular na Altice Labs, que forneceu todo o material necessário, desde acesso às ferramentas tecnológicas e dados. Por sua vez, a Altice Labs está a trabalhar com o Centro Cirúrgico de Coimbra que forneceu os dados necessários para o estudo da automação de diagnósticos oftalmológicos. Palavras-chave: Glaucoma, Inteligência Artificial, Machine Learning, Oftalmologia ix INDEX Copyright............................................................................................................................................ iv Acknowledgements .............................................................................................................................. v Declaration of integrity ........................................................................................................................ vi Abstract............................................................................................................................................. vii Resumo............................................................................................................................................ viii List of abbreviations/Acronyms ......................................................................................................... xiii List of figures .................................................................................................................................... xiv List of tables ..................................................................................................................................... xvi 1. Introduction ................................................................................................................................ 1 2. Literature Review ........................................................................................................................ 4 2.1 Medical Background ............................................................................................................ 4 2.2 Machine Learning ................................................................................................................ 9 2.3 Machine Learning Application in Ophthalmology ................................................................. 23 3. Materials and Methods .............................................................................................................. 27 3.1 Methodology ...................................................................................................................... 27 3.2 Development Environment ................................................................................................. 29 3.3 Procedure .......................................................................................................................... 32 4. Implementation ......................................................................................................................... 36 4.1 Business Understanding .................................................................................................... 36 4.2 Data Acquisition and Transformation .................................................................................. 37 4.2.1 Data Ingestion ............................................................................................................ 37 4.2.2 Data Description ......................................................................................................... 39 4.2.3 Data Exploration ......................................................................................................... 39 4.2.4 Data Preparation ........................................................................................................ 44 xvi LIST OF TABLES Table 1 - Deliveries of medical examinations from CCC ..................................................................... 38 Table 2 – Fields with most nulls on the dataset ................................................................................. 40 Table 3 - Features of the resulting dataset and their origin ................................................................. 46 Table 4 - Features used in Iteration 1 ................................................................................................ 53 Table 5 - Features used in Iteration 2 ................................................................................................ 56 Table 6 - Features used in Iteration 3 ................................................................................................ 57 Table 7 - Features used in Iteration 4 ................................................................................................ 58 Table 8 - Features used in Iteration 5 ................................................................................................ 59 Table 9 - Features used in Iteration 6 ................................................................................................ 60 Table 10 - Results of the algorithm LGBM in Iteration 1, using the Scikit-learn library ......................... 66 Table 11 - Confusion Matrix of the algorithm LGBM in Iteration 1, using the Scikit-learn library ........... 66 Table 12 - Results of the algorithm LR in Iteration 1, using the Scikit-learn library ............................... 66 Table 13 - Confusion Matrix of the algorithm LR in Iteration 1, using the Scikit-learn library ................ 66 Table 14 - Results of the algorithm KNN in Iteration 1, using the Scikit-learn library ........................... 66 Table 15 - Confusion Matrix of the algorithm KNN in Iteration 1, using the Scikit-learn library ............. 66 Table 16 - Results of the algorithm SVM in Iteration 1, using the Scikit-learn library ............................ 67 Table 17 - Confusion Matrix of the algorithm SVM in Iteration 1, using the Scikit-learn library ............. 67 Table 18 - Results of the algorithm RF in Iteration 1, using the Scikit-learn library............................... 67 Table 19 - Confusion Matrix of the algorithm RF in Iteration 1, using the Scikit-learn library ................ 67 Table 20 - Results of the algorithm LGBM in Iteration 2, using the Scikit-learn library ......................... 67 Table 21 - Confusion Matrix of the algorithm LGBM in Iteration 2, using the Scikit-learn library ........... 67 Table 22 - Results of the algorithm LR in Iteration 2, using the Scikit-learn library ............................... 68 Table 23 - Confusion Matrix of the algorithm LR in Iteration 2, using the Scikit-learn library ................ 68 Table 24 - Results of the algorithm KNN in Iteration 2, using the Scikit-learn library ........................... 68 Table 25 - Confusion Matrix of the algorithm KNN in Iteration 2, using the Scikit-learn library ............. 68 Table 26 - Results of the algorithm SVM in Iteration 2, using the Scikit-learn library ............................ 68 Table 27 - Confusion Matrix of the algorithm SVM in Iteration 2, using the Scikit-learn library ............. 68 Table 28 - Results of the algorithm RF in Iteration 2, using the Scikit-learn library............................... 69 Table 29 - Confusion Matrix of the algorithm RF in Iteration 2, using the Scikit-learn library ................ 69 xvii Table 30 - Results of the algorithm LGBM in Iteration 3, using the Scikit-learn library ......................... 69 Table 31 - Confusion Matrix of the algorithm LGBM in Iteration 3, using the Scikit-learn library ........... 69 Table 32 - Results of the algorithm SVM in Iteration 3, using the Scikit-learn library ............................ 69 Table 33 - Confusion Matrix of the algorithm SVM in Iteration 3, using the Scikit-learn library ............. 69 Table 34 - Results of the algorithm RF in Iteration 3, using the Scikit-learn library............................... 70 Table 35 - Confusion Matrix of the algorithm RF in Iteration 3, using the Scikit-learn library ................ 70 Table 36 - Results of the algorithm LGBM in Iteration 4, using the Scikit-learn library ......................... 70 Table 37 - Confusion Matrix of the algorithm LGBM in Iteration 4, using the Scikit-learn library ........... 70 Table 38 - Results of the algorithm SVM in Iteration 4, using the Scikit-learn library ............................ 70 Table 39 - Confusion Matrix of the algorithm SVM in Iteration 4, using the Scikit-learn library ............. 70 Table 40 - Results of the algorithm RF in Iteration 4, using the Scikit-learn library............................... 71 Table 41 - Confusion Matrix of the algorithm RF in Iteration 4, using the Scikit-learn library ................ 71 Table 42 - Results of the algorithm LGBM in Iteration 5, using the Scikit-learn library ......................... 71 Table 43 - Confusion Matrix of the algorithm LGBM in Iteration 5, using the Scikit-learn library ........... 71 Table 44 - Results of the algorithm SVM in Iteration 5, using the Scikit-learn library ............................ 71 Table 45 - Confusion Matrix of the algorithm SVM in Iteration 5, using the Scikit-learn library ............. 71 Table 46 - Results of the algorithm RF in Iteration 5, using the Scikit-learn library............................... 72 Table 47 - Confusion Matrix of the algorithm RF in Iteration 5, using the Scikit-learn library ................ 72 Table 48 - Results of the algorithm LGBM in Iteration 1, using the PyCaret library .............................. 73 Table 49 - Results of the algorithm LR in Iteration 1, using the PyCaret library ................................... 73 Table 50 - Results of the algorithm KNN in Iteration 1, using the PyCaret library ................................ 73 Table 51 - Results of the algorithm SVM in Iteration 1, using the PyCaret library ................................ 73 Table 52 - Results of the algorithm RF in Iteration 1, using the PyCaret library ................................... 73 Table 53 - Results of the algorithm LGBM in Iteration 2, using the PyCaret library .............................. 74 Table 54 - Results of the algorithm LR in Iteration 2, using the PyCaret library ................................... 74 Table 55 - Results of the algorithm KNN in Iteration 2, using the PyCaret library ................................ 74 Table 56 - Results of the algorithm SVM in Iteration 2, using the PyCaret library ................................ 74 Table 57 - Results of the algorithm RF in Iteration 2, using the PyCaret library ................................... 74 Table 58 - Results of the algorithm LGBM in Iteration 3, using the PyCaret library .............................. 75 Table 59 - Results of the algorithm SVM in Iteration 3, using the PyCaret library ................................ 75 Table 60 - Results of the algorithm RF in Iteration 3, using the PyCaret library ................................... 75 Table 61 - Results of the algorithm LGBM in Iteration 4, using the PyCaret library .............................. 75 xviii Table 62 - Results of the algorithm SVM in Iteration 4, using the PyCaret library ................................ 75 Table 63 - Results of the algorithm RF in Iteration 4, using the PyCaret library ................................... 76 Table 64 - Results of the algorithm LGBM in Iteration 5, using the PyCaret library .............................. 76 Table 65 - Results of the algorithm SVM in Iteration 5, using the PyCaret library ................................ 76 Table 66 - Results of the algorithm RF in Iteration 5, using the PyCaret library ................................... 76 Table 67 - Results of the algorithm LGBM in Iteration 6, using the PyCaret library .............................. 76 Table 68 - Results of the algorithm SVM in Iteration 6, using the PyCaret library ................................ 77 Table 69 - Results of the algorithm RF in Iteration 6, using the PyCaret library ................................... 77 Table 70 - Results of model tuning in Iteration 1 ................................................................................ 78 Table 71 - Results of model tuning in Iteration 2 ................................................................................ 78 Table 72 - Results of model tuning in Iteration 3 ................................................................................ 78 Table 73 - Results of model tuning in Iteration 4 ................................................................................ 78 Table 74 - Results of model tuning in Iteration 5 ................................................................................ 78 Table 75 - Results of model tuning in Iteration 6 ................................................................................ 79 1 1. INTRODUCTION Artificial Intelligence is taking over the world in nearly every field and has been dramatically changing our lifestyle over the past years. The health sector, as one of the most important fields, cannot be left behind. Many reasons can explain the need for technological advances in ophthalmology. With population aging becoming a demographic trend worldwide, ocular diseases are increasing at an alarming rate. Existing conventional diagnosis methods are heavily dependent on professional’s experience, which can result in higher misdiagnosis. And with the amount of image data generated rapidly rising, it is more important than ever to have techniques that can analyze and process this data (Lu et al., 2018). On the other hand, the exponential growth of computing power and a better quality of ophthalmic imaging allows for better results. With the usage of ML techniques in medical diagnoses, the detection, surveillance, and treatment of ocular diseases can be improved. These techniques will also aid ophthalmologists in the detection and classification of pathological features of eye diseases. Recent research indicates that Al/ML has the potential to outperform humans in image recognition tasks (Lu et al., 2018). Incorporating Al/ML systems into clinical practice can boost productivity, aid in decision-making, and allow for image analysis automation (Valikodath, 2020). This work emerged in the context of an existing exploratory project at Altice Labs, in collaboration with a clinical entity. Altice Labs is a company part of the Altice Group, located in Aveiro, that develops innovative products and services for the telecommunications and information technology market. In addition to having several offices in Portugal, Altice Labs has a subsidiary company operating in Brazil. For being part of a multinational corporation and due to its credibility based upon 65 years of technological experience in telecommunications, Altice Labs was a suitable option for providing me the necessary support in writing my dissertation. Altice Labs is involved in several projects, and for this project in specific, it is working with Centro Cirúrgico de Coimbra (CCC) that provided the necessary data for the study of the automation of ophthalmologic diagnosis. The data available for the project consists of medical examinations that were turned into useful data through techniques of text extraction, prior to my inclusion in the team. The existent project had the goal of applying ML models to identify which patients have characteristics that reveal glaucoma. With this, it is expected that, at a later stage, we will be able to predict that a patient will be diagnosed with Glaucoma. For that the Altice Labs team conducted an ML project where different features and different algorithms were tested to evaluate which would produce better results. Initially, there were 3 different analyses to try to distinguish the 2 patients: Sick/Not sick, Glaucoma/Other pathologies, Glaucoma/Not glaucoma. The team decided to proceed with Glaucoma/Not glaucoma as it was the one with better results. My work followed on from the work of Altice Labs, I only focused on the analysis of Glaucoma/Not glaucoma. In my work, initially, it was performed research on the state of the art of ocular imaging analysis and interpretation technologies, as well as of AI/ML processes applied to the diagnosis of ocular diseases. This first phase of the project was mainly of literature review to prepare for the project’s development. In a second instance, a specific scenario, among all the eye diseases mentioned in the first phase, was chosen. As we saw above it was Glaucoma and all the work done was focused on this disease. Being an ML project, it followed an appropriate methodology composed of steps such as Business Understanding, Data Acquisition and Transformation, Modeling and Deployment. All the steps of this methodology will be thoroughly analyzed during this document. The main goal of my work was to test different sets of algorithms and features to determine which would produce better results and present these results to Altice Labs. For that it was performed tasks such as Data Collection, Data Preparation and Model Training, where the model was trained and tested using the data collected from the medical examinations. With this research we hope that the diagnosis and treatment of ocular diseases can continue to improve with the help of Machine Learning techniques. This document follows an IMRaD (Introduction, Methods, Results, and Discussion) structure, one of the most prominent formats for the structure of a scientific article. Due to the complexity of a dissertation project, other chapters were added in addition to the ones mentioned above. The document is divided into 9 sections that will be described below. The first section corresponds to the Introduction where it is presented the background and motivations for the project, its objectives, and a presentation of the company. In the second section, it was performed a Literature Review on the two main topics of the dissertation - Ophthalmology and Machine Learning. The Literature Review explores and synthetizes the key concepts of these two topics. In the third section is presented the methodology followed, the development environment and the procedure used to carry out this project. In the fourth section is described the implementation of the methodology presented in the third section. Thus, this section is sub-divided in the steps of the methodology. In the fifth section, are presented the results obtained in the fourth section. The results are presented according to the different evaluation metrics chosen during the Business Understanding stage. In the sixth section, the results obtained are discussed and analyzed. In the seventh section, are presented the conclusions obtained from this work. The eight section contains the references used for the development of the theoretical part of this document. 3 Lastly, the ninth section corresponds to the Appendix, where is presented artifacts that were too long to be included in the main sections, such as tables or prints from the development environment. 4 2. LITERATURE REVIEW In this section, it was performed a Literature Review on the two main topics of the dissertation - Ophthalmology and Machine Learning. The Literature Review explores and synthetizes the key concepts of these two topics. Being a work of medical nature, it is important to explore the medical background first. After that, the topic of Machine Learning will be discussed. 2.1 Medical Background In the next paragraphs will be provided an overview about the pathologies and medical examinations addressed in this study. Starting with pathologies, it will be focused on Glaucoma, Diabetic Retinopathy and Cataracts. Glaucoma: Glaucoma is the most common cause of irreversible blindness in the world (Akkara, 2021). It is characterized by dysfunction of Retinal Ganglion Cells (RGCs), causing gradual damage to the Optic Nerve Head (ONH) and the Retinal Nerve Fiber Layer (RNFL) thickness, and a resultant loss of the Visual Field (VF). This eye condition, referred as the "silent thief of sight", is normally asymptomatic in early stages of the disease which makes it harder to be diagnosed. However, with an early diagnosis and optimal treatment it can effectively be slowed down, minimizing the risk of visual loss, hence the importance of an early diagnosis. There is no known cure for Glaucoma, but it can be controlled through eye drops, pills and, in more severe situations, with laser procedures and surgical operations. This will help to prevent or slow further damage from occurring. Figure 1 - Difference between a Normal Eye (left) and a Glaucomatous Eye (right) utilizing fundus image, retrieved from (Nayak et al., 2009) 5 Diabetic Retinopathy: Diabetic Retinopathy (DR), a complication caused by high blood sugar levels damaging the retina, is the most common and insidious complication of Diabetes Mellitus (Abramoff, 2010). Diabetes Mellitus develops when the pancreas does not produce enough insulin or the body cannot effectively process it, affecting the circulatory system and therefore the retina. When the glucose levels in the blood increase continuously, the blood vessels get severely damaged resulting in blood leak in the eyes which reduces the vision. Normally, DR takes several years to progress to the point where it threatens the vision however, if not diagnosed and treated early it can cause blindness which is incurable at advanced stages. To the present day, a cure to this pathology has not been found. Nevertheless, there are advanced treatments focused on slowing or stopping the progression of the disease that work very well. The ability to treat retinopathy symptoms relies only on early detection and early intervention. The earlier the condition is found, the easier it is to cure and better the chances that the vision will be saved. The most common treatments for retinopathy are managing the blood sugar levels, intravitreal injections and laser surgery. Figure 2 - Detection of DR progression utilizing fundus image, retrieved from (Yun, W. et al., 2008) Cataracts: Cataracts is a dulling or clouding of the lens inside the eye, which is normally clear, causing symptoms such as blurry vision. For the human eye to see, light must pass through a clear lens, located behind the iris. The lens focuses the light so that the brain and the eyes can work together and process the information. If a person suffers from Cataracts, it clouds over the lens preventing the eye from focusing the light and resulting in vision loss. Vision changes depend on the Cataracts location and size. There are several factors that can cause Cataracts such as living in an area with bad air pollution or alcohol and cigarette addiction, however most Cataracts are related to age and systemic 6 diseases. As the world's population grows older, Cataracts are projected to become even more common. Early detection and treatment are essential for increasing patient’s quality of life and lowering healthcare costs. There is no natural cure for Cataracts. Usually, the recommended medical treatments are regular eye examinations to monitor the progression of the disease, lifestyle changes, and ultimately, surgery to remove the diseased lens. Figure 3 - Fundoscopic image of the lesions before cataract extraction in the right eye (A, C) and left eye (B, D), retrieved from (Van Noort, B. et al., 2019) Regarding medical examinations, the ones that will be addressed on this work are Visual Field Tests, Optical Biometry, Fundus Imaging and RNFL Analysis. Visual Field Test: Visual Field Test is an ophthalmological examination performed with a device called Campimeter. It allows to evaluate central and peripheral visual perception, identifying any change or reduction of the visual field, and quantify these same changes by comparing them in future evaluations. This test also helps in the differential diagnosis, to verify the effectiveness of some treatments and to evaluate follow-up visits, verifying the stability or the progression of the lesions. It is indicated in glaucomatous pathologies, neuro-ophthalmology, and retinal degenerations. The sequence of images presented below demonstrates the glaucoma progression through a VF test. 7 Figure 4 - Glaucoma Progression on a VF Test, retrieved from (https://www.glaucoma.org.il/glaucoma/diagnosis/vf-damage) Optical Biometry: Optical Biometry is an ophthalmological examination performed by a technician using a device called Biometer. It is used to take measurements of the eyeball to provide data that is used for the preoperative study of various eye surgeries. It can be performed on people of all age groups, without any contraindication and it only requires the instillation of one drop of anesthetic eye drops. The patient then fixes his eyes on a point determined by the technician so that measurements can be taken with an ultrasonic probe. The probe contacts the cornea, which is anesthetized, making it a painless procedure. The most frequent situation in which biometry is used is in calculating the intraocular lens that will be implanted during Cataract surgery. Figure 5 - Example of an Optical Biometry, retrieved from (https://www.doctor-hill.com/lenstar_haag_streit/lenstar_manual_6c.htm) 14 inconsistencies that may exist. In data integration and normalization, original data collected from diverse sources should be integrated and adapted to a common scale appropriate for thorough comparison evaluation. In feature selection, to optimize the learning process performance, the most important features are chosen. Due to their heterogeneous origin, most data is inconsistent and noisy, therefore it is crucial that data is preprocessed before being fed to the model. The quality of data directly affects the ability of the model to learn. Training, Validation, and Test: To build a reliable ML model and to achieve the best performance possible, the base dataset must be randomly partitioned into two subsets, one to build the model and the other to test the model’s performance. The first dataset is further split into training dataset and validation dataset. Figure 13 - Partitions of the original Dataset into Training, Validation and Test subsets The training dataset is used to fit the parameters of the model and make it learn the hidden features in the data. This set should be as much diversified as possible so that the model is trained to predict any given scenario. The validation dataset is used for parameter selection and to tune the model’s configurations according to those parameters. After the model had been trained in the previous step, the model’s performance is validated to define how well the model was trained. The goal of this step is to prevent the model from overfitting. Several validation techniques can help to estimate unbiased generalized model performances and give a better understanding of how the model was trained. It makes sure that the model is accurately trained and that it outputs the right data. There are several validation methods such as, Train and Test Split, and K-Fold Cross-Validation. In Train and Test Split, the data is split into training data and testing data, holding back the testing data to not expose the model to it, until it is time to test the model. The percentage of data intended to the training data and testing data can be decided at the time of splitting. K-Fold Cross-Validation is similar to the Train and Test Split, except that it splits the data into more than two groups. In this validation method, “K” is used as a placeholder for the number of groups the data will be split into. For example, in the 15 figure below, the data was split into 3 groups. One group is left out of the training data and the model is validated using the group that was left out of the training data. Then, the model is cross validated. Each of the two other groups used as training data are then also used to test the model. Each test and score can give new information about what is working and what is not in your machine learning model. Figure 14 - Example of K-Fold Cross-Validation Finally, the test dataset is a separate set of data from training and validation, that is used to evaluate the performance of the trained model. Since it is an independent set, it provides an unbiased final model performance. Evaluation: After building the best learning model possible, it is required to perform an evaluation of how well the model is performing and if it is good enough to move forward. When the model performs classification predictions, there are four possible outcomes: • True Positives (TP): the total number of correct predictions when the actual class was positive. • True Negatives (TN): the total number of correct predictions when the actual class was negative. • False Positives (FP): the total number of wrong predictions when the actual class was negative. • False Negatives (FN): the total number of wrong predictions when the actual class was positive. For better organization and interpretability of the possible outcomes, a Confusion Matrix can be created. Confusion Matrix is a table that allows visualization of the performance of an algorithm. Each row in the matrix represents an actual class, while each column represents a predicted class. It compares the actual target values with those predicted by the Machine Learning model. The number of correct and incorrect predictions are summarized and divided by each class. 16 Figure 15 - Confusion Matrix From this outcome, it is possible to evaluate the model using several performance metrics: Accuracy: the proportion of positives and negatives that are correctly identified, it represents the ability of the model to correctly identify patients who have the disease and those who do not. The higher the accuracy, the better the classifier. Sensitivity/Recall: the proportion of positives that are correctly identified, also known as the True Positive Rate. It represents the ability of the model to correctly identify patients who have the disease. Specificity: the proportion of negatives that are correctly identified, also known as the True Negative Rate. It represents the ability of the model to correctly identify patients who do not have the disease. Precision: the number of true positives divided by the total number of positive predictions. It represents the proportion of positives that are correctly identified among all positive identified samples. The higher the precision, the lower the false positive rates. 17 F1 score: the harmonic average of the precision and sensitivity, it represents how precise the classifier is, as well as how robust it is. The range for F1 Score is [0, 1] where it reaches its best value at 1 and worst at 0. The higher the F1 Score, the better the performance of the model. Likelihood ratio for positive test results: the probability that a positive test would be expected in a patient divided by the probability that a positive test would be expected in a patient without a disease, it represents the probability of a positive test result occurring compared to a negative one. Likelihood ratio for negative test results: the probability of a patient testing negative who has a disease divided by the probability of a patient testing negative who does not have a disease, it represents the probability of a negative test result occurring compared to a positive one. Receiver Operating Characteristic (ROC) curve: ROC curve is a graph that indicates the performance of a classification model at all classification thresholds. The graph plots two parameters, the True Positive Rate (TPR) and the False Positive Rate (FPR). It shows the trade-off between sensitivity (TPR) and specificity (1 – FPR). The closer the ROC curve is to the upper left corner, the higher the AUC value and the better the performance of the model. Area under the curve (AUC): AUC measures the entire area underneath the ROC curve. It can have any value between 0 and 1 and it is a good indicator of the model’s quality. The AUC indicator measures the accuracy of positive and negative samples at the same time. It tells how much the model is capable of distinguishing between classes, which means the higher the AUC, the better the model is at distinguishing between patients with the disease and no disease. 18 Figure 16 - Example of the ROC curve and AUC, retrieved from (https://towardsdatascience.com/understanding-auc-roc-curve68b2303cc9c5) Machine Learning Algorithms Decision Tree: Decision tree is a type of supervised learning algorithm that is mostly used for classification problems. It can be subclassified as Classification Trees, when the variables are categorical, and Regression Trees, when the variables are continuous. It is a tree-like model in which each node represents a test on an attribute and each branch represents the outcome of the test. The root node represents the entire population and is further split into two or more homogeneous sets, based on the most important attributes. Decision trees do not require any statistical knowledge and are very intuitive to interpret, making it one of the most used algorithms. On the other hand, it is an algorithm with a high tendency to overfit and it is not fit for continuous variables. Random Forest: Random Forest is a supervised machine learning algorithm that is constructed from many individual decision trees. It utilizes ensemble learning, a technique that combines multiple learning algorithms to achieve better predictive performance than could have obtained from any of the individual learning algorithms alone. Each individual tree in the random forest makes its own prediction and the final prediction is done by taking the average of the output from the various trees. The benefit of this algorithm is that although some of the trees may be wrong, many other trees will be right, so together the trees have more chances of being accurate. The precision of the outcome improves as the number of trees grows. Random Forest algorithms solve some of the limitations of decision trees. It reduces the overfitting of datasets and increases precision. 19 Figure 17 - Single Decision Tree vs Random Forest, retrieved from (https://towardsdatascience.com/from-a-single-decision-tree-to-arandom-forest-b9523be65147) Support Vector Machine: Support Vector Machine is a supervised learning algorithm that can be used for regression and classification but is mostly used in classification problems. In this algorithm, each data item is plotted as a point in an N-dimensional space, where N is the number of features you have, with the value of each feature being the value of a particular coordinate. Then, the objective is to find a hyperplane that best differentiates the classes. Support Vectors are the coordinates of each individual observation, and the best hyperplane is the one whose distance to the nearest element of each class is the largest, as demonstrated in the figure below. SVMs algorithms can achieve good values of accuracy with less computation power. Figure 18 - The Optimal Hyperplane in a SVM Algorithm, retrieved from (https://towardsdatascience.com/support-vector-machineintroduction-to-machine-learning-algorithms-934a444fca47) 20 K-Nearest Neighbor: K-Nearest Neighbor is a supervised learning technique that can be used for both classification and regression but is mostly used in classification tasks. KNN works by finding the Euclidean distance between a new data point and all its K neighbors. Given a new data point, the first step of this algorithm is to define the number K of the neighbors we intend to use. Then we calculate the distance between the new data point and its K number of neighbors. We count the number of data points in each category among these K neighbors and designate the new data point to the category with the highest number of neighbors. To identify the best K number, the algorithm is executed several times with different values and the K that reduces the greatest number of errors while maintaining the ability to accurately make predictions is selected. This algorithm is simple to understand and easy to implement. It is also robust when the training data is noisy and large. In return, the computation cost is high due to calculating the distance between the data points for all the training samples and the need to always determine the value of K might sometimes be complex. To more easily understand this algorithm an example is shown below: As illustrated in the image below, assume we have a new data point, and we want to put it in the required category. Figure 19 - New Data Point in a KNN Algorithm, retrieved from (https://www.javatpoint.com/k-nearest-neighbor-algorithm-for-machinelearning) As said before, the first step is to choose the number of K neighbors. For this example, we will choose the k=5. Next, it is calculated the Euclidean distance between the data points, using the formula: 21 Figure 20 - Euclidean Distance between two points in a KNN Algorithm, retrieved from (https://www.javatpoint.com/k-nearest-neighboralgorithm-for-machine-learning) After calculating the Euclidean distance, we now know that we have three nearest neighbors in category A and two nearest neighbors in category B, as illustrated in the image below. Figure 21 - Results of a KNN Algorithm, retrieved from (https://www.javatpoint.com/k-nearest-neighbor-algorithm-for-machine-learning) Therefore, it is safe to assume that the new data point must belong to category A. Linear Regression: Linear Regression is a simple supervised learning algorithm used for regression tasks. It predicts values within a continuous range rather than to classifying them into categories. It establishes a relationship between independent and dependent variables by fitting a best line, also 22 known as regression line. If there is only one independent variable it is called simple linear regression, and if there is more than one independent variable it is called multiple linear regression. Linear regression is a linear model that describes the relationship between the independent variables (input) and the dependent variables (output). It is an algorithm very easy to interpret, to implement and to train but, on the other hand, it is quite sensitive to outliers. Logistic Regression: Despite having “regression” in the name, Logistic Regression is a classification algorithm. It is used to predict the probability of a target variable, thus having its output values between 0 and 1. Generally, the nature of the target is binary, which means the dependent variable will have only two possible types. But Logistic Regression can also be Multinomial, where the dependent variable can have three or more possible unordered types, and Ordinal, where the dependent variable can have three or more possible ordered types. Although it is easy to implement and to interpret, the major limitation of this algorithm is the assumption of linearity between the dependent and the independent variables. Figure 22 - Linear Regression vs Logistic Regression, retrieved from (https://medium.com/@mvanshika25/logistic-regressionee47cc89345f) Light Gradient Boosting Machine: Light Gradient Boosting Machine (LightGBM or LGBM) is a gradient boosting framework based on decision trees and is used for ranking, classification, and many other Machine Learning tasks. Despite being based on tree-based algorithms, LGBM grows trees vertically while other algorithms grow trees horizontally, meaning that LGBM grows tree leaf-wise while other algorithms grow level-wise. Leaf-wise algorithms are capable of reduce more loss than level-wise algorithms. These two approaches can be visualized in the figure below. 23 Due to the ongoing increase of data size, it is becoming harder for traditional ML algorithms to give faster results. LightGBM has advantage over these algorithms due to its high speed, can handle larger sizes of data and takes lower memory to run. Despite all these advantages, it is not advisable to be used on small datasets, since LGBM is sensitive to overfitting and can easily overfit small data. 2.3 Machine Learning Application in Ophthalmology After addressing these two topics individually, it is now necessary to discuss them as a whole. As discussed in the medical background, most ocular diseases can be prevented if diagnosed on time. However, conventional diagnostic methods are failing to meet patients' needs, hence the need for other alternatives. The exponential rise of technologies that we have witnessed in recent years has led to a growth in computational power, resulting in advances in Machine Learning. The goal of this dissertation is to apply these technological advances to ophthalmology so that we can improve the diagnosis and treatment of ocular diseases. Incorporating AI systems into clinical practice can bring many advantages such as enhancing workplace productivity, as well as aid in the clinical decision-making process (Topol, 2019). When comparing the performance between AI systems and human ophthalmologists doing a medical evaluation, AI systems have the potential to achieve much better information integration, data processing, and diagnostic speed. However, these intelligent systems should not be seen as replacements of the human diagnostic, but rather as a tool to help to improve them. There is a general fear that the evolution of AI Figure 23 - How LGBM trees work, retrieved from (https://medium.com/@pushkarmandot/https-medium-compushkarmandot-what-is-lightgbm-how-to-implement-it-how-to-fine-tune-the-parameters-60347819b7fc) 30 Microsoft Teams is a collaboration app built for hybrid work that provides a platform for setting up group conversations, video calls and file sharing. It is the communication tool used in Altice Labs. Jupyter Lab Jupyter Lab is a web-based interactive development environment for notebooks, code, and data. Its flexible interface and modular design allows users to configure and arrange workflows in Data Science and Machine Learning, with extensions to expand functionality. All the programming in this dissertation, since Data Acquisition and Transformation to Modeling, was done through Jupyter Notebooks in Jupyter Lab. Figure 25 - Interface of a notebook on Jupyter Lab MLflow MLflow is a platform to streamline Machine Learning development, including tracking experiments, packaging code into reproducible runs, and sharing and deploying. As it can be seen in the image below, MLflow shows some information like start time, duration and run name, as well as the metrics needed to evaluate the model. It was used to get a more organized view of the results by aggregating them all in one place. 31 Figure 26 - Interface of MLflow MinIO MinIO is a High-Performance Object Storage that can handle unstructured data such as photos, videos, log files, backups, and containers. In this project it was used to store the final dataset after its creation in the Jupyter Notebooks. Figure 27 - Interface of MinIO Python Python is a high-level, interpreted, general-purpose programming language with a design philosophy that emphasizes code readability. One of the reasons for its popularity in Machine Learning projects is the large collection of libraries that users can work with, such as Scikit-Learn, PyCaret and Numpy. All the programming in this project was done using this language. Scikit-learn 32 Scikit-learn is a Machine Learning library for the Python programming language. It features various classification, regression and clustering algorithms and is designed to interoperate with the Python numerical and scientific libraries NumPy and SciPy. It can also be used in preprocessing, model selection and dimensionality reduction tasks. PyCaret PyCaret is an open-source, low-code and end-to-end library in Python that automates Machine Learning workflows. It provides a simple way to do ML tasks such as Exploratory Data Analysis, Data Preprocessing and Model Training. It was introduced with the purpose of validating the results obtained through the Scikit-learn library. 3.3 Procedure The procedure followed to carry out this project will be explained in this section. To begin with, the steps of the methodology that were addressed above theoretically, will be addressed in a practical way in the Implementation section. In there, these steps will be thoroughly analyzed, and all the work done in each one of them will be presented. Business Understanding was a critical stage that provided the necessary information to make projectrelated decisions. Following a review of the literature, it was decided to use Conventional Machine Learning algorithms rather than Deep Learning. The learning of the algorithms can be classified as Supervised Learning and further subclassified as a Classification task. The selection of algorithms (SVM, KNN, Random Forest, LGBM, and Logistic Regression) was also supported by the literature and the Altice Labs team. For each one of the algorithms, it was performed 6 different iterations to evaluate which algorithm and set of features would produce better results. The different iterations are explained in the Model Training section. For every iteration the label used was glaucoma_flag as the goal of this work is to determine whether a patient has glaucoma. All the programming of this project was done in the Python language using Jupyter Notebooks. Python is a trending language in Machine Learning projects due to its large collection of libraries. In that way, the Scikit-learn library was used to perform the project. On a later stage of the project, the library 33 PyCaret was introduced to validate the work done. All the different iterations were performed using the two libraries to have a more reliable result. Bellow it will be presented the different Jupyter Notebooks used during the project, more specifically during the stages of Data Acquisition and Transformation and Modeling. There are a total of 31 notebooks and a brief explanation of each one will be provided. Notebooks from 0.0 to 2.4 correspond to the Data Acquisition and Transformation stage. In these notebooks, tasks such as Dataset Creation and Exploratory Data Analysis were performed. • Notebook 0.0: the first notebook was used to start exploring the tool and its features. It was not produced anything relevant to the project. • Notebook 1.0: this is the notebook where the dataset was created, since reading the tables from the database, apply the transformations needed and exporting the dataset to MinIO. • Notebook 2.1: in this notebook was performed an EDA where we could learn more about the number of nulls in the dataset, its shape, and the distribution of data. • Notebook 2.2: in this notebook was also performed an EDA but with more complex techniques such as PCA and t-SNE. Some of these techniques were just tested but not implemented in the final work. • Notebook 2.3: this notebook was similar to notebook 2.2 but were tested Feature Reduction techniques. • Notebook 2.4: this notebook was similar to notebook 2.3 but it was tested the label eye_sick_label instead of the glaucoma_flag. From the notebook 3.1 until the last one, the notebooks correspond to the Modeling stage. To be better organized, notebooks with the prefix 3 correspond to the first iteration using scikit-learn, notebooks with the prefix 4 correspond to the second iteration using scikit-learn, notebooks with the prefix 5 correspond to the third iteration using scikit-learn, notebooks with the prefix 6 correspond to the fourth iteration scikit-learn, notebooks with the prefix 7 correspond to the fifth iteration scikit-learn, and finally notebooks with the prefix 8 correspond to all the six iterations using the PyCaret library. 34 • Notebook 3.1: This notebook concerns the first iteration where the algorithm LGBM was applied. This iteration was performed using the scikit-learn library. • Notebook 3.2: This notebook concerns the first iteration where the algorithm Logistic Regression was applied. This iteration was performed using the scikit-learn library. • Notebook 3.3: This notebook concerns the first iteration where the algorithm KNN was applied. This iteration was performed using the scikit-learn library. • Notebook 3.4: This notebook concerns the first iteration where the algorithm SVM was applied. This iteration was performed using the scikit-learn library. • Notebook 3.5: This notebook concerns the first iteration where the algorithm Random Forest was applied. This iteration was performed using the scikit-learn library. • Notebook 4.1: This notebook concerns the second iteration where the algorithm LGBM was applied. This iteration was performed using the scikit-learn library. • Notebook 4.2: This notebook concerns the second iteration where the algorithm Logistic Regression was applied. This iteration was performed using the scikit-learn library. • Notebook 4.3: This notebook concerns the second iteration where the algorithm KNN was applied. This iteration was performed using the scikit-learn library. • Notebook 4.4: This notebook concerns the second iteration where the algorithm SVM was applied. This iteration was performed using the scikit-learn library. • Notebook 4.5: This notebook concerns the second iteration where the algorithm Random Forest was applied. This iteration was performed using the scikit-learn library. • Notebook 5.1: This notebook concerns the third iteration where the algorithm LGBM was applied. This iteration was performed using the scikit-learn library. • Notebook 5.2: This notebook concerns the third iteration where the algorithm SVM was applied. This iteration was performed using the scikit-learn library. • Notebook 5.3: This notebook concerns the third iteration where the algorithm Random Forest was applied. This iteration was performed using the scikit-learn library. • Notebook 6.1: This notebook concerns the fourth iteration where the algorithm LGBM was applied. This iteration was performed using the scikit-learn library. 35 • Notebook 6.2: This notebook concerns the fourth iteration where the algorithm SVM was applied. This iteration was performed using the scikit-learn library. • Notebook 6.3: This notebook concerns the fourth iteration where the algorithm Random Forest was applied. This iteration was performed using the scikit-learn library. • Notebook 7.1: This notebook concerns the fifth iteration where the algorithm LGBM was applied. This iteration was performed using the scikit-learn library. • Notebook 7.2: This notebook concerns the fifth iteration where the algorithm SVM was applied. This iteration was performed using the scikit-learn library. • Notebook 7.3: This notebook concerns the fifth iteration where the algorithm Random Forest was applied. This iteration was performed using the scikit-learn library. • Notebook 8.1: This notebook concerns the first iteration using the PyCaret library. • Notebook 8.2: This notebook concerns the second iteration using the PyCaret library. • Notebook 8.3: This notebook concerns the third iteration using the PyCaret library. • Notebook 8.4: This notebook concerns the fourth iteration using the PyCaret library. • Notebook 8.5: This notebook concerns the fifth iteration using the PyCaret library. • Notebook 8.6: This notebook concerns the sixth iteration using the PyCaret library. 36 4. IMPLEMENTATION This section will put into practice the methodology discussed above. All the stages covered theoretically in the Methodology section, such as Business Understanding, Data Acquisition and Transformation, Modeling and Deployment, will now be addressed practically. Each stage of the methodology will have its own sub-sections where all the work done will be explained. 4.1 Business Understanding A correct solution to the problem involves a correct understanding of the problem. Business Understanding stage consists of a precise specification of the problem together with methods of evaluating the achievement of the goal. Before starting a ML project, all team members should have the domain knowledge to formulate the problem correctly. For example, this project in specific is about ophthalmic diseases, a subject that was completely unknown for me until then. Therefore, it is of utmost importance a stage where we try to understand the problem before starting to implement the solution itself. To have a better understanding of the subjects discussed, it was performed a Literature Review, which can be found above in its own section. This stage was extremally important to provide the necessary information to make decisions about important issues of the project. Bellow will be presented some of the most important conclusions of the Literature Review that were used to support the decisions made throughout the project. • “However, they have a significant disadvantage, especially in the medical context. DL algorithms are called black boxes because the networks generate comprehensive features that are too complex to be interpreted by the human mind. These algorithms can achieve better performances but the way they analyze patterns and make decisions is unexplainable, which is not feasible in the healthcare field (Lu et al., 2020).” - This quote was the reason it was decided to use CML algorithms rather than DL algorithms. Despite being able to get better results, in the healthcare field it is given more relevance to interpretability of the results. • “According to most research studies, supervised methods are more used because they can achieve better results (Lu et al., 2018).” - This quote was the reason it was decided to use Supervised Learning rather than Unsupervised Learning. • “Classification is widely used in medical diagnosis, target marketing and credit approval.” - This quote was the reason it was decided to use Classification rather than Regression or Clustering. 37 • “Random Forest and Support Vector Machine are the most used CML technologies in the ophthalmology field (Lu et al., 2018)” - Algorithm selection was done based on the literature and by suggestion of the Altice Labs team. 4.2 Data Acquisition and Transformation Data Acquisition and Transformation has the objective of managing the complete data lifecycle, including ingesting, exploring, preparing, and transforming the data. Therefore, this step is composed of another sub-steps such as Data Ingestion, Data Description, Data Exploration and Data Preparation. Each one of these sub-steps has its own section bellow where they will be better described, and the work done in each one of them will be presented. The creation of a properly clean, transformed, and consolidated dataset is the main goal of this step and its intended output. To this end, it is necessary to go through this entire data lifecycle where any possible inconsistency can be resolved. After the dataset creation, a process of feature selection to improve the model’s performance is also conducted. Large datasets are increasingly common and are often difficult to interpret. To interpret such datasets and prevent them from overfitting, methods are required to reduce their dimensionality by transforming a large set of variables into a smaller one that still preserves most of the information. In the same way, training and testing a very large dataset can be a challenging task. As the performance of a ML algorithm depends heavily on the features contained in the dataset, we should only retain features that actually help our model to learn something. Unnecessary and redundant features slow down the training time of the algorithm and should be removed. Reducing the number of variables of a dataset works as a trade-off between accuracy and interpretability. The goal in dimensionality reduction is to trade a little accuracy for simplicity, because smaller data sets are easier to explore and visualize and make analyzing data much easier and faster for ML algorithms. 4.2.1 Data Ingestion Data sources used in this project are images from real medical examinations, made available by the Centro Cirúrgico de Coimbra. There are 4 different medical examinations: Ocular Tension, Visual Field Test (also referred as PEC due to its Portuguese translation), RNFL Analysis and Optical Biometry. 38 From the beginning of the project to the current date, CCC has delivered medical examinations to Altice Labs in 7 different moments, having in total data from 218 patients, 428 eyes and 479 diagnostics by eye. The data was then stored in a Postgres SQL Database, the database used internally by the Altice Labs team. Table 1 - Deliveries of medical examinations from CCC Delivery Date Description Delivery 0 December 2020 Sample of 6 patients + January 2021 - sample of 141 patients (including 6 previous patients) Delivery 1 July 2021 Sample of 15 patients Delivery 2 August 2021 Sample of 19 patients (one already existing in delivery 1 with equal exams) Delivery 3 October 2021 Sample of 51 patients (three already existing with new exams) Delivery 4 November 2021 Sample of 24 patients Delivery 5 December 2021 Sample of 76 patients Delivery 6 February 2022 Following a request for clarification, missing diagnoses from deliveries 4 and 5 were provided As said before, CCC (Centro Cirúrgico de Coimbra) delivered the data in the form of medical examinations. This raw format has low business value to a Machine Learning project. To transform it into useful data it was needed to extract the text from the images of the exams. I did not participate in 39 this task, as it was already in progress when I joined the Altice Labs team, but a brief explanation will still be provided. The Altice Labs team used mainly two software tools for the text extraction. Opencv for Image Preprocessing and Tesseract for Optical Character Recognition (OCR). These two steps are sequential, Image Preprocessing must be done before Optical Character Recognition to improve the quality of this last step. Some of the tasks of Image Preprocessing include rescaling, binarization, noise removal and thresholding. 4.2.2 Data Description In this section, it is demonstrated an ER Diagram that was designed by the Altice Labs team to illustrate how the entities relate to each other. For further details, in the Appendix it can be found a set of tables containing an individually description of all the entities used in the diagram. The tables are composed of Field, Type, Description, Example and Possible Values. Figure 28 - ER Diagram 4.2.3 Data Exploration 46 RNFL Data The following modifications are applied to the RNFL data: • Uppercases all the string values so that they are uniform The transformations mentioned above were the ones made to each table individually. To the dataset as a whole, were also applied transformations such as dropping nulls and standardization of values. 4.2.5 Resulting Dataset After all the steps performed during the Data Acquisition and Transformation stage, the resulting dataset consists of the following features: Table 3 - Features of the resulting dataset and their origin Features Origin Description patient_id Relation field Patient identifier age Patient Data Age of the patient at the last exam gender Patient Data Gender of the patient label Diagnostics Data Name associated with the pathology other_eye_glau_flag Diagnostics Data Flag (1 or 0) indicating that the other eye has glaucoma glaucoma_flag Diagnostics Data Flag (1 or 0) indicating that the eye has glaucoma eye_sick_label Diagnostics Data Flag (1 or 0) indicating that the eye is affected by a sickness eye Diagnostics Data Eye (left or right) vfi_last PEC Data Last VFI (Visual Field Index) determined for the eye vfi_min PEC Data Lowest VFI determined for the eye vfi_max PEC Data Largest VFI determined for the eye 47 vfi_mean PEC Data Average of the various VFIs determined for the eye vfi_std PEC Data Standard deviation of the various VFIs determined for the eye md_last PEC Data Last MD (Mean deviation) determined for the eye md_min PEC Data Lowest MD determined for the eye md_max PEC Data Largest MD determined for the eye md_mean PEC Data Average of the various MDs determined for the eye md_std PEC Data Standard deviation of the various MDs determined for the eye psd_last PEC Data Last PSD (Pattern standard deviation) determined for the eye psd_min PEC Data Lowest PSD determined for the eye psd_max PEC Data Largest PSD determined for the eye psd_mean PEC Data Average of the various PSDs determined for the eye psd_std PEC Data Standard deviation of the various PSDs determined 48 for the eye last_fl PEC Data Last FL (Fixation Losses) determined for the eye last_ght PEC Data Last GHT (Glaucoma Hemifield Test) determined for the eye pec_qtd PEC Data Quantity of PEC exams in an eye al Biometry Data Axial length (mm) cct Biometry Data Cornea thickness (μm) ad Biometry Data Aqueous depth (mm) acd Biometry Data Anterior chamber depth (mm) lt Biometry Data Lens thickness (mm) k1_value Biometry Data Flat meridian value (D) k1_angle Biometry Data Flat meridian angle (º) k2_value Biometry Data Steep meridian value (D) k2_angle Biometry Data Steep meridian angle (º) ast_value Biometry Data Astigmatism value (D) ast_angle Biometry Data Astigmatism angle (º) wtw Biometry Data White to White (mm) ic_x Biometry Data Iris barycenter x-axis (mm) ic_y Biometry Data Iris barycenter y-axis (mm) pc_x Biometry Data Pupil barycenter x-axis (mm) pc_y Biometry Data Pupil barycenter y-axis (mm) pd Pupil Diameter Data Pupil diameter (mm) ot Optical Tension Data Optical Tension measurement mrw_result RNFL Data Minimum Rim Width 49 mrw_g_value RNFL Data Absolute value for MRW Global mrw_g_perc RNFL Data Percentile ratio of the MRW Global mrw_ts_value RNFL Data Absolute value for MRW Temporal Superior sector mrw_ns_value RNFL Data Absolute value for MRW Nasal Superior sector mrw_t_value RNFL Data Absolute value for MRW Temporal sector mrw_n_value RNFL Data Absolute value for MRW Nasal sector mrw_ti_value RNFL Data Absolute value for MRW Temporal Inferior sector mrw_ni_value RNFL Data Absolute value for MRW Nasal Inferior sector mrw_ts_perc RNFL Data Percentile ratio of the MRW Temporal Superior sector mrw_ns_perc RNFL Data Percentile ratio of the MRW Nasal Superior sector mrw_t_perc RNFL Data Percentile ratio of the MRW Temporal sector mrw_n_perc RNFL Data Percentile ratio of the MRW Nasal sector mrw_ti_perc RNFL Data Percentile ratio of the MRW Temporal Inferior sector mrw_ni_perc RNFL Data Percentile ratio of the MRW Nasal Inferior sector rnflt_result RNFL Data Retinal Nerve Fiber Layer rnflt_g_value RNFL Data Absolute value for RNFLT Global 50 rnflt_g_perc RNFL Data Percentile ratio of the RNFLT Global rnflt_ts_value RNFL Data Absolute value for RNFLT Temporal Superior sector rnflt_ns_value RNFL Data Absolute value for RNFLT Nasal Superior sector rnflt_t_value RNFL Data Absolute value for RNFLT Temporal sector rnflt_n_value RNFL Data Absolute value for RNFLT Nasal sector rnflt_ti_value RNFL Data Absolute value for RNFLT Temporal Inferior sector rnflt_ni_value RNFL Data Absolute value for RNFLT Nasal Inferior sector rnflt_ts_perc RNFL Data Percentile ratio of the RNFLT Temporal Superior sector rnflt_ns_perc RNFL Data Percentile ratio of the RNFLT Nasal Superior sector rnflt_t_perc RNFL Data Percentile ratio of the RNFLT Temporal sector rnflt_n_perc RNFL Data Percentile ratio of the RNFLT Nasal sector rnflt_ti_perc RNFL Data Percentile ratio of the RNFLT Temporal Inferior sector rnflt_ni_perc RNFL Data Percentile ratio of the RNFLT Nasal Inferior sector mrw_value_min RNFL Data Minimum Absolute Value of MRW 51 mrw_value_max RNFL Data Maximum Absolute Value of MRW mrw_perc_min RNFL Data Minimum Percentile Value of MRW mrw_perc_max RNFL Data Maximum Percentile Value of MRW rnflt_value_min RNFL Data Minimum Absolute Value of RNFLT rnflt_value_max RNFL Data Maximum Absolute Value of RNFLT rnflt_perc_min RNFL Data Minimum Percentile Value of RNFLT rnflt_perc_max RNFL Data Maximum Percentile Value of RNFLT rnflt_outsides RNFL Data Quantity of sectors of the eye that had a qualitative evaluation of outside normal limits rnflt_borderlines RNFL Data Quantity of eye sectors that had a qualitative borderline assessment 4.3 Modeling In this section it will be covered the training of the Machine Learning model and all the steps that led to it, from algorithm selection to validation. To train a ML algorithm able to make predictions based on the provided data it is necessary to apply mathematical, computer science, and business knowledge. It is a crucial stage that will determine the quality and accuracy of future predictions in new situations. The process of modeling involves training a ML algorithm to predict the labels from the features, tuning it for the business needs, and validating it. Essentially, the objective of this step is to select, apply, and calibrate several modeling techniques to achieve optimal values. 52 This section is divided in Algorithm Selection, Model Training and Validation. 4.3.1 Algorithm Selection In this step the most appropriated algorithms for training are selected. There are several trade-off points in selecting the best algorithm like model complexity, interpretability, performance, and computer requirements. After a consideration of the different trade-offs and by suggestion of the Altice Labs team, the selected algorithms were KNN, LGBM, Random Forest, Logistic Regression and SVM. The first iterations of Model Training included all the selected algorithms to identify which produces better results. Following iterations only included the best suited models to simplify the work done. The choice for the best suited models were based on the Literature Review and on results obtained in the previous iterations. According to the Literature Review, Random Forest and Support Vector Machine are the most used CML technologies in the ophthalmology field (Lu et al., 2018). Thus, these two algorithms, and LGBM which was obtaining good results, were the algorithms selected for the rest of the iterations. 4.3.2 Model Training Model Training is the phase of the project lifecycle in which training data is provided to a Machine Learning algorithm to learn from. The goal of model training is to create the most accurate mathematical model of the correlation between data features and a target label. The model's performance dictates how well the applications created with it will perform. After completing all the preparatory steps, such as data ingestion, dataset creation, and data transformation, and after selecting the desired algorithms, it is time to apply them and train the models. This step was divided in 6 different iterations that will now be explained. To all these 6 iterations, the label used was glaucoma_flag as the objective of the work is to determine if a patient has glaucoma or not. In this section, it will only be explained theoretically the different iterations. The procedure followed to perform these iterations, since splitting the data into test and train sets, creating the model and evaluate metrics, is described in the Appendix using prints from the Jupyter Notebooks. Iteration 1 53 In the first iteration of Model Training, all the selected algorithms were applied to a dataset containing all the existent features, except for the labels and the patient ID because they do not have business value and even affect the model’s performance. This first iteration, the one with the greatest number of features, will serve as a baseline for the next iterations where different combinations of features will be tested. In the table below it can be found the list of features used in this iteration. Table 4 - Features used in Iteration 1 Feature Origin age Patient Data gender Patient Data eye Diagnostics Data vfi_last PEC Data vfi_min PEC Data vfi_max PEC Data vfi_mean PEC Data vfi_std PEC Data md_last PEC Data md_min PEC Data md_max PEC Data md_mean PEC Data md_std PEC Data psd_last PEC Data psd_min PEC Data psd_max PEC Data psd_mean PEC Data psd_std PEC Data last_fl PEC Data last_ght PEC Data pec_qtd PEC Data al Biometry Data cct Biometry Data ad Biometry Data 54 acd Biometry Data lt Biometry Data k1_value Biometry Data k1_angle Biometry Data k2_value Biometry Data k2_angle Biometry Data ast_value Biometry Data ast_angle Biometry Data wtw Biometry Data ic_x Biometry Data ic_y Biometry Data pc_x Biometry Data pc_y Biometry Data pd Pupil Diameter Data ot Optical Tension Data mrw_result RNFL Data mrw_g_value RNFL Data mrw_g_perc RNFL Data mrw_ts_value RNFL Data mrw_ns_value RNFL Data mrw_t_value RNFL Data mrw_n_value RNFL Data mrw_ti_value RNFL Data mrw_ni_value RNFL Data mrw_ts_perc RNFL Data mrw_ns_perc RNFL Data mrw_t_perc RNFL Data mrw_n_perc RNFL Data mrw_ti_perc RNFL Data mrw_ni_perc RNFL Data rnflt_result RNFL Data 55 rnflt_g_value RNFL Data rnflt_g_perc RNFL Data rnflt_ts_value RNFL Data rnflt_ns_value RNFL Data rnflt_t_value RNFL Data rnflt_n_value RNFL Data rnflt_ti_value RNFL Data rnflt_ni_value RNFL Data rnflt_ts_perc RNFL Data rnflt_ns_perc RNFL Data rnflt_t_perc RNFL Data rnflt_n_perc RNFL Data rnflt_ti_perc RNFL Data rnflt_ni_perc RNFL Data mrw_value_min RNFL Data mrw_value_max RNFL Data mrw_perc_min RNFL Data mrw_perc_max RNFL Data rnflt_value_min RNFL Data rnflt_value_max RNFL Data rnflt_perc_min RNFL Data rnflt_perc_max RNFL Data rnflt_outsides RNFL Data rnflt_borderlines RNFL Data Iteration 2 In the second iteration of Model Training, features that were previously identified, in the Data Exploration phase, as features with a big percentage of nulls were removed. All the algorithms selected were applied in this iteration. In the table below it can be found the list of features used in this iteration. 62 • Ignore Features: It is used to ignore features during model training. For example, one of the features ignored in all iterations was patient ID since it does not represent value to the model training. • Train Size: Proportion of the dataset to be used for training and validation. In this case the proportion used was 0.75 meaning that 75% of the dataset will be used for training and validation. The other 25% of the dataset will be used for testing. • Normalize: It transforms the numeric features by scaling them to a given range. • Fold: Number of folds to be used in cross validation. In this case it was used 3 folds. • Data Split Stratify: Controls stratification during Train Test Split. Stratification ensures that training and test sets have the same proportion of the feature of interest as in the original dataset. • Fix Imbalance: It is used to balance the unequal distribution of target class when training a dataset. Compare Models This function trains and evaluates the performance of the chosen models using cross-validation. The function had 2 parameters. The parameter “sort” had the value of “F1”, meaning that F1 Score would be the metric used to rank the models results, and the parameter “include” where the value was the models we intended to use. For the iteration 1 and 2 all the selected models were used (LGBM, Random Forest, KNN, Logistic Regression and SVM) and for the remaining iterations the models LGBM, Random Forest and SVM were used. The output of this function is a scoring grid with average crossvalidated scores. Tune Model This function tunes the hyperparameters of a given model with the goal of optimizing it. The function was used with 3 Folds and 100 iterations as parameters. Normally, increasing the number of iterations improves the model performance but also increases the training time. The function also had the parameter “optimize” with the value of “F1” since this was the metric we wanted to be optimized Tuning is the process of maximizing a model's performance without overfitting or creating too high of a variance. In Machine Learning, this is accomplished by selecting appropriate hyperparameters. Tuning is usually a trial-and-error process by which you change some hyperparameters, run the algorithm on the data again, then compare its performance on your validation set to determine which set of hyperparameters results in the most accurate model. 63 Evaluate Model This function displays a basic user interface for analyzing the performance of a trained model. It shows several types of plots including AUC, Confusion Matrix, and Feature Importance. 4.3.4 Model Assessment Model Assessment is the process of using different evaluation metrics to understand a ML model’s performance. While evaluating models, selecting the right metric is a critical step of the task. Different applications of the model demand the analysis of different metrics. For some applications it can be more important to analyze the True Positive Rate, while for others is the True Negative Rate. For other applications, it may even be necessary to use a set of metrics to get the whole picture of the problem. From the existent metrics, the ones that were considered in this work were: Accuracy, Precision, Recall, AUC, F1 Score and Confusion Matrix. The metric used to rank the models was F1 Score since it is one of the most complete metrics. Recall was also considered due to its importance in the medical field. This metric is regarded as being among the most important for medical studies, since it is desired to miss as few positive instances as possible (Hicks et al., 2022). Both Recall and Specificity are important metrics in this field of study as they complement each other and are both needed to fully understand a model’s strengths and weaknesses. The ideal situation would be to have both high Recall and high Specificity, but it is important to recognize that these two metrics exist in a state of balance. Increased Recall normally comes at the expense of reduced Specificity. In a situation where a trade-off is needed, we should choose the metric that is more suitable for the problem at hand. As in the medical field the goal is to minimize false-negative results, we prioritize Recall since it measures how often a test correctly generates a positive result for people who have the condition that is being tested for. A test that is highly sensitive (high Recall) will flag almost everyone who has the disease and not generate many false-negative results. 64 4.4 Deployment Deployment is the process of integrating a Machine Learning model in a production environment where it can be used for its intended purpose. Models can be deployed in a wide range of environments, and it involves extensive planning, documentation, and a variety of different tools to complete this process. It also requires coordination between IT teams, software developers, and business professionals to ensure the model works reliably, since there is often a discrepancy between the programming language in which a model is written and the languages of the production system. Being this an academic work, this phase will not be addressed practically, as the model will not be integrated in the Altice Labs production environment. 65 5. RESULTS In this section will be presented the results obtained across the different iterations. These results were evaluated through different metrics that were already described. For each iteration it will be presented both ways that were used to perform it. The first one was using the Scikit Learn library through Jupyter Notebooks and the second one was using the PyCaret library, also through Jupyter Notebooks. 5.1 Scikit-learn To get a more organized view, the results were grouped on MLflow. For each algorithm of each iteration, 10 runs were performed to obtain a more reliable result. The image below shows, as an example, the 10 runs performed for the algorithm LGBM in the Iteration 1 (prints for the other algorithms and iterations can be found in the Appendix). The results of the metrics presented in the tables below are the average of those 10 runs. Figure 36 – Example of the 10 runs performed for the algorithm LGBM in the Iteration 1 66 Iteration 1 • LGBM Table 10 - Results of the algorithm LGBM in Iteration 1, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.845 0.749 0.773 0.827 0.776 Table 11 - Confusion Matrix of the algorithm LGBM in Iteration 1, using the Scikit-learn library 46 (TN) 6 (FP) 5 (FN) 14 (TP) • Logistic Regression Table 12 - Results of the algorithm LR in Iteration 1, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.815 0.689 0.646 0.765 0.824 Table 13 - Confusion Matrix of the algorithm LR in Iteration 1, using the Scikit-learn library 45 (TN) 6 (FP) 8 (FN) 12 (TP) • KNN Table 14 - Results of the algorithm KNN in Iteration 1, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.817 0.750 0.598 0.756 0.804 Table 15 - Confusion Matrix of the algorithm KNN in Iteration 1, using the Scikit-learn library 46 (TN) 4 (FP) 9 (FN) 12 (TP) 67 • SVM Table 16 - Results of the algorithm SVM in Iteration 1, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.849 0.892 0.502 0.737 0.816 Table 17 - Confusion Matrix of the algorithm SVM in Iteration 1, using the Scikit-learn library 51 (TN) 1 (FP) 9 (FN) 10 (TP) • Random Forest Table 18 - Results of the algorithm RF in Iteration 1, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.869 0.826 0.660 0.803 0.892 Table 19 - Confusion Matrix of the algorithm RF in Iteration 1, using the Scikit-learn library 49 (TN) 3 (FP) 6 (FN) 13 (TP) Iteration 2 • LGBM Table 20 - Results of the algorithm LGBM in Iteration 2, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.820 0.701 0.796 0.813 0.744 Table 21 - Confusion Matrix of the algorithm LGBM in Iteration 2, using the Scikit-learn library 42 (TN) 11 (FP) 3 (FN) 15 (TP) 68 • Logistic Regression Table 22 - Results of the algorithm LR in Iteration 2, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.839 0.779 0.699 0.802 0.838 Table 23 - Confusion Matrix of the algorithm LR in Iteration 2, using the Scikit-learn library 50 (TN) 3 (FP) 8 (FN) 10 (TP) • KNN Table 24 - Results of the algorithm KNN in Iteration 2, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.823 0.777 0.639 0.776 0.816 Table 25 - Confusion Matrix of the algorithm KNN in Iteration 2, using the Scikit-learn library 47 (TN) 5 (FP) 8 (FN) 11 (TP) • SVM Table 26 - Results of the algorithm SVM in Iteration 2, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.809 0.783 0.585 0.752 0.840 Table 27 - Confusion Matrix of the algorithm SVM in Iteration 2, using the Scikit-learn library 51 (TN) 2 (FP) 9 (FN) 9 (TP) 69 • Random Forest Table 28 - Results of the algorithm RF in Iteration 2, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.873 0.854 0.740 0.839 0.864 Table 29 - Confusion Matrix of the algorithm RF in Iteration 2, using the Scikit-learn library 47 (TN) 2 (FP) 6 (FN) 16 (TP) Iteration 3 • LGBM Table 30 - Results of the algorithm LGBM in Iteration 3, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.835 0.765 0.772 0.820 0.736 Table 31 - Confusion Matrix of the algorithm LGBM in Iteration 3, using the Scikit-learn library 39 (TN) 10 (FP) 4 (FN) 18 (TP) • SVM Table 32 - Results of the algorithm SVM in Iteration 3, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.795 0.884 0.473 0.719 0.786 Table 33 - Confusion Matrix of the algorithm SVM in Iteration 3, using the Scikit-learn library 49 (TN) 2 (FP) 12 (FN) 8 (TP) 70 • Random Forest Table 34 - Results of the algorithm RF in Iteration 3, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.853 0.851 0.667 0.805 0.886 Table 35 - Confusion Matrix of the algorithm RF in Iteration 3, using the Scikit-learn library 48 (TN) 2 (FP) 6 (FN) 15 (TP) Iteration 4 • LGBM Table 36 - Results of the algorithm LGBM in Iteration 4, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.810 0.701 0.740 0.792 0.767 Table 37 - Confusion Matrix of the algorithm LGBM in Iteration 4, using the Scikit-learn library 43 (TN) 9 (FP) 4 (FN) 15 (TP) • SVM Table 38 - Results of the algorithm SVM in Iteration 4, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.812 0.899 0.530 0.749 0.828 Table 39 - Confusion Matrix of the algorithm SVM in Iteration 4, using the Scikit-learn library 52 (TN) 1 (FP) 9 (FN) 9 (TP) 71 • Random Forest Table 40 - Results of the algorithm RF in Iteration 4, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.860 0.826 0.723 0.825 0.852 Table 41 - Confusion Matrix of the algorithm RF in Iteration 4, using the Scikit-learn library 49 (TN) 3 (FP) 5 (FN) 14 (TP) Iteration 5 • LGBM Table 42 - Results of the algorithm LGBM in Iteration 5, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.801 0.682 0.794 0.801 0.726 Table 43 - Confusion Matrix of the algorithm LGBM in Iteration 5, using the Scikit-learn library 41 (TN) 10 (FP) 5 (FN) 15 (TP) • SVM Table 44 - Results of the algorithm SVM in Iteration 5, using the Scikit-learn library Accuracy Precision Recall AUC F1 Score 0.827 0.834 0.599 0.760 0.828 Table 45 - Confusion Matrix of the algorithm SVM in Iteration 5, using the Scikit-learn library 50 (TN) 3 (FP) 78 Iteration 1 Table 70 - Results of model tuning in Iteration 1 Accuracy Precision Recall AUC F1 Score 0.857 0.792 0.780 0.902 0.793 Iteration 2 Table 71 - Results of model tuning in Iteration 2 Accuracy Precision Recall AUC F1 Score 0.858 0.795 0.777 0.905 0.784 Iteration 3 Table 72 - Results of model tuning in Iteration 3 Accuracy Precision Recall AUC F1 Score 0.848 0.774 0.791 0.888 0.777 Iteration 4 Table 73 - Results of model tuning in Iteration 4 Accuracy Precision Recall AUC F1 Score 0.849 0.778 0.782 0.906 0.777 Iteration 5 Table 74 - Results of model tuning in Iteration 5 Accuracy Precision Recall AUC F1 Score 0.839 0.761 0.761 0.898 0.760 79 Iteration 6 Table 75 - Results of model tuning in Iteration 6 Accuracy Precision Recall AUC F1 Score 0.855 0.780 0.801 0.907 0.788 5.2.3 Evaluate Models After tuning the model’s hyperparameters and having the best possible model, some graphs were plotted using the function evaluate_model(). The graphs plotted were AUC, Confusion Matrix, and Feature Importance, for each one of the iterations. Iteration 1 • AUC Figure 39 - AUC Plot for Iteration 1 80 • Confusion Matrix Figure 40 - Confusion Matrix for Iteration 1 • Feature Importance Figure 41 - Feature Importance Plot for Iteration 1 81 Iteration 2 • AUC Figure 42 - AUC Plot for Iteration 2 • Confusion Matrix Figure 43 - Confusion Matrix for Iteration 2 82 • Feature Importance Figure 44 - Feature Importance Plot for Iteration 2 Iteration 3 • AUC Figure 45 - AUC Plot for Iteration 3 83 • Confusion Matrix Figure 46 - Confusion Matrix for Iteration 3 • Feature Importance Figure 47 - Feature Importance Plot for Iteration 3 84 Iteration 4 • AUC Figure 48 - AUC Plot for Iteration 4 • Confusion Matrix Figure 49 - Confusion Matrix for Iteration 4 85 • Feature Importance Figure 50 - Feature Importance Plot for Iteration 4 Iteration 5 • AUC Figure 51 - AUC Plot for Iteration 5 86 • Confusion Matrix Figure 52 - Confusion Matrix for Iteration 5 • Feature Importance Figure 53 - Feature Importance Plot for Iteration 5 87 Iteration 6 • AUC Figure 54 - AUC Plot for Iteration 6 • Confusion Matrix Figure 55 - Confusion Matrix for Iteration 6 94 Christopher, M., Belghith, A., Weinreb, R. (2018). Retinal Nerve Fiber Layer Features Identified by Unsupervised Machine Learning on Optical Coherence Tomography Scans Predict Glaucoma Progression. Invest Ophthalmol Vis Sci, 59(7), pp. 2748-2756. https://doi.org/10.1167/iovs.1723387 Confirming a Questionable Glaucoma Diagnosis. (2008). Retrieved November 3, 2021, from https://www.ophthalmologymanagement.com/supplements/2008/october-2008/experiencinghigh-definition-spectral-domain-optic/confirming-a-questionable-glaucoma-diagnosis Haleem, M., Han, L., Hemert, J. (2017). A Novel Adaptive Deformable Model for Automated Optic Disc and Cup Segmentation to Aid Glaucoma Diagnosis. J Med Syst. 42(1) https://doi.org/10.1007/s10916-017-0859-4 Hicks, S., Strümke, I., Thambawita, V., Hammou, M., Riegler, M., Halvorsen, P., Parasa, S. (2022). On evaluation metrics for medical applications of artificial intelligence. Scientific Reports, 12(1), pp. 1-9. https://doi.org/10.1038/s41598-022-09954-8 Introduction to Artificial Intelligence in Ophthalmology. (2020). Retrieved November 16, 2021, from https://eyewiki.aao.org/Introduction_to_Artificial_Intelligence_in_Ophthalmology Li, F., Wang, Z., Qu, G. (2018). Automatic differentiation of Glaucoma visual field from nonglaucoma visual filed using deep convolutional neural network. BMC Med Imaging, 18(1), pp. 35. https://doi.org/10.1186/s12880-018-0273-5 Lu, W., Tong, Y., Yu, Y. & Shen, Y. (2020). Application of machine learning in ophthalmic imaging modalities. Eye and Vision, 7(1), pp. 1-15. https://doi.org/10.1186/S40662-020-00183-6 Lu, W., Tong, Y., Yu, Y., Xing, Y., Chen, Changzheng & Shen, Y. (2018). Applications of Artificial Intelligence in Ophthalmology: General Overview. Journal of Ophthalmology. https://doi.org/10.1155/2018/5278196 95 Muhammad, H., Fuchs, T., De Cuir N. (2017). Hybrid Deep Learning on Single Wide-field Optical Coherence tomography Scans Accurately Classifies Glaucoma Suspects. J Glaucoma, 26(12), pp.1086-1094. https://doi.org/10.1097/IJG.0000000000000765 Nriagu, J. (2019). Encyclopedia of Environmental Health, 6(2). Resnikoff, S., Felch, W., Gauthier, T., Spivey, B. (2012). The number of ophthalmologists in practice and training worldwide: a growing gap despite more than 200,000 practitioners. Br J Ophthalmol, 96, pp. 783-787. https://doi.org/10.1136/BJOPHTHALMOL-2011-301378 Tham, Y., Li, X., Wong, T., Quigley, H., Aung, T., Cheng, C. (2014). Global prevalence of glaucoma and projections of glaucoma burden through 2040: a systematic review and meta-analysis. Ophthalmology, 121, pp. 2081-2090. https://doi.org/10.1016/J.OPHTHA.2014.05.013 Ting, D., Ang, M., Mehta, J. (2019). Artificial intelligence-assisted telemedicine platform for cataract screening and management: a potential model of care for global eye health. Br J Ophthalmol, 103(11), pp. 1537-1538. https://doi.org/10.1136/bjophthalmol-2019-315025 Topol, E. J. (2019). High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine, 25(1), pp. 44-56. https://doi.org/10.1038/s41591-018-0300-7 96 APPENDIX Entities Here can be found a set of tables containing an individually description of all the entities used in the diagram. The tables are composed of Field, Type, Description, Example and Possible Values. patients List of all patients with demographic information Field Type Description Example Possible Values id Bigint Patient’s ID 2134087 birth_date Date Birthday date 08-07-1954 gender Text Patient’s gender M M F pathologies List of possible pathologies Field Type Description Example Possible Values id Bigint Diagnostic identifier 0 0 to 30 label Text Diagnostic name Normal "Normal" "Outras Patologias" "Glaucoma" "Suspeito de Glaucoma" "Edema" "Edema NO" "Edema 97 Macular" "Descolamento" "Descolamento Seroso" "Pos Descolamento" "Buraco Lamelar" "Buraco Macular" "Pos Buraco Macular" "Membrana epi" "Pos Membrana" "Patologia NO" "Maculopatia" "Miope" "Miope Retinopatia" "Alteraçoes Perifericas Retinogra" "RD" "DMRI" "Miopia elevada" "Catarata" "Atrofiamento Peripapilar" "Opacidade Corneana" 98 "Pós Vitrectomia" "Queratocone" "Benson Disease" "Opacificação Capsular" "Hemorragia Subhialoideia Macular" p_diags Diagnostic information by patient eye Field Type Description Example Possible Values patient_id Bigint Patient’s Identifier 2134087 location Text Associated eye Left Left Right diagnostic_id Bigint Diagnostic ID 2 0 to 30 delivery_id Bigint Diagnosis/exam delivery ID 3 0 to 5 deliveries Delivery dates, for data versioning control Field Type Description Example Possible Values id Int Delivery Identifier 2 0 to 5 delivery_date Date Delivery date 2021-07-27 99 e_bio This table contains the data from biometry exams, in the most common format. One record per patient/eye. Field Type Description Example Possible Values patient_id Int Patient’s ID 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry Biometry exam_date Date Exam date 17-10-2014 exam_time String Exam time 13:10 exam_duration String Exam duration in minutes 1 Min al Double Axial length (mm) - measured from the corneal tear film to the inner limiting membrane 22,46 cct Double Cornea thickness (µm) - measured form the corneal tear film to the corneal endothelium 538 ad Double Aqueous depth (mm) - measured form the corneal 2,05 100 endothelium to the anterior surface of the lens acd Double Anterior chamber depth (mm) - measured from cornea posterior to the iris (anterior surface of the lens). Is important in estimating the risk of angle closure glaucoma. 2,59 lt Double Lens thickness (mm) - measured form the anterior to the posterior surface of the lens. Specifically in pseudo phakic condition, measurement of implanted lens is not always achievable (e.g., due to light lens 4,50 101 tilt and the nonlight scattering properties of an implant). rt Double Retina thickness (µm) (System constant) - measured only manually by adjusting the measurement gate in the Ascan from the inner limiting membrane to the retinal pigment epithelium. 200 K1_1 Double Flat meridian 1 (D) - K1 and K2 are the two major meridians, determined using the 3mm ring, are 90 degrees from each other. 45,04 K1_2 Double Flat meridian 2 (º) - the degree where the flat meridian was 173 102 determined. It differs from k2º in 90 degrees. k2_1 Double Steep meridian 1 (D) - K1 and K2 are the two major meridians, determined using the 3mm ring 46,26 k2_2 Double Steep meridian 2 (º) - the degree where the flat meridian was determined. It differs from k1º in 90 degrees. 83 ast_1 Double Astigmatism 1 (D) - displayed in Diopters either in plus or minus cylinder notation dependent on the display option chosen in the Biometry preferences 1,23 ast_2 Double Astigmatism 2 (º) 83 n Double Keratometric index - displays 1,3375 103 the keratometric index chosen for this measurement to convert corneal curvature into corneal power wtw Double White to White (mm) - measured as the horizontal diameter of a best fit circle to the iris boarder 11,31 Ic_1 Double Iris barycenter 1 (mm) - provides information of the shift of the iris center relative to the apex in X and Y coordinates -0,27 Ic_2 Double Iris barycenter 2 (mm) - provides information of the shift of the iris center relative to the apex in X and Y coordinates -0,25 pd Double Pupil diameter 3,62 110 boarder pc_1 Double mm 0,56 pc_2 Double mm 154 pc_x Double Pupil barycenter x-axis (mm) -0,29 pc_y Double Pupil barycenter y-axis (mm) -0,08 al Double Axial length (mm) 25,00 delivery_id Int Delivery identifier 3 0 to 5 e_bio_4 This table contains the data from the biometry exams, in Bio Basic format. One record per patient/eye. Field Type Description Example Possible Values patient_id Int Patient’s Identifier 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry_2 Biometry_2 exam_date Date Exam date 17-10-2014 exam_time String Exam time 13:10 cct Double Cornea thickness (µm) - measured from the corneal tear film to the corneal endothelium 538 wtw Double White to White 11,31 111 (mm) - measured as the horizontal diameter of a best fit circle to the iris boarder ad Double Aqueous depth (mm) - measured from the corneal endothelium to the anterior surface of the lens 2,05 acv Double mm3 150.42 aca_dist Double mm 12.42 sts_dist Double mm 12.27 aca_500_1 Double º 10 aca_500_2 Double º 16 ssa_500_1 Double º 16 ssa_500_2 Double º 23 aod_500_1 Double mm 0.15 aod_500_2 Double mm 0.21 tisa_500_1 Double mm2 0.052 tisa_500_2 Double mm2 0.095 aca_750_1 Double º 16 aca_750_2 Double º 21 ssa_750_1 Double º 22 ssa_750_2 Double º 27 aod_750_1 Double mm 0.31 aod_750_2 Double mm 0.37 112 tisa_750_1 Double mm2 0.108 tisa_750_2 Double mm2 0.164 lt Double Lens thickness (mm) - measured form the anterior to the posterior surface of the lens. Specifically, in pseudophakic condition, measurement of implanted lens is not always achievable 4.44 lv Double mm 0.35 pd Double Pupil diameter (mm) - measured at the diameter of a best fit circle to the pupil boarder 4.9 pc_x Double Pupil barycenter X (mm) - provides information of the shift of the pupil center relative to the -0.41 113 apex in X coordinate pc_y Double Pupil barycenter Y (mm) - provides information of the shift of the pupil center relative to the apex in Y coordinate 0.09 delivery_id Int Delivery Identifier 3 0 to 5 e_pec_single This table contains the data from the PEC exams. Field Type Description Example Possible Values patient_id Int Patient’s Identifier 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry_2 PEC_single exam_date Date Exam date 17-10-2014 exam_time String Exam time 13:10 exam_duration String Exam duration 1 Min fix_monitor String Fixation monitor Gaze/Blind Spot Blind Spot Gaze/Blind Spot Gaze Monitor Gaze Track 114 OFF fix_target String Fixation target Central fl_1 Int Fixation losses 1 1 fl_2 Int Fixation losses 2 11 fp Double False positives (%) 0 Fn Double False negatives (%) 2 stimulus String Stimulus III, White background Double Background (ASB) 31.5 strategy String Strategy SITA-Fast pupil_diameter Double Pupil diameter (mm) 4.2 fovea String Fovea OFF vfi Double Visual Field Index (%) 99 0 a 100 md Double Mean deviation (dB) -0.71 -30,91 a 3,11 psd Double Pattern standard deviation (dB) 1.66 0,86 a 15,07 ght String Glaucoma Hemifield Test result for the Borderline Within Normal Limits Borderline Borderline/General 115 exam Reduction Abnormally High Sensitivity Within Normal Limits *** Low Test Reliability *** General Reduction of Sensitivity Outside Normal Limits Outside Normal Limits *** Low Test Reliability *** Delivery_id Int Delivery Identifier 3 0 to 5 e_pec_summary This table contains the data from the PEC exams. Field Type Description Example Possible Values patient_id Int Patient’s Identifier 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry_2 PEC_summary exam_date Date Exam date 17-10-2014 exam_order Int Exam order among the various exams in the file 1 1 2 3 ght String Glaucoma Hemifield Test result for the exam Borderline Within Normal Limits Borderline Borderline/General Reduction 116 Abnormally High Sensitivity Within Normal Limits *** Low Test Reliability *** General Reduction of Sensitivity Outside Normal Limits Outside Normal Limits *** Low Test Reliability *** fl_1 Int Fixation losses 1 1 fl_2 Int Fixation losses 2 11 fp Double False positives (%) 0 fn Double False negatives (%) 2 fovea String Fovea OFF vfi Double Visual Field Index (%) 99 0 a 100 md Double Mean deviation (dB) -0.71 -30,91 a 3,11 psd Double Pattern standard deviation (dB) 1.66 0,86 a 15,07 delivery_id Date Delivery_id 3 0 to 5 e_rnfl Field Type Description Example Possible 117 Values patient_id Int Patient’s Identifier 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry_2 RNFL exam_date Date Exam date 17-10-2014 reference String Reference database version European Descent (2014) c_mrw String Classification MRM Outside Normal Limits Within Normal Limits Borderline Outside Normal Limits mrw_ts_1 Double MRW TS 1 286 mrw_ts_2 Double MRW TS 2 (%) 63 mrw_ns_1 Double MRW NS 1 354 mrw_ns_2 Double MRW NS 2 (%) 72 mrw_t_1 Double MRW T 1 204 mrw_t_2 Double MRW T 2 (%) 59 mrw_g_1 Double MRW G 1 313 mrw_g_2 Double MRW G 2 (%) 76 118 mrw_n_1 Double MRW N 1 345 mrw_n_2 Double MRW N 2 (%) 79 mrw_ti_1 Double MRW TI 1 365 mrw_ti_2 Double MRW TI 2 (%) 82 mrw_ni_1 Double MRW NI 1 402 mrw_ni_2 Double MRW NI 2 (%) 77 c_rnflt String Classification RNFLT Within Normal Limits Within Normal Limits Borderline Outside Normal Limits No Classification Possible rnflt_ts_1 Double RNFLT TS 1 120 rnflt_ts_2 Double RNFLT TS 2 (%) 34 rnflt_ns_1 Double RNFLT NS 1 79 rnflt_ns_2 Double RNFLT NS 2 (%) 23 rnflt_t_1 Double RNFLT T 1 62 rnflt_t_2 Double RNFLT T 2 (%) 35 119 rnflt_g_1 Double RNFLT G 1 86 rnflt_g_2 Double RNFLT G 2 (%) 47 rnflt_n_1 Double RNFLT N 1 65 rnflt_n_2 Double RNFLT N 2 (%) 33 rnflt_ti_1 Double RNFLT TI 1 148 rnflt_ti_2 Double RNFLT TI 2 (%) 75 rnflt_ni_1 Double RNFLT NI 1 112 rnflt_ni_2 Double RNFLT NI 2 (%) 92 delivery_id Int Delivery Identifier 3 0 to 5 e_rnfl_ou This table contains the data from the RNFL exams. Field Type Description Example Possible Values patient_id Int Patient’s Identifier 2134087 eye String Patient´s eye Right Left Right exam_type String Exam type Biometry_2 RNFL exam_date Date Exam date 17-10-2014 reference String Reference database version European Descent (2014) 126 exam_type exam_type exam_type exam_type exam_type Type of exam exam_date exam_date exam_date exam_date exam_date Date of the exam exam_time exam_time exam_time exam_time Hour of the exam exam_duration exam_duration Duration of the exam in minutes al al al al Axial length (mm) cct cct cct cct cct Cornea thickness (µm) ad ad ad ad Aqueous depth (mm) acd acd acd Anterior chamber depth (mm) lt lt lt lt lt Lens thickness (mm) rt rt Retina thickness (µm) k1_value k1_1 r1_2 simk_flat_1 Flat meridian value (D) k1_angle k1_2 r1_3 simk_flat_2 Flat meridian angle (º) k2_value k2_1 r2_2 simk_steep_1 Steep meridian value (D) 127 k2_angle k2_2 r2_3 simk_steep_2 Steep meridian angle (º) ast_value ast_1 ast_1 ast_steep_1 Astigmatism value (D) ast_angle ast_2 ast_2 ast_steep_2 Astigmatism angle (º) n n n Keratometric index wtw wtw wtw wtw wtw White to White (mm) ic_x ic_1 Iris barycenter xaxis (mm) ic_y ic_2 Iris barycenter yaxis (mm) pd pd pd pd Pupil diameter (mm) pc_x pc_x pc_x pc_x Pupil barycenter xaxis (mm) pc_y pc_y pc_y pc_y Pupil barycenter yaxis (mm) r1_1 r1_1 No Description Available. r2_1 r2_1 No Description 128 Available. r_1 r_1 No Description Available. r_2 r_2 No Description Available. simk_avg simk_avg No Description Available. ast_total_value ast_total_1 No Description Available. ast_total_angle ast_total_2 No Description Available. ast_post_value ast_post_1 No Description Available. ast_post_angle ast_post_2 No Description Available. ast_delta_valu ast_delta_1 No Description Available. ast_delta_angle ast_delta_2 No Description Available. z z No Description Available. 129 rms_hoa rms_hoa No Description Available. aqd aqd No Description Available. cct_aqd cct_aqd No Description Available. pc_value pc_1 No Description Available. pc_angle pc_2 No Description Available. acv acv No Description Available. aca_dist aca_dist No Description Available. sts_dist sts_dist No Description Available. aca_500_angle9h aca_500_1 No Description Available. aca_500_angle3h aca_500_2 No Description Available. ssa_500_angle9h ssa_500_1 No 130 Description Available. ssa_500_angle3h ssa_500_2 No Description Available. aod_500_angle9h aod_500_1 No Description Available. aod_500_angle3h aod_500_2 No Description Available. tisa_500_angle9h tisa_500_1 No Description Available. tisa_500_angle3h tisa_500_2 No Description Available. aca_750_angle9h aca_750_1 No Description Available. aca_750_angle3h aca_750_2 No Description Available. ssa_750_angle9h ssa_750_1 No Description Available. ssa_750_angle3h ssa_750_2 No Description Available. aod_750_angle9h aod_750_1 No Description 131 Available. aod_750_angle3h aod_750_2 No Description Available. tisa_750_angle9h tisa_750_1 No Description Available. tisa_750_angle3h tisa_750_2 No Description Available. lv lv No Description Available. PEC Data Read from e_pec_single and e_pec_summary tables in the database. The resulting data is a union of the data in the tables. We have 1 to 3 PEC exams per patient. We consolidated the following features: • 'vfi_last', 'vfi_min', 'vfi_max', 'vfi_mean', 'vfi_std' • 'md_last', 'md_min', 'md_max', 'md_mean', 'md_std' • 'psd_last', 'psd_min', 'psd_max', 'psd_mean', 'psd_std' • 'last_fl', 'last_ght', 'pec_qtd' Field e_pec_single e_pec_summary Description patient_id patient_id patient_id Patient identifier eye eye eye Eye (left or right) exam_type exam_type exam_type Type of exam exam_date exam_date exam_date Date of the exam exam_time exam_time Time of the exam exam_duration exam_duration Duration of the exam fix_monitor fix_monitor Type of fix monitor (Blind Spot, Gaze 132 Track, etc) fix_target fix_target Type of fix target (Central) fl_1 fl_1 fl_1 Fixation Losses numerator fl_2 fl_2 fl_2 Fixation Losses denominator fp fp fp False positives (%) fn fn fn False negatives (%) stimulus stimulus Type of stimulus (III, White) background background Background (ASB) strategy strategy Strategy used (SITAFast) pd pupil_diameter Pupil Diameter (mm) fovea fovea fovea Fovea vfi vfi vfi Visual Field Index (%) md md md Mean deviation (dB) psd psd psd Pattern standard deviation (dB) ght ght ght Glaucoma Hemifield Test result for the exam exam_order '1' exam_order Exam order in the several exams in the file RNFL Data Read from e_rnfl, e_rnfl_cr and e_rnfl_ou tables in the database. The resulting data is a union of the data in the tables. We have 1 to 3 RNFL exams per patient, although most patients only have 1 exam. 133 Original MRW: • 'mrw_result', 'mrw_g_value', 'mrw_g_perc' • 'mrw_ts_value', 'mrw_ns_value', 'mrw_t_value', 'mrw_n_value', 'mrw_ti_value', 'mrw_ni_value' • 'mrw_ts_perc', 'mrw_ns_perc', 'mrw_t_perc', 'mrw_n_perc', 'mrw_ti_perc', 'mrw_ni_perc' MRW summary: • 'mrw_value_min', 'mrw_value_max', 'mrw_perc_min', 'mrw_perc_max' Original RNFLT: • 'rnflt_result', 'rnflt_g_value', 'rnflt_g_perc' • 'rnflt_ts_value', 'rnflt_ns_value', 'rnflt_t_value', 'rnflt_n_value', 'rnflt_ti_value', 'rnflt_ni_value' • 'rnflt_ts_perc', 'rnflt_ns_perc', 'rnflt_t_perc', 'rnflt_n_perc', 'rnflt_ti_perc', 'rnflt_ni_perc' RNFLT summary: • 'rnflt_value_min', 'rnflt_value_max', 'rnflt,_value_min', 'rnflt_value_max' • 'rnflt_outsides', 'rnflt_borderlines' Field e_rnfl e_rnfl_cr e_rnfl_ou Description patient_id patient_id patient_id patient_id Patient identifier eye eye eye eye Eye (left or right) exam_type exam_type exam_type exam_type Type of exam exam_date exam_date exam_date exam_date Date of the exam reference reference reference reference Reference database used for percentiles and status classification c_mrw c_mrw MRW Classification mrw_ts_value mrw_ts_1 Absolute value for MRW 134 Temporal Superior sector mrw_ns_value mrw_ns_1 Absolute value for MRW Nasal Superior sector mrw_t_value mrw_t_1 Absolute value for MRW Temporal sector mrw_g_value mrw_g_1 Absolute value for MRW Global mrw_n_value mrw_n_1 Absolute value for MRW Nasal sector mrw_ti_value mrw_ti_1 Absolute value for MRW Temporal Inferior sector mrw_ni_value mrw_ni_1 Absolute value for MRW Nasal Inferior sector mrw_ts_perc mrw_ts_2 Percentile ratio of the MRW Temporal Superior sector mrw_ns_perc mrw_ns_2 Percentile ratio of the MRW Nasal Superior sector mrw_t_perc mrw_t_2 Percentile ratio of the MRW Temporal sector 135 mrw_g_perc mrw_g_2 Percentile ratio of the MRW Global mrw_n_perc mrw_n_2 Percentile ratio of the MRW Nasal sector mrw_ti_perc mrw_ti_2 Percentile ratio of the MRW Temporal Inferior sector mrw_ni_perc mrw_ni_2 Percentile ratio of the MRW Nasal Inferior sector c_rnflt c_rnflt c_rnflt c_rnflt RNFLT Classification rnflt_ts_value rnflt_ts_1 rnflt_ts_1 rnflt_ts_1 Absolute value for RNFLT Temporal Superior sector rnflt_ns_value rnflt_ns_1 rnflt_ns_1 rnflt_ns_1 Absolute value for RNFLT Nasal Superior sector rnflt_t_value rnflt_t_1 rnflt_t_1 rnflt_t_1 Absolute value for RNFLT Temporal sector rnflt_g_value rnflt_g_1 rnflt_g_1 rnflt_g_1 Absolute value for RNFLT Global rnflt_n_value rnflt_n_1 rnflt_n_1 rnflt_n_1 Absolute value for RNFLT Nasal