scieee AI-readable full text Open interactive document viewer

Enhanced Diagnostic and Interpretable Model for Febrile Diseases Using Explainable AI and Large Language Model

Attai, Kingsley

Full text

1 IGNATIUS AJURU UNIVERSITY OF EDUCATION RUMUOLUMENI, P.M.B. 5047 PORT HARCOURT NIGERIA ENHANCED DIAGNOSTIC AND INTERPRETABLE MODEL FOR FEBRILE DISEASES USING EXPLAINABLE AI AND LARGE LANGUAGE MODEL ATTAI, KINGSLEY FRIDAY B.Sc. (Hons) (ANU, Ghana), M.Sc. (UNIUYO) MATRIC NO: IAUE/2022/COM/PhD/0003 AUGUST 2025 POSTGRADUATE SCHOOL 2 ENHANCED DIAGNOSTIC AND INTERPRETABLE MODEL FOR FEBRILE DISEASES USING EXPLAINABLE AI AND LARGE LANGUAGE MODEL ATTAI, KINGSLEY FRIDAY B.Sc. (Hons) (ANU, Ghana), M.Sc. (UNIUYO) MATRIC NO: IAUE/2022/COM/PhD/0003 THESIS SUBMITTED TO POSTGRADUATE SCHOOL IN PARTIAL FULFILLMENT OF THE REQUIREMENTS FOR THE AWARD OF DEGREE OF DOCTOR OF PHILOSOPHY(PhD) COMPUTER SCIENCE AUGUST 2025 3 DECLARATION I, Attai, Kingsley Friday with Matriculation Number: IAUE/2022/COM/PhD/0003 declare that this thesis on “Enhanced Diagnostic and Interpretable Model for Febrile Diseases Using Explainable AI and Large Language Model” was carried out by me; that this is my original work and it has not been submitted wholly or in part for the award of a degree in any institution. 4 CERTIFICATION IGNATIUS AJURU UNIVERSITY OF EDUCATION POSTGRADUATE SCHOOL ENHANCED DIAGNOSTIC AND INTERPRETABLE MODEL FOR FEBRILE DISEASES USING EXPLAINABLE AI AND LARGE LANGUAGE MODEL BY ATTAI, KINGSLEY FRIDAY B.Sc. (Hons) (ANU, Ghana), M.Sc. (UNIUYO) MATRIC NO: IAUE/2022/COM/PhD/0003 The board of Examiners certifies that this Thesis is accepted in partial fulfilment of the requirements for the degree of Doctor of Philosophy (PhD) in Computer Science. 5 DEDICATION This thesis is dedicated to GOD ALMIGHTY for the strength to complete this work. 6 ACKNOWLEDGEMENTS I am profoundly grateful to my supervisors, Dr. Constance Amannah and Prof. FaithMichael Uzoka, for their exceptional guidance, steadfast support, and unrelenting mentorship, which have been invaluable throughout this work. I extend my deepest gratitude to Dr. ChiefJames P. Michael, Head of the Department of Computer Science. I sincerely appreciate the mentorship of Prof. Peter O. Eke, the Dean of the Faculty of Natural and Applied Sciences; Prof. Ozioma A. Ekpete; Prof. Nathaniel Ojekudo; Prof. P. O. Asagba; Prof. F. E. Onuodu; Prof. W. Nwankwo; Prof. S. Echezona; Dr. P. Spencer; and Dr. P. Nlerum. I am equally grateful to the entire staff of the Department of Computer Science for fostering a supportive and enabling environment that allowed this research to thrive. I owe immense gratitude to Elder and Deaconess F. M. Attai (my beloved parents), Ekerette, and other family members, past and present, whose unconditional love, warmth, prayers, sacrifices, and steadfast support have been the foundation of my success. This research would not have been possible without their enduring encouragement. Special thanks go to Prof. F. M. Uzoka, Prof. O. U. Obot, Dr D. E. Asuquo, Prof. M. E. Ekpenyong, Dr C. Akwaowo, and the remarkable Febra App team for their invaluable support and contributions, which have significantly enriched this research. I am also indebted to my incredible friends: Dr Kitoye E. Okonny, Dr Cornelia Thomas, and Mr Nelson Chuna, whose moral support, stimulating discussions, and insightful suggestions were a source of inspiration and motivation. 7 TABLE OF CONTENTS TITLE PAGE 1 DECLARATION 3 CERTIFICATION 4 DEDICATION 5 ACKNOWLEDGEMENTS 6 TABLE OF CONTENTS 7 LIST OF TABLES 10 LIST OF FIGURES 11 LIST OF ABBREVIATIONS 13 ABSTRACT 16 CHAPTER ONE: INTRODUCTION 1.1. Background to the Study 17 1.2. Statement of the Problem 20 1.3. Aim and Objectives of the Study 22 1.4. Significance of the Study 22 1.5. Scope of the Study 24 CHAPTER TWO: REVIEW OF RELATED LITERATURE 2.1. Theoretical Review 25 2.2. Conceptual Review 28 2.2.1. Introduction to Tropical Febrile Diseases 28 2.2.2 Diagnostic Challenges in Tropical Febrile Diseases 42 2.2.3 Importance of Accuracy in Diagnosis 44 2.2.4 Machine Learning and Explainable AI in Medical Diagnosis 46 2.2.5 Large Language Models in Healthcare 49 2.2.6 Integration of XAI and LLMs in Medical Diagnosis 55 2.2.7 Transparency in AI Models and Medical Diagnosis 58 2.2.8 Ethical Considerations in AI-based Diagnosis 60 2.2.9 User Acceptance and Trust in AI Diagnosis 62 2.2.10 Human-Machine Collaboration in Diagnosis 65 8 2.3 Empirical Studies 66 2.3 Knowledge Gap 120 CHAPTER THREE: SYSTEM ANALYSIS AND DESIGN 3.1 Method Adopted in the Study 121 3.2 Analysis of the Existing System 121 3.2.1 Architecture of the Existing System 122 3.2.2 Constraints of the Existing System 124 3.3 Analysis of the Enhanced Diagnostic System 124 3.3.1 Justification of the Enhanced Diagnostic System 127 3.4 System Model 128 3.4.1 Architecture of the Enhanced Diagnostic System (EDS) 129 3.4.2. High-Level Input Design of the EDS 152 3.4.3. High-Level Process Design of the EDS 154 3.4.4. High-Level Output Design of the EDS 158 3.5. Database Design of the Enhanced Diagnostics System 160 3.6. Security Mechanism of the Enhanced Diagnostic System 164 CHAPTER FOUR: SYSTEM IMPLEMENTATION 4.1 Choice of Implementation Platform 167 4.2 Justification of the Implementation Platform 167 4.3 System Requirements 170 4.3.1 Hardware Requirements 170 4.3.2 Software Requirements 171 4.4 System Testing 171 4.5 Results and Discussion 172 4.6 Documentation 182 4.6.1 The Design Phase 182 4.6.2 The Development Phase 183 4.6.3 The Testing Phase 184 4.6.4 The Deployment Phase 184 4.7. System Setup and User Manual 184 9 CHAPTER FIVE: CONCLUSION AND RECOMMENDATIONS 5.1. Conclusion 188 5.2. Recommendations 189 5.3. Contributions to Knowledge 192 5.4. Suggestions for Further Study 192 REFERENCES 194 APPENDIX A: FEBRILE DISEASE PATIENT DATASET 225 APPENDIX B: RANDOM FOREST AND LIME CODE 229 APPENDIX C: XGBOOST AND LIME CODE 231 APPENDIX D: MLP AND LIME CODE 233 APPENDIX E: SAMPLE PROMPT OF ML AND XAI RESULTS 235 APPENDIX F: ENHANCED DIAGNOSTIC SYSTEM CODE LISTING236 APPENDIX G: ML MODEL RESULTS 283 APPENDIX H: PATIENT SYMPTOM ASSESSMENT & DISEASE CONFIRMATION QUESTIONNAIRE 286 APPENDIX I: CONFIRMED AND EDS RESULTS 288 16 ABSTRACT Tropical febrile diseases pose significant public health challenges in tropical regions, particularly in resource-limited settings, where timely and accurate diagnoses are critical for effective disease management. However, existing diagnostic systems are black-box and often lack interpretability, making it difficult for healthcare practitioners to understand and trust machine learning (ML) predictions. Additionally, the existing system excludes younger patients, limiting its applicability. This study aimed to develop an integrated diagnostic system that enhances interpretability by leveraging machine learning (ML), Explainable AI (XAI), and a large language model (LLM). A dataset comprising 3,914 patient records with 32 symptoms was obtained from a New Frontiers in Research Fund-sponsored project and preprocessed for analysis. Random Forest (RF) was trained with hyperparameter tuning using GridSearchCV and evaluated with 5-fold cross-validation to optimize its performance. The tuned model achieved high diagnostic performance, demonstrating its effectiveness in diagnosing six febrile diseases. To enhance interpretability, a Model Interpretability Framework (MIF) was incorporated, integrating Local Interpretable Model-Agnostic Explanations (LIME) for visual insights and ChatGPT for textual explanations. This combination provided users with a clear understanding of the diagnostic process, addressing the limitations of black-box ML models. The system was implemented using Python in a development environment that included Google Colaboratory, Visual Studio Code, Flet, and MySQL for database management. Agile development methodology and Unified Modeling Language tools were employed to create a user-friendly design for clinicians and patients. The diagnostic system was evaluated using precision, recall, F1-score, and AUC-ROC metrics. The RF model achieved outstanding precision for most diseases, high recall (≥0.97), and an AUC-ROC of 0.99, indicating nearly flawless training performance with few misclassifications. On the test dataset, the model performed better in diagnosing malaria (F1-score = 88%, precision = 85%), urinary tract infection (F1-score = 72%, precision = 80%), and respiratory tract infection (F1-score = 72%, precision = 77%). The evaluated system demonstrated strong predictive performance, correctly identifying 72 out of 99 tested patient cases, with 26 cases misclassified and one case completely missed. It achieved the highest detection rates for Malaria and HIV/AIDS, accurately identifying all cases without false positives. This work enhances diagnostic solutions for tropical healthcare settings by improving explainability and ensuring inclusivity for younger patients, addressing key limitations in existing systems. The integration of ML, XAI, and LLM enhances diagnostic performance and lays the groundwork for future AI-driven healthcare innovations, ultimately improving health outcomes in resource-scarce regions. Healthcare providers should adopt this system to enhance the efficiency of disease diagnosis and extend the system to diagnose additional medical conditions. 17 CHAPTER ONE INTRODUCTION 1.1 Background to the Study Diagnosing tropical febrile diseases in patients can be a huge challenge for healthcare providers due to the confusing symptoms presented by these diseases. Accurate diagnosis of these diseases requires a combination of laboratory testing, clinical evaluation, and, in specific situations, sophisticated diagnostic instruments that may not be easily accessible in some hospitals and health centres in countries with lowto middle-income economies (LMICs). In medical contexts, “febrile” is frequently used to describe conditions characterized by an elevation in body temperature. Diseases with fever as an accompanying symptom are known as febrile diseases, and they are the leading cause of high mortality rates throughout the world's tropical and subtropical regions (World Health Organization [WHO], 2023a; Centres for Disease Control and Prevention [CDC], 2021; Guo et al., 2017). Fever or pyrexia is a surge in the body temperature above the normal diurnal variations of 36.50C to 37.50C and it is often seen as a symptom of an infectious disease. Febrile diseases are prevalent in tropical and subtropical regions due to the high temperature, heavy rainfall, and high humidity, which create favourable conditions for infectious agents. Several factors contribute to the prevalence and thriving of these diseases in tropical and subtropical regions. Environmental factors, biological factors such as the high biodiversity, and social factors such as poor sanitation and vector control, drug resistance, and self-diagnosis contribute to increasing the burden of tropical diseases in these regions. Tropical diseases are infectious and non-infectious diseases, as well as diseases caused by environmental conditions or nutritional deficiencies (Rupali, 2019). Infectious diseases are caused by pathogens, which are harmful agents to a susceptible host (human body), and these pathogens can be viruses, bacteria, fungi, or parasites. Infectious diseases are classified into viral, bacterial, parasitic, and fungal diseases (Goeijenbier et al., 2014). Humans can become infected with these pathogenic or infectious agents through an infected person, animal, vehicle, airborne 18 particle, or vector. A vehicle is anything that is contaminated and can be found in our environment, such as a table, clothing, water, food, tools, etc., while a vector is an animal or insect that spreads an infectious agent from one infected person or animal to another (Ekpenyong et al., 2020). Infectious illnesses are the most common cause of fever, and fever is the patient’s defensive reaction to an infection (Kluger et al., 1998). According to Paul (2024), these infectious diseases are typically categorized into three groups: those that result in high rates of death, those that heavily burden the population with disabilities, and those that, because of their sudden and rapid spread, may have catastrophic repercussions on a worldwide scale. It is projected that between 2030 and 2050, the effects of climate change will result in approximately 250,000 deaths annually due to tropical diseases, diarrhoea, heat stress, and undernutrition (WHO, 2023b). According to Albareedy and Ramadan (2023), Africa has been found to have a high prevalence of both emerging and reemerging tropical diseases, which is one of the main causes of death on the continent. The demographic group for the study is patients above four (4) years old because the data collection instrument employed was not designed to capture relevant symptoms in younger children. Therefore, records of patients under the age of 5 years will be removed from the dataset because of the potential impact of including this age group without proper symptom documentation and the potential consequences of inaccurate data on the machine learning model's performance. The tropical diseases considered in this study include typhoid fever (Marchello et al., 2022), malaria (WHO, 2023), HIV and AIDS (Uwishema et al., 2022a), respiratory tract infection (Woodall et a., 2022), urinary tract infection (Li et al., 2022), and Tuberculosis (Uwishema et al., 2022b). Several factors, such as climate and environmental conditions, poor sanitation and hygiene, overcrowding, urbanization, limited access to healthcare, poverty, and malnutrition, as well as political and economic factors, are responsible for the prevalence of these diseases in tropical regions, particularly Nigeria. Therefore, early diagnosis and treatment of these diseases can prevent several complications, such as reducing the risk of transmission to others, the risk of intestinal perforation in typhoid fever, and ultimately death. According to the WHO (2023a), the initial symptoms of malaria may be mild, similar to many other febrile 19 diseases, and if P. falciparum malaria is not diagnosed and treated early, it can cause severe illness and death in as little as 24 hours. Since tropical diseases present confusable symptoms, it is important to consider the presence of two or more additional conditions that may complicate diagnoses. The presence of comorbidities can complicate the management and treatment of tropical febrile diseases, necessitating a comprehensive strategy to address all of the health issues patients may be facing. Comorbidities should be a cause of concern for healthcare professionals since they can complicate the diagnosis and treatment of tropical febrile diseases. The word "comorbidity" in medicine refers to the coexistence of one or more additional chronic health conditions with a primary disease. The prefix "co" means together, and the word "morbidity" is the presence of an illness or a disease. Diagnosing tropical febrile diseases can be challenging because of overlapping symptoms, a range of clinical presentations, and limited access to qualified medical professionals in lowto middle-income areas. According to Meyer et al. (2013), diagnostics accuracy declines when healthcare providers encounter difficult cases. Therefore, timely differential diagnoses of these diseases by harnessing the power of technology are important to avoid several complications and possible misdiagnoses (Uzoka et al., 2017). The application of technology in healthcare can reduce missed, delayed, or inaccurate diagnoses (El-Kareh et al., 2013). The application of large language models (LLMs) and artificial intelligence (AI) in healthcare could lead to better patient outcomes and revolutionize the solutions to healthcare-related challenges such as handling the high level of imprecision, especially in the diagnosis of illnesses (Alowais et al., 2023; Thirunavukarasu et al., 2023; Briganti, 2023; Yang et al., 2023; Agrawal et al., 2022). AI provides valuable decision support by synthesizing and interpreting complex medical information (Kaul et al., 2023) while LLMs can help with the interpretation of unstructured data, like clinical notes, contributing to evidence-based decisionmaking in patient care (Yang et al., 2022). ML algorithms are seen as "black boxes," casting doubt on their methods of operation and judgmental processes. Because of this uncertainty, machine learning systems have found it challenging to be accepted in sensitive but crucial domains such as healthcare, where they have enormous 20 potential benefits (Linardatos et al., 2020). To address this challenge of interpretability and transparency, eXplainable Artificial Intelligence (XAI) has been developed to enable human users to comprehend and accept the output generated by machine learning algorithms. Consequently, the use of XAI in healthcare can result in better collaboration between AI and medical professionals and increased user confidence in AI systems. 1.2 Statement of the Problem Febrile diseases remain a leading cause of morbidity and mortality worldwide, with the heaviest burden concentrated in lowand middle-income countries (LMICs). In the WHO African Region alone, malaria accounted for an estimated 246 million cases (94% of global cases) and 569,000 deaths (95% of global malaria deaths) in 2023 (WHO, 2023a). Alongside malaria, typhoid fever continues to impose a considerable health burden, with reports of culture-confirmed cases from 42 out of 57 African countries between 1900 and 2018, and with outbreaks increasing in frequency and scale over time (Kim et al., 2019). HIV/AIDS remains highly prevalent in sub-Saharan Africa, where 25.6 million people were living with the virus at the end of 2022, of whom 20.8 million were receiving antiretroviral therapy (WHO, 2023d). Tuberculosis (TB), often co-occurring with HIV, is also widespread, with the continent accounting for one-quarter of all new TB cases globally in 2022 (2.5 million people) and 424,000 TB-related deaths, over one-third of the global total (WHO, 2023e). Respiratory tract infections (RTIs) add further to the burden: a systematic review across Africa between 2013 and 2023 found pooled prevalence rates of 19.9% for rhinovirus, 12.0% for adenovirus, 8.9% for RSV, 5.2% for influenza, and 5.0% for Streptococcus pneumoniae (Nyahoda et al., 2024). Similarly, urinary tract infections (UTIs) affect about 3.7% of the African population, more than double the global average of 1.6% (Mengistu et al., 2023). These statistics reveal that malaria is not an isolated health problem but is closely linked with other febrile and infectious diseases, many of which share overlapping symptoms and risk factors. The combined burden of malaria, typhoid, TB, HIV/AIDS, RTIs, and UTIs is amplified by weak health systems, limited diagnostic 21 infrastructure, and inadequate access to skilled physicians in LMICs. This challenge is further compounded by the migration of experienced healthcare workers to developed countries, leaving community health workers (CHWs) to provide frontline services. While CHWs play an essential role, they depend on simplified protocols such as Standing Orders (SOs) that cannot adequately address differential diagnoses when patients present with overlapping or confusable symptoms. To address these challenges, the Febra Diagnostica (Febra) app was developed as an Android-based mobile health (mHealth) application. Built on an Analytic Hierarchy Process (AHP)-based Multi-Criteria Decision Analysis (MCDA) model, the app provides a structured way to distinguish among febrile diseases by incorporating multiple risk factors and symptoms. While this represents a pioneering effort, its limitations are notable. AHP is a knowledge-driven rather than data-driven approach, meaning it cannot adapt or improve with new data inputs. The system also excludes patients below 17 years of age, reducing its applicability, and its mathematical complexity makes it less transparent to non-technical stakeholders such as CHWs, potentially limiting acceptance and trust. In contrast, machine learning (ML) methods such as Random Forest, Extreme Gradient Boosting, and Multi-layer Perceptron offer powerful alternatives for capturing complex, non-linear relationships in large patient datasets. These models can automatically learn patterns and improve diagnostic accuracy. However, their black-box nature poses a major challenge: predictions are often opaque and difficult to interpret, creating mistrust among healthcare providers and limiting adoption in clinical practice. To address this, the integration of Explainable AI (XAI) techniques such as Local Interpretable Model-Agnostic Explanations (LIME) can provide transparency by showing the contribution of each feature to a diagnostic prediction. Furthermore, coupling XAI outputs with Large Language Models (LLMs) enables explanations to be translated into user-friendly narratives tailored to the knowledge level of different stakeholders, from CHWs to physicians. Despite the growing application of ML in healthcare, there remains a clear knowledge gap in combining ML, XAI, and LLMs for febrile disease diagnosis in resource-poor settings. Existing systems have yet to fully address the dual challenge 22 of diagnostic performance and interpretability in a way that builds trust, supports decision-making, and aligns with global health goals. This study seeks to bridge this gap by developing a data-driven, interpretable, and communicative diagnostic framework that enhances transparency, fosters stakeholder confidence, and ultimately contributes to reducing the burden of febrile diseases in LMICs. 1.3 Aim and Objectives of the Study This study aimed to enhance the interpretability of febrile disease diagnosis using Explainable Artificial Intelligence (XAI) and Large Language Model (LLM). The objectives of the study were to: i. acquire a tropical febrile disease patient dataset from an NFRF-sponsored project, including cases of Enteric Fever, Malaria, HIV and AIDS, Respiratory Tract Infection, Urinary Tract Infection, and Tuberculosis. ii. develop a Disease Diagnostic Model (DDM) using the Random Forest (RF) algorithm, chosen for its robustness to noisy data, ability to handle highdimensional features, and compatibility with interpretability frameworks iii. train and test the DDM using an 80/20 split of the febrile disease dataset. iv. integrate a model interpretability framework (MIF) to enhance the DDM’s transparency. v. implement and evaluate the integrated febrile disease diagnostic model using Python and appropriate performance metrics. 1.4 Significance of the Study This study will provide significant benefits to primary, secondary, and tertiary stakeholders in the healthcare sector by improving the interpretability and accessibility of tropical febrile disease diagnosis. 1.4.1 Primary Stakeholders Medical professionals: doctors, nurses, and community health workers (CHWs), will benefit from an AI-powered diagnostic tool that enhances early detection, reduces the risk of misdiagnosis, and provides interpretable insights into disease predictions. CHWs, especially in resource-limited settings, can leverage this system 23 to make more informed decisions, reducing the burden on referral hospitals, while doctors and nurses can use it as a decision-support tool to enhance treatment and patient outcomes. Healthcare organizations: hospitals, clinics, and research institutions can adopt this diagnostic system to improve workflow efficiency, minimize diagnostic errors, and support disease pattern research. By utilizing XAI and LLMs, healthcare facilities can better manage patients, allocate resources efficiently, and advance medical research on febrile diseases. 1.4.2 Secondary Stakeholders Healthcare Administrators: Hospital administrators and policymakers will gain access to a data-driven approach for improving diagnostic accuracy and optimizing resource allocation. This system can help reduce unnecessary tests, improve patient care quality, and inform policy decisions related to febrile disease management. Medical Researchers: AI-assisted diagnostics will provide valuable insights into disease modeling, predictive analytics, and treatment responses. Researchers in epidemiology and infectious diseases can leverage the dataset and methodologies developed in this study to advance AI applications in healthcare. Healthcare Technology Developers: AI developers and healthcare tech companies can build upon this study to refine AI-driven diagnostic tools, enhance their interpretability, and expand their application to other medical conditions. The system’s framework can serve as a foundation for future AI innovations in healthcare. Public Health Officials: Government agencies responsible for disease control and public health policy will benefit from real-time epidemiological data generated by the AI model. This can support disease surveillance, outbreak management, and informed decision-making for healthcare interventions. Insurance Providers: Health insurance companies can use the system to enhance risk assessment and claims processing, ensuring accurate diagnosis verification. Improved diagnostics will contribute to preventive care, reducing long-term healthcare costs and improving coverage accuracy. 24 1.4.3 Tertiary Stakeholders Healthcare Educators: Medical schools and training institutions can incorporate AI-powered diagnostic tools into their curriculum to train future healthcare professionals on AI-assisted medical decision-making. Healthcare IT Professionals: IT specialists will play a crucial role in implementing and maintaining this AI-driven diagnostic system, ensuring its seamless integration into healthcare infrastructures. Medical Device Manufacturers: Companies producing diagnostic devices can integrate this AI-powered system to improve the interpretability and reliability of their products, enhancing healthcare delivery in both urban and rural settings. Pharmaceutical Companies: Drug manufacturers can utilize AI-driven diagnostic insights to develop targeted treatments for febrile diseases, improving drug efficacy and patient care. This study significantly advances the application of Explainable AI in healthcare, promoting transparency in machine learning-based diagnostics. It provides a replicable framework for integrating XAI and LLMs in disease diagnosis, with potential adaptability to other medical conditions. The findings will inform policymakers, healthcare providers, and researchers, ultimately improving healthcare access, reducing misdiagnosis, lowering costs, and saving lives in tropical and subtropical regions where febrile diseases remain a major health challenge. 1.5 Scope of the Study The research was conducted in the Niger-Delta region of Nigeria, an area characterized by a high prevalence of tropical febrile diseases due to environmental and socio-economic factors. The study acquired a relevant patient dataset from an NFRF-sponsored project for six tropical febrile diseases: Enteric Fever (ENFVR), Malaria (MAL), HIV and AIDS (HVAD), Respiratory Tract Infection (RTI), Urinary Tract Infection (UTI), and Tuberculosis (TB). A Disease Diagnostic Model (DDM) was designed using the Random Forest (RF) algorithm, and the DDM was trained and tested using an 80/20 data split to ensure robustness and generalizability. To 25 address the challenge of interpretability often associated with machine learning models, a Model Interpretability Framework (MIF) was integrated into the system. This allowed for transparent and explainable outputs, making the diagnostic recommendations more understandable to healthcare practitioners. The system was implemented and evaluated using the Python programming language and appropriate performance metrics to validate its effectiveness. The work was conducted between July 2023 and January 2025, following an Agile and data-driven development methodology that emphasized iterative design, continuous feedback, and adaptive refinement of the diagnostic system. 32 selecting the best approach. A timely and precise diagnosis is essential for managing and controlling malaria effectively. b) Typhoid Fever: A bacterial infection that can spread throughout the body and damage numerous organs. If left untreated, it can become fatal or lead to major complications. Typhoid fever is caused by a bacterium called "Salmonella enterica serotype Typhi" (Ashurst et al., 2023) and is spread by contaminated water and food. Symptoms of typhoid fever include fever, headache, abdominal pain, constipation or diarrhea, fatigue, anorexia, nausea, vomiting, and mobility disorder (Nelwan et al., 2023). Some conventional methods for diagnosing typhoid fever include Blood Culture and widal tests (Bector and Kumar, 2023). Blood culture involves isolating Salmonella Typhi bacteria from a blood sample and culturing the blood in an appropriate medium. If bacteria are present, the culture will grow over time. Because of its high specificity, blood culture is regarded as the gold standard for diagnosis; however, because of its time-consuming nature and need for specialized facilities and equipment, results may take several days to obtain. The Widal test is a widely available, relatively simple method that uses a serological test to detect antibodies against Salmonella Typhi antigens (Sapkota et al., 2023). However, it lacks specificity, is prone to false positives and false negatives, and is not advised for routine use. Some advancements in diagnostic technologies for typhoid fever include Polymerase Chain Reaction (PCR), Enzyme-Linked Immunosorbent Assay (ELISA), Typhoid Rapid Tests, Biosensors, and Lateral Flow Assays. PCR amplifies and detects particular Salmonella Typhi DNA sequences using molecular techniques (Mostafa-Mahmoud et al., 2023). Although PCR is very sensitive and, specific, enabling rapid detection, it requires specialized equipment, and trained personnel, which can be relatively expensive. ELISA detects particular antigens or antibodies using immunoassay. The reaction is measured after the patient's blood or serum is added to wells coated with specific antigens or antibodies (Panzner et al., 2022). Quick results and 33 automation potential are provided, but the selection of antigens can affect sensitivity and specificity, which may necessitate laboratory equipment. Typhoid Rapid Tests identify particular antigens or antibodies associated with Salmonella Typhi through immunochromatographic assays (Strobel et al., 2022). Variable sensitivity and specificity may be influenced by factors such as the infection stage, but it yields fast results and is appropriate for field settings without the need for specialized equipment. Biosensors employ a variety of transducers to transform a biological response into an electrical signal that can be used by biosensor devices to identify particular biomolecules linked to Salmonella Typhi (Fathi et al., 2023). Although biosensor technologies are still being developed and validated, they have the potential to detect things quickly and sensitively. Lateral Flow Assays use lateral flow technology to identify particular antigens or antibodies (Frempong et al., 2022). It generates fast results and is appropriate for point-of-care testing, but its sensitivity may be limited and its performance is inconsistent. The best diagnostic approach to use will depend on several variables, including the resources available, the environment (field or laboratory), and the particulars of the diagnostic task. Integrating various techniques could improve the precision of the diagnosis. Typhoid fever management and treatment require prompt and precise diagnosis. c) Human Immunodeficiency Virus (HIV) and Acquired Immune Deficiency Syndrome (AIDS): The human immunodeficiency virus (HIV) targets the body's defenses, particularly the white blood cells known as CD4 cells and these CD4 cells are destroyed by HIV, which reduces patient's resistance to opportunistic infections. HIV can spread through blood, perinatal transmission, and sexual contact. Some of the early HIV symptoms are fever, fatigue, swollen lymph nodes, joint and muscle pain, and headache (Yetman, 2023). The advanced stage of HIV infection known as acquired immunodeficiency syndrome (AIDS) develops when the virus severely 34 compromises the immune system of the body. When a person's CD4 count drops below 200 cells per milliliter of blood (200 cells/mm3), or if they get one or more opportunistic infections, regardless of CD4 count, they are considered to have progressed to AIDS. AIDS-related symptoms include opportunistic Infections, unexplained and significant weight loss, chronic diarrhea or prolonged gastrointestinal issues, persistent cough, and shortness of breath (CDC, 2021). HIV can be diagnosed through rapid diagnostic tests which provide results the same day and make early diagnosis, treatment, and prevention much easier. HIV self-tests are also available in pharmacies and drug-stores which can also be used by individuals to test themselves (CDC, 2022). A qualified and trained health or community worker at a community centre or clinic must perform confirmatory testing since a single test cannot provide a complete HIV-positive diagnosis. Some of the conventional methods for diagnosing HIV/AIDS include Enzyme-Linked Immunosorbent Assay (ELISA) and Western Blot. ELISA employs a serological test to identify HIV-antibody antibodies. When a patient's blood is added to wells coated in HIV antigens, the reaction is noted. A follow-up confirmatory test may be necessary because ELISA has a high sensitivity and is appropriate for large-scale screening, but it may also give false positive results in cases of early infection (Alfie., 2023). The Western Blot is an additional test used to confirm the presence of HIV antibodies. It uses electrophoresis to separate and identify particular HIV proteins. It is used to validate positive ELISA results and has a high specificity, but the process is very complicated and may call for specialized equipment (Guiraud et al., 2023; Korkusuz et al., 2022). Some advancements in diagnostic technologies for HIV/AIDS include Nucleic Acid Testing (NAT), Polymerase Chain Reaction (PCR), FourthGeneration HIV Tests, Point-of-Care Tests (Rapid Tests), Antigen/Antibody Combination Tests, and Home Testing Kits. Nucleic Acid Testing (NAT) uses molecular techniques to detect HIV RNA or DNA. It has high sensitivity in 35 the early detection of infection but requires specialized equipment and is more expensive than antibody tests (Jani and Peter, 2022). PCR amplifies and detects specific HIV DNA or RNA sequences using molecular techniques (White et al., 2022). Early HIV detection is beneficial for early-stage diagnosis; however, it needs specialized equipment, which can be costly. Fourth-generation HIV tests identify early antibodies and viral protein in addition to p24 antigen and HIV antibodies. Confirmatory testing is still required even though this method has a shorter window period and better sensitivity during acute infection. It may also have a higher false-positive rate (Shallal et al., 2022). Point-of-care tests (Rapid Tests) use immunoassays to detect HIV antibodies or antigens. It provides rapid results appropriate for field settings and doesn't require specialized equipment, but may need confirmatory testing due to its variable sensitivity and specificity (Luo et al., 2022; Hsieh et al., 2022). Antigen/Antibody Combination Tests Detects both HIV antibodies and p24 antigen and it is similar to ELISA, with improved sensitivity. This method reduces the window period, used for routine screening but may have possible false positives and confirmatory testing may be needed (White et al., 2022; Whitney et al., 2022). Home Testing Kits are meant for self-testing at home and are comparable to rapid tests (King et al., 2022). After gathering their sample and running the test, users can get the results in a matter of minutes. Although testing is now more accessible and confidential thanks to this method, counseling may not be available to all candidates, and accuracy depends on proper application. The goal of developments in HIV/AIDS diagnostic technologies is to increase test accuracy, shorten the window of opportunity, and improve accessibility. The stage of infection, the resources at hand, and the environment in which testing is done all influence the selected diagnostic approach. Prompt and 36 precise diagnosis is essential for prompt treatment initiation and prevention of further transmission. d) Respiratory Tract Infection (RTI): Any infectious disease affecting the upper or lower respiratory tract is referred to as a respiratory tract infection. The common cold, laryngitis, pharyngitis/tonsillitis, acute rhinitis, acute rhinosinusitis, and acute otitis media are examples of upper respiratory tract infections (URTIs) (Gulmez, 2022; Arroll, 2011). Pneumonia, tracheitis, bronchiolitis, and acute bronchitis are examples of lower respiratory tract infections (LRTIs) (Centre for Clinical Practice [CCP], 2008). Bacteria (such as Streptococcus pneumoniae), viruses (such as influenza and respiratory syncytial virus), and fungi are the causative agents for respiratory tract infections (Calderaro et al., 2022). Common symptoms of RTI include fever, dry or productive cough, shortness of breath, chest pain, and fatigue (Honarmand et al., 2022). Some of the conventional methods for diagnosing RTI include Clinical Assessment, Chest X-ray, Sputum Culture, and Throat Swab and Culture. Clinical assessment entails a physical examination as well as an assessment of clinical symptoms. Medical professionals evaluate symptoms like fever, coughing, and dyspnea and listen for unusual noises coming from the lungs. Although the patient's condition is quickly and initially assessed, the specific pathogen cannot be identified and the patient only exhibits non-specific symptoms. An X-ray of the chest is used to diagnose anomalies like consolidation or infiltrates by visualizing the lungs and surrounding structures. This approach has limited ability to differentiate between bacterial and viral infections, but it helps identify structural abnormalities (Margret et al., 2022; Ait-Nasser and Akhloufi, 2023). The purpose of sputum culture is to identify bacterial pathogens by collecting and cultivating sputum samples (Shen & Sergi, 2023). It takes time, might not 37 work for viral infections, and might be contaminated, but it detects bacterial pathogens and directs antibiotic treatment (Cleveland Clinic, 2023). Throat Swab and Culture are used to identify bacterial pathogens, like streptococcal bacteria, that cause throat infections (Hirsch, 2023). This technique has limited sensitivity for viral infections, but it works well for bacterial infections like strep throat. Some advancements in diagnostic technologies for RTI include Nucleic Acid Amplification Tests (NAATs), Immunofluorescence Assays, Point-of-Care Tests (Rapid Tests), Next-Generation Sequencing (NGS), Biomarker Detection, and Digital PCR/Droplet Digital PCR. Polymerase chain reaction (PCR) is one of the molecular techniques used by NAATs to amplify and detect the genetic material (DNA or RNA) of pathogens (Calderaro et al., 2022). These techniques allow for rapid and accurate detection and multiplexing of pathogens, but they also require specialized equipment and expertise. Fluorescence signals are used in immunofluorescence assays to show which viral antigens are present (Calderaro et al., 2022). It is limited to certain pathogens and antibodies, but it yields fast results and is appropriate for a variety of respiratory viruses. Point-of-care tests identify particular viral antigens or antibodies using straightforward lateral flow tests with visual readouts. Although point-of-care tests yield variable sensitivity and specificity and may need confirmatory testing, they are suitable for field settings, yield rapid results, and do not require specialized equipment (Gentilotti et al., 2022). With Next-Generation Sequencing (NGS), multiple pathogens can be identified at once by sequencing RNA or DNA using high-throughput sequencing of genetic material. Although NGS provides thorough and objective detection and is appropriate for identifying novel pathogens, its complex data analysis and high cost are a drawback (Shao et al., 2022). 38 Biomarker detection uses proteins or other infection-related markers to identify particular biomarkers linked to respiratory infections (Hogendoorn et al., 2022). This approach has variable specificity and might not be able to identify specific pathogens, but it might be able to provide information on treatment response and severity. Digital PCR and Droplet Digital PCR uses precise quantification of nucleic acids for accurate detection of pathogens by partitioning samples into individual reactions, allowing absolute quantification. It yields high sensitivity and precise quantification but requires specialized equipment and expertise (Milosevic et al., 2022; Jiang et al., 2020). The goal of advancements in respiratory tract infection diagnostic technologies is to increase efficiency, precision, and the capacity to identify multiple pathogens at once. The pathogens that are suspected, the extent of the infection, and the resources at hand all influence the diagnostic approach that is selected. For respiratory tract infections to be appropriately managed and treated, a timely and accurate diagnosis is crucial. e) Tuberculosis (TB): Tuberculosis is a communicable infection that typically affects the lungs and Mycobacterium tuberculosis is the bacterium that causes TB. This infection can be transmitted through respiratory droplets or by air (Pöhlker et al., 2023). Although the TB bacteria typically affect the lungs, they can also affect the kidney, spine, and brain, among other parts of the body (Hammami et al., 2022). The common symptoms of TB include Persistent cough for more than three weeks, chest pain, unexplained and significant weight loss, fatigue, fever, and night Sweats (Sharma et al., 2023). Some of the conventional methods for diagnosing TB include Tuberculin Skin Test (TST), Chest X-ray, Sputum Smear Microscopy, and Sputum Culture. TST or Mantoux Test measures the amount of induration (swelling) at the injection site 48–72 hours after administering an intradermal injection of purified protein derivative (PPD) to evaluate the delayed-type hypersensitivity reaction. Easy to use and extensively accessible, but it is 39 unable to distinguish between latent and active tuberculosis and is impacted by prior BCG vaccination. A chest X-ray is a radiographic examination used to find cavities, lung lesions, or other indications of tuberculosis (Sajed et al., 2023). It helps detect pulmonary tuberculosis, but is unable to establish the existence of an active infection or distinguish TB from other lung conditions (Giannelli et al., 2022). Sputum Smear Microscopy examines sputum under a microscope to detect M. tuberculosis in sputum samples using acid-fast bacilli (AFB) staining. It is easy to use and reasonably priced, but it cannot differentiate between different mycobacterial species, has a moderate sensitivity, and may miss low bacterial loads. Sputum Culture is used in the culturing of sputum samples to isolate and identify M. tuberculosis. It is extremely precise, and permits drug susceptibility testing, but necessitates specialized facilities and takes several weeks to produce results (Rimal et al., 2022). Some advancements in diagnostic technologies for TB include Nucleic Acid Amplification Tests (NAATs), Xpert MTB/RIF Assay, Line Probe Assays (LPAs), Cytokine Release Assays (IGRAs), Biosensors and Whole-Genome Sequencing (WGS). NAATs amplify and detect M. tuberculosis DNA using molecular methods, such as polymerase chain reaction (PCR). Although it needs specific tools and knowledge, it yields quick results when identifying drug-resistant strains (Wen et al., 2023). The Xpert MTB/RIF Assay targets rifampicin resistance and M. tuberculosis DNA using automated cartridge-based NAAT. In a matter of hours, it can simultaneously identify TB and resistance to rifampicin; however, its use is expensive and equipment-dependent (Horne et al., 2019). LPAs employ genotypic assays, which identify particular genetic markers linked to drug-resistance mutations in M. tuberculosis. Although it provides quick drug resistance detection, it necessitates laboratory equipment and experience (Diriba et al., 2022). 40 IGRAs use blood tests to measure the release of interferon-gamma in response to M. tuberculosis antigens. Although it distinguishes between latent and active TB, it has drawbacks, including potential false-positive results in some populations and cost considerations (Banaei et al., 2016). Biosensor devices are used to detect specific biomolecules associated with TB. This technique turns a biological response into an electrical signal that can be detected using a variety of transducers. Although biosensor technologies are still being developed and validated, they have the potential to detect things quickly and sensitively (Zhou et al., 2011). WGS provides precise genetic data for strain typing and epidemiological research by employing a thorough sequencing of the whole genome of Mycobacterium tuberculosis. Although it requires advanced sequencing technology and bioinformatics expertise, it has detailed genetic information with high resolution (Zhang et al., 2023). Digital PCR partitions samples into individual reactions, allowing absolute quantification of nucleic acids for accurate detection of M. tuberculosis. It has a high sensitivity and accurate quantification, but it needs specific tools and knowledge (Yang et al., 2017). The goal of improvements in tuberculosis diagnostic technologies is to increase efficiency, precision, and the capacity to identify drug resistance. The suspected clinical presentation, the existence of drug resistance, and the resources at hand all influence the diagnostic method selection. Efficient diagnosis is essential for successful tuberculosis management and control. f) Urinary Tract Infection (UTI): Urinary Tract Infections are common infections caused by bacteria that enter the urethra usually from the skin or the rectum and infect the urinary tract (Flores-Mireles et al., 2015). Although the infections can affect various parts of the urinary tract, bladder infections (cystitis) are the most common type and kidney infection (pyelonephritis) is another form of UTI. Some common symptoms include painful urination, the urge to urinate more often than usual (frequent urination), having an empty 41 bladder but still feeling the need to urinate(urgency), discomfort or pain in the lower abdomen, cloudy or bloody urine, and fever (CDC, 2021). Some of the conventional methods for diagnosing UTI include urinalysis, urine culture, and gram staining. Urinalysis is the process of analyzing urine for physical, chemical, and microscopic features by looking for bacteria, red blood cells, white blood cells, and other abnormalities. The method is easy to use, economical, and gives preliminary information; however, it might not be able to differentiate between distinct pathogens. In urine culture, urine is cultured for the identification of specific bacteria. Although it takes time (24– 48 hours for results) and may be tainted by outside factors (Karah et al., 2020), this method finds the causative organism and directs antibiotic therapy. Gram staining is a technique used to identify bacteria as Gram-positive or Gramnegative by staining their cells according to the properties of their cell walls. It can quickly identify bacteria in clinical specimens, but it is only effective in identifying specific types of bacteria and is not appropriate for all pathogens (Tripathi and Sapra, 2023). Some advancements in diagnostic technologies for UTI include Molecular Testing (PCR), Matrix-Assisted Laser Desorption/Ionization Time-of-Flight Mass Spectrometry (MALDI-TOF MS), Flow Cytometry, Biosensors, NextGeneration Sequencing (NGS) and Digital PCR. Polymerase chain reaction, or PCR, is a technique used in molecular testing to amplify and identify specific pathway DNA sequences. It needs specialized tools and knowledge, but it offers quick results, great accuracy, and pathogen detection for particular diseases. MALDI-TOFMS determines microorganisms based on their mass spectra using bacterial isolates from culture. This technique is mostly employed in laboratory settings and is quick and accurate in identifying pathogens but needs specific equipment (Singhal et al., 2015). Flow Cytometry quantitative analysis of cells in a fluid sample by measuring the size, granularity, and fluorescent properties of cells, including white blood 48 heavily on reinforcement learning (RL), which creates treatment plans using patient data from electronic health records (EHR) (Smith et al., 2023). To optimize long-term rewards, reinforcement learning, a subset of machine learning, enables sequential interactions between an agent and its surroundings (Zhang, 2023). Artificial diagnosis, precision medicine, treatment recommendations, and drug dosage optimization in conditions like acute kidney injury (AKI) or chronic kidney disease (CKD) complications are just a few of the tasks that machine learning has been used for in the medical field (Khezeli et al., 2023; Nambiar et al., 2023). Nevertheless, there are difficulties in applying RL to practical medical applications, particularly when it comes to managing retrospective datasets that are biased and guaranteeing patient safety (Adiga et al., 2023). Notwithstanding these difficulties, reinforcement learning (RL) holds great promise for the healthcare industry. It can automate treatment choices, improve patient care, and improve clinical outcomes by creating dependable simulation environments and incorporating clinical insights. Explainable AI (XAI) refers to a collection of approaches and strategies in machine learning (ML) and artificial intelligence (AI) that seek to explain, reveal, and allow humans to engage with the decision-making processes of AI systems. Its main goal is to create AI systems that can clearly explain their choices and forecasts so that users can comprehend how the model arrived at a given result (Saeed & Omlin, 2023). The three important concepts in XAI are feature importance, local explanations, and model-agnostic techniques. Feature importance describes how much of an influence or contribution each input feature has on the predictions made by the model. The goal of local explanations is to provide context for a machine learning model’s prediction for a particular instance or data point. It explains each prediction in detail so that users can comprehend the reasoning behind a particular choice. Model-agnostic techniques are methods that, regardless of the underlying algorithm, can be used to explain the predictions of different machine learning models (Parisineni & Pal, 2023). These methods are adaptable to different kinds of models and encourage interpretability and flexibility. Model-agnostic approaches 49 improve flexibility by releasing users from the constraints of a specific algorithm, enabling them to apply explanation techniques to a variety of model types. These XAI concepts collectively help machine learning models become more transparent and trustworthy by addressing issues with accountability and interpretability, particularly in critical applications. The popularly used XAI techniques are delineated below a) Local Interpretable Model-agnostic Explanations (LIME): LIME is a framework created to offer understandable justifications for machine learning models' predictions, especially when it comes to specific cases. LIME produces locally accurate explanations for particular cases by varying the input data and tracking how the model's predictions alter. LIME is a community-recognized algorithm that can be used to explain the predictions of any black-box classifier and black-box regressors trained on tabular data (Vinogradova & Myers, 2023). b) SHapley Additive exPlanations (SHAP): SHAP is a framework for analysing machine learning model output, and it helps to comprehend how each feature affects the model's predictions (Al-Najjar et al., 2023) by giving the means to link a model's prediction to each of its features. SHAP values assign contributions to each feature for a given prediction based on cooperative game theory and offer a fair means of allocating "credit" for the model's output (Palar et al., 2023). 2.2.5 Large Language Models in Healthcare Healthcare is one of the many industries where large language models (LLMs) like Generative Pre-trained Transformer 3 (GPT-3) have shown great promise (Peng et al., 2023; Bumgardner et al., 2023). There is both excitement and concern regarding the use of LLMs in healthcare settings because they can answer free-text queries without having received specialized training for the task (Thirunavukarasu et al., 2023). Recent advancements in AI and Deep Learning (DL) have led to the creation of LLMs, which use transformer-based architectures to provide cutting-edge results on a range of language tasks (Chacko & Chacko, 2023). Large volumes of medical text data can be processed and analyzed by these models, which have been trained 50 on enormous datasets and have natural language understanding capabilities. Some common types of large language models are: a) Language Representation Model: Language representation models (LRM) are the foundation of many NLP applications; they are made to comprehend and produce human language. These models include RoBERTa, BERT (Bidirectional Encoder Representations from Transformers), and GPT (Generative Pre-trained Transformer) models (Fajcik et al.,2020). These models can be optimized for particular tasks like text classification and language generation because they have already been pre-trained on large text corpora. b) Zero-Shot Model: A zero-shot model can perform tasks without specific training data; it can generate text for tasks it has never seen before or generalize and make predictions (Meng et al., 2022). GPT-3 is an example of a zero-shot model; it can translate languages, answer questions, and perform a variety of tasks with little to no fine-tuning. c) Multimodal Model: Although multimodal models, like OpenAI's CLIP, can understand and generate content across different modalities, LLMs were originally designed for text content (Carolan et al., 2024). On the other hand, multimodal models, like CLIP, can work with both text and image data. Their ability to associate text with images and vice versa makes them useful for tasks like image captioning and text-based image retrieval. d) Fine-Tuned Or Domain-Specific Models: Even though pre-trained language representation models are flexible, certain tasks or domains may not always yield the best results from them. To boost performance in specific domains, fine-tuned models have undergone additional training on domain-specific data (Bai et al., 2020). For instance, a GPT-3 model could be improved using medical data to develop a medical chatbot tailored to the industry or help with diagnosis. Some examples of predominant large language models are summarized below. a) Bidirectional Encoder Representations from Transformers (BERT): BERT is a pre-trained transformer-based model that can be used in natural language understanding tasks (Devlin et al., 2018). BERT is a deep 51 bidirectional transformer architecture that allows for the multilingual universal language representation of different languages, and contextual embeddings are obtained using pre-trained unlabelled Wikipedia and book corpus (Deepa, 2021). BERT is an advanced and more realistic technique because it recognizes that a document can belong to multiple classes at the same time and has demonstrated exceptional performance in eleven natural language understanding (NLU) tasks (Alaparthi & Mishra, 2020). b) Generative Pre-trained Transformer (GPT) Series: Models like GPT-2 and GPT-3 are part of OpenAI's GPT series. The purpose of these models is to generate natural language. Pre-training and fine-tuning are the two steps in the training process for GPT models. GPT models are trained on enormous volumes of text gathered from various sources, and these models are finetuned on smaller, task-specific datasets after pre-training to conform to particular domains. Fine-tuning preserves the general language understanding acquired during pre-training while enabling the model to specialize in the desired task and enhance its performance on domain-specific data (Kamnis, 2023). GPT-4 is a more sophisticated version of GPT-3 and GPT-3.5, its predecessors. It performs better than the earlier models in terms of creativity, context, and visual comprehension. Users can work together on screenplays, music, technical writing, and other projects with this LLM. In addition to text, images can be entered into GPT-4. Furthermore, GPT-4 is a multilingual model that can respond to thousands of queries in 26 languages, according to OpenAI (Akshay, 2024). c) Text-to-Text Transfer Transformer (T5): Google's T5 is a flexible LLM that was trained with a text-to-text framework (Gupta, 2024). Thanks to its ability to convert input and output formats into a text-to-text format, it can handle various language tasks. In the areas of machine translation, text summarization, text classification, and document generation, T5 has produced cutting-edge outcomes. It is incredibly versatile and effective for a wide range of language-related applications because of its capacity to manage various tasks within a single framework (Akshay, 2024). 52 d) eXtreme Language Understanding (XLNet): Researchers from Google and Carnegie Mellon University created XLNet to address some of the drawbacks of autoregressive models like GPT-3. It makes use of a permutation-based training strategy, which enables the model to take into account every possible word order in pre-training. This makes it possible for XLNet to detect bidirectional dependencies during inference without the need for autoregressive generation (Khan et al., 2020). Impressive results have been shown by XLNet in tasks like sentiment analysis, question and answer, and natural language inference. e) Turing-NLG: Microsoft's Turing-NLG is a potent LLM that specializes in producing conversational responses and to enhance its conversational skills, a vast dataset of exchanges has been utilized for training (Sahoo et al., 2024). Turing-NLG performs well in chatbot applications, offering conversational settings with interactive and contextually relevant responses. f) Pathways Language Model (PaLM): Google AI created a sizable language model called PaLM. The LLM is becoming one of the most potent AI language models since it can train on Google's enormous dataset (Chowdhery et al., 2023). It is a significant advancement in responsible AI and machine learning. Even though PaLM is still in development, it already can comprehend language, answer questions in natural language, and perform machine translation, code generation, summarization, and other creative tasks. PaLM was created with data security and privacy in mind as well. Data can be encrypted and shielded from unwanted access with its help. Because of this, it's perfect for delicate projects like developing safe eCommerce websites and platforms that handle sensitive user data. g) Large Language Model Meta AI(LLaMA): LlaMA is a brand-new, opensource large language model that is presently being worked on by Meta AI. It is intended to be a strong and adaptable LLM that can be applied to a range of tasks, such as reading comprehension, natural language understanding, and query resolution. LlaMA is the outcome of Meta's specific concentration on language learning models for use in teaching (Thirunavukarasu et al., 2023). 53 h) CoHere: CoHere is a large language model created by the same-named Canadian startup. The open-source LLM is proficient in handling a wide range of languages and accents because it was trained on a varied and inclusive dataset (Chen et al., 2023). Furthermore, Cohere's models are better suited for a variety of tasks because they were trained on a sizable and varied corpus of text. i) Gemini: Google AI developed Gemini, an LLM chatbot that was formerly known as BARD. It is trained using a sizable text and code dataset. This enables it to create text, translate between several languages, write code, produce a variety of content, and respond to inquiries with insightful information (Metz & Grant, 2023). Google Search provides access to realworld data for Gemini, one of the top multimodal large language models. This enables it to interpret and respond to a wider range of cues and questions. j) Falcon: The Technology Innovation Institute developed the open-source language model Falcon. It is now the best language model, surpassing Llama on the Hugging Face Open LLM Leaderboard. Trained on a higher-quality dataset that encompasses a vast array of text and code in various languages and dialects, Falcon is an autoregressive model (Kukreja et al., 2024). Additionally, it makes better predictions and processes data more effectively thanks to a more sophisticated architecture. Compared to the best NLP models, this new pre-trained model has learned with 40 billion fewer parameters. k) Claude v1: A sizable language model called Claude v1 was created by the US AI startup Anthropic. It is a flexible AI assistant made especially to make managing, creating, and optimizing websites easier. Claude v1's sophisticated natural language capabilities make it simple for anyone to create, manage, and expand a website without the need for complex technical knowledge. Compared to other LLMs, Claude has a more sophisticated architecture that enables it to process data more quickly and produce more accurate predictions (Turpin et al., 2024). l) Grok: The generative artificial intelligence chatbot Grok was created by xAI. Elon Musk created it as an initiative based on a large language model (LLM) 54 in direct response to ChatGPT's explosive growth, the development of which Musk co-founded with OpenAI (Ding et al., 2024). The chatbot is marketed as having direct access to Twitter (X) and "having a sense of humor". X Premium is required to access it, and beta testing is underway (OptimizeIAS, 2024). m) Med-PaLM: Med-PaLM is intended to deliver superior responses to medical queries. MedLM, a family of foundation models optimized for the healthcare sector, is powered by research models, including our second version, MedPaLM 2. MedLM is now accessible to Google Cloud users experimenting with various applications, ranging from simple jobs to intricate workflows. Med-PaLM leverages the power of Google's extensive language models, which have been assessed using consumer queries, medical research, and examinations in the medical field. The first version of Med-PaLM was preprinted in late 2022, and it was published in Nature in July 2023. This made it the first AI system to pass the U.S. Medical Licensing Examination (USMLE) style questions with a passing score of greater than 60%. According to panels of doctors and users, Med-PaLM also produces precise, beneficial long-form responses to consumer health queries. In March 2023, at Google Health's annual event, The Check Up, Med-PaLM 2 was unveiled. The first to answer questions similar to those on the USMLE at the level of a human expert was Med-PaLM 2. Doctors report significant improvements in the model's long-form responses to consumer medical queries. (Singhal et al., 2023) LLMs are effective tools in a variety of domains, such as healthcare, education, customer service etc and some of its advantages are summarized below (Abdullahi, 2024). a) Increased Efficiency: Due to their capacity for comprehending human language, LLMs are well-suited for handling tedious or repetitive tasks. LLMs are useful for tasks like content creation, writing code, and summarizing vast amounts of information because they can produce human-like text much faster than humans. 55 b) Enhanced Question-Answering Capabilities: An answer-generation machine is another way to characterize LLMs. Experts had to intervene to persuade users that generative AIs would not supplant Google as a search engine because LLMs are so adept at producing precise answers to user inquiries. c) Few-Shot or Zero-Shot Learning: With minimal training examples or no training at all, LLMs are capable of performing tasks and generalizing from available data to infer patterns and predict in novel domains. d) Transfer Learning: Professionals in a variety of industries benefit from LLMs because they can be customized for different tasks, allowing a model to be trained for one task and then used for another with little further training. Large Language Models also present challenges, some of which are summarized below (Abdullahi, 2024). These challenges show how using LLMs requires careful thought and mitigating measures. a) Performance depends on training data: The quality and representativeness of the training data are critical to the effectiveness and precision of LLMs. Since LLMs are only as good as the training data they use, models that are trained on skewed or poor-quality data will undoubtedly yield debatable results. This is a serious potential issue since it can have a negative impact, particularly in delicate fields like law, medicine, or finance where accuracy is crucial. b) Lack of common-sense reasoning: Large language models are remarkably good at language, but they frequently have trouble with common-sense logic. Humans possess common sense by nature; it's one of our innate instincts. However, common sense is not common among LLMs, as they can generate responses that lack context or are factually incorrect, producing outputs that are misleading or nonsensical. c) Ethical concerns: Concerns about potential abuse or malevolent applications arise when using LLMs. The possibility of producing offensive or dangerous content, deepfakes, or impersonations that could be utilized for manipulation or fraud exists. 56 2.2.6 Integration of XAI and LLMs in Medical Diagnosis There are various ways in which the combination of Explainable Artificial Intelligence (XAI) and Large Language Models (LLMs) could improve the diagnostic precision of diseases. There are also numerous benefits to integrating AI and LLMs into clinical pharmacology, but there are also important safety and ethical issues to be aware of (Rubinic et al., 2023). Here are some advantages of this integration: a) Improved Clinical Decision Support: XAI methods can offer clear-cut and understandable insights into how LLMs make decisions. It is simpler for clinicians to trust and depend on the AI system for decision support when they can comprehend how the model arrived at a specific diagnosis (Antoniadi et al., 2021). b) Enhanced Feature Interpretability: XAI techniques can draw attention to the significance of particular traits or terms during the diagnosis process (Pawar et al., 2020). Clinicians can focus on important aspects during patient evaluation for diseases by knowing which clinical symptoms or historical data were important to the model. c) Contextual Understanding: Contextual explanations are given by XAI for the model's predictions (Pawar et al., 2020). XAI can help explain why the model leaned towards a specific diagnosis in the context of tropical febrile diseases, where symptoms may overlap, taking into account the larger clinical context. d) Reduced Diagnostic Ambiguity: Model prediction uncertainty can be estimated with the aid of XAI techniques (Thuy & Benoit, 2023). Clinicians can be informed when the model is unsure or presents with unclear symptoms. This information may lead to additional diagnostic testing or the search for more information. e) Facilitating Clinician-Model Collaboration: Healthcare professionals will be able to easily understand the explanations generated by XAI (Gentile et al., 2021). This makes it easier for doctors and the AI model to collaborate, which 57 improves the effectiveness of incorporating AI-driven insights into diagnosis procedures. f) Error Analysis and Model Improvement: XAI can assist in localizing errors by pinpointing particular situations in which the model might have interpreted data incorrectly (Widyasari et al., 2022). For ongoing model improvement and refinement, this feedback is helpful. g) Transparent Treatment Recommendations: The reasoning behind the model's suggested treatments can be explained by XAI (Pawar et al., 2020). In the case of tropical febrile diseases, this transparency can help physicians comprehend and have confidence in AI-generated treatment plans. h) Education and Training: Healthcare workers may find XAI to be a useful teaching tool (Fiok et al., 2022). It can aid in educating medical professionals about the subtleties of tropical febrile illnesses and enhance their comprehension of the diagnostic reasoning behind the model. i) Compliance with Ethical Standards: Ethical issues about accountability and transparency in AI systems are addressed by the integration of XAI (Arrieta et al., 2020). Clear models support responsible AI applications in healthcare and are consistent with moral principles. j) Patient-Clinician Communication: Better communication between patients and clinicians can be fostered by sharing explanations generated by XAI with patients (Lötsch et al., 2021). When patients understand the rationale behind their diagnosis, their confidence in the medical system grows. There is great potential to improve decision-making processes' interpretability and transparency by integrating XAI and LLMs in medical diagnosis (Zhou et al., 2023). Healthcare professionals can gain valuable insights into the model's decision-making process by combining LLMs' understanding and processing capabilities with XAI techniques. This approach can help improve the trust and acceptance of AI-driven recommendations by providing valuable explanations for diagnostic outcomes (Han et al., 2023). Furthermore, the creation of domain-specific LLMs specifically designed for medical applications, like GatorTron for clinical text processing (Gao et al., 2023), demonstrates the potential of specialized models to improve the 64 2.2.9.3 Other Stakeholders Acceptance The extent to which AI systems abide by current healthcare regulations affects stakeholders like legislators and regulators. Solutions that comply with legal and ethical requirements are more likely to be accepted. a) Economic Viability: Payers and healthcare organizations evaluate whether using AI-assisted diagnosis is economically feasible. Acceptance rates are higher for solutions that provide cost-effectiveness, efficiency gains, sustainability, and better resource utilization (Al-Emran & Griffy-Brown, 2023). b) Interoperability: Integrating AI systems with other health IT systems and the current healthcare infrastructure is essential (Carobene et al., 2023). The stakeholders in healthcare ecosystem management are more accepting when there is interoperability. c) Evidence of Impact: Stakeholders seek proof of how AI affects patient outcomes, healthcare expenses, and system performance as a whole. Acceptance is influenced by lucid examples of successful outcomes (Kelly et al., 2023). d) Ethical Considerations: The ethical aspects of AI use, such as accountability, transparency, and fairness, are crucial (Arrieta et al., 2020). Various stakeholders are more likely to support solutions that address these ethical issues. e) Public Perception: The acceptance of these technologies in healthcare is influenced by public perception and societal attitudes toward artificial intelligence (Kelly et al., 2023). Decisions and policies made by stakeholders can be influenced by a positive public image. A complex interplay of factors about accuracy, transparency, usability, ethics, and societal attitudes determines user acceptance of AI-assisted diagnosis in the healthcare industry. For AI technologies used in medical diagnosis to be successfully integrated and widely accepted, a thorough understanding of these factors is necessary 65 2.2.10 Human-Machine Collaboration in Diagnosis AI systems and medical professionals working together to make decisions have the potential to completely transform healthcare delivery, enhance patient outcomes, and boost overall medical process efficiency. Table 2.1 presents some salient aspects that underscore the possibility of cooperation. Table 2.1. Potential for collaborative decision-making Aspects AI Human Complementary Expertise (Zhang et al., 2022) Rapid analysis of large datasets, pattern recognition, and evidence-based decision support. A comprehension of the patient's background, clinical experience, intuition, and empathy. Improved Diagnostic Accuracy (Reverberi et al., 2022) Healthcare professionals can gain insights from AI systems' analysis of patient data, which includes test results, imaging scans, and medical records. Healthcare professionals make diagnostic decisions by considering patient history, AIgenerated insights, and their clinical expertise. Efficient Workflow Integration (Gu et al., 2023) Facilitating the automation of repetitive tasks, freeing up healthcare professionals to concentrate on more intricate facets of patient care. Supplying a human touch in patient interactions and making sure AI recommendations are in line with the overall patient care strategy. Personalized Treatment Plans (Ali, 2023). Identifying individualized treatment options through data analysis of large datasets, taking into account patient traits, genetics, and response to previous interventions. Treating patients with a focus on their total well-being by considering their preferences, values, and unique situations when designing treatment regimens. Enhanced Predictive Analytics (Lee et al., 2023) Using patient data from the past to forecast the course of a disease, spot possible complications, and issue early alerts. Verifying AI forecasts against the patient's present state, modifying treatment regimens, and making certain that human interaction is preserved with patients. Patient-Centric DecisionMaking (Cabitza et al., 2023) Making recommendations based on evidence that puts patient outcomes and safety first. Making decisions that keep the patient at the centre of care. Continuous Learning and Adaptation (Zhao et al., 2022). Accurate diagnosis can be improved by learning from realworld data and adjusting to advances in medical knowledge. Ensuring that the AI system remains in line with current best practices. Decision Support and Clinical Guidelines (Chen et al., 2022) Providing quick access to the recent clinical recommendations, Contextualizing AI recommendations by using 66 evidence-based treatments, and pertinent medical literature. clinical judgment and the particulars of each patient's case. Improved Resource Allocation (Jain et al., 2023). Finding areas for costeffectiveness and efficiency improvements to help with resource allocation optimization. Making strategic choices based on a deeper comprehension of the objectives of the organization and the needs of the patients. Cooperative decision-making between AI systems and medical professionals can lead to a mutually beneficial partnership in which the advantages of both parties are utilized to enhance patient care, streamline processes, and increase medical knowledge. To fully realize the benefits of this partnership, there needs to be constant communication, mutual trust, and a dedication to a patient-centric methodology. 2.3 Empirical Studies Machine learning is frequently used to analyze medical images to identify anomalies, diagnose diseases like tumors or fractures, evaluate genetic information, find mutations linked to disease, and customize treatment regimens based on unique genetic profiles. Massive clinical data sets, such as electronic health records (EHRs), can also be analyzed by ML algorithms to find trends and forecast the likelihood of developing a disease. ML and XAI have the potential to revolutionize the healthcare industry by providing cutting-edge approaches to patient care, treatment optimization, and diagnosis. Due to the importance of medical decisions and the requirement for interpretability in terms of ethics, XAI has garnered much attention in recent years for integration into clinical decision support systems (Umerenkov et al., 2023). Le et al. (2021) and Moncada-Torres et al. (2021) utilized XAI to predict cancer, while Duell et al. (2021) analyzed electronic health records, and Abeyagunasekera et al. (2021) applied XAI techniques to medical imaging. Alsinglawi et al. (2022) demonstrated the use of XAI in predicting the length of hospital stay for lung cancer patients, and Du et al. (2022) focused on predicting gestational diabetes mellitus. Similarly, Severn et al. (2022) employed XAI for analyses on medical images while Attai et al. (2023) applied XAI techniques to model comorbidities in pregnant women and children. 67 Suh et al. (2020) developed and validated an explainable artificial intelligence-based prostate biopsy decision-supporting tool. The risk calculator was developed and validated by the study using data from 3791 patients. The data was first split into sets for development and validation. Using five-fold cross-validation and hyperparameter tuning after feature selection in the development set, an extreme gradient-boosting algorithm was implemented on the development calculator. The Shapley value was used to calculate the model feature importance. For every calculator validation set, the receiver operating characteristic curve's area under the curve (AUC) was examined. PCa and csPCa were identified in about 1216 (32.7%) and 562 (14.8%) of the patients, respectively. While the data of 948 patients served as a test set, the data of 2843 patients were used for development. Using selection operator regression and the least absolute shrinkage, we chose the variables for each PCa and csPCa risk calculation. In contrast to the csPCa model, which had an AUC of 0.945 (95% CI 0.927–0.963), the final PCa model had an AUC of 0.869 (95% confidence interval [CI] 0.844–0.893). Important variables in the PCa model were discovered to be the prostate-specific antigen (PSA) level, free PSA level, age, prostate volume (both the transitional zone and total), hypoechoic lesions on ultrasonography, and testosterone level. The number of prior biopsies had a negative correlation with PCa risk but no correlation with the risk of csPCa. Before prostate biopsy, the study effectively created and validated a decision-supporting tool that uses XAI to calculate the probability of PCa and csPCa. Peng et al. (2021) proposed an XAI framework to provide both local and global interpretation for auxiliary hepatitis diagnoses while maintaining high prediction accuracy. First, the framework's viability was evaluated using a public hepatitis classification benchmark from UCI. Afterward, both the transparent and black-box machine learning models were used to predict the progression of hepatitis. Clear models like k-nearest neighbour (KNN), decision tree (DT), and logistic regression (LR) were chosen. Random forests (RF), support vector machines (SVM), and eXtreme Gradient Boosting (XGBoost) are examples of black-box models that were chosen. Lastly, the model interpretation of liver disease was enhanced by the use of Partial Dependence Plots (PDP), Local Interpretable Model-agnostic Explanations 68 (LIME), and SHapley Additive exPlanations. The results of the experiments indicated that the complex models perform better than the simple ones. Out of all the models, the developed RF had the highest accuracy (91.9%). The suggested framework, which combines local and global interpretable methods, enhanced the transparency of complex models and provided insight into their conclusions. This helped to improve the prognosis of hepatitis patients and guide treatment decisions. Furthermore, the suggested framework might help clinical data scientists create a more suitable CAD structure. El-Sappagh et al. (2021) developed a model for the diagnosis and progression detection of AD that is both accurate and comprehensible. This model gives doctors precise recommendations along with a series of justifications for each choice. The model specifically incorporates 11 modalities of 1048 subjects, 294 cognitively normal, 254 stable mild cognitive impairment (MCI), 232 progressive MCI, and 268 AD, from the Alzheimer's Disease Neuroimaging Initiative (ADNI) real-world dataset. With random forest as the classifier algorithm, the model is two layers deep. For early AD patient diagnosis, the model performs a multi-class classification in the first layer. To identify potential MCI-to-AD progression within three years of a baseline diagnosis, the model uses binary classification in its second layer. Key markers chosen from a wide range of biological and clinical metrics optimize the model's performance. In terms of explainability, we use the SHapley Additive exPlanations feature attribution framework to provide global and instance-based explanations of the RF classifier for each layer. Furthermore, to provide supplementary explanations for each RF decision in every layer, we implement 22 explainers based on fuzzy rule-based systems and decision trees. To further aid doctors in understanding the forecasts, these explanations are presented in plain language. In the first layer, the designed model achieves 93.95% cross-validation accuracy and an F1-score of 93.94%; in the second layer, it achieves 87.08% crossvalidation accuracy and an F1-score of 87.09%. Thanks to the explanations provided, which are generally consistent with each other and with the AD medical literature, the resulting system is not only accurate but also trustworthy, accountable, and medically applicable. By offering thorough insights into how various modalities 69 affect the risk of AD, the proposed system can contribute to improving clinical understanding of the disease's diagnosis and progression processes. Davagdorj et al. (2021) developed a deep neural network framework based on Deep Shapley Additive Explanations (DeepSHAP) and featuring a feature selection technique for predicting and explaining non-communicable diseases (NCDs) in the US population. To create a precise and understandable decision support system, the DeepSHAP-based DNN framework was furnished with an elastic net (EN) feature selection technique in addition to three distinct sets of important features generated from the NHANES dataset. There are three parts to the suggested framework: Third, the DeepSHAP approach provides two types of model explanation. Firstly, representative features are obtained using the elastic net-based embedded feature selection technique. Secondly, a deep neural network classifier is adjusted with hyperparameters and utilised to train the model using the chosen feature subset. Here, (I) the population-based risk factors influencing the model's prediction are explained; (II) a human-centred explanation of a single instance is attempted. According to the experimental results, the suggested model performs better than several cutting-edge models. Additionally, by offering broad insights into variations in disease risk on a local and global scale, the suggested model enhances medical understanding of NCD diagnosis. As a result, the explainable deep learning framework built on DeepSHAP not only benefits medical decision support systems but also meets practical needs in other fields. Bogdanovic et al. (2022) used an explainable machine learning approach to present a comprehensive understanding of Alzheimer's disease. This study provided a comprehensive analysis of a sizable data set that included lifestyle, cognitive, and medical assessments from over 12,000 people. Given the results, the validity of several established hypotheses has been called into question. The research places a strong emphasis on the significance of using an appropriate experimental design. To properly preprocess the data set, a series of techniques for handling missing data, redundancy, data imbalance, and correlation analysis have been applied. As a result, the XGBoost model has been trained and evaluated with particular attention to the hyperparameter tuning. The Shapley values obtained from the SHAP method were 70 used to explain the model. XGBoost is regarded as being very competitive among those that have been published in the literature because it generated an F1-score of 0.84. But this accomplishment wasn't the paper's primary contribution. The purpose of this study was to evaluate the intelligent model's interpretability on a global and local level and draw insightful conclusions about the hypothesis that had been put forth. These techniques produced a single scheme that shows the values of each feature, whose significance has been verified using Shapley values, as having a positive or negative influence. This plan may be viewed as an extra resource for medical professionals and other specialists who are interested in precisely diagnosing Alzheimer's disease in its early stages. All of the preexisting theories were challenged by the conclusions drawn from the intelligent model's data-driven interpretability. The importance of an explainable machine learning approach that cracks open the mystery and makes the relationships between features and diagnoses visible was demonstrated by this study. Vyas et al. (2022) used interpretable machine learning techniques on structured clinical records to determine the presence and severity of dementia. The techniques discussed in this paper are intended to address two dementia-related issues: (a) basic diagnosis, which is the process of determining whether a person has dementia, and (b) severity diagnosis, which is the process of forecasting both the presence and severity of dementia in an individual. The study used machine learning models based on random forests and decision trees to analyze structured clinical data from an elderly population cohort. These two tasks were formulated as classification problems. To ensure that curation decisions are meaningful, the study implemented a hybrid data curation strategy, involving a dementia expert and machine learning algorithms that categorized individual episodes into a particular dementia class that was used. Moreover, decision trees were used to improve the explainability of choices made by prediction models, enabling medical professionals to determine which patient characteristics are most important and at what threshold level to classify a patient as having dementia. The findings demonstrated that demographic characteristics and baseline math or cognitive tests can accurately predict dementia and its severity. Specifically, our prediction models have achieved an average f1- 71 score of 0.93 for problem (a) and 0.81 for problem (b). Furthermore, the decision trees generated for the two problems strengthen the prediction models' interpretability. This study demonstrated that by analysing different electronic medical record features and cognitive tests from the episodes of the elderly population, it is possible to accurately estimate the presence and severity of dementia disease. In addition, a collection of decision rules could serve as the foundation for a successful patient classification. Predictive features that are pertinent to clinical and screening tests serve as accurate indicators without requiring the computation of scores from widely used cognitive tests like the MMSE and CAMCOG. Not only is it able to recognize significant characteristics, but it can also recognize the rationale behind the classification. Therefore, the predictive ability of machine learning models over carefully selected clinical data is demonstrated, opening the door to a more precise dementia diagnosis. Weng et al. (2022) created a framework that uses an explainable machine learning (ML) model to effectively distinguish between these two diseases. Intestinal data were gathered from Central South University's Second Xiangya Hospital; 160 patients with CD and 40 patients with ITB were included in the investigation. Every patient had an active medical condition. The clinical diagnosis of CD and ITB, as well as the European diagnostic guidelines, were combined with all cases. Nine variables are extracted in total after feature selection: intestinal dilatation, comb sign, bloody stool, PPD, knot, intestinal surgery, ESAT-6, and CFP-10. Also, the study contrasted the traditional statistical methods and machine learning's predictive performance. For the first time, this work also offered insights into the ML model's result using the SHAP method. Models are trained and validated using data from a cohort of 200 patients (CD=160, ITB=40). Findings showed that the XGBoost algorithm performs better than other classifiers in terms of the Matthews correlation coefficient (MCC), sensitivity, specificity, area under the receiver operating characteristic curve (AUC), and precision; these values are 0.891, 0.813, 0.969, 0.867, and 0.801, respectively. More importantly, the SHAP method provides an effective explanation for the prediction outcomes of XGBoost. The suggested 72 framework demonstrated how well interpretable machine learning can separate CD from ITB and provide both a general explanation and a patient-specific explanation. Kibria et al. (2022) proposed an ensemble method that combines an explainable AI and a soft voting classifier to predict the onset of diabetes mellitus. The study employed the Pima Indian diabetes dataset, which comprises 768 cases in total, 268 of which are diabetic, and 500 of which are non-diabetic but have multiple diabetic characteristics. To diagnose diabetes, an ensemble classifier was employed in conjunction with six machine learning algorithms: AdaBoost, XGBoost, logistic regression (LR), support vector machine (SVM), random forest (RF), artificial neural network (ANN), and logistic regression (LR). Shapley additive explanations (SHAP) were used to generate both global and local explanations for each machine learning model. These explanations were displayed in various graph types to aid medical professionals in comprehending the model predictions. With an F1 score of 89% and five-fold cross-validation (CV), the developed weighted ensemble model had a balanced accuracy of 90%. The dataset's classes were balanced using the synthetic minority oversampling technique (SMOTETomek), and the missing values were imputed using the median values. The suggested method can aid in the critical intervention at the earliest stages of the disease and enhance clinical comprehension of a diabetes diagnosis. Islam et al. (2022) proposed an explainable AI model for stroke prediction using EEG data. This study aimed to predict acute stroke in active states by using machine learning (ML) models to classify the ischemic stroke group and the healthy control group. Additionally, the XAI tools Eli5 and LIME were used to identify the key features that go into stroke prediction models and to explain the model's behaviour. In this study, 75 healthy adults without a history of neurological disorders were compared to 48 patients who had been admitted to a hospital due to acute ischemic stroke. EEG was recorded using frontal, central, temporal, and occipital cortical electrodes within three months of the onset of symptoms associated with an ischemic stroke. While walking, working, and reading tasks were being performed, EEG data were being gathered. The Adaptive Gradient Boosting models in the ML approach's results demonstrated about 80% accuracy in classifying the stroke group and the 73 control group. The stroke prediction model's behaviour was explained by Eli5 and LIME, which were also used to interpret the model locally around the prediction. The spectral delta and theta features were highlighted by the Eli5 and LIME interpretable models as local contributors to stroke prediction. The stroke-prediction XAI model was anticipated to aid in post-stroke treatment and recovery, as well as assist healthcare professionals in making more explicable diagnostic decisions, based on the findings of this explainable AI research. Kasani et al. (2023) used explainable machine learning techniques to assess the nutritional status and classify clinical depression. Using publicly available health data from the Korean National Health and Nutrition Examination Survey, this study sought to ascertain whether machine learning-based decision support methods could detect the presence of depression. For explanatory analysis across datasets, two exploration techniques were used: Pearson correlation and uniform manifold approximation and projection. The models were refined using a grid search optimisation with cross-validation in order to classify depression with the highest accuracy. The classifier performances were compared using a number of performance metrics, such as accuracy, precision, recall, F1 score, confusion matrix, areas under the receiver operating characteristic curve and precision-recall curve, and calibration plot. The study also examined the significance of the following features: local interpretable explanations that are based on Shapley additive explanation and model-agnostic explanations, and visualised interpretation that employs ELI5, partial dependence plots, and explanations for prediction at the individual and population levels. The random forest model in the original dataset yielded an accuracy of 86.18% and an area under the curve of 84.96%, while the quantile-based dataset yielded an accuracy of 86.02% and an area under the curve of 85.34% for the XGBoost algorithm. These results indicate that the best model performed well. An additional observation of the relative changes in feature values was provided by the explainable results, which allowed for the identification of the significance of emergent depression risks. This study presented methods for both local and global interpretations to demonstrate the interpretability of ML models. ELI5 and PDP for worldwide interpretability and LIME and SHAP for local interpretation will be used 80 Centres for Disease Control and Prevention website. PPA was able to predict mortality and chronic illness without regard to age. It took only twenty-six variables to predict PPA. The study implemented a precise quantitative associated metric for each variable explaining physiological deviations from age-specific normative data using SHapley Additive exPlanations. Glycated haemoglobin (HbA1c) was one of the variables that showed a significant relative weight in the estimation of PPA. Lastly, distinct ageing trajectories were revealed by clustering profiles of contextualised explanations that were identical, providing opportunities for targeted clinical follow-up. According to these findings, PPA was a reliable, quantifiable, and understandable machine learning-based indicator of individualised health status. Furthermore, the method offered a comprehensive framework that can be applied to various datasets or variables, enabling accurate physiological age estimation. Moreno-Sánchez (2023a) introduced an explainability-driven methodology to determine the optimal heart failure (HF) survival prediction model by weighing predictability and performance. The study uses a dataset of 299 patients with HF to present an explainability analysis and evaluation of two HF survival prediction models. First, the model applies survival analysis, using time and death events as target features; second, it treats the problem as a classification task to predict death. The model made use of an optimisation data workflow pipeline that can choose the best machine learning algorithm and the most advantageous feature set. Additionally, a variety of post hoc methods have been applied to the model's explainability analysis. With a c-index of 0.714 and a balanced accuracy of 0.74 (std 0.03) for the survival analysis and classification approach, respectively, the Survival Gradient Boosting model and Random Forest were the most balanced explainable prediction models. With the addition of "diabetes" for the survival analysis model, the SCI-XAI selected features for the two models like the first, selecting "serum_creatinine," "ejection_fraction," and "sex." Additionally, by ranking "serum_creatinine" and "ejection_fraction" as the most relevant features for the anticipated result, respectively, the application of post hoc XAI techniques also validated common findings from both approaches. The explainable prediction models for heart failure survival that were provided in this study would increase the uptake of clinical 81 prediction models by giving physicians a better understanding of the rationale behind the typically "black-box" AI clinical solutions, enabling them to make more logical and informed decisions. Sharma et al. (2023) proposed a computerized, automated method for creating a machine learning model with explainable artificial intelligence capabilities. They classified CAP A and B phases and phases A sub-phases (A1, A2, A3) using waveletbased Hjorth parameters. To shed light on the model, the study employs feature ranking based on SHAP. The MIT-BIH database's CAP sleep database (CAPSD) provided the experimental dataset used in this investigation. The Sleep Disorders Centre of the Ospedale Maggiore in Parma, Italy provided the 108 people's sleep recordings for the database, which was open to the public. 16 healthy people and 92 patients with a range of sleep disorders were included in these recordings. Of these, 9 patients had insomnia, 5 had narcolepsy, 40 had nocturnal frontal lobe epilepsy (NFLE), 10 had periodic leg movement (PLM), 22 had REM behaviour disorder (RBD), and 4 had sleep-disordered breathing (SBD). Single-channel standardised EEG recordings from patients with nocturnal frontal lobe epilepsy (NFLE), narcolepsy, periodic leg movement disorder (PLM), rapid eye movement behaviour disorder (RBD), and insomnia were used to develop the model. Using ensemble bagged trees (EbagT) and k-nearest neighbours (KNN) classifiers yields the best results. The proposed model successfully classified phases A and B with an average accuracy of 91.6% for healthy subjects and 94.33%, 86.3%, 88.68%, 84.43%, and 88.5% for narcolepsy, RBD, PLM, NFLE, and insomnia subjects, respectively. For the A subphases (A1, A2, A3), the model classified subjects with an average accuracy of 92.85% for healthy subjects and 93.9%, 84.9%, 88.0%, 80.92%, and 89.41% for narcolepsy, RBD, PLM, NFLE, and insomnia subjects, respectively. Using the microstructure of sleep, the suggested method might assist sleep specialists in automatically assessing an individual's quality of sleep. Kırboğa et al. (2023) used explainable white box algorithms to diagnose troponin levels in the COVID-19 process and arrive at a lucid explanation. Troponin data was interpreted in the COVID-19 process using SHAP algorithms, utilizing pandemic data from Erzurum Training and Research Hospital (decision number: 2022/13-145). 82 There were five machine learning algorithms created. AUC (Area Under the Curve) values, recall, F1-score, test accuracies, training, and precision were used to assess the model's performance. Shapley values were used to estimate the importance of each feature by using the highly accurate SHApley Additive exPlanations method on the model. Under the name CVD22, the model made with Streamlit v.3.9 was incorporated into the interface. With values of 1.0, 0.83, 0.86, 0.83, 0.80, and 0.91 in train and test accuracy, precision, F1-score, recall, and AUC values, respectively, the best machine learning model out of the five developed using pandemic data was chosen. Following feature selection and the application of SHAP algorithms to the XGBoost model, the features with the highest significance over the model estimation were found to be DDimer mean, mortality, CKMB (creatine kinase myocardial band), and glucose. With the help of extensive historical datasets and recent developments in explainable artificial intelligence models, it is now possible to successfully predict the future. Therefore, CVD22 can be used as a guide to assist authorities or medical professionals in making timely decisions during the ongoing pandemic. Mridha et al. (2023) presented a machine learning-based automated stroke prediction system. Two different explainable approaches, called SHAP and LIME, were applied to shed light on the black-box machine learning models. Especially in the medical field, well-proven and trustworthy methods for elucidating model decision-making are SHAP and LIME. The study's stroke prediction dataset, which was gathered from Kaggle, comprises 5110 rows and 12 columns. There were 4861 rows with a stroke value of zero and only 249 rows with a value of one, indicating an imbalance in the dataset. The SMOTE method was used to preprocess the data and balance it to improve accuracy. The experiment's results showed that more complex models performed better than simpler ones, with the best model obtaining an accuracy of nearly 91% and the other models achieving an accuracy of 83–91%. The suggested framework could help standardize complex models and gain insight into their decision-making, which can improve stroke care and treatment, and also include global and local explainable methodologies. 83 Moreno-Sánchez (2023b) reported on the creation and assessment of an explainable chronic kidney disease (CKD) prediction model that offers details on how various clinical characteristics of patients influence the early detection of CKD. The dataset, which included 400 patients with some missing values in their features, was gathered from the Apollo Hospital in Karaikudi, India, for about two months in 2015. Ten nominal, three ordinal, eleven numerical, and one target feature (notCKD/CKD) make up each dataset instance. The model was created with an optimisation framework that strikes a balance between explainability and classification accuracy. The primary contribution of the paper is its explicable, data-driven methodology, which provides quantitative insights into the role played by specific clinical features in the early detection of chronic kidney disease (CKD). Consequently, the best explainable prediction model uses three features (specific gravity, hypertension, and haemoglobin) to implement an extreme gradient boosting classifier. It achieves accuracy of 99.2% (standard deviation 0.8) and 97.5% with new, unseen data and a 5-fold cross-validation, respectively. Furthermore, haemoglobin is the most significant factor influencing the prediction, followed by specific gravity and hypertension, according to an explainability analysis. Due to the limited number of features chosen, early CKD diagnosis was now less expensive, suggesting a viable treatment option for developing nations. Li et al. (2023) created a machine learning model that links the identification of CHD and exposure to heavy metals effectively and understandably. The US National Health and Nutrition Examination Survey (US NHANES, 2003–2018) provided the datasets used to examine the relationships between heavy metals and CHD. To detect CHD based on exposure to heavy metals, five machine learning models were created. In addition, the models' strength was evaluated using eleven discriminating characteristics. The model that performed the best was chosen to be identified. Lastly, the features were interpreted using the SHapley Additive exPlanations tool to visualize the decision-making ability of the chosen model. 12,554 people in total met the eligibility requirements for this study. The optimal random forest classifier (RF) for identifying CHD was selected using data from 13 heavy metals (AUC: 0.827; 95%CI: 0.777–0.877; accuracy: 95.9%). The results of the SHAP analysis showed 84 that the model was positively influenced by cesium (1.62), thallium (1.17), antimony (1.63), dimethylarsonic acid (0.91), barium (0.76), arsenous acid (0.79), total arsenic (0.01) in urine, lead (3.58) and cadmium (4.66) in blood, and negatively influenced by cobalt (−0.15), cadmium (−2.93), and uranium (−0.13) in urine. Among US NHANES 2003–2018 participants, the RF model demonstrated efficiency, accuracy, and robustness in detecting a correlation between exposure to heavy metals and CHD. Positive correlations with CHD have been found for cesium, thallium, antimony, dimethylarsonic acid, barium, arsenous acid, and total arsenic in urine; negative correlations have been found for cobalt, cadmium, and uranium in urine. Rajab et al. (2023) presented a transparent approach to malaria diagnosis by offering meaningful interpretations of severe malaria predictions made by machine learning models, using SHAP and LIME. In the study, several models were used, including AdaBoost, Naive Bayes, Explainable Boosting Machines (EBMs), Decision Trees, Logistic Regression, K-means, K-Nearest Neighbour, Support Vector Machine, Naive Gradient Boosting, Random Forest, and Naive Bayes. The study's findings demonstrated that Explainable Boosting Machines and Random Forest had the highest accuracy, at 84%. Additionally, EBM offered a useful clinical comprehension of the characteristics that lead to accurate prediction. After using GridSearchCV to improve prediction accuracy, the LR's accuracy was 81%. Moreover, XGBoost was utilized to estimate the model's skill on fresh data using K-fold validation. XAI improved the interpretations by identifying characteristics that lead to severe malaria. By using these methods, medical practitioners can make more informed decisions and the accuracy of severe malaria predictions can be greatly increased. Ahmad et al. (2024) presented a novel framework that includes Explainable Artificial Intelligence (XAI), a Bagged Tree-based classifier (BTBC), and coefficient and distance correlation feature selection algorithms for effective epileptic seizure detection. The discrete wavelet transform (DWT) is utilised to break down the EEG signals and extract different eigenvalue features of the statistical time domain (STD) as linear and Fractal dimension-based non-linear (FD-NL). The Butterworth filter is used to remove various artifacts. Correlation coefficients with P-value and distance correlation analysis are then used to identify the best features. The Bagged Tree- 85 based classifier (BTBC) then makes use of these features. Among machine learning models using popular Bonn and UCI-EEG benchmark datasets, the proposed model outperforms the others in terms of mitigating overfitting issues and improves the average accuracy by 2% using (CD, E), (AB, CD, E), and (A, B) experimental types. Ultimately, the suggested model's decision-making process was interpreted and explained using SHAP. The findings demonstrate how well the framework can classify ES, which will help patients with brain dysfunctions receive better diagnoses. Alamatsaz et al. (2024) proposed a lightweight hybrid CNN-LSTM explainable model for ECG-based arrhythmia detection The most common and standard diagnostic device for tracking and assessing cardiac electrical signals is the electrocardiogram (ECG). Numerous conditions can affect the human heart, including cardiac arrhythmias. An arrhythmia is an irregular heart rhythm that can be identified with ECG recordings. Severe cases of arrhythmia can result in stroke. Over the past few decades, there has been a lot of interest in the computerized and automated classification and identification of these abnormal heart signals due to the critical nature of early cardiac arrhythmia detection. To detect eight distinct cardiac arrhythmias with high accuracy and normal rhythms, the study used a light Deep Learning approach. The ECG signals were preprocessed using baseline wander removal and resampling methods to use DL techniques. An 11-layer network that combined long short-term memory and convolutional neural networks was used to perform the classification. ECG signals are selected from the two physionet databases, the long-term AF database and the MIT-BIH arrhythmia database, to assess the suggested method. Compared to the majority of state-of-the-art techniques, the CNN and LSTM combination used in the proposed DL framework produced more encouraging results. The mean diagnostic accuracy achieved by the suggested method is 98.24%. To better understand how our model makes predictions, a trained model for arrhythmia classification using a variety of ECG signals was developed and tested using SHAP, the most widely used XAI technique. According to the findings, the characteristics (ECG samples) that have made the biggest contributions to predictions align with the choices made by medical professionals. As a result, the 86 adoption of interpretable models raises clinicians' confidence in AI and reduces the amount of cardiovascular disease misdiagnoses. Wani et al. (2024) proposed a hybrid model for the detection of lung cancer based on clinical data to assist healthcare professionals in making better decisions, explainable AI is used. To detect lung cancer and provide explanations for the predictions, this study presents "DeepXplainer," a novel interpretable hybrid deep learning-based technique. This method is predicated on XGBoost and a convolutional neural network. After "DeepXplainer" has automatically learned the features of the input using its numerous convolutional layers, XGBoost is used for class label prediction. SHAP is an explainable artificial intelligence method used to provide explanations or explainability of the predictions. This approach was used to process the "Survey Lung Cancer" dataset, which is publicly available. In terms of accuracy, sensitivity, F1-score, and other metrics, the suggested approach performed better than the current ones. The suggested approach produced results with an F1-score of 98.08, a sensitivity of 98.71%, and an accuracy of 97.43%. Following the model's exceptionally accurate predictions, each prediction is explicated through the local and global application of explainable artificial intelligence techniques. Numerous metrics have been used to assess the proposed "DeepXplainer," and the findings show that it performs better than the existing benchmarks. By offering explanations for the forecasts, the suggested method might make it easier for medical professionals to identify and treat lung cancer patients. LLMs have enormous potential in medicine; they can be used for anything from enhancing clinical decision-making to increasing diagnostic accuracy (Karabacak & Margetis, 2023). LLMs can extract important information from electronic health records, discharge summaries, and other medical texts. Efficient data retrieval can be facilitated by the ability of large language models to identify and extract critical clinical information, including medication histories, treatment plans, patient demographics, and medical histories. In addition to integrating patient data from multiple sources to help diagnose complex cases by taking a wider view of information, LLMs can help analyze symptoms as reported by patients, giving medical professionals more context for diagnostic decision-making (Umerenkov et 87 al., 2023). LLMs can interpret and summarize laboratory results, diagnostic imaging reports, and pathology reports (Thirunavukarasu et al., 2023) in addition to helping physicians comprehend complex findings and can also help create treatment recommendations based on the most recent clinical guidelines and evidence. Social media and other textual data can be analyzed by LLMs to track public opinion (Törnberg et al., 2023) and spot possible outbreaks or health issues and by searching textual data for patterns suggestive of new health risks, LLMs can aid in the early detection of disease outbreaks. Agbavor and Liang (2022) suggested using large language models to predict dementia from spontaneous speech. The study produced text embedding, a vector representation of the speech transcription that captures the semantic meaning of the input, by utilizing the extensive semantic knowledge encoded in the GPT-3 model. We show that using only speech data, the text embedding can be used to (1) reliably distinguish AD patients from healthy controls, and (2) infer the subject's cognitive testing score. According to the study, text embedding performs on par with existing fine-tuned models and significantly outperforms the traditional acoustic featurebased method. The findings indicated that GPT-3-based text embedding is a workable method for assessing AD straight from speech and may enhance dementia early diagnosis. Almazyad et al. (2023) introduced a novel way to use ChatGPT-4 to improve expert panel discussions during a medical conference. The study examined how well ChatGPT-4 performed in optimizing and summarising the recommendations made by the medical conference panel at the inaugural Pan-Arab Paediatric Palliative Critical Care Hybrid Conference, which took place in Riyadh, Saudi Arabia. First, the AI model optimized scenarios to encourage in-depth conversations; then, the model identified, summarised and contrasted important themes from the panel and audience discussions. This is how ChatGPT-4 was integrated into the discussions. Based on a summary of important themes like good communication, teamwork, patientand family-centered care, trust, and ethical considerations, the findings indicate that ChatGPT-4 successfully supported complex do-not-resuscitate (DNR) conflict resolution. Incorporating ChatGPT-4 into panel discussions about paediatric 88 palliative care has shown potential advantages for improving medical professionals' critical thinking skills. The study recommended that more research is necessary to validate and expand these insights across different contexts and cultures. Yang et al. (2023) examined the use of an optimized model-based outpatient treatment support system for the treatment of diabetes patients and evaluated its possible advantages. The investigation's focus was the ChatGLM model, which was trained using the P-tuning and LoRA fine-tuning techniques. The refined model was then effectively incorporated into the Hospital Information System (HIS). Based on the basic information, primary complaints, medical history, and diagnosis data of each patient, the system generates personalized treatment recommendations, suggestions for laboratory tests, and medication prompts. The results of the experimental testing demonstrated that the improved ChatGLM model can produce precise treatment recommendations based on patient data, along with suggestions for suitable laboratory tests and medication reminders. The model's outputs, however, may come with risks for patients with complicated medical records and are not a complete replacement for outpatient physicians' clinical judgment and decisionmaking skills. The electronic health record (EHR) is the only source of data used by the model, which makes it difficult to fully reconstruct the patient's treatment course and occasionally results in inaccurate assessments of the patient's treatment objectives. Zhang et al. (2023) evaluated ChatGPT and GPT-4 on several medical domain tasks. Nevertheless, none have examined its effectiveness in providing clinical diagnostic support for patients across the entire spectrum of disease presentation, nor evaluated its performance using a large-scale real-world electronic health record database. Using a real-world large electronic health record database, the study conducted two analyses using ChatGPT and GPT-4: one to identify patients with specific medical diagnoses, and the other to assist healthcare providers in the prospective evaluation of hypothetical patients by offering diagnostic support. The findings demonstrate that GPT-4 can achieve up to 96% F1 scores on various disease classification tasks when chain of thought and few-shot prompting are used. In terms of patient assessment, GPT-4 has a three-out-of-four diagnostic accuracy rate. There were, however, 89 references to factually false claims, missing important medical discoveries, suggestions for pointless research, and overtreatment. These problems, along with privacy concerns, render these models unsuitable for clinical use in the real world at this time. In contrast to the configuration of traditional machine learning workflows, prompt engineering requires less data and less time, which highlights its potential for scalability across healthcare applications. Liu et al. (2023) developed and evaluated the efficacy of large language models that had been fine-tuned to generate responses to patient messages sent via an electronic health record patient portal. The authors used an OpenAI API to update physician responses from an open-source dataset into a format with informative paragraphs that offered patient education while emphasizing empathy and professionalism. They did this by using a dataset of messages and responses that were extracted from the patient portal at a large academic medical centre. The model they developed, called CLAIRShort, was based on a pre-trained large language model, called LLaMA-65B. Ten typical patient portal questions in primary care were used to generate responses to assess the fine-tuned models. The datasets were combined to further refine our model (CLAIR-Long). Primary care physicians were asked to rate the empathy, responsiveness, accuracy, and usefulness of generated responses from ChatGPT and our models. Compared to CLAIR-Short responses, CLAIR-Long responses offered more patient education, and they were evaluated similarly to ChatGPT responses, with favourable ratings for accuracy, responsiveness, and empathy, and a neutral rating for usefulness. The results indicated that there was a great deal of promise for improving communication between patients and primary care physicians by using large language models to produce replies to patient messages. Barnard et al. (2023) assessed the potential of large language models as a tool for general user self-diagnosis along with how LLMs could contribute to the dissemination of false information about medical conditions. Although these models' success on medical exams has been used to support their use in medical diagnosis and training, the effects of their unavoidable use as a self-diagnostic tool and their role in disseminating false information about healthcare have not been assessed. A testing methodology that can be applied to assess answers to open-ended questions 96 procedure to maximize its effectiveness. The study comprised patients who were over 40 years old (45%), 20% of whom were younger than 10 years old, 35% of whom were between the ages of 11 and 40 and 80% of the doctors had five years and more experience in diagnosing tropical diseases. The results show that while the FCM's diagnosis was accurate in 85% of cases, the doctors' initial theories were accurate in only 55% of cases. This finding is intriguing since it would seem that the doctors would have better classification accuracy than the FCM system. It's also crucial to note that there was only a negligible correlation (0.14) between the doctor's experience level and the initial hypothesis' accuracy. The study's limitation is its limited generalizability due to the small sample size of test data used. Nathaniel et al. (2017) developed a system to diagnose typhoid fever using fuzzy logic because fuzzy logic can accurately simulate human thought processes. The system demonstrated the ability to receive patient data and symptoms, process the data, and then output a diagnosis to the user. The system has over 200 inference rules and uses twenty-one (21) membership functions as inputs and demonstrated reliability, with an accuracy rate of 97.5%, positioning it as a dependable method for diagnosing typhoid fever. The limitation of the study is that it could only diagnose typhoid fever. Uzoka et al. (2017) employed the AHP model to diagnose tropical febrile diseases. Using the AHP data collection instrument, the study gains firsthand knowledge from medical professionals about the diagnosis of tropical confusable diseases by extracting 15 doctors experience knowledge who specialize in diagnosing tropical diseases. In total, eighteen symptoms were taken into account in the study and a tool was devised to identify the risk factors linked to the diseases that are being studied. Using statistical tests of significance, the following risk factors were identified: working as a street vendor, having poor personal hygiene, traveling to an endemic area, and coming into contact with sick people. The researchers created a modelimplementing Android application and installed it on 35 tablet computers so that doctors in Africa could test the system. Despite the common symptom overlap of tropical diseases, one of the main features of the system is its ability to detect comorbidity. The limitation of the study is that the system was unable to handle 97 diseases that are not included in the model and the system indicates the diagnosis as "other," necessitating additional investigation on the part of the doctor. Hezekiah et al. (2020) created a system of artificial intelligence called ARIS to diagnose typhoid fever. The primary goals were to identify the most important risk factors for Typhoid Fever (Typhoid Responsive Expert System, or TyRes), develop a fuzzy logic base-expert system that can forecast the illness based on symptoms, and utilize TyRes to predict TyF in patients. To gather data, two sets of questionnaires were employed. Patients in 25 hospitals in Lagos, Abeokuta, and Ifo, South-West Nigeria, received 325 copies. To gather information about TyF and its symptoms, another 200 copies were given to human medical experts (HME), which included 140 certified nurses and 70 doctors. Chi-Square was used to analyze the data and determine the primary symptoms that the majority of the HME observed. Matlab 2015a was used to implement TyRes with the primary factors serving as input variables. The development of TyRes involved the use of input variables such as vomiting, loss of appetite, abdominal pain, weakness, and high temperature. A 76% accuracy rate was obtained when predicting TyF in 25 patients by comparing HME predictions with TyRes outcomes and the study concluded that 76% of all TyF predictions can be modelled by TyRes. The limitation of the study is that it could only diagnose typhoid fever. Okagbue et al. (2021) utilized data mining models to diagnose malaria by utilizing 15 symptoms from 337 patients in Nigeria. Weak associations between symptoms and outcomes were discovered after eight ML algorithms were applied to the data. However, a secret pattern was discovered that accurately predicted results according to symptoms, age, and sex. The top model was the Adaboost, which had a 1.8% error rate, 98.2% classification accuracy, and 96.6% precision. The result implies that the Adaboost model could be utilized in malaria-endemic areas to create quick diagnostic tools or decision support systems for malaria diagnosis, thereby lowering misdiagnosis and enhancing public health. Maidabara et al., (2021) developed an effective system for diagnosing patients with typhoid and malaria. Data was gathered from the University of Maiduguri Teaching Hospital over four years, spanning from 2017 to 2020. The study investigated the 98 possible advantages of putting forth a novel model for symptom-based diagnosis and prediction of typhoid and malaria using the Naive Bayes (NB) model, which was implemented in Python. In under thirty minutes, the system provides a real-time diagnosis and does so without requiring a trip to the laboratory for testing. The study employed three distinct algorithms: NB, SVM, and Artificial Neural Network (ANN). The outcomes show that NB and SVM yield the highest diagnostic accuracy of 100%. Muhammad and Varol (2021) used the decision tree classification to develop a predictive diagnostic model for malaria diagnosis. Collected with the symptomatic features of experimental laboratory data from patients, a total of 500 hospital samples were gathered; 300 of which were utilized for training, 150 for testing, and 50 for validation. Demographic information, including age, sex, and tribe, was also gathered at the Maryam Abacha Hospital in Sokoto, Nigeria. These all functioned as input variables for the algorithm. Better pre-processing results were obtained by replacing all missing values with software after the data were preprocessed. By identifying the patients' malaria stages based on their symptoms, a 77% accurate decision algorithm was created. This study also revealed that, contrary to what many other research studies have suggested, malaria does not only kill young people (those under five years old). According to this study, older women are more likely to contract malaria at a severe stage. Because decision trees are so sensitive to even slight variations in the data, small changes in the data can produce entirely different trees. This instability may lead to unreliability in the model. Odion and Ogbonnia (2022), proposed a web-based machine learning system for diagnosing malaria and typhoid fever using the Flask web framework and the XGBoost machine learning algorithm for better classification results. 690 data was collected from a single diagnostic centre located in Nnewi, Anambra state, Nigeria, which included patients' lab test results along with their symptoms for this study. The input variables used were high temperature, headache, weakness, body and Abdominal pain, vomiting, age, gender, and widal count. The binary classifier for malaria has a 99.2% F1-Score and an accuracy of 98.6%, while the multi-class classification for malaria has a 96.8% F1-Score and an accuracy of 97.6%. Similarly, 99 the multi-class classifier for typhoid has an accuracy of 96.1 % and an F1-Score of 95.1 %, while the typhoid binary classifier has a 96.1% accuracy and a 98.5% F1Score. The limitation of the study is that it could only diagnose two diseases (typhoid fever and malaria). Mariki et al. (2022) aimed to demonstrate supervised ML models for diagnosing malaria utilizing demographic information and patient symptoms. The study utilized a malaria diagnosis dataset from two regions in Tanzania, selecting important features to speed up processing and improve model performance. ML classifiers, including SVM, K-Nearest Neighbour, Logistic Regression, Decision Tree, and Random Forest, were employed with k-fold cross-validation. High accuracy was attained by the study; RF in Kilimanjaro reached 95% accuracy, Morogoro 87%, and the combined dataset 82% accuracy. However, the paper does not provide specific quantitative metrics such as recall, precision, or F1 score for the classifiers used. Uzoka et al. (2022) developed an AHP model for the diagnosis of Typhoid fever by mining the experiential knowledge of medical professionals with expertise in treating similar conditions and having firsthand experience with them. The Analytic Hierarchy Process model for typhoid fever diagnosis and the physicians' firsthand experience treating tropical diseases were the study's methods. The AHP model was tested on 2044 patient data, and the model effectively identified whether typhoid fever was present in 78.91% of the cases, indicating the effectiveness of AHP in typhoid fever diagnosis. The study's limitation was that the model was only intended for one febrile illness (typhoid fever), and that techniques like the adaptive AHP approach and linguistic preference relations, which also improve consensus, could have improved the pairwise comparison's consistency. Awotunde et al. (2022) developed a system that utilizes Neuro-Fuzzy Inference System (GENFIS) and Genetic Algorithm (GA) to diagnose typhoid fever and malaria. To train the GENFIS, the GA component first verifies the ideal collection of network features and saves and delivers these features to the right hidden layer nodes. The system is tested and evaluated in the study using the MATLAB environment. The model outperformed some of the current systems, yielding an accuracy performance of 97.2%. The recommended strategy may lessen the main problems 100 with GENFIS systems if it is fully implemented. Additionally, it can be applied to challenging problems in various fields. Obot et al. (2023a) developed a mobile app to aid CHWs in diagnosing, and treating tropical febrile diseases in low-to-middle-income countries due to the acute shortage of experienced physicians in the rural communities in these regions. This work aimed to create a decision support system (DSS) that frontline health workers (FHWs) can use to diagnose febrile illnesses. The methods used in the study were AHP, Agile methodology, and experiential knowledge of physicians. Experiential knowledge was extracted from 62 experts in the field of diagnosing febrile diseases and their related symptoms in secondary and tertiary healthcare facilities in some states in the southern part of Nigeria. 11 illnesses and 50 symptoms were used in this study. 3253 patient data was also obtained from these healthcare facilities by these medical experts, and these patients provided information on the symptoms of the 11 different febrile diseases. The analytic hierarchy process, which is a multi-criteria decision technique, was used to develop the diagnostic engines. The linear consensus models in the AHP engine determine which disease the decision support filters should activate. The AHP diagnostic engines facilitate the identification of suspected illnesses based on the patient's symptoms and risk factors, helping to differentiate between the symptoms of febrile illnesses and recommend the best course of action. The AHP model showed 90% diagnosis accuracy and 91% precision. The limitation of the study is that it relied on the experience and knowledge of 62 doctors. Obot et al. (2023b) created an FCM-based MDSS for febrile diseases to address the issue of patients with febrile illnesses often having limited access to medical care in resource-scarce regions and the issue of inexperienced doctors often struggling to distinguish between the signs and symptoms of febrile diseases. A fuzzy cognitive map model was developed in response to this, to provide medical professionals with a decision-support tool for diagnosing febrile diseases. The model was created by fuzzifying and weighting 2465 datasets from four states in Nigeria's regions that are prone to febrile diseases, with the assistance of 60 medical doctors involved in the process of obtaining the datasets. In this study, 11 diseases and 50 symptoms were used to illustrate the FCM ideas. For the 11 febrile diseases included in the study, the 101 average accuracy of the computations used to predict diagnosis results for the 2465 patients and those diagnosed by the doctors on the scene was 87%. The presence of enteric fever and malaria in over 80% of the datasets is concerning, as the combined percentage of yellow fever, dengue fever, and laser fever in the datasets is only 2% and this imbalance in the dataset contributed to the unsatisfactory results. Another flaw in the study is the early convergence of FCM, which shows up in situations where the equilibrium is reached after as few as six iterations. The study will eventually be expanded to include intuitionistic fuzzy logic and interval type-2 and it suggests conducting additional research excluding dengue, Lassa, and yellow fever from the dataset. The study's drawback is that patient records older than 16 were chosen, depriving patients younger than 16 from using the system. Bhuiyan et al. (2023) proposed predictive models for diagnosing enteric fever, another name for typhoid fever. Deep learning and machine learning techniques were used in the study to develop prediction models. The study demonstrated the effective use of machine learning algorithms, like XGBoost, in diagnosing typhoid fever with a high accuracy rate of 97.87% before clinical trials are conducted. The limitation of the study is that it could only diagnose typhoid fever. Suryani et al. (2023) developed an android-based expert system for diagnosing enteric fever. fuzzy logic algorithm was utilized to reduce the uncertainty and imprecision involved with conventional medical diagnosis techniques. Expert knowledge about the general and clinical symptoms that typhoid fever patients frequently experience, as well as information about the severity of the disease, were incorporated into the development of the application. Direct observation at the hospital and in-person interviews with physicians were used to gather data. After processing the data, the system generated output in the form of suggested solutions and diagnostic findings. Three user levels can be found in the application: Administrators, who can handle user data; Users, who can diagnose their diseases by answering questions in the application; and Experts, who can handle symptom, disease, and solution data. The research findings are available as an application that can be used at any time to serve as a stand-in consultant for the community, providing 102 diagnostic results in the form of negative, positive, or strongly positive typhoid. The limitation of the study is that it could only diagnose typhoid fever. Mariki (2023) used symptomatic and non-symptomatic patient data to develop an ML model for Tanzanian malaria diagnosis. Significant features such as fever, body malaise, visit date, residence area, age, and headache were chosen, and the model was trained using k-fold cross-validation techniques. For the Kilimanjaro, Combined, and Morogoro datasets, respectively, RF and DT methods yielded the highest prediction accuracy rates, at 96%, 99%, and 98%. Using observable symptoms and non-symptomatic variables, the model allows for the prediction of the patient's malarial status before the prescription of antimalarial drugs. The findings have the potential to lower medication resistance and enhance the management of malaria in medical facilities and drug distribution centres. Bhuiyan et al. (2023) created a typhoid fever prediction model using deep learning and machine learning before conducting a clinical trial in Africa. More health industries are using machine learning techniques as a result of the amazing results that machine learning and deep learning have produced for extrapolative analysis. Deep learning and ten different machine learning algorithms were used in this paper to create the model and the XGBoost classifier yielded the best performance with an accuracy of 97.87%. The limitation of the study is that it could only diagnose typhoid fever. In related work, Suryani et al. (2023) developed an Android-based typhoid fever diagnostic application using a fuzzy logic algorithm. The application was developed using data on the severity of typhoid disease as well as expert knowledge about general and clinical symptoms that patients with the illness frequently experience. Direct observation at the hospital and in-person interviews with physicians were used to gather data. After being processed, the data was output in the form of the system's recommended solutions and diagnostic results. The application is divided into three user levels: experts who can process symptom, disease level, and solution data; users who can diagnose their diseases by answering questions in the application; and administrators who can process user data. The research findings are available as an application that can be used at any time to serve as a stand-in consultant for the 103 community, providing diagnostic results in the form of negative, positive, or strongly positive typhoid. The limitation of the study is that medical diagnostics frequently deal with noisy, high-dimensional data, which presents challenges for fuzzy logic algorithms. The diagnosis of typhoid fever requires a number of symptoms, test results, and patient history information. The complexity of creating suitable fuzzy rules and membership functions rises with the number of input variables. Barracloug et al. (2023) present a study aimed at enhancing the accuracy of diagnosing types of malaria by incorporating feature selection strategies and parameter tuning approaches based on AI and ML classifiers into InfoGainAttributeEval. The aim is to build a method that can correctly diagnose types of malaria by leveraging AI and ML classifiers, addressing the challenges associated with traditional methods like blood smear analysis. The study uses 100 features extracted from 4000 samples related to malaria cases. These features are applied to train and assess the performance of the proposed AI system. The methodology involves integrating parameter tuning methods and feature selection techniques with AI and ML classifiers. The study evaluates the performance using ANN, NB, ensemble methods, and RF classifiers. The results indicate that the NB classifier achieved 100% accuracy within a short time, outperforming other classifiers. Additionally, the RF algorithm demonstrated high performance with 100% accuracy in classifying different types of malaria, and the Ensemble methods also achieved 100% accuracy consistently across the feature sets. The study does not explicitly discuss the interpretability or explainability of the model's decisionmaking process. This could be a limitation in terms of understanding how the model arrives at its diagnosis. Apanisile and Ayeni (2023) developed an extended diagnostic system (EDS) to diagnose malaria and typhoid fever using the Naïve Bayes technique. The medical records of patients at the Lagos University teaching hospitals in Lagos State, Nigeria, were used to create a localized clinical database for observed symptoms. Two hundred and fifty (250) records with five (5) attributes, including ailment type, risk level, gender, and symptoms 1 and 2, are included in the dataset. The accuracy of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) 104 were the established metrics used to evaluate the performances of the system. Both typhoid and malaria fevers showed a highly significant improvement in the EDS's performance. A significant limitation of the study is the relatively small dataset size of 250 records. When working with large datasets, where feature independence is more likely to hold, naïve Bayes classifiers perform well. The model might not have enough information to fully represent the subtleties and underlying patterns linked to typhoid and malarial fever due to its small dataset. This may result in overfitting, a phenomenon in which the model learns more specific details and noise from the training set than it does broadly applicable patterns. When a model is overfitted, it becomes less able to generalize to new and unseen data, which could lead to subpar diagnostic performance when used on different patient populations or in actual clinical settings. Furthermore, a small dataset might not fully capture the range of clinical presentations, which could result in incomplete or biased conclusions. Kandula et al. (2023) proposed a method that involves training the VGG19 algorithm on a large dataset of images that detects the presence of malaria parasites. The proposed method identified malaria-infected cells with a 95% accuracy rate, and it may improve the efficacy of malaria diagnosis, particularly in settings with limited resources and limited access to trained medical personnel. The study's main drawback is that, despite the model's high accuracy rate, differences in image quality, staining methods, and equipment between laboratories and healthcare facilities may make it difficult for it to perform well in a variety of clinical settings. In a related study, CNN, ResNet50, and VGG19 models were utilized by Dath et al. (2023) to identify the Plasmodium parasite in thick blood smear images. Compared to other methods, the experimental results show that the VGG19 model performed the best, achieving an accuracy of 98.46%. The study demonstrates how artificial intelligence can increase pathogen detection speed and accuracy, which is more efficient than manual analysis. When a model achieves a 98.46% accuracy rate, it may have overfitted the training dataset, which is a possible limitation of the study. Chattopadhyay (2024) explored the application of ML in grading infectious diseases, focusing on Typhoid fever. The study's objective is to model how a computer can grade Typhoid fever using an ML-based approach that mirrors how novice doctors 105 learn from senior doctors. The data used in the study comprises 'weighted' sign symptoms and matching 'labeled' grades of synthetic cases of Typhoid fever (N = 198, respectively). Ten Machine Learning Classifiers (MLCs) were used to create ten Virtual Junior Clinicians (VJCs) and were trained based on the provided data. The methodology involved training the VJCs with the data, and Each VJC's performance was evaluated using the diagnostic metrics of F1-score, Accuracy, Precision, and Recall. Notably, RFand DT-based Clinical-Based Support Systems (CDSS) achieved an average accuracy of 87%, surpassing human clinicians' accuracy. The study's results show that RF and DT classifiers performed well in grading Typhoid fever. The study delivers insights into choosing the right MLC algorithm for Infectious Disease diagnosis and discusses challenges in implementing MLC-based CDSS in real-world scenarios. Limitations of the study include the synthetic nature of the data used and the assumptions made in mimicking the learning process of novice doctors. Asuquo et al. (2024) developed AHP models for differential diagnosis of febrile conditions. The study aimed to create a multi-criteria decision analysis technique for tropical febrile disease differential diagnosis. Utilizing the Open Data Kit application, data were gathered using two instruments that were verified by domain experts. From one data source, experiential knowledge was extracted from Sixty-two (62) doctors with experience in diagnosing febrile illnesses from both public and private health facilities. The second source of data came from a patient consultation tool that was meant to help doctors gather information about their patients' symptoms, record early diagnoses that included additional testing, and record the results of final diagnoses. The models were developed based on the physicians' experience and the accuracy of the differential diagnosis of febrile diseases was assessed by testing the AHP model's performance using the patients' dataset as a baseline. When comparing the results of the AHP models with the physician's suspected diagnosis, the AHP model performs slightly better. For a variety of diseases, the accuracy of the model ranged from 85.4% to 96.9%, exceeding doctors' predictions for Lassa, Dengue, and Yellow Fevers. The limitation of the study is that the AHP model may be biased and subjective due to the dataset used, and different 112 The process involved classifying the database by computing a Convolutional Neural Network and extracting features using Mel-frequency cepstral coefficients (MFCC). The suggested method, according to the results, has an accuracy of 90.21%, making it a suitable way to quickly classify any respiratory sounds that have been collected from various devices. The limitation of the study is the shortage of medical specialists who can correctly diagnose patients based only on respiratory sounds. Vaid et al. (2020) developed machine learning models based on patient characteristics at admission to forecast the patients' hospital courses over clinically meaningful time horizons. Their analysis involved looking through the electronic health records (EHRs) of patients who were admitted to hospitals within the Mount Sinai Health System in New York City after testing positive for COVID-19. To predict in-hospital mortality and critical events at time windows of 3, 5, 7, and 10 days from admission, Extreme Gradient Boosting was employed along with baseline comparator models. Harmonized electronic health record data for 4098 COVID-19positive patients admitted between March 15 and May 22, 2020, from five hospitals in New York City, comprised our study population. Before or on May 1, the models were externally validated on patients from four other hospitals (n = 2201) before or on May 1, and then prospectively validated on all patients (n = 383) after May 1. The models were initially trained on patients from a single hospital (n = 1514). The study determined the interpretability of the model to determine and prioritize the variables that influence model predictions. With an area under the receiver operating characteristic curve (AUC-ROC) for mortality of 0.89 at 3 days, 0.85 at 5 and 7 days, and 0.84 at 10 days, the XGBoost classifier outperformed baseline models after cross-validation. Critical event prediction was another area where XGBoost excelled, with an AUC-ROC of 0.80 at 3 days, 0.79 at 5 days, 0.80 at 7 days, and 0.81 at 10 days reported. During the external validation, XGBoost was able to predict mortality with an AUC-ROC of 0.88 at 3 days, 0.86 at 5 days, 0.86 at 7 days, and 0.84 at 10 days. AUC-ROC values of 0.78 at 3 days, 0.79 at 5 days, 0.80 at 7 days, and 0.81 at 10 days were also attained by the unimputed XGBoost model. Performance trends on future validation sets were comparable. The strongest predictors of critical events at 7 days were acute kidney injury at admission, elevated LDH, tachypnea, and 113 hyperglycemia; the strongest predictors of mortality were older age, anion gap, and C-reactive protein. Machine learning models for mortality and critical events for patients with COVID-19 at various time horizons were trained and validated by the study. These models found underlying relationships that predicted outcomes and identified at-risk patients. Hu et al. (2021) used a machine learning technique to forecast early patient outcomes for those with severe COVID-19 infection. The predictive models were developed using a dataset of 183 patients from the Sino-French New City Branch of Tongji Hospital, Wuhan, who had a severe COVID-19 infection (115 survivors and 68 nonsurvivors). The features were chosen and the patient outcomes were predicted using machine learning techniques. The performance of the models was compared using the area under the receiver operating characteristic curve (AUROC). The model was validated using 64 patients from the Optical Valley Branch of Tongji Hospital in Wuhan who had severe COVID-19 infection. Between the survivors and nonsurvivors, there were notable differences in the baseline traits and laboratory tests. All five models selected four variables: age, lymphocyte count, d-dimer level, and high-sensitivity C-reactive protein level. Because of its ease of interpretation and simplicity, the logistic regression model was chosen as the final predictive model, despite the models performing similarly. The external validation sets' AUROCs were 0.881. Using a 50% chance of death as the cutoff, the validation set's sensitivity and specificity were 0.839 and 0.794, respectively. The mortality risk can be evaluated using a risk score that is based on the variables that have been chosen. The COVID19 patients' age, high-sensitivity C-reactive protein level, lymphocyte count, and ddimer level at admission are predictive of their prognosis. Mutai et al. (2021) proposed machine learning techniques to determine HIV predictors in sub-Saharan Africa. XGBoost, k-Nearest Neighbors, Elastic Net (EN), RandomForest, Support Vector Machine, and Light Gradient Boosting (LGBT) algorithms were used in training and validating the data. Using population-based HIV Impact Assessment (PHIA) data for 41,939 male and 45,105 female respondents with 30 and 40 variables, respectively, from four sub-Saharan countries, machine learning techniques were used to build models. 80% of the data was used for 114 algorithm training and validation, with the remaining 20% being used for testing. By using the XGBoost algorithm, the f1 scoring mean for males and females was 92% and 90%, respectively, higher than the other five algorithms, suggesting a significant improvement in HIV positivity identification. The model's validity was one of the study's limitations. The training data may had been impacted by the high level of missingness and inconsistency from the self-reported data. Osamor and Okezie (2021) use an extended weighted voting ensemble method to develop a predictive model that would aid in tuberculosis diagnosis. The classification model was created using a combination of Naïve Bayes and Support Vector Machine classifiers. The NCBI GEO site provided a dataset on tuberculosis gene expression, which included 48,803 genes compared to 498 samples. These samples were divided and categorized into 395 samples for other diseases and 103 samples for patients with PTB (Active TB). After removing duplicate genes, there were 25,159 genes left, which were utilized for feature selection and dimensionality calculations. The transcriptional signatures were derived from gene expression data, which was utilized in the developed model. Eighty percent of the dataset was used to train the model, and the remaining twenty percent was used to test the model's accuracy. An enhanced weighted voting ensemble method was employed to complete the model development and enhance the classification accuracy of the individual classifier. This resulted in an accuracy of 0.95, which was higher than the accuracies obtained from the single classifiers. The model was developed using two classifiers, NB and SVM. This suggests that applying ensemble techniques enhances the accuracy of a classification model created by combining individual classifiers. The limitation of the study was that it could only diagnose one febrile illness. Nguyen et al. (2022) proposed a way to diagnose tuberculosis on X-ray images (CXR) automatically using the graph neural network approach. Due to the serious consequences of tuberculosis on patient health and the disease's quick spread, early screening for the illness is vital. Because of their affordability and ease of use, chest X-ray images are frequently utilized as resources for clinical diagnosis of tuberculosis. Machine learning is currently being used in research on ComputerAided Diagnosis (CAD) systems to give physicians analytical, diagnostic, and 115 disease-monitoring tools. Recent years have seen an increase in the use of graph neural networks (GNNs), which provide perfect accuracy in a wide range of fields. The CRX dataset was divided into two categories by us: TB and non-TB. With the suggested model, the study's results were encouraging: accuracy 99.33%, recall 99.07%, precision 99.63%, f1-score 99.35%, and AUC 99.97%. The limitation of the study was that it could only diagnose one febrile illness. Rahhal et al. (2022) used a publicly available symptoms dataset and a machine learning approach to differentiate COVID-19 from other upper respiratory tract infections. They then used Apriori algorithms to identify the most significant relationships between the symptoms. A total of six classifying algorithms, Bagging, Random Forest, Extra Trees, Ada Boost, Stochastic gradient boosting, and voting ensemble, were employed to infer the disease type from its symptoms. According to the experiments, the voting ensemble algorithm (approximately 96.22%) had the highest classification testing accuracy. The study showed, in its conclusion, that the application of an ensemble technique may significantly improve classification accuracy and facilitate the differentiation of COVID-19 from other similar diseases. The limitation of the study was that it could only distinguish COVID-19 from other upper respiratory tract infections. Orjuela-Cañón et al. (2022) proposed machine learning in the loop for tuberculosis diagnosis. Data were gathered at Hospital Santa Clara (HSC) in Bogotá, D.C., Colombia, as part of the TB programme. The data was gathered using the hospital's conventional tuberculosis diagnosis procedure. Data from 233 clinically suspected pulmonary tuberculosis subjects were taken into consideration. These subjects' data had been collected between January 2017 and December 2019. Following the national protocol to diagnose tuberculosis, 184 subjects (79%) had their TB confirmed, and 36 subjects (15%) were found to be disease-free based on smear microscopy, culture, and molecular examination. Thirteen subjects were excluded from consideration due to the lack of information regarding their TB status. The five machine learning models that were employed in the study were artificial neural networks, random forests, logistic regression, and classification trees. The findings demonstrated that artificial neural networks achieve the highest accuracy and 116 sensitivity values, at 0.80 and 0.82, respectively. These findings are superior to smear microscopy, a method that is frequently employed to identify tuberculosis in specific situations. Results show that machine learning in the tuberculosis diagnosis loop can be strengthened with accessible data to function as a substitute diagnosis tool based on data processing in locations with limited health infrastructure. One of the study's limitations is that the data set had a high incidence of tuberculosis, which may have caused bias in the data analysis. To address this, more detailed scenarios involving clinical observation will be needed. In certain instances, TB culture is also regarded as the gold standard for diagnosis, particularly in situations where GenExpert's infrastructure is unavailable. Kaur et al. (2022) proposed a model to predict a patient's dengue fever infection levels by identifying dengue hemorrhagic fever, dengue shock syndrome, and dengue fever. A machine-learning-based tertiary classification technique was employed to determine which group of dengue infections the patient has by collecting and analyzing data, predicting the presence of dengue infections, and estimating risk levels. The system worked in real-time to diagnose patients based on symptomatic and clinical investigations. The patient was warned to seek medical attention in the event of an emergency related to the disease, employing warning indicators that alert them to the possibility of internal bleeding. Based on the World Health Organization's classification system, which includes dengue fever, dengue hemorrhagic fever, and dengue shock syndrome, the suggested model forecasts a patient's infection levels with a notably high accuracy of over 90% as well as high sensitivity and specificity values. The limitation of the study was that it could only predict dengue fever. Abdualgalil et al. (2022) proposed a dengue prediction system to assist doctors in accurately predicting dengue disease. The main goal of this work was to use Efficient Machine Learning Techniques (EMLT) to develop a diagnostic model for the early diagnosis of dengue disease. EMLT-based prediction models for dengue fever were proposed which included five distinct and effective machine learning models: eXtreme Gradient Boosting (XGB), Extra Tree Classifier (ETC), Gradient Boosting Classifier (GBC), K-Nearest Neighbour, and Light Gradient Boosting Machine (LightGBM). The dataset was used to train and evaluate each classifier using the 10- 117 Fold Cross-Validation and Holdout Cross-Validation techniques. Different metrics, including accuracy, F1-sore, recall, precision, AUC, and operating time, were used to assess each model on a test set. According to the results, the ETC model had the best accuracy in 10-fold cross-validation and hold-out, scoring 99.03% and 99.12%, respectively. The finding showed that ETC, which achieved 99.12% accuracy, was the best classifier with high accuracy using the Holdout cross-validation approach. The lengthy time it took for clinical examinations to provide an accurate diagnosis of dengue was the limitation and the requirement for a new diagnostic framework to identify dengue early. Rana et al. (2022) proposed a dengue fever expert system using machine learning analytics (DFES-MLA) to predict dengue fever disease more efficiently considering only the symptomatic features. The study employed Decision Tree and Random Forest (RF) classifiers along with all the data pre-processing steps for Dengue Fever prediction. Fever, joint pain, headaches, vomiting, and other symptoms are typical of dengue; however, if the disease is not identified in time, it can cause severe bleeding, shock, and even death. SMOTE, Borderline SMOTE, ADASYN, Gaussian SMOTE, Decision Tree, and Random Forest classifiers were the methods applied in the study. The dengue dataset's imbalance and the requirement for oversampling techniques to address class imbalance constituted the study's limitations. Veena et al. (2022) proposed a multilabel classification and clinical data analysis for dengue fever prediction. Real patient data with one target value and eighteen attributes were used in the research project, and it was gathered from the General Medicine Department at PESIMSR, Kuppam, Andrapradesh. From January 2020 to June 2021, a total of eighteen months were dedicated to gathering data. The proposed model used the Random Forest with Gridsearch algorithm to tune the hyperparameters for the grid search approach for the prediction to diagnose the class or level of dengue fever. To address the issue of model bias, the 10-fold crossvalidation technique was utilized and the accuracy of classification was positively impacted by hyperparameter tuning. Accuracy, precision, recall, and f1 score are performance metrics that were used in the analysis to compare the suggested model with other machine learning models. The experimental analysis shows that for the 118 training and testing datasets, the RF using GridsearchCV achieved 100% and 98.79% accuracy, respectively. In a related study, a method using supervised machine learning and the influence of seasonality was proposed by Ming et al. (2022) for diagnosing dengue in patients presenting with acute febrile illness. The study presented a supervised machine learning model to determine if dengue or other febrile illnesses (OFI) were diagnosed in patients suffering from acute febrile illnesses, and also examined the impact of seasonality on model performance over time. 8,100 patients were enrolled in the study between October 16, 2010, and December 10, 2014, and 2,240 (27.7%) patients were diagnosed with dengue infection. The enrolled patients had been sick for less than 72 hours with an acute febrile illness. The final diagnosis was predicted using a gradient boosting model (XGBoost) utilizing enrollment-related data on age, sex, hemoglobin, platelet, white blood cell, and lymphocyte counts. Eighty percent of the data were randomly divided into a hold-out set and a training set. The latter was not used in the development of the model. In predicting the final diagnosis, the optimized model derived from training data had an overall median area under the receiver operator curve (AUROC) of 0.86 (interquartile range 0.84–0.86), specificity of 0.92, sensitivity of 0.56, positive predictive value (NPV) of 0.73, negative predictive value (NPV) of 0.84, and Brier score of 0.13. One of the study's limitations is that seasonality and other factors caused the model's performance to fluctuate over time. Hassan et al. (2022) used deep learning and spectroscopic images to develop a model for diagnosing dengue virus infection. 2,000 Raman spectra images were used which 1,200 were human blood sera samples infected with DENV, and 800 were of healthy individuals. The methods used were deep learning and Raman spectroscopy. By using Raman spectroscopic data from human blood sera, the ResNet101 deep learning model is altered by utilizing the transfer learning (TL) concept. With testing data, the system provided 96.0% accuracy in diagnosing DENV infection. Furthermore, compared to other state-of-the-art methods, the developed approach showed a minimum improvement of 6.0% and 7.0% in terms of AUC and Kappa index, respectively. The developed method's sole limitation was that it needs skilled 119 personnel to obtain Raman spectra and feed them to the DENV-TLDNN for diagnosis. Gupta et al. (2023) proposed a diagnosis and prediction model for dengue fever using sentiment analysis and machine learning methods since dengue is a major global health risk. Artificial Neural Network, Decision Tree, Naive Bayes, Random Forest, K-nearest Neighbour, and support vector machine were utilized in the study. Bayesian inferences and support vector machine algorithms were employed in the study to mine viewpoints and extract emotions from text. Even though the results produced by the DT., KNN, SVM, and GNB methods are all better, the R.F. method takes a lot longer to compute because it produces better results. It seems that the R.F. technique is the best option based on the results. Because of this, it has been found that the RF-based diagnostic model is the most suitable for correctly diagnosing dengue fever at an early stage among all of these different machine learning algorithms. These techniques are weak miners of sentiment at the level of the sentence or phrase, and they only function well when the text passage inputs are at the page or paragraph level. They are also weak semantically. Khan and Raza (2023) developed and assessed a Machine Learning-Based Predictive Diagnostic System to identify dengue fever early and classify different forms of dengue fever using machine learning algorithms (Decision Tree, Random Forest, Naive Bayes algorithms). The proposed system was intended to help novice or untrained staff learn and experiment more effectively, as well as aid medical professionals in identifying the disease at an early stage. The 400 records in the dataset used for this study included 15 attributes, and the data was pre-processed to lower noise, incompleteness, and inconsistencies. An accuracy of 95.6% was attained on average by the proposed system, and the highest individual accuracy was recorded by Random Forest, at 98%. The dataset used in the study was limited to 400 records and when compared to other algorithms, Naive Bayes produced lower accuracy. Ramakrishna and Karthikeyan (2023) proposed an Artificial Neural Network-based diagnostic model that uses the Updated Chimp optimization algorithm (UChOA) to detect and predict dengue virus infection. The study used an Updated Chimp Optimization Algorithm-Artificial Neural Network (UChOA-ANN) based 120 diagnostic model to effectively identify dengue-affected blood samples. The ideal weight parameters for an ANN were chosen using UChOA to enhance its performance. According to the findings, the suggested diagnostic model had reached its highest levels of accuracy, specificity, and sensitivity at 94%, 96%, and 92%, respectively. In terms of sensitivity, specificity, accuracy, PPV, NPV, FPR, FNR, and FDR, the suggested diagnostic model has performed better than alternative techniques for identifying blood samples affected by dengue. The study's limitation was the need for lengthy clinical exams to provide an accurate diagnosis. 2.4 Knowledge Gap The existing study discussed in this thesis employed AHP-based models to support the diagnosis of febrile diseases. While this method offered a structured decisionmaking framework, it was limited to patients over the age of 17 and lacked interpretability of the diagnostic results. This limitation reduces its usefulness in various clinical settings where transparent and explainable outcomes are essential for both healthcare providers and patients. Furthermore, the broader body of related literature reviewed has demonstrated the application of machine learning, XAI, or LLMs in various healthcare settings. However, none of these studies have used an integrated framework that combines machine learning for prediction, XAI for interpretability, and LLMs for better reasoning and explanation in diagnosing febrile diseases. This lack of integration highlights a significant gap, as the complexity and overlapping symptoms of febrile illnesses require not only accurate predictions but also transparent, explainable, and context-aware diagnostic support systems. Addressing this gap is essential for advancing healthcare technology by creating reliable and interpretable diagnostic models that can be used across diverse patient populations, ultimately enhancing trust, adoption, and outcomes in diagnosing febrile diseases. 121 CHAPTER THREE SYSTEM ANALYSIS AND DESIGN 3.1 Method Adopted in the Study The study adopted Agile Development and Data-Driven methods to ensure the creation of a robust and user-friendly diagnostic system. Agile development was chosen because it is user-centric, supports incremental delivery, communication, and collaboration, and is well known for its flexibility and adaptability in system development. Data-driven methodologies were selected because they support evidence-based decision-making and improvement through continuous data collection and analysis for constant enhancement and validation of the study’s diagnostic model. Machine learning (ML), Explainable AI (XAI), and Natural Language Processing (NLP) are the specific data-driven methods used in this study. The ML method was used to develop a diagnostic app from trained models on labeled data, which consists of patients' symptoms and diseases, to diagnose tropical diseases. The XAI method used was Local Interpretable Model-agnostic Explanations (LIME). LIME visualizations were utilized to provide a more understandable, trustworthy, and transparent ML model that will facilitate the adoption of the enhanced diagnostic system in healthcare. The NLP method known as Generative Pre-trained Transformer (GPT) was also utilized to provide contextual explanations of diagnostic information, thus further explaining the diagnostic outcomes in natural language. The Unified Modeling Language (UML) tools were used in this study to aid in the modeling, designing, and communication of the enhanced diagnostic system. These UML tools helped produce a reliable, effective, and user-focused system by guaranteeing an Agile framework and a structured yet adaptable approach to creating a machine learning-based diagnostic app. 3.2 Analysis of the Existing System The existing system is Febra Diagnostica (Obot et al, 2023a). It uses analytic hierarchy process (AHP)-based MCDA models to diagnose and treat eleven tropical febrile diseases. The Canadian New Frontiers in Research Fund (NFRF) funded the study, which promotes high-risk/high-reward, transformative, and world-leading 128 Through these explanations, medical professionals are better able to build trust by understanding the underlying causes of each diagnosis. The enhanced diagnostic system is a more efficient and reliable tool for febrile disease diagnoses because of its dual improvements in age inclusivity and diagnostic clarity. 3.4 System Model The enhanced diagnostic system accepts patient personal data, symptoms, and vitals through a mobile device. The local and cloud storage holds the patient data and models' outputs. The ML model processes the patient's symptoms to diagnose the febrile disease or diseases, the XAI model Interprets the results of the ML model while the LLM processes the patient's symptoms, and the XAI explanations to generate natural language explanations for easier decision-making as presented in Figure 3.2. This user-friendly application guarantees accessibility and offers clear explanations with the combination of these models to capture comprehensive diagnostic indicators. This system outperforms the current AHP-based system and extends applicability to a wider range of age groups by providing a scalable, interpretable diagnostic tool. The enhanced diagnostic system architecture; the highlevel input design; the high-level process design, the high-level output design, and the database design are presented to provide an overview of the whole system. 129 Figure 3.2: Enhanced Diagnostic System Model 3.4.1 Architecture of the Enhanced Diagnostic System In the enhanced diagnostic system, the components of the system and their interrelationships are structured. The components comprise medical experts from which patient data were collected, data Preprocessing, diagnostic system, model evaluation, healthcare provider, the patient, the mobile device, and the cloud storage, as illustrated in Figure 3.3. 130 Figure 3.3: Architecture of the Enhanced Diagnostic System 3.4.1.1 Mobile Device The Android-based mobile device serves as the interface between healthcare providers, patients, admin, and the diagnostic system. The mobile device is used by healthcare providers to input patient data such as personal information, patient vitals, and symptoms and receive output such as provisional diagnosis, patient medical record patient history overview, vital signs summary, user management report, and recommendations to healthcare providers as shown in Figure 3.4. The mobile device provides a user-friendly interface for users to enter information, and access the cloudbased system, and serves as the access point for the diagnostic and decision support functionalities. 131 Figure 3.4: Mobile device of the Enhanced Diagnostic System 3.4.1.2 Decision Filter The decision filter mimics a skilled physician's reasoning by appropriately grouping patient vitals and symptoms for diagnostic decisions. The decision filter identifies critical symptoms and ensures that all relevant factors have been considered before finalizing a diagnosis. The input data into the decision filter are personal information, vitals, and symptoms, while the output data are diagnoses and additional recommendations. A flow chart of the decision filter is presented in Figure 3.5, which presents the steps followed by the filter to send appropriate information to the database as well as the diagnostic system for decision-making. 3.4.1.3 Diagnosis and Recommended Treatment The diagnosis and recommended treatment component provides healthcare workers with the patient's diagnosis from the diagnostic system, which comprises the RF diagnosis, LIME interpretation, further explanations by the GPT engine, and recommended treatment. This component holds the diagnostic results from the diagnostic system, patient vitals, and additional recommendations from the decision filter. The recommendation is aligned with medical guidelines and is tailored based on the patient’s history. 132 Figure 3.5: A flowchart of the decision filter 133 3.4.1.4 Cloud Storage The cloud storage holds patient data, model data, diagnostic results, and other system records. It ensures that patient data is stored securely and can be accessed as needed. The application programming interface (API) is the bridge between the enhanced diagnostic app and the cloud architecture, allowing the app to send and retrieve data from the cloud storage, invoke ML models, and receive diagnostic results. The API handles secure data transmission, ensuring privacy and compliance with regulations as shown in Figure 3.6. The API forwards personal information, vitals, and symptoms to the cloud storage service which are saved securely and organized into different tables depending on the type. After the model completes the diagnosis, the evaluation metrics and the diagnostic results are also sent to the cloud storage via the API to keep track of versions and for future reference. Figure 3.6: Diagram of the cloud storage 3.4.1.5 Patient The patient is the source of all medical data, such as symptoms, vitals, and health records. Patients provide symptoms, vitals, and medical history to a healthcare worker for entry into the system using a mobile device. The patient has limited functions, including editing their personal record, viewing medical history summary, and requesting detailed medical records from the healthcare worker. The patient played a key role as one of the end users who provided essential feedback on the system’s usability, accessibility, and relevance. Although not directly involved in the development process, patients participated in user interviews, usability testing, and 134 feedback sessions after system iterations. Their input helped the team understand how the system meets real-world requirements such as ease of use. Their continuous feedback loop is crucial in building a user-centric system that is intuitive and beneficial for the target audience. 3.4.1.6 Healthcare Worker The healthcare worker interacts directly with the system to input personal information, patient history, and examination, including temperature, blood pressure, respiratory rate, height, weight, and symptoms from the patient through the userfriendly interface of the app. The healthcare worker can then interpret and act on the diagnostic results and recommendations from the app for decision-making. The healthcare worker served as a key user and stakeholder who provided valuable insights into the clinical workflows and practical application of the system. They collaborated with the development team to define user stories, ensuring that the system accurately reflects the real-world tasks and challenges they encounter in patient care. The healthcare worker helped validate features such as data input processes, diagnosis generation, and recommendations by offering feedback on their efficiency and usability. During sprint reviews and testing phases, they assess whether the system supports effective and accurate decision-making in a healthcare setting. Their involvement ensured the system was both clinically functional and intuitive for everyday use, bridging the gap between technical development and practical healthcare needs. 3.4.1.7 Medical Experts The medical experts are experienced physicians in tropical febrile diseases from secondary and tertiary healthcare facilities (both public and private) who collected the patients' data from patients with febrile diseases during their clinic days. The medical experts assisted in developing the data collection instrument, identifying key symptoms, and helping define the medical rules that guide the system. They provided valuable insights and initial knowledge for the system development. Their primary responsibility was to provide expert insights into medical practices, patient care, and healthcare workflows, ensuring that the system accurately reflects real-world clinical requirements. They collaborated closely with the system developer to define user 135 stories, validate system requirements, and provide feedback during sprint reviews. The medical expert helped in refining diagnostic algorithms, interpreting clinical data, and ensuring that the system's outputs, such as diagnoses and recommendations, align with medical standards. They also participate in testing phases to validate that the system's functionality supports accurate, practical, and usable healthcare outcomes. By working iteratively within the agile framework, the medical expert ensures the system remains patient-centric and clinically relevant. 3.4.1.8 Data Collection The dataset was collected from a study conducted by a group of computer scientists, doctors, nurses, and community health workers from Nigeria, Canada, the United States of America, and the United Kingdom to develop a multi-disease, multisymptom soft-computing system for early differential diagnosis of tropical diseases by frontline health workers. This study was funded by the New Frontiers in Research Fund (NFRF), which promotes high-risk/high-reward, transformative, and worldleading international research, and was approved in June 2020 with application number 102232. The research involved human participants and was evaluated and granted approval by the Human Research Ethics Board (HREB) of Mount Royal University in Calgary, Alberta, Canada, and ethical clearance of the study was given by the Institutional Health Research Ethical Committee (IHREC) at the University of Uyo Teaching Hospital in Uyo, Akwa Ibom State, Nigeria. The approval to use the dataset in this study was received from the Institute of Health Research and Development, University of Uyo Teaching Hospital, Uyo, Nigeria. The dataset contains 4870 patient records comprising demographic data, patient symptoms, risk factors, suspected diagnoses, further investigation, and confirmed diagnoses (UNIUYO and MRU, 2024). Data exploration was needed once the data had been collected to examine the dataset's structure, types of data, size, and features. It also involved computing basic statistics for numerical data to spot trends, patterns, and outliers. The dataset showed that 4605 patient records were collected during the rainy season, 40 during harmattan, and 225 during the dry season. Table 3.1 presents the descriptive statistics of patients in the dataset, indicating that there were 2175 male and 2695 female patients. 136 Table 3.1 Descriptive statistics of male and female patients in the dataset Age Range Male Female Pregnant Women No Nursing Mothers No < 5years 534 419 1st trimester 139 0-3 months 27 5 years to 12 years 346 323 2nd trimester 184 4-6 months 35 13years to 19 years 150 213 3rd trimester 86 7-9 months 28 20 years to 64 years 1012 1605 Over 9 months 63 65 years and above 133 135 Total 2175 2695 Total 409 Total 153 The data exploration further showed the number of pregnant patients from the first to third trimester and patients who were nursing mothers and their respective months in the dataset. The dataset contained fifty (50) symptoms and eleven (11) suspected and confirmed diagnoses which are highlighted in Table 3.2. Table 3.2 Patient Symptoms and Diseases in the Dataset SN Symptom/Disease Abbreviation SN Symptom/Disease Abbreviation 1 Abdominal pains ABDPN 33 Muscle and body pain MSCBDYPN 2 Back pain BCKPN 34 Mouth ulcer MUTUCR 3 Bitter taste in mouth BITAIM 35 Nausea NUS 4 Bleeding BLDN 36 Night sweats NGTSWT 5 Bloody urine BLDYURN 37 Pain behind the eyes PNBHEYE 6 Catarrh CTRH 38 Upper back pain (loin) UPBCKPN 7 Chest indraw CHSIND 39 Painful urination PNFLURNTN 8 Chest pain CHSPN 40 Peritonitis PERTN 9 Chills and Rigors CHLNRIG 41 Red eyes REDEYE 10 Cloudy urine CLDYURN 42 Red eyes, face, tongue REDEYEFCTNG 11 Constipation CNST 43 Sensitivity to light SENLHT 12 Cough (initial dry) CGHDRY 44 Shock SHK 13 Diarrhea DRH 45 Skin rash SKNRSH 14 Difficulty breathing DIFBRT 46 Sore throat SRTRT 15 Dizziness DIZ 47 Suprapubic pains SPPBPN 16 Dry cough DRYCGH 48 Urinary frequency URNFQC 17 Fatigue FTG 49 Vomiting VMT 18 Fever FVR 50 Wheezing WHZ 137 19 High persistent fever HGPSFVR 51 Malaria MAL 20 High-grade fever HGGDFVR 52 Typhoid fever ENFVR 21 Stepwise rise fever SWRFVR 53 HIV and AIDS HVAD 22 Sudden onset fever SUDONFVR 54 Upper urinary tract infection UPUTI 23 Low-grade fever LWGDFVR 55 Lower urinary tract infection LWUTI 24 Foul breath FOLBRT 56 Upper respiratory tract infection URTI 25 Body itching BDYICH 57 Lower respiratory tract infection LRTI 26 Generalized body pain GENBDYPN 58 Tuberculosis TB 27 Generalized rashes GENRSH 59 Lassa fever LASFVR 28 Headaches HDACH 60 Yellow fever YELFVR 29 Intestinal bleeding and perforation INTBLEPRF 61 Dengue fever DENFVR 30 Joint swelling JNTSWL 31 Lethargy LTG 32 Lymph node swelling LMPNDSWL The patient's symptoms were presented on a five-point scale (1=absent; 2=mild; 3=moderate; 4=severe; 5=very severe) and the diagnoses on a six-point scale (1=absent; 2=very low; 3=low; 4=moderate; 5=high; 6=very high) along with the physician’s level of confidence (a numerical rating scale from 1 to 10) for every symptom and disease. The sample patient dataset is presented in Table 3.3. 3.4.1.9 Data Preprocessing Data processing is a crucial phase that entails feature selection, feature scaling, and data cleaning. Expunging records with missing features, irrelevant data, and columns that were not needed as part of the data-cleaning process. During the data cleaning process, records of patients under the age of five (5) were removed from the data because these patients could not accurately express certain symptoms, and the data collection instrument used did not make provision for certain symptoms of patients 144 Figure 3.7 Random Forest algorithm schematic diagram The RF model was adopted as the main engine of this study for multidisease diagnosis of febrile diseases. This RF model in Equation (3.2) was adopted from a consistency analysis study by Biau et al. (2012) which outlines the formulation of tree splits and their optimization in high-dimensional spaces. For each input 𝑥 (representing a patient's symptoms), the goal is to diagnose a disease 𝑦 =𝑦1,𝑦2,𝑦3,𝑦4,𝑦5,𝑦6 where 𝑦𝑖∈{0,1} indicates whether the 𝑖𝑡ℎ disease is present. 𝑦𝑖=1 𝑛∑𝑇𝑗(𝑥) 𝑛 𝑗=1 (3.2) where: 𝑦𝑖 is the predicted label for the 𝑖𝑡ℎdisease. 𝑇𝑗(𝑥) is the prediction from the 𝑗𝑡ℎdecision tree for input 𝑥. 𝑛 is the total number of trees in the forest. In this multi-label classification, each tree will make a diagnosis for all six diseases, and the final output is the average diagnosis across all trees, which can be thresholded to determine the presence (1) or absence (0) of each disease. ii. Extreme gradient boost (XGBoost) was selected for this study because it is an ensemble technique and an advanced gradCient boosting implementation that is portable, flexible, and efficient, making it a suitable option for disease diagnosis. XGBoost constructs classification trees one 145 at a time, using the residuals from each tree to train the next, and the schematic diagram of the computational process is shown in Figure 3.8. XGBoost makes use of regularisation techniques to improve model generalization and gradient-boosted decision trees as its foundation. It adds weak learners (decision trees) to the ensemble one after the other in a stepwise manner, with each new learner concentrating on fixing the mistakes made by the previous ones. During training, it minimizes a predetermined loss function using the gradient descent optimization technique. Figure 3.8 Extreme gradient boosting algorithm schematic diagram The XGBoost model in Equation (3.3) and (3.4) were adopted from a study by Chen and Guestrin (2016) which utilized gradient boosting to maximize prediction accuracy while controlling regularization. 𝑦𝑖=∑𝑓𝑘(𝑥) 𝐾 𝑘=1 (3.3) where: 𝑦𝑖 is the predicted score for the 𝑖𝑡ℎdisease. 𝐾 is the number of boosting rounds (trees). 𝑓𝑘(𝑥)is the contribution of the 𝐾𝑡ℎ tree in predicting the probability of disease 𝑖 146 Since this work is a multi-label classification task, the loss function is the combination of binary cross-entropy for each disease as presented in Equation (3.4). The symptoms are represented by the input feature space and the model generates six probability scores, each of which indicates the likelihood of a specific disease. 𝐿 =∑ ∑ (𝑦𝑖𝑗 log(𝑦𝑖𝑗)+(1−𝑦𝑖𝑗)log(1−𝑦𝑖𝑗)) 𝑛 𝑗=1 6 𝑖=1 +∑𝛺(𝑓𝑘) 𝐾 𝑘=1 (3.4) where: 𝑦𝑖𝑗 is the true label (0 or 1) for the 𝑖𝑡ℎdisease in the 𝑗𝑡ℎpatient. 𝑦𝑖𝑗 is the predicted probability of disease 𝑖 in patient 𝑗. 𝛺(𝑓𝑘)is the regularization term for the tree 𝑓𝑘 iii. Multi-layered perceptron (MLP) was selected for this study because it is an effective tool for disease diagnosis due to its versatility in managing various data types, capacity to learn from high-dimensional datasets, and ability to model complex relationships. When combined with appropriate training techniques and interpretability methods, MLPs can offer reliable and practical insights for medical diagnosis. It is a feedforward artificial neural network with multiple layers, comprising an input layer, one or more hidden layers, and an output layer as shown in Figure 3.9. Figure 3.9 Multi-layered perceptron architecture 147 The MLP mathematical model in Equation (3.5) was adopted from a study on the formalism of the general mathematical expression of MLP neural networks by Hounmenou et al. (2021). The model comprises an input layer representing patient features(symptoms), multiple hidden layers, and an output layer with six neurons (one for each disease). Each neuron performs a weighted sum of inputs followed by a nonlinear activation function ∅ and uses backpropagation to minimize a loss function. ℎ𝑗=∅(∑𝑤𝑖𝑗𝑥𝑖+𝑏𝑗 𝑛 𝑖=1 ) (3.5) where: ℎ𝑗 is the predicted probability for the 𝑗𝑡ℎ disease. 𝑤𝑖𝑗 and 𝑏𝑗 are the weight matrix and bias vector for the final layer 𝑥𝑖 is the 𝑖𝑡ℎ input symptom, 𝑛 is the number of input symptoms. iv. Local Interpretable Model-agnostic Explanations (LIME) provide interpretability locally by using a more straightforward model to approximate the behaviour of the model around a particular diagnosis. When diagnosing diseases where patient cases may differ greatly from one another, LIME is extremely helpful when generating explanations by aiding healthcare workers in comprehending why a model diagnosed a disease for a particular patient considering their unique symptoms. This localized explanation improves the diagnostic model's accuracy and reliability by helping to spot any anomalies or mistakes in the diagnosis. LIME is model-agnostic and can be used with a variety of ML models, offering versatility and wide applicability in a range of diagnostic scenarios. The LIME model in Equation (3.6) utilized in this study was adopted from a study by Saini and Prasad (2021) which proposed a novel explainable AI method to provide locally interpretable model agnostic explanations. The goal of LIME is to identify an interpretable model 𝑔 that roughly corresponds to the complex model 𝑓 around a given instance 𝑥. 148 𝐿 =𝑎𝑟𝑔𝑚𝑖𝑛𝑔⁡𝐿(𝑓,𝑔,𝜋𝑥+Ω(g)) (3.6) Where 𝐿(𝑓,𝑔,𝜋𝑥)is the loss function that measures how well the interpretable model 𝑔 approximates the complex model 𝑓 around the instance 𝑥. It typically considers a weighted least squares loss. 𝜋𝑥 is a proximity measure that assigns higher weights to samples closer to 𝑥, emphasizing local behaviour. Ω(g) is a regularization term that penalizes the complexity of the interpretable model 𝑔, ensuring that it remains simple. v. Generative Pre-trained Transformer (GPT) greatly improves the diagnosis of febrile diseases by utilizing its strong natural language processing capabilities to evaluate and comprehend intricate medical data. When combined with an ML model, it improves the efficiency of disease diagnosis by providing thorough explanations and justifications for its recommendations, which ultimately lead to better patient outcomes. It leverages the transformer architecture which is the foundational model in natural language processing due to its ability to handle sequential data effectively. The GPT models in Equations (3.7) and (3.8) were adopted from Luo et al. (2022) and Lee (2023). The task is for GPT to generate human-readable explanations based on the predicted labels and probabilities from the ML models. The input to the GPT will be a structured prompt containing the patient's symptoms and ML model results indicating the likelihood of each disease. The goal of GPT is to generate an explanation for each predicted disease based on its probability score and the associated symptoms, incorporating medical knowledge. The GPT model will generate text conditioned on the structured input (ML predictions + symptoms). GPT predicts the next token in a sequence based on previous tokens. The sequence in this study includes: symptoms of the patient, predicted disease probabilities from ML models, and the request for explanation. For a sequence 𝑥1,𝑥2,𝑥3,…,𝑥𝑁 the model generates the next token 𝑥𝑁+1 based on the probability: 149 𝑃(𝑥𝑛+1|𝑥1,𝑥2,𝑥3,…,𝑥𝑁)=𝑠𝑜𝑓𝑡𝑚𝑎𝑥(𝑊0𝑧𝑁+1) (3.7) where 𝑧𝑛+1is the hidden representation at position 𝑁+1. GPT uses self-attention to capture relationships between the symptoms and the predicted disease probabilities: 𝐴𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛(𝑄,𝐾,𝑉)=𝑠𝑜𝑓𝑡𝑚𝑎𝑥(𝑄𝐾𝑇 √𝑑𝑘)𝑉 (3.8) where: 𝑄 =𝑊𝑄ℎ⁡are the queries 𝐾 =𝑊𝐾ℎ⁡are the keys 𝑉 =𝑊𝑉ℎ⁡are the values ℎ is the hidden state from the previous layer 𝑊𝑄, 𝑊𝐾,and 𝑊𝑉 are learnable weight matrices 𝑑𝑘 is the dimension of the keys Here, Q, K, and V represent the symptoms, predicted probabilities, and other relevant medical information. This allows GPT to contextualize why a certain disease was predicted, based on the input features. The prompt that is fed into GPT will be structured as a concatenation of the patient's symptoms, the disease predictions with their probabilities from the ML model, and instruction to generate an explanation. Given this prompt, GPT will use its learned medical knowledge and the context provided by the predictions to generate an interpretative explanation. For each disease, GPT will provide an explanation based on the diagnoses and symptoms. The mathematical formulations of RF, XGBoost, and MLP provide the foundation for predictive modelling, and RF was selected as the main engine of the study while the LIME equation enhances interpretability. Including these equations in this work contextualizes their applicability to the specific problem of diagnosing febrile diseases. This research demonstrates a novel method for using GPT to generate natural language explanations of model predictions, enhancing the transparency and interpretability of machine learning models in healthcare. 150 3.4.1.11 Mathematical Model The mathematical model presented in this section is independent of the specific ML, XAI, and LLM techniques. Therefore, any ML algorithm, XAI technique, and LLM can be implemented using this model. i. Machine Learning Model: The ML algorithm in Equation (3.9) predicts the disease 𝑦 based on input features (symptoms) 𝑋 =𝑥1,𝑥2,𝑥3,…,𝑥32 𝑦 =𝑓𝑀𝐿(𝑋)=𝑓𝑀𝐿(𝑥1,𝑥2,𝑥3,…,𝑥32) (3.9) where: 𝑋 =𝑥1,𝑥2,𝑥3,…,𝑥32 represents the input vector of 32 symptoms. 𝑦 is the predicted disease for each possible disease. The output of this model can be expressed as Equation (3.10): 𝑂𝑀𝐿(𝑋)={𝑃(𝑑1|𝑋),𝑃(𝑑2|𝑋),…,𝑃(𝑑𝑘|𝑋)} (3.10) where: 𝑂𝑀𝐿 is the output of the machine learning model 𝑃(𝑑𝑖|𝑋) is the probability of disease 𝑑𝑖 given the symptoms 𝑋 𝑘 is the total number of diseases (in our case, k=6). ii. Explainable AI Model: The XAI model analyses how each symptom contributes to the diagnosed disease. The output in Equation (3.11) is an explanation of the diagnosis, which is importance scores for each symptom. 𝑂𝑋𝐴𝐼(𝑋,𝑦)=[𝑐1(𝑋,𝑦),𝑐2(𝑋,𝑦),…,𝑐32(𝑋,𝑦)] (3.11) where: 𝑂𝑋𝐴𝐼(𝑋,𝑦) is the output of the explanation of the diagnosis 𝑐𝑖(𝑋,𝑦) is the contribution of symptom 𝑥1 in diagnosing 𝑦 iii. Large Language Model: The LLM in Equation (3.12) interprets both the ML model's diagnosis and the XAI explanation and generates a natural language description 151 𝑂𝐿𝐿𝑀(𝑋)=𝑓𝐿𝐿𝑀(𝑂𝑀𝐿(𝑋),𝑂𝑋𝐴𝐼(𝑋,𝑦)⁡) (3.12) where 𝑂𝐿𝐿𝑀(𝑋) is the LLM's output that uses the ML output 𝑂𝑀𝐿(𝑋) and the explainability output ⁡𝑂𝑋𝐴𝐼(𝑋,𝑦) to form a coherent explanation in natural language, providing healthcare workers with a clear interpretation of both the diagnosis and the reasoning behind it. The outputs of the ML, XAI, and LLM models in Equations (3.10), (3.11), and (3.12), respectively, are combined into a single model in Equation (3.13), which can work with any ML, XAI, or LLM technique. 𝐷(𝑋)=𝑂𝐿𝐿𝑀 (𝑂𝑀𝐿(𝑋),𝑂𝑋𝐴𝐼(𝑋,𝑂𝑀𝐿(𝑋))) (3.13) where: 𝐷(𝑋) is the diagnostic system interpretation in natural language. 𝑂𝐿𝐿𝑀(𝑂𝑀𝐿(𝑋),𝑂𝑋𝐴𝐼(𝑋)) is any LLM that takes the output 𝑂𝑀𝐿(𝑋) from the ML model and the explanation ⁡𝑂𝑋𝐴𝐼(𝑋) from the XAI model to generate a natural language interpretation 𝐷(𝑋). 𝑂𝑀𝐿(𝑋) is any ML algorithm that takes symptoms 𝑋 and diagnose the disease 𝑦. 𝑂𝑋𝐴𝐼(𝑋,𝑦) is any XAI method that explains the diagnosis 𝑦 by analysing the symptoms 𝑋 and providing interpretability in terms of the contributions of each symptom. 3.4.1.12 Model Evaluation The performance of the ML model was assessed using Recall, Precision, and F1 Score because they ensure that the model's diagnoses are accurate and reliable. Recall is important for diagnosing diseases because it shows how well the model can identify patients who truly have the condition, and a high recall rate guarantees that the majority of patients are diagnosed with the disease correctly. Precision is significant because it indicates how accurately the model made positive diagnoses, and a high precision indicates that the majority of patients diagnosed have the condition. The F1 score offers a more thorough assessment of the model's 152 performance by combining Precision and Recall into a single metric. When a model has a high F1-score, it is considered reliable for diagnosing diseases because it has high Precision and high Recall. Recall: Recall measures the proportion of correctly predicted positive observations to all the observations in the actual class. It indicates how well the model can identify positive samples. Recall is important when the cost of false negatives is high, such as in medical screenings, where missing a positive case (false negative) can be critical. 𝑅𝑒𝑐𝑎𝑙𝑙 = True⁡Positives True⁡Positives+False⁡Negatives (3.14) Precision: Precision measures the proportion of correctly predicted positive observations to the total predicted positives. It reflects the accuracy of positive predictions. Precision is crucial when the cost of false positives is high. For example, in medical diagnoses, a false positive could lead to unnecessary treatments. 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = True⁡Positives True⁡Positives+False⁡Positives (3.14) F1 Score: The F1 Score is the harmonic mean of precision and recall. It provides a balance between precision and recall, particularly useful for imbalanced datasets. F1 Score is useful when you need to balance precision and recall, especially in cases where you have an uneven class distribution. 𝐹1⁡𝑆𝑐𝑜𝑟𝑒 =2∗(𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛⁡∗⁡𝑅𝑒𝑐𝑎𝑙𝑙 𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+⁡𝑅𝑒𝑐𝑎𝑙𝑙 ) (3.15) 3.4.2. High-Level Input Design of the Enhanced Diagnostics System The high-level input design of the system accepts input parameters through the touchscreen of a tablet or smartphone into the febrile disease diagnostic system graphic user interface (GUI). The high-level inputs are categorized into personal information, patient vitals, and symptoms as shown in Figure 3.10. 153 Figure 3.10. High-level input design of the proposed system The personal information category captures the vital demographic data of patients and healthcare workers while the patient vitals category captures patients’ vitals such as blood pressure, temperature, etc entered through the system graphic user interface (GUI). The symptoms category captures patients’ symptoms using a slider feature depicting the severity of each symptom. Table 3.8 presents the input parameters, data type, format, and mode of capture for the high-level input design. Table 3.8. High-level input parameters Input parameter Data type Format Mode of capture Personal Information Date of Birth (DOB) Date DD-MM-YYYY Date picker interface to select the date Gender Categorical Predefined categories Dropdown menu Name Text Natural language (plain text) Text input field for first name, middle name (optional), and last name Address String Text Text input field with separate lines for street address, city, and state Mobile phone String Numeric Numeric input field with format validation to ensure proper phone number structure 256 ) self.controls = [ ft.Column([ ft.Text("Patient Registration", size= 20, weight= BOLD), self.form, ft.Text( size= 11, spans= [ ft.TextSpan("By clicking Create account you agree to Recognizes\n"), Link("Terms of use", color= ACCENT2_COLOR), ft.TextSpan(" and "), Link("Private Policy", color= ACCENT2_COLOR) ], text_align= ft.TextAlign.CENTER, ), ], expand= True, horizontal_alignment= ft.CrossAxisAlignment.CENTER, spacing= 10, ) ] 257 # RESULT.PY import flet as ft from utility import * class ResultView(ft.View): def __init__(self) -> None: super().__init__( route= "/result", padding= 40, horizontal_alignment= ft.CrossAxisAlignment.CENTER, bgcolor = BACKGROUND_COLOR, ) self.prediction_ref = ft.Ref[ft.Row]() self.img_ref = ft.Ref[ft.Image]() self.explaination_ref = ft.Ref[ft.Text]() self.controls = [ #Interpretable ft.Text("Explainable Diagnosis Results", size= 18, weight= BOLD, text_align= ft.TextAlign.CENTER), ft.Divider(height= 1, color= ft.Colors. ON_SURFACE), ft.ListView([ ft.Column([ ft.Text( 'Provisional Diagnoses:', weight= BOLD, ), ft.Row( controls=[ self.tags(i) for i in ['Typhoid Fever Likely', 'HIV/AIDS Likely', 'Urinary Tract Infection Likely', 'Respiratory Tract Infection Likely', 'Tuberculosis Likely'] ], wrap= True, ref= self.prediction_ref, visible= False, ), ], spacing= 0), ft.Container( ft.Column([ ft.Image("./plot.png", ref= self.img_ref) ], spacing= 0), bgcolor= 'White', padding= 10, border_radius= 10, expand= True ), ft.Container( ft.Column([ ft.Text( size= 10, ref= self.explaination_ref, text_align= ft.TextAlign.JUSTIFY ), ], spacing= 0), 258 bgcolor= 'White', padding= 10, border_radius= 10, expand= True ), ], spacing= 10, expand= True), ft.Row([ ft.ElevatedButton( "Return Home", # Done expand= True, bgcolor= ACCENT2_COLOR, color= "white", style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(radius= 5), padding= 20 ), on_click= lambda e: self.page.go('/dash') ) ]), ] def tags(self, value): return ft.Container( ft.Text(value, color= ACCENT2_COLOR) , bgcolor= 'white', border= ft.border.all(1, ACCENT2_COLOR), padding= 10, border_radius= 5, ) def did_mount(self): # shift this to the loading page self.results = self.page.session.get('diagnosis_result') predictions = self.results['predictions'] img_path = self.results['img_path'] explanation = self.results['explanation'] self.prediction_ref.current.controls.clear() if not len(predictions): self.prediction_ref.current.controls.append(ft.Text("No Disease")) explanation = "No diesease" else: for i in predictions: # create none self.prediction_ref.current.controls.append(self.tags(i)) self.prediction_ref.current.visible = True self.update() self.explaination_ref.current.value = explanation self.img_ref.current.src = img_path patient = self.page.session.get('Current Patient') hcw = self.page.session.get('authentication') data = { "patient_id": patient["patient_id"], "diagnosed_at": datetime.datetime.now().isoformat(), "hc_provider_id": hcw["hc_provider_id"], "details": f"Predicton: {predictions}\n\nExplaination: {explanation}", "recommendation": "See a doctor" if len(predictions) else 'You\'re fine' } 259 print(data) # connect_backend(self.page, data, "add_diagnosis") return super().did_mount() # SETTINGS.PY import flet as ft from utility import * class SettingsView(ft.View): def __init__(self) -> None: super().__init__( route= "/settings", spacing = 10, bgcolor= BACKGROUND_COLOR, ) self.base_color_picker = ft.Row( controls=[ ft.Container( content=ft.Row( controls=[ ft.Container( bgcolor=i[j], border_radius= 90, expand= True, ) for j in ['background', 'accent', 'text'] ], spacing= 0, ), width= 20, height= 20, border= ft.border.all(4, ft.Colors.SURFACE_TINT), border_radius= 90, on_hover= self.onhover, on_click= self.change_color, data= i, ) for i in THEMES_LIST ], ) self.theme_options = ft.Dropdown( value= "English", options=[ ft.dropdown.Option(i) for i in ["Mandarin Chinese","Spanish","English","Hindi", "Bengali","Portuguese,Russian,Japanese", "Yue Chinese","Vietnamese,Turkish,Wu Chinese", "Marathi","Telugu,Western Punjabi,Korean", "Tamil","Egyptian Arabic,Standard German,French", "Urdu","Javanese,Italian,Iranian Persian", "Gujarati,Hausa,Bhojpuri", "Levantine Arabic", "Southern Min" ] ], 260 dense= True, prefix_icon= ft.Icons.LANGUAGE, focused_border_color= TEXT_COLOR, border_radius= 15, ) self.setting_view = ft.ListView( controls= [ ft.Text('Appearance', weight= BOLD, size= 18, color= TEXT_COLOR), ft.Text('Change how the UI looks and feels in the app.', color= TEXT_COLOR), ft.Divider(1,1, color= ft.Colors.with_opacity(0.2, TEXT_COLOR)), ft.Text('Theme-Colors', weight= BOLD, color= TEXT_COLOR), ft.Text('Update your dashboard with a new style.', size= 13, color= TEXT_COLOR), ft.Column( controls=[ self.base_color_picker, ] ), ft.Divider(1,1, color= ft.Colors.with_opacity(0.2, TEXT_COLOR)), ft.Text('Language', weight= BOLD, color= TEXT_COLOR), ft.Text('Change the language.', size= 13, color= TEXT_COLOR), self.theme_options, ], expand= True, padding= 20, spacing= 10, ) self.controls = [ ft.Row( controls=[ ft.IconButton( ft.Icons.CLOSE, on_click= lambda _: self.page.go('/dash'), icon_color= TEXT_COLOR, style= ft.ButtonStyle( overlay_color= ft.Colors.with_opacity(0.3, TEXT_COLOR), ) ), ft.Text( 'Settings', expand= True, weight= BOLD, color= TEXT_COLOR ), ], ), self.setting_view, ] def onhover(self, e): if e.data == 'false': e.control.border= ft.border.all(4, ft.Colors.SURFACE_TINT) e.control.scale = 1 else: e.control.border= ft.border.all(1, ft.Colors.SURFACE_TINT) e.control.scale = 1.2 self.update() 261 def did_mount(self): self.update() return super().did_mount() def change_color(self, e): save_colors = e.control.data self.page.theme.color_scheme= ft.ColorScheme( primary= save_colors['accent'], primary_container= save_colors['container'], surface= save_colors['background'], on_surface= save_colors['text'], on_surface_variant= save_colors['text'], ) self.page.client_storage.set('theme_color', save_colors) self.page.update() # BIO.PY import flet as ft from utility import * class ProfileTile(ft.Column): def __init__(self, title, content, ref, **kwargs): self.title = title self.content = content super().__init__( controls= [ self.title, self.content, ], ref= ref, **kwargs ) def update_content(self, value): self.content.value = value self.update() @property def _value(self): return self.content.value class BioView(ft.View): def __init__(self) -> None: super().__init__( route= "/bio", spacing = 10, bgcolor= BACKGROUND_COLOR, scroll= ft.ScrollMode.HIDDEN, padding= 20 ) self.base_color_picker = ft.Row( controls=[ ft.Container( content=ft.Row( 262 controls=[ ft.Container( bgcolor=i[j], border_radius= 90, expand= True, ) for j in ['background', 'accent', 'text'] ], spacing= 0, ), width= 20, height= 20, border= ft.border.all(4, ft.Colors.SURFACE_TINT), border_radius= 90, on_hover= self.onhover, on_click= self.change_color, data= i, ) for i in THEMES_LIST ], ) self.theme_options = ft.Dropdown( value= "English", options=[ ft.dropdown.Option(i) for i in ["Mandarin Chinese","Spanish","English","Hindi", "Bengali","Portuguese,Russian,Japanese", "Yue Chinese","Vietnamese,Turkish,Wu Chinese", "Marathi","Telugu,Western Punjabi,Korean", "Tamil","Egyptian Arabic,Standard German,French", "Urdu","Javanese,Italian,Iranian Persian", "Gujarati,Hausa,Bhojpuri", "Levantine Arabic", "Southern Min" ] ], dense= True, prefix_icon= ft.Icons.LANGUAGE, focused_border_color= TEXT_COLOR, border_radius= 15, ) self.setting_view = ft.ListView( controls= [ ft.Text('Appearance', weight= BOLD, size= 18, color= TEXT_COLOR), ft.Text('Change how the UI looks and feels in the app.', color= TEXT_COLOR), ft.Divider(1,1, color= ft.Colors.with_opacity(0.2, TEXT_COLOR)), ft.Text('Theme-Colors', weight= BOLD, color= TEXT_COLOR), ft.Text('Update your dashboard with a new style.', size= 13, color= TEXT_COLOR), ft.Column( controls=[ self.base_color_picker, ] ), ft.Divider(1,1, color= ft.Colors.with_opacity(0.2, TEXT_COLOR)), ft.Text('Language', weight= BOLD, color= TEXT_COLOR), ft.Text('Change the language.', size= 13, color= TEXT_COLOR), self.theme_options, 263 ], expand= True, padding= 20, spacing= 10, ) self.id_ref = ft.Ref[ft.Text]() self.nname_ref = ft.Ref[ft.Text]() self.name_ref = ft.Ref[ProfileTile]() self.email_ref = ft.Ref[ProfileTile]() self.number_ref = ft.Ref[ProfileTile]() self.sex_ref = ft.Ref[ProfileTile]() self.controls = [ ft.Column( controls=[ ft.Row( controls=[ ft.IconButton( ft.Icons.CLOSE, on_click= lambda _: self.page.go('/dash'), icon_color= TEXT_COLOR, style= ft.ButtonStyle( overlay_color= ft.Colors.with_opacity(0.3, TEXT_COLOR), ) ), ft.Text( # 'Bio', "Personal Information", # expand= True, # weight= BOLD, # color= TEXT_COLOR ), ft.Container( content= ft.Text("HCW", color= ACCENT2_COLOR), bgcolor= ft.Colors.with_opacity(0.3, ACCENT2_COLOR), border= ft.border.all(1, ACCENT2_COLOR), border_radius= 5, padding= 5 ), ], alignment= ft.MainAxisAlignment.SPACE_BETWEEN, ), ], ), ft.Row( controls=[ ft.Column( controls= [ ft.Container( content=ft.Container( content= ft.Text( "AD", size= 40, weight= ft.FontWeight.BOLD, color= 'White', 264 ref= self.nname_ref ), alignment= ft.alignment.center, bgcolor= ACCENT2_COLOR, width= 90, height= 90, border_radius= 200 ), border= ft.border.all(2, ACCENT2_COLOR), border_radius= 200, padding= 3, ), ft.Text( "Aks-0000023", ref= self.id_ref, ) ], # expand= True, ), ], alignment= ft.MainAxisAlignment.CENTER ), ft.Text("Personal Info", weight= ft.FontWeight.W_500), ProfileTile( ft.Text("Name"), ft.TextField( hint_text= "Enter your name", prefix_icon= ft.Icons.PERSON ), ref= self.name_ref ), ProfileTile( ft.Text("Email address"), ft.TextField( hint_text= "Enter your email address", prefix_icon= ft.Icons.EMAIL ), ref= self.email_ref ), ProfileTile( ft.Text("Phone number"), ft.TextField( hint_text= "Enter your phone number", prefix_icon= ft.Icons.PHONE ), ref= self.number_ref ), ProfileTile( ft.Text("Gender"), ft.Dropdown( # width=200, label= "Sex", icon= ft.Icons.PERSON_PIN, options=[ ft.dropdown.Option("Male"), ft.dropdown.Option("Female"), 265 ], data= "sex" ), ref= self.sex_ref ), ft.Row( controls=[ ft.ElevatedButton( # add edit logic "Save changes", expand= True, bgcolor= ACCENT2_COLOR, color= "white", style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(radius= 5), padding= 20, overlay_color= ft.Colors.with_opacity( 0.5, "white" ) ), # on_click= lambda _: self.form.bypassed_validate( # lambda e: validate(e, self.page), # data= { # 'email': 'attaikings[email protected]m', # 'password': 'phd@123456', # }, # # nobypass= True, # ), ), ] ), ] def onhover(self, e): if e.data == 'false': e.control.border= ft.border.all(4, ft.Colors.SURFACE_TINT) e.control.scale = 1 else: e.control.border= ft.border.all(1, ft.Colors.SURFACE_TINT) e.control.scale = 1.2 self.update() def did_mount(self): # add cache info = self.page.session.get('authentication') self.id_ref.current.value = f"Aks-{int(info['hc_provider_id']):06d}" NName = [i[0] for i in info['name'].split(' ')] if len(NName) > 2: NName = NName[:1] elif len(NName) == 1: NName *= 2 self.nname_ref.current.value = ''.join(NName).upper() self.name_ref.current.update_content(info["name"]) self.email_ref.current.update_content(info["email"]) self.number_ref.current.update_content(info["mobile_number"]) self.sex_ref.current.update_content(info["sex"]) self.update() return super().did_mount() 272 alertcontrol.excess = self.control_excess def control_excess(self, value= False): self.navigation_bar.visible = value self.floating_action_button.visible = value self.update() def onsearch(self): date = self.calendersearch.date if date: self._body_controls.search(date) else: self._body_controls.clear() self.update() def will_unmount(self): self.overlay.alertcontrol.excess = None return super().will_unmount() # HCWPATIENTREQUEST.PY import flet as ft from utility import * class HCWPatientRequestView(ft.View): def __init__(self) -> None: super().__init__( route= "/HPreq", padding= 20, horizontal_alignment= ft.CrossAxisAlignment.CENTER, bgcolor = BACKGROUND_COLOR, ) color = TEXT_COLOR self.calendersearch = CalenderSearchcontrol(on_change=self.onsearch) page_padding = None #ft.padding.only(right= 15) self._body_controls = ft.ListView( controls= [ ], expand= True, spacing= 10, padding= page_padding, ) self.controls = [ ft.Container( content=ft.Row( controls=[ ft.IconButton( ft.Icons.CHEVRON_LEFT, on_click= lambda _: self.page.go('/dash'), icon_color= color, style= ft.ButtonStyle( overlay_color= ft.Colors.with_opacity(0.3, color) ) ), ft.Text( 'Request Notification', 273 # expand= True, weight= BOLD, size= 17, text_align= ft.TextAlign.CENTER, ), ft.Container(width= 15), ], alignment= ft.MainAxisAlignment.SPACE_BETWEEN, # expand= True, ), padding= page_padding, ), ft.Container( content= self.calendersearch, padding= page_padding ), ft.Container( content=ft.Row( controls=[ ft.Text( 'Request for access to full medical history', weight= BOLD, size= 12, expand= True, ), ] ), padding= page_padding ), self._body_controls, ft.Container( content=ft.Text( 'Unsuccessful responses will be removed automatically after leaving this page.', weight= BOLD, size= 10, expand= True, text_align= ft.TextAlign.CENTER, ), padding= page_padding ), ] def did_mount(self): self.overlay = overlay(self.page) dt_obj = datetime.datetime.strptime("Sat, Sep 7 2024", r"%a, %b %d %Y") items = [{ 'date': dt_obj, 'id': "AKS-00002" }]*10 # patients = connect_backend(page=self.page, url_code= '/get_patient')['message'] for i in items: self._body_controls.controls.append( ft.ElevatedButton( content= ft.Row( controls=[ ft.Icon(ft.Icons.EMAIL, color= TEXT_COLOR), 274 ft.Column( controls=[ ft.Text( i['date'].strftime(r"%a, %b %d %Y, %I:%M %p"), color= TEXT_COLOR, size= 13 ), ft.Text( f"Patient ID: {i['id']}", color= ft.Colors.with_opacity(0.8, TEXT_COLOR) # size= 10, ), ], spacing= 0, ), ] ), data= i, on_click= self.onclick, style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(10), padding= 10, bgcolor= CONTAINER_COLOR, overlay_color= ft.Colors.with_opacity(0.3, BACKGROUND_COLOR) ) ) ) self.messages = self._body_controls.controls.copy() self.update() return super().did_mount() def onclick(self, e): # print(e.control.data) alertcontrol: AlertControl = self.overlay.alertcontrol alertcontrol.change_view( controls=[ ft.Text( "Request for access to full medical history", weight= BOLD, ), ft.Text( "Ajuga Peterben", weight= BOLD, ), ft.Column( controls=[ ft.Text( "Time", weight= BOLD, ), ft.Text("24 Aug, 3:30 pm"), ], spacing= 0, ), ft.Column( controls=[ 275 ft.Text( "Patient Details", weight= BOLD, ), ft.Text("Phone number: 08140147868"), ], spacing= 0, ), ft.Column( controls=[ ft.Text( "Reason", weight= BOLD, ), ft.Text( r"Routine check-ups are essential for maintaining\ overall health and catching potential issues early.\ This visit will allow us to assess your current \ health status, update any vaccinations, review\ medications, and discuss any concerns you may \ have. Regular screenings for blood pressure, \ cholesterol, and glucose levels can help \ prevent chronic conditions like heart disease \ or diabetes. Staying proactive about your health \ is the best way to ensure long-term well-being.".replace(' ', '').replace('\\\n', '') ), ], spacing= 0, ), ], alignment= ft.alignment.top_left, ) alertcontrol.open(True) def onsearch(self): date = self.calendersearch.date if date: self._body_controls.controls = [ i for i in self.messages if i.data['date'] == date] else: self._body_controls.controls = self.messages.copy() self.update() # CREATEAPPOINTMENT.PY import flet as ft from utility import * def validate(data: dict, page: ft.Page): # data = {'patient_last_name': 'ebong', 'patient_first_name': 'chukwu', 'patient_dob': datetime.strptime('09/06/2000', r"%m/%d/%Y").isoformat(), 'patient_email': '200[email protected]', 'patient_mobile_no': '08767594456', 'patient_address': 'bitch', 'patient_gender': 'female', 'nok_phone': '08767594456', 'nok_name': 'bitch gooo', 'nok_address': 'bitch'} result = connect_backend(data=data, page=page, url_code= 'add_patient') 276 if result: page.session.set('Current Patient', result['data']) page.go('/Pdash') errormessage(page, f"Success: {result['message']}") class PatientAppointmentsView(ft.View): def __init__(self) -> None: super().__init__( route= "/Pappoint", padding= ft.padding.symmetric(horizontal= 25, vertical= 40), horizontal_alignment= ft.CrossAxisAlignment.CENTER, bgcolor = BACKGROUND_COLOR, ) color = TEXT_COLOR#CONTAINER_COLOR self.form = FormField([ ft.Text("new appointment",size= 12, weight= BOLD), customTextField( "HCwName".title(), data= f'HCwName', ), customTextField( "PId".title(), data= f'PId', ), customTextField( "type".title(), data= f'type', ), customTextField( # change to date controls "Choose date".title(), data= f'Choose date', ), customTextField( "Reason".title(), type= "multiline", data= f'Reason', ), ft.ElevatedButton( "Register", expand= True, bgcolor= ACCENT2_COLOR, color= "white", style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(radius= 5), padding= 20 ), on_click= lambda _: self.form.validate( lambda e: validate(e, self.page) ), ) ], expand= True, spacing= 10, padding= ft.padding.symmetric(horizontal= 15) ) 277 self.controls = [ ft.Column([ ft.Row( controls=[ ft.IconButton( ft.Icons.CLOSE, on_click= lambda _: self.page.go('/Pdash'), icon_color= color, style= ft.ButtonStyle( overlay_color= ft.Colors.with_opacity(0.3, color) ) ), ft.Text( 'Create Appointment', expand= True, weight= BOLD, size= 20, ), ] ), self.form, ft.Text( "Create appointments for patients here", size= 11, text_align= ft.TextAlign.CENTER, ), ], expand= True, horizontal_alignment= ft.CrossAxisAlignment.CENTER, spacing= 10, ) ] # PATIENTDASHBOARD.PY import flet as ft from utility import * class PatientDashboardView(ft.View): def __init__(self) -> None: super().__init__( route= "/Pdash", padding= 40, horizontal_alignment= ft.CrossAxisAlignment.CENTER, bgcolor = BACKGROUND_COLOR, scroll= ft.ScrollMode.HIDDEN, ) accent_color = ACCENT2_COLOR self.NName = ft.Ref[ft.Text]() self.Name = ft.Ref[ft.Text]() self.Id = ft.Ref[ft.Text]() self.controls = [ ft.Text("Patient Dashboard", size= 20, weight= BOLD), ft.Row([ft.Container( ft.Column( controls=[ ft.Container( 278 ft.Text( 'AP', size= 20, color= 'white', ref= self.NName, text_align= ft.TextAlign.CENTER, ), bgcolor= accent_color, padding= 5, border_radius= 50, width= 40, height= 40, alignment= ft.alignment.center ), ft.Text( 'Ajuga Peterben', weight= BOLD, color= accent_color, ref= self.Name, ), ft.Text( 'Aks-000024', ref= self.Id, color= TEXT_COLOR, ), ], horizontal_alignment= ft.CrossAxisAlignment.CENTER, spacing= 0 ), bgcolor= CONTAINER_COLOR, padding= 10, border_radius= 10, expand= True, )]), ft.Row([ navbuttons( 'Vital Signs', r'.\heart-rate.svg', 'Pulse Rate, Blood Pressure,...', 4 ), navbuttons( 'Provisional Diagnoses', r'.\stethoscope.svg', 'Examinations', 3, lambda e: self.page.go('/Dign') ), ]), ft.Row([ navbuttons( 'Medical History', r'.\medical-report.svg', 'Patient Diagnosis Records', 3, # lambda e: 279 ), navbuttons( 'Patient Appointments', r'.\twenty-two-calendar.svg', 'Create appointments', 4, lambda e: self.page.go('/Pappoint') ), ]), ft.Row([ ft.ElevatedButton( "Back", expand= True, bgcolor= accent_color, color= "white", style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(radius= 5), padding= 20 ), on_click= lambda e: self.page.go('/Plist') ) ]), ] def did_mount(self): data = self.page.session.get('Current Patient') self.Name.current.value = f"{data['first_name']} {data['last_name']}".title() self.NName.current.value = f"{data['first_name'][0]} {data['last_name'][0]}".upper() self.Id.current.value = f"Aks-{int(data['patient_id']):06d}" self.update() return super().did_mount() # PATIENTLIST.PY import flet as ft from utility import * class SearchBar(ft.Container): def __init__(self, onclick= None, on_back= None, on_clear= None): super().__init__( # visible= visible, margin= ft.margin.only(5, 20, 5, 5), ) search_color = ACCENT2_COLOR # h_color = CONTRAST_COLOR self.b_onclick= on_back self.on_clear = on_clear self.back_but = ft.IconButton( ft.Icons.ARROW_BACK, on_click=self.onback, icon_color= search_color, visible= False ) self.input = ft.TextField( expand= True, on_submit= lambda e: onclick(), prefix_icon= ft.Icons.SEARCH, 280 border_radius= 10, dense= True, text_size= 15, hint_text= 'Search for patients', hint_style= ft.TextStyle( color= ft.Colors.with_opacity(0.6, TEXT_COLOR), ), bgcolor= CONTAINER_COLOR, border_width= 0, suffix= ft.IconButton( ft.Icons.CLOSE, icon_size= 11, height= 25, width= 25, on_click= self.clear ), ) self.content = ft.Row( controls=[ self.back_but, self.input, ], alignment= ft.MainAxisAlignment.CENTER, ) def onback(self, e): self.b_onclick() self.clear('') def clear(self, e): self.input.value = "" if self.on_clear != None: self.on_clear() self.update() class PatientTile(ft.ElevatedButton): def __init__(self, _data, on_click, padding=7): super().__init__( content= ft.Row( controls=[ ft.Column( controls=[ ft.Text( f"{_data['first_name']} {_data['last_name']}".title(), size= 17, expand= True, color= TEXT_COLOR, ), ft.Text( f"Medical ID: Aks-{str(_data['patient_id']).zfill(padding)}", size= 12, expand= True, # weight= ft.FontWeight.W_300, color= ft.Colors.with_opacity(0.8, TEXT_COLOR), ), ], expand= True, 281 spacing= 0, alignment= ft.CrossAxisAlignment.START, ), ] ), on_click= on_click, style= ft.ButtonStyle( shape= ft.RoundedRectangleBorder(10), padding= 10, bgcolor= CONTAINER_COLOR, overlay_color= ft.Colors.with_opacity(0.3, BACKGROUND_COLOR), ) ) class PatientListView(ft.View): def __init__(self) -> None: super().__init__( route= "/Plist", padding= 10, horizontal_alignment= ft.CrossAxisAlignment.CENTER, bgcolor = BACKGROUND_COLOR, scroll= ft.ScrollMode.HIDDEN, navigation_bar= customHCWNavbar(1), # appbar= ft.AppBar(), ) self.searchbar = SearchBar() # get all the patientdata self.body = ft.ListView( spacing= 10, padding= 10, expand= True, ) self.controls = [ self.searchbar, self.body, ] def did_mount(self): # add reload function patients = saveload_cahce(self.page, "patientlist", default= []) self.update_list(patients) overlays = overlay(self.page) if not patients: overlays.loadingview.open(True, src="./help.svg" ,text= 'Checking patient list') # change to a decorator self.navigation_bar.visible = False self.update() patients = connect_backend(page=self.page, url_code= '/get_patient')['message'] self.update_list(patients) overlays.loadingview.open(False) # change to a decorator saveload_cahce(self.page, "patientlist", patients) self.navigation_bar.visible = True self.update() return super().did_mount() def update_list(self, patients): if patients: 288 APPENDIX I: CONFIRMED AND EDS RESULTS Patient ID Abdominal pains Headaches Bitter taste in mouth Lethargy Bloody urine Lymph node swelling Catarrh Muscle and body pain Chest indraw Mouth ulcer Chest pain Nausea Chills and rigors Night sweats Constipation Painful urination Cough (initial dry) Sore throat Difficulty breathing Suprapubic pains Dry cough Urinary frequency Fatigue Vomiting Fever Wheezing High-grade fever Stepwise rise fever Low-grade fever Foul breath Generalized body pain Generalized rashes Confirmed Diagnosis EDS Diagnosis Correctly Predicted False Postive False Negative P1 2 2 2 3 1 1 1 4 1 1 1 5 4 1 4 1 1 1 1 1 1 1 3 4 4 1 5 4 3 3 2 1 Typhoid Fever, Malaria Typhoid Fever, Malaria 1 0 0 P2 1 2 5 1 5 1 1 1 1 1 1 2 4 1 1 2 1 1 1 5 1 5 2 4 4 1 5 1 1 1 5 1 Urinary Tract Infection, Malaria Urinary Tract Infection, Malaria 1 0 0 P3 1 1 1 1 3 1 1 1 1 1 1 1 1 1 1 5 1 1 1 2 1 2 1 1 5 1 1 1 1 1 3 1 Urinary Tract Infection Urinary Tract Infection 1 0 0 P4 1 1 1 1 5 1 1 1 1 1 1 1 1 1 1 4 1 1 1 5 1 2 1 1 2 1 1 1 1 1 5 1 Urinary Tract Infection Urinary Tract Infection 1 0 0 P5 4 2 1 3 1 2 1 5 1 4 1 1 1 5 5 1 1 2 1 1 1 1 5 1 1 1 1 4 4 4 4 2 HIV/AIDS, Typhoid Fever HIV/AIDS, Malaria 0 1 0 P6 1 5 4 1 1 1 1 1 1 1 1 2 4 1 1 1 1 1 1 1 1 1 3 4 3 1 4 1 1 1 3 1 Malaria Malaria 1 0 0 P7 1 1 1 1 5 1 1 1 1 1 1 1 1 1 1 2 1 1 1 2 1 2 1 1 2 1 1 1 1 1 5 1 Urinary Tract Infection Malaria, Urinary Tract Infection 0 1 0 P8 1 2 3 1 1 1 1 1 1 1 1 2 4 1 1 1 1 1 1 1 1 1 3 5 2 1 4 1 1 1 2 1 Malaria Malaria 1 0 0 P9 5 5 1 2 5 1 1 5 1 1 1 1 1 1 5 4 1 1 1 4 1 4 1 1 2 1 1 5 3 3 2 1 Urinary Tract Infection, Typhoid Fever Urinary Tract Infection, Malaria 0 1 0 P10 1 1 1 1 5 1 1 1 1 1 1 1 1 1 1 2 1 1 1 3 1 5 1 1 5 1 1 1 1 1 5 1 Urinary Tract Infection Urinary Tract Infection, Malaria 0 1 0 P11 1 5 4 1 1 1 1 1 4 1 3 5 3 5 1 1 1 1 2 1 4 1 4 5 3 1 3 1 1 1 3 1 Malaria, Tuberculosis Malaria, Respiratory Tract Infection 1 0 0 P12 1 5 2 1 2 1 1 1 1 1 1 4 5 1 1 4 1 1 1 4 1 3 4 3 5 1 4 1 1 1 3 1 Malaria, Urinary Tract Infection Malaria, Urinary Tract Infection 1 0 0 P13 1 1 1 1 1 1 1 1 2 1 2 1 1 3 1 1 1 1 2 1 4 1 3 1 1 1 1 1 1 1 1 1 Tuberculosis Respiratory Tract Infection 1 0 0 P14 1 1 1 5 1 4 1 1 1 4 1 1 1 3 1 1 1 2 1 1 1 1 4 1 1 1 1 1 1 1 1 4 HIV/AIDS HIV/AIDS 1 0 0 P15 1 1 1 2 1 2 1 1 1 3 1 1 1 5 1 1 1 4 1 1 1 1 5 1 1 1 1 1 1 1 1 5 HIV/AIDS HIV/AIDS 1 0 0 P16 1 1 1 2 5 4 1 1 1 3 1 1 1 5 1 3 1 3 1 4 1 5 5 1 5 1 1 1 1 1 4 5 Urinary Tract Infection, HIV/AIDS Urinary Tract Infection, HIV/AIDS, Tuberculosis 0 1 0 P17 1 1 1 1 4 1 1 1 1 1 1 1 1 1 1 2 1 1 1 5 1 2 1 1 3 1 1 1 1 1 5 1 Urinary Tract Infection Urinary Tract Infection, Malaria 0 1 0 P18 3 3 1 3 3 1 1 3 1 1 1 1 1 1 2 4 1 1 1 3 1 2 1 1 2 1 1 5 4 3 4 1 Typhoid Fever, Urinary Tract Infection Malaria, Urinary Tract Infection 0 1 0 P19 1 1 1 5 1 5 1 1 1 2 1 1 1 3 1 1 1 4 1 1 1 1 2 1 1 1 1 1 1 1 1 4 HIV/AIDS Malaria, HIV/AIDS 0 1 0 P20 5 4 1 3 1 1 5 2 1 1 1 1 1 1 5 1 2 2 2 1 1 1 1 1 2 4 1 3 4 5 2 1 Respiratory Tract Infection, Typhoid Fever Respiratory Tract Infection, Malaria 0 1 0 P21 1 1 1 1 4 1 1 1 2 1 4 1 1 3 1 2 1 1 2 5 3 2 4 1 3 1 1 1 1 1 4 1 Urinary Tract Infection, Tuberculosis Urinary Tract Infection, Respiratory Tract Infection 1 0 0 P22 1 3 4 1 1 1 4 1 1 1 1 5 2 1 1 1 2 5 4 1 1 1 4 4 3 3 5 1 1 1 3 1 Malaria, Respiratory Tract Infection Malaria, Respiratory Tract Infection 1 0 0 P23 1 3 2 1 1 1 1 1 5 1 3 5 3 4 1 1 1 1 3 1 4 1 2 4 4 1 5 1 1 1 5 1 Malaria, Tuberculosis Malaria 0 1 0 P24 4 2 1 3 1 5 1 4 1 5 1 1 1 3 5 1 1 3 1 1 1 1 2 1 1 1 1 3 4 3 5 4 HIV/AIDS, Typhoid Fever HIV/AIDS, Typhoid Fever, Malaria 0 1 0 P25 3 5 5 3 1 1 1 4 1 1 1 5 4 1 2 1 1 1 1 1 1 1 4 3 2 1 5 5 5 2 2 1 Malaria, Typhoid Fever Malaria, Typhoid Fever 1 0 0 P26 1 4 2 5 1 5 1 1 1 3 1 3 3 5 1 1 1 4 1 1 1 1 2 3 3 1 5 1 1 1 2 5 HIV/AIDS, Malaria HIV/AIDS, Malaria 1 0 0 P27 1 1 1 3 1 2 1 1 1 4 1 1 1 5 1 1 1 5 1 1 1 1 3 1 1 1 1 1 1 1 1 4 HIV/AIDS HIV/AIDS 1 0 0 P28 1 5 2 1 1 1 1 1 1 1 1 4 3 1 1 1 1 1 1 1 1 1 4 3 4 1 3 1 1 1 5 1 Malaria Malaria 1 0 0 P29 1 1 1 1 1 1 1 1 2 1 4 1 1 5 1 1 1 1 3 1 5 1 5 1 1 1 1 1 1 1 1 1 Tuberculosis Tuberculosis, Respiratory Tract Infection 1 0 0 P30 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 3 4 3 1 1 1 1 1 2 2 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P31 5 5 1 2 1 2 1 5 1 5 1 1 1 2 3 1 1 3 1 1 1 1 3 1 1 1 1 4 5 4 3 5 HIV/AIDS, Typhoid Fever HIV/AIDS, Typhoid Fever, Malaria, Tuberculosis 0 1 0 P32 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 5 2 2 1 1 1 1 1 4 4 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P33 1 5 3 1 1 1 1 1 4 1 3 5 5 5 1 1 1 1 5 1 5 1 2 4 4 1 5 1 1 1 2 1 Malaria, Tuberculosis Malaria, Respiratory Tract Infection 1 0 0 P34 1 1 1 4 1 3 1 1 1 5 1 1 1 3 1 1 1 3 1 1 1 1 5 1 1 1 1 1 1 1 1 3 HIV/AIDS HIV/AIDS 1 0 0 P35 2 4 1 3 1 3 1 2 1 5 1 1 1 2 4 1 1 4 1 1 1 1 5 1 1 1 1 5 2 2 5 3 Typhoid Fever, HIV/AIDS Typhoid Fever, HIV/AIDS, Malaria 0 1 0 P36 1 1 1 1 4 1 1 1 4 1 5 1 1 4 1 4 1 1 4 2 4 5 4 1 2 1 1 1 1 1 4 1 Urinary Tract Infection, Tuberculosis Urinary Tract Infection, Respiratory Tract Infection 1 0 0 P37 1 1 1 5 1 4 1 1 1 4 1 1 1 2 1 1 1 3 1 1 1 1 2 1 1 1 1 1 1 1 1 4 HIV/AIDS HIV/AIDS 1 0 0 P38 1 1 1 1 1 1 4 1 4 1 3 1 1 2 1 1 3 2 3 1 5 1 4 1 2 4 1 1 1 1 1 1 Tuberculosis, Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P39 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 5 4 2 1 1 1 1 1 2 5 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P40 1 1 1 1 1 1 5 1 3 1 3 1 1 4 1 1 5 3 2 1 4 1 3 1 4 3 1 1 1 1 1 1 Respiratory Tract Infection, Tuberculosis Respiratory Tract Infection, Tuberculosis 1 0 0 P41 1 1 1 1 5 1 1 1 1 1 1 1 1 1 1 5 1 1 1 4 1 4 1 1 3 1 1 1 1 1 2 1 Urinary Tract Infection No Disease 0 0 1 P42 2 3 1 4 1 1 1 4 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 1 1 1 4 5 3 4 1 Typhoid Fever Typhoid Fever 1 0 0 P43 1 1 1 5 1 5 3 1 1 3 1 1 1 3 1 1 4 3 4 1 1 1 2 1 2 3 1 1 1 1 1 5 HIV/AIDS, Respiratory Tract Infection HIV/AIDS, Respiratory Tract Infection, Tuberculosis 1 0 0 P44 1 1 1 1 1 1 1 1 4 1 3 1 1 5 1 1 1 1 2 1 2 1 5 1 1 1 1 1 1 1 1 1 Tuberculosis Tuberculosis 1 0 0 P45 4 2 1 3 1 1 1 2 5 1 2 1 1 2 3 1 1 1 3 1 5 1 2 1 1 1 1 5 3 5 4 1 Tuberculosis, Typhoid Fever Malaria, Typhoid Fever 0 1 0 P46 1 1 1 1 1 1 2 1 1 1 1 1 1 1 1 1 3 3 3 1 1 1 1 1 4 2 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P47 1 1 1 1 5 1 1 1 5 1 3 1 1 4 1 2 1 1 4 5 2 5 3 1 4 1 1 1 1 1 4 1 Tuberculosis, Urinary Tract Infection Tuberculosis, Urinary Tract Infection, Respiratory Tract Infection 1 0 0 P48 1 1 1 5 5 3 1 1 1 3 1 1 1 5 1 5 1 3 1 2 1 4 3 1 2 1 1 1 1 1 3 3 Urinary Tract Infection, HIV/AIDS Urinary Tract Infection, HIV/AIDS, Tuberculosis 0 1 0 289 P49 1 3 3 1 1 1 1 1 1 1 1 5 2 1 1 1 1 1 1 1 1 1 5 5 5 1 4 1 1 1 4 1 Malaria Malaria 1 0 0 P50 1 4 5 1 1 1 1 1 1 1 1 2 5 1 1 1 1 1 1 1 1 1 3 3 3 1 2 1 1 1 2 1 Malaria Malaria 1 0 0 P51 1 1 1 1 3 1 3 1 1 1 1 1 1 1 1 3 4 5 2 5 1 2 1 1 5 2 1 1 1 1 3 1 Urinary Tract Infection, Respiratory Tract Infection Malaria, Urinary Tract Infection, Respiratory Tract Infection 1 0 0 P52 1 1 1 5 2 2 1 1 1 2 1 1 1 3 1 3 1 3 1 3 1 2 3 1 4 1 1 1 1 1 2 3 Urinary Tract Infection, HIV/AIDS Urinary Tract Infection, HIV/AIDS 1 0 0 P53 1 5 2 1 5 1 1 1 1 1 1 4 2 1 1 5 1 1 1 4 1 4 3 3 4 1 3 1 1 1 3 1 Malaria, Urinary Tract Infection Malaria, Urinary Tract Infection 1 0 0 P54 1 1 1 1 3 1 5 1 1 1 1 1 1 1 1 2 3 5 5 2 1 5 1 1 4 3 1 1 1 1 2 1 Respiratory Tract Infection, Urinary Tract Infection Respiratory Tract Infection, Urinary Tract Infection 1 0 0 P55 5 3 1 4 5 1 1 3 1 1 1 1 1 1 3 4 1 1 1 4 1 5 1 1 5 1 1 3 4 5 3 1 Typhoid Fever, Urinary Tract Infection Malaria, Urinary Tract Infection 0 1 0 P56 1 2 3 1 1 1 3 1 1 1 1 3 3 1 1 1 4 4 3 1 1 1 2 4 4 5 4 1 1 1 3 1 Respiratory Tract Infection, Malaria Respiratory Tract Infection, Malaria 1 0 0 P57 1 1 1 1 1 1 3 1 3 1 4 1 1 4 1 1 4 5 3 1 5 1 4 1 2 4 1 1 1 1 1 1 Tuberculosis, Respiratory Tract Infection Tuberculosis, Respiratory Tract Infection 1 0 0 P58 1 1 1 1 5 1 3 1 1 1 1 1 1 1 1 3 4 2 3 4 1 3 1 1 2 3 1 1 1 1 3 1 Urinary Tract Infection, Respiratory Tract Infection Malaria, Urinary Tract Infection, Respiratory Tract Infection 1 0 0 P59 1 1 1 1 1 1 3 1 1 1 1 1 1 1 1 1 4 3 2 1 1 1 1 1 3 3 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P60 4 5 1 2 1 1 1 3 2 1 4 1 1 3 3 1 1 1 4 1 4 1 3 1 1 1 1 4 4 3 4 1 Tuberculosis, Typhoid Fever Malaria, Typhoid Fever, Respiratory Tract Infection 0 1 0 P61 1 1 1 2 1 5 1 1 5 3 4 1 1 3 1 1 1 3 3 1 3 1 3 1 1 1 1 1 1 1 1 3 HIV/AIDS, Tuberculosis HIV/AIDS, Tuberculosis, Respiratory Tract Infection 1 0 0 P62 1 1 1 5 1 3 2 1 1 4 1 1 1 2 1 1 2 3 2 1 1 1 5 1 5 2 1 1 1 1 1 5 HIV/AIDS, Respiratory Tract Infection HIV/AIDS, Respiratory Tract Infection 1 0 0 P63 1 5 4 1 1 1 1 1 1 1 1 4 5 1 1 1 1 1 1 1 1 1 3 2 2 1 3 1 1 1 2 1 Malaria Malaria 1 0 0 P64 1 1 1 1 1 1 3 1 1 1 1 1 1 1 1 1 5 5 4 1 1 1 1 1 4 4 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P65 1 4 2 1 1 1 1 1 1 1 1 4 3 1 1 1 1 1 1 1 1 1 4 2 2 1 3 1 1 1 2 1 Malaria Malaria 1 0 0 P66 1 4 4 1 1 1 1 1 1 1 1 3 3 1 1 1 1 1 1 1 1 1 5 4 3 1 3 1 1 1 5 1 Malaria Malaria 1 0 0 P67 1 5 2 1 1 1 1 1 1 1 1 2 5 1 1 1 1 1 1 1 1 1 3 4 4 1 4 1 1 1 4 1 Malaria Malaria 1 0 0 P68 1 3 3 1 1 1 1 1 1 1 1 5 5 1 1 1 1 1 1 1 1 1 2 5 3 1 5 1 1 1 2 1 Malaria Malaria 1 0 0 P69 2 4 1 2 1 1 1 4 1 1 1 1 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 5 2 4 5 1 Typhoid Fever Malaria, Typhoid Fever 0 1 0 P70 1 1 1 5 1 3 1 1 1 4 1 1 1 2 1 1 1 2 1 1 1 1 5 1 1 1 1 1 1 1 1 3 HIV/AIDS HIV/AIDS 1 0 0 P71 1 2 2 1 1 1 1 1 1 1 1 5 2 1 1 1 1 1 1 1 1 1 5 2 3 1 4 1 1 1 5 1 Malaria Malaria 1 0 0 P72 1 1 1 3 1 5 1 1 4 2 4 1 1 5 1 1 1 5 5 1 5 1 3 1 1 1 1 1 1 1 1 4 Tuberculosis, HIV/AIDS Tuberculosis, HIV/AIDS, Respiratory Tract Infection 1 0 0 P73 1 5 5 1 1 1 1 1 1 1 1 5 3 1 1 1 1 1 1 1 1 1 2 2 2 1 4 1 1 1 3 1 Malaria Malaria 1 0 0 P74 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 4 5 5 1 1 1 1 1 2 3 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P75 1 1 1 1 1 1 1 1 5 1 5 1 1 4 1 1 1 1 5 1 3 1 2 1 1 1 1 1 1 1 1 1 Tuberculosis Tuberculosis, Respiratory Tract Infection 1 0 0 P76 1 1 1 2 1 4 1 1 1 4 1 1 1 4 1 1 1 5 1 1 1 1 3 1 1 1 1 1 1 1 1 2 HIV/AIDS HIV/AIDS 1 0 0 P77 1 1 1 1 3 1 1 1 1 1 1 1 1 1 1 3 1 1 1 5 1 5 1 1 2 1 1 1 1 1 2 1 Urinary Tract Infection Urinary Tract Infection 1 0 0 P78 1 1 1 2 1 5 1 1 2 3 2 1 1 5 1 1 1 2 4 1 5 1 3 1 1 1 1 1 1 1 1 2 HIV/AIDS, Tuberculosis HIV/AIDS, Tuberculosis, Respiratory Tract Infection 1 0 0 P79 1 1 1 1 3 1 3 1 1 1 1 1 1 1 1 4 4 2 4 3 1 3 1 1 5 5 1 1 1 1 2 1 Respiratory Tract Infection, Urinary Tract Infection Respiratory Tract Infection, Urinary Tract Infection 1 0 0 P80 1 2 4 1 1 1 1 1 1 1 1 5 2 1 1 1 1 1 1 1 1 1 2 2 2 1 4 1 1 1 2 1 Malaria Malaria 1 0 0 P81 1 1 1 1 1 1 2 1 1 1 1 1 1 1 1 1 4 3 3 1 1 1 1 1 4 2 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P82 1 4 4 1 1 1 4 1 1 1 1 2 2 1 1 1 3 3 2 1 1 1 4 4 3 5 3 1 1 1 2 1 Malaria, Respiratory Tract Infection Malaria, Respiratory Tract Infection 1 0 0 P83 4 5 1 3 1 2 1 3 1 3 1 1 1 5 5 1 1 2 1 1 1 1 4 1 1 1 1 3 5 2 2 2 HIV/AIDS, Typhoid Fever HIV/AIDS, Typhoid Fever, Malaria, Tuberculosis 0 1 0 P84 1 1 1 1 4 1 1 1 1 1 1 1 1 1 1 3 1 1 1 2 1 5 1 1 2 1 1 1 1 1 5 1 Urinary Tract Infection Urinary Tract Infection, Malaria 0 1 0 P85 1 1 1 1 1 1 5 1 1 1 1 1 1 1 1 1 2 2 5 1 1 1 1 1 5 2 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P86 3 5 1 4 4 1 1 3 1 1 1 1 1 1 5 3 1 1 1 3 1 3 1 1 2 1 1 5 5 4 4 1 Urinary Tract Infection, Typhoid Fever Urinary Tract Infection, Malaria 0 1 0 P87 2 2 1 4 1 1 1 3 1 1 1 1 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 2 3 5 3 1 Typhoid Fever Typhoid Fever, Malaria 0 1 0 P88 4 3 1 3 1 1 1 3 1 1 1 1 1 1 4 1 1 1 1 1 1 1 1 1 1 1 1 2 4 4 3 1 Typhoid Fever Typhoid Fever, Malaria 0 1 0 P89 1 1 1 1 1 1 1 1 2 1 4 1 1 5 1 1 1 1 4 1 2 1 3 1 1 1 1 1 1 1 1 1 Tuberculosis Tuberculosis, Respiratory Tract Infection 1 0 0 P90 1 2 3 1 1 1 3 1 1 1 1 2 3 1 1 1 3 3 2 1 1 1 5 4 3 2 5 1 1 1 5 1 Malaria, Respiratory Tract Infection Malaria, Respiratory Tract Infection 1 0 0 P91 2 5 1 3 1 3 1 2 1 2 1 1 1 4 3 1 1 3 1 1 1 1 3 1 1 1 1 3 5 5 4 4 Typhoid Fever, HIV/AIDS Typhoid Fever, HIV/AIDS, Malaria, Tuberculosis 0 1 0 P92 3 2 1 3 2 1 1 5 1 1 1 1 1 1 2 3 1 1 1 5 1 5 1 1 3 1 1 5 4 2 5 1 Urinary Tract Infection, Typhoid Fever Urinary Tract Infection, Malaria 0 1 0 P93 1 1 1 1 1 1 1 1 5 1 4 1 1 3 1 1 1 1 5 1 2 1 5 1 1 1 1 1 1 1 1 1 Tuberculosis Respiratory Tract Infection 1 0 0 P94 1 4 4 1 1 1 1 1 2 1 2 4 2 5 1 1 1 1 3 1 4 1 5 4 3 1 3 1 1 1 3 1 Malaria, Tuberculosis Malaria, Tuberculosis 1 0 0 P95 1 1 1 2 1 3 1 1 1 2 1 1 1 2 1 1 1 2 1 1 1 1 3 1 1 1 1 1 1 1 1 3 HIV/AIDS HIV/AIDS 1 0 0 P96 1 1 1 1 1 1 4 1 3 1 4 1 1 3 1 1 2 4 3 1 4 1 4 1 5 4 1 1 1 1 1 1 Respiratory Tract Infection, Tuberculosis Respiratory Tract Infection 1 0 0 P97 1 4 3 1 1 1 3 1 1 1 1 5 3 1 1 1 2 3 4 1 1 1 3 4 5 3 3 1 1 1 4 1 Respiratory Tract Infection, Malaria Respiratory Tract Infection, Malaria 1 0 0 P98 1 1 1 1 1 1 5 1 1 1 1 1 1 1 1 1 5 5 3 1 1 1 1 1 5 5 1 1 1 1 1 1 Respiratory Tract Infection Respiratory Tract Infection 1 0 0 P99 5 5 1 3 1 1 1 5 1 1 1 1 1 1 5 1 1 1 1 1 1 1 1 1 1 1 1 3 2 4 3 1 Typhoid Fever Typhoid Fever, Malaria 0 1 0 72 26 1