Full text
Corresponding author: Mani Abedini Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0. A multi-agent continual learning framework for skin cancer detection leveraging crowdsourced dermoscopic images Mani Abedini * Head of Data, AI and Analytics, AW Rostamani, UAE. World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 Publication history: Received on 25 June 2025; revised on 02 August 2025; accepted on 05 August 2025 Article DOI: https://doi.org/10.30574/wjarr.2025.27.2.2863 Abstract Skin cancer represents one of the most common malignancies globally, making early detection crucial for effective treatment and improved patient outcomes. While dermatologists typically rely on dermoscopy and clinical examinations for diagnosis, recent advances in artificial intelligence, specifically deep learning techniques like convolutional neural networks (CNNs), have shown significant promise in automating skin lesion classification [1,2]. Although CNN models trained on benchmark dermatological datasets such as HAM10000 and ISIC have demonstrated diagnostic accuracies comparable to expert dermatologists, their effectiveness declines when faced with evolving realworld data distributions, a phenomenon known as concept drift [3,4]. To address the limitations associated with static AI models, this paper proposes a novel multi-agent deep learning framework designed for continual learning and adaptive skin lesion diagnosis. The architecture begins with multiple agents trained on trusted expert-annotated datasets, each subsequently specialized by continuous fine-tuning using distinct streams of dermatological images sourced from teledermatology platforms and social media. These crowdsourced datasets capture emerging dermatological conditions, varied imaging technologies, and diverse patient demographics, providing valuable but noisy real-world data. Crucially, the system includes a centralized Supervisor Agent responsible for periodically evaluating the performance of each specialized agent. Once annotated, these validated cases enrich the training datasets, enabling agents to continually adapt to new clinical trends and maintain robust diagnostic accuracy over time. The proposed multi-agent architecture thus integrates continual learning, domain adaptation, and expert oversight, effectively addressing concept drift and advancing practical, scalable AI-driven diagnostic support in dermatology. Keywords: Skin Cancer Detection; Computer vision; Skin Cancer Classification; Image Processing; Deep Learning; Concept Drift; Continual Learning; Domain Adaptation; Multi-Agent Systems; Crowdsourced Data 1. Introduction Excessive exposure to Sunlight’s ultraviolet (UV) rays or other sources of UV rays can damage the DNA of the external layer of the skin (epidermis) cells, resulting in abnormal cell growth and the formation of skin cancer [1,2]. Among many types of skin cancer, Basal Cell Carcinoma (BCC), Squamous Cell Carcinoma (SCC), Melanoma, and Merkel Cell Carcinoma (MCC) are most common types of cancers (Fig 1 shows some examples images of these type of skin cancers). A benign skin tumour is called nevus. Melanoma is the most dangerous type of skin cancer. It develops when melanocytes (the cells that give the skin pigment) start to grow out of control. Fortunately, Melanoma can be cured if detected early. Almost all nevi are not harmful, but some types can become Melanoma.
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 251 Basal Cell Carcinoma (BCC) Squamous Cell Carcinoma (SCC) Merkel Cell Carcinoma (MCC) Melanoma Figure 1 Example images of BCC, SCC, MCC and Melanoma skin cancers Dermatologists typically rely on dermoscopic and clinical examinations to diagnose skin lesions. However, access to expert dermatological evaluation is often limited, prompting the development of computer-aided diagnostic (CAD) tools. Advances in deep learning, particularly convolutional neural networks (CNNs), have significantly improved automated skin lesion classification. Models trained on large, expert-annotated datasets like HAM10000 and the ISIC archive have achieved dermatologist-level performance in classifying common lesion types [3,4]. Segmentation-based methods further enhance diagnostic accuracy by identifying lesion borders and supporting the ABCD dermoscopy rule—assessing asymmetry, border irregularity, color variation, and lesion diameter. These features are critical in differentiating benign from malignant lesions, and segmentation directly contributes to measuring them [8,9]. Despite these advancements, most AI systems remain static, trained once on curated datasets and unable to adapt to evolving real-world data distributions. In clinical practice, lesion characteristics vary across populations, imaging devices, and over time. Moreover, new dermatological conditions can emerge, challenging the generalization of fixed models. Continuously retraining on new data is resource-intensive and not scalable, and domain shift can significantly degrade model performance. In parallel, the proliferation of dermatological images on social media and teledermatology platforms provides a rich, albeit noisy, stream of annotated and unannotated data. Many dermatologists share dermoscopic or clinical images along with diagnostic commentary in online forums, representing a dynamic source of real-world skin lesion examples. Recent work has demonstrated that self-supervised learning on such data—combined with minimal expert-labeled samples—can improve generalization and enable rapid adaptation to new disease categories [57,58]. To address the limitations of static models and leverage community-driven data, we propose a novel multi-agent deep learning framework for continual skin lesion classification. Our architecture begins with agents pre-trained on expertvalidated datasets and fine-tunes them on domain-specific data streams from platforms such as Reddit, Twitter, and dermatology forums. A centralized Supervisor Agent periodically evaluates agents on a reference validation set and ranks their performance. Top-performing agents form a committee that identifies uncertain or novel cases for expert annotation via active learning. Validated cases are added back to the training set, enabling the agents to learn continually and remain aligned with current clinical trends. Our contributions can be summarized as follows: • Multi-Agent Architecture for Continual Learning: We design a novel multi-agent system that enables continuous retraining on incoming dermatology images. To our knowledge, this is one of the first frameworks to harness crowd-sourced expert images in a structured, agent-based manner for skin lesion diagnosis. • Domain-Specific Expert Models: By fine-tuning separate agents on distinct data sources, the system performs implicit domain adaptation. Each agent learns specialized features suited to its source (addressing issues like different image capture devices or demographics), while sharing common foundational knowledge, thereby tackling domain shift challenges. • Supervisor and Committee for Quality Control: We introduce a supervisory mechanism to continually assess agent performance using both trusted and newly curated data. The top agents (committee) perform “query by committee” active learning, identifying cases where the model consensus is high to label the cases for next iteration of training. This ensures that noisy or mis-labelled data from the web do not corrupt the models.
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 252 The remainder of this paper is organized as a typical technical study. In the next Section, Background and Related Work, we review existing skin cancer image datasets, challenges of distribution shift, and approaches in incremental learning, ensemble models, and active learning in medical imaging. Section 3, describes the multi-agent architecture and learning algorithms in detail. It follows by our experiments and discussion. Finally, Conclusion section summarizes our key findings and the significance of continual learning systems in dermatology AI. 2. Related Works 2.1. Skin Lesion Datasets and Benchmark Models Early work on automated skin lesion classification was hampered by limited data – initial studies in the 1990s had only a few hundred images. The creation of larger public datasets has since enabled the training of deep CNNs. One milestone was the ISIC archive (International Skin Imaging Collaboration), which by 2018 hosted over 13,000 dermoscopic images from multiple sources. The ISIC archive aggregates cases with permissive licensing and standardized formats, making it a common benchmark for researchers. However, the ISIC data then (and even now) was dominated by melanocytic lesions (common moles and melanomas), with relatively fewer examples of other skin conditions [3]. Another influential dataset is HAM10000 (“Human Against Machine with 10,000 training images”), published in 2018 [4]. HAM10000 contains 10,015 dermoscopic images representing 7 diagnostic categories of pigmented lesions (including melanoma, various types of nevi, basal cell carcinoma, actinic keratosis, etc.). These images were collected over 20 years using different devices, and each case’s diagnosis was confirmed either by pathology (over 50% of cases) or expert consensus/follow-up. The diversity and quality of HAM10000 made it a valuable training set; it was used in the ISIC 2018 Challenge and has since been cited in thousands of studies. Leveraging widely used dermatological image datasets, numerous deep learning models have been proposed for skin cancer detection, particularly utilizing advanced convolutional neural network (CNN) architectures such as ResNet, EfficientNet, and Inception. These models consistently demonstrate robust performance in binary classification tasks (malignant versus benign lesions), frequently achieving area under the receiver operating characteristic curve (AUC) scores exceeding 0.90. However, performance on multiclass classification tasks—distinguishing multiple lesion categories—is somewhat lower but continues to improve through innovative techniques. Notably, ensemble methods, which combine outputs from multiple CNNs, have demonstrated superior accuracy due to their capacity to capture complementary feature representations. For instance, Halder et al. trained three separate CNN architectures on the HAM10000 dataset and employed a fuzzy rank-based ensemble strategy, achieving an accuracy of 95.14%, significantly outperforming individual models [59]. Over the past decade, deep learning has notably accelerated progress in computer-aided diagnosis (CAD) of skin cancer, primarily through CNN-based approaches applied to curated dermoscopic image collections. Early studies laid the groundwork for these methodologies; for example, Dorj et al. [16] proposed a hybrid approach combining AlexNet for feature extraction with Error-Correcting Output Codes Support Vector Machines (ECOC-SVM), achieving an accuracy of 94%. Subsequently, Ameri et al. [17,18] trained a basic CNN directly on the HAM10000 dataset without explicit lesion segmentation, reporting an accuracy of 84%. Further exploration included Mohapatra et al.’s [20] application of MobileNet to HAM10000, resulting in an accuracy of 80%, later improved by Chaturvedi et al. [21] to 83.1% via strategic hyperparameter tuning and image augmentation. Leveraging widely used dermatological image datasets, numerous deep learning models have been proposed for skin cancer detection, particularly utilizing advanced convolutional neural network (CNN) architectures such as ResNet, EfficientNet, and Inception. These models consistently demonstrate robust performance in binary classification tasks (malignant versus benign lesions), frequently achieving area under the receiver operating characteristic curve (AUC) scores exceeding 0.90. However, performance on multiclass classification tasks—distinguishing multiple lesion categories—is somewhat lower but continues to improve through innovative techniques. Notably, ensemble methods, which combine outputs from multiple CNNs, have demonstrated superior accuracy due to their capacity to capture complementary feature representations. For instance, Halder et al. trained three separate CNN architectures on the HAM10000 dataset and employed a fuzzy rank-based ensemble strategy, achieving an accuracy of 95.14%, significantly outperforming individual models [59]. Over the past decade, deep learning has notably accelerated progress in computer-aided diagnosis (CAD) of skin cancer, primarily through CNN-based approaches applied to curated dermoscopic image collections. Early studies laid the groundwork for these methodologies; for example, Dorj et al. [16] proposed a hybrid approach combining AlexNet for feature extraction with Error-Correcting Output Codes Support Vector Machines (ECOC-SVM), achieving an accuracy of 94%. Subsequently, Ameri et al. [17,18] trained a basic CNN directly on the HAM10000 dataset without explicit lesion
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 253 segmentation, reporting an accuracy of 84%. Further exploration included Mohapatra et al.’s [20] application of MobileNet to HAM10000, resulting in an accuracy of 80%, later improved by Chaturvedi et al. [21] to 83.1% via strategic hyperparameter tuning and image augmentation. Transfer learning has emerged as an effective strategy in skin lesion classification tasks. For example, Garcia et al. [29] utilized a ResNet model initially pretrained on large-scale, non-medical image datasets and subsequently fine-tuned it on dermoscopic images, achieving significant performance improvements. This approach highlights the value of leveraging general image features learned from extensive non-medical datasets as an efficient starting point for medical image analysis, thus reducing reliance on large labeled medical datasets for training from scratch. In addition to classification, skin lesion segmentation has gained increasing research attention due to its capability to enhance diagnostic accuracy by precisely identifying lesion borders. Accurate segmentation facilitates the extraction of morphological features essential to clinical heuristics like the ABCD dermoscopy rule. Benedetti et al. [34] applied InceptionResNetV2 to the HAM10000 dataset and obtained an accuracy of 78.9%. Hatice Catal Reis et al. [35] leveraged the GoogleNet CNN architecture across ISIC datasets (2018–2020), consistently achieving classification accuracies above 90%. Similarly, Bechelli et al. [36] explored multiple CNN architectures, including Xception, VGG16, and ResNet50, for binary classification (benign versus malignant) using both ISIC archive and HAM10000 datasets, demonstrating the robustness of CNN-based segmentation and classification approaches. More recently, advanced foundation models for segmentation have attracted significant interest due to their superior generalization capabilities. The Segment Anything Model (SAM), trained on the large-scale SA-1B dataset containing over 11 million images and 1 billion segmentation masks [37], has shown promising segmentation performance across diverse image domains. Hua et al. [38] demonstrated SAM’s potential for skin lesion segmentation tasks, notably improving segmentation accuracy on the HAM10000 dataset when combined with bounding-box prompts. Extending this approach, Ma et al. [39] retrained SAM specifically on medical images, resulting in MedSAM—a specialized model trained on more than 1.5 million medical image-mask pairs spanning multiple imaging modalities and cancer types. Recent studies by Abedini [8,9] have also validated the effectiveness of SAM and MedSAM in enhancing image classification tasks, further underscoring the promise of foundation models for general-purpose medical image analysis. 2.2. Continual and Incremental Learning in Medical Imaging The need for models that learn continually – incorporating new data without forgetting past knowledge – is widely recognized in medical imaging. Traditional ML pipelines are not built for this; they assume a one-time training on a fixed dataset. If retrained naïvely with new data, neural networks tend to forget what they learned previously (a phenomenon known as catastrophic forgetting). Continual learning (CL) algorithms aim to allow iterative learning on new data while retaining performance on old data. Strategies for CL include: (a) Rehearsal – retaining a buffer of old examples to intermix with new data during training (experience replay); (b) Regularization – adding terms to the loss that prevent important weights from changing too much (e.g., Elastic Weight Consolidation); (c) Dynamic architectures – expanding the model or using separate sub-networks for new tasks; and (d) hybrid approaches. In the context of skin lesion analysis, some recent efforts have emerged. For example, Andrade et al. explored incremental learning for classifying dermatological image modality (clinical vs dermoscopic) and found that an experience replay approach (with 500 stored images) maintained high accuracy (86%) with minimal forgetting [60]. Our approach aligns with the rehearsal and architectural approaches to continual learning. By maintaining multiple agents and not discarding the original training data, we ensure that new training iterations always have a core of “old” knowledge mixed in – i.e., each agent retains access to the base dataset (either by fine-tuning from a pre-trained base model or by explicitly including a subset of base images during retraining). The Supervisor Agent’s evaluation on a fixed reference set also helps detect any forgetting: if an agent’s accuracy on known classes drops, the supervisor can penalize its rank, signaling the need for corrective measures (such as reloading the base model weights or adjusting the training strategy). While static classifiers trained on these datasets demonstrate high accuracy under controlled conditions, performance often declines when applied to new or diverse image domains, highlighting the problem of domain shift. Katharina emphasized that even changes in imaging devices or patient demographics could degrade model generalization [57]. Domain adaptation methods, such as Domain Adversarial Neural Networks (DANN), have been used to mitigate this issue. Gilani et al. reported an 18.47% accuracy improvement on shifted domains using adversarial training techniques [58].
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 254 2.3. Crowdsourced Data and Self-Supervision Leveraging unlabeled or noisy-labeled data from the web is an emerging trend to improve AI models. In dermatology, millions of images are shared in online forums, but they come without guaranteed high-quality labels. A key challenge is how to make use of this wealth of data without being misled by errors. Recent advances in self-supervised learning (SSL) offer one solution: models can be pre-trained on unlabeled images to learn generalizable features, and then finetuned on smaller labeled datasets. The Digital Medicine study by Shen et al. exemplifies this: they collected a large set of unannotated images from health forums, used contrastive SSL to train a feature encoder, and then fine-tuned on a “coarsely” labeled set (where labels from forum users might be less reliable than expert diagnoses) [61]. The model achieved measurable success on a dermatologist-curated test set (45% top-1 accuracy across 22 conditions), and performance improved significantly (to ~50% accuracy) after filtering out noisy labels using a small set of trusted validation images. This filtering was done by clustering embeddings and removing images that were far from cluster centers or had inconsistent labels. Notably, they found that more data is not always better if the data is noisy – a cleaner subset of representative images yielded better accuracy than the full raw set. Their approach validates two important points for our work: First, unannotated images from online sources can indeed enhance model performance when used carefully. The improvement from 42% to 45% accuracy after self-supervised pre-training is evidence that unlabeled community data carries useful information that complements existing datasets. Second, some form of expert validation or cleaning is crucial – in their case, adding just 50 expert-validated images per category (1,100 images total for 22 diseases) raised accuracy by nearly 5 percentage points. In our framework, the committee of agents and Supervisor agent essentially fulfill this filtering/cleaning role by selecting which new images should be sent for expert validation. This is conceptually similar to active learning, where a model identifies examples for an oracle (human) to label that would most improve the model if answered. A classic strategy in active learning is Query by Committee, wherein an ensemble of models votes on unlabeled examples and the ones with highest disagreement are prioritized for labeling. Our committee of top agents emulates this, as disagreement among diverse agents likely indicates an ambiguous or novel case that needs a ground-truth check. Active learning has been applied in medical imaging to reduce annotation costs; for instance, some works integrate it with federated learning for skin lesions, allowing local models to request labels without sharing raw data. While our current scope doesn’t explicitly involve federated learning (all agents could be centrally located since social media data is public), the principle of distributed learning from different data sources is analogous. 2.4. Multi-Model and Agent-Based Systems Ensembles and multi-model systems have long been used to boost predictive performance, as mentioned earlier. Beyond performance, ensembles can also provide more reliable uncertainty estimates – if models unanimously agree, confidence in the prediction is higher; disagreement can signal uncertainty. This is particularly valuable in a safetycritical field like cancer diagnosis. Our multi-agent system can be seen as an ensemble that is distributed across time and data sources. Unlike a standard ensemble where all models train on the same data, here each model has a slightly different training history. This diversity may further enrich the ensemble’s decision-making. There is also a growing interest in agent-based AI systems in general. Recent works (e.g., M3Builder by Feng et al. 2025) have used multiple AI agents to collaborate on complex tasks like automating machine learning workflows. While those agents were orchestrated to handle different tasks (data prep, model training, etc.), our agents are homogeneous in task (all are classifiers) but heterogeneous in experience (each sees different data). The agent metaphor also raises the possibility of modular expansion – new agents can be added for new data streams without disturbing existing ones, and poorly performing agents could even be retired or replaced over time, making the system adaptive and scalable. In practice, there is often a gap when applying models to “images in the wild.” Factors like lighting, zoom, image quality, skin tone diversity, and lesion prevalence can differ greatly outside the curated dataset setting. For instance, model accuracy that is excellent on a test of clinic dermoscopic images might drop when faced with a smartphone photograph of a lesion or a rare subtype not seen in training. This mismatch between training data and real-world data – essentially a domain shift – has been documented as a cause of performance degradation. Katharina highlighted that even different clinics or devices create domain differences, and a classifier trained on one data source may not generalize optimally to another [57]. To address this, researchers have investigated domain adaptation techniques. One approach uses adversarial training to make the model’s feature representations invariant to the domain (source) of the image. Gilani et al. applied a Domain Adversarial Neural Network (DANN) to skin lesion data and achieved an 18.47% accuracy improvement over a baseline when testing on a shifted domain [58]. This underscores the value of adapting models to new distributions. Our proposed multi-agent system inherently performs multi-domain adaptation: each agent finetunes on images from a particular source, learning to compensate for that source’s biases (much like training separate models per domain).
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 255 In summary, the gaps in the current literature that we aim to fill are: (1) implementing a practical continuous learning pipeline for skin lesion classification that directly taps into the stream of data from online expert communities; (2) using a multi-agent (or multi-model) approach to handle domain differences and provide a robust way to decide when to trust new data and when to seek human input; and (3) demonstrating experimentally how such a system can maintain or improve accuracy over time, compared to a conventional static model. The next section details our proposed approach addressing these points. 3. Methodology 3.1. Dataset This study used two publicly available datasets containing Skin Cancer images. The first data set is HAM10000 which contains 10015 images of 7470 lesions, seven categories: Melanocytic Nevi, Melanoma, Benign Keratosis-like Lesions, Basal Cell Carcinoma, Pyogenic Granulomas and Hemorrhage, Actinic Keratoses and Intraepithelial Carcinomae, Dermatofibroma. The data set has border segmentation as well [40,41]. In our experiments we used this data set to train our object detection algorithm to identify the lesion location. The second data set is International Skin Imaging Collaboration (ISIC 2018) contains 2594 dermatologic images and associated ground truth segmentation masks [42,43]. Since 2016, ISIC has conducted annual challenges for the computer science community; since 2016 till today ISIC datasets become the largest publicly available collection of quality controlled dermoscopic images of skin lesions. The objective is to improve melanoma diagnosis crowdsourcing the AI and computer vision enhancement; ISIC is sponsored by the International Society for Digital Imaging of the Skin (ISDIS). 3.2. Pre-processing In computer vision, preprocessing is a critical step, especially for dermoscopic images collected from various clinics and different imaging setups. In our experiments first we applied digital hair removal (DHR) algorithm [44]. To avoid removing any critical patterns we avoid using any noise removal filter. All images are resized to 224 × 224 to be consistent with the input layer of our deep learning models. Image augmentation requires generating a good amount of annotated data so we can retrain the deep neural networks. Since annotating medical images requires to be conducted by experienced medical professionals, data acquisition takes time and very expensive. Generating more images from existing annotated data is the best cost-effective way to overcome the situation. In our experiments, all images were scaled with 1/255. We also allow random rotation between 0 and 45 degrees. The zoom level was between 0.5 and 2; numbers below 1.0 result in zooming out, and numbers bigger than 1.0 will magnify. We also allow random adjustment of brightness. The random noise in brightness will help the network be less sensitive to specific image brightness and try to learn the underlying patterns associated with cancer lesions.. 3.3. Proposed model 3.3.1. System Overview Our proposed system consists of several interacting components, as depicted previously in Figure 2. At a high level, it operates in cycles of training, evaluation, and data selection: Figure 2 High level architecture of Multi Agent Skin Cancer detection, training and continues learning capability
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 256 Base Model Initialization: We start with an Initial Model trained on a Trusted Base Dataset. For this study, we use the combined HAM10000+ISIC dataset as our base. These datasets provide a reliable foundation with expert-confirmed labels. The initial model could be any modern CNN or transformer-based classifier; in our experiments we chose a CNN (EfficientNet-B3) pre-trained on ImageNet and fine-tuned on the base dermoscopy data. This model achieves baseline performance comparable to literature on the base test set (85% balanced accuracy). This Base Model serves as the starting point for all agents, ensuring they all have a strong and consistent initial capability. Agent Deployment to Data Streams: We instantiate multiple copies of the base model to create Agent Models. Each agent is assigned to a particular data source (or a domain). In a real deployment, these sources could be: Agent A for a Twitter account that aggregates dermatologists’ case photos, Agent B for a Instagram dermatologist experts, Agent C for an online forum of a melanoma patient group in the web. In our experiments, we simulate two distinct sources by partitioning a dataset and adding different types of noise/variation to each, to mimic the effect of different “social media” conditions. Each agent continuously receives new images from its source. We assume that along with each image, there may be an initial label provided (for example, the dermatologist who posted it might say “diagnosis: dysplastic nevus” or a forum user might tag it as “melanoma”). These labels are not fully trusted – they are considered noisy labels (though likely more accurate than random, since many posters are experts). Each agent maintains a training dataset consisting of: (i) the original base data (at least a significant subset of it), and (ii) all new cases from its source that have accumulated, with their associated tentative labels. At regular intervals (e.g., weekly or monthly), the agent fine-tunes its model on the current dataset. Supervisor Agent Evaluation: After agents update their models, the Supervisor Agent evaluates them. The Supervisor has a fixed Validation Set that includes two kinds of images: (a) Core validation images from the base dataset (never used in training) to test that the agent still recognizes known classes correctly; and (b) Challenging new images collected from all sources, which have been identified (in previous cycles) as high importance and whose labels have been verified by human experts (more on how these are selected later). This combined validation set is used to compute an accuracy score (or other metrics like F1) for each agent. The Supervisor then ranks the agents by performance. Agents that perform poorly might be suffering from either forgetting or poor adaptation (or might be on a source with very noisy data). High-performing agents presumably have better generalization. We allow the possibility that some sources simply have more relevant data; those agents should rise to the top. Committee of Top Agents: We select the top-K agents (based on validation performance) to form a Committee. In our design, K is a small number (e.g., 3 out of maybe 5-10 agents). The committee is intended to pool the “best knowledge” currently available. Committee members share their newly acquired data among themselves – i.e., the union of new cases from their sources – and each member temporarily evaluates all those cases. They then compare their predictions. For each new image under consideration, the committee either reaches a consensus (all agents agree on the classification with high confidence) or they don’t. Consensus cases that all agents agree on and with high confidence can be accepted as reliably labeled (especially if the agents also agree with the source-provided label). These might be added to a global training set without further review, under the assumption that multiple independent models agreeing reduces the chance of error. On the other hand, controversial cases (where agents disagree or are uncertain) are flagged for human expert review. This is akin to a triage: easy, clear-cut cases are handled autonomously, while difficult cases are escalated to human specialists – a sensible approach in medical AI to ensure safety. Retraining and Knowledge Distillation: Periodically (say after several cycles), we can choose to distill or merge knowledge from the multiple agents. One way is to have a global model update: we could take the union of all data (base + all sources’ new verified data) and train a single model from scratch or fine-tune the base model anew. This global model could then replace the agents’ weights (or serve as a new initialization) if it proves more accurate. However, frequent global retraining might be costly; an alternative is knowledge distillation where the committee’s ensemble predictions on a large set of images are used to train a single model that approximates the ensemble. For now, our framework keeps agents separate and relies on the committee for consensus, leaving global retraining as an occasional maintenance step. 3.4. Evaluation Metrics To evaluate the performance of the proposed methods, four widely used metrics have been measured in our experiments: Accuracy (1), Precision (2), Recall (3), and F1-Score (4). Please see the formula for each metric below: 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝐹𝑃+𝑇𝑁+𝐹𝑁 (1)
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 257 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = 𝑇𝑃 𝑇𝑃+𝐹𝑃 (2) 𝑅𝑒𝑐𝑎𝑙𝑙 = 𝑇𝑃 𝑇𝑃+𝐹𝑁 (3) 𝐹1 − 𝑆𝑐𝑜𝑟𝑒 = 2 × 𝑅𝑒𝑐𝑎𝑙𝑙 ×𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 𝑅𝑒𝑐𝑎𝑙𝑙+𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 (4) TP (True Positive) represents the number of correctly predicted positive cases. TN (True Negative) represents the number of correctly predicted negative cases. FP (True Positive) represents the number of incorrectly predicted positive cases. FN (True Negative) represents the number of incorrectly predicted negative cases. Accuracy is the most commonly used measurement of the accurate model, defined as the portion of actual positive and negative cases over all measured cases. Precision defines the portion of positive predictions that have been correctly identified. Recall, also known as sensitivity or true positive rate, on the other hand, measures the portion of actual positive area that has been correctly predicted as positive. F1-score is the harmonic mean of the precision and recall. 3.5. Model Architecture and Training Details Our implementation uses EfficientNet-B3 as the base CNN architecture, chosen for its good accuracy-efficiency tradeoff for image classification. The model outputs probabilities for each skin lesion category (we focus on melanoma vs nevus vs BCC vs etc., a multi-class problem). We train the base model on HAM10000+ISIC data (approximately 18,000 images total after combining and cleaning overlaps) for 50 epochs, achieving a strong baseline performance (85% balanced multi-class accuracy across seven classes). This base model is saved and then copied to create agent models. New images from streams are initially preprocessed: if they are non-dermoscopic clinical photos (for example, a smartphone picture of a mole on skin), we found it beneficial to apply color normalization and center-cropping to mimic dermoscopic appearance. In a real scenario, one might have separate models for dermoscopic vs clinical images (since they look different), but to keep things simple, we feed all images to the same model architecture and rely on it to generalize. Data augmentation (random flips, zooms, etc.) is applied during training to each agent to account for variability especially in the social media images, which often have diverse backgrounds and lighting. 3.5.1. Handling Noisy Labels and Class Imbalance The labels that come with community-contributed images can be noisy. For example, if a dermatologist posts a case and says “diagnosis: likely seborrheic keratosis,” there is some chance it was actually a melanoma (maybe later confirmed by biopsy). To mitigate learning from incorrect labels, we incorporate a label smoothing approach: when fine-tuning on new data, instead of using the provided label as a hard target, we allow a small probability (e.g., 10%) to be distributed among other classes. This way, if a few labels are wrong, the model doesn’t overly trust them. Another issue is class imbalance. Rare skin cancers (like dermatofibrosarcoma protuberans) might be nearly absent in the base data but could appear in new streams (perhaps exactly because experts share rare cases online disproportionately). Agents might then see a class in their fine-tuning that wasn’t in their initial training, causing a new class introduction problem. We address new classes by expanding the output layer of the model as needed (this is an architectural approach to continual learning – adding new neurons for new classes). The Supervisor can detect if an agent encounters images with labels that were not originally in the model’s classes; the agent will then add those classes and initialize weights for them (we use the base model’s final layer’s biases to guess an initial prior, e.g., very low prior probability for a new class. We maintain class frequency counts and apply dynamic resampling – if an agent’s dataset grows very skewed (which can happen if, say, one source shares mostly melanoma cases), we downsample the majority class or up-weight losses for minority classes to prevent bias.
World Journal of Advanced Research and Reviews, 2025, 27(02), 250-263 258 3.5.2. Committee Decision Process The Top Agents Committee is a crucial component for governance of the system. We set the committee size K = 3 in our tests, meaning the top 3 agents (by validation accuracy) form the committee. During committee voting on new images, we use entropy of the mean prediction as a measure of uncertainty. One might wonder: does the committee ever get it wrong unanimously? It is possible that all agents share a blind spot (since they all originate from the same base model). To catch systematic errors, we ensure the validation set includes some known difficult cases and out-of-distribution examples. If all agents misclassify certain types of lesions, their validation scores suffer, preventing them from all being top-ranked. In future work, an idea could be to include a “devil’s advocate” agent with a different architecture or training paradigm specifically to inject diversity in the committee. In our current design, the diversity comes mainly from the data each agent sees, which we found sufficient to create occasional differences in opinion among agents. 3.5.3. Scalability and Deployment Considerations Our multi-agent approach is inherently parallelizable. Each agent’s training can be done on separate hardware (or sequentially if resources are limited, since agents update on different data). The communication between agents and the supervisor is minimal – just sending evaluation metrics and possibly model checkpoints. For a real-world deployment, one could imagine each hospital or each social media channel running its own agent locally (a bit like federated learning nodes), and a central server acting as the supervisor and integration point for knowledge. Privacy is less of a concern here than in typical federated learning, because most data we consider is openly shared by users (with patient consent assumed, since these are often cases shared for educational purposes). Nevertheless, any patientidentifiable information should be stripped from images and metadata. A potential ethical consideration is that if our system scrapes images from the web, we must ensure it only uses images that were shared publicly and ideally with permission for reuse in research. Collaborating with dermatology communities for data sharing agreements would be ideal to maintain ethical standards. Another practical aspect is the evaluation frequency. We do not want to overburden human experts with constant queries. Thus, the system could be configured to accumulate new data and only perform the committee active learning step once enough new images have gathered or at fixed intervals (e.g., monthly). Also, not every cycle must involve human experts – if the committee finds few or no contentious cases, it can proceed without human input. The threshold for uncertainty can be tuned to control how many queries are made to dermatologists, balancing model autonomy vs. expert oversight. In summary, our methodology provides a blueprint for a continuously improving skin lesion classifier system. By combining techniques from continual learning, ensemble methods, active learning, and domain adaptation, it aims to remain accurate amidst evolving data. The next section describes how we set up experiments to validate this approach in a controlled setting. 3.6. Experiments Designing a rigorous evaluation for a continual learning system is challenging because it involves time-varying data. We describe here the datasets and protocol we used to simulate the scenario of multiple agents learning from separate “social media” streams, along with the metrics used to measure performance. Base Datasets: We used two well-known datasets as the source of trusted data: HAM10000 and the ISIC 2019 Challenge Dataset (which includes ISIC archive images up to 2019). We combined these and removed duplicates (there is some overlap of HAM10000 images in ISIC). The final base training set had 18,384 images across 7 diagnostic categories: Melanoma (1113 images), Melanocytic Nevus (~10,000), Basal Cell Carcinoma (3323), Actinic Keratosis/Bowen’s Disease (867), Dermatofibroma (239), Vascular lesion (253), and Benign Keratosis (2589). The class distribution is imbalanced (e.g., few dermatofibromas), but it reflects real frequencies. We set aside 20% of this dataset as an initial validation and test set (not used in base training). These held-out sets serve to measure base model performance and also form part of the Supervisor’s validation pool (specifically, 500 images were used as the core validation for supervisor scoring, stratified by class). All images are 3-channel color; dermoscopic images were provided at varying resolutions (we resized them to 224×224 for EfficientNet input). Simulated Social Media Streams: To emulate multiple streams of incoming cases, we took additional data from two sources: (1) the Dermatologist’s Instagram Set, and (2) the DermWeb Clinical Images Set. The first is a collection we curated of 500 images shared by dermatologists on public Instagram accounts (mostly dermoscopic images of