Introducing AI practitioners to international standards through online competitions
Abstract
This Comment analyzes a global online landslide classification challenge to examine whether AI competitions can help participants engage with international standards. The results show that while technical innovation was strong, awareness of standards on transparency, sustainability, and bias was limited. The prototype used for this competition can serve as a scalable mechanism to foster the responsible use of AI for disaster risk reduction and beyond.
Full text
npj | naturalhazards Comment https://doi.org/10.1038/s44304-025-00154-8 Introducing AI practitioners to international standards through online competitions Lorenzo Nava, Monique M. Kuglitsch, Jennifer Selby, Amy Bray, Paul Kennedy, Joel Lontsi, Megan Yates, Ahmed Tambal, Ahmed Ilmane, Andrea Manara, Thomas Basikolo, Katharina Weitz, Arif Albayrak, Chester Karwatowski, Rakiya Babamaaji, Silvia Garcia, Elena Xoplaki, Stefan Uhlenbrook, PierrePhilippe Mathieu, Raffaele Albano, Andrea Toreti & Filippo Catani Check for updates This Comment analyzes a global online landslide classification challenge to examine whether AI competitions can help participants engage with international standards. The results show that while technical innovation was strong, awareness of standards on transparency, sustainability, and bias was limited. The prototype used for this competition can serve as a scalable mechanism to foster the responsible use of AI for disaster risk reduction and beyond. Online artificial intelligence (AI) challenges are structured competitions that invite participants to apply machine learning (ML) methods to solve realworld problems using curated datasets. These challenges attract a wide spectrum of participants, from university students to researchers and senior AI engineers, offering a bottom-up approach to training, which democratizes access to data for specific problem statements. They also foster an interactive, community-driven environment that encourages iterative experimentation, peer learning, and innovation. Some platforms that host online challenges focus on engaging participants in specific regions, such as the Global South, while others encourage global collaboration. In both cases, data and problem statements are typically provided by clients (e.g., universities, researcher organizations, or industry), ensuring that competitions remain grounded in practical applications (e.g., identifying natural hazards for disaster management). A leaderboard plus attractive prizes motivate participants to refine their models iteratively while encouraging both competition and collaboration. Because the success of these challenges is typically assessed through quantitative model performance metrics, such as accuracy, other crucial dimensions of AI (e.g., transparency, reproducibility, and long-term social relevance) can be overlooked. As a result, online challenges may risk reinforcing a culture of technical optimization detached from broader considerations of responsible and grounded in reality AI. This emphasis on performance reflects a broader divide between experimental, bottom-up AI development and formal, top-down mechanisms of governance. International standards, for example, aim to ensure reliability, safety, and fairness in AI systems, yet they often remain disconnected from the environments where most practitioners actually learn and build models. In addition, rather than focusing on students and engineers, they are usually catered to policymakers. Adding to the challenge, some of the practitioners might be self-taught because of a broader systemic education and training gap. These documents, which must be created by an international standard developing organization to receive this designation, are critical for: aligning national policies and harmonizing how technologies are used, reducing risk of misapplication, and for stocktaking the use of technology, including national or regional gaps in research and capacities1,2. Although international standards offer guidance on many important concepts, from data bias and privacy to model transparency and trustworthiness, many AI developers are only confronted by them once they are embedded in national policy and the AI developer is producing a tool intended for operational (market) use. However, standards are not just regulatory checkboxes; when observed, they can help prevent real-world harms, build user trust, and improve transferability and interoperability2. Given the importance of these concepts, we believe that anyone who is applying AI should be familiar with them. To enhance their adoption, we propose that online challenges can serve a dual purpose: providing bottomup, hands-on training while introducing AI developers to international standards and cultivating responsible AI. By encouraging AI developers to follow a common workflow and to observe best practices, it is easier to compare the resulting solutions. This Comment explores the potential of online challenges as a standards-fostering mechanism. Through the outcomes of an online landslide challenge that was run from April to August 2025, coordinated by the Global Initiative on Resilience to Natural Hazards through AI Solutions, in partnership with the International Telecommunication Union (ITU), the AI for Good, the European Space Agency (ESA), the World Meteorological Organization (WMO), the Cambridge Complex and Multihazard Research Group (CoMHaz) of the University of Cambridge, and theMachine Intelligence and Slope Stability Lab (MISSLab) of the University of Padua, we propose actionable recommendations to enhance the accessibility of international standards to AI developers. In conjunction, these recommendations form the basis for a global capacity building mechanism that can assist those developing online challenges and hoping to foster the responsible use of AI for disaster risk reduction and beyond. Online challenges: innovation through competition Since 2020, the ITU has organized online challenges3dedicated to advancing the Sustainable Development Goals (SDGs) and drawing over 12,000 participants. As of 2022, the ITU’s AI for Good Summit and the Zindi platform have co-organized such challenges. Governed by platform-wide Terms and Conditions as well as challenge-specific rules that participants must accept npj Natural Hazards | (2025) 2:102 1 1234567890():,; 1234567890():,;
before accessing datasets or submitting solutions, the design of each online challenge also requires careful attention to data security, fairness, and ethical risk. Thus, although those participating in the challenge might not be cognizant of international standards and policy, the platform is already observing AI governance. To explore the potential of online challenges as a standards-fostering mechanism, we organized an online challenge related to the application of AI to respond to landslide disasters. A landslide is the movement of a mass of rock, earth, or debris down a slope under the influence of gravity. Large landslides can destroy or bury entire villages and pose risks to an estimated 8% of the global population4. These hazards occur worldwide in mountainous regions, with particularly high frequency and impact in South Asia along the Himalayan belt, as well as in East Asia, Southeast Asia, and Central and South America5. Compared with other natural hazards, landslides are commonly addressed with AI in the academic literature6, particularly in the areas of landslide susceptibility mapping for land-use planning7, detection for disaster response7, and forecasting for earlywarning systems8. In this online challenge, participants of varying experience levels were asked to develop models that distinguish between satellite images with and without landslides (Fig. 1). The dataset combined Sentinel-2 Level-2 optical imagery (red, green, blue, and near-infrared bands) with Sentinel-1 Ground Range Detected Synthetic Aperture Radar (SAR) data from both ascending and descending passes. Each SAR image included VV and VH polarizations and their temporal differences (ΔVV, ΔVH). Optical imagery offers interpretable spectral information, while SAR provides observations through cloud cover and at night, though typically with lower spatial accuracy9. The task therefore focused on designing methods that exploit the complementary strengths of both data types, using SAR where optical data are unavailable and prioritizing optical inputs where cloud-free conditions allow more accurate detection. Participants were free to use any methods they deemed suitable. A starter notebook and dataset was provided with resources to guide data exploration, processing, and analysis. Multiple submissions (up to 300) were allowed and, although only the final results were logged, conservative estimates suggest each entry required at least three hours of work. For many individualsand teams, this added up tomore than 300 hours over the course of the competition. International standards: guidelines to support responsible AI Despite the many opportunities AI offers, it is essential that national policies safeguard end users by promoting the responsible use of AI and mitigating the risks when systems fail to perform as expected. Furthermore, when national policies are aligned with those of neighboring countries, they can foster interoperability and reduce barriers to cooperation. Here, international standards, documents created by an international standard developing organization [such as the ITU, the International Organization for Standardization (ISO), or International Electrotechnical Commission (IEC)] and containing policy-informing guidelines, provide a key mechanism for this alignment. Complementing this work are the efforts of organizations such as the World Meteorological Organization (WMO) or UNESCO who have developed technical regulations and manuals to guide those using AI in disaster management. Since 2020, the ITU has led efforts to develop international standards at the intersection of AI and disaster risk management. From 2020 to 2024, the ITU/WMO/UNEP Focus Group on AI for Natural Disaster Management, a community of experts and stakeholders coming from the geosciences, computer sciences, social sciences, policy, first and second response, and the humanitarian space, began to lay the groundwork for such standards, publishing technical reports on topics such as data collection and preprocessing, modeling training and evaluation, and operational implementation10. In 2024, the Focus Group transitioned into the Global Initiative on Resilience to Natural Hazards through AI Solutions and established additional partnerships with UN Climate Change (UNFCCC), the Universal Postal Union (UPU), and United Nations Educational, Scientific and Cultural Organization (UNESCO). To align the landslide challenge with international standardization efforts, the best-performing models were required to provide details on how they observed best practices, drawn from the Focus Group, in their methods. This enabled us to compare the solutions in terms of performance and observation of best practices. Weighted scores were provided for each of the sixcategoriesasshowninFig.2. Fig. 1 | Conceptual framework of the landslide image classification challenge. The figure illustrates the integration of Sentinel-1 SAR and Sentinel-2 optical imagery within an AI-based workflow to support landslide image classification. Sources: Sentinel-1, Sentinel-2, Airbus 2025. npj | naturalhazards Comment npj Natural Hazards | (2025) 2:102 2
Thinking outside the box: getting innovators excited about standards In terms of participation, this was the largest AI for Good–Zindi challenge with 972 contributors from over 90 countries. Most participants were from India. Other highly represented countries included Nigeria, Kenya, China, and South Africa. Despite the geographical diversity, there was clear gender disparity, which is consistent with the participation in other AI challenges from the same platform. Although it would be interesting to better understand this disparity, it is unfortunately outside of the scope of this Comment. In total, participants submitted over 8,000 iterations of machine learning solutions, equating to ca. 27,000 hours of collective development effort. The three winning teams adopted ensemble-based deep learning approaches that creatively combined different input representations to capture the spectral and textural signatures of landslides. They explored various image resolutions and input band configurations, leveraging both optical and SAR data to enhance feature diversity and model robustness. In practice, this included dual-branch architectures that processed Sentinel-2 optical bands and Sentinel-1 polarizations separately before merging features through attention mechanisms or learned weighting schemes. Several teams applied modality dropout or cloud-mask-based gating to adaptively balance the contribution of each sensor. Others improved generalization through MixUp and CutMix augmentations, pseudo-labeling of the test set, and cross-fold ensembling of multiple backbones such as EfficientNet, ConvNeXt, and Swin Transformer. The strength of the ensemble itself proved to be one of the key factors in achieving high accuracy, supported by thoughtful preprocessing that reduced SAR speckle, normalized optical reflectance, and exploited temporal differences to maximize the complementarity between data sources. The solution details provided by the participants varied in format, including figures (e.g., visualisation of the explainability generated by XAI methods), text, tables, and code snippets. Overall, categories “Innovation” and “Practicality & Robustness”received the most consideration by the contributors compared to the other categories. In the following, we summarise the evaluators impression of the quality of the participant documentation for each category (Fig. 2): 1. Data & Model Bias: Many critical aspects were overlooked, particularly the various forms of data and model bias that are crucial in disasterrelated AI research. Ignoring systemic, human, or statistical biases can lead to inequitable risk assessments and models that perform well in data-rich regions but fail disastrously in vulnerable or data-poor areas. 2. Model Transparency: While five of the seven finalists reported on the usage of XAI methods, the description of the meaningful insights lacked in-depth reflection. 3. Approach Reusability: Most contestants developed modular frameworks capable of integrating geospatial data to improve model efficiency and quality. This design choice can enhance transferability across disaster types, geographic regions, and contexts with limited data availability. 4. Sustainability & Efficiency: Four of the seven finalists reported measures to reduce training time, and two quantified their energy and carbon usage using tools such as CodeCarbon. While this indicates growing awareness, systematic evaluation of environmental impact remains limited. 5. Innovation: Teams proposed multi-model ensemble strategies with key differences in preprocessing, especially in how they handled optical and SAR inputs. These creative designs achieved outstanding results in capturing the distinct spectral and textural signatures of landslides. 6. Practicality & Robustness: The developed methods proved highly robust and achieved impressive performance. However, they tended to be computationally expensive, some requiring up to three hours to generate predictions on the test set. Lessons learned The high level of engagement with the online challenge highlighted the global interest in applying AI to address natural hazards. Overall, the performance metrics achieved by the models were generally higher than those reported in the literature for similar tasks10. Although the ensembles differed in design, each demonstrated a distinctive and original approach. These approaches contrast with the prevailing trend in the literature, where singlemodel solutions are often preferred due to their faster deployment and greater suitability for low-resource environments. These, however, generally achieve lower performance to the approaches used by the winners of this challenge. As described above, participants were encouraged to consult and integrate relevant standards when designing their approaches. To support this, we provided them with access to the standards and asked them to answer a set of guiding questions. However, the accompanying descriptions varied considerably in content and quality, suggesting that not all AI developers engaged with these concepts in a deep or critical manner. Reasons for this could be that not all the participants were aware that reporting Fig. 2 | Evaluation categories of the AI landslide challenge. npj | naturalhazards Comment npj Natural Hazards | (2025) 2:102 3
would be an important part of the evaluation. Another reason could be a lack of familiarity with these concepts or communication challenges. For example, not all participants in AI challenges have a formal education in AI/ML. This could reflect inequalities in education and training opportunities. As autodidacts, AI developers might not understand the importance of these concepts. For example, in the context of sustainability and efficiency, future challenges could encourage participants to include simple runtime or emissions reporting so efficiency can be compared alongside accuracy. By providing a clearer requirement for documentation, this would strengthen awareness and help balance performance with environmental responsibility. Another issue could be related to reporting itself. Producing written material (in English) for an online challenge is uncommon and might have discouraged some participants, particularly those who are non-native English speakers. An alternative approach to acquire insights into how they observed best practices could be a video recording or an interview format. Based on these findings, Fig. 3proposes actionable recommendations that could be followed by those organizing an online challenge related to natural hazards and intending to bridge innovation with international standardization. To facilitate the implementation of these recommendations, it provides a schematic of a scalable and versatile capacity-building mechanism that can be adapted for different natural hazard use cases. Through following this prototype, online challenges can provide bottomup, hands-on training and guide AI developers to use responsible and internationally recognized methods. It can also plug important gaps in education and training. The insights from this study show that challenges could be a useful mechanism for fostering the use of standards. Through organizing similar studies in the future on other geoscience applications of AI, we can acquire more comprehensive and generalizable insights into how they can be used to foster the adoption of standards. Ultimately, it is our aim to further refine this prototype to shape more specific guidelines. In parallel, we will support complementary efforts to bridge online challenges with standards and education. Recent activities include the publication of a blog on the Zindi website, the hosting of online workshops and webinars (e.g., the hybrid workshop on “Resilience to Natural Hazards through AI Solutions” on 8 May 2025), and the open-access publication of technical reports introducing important concepts and offering best practices1. Lorenzo Nava1,2 , Monique M. Kuglitsch3, Jennifer Selby3, Amy Bray4, Paul Kennedy4, Joel Lontsi4, Megan Yates4, Ahmed Tambal5, Ahmed Ilmane6, Andrea Manara7, Thomas Basikolo7, Katharina Weitz3, Arif Albayrak8,9, Chester Karwatowski10, Rakiya Babamaaji11, Silvia Garcia12, Elena Xoplaki13, Stefan Uhlenbrook14,PierrePhilippe Mathieu15, Raffaele Albano16, Andrea Toreti17 & Filippo Catani18 1Department of Earth Sciences, University of Cambridge, Cambridge, UK. 2Department of Geography, University of Cambridge, Cambridge, UK. 3Fraunhofer Institute for Telecommunications, Heinrich Hertz Institute, Berlin, Germany. 4Zindi, Dover, DE, USA. 5Lal Capital Partners, London, UK; and, Dubai, United Arab Emirates. 6International University of Yamoussoukro, Yamoussoukro, Côte d’Ivoire. 7International Telecommunication Union, Geneva, Switzerland. 8NASA Goddard Space Flight Center, Greenbelt, MD, USA. 9University of Maryland at Baltimore County, Baltimore, MD, USA. 10IBM (Retired), Poughkeepsie, NY, USA. 11Strategic Space Applications Department, National Space Research and Development Agency (NASRDA), Abuja, Nigeria. 12Coordinación de Eléctrica y Computación/Geotecnia, Instituto de Ingeniería Universidad Nacional Autónoma de México, Mexico City, Mexico. 13CMCC Foundation Euro-Mediterranean Center on Climate Change, Bologna, Italy. 14World Meteorological Organization, Geneva, Switzerland. 15ESA/ESRIN, Frascati, Italy. 16Department of Health Sciences, University of Basilicata, Potenza, Italy. 17European Commission Joint Research Centre, Ispra, Italy. 18Machine Intelligence and Slope Stability Laboratory, Department of Geosciences, University of Padua, Padua, Italy. e-mail: [email protected] Fig. 3 | Strategy for enhancing innovation impact and standardization in online challenge frameworks. a Actionable recommendations and bproposed prototype to bridge innovation with international standardization in online challenges. npj | naturalhazards Comment npj Natural Hazards | (2025) 2:102 4
Received: 17 October 2025; Accepted: 7 November 2025; References 1. Oehmen, D. & O’Dwyer Stock, R. Focus Group on Artificial Intelligence for Natural Disaster Management: Standardization Roadmap on Natural Disaster Management—Trends and Gaps in Standardization. ITU-T Focus Group Technical Report (ITU/WMO/UNEP, 2022). 2. Kuglitsch, M. M. et al. AI to the rescue: how to enhance disaster early warning systems with high-tech tools. Nature 634,27–29 (2024). 3. International Telecommunication Union (ITU). Crowdsourcing AI and Machine Learning Solutions for SDGs: ITU AI/ML Challenges 2024 Report (International Telecommunication Union, 2024). 4. Jia, G. et al. A global landslide non-susceptibility map. Geomorphology 389, 107804 (2021). 5. United Nations Office for Disaster Risk Reduction (UNISDR). Landslide hazard and risk assessment. Words into action guidelines: National disaster risk assessment. Hazard specific risk assessment https://www.unisdr.org/files/52828_03landslidehazardandriskassessment. pdf (2017). 6. Kuglitsch, M. et al. Artificial intelligence for disaster risk reduction: opportunities, challenges, and prospects. Bull. World Meteorol. Organ. 71, 1 (2022). 7. Ganesh, B., Vincent, S., Pathan, S. & Garcia Benitez, S. R. Integration of GIS and machine learning techniques for mapping the landslide-prone areas in the state of Goa, India. J. Indian Soc. Remote Sens. 51,https://doi.org/10.1007/s12524-023-01707-y (2023). 8. Nava, L. et al. Brief communication: AI-driven rapid landslide mapping following the 2024 Hualien earthquake in Taiwan. Nat. Hazards Earth Syst. Sci. 25, 2371–2377 (2025). 9. Mondini, A. C., Guzzetti, F. & Melillo, M. Deep learning forecast of rainfall-induced shallow landslides. Nat. Commun. 14, 2466 (2023). 10. Nava, L., Monserrat, O. & Catani, F. Improving landslide detection on SAR data through deep learning. IEEE Geosci. Remote Sens. Lett. 19,1–5 (2022). Acknowledgements The online challenge described in this Comment was coordinated by the Working Group on Educational Materials for Capacity Sharing, within the Global Initiative on Resilience to Natural Hazards through AI Solutions. This work was supported by the European Union’s Horizon Europe research and innovation programme (EU Horizon Europe) as grant MedEWSa (101121192) and ARTEMis (101225852). The dataset used in this challenge was derived from research conducted during Lorenzo Nava’s doctoral studies and the collaborative work of the Machine Intelligence and Slope Stability Laboratory at the University of Padova, whose contributions are gratefully acknowledged. Competing interests The authors declare no competing interests. Additional information Correspondence and requests for materials should be addressed to Lorenzo Nava. Reprints and permissions information is available at http://www.nature.com/reprints Publisher’snoteSpringer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Open Access This article is licensed under a Creative Commons Attribution-NonCommercialNoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creativecommons.org/licenses/by-nc-nd/4.0/. © The Author(s) 2025 npj | naturalhazards Comment npj Natural Hazards | (2025) 2:102 5