This material is based upon work supported by the U.S. Department of Energy, Office of Science, Biological and Environmental Research Program under Award Number DE-SC0023216, the Bridging Barriers Program at The University of Texas at Austin through the Planet Texas 2050 project, and NSF Navigating the New Arctic Award No.1952196 / ©2025 IEEE Connecting Science and Stories: A Human-AI Method for Multi-Modal Knowledge Capture that Integrates Sensors, Community Narratives, and Computational Decision Models Suzanne A. Pierce Decision Support Office, Texas Advanced Computing Ctr The Univ. of Texas Austin Austin, Texas, USA
[email protected] William Mobley Decision Support Office, Texas Advanced Computing Ctr The Univ. of Texas Austin Austin, Texas, USA
[email protected] Vladislav Krendelev Data Management & Collections, Texas Advanced Computing Ctr The Univ. of Texas Austin Austin, Texas, USA
[email protected] Lydia Fletcher Data Management & Collections, Texas Advanced Computing Ctr The Univ. of Texas Austin Austin, Texas, USA
[email protected] Maximiliano Osorio MetaLearn Labs, S.A. Santiago, Chile m[email protected] Abstract— Traditional decision support systems often fail to bridge scientific models with community knowledge, creating an implementation gap. The AI-enabled Modeling project (AIM) presents a novel approach that semantically links computational narratives with geospatial sensor data and physics-based models across domains. By leveraging large language models, semantic model integration, and multi-modal analysis, we establish technical services that enable integration of qualitative and quantitative knowledge systems. Our framework transforms unstructured stakeholder narratives into model configurations, generating culturally relevant representations of past events for complex socio-environmental challenges like hurricane impacts. The AIM project leverages, DataX, a data portal built on reusable science gateway components, incorporating CKAN-based data discovery, semantic variable mapping, interactive geospatially sensed visualization through Potree point cloud rendering, and a model integration platform. We demonstrate how unstructured interview transcripts can be systematically linked to data or physics-based models through a combination of natural language processing, standardized scientific vocabularies, and reusable computational cookbooks. This infrastructure enables transdisciplinary teams to connect lived experiences with observational data and computational simulations for technical integration in complex socio-environmental systems analyses. Keywords— multi-modal information, decision support, semantic integration, AI-enabled modeling I. INTRODUCTION The gap between community knowledge and scientific models represents a fundamental challenge in environmental decision-making. While computational models provide sophisticated predictions, they often fail to incorporate local knowledge and lived experiences that are critical for effective implementation. Conversely, qualitative data from community narratives remains largely disconnected from quantitative modeling frameworks. This paper presents a set of tools that can be accessed via a science gateway infrastructure that address these challenges through three key innovations: (1) creation of a multi-modal data portal supporting narrative collections with associated structured data and models, (2) CKAN-based data discovery services with semantic tagging for scientific variables to make multi-modal data resources and collections findable across services, and (3) AI-enabled linking services that systematically connect problem frames from narratives with data and model resources through computational cookbooks and semantic variable mapping, enhanced by interactive geospatial visualization and containerized model cataloging services. Our approach recognizes that the friction generated at the boundaries between qualitative community knowledge and quantitative data and models can hinder science-based decisionmaking advances. Built on the Planet Texas 2050 DataX portal infrastructure with CKAN-based Data Discovery, the project integrates multiple AI-enabled services including a collection of computational cookbooks called the Sites and Stories App leveraging natural language processing, topic modeling and large language models (LLMs) for narrative analysis, a Remote Sensing App with a Potree viewer for interactive LiDAR point cloud visualization that allows annotation overlays, and a Model INTegration Platform (MINT) for physics-based model integration. This paper presents a description of the prototype architecture and conceptual approaches implemented in the DataX portal through a use case example from the historic Hurricane Beulah event that occurred in the Lower Rio Grande Valley of Texas in 1967 to highlight the mechanics needed to create multi-modal artifacts, compose multi-step workflows, and orchestrate analyses via portal services. This work provides technical implementation of services with capabilities to enable the creation of decision support systems in preparation for future use with real-world communities and decision makers.
II. GATEWAY ARCHITECTURE AND CORE SERVICES A. DataX Foundational Portal and Data Discovery The foundation of our approach is the DataX portal, a reusable deployment of a science gateway architecture originally developed for natural hazards engineering [1]. This infrastructure provides reliable and trustworthy data storage and management, incorporating hierarchical data organization with project-based access controls that enable team collaboration and enable access for data management. TAPIS [2] provides programmatic data access and workflow automation, while direct integration with TACC’s high-performance computing resources allows computationally intensive analyses. This coupling of data storage, authentication, APIs, and compute resources creates a seamless environment for transdisciplinary research teams. The Planet Texas 2050 AIM project extends the base portal with a Comprehensive Knowledge Archive Network (CKAN) [3] system to provide metadata-rich cataloging where each dataset includes standardized descriptors for basic identification, as well as optional metadata entries to make the information interactive with other applications. DataX users are prompted with optional fields to populate, such as spatial and temporal coverage and scientific variable names. This comprehensive metadata framework ensures that heterogeneous data, ranging from observational sensor networks, LiDAR point cloud datasets, and audio recordings of interviews are discoverable through a consistent interface. A key innovation is the semantic tagging framework with variable naming service that creates mappings between community terminology such as “water came up to the house,” geospatial measurements that can be detected with sensors like “flood_depth_meters,” and model parameters such as “max_wse” for maximum water surface elevation. This semantic bridge enables researchers to find relevant data regardless of whether they search using colloquial descriptions or formal scientific terminology and use the data directly in registered models. Moreover, the approach enables the connectivity between people, data, and models. B. Knowledge Capture Services Specifications The infrastructure employs reusable computational cookbook applications [4] that enable multi-modal bundling into coherent collections of analytical tooling, this paper describes workflows using the Sites & Stories App and the Remote Sensing App bundles (Fig.1). Each suite of application services are designed to handle fundamentally different types of information while maintaining interoperability through the DataX portal. Users navigate data ingest, registry, and search via the Data Discovery service based on CKAN and ingest for larger files via DataX. Using cookbooks, users can follow instructions for uploading data while also connecting to the Data Discovery catalog (e.g. CKAN service) to ensure discoverability. In the Hurricane Beulah use case analyses were conducted using Sites & Stories applications that leverage the Natural Language Analysis and Remote Sensing Analysis via the Applications tab in the portal shown in Fig 1. Within the Remote Sensing App users access the Potree geospatial toolkit for handling LiDAR and photogrammetric point cloud data to provide interactive 3D visualization of terrain information. This service enables users to upload point cloud datasets, register them with Data Discovery services, and create interactive visualizations where narrative annotations can be overlaid on terrain data. The Sites & Stories App space includes tools to streamline narrative analysis, such as transcription for unstructured texts, videos, and audio using Whisper [5], custom multi-attribute narrative analysis cookbooks with natural language processing workflows (NLP), BERTopic analyses [6], and foundational LLMs. The approach is reusable across problem settings and based on scientific variable naming techniques that provide semantic bridges between community terminology and scientific vocabularies, the application tab within DataX presents users with a set of AI-enabled services for capturing and analyzing unstructured narratives and placebased knowledge. Users can collect audio interviews then use the DataX applications capability to load the files, register them with Data Discovery, then open tools to transcribe files, use an authenticated space with foundational large language models (LLMs) to explore various prompt strategies, employ a BERTopic workflow to understand key themes in the data, and ultimately apply narrative text analyses to identify, timeline, plotline, character roles, problem frames, and story lines within the reusable tools. C. AI-Enabled Integration Layers The gateway employs three complementary AI technologies that work in concert to bridge qualitative and quantitative data sources with interactive capabilities due to the use of semantic descriptions and naming tools. At the core is the Model INTegration (MINT) platform, a semantic model composition platform that offers model execution and parameterization, standardized model metadata, and catalog services to streamline the loose coupling of physics-based model. Currently, executable models range from hydrology, hydraulics, groundwater, crop, fire, and economic models. MINT uses W3C semantic web standards (RDF, OWL, SWRL, SPARQL, and PROV) to represent model metadata and automate scenario configuration, creating a flexible framework for model composition [7]. Computational cookbooks provide accessible workflows to register Jupyter-based environments and provide headless containers that are registered within MINT and run through TAPIS calls. The Sites & Stories App links narrative data sources via CKAN to cookbooks and MINT where the users complete analyses such as timeline extraction, plot analysis, and problem frame identification [5]. The integration of preconfigured analysis libraries, like scikit-learn, NLTK, and spaCY, within a secure HPC execution environment via Science Gateway and TAPIS services ensures both computational power and data security while maintaining shareable and reproducible analysis workflows.
III. USE CASE: HURRICANE BEULAH The gateway's capabilities are demonstrated through the Hurricane Beulah case study in the Lower Rio Grande Valley, which showcases how historical environmental events can be reconstructed through the integration of diverse data sources and community knowledge. This implementation brings together multiple forms of evidence to create a rich, multi-dimensional understanding of the 1967 hurricane that devastated the region. The case study’s multi-modal data collection approach combines quantitative reconstructions with cultural memory. Historical hurricane data has been reconstructed through ADCIRC coastal surge modeling, providing scientifically rigorous estimates of storm surge heights and flooding extent. These computational results are enriched by oral histories from community members who lived through the hurricane, offering irreplaceable first-person accounts of how the storm unfolded and its impacts on daily life and the shared perspective of a subject matter expert modeler who developed an ADCIRC representation of the Beulah event. The cultural dimension is captured through oral histories that document collective memory of significant events, preserving details about the hurricane’s progression and community response from on-the-ground perspectives. Museum archival photographs and documents provide additional visual evidence and official records that help validate both the model outputs and personal recollections. LiDAR point cloud data of the Rio Grande Valley terrain provides the geospatial foundation for positioning both narrative elements and model outputs within accurate topographic context. All components are stored together as a story-informed collection in the Data Discovery service that is built using DataX, CKAN, TAPIS, and Corral storage at TACC as the underpinning infrastructure. The gateway orchestrates a sophisticated workflow that transforms these diverse inputs into outputs capable of technical integration demonstrations. The process begins with community narratives being captured through the Sites and Stories application during events at the Museum of South Texas, where residents share their memories in comfortable, culturally appropriate settings. Sites & Stories cookbooks are used to processes these narratives through a systematic pipeline: interview audio is transcribed, transcriptions are analyzed using narratology techniques to extract problem frames, which are then positioned semantically through embedding similarity scores (Fig. 1). The narrative processing pipeline identifies temporal patterns, spatial references, and key themes that can be systematically linked to scientific variable names using semantic similarity scores to suggest possible options and positioned on terrain data through the Potree visualization system. The MINT platform integrates physics-based models and accesses ADCIRC storm surge model output data via the Data Discovery service, creating threads of coupled analyses with simulations that capture key elements of interest (see Fig. 1). Semantic services perform the crucial task of linking narrative elements to model parameters, translating community observations such as "water came up to the house" into scientific variables like “flood_depth_meters” that can be used with computational models and positioned at specific coordinates on terrain visualizations. IV. LESSONS LEARNED AND CONCLUSIONS The AIM project demonstrates that meaningful integration of qualitative narratives with quantitative models and geospatial data is achievable through semantic relationships and reusable workflows via cookbook capabilities. The combination of CKAN-based discovery, scientific variable mapping, interactive geospatial visualization, and a model integration platform creates new possibilities for transdisciplinary research, fundamentally changing how communities and scientists can work together to understand complex environmental challenges. Key contributions include developing a reusable architecture for multi-modal data integration that treats diverse data types as first-class citizens within the same infrastructure. The semantic services we've implemented successfully link community knowledge to scientific models and analytical approaches through sophisticated reusable and shareable services. The Potree geospatial visualization capabilities brings observational, or sensed data, into spatially explicit representations of community knowledge. The research team has demonstrated workflows that enable narrative-driven model configuration, showing that stories can indeed inform computational simulations in rigorous, reproducible ways. The Hurricane Beulah case study demonstrates how integrating community narratives with computational models and geospatial terrain data to create a more complete understanding of environmental hazards that honors both scientific analysis and lived experience. The success of this approach lies not just in the technical infrastructure, but in its recognition that effective environmental decision-making requires bridging multiple ways of knowing—from computational simulations to cultural memory, from observational terrain models to personal stories. From a technical perspective, several insights have emerged that can guide future implementations. The scalability achieved by leveraging an existing science gateway infrastructure significantly accelerated our deployment timeline, demonstrating the value of building upon established cyberinfrastructure rather than creating systems from scratch. Interoperability proved crucial to success, with semantic web standards and standardized semantic variable [7] naming enabling seamless integration across disparate models that were never designed to work together. Security concerns, particularly around processing sensitive community narratives, were effectively addressed by running NLP and large language models on TACC infrastructure, ensuring that data never leaves secure institutional boundaries while still benefiting from advanced AI capabilities. Future work will expand the implementation based on variable mapping ontologies, enhance uncertainty quantification methods, improve annotation tools for community engagement, and stakeholder validation of these technical capabilities. The infrastructure provides a foundation for researchers to better integrate local knowledge with scientific modeling to improve access and interoperability with the goal of supporting decision support system design and instantiation. Gateway infrastructure enabled the integration of community knowledge with technical models through semantic linking, AI analysis, and geospatial visualization tools that translate stakeholder perspectives into
model configurations and interactive visualizations. These endto-end capabilities are essential for participatory decision support systems that meaningfully engage communities in climate adaptation planning. ACKNOWLEDGMENT The authors acknowledge funding from the Bridging Barriers Program at The University of Texas at Austin through the Planet Texas 2050 initiative. We thank our collaborators at the Museum of South Texas History, Dr. Francisco Guajardo and Melissa Peña, for invaluable assistance in documenting the Hurricane Beulah event. We also thank Dr. Clint Dawson for the coastal surge modeling of the Beulah case study. REFERENCES [1] Rathje, E. M., Dawson, C., Padgett, J. E., Pinelli, J. P., Stanzione, D., Adair, A., ... & Mosqueda, G. (2017). DesignSafe: New cyberinfrastructure for natural hazards engineering. Natural hazards review, 18(3), 06017001. [2] Stubbs, J., Cardone, R., Packard, M., Jamthe, A., Padhy, S., Terry, S., ... & Jacobs, G. (2021). Tapis: An API platform for reproducible, distributed computational research. In Advances in Information and Communication: Proceedings of the 2021 Future of Information and Communication Conference (FICC), Volume 1 (pp. 878-900). Springer International Publishing. [3] Winn, J. (2013). Open data and the academy: An evaluation of CKAN for research data management. [4] Mobley, Pearson, Osorio, Tijerina, Trueheart, Faust, & Pierce. (2024, October 8). Implementing Reproducible Cookbook Environments for Advanced Analyses on Science Gateways. Science Gateways 2024 (SG24), Bozeman, MT. https://doi.org/10.5281/zenodo.13869305 [5] Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. International Conference on Machine Learning (ICML). [6] Grootendorst, M., (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure, arXiv:2203.05794. [7] Gil, Y., Garijo, D., Khider, D., Knoblock, C. A., Ratnakar, V., Osorio, M., ... & Shu, L. (2021). Artificial intelligence for modeling complex systems: taming the complexity of expert models to improve decision making. ACM Transactions on Interactive Intelligent Systems, 11(2), 1-49. Figure 1. Workflow integrating Hurricane Beulah community narratives with computational models. Interview transcripts (left) undergo narrative analysis to extract problem frames (center), which are mapped to scientific variables and model parameters (right) through semantic services. Semantic elements are then fused with physics-based models and geospatial data to create enriched reports and interactive visualizations, enabling systematic linking of qualitative community knowledge with quantitative computational models.