scieee AI-readable full text Open interactive document viewer

Resources and Tools to Support Integration of Geospatial and Environmental Exposure Data into Epidemiological and Clinical Health Research

Miller, Aubrey; Lane, Kevin; Conway, Mike; Messier, Kyle; Shatz, Maria

Abstract

Slides from the Joint Annual Meeting of the International Society of Exposure Science and the International Society for Environmental Epidemiology 2025 in Atlanta, Georgia. Presentations on: CHORDS Data Ecosystem, Aubrey Miller, NIEHS CAFE Data Repository, Kevin Lane, Boston University CHORDS Data Catalog, Mike Conway, NIEHS Tutorial: Amadeus Software, Kyle Messier, NIEHS Data Standards, Maria Shatz, NIEHS

Full text

National Institutes of Health • U.S. Department of Health and Human Services Resources and Tools to Support Integration of Geospatial and Environmental Exposure Data into Epidemiological and Clinical Health Research Joint Annual Meeting of the International Society of Exposure Science and the International Society for Environmental Epidemiology Atlanta, GA August 17, 2025 National Institutes of Health • U.S. Department of Health and Human Services National Institutes of Health U.S. Department of Health and Human Services Learning Objectives •Increase awareness of existing geospatial environmental exposure data, tools, and resources •Improve understanding of evolving approaches for linking and harmonizing exposure and health data •Considerations for data standards in the context of epidemiologic and clinical research National Institutes of Health U.S. Department of Health and Human Services Workshop Goals •Build a network of researchers working in this space •Solicit feedback on tools and resources •Disseminate resources and opportunities related to the CHORDS and CAFE programs to a wider audience •Join our stakeholder group/community of practice National Institutes of Health U.S. Department of Health and Human Services CHORDS Overview Goal: Strengthen data infrastructure to facilitate research connections between environmental exposures and health outcomes so researchers can 1) identify, analyze, and reduce the health effects associated with disaster-related events (e.g., wildfires) and 2) improve patient and population health outcomes National Institutes of Health U.S. Department of Health and Human Services CHORDS Objectives and Deliverables Objective 1: Web-based data resources catalog Objective 2: Standardized, linked datasets Objective 5: Addressing end user needs Objective 3: Toolkit of software and tutorials Objective 4: Evaluation use case Integrated data platform for health outcomes research National Institutes of Health U.S. Department of Health and Human Services CHORDS Data Platform National Institutes of Health U.S. Department of Health and Human Services Data Catalog: Manually Curated Searchable Geospatial and Population Resources Metadata provided: •Project information •Measures •Spatial characteristics •Temporal characteristics •Data access information •Other: Keywords Citations Publications National Institutes of Health U.S. Department of Health and Human Services Open-Source Software and Tools for Simplifying Environmental Health Data Analysis CHORDS-specific tools •amadeus: Downloads large-scale environmental and weather data and gets it ready for analysis •beethoven: A reproducible and extensible pipeline for air pollution exposure designed for updated and timely releases •chopin: Simplifies parallelization, running many calculations or processes simultaneously, of big environmental exposure data National Institutes of Health U.S. Department of Health and Human Services CHORDS in the Context of the Exposome NIH NIEHS Earth/Environmental Science Related Domains While focused on specific CHORDS use cases, this is an opportunity to think about the broader ecosystem studying the Exposome. • Look ‘inward’ at other parts of NIH •Ease federation with other platforms common in other domains •Be flexible in terms of source and target schema while maintaining as much structure and validation as possible •Think about the provenance and verification of data over time •Prepare for increased use of AI for both metadata extraction and catalog automation as well as the use of AI in data discovery National Institutes of Health U.S. Department of Health and Human Services Designing the Catalog •This is really about the metadata model, first and foremost •With such a broad set of domains, we would have to build out the model as we iterated •Tooling to allow curators to design and communicate design as well as curate the data •Workflows that allowed constant migration •Community standards, ontologies for crossing environment and health not readily available •ECTO - https://pmc.ncbi.nlm.nih.gov/articles/PMC9951428/ •EXO - https://pmc.ncbi.nlm.nih.gov/articles/PMC3314380/ National Institutes of Health U.S. Department of Health and Human Services Designing the System •The care and feeding of the metadata model development and curation was central •The system needed to enable the entire workflow, from design to curation to validation to presentation •The data itself is the catalog, and the system needed flexibility to slice, dice, repurpose and federate •The catalog needed to fit into an expanding, federated and heterogeneous cyberinfrastructure picture National Institutes of Health U.S. Department of Health and Human Services Architectural Design Principles •Avoid building bespoke systems, use existing platforms. •Focus on trans-NIH standards and platforms. •Join communities (standards and open-source). •API first (offer and utilize open API at the foundation). National Institutes of Health U.S. Department of Health and Human Services The CHORDS Portal •https://chordshealth.org/discovery •Gen3 based •Open Source - https://github.com/uc-cdis •GA4GH Compliant •FedRAMP certified •Operated by Gen3 on AWS National Institutes of Health U.S. Department of Health and Human Services Data Model Centric •Above all things, the data model was seen as the key –Structured –Flexible –Validated –Includes controlled vocabularies –Serializable and API Accessible National Institutes of Health U.S. Department of Health and Human Services Data Curation with CEDAR •CEDAR (https://https://cedar.metadatacenter.org/) is being used for human curation of metadata. •CEDAR templates developed by curators, also serving as the Gen3 model specification. National Institutes of Health U.S. Department of Health and Human Services Example CEDAR Template To JSON-LD National Institutes of Health U.S. Department of Health and Human Services CHORDS Phase 2 CEDAR Crosswalk Migrated JSON-LD Data Models Evolve •Migration data flows have been created and must be managed •Note the closed system, without human intervention it cannot be established which data might have disappeared or been updated •Adding automated input from other sources and pushing data to new endpoints creates all sorts of disconnects, stale data issues National Institutes of Health U.S. Department of Health and Human Services CHORDS Phase 2 CEDAR CHORDS Accelerator Core Model Data.gov CDC EPA Other Other Other Cafe Navigator •Add automated cataloging •Create a connector architecture for adding new sources, new sinks, and migration workflows •Automate re-validation and updates •Project the CHORDS catalog into multiple endpoints. •Migration workflows •Revalidation and discovery workflows National Institutes of Health U.S. Department of Health and Human Services Summary •CHORDS is about the data model •CHORDS is about federation •CHORDS is about human and automated curation •Focus on open API and frameworks so that catalog data can be shared and re-purposed •New domains and data types are welcome, and there is a sustainable method for expansion •Core API: https://github.com/NIEHS/accelerator-core •Helm Charts: https://github.com/NIEHS/accelerator-helm Mike Conway NIEHS/NIH [email protected] Deep Patel NIEHS/NIH [email protected] Thanks! Get in touch with any questions! ISEE-ISES 2025 1 BUSPH-HSPH-CAFÉ Research Coordinating Center (U24) 1 Goals of CAFÉ 2 Greg Wellenius (BU) Amruta Nori-Sarma (Harvard) Francesca Dominici (Harvard) Pat Kinney (BU) Danielle Braun (Harvard) Kevin Lane (BU) Benjamin Sovacool (BU) Pam Templer (BU) Jon Levy (BU) Rachel Nethery (Harvard) Rebecca Pearl-Martinez (BU) Gaurab Basu (Harvard) BUSPH-HSPH Climate Change and Health Research Coordinating Center What does CAFÉ provide? Various resources, including: ●CAFÉ University webinars ●Free annual virtual meetings ●Funding for pilot studies ●Research translation lab ●Research Mentore-Mentee Connections ●Youtube tutorials ●Slack channels ●Monthly newsletter ●Data and data management resources https://climatehealthcafe.org/ Find us on Linkedin at Climate & Health CAFÉ Join our Community of Practice to receive our monthly newsletter and Slack channels! CAFÉ Data Management 1. Facilitate sharing and reuse of Climate Health (CCH) data (Dataverse) 2. Develop and promote data science and software tools for processing, linking, and analyzing common CCH data (Github) 3. Provide data management and dissemination guidance 4. Identify CCH research data needs and define common data elements across the community of practice (COP) Note: CAFE is not a data processing/analysis service center! Our focus is on curating commonly used datasets and coding pipelines. CAFÉ: Data Repositories (Dataverse) What are data repositories NIH Data Management Sharing Plans Dataverse Extracting Data Why Create/Use a Dataverse Repository ●Can share data with collaborators or the public. Open content can be accessed directly via the UI or API, and restricted content can be requested using a "request access" feature if enabled by the data depositor. ●Assigns a DOI to every published dataset. The repository records file downloads and views and makes this information available to depositors. Depositors who create collections can also ask or require downloaders to provide information about their data re-use. ●Uses standard-compliant metadata to ensure that dataset metadata can be mapped easily to common metadata schemas to make data more preservable and interoperable. ●Provides a mechanism by which a journal's editors and reviewers can have anonymous access to a dataset or dataverse before it is made public. See Private URLs CAFÉ Dataverse Collection: Contributing Data https://climatecafe.github.io/tutorial.html NIEHS Glossary Keyword Term Developing your CAFÉ Dataverse Collection: Custom Metadata https://www.niehs.nih.gov/health/topics/glossary Findable •Each dataset receives a persistent DOI, making it uniquely identifiable and citable. •Rich metadata standards enhance searchability and indexing in global registries (e.g., DataCite, Google Dataset Search, National Library of Medicine). Accessible •Supports open access to data (with options for restricted access when needed), satisfying NIH’s emphasis on making data broadly available. •Provides a stable, institutionally supported platform with long-term data preservation. Interoperable •Encourages use of standard data formats (e.g., CSV, JSON, Stata, SPSS) that can be reused across systems and disciplines. •Metadata and APIs are compatible with other repositories and tools, supporting machine-readability and integration. Reusable •Users can attach detailed documentation, codebooks, and README files to promote data reuse and understanding. •Enables application of clear data licenses (e.g., CC0, CC BY), fulfilling legal and ethical reuse criteria Benefits of Dataverse to Help You Meet FAIR Data Principles CAFÉ Toolkit Examples: GitHub https://climate-cafe.github.io//intro.html What are data repositories Benefits of using GitHub for code sharing to meet FAIR Data Principles Findable •GitHub repositories are indexed by search engines and can be assigned persistent identifiers (e.g., DOIs), making code discoverable. •Clear project structure and metadata (e.g., README, tags, releases) improve searchability and navigation. Accessible •Code is openly accessible by default (when using public repos), satisfying NIH expectations for open data/code sharing. •GitHub supports easy download, cloning, and API access to repositories. Interoperable •GitHub supports standard code formats and integrates with popular tools and languages (e.g., Python, R, Jupyter, Git). •Enables collaboration and integration with reproducibility platforms (e.g., Binder, CodeOcean). Reusable •Clear version control and documentation support reproducibility and reuse of code by other researchers. •Supports open-source licenses (e.g., MIT, Apache), making reuse legally and ethically straightforward. •Promotes open science through forks, pull requests, and issues, supporting the collaborative spirit of NIH’s open data initiatives. Computational Requirements and Workarounds ●WBGT and UTCI processing involves hourly data for 8+ meteorological variables and is time and computeintensive ●The GitHub repository includes batch scripts to divide the processing and run different regions and time frames in parallel Unfamiliar with Batch Scripting? See the CAFÉ microtutorial! Climate-CAFE/hpc_batch_jobs_micro_tutorial Example Script for Batch Submission to Subset Across Time and Space Computational Requirements and Workarounds ●Batch scripting is an efficient way to separate large spatiotemporal processing but may not be available in all settings ●If you are processing a single county and/or a short time series (less than a year), an example processing pipeline can be found under code_atl ●Loops are used to process across months Scripts 1+2: Downloading ERA5 Data ●Smaller requests move more quickly through Copernicus download queue ●Using batch submission to submit requests will expedite process ○To further speed up this process, the API is directed to save files to a directory that has not been created on your device ○After all requests have been processed through the queue, the directory for saving files is created. ○Log language from the API will be saved to a file on your device so the files can be programmatically downloaded to the newly created directory. Example Log Text which can be used to download files from queue Scripts 1+2: Downloading ERA5 Data ●View the status of your requests by visiting https://cds.climate.copernicus.eu/request s ●Time from queue to completion will vary based on how many users are accessing data ●Once all requests are Complete, run script 1C to download to your device ●Downloading all variables typically takes several hours per year Scripts 5+6: Temporal and Spatial Aggregation ●Script 5 combines the ERA5 raster grid with the input polygon. This can then be used to extract the hourly raster values to the county ●Script 6 includes the daily and county-level aggregation, to create the final outputs for use WBGT and UTCI Next Steps ●Nationwide county products to be posted to CAFÉ Dataverse ●Update package with population weighting to county aggregation ●Enhance quality control procedures to flag values that may be more prone to error due to presence of water bodies, unexpected input values, etc. in line with Spangler et al. ●Use existing R package to demonstrate process for query of ERA5 data with variable, time, and geographic extent inputs ●Across Kenya administrative boundaries, compute daily summary heat measures, including daily max temperature, heat index, and more ERA5 Daily Heat Aggregation OpenStreetMap Road Metric Calculations ❖Objective: Calculate total road lengths by type (e.g., highways, local streets) for municipalities in Mexico, adaptable to other countries ❖Data: Uses OpenStreetMap (OSM) roads from Geofabrik, GADM shapefiles for boundaries, and R scripts to calculate and aggregate road sums. ➢Bash scripts included for fast processing. National Institutes of Health • U.S. Department of Health and Human Services Integration of Geospatial Data into Epidemiological and Clinical Health Research Kyle P Messier ISES/ISEE 2025 Pre-Meeting Workshop 06 Aug 17, 2025 National Institutes of Health U.S. Department of Health and Human Services Outline Geospatial Exposure Data Health Data Integration Challenges Tools to Support Geospatial Calculations and Integrations Amadeus Examples Open-Source Community Questions National Institutes of Health • U.S. Department of Health and Human Services Geospatial Data Sources Environmental and Exposure Variables National Institutes of Health U.S. Department of Health and Human Services •NASA EARTHDATA •Multi-Resolution Land Characteristics Consortium –NLCD •Zenodo –EU Open Research Repository •CAFÉ Dataverse (https://dataverse.harvard.edu/dataverse.xhtml?alias=CAFE) •CHORDS Catalog (https://chordshealth.org/discovery) National Institutes of Health • U.S. Department of Health and Human Services Health Data Population and Person Level Data National Institutes of Health U.S. Department of Health and Human Services •US Census •CDC Places •CDC Wonder National Institutes of Health U.S. Department of Health and Human Services Address or Individual Level Data •Point location of individual with health outcomes •Epidemiological cohorts •Electronic Health Records •Minimum need is 1 spatial location •Better is addresses over life course •Best is detailed tracking •Privacy concerns with address or time-course data National Institutes of Health • U.S. Department of Health and Human Services Data Integration Spatial and Temporal Challenges National Institutes of Health U.S. Department of Health and Human Services National Institutes of Health U.S. Department of Health and Human Services Spatial Geometry Relationships •Source data →Geospatial and exposure data sources •Receptor data →residential addresses, census tracts, etc. •Integrating geospatial data to residential address →Row 1 •Integrating geospatial data to census tracts →Row 3 National Institutes of Health U.S. Department of Health and Human Services National Institutes of Health U.S. Department of Health and Human Services NASA Moderate Resolution Imaging Spectroradiometer (MODIS) •Normalized Difference Vegetation Index (NDVI) •Mean zonal statistics computed for U.S. states and territories •24 lines Maccherone, B., and Frazier, S. (n.d.). Moderate Resolution Imaging Spectroradiometer – Data. https://modis.gsfc.nasa.gov/data/ NASA Earth Observing System Data and Information System. (n.d.). MODIS Terra Vegetation Indices (MOD13A1) [Data set]. NASA Goddard Space Flight Center. https://modis.gsfc.nasa.gov/data/ National Institutes of Health U.S. Department of Health and Human Services Auxiliary functions •Internal functions manage complexities •Various file formats •Native coordinate reference systems •Extraction location geometries –Point or polygon National Institutes of Health U.S. Department of Health and Human Services Interoperability •Output classes are from popular spatial packages •Linking to U.S. Census data •Data analysis pipelines •Beyond R Dowle, M., & Srinivasan, A. (2024). data.table: Extension of ‘data.frame’. R package version 1.14.2. https://cran.r-project.org/package=data.table Hijmans, R. J. (2024). terra: Spatial data analysis package for R. R package version 1.7-19. https://cran.r-project.org/package=terra Miquel, M., & Dorman, M. (2024). targets: Make dynamic and reproducible workflows. R package version 1.0.0. https://cran.r-project.org/package=targets Pebesma, E. J. (2024). sf: Simpl e features for R. R package version 1.0-13. https://cran.r-project.org/package=sf Python Software Foundation. (2024). Python programming language. Version 3.11.5. https://www.python.org/ Wal ker, K. (2024) . tidycensus: Load US census data from the US Census Bureau’s API. R package version 1.0.0. https://cran.r-project.org/package=tidycensus Wal ker, K . (2024) . tigris: R package for loading TIGER/Line shapefiles. R package version 1.0.0. https://cran.r-project.org/package=tigris Wickham, H., Fr ançois, R., Henry, L., & Müll er , K . (2024). dplyr: A grammar of data manipulation. R package version 1.1.2. https://cran.r-project.org/package=dplyr National Institutes of Health U.S. Department of Health and Human Services Test-driven development •Reliability and stability •Error catching during development •Interpretable error messages •Unit tests –URLs return successful response status (download_*) –Correct object classes (process_*; calculate_*) –Identify missing or improper parameters (all functions) •Integration tests –No errors when process_* run with live download_* files + GitHub, Inc. (n.d.). GitHub logo. https://github.com/logos GitHub, Inc. (n.d.). GitHub Invertocat logo. https://github.com/logos Marsner. (n.d.) Why Test-Driven Development (TDD). https://marsner.com/blog/why-test-driven-development-tdd/ Wick ham, H. (2024) . testthat: Get started with testing. R package version 3.1.0. https://cran.r-project.org/package=testthat National Institutes of Health U.S. Department of Health and Human Services National Institutes of Health U.S. Department of Health and Human Services beethoven Building an Extensible, Reproducible, Test-driven, Harmonized, Open-Source, Versioned, Ensemble Model for Air Quality P42 08/20/25 09:00-09:45 Connecting Health Outcomes Research Data Systems (CHORDS) P17 08/18/2025 15:30 – 16:15 chopin Computation of Spatial Data by Hierarchical and Objective Partitioning Inputs for Parallel Processing National Institute of Environmental Health Sciences. (2024, August 14). Climate and Health Outcomes Research Data Systems (CHORDS). https://www.niehs.nih.gov/research/programs/chords Song, I. & Messier, K. (2024). chopin: Computation of Spatial Data by Hierarchical and Objective Partitioning of Inputs for Parallel Processing. R package version 0.9.0. https://docs.ropensci.org/chopin Song, I., Marques, E., Manware, M., et al. (2024). beethoven: Building an Extensible, rEproducible, Test-driven, Harmonized, Open-source, Versioned, ENsemble model for air quality. R package ver sion 0.4.3. https://github.com/NIEHS/beethoven National Institutes of Health U.S. Department of Health and Human Services chopin Convenient parallel geospatial operations via data partitioning and automatic deployment - Regular grid, hierarchical, balanced Custom and common gridding (Uber H3, DGGRID) options Available in CRAN Published in SoftwareX National Institutes of Health U.S. Department of Health and Human Services National Institutes of Health U.S. Department of Health and Human Services https://github.com/NIEHS/urbanheat_pipeline National Institutes of Health U.S. Department of Health and Human Services Dowle, M., & Srinivasan, A. (2024). data.table: Extension of ‘data.frame’. R package version 1.14.2. https://cran.r-project.org/package=data.table GitHub, Inc. (n.d.). GitHub logo. https://github.com/logos GitHub, Inc. (n.d.). GitHub Invertocat logo. https://github.com/logos GNU Project. (2024). Bash (Bourne Again Shell) [Computer software]. Free Software Foundation. https://www.gnu.org/software/bash/ Hijmans, R. J. (2024). terra: Spatial data analysis package for R. R package version 1.7-19. https://cran.r-project.org/package=terra Maccherone, B., and Frazier, S. (n.d.). Moderate Resolution Imaging Spectroradiometer – Data. https://modis.gsfc.nasa.gov/data/ Marsner. (n.d.) Why Test-Driven Development (TDD). https://marsner.com/blog/why-test-driven-development-tdd/ Miquel, M., & Dorman, M. (2024). targets: Make dynamic and reproducible workflows. R package version 1.0.0. https://cran.r-project.org/package=targets NASA Earth Observing System Data and Information System. (n.d.). MODIS Terra Vegetation Indices (MOD13A1) [Data set]. NASA Goddard Space Flight Center. https://modis.gsfc.nasa.gov/data/ National Aeronautics and Space Administration. (n.d.). NASA logo. https://www.nasa.gov/ National Institute of Environmental Health Sciences. (2024, August 14). Climate and Health Outcomes Research Data Systems (CHORDS). https://www.niehs.nih.gov/research/programs/chords National Institute of Environmental Health Sciences. (2024, September 23). Area 4: Data Science and Computational Biology. https://www.niehs.nih.gov/about/strategicplan/research/data-science National Library of Medicine. (n.d.). FAIR Data. https://www.nlm.nih.gov/oet/ed/cde/tutorial/02-200.html National Oceanic and Atmospheric Administration. (n.d.). NOAA logo. https://www.noaa.gov Pebesma, E. J. (2024). sf: Simple features for R. R package version 1.0-13. https://cran.r-project.org/package=sf Python Software Foundation. (2024). Python programming language. Version 3.11.5. https://www.python.org/ R Core Team. (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.r-project.org/ Song, I. & Messier, K. (2024). chopin: Computation of Spatial Data by Hierarchical and Objective Partitioning of Inputs for Parallel Processing. R package version 0.9.0. https://docs.ropensci.org/chopin Song, I., Marques, E., Manware, M., et al. (2024). beethoven: Building an Extensible, rEproducible, Test-driven, Harmonized, Open-source, Versioned, ENsemble model for air quality. R package version 0.4.3. https://github.com/NIEHS/beethoven U.S. Department of the Interior. (n.d.). DOI logo. https://www.doi.gov/ U.S. Environmental Protection Agency. (n.d.). EPA logo. https://www.epa.gov/ U.S. Geological Survey. (n.d.). USGS logo. https://www.usgs.gov/ University of California, Merced. (n.d.). SNRI logo. https://snri.ucmerced.edu/ Walker, K. (2024). tidycensus: Load US census data from the US Census Bureau’s API. R package version 1.0.0. https://cran.r-project.org/package=tidycensus Walker, K. (2024). tigris: R package for loading TIGER/Line shapefiles. R package version 1.0.0. https://cran.r-project.org/package=tigris Wickham, H. (2024). testthat: Get started with testing. R package version 3.1.0. https://cran.r-project.org/package=testthat Wickham, H., François, R., Henry, L., & Müller, K. (2024). dplyr: A grammar of data manipulation. R package version 1.1.2. https://cran.r-project.org/package=dplyr Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18 National Institutes of Health • U.S. Department of Health and Human Services Efforts to standardize geospatial data elements and to create geospatial CDEs Maria Shatz, PhD Office of Data Science National Institute of Environmental Health Sciences (NIEHS) Research Triangle Park, NC, USA e-mail: [email protected] National Institutes of Health U.S. Department of Health and Human Services Conflict of Interest Statement The author declares no conflict of interest. National Institutes of Health U.S. Department of Health and Human Services Examples of geospatial measures for health Geospatial Measure Description Example Data Sources Public Health Use Air Quality Index (AQI), PM2.5, NO ₂, O₃ Measures air pollution levels by location EPA AirNow, NASA satellite remote sensing Identify high - risk areas for respiratory/cardiovascular illness; issue health alerts Proximity to Pollution Sources Distance to highways, animal feed facilities, oil wells, waste sites National Land Cover Database, EPA Toxic Release Inventory Assess exposure risk and plan zoning/regulatory actions Noise Pollution Maps Sound intensity across neighborhoods Environmental monitoring stations noise sensors Study effects on sleep, cardiovascular health, stress Green Space Availability Access to parks, tree canopy coverage Satellite imagery, OpenStreetMap Promote physical activity, mental health benefits Wildfire and smoke exposure Location based exposure to smoke and particle pollution AirNow Fire and Smoke Map, US EPA and Forest Service Enables the public to take protective actions against wildfire smoke. Population Density & Demographics Distribution of population groups (age, income, etc.) Census data Plan resource allocation and emergency response Disease Incidence Heat Maps Spatial distribution of reported cases Public health registries, GIS geocoding Detect clusters/outbreaks for rapid response Flood/Heatwave/Wildfir e Risk Maps Areas prone to climate hazards FEMA flood maps, NOAA climate data, satellite fire tracking Prepare evacuation plans, target public health messaging Water Quality Mapping Location of contamination in rivers, wells, and systems EPA Safe Drinking Water Information System Prevent and mitigate waterborne disease outbreaks National Institutes of Health U.S. Department of Health and Human Services Value of standardized data elements Reason Impact if standardized Impact if Not standardized Consistency PM2.5 in one state means the same as PM2.5 in another Apples -to-oranges comparisons between regions and cohorts Data Integration Easily merge environmental, health, and demographic datasets Incompatible formats, loss of detail Reproducibility Other researchers can verify and replicate findings Conflicting results due to method differences Scalability GIS tools and dashboards work in multiple jurisdictions Costly customizations for each location Public Trust Clear, consistent public health communication Confusion, misinformation, reduced engagement National Institutes of Health U.S. Department of Health and Human Services Core Principles for Standardization •Clear Definitions – Agree on exactly what is measured (e.g., “walkability” defined by X metrics). •Common Units – Use the same measurement units (e.g., micrograms/m³ for PM2.5). •Temporal Alignment – Match timeframes (e.g., daily averages vs. annual means). •Spatial Resolution – Define the geographic scale (e.g., census tract, zip code, 1 km² grid). •Metadata Documentation – Record data sources, collection methods, and processing steps. •Interoperable Formats – Use open, standardized GIS data formats (GeoJSON, shapefiles). National Institutes of Health U.S. Department of Health and Human Services Implementation Path •Adopt Existing Standards – Use established frameworks (CDC’s SVI, WHO air quality guidelines). •Collaborative Governance – Engage diverse stakeholders in setting norms. •Capacity Building – Train staff in GIS, data cleaning, and metadata creation. •Continuous Review – Update standards as technology and health priorities evolve. National Institutes of Health U.S. Department of Health and Human Services Challenges •While there are standards for health data (e.g. ICD-10, LOINC, SNOMED) or for geospatial data (Open Geospatial Consortium, OGD) there are no standard on intersection of the two fields •Multiple existing measures with no community consensus •Researchers are not aware of the existing standards •Variety of data sources and computational methods •Due to the temporal element the values are dynamic; •The presentation should be not only machine-readable but easily reused for computation with different parameters National Institutes of Health U.S. Department of Health and Human Services Approaches to standardization of data elements How do we move forward as a community to improve this situation? NIH proposed creating and endorsing CDEs and publishing them in the CDE repository moving forward National Institutes of Health U.S. Department of Health and Human Services What is a Common Data Element (CDE)? •A CDE is a fixed representation of a variable comprising one defined question •paired with a specified set of allowable responses including value range and units if applicable, •mapped to semantic concepts and codes like UMLS, SNOMED AND •that are used in common across multiple: studies, research sites, initiatives, datasets, etc. Question (variable): What is the participant’s height? Allowable answers: Feet and inches in whole numbers - For example, if all clinical studies supported by an IC were required to collect height data with this question and answer pair, - this data element would be common to all of those studies - it would be a Common Data Element Adopted from CDE WG presentation materials