scieee AI-readable full text Open interactive document viewer

Benchmaking data access and processing performance of OpenEO backends: a reproducible approach

Gohil, Jaykumar Harishbhai; Bhawiyuga, Adhitya; Girgin, Serkan

Abstract

openEO is a community-driven initiative supported by prominent research institutions, as well as cloud service providers and international agencies, that is designed to simplify the access and processing of Earth Observation (EO) data. It offers a standardized interface that allows users to interact with multiple EO data providers, cloud services, and processing engines, facilitating the efficient retrieval, analysis, and visualization of EO data. This facilitates scalable and flexible processing workflows for a broad spectrum of applications. To achieve this, openEO defines a HTTP-based API that enables communication between cloud backends with large EO datasets and frontend applications in an interoperable manner by specifying how to manage users, discover available EO data and processes on backends, execute processes and user-defined functions on backends, and download results. Currently there are different backends available that are based on various big EO data processing frameworks, such as Geotrellis, GRASS GIS/actinia, Google Earth Engine, OpenDataCube, Xarray, and Dask. There are also operational platforms that provide openEO services (i.e. openEO providers) such as Copernicus Data Space Ecosystem, EODC, and VITO. Recently, the openEO Platform also has been launched, providing federated access to various openEO services. Although openEO provides a standard mechanism to access the provided EO data and processing services, available data and processing capabilities differ between the service providers. Data availability is determined by the size of the data archives of the providers, whereas processing capabilities are largely influenced by the specific backend implementations deployed by the operators and their IT infrastructure. Although openEO provides a web portal (openEO Hub), which lists the reported capabilities of the openEO providers, limited information is available on the actual data availability and data access and processing performance of the providers. Benchmarking is crucial for assessing the performance of various openEO backends, because it provides objective and standardized measurements of efficiency, performance, and reliability. By comparing different backends under controlled conditions, it provides valuable information that enables users to select the most suitable backend for their specific needs, based on factors such as processing time and scalability. It can also help to identify strengths and weaknesses, driving improvements and optimizations by backend developers that benefit the entire user community. Eventually, it can support the development of more effective systems, and helps maintain high standards in EO data access and processing workflows. To benchmark the performance of different openEO backends and service providers, we have created an openEO benchmarking framework. Built using a unit testing approach, the framework allows for replicable testing of data availability, data access, and data processing capabilities across various backends. The results are presented in standard formats, providing both detailed reports and summaries. This poster presentation provides a comprehensive overview of the design principles and operational structure of the developed openEO benchmarking framework. It covers the framework's core features, functional capabilities, along with a comprehensive overview of the tests that are designed to assess data availability on different backends and their data access and processing performance. The results of the tests performed are summarized to provide an insight into the current state of the operational openEO providers and their services.

Full text

The openEO Hub highlights the capabilities of different openEO providers, but it offers limited details on actual data availability, data access, and processing performance. To support reliable comparisons between services, we have developed a benchmarking toolkit that enables reproducible testing of data availability, and access and processing performance across multiple backends. The toolkit generates standardized reports to present the results. openEO offers a standardized interface that allows users to interact with various EO data providers, cloud services, and processing engines. A range of backends is available, each built on different data processing frameworks such as GeoTrellis, GRASS GIS/actinia, OpenDataCube, and Google Earth Engine. Operational platforms offering openEO services include the Copernicus Data Space Ecosystem, EODC, and VITO, along with the openEO Platform with federated access to services. Jay Gohil, Adhitya Bhawiyuga, Dr. Serkan Girgin <[email protected]> Center of Expertise in Big Geodata Science (CRIB), Faculty of Geo-information Science and Observation (ITC), University of Twente, The Netherlands Benchmaking data access and processing performance of openEO backends: a reproducible approach SERVICE AVAILABILITY AND LATENCY PROCESSING PERFORMANCE PROCESS AVAILABILITY The toolkit facilitates easy benchmarking of various openEO services, allowing end users to compare analysis results across different backends and platforms. This comparison provides valuable insights into the degree of interoperability between platforms and the potential for achieving consistent, reproducible analysis outcomes. It also aids in detecting issues related to service and data availability, as well as uncovering bugs or implementation inconsistencies that may hinder a federated processing approach. In addition to supporting individual benchmarking studies, the toolkit can be used to regularly compute and report operational performance metrics of openEO platforms. This performance data could be integrated into the openEO Hub, enhancing the current information on service capabilities real-world performance insights. To support this effort, a standardized benchmarking suite, featuring both synthetic and real-world process graphs, along with a reference dataset covering various spatial and temporal scales, is planned as a follow-up activity. openeobench service --url <api_url> --output <out> url,timestamp,response_time,status_code,error_msg,content_size Output: A CSV file (in append mode by default). openeobench process --url <api_url> --output <out> process,level,status,compatibility,reason Output: A CSV file containing summary information and an openEO API compliant JSON file with details. openeobench run --url <api_url> --input <graph.json> --output <out> Output: A folder containing the process graph (JSON), job metadata (JSON), output files (GeoTIFF, etc.). openeobench service-summary <input> --output <out> --format <format> url,availability,avg_response_time,stddev_response_time,latency,latency_stddev Output: A CSV file or Markdown document with summary statistics. openeobench process-summary <input,...> --output <out> --format <format> backend,l1_available,l1_missing,l1_mismatch,l2_available,..., total_mismatch Output: A CSV file or Markdown document with summary information on compliance of backend(s). openeobench run-summary <input,...> --output <out> --format <format> run,time_download,time_download_stddev,processing_time,...,total_time_stddev Output: A CSV file or Markdown document with summary statistics about the runs. openeobench result-summary <input,...> --output <out> --format <format> run,num_files,file_1_min,file_1_max,file_1_avg,...,file_n_stddev Output: A CSV file or Markdown document with summary statistics about the resulting outputs. https://github.com/ITC-CRIB/openeobench Platform Availability Latency (ms /KB) CDSE 100% 0.64 ± 0.12 EO4EU 100% 3.93 ± 1.34 EODC 97% 1.16 ± 0.27 EURAC 100% 3.89 ± 0.52 GEE 100% 1.41 ± 0.17 mundialis 100% 1.76 ± 0.33 openEO Platform 100% 0.74 ± 0.46 Sentinel Hub 100% 1.04 ± 0.06 VITO 100% 0.24 ± 0.06 Example Results The results are derived from data gathered from the /processes endpoint over a five-day period (21–25 June 2025), with a sample size of 1,384 requests per platform. Platform Ver. L1 – Minimal L2 – Advanced L3 – Advanced L4 – A&B Custom CDSE 1.2 55 (0, 0) 44 (4, 4) 18 (5, 26) 4 (1, 4)17 EO4EU 1.2 53 (0, 2) 37 (5, 11) 6 (1, 38) 0 (0, 5) 1 EODC 1.2 55 (0, 0) 47 (7, 1) 33 (17, 11) 4 (4, 1)13 EURAC 1.0 41 (0, 14) 19 (0, 29) 13 (6, 31) 0 (0, 5) 5 GEE 1.0 53 (0, 2) 37 (5, 11) 6 (1, 38) 0 (0, 5) 1 mundialis 1.0 47 (0, 8) 23 (0, 25) 6 (0, 38) 0 (0, 5) 1 openEO Platform 1.2 55 (0, 0) 47 (6, 1) 33 (16, 11) 5 (5, 0)27 Sentinel Hub 1.0 55 (0, 0) 32 (0, 16) 11 (0, 33) 0 (0, 5) 0 VITO 1.2 55 (0, 0) 44 (4, 4) 19 (6, 25) 5 (5, 0)18 Example Results Black: Total number of processes, Blue: Experimental processes, Red: Missing processes Platform Vienna, 2018 Vienna, 2020 Vienna, 2024 Bratislava, 2018 Bratislava, 2020 Bratislava, 2024 CDSE 24 24 24 24 24 22 EO4EU No OIDC authentication EODC 0 0 0 0 0 0 EURAC No suitable data collection available GEE 1 1 1 1 1 1 mundialis Authentication error openEO Platform 24 24 24 24 24 22 Sentinel Hub Authentication error VITO 24 24 24 24 24 22 CDSE STAC 96 96 92 48 48 46 GEE CDSE VITO openEO Platform The values represent the number of GeoTIFF files downloaded for each run, where the scenario is to simply download a datacube for a specific spatiotemporal extent. For comparison, the final row shows the number of STAC items available on the CDSE. The spatial extent for each scenario covers a 10 km radius around Vienna and Bratislava, with Vienna situated at the intersection of Sentinel-2 scenes. Scenarios correspond to two-month periods within the specified years as 01/01 – 28/02 2018, 01/05 – 30/06 2020, and 01/09 – 31/10 2024. Platform Submission (s) Queue (s) Processing (s) Download (s) Total (s) CDSE 0.67 ± 0.60 127.11 ± 65.92 204.24 ± 72.57 9.94 ± 1.22 326.47 ± 103.75 GEE 0.16 ± 0.00 10.69 ± 5.10 0.34 ± 0.08 11.56 ± 5.07 openEO Platform 3.39 ± 6.40 51.63 ± 11.78 210.40 ± 21.60 6.60 ± 8.89 308.18 ± 52.48 VITO 0.65 ± 0.15 58.63 ± 18.72 194.72 ± 26.78 3.52 ± 0.71 281.08 ± 28.78 The values represent the duration of different analysis steps, where the scenario is to reduce a spatiotemporal datacube by a temporal reducer (mean). The spatial extent is 10 km radius around Vienna, whereas the temporal range is a two-month period between 01/04 – 30/06 2024. Example Results