WorldCereal harmonized reference datasets - extended and updated
Boogaard, Hendrik; Pratihast, Arun; Laso Bayasv, Juan Carlos; Karanam, Santosh; Fritz, Steffen; Van Tricht, Kristof; Degerickx, Jeroen; Van Haren, Charlotte
- Publisher
- Zenodo
- Language
- en
Abstract
Within the ESA WorldCereal project we have built a global, community-based and open repository of harmonized reference data on land cover, crop type and irrigation information. These datasets are typically used to construct and validate global cropland and crop type maps. We define reference data as all data which can either be used for calibrating classification algorithms or validating the resulting products. As such, reference data should contain location- and time-specific information about land cover and/or crop type. Reference data can include: In-situ field data gathered through dedicated field surveys Farmers declarations through parcel registration systems Data derived from visual or automated interpretation of very high-resolution satellite imagery or in-situ photographs (e.g. streetview or mapillary) Existing high-quality classified maps based on analysis of satellite imagery Reference data is typically available in different formats depending on the source of the dataset. To ensure multiple reference datasets can be readily combined and can serve as input for a dedicated calibration/validation task, all datasets have been structured, harmonized, annotated and evaluated. The datasets adhere to the same (meta)data standards and formats. See for more information the PDF document 'WorldCereal_Reference_dataset_naming_convention_v2'. Each individual dataset has its own data license defined by the original data holder, which can be retrieved from the accompanying Excel metadata file. The license will be one of the different licenses listed in this publication. The user must check the data license of each of these datasets before starting to use or redistribute the data. Individual datasets have been grouped per year and continent into a separate .zip archive. The title of individual datasets matches the Collection ID as specified in the WorldCereal Reference Data Module, our dedicated user interface for serving reference data to the community. A complete overview on WorldCereal reference data documentation can be found here. Information on WorldCereal reference data harmonization procedures, data standards and legend can be found here: and in the accompanying PDF document 'WorldCereal_Reference_dataset_naming_convention_v2' hosted in this repository.
Full text
Reference dataset naming convention (WorldCereal Reference Data) Document Ref: WorldCereal_Reference_dataset_naming_convention_v2 Version: 2 Creation Date: 2024-05-17 Last Modified: 2025-10-09
Reference dataset naming convention Reference data is typically available in different formats depending on the source of the dataset. To ensure multiple reference datasets can be readily combined and can serve as input for a dedicated calibration/validation task, all datasets have been structured, harmonized, annotated and evaluated. The datasets adhere to the same (meta)data standards and formats. Each dataset is named according to a standardized naming convention, so the origin and contents of the dataset are immediately clear to the user. Elements are: • Year: primary year of the dataset • Region: country code or continent (GO (Globe), AF (Africa), NA (North America), OC (Oceania), AN (Antarctica), AS (Asia), EU (Europe), SA (South America)) • Identifier: unique identifier for the dataset • Data type: point or polygon • Information content: 3-digit code indicating which type of information is included. First digit represents land cover, second represents crop type and third represents irrigation. 0 means absent, 1 means present So, the name of an individual dataset would be as follow: • Vector file containing points: o <year>_<region>_<identifier>_POINT_<information content>.<extension> • Vector file containing polygons: o <year>_<region>_<identifier>_POLY_<information content>.<extension> Each dataset is documented according to a standardized metadata structure, again to enable easy retrieval of information related to origin, history, content and access right of the data. The land cover/crop type labels in each dataset have been defined according to the same generic and hierarchical land cover / crop type legend (see this link). A detailed explanation of the WorldCereal legend can be found in the on-line documentation via this link. Next, each dataset contains the same list of standardized data attributes, and each attribute adhere to predefined formatting standards. Attributes are: • sample_id: unique ID for each individual sample • ewoc_code: land cover and crop type label according to the hierarchical WorldCereal legend • valid_time: a specific date (yyyy-mm-dd), for which the observation is valid. See this link for more information. Note that this PDF document ‘WorldCereal_DerivingValidityTime_v1_1.pdf’ is hosted in this repository. • irrigation_status: presence and type of irrigation according to this legend. • quality_score_lc: a quality score ranging from 0 to 100 indicating the inherent quality of the sample with respect to its land cover label. See this link for more information. Note that this PDF document ‘WorldCereal_ConfidenceScoreCalculations_v1_1.pdf’ is hosted in this repository. • quality_score_ct: a quality score ranging from 0 to 100 indicating the inherent quality of the sample with respect to its crop type label. See this link for more information. Note that this PDF document ‘WorldCereal_ConfidenceScoreCalculations_v1_1.pdf’ is hosted in this repository. • distance_to_road: distance (m) to nearest road infrastructure based on OpenStreetMap. Note this is an attribute added recently so only available for recent point data sets.
The data sets also include specific attributes that are used by the WorldCereal processing line for instance to support sampling of labels or additional information in case data is obtained through interpretation of imagery (latter 5 attributes): • extract: indicates whether for this sample an extraction of model inputs should be done • h3_l3_cell: identifier of the h3 cell (used for sub-sampling) at level 3 resolution • sampling_ewoc_code: crop category label assigned to the sample during sub-sampling of the dataset. • image_time: the date (yyyy-mm-dd) of the imagery used to identify land cover/crop type label. • number_validations: number of people having reviewed this observation. • type_validation,: type of validation used to determine land cover/crop type label (either Expert, NonExpert or Both) • agreement: number of people agreeing • disagreement: number of people disagreeing More information and all details on reference data and the WorldCereal Reference Data Module can be found via this link.