scieee AI-readable full text Open interactive document viewer

1-km Global Arable Land Fraction (GALF) in 2000, 2010, and 2020

Gao, Lun; Lu, Xiaoman; Anderson, Weston; Lesk, Corey; Mazur, Elise; Stolle, Fred; Potapov, Peter; Tubiello, Francesco; Fritz, Steffen; Mao, Jiafu; You, Liangzhi; Ray, Deepak

Abstract

Global Arable Land Fraction (GALF) was produced at a spatial resolution of 1 km by merging five land cover products: MODIS, ESA CCI, GlobeLand30, GLAD, and GLC_FCS30D. The dataset was generated for the years 2000, 2010, and 2020, representing the world’s first global maps specifically dedicated to arable land—defined as areas under temporary agricultural crops (with multiple cropping counted once), temporary meadows for mowing or pasture, and land temporarily fallow. Detailed information on data processing methods and validation results is provided in the accompanying document GALF_readme.pdf.

Full text

1-km Global Arable Land Fraction (GALF) in 2000, 2010, and 2020 Lun Gaoa, b, *, Xiaoman Lub, Weston Andersonc, d, e, Corey Leskf, g, h, Elise Mazuri, Fred Stollei, Peter Potapovi, Francesco N. Tubielloj, Steffen Fritzk, Jiafu Maob, Liangzhi Youl, m, Deepak Raya a Institute on the Environment, University of Minnesota, St. Paul, MN, USA b Environmental Sciences Division, Oak Ridge National Laboratory, Oak Ridge, TN, USA c Earth System Science Interdisciplinary Center, University of Maryland, College Park, MD, USA d NASA Goddard Space Flight Center, Earth Science Division, Greenbelt, MD, USA e Geographical Sciences, University of Maryland, College Park, MD, USA f Department of Earth and Atmospheric Science, University of Quebec in Montreal, Montreal, QC, Canada g Department of Geography, Dartmouth College, Hanover, NH, USA h Neukom Institute for Computational Science, Dartmouth College, Hanover, NH, USA i World Resources Institute, Washington, DC, USA j Statistics Division, Food and Agriculture Organization of the United Nations, Rome 00153, Italy k International Institute for Applied Systems Analysis (IIASA), Schlossplatz 1, A-2361 Laxenburg, Austria l International Food Policy Research Institute, Washington, DC, USA m Digital Agricultural Research Institute, College of Economics and Management, Huazhong Agricultural University, Wuhan, Hubei, China *Corresponding author: [email protected] Data and Background We constructed 1-km Global Arable Land Fraction (GALF) by integrating multiple global satellite-derived land cover products with long-term temporal coverage (≥20 years) and high spatial resolution (≤500 m), selecting only those with consistent temporal characteristics and spatial continuity. The included global products are the International Geosphere-Biosphere Programme (IGBP) from the Moderate Resolution Imaging Spectroradiometer (MODIS, MCD12Q1 v061)1, European Space Agency’s Climate Change Initiative (CCI)2, GlobeLand303, Global Land Analysis and Discovery (GLAD) dataset4, and GLC_FCS30D5. Note that for datasets that do not explicitly distinguish arable land from woody croplands, we hereinafter refer to all such areas collectively as cropland. Table 1. Specifications of five land cover products (MODIS, CCI, GlobeLand30, GLAD, and GLC_FCS30D) and FAO dataset. The definition of arable land used in this study follows the FAO’s definition of arable land, which is highlighted in bold. Dataset Cropland Definition Spatial Res. Temporal Coverage MODIS Cropland (lands covered with temporary crops with harvest cycle less than one year. Includes areas of land temporarily fallow or left idle) and mosaic classes that mix cropland with other land cover types. 500 m 2001-2023 CCI Rainfed cropland, herbaceous cover cropland, tree or shrub cover cropland, irrigated cropland, and mosaic classes that mix cropland with other land cover types. 300 m 1992-2020 GlobeLand30 Category includes paddy fields, irrigated dry land, rain-fed dry land, vegetable land, pasture planting land, greenhouse land, land and mainly for planting crops with fruit trees and other economic trees, as well as tea gardens, coffee gardens and other shrubs. 30 m 2000, 2010, 2020 GLAD Land used for annual and perennial herbaceous crops for human consumption, forage (including hay) and biofuel. Perennial woody crops, permanent pastures and shifting cultivation are excluded from the definition. The fallow length is limited to 4 years for the cropland class. 30 m 2000, 2005, 2010, 2015, 2020 GLC_FCS30D Rainfed cropland, irrigated cropland, herbaceous cover, and tree or shrub cover (orchard). 30 m 1985-2022 FAO Arable land includes areas under temporary agricultural crops (with multiple cropping counted once), temporary meadows for mowing or pasture, and land temporarily fallow. Permanent crops refer to land cultivated with long-term crops that do not require annual replanting (e.g., cocoa, coffee, oil palm, and rubber), as well as land with flowering trees and shrubs (e.g., roses and jasmine) and nurseries, excluding those for forest trees that are classified as forest. Country Level 1961-2023 As shown in Table 1 and discussed previously6, cropland definitions vary across products and deviate from the FAO cropland definition, which includes both arable land and permanent woody crops. This introduces possible bias in our and similar analyses. For instance, three products (CCI, GlobeLand30, and GLC_FCS30D) include permanent crops, mostly woody tree crops (e.g., oil palm, cocoa, coffee, rubber, fruit orchards). Conversely, MODIS and GLAD cropland maps exclude them. Furthermore, it is likely that the annual crops component of all products is in fact aligned more with the FAO definition of temporary crops rather than arable land6. This is because satellite-derived products, by their nature, primarily reflect land cover, whereas cropland represents a combination of land cover and land use shaped by human activities for food production7, which are more difficult to distinguish from space. Relying solely on individual land cover datasets could therefore lead to systematic overor underestimation of cropland extent, particularly in regions dominated by woody crops, such as Southeast Asia. Method To harmonize different products and enhance consistency with the FAO definition of arable land, we developed a synergy approach8–13 for generating GALF. Unlike previous studies8–13, we focused exclusively on arable land rather than including permanent crops for the following reasons: (1) arable land is consistently represented across land cover products and FAO statistics (Table 1), and (2) the harmonization and validation of GALF rely on ground reference samples14–18 that are typically interpreted from satellite or aerial imagery, which could mix woody crops with forests due to their tree-like canopy structure and phenological characteristics. In other words, even if synergized cropland maps were generated following the FAO cropland definition (i.e., including both arable land and permanent crops), it would be difficult to reliably validate the permanent crop component. In constructing GALF, the synergy approach integrates five land cover products with FAO country-level arable land statistics (FAOSTAT)19, considering only the arable land components for products that explicitly distinguish them from woody crops. Woody vegetation was further excluded using the GLAD tree cover mask4. The approach assumes that strong agreement among input layers increases confidence in arable land presence, which generally leads to improved mapping accuracy9–13. Given that most agreement scores can result from multiple combinations of land cover products (Fig. 1), this method first requires ranking the performance of each individual product. Unlike previous studies that rely solely on expert judgment, ground reference samples, or FAOSTAT for ranking9–13, our approach ranks the products based on both their accuracy against ground reference samples and their consistency with FAOSTAT. This is based on the assumption that the product with the highest accuracy against ground reference samples should also show the closest agreement with FAOSTAT country-level arable land area, provided that sufficient validation samples are available and the FAO statistics are reliable. Specifically, for each country, land cover products were first evaluated and ranked by their overall accuracy when at least 20 validation samples were available, considering some countries with small land areas. To avoid overinterpreting marginal differences, only two decimal places of overall accuracy were considered11. In cases where products had identical accuracy or fewer than 20 validation samples were available, the product whose arable land area was closest to the FAOSTAT was assigned a higher rank. In some regions lacking FAOSTAT, rankings were based on continental-scale validation results across nine continental regions (Fig. 2), which were delineated according to administrative boundaries20 and the spatial clustering patterns of global arable land. More information on ground samples and the performance evaluation of land cover products are provided in the section below. Fig. 1. Schematic diagram illustrating possible permutations among three land cover products, along with the ranking sequences derived from the agreement-score and binary-score methods. Product A represents the highest accuracy product, followed by B and C. In this example, scores “2” and “1” result from three different combinations. The primary difference between the two methods is that the agreement-score method prioritizes region 2-3, whereas the binary-score method prioritizes region 1-1 after accounting for regions 3, 2-1, and 2-2. Notably, the divergence between the two methods increases with the number of input land cover products. Combining both methods expands the set of possible permutations, thereby improving the likelihood of identifying the optimal combination with the highest accuracy against ground validation samples and strongest alignment with FAO arable land statistics. Fig. 2. Classification of nine continental regions based on the Large Scale International Boundary (LSIB)20. To reflect the spatial clustering of global arable land, Asia was subdivided into four regions: East Asia, South Asia, Southeast Asia, and Southwest Asia. Meanwhile, Europe, North Asia, and Central Asia were grouped into a single region. Once country-level performance rankings of input land cover products were available, an agreement-ranking score lookup table was established based on binary permutations to prioritize high-ranked regions that best matched reference data (Table 2). Different from previous research that converted binary permutations to either agreement scores8–12 or binary scores13, this study combined both scoring approaches to enhance the likelihood of identifying the optimal configuration that best aligns with ground validation samples and FAO statistics. This is because we found that each scoring method has inherent limitations, but they offer complementary strengths when used together. The agreement-score approach8–12 prioritizes regions where multiple products consistently indicate the presence of arable land (Fig. 1). This method is particularly effective in regions where the input land cover products have comparable accuracy; however, it may underperform in cases where one product significantly outperforms the others. For example, if a single product accurately captures the full extent of arable land, it may mistakenly prioritize regions where multiple lower-performing products indicate arable land over areas correctly identified by the high-performing product alone. Conversely, the binary-score approach13 gives priority to regions identified by the best-performing product but may be less effective when multiple products exhibit similar accuracy (Fig. 1). Thus, combining both approaches increases the diversity of scoring combinations and enhances the ability to identify the configuration that best matches ground validation samples and FAO statistics. In this study, applying the combined method to five input products resulted in a total of 58 unique scoring combinations, whereas using either approach alone yielded only 32 combinations (i.e., 25; see Table 2). Table 2. Scoring scheme for the five input land cover datasets. The agreement-score approach gives precedence to regions where multiple products consistently indicate arable land presence, whereas the binary-score approach prioritizes regions identified by the best products. The agreement and binary scores are identical for values of 0, 1, 2, 29, 30, and 31. Products are ordered by descending accuracy: A, B, C, D, and E. A value of 1 indicates arable land presence and 0 indicates absence. Agreement Score Binary Score A B C D E 31 31 1 1 1 1 1 30 30 1 1 1 1 0 29 29 1 1 1 0 1 28 27 1 1 0 1 1 27 23 1 0 1 1 1 26 15 0 1 1 1 1 25 28 1 1 1 0 0 24 26 1 1 0 1 0 23 22 1 0 1 1 0 22 14 0 1 1 1 0 21 25 1 1 0 0 1 20 21 1 0 1 0 1 19 13 0 1 1 0 1 18 19 1 0 0 1 1 17 11 0 1 0 1 1 16 7 0 0 1 1 1 15 24 1 1 0 0 0 14 20 1 0 1 0 0 13 18 1 0 0 1 0 12 17 1 0 0 0 1 11 12 0 1 1 0 0 10 10 0 1 0 1 0 9 9 0 1 0 0 1 8 5 0 0 1 1 0 7 6 0 0 1 0 1 6 3 0 0 0 1 1 5 16 1 0 0 0 0 4 18 0 1 0 0 0 3 4 0 0 1 0 0 2 2 0 0 0 1 0 1 1 0 0 0 0 1 0 0 0 0 0 0 0 Then, the ground samples and FAOSTAT were used again as benchmarks to identify the best-performing scoring combinations. We found that the combination yielding the highest overall accuracy did not always align with the one most closely matching the FAOSTAT-reported arable land area, particularly in African countries. This discrepancy is likely attributable to relatively low agreement among five land cover products over Africa (Fig. 3) and the inherent uncertainties in both the ground reference and FAOSTAT data21. To balance both criteria, for countries with more than 20 validation samples, we retained only those combinations that either exceeded the overall accuracy of the best individual product or ranked within the top 10% of all combinations based on overall accuracy. Among these, the combination with the closest arable land area to the FAOSTAT was selected. For countries lacking sufficient validation samples, the FAOSTAT alone was used to select the best combination. In cases where FAOSTAT data were unavailable, but more than 20 validation samples were present, ground samples were directly used to identify the optimal combination. For regions lacking both adequate validation samples and FAOSTAT data, the best combination was selected based on the corresponding continental validation results. Fig. 3. Agreement scores among five used cropland products (i.e., MODIS, CCI, GlobeLand30, GLAD, and GLC_FCS30D) at 1-km resolution in 2020. A score of 5 indicates that all products consistently identify the presence of cropland. Finally, the 1-km GALF was computed by averaging arable land weights at 100-m subpixels within areas defined by the optimal scoring combination. Considering that the spatial resolution varies across input land cover datasets, all products were first reprojected into a common coordinate system (WGS84; EPSG:4326) and resampled to a 100-m resolution to facilitate the calculation of arable land fractions at the 1-km scale. In this calculation, 100-m pixels classified as pure arable land by all overlapping products were assigned a weight of 1, while pixels identified as mosaics (i.e., a mix of cropland and other land cover types) by any product were assigned lower weights to reflect their partial arable land composition. Specifically, for CCI, mosaic cropland classes were weighed as 0.75 when cropland fraction exceeded 50%, and 0.25 when it was below 50%. For MODIS, mosaic cropland comprising 40–60% cultivated land was assigned a value of 0.5. In cases where a 100-m pixel was simultaneously labeled as pure arable land and mosaic arable land by different products, the assigned weight was averaged accordingly. Notably, a few small islands were not represented in the continental mask20, for which the arable land fraction was directly calculated based on the GLAD dataset4 due to its similarity to the FAO arable land definition, high spatial resolution, and reliable accuracy, as discussed below. We implemented the procedure using data from 2010 to identify the optimal scoring combination, which was then applied to generate GALF for the years 2000, 2010, and 2020. These three years represent the intersection of temporal coverage across all products, with MODIS land cover from 2001 used as a proxy for 2000. It should be noted that, although the proposed method is expected to produce high-quality arable land maps beyond the capabilities of individual land cover products, its performance strongly depends on the quality of the input datasets and may fail to capture arable land in regions where most products perform poorly. One notable example is greenhouse agriculture, which should ideally be detected by all land cover products but exhibits distinct spectral reflectance compared with other arable land22. We found that the ability of the products to detect greenhouse infrastructures varies considerably. As examined, while smallto medium-sized greenhouse facilities, such as those in Michoacán, Mexico (19.9°N, 102.2°W), are generally captured by all products, only GlobeLand30 and CCI successfully detect the large-scale greenhouse infrastructures (> 500 km2) in Almería, Spain (36.7°N, 2.7°W)23,24. In contrast, other products misclassify these areas as impervious surfaces (GLC_FCS30D), wetlands (GLAD), or non-vegetated land (MODIS). Consequently, the proposed method cannot reliably capture these large-scale greenhouse areas in Spain. To address this issue, we explicitly identified these greenhouse areas across Spain by combining GlobeLand30 and CCI arable land layers with GLC_FCS30D impervious surfaces and GLAD wetlands across Spain, before calculating the 1km arable land fraction. Fig. 4 demonstrates that this combined approach effectively captures the greenhouse areas in southern Spain. On the other hand, in Africa and the Arabian Peninsula, the quality and agreement among the five input cropland products were sometimes very low (Fig. 3). For example, GLC_FCS30D missed an entire scene in Sudan, while CCI misclassified several roads within the Congolian rainforests as cropland. Therefore, we carefully compared GALF with GLAD by evaluating their overall accuracy, area consistency with FAOSTAT statistics, and spatial patterns of arable land distribution. Based on the assessment, we replaced GALF with GLAD in the Republic of the Congo, Democratic Republic of the Congo, South Sudan, Sudan, Somalia, and Saudi Arabia, where GLAD demonstrated comparable accuracy to GALF but provided more realistic spatial patterns. Fig. 4. Detected greenhouses at 1 km in southern Spain in 2020. The mapped greenhouse areas are shown overlaid on Sentinel-2 imagery, where the white regions in the imagery correspond to greenhouse structures. Performance evaluation of input land cover products To support the performance ranking of land cover products in 2010 and identify the best scoring combination, we compiled multiple sources of reference samples (Fig. 5). The year 2010 was selected because it represents the midpoint of the study period and offers the most consistent availability of ground validation data. The reference samples included two datasets from Tsinghua University14,15 and additional samples from the Geo-wiki crowdsourcing platform16 and the U.S. Geological Survey’s Land Change Monitoring, Assessment and Projection (LCMAP)17. The Tsinghua University datasets14,15 were originally developed to support the creation and validation of their 30-m global land cover product (i.e., FROM-GLC), where cropland is defined as arable and tillage land with herbaceous and/or shrub crops—thus more than arable land. The first dataset14 involves 36,352 globally distributed random samples, manually labeled by hundreds of students, researchers, and experts using Google Earth imagery in or around 2010. The second dataset15 contains 38,664 global random sample units spanning the years 1986 to 2010, derived through interpretation of Landsat imagery25, MODIS enhanced vegetation index (EVI)26, and other high-resolution images via Google Earth. For this study, a total of 11,672 samples in 2010 were considered. The Geo-wiki dataset16 records the dominant, secondary, and tertiary land cover types, with cropland defined as “cultivated and managed” or “mosaic of cultivated and managed/natural vegetation”. In this study, locations falling into any of the three cropland-related categories were considered as reference. It comprises 151,942 globally distributed random samples derived from the visual interpretation of Google Earth imagery before 2012. A subset of 12,555 samples from 2010 was selected for analysis. The LCMAP Fig. 7. Comparison of country-level arable land area (km2) estimates from GALF and other products with FAO-reported arable land statistics in 2020, with countries lacking FAO reports excluded from the analysis. The total global arable land area reported by FAO is 1.38 × 107 km2 in 2020. References 1. Friedl, M. A. et al. Global land cover mapping from MODIS: algorithms and early results. Remote Sens. Environ. 83, 287–302 (2002). 2. European Space Agency. Climate Change Initiative Land Cover. (2015). 3. Chen, J. et al. Global land cover mapping at 30m resolution: A POK-based operational approach. ISPRS J. Photogramm. Remote Sens. 103, 7–27 (2015). 4. Potapov, P. et al. The Global 2000-2020 Land Cover and Land Use Change Dataset Derived From the Landsat Archive: First Results. Front. Remote Sens. 3, 856903 (2022). 5. Zhang, X. et al. GLC_FCS30D: the first global 30 m land-cover dynamics monitoring product with a fine classification system for the period from 1985 to 2022 generated using dense-timeseries Landsat imagery and the continuous change-detection method. Earth Syst. Sci. Data 16, 1353–1381 (2024). 6. Tubiello, F. N. et al. Measuring the world’s cropland area. Nat. Food 4, 30–32 (2023). 7. Kerr, J. T. & Cihlar, J. Land use and cover with intensity of agriculture for Canada from satellite and census data. Glob. Ecol. Biogeogr. 12, 161–172 (2003). 8. Jung, M., Henkel, K., Herold, M. & Churkina, G. Exploiting synergies of global land cover products for carbon cycle modeling. Remote Sens. Environ. 101, 534–553 (2006). 9. Ramankutty, N., Evan, A. T., Monfreda, C. & Foley, J. A. Farming the planet: 1. Geographic distribution of global agricultural lands in the year 2000. Glob. Biogeochem. Cycles 22, (2008). 10. Fritz, S. et al. Cropland for sub‐Saharan Africa: A synergistic approach using five land cover data sets. Geophys. Res. Lett. 38, (2011). 11. Fritz, S. et al. Mapping global cropland and field size. Glob. Change Biol. 21, 1980–1992 (2015). 12. Lu, M. et al. A cultivated planet in 2010: 1. the global synergy cropland map. Earth Syst. Sci. Data Discuss. 2020, 1–33 (2020). 13. Tubiello, F. N. et al. A new cropland area database by country circa 2020. Earth Syst. Sci. Data 15, 4997–5015 (2023). 14. Gong, P. et al. Finer resolution observation and monitoring of global land cover: First mapping results with Landsat TM and ETM+ data. Int. J. Remote Sens. 34, 2607–2654 (2013). 15. Zhao, Y. et al. Towards a common validation sample set for global land-cover mapping. Int. J. Remote Sens. 35, 4795–4814 (2014). 16. Fritz, S. et al. A global dataset of crowdsourced land cover and land use reference data. Sci. Data 4, 1–8 (2017). 17. Stehman, S. V., Pengra, B. W., Horton, J. A. & Wellington, D. F. Validation of the US geological survey’s land change monitoring, assessment and projection (LCMAP) collection 1.0 annual land cover products 1985–2017. Remote Sens. Environ. 265, 112646 (2021). 18. Zhao, T. et al. Assessing the accuracy and consistency of six fine-resolution global land cover products using a novel stratified random sampling validation dataset. Remote Sens. 15, 2285 (2023). 19. FAO. FAOSTAT land use [Dataset]. http://www.fao.org/faostat/en/#data/RL (2019). 20. U.S. Department of State, Office of the Geographer. Large Scale International Boundaries (LSIB) – Simplified. (2017). Available at: https://developers.google.com/earthengine/datasets/catalog/USDOS_LSIB_SIMPLE_2017 21. The World Bank. Global Strategy to Improve Agricultural and Rural Statistics. https://www.fao.org/fileadmin/templates/ess/documents/meetings_and_workshops/ICAS5/Ag _Statistics_Strategy_Final.pdf (2011). 22. Tong, X. et al. Global area boom for greenhouse cultivation revealed by satellite mapping. Nat. Food 5, 513–523 (2024). 23. Gómez-Galán, M., Pérez-Alonso, J., Callejón-Ferre, Á.-J. & Sánchez-Hermosilla-López, J. Assessment of postural load during melon cultivation in Mediterranean greenhouses. Sustainability 10, 2729 (2018). 24. Mendoza-Fernández, A. J., Peña-Fernández, A., Molina, L. & Aguilera, P. A. The role of technology in greenhouse agriculture: Towards a sustainable intensification in Campo de Dalías (Almería, Spain). Agronomy 11, 101 (2021). 25. Woodcock, C. E. et al. Free access to Landsat imagery. Sci. VOL 320 1011 (2008). 26. Huete, A. et al. Overview of the radiometric and biophysical performance of the MODIS vegetation indices. Remote Sens. Environ. 83, 195–213 (2002). 27. Di Gregorio, A. Land Cover Classification System: Classification Concepts and User Manual: LCCS. vol. 8 (Food and Agriculture Organization of the United Nations (FAO), 2005).