scieee AI-readable full text Open interactive document viewer

Harnessing GIS, remote sensing, and machine learning for sustainable management and carbon sequestration of non‑timber forest products in Gujarat, India

Mohanta, Agradeep

Abstract

Harnessing GIS, remote sensing, and machine learning for sustainable management and carbon sequestration of non-timber forest products in Gujarat, India

Full text

Vol.: (0123456789) Agroforest Syst (2025) 99:92 https://doi.org/10.1007/s10457-025-01186-9 Harnessing GIS, remote sensing, andmachine learning forsustainable management andcarbon sequestration ofnon‑timber forest products inGujarat, India ShrishtiRajput· AgradeepMohanta· BiplabBanerjee· JayantaDas· SuchiMishra· HariSankar Received: 3 February 2025 / Accepted: 24 March 2025 © The Author(s), under exclusive licence to Springer Nature B.V. 2025 Abstract This study focuses on the sustainable management of non-timber forest products (NTFPs) in the Narmada, Dang, and Panchmahal districts of Gujarat, India, emphasizing carbon sequestration and carbon credits. NTFPs such as medicinal plants, fruits, nuts, and resins play a crucial role in the local economy and biodiversity conservation. Accurate mapping and assessment of these resources are essential for implementing sustainable management and conservation strategies. Advanced spatial analysis techniques, including geographic information systems (GIS), remote sensing, and logistic regression models, were employed to analyze the spatial distribution of NTFPs. High-resolution Sentinel-2 satellite imagery and field survey data were integrated to create detailed spatial maps while, logistic regression models evaluated environmental factors like soil type, elevation, and climatic conditions affecting NTFPs distribution. The study identified that environmental variables such as litter cover, elevation, and NTFPs type are critical in determining the distribution of NTFPs, with NTFPs type accounting for 68% of the variance in distribution patterns. The logistic regression model and GIS expert system predicted NTFPs distribution with 65.43% and 70.37% accuracy, respectively. The GIS expert system demonstrated a higher specificity rate (47.05%) compared to logistic regression (35.29%), indicating its superior ability to predict NTFP absence. Both models were validated using a sample size of 81, with error matrices generated for comparative analysis. Additionally, the research explored the carbon sequestration potential of NTFPs and their implications for carbon credits. The study recommend machine learning algorithms, particularly the random forest (RF) model, for aboveground carbon stock estimation across different NTFPs regions. The RF model showed superior performance with an R2 value exceeding 0.6 across the study areas, compared to multivariate stepwise regression, which had R2 values below 0.4. The RF model’s accuracy was validated through a comparison with actual field data, achieving a root mean square error of 24.72 t hm−2 in NTFPs regions. Carbon stocks were observed to range from 50 to 250 Supplementary Information The online version contains supplementary material available at https:// doi. org/ 10. 1007/ s1045702501186-9. S.Rajput· A.Mohanta(*)· B.Banerjee Faculty ofScience, The Maharaja Sayajirao University ofBaroda, Vadodara390002, India e-mail: [email protected]; agradeepmbotan[email protected] J.Das Department ofGeography, Rampurhat College, Rampurhat, Birbhum731224, India S.Mishra WSP, Houston, TX16200, USA H.Sankar Department ofGeography, Faculty ofArts andHumanities, Kalinga University, Kotni, Chhattisgarh492101, India Agroforest Syst (2025) 99:92 92 Page 2 of 35 Vol:. (1234567890) t hm−2, depending on the region’s ecological characteristics and topographical variations. The research findings highlight the importance of integrating local ecological knowledge with spatial data and advanced modelling techniques to enhance the sustainable management of NTFPs. The potential for carbon credits offers a financial incentive for conserving NTFPs, promoting sustainable harvesting, and ensuring longterm biodiversity conservation. This study provides a replicable framework for NTFPs management and carbon stock estimation, applicable to similar ecological settings worldwide. The integration of GIS, remote sensing, and machine learning methodologies presents a robust approach for policymakers and stakeholders aiming to optimize NTFPs conservation strategies while advocating carbon credit opportunities for economic development. Keywords Ecological modelling· Spatial distribution analysis· Machine learning algorithms· Geographic information systems (GIS)· Vegetation indices Introduction Non-timber products (NTFPs) are integral to the livelihoods of NTFPs-dependent communities and play a vital role in maintaining the biodiversity and sustainability of NTFPs ecosystems. In the Narmada, Dang, and Panchmahal districts of Gujarat, India, NTFPs such as medicinal plants, fruits, nuts, and resins are seasonally harvested and hold substantial economic and ecological significance (Pilania et al. 2015; Pilania etal. 2013; Yadav etal. 2019a, b). Accurate mapping and assessment of carbon sequestration and carbon credits of these resources are critical for the implementation of sustainable NTFPs management and conservation strategies, as demonstrated in earlier works emphasizing the role of carbon accounting in community-based forest management and NTFP conservation. Pandey et al. (2016) studied carbon sequestration potential in community forests with active NTFP management. Nath et al. (2022) highlighted carbon stock estimation in agroforestry systems involving NTFPs. Tewari (1998) focused on the economics of NTFPs and carbon credit opportunities in India’s forest-dependent communities. The sustainable collection and trade of NTFPs provides significant income to local communities, contributing to both the local economy and the conservation of NTFPs resources. Globally, the sustainable harvesting of NTFPs is a critical livelihood strategy that supports economic development while promoting environmental stewardship. However, NTFPs face the threat of over-exploitation and unsustainable ecosystem management practices, emphasizing the need for accurate distribution data to inform sustainable harvest plans. In the context of Gujarat, reliable spatial data on NTFPs is crucial, particularly given the emerging opportunities for carbon credits, which incentivize the conservation of NTFPs by assigning monetary value to carbon sequestration efforts associated with NTFP management (Gupta et al. 2024; Shah 2016). The application of spatial modeling techniques has broadened to include predictions of the distribution patterns of various species and plant communities (Robinson etal. 2008; Nanang and Hauer 2008; Navarrete-Segueda etal. 2021; Asamoah etal. 2024). These models typically emphasize dominant or overstorey species (Zhang etal. 2016), but some have also considered understudied groups such as shrubs (Sardeshpande and Shackleton 2019), ferns (Elliott and Ticktin 2013), and cryptogams (Heilmann-Clausen et al. 2005). Notably, NTFPs products have seldom been the focus of such studies, a gap this research aims to address by modelling the spatial distribution of NTFPs within Gujarat’s varied landscapes. In the context of carbon credits, spatial modeling plays a crucial role in estimating carbon stocks and monitoring the carbon dynamics of NTFPs, with significant contributions from scholars such as Bustamante etal. (2016) and Fadil et al. (2024). These experts have explored the effectiveness of spatial models in carbon credit systems, evaluating how well such models predict carbon sequestration potential in diverse ecosystems. This research intends to extend these methodologies to better understand and enhance the valuation of carbon credits in NTFPs areas, particularly through the lens of NTFPs, thereby contributing a novel perspective to the discourse on environmental conservation and sustainable resource management. Logistic regression is utilized due to its demonstrated effectiveness as a statistical method for binary dependent variables, its straightforward integration with GIS software, and its capability to process complex Agroforest Syst (2025) 99:92 Page 3 of 35 92 Vol.: (0123456789) environmental data. This approach, when combined with a GIS expert system that amalgamates local knowledge and spatial data, forms a robust model for predicting the distribution of NTFPs. Such a model not only facilitates comprehensive analyses but also allows comparisons with logistic regression outputs, thereby identifying crucial environmental factors that influence NTFPs distribution. This is vital for constructing precise models and guiding sustainable management practices. Environmental variables like soil type, moisture levels, elevation, slope, and climatic conditions are integral to the distribution of NTFPs. By incorporating these variables into logistic regression models, areas of high NTFP abundance are predicted, and critical habitats necessitating conservation efforts are identified. The inclusion of local ecological knowledge through the GIS expert system ensures that the models are rooted in practical, field-based insights into NTFPs distribution, as demonstrated in previous studies highlighting the value of integrating indigenous knowledge with geospatial technologies for enhanced resource mapping and management. McCall (2003) seeking good governance in participatory-GIS: a review of processes and governance dimensions in applying GIS to local knowledge. Das etal. (2022) traditional ecological knowledge for sustainable ecosystem management. Robiglio and Mala (2005) integrating local knowledge for participatory mapping of NTFPs using GIS. NTFPs play a vital role in supporting local livelihoods, biodiversity conservation, and carbon sequestration. Accurate estimation of NTFPs carbon storage is essential for sustainable resource management, yet traditional field-based methods are often labor-intensive, costly, and impractical at larger scales (Sharma et al. 2021). To overcome these limitations, remote sensing techniques have gained prominence, leveraging multispectral data to model carbon storage across extensive regions (Zhang etal. 2023). Passive optical remote sensing systems, particularly those utilizing red-edge and greenedge spectral bands, have demonstrated sensitivity to vegetation health and chlorophyll content, which are key indicators for monitoring NTFPs biomass and disturbances (Cadez etal. 2024). Recent studies have successfully applied remote sensing and machine learning for biomass estimation. For instance, Cheng et al. (2024) combined Landsat 7 and hyperspectral data with USDA inventories to estimate NTFPs aboveground biomass in New England, while Cadez etal. (2024) developed automated methods integrating satellite imagery and field data to map NTFPs across Europe. Despite these advancements, research on scalable, regionspecific carbon stock estimation for NTFPs in semi-arid ecosystems like Gujarat remains limited. Moreover, the integration of local ecological knowledge with remote sensing models is underexplored, though it can significantly enhance spatial predictions and management strategies. Research gaps Despite notable progress, the following gaps persist: • Lack of region-specific models for NTFPs carbon stock estimation in dryland ecosystems. • Limited incorporation of local ecological knowledge into remote sensing and GIS-based carbon stock assessments. • Insufficient application of machine learning algorithms tailored to NTFPs distribution and carbon storage in India. Research questions To address these gaps, this study focuses on the following research questions: 1. How can remote sensing data (e.g., Sentinel-2) and machine learning improve the spatial estimation of NTFPs carbon stocks in Gujarat? 2. What is the role of local ecological knowledge in refining NTFPs distribution models? 3. Which environmental factors (e.g., soil type, elevation, moisture) most influence NTFPs distribution and carbon storage? Hypotheses Based on these questions, the study tests the following hypotheses: Agroforest Syst (2025) 99:92 92 Page 4 of 35 Vol:. (1234567890) • H1: Machine learning models integrating Sentinel-2 data significantly enhance the accuracy of NTFPs carbon stock estimation compared to traditional regression methods. • H2: Incorporating local ecological knowledge into GIS-based models improves predictions of NTFPs distribution. • H3: Specific environmental variables (soil type, elevation, moisture) are strong predictors of NTFPs distribution and carbon storage. Study significance By developing an integrated framework that combines remote sensing, machine learning, and local knowledge, this research aims to provide a scalable, cost-effective approach for NTFPs carbon stock estimation. The results will support sustainable NTFPs management and conservation strategies, serving as a model for similar ecosystems globally. Study area The Narmada, Dang, and Panchmahal districts of Gujarat were purposively selected due to their ecological diversity, rich NTFPs resources, and high community dependence on forest products. These areas provide varied environmental conditions ideal for analyzing NTFPs distribution, carbon storage, and developing region-specific sustainable management strategies. Narmada district, positioned at approximately 21.85° N, 73.56° E, lies within the elevational range of 100–900 m above sea level. This region features a subtropical climate with temperatures peaking during hot summers 42 °C and moderate during mild winters 15 °C. The annual rainfall averages between 900 and 1200 mm, primarily occurring during the monsoon season. The soil in Narmada is predominantly lateritic, supporting deciduous NTFPs composed mainly of teak (Tectona grandis) and bamboo (Bambusoideae). These NTFPs are rich sources of NTFPs such as honey, gums, resins, and medicinal herbs. Dang district located around 20.75° N, 73.70° E, varies in elevation from 100 to 1200 m above sea level and is part of the Western Ghats, a biodiversity hotspot. It experiences a tropical climate with heavy rainfall, receiving between 1500 and 2500 mm annually. The soil, primarily red and lateritic, sustains moist deciduous and semi-evergreen NTFPs. Dominant NTFPs species include teak, mahua (Madhuca longifolia), and various bamboos, supporting an array of NTFPs such as bamboo products, honey, and wild edibles crucial for the sustenance of local tribal communities. Panchmahal district, with coordinates at 22.75° N, 73.56° E, showcases elevations ranging from 100 to 600 m above sea level. The climate here is semi-arid to sub-humid, with temperatures that fluctuate significantly between seasons and average annual rainfall between 800 and 1000 mm. The soils are mostly black and fertile in lower elevations and mixed with rocky outcrops in higher areas, supporting dry deciduous NTFPs. These NTFPs, comprising predominantly teak and acacia (Acacia spp.), yield NTFPs such as lac, gums, resins, and medicinal plants. These districts not only display significant ecological diversity but also provide critical habitats for a variety of NTFP-yielding species. The specific climatic conditions, elevation gradients, and soil regions in each district underpin the distinct composition and distribution of NTFPs regions, which are integral to the livelihoods of the local communities and the sustainability of the region’s biodiversity as shown in Fig.1. Vegetative description These districts support varied forest types shaped by their climatic and topographic differences: • Narmada: Dominated by dry and moist deciduous forests, with key species such as teak (Tectona grandis) and bamboo (Bambusoideae), contributing NTFPs like honey, gums, and medicinal plants (Kumar and Biotechnology 2015). • Dang: Home to moist deciduous and semi-evergreen forests, rich in teak, mahua (Madhuca longi‑ folia), and bamboo, yielding products like wild edibles and bamboo crafts (Rathod etal. 2024). • Panchmahal: Characterized by dry deciduous forests with teak and acacia (Acacia spp.), producing lac, resins, and medicinal herbs (Patel etal. 2014). Agroforest Syst (2025) 99:92 Page 5 of 35 92 Vol.: (0123456789) Socio-economic, traditional, and cultural profile These regions are primarily inhabited by tribal communities such as the Bhil, Gamit, and Vasava, whose livelihoods are deeply reliant on forests. • Economy: NTFPs are crucial sources of income through the collection and sale of products like honey, gums, and bamboo crafts. • Traditions: Communities follow age-old sustainable harvesting practices, with traditional ecological knowledge passed through generations to protect forest health. Culture: Festivals, rituals, and daily practices are closely tied to the forest cycle, with ceremonies celebrating harvests of key NTFPs like mahua flowers and honey. Method Sampling design and field measurement Field sampling of NTFPs trees in the study area faced several challenges due to complex terrain and ecological variability. The dissected landscape with steep slopes and uneven elevations made plot accessibility difficult, especially in remote forest patches. Aspect (direction of slope) influenced microclimates, causing variation in species distribution between northand south-facing slopes. Soil conditions ranged from lateritic to rocky and black soils, affecting the growth and density of NTFPs species. Additionally, dense vegetation cover, particularly in Dang’s semi-evergreen forests, restricted movement and visibility during sampling. These factors collectively created logistical difficulties and contributed to the heterogeneous distribution of NTFPs, requiring careful planning to ensure representative and randomized sampling across varying ecological Fig. 1 Study area Agroforest Syst (2025) 99:92 92 Page 6 of 35 Vol:. (1234567890) zones.Moreover, even when probable NTFPs sites have been identified, it is difficult to confirm as the roots and vegetative parts are embedded in the soil, and even during the rare fruiting periods, the aboveground parts of these trees are often obscured by litter, herbs, or mosses. Without preliminary knowledge or experience, locating these plants is not a straightforward task. Thus, our data collection, conducted from August to October 2023, was assisted by local NTFPs collectors. Sample sites were designated as either"presence"where NTFPs trees were found or"absence"in sites identified by collectors as places where specific NTFPs trees had never been found. Two measures were taken to mitigate subjective bias and keep the sampling as close to random as possible. A stratified random sampling approach was used to ensure broad representation of ecological and socioeconomic variability across the study area. • Village Selection: Six villages were randomly selected from forest-dependent settlements within each district, ensuring variation in elevation, forest type, and NTFPs dependency levels. The selection was based on village lists obtained from local forest departments and community records. • Collector Selection: Seven active NTFPs collectors per village were chosen based on their experience, availability, and knowledge. Selection aimed to balance gender, age, and roles in collection to minimize personal bias in site selection. • Sampling Intensity: Sampling covered approximately 10–15% of the total forest area surrounding the selected villages, ensuring that diverse microhabitats and NTFPs species were included while remaining logistically feasible for detailed fieldwork. In each field sample plot, various parameters were recorded including topographic features (elevation, slope, aspect), floristic composition of different strata (tree layer, shrub layer), vegetation height, NTFPs canopy cover, and the presence or absence of target NTFPs trees. The specific variables documented were NTFPs Presence/Absence, NTFPs Type, Elevation (m), Slope (°), Aspect (°), Soil Depth (cm), Litter Cover (%), Average Tree Height (m), Tree Cover (%), Tree Basal Area (m2/ha), and Shrub Cover (%). NTFPs canopy cover and average tree height were visually estimated in a 25 m × 25 m plot. A 5 m × 5m subplot was used for tree cover to balance field efficiency and vegetation structure, as NTFPs trees in the study area are generally small to mediumsized and scattered, making larger plots (like 25 m × 25 m) impractical in rugged terrain (Sharma et al. 2010). For shrubs, a 2 m × 2m subplot was chosen because shrubs are low-stature, densely packed, and smaller plots allow for more accurate and manageable estimates of cover and diversity within the same sampling area (Bhattarai 2017). Using 5m × 5 m for shrubs would have increased overlap with tree layers, complicating distinct layer measurements. This multi-scale subplot design ensured effective and precise sampling of different vegetation layers without excessive time and resource demands (Reddy etal. 2015). A GPS connected with a hand-held computer Garmin eTrex 10 was used to acquire the spatial location of sample points and record each measurement. Site position was identified by an average of 30 readings. A total of 222 field observations were collected across the six selected villages in the study area. Each observation represents a sampling plot where detailed data on NTFPs tree presence or absence were recorded, along with corresponding environmental variables such as elevation, slope, aspect, soil type, and canopy cover. The number 222 was determined to ensure adequate representation of the ecological variability across different forest types, altitudes, and disturbance levels in the Narmada, Dang, and Panchmahal districts. This sample size balances statistical robustness for logistic regression analysis while remaining feasible for fieldwork within the given timeframe and terrain constraints. The table lists the environmental variables and their characteristics, where the presence or absence of NTFPs trees (Presence/Absence) is the dependent variable and the remaining are independent variables as shown in Table1. Data acquisition and preprocessing Remote sensing image acquisition andpreprocessing The research utilizes Sentinel 2 data from the GEE platform database, specifically the USGS Sentinel 2 Agroforest Syst (2025) 99:92 Page 7 of 35 92 Vol.: (0123456789) Level 2, Collection 2, Tier 1 dataset for high-quality satellite image acquisition. To ensure consistency and clarity in our data, we selected the median of all available images from August to October 2023, creating a composite satellite image with cloud cover less than 20%. This period aligns closely with the continuous NTFPs inventory conducted in Gujarat in 2023, spanning from August to October 2023, ensuring minimal temporal discrepancies between the satellite observations and field data. This alignment underscores the synchronicity in data gathering efforts, which is critical for accurate analysis. The Digital Elevation Model (DEM) data leverages the NASA DEM with a resolution of 30 m, a refined version of SRTM data enhanced with auxiliary inputs from ASTER GDEM, ICESat GLAS, and PRISM datasets. This high-resolution DEM is instrumental in our study for generating derived topographical layers such as slope and aspect. These layers, alongside terrain position indicators like channels, ridges, and planes, were processed using the ENVI 6.0 software’s topographic feature module. Our study further incorporates a NTFPs type map derived from a Sentinel 2 image captured in August to October 2023 (path/row: 132/41). The accuracy of this map is pivotal for categorizing NTFPs regions and is complemented by DEM-generated features to create a comprehensive environmental dataset. This dataset, which also includes maps of litter cover and 222 field measurements, forms the backbone of our spatial analysis. However, during the data refinement phase, samples with missing values1 were excluded, resulting in 163 valid samples. These were divided into two subsets: one half (n = 82) was used to develop the logistic regression model, and the other half (n = 81) served for model assessment and validation. This methodological division is designed to ensure robust model training and rigorous validation processes. The study employed GEE and ENVI 6.0 for NTFP mapping. GEE offers cloud-based processing, access to pre-processed Sentinel-2 imagery, and efficient creation of cloud-free composites, ideal for large-scale, time-sensitive analyses. However, it depends on stable internet and has limitations in custom algorithm implementation. ENVI 6.0 provides advanced terrain analysis and precise topographic feature extraction using high-quality NASA DEM data but requires expensive licensing and high computational power. While GEE streamlines data access and processing, ENVI enhances detailed feature extraction, making their combined use effective despite individual platform limitations. Acquisition andprocessing ofground survey data The ground survey data utilized in this study are sourced from the continuous NTFPs resource inventory in Gujarat, conducted by the NTFPs Survey of Table 1 Environmental variables used to analyze NTFPs distribution Variables No. of observations Abbreviation Data type Distribution NTFP Presence/Absence 222 NTFPs P/A Binary Binomial NTFPs Type 216 FT Category – Elevation (m) 175 ELE Numeric Approximately Normal Slope (°) 180 S Numeric Approximately Normal Aspect (°) 174 A Numeric Approximately Normal Soil Depth (cm) 120 SD Numeric Skewed Litter Cover (%) 108 LTC Numeric Skewed Average Tree Height (m) 142 H Numeric Skewed Tree Cover (%) 162 TC Numeric Approximately Normal Tree Basal Area (m2/ha) 119 BA Numeric Skewed Shrub Cover (%) 118 SHC Numeric Skewed 1 Values = At each sampling plot, the measurements of all 11 variables (as detailed in Table 1) must be systematically documented."Bad samples"are those that have missing data for one or more variables. In regression analysis, samples with incomplete data are typically omitted from the analysis. Agroforest Syst (2025) 99:92 92 Page 8 of 35 Vol:. (1234567890) India from September to December 2023. A total of 1072 fixed sample plots, each covering an area of 800 m2, were surveyed. The geographic coordinates for the center points of these sample plots were accurately captured through differential GPS surveys. The collected data encompassed detailed information on dominant tree species, biomass, and origin. This study specifically excluded the carbon storage contributions from shrubs, the understory, herbaceous layer, litter layer, and soil layer. In line with the technical guidelines provided in the ‘Main Technical Specifications for NTFPs Resource Planning and Design Survey’ issued by the Ministry of Environment, NTFPs and Climate Change, the NTFPs classification in this study was refined to include three primaries’ categories: Dang NTFPs, Narmada NTFPs, and Panchamahal NTFPs. Additionally, high-resolution Google Earth imagery from 2023 was employed for visual interpretation in Gujarat. The sampling strategy adopted was random selection, resulting in the choice of 460 nonNTFPs sample points. The dataset used to train the classification model comprises 1532 sample points, combining both field survey data and data obtained from visual interpretation. Calculation ofaboveground carbon storage The calculation of aboveground carbon storage is a critical component in understanding the comprehensive outcomes of NTFPs growth. Research consistently demonstrates a significant correlation between NTFPs biomass and stock volume, establishing a robust basis for estimating NTFPs carbon storage. This correlation is leveraged by employing a direct and efficient method that has been widely adopted in recent years. In this process, the standing volume from the acquired sample plots is multiplied by the Biomass Expansion Factor (BEF) to calculate the aboveground biomass. This biomass is then converted into biomass per hectare (t hm−2). Subsequently, multiplying the biomass per hectare by the corresponding carbon content ratio results in the aboveground carbon stock of the sample plot. The BEF used in this study is sourced from the BEF data for tree species groups provided by (Fararoda et al. 2024; Jalkanen et al. 2005; Petersson et al. 2012) and biomass relationship equations for various dominant tree species are estimated as presented in TableA1 (Online Appendix A1). Conventionally, an average carbon content rate of 0.5 or 0.45 is utilized in most domestic research to estimate NTFPs carbon stocks. However, variations in carbon content coefficients among different tree species groups necessitate a more precise approach. For tree species groups lacking direct carbon content measurements, a reference approach is adopted, substituting carbon content rates from analogous tree species groups (Ferster 2009; Humphrey 2020; UNINVERSITY 2017). Moreover, statistical analysis has indicated the significance of litter cover in predicting the distribution of certain NTFPs characteristics. To enhance the precision of our models, a litter cover map was used using a generalized linear regression model. This model utilized input data including a NTFPs type map, elevation, slope, and aspect. A backward stepwise procedure was implemented to select explanatory variables, setting significance values for “enter” and “remove” at 0.05 and 0.06, respectively. The final model, which incorporated NTFPs type and slope as input data, achieved a high level of explanatory power (R2 = 0.83, n = 108, p < 0.001) and was executed using the"model maker"module in the ERDAS IMAGINE 15.6 software. This comprehensive approach ensures a scientifically robust method for estimating and analyzing NTFPs carbon storage, contributing significantly to NTFPs management and conservation strategies. Modeling NTFP distribution withlogistic regression Logistic regression model Logistic regression was used to describe the relationship between NTFPs presence/absence and environmental predictors. The logistic regression equation transformed the binary presence/absence data into a continuous probability ranging from 0 to 1, with higher values indicating a higher probability of NTFPs presence as represented in below Eq. where Y is the probability of NTFPs presence, xn are the explanatory variables, bn are the coefficients, and xn denotes the exponential function. (1) Y = exp(b 0 +b 1 x 1 +b 2 x 2 + ⋯ +b n x n ) 1+exp(b 0 +b 1 x 1 +b 2 x 2 +⋯+b n x n) Agroforest Syst (2025) 99:92 Page 9 of 35 92 Vol.: (0123456789) Building the logistic regression model The initial selection of variables was conducted based on prior statistical analyses of environmental variables using 82 samples. The optimal model was chosen based on two criteria: the explained variance (Nagelkerke R2) and the goodness-of-fit (Hosmer and Lemeshow test statistic; see SPSS Inc., 2001 for further details) (Hall and Vincenzi 2001). Spatial implementation of logistic regression model The spatial implementation of the model was executed using the"model maker"module within ERDAS IMAGINE 15.6 software. The NTFPs type map was recoded to numerical coefficients representing specific NTFPs regions. For example, pixels corresponding to Sterculia urens were assigned a coefficient value of 0.659. The coefficients for various NTFPs regions are detailed in Table2. The probability surface for NTFPs distribution was subsequently generated as the output of the logistic regression analysis. Presence/absence transformation and sensitivity analysis forthethreshold levels The model’s performance was assessed based on presence or absence, with probability values ranging from 0 to 1 being converted into binary presence/absence data at specified thresholds. While a probability value of 0.5 is commonly used to indicate presence, as suggested by (Manel etal. 1999), this threshold may not be optimal for all scenarios. Therefore, a sensitivity analysis was conducted to evaluate threshold levels from 0.4 to 0.8. Modelling distribution ofNTFPs withaGIS expert system GIS expert system A GIS expert system is a computer-based tool that emulates the decision-making processes of human experts, specifically designed to address challenges related to GIS (Das etal. 2020; Skidmore etal. 1996). The GIS expert system employed in this study was developed within the ENVI-IDL environment by the Natural Resource Department at the International Institute for Geoinformatics Science and Earth Observation, the Netherlands. Bayesian theory, as outlined by Aspinall (1992), Aspinall and Veitch (1993) and Skidmore (1989), Skidmore etal. (1996), served as the inference engine for this system. Within this GIS expert system, the presence of NTFPs at a given location (a hypothesis) is inferred based on the available evidence. Input data and knowledge‑based rule formula‑ tion The selection of environmental variables and the formulation of rules were guided by data availability and the integration of knowledge from multiple sources, including: (a) existing literature; (Tewari 1998, 2000; Yadav etal. 2019a, b) (b) local knowledge obtained through discussions with NTFPs collectors; (c) personal field observations; and (d) results from statistical analyses. In instances where there were discrepancies among these sources, subjective decisions were made based on field expertise. The data layers utilized included: (a) Litter cover; (b) Terrain position; (c) slope; (d) Digital elevation model (DEM); (e) Aspect; and (f) NTFPs Trees type. Figure2 provides detailed probability estimates for these criteria. Table 2 Coefficients for each tree species in the logistic regression model Butea monosperma was chosen as the reference species due to its widespread distribution and significant ecological and economic role in NTFPs ecosystems. It serves as a neutral baseline for comparison; hence its coefficient is assigned a zero value in the logistic regression model (Rai etal. 2021) Tree species Coefficients Family Common name Uses Madhuca indica (Mi) − 7.918 Sapotaceae Mahua Edible oil, medicinal, ritual Goruga pinnata (Gp) − 8.443 Anacardiaceae Indian Sumac Medicinal, dye Boswellia serata (Bs) − 8.807 Burseraceae Indian Frankincense Resin, medicinal Sterculia urens (Su) 0.659 Malvaceae Ghost Tree Gum, medicinal Acacia nilotica (An) 0.204 Fabaceae Babool Gum, fodder, tannin Butea monosperma (Bm) 0 Fabaceae Flame of the Forest Medicinal, dye, fodder Terminalia bellirica (Tb) 0.452 Combretaceae Baheda Medicinal, dye Bauhinia racemosa (Br) − 0.389 Fabaceae Bidi Leaf Tree Fodder, medicinal, fiber Agroforest Syst (2025) 99:92 92 Page 16 of 35 Vol:. (1234567890) The model is expressed by the following shown below Eq: where M represents the probability of NTFPs occurrence, F denotes NTFPs type, LTC corresponds to litter cover, and ELE indicates elevation. The specific coefficients for each NTFPs type are detailed in Table2. Model performance and sensitivity The maps produced by the two models demonstrated some variations Fig. 2. The variations between the two models arise due to differences in their underlying M = exp(−12.616 +F+0.043LTC +0.003ELE) 1 + exp (− 12.616 + F + 0.043LTC + 0.003ELE) methodologies and decision rules. The GIS expert system relies on rule-based knowledge, integrating expert-defined thresholds and relationships among environmental variables, which can result in more conservative predictions, particularly for absence areas. In contrast, the logistic regression model uses statistical relationships based on field data, which may lead to overprediction of presence areas due to its probabilistic nature and sensitivity to imbalanced datasets. These methodological differences influence how each model interprets spatial patterns, causing variations in mapped outputs and accuracy metrics. The GIS expert system achieved a marginally higher overall mapping accuracy compared to the logistic regression model (70.37% versus 65.43%; refer to Tables6 and 7). Nevertheless, the Z-test (for Kappa) statistic indicated no significant difference in Table 6 Kruskal–Wallis ANOVA test of relationship between environmental variables and NTFPs distribution Numbers in brackets are the sample sizes *Denotes the significant test results (p < 0.05) Variable Subclasses × 1 Subclasses × 2 Subclasses × 3 Subclasses × 4 p-values Elevation < = 3400 (60) 3401–3500 (61) 3501–3600 (41) > 3600 (19) *0.0009 Slope < = 10 (26) 11–20 (63) > 20 (91) – *0.032 Aspect < = 90 (25) 91–180 (60) 181–270 (59) 271–360 (30) 0.0061 Surface Depth < = 5 (41) 6–10 (55) > 10 (24) – 0.4899 Litter Cover < = 30 (3) 31–60 (12) 60–80 (33) > 80 (60) *0.0004 Tree Height < = 5 (5) 6–15 (82) 16–25 (40) > 25 (15) *0.0000 Tree Cover < = 20 (6) 21–50 (74) 51–80 (70) > 80 (12) *0.0064 Basal Area < = 30 (59) 31–70 (42) > 70 (8) – 0.0586 Shrub Cover < = 20 (36) 21–60 (62) > 60 (21) – 0.2432 Table 7 Mann–Whitney U test for difference of probabilities of NTFPs occurrence at various subclasses of environmental variables Numbers in brackets are the sample sizes *Denotes the significant test results (p < 0.05) Environmental variable Subclasses 1 Subclasses 2 p values Elevation ≤ 3400 (60) 3401–3500 (61) *0.001 3401–3500 (61) 3501–3600 (41) *0.020 3501–3600 (41) > 3600 (19) 0.101 Tree Height ≤ 5 (5) 6–15 (82) *0.003 6–15 (82) 16–25 (40) *0.001 16–25 (40) > 25 (15) 0.19 Tree Cover ≤ 20 (6) 21–50 (74) *0.017 21–50 (74) 51–80 (70) *0.041 51–80 (70) > 80 (12) *0.011 Litter Cover ≤ 30 (3) 31–60 (12) *0.022 31–60 (12) 61–80 (33) *0.0008 61–80 (33) > 80 (60) *0.000 Slope ≤ 10 (26) 11–20 (63) 0.17 11–20 (63) > 20 (91) 0.93 > 20 (91) ≤ 10 (26) *0.047 Agroforest Syst (2025) 99:92 Page 17 of 35 92 Vol.: (0123456789) accuracy levels between the models (Table 7). Both models exhibited identical sensitivity rates of 87.23% but differed in their specificities: 47.05% for the GIS expert system and 35.29% for the logistic regression model (Table6). This indicates that both models were equally effective in predicting the presence of the subject but were less accurate in predicting its absence. Despite this, the GIS expert system demonstrated a superior capability in predicting absence. It is noteworthy that the logistic regression model predicted a presence of NTFPs in 76.9% of the study area, in contrast to 61.4% as predicted by the GIS expert system. Results from the sensitivity analyses are presented in Fig.3. The accuracy level of the logistic regression model remained relatively stable at approximately 65.43% when the threshold was set below 0.65; however, accuracy decreased sharply above this threshold due to a significant increase in false presence predictions Fig. 4. Consequently, a threshold value of 0.5 is deemed appropriate in Table9. For the GIS expert system, the overall accuracy remained relatively stable when the a priori probabilities for absence ranged between 0.01 and 0.8 as given in Table10. Nonetheless, specificity and sensitivity exhibited considerable fluctuations; these metrics were inversely related, with specificity increasing as sensitivity decreased Fig.5. Classification results and accuracy evaluation Classification results To enhance classification performance, this study conducted comparative experiments across five distinct feature combination schemes in the study area. Initially, seven spectral features were utilized with four classification algorithms. Further experiments incorporated additional factors, which improved the overall accuracy (OA) and Kappa coefficients significantly. Utilizing solely the spectral band features, the overall accuracies for the RF, CART, GBT, and SVM algorithms were recorded at 0.7165, 0.6813, 0.7130, and 0.6127, respectively, with corresponding Kappa coefficients of 0.6147, 0.5672, 0.6068, and 0.4781 in the study area. However, the integration of five features—spectral bands, vegetation index, texture, principal component analysis, and terrain—markedly increased the accuracies to 0.8496, 0.8133, 0.8273, and 0.7315 for RF, CART, GBT, and SVM algorithms, respectively, with Kappa coefficients improving to 0.7646, 0.7186, 0.7320, and 0.6388. The inclusion of vegetation index features demonstrated a positive trend in both OA and Kappa coefficients across the four algorithms. Conversely, the addition of texture features resulted in a decrease in both metrics. The subsequent integration of principal component analysis and terrain features varied in its impact on accuracy and Kappa coefficients. The SVM classifier showed comparatively lower performance with an overall accuracy of 0.7315 and a Kappa coefficient of 0.6388. Among the evaluated algorithms, the RF classifier exhibited the highest performance with an OA of 0.84 and a Kappa coefficient of 0.76. This study thus relied on the RF classification results to estimate the aboveground carbon stock in the NTFPs of Gujarat. The classification outcomes for the four algorithms are presented in Fig.6. Accuracy assessment The classification results of four different classification algorithms are evaluated for precision using the Confusion Matrix method. The results from the Error Matrix for each classification algorithm are presented in Fig. 7. According to Fig. 7 it is evident that the RF algorithm with (Acc: 0.71) outperforms other machine learning algorithms in terms of Cartographic Accuracy and Consumer Accuracy. Therefore, the RF algorithm is more suitable for the classification of NTPFs in the study area than other machine algorithms. Table 8 Nagelkerka R2 and Hosmer and Lemeshow statistic for logistic regression with different input data (n = 82) *Significant goodness of fit (p < 0.05) Variables used Nagelkerka R2 (%) Hosmer and Lemeshow statistic NTFPs Type 68 0.999* Litter Cover 17.3 0.023 Elevation 6.8 0.382* NTFPs Type, Litter Cover 72.2 0.337* NTFPs Type, Litter Cover, Elevation 73.2 0.94* Agroforest Syst (2025) 99:92 92 Page 18 of 35 Vol:. (1234567890) Fig. 3 A–C Presence/absence of NTFPs predicted by: a the logistic regression (threshold = 0.5); b the GIS expert system (a priori for absence = 0.4) for Dang, Panchmahal, Narmada areas NTFPs Agroforest Syst (2025) 99:92 Page 19 of 35 92 Vol.: (0123456789) Fig. 3 (continued) Fig. 4 The accuracy of the logistic regression model varies with different thresholds that distinguish between"presence"and"absence."A decrease in accuracy for thresholds greater than 0.65 is indicative of an increase in false"presence"outcomes Agroforest Syst (2025) 99:92 92 Page 20 of 35 Vol:. (1234567890) Analysis ofmodel variable correlation The Pearson correlation coefficient indicates the degree of linear correlation between two model variables. A higher absolute value of the coefficient signifies a stronger correlation. Among the 80 independent variables selected, 21 exhibit a very significant correlation with carbon storage (0 < p < 0.01). Specifically, carbon storage shows a significant negative correlation with spectral bands b1–b7 at the 0.01 level. In contrast, it is positively correlated with canopy density and elevation at the same significance level. Additionally, carbon storage is significantly correlated with all other independent variables at the 0.01 level as shown in Fig.8. Estimation ofaboveground carbon stock inNTFPs The estimation of aboveground carbon stock in the study area was performed using different models, and their effectiveness was assessed through several statistical metrics, including the coefficient of determination (R2), root mean square error (RMSE), relative RMSE (rRMSE), and mean absolute error (MAE). As depicted in Fig. 9, the model for the trees of NTFPs developed through multivariate stepwise regression shows a relatively low R2 value of 0.23. The corresponding error metrics for this model are also high, with an RMSE of 35.471, an rRMSE of 32.881, and an MAE of 24.442. Similarly, all three NTFPs regions analyzed using this regression approach reveal R2 values below 0.4, reflecting considerable prediction errors and suboptimal model performance. In contrast, the RF regression model demonstrates a significantly improved performance across all NTFPs regions. For the broad-leaved NTFPs, the RF model achieves a higher R2 of 0.663, with reduced error metrics: an RMSE of 24.722, an rRMSE of 27.46, and an MAE of 17.64. This trend of superior performance extends to other NTFPs regions, such as those found in Panchmahal and Dang regions, where the RF and decision tree models exhibit R2 values exceeding 0.6. These results suggest a strong model fit and a greater ability to accurately estimate aboveground carbon stocks. The comparative analysis highlights the clear advantage of the RF model over the multivariate stepwise regression, particularly in handling complex ecological data and interactions. Given its superior efficacy, the RF model was chosen as the preferred method for estimating aboveground carbon storage in the current study area. As illustrated in Fig.9, the contribution percentages to carbon storage vary significantly across the three NTFPs regions, further underscoring the importance of using robust models like RF to capture these variations accurately. This approach not only provides a more reliable estimation but also enhances our understanding of carbon dynamics in different NTFP-dominated landscapes. Table 9 Error matrices of the two models: (a) logistic regression (threshold = 0.5) and (b) GIS expert system (a prior threshold for absence = 0.4) No. of pixels from reference Absence Presence Total (a) Logistic Regression No. of pixels from model Absence 11 5 16 Presence 23 42 65 Total 34 47 81 Overall accuracy 65.43% (b) GIS Expert System Absence 16 6 22 Presence 18 41 59 Total 34 47 81 Overall accuracy 70.37% Table 10 Statistics for evaluation of model performance NS not significant (p < 0.05) Model Overall accuracy (%) Sensitivity (%) Specificity (%) Kappa coefficient (K) Kappa variance Z value Logistic regression 65.43 87.23 35.29 0.2343 0.00223 1.89(NS) Expert system 70.37 87.23 47.05 0.3605 0.00225 – Agroforest Syst (2025) 99:92 Page 21 of 35 92 Vol.: (0123456789) The study assesses model accuracy by comparing actual data from the remaining 30% of plots within the study area against predicted carbon stock values. Figure 10 presents a scatter plot that demonstrates a close alignment between the actual and predicted carbon stock values. This visual evidence confirms the superior performance of RF and Decision Tree regression (DTR) compared to Multivariate Stepwise Regression (MLR) for the dataset under consideration as depicted in Fig.10. To develop these models, 70% of the plot data was utilized for training purposes, employing MLR, RF, and DTR approaches to estimate carbon stock. The remaining 30% of the data was reserved for validating the accuracy of these models. The models were constructed using carbon stock data obtained from various NTFPs regions, incorporating 21 highly correlated predictive factors that significantly influence carbon stock estimation. Both RF and DTR regression models emphasize the importance of input factors by assigning modeling contribution values, as depicted in Fig.10. These values highlight the prioritized influence of each factor on carbon stock prediction, showcasing the capability of RF and DTR models to handle complex relationships within the dataset effectively. This approach reinforces the reliability of RF and DTR regression as robust alternatives to traditional regression methods in carbon stock estimation. For the NTFPs regions, the study employed CART (Classification and Regression Trees) and RF models to estimate forest carbon stocks using remote sensing data due to their capabilities in handling complex, non-linear relationships between variables. The CART model works by dividing the dataset into subsets that minimize variance, creating decision trees that effectively capture these relationships. However, the RF model, an ensemble approach combining multiple decision trees, demonstrated superior performance with higher accuracy and robustness, mitigating overfitting and integrating data from various sources as shown in Fig.11. For different NTFPs regions, the RF model achieved a higher coefficient of determination (R2) values, above (0.666) for Narmada region (0.636), Panchmahal region (0.663), and Dang region (0.638), indicating better fitting performance compared to CART. The study results confirm that RF is the most reliable model for forest carbon stock estimation, providing both accurate predictions and ranked importance of influencing factors, which is crucial for ecological management and policymaking in NTFPs regions. Spatial distribution characteristics ofaboveground carbon stock ofNTFPs trees instudy area In the RF regression analysis, estimated maps of carbon storage in Gujarat’s distinct NTFPs Trees regions have been derived as shown in Fig.12. From these estimates, the carbon stock of the study area is observed to be substantially high, ranging approximately from 100 to 200 t hm−2, particularly in the elevated southern regions of Gujarat, which include Fig. 5 The specificity, sensitivity, and overall accuracy change with an increase in the a priori probability set for absence in a GIS expert system. This also implies a decrease in the a priori probability set for presence, since the sum of the a priori probabilities for absence and presence always equals 1 Agroforest Syst (2025) 99:92 92 Page 22 of 35 Vol:. (1234567890) Panchmahal, Narmada, and Dang. This high carbon stock can be attributed to these areas being situated in a tropical zone with abundant rainfall and prolonged frost-free periods conducive to NTFPs trees in Narmada. Additionally, the complex topography of these regions, which includes diverse NTFPs regions such as seasonal rain NTFPs and montane NTFPs, further contributes to the elevated carbon stocks. In contrast, the carbon stock in Panchmahal region NTFPs is noted to be significantly higher in the northern highlands of Gujarat, particularly in the areas with extensive coverage. Here, carbon stocks are recorded between 100 and 250 t hm−2. However, the accuracy of carbon stock assessments in these regions is affected by limited field data and potential classification errors. The distribution of Dang NTFPs across the study area is comparatively sparse, with carbon stocks consistently observed within the range of approximately 50 to 150 t hm−2. It is evident that Dang NTFPs, owing to their higher biomass per unit volume, possess superior carbon sequestration capabilities compared to coniferous NTFPs. Consequently, NTFPs play a pivotal role in the carbon dynamics of Gujarat’s NTFPs ecosystems. The overall distribution of NTFPs carbon stocks across the study area, as illustrated in Fig. 13, shows a general trend of higher values in the southern regions, decreasing towards the north, with a Fig. 6 Comparison of overall accuracy and Kappa coefficient of different classification algorithms with different features. a Spectral features b Spectral and vegetation index features c Spectral, vegetation index, and texture features d Spectral, vegetation index, texture, and principal component analysis features e Spectral, vegetation index, texture, principal component analysis, and terrain feature Agroforest Syst (2025) 99:92 Page 23 of 35 92 Vol.: (0123456789) decline from west to east. Additionally, the aboveground carbon stock distribution aligned with DEM data, as depicted in Fig.13, indicates that the carbon stock of Narmada NTFPs predominantly occurs at altitudes ranging from 1500 to 2000 m, comprising a significant portion of the total carbon stock. Panchmahal region NTFPs and Dang NTFPs show a higher carbon stock presence at elevations above 2000 m and 2500 m, respectively. This stratification by elevation supports the understanding of carbon stock variations within NTFPs regions across different topographic gradients. Discussion Model performance and its factors In this study, the model’s performance and its influential factors were scrutinized. Both the GIS expert system and logistic regression were utilized to predict the distribution of NTFPs Trees in the study area, using the same initial dataset complemented by three key layers: NTFPs type, litter, and elevation. Additional factors such as slope and aspect were integrated Fig. 7 Confusion matrices of four algorithms Agroforest Syst (2025) 99:92 92 Page 24 of 35 Vol:. (1234567890) based on insights from local experts and NTFPs Collectors, which could explain variations in the model outputs. The GIS expert system slightly outperformed the logistic regression model, improving mapping accuracy by about 5%. However, a Z-test confirmed no significant difference in the accuracy levels between the two methods. The spatial distribution predictions differed visually between the models; the expert system predicted fewer NTFPs presence areas, aligning more closely with field observations, especially at elevations above 300 m where NTFPs Trees presence is unlikely. The performance of these models is influenced by the quality of input data, model constraints, and the methodology used for generating topographic data, such as interpolating DEM from contours and calculating slope and aspect. The GIS system heavily relies on expert knowledge, which requires subjective assessments when sources disagree. For instance, local preferences for a southerly aspect versus empirical evidence suggesting a northerly preference were reconciled by incorporating field knowledge and broader geographical studies. The logistic regression’s limitations stemmed from its inability to Fig. 8 Correlation between carbon stocks and 21 independent variables. Note: b1, b2, b3, b4, b5, b6, and b7 correspond to the 1 st, 2nd, 3rd, 4 th, 5 th, 6 th, and 7 th bands of the Landsat 8 imagery, respectively. NLI stands for Non-linear Vegetation Index; ‘brightness’ and ‘wetness’ represent the tasseled cap transformation factors. ‘Mean’, ‘entropy’, ‘second’, ‘correlation’, ‘variance’, ‘contrast’ and ‘dissimilarity’ respectively represent Mean Texture, Entropy Texture, Second-Order Angular Moment Texture, Correlation Texture, Variance Texture, Contrast Texture, and Dissimilarity Texture Agroforest Syst (2025) 99:92 Page 25 of 35 92 Vol.: (0123456789) accurately model the non-linear relationship between NTFPs trees presence and elevation, overestimating presence at high elevations. Integrating more complex regression functions or exploring rule-based regression models like GARP might mitigate these issues. Both models struggled to predict NTFPs absence, likely due to its prevalence in the study area and the inherent challenges in accurately identifying absence, leading to lower specificity in the results. Further investigation into these aspects is necessary for refining the models and enhancing their predictive accuracy. The observed variations and performance differences between the GIS expert system and logistic regression model in this study are consistent with findings from earlier research. For instance, Pearce et al. (2001) demonstrated that expert-based GIS models often provide better spatial representation in ecological studies when expert knowledge is available, particularly in complex terrains, as they can incorporate nuanced, site-specific insights that statistical models may overlook. Similarly, Store and Jokimäki (2003) highlighted that expert systems are advantageous insituations where environmental relationships are not purely linear and can effectively integrate subjective knowledge with spatial data. On the other hand, logistic regression models have been widely used in species distribution modeling but often face challenges with non-linear ecological processes and overprediction in high-suitability areas. Guisan and Zimmermann (2000) emphasized that while logistic regression is robust for presenceabsence data, it may struggle with complex ecological gradients, such as elevation, which aligns with our findings of overpredicted NTFPs presence at higher altitudes. Furthermore, Pearce and Ferrier (2000) identified that logistic regression models tend to overpredict presence in areas with limited absence data, which may explain the lower specificity observed in this study. Moreover, the difficulty both models faced in accurately predicting absence is also well-documented. Barbet‐Massin et al. (2012) discuss how the scarcity of true absence data in ecological studies often reduces model specificity, especially in presencedominated landscapes. These findings support the Fig. 9 The images presented are Sun burst charts utilized for the comparative analysis of errors in machine learning models across various datasets. In chart (a), two error metrics are depicted: R2 and "RMSE,"both labeled in black and black bold numerals respectively, distributed across distinct concentric rings with corresponding numerical values for diverse datasets. Chart (b) delineates predictive model performance, partitioned into three rings that display MAE and RMSE, expressed in percentages, thereby providing a thorough examination of model evaluations through several statistical measurements. The colors green, blue, and orange represent "Cart," "RF," and "MLR"respectively, highlighting different modeling techniques or data categories within the charts Agroforest Syst (2025) 99:92 92 Page 32 of 35 Vol:. (1234567890) The findings underscore the vital role of spatial modeling and advanced machine learning techniques in informing sustainable NTFPs management, with direct implications for carbon credit schemes and biodiversity conservation. This integrated framework offers a replicable approach for policymakers aiming to balance ecological preservation with socioeconomic benefits. Strengths of the study include the combination of high-resolution satellite imagery with field data and the comparative evaluation of statistical and machine learning models, enhancing the reliability of distribution and carbon stock estimates. However, limitations arise from the moderate sample size (n = 81) and potential seasonal variability in NTFPs availability, which may affect model generalization. Site-specific recommendations include prioritizing conservation efforts in regions with higher carbon stocks and significant NTFPs diversity, promoting community-based management practices, and establishing local carbon credit initiatives to incentivize sustainable harvesting. Future research should focus on expanding temporal data coverage and incorporating socio-economic factors to refine decision-making frameworks for NTFPs management in Gujarat and beyond. Acknowledgements The authors would like to express their sincere gratitude to the local communities and NTFPs collectors in Gujarat, India, whose knowledge and cooperation were invaluable for the field data collection. We also appreciate the support provided by administrative support from the forest department. The authors also acknowledge the helpful comments and suggestions of the anonymous reviewers, which greatly improved the quality of this manuscript. Author contributions AM and BB contributed equally to the conceptualization and design of the study, including the development of research questions and methodology. SR provided crucial support in data collection, particularly in field surveys, and contributed to the analysis of spatial data using GIS tools. JD contributed to the interpretation of results and provided expertise in remote sensing data processing. SM assisted in the formulation of the statistical models and the carbon sequestration analysis, ensuring the accuracy of the results. HS played a key role in reviewing and refining the manuscript, ensuring the integration of ecological knowledge and making significant contributions to the discussion of the findings. Funding The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Data availability The data used in this study are available upon reasonable request from the corresponding author. Declarations Conflict of interest The authors declare no conflict of interest related to the research, authorship, or publication of this article. All authors have contributed to the work in a collaborative manner, and the findings of this study are independent of any commercial or financial interests. References Ai M etal (2021) Optimal subsampling algorithms for big data regressions. Stat Sin 31(2):749–772 Ali I etal (2015) Review of machine learning approaches for biomass and soil moisture retrievals from remote sensing data. Remote Sens 7(12):16398–16421 Andersen PK, Skovgaard LT (2010) Regression with linear predictors. Springer, Berlin Arya A etal (2017) Carbon sequestration analysis of dominant tree species using geo-informatics technology in Gujarat state (INDIA). Int J Environ Geoinform 4(2):79–93 Asamoah O etal (2024) The perception of the locals on the impact of climate variability on non-timber forest products in Ghana. Ecol Front 44(3):489–499 Aspinall R (1992) An inductive modelling procedure based on Bayes’ theorem for analysis of pattern in spatial data. Int J Geogr Inf Syst 6(2):105–121 Aspinall R, Veitch N (1993) Habitat mapping from satellite imagery and wildlife survey data using a Bayesian modeling procedure in a GIS. Photogramm Eng Remote Sens 59(4):537–543 Barbet-Massin M et al (2012) Selecting pseudo-absences for species distribution models: how, where and how many? Methods Ecol Evol 3(2):327–338 Belgiu M et al (2016) Random forest in remote sensing: a review of applications and future directions. ISPRS J Photogramm Remote Sens 114:24–31 Bhattarai KR (2017) Variation of plant species richness at different spatial scales. Bot Orient J Plant Sci 11:49–62 Breiman L (2001) Random forests. Mach Learn 45:5–32 Breiman L (2017) Classification and regression trees. Routledge, London Bustamante MM etal (2016) Toward an integrated monitoring framework to assess the effects of tropical forest degradation and recovery on carbon stocks and biodiversity. Glob Change Biol 22(1):92–109 Cadez L etal (2024) Mapping forest growing stock and its current annual increment using random forest and remote sensing data in Northeast Italy. Forests 15(8):1356 Camps-Valls G etal (2008) Kernel-based framework for multitemporal and multisource remote sensing data classification and change detection. IEEE Trans Geosci Remote Sens 46(6):1822–1835 Chen L et al (2018) Estimation of forest above-ground biomass by geographically weighted regression and machine learning with sentinel imagery. Forests 9(10):582 Agroforest Syst (2025) 99:92 Page 33 of 35 92 Vol.: (0123456789) Cheng F etal (2024) Remote sensing estimation of forest carbon stock based on machine learning algorithms. Forests 15(4):681 Cohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Meas 20(1):37–46 Congalton RG (1991) A review of assessing the accuracy of classifications of remotely sensed data. Remote Sens Environ 37(1):35–46 Danandeh Mehr A (2021) Drought classification using gradient boosting decision tree. Acta Geophys 69(3):909–918 Das AC et al (2020) Integrating an expert system, GIS, and satellite remote sensing to evaluate land suitability for sustainable tea production in Bangladesh. Remote Sens 12(24):4136 Das M et al (2022) Nexus between indigenous ecological knowledge and ecosystem services: a socio-ecological analysis for sustainable ecosystem management. Environ Sci Pollut Res 29(41):61561–61578 Debela T (2017) Species-specific allometric equations for biomass estimation of three selected tree species in Egdu Forest of Oromia National Regional State (Master’s thesis, Addis Ababa University). Addis Ababa Univ Res J 5(2):45–78 Deng W etal (2024) Carbon storages and densities of different ecosystems in Changzhou city, China: subtropical forests, urban green spaces, and wetlands. Forests 15(2):303 Elith J et al (2009) Species distribution models: ecological explanation and prediction across space and time. Annu Rev Ecol Evol Syst 40(1):677–697 Elliott DD, Ticktin T (2013) Epiphytic plants as NTFPs from the forest canopies: priorities for management and conservation. In: Lowman M, Devy S, Ganesh T (eds) Treetops at Risk. Springer, New York, NY. https:// doi. org/ 10. 1007/ 978-146147161-5_ 46 Fadil S etal (2024) An integrating framework for biomass and carbon stock spatialization and dynamics assessment using unmanned aerial vehicle LiDAR (LiDAR UAV) data, landsat imagery, and forest survey data in the mediterranean cork oak forest of Maamora. Land 13(5):688 El-Omairi MA, El Garouani A (2023) A review on advancements in lithological mapping utilizing machine learning algorithms and remote sensing data. Heliyon 9(9). https:// doi. org/ 10. 1016/j. heliy on. 2023. e20168 Fararoda R etal (2024) Improving plot-level above ground biomass estimation in tropical Indian forests. Eco Inform 81:102621 Fentaw AE, Abegaz A (2024) Analyzing land use/land cover changes using google earth engine and random forest algorithm and their implications to the management of land degradation in the Upper Tekeze Basin, Ethiopia. Sci World J 2024(1):3937558 Ferchichi A etal (2022) Forecasting vegetation indices from spatio-temporal remotely sensed data using deep learning-based approaches: a systematic literature review. Eco Inform 68:101552 Ferster CJ (2009) Using inventory based field attributes to characterize carbon stocks and carbon stock changes within eddy-flux covariance tower footprints. University of British Columbia, Vancouver Fielding AH, Bell JF (1997) A review of methods for the assessment of prediction errors in conservation presence/absence models. Environ Conserv 24(1):38–49 Foody GM (2002) Status of land cover classification accuracy assessment. Remote Sens Environ 80(1):185–201 Fox EW etal (2017) Assessing the accuracy and stability of variable selection methods for random forest modeling in ecology. Environ Monit Assess 189:1–20 Freeman EA etal (2016) Random forests and stochastic gradient boosting for predicting tree canopy cover: comparing tuning processes and model performance. Can J for Res 46(3):323–339 Gao Y et al (2018) Comparative analysis of modeling algorithms for forest aboveground biomass estimation in a subtropical region. Remote Sens 10(4):627 Ghimire B et al (2010) Contextual land-cover classification: incorporating spatial dependence in land-cover classification models using random forests and the Getis statistic. Remote Sens Lett 1(1):45–54 Gorelick N etal (2017) Google earth engine: planetary-scale geospatial analysis for everyone. Remote Sens Environ 202:18–27 Guisan A, Zimmermann NE (2000) Predictive habitat distribution models in ecology. Ecol Model 135(2–3):147–186 Gupta SK etal (2024) Optimizing land use for climate mitigation using nature based solution (NBS) strategy: a study on afforestation potential and carbon sequestration in Rajasthan, India. Discov Geosci 2(1):36 Hall S, Vincenzi DA (2001) Software: SPSS for Windows SPSS Inc. 233 S. Wacker Dr., 11th Floor Chicago, IL 60606 (312) 651–3000 www.spss.com. Ergonom des 9(3):27–29 Hardy P-Y etal (2017) Strengthening the resilience of smallscale fisheries: a modeling approach to explore the use of in-shore pelagic resources in Melanesia. Environ Model Softw 96:291–304 Heilmann-Clausen J etal (2005) Cryptogam communities on decaying deciduous wood—does tree species diversity matter? Biodivers Conserv 14:2061–2078 Huang C (2008) Characteristics of carbon stock and its spatial differentiation in the forest ecosystem of Sichuan. Sichuan Agricultural University, Ya’an (in Chinese) Humphrey A (2020) Determination of species abundance, diversity and carbon stocks in Kakamega Forest Ecosystem. Moi University, Kesses Jalkanen A et al (2005) Estimation of the biomass stock of trees in Sweden: comparison of biomass equations and age-dependent biomass expansion factors. Ann for Sci 62(8):845–851 James G etal (2023) Linear regression. An introduction to statistical learning: with applications in python. Springer, Berlin, pp 69–134 Jozdani SE et al (2019) Comparing deep neural networks, ensemble classifiers, and support vector machine algorithms for object-based urban land use/land cover classification. Remote Sens 11(14):1713 Kokaly RF, Clark RN (1999) Spectroscopic determination of leaf biochemistry using band-depth analysis of absorption features and stepwise multiple linear regression. Remote Sens Environ 67(3):267–287 Kumar V (2015) Impact of non timber forest produces (NTFPs) on food and livelihood security: an economic study of Agroforest Syst (2025) 99:92 92 Page 34 of 35 Vol:. (1234567890) tribal economy in Dang’s District of Gujarat, India. Int J Agric Environ Biotechnol 8(2):387–404 Kunwar RM et al (2020) Change in forest and vegetation cover influencing distribution and uses of plants in the Kailash Sacred Landscape, Nepal. Environ Dev Sustain 22(2):1397–1412 Li Y etal (2018) Deep learning for remote sensing image classification: a survey. Wiley Interdiscip Rev Data Min Knowl Discov 8(6):e1264 Manel S etal (1999) Comparing discriminant analysis, neural networks and logistic regression for predicting species distributions: a case study with a Himalayan river bird. Ecol Model 120(2–3):337–347 McCall MK (2003) Seeking good governance in participatoryGIS: a review of processes and governance dimensions in applying GIS to participatory spatial planning. Habitat Int 27(4):549–573 Nanang DM, Hauer GK (2008) Integrating a random utility model for non-timber forest users into a strategic forest planning model. J for Econ 14(2):133–153 Nath PC, Ahmed A, Bania JK etal (2022) Tree diversity and biomass carbon stock along an altitudinal gradient in old-growth secondary semi-evergreen forests in North East India. Trop Ecol 63:20–29. https:// doi. org/ 10. 1007/ s4296502100185-y Navarrete-Segueda A etal (2021) Timber and non-timber forest products in the northernmost Neotropical rainforest: ecological factors unravel their landscape distribution. J Environ Manag 279:111819 Nichol JE, Sarker MLR (2010) Improved biomass estimation using the texture parameters of two high-resolution optical sensors. IEEE Trans Geosci Remote Sens 49(3):930–948 Olofsson P etal (2014) Good practices for estimating area and assessing accuracy of land change. Remote Sens Environ 148:42–57 Pandey SS etal (2016) Assessing the roles of community forestry in climate change mitigation and adaptation: a case study from Nepal. For Ecol Manag 360:400–407 Papazafeiropoulos G (2024) Stepwise regression for increasing the predictive accuracy of artificial neural networks: applications in benchmark and advanced problems. Modelling 5(1):153–179 Patel Y etal (2014) Research article floristic diversity study of Kalol taluka, Panchmahal, Gujarat, India. Int J Recent Sci Res 5(11):2089–2094 Patil P etal (2012) Above ground forest phytomass assessment in southern Gujarat. J Indian Soc Remote Sens 40:37–46 Pearce J, Ferrier S (2000) Evaluating the predictive performance of habitat models developed using logistic regression. Ecol Model 133(3):225–245 Pearce JL etal (2001) Incorporating expert opinion and finescale vegetation mapping into statistical models of faunal distribution. J Appl Ecol 38(2):412–424 Petersson H etal (2012) Individual tree biomass equations or biomass expansion factors for assessment of carbon stock changes in living biomass—a comparative study. For Ecol Manag 270:78–84 Phillips SJ etal (2006) Maximum entropy modeling of species geographic distributions. Ecol Model 190(3–4):231–259 Pilania P et al (2013) 2. Study of floral species diversity of Dahod District (Gujarat) Western India by PK PILANIA 1_RV GUJAR2_PM JOSHI2 AND. Life Sci Leaflets 44:9–18 Pilania P et al (2015) Phytosociological and ethanobotanical study of trees in a tropical dry deciduous forest in Panchmahal district of Gujarat, Western India. The Indian Forester 141(4):422–427 Qian C etal (2021) Estimation of forest aboveground biomass in Karst areas using multi-source remote sensing data and the K-DBN algorithm. Remote Sens 13(24):5030 Rae LF etal (2014) Multiscale impacts of forest degradation through browsing by hyperabundant moose (Alces alces) on songbird assemblages. Divers Distrib 20(4):382–395 Rai A etal (2021) Butea monosperma: a leguminous species for sustainable forestry programmes. Environ Dev Sustain 23(6):8492–8505 Rathod MC et al (2024) Floristic study of Garudeshwar and Nandod Taluka of Narmada District, Gujarat, India. J Plant Sci Res 40(3):426–451 Reddy CS etal (2015) Nationwide classification of forest types of India using remote sensing and GIS. Environ Monit Assess 187:1–30 Rivera-Caicedo JP et al (2017) Hyperspectral dimensionality reduction for biophysical variable statistical retrieval. ISPRS J Photogramm Remote Sens 132:88–101 Robiglio V, Mala W (2005) Integrating local and expert knowledge using participatory mapping and GIS to implement integrated forest management options in Akok, Cameroon. For Chron 81(3):392–397 Robinson EJ et al (2008) Spatial and temporal modeling of community non-timber forest extraction. J Environ Econ Manag 56(3):234–245 Rodriguez-Galiano VF, Chica-Rivas M (2014) Evaluation of different machine learning methods for land cover mapping of a Mediterranean area using multi-seasonal Landsat images and Digital Terrain Models. Photogramm Eng Remote Sens 7(6):492–509 Rodriguez-Galiano V et al (2012a) Incorporating the downscaled Landsat TM thermal band in land-cover classification using random forest. Photogramm Eng Remote Sens 78(2):129–137 Rodriguez-Galiano VF et al (2012b) An assessment of the effectiveness of a random forest classifier for landcover classification. ISPRS J Photogramm Remote Sens 67:93–104 Sardeshpande M, Shackleton C (2019) Wild edible fruits: a systematic review of an under-researched multifunctional NTFP (non-timber forest product). Forests 10(6):467 Senaviratna N, Cooray TA (2019) Diagnosing multicollinearity of logistic regression model. Asian J Probab Stat 5(2):1–9 Shah AN (2016) A study of carbon credit market in India (Gujarat) (Doctoral dissertation, Gujarat Technological University). Gujarat Technological University. https:// www. gtu. ac. in/ uploa ds/ Final% 20The sis% 20-% 20119 99739 2004% 20-% 20Ava ni% 20Shah. pdf Sharma CM etal (2010) Tree diversity and carbon stocks of some major forest types of Garhwal Himalaya, India. For Ecol Manag 260(12):2170–2179 Agroforest Syst (2025) 99:92 Page 35 of 35 92 Vol.: (0123456789) Skidmore AK (1989) An expert system classifies eucalypt forest types using thematic mapper data and a digital terrain model. Photogramm Eng Remote Sens 55(10):1449– 1464. https:// www. asprs. org/ wpconte nt/ uploa ds/ pers/ 1989j ournal/ oct/ 1989_ oct_ 14491464. pdf Skidmore AK etal (1996) An operational GIS expert system for mapping forest soils. Photogramm Eng Remote Sens 62(5):501–511 Store R, Jokimäki J (2003) A GIS-based multi-scale approach to habitat suitability modeling. Ecol Model 169(1):1–15 Store, R., etal. (2001). "Integrating spatial multi-criteria evaluation and expert knowledge for GIS-based habitat suitability modelling." 55(2): 79–93. Stromann O etal (2019) Dimensionality reduction and feature selection for object-based land cover classification based on Sentinel-1 and Sentinel-2 time series using Google Earth Engine. Remote Sens 12(1):76 Tabor JA, Hutchinson CF (1994) Using indigenous knowledge, remote sensing, and GIS for sustainable development. Indigenous Knowl Dev Monit 2(1):2–6 Tewari D (1998) Income and employment generation opportunities and potential of non-timber forest products (NTFPs) a case study of Gujarat, India. J Sustain for 8(2):55–76 Tewari D (2000) Managing non-timber forest products (NTFPs) as an economic resource. J Interdiscip Econ 11(3–4):269–287 Verrelst J etal (2015) Experimental Sentinel-2 LAI estimation using parametric, non-parametric and physical retrieval methods—a comparison. ISPRS J Photogramm Remote Sens 108:260–272 Verrelst J et al (2016) Spectral band selection for vegetation properties retrieval using Gaussian processes regression. Int J Appl Earth Obs Geoinf 52:554–567 Wang H et al (2018) Optimal subsampling for large sample logistic regression. J Am Stat Assoc 113(522):829–844 Wu C etal (2016) Comparison of machine-learning methods for above-ground biomass estimation based on Landsat imagery. J Appl Remote Sens 10(3):035010–035010 Yadav R, Jangid M, Rajpurohit S, Kamboj RD (2019) Estimated quantification and valuation of fuelwood collected from forest areas of Narmada Forest Division of Gujarat State, India. Plant Arch 19(2):2459–2464 Yadav RS, Pandya I, Rajpurohit S, Kamboj RD (2019) Estimated collection and valuation of non-wood forest products in the Dangs North Forest Division of Gujarat State, India. Plant Arch 19(1): 163–166 Yan M et al (2023) Biomass estimation of subtropical arboreal forest at single tree scale based on feature fusion of airborne liDAR data and aerial images. Sustainability 15(2):1676 Yang X etal (2006) Mapping non-wood forest product (matsutake mushrooms) using logistic regression and a GIS expert system. Ecol Model 198(1–2):208–218 Yang Y etal (2021) Testing accuracy of land cover classification algorithms in the Qilian mountains based on gee cloud platform. Remote Sens 13(24):5064 Yu X et al (2017) Deep learning in remote sensing scene classification: a data augmentation enhanced convolutional neural network framework. Gisci Remote Sens 54(5):741–758 Zhang Y et al (2016) Aboveground biomass of understorey vegetation has a negligible or negative association with overstorey tree species diversity in natural forests. Glob Ecol Biogeogr 25(2):141–150 Zhang T etal (2021) Sentinel-2 satellite imagery for urban land cover classification by optimized random forest classifier. Appl Sci 11(2):543 Zhang C et al (2023) Estimating the forest carbon storage of Chongming Eco-Island, China, using multisource remotely sensed data. Remote Sens 15(6):1575 Zhang S etal (2024) Feature selection and regression models for multisource data-based soil salinity prediction: a case study of Minqin Oasis in Arid China. Land 13(6):877 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.