scieee Open visual document viewer

Spectral–Spatial Transformer-based Semantic Segmentation for Large-scale Mapping of Individual Date Palm Trees using Very High-resolution Satellite Data

Al-Ruzouq, Rami; Gibril, Mohamed Barakat A.; Shanableh, Abdallah; Bolcek, Jan; Lamghari, Fouad; Hammour, Nezar Atalla; Al-Keblawy, Ali; Jena, Ratiranjan

Abstract

Date palm plantations in the United Arab Emirates (UAE) are under threat from soil salinity, drought, and date palm weevils. Accordingly, monitoring and conserving date palms are crucial to preserving a vital component of the country’s agricultural heritage, economy, food security, and ecological balance. Previous studies have effectively identified date palm trees using RGB-based aerial and UAV imagery utilizing diverse deep learning methods. However, the utilization of very high-resolution satellite data for delineating individual date palm crowns remains unexplored due to the limited spatial resolution capabilities of existing satellite systems. This study primarily aimed to achieve precise and comprehensive mapping of date palm trees using WorldView-3 (WV-3) satellite data by leveraging the high representational power of the state-of-the-art vision transformers (ViT) in capturing global information from the input data. First, an in-depth analysis assessment of the various transformer-based semantic segmentation architectures, including UperNet with vision transformer and Swin transformer, SegFormer, Mask2Former, and UniFormer, was conducted. Second, the integration of spectral data on the performance of ViTs was evaluated. Moreover, the models’ generalizability and complexity effect on the segmentation effectiveness were assessed. Accordingly, a postprocessing strategy was developed to aid in delineating and counting date palm trees from semantic segmentation outputs. Results demonstrated that integration of WV-3 spectral data into the analysis resulted in a marked improvement in segmentation quality. The UniFormer, UperNet-Swin, and Mask2Former models demonstrated considerable improvements in multispectral data analysis, with increases in mean intersection over union (mIoU) of 2.17% (77.88% mIoU, 86.01% mean F-score [mF-score]), 2% (78.10% mIoU, 86.18% mF-score), and 1.15% (77.36% mIoU, 85.59% mF-score), respectively, compared with their RGB-based results. Evaluations of model transferability also indicated that Mask2Former, UniFormer, and UperNet-Swin transformers efficiently adapted to multispectral data in the Dibba region. These models achieved mIoU scores of 84.36%, 84.25%, and 83.17% and mF-scores of 90.95%, 90.87%, and 90.13%, highlighting their effectiveness and potential for broader regional application. This research highlights the efficacy and feasibility of using ViTs with WV-3 multispectral data for accurate and comprehensive surveying of date palm plantations, enabling the development of palm tree inventories and continuously updating geospatial databases.

Full text

Ecological Indica o s 163 (2024) 112110 A ailable online 8 May 2024 1470-160X/© 2024 The Au ho s. Published by Else ie L d. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by- nc-nd/4.0/). O iginal A icles Spec al–Spa ial ans o me -based seman ic segmen a ion o la ge-scale mapping o indi idual da e palm ees using e y high- esolu ion sa elli e da a Rami Al-Ruzouq a , Mohamed Ba aka A. Gib il a , * , Abdallah Shanableh a , Jan Bolcek a , b , Fouad Lamgha i c , Neza A alla Hammou d , e , Ali El-Keblawy , Ra i anjan Jena a a GIS and Remo e Sensing Cen e , Resea ch Ins i u e o Sciences and Enginee ing, Uni e si y o Sha jah, Sha jah 27272, Uni ed A ab Emi a es b Depa men o Radio Elec onics, Facul y o Elec ical Enginee ing and Communica ion, B no Uni e si y o Technology, B no-K alo o pole 61600, Czech Republic c Fujai ah Resea ch Cen e, Al-Hilal Towe , 3003, P.O. Box 666 Fujai ah, Uni ed A ab Emi a es d Depa men o Applied Physics and As onomy, Facul y o Science, Uni e si y o Sha jah, Sha jah 27272, Uni ed A ab Emi a es e Depa men o Ea h and En i onmen al Sciences, P ince El-Hassan Bin Talal Facul y o Na u al Resou ces & En i onmen , The Hashemi e Uni e si y, Za qa 13133, Jo dan Depa men o Applied Biology, College o Sciences, Uni e si y o Sha jah, Sha jah P.O. Box 2727, Uni ed A ab Emi a es ARTICLE INFO Keywo ds: T ee c own delinea ion Seman ic segmen a ion Vision ans o me s Deep lea ning ABSTRACT Da e palm plan a ions in he Uni ed A ab Emi a es (UAE) a e unde h ea om soil salini y, d ough , and da e palm wee ils. Acco dingly, moni o ing and conse ing da e palms a e c ucial o p ese ing a i al componen o he coun y’s ag icul u al he i age, economy, ood secu i y, and ecological balance. P e ious s udies ha e e ec i ely iden i ied da e palm ees using RGB-based ae ial and UAV image y u ilizing di e se deep lea ning me hods. Howe e , he u iliza ion o e y high- esolu ion sa elli e da a o delinea ing indi idual da e palm c owns emains unexplo ed due o he limi ed spa ial esolu ion capabili ies o exis ing sa elli e sys ems. This s udy p ima ily aimed o achie e p ecise and comp ehensi e mapping o da e palm ees using Wo ldView-3 (WV-3) sa elli e da a by le e aging he high ep esen a ional powe o he s a e-o - he-a ision ans o me s (ViT) in cap u ing global in o ma ion om he inpu da a. Fi s , an in-dep h analysis assessmen o he a ious ans o me -based seman ic segmen a ion a chi ec u es, including Upe Ne wi h ision ans o me and Swin ans o me , SegFo me , Mask2Fo me , and UniFo me , was conduc ed. Second, he in eg a ion o spec al da a on he pe o mance o ViTs was e alua ed. Mo eo e , he models’ gene alizabili y and complexi y e ec on he segmen a ion e ec i eness we e assessed. Acco dingly, a pos p ocessing s a egy was de eloped o aid in delinea ing and coun ing da e palm ees om seman ic segmen a ion ou pu s. Resul s demons a ed ha in e- g a ion o WV-3 spec al da a in o he analysis esul ed in a ma ked imp o emen in segmen a ion quali y. The UniFo me , Upe Ne -Swin, and Mask2Fo me models demons a ed conside able imp o emen s in mul ispec al da a analysis, wi h inc eases in mean in e sec ion o e union (mIoU) o 2.17% (77.88% mIoU, 86.01% mean F- sco e [mF-sco e]), 2% (78.10% mIoU, 86.18% mF-sco e), and 1.15% (77.36% mIoU, 85.59% mF-sco e), espec i ely, compa ed wi h hei RGB-based esul s. E alua ions o model ans e abili y also indica ed ha Mask2Fo me , UniFo me , and Upe Ne -Swin ans o me s e icien ly adap ed o mul ispec al da a in he Dibba egion. These models achie ed mIoU sco es o 84.36%, 84.25%, and 83.17% and mF-sco es o 90.95%, 90.87%, and 90.13%, highligh ing hei e ec i eness and po en ial o b oade egional applica ion. This esea ch highligh s he e icacy and easibili y o using ViTs wi h WV-3 mul ispec al da a o accu a e and comp ehensi e su eying o da e palm plan a ions, enabling he de elopmen o palm ee in en o ies and con inuously upda ing geospa ial da abases. * Co esponding au ho . E-mail add ess: [email p o ec ed] (M.B.A. Gib il). Con en s lis s a ailable a ScienceDi ec Ecological Indica o s jou nal homepage: www.else ie .com/loca e/ecolind h ps://doi.o g/10.1016/j.ecolind.2024.112110 Recei ed 8 Feb ua y 2024; Recei ed in e ised o m 15 Ap il 2024; Accep ed 2 May 2024 Ecological Indica o s 163 (2024) 112110 2 1. In oduc ion 1.1. Backg ound The 2021 s a is ics o he Food and Ag icul u e O ganiza ion (FAO) indica ed ha app oxima ely 9.66 million ons o da e palms we e p oduced, co e ing an a ea o 1.3 million hec a es (FAO, 2023). The Uni ed A ab Emi a es (UAE) is among he op en p oducing coun ies, wi h almos 40 million ees dis ibu ed ac oss he UAE (El-Juhany, 2010). Da e palm ees a e an impo an pa o UAE’s cul u al he i age and a i al ag icul u al esou ce. Thus, accu a e mapping o da e palm plan a ions ac oss he coun y is impe a i e o e ec i ely and sus ain- ably moni o and manage his pi o al ag icul u al asse . T adi ional me hods o da e palm ee in en o y de elopmen , such as on-si e su eys and isual examina ion o ae ial pho og aphs, a e cos ly, labo ious, and ime-consuming. Remo e sensing echniques p o ide a mo e cos -e icien app oach, o e ing de ailed da a wi h e sa ile empo al and spa ial esolu ions (Ha ling e al., 2019). Remo e sensing using unmanned ae ial ehicles (UAVs), inc easingly adop ed in nume ous s udies ocusing on da e palms o i s supe io spa ial and empo al esolu ions and p ecise posi ional accu acy, aids in he p ecise de ec ion and mapping o da e palm ees (Amma e al., 2021; Gib il e al., 2022; Jin asu isak e al., 2022; Gib il e al., 2023). Unlike ea h- obse ing sa elli es, UAVs ha e limi ed a ea co e age and a e con- s ained by wea he condi ions and lying es ic ions (Ozda ici-ok and Ok, 2023). Sa elli e emo e sensing p o ides a comp ehensi e iew, enabling epea ed obse a ions and ex ensi e co e age ac oss as e- gions. Combining sa elli e and UAV da a could esul in a mo e comp ehensi e and in-dep h assessmen o da e palm ees. A di e se ange o machine lea ning algo i hms has been es ablished and u ilized o localize and cha ee c owns using da a de i ed om emo e sensing in au oma ic and semi-au oma ic manne s (Al-Ruzouq e al., 2018; Ghasemi e al., 2022; Ji e al., 2022; Qin e al., 2022; Zhang e al., 2023; H. Zhao e al., 2023). In ecen yea s, compu e ision echniques based on deep lea ning, p ima ily con olu ional neu al ne wo ks (CNNs), ha e become inc easingly p e alen o indi idual ee c own (ITC) delinea ing om e y high- esolu ion (VHR) emo ely sensed images. This p e alence can be a ibu ed o he ema kable capabili y o hese me hods o au oma ically ex ac in ica e high-le el ea u es and complex pa e ns om he inpu image da ase s, subs an- ially enhancing he e icacy and obus ness o he models (Gib il e al., 2022; Zheng e al., 2023). CNNs ha e demons a ed ou s anding pe - o mance in a ious ITC s udies (H. Zhao e al., 2023) due o hei dis inc i e a chi ec u e, which includes localized ecep i e ields, weigh sha ing, and he p ocess o subsampling (Ka enbo n e al., 2021). CNN-based models a e ypically used o unde ake di e en asks o de ec and delinea e ee species. Some widely used asks include objec de ec ion (Zheng e al., 2021; Zheng and Wu, 2022; Cai e al., 2023; Velasquez-camacho and E xega ai, 2023), ins ance (B aga e al., 2020; Yang e al., 2022b; Ball e al., 2023; Hao e al., 2023), and se- man ic segmen a ion (F eudenbe g e al., 2022; Gib il e al., 2023). Nume ous seman ic segmen a ion models based on CNNs, a se o encode –decode deep lea ning a chi ec u es ha a e used o pe o m pixel-wise classi ica ion, has been de eloped and u ilized o ou line ee canopies om di e en ypes o emo ely sensed da a (Bha naga e al., 2020; Cao and Zhang, 2020; Ken sch e al., 2020a; Ken sch e al., 2020b; To es e al., 2020; Wagne e al., 2020; Wu and Mis a, 2020; Xiao e al., 2020; Akca and Pola , 2022; Sun e al., 2023). F eudenbe g e al. (2022) u ilized a wo-s ep app oach based on U-Ne a chi ec u e o ITC om Wo ldView-3 (WV-3) sa elli e da a (acqui ed in Bengalu u, India) and ae ial image y (acqui ed in a densely o es ed a ea nea Ga ow, Ge - many). Thei app oach in ol es ex ac ing ee c owns wi h an in e sec ion-o e -union (IoU) o 71.2 % on he sa elli e da a and an IoU o 81.9 ±2.2% on he ae ial image y. Seman ic segmen a ion a chi- ec u es, wi h di e en backbone ne wo ks, commonly u ilized o ee c own delinea ion in ecen yea s include he U-Ne (F eudenbe g e al., 2019; Ken sch e al., 2020a; Wagne e al., 2019,2020; Liu,Wang,and Wang, 2019; Ken sch e al., 2020b; Schie e e al., 2020) and Deep- LabV3+(Ayhan and Kwan, 2020; Fe ei a e al., 2020; Cheng,Qi,and Cheng, 2021). Gi en ha con olu ions p ocess images by examining local a eas, hei design inhe en ly es ic s hem om g asping he global con ex o an en i e image in a single ope a ion, po en ially diminishing accu acy, especially in scena ios whe e he inpu images exhibi in ica e e- la ionships be ween pixels (Fayad e al., 2024). Recen ly, he ield o emo e sensing has expe ienced an inc easing adop ion o di e se ision ans o me s (ViT) o a a ie y o asks (Abozeid e al., 2022; Fan e al., 2022; Mekhal i e al., 2022; Sun e al., 2022; L. Yang e al., 2022; Zhou e al., 2022; Li e al., 2023; Yi e al., 2023; Zhao e al., 2023b; Jamali e al., 2024). In con as o CNNs, ope a ions wi hin ViTs a e inhe en ly pa allel and sequence-independen , allowing ViTs o e icien ly ga he comp ehensi e con ex ual in o ma ion and a ain a supe io le el o ep esen a ional capabili y (Xia e al., 2022; Zhou e al., 2022). Addi- ionally, encode s cons uc ed wi h ViTs main ain a uni o m dimen- sional ep esen a ion h oughou he p ocessing phases, hus ci cum en ing he necessi y o explici down-sampling ope a ions (Fayad e al., 2024). Recen ly, an inc easing numbe o s udies ha e u ilized a ious ViTs in he ields o plan disease de ec ion (Chang e al., 2024; Li e al., 2024; Rezaei e al., 2024), ui iden i ica ion (Bai e al., 2024; Liu e al., 2024), canopy heigh mapping (Fayad e al., 2024; Tolan e al., 2024), and c own delinea ion s udies. 1.2. Rela ed wo k T ans o me -based models a e inc easingly u ilized in ITC esea ch o a a ie y o isual asks, including ins ance segmen a ion (Gib il e al., 2022; De sch e al., 2023; Fi oze e al., 2023), objec de ec ion (De sch e al., 2022; Zhang e al., 2022), and ee coun ing (Chen and Shang, 2022; Ami kolaee e al., 2023). De sch e al. (2023) u ilized he ans o me -based ins ance segmen a ion app oach de ec ion ans- o me (DETR) o delinea e ITCs in mixed o es a eas u ilizing mul i- spec al UAV images and UAV-de i ed LiDAR da a. The au ho s conduc ed expe imen s wi h wo-channel selec ions – RGB and Colo - In a ed (CIR) – using DETR and achie ed he highes mean F-sco es (mF-sco es) on he CIR da ase , a aining 90 % in coni e ous o es s, 81 % in deciduous, and 84 % in mixed o es s. Zhang e al. (2022) used a Fas e R-CNN wi h a Swin ans o me and ResNe -50 backbone o iden i ying indi idual ees om VHR ae ial images. The s udy exam- ined he pe o mance o he models in di e se en i onmen s and obse ed hei highes pe o mance in s ee a eas. Chen and Shang (2022) in oduced a densi y ans o me designed o he au oma ed coun ing o ees in ae ial pho og aphs and e alua ed i s e icacy agains se e al leading-edge algo i hms. Thei p oposed ans o me -based me hod showed p omising esul s, su passing he pe o mance o a ious exis ing s a e-o - he-a echniques. Ami kolaee e al. (2023) in oduced a semi-supe ised me hodology based on a ans o me a - chi ec u e (encode –decode design) o ee coun ing. Thei ne wo k’s encode , designed o ex ac ea u es a mul iple scales, u ilized a py - amid ision ans o me (Wang e al., 2021). Despi e u ilizing he same quan i y o labeled images, his no el app oach ou pe o med exis ing semi-supe ised and ully-supe ised echniques in pe o mance. Fu e al. (2023) examined he pe o mance o wo algo i hms, namely, adap i e s acking ensemble lea ning (ASEL) and mul i-pa h ision ans o me o dense p edic ion (MPViT), o mang o e species mapping om mul ispec al UAV-based images. The s udy showed ha MPViT ou pe o med ASEL and achie ed 95.5 %–97.3 % o he o e all classi ica ion accu acy. Zhou e al. (2022) p oposed a semi-au oma ic wo k low-based T ansUNe o de ec and cha ac e ize b occoli canopy and heads om he RGB image y and LiDAR da a. Thei esul s showed ha he p oposed ans o me -based amewo k ou pe o med h ee CNN-based and wo shallow lea ning-based echniques. Zhang and Liu (2024) in oduced a dual-b anch ne wo k designed o segmen s ee R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 3 ees using high- esolu ion sa elli e image y. They in eg a ed dila ed con olu ional laye s in pa allel wi h ans o me blocks o imp o e he ne wo k’s capaci y o ex ac ing ea u es ac oss mul iple scales. He e al. (2024) le e aged he CNN-based ne wo k (imp o ed E icicen Ne - V2) and ans o me -based ne wo k (CSwin ans o me ) o cap u e local and global seman ic in o ma ion o ci us ee canopy segmen a- ion om 3D da a ob ained om UAV. Thei app oach showed supe io pe o mance han using some CNN only and some ans o me -based only. A b oad spec um o s udies has ocused on de ec ing palm ees om mul ipla o m ae ial o hopho os using CNNs (Culman e al., 2020; Gib il e al., 2021; Amma e al., 2021; Gib il e al., 2022; Jin asu isak e al., 2022). Gib il e al., (2022,2023, & 2024) demons a ed ha ViTs ou pe o med a ious CNN-based a chi ec u es o he seman ic and ins ance segmen a ion o indi idual da e palms om mul ipla o m ae ial images. Gib il e al. (2022) comp ehensi ely e alua ed a ious ins ance segmen a ion a chi ec u es o iden i y and ou line indi idual palm ees h ough mul iscale d one-based images. The au ho s e alu- a ed a ange o ne wo ks, including Mask R-CNN wi h di e en CNN- based and Swin ans o me backbones, SOLO, SOLO 2, YOLACT, and Mask Sco ing R-CNN. The ans o me -based model, Mask R-CNN wi h Swin ans o me backbones, exhibi ed supe io pe o mance, su pass- ing CNN-based models. Gib il e al. (2023) examined he ealiabili y o a ious ision ans o me o ex ac da e palm ees om mul iscale UAV and ae ial images. Thei esul s showed ha ans o me -based models achi ed an mF-sco e anging om 91.62 % o 92.44 %. Among he e alua ed models, he SegFo me model ou pe o med all o he e alua ed CNN-based models in he mul iscale and he independen es ing da ase s, ollowed by he Upe Ne -Swin ans o me . Gib il e al. (2024) p esen ed a ans o me -based me hod o delinea e, coun i y, and assess he heal h o indi idual da e palm ees om la ge-scale UAV da a. The p oposed me hod is based on he syne gy be ween an imp o ed mul iscale ViT (MViT 2), a ea u e py amid ne wo k, Mask R- CNN, and an imp o ed slicing-aided hype in e ence module. The model, which was ini ially de eloped o delinea e da e palm ees, was hen ine- uned o assess he heal h o he da e palm ees. This a chi- ec u e ou pe o med a ious ins ance segmen a ion models, achie ing an F-sco e o 94.2 % o segmen a ion and 88.4 % o heal h assessmen . P e ious esea ch on mapping and moni o ing palm ees based on deep lea ning has p edominan ly u ilized d one-based and ae ial im- age y wi h spa ial esolu ions anging om 2 cm o 20 cm. Dis inguishing he c owns o palm ees in he e ogeneous landscapes in ensi ies as he da a’s g ound space dis ance (GSD) becomes coa se . This inc eased di icul y can ad e sely a ec he pe o mance o se- man ic segmen a ion models, po en ially hinde ing hei abili y o accu a ely ecognize indi idual da e palm ees. The easibili y o u i- lizing VHR sa elli e image y o ex ac da e palm ees om sa elli e images has ye o be ho oughly in es iga ed. This wo k mainly aims o in oduce an e ec i e deep lea ning me hodology o ex ensi e, de ailed and accu a e egional su eys o da e palm dis ibu ion, u ilizing he VHR WV-3 sa elli e image y ac oss mul iple ci ies. Accu a e mapping and quan i ica ion o indi idual da e palms ac oss egions enables he au oma ed de elopmen o da e palm ee in en o ies and can p o ide insigh s in o hei heal h and gene ic di e si y. These da a a e c ucial o de eloping conse a ion measu es and ensu ing sus ainable cul i a ion in hei na u al habi a s. The speci ic objec i es o his s udy a e o (1) e alua e he pe o mance o di e en deep ViTs in egional su eying o da e palms using WV-3 sa elli e da a, (2) in es iga e he in luence o spec al da a in eg a ion on he e icacy o he e alua ed ViTs, (3) assess model gene alizabili y ac oss a ied geog aphic egions, and (4) de elop a no el pos p ocessing echnique o indi idual da e palm ee delinea ion and quan i ica ion. 2. Ma e ials and me hods 2.1. O e iew The esea ch me hodology encompasses i e main s ages (Fig. 1). Ini ially, WV-3 sa elli e da a iles we e p ep ocessed, including a mo- sphe ic co ec ion, image pansha pening, and da a no maliza ion. The second s age in ol ed manually anno a ing da e palm ees; di iding he da a in o aining, alida ion, and es ing egions; and gene a ing image- mask pai s. The hi d s age consis ed o comp ehensi e expe imen s o assess he e icacy o se e al ad anced ViTs in delinea ing da e palm ees in WV-3 sa elli e image y, u ilizing spa ial and spec al in o ma- ion. The e alua ed ViTs in his s udy a e Upe Ne (Xiao e al., 2018) wi h ision ans o me (Kolesniko e al., 2021) and Swin ans o me (Liu e al., 2021), SegFo me (Xie e al., 2021), Mask2Fo me (Cheng e al., 2022), and UniFo me (Li e al., 2022). T aining and e alua ion o hese models we e conduc ed using di e se da a inpu s, including RGB, a combina ion o RGB wi h NIR1 and NIR2, eigh mul ispec al chan- nels, and an amalgama ion o hese channels wi h NDVI. In he ou h Fig. 1. Me hodological amewo k o his wo k. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 4 s age, he adap abili y o hese models o a ious WV-3 da ase s om di e en loca ions in Dibba, collec ed on di e en da es, was examined o de e mine hei gene aliza ion capabili ies. The inal s ep p oposes a da a pos p ocessing echnique o coun ing and delinea ing indi idual da e palm ees. Fig. 2. Loca ion o he expe imen al si es: (a) UAE emi a es, (b) de ailed zoom o he s udy si es, and (c–e) enhanced zoom le els highligh ing he da e palm ees as seen in mul iple WV-3 images. Fig. 3. Va ia ions in da e palm ees in e ms o c own size, shape, heigh , and he su ounding en i onmen s. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 5 2.2. S udy a eas The expe imen al si es a e si ua ed ac oss a ious ci ies in he UAE, including Fujai ah, Sha jah, and Ajman (Fig. 2a and 2b). Speci ically, he a eas a e loca ed in eas Ajman, Al Zubai , Kalba, Dibba, and Masa i. Collec i ely, hese a eas encompass an a ea o 275 km 2 , dis ibu ed as ollows: Ajman and Sha jah co e an a ea o 38.6 km 2 , Kalba encom- passes an a ea o 93 km 2 , Dibba is 77.1 km 2 , and Masa i spans an a ea o 66.3 km 2 . The s udy si es hos a di e se a ay o ee and ege a ion species, including he da e palm (Phoenix dac yli e a L.), gha (P osopis cine a ia), Mesqui e (P osopis juli lo a), Sid o Ch is ’s ho n jujube (Ziziphus spina-ch is i), umb ella ho n acacia (Acacia o ilis), and neem (Azadi ach a indica), among a ious sh ubs and o he lo a. Fig. 2c–e display close-up iews o he egions ma ked by cyan ci cles, high- ligh ing di e en da e palm ees in WV-3 sa elli e da a. In he p elimina y analysis s age, a comp ehensi e ield campaign was conduc ed o ho oughly unde s and he loca ions, dis ibu ions, and isual cha ac e is ics o da e palm ees. This campaign was also c ucial o gaining insigh in o he s udy a ea’s b oade geog aphic and en i onmen al con ex s, such as ege a ion and ee species’ e ain, a ie y, and su oundings. The A cGIS Field Map mobile applica ion was u ilized o eco d and manage obse a ions h oughou he ield isi s. The coo dina es o he da e palm plan a ions wi hin he s udy a eas we e documen ed. The da e palm ees displayed a conside able a ia- ion in c own shape and size, age, heigh , densi y, and he su ounding landscape cha ac e is ics (Fig. 3). 2.3. Da a acquisi ion and p ep ocessing In his esea ch, a o al o 20 mul ida e VHR WV-3 sa elli e image iles we e u ilized. The WV-3 da ase ea u es eigh spec al bands wi h a GSD o 1.24 m alongside a panch oma ic band wi h a 0.31 m GSD. The WV-3 image y o he di e en si es was acqui ed in Ajman and Sha jah ci y in Ma ch 2021, Kalba in Augus 2022, Dibba in No embe 2019, and Masa i in Sep embe 2019. Table 1 lis s he gene al cha ac e is ics o he WV-3 da a. The p ocessing o he WV-3 da a comp ised se e al s ages. Fi s , a mosphe ic co ec ion was conduc ed using he FLAASH me hod. Second, he spa ial esolu ion o he mul ispec al da a was imp o ed h ough he applica ion o he G am–Schmid pansha pening echnique. The ea e , he bi dep h o he da a unde wen con e sion om 11 bi o 8 bi , s anda dizing he image alues wi hin a 0–255 ange. Conduc ing p ecise, la ge-scale mapping o da e palm ees u ilizing WV-3 sa elli e da a wi h classical machine lea ning echniques, which p ima ily depend on he spec al in o ma ion o he da a, encoun e no able sho comings. Some expec ed challenges include misclassi ica- ion and a dec ease in gene alizabili y, which may esul om sub- s an ial a ia ions in acquisi ion da es and illumina ion condi ions o he acqui ed sa elli e da a. Fig. 4 illus a es he a e age pixel alues o Table 1 Cha ac e is ics o he WV-3 da a. Channel Wa eleng h ange (nm) GSD (m) Panch oma ic 450–800 0.31 Coas al 400–450 1.24 Blue 450–510 1.24 G een 510–580 1.24 Yellow 585–625 1.24 Red 630–690 1.24 Red Edge 705–745 1.24 NIR-1 770–895 1.24 NIR-2 860–1,040 1.24 Fig. 4. A e age pixel alues o comp ehensi e da e palm ees de i ed om a ious channels o he WV-3 da ase . R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 6 di e se da e palm ees, de i ed om a ious channels o he p e- p ocessed WV-3 da ase , based on an analysis o 76,950 indi idual da e palm ees. Fu he mo e, he da e palms wi hin he s udy a ea exhibi ed a wide ange o c own sizes, s uc u es, ages, and heal h s a es, adding ano he laye o complexi y o he mapping p ocess. The e o e, his s udy a emp s o le e age spec al, spa ial, and global con ex ual in- o ma ion h ough he u iliza ion o s a e-o - he-a ViTs o ensu e ac- cu a e la ge-scale mapping o da e palm ees om di e en WV-3 da ase s. 2.4. Da a p epa a ion o seman ic segmen a ion The WV-3 images ea u ing da e palm ees om Ajman and Kalba we e manually ou lined using g ound- u h da a and isual in e p e a- ion o VHR ae ial images o ain and assess di e en ans o me -based deep lea ning models. The p epa a ion o g ound u h da a was Fig. 5. Examples o (a), (b), and (c) image iles, and he co esponding anno a ions (A, B, C) we e selec ed om he aining da ase s. Fig. 6. S eps in ol ed in pos p ocessing o delinea e and enume a e sepa a e da e palm ees. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 7 comp ehensi e and aimed o include la ge a eas wi h di e se a ia ions in da e palm ees, such as palm ype, canopy size, ele a ion, age, and adjacen en i onmen s. Delinea ing smalle da e palm ees was chal- lenging due o he sa elli e da a’s limi ed spa ial esolu ion. The manual e i ica ion o delinea ed palms, se ing as g ound- u h da a, was ho oughly conduc ed by analys s using VHR ae ial and UAV image y om a ious loca ions in he s udy a ea and e ised by mul iple anno- a o s o quali y assu ance. The ec o da a encompassed a o al o 76,950 da e palm ees sca e ed ac oss ag icul u al and u ban land- scapes. The WV-3 da a and i s anno a ions we e c opped in o consis en 512 ×512 image iles. A o al o 16,260 image iles we e dedica ed o de eloping deep lea ning models. Meanwhile, app oxima ely 1800 images we e kep o e alua ing and es ing he models. Fig. 5 displays a se o image iles and hei co esponding labels (masks) chosen om he aining da ase . 2.5. Seman ic segmen a ion a chi ec u es Va ious ans o me -based seman ic segmen a ion a chi ec u es we e subjec ed o aining and e alua ion, u ilizing a di e se a ay o spec al image composi es (i.e., s anda d RGB, an in eg a ion o RGB wi h NIR1 and NIR2 bands, o eigh dis inc mul ispec al channels, and a usion o hese channels wi h he NDVI). The e alua ed seman ic segmen a ion models include Upe Ne (Xiao e al., 2018) based on ViT backbone, Upe Ne based on Swin ans o me (Liu e al., 2021), Mask2Fo me (Cheng e al., 2022), SegFo me (Xie e al., 2021), and UniFo me (Li e al., 2022). The models we e implemen ed u ilizing he PyTo ch (Paszke e al., 2019) and MMsegmen a ion (MMsegmen a ion, 2020) amewo k. The ollowing subsec ions o e a b ie desc ip ion o he a o emen ioned models. 2.5.1. Upe Ne -ViT The Upe Ne , an encode –decode ne wo k o seman ic segmen a- ion de eloped by Xiao e al. (2018), is based on a ea u e py amid ne wo k (FPN) amewo k (Lin e al., 2017). This amewo k in eg a es a py amid pooling module (PPM) om he PSPNe (Zhao e al., 2017). The FPN acili a es a dual p ocess in ol ing he downsampling o ea u e maps h ough a bo om-up pa hway and subsequen upsampling ia a op-down mechanism connec ed h ough la e al linkages. This bo om- up pa hway gene a es ea u e maps ac oss mul iple scales, con ingen on he a chi ec u e o he chosen backbone a chi ec u e. The spa ial dimensions o he ex ac ed ea u es a e e ec i ely doubled by he op- down pa h. This con igu a ion acili a es he amalgama ion o ea u es ha a e low in esolu ion ye ich in seman ic con en wi h hose ha Table 2 Lis o he u ilized hype pa ame e s. Upe Ne based on ViT and Swin ans o me and SegFo me UniFo me Mask2Fo me Op imize SGD AdamW AdamW Loss C oss en opy C oss en opy C oss en opy P e ained weigh s AED20K AED20K AED20K Ini ial lea ning a e 0.01 0.0001 0.0001 Momen um 0.9 0.9 0.9 Weigh decay 0005 0.0001 0.05 Ba ch size 2 2 2 L.R. schedule PolyLR (powe = 0.9, begin =1500, and end =160,000) Linea LR (powe =0.9, begin = 1500, and end = 160,000) PolyLR (powe = 0.9, begin =1500, and end =160,000) Fig. 7. (a) T aining ime, (b) loss g aph, (c) mIoU, and (d) mF-sco e. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 8 possess high esolu ion bu compa a i ely weake seman ic s eng h. The e ec i eness o Upe Ne , u ilizing ViT and Swin ans o me back- bones, was assessed in his s udy. The s uc u e o ViT comp ises an embedding laye , a ans o me encode wi h se e al iden ical laye s, and a head classi ie . Fi s , ViT spli s an image in o uni o m squa e pa ches (i.e., 16 ×16 o 32 ×32 sizes), wi h each pa ch p ocessed as a oken. These pa ches a e hen con e ed in o 1D ec o sequences ia a lea nable linea p ojec ion o dimension E, and posi ional embeddings a e added o e ain hei loca ion in o ma ion. This sequence is subsequen ly inpu ed in o he ans o me encode . E e y ans o me encode con ains a mul i-head sel -a en ion (MSA) block and a dense eed o wa d block. A esidual connec ion suppo s each sublaye , ollowed by laye no maliza ion o he ou pu . The MSA block comp ises ou componen s: an ini ial linea laye , a sel -a en ion laye , a conca ena ion laye ha me ges he ou - pu s om a ious a en ion heads, and a inal linea laye . Sel -a en ion in ol es calcula ing a en ion sco es e lec ing he pa ch embeddings’ ela ionships. A posi ion-wise eed o wa d ne wo k is applied o each pa ch embedding by le e aging he sel -a en ion mechanism. The ans o med encode ’s ou pu is hen ed in o he classi ica ion head, which pe o ms he classi ica ion o he inpu image, u ilizing he ea u e ep esen a ions de i ed om he encoded pa ch embeddings. 2.5.2. Upe Ne -Swin ans o me The Swin ans o me , a hie a chical ans o me s uc u ed in mul iple s ages, u ilizes a “shi ed window” echnique o de i a e i s ep esen a ions. This app oach allows he ex ac ion o ea u es a a ious le els, he eby e icien ly cap u ing he ex ensi e dependencies inhe en in he da ase (Liu e al. 2021). The “shi ed windowing” echnique enhances compu a ional e iciency by con ining he sel - a en ion mechanism o dis inc , non-o e lapping local windows. This me hod p omo es he es ablishmen o in e connec ions be ween hese windows. The Swin ans o me ’s ou -s age design encompasses a ious p ocesses, including mul iscale ea u e map ex ac ion, pa ch pa i ioning, me ging o pa ches, linea embedding, and implemen ing Swin ans o me blocks. The pa ch pa i ion module ini ially di ides he o iginal image in o dis inc , non-o e lapping pa ches, wi h each pa ch being p ocessed as an indi idual oken. These okens combine aw pixel da a om mul iple channels o o m hei ea u e ep esen a ion. This aw ea u e is hen ans o med h ough a linea embedding laye in o a ea u e ec o o a speci ied dimension (C). Finally, hese pa ch okens unde go p ocessing by a se ies o ans o me blocks, aiding he de elopmen o ea u e ep esen a ions. The Swin ans o me block comp ises h ee main componen s: window MSA, shi ed window MSA, and a mul ilaye pe cep on (MLP). A e each s age, he a chi ec u e uses a pa ch me ging laye o educe he oken coun and c ea e a hi- e a chical ep esen a ion. This model’s “ iny” a ian was chosen as he ounda ional a chi ec u e o he Upe Ne amewo k, which consis s o {2, 2, 6, 2} laye s ac oss i s ou s ages, and i ope a es wi h a ea u e dimension (C) o 96. The eade s may e e o Liu e al. (2021) o addi ional de ails abou he design and a ia ions o he Swin ans o me . 2.5.3. Mask2Fo me The Mask2Fo me (Cheng e al., 2022) is a uni e sal a chi ec u e o e sa ile asks, such as seman ic segmen a ion, panop ic segmen a ion, and ins ance segmen a ion. Mask2Fo me is cons uc ed using a simple me a-a chi ec u e encompassing a backbone ne wo k, a pixel decode , and a ans o me decode . The backbone’s a chi ec u e can be designed using ei he CNN (i.e., esidual lea ning ne wo ks) o ans o me -based amewo ks. In his wo k, he adop ed backbone ne wo k was based on he Swin ans o me (Liu e al., 2021), conside ing i s e iciency in cap u ing global and local ea u es. The pixel decode o Mask2Fo me u ilizes a mul iscale de o mable a en ion ans o me (MSDe o mA n) (Zhu e al. 2020) o e ec i ely inco po a e low- and high- esolu ion ea u es, op imizing he compu a ional e iciency. The ans o me decode in Mask2 o me uses a masked a en ion mechanism, ocusing en i ely on local ea u es su ounding he p edic ed segmen s a he han p ocessing he en i e ea u e map. In his s udy, he “ iny” e sion o he Swin ans o me was selec ed o se e as he backbone a chi- ec u e o he Mask2 o me . Reade s can e e o he wo k by Cheng e al. (2022) o mo e insigh s in o he design and a ia ions o he Mask2Fo me . 2.5.4. SegFo me SegFo me (Xie e al., 2021), which is a obus and e ec i e ans o me -based seman ic segmen a ion a chi ec u e, has demon- s a ed ema kable pe o mance in a ious emo e sensing applica ions (Tang e al., 2022; Gib il e al., 2023; Gonçal es e al., 2023; Jiang e al., 2023). SegFo me is designed using encode –decode design, uniquely in eg a ing T ans o me s wi h a ligh weigh MLP-based decode . The encode ex ac s mul iscale ea u es using ou dis inc ans o me blocks. Each block en ails h ee in eg al modules: an e icien sel - a en ion mechanism, a mixed eed o wa d ne wo k (FFN) (known as Mix-FFN), and me ging blocks o o e lapping pa ches. Con a y o he con en ional ViT ha u ilizes ixed esolu ion posi ion encodings (P.E.s) o posi ional in o ma ion in eg a ion, he SegFo me adop s con olu- ional laye s in i s FFN o a da a-dependen app oach o posi ional encoding. The decode uses mul iscale in o ma ion, encompassing local and global de ails, o accu a ely p edic he inal segmen a ion ou - comes. This s udy selec ed a mix ans o me (MiT) encode based on B3 (MiT-B3). Fu he in o ma ion can be ound in he s udy by Xie e al. (2021). 2.5.5. UniFo me The uni ied ans o me (Li e al., 2022) le e ages he capabili ies o CNNs and ans o me s o mi iga e each app oach’s limi a ions, achie ing an op imal balance be ween compu a ional e iciency and accu acy. A basic ans o me o ma (Vaswani e al., 2017) is u ilized and cus omized o e icien and e ec i e lea ning o spa io empo al ep esen a ions. The UniFo me block consis s o h ee main modules: dynamic posi ion embedding (DPE), mul i-head ela ion agg ega o Table 3 Quan i a i e e alua ion o he ans o me -based seman ic segmen a ion models. Me ics RGB (%) RGB þ NIR1 þ NIR2 (%) Eigh MS bands (%) Eigh MS bands þ NDVI (%) SegFo me -MiT- B3 mAcc 85.29 85.69 86.89 85.60 mIoU 76.16 76.54 76.62 76.47 mP ecision 83.92 84.16 83.27 84.15 mRecall 85.29 85.69 86.89 85.60 mF-sco e 84.59 84.91 84.98 84.86 Upe Ne -ViT mAcc 83.41 74.29 78.06 78.13 mIoU 75.04 68.99 71.57 71.64 mP ecision 83.87 83.09 83.46 83.51 mRecall 83.41 74.29 78.06 78.13 mF-sco e 83.64 77.98 80.51 80.57 Upe Ne -Swin- iny mAcc 85.84 86.81 86.93 85.64 mIoU 76.12 76.78 78.10 76.65 mP ecision 83.38 83.57 85.46 84.38 mRecall 85.84 86.81 86.93 85.64 mF-sco e 84.56 85.12 86.18 85.0 Mask2Fo me - Swin- iny mAcc 85.11 84.23 87.54 85.41 mIoU 76.21 75.28 77.36 75.73 mP ecision 84.17 83.48 83.85 83.14 mRecall 85.11 84.23 87.54 85.41 mF-sco e 84.63 83.85 85.59 84.24 UniFo me mAcc 83.58 86.86 88.30 86.33 mIoU 75.71 77.51 77.88 77.34 mP ecision 84.86 84.62 84.00 84.83 mRecall 83.58 86.86 88.3 86.33 mF-sco e 84.21 85.70 86.01 85.56 R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 9 (MHRA), and eed o wa d ne wo k (FFN). The DPE ac i ely in- co po a es 3D posi ional da a in o each oken, e icien ly le e aging he spa io empo al sequence o he okens o ideo modeling pu poses. The MHRA combines each oken wi h i s con ex ual coun e pa s, which handles local edundancy and global dependency in ideos h ough i s adap able app oach o oken a ini y lea ning ac oss shallow and deep laye s. Al hough he local MHRA in he shallow laye s signi ican ly lessens he compu a ional bu den, he global MHRA in he deepe laye s ocuses on comp ehensi ely lea ning global oken ela ionships. The eade may e e o Li e al. (2022) o he de ailed in o ma ion abou he UniFo me a chi ec u e. 2.6. Accu acy me ics This s udy e alua ed he e icacy o a ious seman ic segmen a ion a chi ec u es h ough mul iple accu acy me ics. These me ics de e - mine he cong uence be ween he e e ence ee c owns and he p e- dic ed delinea ed ee c owns by ViTs. The combina ion o he a ious seman ic segmen a ion e alua ion me ics p o ides a ho ough unde - s anding o he esul s o da e palm ee mapping, highligh ing he e o s and limi a ions o he di e en models. The u ilized me ics include mIoU, p ecision, ecall, and mF-sco e and a e de ailed in Equa ions (1)– (6). The IoU and F-sco e a e commonly u ilized me ics in ITC s udies o assess he le el o ag eemen be ween deep lea ning model esul s and g ound- u h da a. These me ics ange om ze o o one o hund ed, Fig. 8. Cu a ed selec ion o se en mul ispec al images (a–g) andomly chosen om he es ing da a, accompanied by hei co esponding e e ence da a and he expe imen al esul s. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 16 managemen and moni o ing da e palm ees. CRediT au ho ship con ibu ion s a emen Rami Al-Ruzouq: W i ing – e iew & edi ing, W i ing – o iginal d a , Supe ision, P ojec adminis a ion, Me hodology, In es iga ion, Fo mal analysis, Concep ualiza ion. Mohamed Ba aka A. Gib il: W i ing – e iew & edi ing, W i ing – o iginal d a , Visualiza ion, Valida ion, So wa e, Me hodology, Fo mal analysis, Da a cu a ion, Concep ualiza ion. Abdallah Shanableh: W i ing – e iew & edi ing, Supe ision, Resou ces, Concep ualiza ion. Jan Bolcek: W i ing – o iginal d a , Valida ion, So wa e, Me hodology, Fo mal analysis. Fouad Lamgha i: W i ing – e iew & edi ing, Supe ision, P ojec adminis a ion, Funding acquisi ion, Concep ualiza ion. Neza A alla Hammou : W i ing – e iew & edi ing, Supe ision, Concep ualiza ion. Ali El-Keblawy: W i ing – e iew & edi ing, W i ing – o iginal d a , Supe ision, Concep ualiza ion. Ra i anjan Jena: W i ing – e iew & edi ing, Me hodology, Fo mal analysis. Decla a ion o compe ing in e es The au ho s decla e ha hey ha e no known compe ing inancial in e es s o pe sonal ela ionships ha could ha e appea ed o in luence he wo k epo ed in his pape . Da a a ailabili y The like o he code is gi en in he manusc ip Acknowledgmen s The au ho s exp ess hei app ecia ion o he Fujai ah Resea ch Cen e (FRC) o i s inancial suppo (Funded P ojec Numbe : 133049) and o he Uni e si y o Sha jah o o e ing esea ch acili ies. Re e ences “Food and Ag icul u e O ganiza ion.” FAOSTAT [In e ne ]. [accessed 2021 Ma 9]. h p://www. ao.o g/ aos a /en/#da a/QC. Abozeid, A., Alanazi, R., Elhadad, A., Taloba, A.I., Abd El-Aziz, R.M., 2022. A la ge-scale da ase and deep lea ning model o de ec ing and coun ing oli e ees in sa elli e image y. Compu In ell Neu osci. 2022 h ps://doi.o g/10.1155/2022/1549842. Akca S, Pola N. 2022. Seman ic segmen a ion and quan i ica ion o ees in an o cha d using UAV o hopho o. Ea h Sci In o ma ics 2022 154 [In e ne ]. [accessed 2022 No 28] 15(4):2265–2274. h ps://doi.o g/10.1007/S12145-022-00871-Y. Al-Ruzouq, R., Shanableh, A., Ba aka , A., Gib il, M., AL-Mansoo i, S., 2018. Image segmen a ion pa ame e selec ion and an colony op imiza ion o da e palm ee de ec ion and mapping om e y-high-spa ial- esolu ion ae ial image y. Remo e Sens [in e ne ]. 10 (9), 1413. h ps://doi.o g/10.3390/ s10091413. Ami kolaee, H.A., Shi, M., Membe , S., Mulligan, M., 2023. T eeFo me : a semi- supe ised ans o me -based amewo k o ee coun ing om a single. IEEE T ans Geosci Remo e Sens. 61, 1–15. h ps://doi.o g/10.1109/TGRS.2023.3295802. Amma , A., Koubaa, A., Benjdi a, B., 2021. Deep-lea ning-based au oma ed palm ee coun ing and geoloca ion in la ge a ms om ae ial geo agged images. Ag onomy [in e ne ]. 11 (8), 1458. h ps://doi.o g/10.3390/ag onomy11081458. Ayhan, B., Kwan, C., 2020. T ee, sh ub, and g ass classi ica ion using only RGB images. Remo e Sens. 12 (8) h ps://doi.o g/10.3390/RS12081333. Bai, Y., Yu, J., Yang, S., Ning, J., 2024. An imp o ed YOLO algo i hm o de ec ing lowe s and ui s on s awbe y seedlings. Biosys Eng [in e ne ]. 237(June 2023): 1–12 h ps://doi.o g/10.1016/j.biosys emseng.2023.11.008. Ball JGC, Hickman SHM, Jackson TD, Koay XJ, Hi s J, Jay W, A che M, Coomes DA. 2023. Accu a e delinea ion o indi idual ee c owns in opical o es s om ae ial RGB image y using Mask R-CNN. :1–14. h ps://doi.o g/10.1002/ se2.332. Bha naga , S., Gill, L., Ghosh, B., 2020. D one image segmen a ion using machine and deep lea ning o mapping aised bog ege a ion communi ies. Remo e Sens. 12 (16) h ps://doi.o g/10.3390/RS12162602. B aga, J.R.G., Pe ipa o, V., Dalagnol, R., Fe ei a, M.P., Ta abalka, Y., A ag˜ ao, L.E.O.C., de Campos Velho, H.F., Shiguemo i, E.H., Wagne , F.H., 2020. T ee c own delinea ion algo i hm based on a con olu ional neu al ne wo k. Remo e Sens. 12 (8), 1–27. h ps://doi.o g/10.3390/RS12081288. Cai C, Xu H, Chen S, Yang L, Weng Y, Huang S, Dong C, Lou X. 2023. T ee Recogni ion and C own Wid h Ex ac ion Based on No el Fas e -RCNN in a Dense Loblolly Pine En i onmen . Cao, K., Zhang, X., 2020. An imp o ed Res-UNe model o ee species classi ica ion using ai bo ne high- esolu ion images. Remo e Sens. 12 (7) h ps://doi.o g/ 10.3390/ s12071128. Chang, B., Wang, Y., Zhao, X., Li, G., Yuan, P., 2024. A gene al-pu pose edge- ea u e guidance module o enhance ision ans o me s o plan disease iden i ica ion. Expe Sys Appl [in e ne ]. 237, 121638 h ps://doi.o g/10.1016/j. eswa.2023.121638. Chen, G., Shang, Y., 2022. T ans o me o ee coun ing in ae ial images. Remo e Sens. 14 (3), 476. h ps://doi.o g/10.3390/ s14030476. Cheng B, Mis a I, Schwing AG, Ki illo A, Gi dha R. 2022. Masked-a en ion Mask T ans o me o Uni e sal Image Segmen a ion. P oc IEEE Compu Soc Con Compu Vis Pa e n Recogni . 2022-June:1280–1289. h ps://doi.o g/10.1109/ CVPR52688.2022.00135. Cheng, Z., Qi, L., Cheng, Y., 2021. Che y ee c own ex ac ion om na u al o cha d images wi h complex backg ounds. Ag icul u e [in e ne ]. 11 (5), 431. h ps://doi. o g/10.3390/ag icul u e11050431. Culman, M., Delalieux, S., Van T ich , K., 2020. Indi idual palm ee de ec ion using deep lea ning on RGB image y o suppo ee in en o y. Remo e Sens. 12 (21), 1–31. h ps://doi.o g/10.3390/ s12213476. De sch, S., K zys ek, P., Heu ich, M., Resou ces, N., Fo es , B., Pa k, N., Moni o ing, N.P., Technology, I., Sys ems, I., 2022. NOVEL SINGLE TREE DETECTION BY TRANSFORMERS USING UAV-BASED MULTISPECTRAL IMAGERY, XLIII(June): 6–11. De sch, S., Sch¨ o l, A., K zys ek, P., Heu ich, M., 2023. Towa ds comple e ee c own delinea ion by ins ance segmen a ion wi h Mask R-CNN and DETR using UAV-based mul ispec al image y and lida da a. ISPRS Open J Pho og amm Remo e Sens. 8, 100037. El-Juhany, L., 2010. Deg ada ion o da e palm ees and da e p oduc ion in A ab coun ies: causes and po en ial ehabili a ion. Aus J Basic Appl Sci [in e ne ]. 4 (8), 3998–4010. h ps://doi.o g/10.1016/j. c .2011.04.017. Fan, F., Zeng, X., Wei, S., Zhang, H., Tang, D., Shi, J., Zhang, X., 2022. E icien ins ance segmen a ion pa adigm o in e p e ing SAR and op ical images. Remo e Sens. 14 (3), 1–22. h ps://doi.o g/10.3390/ s14030531. Fayad, I., Ciais, P., Schwa z, M., Wigne on, J.P., Baghdadi, N., de T uchis, A., d’Asp emon , A., F appa , F., Saa chi, S., Sean, E., e al., 2024. Hy-TeC: a hyb id ision ans o me model o high- esolu ion and la ge-scale mapping o canopy heigh . Remo e Sens En i on. 302 (Decembe 2023) h ps://doi.o g/10.1016/j. se.2023.113945. Fe ei a MP, Almeida DRA de, Papa D de A, Mine ino JBS, Ve as HFP, Fo mighie i A, San os CAN, Fe ei a MAD, Figuei edo EO, Fe ei a EJL. 2020. Indi idual ee de ec ion and species classi ica ion o Amazonian palms using UAV images and deep lea ning. Fo Ecol Manage [In e ne ]. 475(Ap il):118397. h ps://doi.o g/10.1016/ j. o eco.2020.118397. Fi oze A, Wing en C, Yeh RA, Benes B, Aliaga D. 2023. T ee Ins ance Segmen a ion Wi h Tempo al Con ou G aph. In: P oc IEEE/CVF Con Compu Vis Pa e n Recogni . [place unknown]; p. 2193–2202. F eudenbe g, M., N¨ olke, N., Agos ini, A., U ban, K., W¨ o g¨ o e , F., Kleinn, C., 2019. La ge scale palm ee de ec ion in high esolu ion sa elli e images using U-Ne . Remo e Sens. 11 (3), 1–18. h ps://doi.o g/10.3390/ s11030312. F eudenbe g, M., Magdon, P., N¨ olke, N., 2022. Indi idual ee c own delinea ion in high- esolu ion emo e sensing images based on U-Ne . Neu al Compu Appl. 34 (24), 22197–22207. h ps://doi.o g/10.1007/s00521-022-07640-4. Fu, B., He, X., Liang, Y., Deng, T., Li, H., He, H., Jia, M., Fan, D., Wang, F., 2023. Examina ion o he pe o mance o ASEL and MPViT algo i hms o classi ying mang o e species o mul iple na u al ese es o Beibu Gul , sou h China. Ecol Indic. 154 (Augus ) h ps://doi.o g/10.1016/j.ecolind.2023.110870. Ghasemi, M., La i i, H., Pou hashemi, M., 2022. A no el me hod o de ec ing and delinea ing coppice ees in UAV images o moni o ee decline. Remo e Sens. 14 (23), 5910. Gib il MBA, Sha i HZM, Shanableh A, Al-Ruzouq R, Wayayok A, Hashim SJ bin, Sachi MS. 2022. Deep con olu ional neu al ne wo ks and Swin ans o me -based amewo ks o indi idual da e palm ee de ec ion and mapping om la ge-scale UAV images. Geoca o In [In e ne ]. 37(27):18569–18599. h ps://doi.o g/ 10.1080/10106049.2022.2142966. Gib il MBA, Sha i HZM, Shanableh A, Al-Ruzouq R, bin Hashim SJ, Wayayok A, Sachi MS. 2024. La ge-scale assessmen o da e palm plan a ions based on UAV emo e sensing and mul iscale ision ans o me . Remo e Sens Appl Soc En i on [In e ne ].:101195. h ps://doi.o g/h ps://doi.o g/10.1016/j. sase.2024.101195. Gib il, M.B.A., Sha i, H.Z.M., Shanableh, A., Al-Ruzouq, R., Wayayok, A., Hashim, S.J., 2021. Deep con olu ional neu al ne wo k o la ge-scale da e palm ee mapping om ua -based images. Remo e Sens. 13 (14), 1–24. h ps://doi.o g/10.3390/ s13142787. Gib il, M.B.A., Sha i, H.Z.M., Al-Ruzouq, R., Shanableh, A., Nahas, F., Al, M.S., 2023. La ge-scale da e palm ee segmen a ion om mul iscale UAV-based and ae ial images using deep ision ans o me s. D ones. 7 (2) h ps://doi.o g/10.3390/ d ones7020093. Gonçal es DN, Ma ca o J, Ca ilho AC, Acos a PR, Ramos APM, Gomes FDG, Osco LP, da Rosa Oli ei a M, Ma ins JAC, Damasceno GA, e al. 2023. T ans o me s o mapping bu ned a eas in B azilian Pan anal and Amazon wi h Plane Scope image y. In J Appl Ea h Obs Geoin [In e ne ]. 116(Decembe 2022):103151. h ps://doi.o g/ 10.1016/j.jag.2022.103151. Hao, Z., Lin, L., Pos , C.J., Mikhailo a, E.A., Yu, K., 2023. The co-e ec o image esolu ion and c own size on deep lea ning o indi idual ee de ec ion and delinea ion. In J Digi Ea h [In e ne ].:3753–3771. h ps://doi.o g/10.1080/ 17538947.2023.2257636. R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 17 Ha ling, S., Sagan, V., Sidike, P., Maimai ijiang, M., Ca on, J., 2019. U ban ee species classi ica ion using a wo ld iew-2/3 and liDAR da a usion app oach and deep lea ning. Senso s (swi ze land). 19 (6), 1–23. h ps://doi.o g/10.3390/s19061284. He, H., Zhou, F., Xia, Y., Chen, M., Chen, T., 2024. Pa allel usion neu al ne wo k conside ing local and global seman ic in o ma ion o ci us ee canopy segmen a ion. IEEE J Sel Top Appl Ea h Obs Remo e Sens. 17, 1535–1549. h ps:// doi.o g/10.1109/JSTARS.2023.3339290. Jamali, A., Roy, S.K., Hashemi Beni, L., P adhan, B., Li, J., Ghamisi, P., 2024. Residual wa e ision U-Ne o lood mapping using dual pola iza ion Sen inel-1 SAR image y. In J Appl Ea h Obs Geoin [in e ne ]. 127, 103662 h ps://doi.o g/ 10.1016/j.jag.2024.103662. Ji, Y., Yan, E., Yin, X., Song, Y., Wei, W., Mo, D., 2022. Au oma ed ex ac ion o Camellia olei e a c own using unmanned ae ial ehicle isible images and he ResU-Ne deep lea ning model. F on Plan Sci. 13 (Augus ), 1–13. h ps://doi.o g/10.3389/ pls.2022.958940. Jiang, K., A zaal, U., Lee, J., 2023. T ans o me -based weed segmen a ion o g ass managemen . Senso s. 23 (1), 1–15. h ps://doi.o g/10.3390/s23010065. Jin asu isak T, Edi isinghe E, Elba ay A. 2022. Deep neu al ne wo k based da e palm ee de ec ion in d one image y. Compu Elec on Ag ic [In e ne ]. 192(Ap il 2021): 106560. h ps://doi.o g/10.1016/j.compag.2021.106560. Ka enbo n, T., Lei lo , J., Schie e , F., Hinz, S., 2021. Re iew on con olu ional neu al ne wo ks (CNN) in ege a ion emo e sensing. ISPRS J Pho og amm Remo e Sens. 173, 24–49. Ken sch, S., Cace es, M.L.L., Se ano, D., Rou e, F., Diez, Y., 2020a. Compu e ision and deep lea ning echniques o he analysis o d one-acqui ed o es images, a ans e lea ning s udy. Remo e Sens. 12 (8), 1–19. h ps://doi.o g/10.3390/RS12081287. Ken sch S, Ka a siolis S, Kamila is A, Tomha e L, Lopez Cace es ML. 2020. Iden i ica ion o T ee Species in Japanese Fo es s based on Ae ial Pho og aphy and Deep Lea ning. a Xi . h ps://doi.o g/10.1007/978-3-030-61969-5_18. Kolesniko , A., Doso i skiy, A., Weissenbo n, D., Heigold, G., Uszko ei , J., Beye , L., Minde e , M., Dehghani, M., Houlsby, N., Gelly, S., e al., 2021. An image is wo h 16x16 wo ds: T ans o me s o image ecogni ion a scale. In: [place Unknown]. Li, R., Chen, T., Liu, Y., Jiang, H., 2023. CoupleUNe : Swin T ans o me coupling CNNs makes s ong con ex ual encode s o VHR image oad ex ac ion. In J Remo e Sens. 44 (18), 5788–5813. Li K, Wang Y, Gao P, Song G, Liu Y, Li H, Qiao Y. 2022. Uni o me : Uni ied ans o me o e icien spa io empo al ep esen a ion lea ning. a Xi P ep a Xi 220104676. Li, Y., Ma, L., Sun, N., 2024. A bilinea ans o me in e ac i e neu al ne wo ks-based app oach o ine-g ained ecogni ion and p o ec ion o plan diseases o ga dening design. C op P o [in e ne ]. 180, 106660 h ps://doi.o g/10.1016/j. c op o.2024.106660. Lin, T.-Y., Doll´ a , P., Gi shick, R., He, K., Ha iha an, B., Belongie, S., 2017. Fea u e py amid ne wo ks o objec de ec ion. In: P oc IEEE Con Compu is Pa e n Recogni . Honolulu, HI, USA, pp. 2117–2125. Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B. 2021. Swin T ans o me : Hie a chical Vision T ans o me using Shi ed Windows. In: 2021 IEEE/CVF In Con Compu Vis. Mon eal, QC, Canada: IEEE; p. 9992–10002. h ps://doi.o g/10.1109/ ICCV48922.2021.00986. Liu, J., Wang, X., Wang, T., 2019. Classi ica ion o ee species and s ock olume es ima ion in g ound o es images using Deep Lea ning. Compu Elec on Ag ic [in e ne ]. 166 (May), 105012 h ps://doi.o g/10.1016/j.compag.2019.105012. Liu, H., Wang, X., Zhao, F., Yu, F., Lin, P., Gan, Y., Ren, X., Chen, Y., Tu, J., 2024. Upg ading swin-B ans o me -based model o accu a ely iden i ying ipe s awbe ies by coupling ask-aligned one-s age objec de ec ion mechanism. Compu Elec on Ag ic [in e ne ]. 218, 108674 h ps://doi.o g/10.1016/j. compag.2024.108674. Mekhal i, M.L., Nicolo, C., Bazi, Y., Al, R.MM., Alsha i , N.A., Al, M.E., 2022. Con as ing YOLO 5, ans o me , and e icien de de ec o s o c op ci cle de ec ion in dese . IEEE Geosci Remo e Sens Le . 19, 19–23. h ps://doi.o g/10.1109/ LGRS.2021.3085139. Con ibu o s Mms. 2020. {MMSegmen a ion}: OpenMMLab Seman ic Segmen a ion Toolbox and Benchma k. Ozda ici-ok, A., Ok, A.O., 2023. Scien ia Ho icul u ae Using emo e sensing o iden i y indi idual ee species in o cha ds : Sci Ho ic (Ams e dam). [in e ne ]. 321 (10), 112333 h ps://doi.o g/10.1016/j.scien a.2023.112333. Paszke, A., G oss, S., Massa, F., Le e , A., B adbu y, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., An iga, L., e al., 2019. Py o ch: An impe a i e s yle, high- pe o mance deep lea ning lib a y. Ad Neu al In P ocess Sys . 32. Qin, H., Zhou, W., Yao, Y., Wang, W., 2022. Indi idual ee segmen a ion and ee species classi ica ion in sub opical b oadlea o es s using UAV-based LiDAR, hype spec al, and ul ahigh- esolu ion RGB da a. Remo e Sens En i on. 280, 113143. Rezaei, M., Diepe een, D., Laga, H., Jones, M.G.K., Sohel, F., 2024. Plan disease ecogni ion in a low da a scena io using ew-sho lea ning. Compu Elec on Ag ic [in e ne ]. 219, 108812 h ps://doi.o g/10.1016/j.compag.2024.108812. Schie e , F., Ka enbo n, T., F ick, A., F ey, J., Schall, P., Koch, B., Schmid lein, S., 2020. Mapping o es ee species in high esolu ion UAV-based RGB-image y by means o con olu ional neu al ne wo ks. ISPRS J Pho og amm Remo e Sens [in e ne ]. 170 (No embe ), 205–215. h ps://doi.o g/10.1016/j.isp sjp s.2020.10.015. Sel a aju RR, Cogswell M, Das A, Vedan am R, Pa ikh D, Ba a D. 2017. G ad-cam: Visual explana ions om deep ne wo ks ia g adien -based localiza ion. In: P oc IEEE In Con Compu Vis. [place unknown]; p. 618–626. Sun, Y., Li, Z., He, H., Guo, L., Zhang, X., Xin, Q., 2022. Coun ing ees in a sub opical mega ci y using he ins ance segmen a ion me hod. In J Appl Ea h Obs Geoin . 106, 102662 h ps://doi.o g/10.1016/j.jag.2021.102662. Sun, Y., Hao, Z., Guo, Z., Liu, Z., Jh., 2023. De ec ion and mapping o ches nu using deep lea ning om high- esolu ion UAV-Based RGB image . Remo e Sens.:1–18. Suzuki, S., 1985. Topological s uc u al analysis o digi ized bina y images by bo de ollowing. Compu Vision, G aph Image P ocess. 30 (1), 32–46. Tang, X., Tu, Z., Wang, Y., Liu, M., Li, D., Fan, X., 2022. Au oma ic de ec ion o coseismic landslides using a new ans o me me hod. Remo e Sens. 14 (12), 1–19. h ps://doi. o g/10.3390/ s14122884. Tolan J, Yang HI, Nosa zewski B, Couai on G, Vo H V., B and J, Spo e J, Majumda S, Haziza D, Vama aju J, e al. 2024. Ve y high esolu ion canopy heigh maps om RGB image y using sel -supe ised ision ans o me and con olu ional decode ained on ae ial lida . Remo e Sens En i on [In e ne ]. 300(Ap il 2023):113888. h ps://doi.o g/10.1016/j. se.2023.113888. To es, D.L., Fei osa, R.Q., Happ, P.N., La Rosa, L.E.C., Junio , J.M., Ma ins, J., B essan, P.O., Gonçal es, W.N., Liesenbe g, V., 2020. Applying ully con olu ional a chi ec u es o seman ic segmen a ion o a single ee species in u ban en i onmen on high esolu ion UAV op ical image y. Senso s (swi ze land). 20 (2), 1–20. h ps://doi.o g/10.3390/s20020563. Ulku, I., Akagündüz, E., Ghamisi, P., Membe , S., 2022. Deep Seman ic Segmen a ion o T ees Using Mul ispec al Images. 15, 7589–7604. h ps://doi.o g/10.1109/ JSTARS.2022.3203145. Vaswani A, Shazee N, Pa ma N, Uszko ei J, Jones L, Gomez AN, Kaise Ł, Polosukhin I. 2017. A en ion is all you need. In: Ad Neu al In P ocess Sys . Vol. 30. [place unknown]; p. 5998–6008. Velasquez-camacho, L., E xega ai, M., 2023. Deep Lea ning algo i hms o u ban ee de ec ion and geoloca ion wi h high- esolu ion ae ial, sa elli e, and g ound-le el images. Compu En i on U ban Sys [in e ne ]. 105 (Augus ), 102025 h ps://doi. o g/10.1016/j.compen u bsys.2023.102025. Wagne , F.H., Sanchez, A., Ta abalka, Y., Lo e, R.G., Fe ei a, M.P., Aida , M.P.M., Gloo , E., Phillips, O.L., A ag˜ ao, L.E.O.C., 2019. Using he U-ne con olu ional ne wo k o map o es ypes and dis u bance in he A lan ic ain o es wi h e y high esolu ion images. Remo e Sens Ecol Conse . 5 (4), 360–375. h ps://doi.o g/ 10.1002/ se2.111. Wagne , F.H., Sanchez, A., Aida , M.P.M., Rochelle, A.L.C., Ta abalka, Y., Fonseca, M.G., Phillips, O.L., Gloo , E., A ag˜ ao, L.E.O.C., 2020. Mapping A lan ic ain o es deg ada ion and egene a ion his o y wi h indica o species using con olu ional ne wo k. PLoS One. 15 (2), e0229448. Wang W, Xie E, Li X, Fan D-P, Song K, Liang D, Lu T, Luo P, Shao L. 2021. Py amid Vision T ans o me : A Ve sa ile Backbone o Dense P edic ion wi hou Con olu ions [In e ne ]. :568–578. h p://a xi .o g/abs/2102.12122. Wu, Y., Mis a, S., 2020. In elligen image segmen a ion o o ganic- ich shales using andom o es , wa ele ans o m, and hessian ma ix. IEEE Geosci Remo e Sens Le . 17 (7), 1144–1147. h ps://doi.o g/10.1109/LGRS.2019.2943849. Xia Z, Pan X, Song S, Li LE, Huang G. 2022. Vision T ans o me wi h De o mable A en ion [In e ne ]. h p://a xi .o g/abs/2201.00520. Xiao T, Liu Y, Zhou B, Jiang Y, Sun J. 2018. Uni ied Pe cep ual Pa sing o Scene Unde s anding. Lec No es Compu Sci (including Subse Lec No es A i In ell Lec No es Bioin o ma ics). 11209 LNCS:432–448. h ps://doi.o g/10.1007/978-3-030- 01228-1_26. Xiao, C., Qin, R., Huang, X., 2020. T ee op de ec ion using con olu ional neu al ne wo ks ained h ough au oma ically gene a ed pseudo labels. In J Remo e Sens [in e ne ]. 41 (8), 3010–3030. h ps://doi.o g/10.1080/01431161.2019.1698075. Xie, E., Wang, W., Yu, Z., Anandkuma , A., Al a ez, J.M., Luo, P., 2021. SegFo me : simple and e icien design o seman ic segmen a ion wi h ans o me s. Ad Neu al In P ocess Sys . 15, 12077–12090. Yang, M., Mou, Y., Liu, S., Meng, Y., Liu, Z., Li, P., Xiang, W., Zhou, X., Peng, C., 2022b. De ec ing and mapping ee c owns based on con olu ional neu al ne wo k and Google Ea h images. In J Appl Ea h Obs Geoin . 108(Augus , 2021). h ps://doi. o g/10.1016/j.jag.2022.102764. Yang, L., Wang, X., Zhai, J., 2022a. Wa e line ex ac ion o a i icial coas wi h ision ans o me s. F on En i on Sci. 10 (Feb ua y), 1–12. h ps://doi.o g/10.3389/ en s.2022.799250. Yi, S., Liu, X., Li, J., Chen, L., 2023. UAV o me : a composi e ans o me ne wo k o u ban scene segmen a ion o UAV images. Pa e n Recogni . 133 h ps://doi.o g/ 10.1016/j.pa cog.2022.109019. Zhang, T., Hu, D., Wu, C., Liu, Y., Yang, J., Tang, K., 2023. La ge-scale apple o cha d mapping om mul i-sou ce da a using he seman ic segmen a ion model wi h image- o-image ansla ion and ans e lea ning. Compu Elec on Ag ic. 213, 108204. Zhang, L., Lin, H., Wang, F., 2022. Indi idual ee de ec ion based on high- esolu ion RGB images o u ban o es y applica ions. IEEE Access. 10, 46589–46598. h ps:// doi.o g/10.1109/ACCESS.2022.3171585. Zhang, H., Liu, S., 2024. Double-b anch mul i-scale con ex ual ne wo k: a model o mul i-scale s ee ee segmen a ion in high- esolu ion emo e sensing images. Senso s. 24 (4), 1110. h ps://doi.o g/10.3390/s24041110. Zhao, J., Be ge, T.W., Geipel, J., 2023b. T ans o me in UAV image-based weed mapping. Remo e Sens. 15 (21), 5165. Zhao H, Shi J, Qi X, Wang X, Jia J. 2017. Py amid scene pa sing ne wo k. In: P oc - 30 h IEEE Con Compu Vis Pa e n Recogni ion, CVPR 2017. Vol. 2017-Janua. [place unknown]; p. 6230–6239. h ps://doi.o g/10.1109/CVPR.2017.660. Zhao, H., Mo gen o h, J., Pea se, G., Schindle , J., 2023a. A Sys ema ic e iew o indi idual ee c own de ec ion and delinea ion wi h con olu ional neu al ne wo ks (CNN). Cu o Repo s [in e ne ]. 9 (3), 149–170. h ps://doi.o g/10.1007/ s40725-023-00184-3. Zheng, J., Wu, W., Yu, L., Fu, H., 2021. COCONUT TREES DETECTION ON THE TENARUNGA USING HIGH-RESOLUTION SATELLITE IMAGES AND DEEP LEARNING Minis y o Educa ion Key Labo a o y o Ea h Sys em Modeling, and R. Al-Ruzouq e al. Ecological Indica o s 163 (2024) 112110 18 Depa men o Ea h Sys em Science, Tsinghua Uni e si y, Beijing 100084, China Na ion. In Geosci Remo e Sens. Symp.:6512–6515. Zheng, Y., Wu, G., 2022. YOLO 4-Li e – Based U ban Plan a ion T ee De ec ion and Posi ioning wi h High-Resolu ion Remo e Sensing Image y. 9(Janua y):1–12 h ps://doi.o g/10.3389/ en s.2021.756227. Zheng, J., Yuan, S., Wu, W., Li, W., Yu, L., Fu, H., Coomes, D., 2023. Su eying coconu ees using high- esolu ion sa elli e image y in emo e a olls o he Paci ic Ocean. Remo e Sens En i on. 287 h ps://doi.o g/10.1016/j. se.2023.113485. Zhou, C., Ye, H., Sun, D., Yue, J., Yang, G., Hu, J., 2022. An au oma ed, high- pe o mance app oach o de ec ing and cha ac e izing b occoli based on UAV emo e-sensing and ans o me s: A case s udy om Haining, China. In J Appl Ea h Obs Geoin [in e ne ]. 114 (July), 103055 h ps://doi.o g/10.1016/j. jag.2022.103055. Zhou, B., Zhao, H., Puig, X., Fidle , S., Ba iuso, A., To alba, A., 2017. Scene Pa sing h ough ADE20K Da ase . In: In: P oc IEEE Con Compu is Pa e n Recogni . [place Unknown], p. p. page 4.. Zhu X, Su W, Lu L, Li B, Wang X, Dai J. 2020. De o mable DETR: De o mable T ans o me s o End- o-End Objec De ec ion [In e ne ]. :1–16. h p://a xi .o g/ abs/2010.04159. R. Al-Ruzouq e al.