Full text
Ecological Indica o s 163 (2024) 112110
A ailable online 8 May 2024
1470-160X/© 2024 The Au ho s. Published by Else ie L d. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by-
nc-nd/4.0/).
O iginal A icles
Spec al–Spa ial ans o me -based seman ic segmen a ion o la ge-scale
mapping o indi idual da e palm ees using e y high- esolu ion
sa elli e da a
Rami Al-Ruzouq
a
, Mohamed Ba aka A. Gib il
a
,
*
, Abdallah Shanableh
a
, Jan Bolcek
a
,
b
,
Fouad Lamgha i
c
, Neza A alla Hammou
d
,
e
, Ali El-Keblawy
, Ra i anjan Jena
a
a
GIS and Remo e Sensing Cen e , Resea ch Ins i u e o Sciences and Enginee ing, Uni e si y o Sha jah, Sha jah 27272, Uni ed A ab Emi a es
b
Depa men o Radio Elec onics, Facul y o Elec ical Enginee ing and Communica ion, B no Uni e si y o Technology, B no-K alo o pole 61600, Czech Republic
c
Fujai ah Resea ch Cen e, Al-Hilal Towe , 3003, P.O. Box 666 Fujai ah, Uni ed A ab Emi a es
d
Depa men o Applied Physics and As onomy, Facul y o Science, Uni e si y o Sha jah, Sha jah 27272, Uni ed A ab Emi a es
e
Depa men o Ea h and En i onmen al Sciences, P ince El-Hassan Bin Talal Facul y o Na u al Resou ces & En i onmen , The Hashemi e Uni e si y, Za qa 13133,
Jo dan
Depa men o Applied Biology, College o Sciences, Uni e si y o Sha jah, Sha jah P.O. Box 2727, Uni ed A ab Emi a es
ARTICLE INFO
Keywo ds:
T ee c own delinea ion
Seman ic segmen a ion
Vision ans o me s
Deep lea ning
ABSTRACT
Da e palm plan a ions in he Uni ed A ab Emi a es (UAE) a e unde h ea om soil salini y, d ough , and da e
palm wee ils. Acco dingly, moni o ing and conse ing da e palms a e c ucial o p ese ing a i al componen o
he coun y’s ag icul u al he i age, economy, ood secu i y, and ecological balance. P e ious s udies ha e
e ec i ely iden i ied da e palm ees using RGB-based ae ial and UAV image y u ilizing di e se deep lea ning
me hods. Howe e , he u iliza ion o e y high- esolu ion sa elli e da a o delinea ing indi idual da e palm
c owns emains unexplo ed due o he limi ed spa ial esolu ion capabili ies o exis ing sa elli e sys ems. This
s udy p ima ily aimed o achie e p ecise and comp ehensi e mapping o da e palm ees using Wo ldView-3
(WV-3) sa elli e da a by le e aging he high ep esen a ional powe o he s a e-o - he-a ision ans o me s
(ViT) in cap u ing global in o ma ion om he inpu da a. Fi s , an in-dep h analysis assessmen o he a ious
ans o me -based seman ic segmen a ion a chi ec u es, including Upe Ne wi h ision ans o me and Swin
ans o me , SegFo me , Mask2Fo me , and UniFo me , was conduc ed. Second, he in eg a ion o spec al da a
on he pe o mance o ViTs was e alua ed. Mo eo e , he models’ gene alizabili y and complexi y e ec on he
segmen a ion e ec i eness we e assessed. Acco dingly, a pos p ocessing s a egy was de eloped o aid in
delinea ing and coun ing da e palm ees om seman ic segmen a ion ou pu s. Resul s demons a ed ha in e-
g a ion o WV-3 spec al da a in o he analysis esul ed in a ma ked imp o emen in segmen a ion quali y. The
UniFo me , Upe Ne -Swin, and Mask2Fo me models demons a ed conside able imp o emen s in mul ispec al
da a analysis, wi h inc eases in mean in e sec ion o e union (mIoU) o 2.17% (77.88% mIoU, 86.01% mean F-
sco e [mF-sco e]), 2% (78.10% mIoU, 86.18% mF-sco e), and 1.15% (77.36% mIoU, 85.59% mF-sco e),
espec i ely, compa ed wi h hei RGB-based esul s. E alua ions o model ans e abili y also indica ed ha
Mask2Fo me , UniFo me , and Upe Ne -Swin ans o me s e icien ly adap ed o mul ispec al da a in he Dibba
egion. These models achie ed mIoU sco es o 84.36%, 84.25%, and 83.17% and mF-sco es o 90.95%, 90.87%,
and 90.13%, highligh ing hei e ec i eness and po en ial o b oade egional applica ion. This esea ch
highligh s he e icacy and easibili y o using ViTs wi h WV-3 mul ispec al da a o accu a e and comp ehensi e
su eying o da e palm plan a ions, enabling he de elopmen o palm ee in en o ies and con inuously
upda ing geospa ial da abases.
* Co esponding au ho .
E-mail add ess: [email p o ec ed] (M.B.A. Gib il).
Con en s lis s a ailable a ScienceDi ec
Ecological Indica o s
jou nal homepage: www.else ie .com/loca e/ecolind
h ps://doi.o g/10.1016/j.ecolind.2024.112110
Recei ed 8 Feb ua y 2024; Recei ed in e ised o m 15 Ap il 2024; Accep ed 2 May 2024
Ecological Indica o s 163 (2024) 112110
2
1. In oduc ion
1.1. Backg ound
The 2021 s a is ics o he Food and Ag icul u e O ganiza ion (FAO)
indica ed ha app oxima ely 9.66 million ons o da e palms we e
p oduced, co e ing an a ea o 1.3 million hec a es (FAO, 2023). The
Uni ed A ab Emi a es (UAE) is among he op en p oducing coun ies,
wi h almos 40 million ees dis ibu ed ac oss he UAE (El-Juhany,
2010). Da e palm ees a e an impo an pa o UAE’s cul u al he i age
and a i al ag icul u al esou ce. Thus, accu a e mapping o da e palm
plan a ions ac oss he coun y is impe a i e o e ec i ely and sus ain-
ably moni o and manage his pi o al ag icul u al asse .
T adi ional me hods o da e palm ee in en o y de elopmen , such
as on-si e su eys and isual examina ion o ae ial pho og aphs, a e
cos ly, labo ious, and ime-consuming. Remo e sensing echniques
p o ide a mo e cos -e icien app oach, o e ing de ailed da a wi h
e sa ile empo al and spa ial esolu ions (Ha ling e al., 2019). Remo e
sensing using unmanned ae ial ehicles (UAVs), inc easingly adop ed in
nume ous s udies ocusing on da e palms o i s supe io spa ial and
empo al esolu ions and p ecise posi ional accu acy, aids in he p ecise
de ec ion and mapping o da e palm ees (Amma e al., 2021; Gib il
e al., 2022; Jin asu isak e al., 2022; Gib il e al., 2023). Unlike ea h-
obse ing sa elli es, UAVs ha e limi ed a ea co e age and a e con-
s ained by wea he condi ions and lying es ic ions (Ozda ici-ok and
Ok, 2023). Sa elli e emo e sensing p o ides a comp ehensi e iew,
enabling epea ed obse a ions and ex ensi e co e age ac oss as e-
gions. Combining sa elli e and UAV da a could esul in a mo e
comp ehensi e and in-dep h assessmen o da e palm ees.
A di e se ange o machine lea ning algo i hms has been es ablished
and u ilized o localize and cha ee c owns using da a de i ed om
emo e sensing in au oma ic and semi-au oma ic manne s (Al-Ruzouq
e al., 2018; Ghasemi e al., 2022; Ji e al., 2022; Qin e al., 2022; Zhang
e al., 2023; H. Zhao e al., 2023). In ecen yea s, compu e ision
echniques based on deep lea ning, p ima ily con olu ional neu al
ne wo ks (CNNs), ha e become inc easingly p e alen o indi idual
ee c own (ITC) delinea ing om e y high- esolu ion (VHR) emo ely
sensed images. This p e alence can be a ibu ed o he ema kable
capabili y o hese me hods o au oma ically ex ac in ica e high-le el
ea u es and complex pa e ns om he inpu image da ase s, subs an-
ially enhancing he e icacy and obus ness o he models (Gib il e al.,
2022; Zheng e al., 2023). CNNs ha e demons a ed ou s anding pe -
o mance in a ious ITC s udies (H. Zhao e al., 2023) due o hei
dis inc i e a chi ec u e, which includes localized ecep i e ields,
weigh sha ing, and he p ocess o subsampling (Ka enbo n e al.,
2021). CNN-based models a e ypically used o unde ake di e en asks
o de ec and delinea e ee species. Some widely used asks include
objec de ec ion (Zheng e al., 2021; Zheng and Wu, 2022; Cai e al.,
2023; Velasquez-camacho and E xega ai, 2023), ins ance (B aga e al.,
2020; Yang e al., 2022b; Ball e al., 2023; Hao e al., 2023), and se-
man ic segmen a ion (F eudenbe g e al., 2022; Gib il e al., 2023).
Nume ous seman ic segmen a ion models based on CNNs, a se o
encode –decode deep lea ning a chi ec u es ha a e used o pe o m
pixel-wise classi ica ion, has been de eloped and u ilized o ou line ee
canopies om di e en ypes o emo ely sensed da a (Bha naga e al.,
2020; Cao and Zhang, 2020; Ken sch e al., 2020a; Ken sch e al., 2020b;
To es e al., 2020; Wagne e al., 2020; Wu and Mis a, 2020; Xiao e al.,
2020; Akca and Pola , 2022; Sun e al., 2023). F eudenbe g e al. (2022)
u ilized a wo-s ep app oach based on U-Ne a chi ec u e o ITC om
Wo ldView-3 (WV-3) sa elli e da a (acqui ed in Bengalu u, India) and
ae ial image y (acqui ed in a densely o es ed a ea nea Ga ow, Ge -
many). Thei app oach in ol es ex ac ing ee c owns wi h an
in e sec ion-o e -union (IoU) o 71.2 % on he sa elli e da a and an IoU
o 81.9 ±2.2% on he ae ial image y. Seman ic segmen a ion a chi-
ec u es, wi h di e en backbone ne wo ks, commonly u ilized o ee
c own delinea ion in ecen yea s include he U-Ne (F eudenbe g e al.,
2019; Ken sch e al., 2020a; Wagne e al., 2019,2020; Liu,Wang,and
Wang, 2019; Ken sch e al., 2020b; Schie e e al., 2020) and Deep-
LabV3+(Ayhan and Kwan, 2020; Fe ei a e al., 2020; Cheng,Qi,and
Cheng, 2021).
Gi en ha con olu ions p ocess images by examining local a eas,
hei design inhe en ly es ic s hem om g asping he global con ex o
an en i e image in a single ope a ion, po en ially diminishing accu acy,
especially in scena ios whe e he inpu images exhibi in ica e e-
la ionships be ween pixels (Fayad e al., 2024). Recen ly, he ield o
emo e sensing has expe ienced an inc easing adop ion o di e se ision
ans o me s (ViT) o a a ie y o asks (Abozeid e al., 2022; Fan e al.,
2022; Mekhal i e al., 2022; Sun e al., 2022; L. Yang e al., 2022; Zhou
e al., 2022; Li e al., 2023; Yi e al., 2023; Zhao e al., 2023b; Jamali
e al., 2024). In con as o CNNs, ope a ions wi hin ViTs a e inhe en ly
pa allel and sequence-independen , allowing ViTs o e icien ly ga he
comp ehensi e con ex ual in o ma ion and a ain a supe io le el o
ep esen a ional capabili y (Xia e al., 2022; Zhou e al., 2022). Addi-
ionally, encode s cons uc ed wi h ViTs main ain a uni o m dimen-
sional ep esen a ion h oughou he p ocessing phases, hus
ci cum en ing he necessi y o explici down-sampling ope a ions
(Fayad e al., 2024). Recen ly, an inc easing numbe o s udies ha e
u ilized a ious ViTs in he ields o plan disease de ec ion (Chang e al.,
2024; Li e al., 2024; Rezaei e al., 2024), ui iden i ica ion (Bai e al.,
2024; Liu e al., 2024), canopy heigh mapping (Fayad e al., 2024;
Tolan e al., 2024), and c own delinea ion s udies.
1.2. Rela ed wo k
T ans o me -based models a e inc easingly u ilized in ITC esea ch
o a a ie y o isual asks, including ins ance segmen a ion (Gib il
e al., 2022; De sch e al., 2023; Fi oze e al., 2023), objec de ec ion
(De sch e al., 2022; Zhang e al., 2022), and ee coun ing (Chen and
Shang, 2022; Ami kolaee e al., 2023). De sch e al. (2023) u ilized he
ans o me -based ins ance segmen a ion app oach de ec ion ans-
o me (DETR) o delinea e ITCs in mixed o es a eas u ilizing mul i-
spec al UAV images and UAV-de i ed LiDAR da a. The au ho s
conduc ed expe imen s wi h wo-channel selec ions – RGB and Colo -
In a ed (CIR) – using DETR and achie ed he highes mean F-sco es
(mF-sco es) on he CIR da ase , a aining 90 % in coni e ous o es s, 81
% in deciduous, and 84 % in mixed o es s. Zhang e al. (2022) used a
Fas e R-CNN wi h a Swin ans o me and ResNe -50 backbone o
iden i ying indi idual ees om VHR ae ial images. The s udy exam-
ined he pe o mance o he models in di e se en i onmen s and
obse ed hei highes pe o mance in s ee a eas. Chen and Shang
(2022) in oduced a densi y ans o me designed o he au oma ed
coun ing o ees in ae ial pho og aphs and e alua ed i s e icacy agains
se e al leading-edge algo i hms. Thei p oposed ans o me -based
me hod showed p omising esul s, su passing he pe o mance o
a ious exis ing s a e-o - he-a echniques. Ami kolaee e al. (2023)
in oduced a semi-supe ised me hodology based on a ans o me a -
chi ec u e (encode –decode design) o ee coun ing. Thei ne wo k’s
encode , designed o ex ac ea u es a mul iple scales, u ilized a py -
amid ision ans o me (Wang e al., 2021). Despi e u ilizing he same
quan i y o labeled images, his no el app oach ou pe o med exis ing
semi-supe ised and ully-supe ised echniques in pe o mance.
Fu e al. (2023) examined he pe o mance o wo algo i hms,
namely, adap i e s acking ensemble lea ning (ASEL) and mul i-pa h
ision ans o me o dense p edic ion (MPViT), o mang o e species
mapping om mul ispec al UAV-based images. The s udy showed ha
MPViT ou pe o med ASEL and achie ed 95.5 %–97.3 % o he o e all
classi ica ion accu acy. Zhou e al. (2022) p oposed a semi-au oma ic
wo k low-based T ansUNe o de ec and cha ac e ize b occoli canopy
and heads om he RGB image y and LiDAR da a. Thei esul s showed
ha he p oposed ans o me -based amewo k ou pe o med h ee
CNN-based and wo shallow lea ning-based echniques. Zhang and Liu
(2024) in oduced a dual-b anch ne wo k designed o segmen s ee
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
3
ees using high- esolu ion sa elli e image y. They in eg a ed dila ed
con olu ional laye s in pa allel wi h ans o me blocks o imp o e he
ne wo k’s capaci y o ex ac ing ea u es ac oss mul iple scales. He
e al. (2024) le e aged he CNN-based ne wo k (imp o ed E icicen Ne -
V2) and ans o me -based ne wo k (CSwin ans o me ) o cap u e
local and global seman ic in o ma ion o ci us ee canopy segmen a-
ion om 3D da a ob ained om UAV. Thei app oach showed supe io
pe o mance han using some CNN only and some ans o me -based
only.
A b oad spec um o s udies has ocused on de ec ing palm ees om
mul ipla o m ae ial o hopho os using CNNs (Culman e al., 2020;
Gib il e al., 2021; Amma e al., 2021; Gib il e al., 2022; Jin asu isak
e al., 2022). Gib il e al., (2022,2023, & 2024) demons a ed ha ViTs
ou pe o med a ious CNN-based a chi ec u es o he seman ic and
ins ance segmen a ion o indi idual da e palms om mul ipla o m
ae ial images. Gib il e al. (2022) comp ehensi ely e alua ed a ious
ins ance segmen a ion a chi ec u es o iden i y and ou line indi idual
palm ees h ough mul iscale d one-based images. The au ho s e alu-
a ed a ange o ne wo ks, including Mask R-CNN wi h di e en CNN-
based and Swin ans o me backbones, SOLO, SOLO 2, YOLACT, and
Mask Sco ing R-CNN. The ans o me -based model, Mask R-CNN wi h
Swin ans o me backbones, exhibi ed supe io pe o mance, su pass-
ing CNN-based models. Gib il e al. (2023) examined he ealiabili y o
a ious ision ans o me o ex ac da e palm ees om mul iscale
UAV and ae ial images. Thei esul s showed ha ans o me -based
models achi ed an mF-sco e anging om 91.62 % o 92.44 %.
Among he e alua ed models, he SegFo me model ou pe o med all o
he e alua ed CNN-based models in he mul iscale and he independen
es ing da ase s, ollowed by he Upe Ne -Swin ans o me . Gib il e al.
(2024) p esen ed a ans o me -based me hod o delinea e, coun i y,
and assess he heal h o indi idual da e palm ees om la ge-scale UAV
da a. The p oposed me hod is based on he syne gy be ween an
imp o ed mul iscale ViT (MViT 2), a ea u e py amid ne wo k, Mask R-
CNN, and an imp o ed slicing-aided hype in e ence module. The
model, which was ini ially de eloped o delinea e da e palm ees, was
hen ine- uned o assess he heal h o he da e palm ees. This a chi-
ec u e ou pe o med a ious ins ance segmen a ion models, achie ing
an F-sco e o 94.2 % o segmen a ion and 88.4 % o heal h assessmen .
P e ious esea ch on mapping and moni o ing palm ees based on
deep lea ning has p edominan ly u ilized d one-based and ae ial im-
age y wi h spa ial esolu ions anging om 2 cm o 20 cm.
Dis inguishing he c owns o palm ees in he e ogeneous landscapes
in ensi ies as he da a’s g ound space dis ance (GSD) becomes coa se .
This inc eased di icul y can ad e sely a ec he pe o mance o se-
man ic segmen a ion models, po en ially hinde ing hei abili y o
accu a ely ecognize indi idual da e palm ees. The easibili y o u i-
lizing VHR sa elli e image y o ex ac da e palm ees om sa elli e
images has ye o be ho oughly in es iga ed. This wo k mainly aims o
in oduce an e ec i e deep lea ning me hodology o ex ensi e, de ailed
and accu a e egional su eys o da e palm dis ibu ion, u ilizing he
VHR WV-3 sa elli e image y ac oss mul iple ci ies. Accu a e mapping
and quan i ica ion o indi idual da e palms ac oss egions enables he
au oma ed de elopmen o da e palm ee in en o ies and can p o ide
insigh s in o hei heal h and gene ic di e si y. These da a a e c ucial o
de eloping conse a ion measu es and ensu ing sus ainable cul i a ion
in hei na u al habi a s. The speci ic objec i es o his s udy a e o (1)
e alua e he pe o mance o di e en deep ViTs in egional su eying o
da e palms using WV-3 sa elli e da a, (2) in es iga e he in luence o
spec al da a in eg a ion on he e icacy o he e alua ed ViTs, (3) assess
model gene alizabili y ac oss a ied geog aphic egions, and (4)
de elop a no el pos p ocessing echnique o indi idual da e palm ee
delinea ion and quan i ica ion.
2. Ma e ials and me hods
2.1. O e iew
The esea ch me hodology encompasses i e main s ages (Fig. 1).
Ini ially, WV-3 sa elli e da a iles we e p ep ocessed, including a mo-
sphe ic co ec ion, image pansha pening, and da a no maliza ion. The
second s age in ol ed manually anno a ing da e palm ees; di iding he
da a in o aining, alida ion, and es ing egions; and gene a ing image-
mask pai s. The hi d s age consis ed o comp ehensi e expe imen s o
assess he e icacy o se e al ad anced ViTs in delinea ing da e palm
ees in WV-3 sa elli e image y, u ilizing spa ial and spec al in o ma-
ion. The e alua ed ViTs in his s udy a e Upe Ne (Xiao e al., 2018)
wi h ision ans o me (Kolesniko e al., 2021) and Swin ans o me
(Liu e al., 2021), SegFo me (Xie e al., 2021), Mask2Fo me (Cheng
e al., 2022), and UniFo me (Li e al., 2022). T aining and e alua ion o
hese models we e conduc ed using di e se da a inpu s, including RGB,
a combina ion o RGB wi h NIR1 and NIR2, eigh mul ispec al chan-
nels, and an amalgama ion o hese channels wi h NDVI. In he ou h
Fig. 1. Me hodological amewo k o his wo k.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
4
s age, he adap abili y o hese models o a ious WV-3 da ase s om
di e en loca ions in Dibba, collec ed on di e en da es, was examined
o de e mine hei gene aliza ion capabili ies. The inal s ep p oposes a
da a pos p ocessing echnique o coun ing and delinea ing indi idual
da e palm ees.
Fig. 2. Loca ion o he expe imen al si es: (a) UAE emi a es, (b) de ailed zoom o he s udy si es, and (c–e) enhanced zoom le els highligh ing he da e palm ees as
seen in mul iple WV-3 images.
Fig. 3. Va ia ions in da e palm ees in e ms o c own size, shape, heigh , and he su ounding en i onmen s.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
5
2.2. S udy a eas
The expe imen al si es a e si ua ed ac oss a ious ci ies in he UAE,
including Fujai ah, Sha jah, and Ajman (Fig. 2a and 2b). Speci ically,
he a eas a e loca ed in eas Ajman, Al Zubai , Kalba, Dibba, and Masa i.
Collec i ely, hese a eas encompass an a ea o 275 km
2
, dis ibu ed as
ollows: Ajman and Sha jah co e an a ea o 38.6 km
2
, Kalba encom-
passes an a ea o 93 km
2
, Dibba is 77.1 km
2
, and Masa i spans an a ea o
66.3 km
2
. The s udy si es hos a di e se a ay o ee and ege a ion
species, including he da e palm (Phoenix dac yli e a L.), gha (P osopis
cine a ia), Mesqui e (P osopis juli lo a), Sid o Ch is ’s ho n jujube
(Ziziphus spina-ch is i), umb ella ho n acacia (Acacia o ilis), and neem
(Azadi ach a indica), among a ious sh ubs and o he lo a. Fig. 2c–e
display close-up iews o he egions ma ked by cyan ci cles, high-
ligh ing di e en da e palm ees in WV-3 sa elli e da a.
In he p elimina y analysis s age, a comp ehensi e ield campaign
was conduc ed o ho oughly unde s and he loca ions, dis ibu ions,
and isual cha ac e is ics o da e palm ees. This campaign was also
c ucial o gaining insigh in o he s udy a ea’s b oade geog aphic and
en i onmen al con ex s, such as ege a ion and ee species’ e ain,
a ie y, and su oundings. The A cGIS Field Map mobile applica ion was
u ilized o eco d and manage obse a ions h oughou he ield isi s.
The coo dina es o he da e palm plan a ions wi hin he s udy a eas
we e documen ed. The da e palm ees displayed a conside able a ia-
ion in c own shape and size, age, heigh , densi y, and he su ounding
landscape cha ac e is ics (Fig. 3).
2.3. Da a acquisi ion and p ep ocessing
In his esea ch, a o al o 20 mul ida e VHR WV-3 sa elli e image
iles we e u ilized. The WV-3 da ase ea u es eigh spec al bands wi h a
GSD o 1.24 m alongside a panch oma ic band wi h a 0.31 m GSD. The
WV-3 image y o he di e en si es was acqui ed in Ajman and Sha jah
ci y in Ma ch 2021, Kalba in Augus 2022, Dibba in No embe 2019, and
Masa i in Sep embe 2019. Table 1 lis s he gene al cha ac e is ics o he
WV-3 da a. The p ocessing o he WV-3 da a comp ised se e al s ages.
Fi s , a mosphe ic co ec ion was conduc ed using he FLAASH me hod.
Second, he spa ial esolu ion o he mul ispec al da a was imp o ed
h ough he applica ion o he G am–Schmid pansha pening echnique.
The ea e , he bi dep h o he da a unde wen con e sion om 11 bi o
8 bi , s anda dizing he image alues wi hin a 0–255 ange.
Conduc ing p ecise, la ge-scale mapping o da e palm ees u ilizing
WV-3 sa elli e da a wi h classical machine lea ning echniques, which
p ima ily depend on he spec al in o ma ion o he da a, encoun e
no able sho comings. Some expec ed challenges include misclassi ica-
ion and a dec ease in gene alizabili y, which may esul om sub-
s an ial a ia ions in acquisi ion da es and illumina ion condi ions o he
acqui ed sa elli e da a. Fig. 4 illus a es he a e age pixel alues o
Table 1
Cha ac e is ics o he WV-3 da a.
Channel Wa eleng h ange (nm) GSD (m)
Panch oma ic 450–800 0.31
Coas al 400–450 1.24
Blue 450–510 1.24
G een 510–580 1.24
Yellow 585–625 1.24
Red 630–690 1.24
Red Edge 705–745 1.24
NIR-1 770–895 1.24
NIR-2 860–1,040 1.24
Fig. 4. A e age pixel alues o comp ehensi e da e palm ees de i ed om a ious channels o he WV-3 da ase .
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
6
di e se da e palm ees, de i ed om a ious channels o he p e-
p ocessed WV-3 da ase , based on an analysis o 76,950 indi idual da e
palm ees. Fu he mo e, he da e palms wi hin he s udy a ea exhibi ed
a wide ange o c own sizes, s uc u es, ages, and heal h s a es, adding
ano he laye o complexi y o he mapping p ocess. The e o e, his
s udy a emp s o le e age spec al, spa ial, and global con ex ual in-
o ma ion h ough he u iliza ion o s a e-o - he-a ViTs o ensu e ac-
cu a e la ge-scale mapping o da e palm ees om di e en WV-3
da ase s.
2.4. Da a p epa a ion o seman ic segmen a ion
The WV-3 images ea u ing da e palm ees om Ajman and Kalba
we e manually ou lined using g ound- u h da a and isual in e p e a-
ion o VHR ae ial images o ain and assess di e en ans o me -based
deep lea ning models. The p epa a ion o g ound u h da a was
Fig. 5. Examples o (a), (b), and (c) image iles, and he co esponding anno a ions (A, B, C) we e selec ed om he aining da ase s.
Fig. 6. S eps in ol ed in pos p ocessing o delinea e and enume a e sepa a e da e palm ees.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
7
comp ehensi e and aimed o include la ge a eas wi h di e se a ia ions
in da e palm ees, such as palm ype, canopy size, ele a ion, age, and
adjacen en i onmen s. Delinea ing smalle da e palm ees was chal-
lenging due o he sa elli e da a’s limi ed spa ial esolu ion. The manual
e i ica ion o delinea ed palms, se ing as g ound- u h da a, was
ho oughly conduc ed by analys s using VHR ae ial and UAV image y
om a ious loca ions in he s udy a ea and e ised by mul iple anno-
a o s o quali y assu ance. The ec o da a encompassed a o al o
76,950 da e palm ees sca e ed ac oss ag icul u al and u ban land-
scapes. The WV-3 da a and i s anno a ions we e c opped in o consis en
512 ×512 image iles. A o al o 16,260 image iles we e dedica ed o
de eloping deep lea ning models. Meanwhile, app oxima ely 1800
images we e kep o e alua ing and es ing he models. Fig. 5 displays a
se o image iles and hei co esponding labels (masks) chosen om he
aining da ase .
2.5. Seman ic segmen a ion a chi ec u es
Va ious ans o me -based seman ic segmen a ion a chi ec u es
we e subjec ed o aining and e alua ion, u ilizing a di e se a ay o
spec al image composi es (i.e., s anda d RGB, an in eg a ion o RGB
wi h NIR1 and NIR2 bands, o eigh dis inc mul ispec al channels, and
a usion o hese channels wi h he NDVI). The e alua ed seman ic
segmen a ion models include Upe Ne (Xiao e al., 2018) based on ViT
backbone, Upe Ne based on Swin ans o me (Liu e al., 2021),
Mask2Fo me (Cheng e al., 2022), SegFo me (Xie e al., 2021), and
UniFo me (Li e al., 2022). The models we e implemen ed u ilizing he
PyTo ch (Paszke e al., 2019) and MMsegmen a ion (MMsegmen a ion,
2020) amewo k. The ollowing subsec ions o e a b ie desc ip ion o
he a o emen ioned models.
2.5.1. Upe Ne -ViT
The Upe Ne , an encode –decode ne wo k o seman ic segmen a-
ion de eloped by Xiao e al. (2018), is based on a ea u e py amid
ne wo k (FPN) amewo k (Lin e al., 2017). This amewo k in eg a es a
py amid pooling module (PPM) om he PSPNe (Zhao e al., 2017). The
FPN acili a es a dual p ocess in ol ing he downsampling o ea u e
maps h ough a bo om-up pa hway and subsequen upsampling ia a
op-down mechanism connec ed h ough la e al linkages. This bo om-
up pa hway gene a es ea u e maps ac oss mul iple scales, con ingen
on he a chi ec u e o he chosen backbone a chi ec u e. The spa ial
dimensions o he ex ac ed ea u es a e e ec i ely doubled by he op-
down pa h. This con igu a ion acili a es he amalgama ion o ea u es
ha a e low in esolu ion ye ich in seman ic con en wi h hose ha
Table 2
Lis o he u ilized hype pa ame e s.
Upe Ne based on
ViT and Swin
ans o me and
SegFo me
UniFo me Mask2Fo me
Op imize SGD AdamW AdamW
Loss C oss en opy C oss en opy C oss en opy
P e ained
weigh s
AED20K AED20K AED20K
Ini ial
lea ning
a e
0.01 0.0001 0.0001
Momen um 0.9 0.9 0.9
Weigh
decay
0005 0.0001 0.05
Ba ch size 2 2 2
L.R.
schedule
PolyLR (powe =
0.9, begin =1500,
and end =160,000)
Linea LR (powe
=0.9, begin =
1500, and end =
160,000)
PolyLR (powe =
0.9, begin =1500,
and end =160,000)
Fig. 7. (a) T aining ime, (b) loss g aph, (c) mIoU, and (d) mF-sco e.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
8
possess high esolu ion bu compa a i ely weake seman ic s eng h.
The e ec i eness o Upe Ne , u ilizing ViT and Swin ans o me back-
bones, was assessed in his s udy.
The s uc u e o ViT comp ises an embedding laye , a ans o me
encode wi h se e al iden ical laye s, and a head classi ie . Fi s , ViT
spli s an image in o uni o m squa e pa ches (i.e., 16 ×16 o 32 ×32
sizes), wi h each pa ch p ocessed as a oken. These pa ches a e hen
con e ed in o 1D ec o sequences ia a lea nable linea p ojec ion o
dimension E, and posi ional embeddings a e added o e ain hei
loca ion in o ma ion. This sequence is subsequen ly inpu ed in o he
ans o me encode . E e y ans o me encode con ains a mul i-head
sel -a en ion (MSA) block and a dense eed o wa d block. A esidual
connec ion suppo s each sublaye , ollowed by laye no maliza ion o
he ou pu . The MSA block comp ises ou componen s: an ini ial linea
laye , a sel -a en ion laye , a conca ena ion laye ha me ges he ou -
pu s om a ious a en ion heads, and a inal linea laye . Sel -a en ion
in ol es calcula ing a en ion sco es e lec ing he pa ch embeddings’
ela ionships. A posi ion-wise eed o wa d ne wo k is applied o each
pa ch embedding by le e aging he sel -a en ion mechanism. The
ans o med encode ’s ou pu is hen ed in o he classi ica ion head,
which pe o ms he classi ica ion o he inpu image, u ilizing he
ea u e ep esen a ions de i ed om he encoded pa ch embeddings.
2.5.2. Upe Ne -Swin ans o me
The Swin ans o me , a hie a chical ans o me s uc u ed in
mul iple s ages, u ilizes a “shi ed window” echnique o de i a e i s
ep esen a ions. This app oach allows he ex ac ion o ea u es a
a ious le els, he eby e icien ly cap u ing he ex ensi e dependencies
inhe en in he da ase (Liu e al. 2021). The “shi ed windowing”
echnique enhances compu a ional e iciency by con ining he sel -
a en ion mechanism o dis inc , non-o e lapping local windows. This
me hod p omo es he es ablishmen o in e connec ions be ween hese
windows. The Swin ans o me ’s ou -s age design encompasses
a ious p ocesses, including mul iscale ea u e map ex ac ion, pa ch
pa i ioning, me ging o pa ches, linea embedding, and implemen ing
Swin ans o me blocks. The pa ch pa i ion module ini ially di ides
he o iginal image in o dis inc , non-o e lapping pa ches, wi h each
pa ch being p ocessed as an indi idual oken. These okens combine aw
pixel da a om mul iple channels o o m hei ea u e ep esen a ion.
This aw ea u e is hen ans o med h ough a linea embedding laye
in o a ea u e ec o o a speci ied dimension (C). Finally, hese pa ch
okens unde go p ocessing by a se ies o ans o me blocks, aiding he
de elopmen o ea u e ep esen a ions. The Swin ans o me block
comp ises h ee main componen s: window MSA, shi ed window MSA,
and a mul ilaye pe cep on (MLP). A e each s age, he a chi ec u e
uses a pa ch me ging laye o educe he oken coun and c ea e a hi-
e a chical ep esen a ion. This model’s “ iny” a ian was chosen as he
ounda ional a chi ec u e o he Upe Ne amewo k, which consis s o
{2, 2, 6, 2} laye s ac oss i s ou s ages, and i ope a es wi h a ea u e
dimension (C) o 96. The eade s may e e o Liu e al. (2021) o
addi ional de ails abou he design and a ia ions o he Swin
ans o me .
2.5.3. Mask2Fo me
The Mask2Fo me (Cheng e al., 2022) is a uni e sal a chi ec u e o
e sa ile asks, such as seman ic segmen a ion, panop ic segmen a ion,
and ins ance segmen a ion. Mask2Fo me is cons uc ed using a simple
me a-a chi ec u e encompassing a backbone ne wo k, a pixel decode ,
and a ans o me decode . The backbone’s a chi ec u e can be designed
using ei he CNN (i.e., esidual lea ning ne wo ks) o ans o me -based
amewo ks. In his wo k, he adop ed backbone ne wo k was based on
he Swin ans o me (Liu e al., 2021), conside ing i s e iciency in
cap u ing global and local ea u es. The pixel decode o Mask2Fo me
u ilizes a mul iscale de o mable a en ion ans o me (MSDe o mA n)
(Zhu e al. 2020) o e ec i ely inco po a e low- and high- esolu ion
ea u es, op imizing he compu a ional e iciency. The ans o me
decode in Mask2 o me uses a masked a en ion mechanism, ocusing
en i ely on local ea u es su ounding he p edic ed segmen s a he
han p ocessing he en i e ea u e map. In his s udy, he “ iny” e sion
o he Swin ans o me was selec ed o se e as he backbone a chi-
ec u e o he Mask2 o me . Reade s can e e o he wo k by Cheng
e al. (2022) o mo e insigh s in o he design and a ia ions o he
Mask2Fo me .
2.5.4. SegFo me
SegFo me (Xie e al., 2021), which is a obus and e ec i e
ans o me -based seman ic segmen a ion a chi ec u e, has demon-
s a ed ema kable pe o mance in a ious emo e sensing applica ions
(Tang e al., 2022; Gib il e al., 2023; Gonçal es e al., 2023; Jiang e al.,
2023). SegFo me is designed using encode –decode design, uniquely
in eg a ing T ans o me s wi h a ligh weigh MLP-based decode . The
encode ex ac s mul iscale ea u es using ou dis inc ans o me
blocks. Each block en ails h ee in eg al modules: an e icien sel -
a en ion mechanism, a mixed eed o wa d ne wo k (FFN) (known as
Mix-FFN), and me ging blocks o o e lapping pa ches. Con a y o he
con en ional ViT ha u ilizes ixed esolu ion posi ion encodings (P.E.s)
o posi ional in o ma ion in eg a ion, he SegFo me adop s con olu-
ional laye s in i s FFN o a da a-dependen app oach o posi ional
encoding. The decode uses mul iscale in o ma ion, encompassing local
and global de ails, o accu a ely p edic he inal segmen a ion ou -
comes. This s udy selec ed a mix ans o me (MiT) encode based on B3
(MiT-B3). Fu he in o ma ion can be ound in he s udy by Xie e al.
(2021).
2.5.5. UniFo me
The uni ied ans o me (Li e al., 2022) le e ages he capabili ies o
CNNs and ans o me s o mi iga e each app oach’s limi a ions,
achie ing an op imal balance be ween compu a ional e iciency and
accu acy. A basic ans o me o ma (Vaswani e al., 2017) is u ilized
and cus omized o e icien and e ec i e lea ning o spa io empo al
ep esen a ions. The UniFo me block consis s o h ee main modules:
dynamic posi ion embedding (DPE), mul i-head ela ion agg ega o
Table 3
Quan i a i e e alua ion o he ans o me -based seman ic segmen a ion
models.
Me ics RGB
(%)
RGB þ
NIR1 þ
NIR2
(%)
Eigh MS
bands
(%)
Eigh MS
bands þ
NDVI (%)
SegFo me -MiT-
B3
mAcc 85.29 85.69 86.89 85.60
mIoU 76.16 76.54 76.62 76.47
mP ecision 83.92 84.16 83.27 84.15
mRecall 85.29 85.69 86.89 85.60
mF-sco e 84.59 84.91 84.98 84.86
Upe Ne -ViT mAcc 83.41 74.29 78.06 78.13
mIoU 75.04 68.99 71.57 71.64
mP ecision 83.87 83.09 83.46 83.51
mRecall 83.41 74.29 78.06 78.13
mF-sco e 83.64 77.98 80.51 80.57
Upe Ne -Swin-
iny
mAcc 85.84 86.81 86.93 85.64
mIoU 76.12 76.78 78.10 76.65
mP ecision 83.38 83.57 85.46 84.38
mRecall 85.84 86.81 86.93 85.64
mF-sco e 84.56 85.12 86.18 85.0
Mask2Fo me -
Swin- iny
mAcc 85.11 84.23 87.54 85.41
mIoU 76.21 75.28 77.36 75.73
mP ecision 84.17 83.48 83.85 83.14
mRecall 85.11 84.23 87.54 85.41
mF-sco e 84.63 83.85 85.59 84.24
UniFo me mAcc 83.58 86.86 88.30 86.33
mIoU 75.71 77.51 77.88 77.34
mP ecision 84.86 84.62 84.00 84.83
mRecall 83.58 86.86 88.3 86.33
mF-sco e 84.21 85.70 86.01 85.56
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
9
(MHRA), and eed o wa d ne wo k (FFN). The DPE ac i ely in-
co po a es 3D posi ional da a in o each oken, e icien ly le e aging he
spa io empo al sequence o he okens o ideo modeling pu poses. The
MHRA combines each oken wi h i s con ex ual coun e pa s, which
handles local edundancy and global dependency in ideos h ough i s
adap able app oach o oken a ini y lea ning ac oss shallow and deep
laye s. Al hough he local MHRA in he shallow laye s signi ican ly
lessens he compu a ional bu den, he global MHRA in he deepe laye s
ocuses on comp ehensi ely lea ning global oken ela ionships. The
eade may e e o Li e al. (2022) o he de ailed in o ma ion abou he
UniFo me a chi ec u e.
2.6. Accu acy me ics
This s udy e alua ed he e icacy o a ious seman ic segmen a ion
a chi ec u es h ough mul iple accu acy me ics. These me ics de e -
mine he cong uence be ween he e e ence ee c owns and he p e-
dic ed delinea ed ee c owns by ViTs. The combina ion o he a ious
seman ic segmen a ion e alua ion me ics p o ides a ho ough unde -
s anding o he esul s o da e palm ee mapping, highligh ing he e o s
and limi a ions o he di e en models. The u ilized me ics include
mIoU, p ecision, ecall, and mF-sco e and a e de ailed in Equa ions (1)–
(6). The IoU and F-sco e a e commonly u ilized me ics in ITC s udies o
assess he le el o ag eemen be ween deep lea ning model esul s and
g ound- u h da a. These me ics ange om ze o o one o hund ed,
Fig. 8. Cu a ed selec ion o se en mul ispec al images (a–g) andomly chosen om he es ing da a, accompanied by hei co esponding e e ence da a and he
expe imen al esul s.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
16
managemen and moni o ing da e palm ees.
CRediT au ho ship con ibu ion s a emen
Rami Al-Ruzouq: W i ing – e iew & edi ing, W i ing – o iginal
d a , Supe ision, P ojec adminis a ion, Me hodology, In es iga ion,
Fo mal analysis, Concep ualiza ion. Mohamed Ba aka A. Gib il:
W i ing – e iew & edi ing, W i ing – o iginal d a , Visualiza ion,
Valida ion, So wa e, Me hodology, Fo mal analysis, Da a cu a ion,
Concep ualiza ion. Abdallah Shanableh: W i ing – e iew & edi ing,
Supe ision, Resou ces, Concep ualiza ion. Jan Bolcek: W i ing –
o iginal d a , Valida ion, So wa e, Me hodology, Fo mal analysis.
Fouad Lamgha i: W i ing – e iew & edi ing, Supe ision, P ojec
adminis a ion, Funding acquisi ion, Concep ualiza ion. Neza A alla
Hammou : W i ing – e iew & edi ing, Supe ision, Concep ualiza ion.
Ali El-Keblawy: W i ing – e iew & edi ing, W i ing – o iginal d a ,
Supe ision, Concep ualiza ion. Ra i anjan Jena: W i ing – e iew &
edi ing, Me hodology, Fo mal analysis.
Decla a ion o compe ing in e es
The au ho s decla e ha hey ha e no known compe ing inancial
in e es s o pe sonal ela ionships ha could ha e appea ed o in luence
he wo k epo ed in his pape .
Da a a ailabili y
The like o he code is gi en in he manusc ip
Acknowledgmen s
The au ho s exp ess hei app ecia ion o he Fujai ah Resea ch
Cen e (FRC) o i s inancial suppo (Funded P ojec Numbe : 133049)
and o he Uni e si y o Sha jah o o e ing esea ch acili ies.
Re e ences
“Food and Ag icul u e O ganiza ion.” FAOSTAT [In e ne ]. [accessed 2021 Ma 9].
h p://www. ao.o g/ aos a /en/#da a/QC.
Abozeid, A., Alanazi, R., Elhadad, A., Taloba, A.I., Abd El-Aziz, R.M., 2022. A la ge-scale
da ase and deep lea ning model o de ec ing and coun ing oli e ees in sa elli e
image y. Compu In ell Neu osci. 2022 h ps://doi.o g/10.1155/2022/1549842.
Akca S, Pola N. 2022. Seman ic segmen a ion and quan i ica ion o ees in an o cha d
using UAV o hopho o. Ea h Sci In o ma ics 2022 154 [In e ne ]. [accessed 2022
No 28] 15(4):2265–2274. h ps://doi.o g/10.1007/S12145-022-00871-Y.
Al-Ruzouq, R., Shanableh, A., Ba aka , A., Gib il, M., AL-Mansoo i, S., 2018. Image
segmen a ion pa ame e selec ion and an colony op imiza ion o da e palm ee
de ec ion and mapping om e y-high-spa ial- esolu ion ae ial image y. Remo e
Sens [in e ne ]. 10 (9), 1413. h ps://doi.o g/10.3390/ s10091413.
Ami kolaee, H.A., Shi, M., Membe , S., Mulligan, M., 2023. T eeFo me : a semi-
supe ised ans o me -based amewo k o ee coun ing om a single. IEEE T ans
Geosci Remo e Sens. 61, 1–15. h ps://doi.o g/10.1109/TGRS.2023.3295802.
Amma , A., Koubaa, A., Benjdi a, B., 2021. Deep-lea ning-based au oma ed palm ee
coun ing and geoloca ion in la ge a ms om ae ial geo agged images. Ag onomy
[in e ne ]. 11 (8), 1458. h ps://doi.o g/10.3390/ag onomy11081458.
Ayhan, B., Kwan, C., 2020. T ee, sh ub, and g ass classi ica ion using only RGB images.
Remo e Sens. 12 (8) h ps://doi.o g/10.3390/RS12081333.
Bai, Y., Yu, J., Yang, S., Ning, J., 2024. An imp o ed YOLO algo i hm o de ec ing
lowe s and ui s on s awbe y seedlings. Biosys Eng [in e ne ]. 237(June 2023):
1–12 h ps://doi.o g/10.1016/j.biosys emseng.2023.11.008.
Ball JGC, Hickman SHM, Jackson TD, Koay XJ, Hi s J, Jay W, A che M, Coomes DA.
2023. Accu a e delinea ion o indi idual ee c owns in opical o es s om ae ial
RGB image y using Mask R-CNN. :1–14. h ps://doi.o g/10.1002/ se2.332.
Bha naga , S., Gill, L., Ghosh, B., 2020. D one image segmen a ion using machine and
deep lea ning o mapping aised bog ege a ion communi ies. Remo e Sens. 12 (16)
h ps://doi.o g/10.3390/RS12162602.
B aga, J.R.G., Pe ipa o, V., Dalagnol, R., Fe ei a, M.P., Ta abalka, Y., A ag˜
ao, L.E.O.C.,
de Campos Velho, H.F., Shiguemo i, E.H., Wagne , F.H., 2020. T ee c own
delinea ion algo i hm based on a con olu ional neu al ne wo k. Remo e Sens. 12 (8),
1–27. h ps://doi.o g/10.3390/RS12081288.
Cai C, Xu H, Chen S, Yang L, Weng Y, Huang S, Dong C, Lou X. 2023. T ee Recogni ion
and C own Wid h Ex ac ion Based on No el Fas e -RCNN in a Dense Loblolly Pine
En i onmen .
Cao, K., Zhang, X., 2020. An imp o ed Res-UNe model o ee species classi ica ion
using ai bo ne high- esolu ion images. Remo e Sens. 12 (7) h ps://doi.o g/
10.3390/ s12071128.
Chang, B., Wang, Y., Zhao, X., Li, G., Yuan, P., 2024. A gene al-pu pose edge- ea u e
guidance module o enhance ision ans o me s o plan disease iden i ica ion.
Expe Sys Appl [in e ne ]. 237, 121638 h ps://doi.o g/10.1016/j.
eswa.2023.121638.
Chen, G., Shang, Y., 2022. T ans o me o ee coun ing in ae ial images. Remo e Sens.
14 (3), 476. h ps://doi.o g/10.3390/ s14030476.
Cheng B, Mis a I, Schwing AG, Ki illo A, Gi dha R. 2022. Masked-a en ion Mask
T ans o me o Uni e sal Image Segmen a ion. P oc IEEE Compu Soc Con Compu
Vis Pa e n Recogni . 2022-June:1280–1289. h ps://doi.o g/10.1109/
CVPR52688.2022.00135.
Cheng, Z., Qi, L., Cheng, Y., 2021. Che y ee c own ex ac ion om na u al o cha d
images wi h complex backg ounds. Ag icul u e [in e ne ]. 11 (5), 431. h ps://doi.
o g/10.3390/ag icul u e11050431.
Culman, M., Delalieux, S., Van T ich , K., 2020. Indi idual palm ee de ec ion using
deep lea ning on RGB image y o suppo ee in en o y. Remo e Sens. 12 (21),
1–31. h ps://doi.o g/10.3390/ s12213476.
De sch, S., K zys ek, P., Heu ich, M., Resou ces, N., Fo es , B., Pa k, N., Moni o ing, N.P.,
Technology, I., Sys ems, I., 2022. NOVEL SINGLE TREE DETECTION BY
TRANSFORMERS USING UAV-BASED MULTISPECTRAL IMAGERY, XLIII(June):
6–11.
De sch, S., Sch¨
o l, A., K zys ek, P., Heu ich, M., 2023. Towa ds comple e ee c own
delinea ion by ins ance segmen a ion wi h Mask R-CNN and DETR using UAV-based
mul ispec al image y and lida da a. ISPRS Open J Pho og amm Remo e Sens. 8,
100037.
El-Juhany, L., 2010. Deg ada ion o da e palm ees and da e p oduc ion in A ab
coun ies: causes and po en ial ehabili a ion. Aus J Basic Appl Sci [in e ne ]. 4 (8),
3998–4010. h ps://doi.o g/10.1016/j. c .2011.04.017.
Fan, F., Zeng, X., Wei, S., Zhang, H., Tang, D., Shi, J., Zhang, X., 2022. E icien ins ance
segmen a ion pa adigm o in e p e ing SAR and op ical images. Remo e Sens. 14
(3), 1–22. h ps://doi.o g/10.3390/ s14030531.
Fayad, I., Ciais, P., Schwa z, M., Wigne on, J.P., Baghdadi, N., de T uchis, A.,
d’Asp emon , A., F appa , F., Saa chi, S., Sean, E., e al., 2024. Hy-TeC: a hyb id
ision ans o me model o high- esolu ion and la ge-scale mapping o canopy
heigh . Remo e Sens En i on. 302 (Decembe 2023) h ps://doi.o g/10.1016/j.
se.2023.113945.
Fe ei a MP, Almeida DRA de, Papa D de A, Mine ino JBS, Ve as HFP, Fo mighie i A,
San os CAN, Fe ei a MAD, Figuei edo EO, Fe ei a EJL. 2020. Indi idual ee
de ec ion and species classi ica ion o Amazonian palms using UAV images and deep
lea ning. Fo Ecol Manage [In e ne ]. 475(Ap il):118397. h ps://doi.o g/10.1016/
j. o eco.2020.118397.
Fi oze A, Wing en C, Yeh RA, Benes B, Aliaga D. 2023. T ee Ins ance Segmen a ion Wi h
Tempo al Con ou G aph. In: P oc IEEE/CVF Con Compu Vis Pa e n Recogni .
[place unknown]; p. 2193–2202.
F eudenbe g, M., N¨
olke, N., Agos ini, A., U ban, K., W¨
o g¨
o e , F., Kleinn, C., 2019. La ge
scale palm ee de ec ion in high esolu ion sa elli e images using U-Ne . Remo e
Sens. 11 (3), 1–18. h ps://doi.o g/10.3390/ s11030312.
F eudenbe g, M., Magdon, P., N¨
olke, N., 2022. Indi idual ee c own delinea ion in high-
esolu ion emo e sensing images based on U-Ne . Neu al Compu Appl. 34 (24),
22197–22207. h ps://doi.o g/10.1007/s00521-022-07640-4.
Fu, B., He, X., Liang, Y., Deng, T., Li, H., He, H., Jia, M., Fan, D., Wang, F., 2023.
Examina ion o he pe o mance o ASEL and MPViT algo i hms o classi ying
mang o e species o mul iple na u al ese es o Beibu Gul , sou h China. Ecol Indic.
154 (Augus ) h ps://doi.o g/10.1016/j.ecolind.2023.110870.
Ghasemi, M., La i i, H., Pou hashemi, M., 2022. A no el me hod o de ec ing and
delinea ing coppice ees in UAV images o moni o ee decline. Remo e Sens. 14
(23), 5910.
Gib il MBA, Sha i HZM, Shanableh A, Al-Ruzouq R, Wayayok A, Hashim SJ bin, Sachi
MS. 2022. Deep con olu ional neu al ne wo ks and Swin ans o me -based
amewo ks o indi idual da e palm ee de ec ion and mapping om la ge-scale
UAV images. Geoca o In [In e ne ]. 37(27):18569–18599. h ps://doi.o g/
10.1080/10106049.2022.2142966.
Gib il MBA, Sha i HZM, Shanableh A, Al-Ruzouq R, bin Hashim SJ, Wayayok A, Sachi
MS. 2024. La ge-scale assessmen o da e palm plan a ions based on UAV emo e
sensing and mul iscale ision ans o me . Remo e Sens Appl Soc En i on
[In e ne ].:101195. h ps://doi.o g/h ps://doi.o g/10.1016/j. sase.2024.101195.
Gib il, M.B.A., Sha i, H.Z.M., Shanableh, A., Al-Ruzouq, R., Wayayok, A., Hashim, S.J.,
2021. Deep con olu ional neu al ne wo k o la ge-scale da e palm ee mapping
om ua -based images. Remo e Sens. 13 (14), 1–24. h ps://doi.o g/10.3390/
s13142787.
Gib il, M.B.A., Sha i, H.Z.M., Al-Ruzouq, R., Shanableh, A., Nahas, F., Al, M.S., 2023.
La ge-scale da e palm ee segmen a ion om mul iscale UAV-based and ae ial
images using deep ision ans o me s. D ones. 7 (2) h ps://doi.o g/10.3390/
d ones7020093.
Gonçal es DN, Ma ca o J, Ca ilho AC, Acos a PR, Ramos APM, Gomes FDG, Osco LP, da
Rosa Oli ei a M, Ma ins JAC, Damasceno GA, e al. 2023. T ans o me s o mapping
bu ned a eas in B azilian Pan anal and Amazon wi h Plane Scope image y. In J
Appl Ea h Obs Geoin [In e ne ]. 116(Decembe 2022):103151. h ps://doi.o g/
10.1016/j.jag.2022.103151.
Hao, Z., Lin, L., Pos , C.J., Mikhailo a, E.A., Yu, K., 2023. The co-e ec o image
esolu ion and c own size on deep lea ning o indi idual ee de ec ion and
delinea ion. In J Digi Ea h [In e ne ].:3753–3771. h ps://doi.o g/10.1080/
17538947.2023.2257636.
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
17
Ha ling, S., Sagan, V., Sidike, P., Maimai ijiang, M., Ca on, J., 2019. U ban ee species
classi ica ion using a wo ld iew-2/3 and liDAR da a usion app oach and deep
lea ning. Senso s (swi ze land). 19 (6), 1–23. h ps://doi.o g/10.3390/s19061284.
He, H., Zhou, F., Xia, Y., Chen, M., Chen, T., 2024. Pa allel usion neu al ne wo k
conside ing local and global seman ic in o ma ion o ci us ee canopy
segmen a ion. IEEE J Sel Top Appl Ea h Obs Remo e Sens. 17, 1535–1549. h ps://
doi.o g/10.1109/JSTARS.2023.3339290.
Jamali, A., Roy, S.K., Hashemi Beni, L., P adhan, B., Li, J., Ghamisi, P., 2024. Residual
wa e ision U-Ne o lood mapping using dual pola iza ion Sen inel-1 SAR
image y. In J Appl Ea h Obs Geoin [in e ne ]. 127, 103662 h ps://doi.o g/
10.1016/j.jag.2024.103662.
Ji, Y., Yan, E., Yin, X., Song, Y., Wei, W., Mo, D., 2022. Au oma ed ex ac ion o Camellia
olei e a c own using unmanned ae ial ehicle isible images and he ResU-Ne deep
lea ning model. F on Plan Sci. 13 (Augus ), 1–13. h ps://doi.o g/10.3389/
pls.2022.958940.
Jiang, K., A zaal, U., Lee, J., 2023. T ans o me -based weed segmen a ion o g ass
managemen . Senso s. 23 (1), 1–15. h ps://doi.o g/10.3390/s23010065.
Jin asu isak T, Edi isinghe E, Elba ay A. 2022. Deep neu al ne wo k based da e palm
ee de ec ion in d one image y. Compu Elec on Ag ic [In e ne ]. 192(Ap il 2021):
106560. h ps://doi.o g/10.1016/j.compag.2021.106560.
Ka enbo n, T., Lei lo , J., Schie e , F., Hinz, S., 2021. Re iew on con olu ional neu al
ne wo ks (CNN) in ege a ion emo e sensing. ISPRS J Pho og amm Remo e Sens.
173, 24–49.
Ken sch, S., Cace es, M.L.L., Se ano, D., Rou e, F., Diez, Y., 2020a. Compu e ision and
deep lea ning echniques o he analysis o d one-acqui ed o es images, a ans e
lea ning s udy. Remo e Sens. 12 (8), 1–19. h ps://doi.o g/10.3390/RS12081287.
Ken sch S, Ka a siolis S, Kamila is A, Tomha e L, Lopez Cace es ML. 2020. Iden i ica ion
o T ee Species in Japanese Fo es s based on Ae ial Pho og aphy and Deep Lea ning.
a Xi . h ps://doi.o g/10.1007/978-3-030-61969-5_18.
Kolesniko , A., Doso i skiy, A., Weissenbo n, D., Heigold, G., Uszko ei , J., Beye , L.,
Minde e , M., Dehghani, M., Houlsby, N., Gelly, S., e al., 2021. An image is wo h
16x16 wo ds: T ans o me s o image ecogni ion a scale. In: [place Unknown].
Li, R., Chen, T., Liu, Y., Jiang, H., 2023. CoupleUNe : Swin T ans o me coupling CNNs
makes s ong con ex ual encode s o VHR image oad ex ac ion. In J Remo e Sens.
44 (18), 5788–5813.
Li K, Wang Y, Gao P, Song G, Liu Y, Li H, Qiao Y. 2022. Uni o me : Uni ied ans o me
o e icien spa io empo al ep esen a ion lea ning. a Xi P ep a Xi 220104676.
Li, Y., Ma, L., Sun, N., 2024. A bilinea ans o me in e ac i e neu al ne wo ks-based
app oach o ine-g ained ecogni ion and p o ec ion o plan diseases o ga dening
design. C op P o [in e ne ]. 180, 106660 h ps://doi.o g/10.1016/j.
c op o.2024.106660.
Lin, T.-Y., Doll´
a , P., Gi shick, R., He, K., Ha iha an, B., Belongie, S., 2017. Fea u e
py amid ne wo ks o objec de ec ion. In: P oc IEEE Con Compu is Pa e n
Recogni . Honolulu, HI, USA, pp. 2117–2125.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B. 2021. Swin T ans o me :
Hie a chical Vision T ans o me using Shi ed Windows. In: 2021 IEEE/CVF In Con
Compu Vis. Mon eal, QC, Canada: IEEE; p. 9992–10002. h ps://doi.o g/10.1109/
ICCV48922.2021.00986.
Liu, J., Wang, X., Wang, T., 2019. Classi ica ion o ee species and s ock olume
es ima ion in g ound o es images using Deep Lea ning. Compu Elec on Ag ic
[in e ne ]. 166 (May), 105012 h ps://doi.o g/10.1016/j.compag.2019.105012.
Liu, H., Wang, X., Zhao, F., Yu, F., Lin, P., Gan, Y., Ren, X., Chen, Y., Tu, J., 2024.
Upg ading swin-B ans o me -based model o accu a ely iden i ying ipe
s awbe ies by coupling ask-aligned one-s age objec de ec ion mechanism.
Compu Elec on Ag ic [in e ne ]. 218, 108674 h ps://doi.o g/10.1016/j.
compag.2024.108674.
Mekhal i, M.L., Nicolo, C., Bazi, Y., Al, R.MM., Alsha i , N.A., Al, M.E., 2022. Con as ing
YOLO 5, ans o me , and e icien de de ec o s o c op ci cle de ec ion in dese .
IEEE Geosci Remo e Sens Le . 19, 19–23. h ps://doi.o g/10.1109/
LGRS.2021.3085139.
Con ibu o s Mms. 2020. {MMSegmen a ion}: OpenMMLab Seman ic Segmen a ion
Toolbox and Benchma k.
Ozda ici-ok, A., Ok, A.O., 2023. Scien ia Ho icul u ae Using emo e sensing o iden i y
indi idual ee species in o cha ds : Sci Ho ic (Ams e dam). [in e ne ]. 321 (10),
112333 h ps://doi.o g/10.1016/j.scien a.2023.112333.
Paszke, A., G oss, S., Massa, F., Le e , A., B adbu y, J., Chanan, G., Killeen, T., Lin, Z.,
Gimelshein, N., An iga, L., e al., 2019. Py o ch: An impe a i e s yle, high-
pe o mance deep lea ning lib a y. Ad Neu al In P ocess Sys . 32.
Qin, H., Zhou, W., Yao, Y., Wang, W., 2022. Indi idual ee segmen a ion and ee
species classi ica ion in sub opical b oadlea o es s using UAV-based LiDAR,
hype spec al, and ul ahigh- esolu ion RGB da a. Remo e Sens En i on. 280,
113143.
Rezaei, M., Diepe een, D., Laga, H., Jones, M.G.K., Sohel, F., 2024. Plan disease
ecogni ion in a low da a scena io using ew-sho lea ning. Compu Elec on Ag ic
[in e ne ]. 219, 108812 h ps://doi.o g/10.1016/j.compag.2024.108812.
Schie e , F., Ka enbo n, T., F ick, A., F ey, J., Schall, P., Koch, B., Schmid lein, S., 2020.
Mapping o es ee species in high esolu ion UAV-based RGB-image y by means o
con olu ional neu al ne wo ks. ISPRS J Pho og amm Remo e Sens [in e ne ]. 170
(No embe ), 205–215. h ps://doi.o g/10.1016/j.isp sjp s.2020.10.015.
Sel a aju RR, Cogswell M, Das A, Vedan am R, Pa ikh D, Ba a D. 2017. G ad-cam:
Visual explana ions om deep ne wo ks ia g adien -based localiza ion. In: P oc
IEEE In Con Compu Vis. [place unknown]; p. 618–626.
Sun, Y., Li, Z., He, H., Guo, L., Zhang, X., Xin, Q., 2022. Coun ing ees in a sub opical
mega ci y using he ins ance segmen a ion me hod. In J Appl Ea h Obs Geoin . 106,
102662 h ps://doi.o g/10.1016/j.jag.2021.102662.
Sun, Y., Hao, Z., Guo, Z., Liu, Z., Jh., 2023. De ec ion and mapping o ches nu using
deep lea ning om high- esolu ion UAV-Based RGB image . Remo e Sens.:1–18.
Suzuki, S., 1985. Topological s uc u al analysis o digi ized bina y images by bo de
ollowing. Compu Vision, G aph Image P ocess. 30 (1), 32–46.
Tang, X., Tu, Z., Wang, Y., Liu, M., Li, D., Fan, X., 2022. Au oma ic de ec ion o coseismic
landslides using a new ans o me me hod. Remo e Sens. 14 (12), 1–19. h ps://doi.
o g/10.3390/ s14122884.
Tolan J, Yang HI, Nosa zewski B, Couai on G, Vo H V., B and J, Spo e J, Majumda S,
Haziza D, Vama aju J, e al. 2024. Ve y high esolu ion canopy heigh maps om
RGB image y using sel -supe ised ision ans o me and con olu ional decode
ained on ae ial lida . Remo e Sens En i on [In e ne ]. 300(Ap il 2023):113888.
h ps://doi.o g/10.1016/j. se.2023.113888.
To es, D.L., Fei osa, R.Q., Happ, P.N., La Rosa, L.E.C., Junio , J.M., Ma ins, J.,
B essan, P.O., Gonçal es, W.N., Liesenbe g, V., 2020. Applying ully con olu ional
a chi ec u es o seman ic segmen a ion o a single ee species in u ban
en i onmen on high esolu ion UAV op ical image y. Senso s (swi ze land). 20 (2),
1–20. h ps://doi.o g/10.3390/s20020563.
Ulku, I., Akagündüz, E., Ghamisi, P., Membe , S., 2022. Deep Seman ic Segmen a ion o
T ees Using Mul ispec al Images. 15, 7589–7604. h ps://doi.o g/10.1109/
JSTARS.2022.3203145.
Vaswani A, Shazee N, Pa ma N, Uszko ei J, Jones L, Gomez AN, Kaise Ł, Polosukhin I.
2017. A en ion is all you need. In: Ad Neu al In P ocess Sys . Vol. 30. [place
unknown]; p. 5998–6008.
Velasquez-camacho, L., E xega ai, M., 2023. Deep Lea ning algo i hms o u ban ee
de ec ion and geoloca ion wi h high- esolu ion ae ial, sa elli e, and g ound-le el
images. Compu En i on U ban Sys [in e ne ]. 105 (Augus ), 102025 h ps://doi.
o g/10.1016/j.compen u bsys.2023.102025.
Wagne , F.H., Sanchez, A., Ta abalka, Y., Lo e, R.G., Fe ei a, M.P., Aida , M.P.M.,
Gloo , E., Phillips, O.L., A ag˜
ao, L.E.O.C., 2019. Using he U-ne con olu ional
ne wo k o map o es ypes and dis u bance in he A lan ic ain o es wi h e y high
esolu ion images. Remo e Sens Ecol Conse . 5 (4), 360–375. h ps://doi.o g/
10.1002/ se2.111.
Wagne , F.H., Sanchez, A., Aida , M.P.M., Rochelle, A.L.C., Ta abalka, Y., Fonseca, M.G.,
Phillips, O.L., Gloo , E., A ag˜
ao, L.E.O.C., 2020. Mapping A lan ic ain o es
deg ada ion and egene a ion his o y wi h indica o species using con olu ional
ne wo k. PLoS One. 15 (2), e0229448.
Wang W, Xie E, Li X, Fan D-P, Song K, Liang D, Lu T, Luo P, Shao L. 2021. Py amid Vision
T ans o me : A Ve sa ile Backbone o Dense P edic ion wi hou Con olu ions
[In e ne ]. :568–578. h p://a xi .o g/abs/2102.12122.
Wu, Y., Mis a, S., 2020. In elligen image segmen a ion o o ganic- ich shales using
andom o es , wa ele ans o m, and hessian ma ix. IEEE Geosci Remo e Sens Le .
17 (7), 1144–1147. h ps://doi.o g/10.1109/LGRS.2019.2943849.
Xia Z, Pan X, Song S, Li LE, Huang G. 2022. Vision T ans o me wi h De o mable
A en ion [In e ne ]. h p://a xi .o g/abs/2201.00520.
Xiao T, Liu Y, Zhou B, Jiang Y, Sun J. 2018. Uni ied Pe cep ual Pa sing o Scene
Unde s anding. Lec No es Compu Sci (including Subse Lec No es A i In ell Lec
No es Bioin o ma ics). 11209 LNCS:432–448. h ps://doi.o g/10.1007/978-3-030-
01228-1_26.
Xiao, C., Qin, R., Huang, X., 2020. T ee op de ec ion using con olu ional neu al
ne wo ks ained h ough au oma ically gene a ed pseudo labels. In J Remo e Sens
[in e ne ]. 41 (8), 3010–3030. h ps://doi.o g/10.1080/01431161.2019.1698075.
Xie, E., Wang, W., Yu, Z., Anandkuma , A., Al a ez, J.M., Luo, P., 2021. SegFo me :
simple and e icien design o seman ic segmen a ion wi h ans o me s. Ad Neu al
In P ocess Sys . 15, 12077–12090.
Yang, M., Mou, Y., Liu, S., Meng, Y., Liu, Z., Li, P., Xiang, W., Zhou, X., Peng, C., 2022b.
De ec ing and mapping ee c owns based on con olu ional neu al ne wo k and
Google Ea h images. In J Appl Ea h Obs Geoin . 108(Augus , 2021). h ps://doi.
o g/10.1016/j.jag.2022.102764.
Yang, L., Wang, X., Zhai, J., 2022a. Wa e line ex ac ion o a i icial coas wi h ision
ans o me s. F on En i on Sci. 10 (Feb ua y), 1–12. h ps://doi.o g/10.3389/
en s.2022.799250.
Yi, S., Liu, X., Li, J., Chen, L., 2023. UAV o me : a composi e ans o me ne wo k o
u ban scene segmen a ion o UAV images. Pa e n Recogni . 133 h ps://doi.o g/
10.1016/j.pa cog.2022.109019.
Zhang, T., Hu, D., Wu, C., Liu, Y., Yang, J., Tang, K., 2023. La ge-scale apple o cha d
mapping om mul i-sou ce da a using he seman ic segmen a ion model wi h image-
o-image ansla ion and ans e lea ning. Compu Elec on Ag ic. 213, 108204.
Zhang, L., Lin, H., Wang, F., 2022. Indi idual ee de ec ion based on high- esolu ion
RGB images o u ban o es y applica ions. IEEE Access. 10, 46589–46598. h ps://
doi.o g/10.1109/ACCESS.2022.3171585.
Zhang, H., Liu, S., 2024. Double-b anch mul i-scale con ex ual ne wo k: a model o
mul i-scale s ee ee segmen a ion in high- esolu ion emo e sensing images.
Senso s. 24 (4), 1110. h ps://doi.o g/10.3390/s24041110.
Zhao, J., Be ge, T.W., Geipel, J., 2023b. T ans o me in UAV image-based weed
mapping. Remo e Sens. 15 (21), 5165.
Zhao H, Shi J, Qi X, Wang X, Jia J. 2017. Py amid scene pa sing ne wo k. In: P oc - 30 h
IEEE Con Compu Vis Pa e n Recogni ion, CVPR 2017. Vol. 2017-Janua. [place
unknown]; p. 6230–6239. h ps://doi.o g/10.1109/CVPR.2017.660.
Zhao, H., Mo gen o h, J., Pea se, G., Schindle , J., 2023a. A Sys ema ic e iew o
indi idual ee c own de ec ion and delinea ion wi h con olu ional neu al ne wo ks
(CNN). Cu o Repo s [in e ne ]. 9 (3), 149–170. h ps://doi.o g/10.1007/
s40725-023-00184-3.
Zheng, J., Wu, W., Yu, L., Fu, H., 2021. COCONUT TREES DETECTION ON THE
TENARUNGA USING HIGH-RESOLUTION SATELLITE IMAGES AND DEEP
LEARNING Minis y o Educa ion Key Labo a o y o Ea h Sys em Modeling, and
R. Al-Ruzouq e al.
Ecological Indica o s 163 (2024) 112110
18
Depa men o Ea h Sys em Science, Tsinghua Uni e si y, Beijing 100084, China
Na ion. In Geosci Remo e Sens. Symp.:6512–6515.
Zheng, Y., Wu, G., 2022. YOLO 4-Li e – Based U ban Plan a ion T ee De ec ion and
Posi ioning wi h High-Resolu ion Remo e Sensing Image y. 9(Janua y):1–12
h ps://doi.o g/10.3389/ en s.2021.756227.
Zheng, J., Yuan, S., Wu, W., Li, W., Yu, L., Fu, H., Coomes, D., 2023. Su eying coconu
ees using high- esolu ion sa elli e image y in emo e a olls o he Paci ic Ocean.
Remo e Sens En i on. 287 h ps://doi.o g/10.1016/j. se.2023.113485.
Zhou, C., Ye, H., Sun, D., Yue, J., Yang, G., Hu, J., 2022. An au oma ed, high-
pe o mance app oach o de ec ing and cha ac e izing b occoli based on UAV
emo e-sensing and ans o me s: A case s udy om Haining, China. In J Appl Ea h
Obs Geoin [in e ne ]. 114 (July), 103055 h ps://doi.o g/10.1016/j.
jag.2022.103055.
Zhou, B., Zhao, H., Puig, X., Fidle , S., Ba iuso, A., To alba, A., 2017. Scene Pa sing
h ough ADE20K Da ase . In: In: P oc IEEE Con Compu is Pa e n Recogni . [place
Unknown], p. p. page 4..
Zhu X, Su W, Lu L, Li B, Wang X, Dai J. 2020. De o mable DETR: De o mable
T ans o me s o End- o-End Objec De ec ion [In e ne ]. :1–16. h p://a xi .o g/
abs/2010.04159.
R. Al-Ruzouq e al.