Recei ed 5 June 2025, accep ed 26 June 2025, da e o publica ion 21 July 2025, da e o cu en e sion 29 July 2025.
Digi al Objec Iden i ie 10.1109/ACCESS.2025.3591272
Compa a i e Analysis o Deep Lea ning-Based
Fea u e Ex ac o s o Change De ec ion
in Au omo i e Rada Maps
HARIHARA BHARATHY SWAMINATHAN 1, ARON SOMMER2, URI IURGEL2,
ANDREAS BECKER 3, (Membe , IEEE), AND MARTIN ATZMUELLER 1,4
1Seman ic In o ma ion Sys ems G oup, Osnab ück Uni e si y, 49074 Osnab ück, Ge many
2Ap i Se ices Deu schland GmbH, 42119 Wuppe al, Ge many
3Facul y o In o ma ion Technology, Fachhochschule Do mund, 44139 Do mund, Ge many
4Ge man Resea ch Cen e o A i icial In elligence (DFKI), 49084 Osnab ück, Ge many
Co esponding au ho : Ha iha a Bha a hy Swamina han (hsw[email p o ec ed])
ABSTRACT The Siamese ne wo k a chi ec u e has been applied by deep lea ning p ac i ione s o
ind simila i ies be ween images. In he domain o au onomous d i ing, his ne wo k con igu a ion has
ecen ly gained a en ion o sol ing he change de ec ion ask, which in ol es iden i ying changes in
a p e iously known map o a ehicle’s en i onmen . This is i al, as such de ia ions may comp omise
he accu acy and eliabili y o he map, which is essen ial o he ehicle’s abili y o localize i sel and
na iga e e ec i ely. In his pape , we p esen a se o expe imen s in ol ing s a e-o - he-a deep lea ning
a chi ec u es based on bo h con olu ion (CNN) and a en ion mechanisms such as AlexNe , GoogLeNe ,
VGG, ResNe , Vision T ans o me , and Shi ed Windows T ans o me as possible candida es o he ea u e
ex ac o backbone module in he Siamese a chi ec u e o de ec changes caused by he disappea ance and
appea ance o cons uc ion zones. Also, we e alua e he pe o mance o hese a chi ec u es using ine-
uning, i. e., ini ializing he con olu ional laye s wi h p e- ained weigh s. In ou expe imen a ion, he bes
esul s we e ob ained using VGG16 (CNN), especially when i was ini ialized using p e- ained weigh s
om he ImageNe -1K da ase . In pa icula , VGG16 wi h an a e age F1 sco e o 92% on highway da ase s
ou pe o med he baseline esidual ne wo k composed o ResNe 18 con olu ions by abou 13.5%.
INDEX TERMS Change de ec ion, au omo i e ada , occupancy maps, siamese ne wo ks.
I. INTRODUCTION
The pe o mance o s a e-o - he-a deep neu al ne wo ks
has been e alua ed in he pas on he basis o hei sco es
in he classi ica ion o benchma k da ase s such as he
ImageNe . In his pape , we ocus on ea u e ex ac o s
in a speci ic domain, i. e., au omo i e ada . In pa icula ,
we e alua e he pe o mance o ea u e ex ac o s such
as AlexNe , GoogLeNe , VGG, ResNe , Vision T ans-
o me (ViT) and Shi ed Windows (SWIN) T ans o me
o he pu pose o change de ec ion on au omo i e ada -
based maps. In his wo k, he ask is amed as a
bina y classi ica ion p oblem, in which he model p edic s
The associa e edi o coo dina ing he e iew o his manusc ip and
app o ing i o publica ion was Fab izio San i .
whe he he e e ence map has di e ged om he cu en
en i onmen . These p edic ions help assess he accu acy
and ele ance o he e e ence map and indica e when e-
mapping is necessa y due o de ec ed changes in he ehicle’s
en i onmen .
As ou main con ibu ions, we p esen esul s on how
o de ec changes caused by he disappea ance and appea -
ance o cons uc ion zones. Fu he mo e, we e alua e he
pe o mance o hese ea u e ex ac o s a e ine- uning,
i. e., ini ializing he con olu ional laye s wi h p e- ained
weigh s. Fo he pu pose o change de ec ion, we use a
Siamese ne wo k con igu a ion, which consis s o wo deep
neu al ne wo k (DNN) sub-ne wo ks as backbones o ea u e
ex ac ion om an image pai , ollowed by a decision head o
image simila i y lea ning.
VOLUME 13, 2025
2025 The Au ho s. This wo k is licensed unde a C ea i e Commons A ibu ion 4.0 License.
Fo mo e in o ma ion, see h ps://c ea i ecommons.o g/licenses/by/4.0/ 130629
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
In ou expe imen s, we ocus on he domain o au omo i e
ada , o which ealis ic open benchma k da ase s a e a he
sca ce and need o be ailo ed o he speci ic applica ion
domain. The e o e, as a da ase , we use eal-li e au omo i e
ada map image pai s gene a ed o e a one-yea pe iod
a ound he Ge man ci y o Wuppe al, pa icula ly on
highways wi h cons uc ion zones.
Ou main indings a e summa ized as ollows:
1) Siamese ne wo ks con igu ed wi h VGG- amily o
ea u e ex ac o s ou pe o med he baseline con-
igu a ion wi h a esidual ne wo k i. e., ResNe 18.
Speci ically, he VGG16 backbone achie ed an F1
sco e inc ease o abou 13.5% on a e age on es
da ase s depic ing disappea ing and appea ing highway
cons uc ion zones.
2) By ini ializing he con olu ion laye s in VGG using
p e- ained weigh s ob ained om classi ica ion o
ImageNe -1K da ase , we o e all achie ed as e con-
e gence h ough he p ocess o ine- uning.
The es o he pape is s uc u ed as ollows: Sec ion II
discusses ela ed wo k. A e ha , Sec ion III p esen s he
change de ec ion da ase along wi h he applied e alua ion
me ics. Nex , Sec ion IV and Sec ion Vdiscusses ou
expe imen a ion and esul s in de ail, espec i ely. Finally,
Sec ion VI concludes wi h a summa y and in e es ing
di ec ions o u u e wo k.
II. RELATED WORK
In he pas , Siamese ne wo ks [1] ha e been ex ensi ely
s udied o a ious use cases, anging om signa u e
e i ica ion [2], ace e i ica ion [3] o change de ec ion in
emo e sensing applica ions [4].
Recen ly, Swamina han e al. ha e p oposed a Siamese
ne wo k-based classi ie o de ec ing changes due o he
appea ance and disappea ance o cons uc ion zones on ada
maps o he en i onmen [5]. As backbones, hey used
he con olu ion laye s o ResNe -18 as a ea u e ex ac o .
As an inpu o he mul i-laye pe cep on (MLP) ne wo k
o he decision module, ea u e ec o s om he CNN
backbones we e conca ena ed. The in e media e laye s o he
MLP had a size o 256 and 1, subsequen ly ollowed by a
sigmoid ac i a ion a he end, ained using a bina y c oss-
en opy (BCE) loss unc ion. This con igu a ion will se e as
he baseline o ou wo k and is shown in Figu e 1.
FIGURE 1. The baseline siamese ne wo k a chi ec u e composed o a
ResNe -18 ea u e ex ac o and MLP ne wo k as a decision head.
O e all, o he bes o he au ho s’ knowledge, a compa a-
i e analysis o exis ing s a e-o - he-a deep neu al ne wo k
a chi ec u es as candida es o a ea u e ex ac ion backbone
in a Siamese ne wo k con igu a ion has no been ca ied ou
in he li e a u e be o e.
III. DATASET AND EVALUATION METRICS
A. CHANGE DETECTION DATASET
Fo he pu pose o aining and es ing, we use he
change de ec ion da ase [5] p esen ed p e iously by
Swamina han e al. The da ase includes ada maps show-
ing changes such as he appea ance and disappea ance
o cons uc ion zones. These maps we e collec ed o e a yea
a ound Wuppe al, Ge many, wi h a ocus on a highway
connec ing Wuppe al and Düsseldo , as shown in Figu e 2.
A nega i e sample (Figu e 2a) shows only i ele an o mino
changes (ma ked wi h o als), while posi i e samples show
cons uc ion zones appea ing o disappea ing, along wi h
some un ela ed changes, c . Figu es 2b and 2c. Du ing da a
collec ion, ce ain lanes we e closed due o cons uc ion, and
he es ehicle was d i en h ough a nea by lane on he
highway.
These maps we e gene a ed using measu emen s o co ne
ada senso s i ed o he es ehicle. Fu he mo e, a high
p ecision DGPS measu emen uni i.e. he POS LV-220
posi ioning sys em om Applanix co po a ion was used o
ob ain p ecise posi ioning o he ehicle. Each sample o his
da ase consis s o pe ec ly aligned bi d’s eye- iew (BEV)
ada map image pai s ob ained om di e en measu emen
d i es o he same en i onmen . Wi hin a speci ic da a
sample, e ms such as a ‘‘ e e ence map’’ and a ‘‘cu en
map’’ a e used o deno e images ha we e cap u ed a
di e en poin s in ime.
We di ide he aining and es ing samples, ensu ing
ha a cons uc ion zone shown du ing he aining p ocess
is no pa o he es ing one. Fo ou expe imen s, he
image pai s comp ising he appea ance and disappea ance
o cons uc ion zones in he ou e owa ds Düsseldo a e
conside ed as pa o he aining da ase and he image pai s
comp ising o appea ing (A) and disappea ing cons uc ion
zones (D) in he ou e owa ds Wuppe al as pa o
he es ing da ase . In e ms o size and dis ibu ion, he
aining se con ained a o al o 38,353 samples wi h 55.56%
nega i es and 44.43% posi i es and he es ing se con ained
a o al o 8967 samples wi h 74.18% nega i es and 25.82%
posi i es.
B. EVALUATION METRICS
In eal-wo ld si ua ions, when conside ing he change de ec-
ion da ase , he e e ence map ypically ma ches he cu en
senso measu emen s, and de ia ions om a e e ence map
a ely occu . Thus, a co esponding change de ec ion da ase
ea u es a a he la ge amoun o non-changes (nega i es)
s. ins ances o changes (posi i es). Since accu acy can be
a misleading pe o mance me ic o imbalanced da ase s,
we apply he F1 sco es o e alua ing he p edic ions
130630 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 2. Ins ances o a non change, an appea ing and a disappea ing cons uc ion zone om he
change de ec ion da ase . Rele an changes i.e. cons uc ion zones a e anno a ed wi h whi e boxes and
appa en changes a e anno a ed in o als. No e: Only ele an changes we e aken in o accoun o
labeling.
o he classi ie . Table 1p esen s he componen s o he
con usion ma ix o e alua ing classi ica ion pe o mance.
In o de o e alua e model pe o mance, we epo
he a e age F1 sco e ac oss wo es da ase s ea u ing
disappea ing and appea ing cons uc ion zones on highways.
Addi ionally, we conside he numbe o Floa ing Poin
Ope a ions (FLOPs), a commonly used me ic o assessing
compu a ional complexi y.
VOLUME 13, 2025 130631
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
TABLE 1. Con usion ma ix / change de ec ion p oblem.
IV. EXPERIMENTS
Below, we p esen he expe imen al se ups in ol ing s a e-o -
he-a deep ne wo ks such as AlexNe , GoogLeNe , VGG,
ResNe , ViT and SWIN ans o me as ea u e ex ac o s.
A. IMPLEMENTATION DETAILS
We ained all he models using an adam op imize [6] o
15 epochs. Fo aining, we use mini-ba ches o 16 image
pai s, he Bina y C oss-En opy (BCE) loss unc ion as
o malized in Equa ion (1).
LBCE = − 1
N
N
X
i=1
yi·log(p(yi)) +(1 −yi)·log(1 −p(yi))
(1)
Fo anilla model implemen a ions and p e- ained
weigh s, we apply he o ch ision subpackage om py o ch.
Usually, he inpu images con ained a single channel and
spa ial dimensions o 256 ×256. Any o he changes
o he inpu a e de ined in he desc ip ion pa o he
co esponding backbone. Due o a change on he o e all
ne wo k con igu a ion, he sea ch o an op imal lea ning a e
was pe o med indi idually o he di e en backbones. Fo
expe imen s inco po a ing p e- ained ImageNe weigh s, he
lea ning a e was se o 0.0001, which was lowe han he
lea ning a e used o aining models om sc a ch.
B. AlexNe
AlexNe [7] is a deep ne wo k composed o i e con olu ion
laye s and h ee max-pooling laye s. Fo he i s wo
con olu ional laye s, ke nels o size 11 ×11 and 5 ×5 a e
used alongside ke nels o size 3 ×3 o deepe laye s. Each
b anch o he siamese ne wo k p ocesses an inpu image o
gi e an ou pu olume con aining 256 ou pu channels wi h
spa ial dimensions o 6 ×6. In Figu e 3, we ha e shown he
con olu ion pa o his ea u e ex ac o , a pa o he anilla
AlexNe . Subsequen ly, we la en his 3D olume and eed i
as inpu o a ully connec ed (FC) laye wi h 512 ou pu nodes
o comple e he p ocess o ea u e ex ac ion.
C. GoogLeNe
The anilla GoogLeNe [8], shown in Figu e 5, con ains a
o al o nine incep ion blocks ollowed by a global a e age
pooling laye a he end. As shown in Figu e 4, each
incep ion block con ains con olu ions wi h ke nels such as
1×1, 3 ×3, 5 ×5 and 3 ×3 max pooling o le he
ne wo k decide he bes ke nel con igu a ion o he p oblem.
Inside an incep ion block, hese con olu ions a e ca ied ou
FIGURE 3. AlexNe a chi ec u e, c . [7].
in a pa allel ashion and he ea u e maps gene a ed a e
conca ena ed oge he a he end. In e e y pa allel b anch o
he incep ion block, 1 ×1 con olu ion educes he numbe o
inpu channels o la ge con olu ions such as 3 ×3 and 5 ×5.
Global a e age pooling ope a ion ins ead o he con en ional
ully-connec ed laye s esul s in educ ion o he numbe o
ainable pa ame e s and compu a ion cos in GoogLeNe .
As a esul o global a e age pooling a he end, we ex ac
an one-dimensional ec o o 1024 ea u es om each single
channel image inpu .
FIGURE 4. Incep ion module o GoogLeNe , c . [8].
FIGURE 5. The GoogLeNe a chi ec u e, c . [8].
D. VGG
The VGG [9] ne wo k is composed o mul iple succes-
si e smalle 3 ×3 ke nels (s ide 1) han he use o
11 ×11, 7 ×7 and 5 ×5 ke nels. Fo ins ance, a single
5×5 con olu ion laye con aining 25 lea nable pa ame e s
could be eplaced by wo 3 ×3 con olu ion laye s wi h
18 ainable pa ame e s. The p ima y building block o he
VGG consis s o a sequence o con olu ions wi h 3×3 ke nels
(s ide 1, padding 1) ollowed by a 2 ×2 max-pooling
laye (s ide 2). Th ough he a angemen o VGG blocks
in cascades, ou di e en a ian s o he VGG a chi ec u e
a e possible namely VGG11, VGG13, VGG16 and VGG19,
whose con olu ion s ems and bodies we e conside ed o
ou expe imen s. Finally, a global a e age pooling laye
130632 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
was a ached o he las con olu ion laye o ex ac a
o al o 512 ea u es om each single channel image. The
expe imen al se up con aining hese a ian s along wi h hei
co esponding con olu ion laye s a e depic ed in Figu e 6.
FIGURE 6. Mul iple a ian s o VGG a chi ec u e, c . [9].
E. ResNe
Residual ne wo ks o ResNe s [10] use skip connec ion
be ween laye s, ac ing as a highway ha connec s ea lie
laye s o a mo e deepe laye s o a CNN due o he exploding
and anishing g adien p oblem encoun e ed du ing aining
o deep ne wo ks. A ResNe is made up o a numbe o
esidual blocks con aining such skip connec ions, a anged
one a e he o he . A compa ison be ween adi ional CNN
a chi ec u e and a ResNe block wi h a skip connec ion is
shown in Figu e 7. Fo he expe imen s, con olu ion laye s
o he a ian s ResNe -18, ResNe -34 and ResNe -50 wi h
a global a e age pooling laye a he end, as desc ibed
in Figu e 8we e conside ed o ea u e ex ac ion. While
a ian s such as ResNe -18, ResNe -34 ex ac ed a o al
o 512 ea u es om a single channel image, ResNe -50
ex ac ed 2048 ea u es om each single channel inpu image.
F. VISION TRANSFORMER (ViT)
In ViT [11], an inpu image is di ided in o se ies o squa e
pa ches, ans o med in o okens and encoded using blocks
ha implemen a do -p oduc based a en ion mechanism [12]
o cap u e ela ionship be ween he pa ches. Fo ou expe i-
men s, we use ViT-B/16 o p ocess h ee channel images wi h
spa ial dimensions 224 ×224, di ided in o mul iple squa e
pa ches o size 16 ×16 ×3 and ans o med in o okens
o dimension 196 ×768, whe e 196 ep esen s sequence
leng h and 768 is he hidden dimension. The ans o me
encode consis s o wel e blocks in cascade. Each encode
FIGURE 7. Compa ison o a adi ional deep CNN block and he Residual
block in ResNe wi h skip connec ions, c . [10].
FIGURE 8. Laye -wise compa ison o ResNe -18, ResNe -34 and
ResNe -50 a chi ec u es, c . [10].
block consis s o a mul i-head a en ion module ollowed
by a Mul i-laye pe cep on (MLP) ne wo k. As shown
in Figu e 9, he Vision T ans o me p ocesses image pa ches
as okens, simila o wo ds in NLP ans o me s. Fo ea u e
ex ac ion, we calcula e an a e age o all pa ches ins ead o
aking he ea u es om he class oken. The e o e, we ob ain
a o al o 768 ea u es om each inpu image.
G. SWIN TRANSFORMER (SWIN)
Fo ou expe imen s, we use he Swin-V2-S [13] a ian
o p ocess h ee channel images o spa ial dimensions
256 ×256. Simila o ViT, an image passed h ough a
Swin T ans o me is di ided in o non-o e lapping pa ches.
Howe e , a scaled cosine a en ion mechanism eplaces he
do -p oduc a en ion used in ViT. Ins ead o applying sel -
a en ion o e all pa ches a once, he Swin T ans o me
compu es a en ion wi hin small windows o he image. These
windows a e shi ed be ween laye s o allow in o ma ion o
low ac oss di e en egions.
VOLUME 13, 2025 130633
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 9. Vision T ans o me a chi ec u e, c . [11].
The a chi ec u e o SWIN ollows a hie a chical design
wi h mul iple s ages, whe e each s age educes he image
esolu ion and inc eases he ea u e dimension. This s uc u e
helps he model cap u e bo h local and global in o ma ion
e icien ly, as shown in Figu e 10. Apa om emo ing he
inal classi ica ion head, no o he changes we e made o
he s anda d Swin T ans o me o ou expe imen s. As he
encode ou pu , we ob ain a o al o 768 ea u es om each
inpu image.
FIGURE 10. Swin T ans o me a chi ec u e, c . [13].
H. VISION TRANSFORMER (ViT) WITH EARLY
CONVOLUTIONS
We explo ed he impac o ea ly con olu ional laye s on
Vision T ans o me s (ViTs), inspi ed by he pape ‘‘Ea ly
Con olu ions Help T ans o me s See Be e ’’ by Xiao [14].
Ins ead o he s anda d ViT app oach o simply spli ing an
image in o 16×16 pa ches as p oposed by [11], we in eg a ed
a ‘‘con olu ional s em’’ a he beginning. This s em uses
smalle , successi e 3×3 ke nels (like hose ound in VGG
ne wo ks) o p ocess he 224×224 inpu images. A he s em’s
end, a 1×1 ke nel ensu es he ou pu co ec ly ma ches he
768-dimensional inpu expec ed by he T ans o me encode .
We hen eshape hese ea u e maps in o a sequence be o e
eeding hem in o he i s block o a ViT-B/16 ans o me .
FIGURE 11. Image ea u e ex ac ion using a hyb id a chi ec u e
consis ing o a con olu ional s em and a ViT-B/16 ans o me , c . [14].
Figu e 11 illus a es he o wa d pass o ea u e ex ac ion
om an image inpu .
We es ed wo dis inc con olu ional s em con igu a ions:
1) VGG16-Inspi ed Con olu ional S em: This i s se up
uses a deep con olu ional s em designed o mimic
he a chi ec u e o VGG16. I consis s o hi een
3×3 con olu ional laye s, con igu ed wi h inc easing
ou pu channels (like 64, 128, 256, 512, e c.), simila
o he o iginal VGG16 ne wo k shown in Figu e 12.
This makes i a ela i ely hea yweigh op ion, packed
wi h many lea nable pa ame e s. When used wi hin
a Siamese ne wo k (whe e wo iden ical b anches
p ocess sepa a e inpu s), he ea u es ex ac ed om
bo h b anches a e conca ena ed be o e being ed in o
he decision head’s Mul i-Laye Pe cep on (MLP).
FIGURE 12. VGG16-Inspi ed Con olu ional S em, c . [9].
2) Ligh weigh Con olu ional S em: In con as , ou
second con igu a ion, shown in Figu e 13, ea u es
a much ligh e con olu ional s em. This e sion has
only ou 3×3 con olu ional laye s. Each o hese
laye s is ollowed by a ba ch no maliza ion (BN) laye ,
a ec i ied linea uni (ReLU) ac i a ion, and a max-
pooling laye . Simila o he VGG16-inspi ed s em, he
ou pu channels o hese ou laye s a e sequen ially
se o 64, 128, 256, and 512. This design has ewe
lea nable pa ame e s compa ed o he VGG16-inspi ed
s em. Ins ead o conca ena ing he ea u es om bo h
b anches, i akes he squa ed di e ence be ween hem.
This speci ic ope a ion esul s in he i s laye o
he decision head’s MLP ha ing only 768 nodes,
as i s p ocessing he di e ence di ec ly a he han a
combined inpu .
In Table 2, we ha e lis ed all combina ions o ea u e
ex ac o s and he co esponding decision heads composed
o a MLP ne wo k, conside ed o he expe imen s.
V. RESULTS
In his sec ion, we p esen he a e age F1 sco es ob ained
by he siamese ne wo k con igu a ions lis ed in Table 2
130634 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 13. Ligh weigh Con olu ional S em, c . [9].
TABLE 2. Fea u e ex ac o s and decision heads con igu ed in a Siamese
ne wo k.
on he da ase p esen ed p e iously in Sec ion III.
Addi ionally, we also discuss no able conclusions om
expe imen s whe e he en i e ne wo k was ained om
sc a ch as well as ine- uning he ne wo k using weigh s o
backbone laye s om p e- ained ImageNe 1k [15] da ase .
Fo he pu pose o inding he bes model, we p esen
he a e age F1 sco es ob ained on highway es da ase s
in Figu e 14 and amoun o Floa ing Poin Ope a ions
(FLOPs) in Figu e 15.
FIGURE 14. A e age F1 sco e ob ained by siamese ne wo ks con igu ed
wi h a ious s a e-o - he-a ea u e ex ac o s on highway change
de ec ion da ase s. The expe imen al esul s whe e ine- uning was
pe o med wi h p e- ained weigh s o ImageNe 1k da ase a e
ep esen ed using (p) nex o he backbone.
On he basis o he F1 sco es p esen ed in Figu e 14,
we obse e ha con olu ion based ea u e ex ac o s such
FIGURE 15. Amoun o Floa ing Poin Ope a ions (FLOPs) in siamese
ne wo ks con igu ed wi h a ious s a e-o - he-a backbone ea u e
ex ac o s.
as he VGG and ResNe 18 ou pe o m he pu ely a en ion-
based ViT-B/16 and SWIN. E en hough his end is coun e -
in ui i e, we can a ibu e i o size o he aining da ase .
Addi ionally, Vision ans o me s such as ViT o SWIN
ake a e y long ime o con e ge and a e highly-sensi i e
o he choice o he lea ning a e, exhibi ing subs anda d
op imizabili y. In con as , he use o ea ly 3×3 con olu ions
in con igu a ions such as VGG16 con olu ional S em +
ViT-B/16 and Ligh -weigh con olu ional S em +ViT-B/16
has enabled o quicke con e gence, obus ness o he
espec i e lea ning a e choice (0.0001) as well as he choice
o he applied op imize (Adam).
O e all, VGG ea u e ex ac o s like VGG11(p),
VGG13(p), VGG16(p) and VGG19(p) ha e ou pe o med
he baseline ResNe 18 due o he use o mul iple successi e
smalle 3 ×3 ke nels wi h s ide 1.
E en hough we un in o di icul ies aining a deep,
memo y in ensi e ne wo k om sc a ch wi hou any skip
connec ions like VGG, ini ializing i s con olu ion laye s
using p e- ained weigh s enables o as e con e gence.
This p ocess o ine- uning a ea u e ex ac o ha has
al eady been ained o ex ac ion o ea u es om a la ge
ImageNe 1K da ase is mo e e icien han aining i om
sc a ch.
VGG16(p) wi h close o 92% F1 sco e and ≈79.8 GFLOPs
is conside ed as he bes backbone choice o change
de ec ion. An inc ease in model pe o mance by ≈13.5% o e
he baseline con igu a ion wi h ResNe 18 (81% F1 sco e)
as ea u e ex ac o is obse ed om hei co esponding
a e age F1 sco es on he es da ase s.
Fo he pu pose o unde s anding he con ibu ion om
indi idual laye s o VGG16(p) owa ds model pe o mance,
we conduc ed expe imen s by eezing hem and aining he
ne wo k o a o al o 15 epochs. The a e age F1 sco es
ob ained by he siamese ne wo k wi h hei co esponding
ine- uning se ings a e gi en in Table 3.
Table 3indica es ha eezing deepe laye s such
as 18 o 25 o he VGG16(p) esul ed in a signi ican d op
in model pe o mance as compa ed o eezing ea lie laye s
such as laye s 0 o 6. When he comple e backbone is ozen
VOLUME 13, 2025 130635
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
TABLE 3. A e age F1 sco e ob ained by he siamese ne wo k wi h VGG16
ea u e ex ac o on highway change de ec ion es da ase s du ing
ine- uning.
and only he weigh s o FC laye s a e upda ed du ing aining
he ne wo k o ans e lea ning, we ob ain an a e age F1
sco e o 68% only. Hence, he p ocess o ine- uning he
model is p o en o be mo e e ec i e han ans e lea ning.
Addi ionally, he compu a ional easibili y o eal- ime
deploymen was assessed by pe o ming pos - aining s a ic
quan iza ion using he PyTo ch quan iza ion API, esul ing
in a model size o 14.4 MB and an in e ence ime o
app oxima ely 188 milliseconds on he a ge ha dwa e,
which ea u ed an AMD Ryzen 7 2700X Eigh -Co e CPU
wi h an x86-based a chi ec u e—demons a ing he model’s
sui abili y o eal- ime applica ions.
A. VISUALIZING MODEL PREDICTIONS WITH G adCAM
Since DNNs such as CNNs a e o en ea ed as black
boxes, G adCAM [16] was in oduced o p o iding isual
explana ions ha help o unde s anding why a model makes
a ce ain p edic ion. This echnique helps by gene a ing a
coa se ac i a ion map highligh ing which pa s o he inpu
image con ibu ed mos o a model’s p edic ion.
FIGURE 16. Inpu maps and hei co esponding G adCAM hea maps o
p edic ions ob ained om he siamese ne wo k a chi ec u e wi h VGG-16
backbone ea u e ex ac o . No e: The highligh ed egions o he
hea maps deno e he egions on he inpu s which con ibu ed he mos
owa ds a p edic ion.
Acco ding o he G adCAM ac i a ion maps o co ec ly
p edic ed posi i e (TP) shown in Figu e 16a, he ne wo k
is ained o igno e i ele an changes and ocus di ec ly on
hose po ions o he inpu map pai s whe e ele an changes
due o appea ance o cons uc ion zones a e loca ed. I can
also be obse ed ha he model is also unable o ocus
on cons uc ion zones loca ed away om he ehicle
on Figu e 16b. Fo he ue nega i e (TN) p edic ion shown
in Figu e 16c, he ne wo k ocuses on hose po ions o he
images whe e he e a e some amoun o isible changes,
which a e no pa o he cons uc ion zones and hus con-
side ed as i ele an . The occu ence o alse posi i es (FP)
is a ibu ed mainly o inco ec labeling as shown by he
G adCAM ac i a ion maps o inco ec ly p edic ed posi i e
in Figu e 16c. He e, he disappea ance o a cons uc ion zone
is no labeled co ec ly as a Change. Finally, o an example o
an inco ec ly classi ied nega i e (FN) shown in Figu e 16d,
he ea u e ex ac o ails o ocus on hose po ions o he
inpu map pai s whe e he e a e ele an changes due o he
disappea ance o a cons uc ion zone.
VI. CONCLUSION
In his wo k, we p esen ed a compa a i e analysis on s a e-
o - he-a deep neu al ne wo ks using bo h con olu ion and
a en ion mechanisms as candida es o backbone ea u e
ex ac ion in a siamese ne wo k con igu a ion o he
pu pose o change de ec ion. A VGG16 backbone managed
o ou pe o m he baseline siamese ne wo k con igu a ion
which used ResNe 18 by abou 13.5%. In o de o ob ain bes
esul s, ine- uning using p e- ained weigh s o con olu ion
laye s o VGG was pe o med. Acco ding o hese esul s,
we highly ecommend he use o VGG-blocks con aining
mul iple successi e smalle 3 ×3 ke nels wi h s ide 1
as a backbone ea u e ex ac o o change de ec ion on
au omo i e ada based en i onmen maps. Fu he mo e, he
pe o mance o a en ion-based backbones such as ViT and
SWIN we e also e alua ed and ound o be in e io han hose
o hei con olu ion mechanism coun e pa s. Howe e , he
o e all pe o mance o an a en ion-based ea u e ex ac ion
mechanism such as ViT was imp o ed ia he help o ea ly
con olu ions wi h mul iple successi e smalle 3 ×3 ke nels.
A signi ican challenge in add essing he change de ec ion
ask is he limi ed a ailabili y o da ase s ha cap u e a di e se
ange o changes, including hose a ising om seasonal
a ia ions and dynamic a ic condi ions. Fo u u e wo k,
i would be bene icial o pe o m a domain shi analysis o
assess he model’s pe o mance on scenes om p e iously
unobse ed loca ions o en i onmen al condi ions. Fu he -
mo e, he lack o seman ic anno a ions poses an addi ional
hu dle, as hei p esence could enable models o be ained
o mo e ine-g ained, pixel-le el change de ec ion. In u u e
wo k, we aim o ex end ou modeling app oach and analysis,
also a ge ing ou -o -dis ibu ion cases, wi h me hodological
and a chi ec u al e inemen s. He e, also ew-sho lea ning
me hods p o ide in e es ing u u e di ec ions, e.g., [17],[18].
REFERENCES
[1] Y. Li, C. L. P. Chen, and T. Zhang, ‘‘A su ey on Siamese ne wo k:
Me hodologies, applica ions, and oppo uni ies,’’ IEEE T ans. A i . In ell.,
ol. 3, no. 6, pp. 994–1014, Dec. 2022.
[2] J. B omley, I. Guyon, Y. LeCun, E. Säckinge , and R. Shah, ‘‘Signa u e
e i ica ion using a ‘Siamese’ ime delay neu al ne wo k,’’ in P oc. Ad .
Neu al In . P ocess. Sys ., ol. 6, 1993, pp. 737–744.
130636 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
[3] S. Chop a, R. Hadsell, and Y. LeCun, ‘‘Lea ning a simila i y me ic
disc imina i ely, wi h applica ion o ace e i ica ion,’’ in P oc. IEEE
Compu . Soc. Con . Compu . Vis. Pa e n Recogni . (CVPR), ol. 1,
Sep. 2005, pp. 539–546.
[4] L. Mou, M. Schmi , Y. Wang, and X. Xiang Zhu, ‘‘A CNN o he
iden i ica ion o co esponding pa ches in SAR and op ical image y
o u ban scenes,’’ in P oc. Join U ban Remo e Sens. E en (JURSE),
Ma . 2017, pp. 1–4.
[5] H. B. Swamina han, A. Somme , U. Iu gel, A. Becke , and M. A zmuelle ,
‘‘Change de ec ion in au omo i e ada based occupancy maps using
Siamese ne wo ks,’’ in P oc. In . Rada Symp., Jul. 2024, pp. 56–61.
[6] P. Qi, W. Zhou, and J. Han, ‘‘A me hod o s ochas ic L-BFGS
op imiza ion,’’ in P oc. IEEE 2nd In . Con . Cloud Compu . Big Da a Anal.
(ICCCBDA), Ap . 2017, pp. 156–160.
[7] A. K izhe sky, I. Su ske e , and G. E. Hin on, ‘‘ImageNe classi ica ion
wi h deep con olu ional neu al ne wo ks,’’ in P oc. Ad . Neu al In .
P ocess. Sys ., ol. 60, 2017, pp. 84–90.
[8] C. Szegedy, W. Liu, Y. Jia, P. Se mane , S. Reed, D. Anguelo , D. E han,
V. Vanhoucke, and A. Rabino ich, ‘‘Going deepe wi h con olu ions,’’
in P oc. IEEE Con . Compu . Vis. Pa e n Recogni . (CVPR), Jun. 2015,
pp. 1–9.
[9] K. Simonyan and A. Zisse man, ‘‘Ve y deep con olu ional ne wo ks o
la ge-scale image ecogni ion,’’ 2014, a Xi :1409.1556.
[10] K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Deep esidual lea ning o image
ecogni ion,’’ in P oc. IEEE Con . Compu . Vis. Pa e n Recogni . (CVPR),
Jun. 2016, pp. 770–778.
[11] A. Doso i skiy, L. Beye , A. Kolesniko , D. Weissenbo n, X. Zhai,
T. Un e hine , M. Dehghani, M. Minde e , G. Heigold, S. Gelly, J. Uszko-
ei , and N. Houlsby, ‘‘An image is wo h 16x16 wo ds: T ans o me s o
image ecogni ion a scale,’’ 2020, a Xi :2010.11929.
[12] A. Vaswani, N. Shazee , N. Pa ma , J. Uszko ei , L. Jones, A. N. Gomez,
Ł. Kaise , and I. Polosukhin, ‘‘A en ion is all you need,’’ in P oc. Ad .
Neu al In . P ocess. Sys ., ol. 30, 2017, pp. 5998–6008.
[13] Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang,
L. Dong, F. Wei, and B. Guo, ‘‘Swin ans o me 2: Scaling up capaci y
and esolu ion,’’ in P oc. IEEE/CVF Con . Compu . Vis. Pa e n Recogni .
(CVPR), Jun. 2022, pp. 12009–12019.
[14] T. Xiao, M. Singh, E. Min un, T. Da ell, P. Dollá , and R. Gi shick,
‘‘Ea ly con olu ions help ans o me s see be e ,’’ in P oc. Ad . Neu al
In . P ocess. Sys ., 2021, pp. 30392–30400.
[15] J. Deng, W. Dong, R. Soche , L.-J. Li, K. Li, and L. Fei-Fei, ‘‘ImageNe :
A la ge-scale hie a chical image da abase,’’ in P oc. IEEE Con . Compu .
Vis. Pa e n Recogni ., Jun. 2009, pp. 248–255.
[16] R. R. Sel a aju, M. Cogswell, A. Das, R. Vedan am, D. Pa ikh, and
D. Ba a, ‘‘G ad-CAM: Visual explana ions om deep ne wo ks ia
g adien -based localiza ion,’’ in P oc. IEEE In . Con . Compu . Vis.
(ICCV), Oc . 2017, pp. 618–626.
[17] G. Cheng, L. Cai, C. Lang, X. Yao, J. Chen, L. Guo, and J. Han, ‘‘SPNe :
Siamese-p o o ype ne wo k o ew-sho emo e sensing image scene
classi ica ion,’’ IEEE T ans. Geosci. Remo e Sens., ol. 60, 2022.
[18] J. J. Vale o-Mas, A. J. Gallego, and J. R. Rico-Juan, ‘‘An o e iew
o ensemble and ea u e lea ning in ew-sho image classi ica ion
using Siamese ne wo ks,’’ Mul imedia Tools Appl., ol. 83, no. 7,
pp. 19929–19952, Jul. 2023.
HARIHARA BHARATHY SWAMINATHAN was
bo n in Chennai, India, in 1993. He ecei ed he
B.Tech. deg ee in elec onics and communica-
ion enginee ing om he B. S. Abdu Rahman
C escen Ins i u e o Science and Technology,
Tamil Nadu, India, in 2014, and he M.Eng.
deg ee in embedded sys ems o mecha onics
om he Uni e si y o Applied Sciences (FH),
Do mund, Ge many, in 2020. He is cu en ly
pu suing he Ph.D. deg ee wi h he Seman ic
In o ma ion Sys ems (SIS) G oup, Osnab ück Uni e si y, Ge many. His
esea ch in e es s include applica ion o machine lea ning and deep lea ning
o en i onmen pe cep ion o sel -d i ing ca s, pa icula ly using da a om
au omo i e ada s.
ARON SOMMER was bo n in Be lin, Ge many,
in 1986. He ecei ed he Dipl.-Ma h. Techn.
deg ee in Technoma hema ics wi h a echnical
backg ound in communica ions enginee ing om
Ka ls uhe Ins i u e o Technology (KIT), Ka l-
s uhe, Ge many, in 2013, and he Ph.D. deg ee
om he Ins i u e o In o ma ion P ocessing, Leib-
niz Uni e si y Hanno e , in 2019, whe e he was
in ol ed in a p ojec on syn he ic ape u e ada .
Since 2020, he has been he Rada Algo i hm
Expe o au omo i e applica ions wi h Ap i .
URI IURGEL ecei ed he Dipl.-Ing. deg ee in
elec ical enginee ing (wi h a ocus on in o ma ion
echnology) om he Uni e si y o Duisbu g,
Ge many, in 2000, and he D .-Ing. deg ee om
he Technical Uni e si y o Munich, in 2006.
Since 2005, he has been wi h Delphi/Ap i ,
cu en ly holding he posi ion o he Manage
Rada Pe cep ion and a Senio Expe Senso
Algo i hms. His esea ch in e es s include a i-
icial in elligence, in o ma ion e ie al, na u al
language p ocessing, algo i hms o en i onmen pe cep ion wi h came a and
ada , and ada signal p ocessing, and speci ically ela es o applica ions o
au onomous d i ing and ad anced d i e assis ance sys ems.
ANDREAS BECKER (Membe , IEEE) was bo n
in Wuppe al, Ge many, in 1975. He ecei ed
he Dipl.-Ing. deg ee in elec ical enginee ing
and he D .-Ing. deg ee om he Uni e si y o
Wuppe al, in 2000 and 2006, espec i ely. Du ing
he doc o al deg ee, his esea ch was ocused
on he nume ical simula ion o elec omagne ic
p oblems, including he analysis o bo ehole
ada s. F om 2007 o 2016, he was wi h Hella
GmbH & Co. KGaA, and Delphi Deu schland
GmbH (now Ap i ), e en ually achie ing he posi ion o he Rada Technical
Manage . Since 2017, he has been a P o esso o in o ma ion echnology
wi h he Uni e si y o Applied Sciences Do mund, Do mund, Ge many. His
esea ch in e es includes pe cep ion and con ol sys ems o mobile obo s.
MARTIN ATZMUELLER is cu en ly a Full
P o esso wi h he Ins i u e o Compu e Science,
Osnab ück Uni e si y, Ge many, whe e he holds
he ROSEN-G oup-Endowed Chai o seman-
ic in o ma ion sys ems. He is he Founding
Spokespe son wi h he Join Labo a o y on A i-
icial In elligence and Da a Science, a Founding
Membe wi h he Resea ch Uni Da a Science,
Osnab ück Uni e si y, and he Scien i ic Di ec o
wi h he Resea ch Depa men Coope a i e and
Au onomous Sys ems, Ge man Resea ch Cen e o A i icial In elli-
gence (DFKI). His esea ch in e es s include a i icial in elligence (AI), da a
science, and in eg a i e AI sys ems, whe e his pa icula esea ch in e es s
include modeling complex da a, explainable AI, in e p e able lea ning,
machine pe cep ion, and seman ic in e p e a ion, and ela es o applica ions
in complex in eg a i e AI sys ems, especially obo con ol and senso -based
AI sys ems.
VOLUME 13, 2025 130637