scieee Science in your language
[en] (orig)

Comparative Analysis of Deep Learning-Based Feature Extractors for Change Detection in Automotive Radar Maps

Abstract

The Siamese network architecture has been applied by deep learning practitioners to find similarities between images. In the domain of autonomous driving, this network configuration has recently gained attention for solving the change detection task, which involves identifying changes in a previously known map of a vehicle’s environment. This is vital, as such deviations may compromise the accuracy and reliability of the map, which is essential for the vehicle’s ability to localize itself and navigate effectively. In this paper, we present a set of experiments involving state-of-the-art deep learning architectures based on both convolution (CNN) and attention mechanisms such as AlexNet, GoogLeNet, VGG, ResNet, Vision Transformer, and Shifted Windows Transformer as possible candidates for the feature extractor backbone module in the Siamese architecture to detect changes caused by the disappearance and appearance of construction zones. Also, we evaluate the performance of these architectures using fine-tuning, i.e., initializing the convolutional layers with pre-trained weights. In our experimentation, the best results were obtained using VGG16 (CNN), especially when it was initialized using pre-trained weights from the ImageNet-1K dataset. In particular, VGG16 with an average F1 score of 92% on highway datasets outperformed the baseline residual network composed of ResNet18 convolutions by about 13.5%.

Read accessible full text

Comparative Analysis of Deep Learning-Based Feature Extractors for Change Detection in Automotive Radar Maps

Author: Swaminathan, Harihara Bharathy,Sommer, Aron,Iurgel, Uri,Becker, Andreas,Atzmueller, Martin
Year: 2025
DOI: 10.48693/879
Source: https://osnadocs.ub.uni-osnabrueck.de/bitstream/ds-2026022014574/1/Swaminathan_etal_IEEEAccess_13_130629-130637_2025.pdf
Recei ed 5 June 2025, accep ed 26 June 2025, da e o publica ion 21 July 2025, da e o cu en e sion 29 July 2025.
Digi al Objec Iden i ie 10.1109/ACCESS.2025.3591272
Compa a i e Analysis o Deep Lea ning-Based
Fea u e Ex ac o s o Change De ec ion
in Au omo i e Rada Maps
HARIHARA BHARATHY SWAMINATHAN 1, ARON SOMMER2, URI IURGEL2,
ANDREAS BECKER 3, (Membe , IEEE), AND MARTIN ATZMUELLER 1,4
1Seman ic In o ma ion Sys ems G oup, Osnab ück Uni e si y, 49074 Osnab ück, Ge many
2Ap i Se ices Deu schland GmbH, 42119 Wuppe al, Ge many
3Facul y o In o ma ion Technology, Fachhochschule Do mund, 44139 Do mund, Ge many
4Ge man Resea ch Cen e o A i icial In elligence (DFKI), 49084 Osnab ück, Ge many
Co esponding au ho : Ha iha a Bha a hy Swamina han (hsw[email p o ec ed])
ABSTRACT The Siamese ne wo k a chi ec u e has been applied by deep lea ning p ac i ione s o
ind simila i ies be ween images. In he domain o au onomous d i ing, his ne wo k con igu a ion has
ecen ly gained a en ion o sol ing he change de ec ion ask, which in ol es iden i ying changes in
a p e iously known map o a ehicle’s en i onmen . This is i al, as such de ia ions may comp omise
he accu acy and eliabili y o he map, which is essen ial o he ehicle’s abili y o localize i sel and
na iga e e ec i ely. In his pape , we p esen a se o expe imen s in ol ing s a e-o - he-a deep lea ning
a chi ec u es based on bo h con olu ion (CNN) and a en ion mechanisms such as AlexNe , GoogLeNe ,
VGG, ResNe , Vision T ans o me , and Shi ed Windows T ans o me as possible candida es o he ea u e
ex ac o backbone module in he Siamese a chi ec u e o de ec changes caused by he disappea ance and
appea ance o cons uc ion zones. Also, we e alua e he pe o mance o hese a chi ec u es using ine-
uning, i. e., ini ializing he con olu ional laye s wi h p e- ained weigh s. In ou expe imen a ion, he bes
esul s we e ob ained using VGG16 (CNN), especially when i was ini ialized using p e- ained weigh s
om he ImageNe -1K da ase . In pa icula , VGG16 wi h an a e age F1 sco e o 92% on highway da ase s
ou pe o med he baseline esidual ne wo k composed o ResNe 18 con olu ions by abou 13.5%.
INDEX TERMS Change de ec ion, au omo i e ada , occupancy maps, siamese ne wo ks.
I. INTRODUCTION
The pe o mance o s a e-o - he-a deep neu al ne wo ks
has been e alua ed in he pas on he basis o hei sco es
in he classi ica ion o benchma k da ase s such as he
ImageNe . In his pape , we ocus on ea u e ex ac o s
in a speci ic domain, i. e., au omo i e ada . In pa icula ,
we e alua e he pe o mance o ea u e ex ac o s such
as AlexNe , GoogLeNe , VGG, ResNe , Vision T ans-
o me (ViT) and Shi ed Windows (SWIN) T ans o me
o he pu pose o change de ec ion on au omo i e ada -
based maps. In his wo k, he ask is amed as a
bina y classi ica ion p oblem, in which he model p edic s
The associa e edi o coo dina ing he e iew o his manusc ip and
app o ing i o publica ion was Fab izio San i .
whe he he e e ence map has di e ged om he cu en
en i onmen . These p edic ions help assess he accu acy
and ele ance o he e e ence map and indica e when e-
mapping is necessa y due o de ec ed changes in he ehicle’s
en i onmen .
As ou main con ibu ions, we p esen esul s on how
o de ec changes caused by he disappea ance and appea -
ance o cons uc ion zones. Fu he mo e, we e alua e he
pe o mance o hese ea u e ex ac o s a e ine- uning,
i. e., ini ializing he con olu ional laye s wi h p e- ained
weigh s. Fo he pu pose o change de ec ion, we use a
Siamese ne wo k con igu a ion, which consis s o wo deep
neu al ne wo k (DNN) sub-ne wo ks as backbones o ea u e
ex ac ion om an image pai , ollowed by a decision head o
image simila i y lea ning.
VOLUME 13, 2025
2025 The Au ho s. This wo k is licensed unde a C ea i e Commons A ibu ion 4.0 License.
Fo mo e in o ma ion, see h ps://c ea i ecommons.o g/licenses/by/4.0/ 130629
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
In ou expe imen s, we ocus on he domain o au omo i e
ada , o which ealis ic open benchma k da ase s a e a he
sca ce and need o be ailo ed o he speci ic applica ion
domain. The e o e, as a da ase , we use eal-li e au omo i e
ada map image pai s gene a ed o e a one-yea pe iod
a ound he Ge man ci y o Wuppe al, pa icula ly on
highways wi h cons uc ion zones.
Ou main indings a e summa ized as ollows:
1) Siamese ne wo ks con igu ed wi h VGG- amily o
ea u e ex ac o s ou pe o med he baseline con-
igu a ion wi h a esidual ne wo k i. e., ResNe 18.
Speci ically, he VGG16 backbone achie ed an F1
sco e inc ease o abou 13.5% on a e age on es
da ase s depic ing disappea ing and appea ing highway
cons uc ion zones.
2) By ini ializing he con olu ion laye s in VGG using
p e- ained weigh s ob ained om classi ica ion o
ImageNe -1K da ase , we o e all achie ed as e con-
e gence h ough he p ocess o ine- uning.
The es o he pape is s uc u ed as ollows: Sec ion II
discusses ela ed wo k. A e ha , Sec ion III p esen s he
change de ec ion da ase along wi h he applied e alua ion
me ics. Nex , Sec ion IV and Sec ion Vdiscusses ou
expe imen a ion and esul s in de ail, espec i ely. Finally,
Sec ion VI concludes wi h a summa y and in e es ing
di ec ions o u u e wo k.
II. RELATED WORK
In he pas , Siamese ne wo ks [1] ha e been ex ensi ely
s udied o a ious use cases, anging om signa u e
e i ica ion [2], ace e i ica ion [3] o change de ec ion in
emo e sensing applica ions [4].
Recen ly, Swamina han e al. ha e p oposed a Siamese
ne wo k-based classi ie o de ec ing changes due o he
appea ance and disappea ance o cons uc ion zones on ada
maps o he en i onmen [5]. As backbones, hey used
he con olu ion laye s o ResNe -18 as a ea u e ex ac o .
As an inpu o he mul i-laye pe cep on (MLP) ne wo k
o he decision module, ea u e ec o s om he CNN
backbones we e conca ena ed. The in e media e laye s o he
MLP had a size o 256 and 1, subsequen ly ollowed by a
sigmoid ac i a ion a he end, ained using a bina y c oss-
en opy (BCE) loss unc ion. This con igu a ion will se e as
he baseline o ou wo k and is shown in Figu e 1.
FIGURE 1. The baseline siamese ne wo k a chi ec u e composed o a
ResNe -18 ea u e ex ac o and MLP ne wo k as a decision head.
O e all, o he bes o he au ho s’ knowledge, a compa a-
i e analysis o exis ing s a e-o - he-a deep neu al ne wo k
a chi ec u es as candida es o a ea u e ex ac ion backbone
in a Siamese ne wo k con igu a ion has no been ca ied ou
in he li e a u e be o e.
III. DATASET AND EVALUATION METRICS
A. CHANGE DETECTION DATASET
Fo he pu pose o aining and es ing, we use he
change de ec ion da ase [5] p esen ed p e iously by
Swamina han e al. The da ase includes ada maps show-
ing changes such as he appea ance and disappea ance
o cons uc ion zones. These maps we e collec ed o e a yea
a ound Wuppe al, Ge many, wi h a ocus on a highway
connec ing Wuppe al and Düsseldo , as shown in Figu e 2.
A nega i e sample (Figu e 2a) shows only i ele an o mino
changes (ma ked wi h o als), while posi i e samples show
cons uc ion zones appea ing o disappea ing, along wi h
some un ela ed changes, c . Figu es 2b and 2c. Du ing da a
collec ion, ce ain lanes we e closed due o cons uc ion, and
he es ehicle was d i en h ough a nea by lane on he
highway.
These maps we e gene a ed using measu emen s o co ne
ada senso s i ed o he es ehicle. Fu he mo e, a high
p ecision DGPS measu emen uni i.e. he POS LV-220
posi ioning sys em om Applanix co po a ion was used o
ob ain p ecise posi ioning o he ehicle. Each sample o his
da ase consis s o pe ec ly aligned bi d’s eye- iew (BEV)
ada map image pai s ob ained om di e en measu emen
d i es o he same en i onmen . Wi hin a speci ic da a
sample, e ms such as a ‘‘ e e ence map’’ and a ‘‘cu en
map’’ a e used o deno e images ha we e cap u ed a
di e en poin s in ime.
We di ide he aining and es ing samples, ensu ing
ha a cons uc ion zone shown du ing he aining p ocess
is no pa o he es ing one. Fo ou expe imen s, he
image pai s comp ising he appea ance and disappea ance
o cons uc ion zones in he ou e owa ds Düsseldo a e
conside ed as pa o he aining da ase and he image pai s
comp ising o appea ing (A) and disappea ing cons uc ion
zones (D) in he ou e owa ds Wuppe al as pa o
he es ing da ase . In e ms o size and dis ibu ion, he
aining se con ained a o al o 38,353 samples wi h 55.56%
nega i es and 44.43% posi i es and he es ing se con ained
a o al o 8967 samples wi h 74.18% nega i es and 25.82%
posi i es.
B. EVALUATION METRICS
In eal-wo ld si ua ions, when conside ing he change de ec-
ion da ase , he e e ence map ypically ma ches he cu en
senso measu emen s, and de ia ions om a e e ence map
a ely occu . Thus, a co esponding change de ec ion da ase
ea u es a a he la ge amoun o non-changes (nega i es)
s. ins ances o changes (posi i es). Since accu acy can be
a misleading pe o mance me ic o imbalanced da ase s,
we apply he F1 sco es o e alua ing he p edic ions
130630 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 2. Ins ances o a non change, an appea ing and a disappea ing cons uc ion zone om he
change de ec ion da ase . Rele an changes i.e. cons uc ion zones a e anno a ed wi h whi e boxes and
appa en changes a e anno a ed in o als. No e: Only ele an changes we e aken in o accoun o
labeling.
o he classi ie . Table 1p esen s he componen s o he
con usion ma ix o e alua ing classi ica ion pe o mance.
In o de o e alua e model pe o mance, we epo
he a e age F1 sco e ac oss wo es da ase s ea u ing
disappea ing and appea ing cons uc ion zones on highways.
Addi ionally, we conside he numbe o Floa ing Poin
Ope a ions (FLOPs), a commonly used me ic o assessing
compu a ional complexi y.
VOLUME 13, 2025 130631
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
TABLE 1. Con usion ma ix / change de ec ion p oblem.
IV. EXPERIMENTS
Below, we p esen he expe imen al se ups in ol ing s a e-o -
he-a deep ne wo ks such as AlexNe , GoogLeNe , VGG,
ResNe , ViT and SWIN ans o me as ea u e ex ac o s.
A. IMPLEMENTATION DETAILS
We ained all he models using an adam op imize [6] o
15 epochs. Fo aining, we use mini-ba ches o 16 image
pai s, he Bina y C oss-En opy (BCE) loss unc ion as
o malized in Equa ion (1).
LBCE = − 1
N
N
X
i=1
yi·log(p(yi)) +(1 −yi)·log(1 −p(yi))
(1)
Fo anilla model implemen a ions and p e- ained
weigh s, we apply he o ch ision subpackage om py o ch.
Usually, he inpu images con ained a single channel and
spa ial dimensions o 256 ×256. Any o he changes
o he inpu a e de ined in he desc ip ion pa o he
co esponding backbone. Due o a change on he o e all
ne wo k con igu a ion, he sea ch o an op imal lea ning a e
was pe o med indi idually o he di e en backbones. Fo
expe imen s inco po a ing p e- ained ImageNe weigh s, he
lea ning a e was se o 0.0001, which was lowe han he
lea ning a e used o aining models om sc a ch.
B. AlexNe
AlexNe [7] is a deep ne wo k composed o i e con olu ion
laye s and h ee max-pooling laye s. Fo he i s wo
con olu ional laye s, ke nels o size 11 ×11 and 5 ×5 a e
used alongside ke nels o size 3 ×3 o deepe laye s. Each
b anch o he siamese ne wo k p ocesses an inpu image o
gi e an ou pu olume con aining 256 ou pu channels wi h
spa ial dimensions o 6 ×6. In Figu e 3, we ha e shown he
con olu ion pa o his ea u e ex ac o , a pa o he anilla
AlexNe . Subsequen ly, we la en his 3D olume and eed i
as inpu o a ully connec ed (FC) laye wi h 512 ou pu nodes
o comple e he p ocess o ea u e ex ac ion.
C. GoogLeNe
The anilla GoogLeNe [8], shown in Figu e 5, con ains a
o al o nine incep ion blocks ollowed by a global a e age
pooling laye a he end. As shown in Figu e 4, each
incep ion block con ains con olu ions wi h ke nels such as
1×1, 3 ×3, 5 ×5 and 3 ×3 max pooling o le he
ne wo k decide he bes ke nel con igu a ion o he p oblem.
Inside an incep ion block, hese con olu ions a e ca ied ou
FIGURE 3. AlexNe a chi ec u e, c . [7].
in a pa allel ashion and he ea u e maps gene a ed a e
conca ena ed oge he a he end. In e e y pa allel b anch o
he incep ion block, 1 ×1 con olu ion educes he numbe o
inpu channels o la ge con olu ions such as 3 ×3 and 5 ×5.
Global a e age pooling ope a ion ins ead o he con en ional
ully-connec ed laye s esul s in educ ion o he numbe o
ainable pa ame e s and compu a ion cos in GoogLeNe .
As a esul o global a e age pooling a he end, we ex ac
an one-dimensional ec o o 1024 ea u es om each single
channel image inpu .
FIGURE 4. Incep ion module o GoogLeNe , c . [8].
FIGURE 5. The GoogLeNe a chi ec u e, c . [8].
D. VGG
The VGG [9] ne wo k is composed o mul iple succes-
si e smalle 3 ×3 ke nels (s ide 1) han he use o
11 ×11, 7 ×7 and 5 ×5 ke nels. Fo ins ance, a single
5×5 con olu ion laye con aining 25 lea nable pa ame e s
could be eplaced by wo 3 ×3 con olu ion laye s wi h
18 ainable pa ame e s. The p ima y building block o he
VGG consis s o a sequence o con olu ions wi h 3×3 ke nels
(s ide 1, padding 1) ollowed by a 2 ×2 max-pooling
laye (s ide 2). Th ough he a angemen o VGG blocks
in cascades, ou di e en a ian s o he VGG a chi ec u e
a e possible namely VGG11, VGG13, VGG16 and VGG19,
whose con olu ion s ems and bodies we e conside ed o
ou expe imen s. Finally, a global a e age pooling laye
130632 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
was a ached o he las con olu ion laye o ex ac a
o al o 512 ea u es om each single channel image. The
expe imen al se up con aining hese a ian s along wi h hei
co esponding con olu ion laye s a e depic ed in Figu e 6.
FIGURE 6. Mul iple a ian s o VGG a chi ec u e, c . [9].
E. ResNe
Residual ne wo ks o ResNe s [10] use skip connec ion
be ween laye s, ac ing as a highway ha connec s ea lie
laye s o a mo e deepe laye s o a CNN due o he exploding
and anishing g adien p oblem encoun e ed du ing aining
o deep ne wo ks. A ResNe is made up o a numbe o
esidual blocks con aining such skip connec ions, a anged
one a e he o he . A compa ison be ween adi ional CNN
a chi ec u e and a ResNe block wi h a skip connec ion is
shown in Figu e 7. Fo he expe imen s, con olu ion laye s
o he a ian s ResNe -18, ResNe -34 and ResNe -50 wi h
a global a e age pooling laye a he end, as desc ibed
in Figu e 8we e conside ed o ea u e ex ac ion. While
a ian s such as ResNe -18, ResNe -34 ex ac ed a o al
o 512 ea u es om a single channel image, ResNe -50
ex ac ed 2048 ea u es om each single channel inpu image.
F. VISION TRANSFORMER (ViT)
In ViT [11], an inpu image is di ided in o se ies o squa e
pa ches, ans o med in o okens and encoded using blocks
ha implemen a do -p oduc based a en ion mechanism [12]
o cap u e ela ionship be ween he pa ches. Fo ou expe i-
men s, we use ViT-B/16 o p ocess h ee channel images wi h
spa ial dimensions 224 ×224, di ided in o mul iple squa e
pa ches o size 16 ×16 ×3 and ans o med in o okens
o dimension 196 ×768, whe e 196 ep esen s sequence
leng h and 768 is he hidden dimension. The ans o me
encode consis s o wel e blocks in cascade. Each encode
FIGURE 7. Compa ison o a adi ional deep CNN block and he Residual
block in ResNe wi h skip connec ions, c . [10].
FIGURE 8. Laye -wise compa ison o ResNe -18, ResNe -34 and
ResNe -50 a chi ec u es, c . [10].
block consis s o a mul i-head a en ion module ollowed
by a Mul i-laye pe cep on (MLP) ne wo k. As shown
in Figu e 9, he Vision T ans o me p ocesses image pa ches
as okens, simila o wo ds in NLP ans o me s. Fo ea u e
ex ac ion, we calcula e an a e age o all pa ches ins ead o
aking he ea u es om he class oken. The e o e, we ob ain
a o al o 768 ea u es om each inpu image.
G. SWIN TRANSFORMER (SWIN)
Fo ou expe imen s, we use he Swin-V2-S [13] a ian
o p ocess h ee channel images o spa ial dimensions
256 ×256. Simila o ViT, an image passed h ough a
Swin T ans o me is di ided in o non-o e lapping pa ches.
Howe e , a scaled cosine a en ion mechanism eplaces he
do -p oduc a en ion used in ViT. Ins ead o applying sel -
a en ion o e all pa ches a once, he Swin T ans o me
compu es a en ion wi hin small windows o he image. These
windows a e shi ed be ween laye s o allow in o ma ion o
low ac oss di e en egions.
VOLUME 13, 2025 130633

H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 9. Vision T ans o me a chi ec u e, c . [11].
The a chi ec u e o SWIN ollows a hie a chical design
wi h mul iple s ages, whe e each s age educes he image
esolu ion and inc eases he ea u e dimension. This s uc u e
helps he model cap u e bo h local and global in o ma ion
e icien ly, as shown in Figu e 10. Apa om emo ing he
inal classi ica ion head, no o he changes we e made o
he s anda d Swin T ans o me o ou expe imen s. As he
encode ou pu , we ob ain a o al o 768 ea u es om each
inpu image.
FIGURE 10. Swin T ans o me a chi ec u e, c . [13].
H. VISION TRANSFORMER (ViT) WITH EARLY
CONVOLUTIONS
We explo ed he impac o ea ly con olu ional laye s on
Vision T ans o me s (ViTs), inspi ed by he pape ‘‘Ea ly
Con olu ions Help T ans o me s See Be e ’’ by Xiao [14].
Ins ead o he s anda d ViT app oach o simply spli ing an
image in o 16×16 pa ches as p oposed by [11], we in eg a ed
a ‘‘con olu ional s em’’ a he beginning. This s em uses
smalle , successi e 3×3 ke nels (like hose ound in VGG
ne wo ks) o p ocess he 224×224 inpu images. A he s em’s
end, a 1×1 ke nel ensu es he ou pu co ec ly ma ches he
768-dimensional inpu expec ed by he T ans o me encode .
We hen eshape hese ea u e maps in o a sequence be o e
eeding hem in o he i s block o a ViT-B/16 ans o me .
FIGURE 11. Image ea u e ex ac ion using a hyb id a chi ec u e
consis ing o a con olu ional s em and a ViT-B/16 ans o me , c . [14].
Figu e 11 illus a es he o wa d pass o ea u e ex ac ion
om an image inpu .
We es ed wo dis inc con olu ional s em con igu a ions:
1) VGG16-Inspi ed Con olu ional S em: This i s se up
uses a deep con olu ional s em designed o mimic
he a chi ec u e o VGG16. I consis s o hi een
3×3 con olu ional laye s, con igu ed wi h inc easing
ou pu channels (like 64, 128, 256, 512, e c.), simila
o he o iginal VGG16 ne wo k shown in Figu e 12.
This makes i a ela i ely hea yweigh op ion, packed
wi h many lea nable pa ame e s. When used wi hin
a Siamese ne wo k (whe e wo iden ical b anches
p ocess sepa a e inpu s), he ea u es ex ac ed om
bo h b anches a e conca ena ed be o e being ed in o
he decision head’s Mul i-Laye Pe cep on (MLP).
FIGURE 12. VGG16-Inspi ed Con olu ional S em, c . [9].
2) Ligh weigh Con olu ional S em: In con as , ou
second con igu a ion, shown in Figu e 13, ea u es
a much ligh e con olu ional s em. This e sion has
only ou 3×3 con olu ional laye s. Each o hese
laye s is ollowed by a ba ch no maliza ion (BN) laye ,
a ec i ied linea uni (ReLU) ac i a ion, and a max-
pooling laye . Simila o he VGG16-inspi ed s em, he
ou pu channels o hese ou laye s a e sequen ially
se o 64, 128, 256, and 512. This design has ewe
lea nable pa ame e s compa ed o he VGG16-inspi ed
s em. Ins ead o conca ena ing he ea u es om bo h
b anches, i akes he squa ed di e ence be ween hem.
This speci ic ope a ion esul s in he i s laye o
he decision head’s MLP ha ing only 768 nodes,
as i s p ocessing he di e ence di ec ly a he han a
combined inpu .
In Table 2, we ha e lis ed all combina ions o ea u e
ex ac o s and he co esponding decision heads composed
o a MLP ne wo k, conside ed o he expe imen s.
V. RESULTS
In his sec ion, we p esen he a e age F1 sco es ob ained
by he siamese ne wo k con igu a ions lis ed in Table 2
130634 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
FIGURE 13. Ligh weigh Con olu ional S em, c . [9].
TABLE 2. Fea u e ex ac o s and decision heads con igu ed in a Siamese
ne wo k.
on he da ase p esen ed p e iously in Sec ion III.
Addi ionally, we also discuss no able conclusions om
expe imen s whe e he en i e ne wo k was ained om
sc a ch as well as ine- uning he ne wo k using weigh s o
backbone laye s om p e- ained ImageNe 1k [15] da ase .
Fo he pu pose o inding he bes model, we p esen
he a e age F1 sco es ob ained on highway es da ase s
in Figu e 14 and amoun o Floa ing Poin Ope a ions
(FLOPs) in Figu e 15.
FIGURE 14. A e age F1 sco e ob ained by siamese ne wo ks con igu ed
wi h a ious s a e-o - he-a ea u e ex ac o s on highway change
de ec ion da ase s. The expe imen al esul s whe e ine- uning was
pe o med wi h p e- ained weigh s o ImageNe 1k da ase a e
ep esen ed using (p) nex o he backbone.
On he basis o he F1 sco es p esen ed in Figu e 14,
we obse e ha con olu ion based ea u e ex ac o s such
FIGURE 15. Amoun o Floa ing Poin Ope a ions (FLOPs) in siamese
ne wo ks con igu ed wi h a ious s a e-o - he-a backbone ea u e
ex ac o s.
as he VGG and ResNe 18 ou pe o m he pu ely a en ion-
based ViT-B/16 and SWIN. E en hough his end is coun e -
in ui i e, we can a ibu e i o size o he aining da ase .
Addi ionally, Vision ans o me s such as ViT o SWIN
ake a e y long ime o con e ge and a e highly-sensi i e
o he choice o he lea ning a e, exhibi ing subs anda d
op imizabili y. In con as , he use o ea ly 3×3 con olu ions
in con igu a ions such as VGG16 con olu ional S em +
ViT-B/16 and Ligh -weigh con olu ional S em +ViT-B/16
has enabled o quicke con e gence, obus ness o he
espec i e lea ning a e choice (0.0001) as well as he choice
o he applied op imize (Adam).
O e all, VGG ea u e ex ac o s like VGG11(p),
VGG13(p), VGG16(p) and VGG19(p) ha e ou pe o med
he baseline ResNe 18 due o he use o mul iple successi e
smalle 3 ×3 ke nels wi h s ide 1.
E en hough we un in o di icul ies aining a deep,
memo y in ensi e ne wo k om sc a ch wi hou any skip
connec ions like VGG, ini ializing i s con olu ion laye s
using p e- ained weigh s enables o as e con e gence.
This p ocess o ine- uning a ea u e ex ac o ha has
al eady been ained o ex ac ion o ea u es om a la ge
ImageNe 1K da ase is mo e e icien han aining i om
sc a ch.
VGG16(p) wi h close o 92% F1 sco e and ≈79.8 GFLOPs
is conside ed as he bes backbone choice o change
de ec ion. An inc ease in model pe o mance by ≈13.5% o e
he baseline con igu a ion wi h ResNe 18 (81% F1 sco e)
as ea u e ex ac o is obse ed om hei co esponding
a e age F1 sco es on he es da ase s.
Fo he pu pose o unde s anding he con ibu ion om
indi idual laye s o VGG16(p) owa ds model pe o mance,
we conduc ed expe imen s by eezing hem and aining he
ne wo k o a o al o 15 epochs. The a e age F1 sco es
ob ained by he siamese ne wo k wi h hei co esponding
ine- uning se ings a e gi en in Table 3.
Table 3indica es ha eezing deepe laye s such
as 18 o 25 o he VGG16(p) esul ed in a signi ican d op
in model pe o mance as compa ed o eezing ea lie laye s
such as laye s 0 o 6. When he comple e backbone is ozen
VOLUME 13, 2025 130635
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
TABLE 3. A e age F1 sco e ob ained by he siamese ne wo k wi h VGG16
ea u e ex ac o on highway change de ec ion es da ase s du ing
ine- uning.
and only he weigh s o FC laye s a e upda ed du ing aining
he ne wo k o ans e lea ning, we ob ain an a e age F1
sco e o 68% only. Hence, he p ocess o ine- uning he
model is p o en o be mo e e ec i e han ans e lea ning.
Addi ionally, he compu a ional easibili y o eal- ime
deploymen was assessed by pe o ming pos - aining s a ic
quan iza ion using he PyTo ch quan iza ion API, esul ing
in a model size o 14.4 MB and an in e ence ime o
app oxima ely 188 milliseconds on he a ge ha dwa e,
which ea u ed an AMD Ryzen 7 2700X Eigh -Co e CPU
wi h an x86-based a chi ec u e—demons a ing he model’s
sui abili y o eal- ime applica ions.
A. VISUALIZING MODEL PREDICTIONS WITH G adCAM
Since DNNs such as CNNs a e o en ea ed as black
boxes, G adCAM [16] was in oduced o p o iding isual
explana ions ha help o unde s anding why a model makes
a ce ain p edic ion. This echnique helps by gene a ing a
coa se ac i a ion map highligh ing which pa s o he inpu
image con ibu ed mos o a model’s p edic ion.
FIGURE 16. Inpu maps and hei co esponding G adCAM hea maps o
p edic ions ob ained om he siamese ne wo k a chi ec u e wi h VGG-16
backbone ea u e ex ac o . No e: The highligh ed egions o he
hea maps deno e he egions on he inpu s which con ibu ed he mos
owa ds a p edic ion.
Acco ding o he G adCAM ac i a ion maps o co ec ly
p edic ed posi i e (TP) shown in Figu e 16a, he ne wo k
is ained o igno e i ele an changes and ocus di ec ly on
hose po ions o he inpu map pai s whe e ele an changes
due o appea ance o cons uc ion zones a e loca ed. I can
also be obse ed ha he model is also unable o ocus
on cons uc ion zones loca ed away om he ehicle
on Figu e 16b. Fo he ue nega i e (TN) p edic ion shown
in Figu e 16c, he ne wo k ocuses on hose po ions o he
images whe e he e a e some amoun o isible changes,
which a e no pa o he cons uc ion zones and hus con-
side ed as i ele an . The occu ence o alse posi i es (FP)
is a ibu ed mainly o inco ec labeling as shown by he
G adCAM ac i a ion maps o inco ec ly p edic ed posi i e
in Figu e 16c. He e, he disappea ance o a cons uc ion zone
is no labeled co ec ly as a Change. Finally, o an example o
an inco ec ly classi ied nega i e (FN) shown in Figu e 16d,
he ea u e ex ac o ails o ocus on hose po ions o he
inpu map pai s whe e he e a e ele an changes due o he
disappea ance o a cons uc ion zone.
VI. CONCLUSION
In his wo k, we p esen ed a compa a i e analysis on s a e-
o - he-a deep neu al ne wo ks using bo h con olu ion and
a en ion mechanisms as candida es o backbone ea u e
ex ac ion in a siamese ne wo k con igu a ion o he
pu pose o change de ec ion. A VGG16 backbone managed
o ou pe o m he baseline siamese ne wo k con igu a ion
which used ResNe 18 by abou 13.5%. In o de o ob ain bes
esul s, ine- uning using p e- ained weigh s o con olu ion
laye s o VGG was pe o med. Acco ding o hese esul s,
we highly ecommend he use o VGG-blocks con aining
mul iple successi e smalle 3 ×3 ke nels wi h s ide 1
as a backbone ea u e ex ac o o change de ec ion on
au omo i e ada based en i onmen maps. Fu he mo e, he
pe o mance o a en ion-based backbones such as ViT and
SWIN we e also e alua ed and ound o be in e io han hose
o hei con olu ion mechanism coun e pa s. Howe e , he
o e all pe o mance o an a en ion-based ea u e ex ac ion
mechanism such as ViT was imp o ed ia he help o ea ly
con olu ions wi h mul iple successi e smalle 3 ×3 ke nels.
A signi ican challenge in add essing he change de ec ion
ask is he limi ed a ailabili y o da ase s ha cap u e a di e se
ange o changes, including hose a ising om seasonal
a ia ions and dynamic a ic condi ions. Fo u u e wo k,
i would be bene icial o pe o m a domain shi analysis o
assess he model’s pe o mance on scenes om p e iously
unobse ed loca ions o en i onmen al condi ions. Fu he -
mo e, he lack o seman ic anno a ions poses an addi ional
hu dle, as hei p esence could enable models o be ained
o mo e ine-g ained, pixel-le el change de ec ion. In u u e
wo k, we aim o ex end ou modeling app oach and analysis,
also a ge ing ou -o -dis ibu ion cases, wi h me hodological
and a chi ec u al e inemen s. He e, also ew-sho lea ning
me hods p o ide in e es ing u u e di ec ions, e.g., [17],[18].
REFERENCES
[1] Y. Li, C. L. P. Chen, and T. Zhang, ‘‘A su ey on Siamese ne wo k:
Me hodologies, applica ions, and oppo uni ies,’’ IEEE T ans. A i . In ell.,
ol. 3, no. 6, pp. 994–1014, Dec. 2022.
[2] J. B omley, I. Guyon, Y. LeCun, E. Säckinge , and R. Shah, ‘‘Signa u e
e i ica ion using a ‘Siamese’ ime delay neu al ne wo k,’’ in P oc. Ad .
Neu al In . P ocess. Sys ., ol. 6, 1993, pp. 737–744.
130636 VOLUME 13, 2025
H. B. Swamina han e al.: Compa a i e Analysis o Deep Lea ning-Based Fea u e Ex ac o s
[3] S. Chop a, R. Hadsell, and Y. LeCun, ‘‘Lea ning a simila i y me ic
disc imina i ely, wi h applica ion o ace e i ica ion,’’ in P oc. IEEE
Compu . Soc. Con . Compu . Vis. Pa e n Recogni . (CVPR), ol. 1,
Sep. 2005, pp. 539–546.
[4] L. Mou, M. Schmi , Y. Wang, and X. Xiang Zhu, ‘‘A CNN o he
iden i ica ion o co esponding pa ches in SAR and op ical image y
o u ban scenes,’’ in P oc. Join U ban Remo e Sens. E en (JURSE),
Ma . 2017, pp. 1–4.
[5] H. B. Swamina han, A. Somme , U. Iu gel, A. Becke , and M. A zmuelle ,
‘‘Change de ec ion in au omo i e ada based occupancy maps using
Siamese ne wo ks,’’ in P oc. In . Rada Symp., Jul. 2024, pp. 56–61.
[6] P. Qi, W. Zhou, and J. Han, ‘‘A me hod o s ochas ic L-BFGS
op imiza ion,’’ in P oc. IEEE 2nd In . Con . Cloud Compu . Big Da a Anal.
(ICCCBDA), Ap . 2017, pp. 156–160.
[7] A. K izhe sky, I. Su ske e , and G. E. Hin on, ‘‘ImageNe classi ica ion
wi h deep con olu ional neu al ne wo ks,’’ in P oc. Ad . Neu al In .
P ocess. Sys ., ol. 60, 2017, pp. 84–90.
[8] C. Szegedy, W. Liu, Y. Jia, P. Se mane , S. Reed, D. Anguelo , D. E han,
V. Vanhoucke, and A. Rabino ich, ‘‘Going deepe wi h con olu ions,’’
in P oc. IEEE Con . Compu . Vis. Pa e n Recogni . (CVPR), Jun. 2015,
pp. 1–9.
[9] K. Simonyan and A. Zisse man, ‘‘Ve y deep con olu ional ne wo ks o
la ge-scale image ecogni ion,’’ 2014, a Xi :1409.1556.
[10] K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Deep esidual lea ning o image
ecogni ion,’’ in P oc. IEEE Con . Compu . Vis. Pa e n Recogni . (CVPR),
Jun. 2016, pp. 770–778.
[11] A. Doso i skiy, L. Beye , A. Kolesniko , D. Weissenbo n, X. Zhai,
T. Un e hine , M. Dehghani, M. Minde e , G. Heigold, S. Gelly, J. Uszko-
ei , and N. Houlsby, ‘‘An image is wo h 16x16 wo ds: T ans o me s o
image ecogni ion a scale,’’ 2020, a Xi :2010.11929.
[12] A. Vaswani, N. Shazee , N. Pa ma , J. Uszko ei , L. Jones, A. N. Gomez,
Ł. Kaise , and I. Polosukhin, ‘‘A en ion is all you need,’’ in P oc. Ad .
Neu al In . P ocess. Sys ., ol. 30, 2017, pp. 5998–6008.
[13] Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang,
L. Dong, F. Wei, and B. Guo, ‘‘Swin ans o me 2: Scaling up capaci y
and esolu ion,’’ in P oc. IEEE/CVF Con . Compu . Vis. Pa e n Recogni .
(CVPR), Jun. 2022, pp. 12009–12019.
[14] T. Xiao, M. Singh, E. Min un, T. Da ell, P. Dollá , and R. Gi shick,
‘‘Ea ly con olu ions help ans o me s see be e ,’’ in P oc. Ad . Neu al
In . P ocess. Sys ., 2021, pp. 30392–30400.
[15] J. Deng, W. Dong, R. Soche , L.-J. Li, K. Li, and L. Fei-Fei, ‘‘ImageNe :
A la ge-scale hie a chical image da abase,’’ in P oc. IEEE Con . Compu .
Vis. Pa e n Recogni ., Jun. 2009, pp. 248–255.
[16] R. R. Sel a aju, M. Cogswell, A. Das, R. Vedan am, D. Pa ikh, and
D. Ba a, ‘‘G ad-CAM: Visual explana ions om deep ne wo ks ia
g adien -based localiza ion,’’ in P oc. IEEE In . Con . Compu . Vis.
(ICCV), Oc . 2017, pp. 618–626.
[17] G. Cheng, L. Cai, C. Lang, X. Yao, J. Chen, L. Guo, and J. Han, ‘‘SPNe :
Siamese-p o o ype ne wo k o ew-sho emo e sensing image scene
classi ica ion,’’ IEEE T ans. Geosci. Remo e Sens., ol. 60, 2022.
[18] J. J. Vale o-Mas, A. J. Gallego, and J. R. Rico-Juan, ‘‘An o e iew
o ensemble and ea u e lea ning in ew-sho image classi ica ion
using Siamese ne wo ks,’’ Mul imedia Tools Appl., ol. 83, no. 7,
pp. 19929–19952, Jul. 2023.
HARIHARA BHARATHY SWAMINATHAN was
bo n in Chennai, India, in 1993. He ecei ed he
B.Tech. deg ee in elec onics and communica-
ion enginee ing om he B. S. Abdu Rahman
C escen Ins i u e o Science and Technology,
Tamil Nadu, India, in 2014, and he M.Eng.
deg ee in embedded sys ems o mecha onics
om he Uni e si y o Applied Sciences (FH),
Do mund, Ge many, in 2020. He is cu en ly
pu suing he Ph.D. deg ee wi h he Seman ic
In o ma ion Sys ems (SIS) G oup, Osnab ück Uni e si y, Ge many. His
esea ch in e es s include applica ion o machine lea ning and deep lea ning
o en i onmen pe cep ion o sel -d i ing ca s, pa icula ly using da a om
au omo i e ada s.
ARON SOMMER was bo n in Be lin, Ge many,
in 1986. He ecei ed he Dipl.-Ma h. Techn.
deg ee in Technoma hema ics wi h a echnical
backg ound in communica ions enginee ing om
Ka ls uhe Ins i u e o Technology (KIT), Ka l-
s uhe, Ge many, in 2013, and he Ph.D. deg ee
om he Ins i u e o In o ma ion P ocessing, Leib-
niz Uni e si y Hanno e , in 2019, whe e he was
in ol ed in a p ojec on syn he ic ape u e ada .
Since 2020, he has been he Rada Algo i hm
Expe o au omo i e applica ions wi h Ap i .
URI IURGEL ecei ed he Dipl.-Ing. deg ee in
elec ical enginee ing (wi h a ocus on in o ma ion
echnology) om he Uni e si y o Duisbu g,
Ge many, in 2000, and he D .-Ing. deg ee om
he Technical Uni e si y o Munich, in 2006.
Since 2005, he has been wi h Delphi/Ap i ,
cu en ly holding he posi ion o he Manage
Rada Pe cep ion and a Senio Expe Senso
Algo i hms. His esea ch in e es s include a i-
icial in elligence, in o ma ion e ie al, na u al
language p ocessing, algo i hms o en i onmen pe cep ion wi h came a and
ada , and ada signal p ocessing, and speci ically ela es o applica ions o
au onomous d i ing and ad anced d i e assis ance sys ems.
ANDREAS BECKER (Membe , IEEE) was bo n
in Wuppe al, Ge many, in 1975. He ecei ed
he Dipl.-Ing. deg ee in elec ical enginee ing
and he D .-Ing. deg ee om he Uni e si y o
Wuppe al, in 2000 and 2006, espec i ely. Du ing
he doc o al deg ee, his esea ch was ocused
on he nume ical simula ion o elec omagne ic
p oblems, including he analysis o bo ehole
ada s. F om 2007 o 2016, he was wi h Hella
GmbH & Co. KGaA, and Delphi Deu schland
GmbH (now Ap i ), e en ually achie ing he posi ion o he Rada Technical
Manage . Since 2017, he has been a P o esso o in o ma ion echnology
wi h he Uni e si y o Applied Sciences Do mund, Do mund, Ge many. His
esea ch in e es includes pe cep ion and con ol sys ems o mobile obo s.
MARTIN ATZMUELLER is cu en ly a Full
P o esso wi h he Ins i u e o Compu e Science,
Osnab ück Uni e si y, Ge many, whe e he holds
he ROSEN-G oup-Endowed Chai o seman-
ic in o ma ion sys ems. He is he Founding
Spokespe son wi h he Join Labo a o y on A i-
icial In elligence and Da a Science, a Founding
Membe wi h he Resea ch Uni Da a Science,
Osnab ück Uni e si y, and he Scien i ic Di ec o
wi h he Resea ch Depa men Coope a i e and
Au onomous Sys ems, Ge man Resea ch Cen e o A i icial In elli-
gence (DFKI). His esea ch in e es s include a i icial in elligence (AI), da a
science, and in eg a i e AI sys ems, whe e his pa icula esea ch in e es s
include modeling complex da a, explainable AI, in e p e able lea ning,
machine pe cep ion, and seman ic in e p e a ion, and ela es o applica ions
in complex in eg a i e AI sys ems, especially obo con ol and senso -based
AI sys ems.
VOLUME 13, 2025 130637