scieee Science in your language
[en] (orig)

Independent Channel Residual Convolutional Network for Gunshot Detection

Abstract

The main purpose of this work is to propose a robust approach for dangerous sound events detection (e.g. gunshots) to improve recent surveillance systems. Despite the fact that the detection and classification of different sound events has a long history in signal processing, the analysis of environmental sounds is still challenging. The most recent works aim to prefer the time-frequency 2-D representation of sound as input to feed convolutional neural networks. This paper includes an analysis of known architectures as well as a newly proposed Independent Channel Residual Convolutional Network architecture based on standard residual blocks. Our approach consists of processing three different types of features in the individual channels. The UrbanSound8k and the Free Firearm Sound Library audio datasets are used for training and testing data generation, achieving a 98 % F1 score. The model was also evaluated in the wild using manually annotated movie audio track, achieving a 44 % F1 score, which is not too high but still better than other state-of-the-art techniques.

Read accessible full text

Independent Channel Residual Convolutional Network for Gunshot Detection

Author: Bajzík, Jakub; Přinosil, Jiří; Jarina, Roman; Mekyska, Jiří
Publisher: Science and Information Organization
Year: 2022
DOI: 10.14569/IJACSA.2022.01304108
Source: https://dspace.vut.cz/bitstreams/b44e63d1-b067-4119-84c0-6e669b332e85/download
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
Independen Channel Residual Con olu ional
Ne wo k o Gunsho De ec ion
Jakub Bajzik1, Ji i P inosil2, Roman Ja ina3, Ji i Mekyska4
Dep . o Mecha onics and Elec onics, Uni e si y o Zilina
Zilina 010 26, Slo akia1
Dep . o Telecommunica ions, B no Uni e si y o Technology
601 90 B no, Czech Republic2,4
Dep . o Mul imedia and In o ma ion and Communica ion Technology
Uni e si y o Zilina, Zilina 010 26, Slo akia3
Abs ac —The main pu pose o his wo k is o p opose
a obus app oach o dange ous sound e en s de ec ion (e.g.
gunsho s) o imp o e ecen su eillance sys ems. Despi e he
ac ha he de ec ion and classi ica ion o di e en sound
e en s has a long his o y in signal p ocessing, he analysis o
en i onmen al sounds is s ill challenging. The mos ecen wo ks
aim o p e e he ime- equency 2-D ep esen a ion o sound
as inpu o eed con olu ional neu al ne wo ks. This pape
includes an analysis o known a chi ec u es as well as a newly
p oposed Independen Channel Residual Con olu ional Ne wo k
a chi ec u e based on s anda d esidual blocks. Ou app oach
consis s o p ocessing h ee di e en ypes o ea u es in he
indi idual channels. The U banSound8k and he F ee Fi ea m
Sound Lib a y audio da ase s a e used o aining and es ing
da a gene a ion, achie ing a 98 % F1 sco e. The model was also
e alua ed in he wild using manually anno a ed mo ie audio
ack, achie ing a 44 % F1 sco e, which is no oo high bu s ill
be e han o he s a e-o - he-a echniques.
Keywo ds—Acous ic signal p ocessing; gunsho de ec ion sys-
ems; audio signal analysis; machine lea ning; deep lea ning;
esidual ne wo ks
I. INTRODUCTION
In he ield o signal p ocessing, he audio da a analysis
akes an ex ensi e pa , which is cons an ly s udied. Many
machine lea ning-based algo i hms we e p oposed o sol ing
asks such as classi ica ion, segmen a ion, and denoising. In
many cases, he me hods a e adap ed o a speci ic ype o
sound, mainly speech and music, which ake an ex ensi e
pa in he esea ch. On he o he hand, he en i onmen al
sounds a e uns uc u ed, and i is challenging o gene alize
hei na u e. Many en i onmen al sound analysis applica ions,
anging om u ban moni o ing [1] o IoT [2] and su eillance
sys em [3], [4], [5], ha e been de eloped wi hin he pas yea s.
Howe e , in ecen yea s i has become a opical ask o
use known echniques o classi y en i onmen al sounds like
explosion, gunsho , si en, ca ala m, baby c ying, window
b eakage and o he e en s associa ed wi h po en ial dange
[6]. Usage o lea ning algo i hms ends o inc ease pe sonal
sa e y. The possible implemen a ions a e in-home o indus ial
p o ec ion sys ems, in ca s o ale dea o poo ly hea ing
d i e s o he si en, in homes o ale a pa en o a c ying child,
and in a wide ange o assis i e de ices, especially o dea
people. Recen mode n su eillance sys ems o isk p e en ion
pu poses ocuses mainly on he analysis o ideo signals om
came as using ad anced compu e ision echniques [7], [8],
[9]. Howe e , he analysis o audio signals has conside able
po en ial in hese sys ems as well. Especially gunsho de ec ion
echnologies ha e been inc easingly adop ed by law en o ce-
men agencies o mapping he spa ial and empo al pa e ns
o gun iolence [10]. The gene al p oblem o gun de ec ion
echnologies is a high a e o alse ala ms esul ing in he was e
o police esou ces when esponding o hose alse ale s [11].
F om a p ac ical poin o iew, i is necessa y o minimize he
amoun o alse-posi i e p edic ions. The e o e, when se ing
he ope a ing poin o he sys em in p ac ical applica ions, he
alse-posi i e a e mus be aken in o accoun .
The main objec i e o ou wo k is o p opose a me hod o
de ec ing dange - ela ed audio e en s (gunsho s) ha achie e
high speci ici y in eal condi ions. We also aim o explo e
se e al ypes o ea u e spaces and neu al a chi ec u es and
discuss, wha kind o se up is he mos sui able o such an
applica ion.
The s anda d app oach o sound analysis is o collec a
single ec o o ea u es. The o en used classi ie s a e Sup-
po Vec o Machines (SVM), Deep Neu al Ne wo ks (DNN),
o mul ilaye pe cep ons. When p ocessing en i onmen al
sounds, he uni o m s uc u e can no be expec ed as in speech
o music, whe e he signal con ains a ha monic s uc u e o
epe i ions. The ea u es ha pe o m well in speci ic applica-
ions may be inadequa e o sounds wi h o he na u e and ice
e sa.
In ou wo k we a e using se e al 2-dimensional sound ep-
esen a ions such as spec og ams as well as he s anda d 1-D
app oach and analyse he pe o mance in he gunsho de ec ion
applica ion. In addi ion o he equen ly used spec og am
[12], [13], [14], also new isualiza ions and ad anced me hods
o ea u e p ocessing a e used. The main con ibu ions o ou
wo k a e:
•Explo ing he sui abili y o se e al s a e-o - he-a
con olu ional ne wo ks based app oaches o gunsho
de ec ion.
•P oposing he con olu ional model o boos ing he
pe o mance on 2-dimensional independen ea u e
spaces.
The signal is ans o med in o h ee independen audio ea u e
se s o ming an ”RGB image” ha is sui able o p ocessing
www.ijacsa. hesai.o g 950 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
by common 2-D con olu ional ne wo ks used o image p o-
cessing. Unlike image p ocessing, whe e colo channels a e
highly co ela ed and p ocessed oge he a e he i s DNN
laye , he h ee audio ea u e se s a e p ocessed independen ly
by he i s wo DNN blocks. Ou p oposed a chi ec u e is
based on s anda d esidual uni s. The impo an ask is no only
inc easing he numbe o ue p edic ions bu also educing he
numbe o alse-posi i e gunsho p edic ions.
II. RELATED WORK
O e he yea s se e al wo ks dealing wi h gunsho s de-
ec ion om audio signal ha e been published. Mos o hem
a e based on he ex ac ion o handc a ed acous ic ea u es
and he use o machine lea ning echniques o he ask o
classi ica ion. A combina ion o 7 Linea P edic i e Coding
(LPC) coe icien s and 13 Mel F equency Ceps al Coe icien s
(MFCC) wi h SVM classi ie is used in wo k [15] p o iding
8 % alse ala ms a e on a cus om da ase . The au ho o [16]
ex ended he p e ious ea u e se by Linea P edic i e Coding
Ceps al (LPCC) and au o-co ela ion coe icien s eaching
82 % accu acy and 70 % p ecision on a combina ion o public
a ailable da ase s. The same au ho hen e alua ed he e ec
o indi idual ea u es on he accu acy o he classi ica ion
ask [17], conside ing he ea u es wi h he bes sco e o
be he i s i e coe icien s o he 24 h o de LPC. A la ge
se o a ious acous ic ea u es wi h Hidden Ma ko Model
(HMM) and Vi e bi decode is used in EAR-TUKE sys em
[18] o de ec ing gunsho s and glass b eaking e en s wi h
98 % accu acy in eco ds wi h Signal o Noise Ra io (SNR)
≥20 dB. In case o mic ophones a ays i is possible o
use a wo s age me hodology comp ising o a Blind Sys em
Iden i ica ion and Decon olu ion (BSID) s age ollowed by a
SVM-based classi ica ion [19] o gunsho de ec ion in a noisy
u ban en i onmen . In [20] a me hod o classi ying impulsi e
sounds based on a Weigh ed Majo i y Vo ing (WMV) s a egy
is desc ibed. In [21] Con olu ional Neu al Ne wo k (CNN)
wi h empo al and spec al ea u es is used o gunsho sound
ca ego ies classi ica ion (pis ol, i le and sho gun o di e en
calib es) eaching o e 90 % accu acy. Ano he CNN app oach
de ec s gunsho s wi h 99 % accu acy and low alse ala m
a e using he ResNe a chi ec u e [22]. Mos o he abo e
app oaches wo k wi h da abase eco dings ha con ain a low
le el o en i onmen al noise. In he case o eal applica ions,
his condi ion can ha dly be me . Fo his eason, dealing wi h
he de ec ion and classi ica ion o en i onmen al noisy sounds
is impo an .
The da ase s o en i onmen al sounds a e made mainly
o lea ning algo i hms ha pe o m En i onmen al Sound
Classi ica ion (ESC) ask. One o he widely used da ase s o
ESC is he U banSound8K [23], u he desc ibed in Sec ion
III-E.
The signal ep esen a ion o audio is ela ed o he a -
chi ec u e o he lea ning algo i hm and lea ning objec i e.
The s anda d p ocess ha ollows he app oaches om speech
and music analysis is o collec a single ec o o ea u es.
Sho - e m o long- e m ea u es may no always gene alize
he uns uc u ed na u e o en i onmen al sounds.
The di ec solu ion o he ea u e ex ac ion p oblem is o
build a model ha ope a es on he aw audio signals di ec ly.
The 1-D CNNs can handle he in e nal ep esen a ion o he
inpu signal, which allows end- o-end usage. The impo an
ad an age is ha he e is no need o ans o m o p e-p ocess
he da a, and such a model can adap o a a ie y o audio
signals. In he s udies [24], [25], he i s end- o-end ESC
a chi ec u e called En Ne was p oposed in e sions 1 and
2. Ano he 1-D a chi ec u e was p oposed in [26], whe e he
audio signal was p ocessed a di e en ime scales. The s udy
[27] p esen s an end- o-end 1-D con olu ional ne wo k ha
has ewe pa ame e s compa ed o dense 2-D con olu ional
neu al ne wo ks and does no equi e a la ge amoun o aining
da a. I eaches 87 % mean accu acy on he U banSound8K
da ase wi h andom weigh s ini ializa ion and 89 % wi h
weigh s ini ializa ion by Gamma one il e bank coe icien s
[28] syn hesizing an impulse esponse om ne e cells in he
audi o y ibe [29].
The 2-D CNN models ope a e on he p e-compu ed ea u e
ep esen a ions ob ained by a ixed p ocess o ex ac ion. The
s udy [13] is he i s which deals wi h ESR using CNN ained
on mel-scaled spec og ams. Such an app oach is ex ended in
s udy [12], whe e di e en augmen a ion me hods a e used.
In wo k [14], he au ho s p esen ed he model ESResNe
based on STFT, ha ou pe o ms ecen ly known app oaches
wi h ESC da ase s achie ing 82% accu acy when ained om
sc a ch and 85% accu acy wi h ImageNe weigh s ini aliza ion
on U banSound8K da ase . Since he Piczak’s wo k [13], he
esea ch end in en i onmen al sound analysis seems o be he
usage o 2-D ea u e ma ices o eeding 2-D CNNs [30].
III. MATERIALS AND METHODS
A. Signal Model and P oblem Fo mula ion
Following he ecen s udies [12], [13], [14], [31], [32],
[33], [34], he mos sui able se up o en i onmen al sound
ecogni ion employs a 2-D con olu ional neu al ne wo k ed
by a Time-F equency ep esen a ion o he audio signal ( u he
men ioned as audio ea u es). In o de o be p ocessed by a
2-D con olu ional ne wo k, hese audio ea u es need o be
con e ed in o a sui able uni o m 2-D ep esen a ion. Based
on his 2-D ep esen a ion, i is hen necessa y o choose
an op imal a chi ec u e o he con olu ional ne wo k o he
classi ica ion ask. The whole wo k low om audio samples
o p edic ions is depic ed in Fig. 1.
Fig. 1. Fea u e Ex ac ion and Classi ica ion Wo k low.
B. Audio Fea u es
In he case o audio signal p ocessing, he e is a la ge num-
be o a ious audio ea u es. The Log-Mel Spec og ams (LM
Spec) and Mel F equency Ceps al Coe icien s (MFCC) a e
among he mos commonly used audio ea u es. In ou wo k,
we addi ionally include he Sel -Simila i y Ma ix (SSM),
www.ijacsa. hesai.o g 951 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
which is equen ly used o analyze he global s uc u e o
musical wo ks, o he audio ea u e lis . The hype pa ame e s
o ea u e ex ac ion a e de ailed in Table I.
1) Log-Mel Spec og am: The spec og am is he mos
commonly used audio signal isualiza ion. I shows he e-
quency spec um change o e ime. A Sho Time Fou ie
T ans o m (STFT) is used o con e he signal om ime o
equency domain. Addi ional mel- equency scale ans o m
1 is applied o emb ace he psychoacous ic knowledge.
mel = 2595 ·log10 1 + Hz
700 (1)
2) Mel F equency Ceps al Coe icien : MFCC e lec s
he non-linea and masking psychoacous ic cha ac e is ics o
human hea ing. MFCC coe icien s a e ob ained by mul iplying
he signal spec um by a mel-scale dis ibu ed il e bank,
loga i hm and Disc e e Cosine T ans o m (DCT).
3) Sel -Simila i y Ma ix: SSM is he measu e o sel -
simila i y o he signal based on dis ances. We use a sel -
simila i y ma ix o display signal co ela ion. To isualize he
sel -simila i y, we use a ma ix Sde ined by Equa ion 2. The
ma ix dimensions N×Ndepend on he numbe o signal
samples. Fo educing he compu a ional complexi y, a sel -
simila i y ma ix Sis compu ed on downsampled en elope
s= (s1, s2, s3, ..., sN)o inpu audio signal, ob ained using
he Hilbe ans o m. As a measu e o simila i y we a e using
he absolu e dis ance.
S(i, j) = |si−sj|i, j = 1, ..., N (2)
The e ical and ho izon al axes ep esen he ime sequence.
The ma ix is symme ical by he main diagonal whe e he
simila i y is maximal.
TABLE I. HYPERPARAMETERS FOR FEATURE EXTRACTION
Fea u es FFT leng h Banks Window
Log-mel spec og am 2048 256 Hamming
MFCC 2048 20 Hamming
Sel -simila i y - - -
C. 2-D Fea u e Rep esen a ion
Audio ea u es a e ex ac ed as 2-D ma ices and aligned
o con olu ional neu al ne wo k inpu . Since mos 2D con-
olu ional ne wo k a chi ec u es we e p ima ily designed o
image p ocessing, hey expec 3 se s o 2-D ea u e ma ices
a he inpu (an analogy o RGB channels o images).
1) Band-Spli ed Spec og am (BS Spec): In ou expe i-
men s, we a e using he log-mel spec og am spli o h ee e-
quency bands (high, middle, low) aligned wi h he RGB colo
channels (each band as one colo channel). The band cu ing
equencies depends he on maximal equency max = s
2
gi en by he sampling a e (Fig. 2).
The same p inciple was used in s udy [14], whe e au ho s
explain he usage o he band-spli ed spec og am o a oiding
edundancy. O he solu ions a e eplica ing he spec og am o
passing ze os.
Fig. 2. Band-spli Log-mel Spec og am and Resul ing RGB Image.
2) Independen Fea u e Spaces (IFS): The combina ion o
he log-mel spec og am, MFCC and sel -simila i y ma ix
ep esen s he independen ea u e spaces. The simila me hod
was used in s udy [32]. We assume ha he MFCC and SSM
will help o classi y non-impulsi e backg ound sounds. The
ha monici y o he gunsho signal is low, so he SSM is almos
emp y, while he backg ound noise esul s in a isible g id.
Th ee ea u e ma ices a e o e lapped in ma ched ime po-
si ions. I means, ha he x-axis esolu ions a e app oxima ely
he same o all ma ices. Howe e , on he y-axis we ha e
di e en dimensions when using spec og am ( equency),
MFCC (mel banks) and SSM ( ime) (Fig. 3).
(a) (b) (c) (d)
Fig. 3. Independen Fea u e Spaces as RGB Image Channels. (a) Red
Channel, Log-mel Spec og am. (b) G een Channel, MFCCs. (c) Blue
Channel, Sel -simila i y Ma ix. (d) Resul ing RGB Image.
D. Con olu ional Neu al Ne wo ks A chi ec u es
The mos widely used con olu ional neu al ne wo k models
o 2-D ea u e space classi ica ion a e based on he esidual
ne wo k a chi ec u e (ResNe ). Howe e , he e is also an
app oach ha uses only a 1-dimensional con olu ional ne wo k
ed di ec ly by a aw audio signal o classi y en i onmen al
sounds. In addi ion, we include a cus om app oach based on
esidual ne wo ks whe e indi idual channels a e p ocessed
independen ly.
1) Residual Ne wo ks: Following ecen s udies [14], he
esidual models pe o m well on en i onmen al sound clas-
si ica ion. The esidual ne wo k was designed as a ne wo k
in a ne wo k, which means ha he lowe laye ’s inpu s a e
connec ed o he ou pu s o he wo highe laye s. The example
o he s anda d esidual block is shown in Fig. 4.
The skip connec ions de ined as
y=F(x) + x(3)
a e also called sho cu connec ions. The unc ion F ep-
esen s he con olu ion ope a ions. The sho cu connec ions
help o elimina e he p oblem o anishing g adien in deep
neu al ne wo ks. The au ho s o [35] designed ResNe s wi h
a di e en numbe o laye s, speci ically 18, 34, 50, 101 and
152.
www.ijacsa. hesai.o g 952 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
Fig. 4. S anda d Residual Block [35].
2) End- o-End Classi ica ion using a 1-D CNN: In his
app oach, a 1-D con olu ional (1-D CNN) neu al ne wo k
lea ns low-le el and high-le el in o ma ion di ec ly om he
audio signal wa e o m. Since he size o he inpu da a o
he amoun o da a is in imbalance, i is no ecommended
o use oo deep con olu ional ne wo k a chi ec u es o a oid
signi ican o e i ing. The s udy [27] p esen s he op imal
a chi ec u e wi h espec o he sampling equency o he
audio signal a di e en audio signal leng hs. Fo he sampling
equency o 16 kHz conside ed u he in his pape , i has
been shown ha he bes sco e is achie ed a he signal leng h
o 1 second. The co esponding CNN a chi ec u e is shown
in Table II consis ing o 4 Con olu ional Laye s (CL), 2
Pooling Laye s (PL) and 2 Fully Connec ed laye s (FC). The
Rec i ied Linea Uni (ReLU) ac i a ion unc ion is used o all
laye s, excep o he ou pu laye whe e he so max ac i a ion
unc ion is used wi h he ou pu size equal o he numbe o
classes beeing classi ied.
TABLE II. ARCHITECTURE OF 1-D CNN FOR 16 KHZSAMPLING RATE
AND AUDIO LENGTH OF 1 SECOND [27]
CL1 PL1 CL2 PL2 CL3 CL4 FC1 FC2
Dimension 7969 996 483 60 23 8 128 64
Fil e s coun 16 16 32 32 64 128 - -
Fil e s size 64 8 32 8 16 8 - -
S ide size 2 8 2 8 2 2 - -
3) Independen Channel Residual Con olu ional Ne wo k:
We p opose he Independen Channel Residual Con olu ional
Ne wo k (ICRCN), whe e he inpu RGB image is di ided o
he 2-D ma ices in indi idual channels. The ea u e ma ices
sha e he dimension o he x-axis ( ime) bu no he y-
axis ( equency, mel banks, ime). The sys em ha combines
di e en isual ep esen a ions may su e , when he ea u es
a e combined as one inpu image. The e o e, we build he
esidual con olu ional ne wo k, ha p ocesses di e en audio
isualiza ions sepa a ely. The whole model a chi ec u e is
shown in Table III.
The model inpu is a h ee channel RGB image. The
sepa a e channels con ain esidual blocks, whe e he numbe
o il e s is 32. The ea u e dimensions me ging is made
a e he second esidual block. F om his poin , he ea u es
a e p ocessed as in s anda d esidual con olu ional ne wo ks.
The las con olu ional block consis s o 512 il e s and i is
ollowed by he classi ica ion laye . The p oposed a chi ec u e
is buil up om s anda d esidual blocks, as desc ibed in
TABLE III. PROPOSED INDEPENDENT CHANNEL RESIDUAL
CONVOLUTIONAL NETWORK
Ou pu size ICRCN blocks
(224, 224, 3) Inpu RGB image
(112, 112, 32) 7×7, 32 7×7, 32 7×7, 32
(56, 56, 32) 3x3, 32
3x3, 32 ×23x3, 32
3x3, 32 ×23x3, 32
3x3, 32 ×2
(28, 28, 64) 3x3, 64
3x3, 64 ×23x3, 64
3x3, 64 ×23x3, 64
3x3, 64 ×2
(28, 28, 192) Conca ena ion
(14, 14, 256) 3x3, 256
3x3, 256 ×2
(7, 7, 512) 3x3, 512
3x3, 512 ×2
(512) Global a e age pooling
(2) Dense 2 + so max
III-D1. The p oposed a chi ec u e is compa ed o s anda d
esidual ne wo ks ResNe 50 and 1-D CNN in Table IV.
TABLE IV. COMPARISON OF ARCHITECTURES COMPLEXITY
A chi ec u e Numbe o ainable pa ame e s
1-D CNN 256k
ResNe 50 23.5M
ICRCN 11M
E. Da ase s
One o he mos widely known da ase s o en i onmen al
sound classi ica ion is he U banSound8K [23]. In his wo k
we build ou own gunsho de ec ion da ase as a combina ion
o he U banSound8k and The F ee Fi ea m Sound Lib a y
[36].
•U banSound8k - 8732 acks o 10 classes (ai
condi ione , ca ho n, child en playing, dog ba king,
d illing, engine idling, gunsho , jackhamme , si en,
s ee music), wi h a ying sampling equency ile o
ile.
•The F ee Fi ea m Sound Lib a y - 2200 acks o
gunsho s in a noise ee en i onmen including hand-
guns (pis ols, e ol e s, semi-au oma ic pis ols), i les
(le e -ac ion, semi-au oma ic, ully au oma ic, ma-
chine guns, e c.) and sho guns. Reco dings a e in loss-
less wa o ma wi h sampling equency 44.1 kHz.
The examples om The F ee Fi ea m Sound Lib a y we e
combined wi h gunsho s om he U banSound da ase . The
o he sounds om U banSound we e used as nega i e exam-
ples ( andom backg ound). Since U banSound eco ds a e no
equal in leng h, he window size o ou examples loa s in he
ange om 1 s o 4 s.
In he ield o supe ised machine lea ning, he pe o -
mance o algo i hms s ongly depends on he quali y o he
da ase . To some ex en , i is possible o simula e a big da ase
using augmen a ion me hods. Howe e , his app oach will
ne e be as good as he expansion o he da ase by eal
examples. The ollowing augmen a ion echniques a e applied
o posi i e examples (gunsho s) while aining.
www.ijacsa. hesai.o g 953 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
•Random backg ound mixing - The gunsho we e
andomly mixed wi h backg ound noise in a andom
SNR om 0 dB o 20 dB.
•Random ime shi - The onse posi ions o gunsho s
wi hin he p ocessing window we e chosen andomly
om 0 o 0.8 % o he window leng h.
•Random Gaussian noise addi ion - SNR om 60 dB
o 100 dB.
F. E alua ion Me ics
The so max ou pu ac i a ion unc ion di ec ly indica es
he p obabili y o he example belonging o a ce ain class. The
so max ac i a ion also gua an ees ha he sum o p obabili ies
o e classes (backg ound and gunsho ) is one. The examples
a e classi ied as posi i e o nega i e acco ding o highe o
lowe ac i a ion o ou pu neu ons. The con usion ma ix can
be cons uc ed using p edic ed and ue labels. The ea e , he
pe o mance is e alua ed ia known me ics.
Fo p ac ical implemen a ion o he sys em o dange ous
sound de ec ion (gunsho s, explosion, ...) i is necessa y o
minimize he amoun o alse-posi i e p edic ions. In his case,
he high accu acy alue may be a li le bi con using and hus
we ha e o choose a me ic o ele an e alua ion. The e o e,
in ou wo k we use he ollowing e alua ion me ics:
•Accu acy: I e lec s he o e all algo i hm pe o -
mance, so i also akes in o accoun he ue nega i e
p edic ions (TN). The accu acy is high when he num-
be o ue posi i e (TP) and ue nega i e p edic ions
is la ge. False posi i es (FP) and alse nega i es (FN)
a e p edic ion e o s.
Accu acy =T P +T N
T N +T P +F N +F P (4)
•Sensi i i y: The ela i e amoun o ue posi i e p e-
dic ions agains all posi i e examples. I is o en called
ecall o T ue Posi i e Ra e (TPR).
Sensi i i y =T P
T P +F N (5)
•Speci ici y: The ela i e amoun o ue nega i e p e-
dic ions agains all nega i e examples.
Speci ici y =T N
F P +T N (6)
•False-Posi i e Ra e (FPR): Re lec s he ela i e
amoun o alse posi i es agains all nega i e exam-
ples.
FPR =F P
F P +T N (7)
•F1 sco e: The balance alue be ween sensi i i y and
p ecision, whe e p ecision is he amoun o ue pos-
i i es agains all posi i e p edic ions. This means ha
he F1 sco e is no dis o ed by a la ge numbe o
ue nega i e p edic ions and may be conside ed as
decisi e, when es ing da a a e unbalanced.
F1 = 2T P
2T P +F P +F N (8)
•De ec ion E o T adeo - DET cu e isualizes
he alse posi i e a e s. he alse nega i e a e.
S anda dly, he axes a e scaled non-linea ly. Unlike
he ROC cu e, he DET cu es a e mo e linea and
si ua ed in mos o he plo a ea.
•Equal E o Ra e - EER is de ined as he poin in
DET o ROC whe e he e o s a e equal. The lowe
EER alue e lec s be e pe o mance o he sys em.
IV. RESULTS
All he ne wo ks we e ained as long as alida ion loss was
dec easing. Ea ly s op was applied and only he bes model was
sa ed. Adam op imize was used while aining he models and
he ca ego ical c oss-en opy loss was compu ed in each ba ch
o 16 examples. Ha dwa e used is NVIDIA GeFo ce GTX
1650, In el Co e i9-10900X CPU 3.7 GHz, RAM 64 GB. Fig. 5
shows he aining and alida ion his o y o ou ICRCN model
when using 16 kHz sampling equency.
Fig. 5. T aining and Valida ion His o y o ICRCN Model.
We use Tenso low and Ke as o building and aining he
models. Fo augmen ing he audio we use he Audiomen a ions
lib a y. The Sci-ki lea n is used o e alua ion using desc ibed
me ics.
A. 2-D Con olu ional Ne wo ks
The aining audio examples we e gene a ed in h ee sam-
pling equencies o es he e ec o sampling on sys em
pe o mance. The o e all F1 sco e and FPR is e alua ed in
Fig. 6. As he es subse is balanced, he F1 sco e and o e all
accu acy me ics a e e y simila . As seen, he high F1 sco e
alue does no necessa ily gua an ee good pe o mance in he
educ ion o alse-posi i es.
When looking a he F1 sco e o indi idual models, he
di e ence be ween 44.1 kHz and 16 kHz is no signi ican .
In se e al cases, he sco e d ops signi ican ly a he 8 kHz
sampling. The signi ican d op o F1 sco e is seen in case
o ICRCN ained on band-spli spec og ams. Table V, Table
VI, and Table VII show also he change in sensi i i y and
speci ici y o e all combina ions.
TABLE V. TESTING RESULTS FOR SAMPLING FREQUENCY 8KHZ
Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%]
ICRCN
IFS 93.1 94.8 91.3
BS Spec 85.1 74.0 100.0
ResNe 50
IFS 95.3 98.4 91.9
LM Spec 96.0 93.2 99.0
BS Spec 96.9 97.1 96.7
MFCC 84.8 91.5 75.8
SSM 96.0 94.2 98.1
www.ijacsa. hesai.o g 954 |Page

(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
(a) (b)
(c) (d)
(e) ( )
Fig. 6. Resul s o Di e en Sampling F equencies. (a) F1 Sco e, 8 kHz. (b)
FPR, 8 kHz. (c) F1 Sco e, 16 kHz. (d) FPR, 16 kHz. (e) F1 Sco e, 44.1 kHz.
( ) FPR, 44.1 kHz.
TABLE VI. TESTING RESULTS FOR SAMPLING FREQUENCY 16 KHZ
Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%]
ICRCN
IFS 97.9 98.4 97.3
BS Spec 93.9 99.6 87.4
ResNe 50
IFS 95.9 95.5 96.3
LM Spec 97.8 98.8 96.7
BS Spec 97.1 99.0 95.2
MFCC 87.0 93.8 78.1
SSM 95.9 95.5 96.3
TABLE VII. TESTING RESULTS FOR SAMPLING FREQUENCY 44.1 KHZ
Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%]
ICRCN
IFS 98.7 98.4 99.0
BS Spec 97.3 99.6 95.0
ResNe 50
IFS 97.2 95.2 99.4
LM Spec 98.1 97.5 98.8
BS Spec 99.0 100.0 98.1
MFCC 60.9 44.4 98.6
SSM 95.9 94.6 97.3
As he so max ou pu ac i a ion unc ion is used, he
ou pu gi es he p obabili ies o he example belonging o a
ce ain class. Ne e helss, he ope a ing poin can be mo ed
by changing he h eshold o posi i e p edic ions. Choosing
he ope a ing poin is s ongly ela ed o he applica ion o
he de eloped sys em and he e a e di e en ways how o se
up he ope a ing poin . The use ul isualiza ions a e he DET
cu es (Fig. 7), which end o be mo e linea and highligh
he di e ences in he ope a ing egion mo e clea ly han he
ROC cu es. The EER alue ep esen s he ope a ing poin ,
whe e he sys em pe o ms equal in e ms o he alse-posi i e
a e and alse-nega i e a e. The alse-posi i e a e can be
in e p e ed as he alse ala m p obabili y and alse-nega i e
a e is he p obabili y o missing he posi i e de ec ion.
(a) (b)
(c) (d)
(e) ( )
Fig. 7. De ec ion E o T adeo Cu es o Di e en Models and Sampling
F equencies. (a) ICRCN, 8 kHz. (b) ResNe 50, 8 kHz. (c) ICRCN, 16 kHz.
(d) ResNe 50, 16 kHz. (e) ICRCN, 44.1 kHz. ( ) ResNe 50, 44.1 kHz.
As seen in Table 7, he EER e alua ion sligh ly co ela es
wi h he o e all F1 sco e o he sys em. Though he e is
no e iden eg ess o e he pa ame e s, we can highligh
se e al indings om he esul s. Using he o iginal sampling
equency when he da a a e clea and de ailed, all a chi ec u es
model he da a well. The mo e ainable pa ame e s seem o be
use ul as he sys em losses he in o ma ion by audio downsam-
pling. A 16 kHz sampling, he pe o mance o a chi ec u es is
closely compa able. The p oposed ICRCN pe o ms bes wi h
IFS ea u e combina ion. Acco ding o s udy [32], he single
SSM pe o ms wo s on en i onmen al sounds classi ica ion.
Ou esul s showed, ha o gunsho de ec ion, using he SSM
only may leads o be e sco es han single MFCC, which
su p isingly pe o ms wo s .
www.ijacsa. hesai.o g 955 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
B. 1-D Con olu ional Ne wo k
In his pa o wo k, we compa e he 1-D CNN model o
he bes 2-D models. Fo his compa ison we conside only
he 16 kHz sampling equency. As he bes candida es om
he p e ious es we choose wo model- ea u es combina ions,
namely ICRCN-IFS and ResNe 50-LM Spec. The compa ison
is shown in Table VIII. The 1-D CNN model was ained on
he same da ase as 2-D models.
TABLE VIII. THE RESULTS FOR 1-D CNN MODEL IN COMPARISON TO
THE BEST 2-D MODELS
Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%]
ICRCN IFS 97.9 98.4 97.3
ResNe 50 LM Spec 97.8 98.8 96.7
1-D CNN - 98.1 98.2 98.1
C. E alua ion o Pe o mance in he Wild
The esul s on a i icially gene a ed es subse s may no
always mee he eal pe o mance in he wild. The e o e, we
e alua e he models on a eal audio ack om ac ion mo ie
o simula e he eal wo ld condi ions. On his s age we use
only 16 kHz sampling equency as he mos sui able based on
p e ious esul s. The audio ack om he mo ie John Wick
(2014) was manually anno a ed. We use a 4 seconds window
leng h wi h 2 seconds o e lap and hi d-o de median il e ing
o he CNN ou pu . Each segmen is labeled as posi i e i he e
is an occu ence o a gunsho . The numbe s o posi i e and
nega i e segmen s is e y unbalanced, he e o e we use he
F1 sco e as he e alua ion me ic. The e alua ion esul s a e
shown in Table IX.
TABLE IX. EVALUATION ON REAL DATA IN THE WILD
Model Fea u es Tes E alua ion
F1[%] F1[%] Sensi i i y[%] Speci ici y[%]
ICRCN BS Spec 93.9 14.5 98.2 10.5
IFS 97.9 44.3 59.2 91.6
ResNe 50
LM Spec 97.8 26.4 89.0 62.5
MFCC 87.0 16.5 47.2 67.0
SSM 95.9 5.5 5.5 92.7
BS Spec 97.1 17.6 98.2 29.2
IFS 95.9 16.2 95.4 24.3
1-D CNN - 98.1 4.2 7.8 96.5
As shown in Table IX, he pe o mance on eal da a d ops
signi ican ly. Howe e , he e alua ion es sco es co ela e in a
ela i e way. The es sco e di e ence is sligh in mos cases,
while he e alua ion sco e di e ence is no iceable. We can see,
ha single SSM ou pe o ms he MFCC ea u es in es esul s,
bu ails in e alua ion, whe e he sensi i i y and speci ici y a e
e y unbalanced. The esul s show, ha he p oposed ICRCN
ne wo k a chi ec u e was able o model he in a da a obus
way and ou pe o ms he s anda d esidual ne wo k ResNe 50
e en using less ainable pa ame e s. In he case o 1-D CNN,
he bad inal sco e was p obably caused by he low numbe
o pa ame e s o he con olu ional neu al ne wo k. Thus he
ne wo k was no able o gene alize he ex ac ed ea u es
su icien ly o deal wi h he high a iabili y o eal audio da a.
The p e ious expe imen mainly ocused on he abili y
o he p oposed algo i hm o de ec gunsho sounds in he
simula ed eal eco ding. Howe e , in he case o su eillance
sys ems, in addi ion o accu acy, a low equency o alse
ala ms is also equi ed. This alue canno be de i ed om
he abo e expe imen because he sound acks o he mo ies
a e sound-exposed. Thus, in his case, eco dings om he
QUT-NOISE da abase [37], which is he only sui able publicly
a ailable da abase con aining con inuous eco dings o ci y
sounds, we e used o de e mine he alse e o a e. Since he
o al eco ding ime is only a ew hou s, his da abase was
supplemen ed wi h cus om eco dings om a busy ci y s ee .
Wi hin hese eco dings, he p oposed algo i hm did no de ec
a single alse gunsho e en .
D. P ocessing Time
In Table X, we showed he mean p ocessing ime o each
o he ea u es used in he expe imen s. The gene a ion o h ee
ea u e ma ices akes wice as much ime han gene a ion o
a single spec og am. This mus be aken in o accoun i he
de ec o is implemen ed on he sys em wi h limi ed p ocessing
capaci y.
TABLE X. FEATURE EXTRACTION TIME
Sampling equency [kHz] Fea u es Time [ms]
8Spec 5.02
MFCC 2.68
SSM 2.31
16 Spec 6.23
MFCC 3.62
SSM 3.71
44.1 Spec 10.53
MFCC 6.9
SSM 10.58
When compa ing p ocessing ime be ween mul iple sam-
pling equencies, he good comp omise is o p e e he 16 kHz
sampling. I o e s he compa able pe o mance o o iginal
sampling while he ea u e ex ac ion ime is almos hal ed.
V. DISCUSSION
When compa ing he esidual a chi ec u es and ea u es
using es ing da a, he e is no a signi ican bes se up. The
e alua ion on eal da a showed he bene i o independen
ea u es p ocessing, whe e he di e ence in sco e is obse -
able. The s anda d ResNe 50 pe o ms well wi h a single
spec og am. The ea u e combina ion wo ks bes wi h he
p oposed a chi ec u e wi h independen channels. The e o e,
we p e e using he p oposed ICRCN a chi ec u e and IFS
o gunsho de ec ion. Taking he ea u e ex ac ion imes in o
accoun , we p e e using he 16 kHz sampling.
We assume, ha he ICRCN model ed by IFS ma ices is
he mos sui able o gunsho de ec ion sys ems, ou pe o ming
he o he s a e-o - he-a app oaches (in e ms o F1 sco e).
The combina ion o spec og am, MFCC, and SSM was used
in s udy [32], wi h conclusion ha i does no imp o e he
accu acy o en i onmen al sounds classi ica ion. We assume,
ha he ea u e combina ion can be bene icial o gunsho
de ec ion applica ions and has a po en ial o lowe he alse-
posi i e a e. The esul s showed, ha he e is a bene i o
spli ing he con olu ional channels. The eal pe o mance o
www.ijacsa. hesai.o g 956 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
he sys em ela es on he ope a ing poin . In some applica ions
he speci ici y may be a mo e impo an sco e han sensi i i y,
o ice e sa. We isualized he DET cu es, as hey can help
o i he equi emen s o he sys em o pa icula applica ions.
VI. CONCLUSION
In his pape , we analyzed he usage o se e al con-
olu ional a chi ec u es o he gunsho de ec ion ask. We
p oposed a new a chi ec u e and compa ed i s pe o mance
o s anda d ResNe 50 and 1-D a chi ec u es. The e ec o
di e en ea u es on he esul ing pe o mance was es ed using
se e al Time-F equency audio ep esen a ions. Fo aining,
alida ion and es ing we collec ed he gunsho s and andom
backg ounds audio clips om public da ase s. In he aining
phase, he s anda d augmen a ion me hods we e used. Finally,
we simula ed he eal wo ld condi ions by e alua ion o a eal
audio ack om he ac ion mo ie.
We achie ed he goal o ou wo k by p oposing a sys em
o gunsho audio e en s de ec ion. Ou ICRCN app oach is
able o ope a e in noisy en i onmen s wi h high speci ici y
(sligh ly o e 90 %), by main aining ai sensi i i y (almos
60 %). The De ec ion E o T adeo analysis showed, ha
he eal pe o mance o he sys em s ongly depends on he
e o ole ance and equi emen s. The ope a ing poin should
be selec ed o speci ic applica ions. On ou es da a, he
missed de ec ion p obabili y is abou 10 % when he alse
ala m p obabili y is as minimal as possible. We expec he
simila esul s also o explosion de ec ion. Such a sys ems
can be implemen ed in a complex applica ion oge he wi h
a smoke o i e de ec o . Due o he ac ha he p oposed
app oach has a ela i ely low compu a ional ime, i can be
easily in eg a ed in o exis ing su eillance sys ems wi hou he
need o in es in expensi e compu ing se e s o ope a e in eal
ime.
ACKNOWLEDGMENT
This wo k was suppo ed by he Slo ak G an Agency
KEGA unde con ac no. KEGA 008ZU-4/2021.
REFERENCES
[1] J. P. Bello, C. Sil a, O. No , R. L. DuBois, A. A o a, J. Sala-
mon, C. Mydla z, and H. Do aiswamy, “SONYC: A Sys em o
he Moni o ing, Analysis and Mi iga ion o U ban Noise Pollu ion,”
a Xi :1805.00889 [cs, eess], May 2018, a Xi : 1805.00889.
[2] S. C. Tan and A. Abd Mana , “Cha ac e iza ion o in e ne o hings (io )
powe ed-acous ics senso o indoo su eillance sound classi ica ion,”
in 2021 IEEE In e na ional Con e ence on Senso s and Nano echnology
(SENNANO). IEEE, 2021, pp. 109–112.
[3] F. G a and M. G ube , “Rapid Inciden De ec ion in Tunnels h ough
Acous ic Moni o ing – Ope a ing Expe iences in Aus ian Road Tun-
nels,” G az, Jan. 2018, p. 8.
[4] C. Wa kins, L. G een Maze olle, D. Rogan, and J. F ank, “Technological
app oaches o con olling andom gun i e: Resul s o a gunsho de ec ion
sys em ield es ,” Policing: An In e na ional Jou nal o Police S a egies
& Managemen , ol. 25, no. 2, pp. 345–370, Jun. 2002.
[5] G. Valenzise, L. Ge osa, M. Tagliasacchi, F. An onacci, and A. Sa i,
“Sc eam and gunsho de ec ion and localiza ion o audio-su eillance
sys ems,” Oc . 2007, pp. 21–26.
[6] M. Sigmund and M. H abina, “E icien ea u e se de eloped o acous-
ic gunsho de ec ion in open space,” Elek onika i Elek o echnika,
ol. 27, no. 4, pp. 62–68, 2021.
[7] H. Luo, J. Liu, W. Fang, P. E. Lo e, Q. Yu, and Z. Lu, “Real- ime
sma ideo su eillance o manage sa e y: A case s udy o a anspo
mega-p ojec ,” Ad anced Enginee ing In o ma ics, ol. 45, p. 101100,
2020.
[8] N. Khalid, M. Gochoo, A. Jalal, and K. Kim, “Modeling wo-pe son
segmen a ion and locomo ion o s e eoscopic ac ion iden i ica ion: A
sus ainable ideo su eillance sys em,” Sus ainabili y, ol. 13, no. 2,
2021.
[9] P. K.-Y. Wong, H. Luo, M. Wang, P. H. Leung, and J. C. Cheng, “Recog-
ni ion o pedes ian ajec o ies and a ibu es wi h compu e ision and
deep lea ning echniques,” Ad anced Enginee ing In o ma ics, ol. 49,
p. 101356, 2021.
[10] W. Renda and C. H. Zhang, “Compa a i e analysis o i ea m discha ge
eco ded by gunsho de ec ion echnology and calls o se ice in
louis ille, ken ucky,” ISPRS In e na ional Jou nal o Geo-In o ma ion,
ol. 8, no. 6, p. 275, 2019.
[11] J. H. Ra cli e, M. La anzio, G. Kikuchi, and K. Thomas, “A pa ially
andomized ield expe imen on he e ec o an acous ic gunsho
de ec ion sys em on police inciden epo s,” Jou nal o Expe imen al
C iminology, ol. 15, no. 1, pp. 67–76, 2019.
[12] J. Salamon and J. P. Bello, “Deep Con olu ional Neu al Ne wo ks
and Da a Augmen a ion o En i onmen al Sound Classi ica ion,” IEEE
Signal P ocessing Le e s, ol. 24, no. 3, pp. 279–283, Ma . 2017, a Xi :
1608.04363.
[13] K. J. Piczak, “En i onmen al sound classi ica ion wi h con olu ional
neu al ne wo ks,” in 2015 IEEE 25 h In e na ional Wo kshop on Ma-
chine Lea ning o Signal P ocessing (MLSP). Bos on, MA, USA:
IEEE, Sep. 2015, pp. 1–6.
[14] A. Guzho , F. Raue, J. Hees, and A. Dengel, “ESResNe : En i-
onmen al Sound Classi ica ion Based on Visual Domain Models,”
a Xi :2004.07301 [cs, eess], Ap . 2020, a Xi : 2004.07301.
[15] T. Ahmed, M. Uppal, and A. Muhammad, “Imp o ing e iciency and
eliabili y o gunsho de ec ion sys ems,” in 2013 IEEE In e na ional
Con e ence on Acous ics, Speech and Signal P ocessing. IEEE, 2013,
pp. 513–517.
[16] M. H abina and M. Sigmund, “Gunsho ecogni ion using low le el
ea u es in he ime domain,” in 2018 28 h In e na ional Con e ence
Radioelek onika (RADIOELEKTRONIKA). IEEE, 2018, pp. 1–5.
[17] M. H abina, “Analysis o linea p edic i e coe icien s o gunsho
de ec ion based on neu al ne wo ks,” in 2017 IEEE 26 h In e na ional
Symposium on Indus ial Elec onics (ISIE). IEEE, 2017, pp. 1961–
1965.
[18] M. Lojka, M. Ple a, E. Kik o ´
a, J. Juh´
a , and A. ˇ
Ciˇ
zm´
a , “E icien
acous ic de ec o o gunsho s and glass b eaking,” Mul imedia Tools
and Applica ions, ol. 75, no. 17, pp. 10 441–10 469, 2016.
[19] A. A. Shiekh, M. Tahi , and M. Uppal, “Accu a e gunsho de ec ion in
u ban en i onmen s using blind decon olu ion,” in 2017 In e na ional
Mul i- opic Con e ence (INMIC). IEEE, 2017, pp. 1–4.
[20] A. Suliman, B. Oma o , and Z. Dosbaye , “De ec ion o impulsi e
sounds in s eam o audio signals,” in 2020 8 h In e na ional Con e ence
on In o ma ion Technology and Mul imedia (ICIMU). IEEE, 2020, pp.
283–287.
[21] S. Raponi, I. Ali, and G. Olige i, “Sound o guns: digi al o ensics
o gun audio samples mee s a i icial in elligence,” a Xi p ep in
a Xi :2004.07948, 2020.
[22] J. Bajzik, J. P inosil, and D. Konia , “Gunsho de ec ion using con-
olu ional neu al ne wo ks,” in 2020 24 h In e na ional Con e ence
Elec onics. IEEE, 2020, pp. 1–5.
[23] J. Salamon, C. Jacoby, and J. P. Bello, “A Da ase and Taxonomy
o U ban Sound Resea ch,” in P oceedings o he ACM In e na ional
Con e ence on Mul imedia - MM ’14. O lando, Flo ida, USA: ACM
P ess, 2014, pp. 1041–1044.
[24] Y. Tokozume and T. Ha ada, “Lea ning en i onmen al sounds wi h
end- o-end con olu ional neu al ne wo k,” in 2017 IEEE In e na ional
Con e ence on Acous ics, Speech and Signal P ocessing (ICASSP).
New O leans, LA: IEEE, Ma . 2017, pp. 2721–2725.
[25] Y. Tokozume, Y. Ushiku, and T. Ha ada, “Lea ning om Be ween-class
Examples o Deep Sound Recogni ion,” a Xi :1711.10282 [cs, eess,
s a ], Feb. 2018, a Xi : 1711.10282.
www.ijacsa. hesai.o g 957 |Page
(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions,
Vol. 13, No. 4, 2022
[26] B. Zhu, K. Xu, D. Wang, L. Zhang, B. Li, and Y. Peng, “En i-
onmen al Sound Classi ica ion Based on Mul i- empo al Resolu ion
Con olu ional Neu al Ne wo k Combining wi h Mul i-le el Fea u es,”
a Xi :1805.09752 [cs, eess], Jun. 2018, a Xi : 1805.09752.
[27] S. Abdoli, P. Ca dinal, and A. L. Koe ich, “End- o-End En i onmen-
al Sound Classi ica ion using a 1D Con olu ional Neu al Ne wo k,”
a Xi :1904.08990 [cs, s a ], Ap . 2019, a Xi : 1904.08990.
[28] A. G. Ka siamis, E. M. D akakis, and R. F. Lyon, “P ac ical gamma one-
like il e s o audi o y p ocessing,” EURASIP Jou nal on Audio,
Speech, and Music P ocessing, ol. 2007, pp. 1–15, 2007.
[29] H. Pa k and C. D. Yoo, “Cnn-based lea nable gamma one il e bank and
equal-loudness no maliza ion o en i onmen al sound classi ica ion,”
IEEE Signal P ocessing Le e s, ol. 27, pp. 411–415, 2020.
[30] I. Papadimi iou, A. Va eiadis, A. Lalas, K. Vo is, and D. Tzo-
a as, “Audio-based e en de ec ion a di e en sn se ings using
wo-dimensional spec og am magni ude ep esen a ions,” Elec onics,
ol. 9, no. 10, p. 1593, 2020.
[31] N. Takahashi, M. Gygli, B. P is e , and L. Van Gool, “Deep Con o-
lu ional Neu al Ne wo ks and Da a Augmen a ion o Acous ic E en
De ec ion,” a Xi :1604.07160 [cs], Dec. 2016, a Xi : 1604.07160.
[32] V. Boddapa i, A. Pe e , J. Rasmusson, and L. Lundbe g, “Classi ying
en i onmen al sounds using image ecogni ion ne wo ks,” P ocedia
Compu e Science, ol. 112, pp. 2048–2056, 2017.
[33] Z. Zhang, S. Xu, S. Cao, and S. Zhang, “Deep Con olu ional Neu-
al Ne wo k wi h Mixup o En i onmen al Sound Classi ica ion,”
a Xi :1808.08405 [cs, eess], Aug. 2018, a Xi : 1808.08405.
[34] H. Wang, Y. Zou, D. Chong, and W. Wang, “En i onmen al Sound Clas-
si ica ion wi h Pa allel Tempo al-spec al A en ion,” a Xi :1912.06808
[cs, eess], May 2020, a Xi : 1912.06808.
[35] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Lea ning
o Image Recogni ion,” a Xi :1512.03385 [cs], Dec. 2015, a Xi :
1512.03385.
[36] “Sound E ec s Lib a y,” Ma . 2013.
[37] D. Dean, S. S idha an, R. Vog , and M. Mason, “G he qu -noise- imi
co pus o e alua ion o oice ac i i y de ec ion algo i hms,” in 2010
11 h Annual Con e ence o he In e na ional Speech Communica ion
Associa ion, 2010, pp. 3110–3113.
www.ijacsa. hesai.o g 958 |Page