scieee Open visual document viewer

Independent Channel Residual Convolutional Network for Gunshot Detection

Bajzík, Jakub; Přinosil, Jiří; Jarina, Roman; Mekyska, Jiří

Abstract

The main purpose of this work is to propose a robust approach for dangerous sound events detection (e.g. gunshots) to improve recent surveillance systems. Despite the fact that the detection and classification of different sound events has a long history in signal processing, the analysis of environmental sounds is still challenging. The most recent works aim to prefer the time-frequency 2-D representation of sound as input to feed convolutional neural networks. This paper includes an analysis of known architectures as well as a newly proposed Independent Channel Residual Convolutional Network architecture based on standard residual blocks. Our approach consists of processing three different types of features in the individual channels. The UrbanSound8k and the Free Firearm Sound Library audio datasets are used for training and testing data generation, achieving a 98 % F1 score. The model was also evaluated in the wild using manually annotated movie audio track, achieving a 44 % F1 score, which is not too high but still better than other state-of-the-art techniques.

Full text

(IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 Independen Channel Residual Con olu ional Ne wo k o Gunsho De ec ion Jakub Bajzik1, Ji i P inosil2, Roman Ja ina3, Ji i Mekyska4 Dep . o Mecha onics and Elec onics, Uni e si y o Zilina Zilina 010 26, Slo akia1 Dep . o Telecommunica ions, B no Uni e si y o Technology 601 90 B no, Czech Republic2,4 Dep . o Mul imedia and In o ma ion and Communica ion Technology Uni e si y o Zilina, Zilina 010 26, Slo akia3 Abs ac —The main pu pose o his wo k is o p opose a obus app oach o dange ous sound e en s de ec ion (e.g. gunsho s) o imp o e ecen su eillance sys ems. Despi e he ac ha he de ec ion and classi ica ion o di e en sound e en s has a long his o y in signal p ocessing, he analysis o en i onmen al sounds is s ill challenging. The mos ecen wo ks aim o p e e he ime- equency 2-D ep esen a ion o sound as inpu o eed con olu ional neu al ne wo ks. This pape includes an analysis o known a chi ec u es as well as a newly p oposed Independen Channel Residual Con olu ional Ne wo k a chi ec u e based on s anda d esidual blocks. Ou app oach consis s o p ocessing h ee di e en ypes o ea u es in he indi idual channels. The U banSound8k and he F ee Fi ea m Sound Lib a y audio da ase s a e used o aining and es ing da a gene a ion, achie ing a 98 % F1 sco e. The model was also e alua ed in he wild using manually anno a ed mo ie audio ack, achie ing a 44 % F1 sco e, which is no oo high bu s ill be e han o he s a e-o - he-a echniques. Keywo ds—Acous ic signal p ocessing; gunsho de ec ion sys- ems; audio signal analysis; machine lea ning; deep lea ning; esidual ne wo ks I. INTRODUCTION In he ield o signal p ocessing, he audio da a analysis akes an ex ensi e pa , which is cons an ly s udied. Many machine lea ning-based algo i hms we e p oposed o sol ing asks such as classi ica ion, segmen a ion, and denoising. In many cases, he me hods a e adap ed o a speci ic ype o sound, mainly speech and music, which ake an ex ensi e pa in he esea ch. On he o he hand, he en i onmen al sounds a e uns uc u ed, and i is challenging o gene alize hei na u e. Many en i onmen al sound analysis applica ions, anging om u ban moni o ing [1] o IoT [2] and su eillance sys em [3], [4], [5], ha e been de eloped wi hin he pas yea s. Howe e , in ecen yea s i has become a opical ask o use known echniques o classi y en i onmen al sounds like explosion, gunsho , si en, ca ala m, baby c ying, window b eakage and o he e en s associa ed wi h po en ial dange [6]. Usage o lea ning algo i hms ends o inc ease pe sonal sa e y. The possible implemen a ions a e in-home o indus ial p o ec ion sys ems, in ca s o ale dea o poo ly hea ing d i e s o he si en, in homes o ale a pa en o a c ying child, and in a wide ange o assis i e de ices, especially o dea people. Recen mode n su eillance sys ems o isk p e en ion pu poses ocuses mainly on he analysis o ideo signals om came as using ad anced compu e ision echniques [7], [8], [9]. Howe e , he analysis o audio signals has conside able po en ial in hese sys ems as well. Especially gunsho de ec ion echnologies ha e been inc easingly adop ed by law en o ce- men agencies o mapping he spa ial and empo al pa e ns o gun iolence [10]. The gene al p oblem o gun de ec ion echnologies is a high a e o alse ala ms esul ing in he was e o police esou ces when esponding o hose alse ale s [11]. F om a p ac ical poin o iew, i is necessa y o minimize he amoun o alse-posi i e p edic ions. The e o e, when se ing he ope a ing poin o he sys em in p ac ical applica ions, he alse-posi i e a e mus be aken in o accoun . The main objec i e o ou wo k is o p opose a me hod o de ec ing dange - ela ed audio e en s (gunsho s) ha achie e high speci ici y in eal condi ions. We also aim o explo e se e al ypes o ea u e spaces and neu al a chi ec u es and discuss, wha kind o se up is he mos sui able o such an applica ion. The s anda d app oach o sound analysis is o collec a single ec o o ea u es. The o en used classi ie s a e Sup- po Vec o Machines (SVM), Deep Neu al Ne wo ks (DNN), o mul ilaye pe cep ons. When p ocessing en i onmen al sounds, he uni o m s uc u e can no be expec ed as in speech o music, whe e he signal con ains a ha monic s uc u e o epe i ions. The ea u es ha pe o m well in speci ic applica- ions may be inadequa e o sounds wi h o he na u e and ice e sa. In ou wo k we a e using se e al 2-dimensional sound ep- esen a ions such as spec og ams as well as he s anda d 1-D app oach and analyse he pe o mance in he gunsho de ec ion applica ion. In addi ion o he equen ly used spec og am [12], [13], [14], also new isualiza ions and ad anced me hods o ea u e p ocessing a e used. The main con ibu ions o ou wo k a e: •Explo ing he sui abili y o se e al s a e-o - he-a con olu ional ne wo ks based app oaches o gunsho de ec ion. •P oposing he con olu ional model o boos ing he pe o mance on 2-dimensional independen ea u e spaces. The signal is ans o med in o h ee independen audio ea u e se s o ming an ”RGB image” ha is sui able o p ocessing www.ijacsa. hesai.o g 950 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 by common 2-D con olu ional ne wo ks used o image p o- cessing. Unlike image p ocessing, whe e colo channels a e highly co ela ed and p ocessed oge he a e he i s DNN laye , he h ee audio ea u e se s a e p ocessed independen ly by he i s wo DNN blocks. Ou p oposed a chi ec u e is based on s anda d esidual uni s. The impo an ask is no only inc easing he numbe o ue p edic ions bu also educing he numbe o alse-posi i e gunsho p edic ions. II. RELATED WORK O e he yea s se e al wo ks dealing wi h gunsho s de- ec ion om audio signal ha e been published. Mos o hem a e based on he ex ac ion o handc a ed acous ic ea u es and he use o machine lea ning echniques o he ask o classi ica ion. A combina ion o 7 Linea P edic i e Coding (LPC) coe icien s and 13 Mel F equency Ceps al Coe icien s (MFCC) wi h SVM classi ie is used in wo k [15] p o iding 8 % alse ala ms a e on a cus om da ase . The au ho o [16] ex ended he p e ious ea u e se by Linea P edic i e Coding Ceps al (LPCC) and au o-co ela ion coe icien s eaching 82 % accu acy and 70 % p ecision on a combina ion o public a ailable da ase s. The same au ho hen e alua ed he e ec o indi idual ea u es on he accu acy o he classi ica ion ask [17], conside ing he ea u es wi h he bes sco e o be he i s i e coe icien s o he 24 h o de LPC. A la ge se o a ious acous ic ea u es wi h Hidden Ma ko Model (HMM) and Vi e bi decode is used in EAR-TUKE sys em [18] o de ec ing gunsho s and glass b eaking e en s wi h 98 % accu acy in eco ds wi h Signal o Noise Ra io (SNR) ≥20 dB. In case o mic ophones a ays i is possible o use a wo s age me hodology comp ising o a Blind Sys em Iden i ica ion and Decon olu ion (BSID) s age ollowed by a SVM-based classi ica ion [19] o gunsho de ec ion in a noisy u ban en i onmen . In [20] a me hod o classi ying impulsi e sounds based on a Weigh ed Majo i y Vo ing (WMV) s a egy is desc ibed. In [21] Con olu ional Neu al Ne wo k (CNN) wi h empo al and spec al ea u es is used o gunsho sound ca ego ies classi ica ion (pis ol, i le and sho gun o di e en calib es) eaching o e 90 % accu acy. Ano he CNN app oach de ec s gunsho s wi h 99 % accu acy and low alse ala m a e using he ResNe a chi ec u e [22]. Mos o he abo e app oaches wo k wi h da abase eco dings ha con ain a low le el o en i onmen al noise. In he case o eal applica ions, his condi ion can ha dly be me . Fo his eason, dealing wi h he de ec ion and classi ica ion o en i onmen al noisy sounds is impo an . The da ase s o en i onmen al sounds a e made mainly o lea ning algo i hms ha pe o m En i onmen al Sound Classi ica ion (ESC) ask. One o he widely used da ase s o ESC is he U banSound8K [23], u he desc ibed in Sec ion III-E. The signal ep esen a ion o audio is ela ed o he a - chi ec u e o he lea ning algo i hm and lea ning objec i e. The s anda d p ocess ha ollows he app oaches om speech and music analysis is o collec a single ec o o ea u es. Sho - e m o long- e m ea u es may no always gene alize he uns uc u ed na u e o en i onmen al sounds. The di ec solu ion o he ea u e ex ac ion p oblem is o build a model ha ope a es on he aw audio signals di ec ly. The 1-D CNNs can handle he in e nal ep esen a ion o he inpu signal, which allows end- o-end usage. The impo an ad an age is ha he e is no need o ans o m o p e-p ocess he da a, and such a model can adap o a a ie y o audio signals. In he s udies [24], [25], he i s end- o-end ESC a chi ec u e called En Ne was p oposed in e sions 1 and 2. Ano he 1-D a chi ec u e was p oposed in [26], whe e he audio signal was p ocessed a di e en ime scales. The s udy [27] p esen s an end- o-end 1-D con olu ional ne wo k ha has ewe pa ame e s compa ed o dense 2-D con olu ional neu al ne wo ks and does no equi e a la ge amoun o aining da a. I eaches 87 % mean accu acy on he U banSound8K da ase wi h andom weigh s ini ializa ion and 89 % wi h weigh s ini ializa ion by Gamma one il e bank coe icien s [28] syn hesizing an impulse esponse om ne e cells in he audi o y ibe [29]. The 2-D CNN models ope a e on he p e-compu ed ea u e ep esen a ions ob ained by a ixed p ocess o ex ac ion. The s udy [13] is he i s which deals wi h ESR using CNN ained on mel-scaled spec og ams. Such an app oach is ex ended in s udy [12], whe e di e en augmen a ion me hods a e used. In wo k [14], he au ho s p esen ed he model ESResNe based on STFT, ha ou pe o ms ecen ly known app oaches wi h ESC da ase s achie ing 82% accu acy when ained om sc a ch and 85% accu acy wi h ImageNe weigh s ini aliza ion on U banSound8K da ase . Since he Piczak’s wo k [13], he esea ch end in en i onmen al sound analysis seems o be he usage o 2-D ea u e ma ices o eeding 2-D CNNs [30]. III. MATERIALS AND METHODS A. Signal Model and P oblem Fo mula ion Following he ecen s udies [12], [13], [14], [31], [32], [33], [34], he mos sui able se up o en i onmen al sound ecogni ion employs a 2-D con olu ional neu al ne wo k ed by a Time-F equency ep esen a ion o he audio signal ( u he men ioned as audio ea u es). In o de o be p ocessed by a 2-D con olu ional ne wo k, hese audio ea u es need o be con e ed in o a sui able uni o m 2-D ep esen a ion. Based on his 2-D ep esen a ion, i is hen necessa y o choose an op imal a chi ec u e o he con olu ional ne wo k o he classi ica ion ask. The whole wo k low om audio samples o p edic ions is depic ed in Fig. 1. Fig. 1. Fea u e Ex ac ion and Classi ica ion Wo k low. B. Audio Fea u es In he case o audio signal p ocessing, he e is a la ge num- be o a ious audio ea u es. The Log-Mel Spec og ams (LM Spec) and Mel F equency Ceps al Coe icien s (MFCC) a e among he mos commonly used audio ea u es. In ou wo k, we addi ionally include he Sel -Simila i y Ma ix (SSM), www.ijacsa. hesai.o g 951 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 which is equen ly used o analyze he global s uc u e o musical wo ks, o he audio ea u e lis . The hype pa ame e s o ea u e ex ac ion a e de ailed in Table I. 1) Log-Mel Spec og am: The spec og am is he mos commonly used audio signal isualiza ion. I shows he e- quency spec um change o e ime. A Sho Time Fou ie T ans o m (STFT) is used o con e he signal om ime o equency domain. Addi ional mel- equency scale ans o m 1 is applied o emb ace he psychoacous ic knowledge. mel = 2595 ·log10 1 + Hz 700 (1) 2) Mel F equency Ceps al Coe icien : MFCC e lec s he non-linea and masking psychoacous ic cha ac e is ics o human hea ing. MFCC coe icien s a e ob ained by mul iplying he signal spec um by a mel-scale dis ibu ed il e bank, loga i hm and Disc e e Cosine T ans o m (DCT). 3) Sel -Simila i y Ma ix: SSM is he measu e o sel - simila i y o he signal based on dis ances. We use a sel - simila i y ma ix o display signal co ela ion. To isualize he sel -simila i y, we use a ma ix Sde ined by Equa ion 2. The ma ix dimensions N×Ndepend on he numbe o signal samples. Fo educing he compu a ional complexi y, a sel - simila i y ma ix Sis compu ed on downsampled en elope s= (s1, s2, s3, ..., sN)o inpu audio signal, ob ained using he Hilbe ans o m. As a measu e o simila i y we a e using he absolu e dis ance. S(i, j) = |si−sj|i, j = 1, ..., N (2) The e ical and ho izon al axes ep esen he ime sequence. The ma ix is symme ical by he main diagonal whe e he simila i y is maximal. TABLE I. HYPERPARAMETERS FOR FEATURE EXTRACTION Fea u es FFT leng h Banks Window Log-mel spec og am 2048 256 Hamming MFCC 2048 20 Hamming Sel -simila i y - - - C. 2-D Fea u e Rep esen a ion Audio ea u es a e ex ac ed as 2-D ma ices and aligned o con olu ional neu al ne wo k inpu . Since mos 2D con- olu ional ne wo k a chi ec u es we e p ima ily designed o image p ocessing, hey expec 3 se s o 2-D ea u e ma ices a he inpu (an analogy o RGB channels o images). 1) Band-Spli ed Spec og am (BS Spec): In ou expe i- men s, we a e using he log-mel spec og am spli o h ee e- quency bands (high, middle, low) aligned wi h he RGB colo channels (each band as one colo channel). The band cu ing equencies depends he on maximal equency max = s 2 gi en by he sampling a e (Fig. 2). The same p inciple was used in s udy [14], whe e au ho s explain he usage o he band-spli ed spec og am o a oiding edundancy. O he solu ions a e eplica ing he spec og am o passing ze os. Fig. 2. Band-spli Log-mel Spec og am and Resul ing RGB Image. 2) Independen Fea u e Spaces (IFS): The combina ion o he log-mel spec og am, MFCC and sel -simila i y ma ix ep esen s he independen ea u e spaces. The simila me hod was used in s udy [32]. We assume ha he MFCC and SSM will help o classi y non-impulsi e backg ound sounds. The ha monici y o he gunsho signal is low, so he SSM is almos emp y, while he backg ound noise esul s in a isible g id. Th ee ea u e ma ices a e o e lapped in ma ched ime po- si ions. I means, ha he x-axis esolu ions a e app oxima ely he same o all ma ices. Howe e , on he y-axis we ha e di e en dimensions when using spec og am ( equency), MFCC (mel banks) and SSM ( ime) (Fig. 3). (a) (b) (c) (d) Fig. 3. Independen Fea u e Spaces as RGB Image Channels. (a) Red Channel, Log-mel Spec og am. (b) G een Channel, MFCCs. (c) Blue Channel, Sel -simila i y Ma ix. (d) Resul ing RGB Image. D. Con olu ional Neu al Ne wo ks A chi ec u es The mos widely used con olu ional neu al ne wo k models o 2-D ea u e space classi ica ion a e based on he esidual ne wo k a chi ec u e (ResNe ). Howe e , he e is also an app oach ha uses only a 1-dimensional con olu ional ne wo k ed di ec ly by a aw audio signal o classi y en i onmen al sounds. In addi ion, we include a cus om app oach based on esidual ne wo ks whe e indi idual channels a e p ocessed independen ly. 1) Residual Ne wo ks: Following ecen s udies [14], he esidual models pe o m well on en i onmen al sound clas- si ica ion. The esidual ne wo k was designed as a ne wo k in a ne wo k, which means ha he lowe laye ’s inpu s a e connec ed o he ou pu s o he wo highe laye s. The example o he s anda d esidual block is shown in Fig. 4. The skip connec ions de ined as y=F(x) + x(3) a e also called sho cu connec ions. The unc ion F ep- esen s he con olu ion ope a ions. The sho cu connec ions help o elimina e he p oblem o anishing g adien in deep neu al ne wo ks. The au ho s o [35] designed ResNe s wi h a di e en numbe o laye s, speci ically 18, 34, 50, 101 and 152. www.ijacsa. hesai.o g 952 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 Fig. 4. S anda d Residual Block [35]. 2) End- o-End Classi ica ion using a 1-D CNN: In his app oach, a 1-D con olu ional (1-D CNN) neu al ne wo k lea ns low-le el and high-le el in o ma ion di ec ly om he audio signal wa e o m. Since he size o he inpu da a o he amoun o da a is in imbalance, i is no ecommended o use oo deep con olu ional ne wo k a chi ec u es o a oid signi ican o e i ing. The s udy [27] p esen s he op imal a chi ec u e wi h espec o he sampling equency o he audio signal a di e en audio signal leng hs. Fo he sampling equency o 16 kHz conside ed u he in his pape , i has been shown ha he bes sco e is achie ed a he signal leng h o 1 second. The co esponding CNN a chi ec u e is shown in Table II consis ing o 4 Con olu ional Laye s (CL), 2 Pooling Laye s (PL) and 2 Fully Connec ed laye s (FC). The Rec i ied Linea Uni (ReLU) ac i a ion unc ion is used o all laye s, excep o he ou pu laye whe e he so max ac i a ion unc ion is used wi h he ou pu size equal o he numbe o classes beeing classi ied. TABLE II. ARCHITECTURE OF 1-D CNN FOR 16 KHZSAMPLING RATE AND AUDIO LENGTH OF 1 SECOND [27] CL1 PL1 CL2 PL2 CL3 CL4 FC1 FC2 Dimension 7969 996 483 60 23 8 128 64 Fil e s coun 16 16 32 32 64 128 - - Fil e s size 64 8 32 8 16 8 - - S ide size 2 8 2 8 2 2 - - 3) Independen Channel Residual Con olu ional Ne wo k: We p opose he Independen Channel Residual Con olu ional Ne wo k (ICRCN), whe e he inpu RGB image is di ided o he 2-D ma ices in indi idual channels. The ea u e ma ices sha e he dimension o he x-axis ( ime) bu no he y- axis ( equency, mel banks, ime). The sys em ha combines di e en isual ep esen a ions may su e , when he ea u es a e combined as one inpu image. The e o e, we build he esidual con olu ional ne wo k, ha p ocesses di e en audio isualiza ions sepa a ely. The whole model a chi ec u e is shown in Table III. The model inpu is a h ee channel RGB image. The sepa a e channels con ain esidual blocks, whe e he numbe o il e s is 32. The ea u e dimensions me ging is made a e he second esidual block. F om his poin , he ea u es a e p ocessed as in s anda d esidual con olu ional ne wo ks. The las con olu ional block consis s o 512 il e s and i is ollowed by he classi ica ion laye . The p oposed a chi ec u e is buil up om s anda d esidual blocks, as desc ibed in TABLE III. PROPOSED INDEPENDENT CHANNEL RESIDUAL CONVOLUTIONAL NETWORK Ou pu size ICRCN blocks (224, 224, 3) Inpu RGB image (112, 112, 32) 7×7, 32 7×7, 32 7×7, 32 (56, 56, 32) 3x3, 32 3x3, 32 ×23x3, 32 3x3, 32 ×23x3, 32 3x3, 32 ×2 (28, 28, 64) 3x3, 64 3x3, 64 ×23x3, 64 3x3, 64 ×23x3, 64 3x3, 64 ×2 (28, 28, 192) Conca ena ion (14, 14, 256) 3x3, 256 3x3, 256 ×2 (7, 7, 512) 3x3, 512 3x3, 512 ×2 (512) Global a e age pooling (2) Dense 2 + so max III-D1. The p oposed a chi ec u e is compa ed o s anda d esidual ne wo ks ResNe 50 and 1-D CNN in Table IV. TABLE IV. COMPARISON OF ARCHITECTURES COMPLEXITY A chi ec u e Numbe o ainable pa ame e s 1-D CNN 256k ResNe 50 23.5M ICRCN 11M E. Da ase s One o he mos widely known da ase s o en i onmen al sound classi ica ion is he U banSound8K [23]. In his wo k we build ou own gunsho de ec ion da ase as a combina ion o he U banSound8k and The F ee Fi ea m Sound Lib a y [36]. •U banSound8k - 8732 acks o 10 classes (ai condi ione , ca ho n, child en playing, dog ba king, d illing, engine idling, gunsho , jackhamme , si en, s ee music), wi h a ying sampling equency ile o ile. •The F ee Fi ea m Sound Lib a y - 2200 acks o gunsho s in a noise ee en i onmen including hand- guns (pis ols, e ol e s, semi-au oma ic pis ols), i les (le e -ac ion, semi-au oma ic, ully au oma ic, ma- chine guns, e c.) and sho guns. Reco dings a e in loss- less wa o ma wi h sampling equency 44.1 kHz. The examples om The F ee Fi ea m Sound Lib a y we e combined wi h gunsho s om he U banSound da ase . The o he sounds om U banSound we e used as nega i e exam- ples ( andom backg ound). Since U banSound eco ds a e no equal in leng h, he window size o ou examples loa s in he ange om 1 s o 4 s. In he ield o supe ised machine lea ning, he pe o - mance o algo i hms s ongly depends on he quali y o he da ase . To some ex en , i is possible o simula e a big da ase using augmen a ion me hods. Howe e , his app oach will ne e be as good as he expansion o he da ase by eal examples. The ollowing augmen a ion echniques a e applied o posi i e examples (gunsho s) while aining. www.ijacsa. hesai.o g 953 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 •Random backg ound mixing - The gunsho we e andomly mixed wi h backg ound noise in a andom SNR om 0 dB o 20 dB. •Random ime shi - The onse posi ions o gunsho s wi hin he p ocessing window we e chosen andomly om 0 o 0.8 % o he window leng h. •Random Gaussian noise addi ion - SNR om 60 dB o 100 dB. F. E alua ion Me ics The so max ou pu ac i a ion unc ion di ec ly indica es he p obabili y o he example belonging o a ce ain class. The so max ac i a ion also gua an ees ha he sum o p obabili ies o e classes (backg ound and gunsho ) is one. The examples a e classi ied as posi i e o nega i e acco ding o highe o lowe ac i a ion o ou pu neu ons. The con usion ma ix can be cons uc ed using p edic ed and ue labels. The ea e , he pe o mance is e alua ed ia known me ics. Fo p ac ical implemen a ion o he sys em o dange ous sound de ec ion (gunsho s, explosion, ...) i is necessa y o minimize he amoun o alse-posi i e p edic ions. In his case, he high accu acy alue may be a li le bi con using and hus we ha e o choose a me ic o ele an e alua ion. The e o e, in ou wo k we use he ollowing e alua ion me ics: •Accu acy: I e lec s he o e all algo i hm pe o - mance, so i also akes in o accoun he ue nega i e p edic ions (TN). The accu acy is high when he num- be o ue posi i e (TP) and ue nega i e p edic ions is la ge. False posi i es (FP) and alse nega i es (FN) a e p edic ion e o s. Accu acy =T P +T N T N +T P +F N +F P (4) •Sensi i i y: The ela i e amoun o ue posi i e p e- dic ions agains all posi i e examples. I is o en called ecall o T ue Posi i e Ra e (TPR). Sensi i i y =T P T P +F N (5) •Speci ici y: The ela i e amoun o ue nega i e p e- dic ions agains all nega i e examples. Speci ici y =T N F P +T N (6) •False-Posi i e Ra e (FPR): Re lec s he ela i e amoun o alse posi i es agains all nega i e exam- ples. FPR =F P F P +T N (7) •F1 sco e: The balance alue be ween sensi i i y and p ecision, whe e p ecision is he amoun o ue pos- i i es agains all posi i e p edic ions. This means ha he F1 sco e is no dis o ed by a la ge numbe o ue nega i e p edic ions and may be conside ed as decisi e, when es ing da a a e unbalanced. F1 = 2T P 2T P +F P +F N (8) •De ec ion E o T adeo - DET cu e isualizes he alse posi i e a e s. he alse nega i e a e. S anda dly, he axes a e scaled non-linea ly. Unlike he ROC cu e, he DET cu es a e mo e linea and si ua ed in mos o he plo a ea. •Equal E o Ra e - EER is de ined as he poin in DET o ROC whe e he e o s a e equal. The lowe EER alue e lec s be e pe o mance o he sys em. IV. RESULTS All he ne wo ks we e ained as long as alida ion loss was dec easing. Ea ly s op was applied and only he bes model was sa ed. Adam op imize was used while aining he models and he ca ego ical c oss-en opy loss was compu ed in each ba ch o 16 examples. Ha dwa e used is NVIDIA GeFo ce GTX 1650, In el Co e i9-10900X CPU 3.7 GHz, RAM 64 GB. Fig. 5 shows he aining and alida ion his o y o ou ICRCN model when using 16 kHz sampling equency. Fig. 5. T aining and Valida ion His o y o ICRCN Model. We use Tenso low and Ke as o building and aining he models. Fo augmen ing he audio we use he Audiomen a ions lib a y. The Sci-ki lea n is used o e alua ion using desc ibed me ics. A. 2-D Con olu ional Ne wo ks The aining audio examples we e gene a ed in h ee sam- pling equencies o es he e ec o sampling on sys em pe o mance. The o e all F1 sco e and FPR is e alua ed in Fig. 6. As he es subse is balanced, he F1 sco e and o e all accu acy me ics a e e y simila . As seen, he high F1 sco e alue does no necessa ily gua an ee good pe o mance in he educ ion o alse-posi i es. When looking a he F1 sco e o indi idual models, he di e ence be ween 44.1 kHz and 16 kHz is no signi ican . In se e al cases, he sco e d ops signi ican ly a he 8 kHz sampling. The signi ican d op o F1 sco e is seen in case o ICRCN ained on band-spli spec og ams. Table V, Table VI, and Table VII show also he change in sensi i i y and speci ici y o e all combina ions. TABLE V. TESTING RESULTS FOR SAMPLING FREQUENCY 8KHZ Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%] ICRCN IFS 93.1 94.8 91.3 BS Spec 85.1 74.0 100.0 ResNe 50 IFS 95.3 98.4 91.9 LM Spec 96.0 93.2 99.0 BS Spec 96.9 97.1 96.7 MFCC 84.8 91.5 75.8 SSM 96.0 94.2 98.1 www.ijacsa. hesai.o g 954 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 (a) (b) (c) (d) (e) ( ) Fig. 6. Resul s o Di e en Sampling F equencies. (a) F1 Sco e, 8 kHz. (b) FPR, 8 kHz. (c) F1 Sco e, 16 kHz. (d) FPR, 16 kHz. (e) F1 Sco e, 44.1 kHz. ( ) FPR, 44.1 kHz. TABLE VI. TESTING RESULTS FOR SAMPLING FREQUENCY 16 KHZ Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%] ICRCN IFS 97.9 98.4 97.3 BS Spec 93.9 99.6 87.4 ResNe 50 IFS 95.9 95.5 96.3 LM Spec 97.8 98.8 96.7 BS Spec 97.1 99.0 95.2 MFCC 87.0 93.8 78.1 SSM 95.9 95.5 96.3 TABLE VII. TESTING RESULTS FOR SAMPLING FREQUENCY 44.1 KHZ Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%] ICRCN IFS 98.7 98.4 99.0 BS Spec 97.3 99.6 95.0 ResNe 50 IFS 97.2 95.2 99.4 LM Spec 98.1 97.5 98.8 BS Spec 99.0 100.0 98.1 MFCC 60.9 44.4 98.6 SSM 95.9 94.6 97.3 As he so max ou pu ac i a ion unc ion is used, he ou pu gi es he p obabili ies o he example belonging o a ce ain class. Ne e helss, he ope a ing poin can be mo ed by changing he h eshold o posi i e p edic ions. Choosing he ope a ing poin is s ongly ela ed o he applica ion o he de eloped sys em and he e a e di e en ways how o se up he ope a ing poin . The use ul isualiza ions a e he DET cu es (Fig. 7), which end o be mo e linea and highligh he di e ences in he ope a ing egion mo e clea ly han he ROC cu es. The EER alue ep esen s he ope a ing poin , whe e he sys em pe o ms equal in e ms o he alse-posi i e a e and alse-nega i e a e. The alse-posi i e a e can be in e p e ed as he alse ala m p obabili y and alse-nega i e a e is he p obabili y o missing he posi i e de ec ion. (a) (b) (c) (d) (e) ( ) Fig. 7. De ec ion E o T adeo Cu es o Di e en Models and Sampling F equencies. (a) ICRCN, 8 kHz. (b) ResNe 50, 8 kHz. (c) ICRCN, 16 kHz. (d) ResNe 50, 16 kHz. (e) ICRCN, 44.1 kHz. ( ) ResNe 50, 44.1 kHz. As seen in Table 7, he EER e alua ion sligh ly co ela es wi h he o e all F1 sco e o he sys em. Though he e is no e iden eg ess o e he pa ame e s, we can highligh se e al indings om he esul s. Using he o iginal sampling equency when he da a a e clea and de ailed, all a chi ec u es model he da a well. The mo e ainable pa ame e s seem o be use ul as he sys em losses he in o ma ion by audio downsam- pling. A 16 kHz sampling, he pe o mance o a chi ec u es is closely compa able. The p oposed ICRCN pe o ms bes wi h IFS ea u e combina ion. Acco ding o s udy [32], he single SSM pe o ms wo s on en i onmen al sounds classi ica ion. Ou esul s showed, ha o gunsho de ec ion, using he SSM only may leads o be e sco es han single MFCC, which su p isingly pe o ms wo s . www.ijacsa. hesai.o g 955 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 B. 1-D Con olu ional Ne wo k In his pa o wo k, we compa e he 1-D CNN model o he bes 2-D models. Fo his compa ison we conside only he 16 kHz sampling equency. As he bes candida es om he p e ious es we choose wo model- ea u es combina ions, namely ICRCN-IFS and ResNe 50-LM Spec. The compa ison is shown in Table VIII. The 1-D CNN model was ained on he same da ase as 2-D models. TABLE VIII. THE RESULTS FOR 1-D CNN MODEL IN COMPARISON TO THE BEST 2-D MODELS Model Fea u es F1 [%] Sensi i i y [%] Speci ici y [%] ICRCN IFS 97.9 98.4 97.3 ResNe 50 LM Spec 97.8 98.8 96.7 1-D CNN - 98.1 98.2 98.1 C. E alua ion o Pe o mance in he Wild The esul s on a i icially gene a ed es subse s may no always mee he eal pe o mance in he wild. The e o e, we e alua e he models on a eal audio ack om ac ion mo ie o simula e he eal wo ld condi ions. On his s age we use only 16 kHz sampling equency as he mos sui able based on p e ious esul s. The audio ack om he mo ie John Wick (2014) was manually anno a ed. We use a 4 seconds window leng h wi h 2 seconds o e lap and hi d-o de median il e ing o he CNN ou pu . Each segmen is labeled as posi i e i he e is an occu ence o a gunsho . The numbe s o posi i e and nega i e segmen s is e y unbalanced, he e o e we use he F1 sco e as he e alua ion me ic. The e alua ion esul s a e shown in Table IX. TABLE IX. EVALUATION ON REAL DATA IN THE WILD Model Fea u es Tes E alua ion F1[%] F1[%] Sensi i i y[%] Speci ici y[%] ICRCN BS Spec 93.9 14.5 98.2 10.5 IFS 97.9 44.3 59.2 91.6 ResNe 50 LM Spec 97.8 26.4 89.0 62.5 MFCC 87.0 16.5 47.2 67.0 SSM 95.9 5.5 5.5 92.7 BS Spec 97.1 17.6 98.2 29.2 IFS 95.9 16.2 95.4 24.3 1-D CNN - 98.1 4.2 7.8 96.5 As shown in Table IX, he pe o mance on eal da a d ops signi ican ly. Howe e , he e alua ion es sco es co ela e in a ela i e way. The es sco e di e ence is sligh in mos cases, while he e alua ion sco e di e ence is no iceable. We can see, ha single SSM ou pe o ms he MFCC ea u es in es esul s, bu ails in e alua ion, whe e he sensi i i y and speci ici y a e e y unbalanced. The esul s show, ha he p oposed ICRCN ne wo k a chi ec u e was able o model he in a da a obus way and ou pe o ms he s anda d esidual ne wo k ResNe 50 e en using less ainable pa ame e s. In he case o 1-D CNN, he bad inal sco e was p obably caused by he low numbe o pa ame e s o he con olu ional neu al ne wo k. Thus he ne wo k was no able o gene alize he ex ac ed ea u es su icien ly o deal wi h he high a iabili y o eal audio da a. The p e ious expe imen mainly ocused on he abili y o he p oposed algo i hm o de ec gunsho sounds in he simula ed eal eco ding. Howe e , in he case o su eillance sys ems, in addi ion o accu acy, a low equency o alse ala ms is also equi ed. This alue canno be de i ed om he abo e expe imen because he sound acks o he mo ies a e sound-exposed. Thus, in his case, eco dings om he QUT-NOISE da abase [37], which is he only sui able publicly a ailable da abase con aining con inuous eco dings o ci y sounds, we e used o de e mine he alse e o a e. Since he o al eco ding ime is only a ew hou s, his da abase was supplemen ed wi h cus om eco dings om a busy ci y s ee . Wi hin hese eco dings, he p oposed algo i hm did no de ec a single alse gunsho e en . D. P ocessing Time In Table X, we showed he mean p ocessing ime o each o he ea u es used in he expe imen s. The gene a ion o h ee ea u e ma ices akes wice as much ime han gene a ion o a single spec og am. This mus be aken in o accoun i he de ec o is implemen ed on he sys em wi h limi ed p ocessing capaci y. TABLE X. FEATURE EXTRACTION TIME Sampling equency [kHz] Fea u es Time [ms] 8Spec 5.02 MFCC 2.68 SSM 2.31 16 Spec 6.23 MFCC 3.62 SSM 3.71 44.1 Spec 10.53 MFCC 6.9 SSM 10.58 When compa ing p ocessing ime be ween mul iple sam- pling equencies, he good comp omise is o p e e he 16 kHz sampling. I o e s he compa able pe o mance o o iginal sampling while he ea u e ex ac ion ime is almos hal ed. V. DISCUSSION When compa ing he esidual a chi ec u es and ea u es using es ing da a, he e is no a signi ican bes se up. The e alua ion on eal da a showed he bene i o independen ea u es p ocessing, whe e he di e ence in sco e is obse - able. The s anda d ResNe 50 pe o ms well wi h a single spec og am. The ea u e combina ion wo ks bes wi h he p oposed a chi ec u e wi h independen channels. The e o e, we p e e using he p oposed ICRCN a chi ec u e and IFS o gunsho de ec ion. Taking he ea u e ex ac ion imes in o accoun , we p e e using he 16 kHz sampling. We assume, ha he ICRCN model ed by IFS ma ices is he mos sui able o gunsho de ec ion sys ems, ou pe o ming he o he s a e-o - he-a app oaches (in e ms o F1 sco e). The combina ion o spec og am, MFCC, and SSM was used in s udy [32], wi h conclusion ha i does no imp o e he accu acy o en i onmen al sounds classi ica ion. We assume, ha he ea u e combina ion can be bene icial o gunsho de ec ion applica ions and has a po en ial o lowe he alse- posi i e a e. The esul s showed, ha he e is a bene i o spli ing he con olu ional channels. The eal pe o mance o www.ijacsa. hesai.o g 956 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 he sys em ela es on he ope a ing poin . In some applica ions he speci ici y may be a mo e impo an sco e han sensi i i y, o ice e sa. We isualized he DET cu es, as hey can help o i he equi emen s o he sys em o pa icula applica ions. VI. CONCLUSION In his pape , we analyzed he usage o se e al con- olu ional a chi ec u es o he gunsho de ec ion ask. We p oposed a new a chi ec u e and compa ed i s pe o mance o s anda d ResNe 50 and 1-D a chi ec u es. The e ec o di e en ea u es on he esul ing pe o mance was es ed using se e al Time-F equency audio ep esen a ions. Fo aining, alida ion and es ing we collec ed he gunsho s and andom backg ounds audio clips om public da ase s. In he aining phase, he s anda d augmen a ion me hods we e used. Finally, we simula ed he eal wo ld condi ions by e alua ion o a eal audio ack om he ac ion mo ie. We achie ed he goal o ou wo k by p oposing a sys em o gunsho audio e en s de ec ion. Ou ICRCN app oach is able o ope a e in noisy en i onmen s wi h high speci ici y (sligh ly o e 90 %), by main aining ai sensi i i y (almos 60 %). The De ec ion E o T adeo analysis showed, ha he eal pe o mance o he sys em s ongly depends on he e o ole ance and equi emen s. The ope a ing poin should be selec ed o speci ic applica ions. On ou es da a, he missed de ec ion p obabili y is abou 10 % when he alse ala m p obabili y is as minimal as possible. We expec he simila esul s also o explosion de ec ion. Such a sys ems can be implemen ed in a complex applica ion oge he wi h a smoke o i e de ec o . Due o he ac ha he p oposed app oach has a ela i ely low compu a ional ime, i can be easily in eg a ed in o exis ing su eillance sys ems wi hou he need o in es in expensi e compu ing se e s o ope a e in eal ime. ACKNOWLEDGMENT This wo k was suppo ed by he Slo ak G an Agency KEGA unde con ac no. KEGA 008ZU-4/2021. REFERENCES [1] J. P. Bello, C. Sil a, O. No , R. L. DuBois, A. A o a, J. Sala- mon, C. Mydla z, and H. Do aiswamy, “SONYC: A Sys em o he Moni o ing, Analysis and Mi iga ion o U ban Noise Pollu ion,” a Xi :1805.00889 [cs, eess], May 2018, a Xi : 1805.00889. [2] S. C. Tan and A. Abd Mana , “Cha ac e iza ion o in e ne o hings (io ) powe ed-acous ics senso o indoo su eillance sound classi ica ion,” in 2021 IEEE In e na ional Con e ence on Senso s and Nano echnology (SENNANO). IEEE, 2021, pp. 109–112. [3] F. G a and M. G ube , “Rapid Inciden De ec ion in Tunnels h ough Acous ic Moni o ing – Ope a ing Expe iences in Aus ian Road Tun- nels,” G az, Jan. 2018, p. 8. [4] C. Wa kins, L. G een Maze olle, D. Rogan, and J. F ank, “Technological app oaches o con olling andom gun i e: Resul s o a gunsho de ec ion sys em ield es ,” Policing: An In e na ional Jou nal o Police S a egies & Managemen , ol. 25, no. 2, pp. 345–370, Jun. 2002. [5] G. Valenzise, L. Ge osa, M. Tagliasacchi, F. An onacci, and A. Sa i, “Sc eam and gunsho de ec ion and localiza ion o audio-su eillance sys ems,” Oc . 2007, pp. 21–26. [6] M. Sigmund and M. H abina, “E icien ea u e se de eloped o acous- ic gunsho de ec ion in open space,” Elek onika i Elek o echnika, ol. 27, no. 4, pp. 62–68, 2021. [7] H. Luo, J. Liu, W. Fang, P. E. Lo e, Q. Yu, and Z. Lu, “Real- ime sma ideo su eillance o manage sa e y: A case s udy o a anspo mega-p ojec ,” Ad anced Enginee ing In o ma ics, ol. 45, p. 101100, 2020. [8] N. Khalid, M. Gochoo, A. Jalal, and K. Kim, “Modeling wo-pe son segmen a ion and locomo ion o s e eoscopic ac ion iden i ica ion: A sus ainable ideo su eillance sys em,” Sus ainabili y, ol. 13, no. 2, 2021. [9] P. K.-Y. Wong, H. Luo, M. Wang, P. H. Leung, and J. C. Cheng, “Recog- ni ion o pedes ian ajec o ies and a ibu es wi h compu e ision and deep lea ning echniques,” Ad anced Enginee ing In o ma ics, ol. 49, p. 101356, 2021. [10] W. Renda and C. H. Zhang, “Compa a i e analysis o i ea m discha ge eco ded by gunsho de ec ion echnology and calls o se ice in louis ille, ken ucky,” ISPRS In e na ional Jou nal o Geo-In o ma ion, ol. 8, no. 6, p. 275, 2019. [11] J. H. Ra cli e, M. La anzio, G. Kikuchi, and K. Thomas, “A pa ially andomized ield expe imen on he e ec o an acous ic gunsho de ec ion sys em on police inciden epo s,” Jou nal o Expe imen al C iminology, ol. 15, no. 1, pp. 67–76, 2019. [12] J. Salamon and J. P. Bello, “Deep Con olu ional Neu al Ne wo ks and Da a Augmen a ion o En i onmen al Sound Classi ica ion,” IEEE Signal P ocessing Le e s, ol. 24, no. 3, pp. 279–283, Ma . 2017, a Xi : 1608.04363. [13] K. J. Piczak, “En i onmen al sound classi ica ion wi h con olu ional neu al ne wo ks,” in 2015 IEEE 25 h In e na ional Wo kshop on Ma- chine Lea ning o Signal P ocessing (MLSP). Bos on, MA, USA: IEEE, Sep. 2015, pp. 1–6. [14] A. Guzho , F. Raue, J. Hees, and A. Dengel, “ESResNe : En i- onmen al Sound Classi ica ion Based on Visual Domain Models,” a Xi :2004.07301 [cs, eess], Ap . 2020, a Xi : 2004.07301. [15] T. Ahmed, M. Uppal, and A. Muhammad, “Imp o ing e iciency and eliabili y o gunsho de ec ion sys ems,” in 2013 IEEE In e na ional Con e ence on Acous ics, Speech and Signal P ocessing. IEEE, 2013, pp. 513–517. [16] M. H abina and M. Sigmund, “Gunsho ecogni ion using low le el ea u es in he ime domain,” in 2018 28 h In e na ional Con e ence Radioelek onika (RADIOELEKTRONIKA). IEEE, 2018, pp. 1–5. [17] M. H abina, “Analysis o linea p edic i e coe icien s o gunsho de ec ion based on neu al ne wo ks,” in 2017 IEEE 26 h In e na ional Symposium on Indus ial Elec onics (ISIE). IEEE, 2017, pp. 1961– 1965. [18] M. Lojka, M. Ple a, E. Kik o ´ a, J. Juh´ a , and A. ˇ Ciˇ zm´ a , “E icien acous ic de ec o o gunsho s and glass b eaking,” Mul imedia Tools and Applica ions, ol. 75, no. 17, pp. 10 441–10 469, 2016. [19] A. A. Shiekh, M. Tahi , and M. Uppal, “Accu a e gunsho de ec ion in u ban en i onmen s using blind decon olu ion,” in 2017 In e na ional Mul i- opic Con e ence (INMIC). IEEE, 2017, pp. 1–4. [20] A. Suliman, B. Oma o , and Z. Dosbaye , “De ec ion o impulsi e sounds in s eam o audio signals,” in 2020 8 h In e na ional Con e ence on In o ma ion Technology and Mul imedia (ICIMU). IEEE, 2020, pp. 283–287. [21] S. Raponi, I. Ali, and G. Olige i, “Sound o guns: digi al o ensics o gun audio samples mee s a i icial in elligence,” a Xi p ep in a Xi :2004.07948, 2020. [22] J. Bajzik, J. P inosil, and D. Konia , “Gunsho de ec ion using con- olu ional neu al ne wo ks,” in 2020 24 h In e na ional Con e ence Elec onics. IEEE, 2020, pp. 1–5. [23] J. Salamon, C. Jacoby, and J. P. Bello, “A Da ase and Taxonomy o U ban Sound Resea ch,” in P oceedings o he ACM In e na ional Con e ence on Mul imedia - MM ’14. O lando, Flo ida, USA: ACM P ess, 2014, pp. 1041–1044. [24] Y. Tokozume and T. Ha ada, “Lea ning en i onmen al sounds wi h end- o-end con olu ional neu al ne wo k,” in 2017 IEEE In e na ional Con e ence on Acous ics, Speech and Signal P ocessing (ICASSP). New O leans, LA: IEEE, Ma . 2017, pp. 2721–2725. [25] Y. Tokozume, Y. Ushiku, and T. Ha ada, “Lea ning om Be ween-class Examples o Deep Sound Recogni ion,” a Xi :1711.10282 [cs, eess, s a ], Feb. 2018, a Xi : 1711.10282. www.ijacsa. hesai.o g 957 |Page (IJACSA) In e na ional Jou nal o Ad anced Compu e Science and Applica ions, Vol. 13, No. 4, 2022 [26] B. Zhu, K. Xu, D. Wang, L. Zhang, B. Li, and Y. Peng, “En i- onmen al Sound Classi ica ion Based on Mul i- empo al Resolu ion Con olu ional Neu al Ne wo k Combining wi h Mul i-le el Fea u es,” a Xi :1805.09752 [cs, eess], Jun. 2018, a Xi : 1805.09752. [27] S. Abdoli, P. Ca dinal, and A. L. Koe ich, “End- o-End En i onmen- al Sound Classi ica ion using a 1D Con olu ional Neu al Ne wo k,” a Xi :1904.08990 [cs, s a ], Ap . 2019, a Xi : 1904.08990. [28] A. G. Ka siamis, E. M. D akakis, and R. F. Lyon, “P ac ical gamma one- like il e s o audi o y p ocessing,” EURASIP Jou nal on Audio, Speech, and Music P ocessing, ol. 2007, pp. 1–15, 2007. [29] H. Pa k and C. D. Yoo, “Cnn-based lea nable gamma one il e bank and equal-loudness no maliza ion o en i onmen al sound classi ica ion,” IEEE Signal P ocessing Le e s, ol. 27, pp. 411–415, 2020. [30] I. Papadimi iou, A. Va eiadis, A. Lalas, K. Vo is, and D. Tzo- a as, “Audio-based e en de ec ion a di e en sn se ings using wo-dimensional spec og am magni ude ep esen a ions,” Elec onics, ol. 9, no. 10, p. 1593, 2020. [31] N. Takahashi, M. Gygli, B. P is e , and L. Van Gool, “Deep Con o- lu ional Neu al Ne wo ks and Da a Augmen a ion o Acous ic E en De ec ion,” a Xi :1604.07160 [cs], Dec. 2016, a Xi : 1604.07160. [32] V. Boddapa i, A. Pe e , J. Rasmusson, and L. Lundbe g, “Classi ying en i onmen al sounds using image ecogni ion ne wo ks,” P ocedia Compu e Science, ol. 112, pp. 2048–2056, 2017. [33] Z. Zhang, S. Xu, S. Cao, and S. Zhang, “Deep Con olu ional Neu- al Ne wo k wi h Mixup o En i onmen al Sound Classi ica ion,” a Xi :1808.08405 [cs, eess], Aug. 2018, a Xi : 1808.08405. [34] H. Wang, Y. Zou, D. Chong, and W. Wang, “En i onmen al Sound Clas- si ica ion wi h Pa allel Tempo al-spec al A en ion,” a Xi :1912.06808 [cs, eess], May 2020, a Xi : 1912.06808. [35] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Lea ning o Image Recogni ion,” a Xi :1512.03385 [cs], Dec. 2015, a Xi : 1512.03385. [36] “Sound E ec s Lib a y,” Ma . 2013. [37] D. Dean, S. S idha an, R. Vog , and M. Mason, “G he qu -noise- imi co pus o e alua ion o oice ac i i y de ec ion algo i hms,” in 2010 11 h Annual Con e ence o he In e na ional Speech Communica ion Associa ion, 2010, pp. 3110–3113. www.ijacsa. hesai.o g 958 |Page