FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO
Ha monic Change De ec ion om
Musical Audio
Ped o Ramoneda F anco
Mes ado In eg ado em Engenha ia In o má ica e Compu ação
Supe iso : Gilbe o Be na des
July 28, 2020
Ha monic Change De ec ion om Musical Audio
Ped o Ramoneda F anco
Mes ado In eg ado em Engenha ia In o má ica e Compu ação
July 28, 2020
Abs ac
In his disse a ion, we ad ance an enhanced me hod o compu ing Ha e e al.’s [31] Ha monic
Change De ec ion Func ion (HCDF). HCDF aims o de ec ha monic ansi ions in musical audio
signals. HCDF is c ucial bo h o he cho d ecogni ion in Music In o ma ion Re ie al (MIR)
and a wide ange o c ea i e applica ions. In ligh o ecen ad ances in ha monic desc ip ion
and ans o ma ion, we depa om he o iginal a chi ec u e o Ha e e al.’s HCDF, o e isi
each one o i s componen blocks, which a e e alua ed using an exhaus i e g id sea ch aimed o
iden i y op imal pa ame e s ac oss ou la ge s yle-speci ic musical da ase s. Ou esul s show
ha he newly p oposed me hods and pa ame e op imiza ion imp o e he de ec ion o ha monic
changes, by 5.57% ( -sco e) wi h espec o p e ious me hods. Fu he mo e, while gua an eeing
ecall alues a >99%, ou me hod imp o es p ecision by 6.28%. Aiming o le e age no el
s a egies o eal- ime ha monic-con en audio p ocessing, he op imized HCDF is made a ailable
o Ja asc ip and he MAX and Pu e Da a mul imedia p og amming en i onmen s. Mo eo e , all
he da a as well as he Py hon code used o gene a e hem, a e made a ailable.
Keywo ds: Ha monic changes, Musical audio segmen a ion, Ha monic-con en desc ip ion, Mu-
sic in o ma ion e ie al
i
ii
Resumo
Nes a disse ação, a ançamos um mé odo melho ado pa a compu a Ha e e al.’s [31] Ha monic
Change De ec ion Func ion (HCDF). O HCDF isa de ec a ansições ha mónicas em sinais áudio
musicais, undamen ais pa a a a e a de es ima i a au omá ica de aco des den o da Recupe ação
de In o mação Musical (MIR) pa a um as o âmbi o de aplicações c ia i as que ão desde a ha -
monização a ans o mações ha mónicas. À luz dos ecen es a anços na desc ição ha mónica e
ans o mação do audio musical, pa imos da a qui ec u a o iginal do HCDF de Ha e e al. pa a
e isi a cada um dos seus componen es, que são a aliados u ilizando uma pesquisa exaus i a em
g elha pa a iden i ica pa âme os óp imos em qua o g andes conjun os de dados musicais especí-
icos de es ilos. Os nossos esul ados mos am que os no os mé odos p opos os e a op imização
de pa âme os melho am a de ecção de al e ações ha mónicas. Po 5,57% ( -sco e) em elação
aos mé odos an e io es. Além disso, embo a ga an indo alo es de eco dação a >99%, o nosso
ou o mé odo melho a a p ecisão em 6,28%. Com o objec i o de ap o ei a no as es a égias
pa a o p ocessamen o de áudio com con eúdo ha mónico em empo eal, o HCDF op imizado é
disponibilizado pa a Ja asc ip e os ambien es de p og amação mul imédia MAX e Pu e Da a.
Além disso, es ão disponí eis odos os esul ados e o código py hon pa a ge a odos eles.
Keywo ds: Ha monic changes, Musical audio segmen a ion, Ha monic-con en desc ip ion, Mu-
sic in o ma ion e ie al
iii
i
Acknowledgemen s
Fi s o all, I wan o hank my mum, my main music educa o , o all my li e. How can you each
an ins umen o a child as a game, wi hou he child ega ding? I ind i inc edible. When he
child unde s ood i , i was oo la e. Thank you and dad o suppo ing me in e e y single momen
o my li e. I don’ know wha I’d do wi hou you.
Many people would say I shouldn’ pu his he e, because in he u u e his dedica ion may
be blown away. Howe e , i will be a eminde o e e y hing we ha e li ed. Thanks o Ma ia
Valen ina o suppo ing and inspi ing me while I was doing all his wo k. Especially du ing
he qua an ine, I know ha some imes I am unbea able. I hope ha you ne e ge i ed o my
unbea ably because I don’ hink I will ge i ed. And hank you also o all he help you ga e me
in making he g aphics o his disse a ion and he pape .
I especially wan o hank Gilbe o Be na des. You suddenly ecei ed an E asmus s uden om
ano he uni e si y who spoke english badly and abou which you had no obliga ion. And you ga e
him a e y cool p ojec , you helped him in e e y hing you could and mo e, and you ga e him all
he oppo uni ies he needed. Ha s o , I’ e lea ned mo e om you han mos o my p o esso s. I
I’m e e a eache , I wan o be like you. E en i we’ e no oge he nex yea , I know I ha e a
men o and a iend in Po o.
Thank you o my amily in special o my g andma and Valen ina o s ay always he e wi h me.
To my o he amily, o all my childhood iends who a e my iends now, hanks o God. Fo all,
hese yea s, o suppo ing me in e e y momen . Fo ole a ing me, o s udying wo deg ees a he
same ime, o you jokes and o you s icke s. Thanks in pa icula o Albe o, Edu, And és and
Ca los o helping me wi h my bad English in his disse a ion. Also o all he iends I’ e made
his yea , Ko eans, Lau a and Pollo, my li le kid, because ou s o y doesn’ end he e. This yea
has been a g ea yea .
Las bu no leas , Thank you o all he eache s who ha e augh me h oughou my li e. To
hose a he p o essional conse a o y and he school. A lo o enginee ing eache s and e y ew
om he supe io conse a o y. Fo all hose who ha e lea ned and made me who I am, as a
musician and an enginee . Thanks o he Eu opean Union o belie ing ha he E asmus do se e
a pu pose. And hanks o music and a , o making me enjoy mysel and being he engine ha
mo es he wo ld.
Ped o Ramoneda
xii LIST OF FIGURES
4.7 Subjec i i y in anno a ions. Beginning o The Bea les’ song Please Please Me.
Sco e and HCDF diag am o e spec og am. F om 2 o 11 e e y bounda y is a
alse posi i e. The numbe s a e in he same empo al posi ion on he sco e as on
he spec og am. .................................. 50
4.8 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF .................................... 51
4.9 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF wi hou HPSS me hod in p ep ocessing block ............ 52
4.10 HCDF diag am. Fi s 60 seconds o Please Please Me . Bes -sco e esul o g id
sea ch HCDF wi hou onal model. ........................ 52
4.11 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF wi h W(h)as a onal model. ..................... 53
4.12 HCDF diag am. Fi s 60 seconds o Please Please Me. HCDF wi h pa ame e i-
za ion adop ed om bes ecall esul o g id sea ch. ............... 54
4.13 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes ecall esul o g id
sea ch HCDF wi hou HPSS p ep ocessing block. ................. 54
4.14 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes ecall anked HCDF
pa ame iza ion wi hou Tonal Space. ....................... 55
4.15 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes ecall esul o g id
sea ch HCDF wi h W(h)as a onal model. ..................... 56
4.16 HCDF diag am. Fi s 60 seconds o Please Please Me. Bes ecall anked on g id
sea ch HCDF wi h Cosine ξcos
n} as a cen oid dis ance. .............. 56
5.1 Some music indus y isualiza ion sys ems. ..................... 61
5.2 D’acco d desk op in e ace. Diag am ex ac ed om Gilbe o pe sonal si e. ... 62
6.1 E olu ion o he ch omag ams depending on he a e o ha monic change. .... 65
Lis o Tables
2.1 Numbe o semi ones and co esponding in e al name. .............. 8
4.1 Algo i hms and pa ame e s e alua ed in he g id sea ch me hod ......... 39
4.2 HCDF e alua ion indica o s o he bes -sco e and ecall me hods ac oss he ou
da ase s unde s udy, o which we compa e wi h p e ious me hods. The a e age is
compu ed o e all pieces. .............................. 48
4.3 Expe imen al esul s o HCDF peaks compa ed wi h hand labelled cho d changes
o 16 Bea les Songs (songs a anged in ch onological o de o elease da e).
Please Please Me (1), Do You Wan To Know A Sec e (2), All My Lo ing (3), Till
The e Was You (4), A Ha d Day’sDay’s Nigh (5), I I Fell (6), Eigh Days A Week
(7), E e y Li le Thing (8), Help! (9), Yes e day (10), D i e My Ca (11), Michelle
(12), Eleano Rigby (13), He e The e And E e ywhe e (14), Lucy In The Sky Wi h
Diamonds (15), Being Fo The Bene i O M Ki e (16). .............. 49
xiii
xi LIST OF TABLES
Abb e ia ions
HCDF Ha monic Change De ec ion Func ion
SMC Sound and Music Compu ing
MIR Music In o ma ion Re ie al
ACR Au oma ic Cho d Recogni ion
HPSS ha monic pe cussi e sou ce sepa a ion
x
Chap e 1
In oduc ion
The sea ch o a ma hema ical basis in musical ha mony has been a cons an ques o housands
o yea s [23,67,68,69,44]. Ancien G eeks [26] belie ed music could be used as a ool o unde -
s and he na u al wo ld, an app oach ha has been sus ained e e since1. O e he las wo decades,
in o ma ion echnologies ha e had a subs an ial impac on he newes indings. The compu a ional
p ocessing o ha mony has become a s anda d me hodology ac oss many disciplines, om mu-
sicology o audio signal p ocessing. The opic unde discussion in his disse a ion, ha monic
changes when cho ds change in ime a e unde s ood om a Wes e n pe spec i e al hough his in-
o ma ion is simple in symbolic mani es a ions o music, such as a musical sco e, he e ie al o
such in o ma ion om audio poses signi ican challenges [35].
1.1 Mo i a ion
In my longs anding classical piano s udies, polyphony, and in pa icula ha mony, has been ins u-
men al in my o mal educa ion. Ha mony can be ega ded as a ounda ion pa o music heo y,
composi ion, o mal analysis, ea aining, and choi , o ci e a ew.
On a b oade scope, ha mony is a cen al concep in Wes e n music. I is dis inc i e o he
compose s’ and a is s’ s yle, and as such, i s impo ance while pe o ming, imp o ising, and
composing mus be ecognized. In he cu en iew o compu a ional c ea i e asks, whe e he
compu e is assuming a somewha ac i e ole, ha monic change de ec ion o , in o he wo ds, audio
signal segmen a ion a he change o a cho d, is he i s phase o add essing a compu a ional
unde s anding o ha mony.
Ha monic change de ec ion has been add essed in he li e a u e o he pas wo decades, in
Sound and Music Compu ing (SMC) and Music In o ma ion Re ie al (MIR). Hainswo h e al.
[28] and Ha e e al. [31] a e he main con ibu ions o his p oblem. This wo k con inues a line o
1P ojec Cosmos ies o model musical s uc u es embedded in ca diac a hy hmias h p://cosmos.i cam.
1
2In oduc ion
esea ch o he Sound and Music Compu ing Lab a FEUP/INESC TEC on ha monic modelling,
namely on he pe cep ually-inspi ed ha monic ep esen a ion Tonal In e al Space [4,6,7,61].
In Music In o ma ion Re ie al (MIR), he ield o scien i ic esea ch in which his disse a ion
can be amed, Ha monic Change De ec ion is a undamen al block in much Au oma ic Change
De ec ion Recogni ion (ACR) algo i hms. Hence, i is c ucial o many asks such as Co e Song
Iden i ica ion, S uc u al Segmen a ion, Gen e Classi ica ion o Music Gene a ion. Mo eo e , i is
likely o p o ide many u u e applica ions, anging om echnical o c ea i e asks.
1.2 Objec i es
The main objec i e o his disse a ion is o e isi ha monic change de ec ion om musical audio,
a undamen al phase in he classical MIR ask, au oma ic cho d ecogni ion. Fu he mo e, i can
suppo a ious c ea i e- ela ed asks associa ed wi h ha monisa ion, and o he ha monic- ela ed
audio ans o ma ions, whe e a cho d segmen a ion is needed. Aiming o imp o e cu en s a e-
o - he-a me hods o ha monic change de ec ion, we will ollow h ee speci ic goals:
1. To in es iga e he use o he Tonal In e al Space as a pe cep ually-enhanced ep esen a ion
o cho d dis ances.
2. To pe o m a c i ical quan i a i e assessmen o he s a e-o - he-a me hods o he mul iple
componen modules o he HCDF by Ha e e al. [31], namely inspec ing me hods han ha e
been ecen ly p oposed in he li e a u e, such as no el ch omag am ep esen a ions and he
adop ion o ha monic pe cussi e sou ce sepa a ion.
3. Focusing on pa icula ly challenging musical examples and gen es and ying o a oid o e -
i ing p oblems. In he p e ious esea ches [31,21], only he 16 songs p esen ed in he i s
a icle [21] we e aken in o accoun . The use o small da ase s can lead o signi ican biases,
making i impossible o gene alise he model o any musical audio, e en o gene alise he
model o musical audio wi h he same gen e as he 16 songs.
4. To explo e isualisa ions and mining echniques which can p omo e be e compu a ional
musicology and e ie al asks in la ge audio da ase s.
1.3 App oach
To mee he objec i e de ined in Sec ion 1.2, we es ablished a me hodological pa h which co e s
he ollowing phases. To add ess he pu pose, i s o all, we will e isi he HCDF p oposed by
Ha e e al. [31]. Speci ically, we conside mul iple pa ame e isa ions in i s componen modules,
and we examine he implica ion o di e en onal spaces. In pa icula , he ecen ly p oposed Tonal
In e al Space by Be na des e al. [4] in imp o ing a pe cep ual spa ial ep esen a ion o onal
pi ch. Se e al elemen s o Ha e e al.’s o iginal pipeline ha e been imp o ed du ing he las 15
yea s, and in his disse a ion, we es hem, as well as we y o b ing possible new imp o emen s.
1.4 S uc u e o he disse a ion 3
Mo eo e , Examining HCDF on di e en gen es and wi h a ious se ings could p oduce new
esea ch ou pu s and allow us o unde s and be e he me hod. Da ase s o di e en gen es we e
p oposed by Ha e e al. 15 yea s ago. Gi en ha i is possible o unde s and how HCDF wo ks
ac oss a ious gen es, we belie e i is possible o de elop a obus gene alized HCDF me hod.
Ano he objec i e is compa ing di e en onal spaces and hei pe o mance in di e en sce-
na ios and analysing how he di e en blocks o he pipeline and pa ame e iza ions ela e o each
o he . As i is s a ed in objec i e 3, i is impo an o make a c i ical analysis om a mo e mu-
sical pe spec i e o how he algo i hm wo ks. We use a musicological app oach wi h he aim o
unde s and concep s al eady exis ing in wes e n music.
To sum up, in gene al, he conduc ed esea ch seeks o asce ain whe he he Tonal In e al
Space, a no el ma hema ical amewo k which cap u es pe cep ual a ini y in onal pi ch as dis-
ances, ou pe o ms he cu en s a e o he a in ha monic change segmen a ion and unc ion
analysis.
1.4 S uc u e o he disse a ion
The ollowing pa ag aphs y o gi e an o e iew o how his disse a ion will be o ganised. In his
i s chap e 1, we ha e made a small app oxima ion o why and how we a e going o app oach o
ha monic change de ec ion.
In Chap e 2, we will deal wi h e e y hing ha has been done in he ield o he s udy o
ha monic change, s a ing wi h he app oaches om he Music Technology ield. And all he
backg ound on which i is based, ha is, echnological, musicological o cogni i e. The p e ious
p oposals and s uc u es will also be deal wi h o y o analyse ha mony au oma ically.
In Chap e 3, we will analyse all he new possibili ies ha we p opose o ex end he pipeline
p oposed by Ha e e al [31].
In Chap e 4shows how he p oblem has been sol ed om a mo e echnical pe spec i e. A e
his, we will analyse he esul s ob ained om a da a mining and musicology pe spec i e.
In Chap e 5, we p opose applica ion scena ios o ha monic change de ec ion. Finally, in
Chap e 6, we e iew e e y hing ha has been done in his disse a ion, and we p opose wha is
going o be done in he u u e.
4In oduc ion
Chap e 2
S a e o he A
The compu a ional p ocessing o musical ha mony has been a opic o in e es since he ea ly days
o compu e music [87,43,81]. O e he las 20 yea s, we ha e wi nessed he c ea ion o many
associa ions and communi ies a he in e sec ion o compu ing and music, such as as he Sound
and Music Compu ing Ne wo k and he Socie y o Musical In o ma ion Re ie al. No only hese
socie ies ha e as ly inc emen ed he isibili y o his a ea, bu also ha e p omo ed in e na ional
enues o he dissemina ion and discussion o esea ch ou pu s. Like many o he a eas, he g ow h
o music echnology is due mainly o he e olu ion o compu ing, in o ma ion echnologies, signal
p ocessing, a i icial in elligence, machine lea ning and, in ecen yea s, deep lea ning. On he
o he hand, he longs anding musicological adi ion has equally esponded wi h di e en me hods
and app oaches wi h access o compu a ional ad ances. Compu a ional musicology is a b anch o
he ea ly musicological a ea ha exis s a he in e sec ion be ween music and echnology [20,48,
17].
2.1 S a e o he Field
In he ea ly 1950s, a dedica ed g oup o compose s, enginee s and scien is s began o explo e
he use o new digi al echnologies o c ea e new ypes o music. Music and echnology ha e
had a ui ul his o y e e since. Many e ms, such as compu e music and music echnology,
ha e been used o wha oday is commonly e e ed o as sound and music compu ing [74].
In 1974, i was es ablished he In e na ional Compu e Music Associa ion and he In e na ional
Compu e Music Con e ence. In 1977, he Compu e Music Jou nal was ounded. The Cen e
o Compu e Resea ch in Music and Acous ics (CCRMA) a S an o d was c ea ed in he mid
1970s, and he Ins i u e o Resea ch and Coo dina ion Acous ic/Music (IRCAM) in Pa is was
es ablished sho ly a e , as he p incipal depa men o he Pompidou Cen e . These wo cen es,
oge he wi h he nowadays inac i e composi ion depa men o P ince on Uni e si y, lead esea ch
in music echnology h oughou he 20 h cen u y. O e he pas 25 yea s, many o he cen es ha e
5
12 S a e o he A
• I he slash symbol is be ween wo no es "/" and means ha he lowes no e is di e en om
he oo . Fo example, C/F indica es ha a C majo iad wi h an F added o he bass should
be played.
Figu e 2.4: Di e en no a ions compa ison.
Cho d P og essions
In a musical composi ion, a cho d p og ession o ha monic p og ession is a succession o
cho ds. Cho d p og essions a e he ounda ion o ha mony in Wes e n musical adi ion om he
common p ac ise e a o Classical music o he 21s cen u y. Cho d p og essions a e he ounda ion
o Wes e n popula music s yles (e.g., pop music, ock music) and adi ional music (e.g., blues and
jazz). The bes way o s udy ha monic p og ession is o conside p og essions in g oups acco ding
o he in e al p oduced by he oo s o wo adjacen cho ds. The ollowing gene al ca ego ies will
o m he basis o ou s udy o ha monic p og ession.
Undoub edly he mos common o all ha monic p og essions is he ci cle p og ession, shown
in Figu e 2.5. Mo e han any o he , his p og ession has he capabili y o de e mining a onali y,
gi ing di ec ion and h us , and p o iding o de in a sec ion o ph ase o music. I is indeed he
basis o all ha monic p og ession [3].
2.2 Fundamen als o Ha mony: A Music Theo y Pe spec i e 13
Figu e 2.5: T adi ional ci cle o i hs diag am.
Two concep s widely used in his disse a ion a e ha monic change and ha monic hy hm.
On he one hand, ha monic change e e s o ansi ions om one cho d o ano he on musical
audio. On he o he hand, ha monic hy hm is he a e a which he cho ds change, in a musical
composi ion. I can be conside ed om a me ely ime pe spec i e o om he ela ion o cho ds
pe no e.
Keys
In music heo y, he key o a piece is he g oup o ones, o scale, which o ms he basis o a
musical composi ion om classical music, h ough pop o ock music o adi ional music. F om
a educed pe spec i e and in ela ion o he cho ds o he 18 h cen u y onwa ds, he key o a piece
has been explained as he esul o a educ ion o he h ee main cho ds, he onic, he dominan
and he subdominan .
The key is composed o he onic no e and i s co esponding cho ds also called he onic
cho d. They p o ide a subjec i e sense o a i al and es , and also ha e a ela ionship o i s
nea by onali ies, i s co esponding cho ds, and he onali ies and cho ds ou side he g oup. No es
and cho ds di e en om he onic c ea e di e en ensions, which a e esol ed when he no e o
he onic cho d e u ns.
14 S a e o he A
In he Wes e n music adi ion, he onali y can be in he majo o mino mode. Popula songs
a e usually in a pa icula majo o mino key, and so is classical music du ing he pe iod o
common p ac ice, a ound 1650-1900. Longe pieces in he classical epe oi e may ha e sec ions
in con as ing keys.
Modula ions
In music, modula ion is he change om one onali y ( onic o onal cen e) o ano he . This
may o may no be accompanied by a change in he key o he onali y, modula ions s uc u e
he o m o many pieces, as well as add in e es . T ea ing a cho d as a onic o no mo e han
one ph ase is no conside ed a modula ion. Modula ion desc ibes he p ocess in which a piece o
music changes om one key o ano he comple ely.
When you s a w i ing a piece o music, one o he i s hings you do is choosing a key o
compose. This choice o onali y de e mines which scale is used, how many sha ps and la s he e
a e, and which cho ds can be used. This onali y is some imes called he "s a ing onali y".
Many songs and pieces emain in his s a ing key and do no change. Howe e , o make a
piece mo e in e es ing compose s o en change o a di e en key a some poin in he piece. This
change is a modula ion.
Musical S uc u e
Fo m in music is he esul o he in e ac ion o all s uc u al elemen s. I consis s o by
di e en ph ases and pe iods, bu s uc u e e e s as a whole, he o ganisa ion o a comple e com-
posi ion, and akes in o accoun ha monic, hy hmic and melodic elemen s.
A piece o music can gene ally be di ided in o wo o mo e main sec ions, and he bounda ies
be ween hese sec ions a e called o mal di isions. The o mal di isions a e he esul o s ong
ha monic and melodic cadences and hy hmic ac o s and ha e been widely s udied he cen u ies.
Fo mal di isions de ine he sec ions o a composi ion, and hese sec ions a e labelled wi h capi al
le e s: A, B, C, e c. I a musical sec ion is epea ed, he same le e is used: A, A, B, B, e c., and
i i con ains simila ma e ial, i is designa ed by adding cousins o he p e ious le e : A, A’, A",
e c. Popula songs, song and ock usually ha e a cho us ha is epea ed o example A, A’, A",
e c., and se e al di e en s anzas B, C, C’, D, e c.
2.3 Audio Con en -Based P ocessing
Music can be ep esen ed di e en ly, om musical sco es, a a symbolic le el, o he egis a ion
o i s pe o mance, in audio o ma s. Audio con en -based p ocessing aims o add ess he la e
case and ex ac o abs ac in o ma ion, named desc ip ions o ea u es, om aw audio. Au-
dio Con en -based p ocessing suppo s many audio-d i en applica ions, such as au oma ic speech
ecogni ion, audio segmen a ion, music ecommenda ion o en i onmen al sound e ie al. Ac-
cessing in o ma ion by p ocessing he con en o an audio signal has been add essed by many
disciplines om musicology o audio signal p ocessing o ind a new way o modelling aw audio
2.3 Audio Con en -Based P ocessing 15
as a mo e complex phenomenon. Con en is a e y used e m in mul imedia e ie al. In his chap-
e , con en om a music pe spec i e is he cen al axis. The o he e m analysed in his chap e is
p ocessing om a digi al signal p ocessing pe spec i e.
Fi s o all, we a e going o ocus on he con en e m. In o ma ion science and linguis ics o -
e he meaning o e m con en . Bu we a e going o ocus on a Mul imedia pe spec i e co e on
se e al pas esea ch [62,49]. The Socie y o Mo ion Pic u e and Tele ision Enginee s (SMPTE)
and he Eu opean B oadcas ing Union (EBU) de ined con en as he combina ion o wo concep s:
me ada a and essence. The essence is he aw p og am ma e ial i sel , he da a ha di ec ly en-
codes images, ex , ideo, e c. o as in ou case aw audio. Essence in o ma ion can be encoded
o di ec ly ep esen he ac ual message, and i is usually p esen ed o de ed by ime. On he
o he hand, me ada a is he desc ip o s o he essence and i s a ie ies. SMPTE/EBU p esen ed a
classi ica ion o me ada a:
• Essen ial (me a-in o ma ion ha is necessa y o ep oduce he essence, like he incompa -
ibili ies, he numbe o audio channels, he Unique Iden i ie , he ideo o ma which is
encoded aw da a, e c.).
• Access (who can access o he essence, as legal access, i.e. copy igh and licenses);
• Pa ame ic (how he essence ha e been cap u ed, numbe o cap u e machine, ypes, loca ion
).
• Rela ional (how o synch onize essence encoding, e.g. ime-code).
• Desc ip i e (desc ip o s o he essence ha can allow use s ca alogue, sea ch, e ie al and
adminis a ion o con en ).
O he classi ica ion p esen ed by he Na ional In o ma ion S anda ds O ganisa ion only con-
side s h ee di e en me ada a ypes:
• Desc ip i e me ada a, which desc ibes he essence.
• S uc u al me ada a, which desc ibes how pa s o he essence ha e ela ionships, such as
ideo and audio.
• Adminis a i e me ada a, which desc ibes managemen in o ma ion, pe missions, he ie o
secu i y, da e o c ea ion, who ha e accessed i .
In gene al, me a-da a is all he in o ma ion ha can be e ie ed om a media essence. In
o he wo ds, any in o ma ion ha can be anno a ed o ex ac ed o e a music piece in a mean-
ing ul way. MPEG-7 s anda d has a me ada a slo which is de ined as a con en desc ip o , "a
dis inc i e cha ac e is ic o he da a which signi ies some hing o somebody" [39]. This app oach
o music analysis has a e y big p oblem [79,78], he seman ic-gap p oblem, which a ises om
he disc epancy abou ex ac ed me ada a and how is he me ada a concep pe cei ed by one use
16 S a e o he A
in a gi en si ua ion. Tha is why i is essen ial o keep he da a mos objec i ely o ins ancing; i
depends on he ci cums ance. A way o sol e he seman ic-gap p oblem is o di ide me ada a in a
hie a chy [10], low-le el, mid-le el and high-le el. I is easy o unde s and how di e en a con en
desc ip ion o a music piece is i he a ge ed use was an ama eu lis ene o an expe musicolo-
gis . The educa ion o his designe can bias e en a low-le el desc ip o such as he Spec al lux,
and ano he digi al audio designe would ha e designed i di e en ly.
• Low-le el desc ip o can be compu ed om aw da a indi ec o de i ed way (i.e. a e
signal ans o ma ions like Fou ie o Wa ele ans o ms, a e s a is ical p ocessing like
a e aging, a e alue quan isa ion like he assignmen o a disc e e no e name o a gi en
se ies o pi ch alues, e c.). Mos o he low-le el desc ip o s don’ ha e he sense om a
musical pe spec i e. Howe e , hey can be easily compu ed om a compu e .
• Mid-le el desc ip o s equi e and induc ion om alida ed da a o an abs ac concep . In
his ca ego y would be ideas ypically musical like cho ds o ha monic change de i ed om
o he desc ip o s, o a Hidden Ma ko Model o a Deep Lea ning model ha segmen a song
by imb e simila i ies. Machine lea ning, Deep lea ning and s a is ical modelling make mid-
le el desc ip o s possible. Mid-le el desc ip o s a e also some imes e e ed o as objec -
cen ed desc ip o s.
• The di e ence om low-le el and mid-le el desc ip o s o high-le el desc ip o s is ha he
is a subjec i i y concep mix u e wi h he echnical concep . Fo example, piano and o e
depend on he es o he piece dynamics and how he lis ene pe cei es i . Mo e abs ac
concep s can be e ie ed om essen ia as agg essi i y, melancholy, dissonance, beau y, e c.
High-le el desc ip o s a e named use -cen ed desc ip o s oo.
Due o a wide a ie y o desc ip o s, ha e been p esen ed some amewo ks o build i . The
Dublin Co e and MPEG- 7 a e cu en ly he mos ele an s anda ds o music con en desc ip o s.
The o he big pa o audio-con en p ocessing is p ocessing. The o mal de ini ion is o sub-
jec a hing o a p ocess o elabo a ion o ans o ma ion; howe e , usually deno es a unc ional
o compu a ional app oach o sol e some scien i ic p oblems au oma ically. "Signal p ocessing"
is he cen al p ocessing in audio o p ocessing and modelling i , he e is ano he p ocessing in-
ol ed as language p ocessing, isual p ocessing, speech p ocessing, o knowledge p ocessing. I
should be no ed ha he di e ence be ween p esc ip ion and algo i hmic is unc ional; in p ocess-
ing, he c i ical hing is no how i is done bu wha is achie ed by doing i . And all he p ocessing
o music is a syne gy be ween Signal P ocessing, A i icial In elligence, In o ma ion Re ie al,
and Cogni i e Science o be able o p ocess and model music.
Audio Con en -Based P ocessing has been a g owing ield in music echnology [49]. The e a e
many di e en desc ip o s de eloped o unde s anding audio as Tempo al desc ip o s, Physical e-
quency desc ip o s, pe cep ual equency desc ip o s, Ceps al desc ip o s, Rhy hm (modula ion
equency desc ip o s) eigendomain desc ip o s o phase space desc ip o s. Mo eo e , he s a e
2.4 Ha monic Desc ip ion o Musical Audio 17
o esea ch is ma u e and some da abases can be use ul in pe cep ual esea ch, in music in o ma-
ion e ie al and machine-lea ning app oaches o con en -based e ie al in la ge sound da abases
as some ypes o desc ip o s as imb e desc ip o s [60]. In addi ion, he e a e so wa e lib a ies
wi h se e al audio-con en p ocessing ools, such as essen ia [11], madmom [9] and lib osa [47],
consolida ed and suppo ed by a la ge communi y, many o hem esea ch cen es.
This disse a ion co e s wo pe spec i es o audio con en -based p ocessing. The i s is a mu-
sicological iew. The musicological angle ies o ollow he Wes e n musical concep s de eloped
o yea s. The second is compu a ional. Compu a ional pe spec i e e e s o con en -based audio
p ocessing s a egies, and his means ha au oma ically gi e in o ma ion abou he musical signal
hanks o desc ip o s. Musically mo i a ed ones y o imi a e wes e n musical concep s.
2.4 Ha monic Desc ip ion o Musical Audio
Ha mony is a e y abs ac concep de eloped om human pe cep ion [83]. Desc ip o s ha e o be
meaning ul in o de o gene a e aluable in o ma ion. Melody (sequence o single pi ches), un-
damen al equency, pi ch classes and cho ds (simul aneous combina ions o pi ches), and cho d
p og essions, ha mony and key ( empo al combina ions o cho ds) a e desc ip o s e y use ul o
unde s anding ha mony [49].
The sea ch o a obus ha monic sound ep esen a ion has been con inuous o he las 50
yea s. Ha monic ea u es a e noisy om a lo o pe spec i es. Musical ins umen s p oduce o e -
ones summed o F0; pe cussi e ins umen s and di e en imb es and ins umen a ion p oduce
pe e sions in he ep esen a ion o ha mony.
Ha monic con en -based audio desc ip o s a e mainly e ical (i.e., cho ds), howe e , hey
ha e a deep ela ionship wi h ho izon al pa (i.e., melodic and oice-leading). Fu he mo e,
highe -le el desc ip o s, such as he concep o musical key o conno a ions and ensions, a e
in a s ong ela ionship wi h mid-le el ha monic desc ip o s. The i s a emp s o add ess he
ha monic domain was om symbolic mani es a ions o music such as hose encoded using he
MIDI s anda d [16]. Howe e , he e a e oo many di e ences be ween he symbolic domain and
audio domain. The la e equi es dedica ed me hods. Fu he mo e, a emp s o ansla e om he
audio domain o symbolic ha e no been p oli ic [82]. Mo eo e , MIDI esea ch is e y biased
owa ds piano ansc ip ion. The sea ch o a obus se o desc ip o s and s uc u es abou sound,
polyphony and ha mony is cen al o he many MIR asks. Fu he mo e, esea ch o link ha monic
desc ip o s o seman ic in o ma ion (high-le el desc ip o s) o e en o he ypes o in o ma ion
ha e been essen ial [89].
Ha monic con en -based audio desc ip o s on which we ha e ocused in his disse a ion a e
Mid-le el desc ip o s, such as ch omag am, cho ds, Tonne z based desc ip o s and ha monic bound-
a ies. Ch omag ams a e one o he mos used me hods in he MIR ield. They a e a simpli ica ion
o cho ds, and di e en ch omag ams ha e been p oposed in he li e a u e. Fu he mo e, he Ha -
monic Ne wo k o Tonne z is a ypical ep esen a ion o he ela ionships be ween ha monies in
musicology, ypically a ibu ed o Eule , used by music heo is s such as Riemann and Oe ingen
18 S a e o he A
in he 17 h cen u y [71,70,57]. The gene aliza ion o his led o he spi al a ay c ea ed by Elaine
Chew, a ma hema ical model o y o unde s and onali y.
2.4.1 Spec og am
Figu e 2.6: Example o ypical spec og am. Compu ed wi h lib osa [47].
The spec og am is he signal ep esen a ion o spli ing di e en anges o equencies as a his-
og am. The esul is a h ee-dimensional g aph ha ep esen s he ene gy o he equency con en
o he signal as i a ies o e ime, see Figu e 2.6. I is a undamen al pilla ying o ep esen au-
dio. Ha mony is an example oo, in pa icula , he i s s ep o some ch omag ams as is explained
in he nex sec ion.
2.4.2 Ch omag am
In wha ollows we de ail he mos used ch omag ams. I includes om simple and ea lie me h-
ods o ch omag am compu a ion such as CSTFT [22] o CCQT [15], ha di ec ly map he window-
based spec al analysis om he p ep ocessing block in o 12 ch oma elemen s, o mo e obus
ch omag ams such as he CNNLS [42] o CHPCP [27], which include addi ional p ocessing o en-
hancing he ansc ip ion o he unning o he ep esen a ion. Figu e 2.8 shows he ou ch oma-
g ams adap ed o a majo cho d. Mo eo e , o he ypes o ch omag ams a e men ioned in he
ollowing sec ion.
One o he mos p ominen ha monic-con en audio desc ip o s is he ch omag am, which
accumula es he ene gy o an audio signal as 12-elemen ec o s, as is shown in Figu e 2.7, ep-
esen ing he 12 no es o he ch oma ic scale ac oss all oc a es. I was i s p oposed in 1999 by
Fujishima [25]. Many algo i hms o ch oma compu a ion ha e ollowed, such as pi ch class p o-
iles [25], ha monic pi ch class p o iles (HPCP) [27], he CRP ch oma [52] o he NNLS ch oma
[42]. Las Deep Lea ning ch omag ams ha e a emp ed o ge a bina y ch omag am wi h only he
no es han could be ac i a ed. Thei di e ences s em mos ly om di e en deg ees o in a iance
o a pa icula musical a ibu e (e.g., imb e) o some enhanced le el o pe o mance owa ds a
symbolic-based ep esen a ion (e.g. ha monics o ansien noise).
Sho Time Fou ie T ans o m ch omag am CSTFT [22] esul s om mapping and accu-
mula ing he ene gy om he equency bins o a sho - ime Fou ie ans o m ep esen a ion in o
hei co esponding pi ch class.
2.4 Ha monic Desc ip ion o Musical Audio 19
Figu e 2.7: Chomag am. A 12-pi ch-class ci cula a ay o modelling cho ds and ha mony.
Cons an -Q T ans o m ch omag am CCQT depa s om he loga i hmic scaled Cons an -Q
ans o m as he base ep esen a ion o he mapping spec al bins o he 12 pi ch classes o a
ch omag am. Due o i s loga i hmically spaced spec al basis, his ch omag am is be e aligned
wi h human pe cep ion. Howe e , he cons an -Q ans o m [15] is hea ie o compu e when com-
pa ed wi h he as Fou ie ans o m (FFT) used in CSTFT. Compu ing his algo i hm wi h g ea e
obus ness in ol es he applica ion o CQT ( ia FFT) oc a e-by-oc a e, using lowpass il e ed
and downsampled esul s o sequen ially lowe pi ches. On he CCQT, usually, he subsampled
me hod and he di ec FFT me hod a e combined. In he end, he p ocess is ca ied ou wi h all he
necessa y equencies.
Ha monic Pi ch Class P o ile ch omag am CHPCP [27] is uning-independen and disca ds
he p esence o noisy ansien s by depa ing om a sinusoidal componen analysis o he audio
signal. The CHPCP makes a i s app oxima ion o he uning, based on he Wes e n well- empe ed
sys em, o ha e a e e ence equency. Fu he mo e, he accumula ed ene gy in each pi ch class o
he CHPCP is compu ed om he undamen al equency and ha monically aligned pa ial equen-
cies. This p ocedu e aligns wi h human pe cep ion as ou hea ing sys em ypically uses hese
pa ials in o a unique audi o y image.
Non-Nega i e Leas -Squa es ch omag am CNNLS [42]uses a Non-Nega i e Leas -Squa es
(NNLS) op imiza ion algo i hm o app oxima e he ansc ip ion o he no es be o e he ch oma
compu a ion. In he beginning, a log- equency DFT spec um is compu ed. La e , he global
equal- empe ed uning equency o all he piece is es ima ed om he spec og am. Then, he
p e ious log- equency spec og am is ecalcula ed using linea in e pola ion and aking in o ac-
coun he global equency. A e ha , he spec um backg ound is calcula ed and emo ed om
20 S a e o he A
Figu e 2.8: Di e en ch omag ams o an A majo cho d.
he abo e spec um, as indica ed in [42]. A e he NNLS decomposi ion, he spec um men ioned
abo e is mapped in o a 12 bin ch oma. The e o e, his ch omag am should no ha e any non- onal
componen s like ansien noise, since i is calcula ed in a symbolic way.
Deep Lea ning ch omag am CDEEP [42]. In he las ew yea s he e ha e been se e al a -
emp s o do a ch omag am wi h deep lea ning. Good examples a e he deep ch oma ex ac o [37]
and c ema-pcp [46].
Ch oma DCT-Reduced log Pi ch [52]. The gene al idea is o disca d imb e- ela ed in o -
ma ion wi h ce ain Mel- equency ceps al coe icien s (MFCCs) me hods. I ’s p o en ha he
lowe MFCCs ha e a big ela ionship o he aspec o imb e. Then, his me hod disca ds he pa
o he signal ha is closely ela ed o imb e (MFCC). These ec o s a e named as CRP (Ch oma
DCT-Reduced log Pi ch) ea u es.
Ch oma Ene gy No malized S a is ics [52]. I is a ch omag am based on sho - ime s a is-
ics o e ene gy dis ibu ions bands, ha c ea es a high le el o abs ac ion abou ch omag am.
CENS (Ch oma Ene gy No malized S a is ics) cons i u e a amily o scalable and obus audio
ea u es which ha e i s been in oduced in [53]. These ea u es s ongly co ela e o he sho -
ime ha monic con en o he unde lying audio signal, and hey a e e y obus o dynamics, im-
b e, a icula ion, execu ion o no e g oups, and empo al mic o-de ia ions. Mo eo e , his ype o
ch omag am is one o he mos e icien on ime cos .
Bea synch onous ch omag am [1]. Ano he al e na i e o ch omag am ha in ol es he
use o cen oid calcula ions o localise peaks on he ime- equency su ace o spec ally spa se
2.4 Ha monic Desc ip ion o Musical Audio 21
signals, as i is explained in [1], p o iding imp o ed ampli ude and equency es ima es. This
e ined me hod es ima es an accu a e compu a ion o he associa ed ch oma alues.
2.4.3 Tonal Spaces
Se e al onal pi ch spaces ha e been p esen ed in he li e a u e since he Eule onne z [23]. These
onal spaces allow us o measu e dis ances be ween se s o pi ches. This dis ance is calcula ed
based on how p oxima e he se s o pi ches a e pe cei ed, acco ding o Wes e n musical adi ion.
Pi ches, cho ds, and egions (o keys) a e o he models o desc ibe he ela ionship be ween cho ds
ha a e used as desc ip o s oo.
Mos a e based on Tonne z, he his o ical inspi a ion o all o hem. I is a concep ual la ice
diag am ep esen ing onal space i s desc ibed by Leonha d Eule in 1739 [23]. As shown in
Figu e 2.9, wi h he Tonne z g aph is sui able o ep esen mos o he cho d ypes, modelling o e
his space ela ionships (in e als) be ween di e en pi ches.
Figu e 2.9: T adi ional Tonne z ep esen a ion.
Tonne z is a g aph model ha deno es, acco ding o he classic Wes e n ha mony, he p oximi y
be ween no es. I allows isualizing mos o he cho ds in a geome ical way. The a he away
one is in he g aph, he mo e pe cep ually i is acco ding o he classical ha mony. The ho izon al
lines deno e pe ec i hs, he down igh diagonals deno e majo hi ds and up igh diagonals,
mino hi ds, as shown in e 2.9.
La e , o he mo e complex s uc u es ha modelled mo e cha ac e is ics o ha mony we e
p oposed. One o hem is Shepha d Helix [77] model, shown in Figu e 2.10. Subsequen ly,
models explained in he nex sec ions a e inspi ed by Sepha d helix.
28 S a e o he A
P ep ocessingAudio
Ch omag am
Tonal space
Smoo hing and
dis ance
calcula ion
HCDF
Figu e 2.18: P ocessing blocks o he HCDF.
esea ch assumes ha abo e 5 kHz ha monic con en is no in e es ing gene ally.
2.5.2 De ec ing ha monic change in musical audio
The Ha monic Change De ec ion Func ion o Ha e e al. [31] aims o ex ac cho d bounda ies
om musical audio, aking ad an age o he ac ha dis ances can be measu ed be ween mapped
audios on he model p oposed in Ha e e al’s pape . The pi ch space model p ojec s audio in-
o ma ion as Tonal Cen oid Space ci ed abo e 2.4.3. Then i is heo e ically possible o measu e
he dis ance be weeen onal cen oids, i he dis ance is wide enough hen a ha monic chage has
occu ed.
Figu e 2.18 shows he ou p ocessing blocks o he Ha e e Al. HCDF and he da a lux
be ween hem. Ha e e al. [31] HCDF p ep ocessing block includes he spec al analysis o 8192
sample windows (≈743msec) wi h a 50% o e lap om a mono musical audio inpu a 11025 Hz
sample a e. Cons an -Q spec al analysis is adop ed wi h 36 bins-pe -oc a e ac oss he 110-3520
Hz equency ange.
A 12-elemen ch omag am is hen ex ac ed wi h he me hod desc ibed in [30]. Ch omag ams
a e hen mapped o a onal space, o enhance he pe cep ual ep esen a ion o he ha monic con en
as dis ances. The HCDF a a gi en ime ame nis hen compu ed using he Euclidean dis ance
be ween he ime ame n−1 and n+1. Smoo hed by a 17-elemen Gaussian unc ion wi h
5≤σ≤20. Finally, as shown in Figu e 2.17 peaks in he HCDF a e conside ed ha monic changes.
2.5 Ha monic Change De ec ion 29
2.5.3 Ha monic Change De ec ion o Musical Cho ds Segmen a ion
Degani e al. esea ch [21] shows ha pa ame e iza ions o di e en audio ea u es wi h new me h-
ods p oposed in he las yea s ha e a signi ican pe o mance impac . In pa icula , in his esea ch
cosine dis ance, Non-Nega i e Leas -Squa es CNNLS ch omag am and Ch oma DCT-Reduced log
Pi ch a e es ed among Ha e e al. i s pipeline [31]. This esea ch ou pu s e y good esul s.
Howe e , h esholding is used wi h a cons an alue o e he HCDF gene a ed by a li le se o
anno a ed songs. I Degani e al. HCDF is es ed on o he ypes o songs hen he model does no
gene alize well, pe o ming wo se han Ha e e al.’s HCDF.
30 S a e o he A
Chap e 3
Re isi ing Ha monic Change De ec ion
Figu e 3.1: HCDF diag am.
In his Chap e , we e isi each o he componen blocks o he HCDF (Figu e 3.1) as de ined by
Ha e e al. [31] in ligh o he ecen ad ances in ha monic-con en desc ip ion and ans o ma ion,
as well as hei ela ed signal p ocessing me hods. A comp ehensi e lis o di e en algo i hms
pe block is conside ed in each o he sec ions o his chap e .
31
32 Re isi ing Ha monic Change De ec ion
3.1 P ep ocessing
The p ep ocessing s age in he HCDF is esponsible o c ea ing a spec al ep esen a ion om a
ime-domain audio signal (i.e., audio wa e o m). The audio unde conside a ion can be ei he a
single song in a digi al audio o ma , such as a .WAV o .AIFF, o can be a collec ion o mul iple
audio iles. In his sec ion, we de ail he spec al analysis and i s pa ame e iza ion (e.g., window-
ing, o e lap) as well as a il e ing p ocessing s age which aims o enhance he ha monic con en
in he digi al ep esen a ion using he Ha monic-Pe cussi e Sou ce Sepa a ion (HPSS) [24]. Fig-
u e 3.2 shows he modula algo i hmic s uc u e o he p ep ocessing module.
Figu e 3.2: P ep ocessing block diag am.
F om he audio wa e o m ep esen a ion, spec al analysis is ypically done by applying he
as Fou ie ans o m [55], and in some pa icula cases, no ably when equi ed by he ch oma-
g am compu a ion, Cons an -Q spec al analysis can be adop ed. The Cons an -Q ans o ma ion
is be e adap ed o human pe cep ion as i s equency ep esen a ion is loga i hmic. The e o e, he
lowe no es a e accumula ed in a ew equencies while he highe he equency he less he no es
a e accumula ed. This means ha i he same esolu ion is used in he low equencies he esolu-
ion is poo and in he high equencies i is excessi e. Since he ou pu o he ans o ma ion is
e ec i ely ampli ude/phase e sus loga i hmic equency, ewe equency in e als a e equi ed
o co e a gi en ange e ec i ely, and his is use ul when equencies span se e al oc a es. In
Figu e 3.3 one can see he di e ences om a isual pe spec i e he di e ences be ween FFT and
Cons an -Q ans o m.
3.1 P ep ocessing 33
Figu e 3.3: FFT (le ) spec og am and Cons an -Q ans o m ( igh ) spec og am. Compu ed
om lib osa [47] umpe audio example.
The ange o human hea ing co e s app oxima ely en oc a es om 20 Hz o abou 20 kHz.
The e o e in music p ocessing only his ange o da a has o be analyzed. The ans o ma ion
exhibi s a educ ion in he equency esolu ion o highe equency bins, which makes sense
om an audi o y pe spec i e. As we ha e al edy men ioned, he Cons an -Q ans o ma ion has
a highe esolu ion a lowe equencies, which also comes close o human pe cep ion. Fo he
lowes no es o a piano (abou 30 Hz), a semi one is a di e ence o abou 1.5 Hz, while o he
highes no es o he piano, a he 5 kHz ange, a semi one can be a di e ence o abou 200 Hz.
The e o e, o musical da a he exponen ial equency esolu ion o he Cons an Q ans o ma ion
has ad an ages among he FFT.
In addi ion, he di e en ones s uc u e a cha ac e is ic pa e n in he Cons an Q ans o m.
Assuming he same ela i e powe o each ha monic, i he F0 changes, he ela i e posi ion is
cons an . This can help o segmen ins umen s bu also o ind he undamen al sequences.
In ela ion o he Fou ie ans o m, he implemen a ion o his ans o ma ion is mo e com-
plica ed. This is because each in e al equi es a ce ain sample a e, so he window unc ion is
adap i e and has o a y acco ding o he in e al. Also, since he scale is loga i hmic he e is no
ze o/DC equency, and his can be a p oblem in ce ain cases.
T ansien noise [80] is one o he bigges p oblems [29] in HCDF algo i hms [28,31]. This
p oduces sudden changes in he HCDF, as Ha e men ioned on his hesis [29]. A signal is said o
ha e a ansien noise when i s Fou ie expansion equi es an in ini e numbe o sinusoids [80].
On he o he hand, any signal exp essible as a ini e numbe o sinusoids is called as a s eady-s a e
signal. When he e is wa e o m discon inui y he e is a ansien . Howe e , in digi al audio do-
main, o de ine ansien s is a he di icul . As s a ed in [80], one can ask which sounds should be
"s e ched" and which should be ansla ed in ime when a signal is "slowed down"? In he case
o speech, o example, sho consonan s would be conside ed ansien s, while owels and sibi-
lan s such as "ssss" would be conside ed s eady-s a e signals. Pe cussi e sounds and o e ones a e
gene ally conside ed ansien s. Mo e gene ally, almos any a ack is conside ed a ansien [80].
34 Re isi ing Ha monic Change De ec ion
To his end, we conside he adop ion o he Ha monic Pe cussi e Sou ce Sepa a ion (HPSS)
algo i hm using median il e ing [24] o decompose an audio signal in o wo ha monic and pe -
cussi e sou ces, om which we uniquely adop he o me o u he p ocessing. This s ep aims
o exclude ansien equencies ha ha e a nega i e impac on p o iding an op imal ansc ip ion
o he audio signal’s ha monic con en . We adop his algo i hm a e donwnsampling he musical
audio and p io o he spec al audio analysis.
Di e en pa ame e s o sample a e s, window size , and o e lap o, a e conside ed in he
p ep ocessing o he audio iles. On he one hand, la e pa ame e s ha e an impac on he e i-
ciency o he sys em by changing he amoun o audio sample da a o be p ocessed. Lowe s, and
highe alues and o esul in mo e e icien HCDF compu a ion. On he o he hand, p e ious
s udies ha e shown ha hese pa ame e s can ha e an in luence on he ou comes o mul iple asks.
Windowing and downsampling a e echniques widely used [51] o simpli y and s anda dize ep-
esen a ions. They can be ega ded as a bandwise lowpass il e ing, which a e used o a enua e
as luc ua ions in he ea u es ep esen a ion. Downsampling is o en used o e ec i ely inc ease
equency esolu ion a lowe equencies, al hough one has o keep in mind ha he maximum
equency analyzed will be hal o he sample a e by Nyquis [56]. O en, hese echniques a e
also used o educe he cos o compu ing in exchange o he loss o in o ma ion.
3.2 Ch omag am
Chomag ams a e pe asi e ac oss mos ha monic-based audio-con en desc ip ion. They encom-
pass mul iple a ian s d i en om a numbe o p oposed algo i hms, as we de ailed in Sec ion
2.4.2. In he con ex o ou wo k on HCDF, we ha e selec ed ou ep esen a i e ch omag ams:
Cons an -Q T ans o m ch omag am CCQT, Non-Nega i e Leas -Squa es ch omag am CNNLS [42],
Ha monic Pi ch Class P o ile ch omag am CHPCP [27] and Sho Time Fou ie T ans o m ch o-
mag am CSTFT [22]. We ha e ejec ed Bea synch onous ch omag am [1] because we could no
ind an open sou ce implemen a ion a ailable. Ch oma DCT-Reduced log Pi ch [52] and Ch oma
Ene gy No malized S a is ics [52] ha e been ejec ed because hey can be pa ame e ized as he
o he ch omag ams, due o hem being implemen ed wi h a ix il e banks. The ch omag ams
implemen ed wi h deep lea ning ha e no been included in he s udy due o empi ical cha ac e
o he g id sea ch, da a gene a ed in his disse a ion is e y la ge, and a song p ocessed by deep
lea ning akes abou 160 seconds o compu ing ime [89].
In chap e 2, a high le el explana ion is gi en o wha he di e en ch omag ams 2.4.2 a e and
wha hey a e o , om a high le el. In pa icula hose used in his disse a ion: CSTFT,CCQT,
CHPCP and CNNLS. Hence, in his sec ion he wo king o ch omag ams will be de ailed om a low
le el pe spec i e.
All he ch omag am used in ou esea ch ha e some common pa s. These ch omag ams a e
implemen ed as a lineal digi al signal p ocessing pipeline. The inpu o his me hod is a digi al
audio ha may ha e been p ep ocessed be o e 3.1. The ou pu is a ch omag am o 12 pi ch classes
whe e he di e en audio equencies ha e been classi ied.
3.2 Ch omag am 35
Figu e 3.4: Sho Fou ie T ans o m ch omag am CSTFT pipeline diag am.
The Sho Time Fou ie T ans o m ch omag am CSTFT has been widely used in he las 20
yea s, in [25] a ull explana ion can be ound. The lineal pipeline o his ch omag am is shown
in 3.4. The i s block o his ch omag am ope a es as a spec og am, aking an audio inpu and
gene a ing a ma ix o sho - ime spec um ames. The second block uses immedia e equency
es ima es om he spec og am o ob ain he ch oma p o iles. Ano he op ion, wi h less compu-
a ional complexi y, is mapping each STFT bin di ec ly o ch oma classes, a e selec ing spec al
peaks. In he las block, each class p o ile o he ch omag am is no malized.
Figu e 3.5: Cons an -Q T ans o m ch omag am CCQT pipeline diag am.
The Cons an -Q T ans o m ch omag am CCQT was p oposed by B own e al. [15]. The lineal
pipeline o his ch omag am is shown in 3.4 and is e y simila o he p e ious ch omag am, CSTFT.
The i s block compu es a CQT spec og am, in p e ious sec ion 3.1 he di e ence be ween his
ype o ans o m and o he ypes o Fou ie ans o ms is explained and his musical ad an ages.
This i s block can be calcula ed by a ull Cons an -Q ans o m o by a hyb id one, wi h a lowe
compu a ional cos . The ou pu is also a ma ix o Cons an -Q spec um ames. The second block
36 Re isi ing Ha monic Change De ec ion
is, as he p e ious ch omag am, mapping he spec og am in o pi ch class p o iles. In he las
block, as in he CSTFT pipeline, class p o iles o he ch omag am a e no malized.
Figu e 3.6: Ha monic Pi ch Class P o ile ch omag am CHPCP pipeline diag am.
The Ha monic Pi ch Class P o ile ch omag am CHPCP was p oposed by Emilia Gomez e
Al.[27], he pipeline o his ch omag am is shown in Figu e 3.6. On he i s block he Cons an -Q
spec og am is compu ed om a unning ecuency. Then, he second block emo es edundan
in o ma ion, i s disca ding equencies beyond a low pass h eshold and a highpass h eshold
and la e h esholding he spec og am wi h a global and local ( ame-wise) h eshold. The hi d
pipeline blocks o compu e he peak in e pola ion o ob ain spec al peaks wi h a high esolu ion.
Then, in he ou h block, each o hose peaks a e assigned o he pi ches o he oc a es ha we e
no disca ded. In he i h block, one ies o elimina e he ene gy edundancies p oduced by he
ha monic pa ials. And inally, all he oc a es a e mapped in one oc a e, he Pi ch Class P o iles.
Figu e 3.7: Non-Nega i e Leas -Squa es ch omag am CNNLS pipeline diag am.
3.3 Tonal Space 37
The las ch omag am is he Non-Nega i e Leas -Squa es ch omag am CNNLS. I was p oposed
by Mauch e al. [42] and on ha esea ch a ull desc ip ion o he pipeline shown in he Figu e 3.7
can be ound. The i s block calcula es uning om NNLS, using he angle o he complex num-
be de ined by he cumula i e mean o eal and imagina y alues. In he second block is compu ed
he spec og am, using he uning es ima ed alue, an app oach o undamen al equency o pe -
o ming linea in e pola ion on he exis ing log- equency spec og am wi h all he pa ials. In he
hi d block, a Semi one-spaced log- equency spec um de i ed om he uned log- eq spec um
abo e is compu ed. The spec um is in e ed using a non-nega i e leas squa es algo i hm. In
he las block, h ee di e en kinds o ch omag am a e calcula ed, a eble ch omag am in highe
equencies, a bass ch omag am in lowe equencies, and a combina ion o bo h.
3.3 Tonal Space
Tonal In e al Vec o s, T(k), [4] ex end he Tonal Space by Ha e e al.’s [31] o he en i e se o
complemen a y in e als wi hin he 12 ch oma ic no es o he equal empe ed pi ch class space
(i.e. all complemen a y in e als esul ing in he 12-elemen s ch omag am ep esen a ion space).
The ex ended ec o space p ojec s he mos salien pi ch le els o onal Wes e n music – pi ches,
cho ds and keys – as unique loca ions in he space. The esul ing spa ial loca ion o T(k)en-
su es ha pe cep ually- ela ed pi ch wi hin he Wes e n onal music con ex co espond o small
Euclidean dis ances [4].
Depa ing om a ch omag am ep esen a ion, C, we compu e a 12-dimensional T(k)as he L1
no malized disc e e Fou ie ans o m (DFT), such ha :
T(k) = w?(k)
N−1
∑
n=0
¯
C(m)e−j2πkm
M,k∈Zwi h ¯
C(m) = C(m)
∑M−1
n=0C(m)(3.1)
whe e M=12 is he dimension o he ch omag am, C;kis se o 1 ≤k≤6 since he emain-
ing coe icien s a e symme ic; T(k)uses ¯
C(m)which is C(m)no malized by he DC com-
ponen T(0) = ∑M−1
n=0C(m) o allow he ep esen a ion and compa ison o di e en hie a chical
le els o onal pi ch [4]; and w?(k)a e weigh s which egula e he impo ance o each coe -
icien (o in e p e ed musical in e al) in T(k). Th ee se s o weigh s a e adop ed. ws(k) =
{2,11,17,16,19,7}was p oposed o ch omag ams d i en om symbolic musical mani es a-
ions [4], wa(k) = {3,8,11.5,15,14.5,7.5}was p oposed o musical audio [7] and wh(k) =
{0,0,1,0.5,1,0}p o ides Ha e e al.’s [31]. The wo o me se o weigh s esul om empi -
ical consonance a ings o dyads, used o adjus he con ibu ion o each dimension ko he space
(o in e p e ed musical in e al), making i a pe cep ually ele an space in compa ison o i s non-
weigh ed e sion [4].
44 E alua ion
me ics o he HCDF a e sa ed in a human- eadable o ma , namely using JSON o ma .
Figu e 4.4: Da abase en i y/ ela ionship model.
All he da a is sen o a Pos g e SQL da abase, whe e i is s o ed in a s uc u ed manne in a
simple da abase schema, as seen in he en i y- ela ionship model ( Figu e 4.4).
Figu e 4.5: Final da abase diag am.
Final HCDF da ase is a ailable a : h ps://d i e.google.com/d i e/ olde s/
1a-SzgmqP 7DmSnNaXAZM8b1MuB 1Id1A?usp=sha ing. I is compound by a README.md
whe e i s s uc u e is explained and a di ec o y ile whe e all he da a is a ailable. Each ile o da a
4.2 Technical A chi ec u e: Implemen ing an E icien G id-Sea ch Me hod 45
includes he pa ame iza ion esul s o each es ed song. To easily pa se he da a, each ile is
named desc ip i ely. The name ile adop ed is composed by he pa ame iza ion es es. On he
ile name, each pa ame e o he pa ame e iza ion is sepa a ed by h cha ac e ".", wi h he o de :
HPSS, onal model, ch oma, sample a e, window size, o e lap, blu and dis ance. Fo example,
in Code 4.2 he ile name would be hcd .8000.1024.0.hpcp.T(w h).sigma11.euclidean.json. Inside
he JSON ile, he i s eigh en ies a e pa ame e s wi h possible alues de ined in Table 4.1 wi h
he name: sample a e, window size, o e lap, HPSS, onal model, ch oma, σsmoo hing il e , and
dis ance. The JSON ile includes he pa ame iza ion alues adop ed, ollowed by a lis o songs
wi h hei espec i e HCDF e alua ion me ics (i.e., -sco e, ecall, p ecision and he ha monic
changes). To al size o he esul s amoun o 128 GB.
{
"hpss": alse,
" onal_model":"T(w_h)",
"ch oma":"hpcp",
"sample_ a e":8000,
"window_size":"1024",
"o e lap":0,
"blu ":" ull",
"dis ance":11,
" esul s": [
{
"song":"01_ -_A_Ha d_Day's_Nigh ",
"ha monic_change": [
0.0,
1.1542631509874295,
...
144.57145966117554,
146.68760877131916
],
" _sco e":0.524390243902439,
"p ecision":0.6825396825396826,
" ecall":0.42574257425742573
},
{
"song":"01 A Kind O Magic",
"ha monic_change": [
0.0,
0.37655172413793103,
...
250.21862068965518,
46 E alua ion
253.13689655172413,
257.93793103448274
],
" _sco e":0.33136094674556216,
"p ecision":0.2978723404255319,
" ecall":0.37333333333333335
},
...
]
}
Visualisa ion sc ip s.
Abo e all, di e se unc ions ha e been implemen ed o ind he di e en ela ionships be ween
he hype pa ame e s and be ween hese hype pa ame e s and speci ic da a, such as he mode o key
o a song. Va ious p og ams ha e also been buil o be able o isualize ei he he HCDF o se e al
HCDFs a he same ime and be able o compa e hem.
Di e en sc ip s ha e been implemen ed, o pe o m he unc ions explained in he ollowing
pa ag aph. To measu e he a e o ha monic change in all he di e en songs analysed and c ea e
a ious isualisualiza ions om he da a o a e o ha monic change. To se h esholds o he
HCDF o o e i i as done in p e ious a icles o o es adap i e h esholding ideas. O o ix
a ious bugs ha we had du ing he p ocess.
Suppo sc ip s.
Sc ip s ha e been made o popula e he da abase wi h da a p o ided by he da ase and ad-
minis e i la e . Tools ha e been implemen ed o sa e backups o he inal da a o o moni o
how much da a was missing. Sc ip s ha e also been c ea ed o moni o i any es had no been
pe o med co ec ly.
The dis ibu ed sys em has been main ained h ough sc ip ing and moni o ing. To be able o
access each o he se e s. To moni o hei space a ailabili y, he lab ne wo k s a us and he
sys em s a us. To eboo he sys em o jus a speci ic node. O o ee up memo y.
4.3 Resul s
The comple e se o esul s o he g id sea ch me hod can be accessed online a : h ps://
d i e.google.com/d i e/ olde s/1a-SzgmqP 7DmSnNaXAZM8b1MuB 1Id1A?usp=
sha ing. In his sec ion, we p esen and discuss he esul s o he algo i hms and pa ame e s
which exhibi highe -sco e and ecall, as hey play a p ominen ole in suppo ing applica ions
wi hin con en -based digi al audio p ocessing (e.g. analysis, e ie al, and ans o ma ion), in
de imen o highe p ecision. Highe -measu e p o ides a balanced p edic ion o cho d bound-
a ies, ele an o c ea i e applica ions as ha moniza ion [38], adap i e digi al audio e ec s [61]
o gene a i e isuals om music [14]. On he o he hand, highe ecall gua an ees an excellen
4.3 Resul s 47
Figu e 4.6: Typical alse posi i es caused by bass and pe cussi e sounds.
esolu ion as a p ep ocessing segmen a ion s age o asks such as au oma ic cho d ecogni ion
(ACR) [58,88].
Table 4.2 shows he F,Pand Re alua ion me ics o he o e all collec ion unde conside a ion
and o each da ase . These me ics a e shown o he bes -sco e and ecall esul s a ising om
he op imal combina ion o algo i hms and pa ame e s in he g id sea ch me hod. Fo compa ison
pu poses, we include in Table 4.2 he esul s om he Ha e e al. ’s me hod as p esen ed in [31],
which we e alua ed in he g id sea ch me hod in e ms o bes he Gaussian il e ing σpa ame e
only, as i emained open in he o iginal con ibu ion.
The o al numbe o combina ions ac oss all algo i hms and pa ame e s condi ions adop ed in
each ins ance o he g id sea ch me hod is 8890560. The bes -sco e esul s adop s a sample a e
o s=8000 Hz, a window size o =1024 and a window o e lap o o=50%. Musical audio
is p ep ocessed using HPSS. Ha monic con en ep esen a ion esul s om CNNLS ch omag am
u he encoded and p ojec ed on he onal model T(k)wi h wa. Fo compu ing he HCDF ξ, we
adop he Euclidean dis ance ξeucl
na e Gaussian smoo hing wi h σ=5. The bes ecall esul s
adop s a sample a e o s=44.100 Hz, a window size o =2048 and a o=25% o e lap.
P ocessed ha monic con en ep esen a ion esul s om he ha monic sou ce only om he HPSS
algo i hm, which is u he encoded as an CSTFT ch omag am and p ojec ed on he onal model
T(k)wi h wa. Fo compu ing he HCDF ξ, we adop he Euclidean dis ance ξeucl
na e Gaussian
smoo hing wi h σ=17.
The bes -sco e esul s in he en i e collec ion o ou da ase s imp o e by 5,57% in compa i-
son wi h he p e ious Ha e e al. [31] me hod, wi h no iceable 6,28% gains in e ms o ecall. The
indi idual da ase esul s show conside able imp o emen s in he Queen and The Bea les da ase s,
sugges ing ha in lowe a es o ha monic change a e simple o achie e in his ype o gen es. I
MTG-JAAH da ase is no aken in o accoun in he scena io o he compu ing he bes -sco e, ou
me hod imp o es by 8.29% agains Ha e e al. me hod.
The adop ion o he onal model T(k)wi h wahas shown o imp o e he esul s ac oss all
48 E alua ion
Table 4.2: HCDF e alua ion indica o s o he bes -sco e and ecall me hods ac oss he ou
da ase s unde s udy, o which we compa e wi h p e ious me hods. The a e age is compu ed o e
all pieces.
Bes -sco e P oposed Bes Recall P oposed
Da ase i le F P R F P R
MTG-JAAH 63,3% 62,7% 72,2% 49,9% 35,8% 98,4%
Queen 65,2% 60,1% 75,8% 43,3% 28,5% 98,8%
The Bea les 69,1% 62,8% 81,9% 42,5% 27,7% 99,5%
Zweieck 68,6% 61,7% 81,0% 43,3% 28,0% 99,1%
A e age 66,4% 62,4% 77,4% 45,6% 31,1% 99,0%
Bes -sco e Ha e Bes Recall Ha e
Da ase i le F P R F P R
MTG-JAAH 61,7% 58,4% 74,5% 43,5% 29,6% 99,2%
Queen 57,9% 49,7% 74,71% 37,0% 23,3% 99,7%
The Bea les 60,5% 51,0% 79,9% 36,6% 22,9% 99,8%
Zweieck 61,0% 50,8% 79,3% 35,9% 22,2% 99,4%
A e age 60,8% 53,9% 77,2% 39,4%25,6%99,5%
F -sco e Pp ecision R ecall
da ase s, as hey p ominen ly appea in he bes - anked -sco e and ecall esul s om he g id
sea ch (75 and 58 ou o he 100 bes - anked pa ame e iza ions, espec i ely). The app oxima e
ansc ip ion p o ided by he CNNLS ch omag am is equally iden i ied in he op- anked g id -
sco e sea ch esul s (98 ou o 100 bes - anked g id sea ch’s pa ame e iza ions).
Median- il e ing ha monic pe cussi e sou ce sepa a ion il e [24] imp o es he -sco e mea-
su e (60 ou o 100 bes - anked g id sea ch’s pa ame e iza ions). In pa icula , by up o 2% on he
60 ci ed bes anked -sco e g id sea ch esul s wi h HPSS.
Window o e lap o 50% o e lap p o ide enhanced esul s. Sample a es o 8000 samples
pe second a e enough o ge ing good HCDF pe o mance. This allows, by being smalle he
window, ha HCDF has a compu a ional cos in less ime (mo e e icien ).
Ha e e al. [31] HCDF pa ame iza ion p oposed 15 yea s ago sco es wi h lowe esul s, as
shown in Table 4.3 han he o he algo i hms.
Ou bes esul s co obo a e Ha e e al. [31] indings ha he adop ion o a onal model con-
ibu es o a high ecall sco e Ra he expense o p ecision P. Ou esul s sugges ha he new
algo i hms and pa ame e s a e be e a disc imina ing signi ican ha monic changes in he signal,
while signi ican ly educing he numbe o alse posi i es in he HCDF ξ. Fu he mo e, as sug-
ges ed in [31], hese low p ecision sco es Pcan be explained by he ac ha he g ound u h
4.3 Resul s 49
Table 4.3: Expe imen al esul s o HCDF peaks compa ed wi h hand labelled cho d changes o
16 Bea les Songs (songs a anged in ch onological o de o elease da e). Please Please Me (1),
Do You Wan To Know A Sec e (2), All My Lo ing (3), Till The e Was You (4), A Ha d Day’sDay’s
Nigh (5), I I Fell (6), Eigh Days A Week (7), E e y Li le Thing (8), Help! (9), Yes e day (10),
D i e My Ca (11), Michelle (12), Eleano Rigby (13), He e The e And E e ywhe e (14), Lucy In
The Sky Wi h Diamonds (15), Being Fo The Bene i O M Ki e (16).
Bes P ecision P oposed Bes Recall P oposed Bes Ha e
Song # F P R F P R F P R
1 76% 74% 77% 52% 35% 100% 65% 53% 87%
2 73% 89% 62% 70% 54% 98% 75% 72% 80%
3 79% 74% 84% 44% 28% 100% 63% 50% 86%
4 77% 78% 76% 56% 39% 100% 71% 59% 90%
5 81% 78% 83% 50% 34% 100% 54% 45% 70%
6 87% 86% 72% 56% 39% 100% 65% 79% 53%
7 75% 74% 76% 48% 32% 100% 61% 49% 82%
8 74% 80% 69% 55% 38% 94% 56% 49% 67%
9 65% 53% 84% 35% 21% 100% 41% 29% 74%
10 70% 73% 67% 62% 45% 98% 73% 64% 86%
11 72% 67% 78% 45% 29% 100% 62% 49% 86%
12 76% 71% 81% 47% 31% 98% 66% 53% 90%
13 75% 63% 90% 35% 21% 100% 47% 34% 81%
14 81% 77% 84% 55% 38% 100% 72% 65% 83%
15 78% 76% 79% 47% 31% 99% 63% 50% 88%
16 84% 86% 82% 54% 37% 100% 79% 68% 95%
A e age 75% 74% 78% 51% 34% 99% 64.9% 53% 84%
F -sco e Pp ecision R ecall
50 E alua ion
Figu e 4.7: Subjec i i y in anno a ions. Beginning o The Bea les’ song Please Please Me. Sco e
and HCDF diag am o e spec og am. F om 2 o 11 e e y bounda y is a alse posi i e. The
numbe s a e in he same empo al posi ion on he sco e as on he spec og am.
anno a ions only label cho d changes. Howe e , HCDF ξis also sensi i e o changes in ha monic
con en caused by s ong melody o bass line mo emen s ha include non-cho d ones. Thus,
a high numbe o alse posi i es is o be expec ed o his expe imen and does no necessa ily
deno e a pe cep ual phenomenon.
As shown in [54] and [36], besides he high complexi y o polyphonic music and he subjec-
i i y o i s anno a ions, he adop ed g ound u h is no exac ly p one o he ask a hand om a
pe cep ual iewpoin , ye no only o compa a i e, and legacy pu poses a e s ill adop ed in ou
wo k, bu also due o he impo ance o HCDF in suppo ing ACD sys ems. Mos o he alse
posi i es in he HCDF ξcan be seen in Figu e 4.7, one could a gue ha he ha monic change is
co ec when seen om a non- unc ional pe spec i e. Indeed, a slice by slice analysis o he piece
would yield he same esul s as he HCDF ξ: al hough unc ionally he e is a sus ained E cho d,
he no es ha a e being played a each poin desc ibe a succession o di e en cho ds, as one can
see in Figu e 4.7. When he bass line mo es a lo , i can gene a e alse posi i es, as shown in
Figu e 4.6. Some imes simila e o s can occu when qui e la ge a ia ions o pe cussi e sounds
a e combined wi h he ha mony. This is sol ed mainly by he HPSS il e e en in ch omag ams
ha should be obus o bass wi h a la ge numbe o ha monics.
I i is analyzed a small and biased da ase 4.3 as which was analyzed by Ha e e al. [31]
due o no exis ence o da ase s 15 yea s ago, esul s a e e y imp essi e. Tonal in e al space
combined wi h CNNLS ou pe o ms Ha e e al. esul s in p ecision o ien ed algo i hm and in ecall
o ien ed algo i hm.
4.4 Visualizing Ha monic Change De ec ion Func ion 51
4.4 Visualizing Ha monic Change De ec ion Func ion
In his sec ion, we will discuss he HCDF depa ing om i s g aphical ep esen a ion aiming o
be e g asp he implica ion o he mul iple algo i hms and pa ame e iza ions p oposed. To his
end, mul iple isualiza ions o he 60 i s seconds o song Please Please Me by The Bea les
using di e en HCDF condi ions a e plo ed.
4.4.1 Visualizing -sco e esul s
Figu e 4.8: HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF
Figu e 4.8 shows he HCDF o he bes -sco e esul om he g id sea ch 4.3. The pa ame e i-
za ion whose ou pu om he g id sea ch ha e been selec ed and applied o Please Please Me. As
i can see, in he abo e-men ioned HCDF diag am, he peaks ep esen ha monic changes. I he
peaks a e highe (o g ea e magni ude), i means ha he ha monic change is highe om a onal
pe spec i e. As i can be easily seen in he onne z diag am 2.9 hi ds and i hs a e close o o he
in e als be ween cho ds, mo eo e , Tonal In e al Space p oposed by Gilbe o e Al. [4] expands
he in e als ecognized, as is explained in Chap e 2. Mo eo e , F-sco e measu e e alua ed wi h
g ound- u h anno a ions is 75,3% wi h a balanced p ecision/ ecall; p ecision is 73,4% and ecall
is 77.2%.
52 E alua ion
Figu e 4.9: HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF wi hou HPSS me hod in p ep ocessing block
Figu e 4.9 shows an HCDF wi h a pa ame e iza ion om he bes -sco e esul wi hou HPSS.
This allows us o isually compa e he beha iou o he HCDF when adop ing he HPSS. The
HCDF pa ame e iza ions, adop s a sample a e o s=8000 Hz, a window size o =1024 and
a window o e lap o o=50%. Ha monic con en ep esen a ion esul s om CNNLS ch omag am
u he encoded and p ojec ed on he onal model T(k)wi h wa. Fo compu ing he HCDF ξ, we
adop he Euclidean dis ance ξeucl
na e Gaussian smoo hing wi h σ=5.
When HPSS is adop ed, we can obse e ha he magni ude o he peaks is educed, mos
p obably as he esul o he lack o pe cussi e ansien s. Despi e he esul s imp o emen o
2% when adop ing HPSS, in his pa icula song -sco e is 76,4%, p ecision 71,4% and ecall is
82.2%. Al hough he esul s a e be e he di e ence be ween p ecision and ecall is highe .
Figu e 4.10: HCDF diag am. Fi s 60 seconds o Please Please Me . Bes -sco e esul o g id
sea ch HCDF wi hou onal model.
Figu e 4.10 shows he bes esul o -sco e wi hou onal space. Smoo hing and dis ances
a e compu ed di ec ly on he ch omag am. This HCDF pa ame e iza ions adop s a sample a e o
s=8000 Hz, a window size o =1024 and a windows o o=50% o e lap. Musical audio is
4.4 Visualizing Ha monic Change De ec ion Func ion 53
p ep ocessed using HPSS. Ha monic con en ep esen a ion esul s om CNNLS. Fo compu ing
he HCDF ξ, we adop he Euclidean dis ance ξeucl
na e Gaussian smoo hing wi h σ=5.
Con e sely o Figu es 4.9 and 4.8, he peaks in Figu e 4.10 a e e y homogeneous in e ms
o magni ude. F-sco e is 72,1%, p ecision 67,1% and ecall is 77.2%. The -sco e is 10%, lowe
han p e ious pa ame e iza ions (o isualiza ions), and he di e ence be ween p ecision/ ecall is
h ee imes wide han Figu e 4.8. Ch omag ams wi h a low compu a ional cos likeCCQT o CSTFT
pe o m well unde his se o pa ame e s.
Figu e 4.11: HCDF diag am. Fi s 60 seconds o Please Please Me. Bes -sco e esul o g id
sea ch HCDF wi h W(h)as a onal model.
Figu e 4.11 adop s he pa ame e iza ion o iginally p oposed in Ha e e al. In pa icula , i
adop s a sample a e o s=44100 Hz, a window size o =2048 and a window o e lap o
o=25%. Musical audio is p ep ocessed wi hou HPSS. Ha monic con en ep esen a ion esul s
om CSTFT ch omag am u he encoded and p ojec ed on he onal model T(k)wi h wh. Fo
compu ing he HCDF ξ, we adop he Euclidean dis ance ξeucl
na e Gaussian smoo hing wi h
σ=17.
The esul ing HCDF has in Figu e 4.11 has less peak magni ude a iance han onal spaces
based on Tonal In e al Space w(a), as shown in he Figu e 4.9 and 4.8. We belie e ha hese
esul s s em om he added in e allic ela ions in he onal space T(k)wi h wso wa, which
p omo e enhanced di e ence in iadic ha mony abo e he se en hs cho ds. In his case -sco e
equals 63,7%, p ecision 55,6% and ecall 74.68%.
60 Applica ions o HCDF
wi h his in o ma ion [33,13]. Ano he a ge is o ge an ACR sys em close o unc ional ha mony
analysis, which has a highe in o ma ion densi y [72], al hough his ep esen a ion does no ha e
demons a ed bene i s in ACR.
HCDF is usually adop ed ea ly in he ACR algo i hmic pipeline. I aims o segmen he signal
in o s uc u ally-awa e uni s (o cho d bounda ies). Recen s a e-o - he-a me hods o ACR use
deep-lea ning a chi ec u es wi h HCDF [88]. Rela ed ACR li e a u e o ou wo k can be ound in
he wo k [34], which adop he Tonne z as a base ep esen a ion, in a simila ashion as he onal
space he e p oposed.
5.2 C ea i e applica ions
In his sec ion, wo c ea i e applica ions o HCDF will be p esen ed: Audio Visualiza ions and
Musical Audio Ha moniza ion. In o de o gi e a obus idea o he mul iple applica ions ha
HCDF can ha e.
5.2.1 Audio Visualiza ions
Visualiza ions o music ha e been pe asi e ac oss he widely popula music playe s. mos hu-
mans ecognize isualiza ions o Windows media playe o SoundCloud, as is shown in Figu e 5.1.
Howe e , o c ea e hese isualiza ions ha change wi h music, usually, hy hm con en has been
used. Wi h he analysis o pa ame e iza ion and pe o mance o e di e en ha monic a ibu es
as hose de ailed in his disse a ion, i could be possible o expand hese music isualisa ions
pla o ms. Fo example, HCDF peak magni udes, deno ing ha monically p oximi y ac oss ime,
can p o ide ha monically-awa e segmen a ion and pe cep ually a ini y ac oss ime ha can be
mapped o isual se ings.
5.2 C ea i e applica ions 61
Figu e 5.1: Some music indus y isualiza ion sys ems.
5.2.2 Musical Audio Ha moniza ion: D’acco d
Ano he c ea i e applica ion ha can make use o HCDF is he au oma ic ha moniza ion o music.
D’acco d is a ep esen a i e example o a gene a i e music sys em o c ea ing ha monically com-
pa ible accompanimen s [5]. I ha monizes musical audio by accompanimen s wi h use -speci ied
numbe o oices, ins umen a ion and complexi y. On he backend o he sys em i elies he
onal space T(k) o ha monic-awa e segmen a ion, key de ec ion and ha moniza ion. D’acco d
was o iginally de eloped o Able on Li e, MAX, and Pu e Da a. Figu e 5.2 shown he in e ace
o D’acco d in Pu e Da a.
62 Applica ions o HCDF
Figu e 5.2: D’acco d desk op in e ace. Diag am ex ac ed om Gilbe o pe sonal si e.
An ongoing e ision o he ool is unde de elopmen as a web-based applica ion. The p o-
posed me hod, and i s ja asc ip implemen a ion, is suppo ing he ha monic change de ec ion
om musical audio in he b owse .
Chap e 6
Conclusions
In his disse a ion, we e isi ed ha monic change de ec ion in ligh o ecen ad ances in ha monic
desc ip ion and ans o ma ion. S emming om Ha e e al.’s [31] HCDF, we p oposed no el
algo i hms o each o hei p ocessing blocks. We exhaus i ely inspec ed he pa ame e s ha bes
pe o ms ac oss he mul iple algo i hm combina ions using a g id sea ch me hod. We e alua ed he
p oposed algo i hms and i s pa ame e iza ion on ou s yle-speci ic da ase s (The Bea les, Queen,
Zweieck and MTG-JAAH).
Di e en p ep ocessing s a egies ha e been applied o a y he sample a e s, windows size
, and o e lap o, as well he adop ion o ha monic-only signal decomposi ion om he median-
il e ing HPSS algo i hm. Fou di e en ch omag ams CSTFT,CCQT,CNNLS and CHPCP and he
ecen ly p oposed onal space T(k)wi h weigh s wa,wsand whha e been es ed. Finally, Eu-
clidean ξeucl
nand cosine ξcos
ndis ance me ics and 1 ≤σ≤17 alues in Gaussian il e ing ha e
been e alua ed in he HCDF pe o mance.
Ou esul s showed ha he newly p oposed algo i hms and pa ame e s imp o ed p e ious
HCDF compu a ion in de ec ing ha monic changes. The adop ion o he Tonal Space T(k)wi h
waand CNNLS imp o ed -sco e and ecall by 5.57% and 6.28%, espec i ely. An assessmen o
he HCDF ξac oss a signi ican ly la ge numbe o musical examples om mul iple s yles no
only imp o es he gene ali y o he me hod o unknown musical audio sou ces bu also sugges s
he link be ween ha monic change a es and pa ame e iza ion o he HCDF.
6.1 O iginal Con ibu ions and Model Implemen a ions
Wi h he objec i e o dissemina ing ou con ibu ion o imp o ing HCDF, we made a ailable he
esul ing HCDF o p ocessing he bes -sco e and ecall esul s. Fu he mo e, he models ha e
63
64 Conclusions
been de eloped as unc ions o eal- ime audio p ocessing in Pu e Da a [64] and o o line p o-
cessing in Py hon 1, Ja asc ip 2. A Ja asc ip implemen a ion o TIVlib 3has been made a ailable.
Fu he mo e, we also dis ibu e a well-s uc u ed da ase con aining all he esul ing da a om he
g id sea ch me hod. I has been uploaded o Google D i e4in a legible o ma , JSON.
6.2 Fu he wo k
In u u e wo k, we aim o ackle he adop ion o pa ame e s pe gen e (o o di e en a es o ha -
monic changes). Mo eo e , we belie e ha an adap i e h eshold o de ec ing he peaks (i.e. ha -
monic changes) can imp o e he peak picking phase in de ec ing cho d bounda ies in he HCDF.
Ul ima ely, his should imp o e p ecision by elimina ing alse posi i es.
In Fu u e esea ch, we would like o imp o e he use o he Tonal In e al Space. Fo example,
o ca y ou Au oma ic Change Recogni ion me hods on he basis o p e ious simila esea ch [34].
We will de elop a compu a ionally-e icien ha monic hy hm de ec o . The ha monic hy hm
o a e o ha monic change is e y c ucial. A compu e ision app oach could be used, songs’
ch omag ams would be as a 2d image wi h g ey colou s (ene gy ha e only one dimension). A
i s sigh i is possible o obse e a ela ionship be ween he ha monic hy hm and he ch omag am
in a song, as is shown in Figu e 6.1.
1A : h ps://gi hub.com/PRamoneda/HCDF.
2A : h ps://gi hub.com/PRamoneda/HCDF.js.
3A : h ps://gi hub.com/PRamoneda/TIVlib.js.
4A : h ps://d i e.google.com/d i e/ olde s/1a-SzgmqP 7DmSnNaXAZM8b1MuB 1Id1A?usp=
sha ing.
6.2 Fu he wo k 65
Figu e 6.1: E olu ion o he ch omag ams depending on he a e o ha monic change.
I would be bene icial o apply neu al ne wo ks o audio T ansien /S eady-S a e sepa a ion o
p ocessing he musical audio be o e applying HCDF. I would also be in e es ing o esea ch he
possibili y o de eloping a de ec o o spu ious peaks in HCDF.
Some ea ly expe imen s on he adop ion o h esholding echniques in he peak picking s age
ha e been pu sued. One o hem is o use a double ha monic change de ec ion unc ion. Fi s one
was op imized o gene a e many changes, and hen he second one was o e he sum o he ch oma-
g ams o he bounda ies ound by he i s one HCDF. This second HCDF op imized o ge he bes
possible esul s. We would use di e en algo i hms in each o he blocks o each o he wo algo-
i hms. We unde s ood ha he o e lap o he wo solu ions would ge a be e solu ion. Howe e ,
he esul s o he expe imen al es s we e no good. Las bu no leas , an op imal adap i e unc ion
could be de eloped h ough deep lea ning, machine lea ning o ein o cemen lea ning combined
wi h digi al signal p ocessing. This adap i e unc ion would imp o e signi ican ly HCDF and we
will keep ou esea ch in his di ec ion.
66 Conclusions
Appendix A
Anexo
A.1 Py hon lib a y
A.1.1 README
# Ha monic Change De ec ion Func ion (HCDF) lib a y
This lib a y is used o compu e HCDF. [He e]() is desc ibed he algo i hm in de ail. As many o he solu ions o e lap pa illy, all he algo i hm da a compu ed in he di e en blocks is sa ed in ou olde in o de o no compu e same blocks pa ame e iza ion wo imes.
## Ins alla ion
Ins all Vamp-plugins:
- (NNLSch oma) h p://www.isophonics.ne /nnls-ch oma
- (HPCPch oma) h p://m g.up .edu/ echnologies/hpcp
Ins all dependencies:
```BASH
pip3 ins all eque imen s. x
```
## Usage
The lib a y can be impo ed as a module wi h `impo HCDF`. All he unc ions han begins by ge a e blocks om HCDF
unc ion. The es a e auxilia unc ions.
HCDF.py ac as sc ip allowing he use o p in o he console Ha monic Change De ec ion Func ion (HCDF)
ocus on maximizing ecall o -sco e. I is assumed ha he i s command line a gumen is
he name ile o he audio ile loca ed in audio_ iles and he second one is, i is ocus on
ecall o -sco e.
### Example o use
67
68 Anexo
Wi h a ge on maximize -sco e:
```BASH
py hon3 -sco e ile/name
```
Wi h a ge on maximize ecall:
```BASH
py hon3 ecall ile/name
```
## Re e ences
h ps://lib osa.gi hub.io
h ps:// amp-plugins.o g
h ps://gi hub.com/a ami es/TIVlib
A.1.2 CODE
"""HCDF py hon implemen a ion
Au ho : Ped o Ramoneda F anco
Yea : 2020
This sc ip allows he use o p in o he console Ha monic Change De ec ion Func ion (HCDF)
ocus pe o mance on ecall o -sco e. I is assumed ha he i s command line a gumen is
he name ile o he audio ile loca ed in audio_ iles and he second one is i is ocus on
ecall o -sco e.
This ool accep s comma sepa a ed alue iles (.cs ) as well as excel
(.xls, .xlsx) iles.
This sc ip equi es ha `se up.py` eque imen s be ins alled wi hin he Py hon
en i onmen you a e unning his sc ip in. Mo e o e i is a need ins al amp plugins
NNLS and HPCP
This ile can also be impo ed as a module. All he unc ions han begins by ge a e blocks om HCDF
unc ion. The es a e auxilia .
A.1 Py hon lib a y 69
"""
impo os
om os impo pa h
impo sys
om TIVlib impo TIV
impo lib osa
impo numpy
numpy.se _p in op ions( h eshold=sys.maxsize)
om lib osa impo display
om lib osa. ea u e impo ch oma_cq , onne z, ch oma_cens, ch oma_s
om lib osa. il e s impo ge _window
om scipy.ndimage. il e s impo gaussian_ il e
impo ma plo lib.pyplo as pl
om as opy.con olu ion impo con ol e, Gaussian1DKe nel
om scipy.spa ial.dis ance impo cosine
om io x impo load_ eal_onse , load_bina y, ge _name_ha monic_change, sa e_bina y, ge _name_ch omag am,
ge _name_ onal_model, ge _name_gaussian_blu , ge _name_audio
impo amp
de ge _dis ance(cen oids, dis ):
"""
Re u ns he quan i y o cen oids pe second
Pa ame e s
----------
cen oids : lis o loa s
The ile loca ion o he sp eadshee
s : bool
A lag used o p in he columns o he console (de aul is False)
Re u ns
-------
loa
76 Anexo
e u n doce_bins_ uned_ch oma
de ch omag am(hpss, name_ ile, y, s , ch oma):
"""
w appe o ge _ch omag am o sa e all esul s o u u e same calcula ions
Pa ame e s
----------
hpss : bool
ue o alse depends on hpss block
name_ ile: s
name o he ile ha is being compu ed
y : numbe > 0 [scala ]
audio
s : numbe > 0 [scala ]
a ge sampling a e
ch oma: s
ch oma-sample a e- amesize-o e lap
Re u ns
-------
lis o ch omag ams
"""
name_ch omag am =ge _name_ch omag am(name_ ile, hpss, ch oma)
i pa h.exis s(name_ch omag am):
dic =load_bina y(name_ch omag am)
else:
# i mu ex_global.mu ex is no None:
# mu ex_global.mu ex.acqui e()
doce_bins_ uned_ch oma =ge _ch omag am(y, s , ch oma)
# i mu ex_global.mu ex is no None:
# mu ex_global.mu ex. elease()
dic ={'doce_bins_ uned_ch oma': doce_bins_ uned_ch oma}
# dic_sa e = {'doce_bins_ uned_ch oma': doce_bins_ uned_ch oma. olis ()}
A.1 Py hon lib a y 77
sa e_bina y(dic, name_ch omag am)
# sa e_json(dic_sa e, name_ch omag am + '.json')
e u n dic['doce_bins_ uned_ch oma']
de ge _ onal_cen oid_ ans o m(y, s , onal_model, doce_bins_ uned_ch oma):
"""
e u ns cen oids om onal model
Pa ame e s
----------
hpss : bool
ue o alse depends on hpss block
name_ ile: s
name o he ile ha is being compu ed
y : numbe > 0 [scala ]
audio
s : numbe > 0 [scala ]
a ge sampling a e
ch oma: s
ch oma-sample a e- amesize-o e lap
onal_model: s op ional
Tonal model block ype. "TIV2" o Tonal In e al space ocus on audio. "TIV2" o audio. "TIV2_Symb" o symbolic da a.
" onne z" o ha e cen oids ap oach. De aul TIV2
doce_bins_ uned_ch oma: lis
lis o ch oma ec o s
Re u ns
-------
lis o onal cen oids ec o s
"""
cen oid_ ec o =None
i onal_model == ' onne z':
cen oid_ ec o = onne z(y=y, s =s , ch oma=doce_bins_ uned_ch oma)
78 Anexo
eli onal_model == 'TIV2':
cen oid_ ec o = onal_in e al_space(doce_bins_ uned_ch oma)
eli onal_model == 'TIV2_symb':
cen oid_ ec o = onal_in e al_space(doce_bins_ uned_ch oma, symbolic=T ue)
e u n cen oid_ ec o
de onal_cen oid_ ans o m(hpss, ch oma, name_ ile, y, s , onal_model, doce_bins_ uned_ch oma):
"""
w appe o onal cen oid ans o m o sa e all esul s o u u e same calcula ions
Pa ame e s
----------
hpss : bool
ue o alse depends on hpss block
name_ ile: s
name o he ile ha is being compu ed
y : numbe > 0 [scala ]
audio
s : numbe > 0 [scala ]
a ge sampling a e
ch oma: s
ch oma-sample a e- amesize-o e lap
onal_model: s op ional
Tonal model block ype. "TIV2" o Tonal In e al space ocus on audio. "TIV2" o audio. "TIV2_Symb" o symbolic da a.
" onne z" o ha e cen oids ap oach. De aul TIV2
doce_bins_ uned_ch oma: lis
lis o ch oma ec o s
Re u ns
-------
lis o onal cen oids ec o s
"""
name_ onal_model =ge _name_ onal_model(name_ ile, hpss, ch oma, onal_model)
A.1 Py hon lib a y 79
i onal_model == 'wi hou _ c':
dic ={'cen oid_ ec o ': doce_bins_ uned_ch oma}
else:
i pa h.exis s(name_ onal_model):
dic =load_bina y(name_ onal_model)
else:
cen oid_ ec o =ge _ onal_cen oid_ ans o m(y, s , onal_model, doce_bins_ uned_ch oma)
dic ={'cen oid_ ec o ': cen oid_ ec o }
sa e_bina y(dic, name_ onal_model)
e u n dic['cen oid_ ec o ']
de ge _gaussian_blu (cen oid_ ec o , blu , sigma):
"""
Apply gaussian smoo hing o onal model cen oids
Pa ame e s
----------
cen oid_ ec o : lis
onal cen oids o he onal model
sigma: numbe (scala > 0) op ional
sigma o gaussian smoo hing alue. De aul 11
Re u ns
-------
lis
cen oids blu ed by gassuian smoo hing
"""
i blu == ' ull':
cen oid_ ec o =gaussian_ il e (cen oid_ ec o , sigma=sigma)
eli blu == '17-poin s':
gauss_ke nel =Gaussian1DKe nel(17)
i= 0
o cen oid in cen oid_ ec o :
cen oid =con ol e(cen oid, gauss_ke nel)
cen oid_ ec o [i] =cen oid
e u n numpy.a ay(cen oid_ ec o )
80 Anexo
de gaussian_blu (hpss, ch oma, onal_model, name_ ile, cen oid_ ec o , log_comp esion, blu , sigma):
"""
W appe o ge _gaussian_blu o sa e all esul s o u u e same calcula ions. I pa ame e iza ion
ha e been compu ed be o e ge _gaussian_blu is no compu ed.
Pa ame e s
----------
name_ ile: s
name o he ile ha is being compu ed
hpss: bool op ional
ue o alse depends is ha monic pe cussi e sou ce sepa a ion (hpss) block wan s o be compu ed. De aul False.
s : numbe > 0 [scala ]
a ge sampling a e
ch oma: s op ional
"ch oma-sample a e- amesize-o e lap"
ch oma can be "CQT","NNLS", "STFT", "CENS" o "HPCP"
sample a e as a numbe scala
ame size as a numbe scala
o e lap numbe ha a windows is di ided
onal_model: s op ional
Tonal model block ype. "TIV2" o Tonal In e al space ocus on audio. "TIV2" o audio. "TIV2_Symb" o symbolic da a.
" onne z" o ha e cen oids ap oach. De aul TIV2
cen oid_ ec o : lis
onal cen oids o he onal model
sigma: numbe (scala > 0) op ional
sigma o gaussian smoo hing alue. De aul 11
Re u ns
-------
lis
sample o audio
"""
gaussian_blu =ge _name_gaussian_blu (name_ ile, hpss, ch oma, onal_model, blu , sigma, log_comp esion)
A.1 Py hon lib a y 81
i pa h.exis s(gaussian_blu ):
dic =load_bina y(gaussian_blu )
else:
cen oid_ ec o =ge _gaussian_blu (cen oid_ ec o , blu , sigma)
dic ={'cen oid_ ec o ': cen oid_ ec o }
# dic_sa e = {'cen oid_ ec o ': cen oid_ ec o . olis ()}
sa e_bina y(dic, gaussian_blu )
# sa e_json(dic_sa e, gaussian_blu + '.json')
e u n dic['cen oid_ ec o ']
de ge _audio( ilename, hpss, s ):
"""
Ge audio as lis
Pa ame e s
----------
ilename: s
name o he ile ha is being compu ed wi ou o ma ex ension
hpss: bool op ional
ue o alse depends is ha monic pe cussi e sou ce sepa a ion (hpss) block wan s o be compu ed. De aul False.
s : numbe > 0 [scala ]
a ge sampling a e
Re u ns
-------
lis
sample o audio
"""
y, s =lib osa.load( ilename, s =s , mono=T ue)
i hpss:
y=lib osa.e ec s.ha monic(y)
e u n y, s
de audio( ilename, name_ ile, hpss, s ):
"""
W appe o ge audio o sa e all esul s o u u e same calcula ions. I pa ame e iza ion
82 Anexo
ha e been compu ed be o e ge audio is no compu ed.
Pa ame e s
----------
ilename: s
name o he ile ha is being compu ed wi ou o ma ex ension
name_ ile: s
name o he ile ha is being compu ed
hpss: bool op ional
ue o alse depends is ha monic pe cussi e sou ce sepa a ion (hpss) block wan s o be compu ed. De aul False.
s : numbe > 0 [scala ]
a ge sampling a e
Re u ns
-------
lis
sample o audio
"""
name_audio =ge _name_audio(name_ ile, hpss, s )
i pa h.exis s(name_audio):
dic =load_bina y(name_audio)
else:
y, s =ge _audio( ilename, hpss, s )
dic ={'y':y,'s ': s }
# dic_sa e = {'y': y. olis (), 's ': s }
sa e_bina y(dic, name_audio)
# sa e_json(dic_sa e, name_audio + '.json')
e u n dic['y'], dic['s ']
de ge _ha monic_change( ilename: s , name_ ile: s , hpss: bool =False, onal_model: s ='TIV2',
ch oma: s ='cq ',
blu : s =' ull', sigma: in = 11, log_comp esion: s ='none', dis : s ='euclidean'):
"""
Compu es Ha monic Change De ec ion Func ion
Pa ame e s
A.1 Py hon lib a y 83
----------
ilename: s
name o he ile ha is being compu ed wi ou o ma ex ension
name_ ile: s
name o he ile ha is being compu ed
hpss : bool op ional
ue o alse depends is ha monic pe cussi e sou ce sepa a ion (hpss) block wan s o be compu ed. De aul False.
onal_model: s op ional
Tonal model block ype. "TIV2" o Tonal In e al space ocus on audio. "TIV2" o audio. "TIV2_Symb" o symbolic da a.
" onne z" o ha e cen oids ap oach. De aul TIV2
ch oma: s op ional
"ch oma-sample a e- amesize-o e lap"
ch oma can be "CQT","NNLS", "STFT", "CENS" o "HPCP"
sample a e as a numbe scala
ame size as a numbe scala
o e lap numbe ha a windows is di ided
sigma: numbe (scala > 0) op ional
sigma o gaussian smoo hing alue. De aul 11
dis ance: s op ional
ype o dis ance measu e used. Types can be "euclidean" o euclidean dis ance and "cosine" o cosine dis ance. De aul "euclidean".
Re u ns
-------
lis
ha monic changes ( he peaks) on he song de ec ed
lis
HCDF unc ion alues
numbe
windows size
"""
# audio
y, s =audio( ilename, name_ ile, hpss, ge _pa ame e s_ch oma(ch oma)["s "])
84 Anexo
# ch oma
doce_bins_ uned_ch oma =ch omag am(hpss, name_ ile, y, s , ch oma)
# onal_model
cen oid_ ec o = onal_cen oid_ ans o m(hpss, ch oma, name_ ile, y, s , onal_model, doce_bins_ uned_ch oma)
# blu
cen oid_ ec o _blu ed =gaussian_blu (hpss, ch oma, onal_model, name_ ile, cen oid_ ec o , log_comp esion, blu ,
sigma)
# ha monic dis ance and calcula e peaks
ha monic_ unc ion =ge _dis ance(cen oid_ ec o _blu ed, dis )
windows_size =cen oids_pe _second(y, s , cen oid_ ec o _blu ed)
changes, cen oid_changes =ge _peaks_hcd (ha monic_ unc ion, cen oid_ ec o _blu ed, 0, windows_size,
cen oid_ ec o )
e u n changes, ha monic_ unc ion, windows_size, numpy.a ay(cen oid_changes)
de ha monic_change( ilename: s , name_ ile: s , hpss: bool =False, onal_model: s ='TIV2', ch oma: s ='cq ',
blu : s =' ull', sigma: in = 11, log_comp esion: s ='none', dis ance: s ='euclidean'):
"""
W appe o ha monic change de ec ion unc ion o sa e all esul s o u u e same calcula ions. I pa ame e iza ion
ha e been compu ed be o e HCDF is no compu ed.
Pa ame e s
----------
ilename: s
name o he ile ha is being compu ed wi ou o ma ex ension
name_ ile: s
name o he ile ha is being compu ed
hpss : bool op ional
ue o alse depends is ha monic pe cussi e sou ce sepa a ion (hpss) block wan s o be compu ed. De aul False.
onal_model: s op ional
Tonal model block ype. "TIV2" o Tonal In e al space ocus on audio. "TIV2" o audio. "TIV2_Symb" o symbolic da a.
" onne z" o ha e cen oids ap oach. De aul TIV2
A.1 Py hon lib a y 85
ch oma: s op ional
"ch oma-sample a e- amesize-o e lap"
ch oma can be "CQT","NNLS", "STFT", "CENS" o "HPCP"
sample a e as a numbe scala
ame size as a numbe scala
o e lap numbe ha a windows is di ided
sigma: numbe (scala > 0) op ional
sigma o gaussian smoo hing alue. De aul 11
dis ance: s op ional
ype o dis ance measu e used. Types can be "euclidean" o euclidean dis ance and "cosine" o cosine dis ance. De aul "euclidean".
Re u ns
-------
lis
ha monic changes ( he peaks) on he song de ec ed
lis
HCDF unc ion alues
numbe
windows size
"""
cen oid_changes =[]
check_pa ame e s(ch oma, blu , onal_model, log_comp esion, dis ance)
name_ha monic_change =ge _name_ha monic_change(name_ ile, hpss, onal_model, ch oma, blu , sigma, log_comp esion,
dis ance)
i pa h.exis s(name_ha monic_change):
dic =load_bina y(name_ha monic_change)
else:
changes, ha monic_ unc ion, windows_size, cen oid_changes =ge _ha monic_change( ilename, name_ ile, hpss,
onal_model, ch oma,
blu , sigma, log_comp esion,
dis ance)
dic ={'changes': changes, 'ha monic_ unc ion': ha monic_ unc ion, 'windows_size': windows_size}
sa e_bina y(dic, name_ha monic_change)
e u n dic['changes'], dic['ha monic_ unc ion'], dic['windows_size']
92 Anexo
i (Ma h.abs( equencySignal_im) <ze o) {
equencySignal_im = 0;
}
// A e age con ibu ion a his equency.
// complex( ecuencySignal) / N
equencySignal_ e =( equencySignal_ e *N) /(N*N);
equencySignal_im =( equencySignal_im *N) /(N*N);
// Add cu en equency signal o he lis o compound signals.
signals.push( equencySignal_ e);
signals.push( equencySignal_im);
}
e u n signals;
}
unc ion di ision( ec o , ene gy){
o ( a i= ec o .leng h - 1;i>= 0; i--) {
ec o [i] = ec o [i]/ene gy;
}
e u n ec o ;
}
unc ion mul iply( ec o A, ec o B){
a ans =new A ay(12);
o ( a i= ec o A.leng h - 1;i>= 0; i--) {
ans[i] = ec o A[i] * ec o B[i]
}
e u n ans;
}
unc ion TIV(pcp, weigh s){
// Tonal In e al Vec o s
le =DFT(pcp);
le ene gy = [0];
A.2 Ja asc ip lib a y 93
le ec o = .slice(2,14);
i (weigh s === "symbolic"){
le weigh s_symbolic =[2,2,11,11,17,17,16,16,19,19,7,7]
ec o =mul iply(di ision( ec o , ene gy), weigh s_symbolic);
}
else i (weigh s === "audio"){
le weigh s_audio =[3,3,8,8,11.5,11.5,15,15,14.5,14.5,7.5,7.5];
ec o =mul iply(di ision( ec o , ene gy), weigh s_audio);
}
else i (weigh s === "ha e"){
le wei h s_ha e =[0,0,0,0,1,1,0.5,0.5,1,1,0,0];
ec o =mul iply(di ision( ec o , ene gy), wei h s_ha e);
}
e u n ec o ;
}
/*
* e u ns onal in e al space om a ec o o ch omag ams
*
*Pa ame e s
*----------
*ch oma : lis
*lis o ch omag ams
*
*weigh s: s
*"audio", "symbolic" o "ha e"
*
*Re u ns
*-------
*lis o onal in e al space ec o s
*/
unc ion onal_in e al_space(ch oma, weigh s="audio"){
// Tonal In e al Space
le cen oid_ ec o =[];
o ( a i= 0;i<ch oma.leng h; i++){
le each_ch oma =ch oma[i];
le cen oid =[0,0,0,0,0,0,0,0,0,0,0,0];
i (!e e y hing_is_ze o(each_ch oma)){
cen oid =TIV(each_ch oma, weigh s)
94 Anexo
}
cen oid_ ec o .push(cen oid);
}
e u n cen oid_ ec o ;
}
unc ion a g ( ) {
e u n . educe((a,b) => a+b, 0)/ .leng h;
}
/*
*Apply gaussian smoo hing o onal model cen oids
*Pa ame e s
*----------
* ec o : lis
* onal cen oids o he onal model
*sigma: numbe (scala > 0) op ional
*sigma o gaussian smoo hing alue.
*Re u ns
*-------
*lis
*cen oids blu ed by gassuian smoo hing
*/
unc ion gaussian_smoo hing_ ec o ( ec o , sigma) {
a _a g =a g( ec o )*sigma;
a e =A ay( ec o .leng h);
o ( a i= 0;i< ec o .leng h; i++) {
( unc ion () {
a p e =i>0? e [i-1]: ec o [i];
a nex =i< ec o .leng h ? ec o [i] : ec o [i-1];
e [i] =a g([ _a g, a g([p e , ec o [i], nex ])]);
})();
}
e u n e ;
}
A.2 Ja asc ip lib a y 95
unc ion gaussian_smoo hing( is, sigma){
a ans =[];
o ( a i= is.leng h - 1;i>= 0; i--) {
ans.push(gaussian_smoo hing_ ec o ( is[i], sigma));
}
e u n ans;
}
/*
*Re u ns he quan i y o cen oids pe second
*Pa ame e s
*----------
*cen oids : lis o loa s
*The ile loca ion o he sp eadshee
*Re u ns
*-------
* loa
*cen oids pe second
*/
unc ion dis ance(cen oids){
a ans =[0];
o ( a i= 1;i<cen oids.leng h - 1; i++) {
a sum = 0;
o ( a j= 1;j<cen oids[i].leng h - 1; j++) {
sum += Ma h.pow((cen oids[i][j + 1]-cen oids[i][j - 1]), 2)
}
sum =Ma h.sq (sum)
ans.push(sum);
}
e u n ans;
}
/*
*Re u ns he quan i y o cen oids pe second
*
*Pa ame e s
*----------
96 Anexo
*y : lis o loa s
*The ile loca ion o he sp eadshee
*s : bool
*A lag used o p in he columns o he console (de aul is False)
*
*Re u ns
*-------
* loa
*cen oids pe second
*/
unc ion cen oids_pe _second(y, s , cen oids){
e u n s *cen oids.leng h /y.leng h;
}
unc ion peaks(hcd _ unc ion, a e_cen oids_second){
le changes =[0];
o ( a i= 0;i<hcd _ unc ion.leng h; i++) {
i (hcd _ unc ion[i - 1]<hcd _ unc ion[i] && hcd _ unc ion[i + 1]<hcd _ unc ion[i]){
changes.push(i / a e_cen oids_second)
}
}
e u n changes;
}
/*
*Compu es Ha monic Change De ec ion Func ion
*Pa ame e s
*----------
*id_audio: s
*id o HTML elemen <audio>
*Re u ns
*-------
*lis
*ha monic changes ( he peaks) on he song de ec ed
*/
expo async unc ion HCDF(id_audio) {
A.2 Ja asc ip lib a y 97
le audioURL =documen .ge Elemen ById(id_audio).cu en S c;
console.log(audioURL);
// load audio ile om an u l
le audioDa a =awai essen ia.ge AudioChannelDa aF omURL(audioURL, audioC x, 0);
i (isCompu ed) { plo Ch oma.des oy(); };
cons ameSize = 2048;
cons hopSize = 512;
cons sampleRa e = 8000;
console.log("audio an es downsampling", audioDa a);
audioDa a =downsample(audioDa a, 44100, sampleRa e);
console.log("audio despues downsampling", audioDa a);
le ames =essen ia.F ameGene a o (audioDa a,
ameSize,
hopSize)
le ch oma =ch omaNNLS( ames, ameSize, hopSize, sampleRa e);
console.log("ch oma", ch oma);
le ch oma =[[0.0,1.0,0.0,1.0,0.0,1.0,0.0,0.0,0.0,1.0,0.0,0.0]]
le onal_cen oids = onal_in e al_space(ch oma, "symbolic");
console.log(" onal cen oids", onal_cen oids);
le smoo hed_cen oids =gaussian_smoo hing( onal_cen oids, 5);
console.log("gaussian smoo hing", onal_cen oids);
le ha monic_ unc ion =dis ance(smoo hed_cen oids);
console.log("dis ance", dis ance);
le cps =cen oids_pe _second(audioDa a, sampleRa e, smoo hed_cen oids);
le ha monic_changes =peaks(ha monic_ unc ion, cps);
console.log("ha monic_changes", ha monic_changes);
e u n awai ha monic_changes;
98 Anexo
}
/*
*Func ion o loading essen ia wasm module
*
*/
expo async unc ion loadEssen ia(){
// Now le 's load he essen ia wasm back-end, i so c ea e UI elemen s o compu ing ea u es
Essen iaModule(). hen(async unc ion(WasmModule) {
essen ia =new Essen ia(WasmModule);
});
};
expo de aul {HCDF, loadEssen ia}
Re e ences
[1] Ma k A Ba sch and G ego y H Wake ield. Audio humbnailing o popula music using
ch oma-based ep esen a ions. IEEE T ansac ions on mul imedia, 7(1):96–104, 2005.
[2] Juan Pablo Bello and Je emy Pickens. A obus mid-le el ep esen a ion o ha monic con-
en in music signals. In ISMIR, olume 5, pages 304–311. Ci esee , 2005.
[3] B uce Benwa d. Music in Theo y and P ac ice Volume 1, olume 1. McG aw-Hill Highe
Educa ion, 2014.
[4] Gilbe o Be na des, Diogo Cocha o, Ma celo Cae ano, Ca los Guedes, and Ma hew EP
Da ies. A mul i-le el onal in e al space o modelling pi ch ela edness and musical con-
sonance. Jou nal o New Music Resea ch, 45(4):281–294, 2016.
[5] Gilbe o Be na des, Diogo Cocha o, Ca los Guedes, and Ma hew EP Da ies. Ha mony
gene a ion d i en by a pe cep ually mo i a ed onal in e al space. Compu e s in En e ain-
men (CIE), 14(2):1–21, 2016.
[6] Gilbe o Be na des, Ma hew EP Da ies, and Ca los Guedes. Au oma ic musical key es ima-
ion wi h adap i e mode bias. In 2017 IEEE In e na ional Con e ence on Acous ics, Speech
and Signal P ocessing (ICASSP), pages 316–320. IEEE, 2017.
[7] Gilbe o Be na des, Ma hew EP Da ies, and Ca los Guedes. A hie a chical ha monic mix-
ing me hod. In In e na ional Symposium on Compu e Music Mul idisciplina y Resea ch,
pages 151–170. Sp inge , 2017.
[8] Nicola Be na dini and Gio anni De Poli. The sound and music compu ing ield: p esen and
u u e. Jou nal o New Music Resea ch, 36(3):143–148, 2007.
[9] Sebas ian Böck, Filip Ko zeniowski, Jan Schlü e , Flo ian K ebs, and Ge ha d Widme . Mad-
mom: A new py hon audio and music signal p ocessing lib a y. In P oceedings o he 24 h
ACM in e na ional con e ence on Mul imedia, pages 1174–1178, 2016.
[10] Dmi y Bogdano , Joan Se a, Nicolas Wack, and Pe ec o He e a. F om low-le el o high-
le el: Compa a i e s udy o music simila i y measu es. In 2009 11 h IEEE In e na ional
Symposium on Mul imedia, pages 453–458. IEEE, 2009.
[11] Dmi y Bogdano , Nicolas Wack, Emilia Gómez Gu ié ez, Sankalp Gula i, He e a Boye ,
Osca Mayo , Ge a d Roma T epa , Jus in Salamon, José Rica do Zapa a González, Xa ie
Se a, e al. Essen ia: An audio analysis lib a y o music in o ma ion e ie al. In B i o A,
Gouyon F, Dixon S, edi o s. 14 h Con e ence o he In e na ional Socie y o Music In o ma-
ion Re ie al (ISMIR); 2013 No 4-8; Cu i iba, B azil.[place unknown]: ISMIR; 2013. p.
493-8. In e na ional Socie y o Music In o ma ion Re ie al (ISMIR), 2013.
99
100 REFERENCES
[12] Michael Bos ock, Vadim Ogie e sky, and Je ey Hee . D3da a-d i en documen s. IEEE
ansac ions on isualiza ion and compu e g aphics, 17(12):2301–2309, 2011.
[13] Nicolas Boulange -Lewandowski, Yoshua Bengio, and Pascal Vincen . Audio cho d ecog-
ni ion wi h ecu en neu al ne wo ks. In ISMIR, pages 335–340. Ci esee , 2013.
[14] And ew B ock, Je Donahue, and Ka en Simonyan. La ge scale gan aining o high ideli y
na u al image syn hesis. a Xi p ep in a Xi :1809.11096, 2018.
[15] Judi h C B own and Mille S Pucke e. An e icien algo i hm o he calcula ion o a cons an
q ans o m. The Jou nal o he Acous ical Socie y o Ame ica, 92(5):2698–2701, 1992.
[16] Emilios Cambou opoulos. F om midi o adi ional musical no a ion. In P oceedings o he
AAAI Wo kshop on A i icial In elligence and Music: Towa ds Fo mal Models o Composi-
ion, Pe o mance and Analysis, olume 30, 2000.
[17] Soubhik Chak abo y, Gue ino Mazzola, Swa ima Tewa i, and Moujhu i Pa a. Compu a-
ional musicology in Hindus ani music. Sp inge , 2014.
[18] Ruo eng Chen, Weibin Shen, Ajay S ini asamu hy, and Pa ag Cho dia. Cho d ecogni ion
using du a ion-explici hidden ma ko models. In ISMIR, pages 445–450. Ci esee , 2012.
[19] Elaine Chew. The spi al a ay: An algo i hm o de e mining key bounda ies. In In e na-
ional Con e ence on Music and A i icial In elligence, pages 18–31. Sp inge , 2002.
[20] E ic Cla ke and Nicholas Cook. Empi ical musicology: Aims, me hods, p ospec s. Ox o d
Uni e si y P ess, 2004.
[21] Alessio Degani, Ma co Dalai, Ricca do Leona di, and Pie angelo Miglio a i. Ha monic
change de ec ion o musical cho ds segmen a ion. In 2015 IEEE In e na ional Con e ence
on Mul imedia and Expo (ICME), pages 1–6. IEEE, 2015.
[22] Dan Ellis. Ch oma Fea u e Analysis and Syn hesis. A ailable a
h ps://lab osa.ee.columbia.edu/ma lab/ch oma-ansyn/, accessed Ap il 18, 2020.
[23] Leonha d Eule . Ten amen no ae heo iae musicae ex ce issismis ha moniae p incipiis dilu-
cide exposi ae. Sain Pe e sbu g Academy, 1739.
[24] De y Fi zge ald. Ha monic/pe cussi e sepa a ion using median il e ing. In P oc. o DAFX,
olume 10, 2010.
[25] Takuya Fujishima. Real- ime cho d ecogni ion o musical sound: A sys em using common
lisp music. P oc. ICMC, Oc . 1999, pages 464–467, 1999.
[26] Ba ba a R Gaizauskas. The ha mony o he sphe es. Jou nal o he Royal As onomical
Socie y o Canada, 68:146, 1974.
[27] Emilia Gómez. Tonal desc ip ion o music audio signals. Depa men o In o ma ion and
Communica ion Technologies, 2006.
[28] S ephen W Hainswo h, Malcolm D Macleod, e al. Onse de ec ion in musical audio signals.
In ICMC, 2003.
[29] Ch is ophe Ha e. Towa ds au oma ic ex ac ion o ha mony in o ma ion om music sig-
nals. PhD hesis, 2010.
REFERENCES 101
[30] Ch is ophe Ha e and Ma k Sandle . Au oma ic cho d iden i ca ion using a quan ised ch o-
mag am. In Audio Enginee ing Socie y Con en ion 118. Audio Enginee ing Socie y, 2005.
[31] Ch is ophe Ha e, Ma k Sandle , and Ma in Gasse . De ec ing ha monic change in musical
audio. In P oceedings o he 1s ACM wo kshop on Audio and music compu ing mul imedia,
pages 21–26, 2006.
[32] Douglas M Hawkins. The p oblem o o e i ing. Jou nal o chemical in o ma ion and
compu e sciences, 44(1):1–12, 2004.
[33] E ic J Humph ey and Juan P Bello. Re hinking au oma ic cho d ecogni ion wi h con olu-
ional neu al ne wo ks. In 2012 11 h In e na ional Con e ence on Machine Lea ning and
Applica ions, olume 2, pages 357–362. IEEE, 2012.
[34] E ic J Humph ey, Taemin Cho, and Juan P Bello. Lea ning a obus onne z-space ans-
o m o au oma ic cho d ecogni ion. In 2012 IEEE In e na ional Con e ence on Acous ics,
Speech and Signal P ocessing (ICASSP), pages 453–456. IEEE, 2012.
[35] Elyo Kodi o , Sejin Han, Guee-Sang Lee, and YoungChul Kim. Music wi h ha mony:
Cho d sepa a ion and ecogni ion in p in ed music sco e images. In P oceedings o he
8 h In e na ional Con e ence on Ubiqui ous In o ma ion Managemen and Communica ion,
ICUIMC ’14, New Yo k, NY, USA, 2014. Associa ion o Compu ing Machine y.
[36] Hend ik Vincen Koops, W Bas de Haas, John Ashley Bu goyne, Je oen B ansen, and Anja
Volk. Ha monic subjec i i y in popula music, 2017.
[37] Filip Ko zeniowski and Ge ha d Widme . Fea u e lea ning o cho d ecogni ion: The deep
ch oma ex ac o . a Xi p ep in a Xi :1612.05065, 2016.
[38] Ma hieu Lag ange, G aham Pe ci al, and Geo ge Tzane akis. Adap i e ha moniza ion and
pi ch co ec ion o polyphonic audio using spec al clus e ing. In P oceedings o DAFx,
pages 1–4, 2007.
[39] Bangalo e S Manjuna h, Philippe Salembie , and Thomas Siko a. In oduc ion o MPEG-7:
mul imedia con en desc ip ion in e ace. John Wiley & Sons, 2002.
[40] Ma hias Mauch, Ch is Cannam, Ma hew Da ies, Simon Dixon, Ch is ophe Ha e, Se ki
Kolozali, Dan Tidha , and Ma k Sandle . Om as2 me ada a p ojec 2009. In P oc. o 10 h
In e na ional Con e ence on Music In o ma ion Re ie al, page 1, 2009.
[41] Ma hias Mauch and Simon Dixon. Simul aneous es ima ion o cho ds and musical con ex
om audio. IEEE T ansac ions on Audio, Speech, and Language P ocessing, 18(6):1280–
1289, 2009.
[42] Ma hias Mauch and Simon Dixon. App oxima e no e ansc ip ion o he imp o ed iden-
i ica ion o di icul cho ds. In P oceedings o he 11 h In e na ional Socie y o Music
In o ma ion Re ie al Con e ence (ISMIR 2010), 2010.
[43] HJJ MAXWELL. An a i icial in elligence app oach o compu e -implemen ed analysis o
ha mony in onal music. 1986.
[44] Gue ino Mazzola. The opos o music: geome ic logic o concep s, heo y, and pe o mance.
Bi khäuse , 2012.