2021
77
Ignacio Viñals Bailo
Ad ances in Subspace-
based Solu ions o
Dia iza ion in he
B oadcas Domain
Di ec o /es
O ega Giménez, Al onso
© Uni e sidad de Za agoza
Se icio de Publicaciones
ISSN 2254-7606
Ignacio Viñals Bailo
ADVANCES IN SUBSPACE-BASED SOLUTIONS
FOR DIARIZATION IN THE BROADCAST DOMAIN
Di ec o /es
O ega Giménez, Al onso
Tesis Doc o al
Au o
2020
UNIVERSIDAD DE ZARAGOZA
Escuela de Doc o ado
P og ama de Doc o ado en Tecnologías de la In o mación y
Comunicaciones en Redes Mó iles
Reposi o io de la Uni e sidad de Za agoza – Zaguan h p://zaguan.uniza .es
UNIVERSIDAD DE ZARAGOZA
TESIS DOCTORAL - INGENIERÍA DE TELECOMUNICACIÓN
Ad ances in Subspace-based Solu ions o
Dia iza ion in he B oadcas Domain
Au ho :
Ignacio Viñals Bailo
Supe iso :
Al onso O ega Giménez
DEPARTAMENTO DE INGENIERÍA ELECTRÓNICA Y COMUNICACIONES
ESCUELA DE INGENIERÍA Y ARQUITECTURA
Ap il, 2020
A mis pad es
The human oice is
he mos pe ec ins umen o all.
A o Pä
The mos impo an ques ions o li e a e,
o he mos pa ,
eally only p oblems o p obabili y
Pie e-Simon Laplace
Resea ch is c ea ing new knowledge.
Neil A ms ong
The human b ain is an inc edible
pa e n-ma ching machine.
Je Bezos
4
Acknowledgemen s
I has been i e long yea s since I made he decision o s a a PhD p og amme. Along all his
ime I ha e been o una e enough o mee , collabo a e and be helped as well as suppo ed by
many people, wi hou whom his hesis would no be a ailable oday. These lines a e dedica ed
o all o hem.
Fi s and o emos , I wan o dedica e some lines o my pa en s. I wan o hank hem o
being on my side om he e y beginning. They always o e ed me hei suppo since I decided
o become a esea che . Fo he las i e yea s hey ha e chee ed me up du ing bad imes and
kep my ee on he g ound du ing hose limi ed success ul occasions. This wo k could no be
possible wi hou hei con ibu ion.
Besides, I also mus hank Al onso O ega o gi ing me he chance o g ow up p o es-
sionally and pe sonally. He ga e me he chance o disco e he esea che ca ee when I was
unde g adua e, and o e ed me he oppo uni y o keep on de eloping mysel wi h ViVoLAB
g oup. Fo he las i e yea s he has become a iend apa om a supe iso , who guided me
along his di icul lea ning p ocess. By means o ou mee ings he helped me o disco e some
o he bes ideas while o he imes he simply made me awa e ha some imes I could no see he
o es o he ees.
Apa om Al onso O ega, ViVoLAB g oup is also ull o wonde ul people who dese e
hei own men ion. Some o my kindes memo ies along my PhD yea s include Edua do Lleida.
He always ied o build a g oup based on iendship ela ionships a he han simply p o es-
sional ones. Besides, i is p aisewo hy how his ea men also included ViVoLAB alumni all
o e he wo ld. Ano he impo an collabo a o in his hesis is An onio Miguel. His aluable
expe ise in subspace models and neu al ne wo ks we e capi al o he wo k done in his hesis.
Besides, he heo e ical discussions we had abou hese opics we e also e y en iching, opening
my eyes abou unseen lines o esea ch o explo e.
I also wan o acknowledge Johns Hopkins uni e si y s a , specially Najim Dehak and his
5
Con en s
1 In oduc ion 1
1.1 Mo i a ion o he wo k . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Objec i es and Me hodology . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.3 Thesis o ganiza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
I Dia iza ion Basic Knowledge 7
2 Dia iza ion S a e o he A 9
2.1 In oduc ion .................................... 9
2.2 Main dia iza ion s a egies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.2.1 Bo om-Up dia iza ion sys ems . . . . . . . . . . . . . . . . . . . . . 12
2.3 Acous ic ea u es o dia iza ion . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.4 Audio segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.4.1 Me ic-based segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . 17
2.4.1.1 Bayesian In o ma ion C i e ion (BIC) . . . . . . . . . . . . . 18
2.4.1.2 Kullback-Leible Di e gence (KL) . . . . . . . . . . . . . . 19
2.4.1.3 Deep Neu al Ne wo ks (DNNs) . . . . . . . . . . . . . . . . 20
2.4.2 Model-based segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . 20
2.5 Speake cha ac e iza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.5.1 Ea ly days . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.5.2 Model-based ep esen a ions . . . . . . . . . . . . . . . . . . . . . . . 22
2.5.2.1 Gaussian Mix u e Models (GMMs) . . . . . . . . . . . . . . 22
2.5.2.2 Suppo Vec o Machines (SVM) . . . . . . . . . . . . . . . 23
2.5.2.3 Join Fac o Analysis (JFA) . . . . . . . . . . . . . . . . . . 24
iii
CONTENTS
2.5.3 Embedded ep esen a ions . . . . . . . . . . . . . . . . . . . . . . . . 25
2.5.3.1 I- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.5.3.2 Hyb id i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . 26
2.5.3.3 DNN embeddings . . . . . . . . . . . . . . . . . . . . . . . 27
2.5.4 P obabilis ic Linea Disc iminan Analysis (PLDA) . . . . . . . . . . . 27
2.6 Clus e ing ..................................... 28
2.6.1 Hie a chical clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . 32
2.6.2 S a is ical app oaches . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.6.3 O he al e na i es . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
2.7 Pe o mance me ics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3 Analysis o Dia iza ion in B oadcas Da a 39
3.1 The dia iza ion e e ence sys em . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.2 Analysis o b oadcas da a . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2.1 Mul i-Gen e B oadcas Challenge 2015 (MGB 2015) . . . . . . . . . . 42
3.2.2 Albayzín 2018 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.2.3 Acous ic a iabili y . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.2.4 Va iabili y in he speake dis ibu ion . . . . . . . . . . . . . . . . . . 46
3.3 E alua ion o pe o mance o he dia iza ion e e ence sys em . . . . . . . . . 48
3.3.1 E alua ion o pe o mance in MGB 2015 . . . . . . . . . . . . . . . . 48
3.3.2 E alua ion o pe o mance in Albayzín 2018 . . . . . . . . . . . . . . 50
3.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
3.4.1 The clus e ing app oxima ion . . . . . . . . . . . . . . . . . . . . . . 52
3.4.2 The quali y o he embeddings . . . . . . . . . . . . . . . . . . . . . . 52
3.4.3 The domain misma ch p oblem . . . . . . . . . . . . . . . . . . . . . 53
II The Clus e ing P oblem 55
4 Clus e ing by means o Fully Bayesian PLDA 57
4.1 The Fully Bayesian PLDA clus e ing solu ion . . . . . . . . . . . . . . . . . . 57
4.1.1 The Fully Bayesian PLDA (FBPLDA) model . . . . . . . . . . . . . . 57
4.1.2 The clus e ing p ocedu e . . . . . . . . . . . . . . . . . . . . . . . . . 60
4.1.3 Dia iza ion using he FBPLDA model . . . . . . . . . . . . . . . . . . 62
4.2 Analysis o FBPLDA pe o mance . . . . . . . . . . . . . . . . . . . . . . . . 65
4.2.1 Ini ializa ion impac . . . . . . . . . . . . . . . . . . . . . . . . . . . 65
i
CONTENTS
4.2.2 In e ence o he numbe o speake s . . . . . . . . . . . . . . . . . . . 66
4.2.3 Numbe o speake s s DER . . . . . . . . . . . . . . . . . . . . . . . 69
4.2.4 Numbe o speake s s ELBO . . . . . . . . . . . . . . . . . . . . . . 71
4.3 Al e na i e ini ializa ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
4.3.1 Compu a ionally e icien ini ializa ion . . . . . . . . . . . . . . . . . 73
4.3.2 ELBO-based ini ializa ion choice c i e ion . . . . . . . . . . . . . . . 74
4.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
5 Unce ain y P opaga ion o Dia iza ion 79
5.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
5.2 PLDA wi h Unce ain y P opaga ion (PLDAUP) . . . . . . . . . . . . . . . . . 80
5.2.1 PLDAUP in speake ecogni ion . . . . . . . . . . . . . . . . . . . . . 82
5.2.2 PLDAUP in speake clus e ing . . . . . . . . . . . . . . . . . . . . . . 86
5.3 FBPLDA wi h Unce ain y P opaga ion (FBPLDAUP) . . . . . . . . . . . . . 89
5.4 Dia iza ion o b oadcas da a wi h FBPLDAUP . . . . . . . . . . . . . . . . . 91
5.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
6 T ee-Based Clus e ing App oaches 97
6.1 T ee-based poin o iew o clus e ing . . . . . . . . . . . . . . . . . . . . . . 98
6.2 PLDA ee-based clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
6.2.1 PLDA-based model . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
6.2.2 M-algo i hm op imiza ion . . . . . . . . . . . . . . . . . . . . . . . . 103
6.3 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
6.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112
III The Speake Rep esen a ion P oblem 113
7 S udy o embeddings o sho u e ances 115
7.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
7.2 Sho u e ances as occluded u e ances . . . . . . . . . . . . . . . . . . . . . 116
7.3 Fo mula ion o he embedding ex ac ion wi h sho u e ances . . . . . . . . . 118
7.3.1 Gene al case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
7.3.2 i- ec o embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
7.3.3 Sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
7.4 E ec s o he sho u e ances in i- ec o s . . . . . . . . . . . . . . . . . . . . 122
CONTENTS
7.5 Expe imen s & Resul s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5.1 Expe imen al se up . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5.2 Baseline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
7.5.3 Reduc ion o he misma ch in α: Phone ic balance . . . . . . . . . . . 128
7.5.4 En ollmen - es dis ance s log-likelihood a io (ll ) . . . . . . . . . . . 131
7.5.5 En ollmen - es dis ance s pe o mance (EER and minDCF) . . . . . . 133
7.5.6 Long-sho s Equalized Sho -Sho . . . . . . . . . . . . . . . . . . . 134
7.6 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
8 DNNs embeddings o Dia iza ion 137
8.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
8.2 Hyb id i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
8.2.1 Bo leneck Fea u es (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 139
8.2.2 Phone ic i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
8.3 X- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
8.4 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144
8.4.1 Bo leneck Fea u es (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 145
8.4.2 Phone ic i- ec o s & x- ec o s . . . . . . . . . . . . . . . . . . . . . . 147
8.4.2.1 Speake ecogni ion . . . . . . . . . . . . . . . . . . . . . . 148
8.4.2.2 B oadcas dia iza ion . . . . . . . . . . . . . . . . . . . . . 150
8.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
IV The Model Adap a ion P oblem 155
9 Da a-E icien Domain Adap a ion o PLDA Models 157
9.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
9.2 Me hods o domain misma ch educ ion . . . . . . . . . . . . . . . . . . . . . 158
9.3 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
9.3.1 Independen unsupe ised adap a ion . . . . . . . . . . . . . . . . . . 162
9.3.2 Longi udinal unsupe ised adap a ion . . . . . . . . . . . . . . . . . . 163
9.3.3 Use o in-domain labeled da a and semi-supe ised adap a ion . . . . . 164
9.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166
i
CONTENTS
V Conclusions & Fu u e Wo k 167
10 Conclusions & Fu u e wo k 169
10.1 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
10.1.1 The clus e ing ask . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
10.1.2 The speake cha ac e iza ion s age . . . . . . . . . . . . . . . . . . . . 170
10.1.3 Unsupe ised domain adap a ion esea ch . . . . . . . . . . . . . . . . 171
10.2 Scien i ic Con ibu ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.1 Book chap e s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.2 Pape s published in jou nals included in he Jou nal Ci a ion Repo s
(JCR) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.3 Con e ence p oceedings . . . . . . . . . . . . . . . . . . . . . . . . . 173
10.3 Fu u e Wo k . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
VI Appendix 175
A Fully Bayesian PLDA wi h Unce ain y P opaga ion I
A.1 De ini ions ..................................... I
A.2 Da a ........................................ II
A.3 Da a condi ional likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . III
A.3.1 P(Φi|yi,Xi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . III
A.3.2 P(Xi|yi,Θi,Φi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . IV
A.3.3 P(yi|Φi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . . . IV
A.4 Va ia ional app oach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V
A.4.1 Join p obabili y . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V
A.4.2 Va ia ional Bayes app oxima ion . . . . . . . . . . . . . . . . . . . . . VI
A.4.3 Op imal de ini ion o q∗(Y,X). . . . . . . . . . . . . . . . . . . . . VI
A.4.4 Op imal de ini ion o q∗(Θ) . . . . . . . . . . . . . . . . . . . . . . . VII
A.4.5 Op imal de ini ion o q∗(πθ). . . . . . . . . . . . . . . . . . . . . . . VIII
A.4.6 op imal de ini ion o q∗˜
V. . . . . . . . . . . . . . . . . . . . . . . VIII
A.4.7 Op imal de ini ion o q∗(W). . . . . . . . . . . . . . . . . . . . . . . X
A.4.8 Op imal de ini ion o q∗(ε). . . . . . . . . . . . . . . . . . . . . . . XI
A.4.9 Necessa y Expec a ions . . . . . . . . . . . . . . . . . . . . . . . . . XI
A.4.10 Va ia ional Lowe Bound . . . . . . . . . . . . . . . . . . . . . . . . . XIV
A.5 Hype pa ame e op imiza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . XV
ii
CONTENTS
iii
Lis o Figu es
1.1 Example o dia iza ion esul s . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.2 Concep ual map o he s udied opics in his Thesis . . . . . . . . . . . . . . . 5
2.1 Schema ic o Bo om-Up and Top-Down dia iza ion . . . . . . . . . . . . . . . 11
2.2 Gene al schema ic o a dia iza ion sys em . . . . . . . . . . . . . . . . . . . . 12
2.3 Schema ic o he MFCC ex ac ion pipeline . . . . . . . . . . . . . . . . . . . 14
2.4 Scheme o a sliding window me ic based segmen a ion . . . . . . . . . . . . 18
2.5 Schema ic o an Agglome a i e Hie a chical Clus e ing (AHC) pe o mance . 32
3.1 Schema ic o ou baseline dia iza ion sys em . . . . . . . . . . . . . . . . . . . 40
3.2 Sec ion a iabili y example. Fo 100 i s embeddings om a Sp ingwa ch
episode wi h SPLDA pai wise LLR simila i y me ic and G ound u h ela-
ionship. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
3.3 Va iabili y in he speake dis ibu ion o MGB 2015, numbe o speake s pe
show and he p opo ion o speech o he mos ac i e speake pe show. . . . . 47
3.4 Va iabili y in he speake dis ibu ion o Albayzín 2018, numbe o speake s
pe show. and p opo ion o speech o he mos ac i e speake pe show. . . . . 48
3.5 Dis ibu ion o speech pe speake o wo episodes: An episode wi h a domi-
nan speake and an episode wi h a mo e e en speech dis ibu ion . . . . . . . . 49
4.1 Bayesian ne wo k o he Fully Bayesian PLDA . . . . . . . . . . . . . . . . . 58
4.2 Clus e ing schema ic based on label ini ializa ion and FBPLDA esegmen a ion 61
4.3 Schema ic o he dia iza ion sys em based on he FBPLDA esegmen a ion . . 62
4.4 Analysis o ∆I=IORACLE −IHY P o shows in MGB 2015 wi h AHC and
FBPLDA esegmen a ion dia iza ion sys ems. . . . . . . . . . . . . . . . . . . 64
4.5 5-le el dend og am example. . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
ix
LIST OF FIGURES
4.6 Inpu /ou pu ela ionship o he numbe o speake s wi h FBPLDA esegmen-
a ion. ....................................... 67
4.7 Los speake s acco ding o he ela i e numbe o speake s ∆Iin he ini ial
pa i ion Θ0.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
4.8 DER (%) esul s o a) AHC and b) FBPLDA in e ms o he ela i e numbe o
speake s ∆I.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
4.9 Dis ibu ion o he ini ializa ion wi h bes DER in e ms o he ela i e numbe
o speake s. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
4.10 Dis ibu ion o he ini ializa ion wi h bounded DER, a) 1% and b) 3%, in e ms
o he ela i e numbe o speake s. . . . . . . . . . . . . . . . . . . . . . . . . 72
4.11 Dis ibu ion o he pa i ion wi h bes ELBO in e ms o ela i e speake s. . . . 73
4.12 Schema ic o dia iza ion based on he simul aneous e alua ion o K di e en
ini ializa ions. The inal pa i ion is selec ed by means o PELBO. . . . . . . . 75
5.1 Bayesian ne wo k o PLDA wi h Unce ain y P opaga ion (PLDAUP) . . . . . 81
5.2 DET cu es wi h SPLDA o SRE10 co ex -co eex de 5 emale wi h in ol ed
sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
5.3 DET cu es wi h PLDAUP o SRE10 co ex -co eex de 5 emale wi h in ol ed
sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
5.4 Impu i y esul s o SPLDA and PLDAUP in SRE10 co eex -co eex de 5 e-
male chopped . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
5.5 Impu i y esul s o a) SPLDA and b) PLDAUP in SRE10 co eex -co eex de 5
emale chopped aining wi h sho u e ances . . . . . . . . . . . . . . . . . . 88
5.6 Bayesian ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opaga ion . 90
5.7 His og am o DER a ia ions be ween SPLDA and PLDAUP in MGB 2015 da a. 92
5.8 His og am o a) clus e and b) speake impu i ies a ia ions be ween SPLDA
and PLDAUP ini ializa ions in MGB 2015 da a. . . . . . . . . . . . . . . . . . 92
6.1 4-le el ee clus e ing example . . . . . . . . . . . . . . . . . . . . . . . . . . 99
6.2 PLDA ee-based clus e ing Bayesian Ne wo k . . . . . . . . . . . . . . . . . 103
6.3 M-algo i hm example o a clus e ing ee o dep h 4 . . . . . . . . . . . . . . 104
6.4 Es ima ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4 . . . 105
6.5 Maximiza ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4 . . 106
6.6 Analysis pe show o ∆Iand DER(%) o AHC, FBPLDA and PLDA ee-
based clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
x
LIST OF FIGURES
6.7 DER (%) esul s o he PLDA ee-based clus e ing wi h M-algo i hm in Al-
bayzín 2018 in e ms o δ,ζand M. . . . . . . . . . . . . . . . . . . . . . . 109
6.8 DER ela i e esul s be ween Random o de and Time o de . . . . . . . . . . 111
7.1 Scena io o in e es . a) U e ances ed and blue in he ea u e domain, wi h he
UBM componen s in g een. b) U e ances ed and blue in he i- ec o domain.
c) P ojec ions o he GMM componen s in he i- ec o domain o u e ances . 123
7.2 Compa ison o pos e io dis ibu ion o he i- ec o s wi h e e ence phoneme
dis ibu ion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
7.3 Compa ison o pos e io dis ibu ion o i- ec o s wi h modi ica ions in he phoneme
dis ibu ion α. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
7.4 Compa ison o pos e io dis ibu ion o i- ec o s when wo phonemes a e no
con ibu ing and αc= 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5 DET cu es o he scena ios Long-Long (blue), Long-Sho and Sho -Sho
Random ( ed con inuous and dashed line espec i e), Long-Sho and Sho -
Sho Balanced (g een con inuous and dashed line espec i e) o SRE10 "co eex -
co eex de 5 emale" expe imen . . . . . . . . . . . . . . . . . . . . . . . . . 130
7.6 T ial sco e in e ms o KL2 dis ance o he whole da a pool. Rep esen ed he
mean and he mean plus/minus he s anda d de ia ion . . . . . . . . . . . . . . 132
7.7 E alua ion me ics, EER (a) and minDCF (b) in e ms o he KL2 dis ance. . . 133
7.8 DET cu es o he scena ios Long-Sho Random and Sho -Sho Equalized
in SRE10 "co eex -co eex de 5 emale" . . . . . . . . . . . . . . . . . . . . . 135
7.9 No malized dis ibu ion o sco es o Ta ge (blue) and Non- a ge ( ed) ials
o scena ios Long-Sho (con inuous line) and Sho -Sho Equalized (dashed
line).Expe imen ca ied ou wi h SRE10 "co eex -co eex de 5 emale". . . . . 135
8.1 Example o a Bo leneck Fea u e ex ac o DNN . . . . . . . . . . . . . . . . . 140
8.2 BNF pipeline om he o iginal MFCCs up o Baum Welch s a is ics . . . . . . 140
8.3 Bayesian ne wo k o he phone ic i- ec o . . . . . . . . . . . . . . . . . . . . 142
8.4 Phone ic i- ec o pipeline om he o iginal MFCCs up o Baum Welch s a is ics 143
8.5 X- ec o a chi ec u e schema ic . . . . . . . . . . . . . . . . . . . . . . . . . 143
8.6 DET cu es o x- ec o s in SRE10 wi h long and sho u e ances . . . . . . . 150
8.7 Dis ibu ion o he embedding i s componen in s anda d i- ec o s, phone ic
i- ec o s and x- ec o s o he aining co pus in Albayzín 2018 . . . . . . . . 153
9.1 Schema ic o he supe ised and unsupe ised adap a ion . . . . . . . . . . . 159
xi
LIST OF FIGURES
9.2 Schema ic o unsupe ised independen adap a ion o he episodes n−1,n
and n+ 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
9.3 Schema ic o unsupe ised longi udinal adap a ion o he episodes n−1,n
and n+ 1.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
9.4 Semi-supe ised adap a ion s a egy based on he unsupe ised independen
adap a ion app oach o he episodes n−1,nand n+ 1.. . . . . . . . . . . . 160
9.5 Semi-supe ised adap a ion s a egy based on he longi udinal unsupe ised
adap a ion app oach o he episodes n−1,nand n+ 1 . . . . . . . . . . . . 160
9.6 ∆DER (%) pe o mance episode by episode o he wo shows o he e alua-
ion se . De ined as ∆DER = (DERINDEP −DERLONG). AHC e e s o he
Agglome a i e clus e ing pseudo-speake labels. . . . . . . . . . . . . . . . . . 164
A.1 Bayesian Ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opaga ion . II
xii
Objec i es and Me hodology
o each domain.
1.2 Objec i es and Me hodology
The objec i es o his hesis a e he imp o emen o dia iza ion capabili ies so ha sys ems
could wi hs and he ha m ul condi ions o he b oadcas domain. These e olu ions should be
in eg a ed in a single sys em, obus enough o deal wi h any so o audio om he s udied
en i onmen . The e o e, we should analyze possible e olu ions in he p e iously desc ibed
h ee lines o esea ch.
Rega ding o he speake cha ac e iza ion p oblem, we wan o imp o e he ex ac ion o
he speake ep esen a ions, ob aining e icien and disc imina i e cha ac e iza ions o he
in ol ed speake s. Thus, we i s seek a deepe unde s anding abou he s a e-o - he-a mod-
elling echniques based on subspace p ojec ion. Once his knowledge is is acqui ed, i will le
us explo e he limi a ions o hese echnologies, as well as p opose new app oaches designed
acco dingly.
Wi h espec o he g ouping ask, ou goal is he imp o emen o he clus e ing echniques
es ima ing he dia iza ion pa i ions. Fo his pu pose, we make use o subspace-based ech-
niques, specially PLDA, explo ing di e en a chi ec u es and s a egies.
Finally, we also mus deal wi h he domain a iabili y. In his a ea we will y o p o ide
ools and s a egies capable o dec easing he deg ada ion o domain misma ch in ci cums ances
whe e in-domain da a is sca ce o una ailable. In o de o each his goal we will deal wi h
he domain adap a ion p oblem by explo ing he in e ence o unsupe isedly-c a ed pseudo-
speake labels, ob ained om he audio o dia ize. These labels should be la e used o speci i-
cally adap he ou -o -domain model o he e alua ion audio.
1.3 Thesis o ganiza ion
The ou line o his hesis is e y o ien ed o he di e en challenges we p e iously desc ibed.
Fo his eason, his wo k is di ided in i e main pa s, as shown in he concep ual map in
Fig. 1.2:
•Basic Knowledge: This pa is dedica ed o p esen he dia iza ion p oblem and an
o e iew o he al eady p oposed echniques in he s a e o he a (Chap e 2). Mo e-
o e , his pa also s a s he expe imen al ac i i y, analyzing he cha ac e is ics o he
b oadcas domain and he pe o mance o a baseline dia iza ion sys em (Chap e 3).
4
Chap e 1. In oduc ion
Dia iza ion
Basic
Knowledge
S a e o
he a
Baseline
dia iza ion
Sys em
Model
Adap a ion
Speake
Clus e ing
FBPLDA
Clus e ing
T ee-based
Clus e ing
Unce ain y
P opa-
ga ion
Conclusions
&
Fu u e Wo k
Speake
Rep esen a ion
Embeddings
in Sho
U e ances
DNN
Embeddings
Figu e 1.2: Concep ual map o he s udied opics in his Thesis
5
Thesis o ganiza ion
•Speake Clus e ing: This pa is ocused on he di e en ools o imp o e he pe o -
mance o he clus e ing s age. Fi s , we analyze he pe o mance o he Fully Bayesian
P obabilis ic Linea Disc iminan Analysis (FBPLDA) model, dealing wi h i s weak-
nesses (Chap e 4). Chap e 5upda es he FBPLDA mode including he concep o Unce -
ain y P opaga ion (FBPLDAUP). Finally, in Chap e 6we p esen a o ally independen
clus e ing solu ion by means o a ee-based app oach.
•Speake Rep esen a ion This pa o he hesis pays a en ion o he way speake in o -
ma ion is ex ac ed om an audio u e ance and compac ed in o a condensed ep esen-
a ion, he embedding. Fi s , we s udy he s anda d app oxima ion o his in o ma ion
ex ac ion, analyzing i s impac on sho u e ances (Chap e 7). La e on, we make use
o he lea n conclusions, applying hem on he ob en ion o DNN-based embeddings o
dia iza ion (Chap e 8).
•Model Adap a ion: This pa , consis ing on Chap e 9, wo ks on he unsupe ised ex-
ac ion o in-domain in o ma ion, sui able o he adap a ion o ou -domain labels. This
•Summa y: This inal pa summa izes he conclusions o all he di e en pa s o he
hesis and p oposes how his esea ch could be ollowed in he u u e (Chap e 10)
6
Pa I
Dia iza ion Basic Knowledge
7
Chap e 2
Dia iza ion S a e o he A
The objec i e o his chap e is he e ision o he s a e o he a in di-
a iza ion. Fo his pu pose, we ake in o accoun impo an e iews such as
[Angue a e al., 2012][T an e and Reynolds, 2006]. Ou i s goal is he iden i ica ion o
he main domains in which dia iza ion has been applied. This di e en ia ion helps unde s and-
ing he e olu ion o dia iza ion echnologies. This knowledge allows he in oduc ion o he
wo main app oxima ions o dia iza ion. Then, we explain in de ail he unc ional blocks o
he mos popula dia iza ion app oach. Finally, he las pa o he chap e includes a e iew
abou how o measu e dia iza ion pe o mance.
2.1 In oduc ion
The dia iza ion ask includes all he echniques and p ocedu es needed o di e en ia e he con-
ibu ions o speake s gi en an audio. In he mos gene al case, dia iza ion wo ks in an unsupe -
ised way, i.e. wi hou p io knowledge abou he in ol ed speake s no i s numbe . Howe e ,
dia iza ion can ge bene i ed by means o he knowledge o hese cha ac e is ics, usually sim-
pli ying he p oblem. His o ically, dia iza ion esea ch has ocused on h ee main domains o
in e es :
•Telephone channel domain. This en i onmen in ol es he analysis o elephone con e -
sa ions, cha ac e ized by he p esence o ew speake s, usually wo, and con e sa ional
speech wi h sho in e en ions. Mo eo e , elephone con ex usually conside s close- o-
mou h mic ophones and es ic ed a p io i known channel condi ions.
•B oadcas domain. This condi ion includes audios om mass media b oadcas e s (TV,
adio, VoD, e c.). The mos impo an ea u e in b oadcas da a is he la ge a iabili y
9
In oduc ion
o condi ions. The a iabili y in he numbe o speake s is almos un es ic ed: om
3-4 up o 100 di e en speake s pe hou o con en , depending on he show. The e is
also a iabili y in he ype o speech: while some shows con ain mo e ead speech, e.g.
he news, o he s ha dly e e include i , being mainly composed o con e sa ional speech,
such as alk-shows. This cha ac e is ic has g ea ele ance in dia iza ion due o he leng h
o he in e en ions. Whils con e sa ional speech usually consis s o sho in e en ions
in o de o main ain he con e sa ion low, ead speech can gene a e longe u ns due o
he absence o eedback. Mo eo e , b oadcas audio also p esen s a iabili y o acous ic
scena ios, such as s udio and ou doo s, each one wi h i s own acous ic cha ac e is ics.
Finally, excep o li e con en , speech signal usually main ains high Signal o Noise
Ra io (SNR), al hough e y o en speech is pa ially occluded by complemen a y acous ic
addi ions such as music, and noises like canned laugh e and applauses.
•Mee ings domain. This scena io implies eco dings om mee ing ooms, whe e an un-
de e mined numbe o people is eco ded om one o mul iple mic ophones. Hence,
eco dings om his domain mainly include con e sa ional speech. In his domain eco d-
ing condi ions a e also e y ele an . Despi e he ac ha close- o-mou h mic ophones
can be used, mo e o en omnidi ec ional mic ophone a ays a e conside ed. These a ays
can be loca ed in a single poin , e.g. on op o he con e ence able, o sp ead along he
oom. Rega dless o he mic ophone loca ions, he dis ance be ween speake and mic o-
phone canno be igno ed. This dis ance is esponsible o no iceable channel e ec s in
he speech p opaga ion up o he mic ophones, including deg ada ions as e e be a ion.
Besides, he s a iona i y o his ansmission channel canno be gua an eed, a ec ed by
he ela i e mo emen s be ween speake and mic ophone. Finally, hese channel e ec s
usually imply powe losses o he signal, making speech quali y mo e sensi i e o noises.
The his o ic e olu ion o dia iza ion o iginally s a ed in he elephone domain. Due o
i s cha ac e is ics his domain p o ided he mos es ic ed e sion o he dia iza ion p ob-
lem. Besides, he e was a g ea in e es o dia iza ion solu ions included in speake ecog-
ni ion applica ions. This is why since 1996 dia iza ion was pa o NIST SRE e alua ions
[P zybocki and Ma in, 2004]. Only a e speake ecogni ion e ol ed i s ools in e ms o accu-
acy and obus ness, dia iza ion was able o expo i s knowledge o al e na i e domains, as in
Rich T ansc ip ion (RT) e alua ions [Ga o olo e al., 2002], whe e al e na i e domains (b oad-
cas news and mee ings) complemen ed he con e sa ional elephone speech.
10
Chap e 2. Dia iza ion S a e o he A
SPK 1 SPK 2 SPK 3 SPK 4 SPK 1
COARSEST PARTITION
FINEST PARTITION BOTTOM-UP
TOP-DOWN
Figu e 2.1: Schema ic o Bo om-Up and Top-Down dia iza ion
2.2 Main dia iza ion s a egies
Along li e a u e se e al op ions ha e been p oposed o he ob en ion o he dia iza ion labels.
Howe e , mos o hese con ibu ions can be g ouped in o wo main concep ual app oaches:
Bo om-Up and Top-Down dia iza ion s a egies. Fig. 2.1 illus a es bo h dia iza ion app oaches
in o de o ob ain he same dia iza ion labels.
•Bo om-Up. The gi en audio is i s di ided in o indi idual segmen s, in which a single
speake is assumed o be p esen . Then, hese segmen s a e clus e ed so all blocks om
he same speake a e agged wi h he same label.
•Top-Down. This al e na i e conside s he opposi e s a ing poin . This app oach s a s
conside ing a single speake esponsible o all he audio. A e wa ds, he ini ial clus e is
di ided ying o ma ch each inal clus e wi h a eal speake in he audio.
In spi e o hei opposi e app oach, bo h s a egies need o sol e he same wo challenges:
De e mining whe he some pa o he audio con ains speech om a single speake and inding
he bounda ies i necessa y. Despi e he appa en simplici y o bo h asks, hei de elopmen
o eal applica ions has equi ed se e al con ibu ions in he li e a u e. Ne e heless, bo h asks
a e s ill a o being o ally sol ed.
Despi e bo h Bo om-Up and Top-Down app oaches a e equally alid, hey a e no simila ly
popula . While bo h op ions ha e been de eloped along mul iple publica ions, in ecen yea s
11
Main dia iza ion s a egies
Figu e 2.2: Gene al schema ic o a dia iza ion sys em
he Bo om-Up s a egy has gained much mo e awa eness han he Top-Down coun e pa . A
eason o his popula i y is he i among he la es imp o emen s in speake ecogni ion and
he Bo om-Up dia iza ion pipeline, making hei inclusion s aigh o wa d. Unde hese ci -
cums ances, Bo om-Up dia iza ion has aken i s pe o mance o unp eceden le els o quali y.
In consequence, his op ion has ecen ly gained popula i y becoming he s anda d dia iza ion
app oach nowadays.
2.2.1 Bo om-Up dia iza ion sys ems
The popula i y o Bo om-Up dia iza ion has inspi ed he de elopmen o a s anda d a chi ec-
u e, which we p esen in Fig. 2.2. This schema ic desc ibes he s anda d conside ed blocks o
ans o m he inpu aw audio in o he desi ed inal labels. The unc ionali y o each block is
desc ibed as ollows:
•Acous ic Fea u e Ex ac ion. Raw speech audio is a e y complex signal wi h many
so s o in o ma ion. While some o hem a e aluable depending on he applica ion
(speake , speech, language, e c.), o he s a e no o in e es (channel, noises, e c.) because
hey can al e ou es ima es. The ea u e ex ac ion s ep aims o ans o m he aw signal
in o a ai h ul bu compac ep esen a ion o he acous ic in o ma ion, simpli ying he
access o ou a ge in o ma ion and compensa ing hose ha m ul deg ada ions.
•Segmen a ion. Gene ally speaking, segmen a ion is he ask o di iding an audio in o
pieces acco ding o an a ibu e, which should emain homogeneous along he o al leng h
o each piece. Focusing on dia iza ion, he di ision a ibu e is he speake iden i y. Thus,
he goal o dia iza ion segmen a ion is he di ision o a gi en audio in o segmen s whe e
a single speake is p esen in hem. This sys em mus exploi he homogenei y o da a
in sho pe iods o ime o ind he bounda ies be ween speake s. An ideal segmen a ion
s ep should p o ide he ime ma ks o he di e en speake in e en ions in an audio.
12
Chap e 2. Dia iza ion S a e o he A
•Speake Cha ac e iza ion. Speake cha ac e iza ion is a high-le el in o ma ion ex ac-
ion which collec s he speake in o ma ion om he acous ic ea u es. In o de o p op-
e ly do so, i equi es wo king wi h audio om a single speake . Thanks o his equi e-
men , highly e ol ed echniques wo k along he gi en inpu segmen s, enhancing hei
speake disc imina i e p ope ies while compensa ing he ha m ul a iabili ies. Mo eo e ,
his p ocess usually con e s a iable-leng h segmen s in o ixed-dimension compac ep-
esen a ions, mo e sui able o pos p ocessing.
•Clus e ing. The ou pu o he segmen a ion s ep is a se o acous ic agmen s wi h a
single speake in each o hem. Howe e , he same speake may ha e p oduced mo e
han one segmen . The clus e ing s age is esponsible o g ouping all hose segmen s
om he same speake and label hem wi h a unique ag. Fo his pu pose, clus e ing
akes he segmen ep esen a ions as inpu , gene a ing he dia iza ion labels as ou pu .
•Resegmen a ion. Resegmen a ion is an op ional ex a segmen a ion s ep o e ine he
ini ial segmen a ion bounda ies. This ex a bo de uning akes ad an age o he in e ed
clus e ing ou pu , wi h an accu a e knowledge abou he e alua ion audio. Resegmen-
a ion ou pu may be conside ed as dia iza ion labels o be edback in o he sys em o
u he e ining.
2.3 Acous ic ea u es o dia iza ion
In o de o di e en ia e speake s, dia iza ion sys ems equi e a subsys em capable o p o iding
disc imina i e cha ac e is ics a each ime s ep o an audio. These cha ac e is ics, also known
as ea u es, should maximize hei classi ica ion capabili ies. Fo his eason, hey y o ep e-
sen he audio in o ma ion in a ac able manne , simpli ying he in o ma ion ga he ing while
educing ha m ul so s o a iabili y (noise, channel in o ma ion, e c.) meanwhile. Mo eo e ,
ea u e ex ac ion is he i s dia iza ion block, hence no assump ion like numbe o speake s
no hei iden i y, speake ansi ions, e c. can be done.
The mos popula ea u es so a a e hose commonly known as sho - e m acous ic ea u es.
O iginally designed o speech ecogni ion, hese ea u es ca y ou a spec al analysis o he
aw signal while inspi ed by bo h he human p oduc ion and pe cep ion sys ems. Because he
speech signal is no s a iona y, his analysis mus be pe o med in sho analysis windows.
The mos popula ea u es a e he Mel F equency Ceps al Coe icien s (MFCCs), o iginally
p esen ed in [Da is and Me mels ein, 1980]. These ea u es p opose a sho - ime analysis o he
13
Audio segmen a ion
DKL(P||Q) = X
x
P(x) log P(x)
Q(x)(2.6)
Un o una ely, i s o iginal de ini ion is no symme ic, i.e., he KL di e gence o Q wi h
espec o P (DKL(P||Q)) may no be he same as he di e gence o P wi h espec o Q
(DKL(Q||P)). In consequence a symme ized e sion, known as KL2 di e gence, is used in-
s ead. This di e gence o dis ibu ions P and Q is de ined as:
DKL2(P||Q) = DKL(P||Q) + DKL(Q||P)(2.7)
Mo ing o dia iza ion, his dis ibu ion is conside ed in segmen a ion in [Siegle e al., 1997]
[Delacou and Wellekens, 2000].
2.4.1.3 Deep Neu al Ne wo ks (DNNs)
Thanks o he e olu ion o neu al ne wo ks many o he asks p e iously ca ied ou by o he
means, such as s a is ics, a e now pe o med by his echnology. Rega ding segmen a ion, some
con ibu ions ha e a emp ed he inclusion o DNNs in his ask.
In [Gup a, 2015] DNNs a e used as classi ie s. The hypo he ical bounda y ame is s acked
along i s con ex window, eeding a monoli hic DNN consis ing o eed o wa d laye s. The
inal laye classi ies he bounda y ame as eal o no . Mo eo e , a likelihood measu e can be
ob ained in he p ocess.
By con as , DNN eg ession capabili ies can also been applied. In [H uz and Zajic, 2017]
he neu al ne wo k mus ca y ou he eg ession o he ansi ion p obabili y, so ened du ing
aining. Fo his pu pose, inpu da a is ea ed by means o s acks o con olu ional neu al
ne wo ks.
In bo h cases, DNNs wo k as s andalone sys ems. Howe e , bo h a chi ec u es i he gi en
mo e gene al de ini ion, whe e a neu al ne wo k p o ides a me ic o a ixed-leng h analysis
window and compa ed agains a h eshold.
2.4.2 Model-based segmen a ion
Despi e he ac ha me ic-based segmen a ions a e he mos popula ones, o he al e na i es
ha e also been p oposed. Conside ing model-based segmen a ions, [Li e al., 2009] conside s
Hidden Ma ko Models (HMMs) o segmen a ion. The model ep esen s each class by means
o a 64-Gaussian GMM. This concep is e ol ed in [Diez e al., 2018], whe e classes a e ep e-
sen ed wi h ied GMMs, mo e sui able o speake ep esen a ion.
20
Chap e 2. Dia iza ion S a e o he A
Finally, some sys ems [Ga cia-Rome o e al., 2017][Diez e al., 2019] wo k in e ms o a
coa se SbC app oach. They p e e wo king wi h e y sho ixed-leng h (a ound 1.5 seconds)
segmen s, no aking ca e o bounda ies. These sys ems ely on he la es speake cha ac e i-
za ion echniques, which ha e e ol ed o p o ide obus enough ep esen a ions when wo king
wi h e y sho segmen s. By doing so, hey alle ia e he compu a ional cos s while only in o-
ducing a small p opo ion o co up ed segmen s: as many deg aded segmen s as eal bound-
a ies. Besides, hese sys ems usually coun wi h esegmen a ion sys ems o elimina e he gene -
a ed dis o ions once speake models a e a ailable.
2.5 Speake cha ac e iza ion
The na u e o speech makes his in o ma ion o ha e a sequen ial na u e. Human beings con-
ca ena e mul iple sounds o ansmi he desi ed in o ma ion. Howe e , he e is no limi a-
ion in e ms o i s leng h no he message. I can ei he be a la ge speech o a sho eply
o a closed ques ion, i.e., "yes" o "no". Mo eo e , i can include all he acous ic uni s o
jus a es ic ed se . Speake ecogni ion echnologies should p o ide a ool o obus ly en-
code he iden i y o he in ol ed speake ega dless o he in a-speake a iabili y, i.e. he
a iabili y wi hin all he possible u e ances om he same speake . Some e iews such as
[Fu ui, 2004][Kinnunen and Li, 2010] p o ide a good o e iew abou he e olu ion o hese
echnologies.
2.5.1 Ea ly days
Some o he i s success ul speake ecogni ion sys ems we e based on he co ela ion o
spec og ams [P uzansky, 1963]. This idea was la e e ol ed o ake in o accoun he o -
man analysis [Dodding on, 1971]. Because hese echniques we e no powe ul enough o
deal wi h ex -independen ecogni ion, some al e na i es we e explo ed o he ollowing
decade. Some p oposals du ing hose yea s a e he ins an aneous spec a co a iance ma ix
[Li and Hughes, 1974], spec um and undamen al equency his og ams [Beek e al., 1977] o
linea p edic ion coe icien s [Sambu , 1972]. The ollowing g ea e olu ion appea ed wi h he
conside a ion o empla e models: Dynamic Time Wa ping (DTW) [Fu ui, 1981] and Vec o
Quan iza ion (VQ) [Rosenbe g and Soong, 1987] [Soong e al., 1985], which p oposes sho
ime ea u e ec o s comp essed in codebooks. This p inciple was la e e ol ed as long as
ma ix quan izie s o mul i- ame we e also p oposed [Juang, 1990].
21
Speake cha ac e iza ion
2.5.2 Model-based ep esen a ions
In he 80s, a g ea e olu ion in he cha ac e iza ion philosophy was p oposed. Ra he han con-
side ing speech as a de e minis ic p ocess whe e ea u es could be measu ed, s a e-o - he-a
con ibu ions s a ed o de ine s a is ical models as gene a o s o speech. Mo eo e , hese gen-
e a o s we e o en designed only aking in o accoun he acous ic in o ma ion, no conside ing
high-le el c a ed ea u es.
The gene a i e so o solu ion has many ad an ages. Fi s , all segmen s a e supposedly
gene a ed by a known dis ibu ion, a pa ame ic solu ion pe ec ly desc ibed by a closed se o
pa ame e s, some o hem speake dependen . Besides, his so o solu ion allows he same
model o wo k wi h a iable-leng h segmen s while p o iding a ixed-dimension speake ep-
esen a ion. Finally, s a is ical solu ions can also p o ide p o ec ion agains di e en ypes o
andomness associa ed wi h he oice (phone ic a iabili y, noises, e c.).
When choosing he dis ibu ion o be e ep esen speake s, Gaussian dis ibu ions a e usu-
ally aken in o accoun . Ve y well known among s a is icians, Gaussian dis ibu ions ha e wo -
hy p ope ies. Howe e , Gaussian dis ibu ions a e oo simple o p ope ly ep esen all he
a iabili y and condi ions in speech. Thus, combina ions o hem, Gaussian Mix u e Models
(GMMs) a e conside ed ins ead. This app oach is conside ed unde he assump ion ha a linea
combina ion o enough Gaussians should be able o ep oduce any dis ibu ion.
2.5.2.1 Gaussian Mix u e Models (GMMs)
Gaussian Mix u e Models a e gene a i e s a is ical models i s in oduced in speake ecogni-
ion in [Reynolds and Rose, 1995]. They a e composed by he weigh ed sum o CGaussian
componen s, each one wi h i s own weigh πc, mean ec o µcand co a iance ma ix Σc, be-
ing c= 1..C. Thus, he sequence O={o1, ..., on, ..., oN}gene a ed by a GMM has been
andomly d awn as:
P(O|M) =
N
Y
n=1
C
X
c
πcN(on|µc,Σc)(2.8)
The e alua ion o hese sys ems wo ked in e ms o a loglikelihood a io. Two loglikelihood
e ms we e conside ed, bo h aking in o accoun he es audio audio es bu conside ing wo
di e en models: A model o he claimed en oll speake (Men oll) and a model ep esen ing
speake s excep o ou en ollmen one (Men oll).
ll = ln P(audio es |Men oll)
P(audio es |Men oll)(2.9)
22
Chap e 2. Dia iza ion S a e o he A
While Men oll was s aigh o wa d, he de ini ion o Men oll was no so clea . Many sys ems
wo ked wi h a pool o coho models, chosen o each ial acco ding o di e en c i e ia.
Then, his idea was e ol ed in [Reynolds e al., 2000], which p oposes he GMM-UBM
pa adigm. Fi s , his con ibu ion in eg a es he coho o al e na i e speake s in o a single
model, esponsible o ep esen he o al a iabili y o he acous ic da a. This gene al model, a
la ge GMM ained wi h se e al speake s, is known as Uni e sal Backg ound Model o UBM.
Fu he mo e, ins ead o building om sc a ch indi idual en oll models Men oll, i p oposes he
op ion o hei cons uc ion as a MAP adap a ion om he UBM, speci ically an adap a ion o
he componen means. The ob ained bene i s a e a igh e coupling be ween models, and as e
sco ing echniques. In his scena io, he p oposed ll was:
ll = ln P(audio es |Men oll)
P(audio es |MUBM)(2.10)
2.5.2.2 Suppo Vec o Machines (SVM)
The GMM-UBM s a egy became a miles one in speake e i ica ion, specially ega ding he
way o model speake s. Howe e , al e na i e sco ing app oaches we e a emp ed. Wi hin
his line o esea ch Suppo Vec o Machines (SVMs) we e p oposed o speake ecogni ion
[Campbell e al., 2006a], leading o he SVM-GMM s a egy.
Suppo Vec o Machines a e bina y classi ie s ha p ojec he inpu da a in o a high-
dimensional space whe e a hype plane sepa a es he wo classes. The e alua ion in SVMs
is de ined as ollows:
(x) =
C
X
c=1
ycαcK(xc, x) + b(2.11)
whe e (x)s ands o he dis ance o he u e ance xwi h espec o he hype plane. xc,αc
and yc ep esen he suppo ec o s, weigh s and labels espec i ely, wi h he es ic ion ha
PC
c=1 ycαc= 0 and αc>0. The labels yc ake he alue +1 o one class and −1 o he o he
one. Besides b ep esen s he hype plane bias. Finally, K(·,·)s ands o he ke nel unc ion, e-
sponsible o p ojec ing he da a in o he high-dimension space and calcula ing dis ances e ms
I K(·,·)is es ic ed o sa is y he Me ce condi ion, he Ke nel condi ion can be exp essed as
an inne p oduc as:
K(x, y) =< g(x), g(y)>(2.12)
whe e g(·)is he ans o ma ion in o he highly dimensional space.
SVMs a e ained by a maximum ma gin s a egy. This ype o aining mus iden i y a
hype plane which p ope ly classi ies he aining elemen s while sa is ying he ollowing e-
23
Speake cha ac e iza ion
s ic ion: he chosen hype plane mus keep he maximum dis ance wi h espec o he aining
popula ions o bo h classes. This eques o ces he hype plane o p o ide he maximum ma gin
p o ec ion agains spu ious da a du ing e alua ion.
The inclusion o SVMs in speake ecogni ion [Campbell e al., 2006a] was ca ied ou by
p oposing a ke nel ha bounds he KL di e gence. In ou scena io, KL di e gence measu es he
dis ance be ween u e ances Oaand Ob, modeled by GMMs Maand Mb espec i ely. Bo h
GMMs a e ob ained acco ding o he GMM-UBM pa adigm, hence hey sha e he componen
weigh s πcand he componen co a iance ma ices Σc, only di e ing a he componen means
µc. The ke nel accomplishing his eques is:
K(Oa,Ob) =
C
X
c=1 √πcΣ1/2
cµa
c√πcΣ1/2
cµb
c(2.13)
In consequence, he ke nel unc ion can be in e p e ed as he inne p oduc o he wo GMM
supe ec o s, a conca ena ion o he GMM means unde going a diagonal scaling. Applied o ou
p e ious de ini ion o SVMs, he en ollmen supe ec o cons i u es he se o suppo ec o s
and he es supe ec o plays he ole o e alua ed u e ance, deciding whe he i comes om
he en ollmen speake .
The GMM-SVM pa adigm was complemen ed wi h he Nuisance A ibu e P ojec ion
(NAP) concep . This idea, in oduced in [Campbell e al., 2006b], conside ed he compensa ion
o he in a-speake a iabili y p esen in he supe ec o s. This compensa ion is pe o med by
es ima ing a low ank ma ix U, also known as eigen-channels ma ix, which de ines he in a-
speake a iabili y wi hin he supe ec o space. Once modeled, a ma ix P=I−UUTcan be
in oduced in he ke nel unc ion al eady seen in eq (2.13):
K(Oa,Ob) =
C
X
c=1 √πcΣ1/2
cµa
cP√πcΣ1/2
cµb
c(2.14)
2.5.2.3 Join Fac o Analysis (JFA)
Join Fac o Analysis (JFA) [Kenny, 2005] is an e olu ion o he GMM ep esen a ions, paying
special a en ion o wo concep s de eloped wi h he GMM-SVM app oach: supe ec o s and
subspaces o ce ain a iabili ies. Taking bo h concep s in o accoun he JFA me hodology
e ol es de he GMM-UBM pa adigm decomposing he adap ed GMM supe ec o as a sum o
e ms:
µj=µUBM +Vyi+Uxk+Dzj(2.15)
24
Chap e 2. Dia iza ion S a e o he A
whe e µjis he adap ed supe ec o mean o u e ance j.µUBM ep esen s he mean supe ec o
om he UBM model. The e m Vyiis he speake dependen e m. Vis a low ank ma ix
desc ibing he subspace o he in e -speake a iabili y while yiis a ied la en a iable, i.e.
a la en a iable whose alue is he same o all u e ances om speake i, esponsible o
he u e ance. Simila ly, we ha e he e m Uxko channel e m. Uis a low ank ma ix
desc ibing he channel a iabili y space and xkis he ied la en a iable o all u e ances wi h
he same channel k, including u e ance j. Finally, we ha e he e m Dzj, which mus explain
he emaining a iabili y. Fo his pu pose, Dis a diagonal ma ix and zja la en a iable unique
o he u e ance. All he h ee la en a iables, yi,xkand zja e S anda d No mal dis ibu ed.
2.5.3 Embedded ep esen a ions
The ollowing la ge e olu ion implied he imp o emen o he al eady p oposed models, bu
also a new me hodology. On he one hand, models including la en a iables o map he speake
in o ma ion signi ican ly imp o ed he pe o mance. On he o he hand, he app oach o he
GMM-SVM showed ha in o ma ion could be ex ac ed om he models and independen ly
ea ed. The combina ion o bo h ideas c ea ed he embedding pa adigm, de ining models ha
cons ain he speake in o ma ion in o a es ic ed space whe e a la en a iable should explain
each speake . F om hese la en a iables we could ex ac compac ep esen a ions, oicep in s
o each speake , also known as embeddings.
Embeddings o e se e al ad an ages compa ed wi h p e ious app oaches. Once embed-
dings a e ex ac ed, hey can be decoupled om he o iginal ex ac ion me hod, simpli ying
hei s o age. Mo eo e , his decoupling makes impossible he e u n o he o iginal audio, gua -
an eeing p i acy. Finally, he ob ained embeddings can be pos p ocessed by al e na i e me hods,
also known as backends. In ac , he cu en speake ecogni ion s a e o he a , om which
mos o hese echnologies a e concei ed, is domina ed by he embedding-backend pipeline.
In he ollowing lines some o he mos popula embeddings a e p esen ed, and one o he
mos popula backends, he P obabilis ic Linea Disc iminan Analysis (PLDA) is explained
a e wa ds.
2.5.3.1 I- ec o s
I- ec o s [Dehak e al., 2011] a e a di ec e olu ion o he JFA modeling. Ra he han di e -
en ia ing be ween speake and channel ac o s, i- ec o model in eg a es hem in o he o al
a iabili y subspace. This usion makes he la en a iable s o e bo h speake and channel in-
o ma ion oge he . Mo eo e , la en a iables a e no linked among u e ances anymo e, being
25
Speake cha ac e iza ion
only ied along he samples om he u e ance. Besides, his model no longe conside s a
esidual a iabili y e m. In consequence, he u e ance j, consis ing o he sequence o ames
O={o1, ..., on, ..., oN}, is now modeled by a GMM whose mean supe ec o µjis de ined as:
µj=µUBM +Twj(2.16)
whe e µUBM again desc ibes he UBM mean supe ec o . Ts ands o a low ank ma ix de-
sc ibing he o al a iabili y subspace and wjis he la en a iable depending on he u e ance.
The men ioned model s ill can be e alua ed in e ms o likelihoods as in JFA. Ne e heless,
his echnology e ol ed o become a oicep in ex ac o . The commonly used i- ec o is he
mean o he pos e io dis ibu ion o he la en wjgi en he u e ance j. This dis ibu ion is
Gaussian and de ined as:
wj∼N (wj|µw,Σw) = Nwj|L−1
wΓw,L−1
w(2.17)
Γw=
C
X
c=1
TT
cΣc
Nj
X
n=1
γnc (on−µc) =
C
X
c=1
TT
cΣcFc(2.18)
Lw=I+
C
X
c=1
TT
c
Nj
X
n=1
γncΣcTc=I+
C
X
c=1
TT
cNcΣcTc(2.19)
whe e µw ep esen s he mean o he pos e io dis ibu ion and Σwis i s co a iance. These
e ms a e cons uc ed in e ms o Tc, he subma ix om Tdesc ibing he con ibu ion o he
c h Gaussian componen , and Σc, he co a iance ma ix o he c h componen in he UBM. The
in o ma ion o he u e ance is con ained in Ncand Fc, he ze o h and cen e ed i s o de Baum
Welch s a is ics o he c h componen espec i ely. Finally, Nj ep esen s he o al amoun o
samples in he u e ance j. Bo h o hem a e ob ained in e ms o he esponsibili ies γnc, he
p obabili y o he n h sample on o be d awn om componen co he GMM-UBM.
2.5.3.2 Hyb id i- ec o s
The la es g ea e olu ion o neu al ne wo ks, a ec ing bo h so wa e and ha dwa e, has be-
come a miles one along mos a i icial in elligence asks. This e olu ion also eached speech
echnologies [Hin on e al., 2012], including speake cha ac e iza ion. In hese asks, a i s ,
his acquisi ion o he new app oaches was smoo h, complemen ing exis ing s a e-o - he-a
echnologies.
A p oposed inclusion o DNNs in i- ec o s was p esen ed in [Lei e al., 2014] as hyb id i-
ec o s. The i- ec o ex ac o p inciple is he same, i.e. i explo es how an u e ance speci ic
model di e s om a UBM due o he unique cha ac e is ics o he u e ance. Howe e , he
26
Chap e 2. Dia iza ion S a e o he A
UBM is no a GMM anymo e. Now his ole is played by a DNN, disc imina i ely ained
o disce n phoneme senones. This neu al ne wo k is now in cha ge o he esponsibili ies γnc
equi ed o compu e he Baum Welch s a is ics Ncand Fc, he unique inpu o i- ec o aining.
Howe e , due o he ac ha no GMM-UBM is in ol ed, γnc now ep esen s he p obabili y o
he ea u e ame on o con ain he he senon cins ead.
An al e na i e p oposal a e phone ic i- ec o s [Viñals e al., 2019d]. This p oposal se s an
o iginal i- ec o model in which he GMM-UBM esponsibili y depends on a p io ac i a ion,
con olled by a DNN phoneme classi ie . Unde his app oach, he se o C componen s is
decomposed in mul iple subse s, each one esponding o indi idual phonemes. By doing so,
pa icula phoneme models become mo e speci ic while educing acous ic unce ain ies.
2.5.3.3 DNN embeddings
The imp o emen s o hyb id i- ec o s we e ou s anding, ou pe o ming pas echnologies. The
esul s in [Sadjadi e al., 2016] p esen ed an unp eceden pe o mance combining DNN pos e i-
o s wi h BNFs. Howe e , echnologies we e s ill su e ing om i- ec o s laws.
The p oposed e olu ion was a cu ing-edge idea. Ra he han e ol ing he gene a i e i-
ec o model, i ains a o ally disc imina i e DNN. In [Snyde e al., 2016] x- ec o s we e
p oposed ollowing his idea: A neu al ne wo k is ain o classi y an audio among a closed se
o speake s. The inpu ea u es i s unde go mul iple ame-le el non-linea ans o ma ions
and hen hey a e pooled in o an u e ance p ojec ion. This p ojec ion goes h ough u e ance-
le el non-linea ans o ma ions be o e i s classi ica ion. The ne wo k is ained o ecognize
a la ge pool o speake s by means o c oss en opy. Gi en a ained ne wo k, he embeddings
also known as x- ec o s a e ex ac ed du ing he o wa d p opaga ion o he in o ma ion, in he
u e ance-le el ans o ma ions.
The g ea pe o mance o x- ec o s has encou aged he communi y o e ol e
o DNNs. Now mul iple al e na i es o x- ec o s a e a ailable, including LSTM
based a chi ec u es [Wang e al., 2018], Wide Residual Ne wo k based embeddings
[Villalba e al., 2019][Viñals e al., 2019d] o e en expanded x- ec o s [Villalba e al., 2019].
2.5.4 P obabilis ic Linea Disc iminan Analysis (PLDA)
PLDA is a s a is ical linea backend. De ined in [P ince and Elde , 2007] as a gene a i e model,
PLDA applies he subspace concep al eady conside ed in JFA, assuming he embedding φjas
a sum o a iabili y e ms:
φj=µ+Vyi+Uxj+ǫj(2.20)
27
Clus e ing
whe e Vyi ep esen s he speake a iabili y e m and Uxj he u e ance a iabili y coun e pa .
Bo h e ms consis o low ank ma ices (Vand U espec i ely), which de ine subspaces o he
la en a iables yiand xj espec i ely. Whe eas he speake la en a iable yiis ied along all
u e ances wi h he same speake , xjis pa icula o each embedding j. We conside hese la en
a iables, yiand xj, o be s anda d no mal dis ibu ed. Addi ionally, he model also includes
an ex a a iabili y e m ǫj o explain he esidual a iabili y in each pa icula embedding. ǫjis
modeled by means o a ze o-mean Gaussian dis ibu ion and diagonal co a iance ma ix D−1.
Finally, µis he cons an speake independen e m.
Al hough his model o e s a closed- o m solu ion, when i s ly applied on embeddings (i-
ec o s a ha ime), i s pe o mance was no signi ica i ely be e . I equi es embeddings
o be Gaussian in o de o p ope ly ob ain i s imp o emen , al hough he ex ac ed i- ec o s
we e a om his dis ibu ion. The mos popula solu ion o his issue is leng h no maliza ion
[Ga cia-Rome o and Espy-Wilson, 2011]. Embeddings, be o e eeding he PLDA model, a e
o ced o eassu e ha i s Euclidean no m is equal o one. This p ocess p ojec s he inpu
embeddings in o a hype sphe e o adius equal o 1. Be o e leng h-no maliza ion, embeddings
should be cen e ed and whi ened. By doing so, he esul ing embeddings a e sp ead along he
hype sphe e a he han being concen a ed in es ic ed egions o he hype sphe e, leading o
mo e disc imina i e capabili ies o he sys ems.
Mul iple al e na i es ha e appea ed o he o iginal PLDA model. The Simpli ied PLDA
(SPLDA) uses he channel and esidual e ms. Ano he al e na i e is he Disc imina i e PLDA
[Cumani e al., 2013a], which ains he same model in a disc imina i e manne . An impo -
an al e na i e is he Hea y-Tailed PLDA (HTPLDA) [Kenny, 2010]. This model was p o-
posed be o e leng h-no maliza ion as a way o deal wi h non-Gaussian embeddings by modi-
ying he p io dis ibu ions. Howe e , a e leng h-no maliza ion i s compu a ional complex-
i y discou aged i s usage. Ne e heless, wi h he ad en o DNN embeddings, a mo e non-
Gaussian han i- ec o s, HTPLDA p o ides small bene i s wi h espec o o he al e na i es
[B umme e al., 2018].
2.6 Clus e ing
The clus e ing s age in a Bo om-Up dia iza ion a chi ec u e is esponsible o he ga he ing o
he acous ic agmen s in e ms o hei speake . This du y can al e na i ely be conside ed as a
labeling ask. Being he audio o Nacous ic segmen s ep esen ed by he se Φ={φ1, .., φN}
o speake ep esen a ions o embeddings, clus e ing mus in e a pa i ion, a se o labels Θ =
{θ1, .., θN}, so ha hose segmen s om he same speake sha e a common label.
28
Chap e 2. Dia iza ion S a e o he A
Table 2.1: Bell numbe Bin e ms o he numbe o elemen s o clus e
Numbe o segmen s NNumbe o pa i ions B
1 1
22
3 5
415
552
6 203
... ...
10 115975
... ...
20 51724158235372
To do so, we i s equi e a measu e o de e mine how a ce ain pa i ion Θ i s he se o
embeddings Φ. This me ic may ha e mul iple na u es, e.g. s a is ical, g aphs, ke nels, e c.
Then, gi en he chosen me ic we mus ind he pa i ion wi h he bes me ic alue. Un o -
una ely, ega dless o he me ic hey all sha e he same di icul y: The bes pa i ion is only
gua an eed o be ob ained i all possible pa i ions a e analyzed, jus choosing he one wi h he
bes me ic. This op ion is usually e e ed as b u e- o ce app oach. Un o una ely, s udies such
as [B umme and de Villie s, 2010] e eal ha he numbe o pa i ions inc ease e y as as
long as he numbe o segmen s o clus e N ises. In ac , excep o e y low alues o N,
b u e- o ce app oaches a e in gene al no iable.
Gi en an audio wi h Nsegmen s, he o al numbe o possible pa i ions is desc ibed by he
Bell numbe B. This numbe , applicable o any clus e ing ask, ep esen s he o al numbe
o independen possible g ouping a angemen s, in ou case pa i ions, o a se o Nelemen s.
This numbe is de ined by means o a ecu en ela ion:
BN+1 =
N
X
n=0 N
nBn(2.21)
B0=1 (2.22)
This numbe inc eases e y as as long as he alue o N ises. In Table 2.1 alues o low
alues o Na e shown.
Acco ding o Table 2.1, e en e y low alues o Nimply huge numbe o candida e pa i-
ions. I we conside e en highe alues o N, e.g. 100 o 200 segmen s o a one-hou TV
show, he numbe o candida e hypo heses o compa e becomes in ac able. Fo una ely, no all
29
Pe o mance me ics
•MISS ERROR (MISS). Speech audio inco ec ly labeled as non-speech. This e m also
measu es he Voice Ac i i y De ec ion (VAD) pe o mance.
•FALSE ALARM ERROR (FA). Non-speech audio in which a speake is conside ed o
be p esen . Voice Ac i i y De ec ion (VAD) pe o mance is a ec ed by his e m as well.
•SPEAKER ERROR (SPK). Speech misclassi ied as gene a ed by an al e na i e speake .
•OVERLAP ERROR (OV). Pe iods o ime when mul iple speake s a e simul aneously
alking. This e o in ol es he es ima ion o he numbe o speake s (unde es ima ion
when no all speake s a e de ec ed and o e es ima ion when non-p esen speake s a e
also labeled) as well as he misclassi ica ion o he in ol ed speake s.
Rega ding he ou e o e ms, clea ly wo o hem, Miss E o and False Ala m, a e e-
la ed wi h segmen a ion, specially he VAD s ep. Wi h espec o he speake and o e lap e o s,
hese e ms mainly depend on he clus e ing s age. Howe e , hey a e ea ed di e en ly. Ac-
co ding o i s de ini ion, he o e lap e o in ol es any in e ence e o when mul iple speake s
a e alking. Thus, i measu es miss, alse ala m and speake e o s o hese pe iods o ime.
Hence o e lap is he mos challenging e o e m, wi h se e al p oposed con ibu ions abou i s
de ec ion (e.g. [O e son and Os endo , 2007][Zelenák and He nando, 2012]) bu wi hou any
unc ional solu ion ye . The e o e, in ce ain e alua ions his e m is ob ia ed o pe o mance
compa isons.
Due o he ac ha e o s a e non-o e lapped, DER can be decomposed on mul iple e ms,
each one e alua ing he deg ada ion due o each ype o e o . The al e na i e de ini ion o DER
is:
DER =LMISS +LFA +LSPK +LOV
L o al
(2.31)
=EMISS +EFA +ESPK +EOV (2.32)
whe e EMISS,EFA,ESPK and EOV a e he DER e o e ms o miss, alse ala m, speake and
o e lap causes espec i ely.
Despi e all i s bene i s ega ding simplici y and decomposi ion o e o , DER p esen s s ong
limi a ions. Ob ia ing he miss e o and alse ala m e ms, di ec ly ela ed wi h he VAD
pe o mance, no u he knowledge can be in e ed om DER abou he speake e o . A simila
sco e is ob ained i some amoun o audio is misclassi ied, ega dless o how many speake s
a e a ec ed. Besides, DER conside s all he audio uni o mly ele an , and e o s in ol ing
36
Chap e 2. Dia iza ion S a e o he A
he same amoun o audio a e equally ha m ul. Howe e , in eal li e nei he he speake s no
hei speech a e equally aluable, in alida ing his conside a ion. This is specially ele an
when speake s do no con ibu e o he audio wi h he same amoun o speech. Fo example,
some e o s may be i ele an o he mos alka i e speake s bu a mo e signi ican o hose
speake s con ibu ing wi h much less speech. The e o e, some c i ics abou DER me ic a e
a ising while he communi y is eage o inding an al e na i e sco e.
In ecen imes some al e na i e me ics ha e also been p oposed o dia iza ion asks. The
Mu ual In o ma ion (MI) me ic was de ined in DIHARD 2018, a dia iza ion e alua ion in
di icul condi ions. The idea behind MI is measu ing how much in o ma ion we ha e abou he
eal labels p o ided ou hypo hesed pa i ion. The p oposed me ic was p oposed o s udy and
was complemen ed by DER, which managed he leade boa d. The me ic is de ined as:
MI =
R
X
i=1
S
X
j=1
nij
Nlog2
nijN
isj
(2.33)
whe e R ep esen s he numbe o clus e s in he e e ence wi h idu a ion each, and Ss ands
o he numbe o hypo hesized clus e s, each one wi h du a ion sj. Besides, he e m nij is he
amoun o speech assigned o he speake iin he e e ence and o he clus e jin he hypo hesis.
Finally, Nsymbolizes he o al amoun o speech.
Ano he al e na i e is he Jacca d E o Ra e (JER). This me ic was p oposed as al e na i e
o DER in DIHARD 2019. The i s s ep in he e alua ion is a mapping among he Rclus e s in
he e e ence and Sclus e s in he hypo hesed pa i ion. This mapping is ca ied ou acco ding
o he Hunga ian algo i hm, so each clus e in he e e ence will be mapped o a mos one
clus e o he hypo hesis and ice e sa. Then o each speake in he e e ence we es ima e:
JER e =F A +MISS
T OT AL (2.34)
whe e T OT AL ep esen s is he amoun o audio p esen in bo h he e e ence clus e e and
he mapped coun e pa . I he e was no pai ed clus e , i s alue would be he o al amoun o
speech o speake e .F A s ands o he o al amoun o speech no p esen in he e e ence
clus e e bu conside ed as pa o he pai ed g ouping. I s alue is 0i no mapping o e
was ca ied ou . Finally, MISS is he amoun o speech p esen in speake e bu no included
in he mapped coun e pa . I e speake has no pai ed clus e , i s alue is equal o T OT AL.
Ha ing de ined he indi idual e ms JER e , he Jacca d e o a e o a eco ding is he
a e age o speci ic Jacca d e o a es:
JER =1
RX
e
JER e (2.35)
37
Pe o mance me ics
Rega dless o he used me ics, hey do no p o ide any clue abou he easons o he
misclassi ica ion o audio. Hence al e na i e me ics should be help ul o be e unde s and
he speake e o . This e o is mainly gene a ed in he clus e ing block, hus clus e ing
me ics, such as he complemen a y clus e ing and speake impu i ies, well desc ibed in
[ an Leeuwen, 2010] a e sui able o his ask.
The clus e impu i y ep esen s how well he clus e s om a hypo hesized pa i ion con ain
audio om a single speake . De ined in e ms o i s clus e pu i y coun e pa , clus e impu i y
is minimized as long as he ob ained clus e s con ain audio om a single speake . Howe e , i is
no obliga o y ha clus e s om he same speake sha e he same label, in alida ing he me ic
o dia iza ion.
Simila ly o he clus e impu i y concep , we can also de ine he speake impu i y. This
new concep desc ibes how well he speech om a speake is agged wi h a single label, and
i is minimized as long as mo e and mo e da a om one speake only equi es a single label.
This me ic is also in alid o dia iza ion because mul iple speake s in he same clus e do no
deg ade he inal sco e.
38
Chap e 3
Analysis o Dia iza ion in B oadcas Da a
Dia iza ion in b oadcas is a complex ask, composed o a la ge numbe o sub asks wo king
oge he , as seen in Chap e 2. Whils mos o he p e iously desc ibed echniques wo k well
in es ic ed condi ions as he elephone domain, dia iza ion in b oadcas da a equi es many
pa icula i ies o be aken in o conside a ion. In his chap e we analyze he b oadcas domain,
emphasizing i s pa icula i ies. Fo his pu pose, we i s in oduce a e e ence dia iza ion sys-
em. This sys em will se e o explo e he wide a iabili y along b oadcas da a a e wa ds. In
his analysis we co e bo h quan i a i e and quali a i e esul s, and s udy how esul s a e a -
ec ed by his unce ain y. Finally, acco ding o he analysis and ob ained esul s, we sugges
he di e en lines o esea ch, some o hem ea ed along his hesis.
3.1 The dia iza ion e e ence sys em
The e e ence sys em conside ed o his analysis is an AHC-based dia iza ion app oach, e y
common in he li e a u e as baseline sys em. This a chi ec u e ollows a Bo om-Up app oach,
i s di iding he aw signal in o segmen s, which a e la e clus e ed acco ding o hei speake
iden i y. Fig. 3.1 illus a es he basic a chi ec u e o he sys em.
In he ollowing lines we explain in de ail he se up o each elemen in ou sys em.
•Fea u e Ex ac ion Fo he audio ans o ma ion in o ea u es, we s ic ly conside
MFCCs, s anda d ea u es in he s a e o he a . Ou MFCC se up includes a 32-band
Mel il e bank, and a inal coe icien educ ion, only conside ing coe icien s C1-C20.
The ene gy in o ma ion is disca ded oo. The in e ed s eam o ea u e ec o s does no
include de i a i es, and unde goes a no maliza ion o i s mean and a iance (CMVN).
•Segmen a ion The ob ained s eam o ea u e ec o s a e he inpu o he segmen a ion
39
The dia iza ion e e ence sys em
Figu e 3.1: Schema ic o ou baseline dia iza ion sys em
s age. In he e e ence sys em he segmen a ion s ep is di ided in o wo independen
sub asks, Voice Ac i i y De ec ion (VAD) o di e en ia e speech/non-speech and Speake
Change Poin De ec ion (SCPD) o ob ain he speake u ns.
– Voice Ac i i y De ec ion (VAD) The in e ence o he VAD mask is done by means
o [Viñals e al., 2018a], using a segmen a ion-by-classi ica ion app oach, which
wo ks in e ms o DNNs. A 2-laye BLSTM DNN, wi h 256 neu ons pe laye ,
is aken in o accoun . Each elemen in he second BLSTM ou pu sequence is p o-
jec ed in o a bina y deciso . Thus, we in e one VAD label pe inpu ame. This
laye is ained and e alua ed in 3-second analysis windows. Whene e he audio
exceeds he window dimension, a sliding window analysis is ca ied ou . This anal-
ysis implies a 3-second window and 2.5 seconds o wa d s ep. Fo he 0.5-second
o e lap pe iod he in e ence wo ks as ollows. he i s 0.25 seconds a e ob ained
om he p eceding window while he emaining 0.25 seconds a e labeled wi h he
in e ence om he ollowing window. This o e lapping design choice is made o
a oid undesi ed windowing e ec s, specially nea he a i icial window bo de s.
– Speake Change Poin De ec ion (SCPD) Fo he SCPD ask we ely on well
known echniques. In his case we op o a SCPD hyb id solu ion led by a me ic-
40
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
based segmen a ion, in pa icula ∆BIC conside ing Gaussian dis ibu ions wi h ull
co a iance ma ix. We ope a e in a sliding window egime ollowing he desc ip ion
in Sec ion 2.4. We make use o an analysis window wi h a minimum leng h o h ee
seconds, and a 0.25-second accumula i e window expansion whene e he window
does no con ain any bounda y. Rega ding he hype pa ame e λ, i is adjus ed ac-
co ding o hose esul s ob ained du ing he de elopmen phase. The me ic-based
solu ion is combined wi h a silence-based s a egy, which assumes all speech/non-
speech ansi ions o be speake bo de s. In ac , hese bo de s a e used as ancho s
o he ∆BIC segmen a ion.
•Speake Cha ac e iza ion The es ima ed segmen s a e hen con e ed in o compac
ep esen a ions, each one summa izing he pa icula ea u e s eam o each segmen .
Among he mul iple op ions desc ibed in Sec ion 2.5, ou choice o he ype o ep e-
sen a ion is he i- ec o . In ou sys em i- ec o s a e in e ed by means o an ex ac o o
256 Gaussians and 100-dimension o al a iabili y ma ix T. The ex ac ed i- ec o s a e
cen e ed, whi ened and leng h-no malized be o e eeding he clus e ing s age.
•Clus e ing The clus e ing s age in ou e e ence sys em is cons uc ed a ound an AHC
app oach, using SPLDA pai wise log-likelihood a io (LLR) as me ic. Ra he han con-
side ing he o iginal AHC ha exhaus i ely e alua ing all simila i ies among clus e s,
we ollow a simpli ica ion desc ibed in [ an Leeuwen, 2010]. This simpli ica ion only
eques s he es ima ion o he ini ial pai wise simila i y among he embeddings. Then,
a each usion i e a ion we app oxima e he exhaus i e simila i ies by app oxima ions
conside ing he al eady es ima ed alues. Among he di e en op ions o ca y ou he
app oxima ion, we op o he UPGMA (unweigh ed pai g oup me hod wi h a i hme ic
mean) app oach [Sokal and Michene , 1958]. Rega ding he clus e ing s op c i e ion, i
is done by means o a h eshold expe imen ally adjus ed du ing de elopmen .
3.2 Analysis o b oadcas da a
B oadcas da a is a ype o domain specially cha ac e ized by i s wide a iabili y. Whene e no
es ic ions a e applied ega ding he shows o analysis, speech p ocessing asks mus be obus
enough o wi hs and a g ea ange o condi ions. Among he di e en asks a ec ed by his
a iabili y we mus ake in o accoun dia iza ion.
The e a e many easons o b oadcas da a o be so mu able. F om di e en eco ding loca-
ions like s udio and ou doo , o he conside ed equipmen . Fu he mo e, ex a ac o s should
41
Analysis o b oadcas da a
be conside ed, such as he acous ic addons, i.e. acous ic a i ac s like laugh e o applause, ha
co up he audio signal. Fo dia iza ion pu poses we will pay a en ion o his a iabili y along
he clus e ing s age. P e ious blocks can be in e p e ed as high-quali y ea u e ex ac o s, hus
clus e ing mus p o ide he knowledge o p ope ly g oup oge he hei ep esen a ions in o de
o ob ain he inal labels. Du ing clus e ing we mus deal wi h wo main ypes o a iabili y: he
one p esen in he acous ic ep esen a ions, he acous ic a iabili y, and he one a ailable in he
inal speake labels, he speake dis ibu ion a iabili y.
The e ec s o bo h ypes o a iabili y di e , specially due o hei in luence along he di-
a iza ion pipeline. The acous ic a iabili y is consequence o he di e en acous ic condi ions
along he di e en audios o in e es . These di e en condi ions cause embeddings om he
same speake o be less homogeneous, making hem less obus o clus e ing pu poses. In con-
sequence, his a iabili y may be esponsible o a deg ada ion o he dia iza ion pe o mance.
Rega ding he speake dis ibu ion a iabili y, we a e e e ing o he numbe o speake s in
an audio and how much speech hey con ibu e wi h. E o s in he es ima ion o he numbe o
speake s a e highly impo an because all he speech p oduced by he a ec ed o a o s will be
misclassi ied. Besides, he la ge is he ange o possible speake s o an audio, he la ge a e
he po en ial e o s in his es ima ion and mo e audio is usually in ol ed. A g ea ac o o ake
in o accoun in his es ima ion is he dis ibu ion o speech along he di e en speake s. Those
alka i e o a o s wi h se e al con ibu ions ha e enough audio o be obus ly modeled and small
misclassi ica ions ha e negligible e ec s on hem. By con as , hose speake s wi h e y ew
speech a e weakly ep esen ed and small e o s may cause hei loss.
We now p esen an analysis abou hese wo ypes o sou ces o a iabili y in b oadcas
da a. Fo his pu pose, we will ake in o conside a ion wo la ge da ase s: Mul i-Gen e B oad-
cas Challenge 2015[Bell e al., 2015] and Albayzín 2018 [O ega e al., 2018]. Bo h da ase s
include a la ge amoun o b oadcas audio con en co e ing a wide a iabili y o shows, gen es,
languages and media.
3.2.1 Mul i-Gen e B oadcas Challenge 2015 (MGB 2015)
This da ase was eleased o he Mul i-Gen e B oadcas challenge in 2015 [Bell e al., 2015].
This challenge aims a p ocessing asks in he b oadcas domain, including ASR, alignmen
and dia iza ion. The da ase consis s o app oxima ely 1600 hou s o B oadcas audio collec ed
om B i ish B oadcas ing Co po a ion (BBC) along ou o i s channels. The o al amoun o
audio in ol es a ound 1200 episodes om 500 di e en shows. All his audio is di ided in o
h ee subse s: ain, longi udinal de elopmen and e alua ion. The subse di ision ies no o
42
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
sha e di ec knowledge among he subse s, hus all episodes om a show a e placed in he same
subse . Aside he audio, he h ee subse s we e dis ibu ed wi h dia iza ion labels. Howe e , he
label accu acy is no uni o m among subse s. Whils he aining subse includes he o iginally
b oadcas sub i les e ined by a ligh ly-supe ised ASR alignmen as me ada a, de elopmen
and e alua ion labels a e manual anno a ions. Finally, MGB 2015 also con ains an ex a subse ,
namely de elopmen , eleased o he ASR e alua ion. This subse consis s o 28 hou s om 47
shows, wi h manually anno a ed VAD ma ks.
3.2.2 Albayzín 2018
Albayzín 2018 is he la es edi ion o he Albayzín e alua ions, he a emp om Red Temá ica
de Tecnologías del Habla (RTTH) o he e olu ion o speech echnologies in hose languages
spoken in he Ibe ian Peninsula. Rega ding dia iza ion, 2018 is he hi d edi ion a e hose hold
in 2010 and 2016.
Fo he 2018 edi ion he e alua ion consis s o app oxima ely 600 hou s o da a om mass
media domain, co e ing wo di e en languages (Spanish and Ca alan) and wo di e en mass
media (TV and adio). The whole da ase was composed by h ee di e en subse s, acqui ed
along he di e en edi ions: F om 2010 edi ion we ha e a ailable 84 labeled hou s o audio
om 3/24 TV channel in Ca alan. These da a a e complemen ed by 2016 da a: 23 hou s o
manually anno a ed audio om b oadcas adio signal om Co po ación A agonesa de Radio
y Tele isión (CARTV) in Spanish. Finally, 2018 edi ion also adds a ound 400 hou s om
b oadcas con en om Radio Tele isión Española (RTVE). The e alua ion di ides he pool o
da a as ollows: Fo aining and de elopmen bo h 3/24 and CARTV a e a ailable, as well as
10 hou s om RTVE wi h manual anno a ions. E alua ion da a consis s o 40 hou s om RTVE
subse .
3.2.3 Acous ic a iabili y
The ichness o audio con en makes bo h MGB 2015 and Albayzín 2018 a sui able choice
o analyze a iabili y in B oadcas da a, s udying how di e en ac o s a ec he pe o mance
o dia iza ion sys ems. Fo he audio a iabili y we will make use o he dia iza ion sys em
pa adigm, specially he SPLDA model.
Acco ding o he PLDA pa adigm, SPLDA models he a iabili y along i s aining da a by
p ojec ing i in wo subspaces, he in e -speake space de ined by he ma ix VVTand he in a-
speake space desc ibed by ma ix W−1. Mo eo e , due o he ac ha bo h ma ices ep esen
co a iances as well, hei analysis can lead o in e es ing in o ma ion abou he a iabili y om
43
Analysis o b oadcas da a
Da ase VVTSubspace (W−1)Subspace
Telephone
SRE 0.50 0.49
B oadcas
MGB 2015 0.19 0.86
Albayzín 2018 0.49 0.58
Table 3.1: T ace analysis o PLDA in e -speake (VVT) and in a-speake (W−1)
subspaces
each ype, in e -speake and in a-speake .
In Table 3.1 we ca y ou an analysis o bo h in e -speake and in a-speake subspaces in
e ms o hei co a iance ma ices. Fo his analysis we s udy he ace o bo h VVTand W−1
ma ices o ou wo b oadcas da ase s o in e es , MGB 2015 and Albayzín 2018. This s udy
app oxima es he o al a iabili y wi hin each subspace, equi alen o add he a iabili y along
each dimension o he subspace as i hey we e independen . This s udy is complemen ed by a
simila analysis o elephone channel da a, which plays he ole o baseline. This baseline anal-
ysis conside s SRE da a, cons uc ing ou models wi h exce p s om SRE04, SRE05, SRE06
and SRE08.
The esul s illus a ed in Fig. 3.1 show a g ea misma ch in e ms o condi ions be ween
elephone and b oadcas da a. While elephone channel p esen s a simila a iabili y con ained
in bo h subspaces, ou b oadcas da abases show a leas a ound 18% ela i e ex a in a-speake
a iabili y. This measu e inc eases up o a 352% ela i e ex a a iabili y in MGB da ase . Thus,
when conside ing he b oadcas domain, we mus ake in o accoun he ollowing ques ion:
Do simila embeddings sha e he same speake s o jus analogous acous ic condi ions?
Fu he mo e, his in a-speake a iabili y is no only caused by di e ences among shows.
In ac , b oadcas da a p esen s a high wi hin-episode a iabili y. This so o a iabili y co e-
sponds o he di e en condi ions in which he audio is eco ded, e.g. he eco ding loca ion
(s udio, ou doo s, e c.), he in ol ed ma e ial (mic ophones, pos p ocessing, ...), and acous-
ic addi ions (laugh e , applauses, e c.). Besides, he speech signal is highly a ec ed by he
p esence o emo ional speech, i.e. he ans o ma ion o he oice in o de o ansmi ex a
in o ma ion such as shou ing (w a h), whispe ing ( ea ) o whining (pain). Rega dless o he
na u e o he a iabili y, i is usually e y co ela ed along ime, emaining he acous ic cha -
ac e is ics s able du ing pe iods o ime ha can con ain mul iple in e en ions om di e en
44
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
❆✁✂✄☎✆ ✝✆✞✆✟✠✡✆☎☛
✶☞ ✷☞ ✸☞ ✹☞ ✺ ☞ ✻ ☞ ✼☞ ✽ ☞ ✾ ☞ ✶☞☞
✶☞
✷☞
✸☞
✹☞
✺☞
✻☞
✼☞
✽☞
✾☞
✶☞☞
❙
❡
❣
♠
❡
♥
❥
✌✍✎✏✍✑✒ ✓
❙✁✂✄✁☎ ✆☎✝✞✟✠ ✡☎✞☛☞
✶✌ ✷✌ ✸✌ ✹✌ ✺ ✌ ✻ ✌ ✼✌ ✽ ✌ ✾ ✌ ✶✌✌
✶✌
✷✌
✸✌
✹✌
✺✌
✻✌
✼✌
✽✌
✾✌
✶✌✌
✍
❡
❣
♠
❡
♥
❥
✎✏✑✒✏✓✔ ✕
a) Pai wise LLR b) G ound u h mask
Figu e 3.2: Sec ion a iabili y example. Fo 100 i s embeddings om a Sp ingwa ch
episode a) SPLDA pai wise LLR simila i y me ic b) G ound u h ela ionship.
speake s. These pe iods wi h s able condi ions usually co espond o he di e en sec ions o a
show. A clea example could be he news, whe e in e en ions om he news eade s in s udio
condi ions a e in e lea ed wi h ou doo connec ions.
In o de o expose his a iabili y we again make use o ou e e ence dia iza ion sys em,
s udying he AHC simila i y ma ix cons uc ed by means o PLDA pai wise log-likelihood a-
io. This ma ix should con ain highe alues o hose elemen s compa ing embeddings om
he same speake , ega dless o he acous ic condi ions. In Fig. 3.2 we illus a e he acous ic
simila i y ma ix among he 100 i s de ec ed segmen s om an episode o he TV show Sp ing-
wa ch, om MGB 2015. In his analysis we co e an app oxima e 25% o he o al de ec ed
in e en ions in he episode, balancing he adeo be ween gene aliza ion and isualiza ion
capabili ies. The segmen s a e s udied in ch onological o de o imeline comp ehension. The
in o ma ion includes wo pa s, he acous ic simila i y and he g ound u h mask. Fo he acous-
ic simila i y we make use o he PLDA pai wise LLR, whe e he elemen ij e eals how simila
a e he embedding iand j. Ligh e colo s indica e highe speake simila i ies and da ke colo s
less p obabili y o sha e he same speake . Rega ding he g ound u h mask, he ij posi ion
in he igu e is whi e i bo h embeddings, iand j, ha e he same speake label, being black
o he wise.
The wo images shown in Fig. 3.2 e eal he capabili ies do disce n be ween speake and
acous ic condi ions a e limi ed. We i s analyze he speake labels o he las 50 embeddings.
Acco ding o he g ound u h, wo speake s a e esponsible o a sequence o in e lea ed u -
e ances, as in a dialog. Howe e , he LLR sco es did no ealize abou ha , p o iding an
45
Conclusions
3.4 Conclusions
The esul s ob ained along he p esen chap e ha e e ealed di e en ac o s o he inhe en
a iabili y in b oadcas da a. In o de o deal wi h he de ec ed unce ain ies, dia iza ion sys ems
mus wo k in he ollowing elemen s:
3.4.1 The clus e ing app oxima ion
Ou e e ence sys em bases i s dia iza ion choices acco ding o an agglome a i e a chi ec u e.
This a chi ec u e is well-known in he communi y due o i s simplici y. Ne e heless, mo e
e ol ed solu ions could ob ain be e dia iza ion esul s. The choice o an al e na i e clus e ing
p ocedu e equi es ha some conside a ions mus be aken in o accoun .
In i s place we mus ake ca e o he me ic o de e mine he quali y o he pa i ion. Resul s
in Fig. 3.2 illus a e he g ea in luence o channel e ec s in he PLDA LLR. Imp o emen s
abou he modeliza ion o he in a-speake a iabili y should lead o g ea bene i s. Mo eo e ,
we can also wo k in he pa i ion measu emen , combining he local in o ma ion conside ed
in AHC (pai wise simila i y) wi h a mo e gene al poin o iew. Thus, al e na i e clus e ing
app oaches should ake in o conside a ion he implica ions o some o he clus e ing choices in
he decision-making p ocess.
Ano he poin o ocus on is he s op c i e ion. Many o he al e na i e clus e ings simul-
aneously wo k wi h mul iple pa i ions, which con ain a wide ange o speake s. Whene e
compa ing hypo heses, biased measu emen s mus be compensa ed p io o i s compa ison in
o de o p e en signi ican deg ada ions.
3.4.2 The quali y o he embeddings
The obse ed undesi ed a iabili y shown in Fig. 3.2 may no be exclusi ely compensa ed du -
ing clus e ing, bu also du ing he embedding ex ac ion. In ac , he mo e disc imina i e is he
in o ma ion in he embeddings, he be e will be i s pe o mance du ing he clus e ing s age.
Un o una ely, embeddings include mo e a iabili y ha is inhe en o b oadcas da a. We
a e e e ing o mo e gene al a iabili y e ms, such as phone ic a iabili y and sho segmen s.
Any imp o emen in he managemen o hese wo a iabili ies would lead o a gene al imp o e-
men in all ypes o dia iza ion, as well as in speake ecogni ion.
52
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
3.4.3 The domain misma ch p oblem
Finally, we mus also co e he domain misma ch. E en i clus e ing sys ems could co e he
a iabili y in di e en domains, hei pa icula cha ac e is ics should equi e some indi idual
adap a ion o an op imal pe o mance.
Fo his eason, we can make use o he domain adap a ion echniques, i.e. adap he p o-
posed solu ion o each one o he domains o in e es . Ano he solu ion could be he opposi e,
ans o ming he e alua ion audio o i he aining condi ions. Wha e e is he solu ion, we
mus ace ano he issue: b oadcas da a includes se e al shows and gen es, hus in-domain da a
o each o hem may be limi ed o jus una ailable.
53
Conclusions
54
Pa II
The Clus e ing P oblem
55
Chap e 4
Clus e ing by means o Fully Bayesian PLDA
The esul s ob ained in Sec ion 3.3 ha e shown he limi a ions o he baseline dia iza ion sys-
em, specially conce ning he clus e ing s age based on an AHC solu ion. The e o e, his poo
pe o mance mo i a es he sea ch o al e na i e clus e ing op ions. One o he main d awbacks
o AHC is he use o local decisions, i.e. decisions aking in o accoun e y li le in o ma ion, as
he pai wise loglikelihood a ios be ween wo single embeddings. Thus, we would a he p e e
a clus e ing me hod whose me ic e alua es he o e all pa i ion. Ano he eques is ha he
op imiza ion p ocess simul aneously op imizes all labels. A solu ion i ing bo h equi emen s
is he clus e ing by means o Fully Bayesian PLDA.
4.1 The Fully Bayesian PLDA clus e ing solu ion
This p oposal o clus e ing was i s p oposed in [Villalba and Lleida, 2014] as an unsupe ised
clus e ing o model adap a ion. Along he ollowing lines we will de ine he model and ex-
plain how i can be used o clus e ing asks, including dia iza ion. In his p ocess we will pay
a en ion o i s Va ia ional Bayes (VB) decomposi ion, key poin in his app oach.
4.1.1 The Fully Bayesian PLDA (FBPLDA) model
The Fully Bayesian PLDA model [Villalba and Lleida, 2014] is a gene a i e s a is ical model
which desc ibes he inpu embeddings in e ms o la en a iables, some o hem ied along all
embeddings om he same speake . Based on he Simpli ied PLDA, he FBPLDA also desc ibes
he embedding φj om he i h speake as:
φj=µ+Vyi+ǫj(4.1)
57
The Fully Bayesian PLDA clus e ing solu ion
µW V
ε
φjyi
θj
πθ
Ni
I
Figu e 4.1: Bayesian ne wo k o he Fully Bayesian PLDA
whe e µs ands o he speake independen e m. V ep esen s he low-dimension ma ix de in-
ing he speake subspace. yiis he speake la en a iable, s anda d no mal dis ibu ed and
common o all embeddings om he speake i. Finally, he emaining unexplained a iabili y
in he embedding jis included by he e m ǫj, which is modeled by means o a ze o mean
Gaussian wi h co a iance W.
The e olu ion o he Fully Bayesian PLDA is ha , in con as o SPLDA, speake assign-
men s o bo h aining and e alua ion a e unknown, using la en a iables ins ead. Thus, a
se o Nembeddings Φ={φ1, ..., φj, ..., φN}is explained by a se o Icandida e speake s,
each one modeled by a speake la en a iable yi om he se Y={y1, ..., yi, ..., yI}. In
o de o map each embedding o i s gene a o speake , we conside he se o la en a iables
Θ = {θ1, ..., θj, ..., θN}. Each one o he θjla en a iables ollows a mul inomial dis ibu ion,
which p oduces a one-ho sample wi h I alues (θj={θ1j, ..., .θij, ..., θIj}). Each one o hese
alues θij ep esen s he assignmen o he u e ance j o he i h speake . Thus, θjwill ha e i s
componen θij equal o one when he i h speake is esponsible o he embedding j, being ze o
o he wise. Taking his assignmen in o accoun , we can model Φin e ms o Yand Θas:
P(Φ|Y,Θ) =
N
Y
j=1
I
Y
i=1 Nφj|µ+Vyi,W−1θij (4.2)
Due o he Bayesian app oach, he speake labels Θa p io i ollow a mul inomial dis ibu-
ion. This dis ibu ion is complemen ed by i s own p io , πθ, which explains he mul inomial
weigh s acco ding o Di ichle dis ibu ion. Besides, he desc ibed Fully Bayesian PLDA p o-
poses an ex a e olu ion. Ins ead o conside ing poin es ima ions o he model pa ame e s (µ,
Vand W), his e olu ion assumes hem o be la en a iables as well. While he mean µand
he columns o he speake ma ix Va e ea ed wi h a Gaussian p io , Wis modeled in e ms
o a Wisha dis ibu ion. Finally, he model also includes a p io a iable ε o he a iable V.
The Bayesian ne wo k desc ibing he whole model is illus a ed in Fig. 4.1.
58
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
The aining p ocedu e o his model is no nea ly as simple as o he SPLDA. The ain-
ing o he la e model wo ks in e ms o he Expec a ion Maximiza ion (EM) algo i hm. This
algo i hm equi es he es ima ion o he pos e io dis ibu ion o each one o he la en a i-
ables in he model (E s ep), upda ing he poin -es ima ion model pa ame e s o maximize he
loglikelihood (M s ep). Howe e , in he FBPLDA model a closed- o m solu ion o each
pos e io is no possible, hus E s ep canno be pe o med. The e o e, he o iginal wo k
[Villalba and Lleida, 2014] also p oposes an al e na i e aining s a egy by means o he Va ia-
ional Bayes (VB) [A ias, 1999][Bishop, 2006].
Va ia ional Bayes is an app oxima ion me hod ha allows o mimic he EM algo i hm
by a a ia ional equi alen . Gi en a model depending on he se o la en a iables Z=
{Z1, ..., Zh, ..., ZH}, VB app oxima es he pos e io dis ibu ion P(Z|Φ)by a ac o ial dis i-
bu ion q(Z) = QH
h=1 q(Zh). Each one o he ob ained ac o s q(Zh)is an app oxima ion o he
eal pos e io dis ibu ion P(Zh|Φ). In o de o ob ain he bes app oxima ion ollowing he
ac o ial es ic ions each ac o q(Zi)mus ollow a dis ibu ion ollowing he ela ionship:
ln q(Zh) = E∀Z , 6=h[ln P(Z,Φ)] (4.3)
Un o una ely, he app oxima ion by means o a ac o ial dis ibu ion has limi a ions. De-
spi e he ac ha he ob ained ac o dis ibu ions q(Zh)only depend on one o he la en
a iables Zh, hey a e no comple ely independen . Taking in o accoun eq. (4.3), some de-
pendencies emain, being each ac o cons uc ed on op o he expec ed alues om he o he
ac o s.
A colla e al e ec o he Va ia ional Bayes app oxima ion is ha he loglikelihood o he eal
model is no longe a sui able me ic. These ype o solu ions wo ks in e ms o he E idence
Lowe Bound (ELBO o L).
L(Φ) = Zq(Z) ln P(Φ,Z)
q(Z)dZ(4.4)
Bo h ELBO L(Φ)and loglikelihood ln P(Φ)a e in e connec ed. In ac , he loglikelihood
e m is he sum o he ELBO e m plus he KL di e gence be ween he ac o ial dis ibu ion
q(Z)and he eal pos e io dis ibu ion P(Z|Φ). We can exp ess his as:
ln P(Φ) = L(Φ) + KL (q(Z)||P(Z|Φ)) (4.5)
In ac , he ELBO and KL e ms a e in e connec ed. The maximiza ion o ELBO makes KL
di e gence o be educed, be e app oxima ing P(Z|Φ)by means o q(Z). As long as his
app oxima ion is mo e accu a e, ou ELBO e m will be a mo e eliable ep esen a ion o he
loglikelihood om he o iginal model.
59
The Fully Bayesian PLDA clus e ing solu ion
Mo ing o he speci ic case o he Fully Bayesian PLDA, he p oposed decomposi ion o
ac o s is desc ibed as ollows:
P(Y,Θ, πθ,µ,V,W,ε|Φ) = q(Y)q(Θ) q(πΘ)q(µ)q(V)q(W)q(ε)(4.6)
Thus, o aining pu poses we can now pe o m an analogous al e na i e o he EM algo-
i hm, now maximizing ELBO. The a ia ional equi alen o he E s ep mus i e a i ely upda e
he di e en ac o s o ob ain he pos e io dis ibu ions, and he analogous M s ep will p oceed
o he poin es ima ion upda e. This p ocess is epea ed un il con e gence.
4.1.2 The clus e ing p ocedu e
The clus e ing echnique by means o he FBPLDA model p oposed in
[Villalba and Lleida, 2014] has a s a is ical backg ound. This app oach assumes ha di-
a iza ion labels Θdia a e hose ha bes explain he embeddings Φ. Then, he way we should
compa e how pa i ions Θexplain he da a Φis he p obabili y P(Θ|Φ), as desc ibed in
Sec ion 2.6.2.
Θdia = a g max
Θ
P(Θ|Φ) = a g max
ΘZP(Z′,Θ|Φ)dZ′(4.7)
whe e Z′ ep esen s he se o all la en a iables in he model (Y,πθ,µ,V,Wand ε) excep
o Θ.
The applica ion o his app oach o he FBPLDA model is no s aigh o wa d. The same
di icul ies du ing aining a e p esen in his app oach, hus we again mus ely on ou VB
decomposi ion. In consequence, we mus wo k in e ms o app oxima ions as ollows:
Θdia = a g max
ΘZq(Y)q(Θ) q(πθ)q(µ)q(V)q(W)q(ε)dZ′= a g max
Θ
q(Θ) (4.8)
Despi e ha ing simpli ied he clus e ing s ep o he maximiza ion o a single ac o q(Θ),
he same di icul ies as du ing aining emain. The p oposed ac o s a e s ill in e connec ed
by means o expec a ions, so he op imiza ion o he ac o q(Θ) needs o he ac o s o be
op imized as well. Un o una ely, hese o he ac o s also depend on q(Θ). Fo his eason
we wo k in e ms o an i e a i e upda e o ac o s in which, s a ing om an ini ial s a e, we
each a maximum ELBO. This i e a i e p ocess is simila o he conside ed EM p ocedu e
du ing he SPLDA aining. Howe e , in his occasion no poin es ima ion equi es upda e
( hese only a ec he model pa ame e s µ,V,Wand ε), hus we only conside he E s ep.
60
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
Figu e 4.2: Clus e ing schema ic based on label ini ializa ion and FBPLDA eseg-
men a ion
Mo eo e , because we assume he model pa ame e s µ,V,Wand ε o be pe ec ly uned, we
exclusi ely ee alua e q(Y),q(Θ) and q(πΘ). This clus e ing p ocedu e can be in e p e ed as
a wo-s ep sea ch: The se o embeddings Φis dis ibu ed along a se o clus e s Ydu ing he
upda e o he ac o q(θ). Then, he same clus e s Ya e ee alua ed in e ms o he ecen ly
es ima ed Θdu ing he es ima ion o q(Y). This i e a i e p ocess may be easily unde s ood as
ollows: A he beginning o each i e a ion, acco ding o he cu en alue o he speake labels
Θwe cha ac e ize each one o he conside ed Iclus e s. This cha ac e iza ion is done by he
ee alua ion o he speake la en a iables Ydu ing he upda e o he ac o q(Y). Once he
clus e s a e ede ined, each embedding is hen assigned o he mos likely clus e when q(Θ) is
ee alua ed again.
The main disad an age o his app oach is he need o some ini ializa ion Θ0. This ini-
ializa ion can be ob ained in se e al ways, ei he om some p io knowledge o mo e o en
elying on he same embeddings Φ. Hence he p oposed clus e ing s age ollows he schema ic
ep esen ed in Fig. 4.2.
This clus e ing s a egy can be in e p e ed as a wo-s ep clus e ing: A i s block is in cha ge
o ob aining an ini ial pa i ion Θ0, which is e ined a e wa ds by he FBPLDA clus e ing
app oach. Apa om he bene i s due o label eassignmen , his eclus e ing by means o he
FBPLDA o e s ano he ad an age: an es ima ion abou he speake numbe . q(Θ) dis ibu es
he embeddings Φalong Ia p io i candida e speake s. Ne e heless, i is no obliga o y ha
all candida es gene a e a leas one embedding. Those candida e speake s wi hou assigned
embeddings could be elimina ed as pa o he s op c i e ion.
61
Analysis o FBPLDA pe o mance
e s in he ini ial pa i ion Θ0and hose ob ained a e he esegmen a ion. Whene e Θ0con ains
less speake s han he g ound u h (posi i e ela i e speake s), he FBPLDA eclus e ing does
no disca d any single speake , only eassigning he embeddings along he di e en a ailable
candida e speake s. By con as , whene e he ini ializa ion con ains mo e speake s han he
g ound u h numbe , he algo i hm s a s disca ding ew speake s (1-3 speake s on a e age)
al hough he ejec ion o ex a speake s is no enough, signi ican ly o e es ima ing he numbe
o speake s in an audio. In consequence, a bad es ima ion o he speake numbe is di icul o
be ixed by his esegmen a ion.
Apa om he numbe o speake s, dia iza ion is a ec ed by o he ac o s. Ano he impo -
an conside a ion o ake in o accoun is he chance o losing eal speake s. In many occasions
he wo hy speake s a e no hose who speak he mos bu hose wi h ew bu ele an in e en-
ions. An example may be he alk shows, whe e he mos alka i e indi idual is he mode a o
despi e he eal aluable con ibu ions come om he emaining speake s. Rega ding he FB-
PLDA solu ion, ou expe imen al wo k has e ealed ha i is usually eluc an o conside small
se s o embeddings (some imes a single one) as an independen speake , op ing o assuming
hem as spu ious da a om a much la ge clus e . This end o using small clus e s wi h la ge
ones is mo e ele an as long as he balance o da a becomes odde . By con as , when wo
clus e s o simila size p esen audio om he same speake , he algo i hm is unlikely o use
hem oge he .
These limi a ions abou how he VB solu ion handles he esegmen a ion specially a ec s
o low- alka i e speake s. They con ibu e e y li le o he eal audio, bu depending on he
applica ion, hei loss is no a o dable. The e o e, we s udy he numbe o los speake s acco d-
ing o he ini ial pa i ion. The ob ained esul s a e shown in Fig. 4.7. Again he pa i ions a e
iden i ied in e ms o ela i e numbe o speake s ∆I. The esul s a e also shown in e ms o
he i s , second and hi d qua ile.
Acco ding o Fig. 4.7 clea subclus e ing ini ializa ions (we assume up o 20 ex a speake s)
lead o he loss o 2-3 speake s on a e age. This end seems s eady a e 12 ex a speake s in
ou ini ial pa i ion Θ0. This esul is specially in e es ing when compa ed wi h Fig. 4.6, which
shows a g owing o e es ima ion o he speake numbe , p opo ional o hose p esen in he
ini ial pa i ion. The combina ion o bo h sou ces o in o ma ion leads o he conclusion ha we
a e usually losing eal speake s, and mos o he o e es ima ion o speake s is a consequence o
he unde clus e ing o he emaining ones.
68
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✵
✺
✶✵
✶✺
✷✵
✷✺
✸✵
✲
✷✵ ✲
✶ ✲
✶✷ ✲✁ ✲✂ ✵✂ ✁ ✶✷ ✶ ✷ ✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✄ ■
▲
♦
☎
✆
✝
✞
✟
✠
✡
✟
☛
☎
☞✌✍ ✎ ✏✑ ✒✓ ✔✒✕✍ ❢✌ ✕ ❱❇ ✍✌✖ ✉ ✎✗ ✌ ♥
Figu e 4.7: Los speake s acco ding o he ela i e numbe o speake s ∆Iin he
ini ial pa i ion Θ0. Resul s ob ained wi h Albayzín 2018.
4.2.3 Numbe o speake s s DER
Dia iza ion pe o mance does no exclusi ely depend on he in e ed numbe o speake s. Ac-
ually, a good es ima ion abou he numbe o speake s is no always ep esen a i e o a good
dia iza ion in e ms o DER. This is because DER is highly dependen on a p ope iden i ica-
ion o he la ges clus e s in he analysis audio. Thus, as long as eally alka i e speake s a e
p ope ly clus e ed, any ea men o non- alka i e speake s may be conside ed bene icial, e en
hei disca d.
Ou nex expe imen s udies he ela ionship be ween he ini ializa ion and he DER pe o -
mance measu e. Fo his expe imen we ha e e alua ed he in e ed pa i ions ob ained om
he FBPLDA eclus e ing o mul iple le els o he AHC dend og am. The ob ained esul s a e
illus a ed in Fig. 4.8, ep esen ing he ob ained DER in e ms o he ela i e numbe o speake s
in he ini ializa ion Θ0. The ep esen ed in o ma ion simul aneously analyzes he ini ializa ion
AHC sys em (Fig. 4.8a) as well as he AHC block ollowed by he eclus e ing s age (Fig. 4.8b).
The esul s in Fig. 4.8 show ha any unde es ima ion abou he numbe o speake s is e y
ha m ul o bo h dia iza ions, AHC and he FBPLDA eclus e ing. This deg ada ion is mo e
se e e as long as he unde es ima ion inc eases. These esul s a e easonable om he DER
pe spec i e, because se e e losses o speake s will de ini ely cause he misclassi ica ion o e y
69
Analysis o FBPLDA pe o mance
✵
✶✵
✷✵
✸✵
✹✵
✺✵
✻✵
✼✵
✽✵
✲✷✵✲
✶✻ ✲
✶
✷ ✲✽ ✲
✹ ✵ ✹ ✽✶✷✶✻ ✷✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✁■
❉
❊
✭
✪
✮
✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✠♦✡✉☛☞♦♥
✵
✶✵
✷✵
✸✵
✹✵
✺✵
✻✵
✼✵
✽✵
✲✷✵✲
✶✻ ✲
✶
✷ ✲✽ ✲
✹ ✵ ✹ ✽✶✷✶✻ ✷✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✁■
❉
❊
✭
✪
✮
✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✰ ❱❇ ✠♦✡✉☛☞♦♥
a) AHC esul s b) AHC + FBPLDA esul s
Figu e 4.8: DER (%) esul s o a) AHC and b) FBPLDA in e ms o he ela i e
numbe o speake s ∆I. Resul s ob ained om Albayzín 2018, indica ing he i s ,
second and hi d qua ile pe bin.
alka i e speake s. In e es ingly, acco ding o Fig. 4.8 when an o e es ima ion abou he num-
be o speake s should happen o occu , he end o DER is no so deg aded, wi h small loses
o pe o mance in AHC and no no iceable deg ada ion when he FBPLDA eclus e ing is ap-
plied. In eal li e applica ions his o e es ima ion scena io may be wo hy enough, specially
conside ing semi-supe ised applica ions. By means o au oma ic echniques dia iza ion labels
wi h mul iple pu e clus e s pe speake could be easily ob ained, only equi ing li le human
supe ision o ma ch hose clus e s wi h a common speake . This manual wo k is simple han
cleaning clus e s wi h mul iple speake s, i.e. he scena io in which an unde es ima ion o he
speake numbe is done.
An al e na i e analysis is he sea ch o he pa i ion ha p o ides he bes dia iza ion esul .
Fo his analysis each episode has been dia ized wi h mul iple ini ial pa i ions Θ0, all ob ained
om he AHC dend og am. The esul s, shown in Fig. 4.9, a e compa ed in e ms o he ela i e
speake o he ini ializa ion wi h espec o he g ound u h.
Acco ding o Fig. 4.9, dia iza ion esul s end o p e e an o e es ima ion o he numbe
o speake s, in e ing mo e speake s han hose p esen in he e e ence labels. Mo eo e , he
o e es ima ion can be signi ica i e, wi h many episodes (mo e han 90% o he episodes) wi h
a leas 5 ex a speake s ob aining he bes DER esul s. These esul s i wi h hose p e iously
ob ained in Fig. 4.8, which showed ha an unde es ima ion o he speake s would lead o sig-
ni ican deg ada ions.
The choice o he bes ini ializa ion is a g ea challenge. The choice o he ini ial pa i ion
leading o he minimum DER may be di icul o e en impossible. E en i his op ion was easi-
70
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✵
✷✵
✸✵
✹✵
✁✵
✻✵
✼✵
❘❛♥❦ ♦❢ ❘❡ ❧ ❛ ✐✈❡ ❙♣ ❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎ ③✑✏✎✓ ✍ ✔✓✕ ❇✖✗✏ ❉❊✘
Figu e 4.9: Dis ibu ion o he ini ializa ion wi h bes DER in e ms o he ela i e
numbe o speake s. Resul s ob ained om Albayzín 2018 and p esen ed acco ding o
5 bins.
ble, i can imply such an una o dable compu a ional cos . The e o e, we migh some imes seek
a adeo , assuming ce ain deg ada ions in pe o mance i a la ge simpli ica ion o he sys ems
is achie ed. In Fig. 4.10 we analyze he p opo ion o ini ial pa i ions whose dia iza ion ou pu
di e s om he bes esul up o a maximum bound. This igu e is composed o wo di e en
dis ibu ions, Fig. 4.10a showing he chance o a 1% DER bound and 4.10b illus a ing he
dis ibu ion o a 3% DER bound.
Fig. 4.10 illus a es ha he p obabili y o a pa i ion o be unde a DER deg ada ion bound
is e y educed. Only an app oxima e 17% o he ini ializa ions each unde he 1% DER bound,
being o e 33% wi h a highe bound (3% DER).
4.2.4 Numbe o speake s s ELBO
In hese lines we wan o explo e ma ke s o de e mine whe he we a e wo king wi h a good
pa i ion. E en i a closed se o pa i ions is p o ided, e.g. he mul iple le els o he AHC
dend og am, we need an unsupe ised op ion o compa e hem and decide which one is ou
bes op ion.
Because we a e aking in o accoun a s a is ical solu ion, a ai op ion should be he loglike-
71
Al e na i e ini ializa ions
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✁
✵
✁
✷✵
✷✁
✸✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐ ✈❡ ❙♣❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✶✙ ❇✓✚✍✛
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✁
✵
✁
✷✵
✷✁
✸✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐ ✈❡ ❙♣❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✙✚ ❇✓✛✍✜
1% Bound 3% Bound
Figu e 4.10: Dis ibu ion o he ini ializa ion wi h bounded DER, a) 1% and b) 3%,
in e ms o he ela i e numbe o speake s. Resul s ob ained om Albayzín 2018 and
p esen ed acco ding o 5 bins.
lihood. Ac ually, because ou model is sol ed by means o a Va ia ional Bayes app oxima ion,
we should conside ELBO ins ead. Conside ed as an app oxima ion o loglikelihood, ELBO is
s ill a good ep esen a i e numbe abou how well he pa i ion ep esen s he inpu da a Φ.
In he nex expe imen we s udy he eliabili y o ELBO as pa i ion selec ion c i e ion,
explo ing which ini ializa ion ob ains he bes ELBO a e FBPLDA eclus e ing is pe o med.
This esul will indica e us which esul s a e mo e s a is ically eliable. Due o ange issues, we
ep esen ou esul s in Fig. 4.11 as a his og am illus a ing he dis ibu ion o maximum ELBO
in e ms o he ela i e speake numbe o he ini ializa ion Θ0.
Acco ding o he esul s in Fig. 4.11, ELBO is a g ea indica o abou he numbe o speake s,
op ing o small de ia ions (±5 speake s) wi h espec he g ound u h in almos 50% o he
in ol ed da a. Howe e , ano he 25% o pa i ions op imize ELBO by o e clus e ing up o 15
speake s. This is specially undesi able when conside ing Fig. 4.10, which eques s he opposi e
(unde clus e ing) o a be e dia iza ion pe o mance.
4.3 Al e na i e ini ializa ions
In he p e ious lines we ha e s udied he g ea impac o he ini ializa ion on he pe o mance
in he FBPLDA eclus e ing solu ion. I s in luence ex ends o mul iple ac o s such as o e -
all quali y, he es ima ed numbe o speake s, how many speake s we may lose o he o e all
ELBO.
72
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✵
✷✵
✸✵
✹✵
✁✵
✻✵
✼✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐✈❡ ❙♣ ❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
❊▲❇❖ ✌❝❝✍✎✏✑✒❣ ✓✍ ✔✒✑✓✑✌✕✑③✌✓✑✍✒
Figu e 4.11: Dis ibu ion o he pa i ion wi h bes ELBO in e ms o ela i e speake s.
Resul s ob ained wi h Albayzín 2018 and ep esen ed acco ding o 5 bins.
All his acqui ed knowledge was ex ac ed o be e deal wi h he ini ializa ion issue. We
expec o ind an al e na i e ini ializa ion app oach wi h espec o ou i s FBPLDA app oach
(Table 4.1), whe e a h eshold de e mines he le el o he dend og am in he AHC, e ined
a e wa ds by he FBPLDA.
Acco ding o he al eady seen in o ma ion, we illus a e wo di e en app oaches. The i s
one seeks an e icien adeo be ween DER imp o emen and compu a ional cos . The second
op ion ies o each he bes possible esul s despi e alling in o mo e elabo a ed s a egies wi h
a highe compu a ional cos .
4.3.1 Compu a ionally e icien ini ializa ion
The compu a ionally e icien ini ializa ion app oach was de eloped acco ding o he in o ma-
ion in Fig. 4.10. The illus a ed in o ma ion e eals ha hose ini ial pa i ions whose eseg-
men a ion di e s om he bes esul below a bound a e p one o con ain signi ican ly mo e
speake s han he uned ini ializa ion.
Ou i s al e na i e simply p oposes assuming an ini ializa ion whose numbe o speake s is
gua an eed o o e come he g ound u h, hus exploi ing his ci cums ance. By doing his, we
assume an ini ializa ion in which he AHC algo i hm is mo e unlikely o ha e made signi ican
e o s, and exploi he s op c i e ia om he FBPLDA algo i hm. The g ea bene i o his
app oach is ha is compu a ionally e icien , equi ing as much ime as ou FBPLDA baseline.
In Table 4.4 we analyze he impac o his app oach conside ing di e en alues o he
uppe numbe o speake s. The expe imen includes bo h de elopmen and es subse s om
73
Al e na i e ini ializa ions
Expe imen De . DER(%) E al. DER(%)
MGB 2015
50 Speake s 26.08 42.22
75 Speake s 26.13 41.37
100 Speake s 26.55 42.25
200 Speake s 29.68 45.93
300 Speake s 29.99 44.95
Fines pa i ion 29.24 44.76
Albayzín 2018
50 Speake s 16.48 18.36
100 Speake s 22.60 25.38
150 Speake s 25.94 28.14
Fines pa i ion 60.48 62.54
Table 4.4: DER (%) esul s om AHC ini ializa ion wi h a maximum numbe o
speake s. Included mul iple maxima and he ines pa i ion wi h one segmen pe
clus e . Resul s shown o de elopmen and es subse s om MGB 2015 and Albayzín
2018 co po a.
bo h MGB 2015 and Albayzín 2018. Apa om ixed numbe o speake s o all episodes,
ou esul s also include he case o he ines AHC pa i ion, i.e. one embedding pe candida e
speake , as limi case.
The esul s in Table 4.4 show ha a es ic ed numbe o speake s in bo h da ase s may sig-
ni ican ly imp o e he esul s, specially compa ed wi h Table 4.1. This may be a consequence o
he di e en audio cha ac e is ics among shows, making h esholds on op o simila i y me ics
inapp opia e. Howe e , i is impo an o no ice ha no all ini ializa ions wi h a highe numbe
a e equally use ul. As long as he alue becomes highe , e o a ios s a appea ing. This deg a-
da ion can be aken o he limi conside ing he ines ini ializa ion. The e o e, his app oach
needs o be highe han he g ound u h alue and no oo high o alling in o deg ading he
pe o mance. While in ou app oach we assume he same alue o all episodes, hei du a ion
is no he same. Thus, mo e elabo a ed al e na i es may adjus his alue acco ding o he audio
du a ion.
4.3.2 ELBO-based ini ializa ion choice c i e ion
The s a is ical na u e o he FBPLDA clus e ing solu ion makes easonable he use o al e na i e
ini ializa ion choices. F om a s a is ical poin o iew, he bes pa i ion ΘDIAR should be he
one ha bes explains he gi en da a Φ, as shown in eq. 2.28. The adi ional way o measu e
74
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
Figu e 4.12: Schema ic o dia iza ion based on he simul aneous e alua ion o K
di e en ini ializa ions. The inal pa i ion is selec ed by means o PELBO.
how well some model ep esen s he gi en da a is by means o he pos e io loglikelihood.
When adap ing his app oach o he FBPLDA model we mus deal wi h he VB na u e o
ou solu ion. This solu ion subs i u es he o iginal likelihood e m by he ELBO e m. Thus,
ollowing a simila app oach ou bes dia iza ion pa i ion should be he one ha maximizes he
ELBO L e m as ollows:
ΘDIAR = a g max
ΘL(Θ,Φ)(4.9)
Ne e heless, his idea is no comple e ye . When conside ing mul iple ini ializa ions, no
all o hem ini ially suppose he same numbe o speake s. Hence he highe he numbe o
speake s, he mo e likely his da a can o e i o he e alua ion da a. The e o e, as p esen ed
in [Viñals e al., 2018a], a penalized ELBO (PELBO) e m is p oposed ins ead. This app oach
is inspi ed in BIC, in whe e he likelihood o he model is penalized in e ms o he modelling
capabili ies. Then, he bes labels can be ob ained as:
ΘDIAR = a g max
Θ
PELBO(Θ,Φ) = a g max
Θ
(L(Θ,Φ)−λQ(Θ)) (4.10)
whe e Q(Θ) ep esen s he conside ed excess o modeling capabili ies due o he o al amoun
o speake s in he pa i ion Θ. This e m is mul iplied by a ine uning pa ame e λ.
This pa i ion choice me hod can be applied igh a e he AHC algo i hm, in o de o
choose a single pa i ion. Howe e , due o he g ea e ining p ope ies o he FBPLDA eclus-
e ing, we op ed o choosing among eclus e ed pa i ions. The e o e, a se o Kdi e en ini-
ializa ions a e simul aneously eclus e ed, choosing among hem he inal pa i ion a e wa ds.
This clus e ing s uc u e is ep esen ed in Fig. 4.12.
In Table 4.5 we illus a e he esul s ob ained wi h his algo i hm. The esul s include DER
ma ks o bo h MGB 2015 and Albayzín 2018 da ase s, including bo h de elopmen and es
subse s. Two di e en esul s a e included, a single esul exclusi ely in e ms o ELBO, and
he penalized ELBO as well.
75
Conclusions
Expe imen De . DER(%) E al. DER(%)
MGB 2015
ELBO 26.82 39.12
PELBO 25.95 39.88
Albayzín 2018
ELBO 14.48 17.77
PELBO 13.90 17.79
Table 4.5: DER (%) esul s o ELBO and PELBO ini ializa ion choice. Resul s
shown o de elopmen and es subse s om bo h MGB 2015 and Albayzín 2018
co po a
The ob ained esul s a e signi ican ly be e han ou baseline sys em, and also o e comes
he p e iously desc ibed e icien solu ion. Mo eo e , bo h ELBO and penalized ELBO show
e y simila esul s, illus a ing he obus ness o he app oach. Un o una ely, while penal-
ized ELBO helps o imp o e simple ELBO du ing aining, in e alua ion sligh ly deg ades in
pe o mance, pa ially illus a ing he high domain misma ch be ween shows.
4.4 Conclusions
In his chap e we ha e explo ed he addi ion o he FBPLDA eclus e ing o he baseline di-
a iza ion sys em desc ibed in Sec ion 3.1. Acco ding o he ob ained esul s, he pe o mance
in bo h b oadcas dia iza ion da ase s has been signi ican ly imp o ed due o his block. These
imp o emen s a ec bo h he dia iza ion me ic DER as well as he es ima ion o he speake
numbe .
Mo eo e , we ha e explo ed some o he FBPLDA limi a ions. Ou analysis has e ealed
a g ea dependence o pe o mance acco ding o he ini ializa ion. Thus, he be e he ini ial-
iza ion, he be e is he label e inemen by means o FBPLDA. Besides, ou s udy also e eals
ha ini ializa ions should be e o e es ima e he numbe o speake s (a ound 10 ex a speake s
compa ed o he o acle alue) so ha FBPLDA p o ided he bes labels. Howe e , despi e he
imp o emen s ob ained in his o e es ima ion scena io, we s ill mus assume deg ada ions as
he loss o low alka i e speake s. Fu he mo e, ou analysis has de e mined ha he E idence
Lowe Bound (ELBO) seems a easonable indica o o in e he numbe o eal speake s wi hin
an audio.
Finally, we ha e explo ed he join collabo a ion o he AHC ini ializa ion and he FBPLDA
eclus e ing a he han assuming hem as independen blocks. Thus, we p oposed wo success-
76
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
ul app oaches whe e a limi ed numbe o le els o he AHC dend og am a e e ined by he
FBPLDA, which is also esponsible o he s op c i e ia. Whils ou i s app oach explo ed an
e icien sea ch by assuming a single ini ial pa i ion eassu ing an o e es ima ion o i s speake
numbe , ou second s a egy ca ies ou a simul aneous eclus e ing o mul iple ini ializa ions,
op ing o one o he ob ained pa i ions acco ding o ELBO. Bo h o hem ha e demons a ed
ha FBPLDA eclus e ing is a mo e powe ul s op c i e ia han AHC, ob aining signi ican im-
p o emen s. Wi h espec o he compa ison be ween hem, he ELBO s op c i e ia is able o
ou pe o m he o e es ima ion s a egy, a he cos o inc easing he compu a ional cos s.
77
PLDA wi h Unce ain y P opaga ion (PLDAUP)
Expe imen EER (%) minDCF
Long-Long 3.37 0.161
Long-Sho 5.98 0.291
Sho -Long 5.98 0.283
Sho -Sho 8.76 0.403
Table 5.1: EER (%) and minDCF esul s wi h SPLDA o SRE10 co eex -co eex
de 5 emale wi h in ol ed sho u e ances
Sho ) o each ole (en ollmen and es ). In Table 5.1 we ep esen he pe o mance o each
o hese combina ions. The pe o mance is e alua ed acco ding o wo di e en me ics, Equal
E o Ra e (EER) and minimum De ec ion Cos Func ion (minDCF).
Acco ding o he ob ained esul s, we obse e a se e e deg ada ion in pe o mance as long
as sho u e ances a e conside ed. These deg ada ions a e no iceable when sho u e ances
a e in ol ed, ega dless o hei ole. I bo h oles, en ollmen and es , a e played by sho
u e ances he deg ada ion is much mo e signi ica i e. Fu he mo e, he seen deg ada ion is no-
iceable in bo h me ics, EER and minDCF. This in o ma ion is complemen ed by DET cu es,
shown in Fig. 5.2.
DET cu es in Fig. 5.2 con i m hose esul s ob ained in Table 5.1, showing clea di e ences
in pe o mance when sho u e ances play any ole in e i ica ion. Besides, his deg ada ion
is mo e no iceable when sho u e ances simul aneously play he en ollmen and es oles.
Finally, his beha iou is consis en along he whole cu es, ega dless o he ope a ion poin .
The p e ious expe imen illus a es he impac o sho u e ances in s a e-o - he-a ech-
nologies and jus i ies he sea ch o al e na i e echniques as PLDAUP. This app oach is e alu-
a ed in ou nex expe imen , sco ing he same ials wi h long and sho u e ances. Howe e , in
his expe imen we es ic ou e alua ion o wo o he p e ious scena ios: Long-Sho (o iginal
long u e ance as en ollmen and sho u e ance as es ) and Sho -Sho (bo h en ollmen and
es a e sho u e ances). These wo condi ions a e e alua ed wi h ou new PLDAUP model,
unde going wo di e en al e na i es o mimic leng h-no maliza ion: scala no maliza ion and
unscen ans o ma ion. The ob ained esul s a e shown in Table 5.2, including bo h EER and
minDCF esul s.
Those esul s in Table 5.2 illus a e he bene i s due o he inclusion o Unce ain y P opa-
ga ion. Howe e , he ob ained imp o emen s do no a ec he pe o mance in he same way.
While EER is clea ly mo e a ec ed (a leas 12% ela i e imp o emen s), minDCF bene i s
a e mo e negligible. Besides, bene i s a e ob ained ega dless o he leng h-no maliza ion ap-
p oxima ion o ma ices, al hough minimum ex a imp o emen s a e ob ained wi h unscen
84
Chap e 5. Unce ain y P opaga ion o Dia iza ion
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠✆✝✞✟
✆✝✞✟✠✡☛✝☞✌
✡☛✝☞✌✠✆✝✞✟
✡☛✝☞✌✠✡☛✝☞✌
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✍✎ ✏❡ ❆✎ ✍✑♠ ♣✑✒✓ ✍✓✔✎ ✔ ✕✖ ✗✘✙
❉❊❚ ❝✉✚✈✛✜ ❢✢✚ ◆■❙❚ ❙❘❊✶ ✣
Figu e 5.2: DET cu es wi h SPLDA o SRE10 co ex -co eex de 5 emale wi h
in ol ed sho u e ances
Expe imen Long-Sho Sho -Sho
EER (%) minDCF EER (%) minDCF
SPLDA 5.98 0.291 8.76 0.403
PLDAUP scala 5.11 0.261 7.72 0.389
PLDAUP unscen 5.08 0.269 7.67 0.385
Table 5.2: EER (%) and minDCF esul s wi h PLDAUP o SRE10 co eex -co eex
de 5 emale wi h in ol ed sho u e ances
85
PLDA wi h Unce ain y P opaga ion (PLDAUP)
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠
✝✞✟✠✡
✞☛✝
✝✞✟✠☛☞✌✍✎☞✏☛✝
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜
❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠
✝✞✟✠✡
✞☛✝
✝✞✟✠☛☞✌✍✎☞✏☛✝
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜
❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧
a) Long-Sho expe imen b) Sho -Sho Expe imen
Figu e 5.3: DET cu es wi h PLDAUP o SRE10 co ex -co eex de 5 emale wi h
in ol ed sho u e ances
ans o ma ions. These conside a ions can also be obse ed in Fig. 5.3, whe e DET cu es a e
shown.
DET cu es explain he di e ences be ween minDCF and EER. PLDAUP seems o wo k
simila ly o SPLDA in hose egions o he DET cu e highly penalizing missing a ge ials.
By con as , in hose egions o he cu e o high alse ala m is whe e PLDAUP ob ains he
highes imp o emen s.
5.2.2 PLDAUP in speake clus e ing
The inclusion o he unce ain y p opaga ion in he PLDA o speake e i ica ion has led o
small imp o emen s, despi e no being such a e olu ion. Howe e , dia iza ion is a sligh ly
di e en ask. Ra he han making independen decisions, dia iza ion is he esul o a la ge se
o choices depending on each o he . Thus, small indi idual imp o emen s may esul in o an
accumula ion o bene i s.
This is why we wan o e alua e he unce ain y p opaga ion capabili ies in ou clus e ing
s ep, he s age whe e we can easily in eg a e ou PLDAUP model. As a i s app oxima ion we
do no wo k in dia iza ion ye bu in a speake clus e ing ask, wo king wi h sho u e ances.
Fo his eason, we emain wo king in he elephone channel domain, using he same chopped
subse p e iously used in speake e i ica ion. The o al amoun o in ol ed audios is 2740
u e ances, con aining 232 di e en speake s.
86
Chap e 5. Unce ain y P opaga ion o Dia iza ion
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁
✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
❙✂✄☎✆ ✝✞
✂✄☎✆P✄✟✂ ✝✞
✂✄☎✆✠✡☛☞✌✡✍
✟✂ ✝✞
❙✂✄☎✆ ❙✞
✂✄☎✆P✄✟✂ ❙✞
✂✄☎✆✠✡☛☞✌✡✍
✟✂ ❙✞
◆
✠✎✏
✌✑ ✒✓ ☛✔
✌✕✖✌
✑☛
❈❧ ✉s ❡ s
■
♠
♣
✗
✘
✐
✙
②
✭
✪
✮
✚✛✜✢✣✤✥✦ ♦❢ ✧★▲❉❆ ✈✩ ★▲ ❉❆❯★
Figu e 5.4: Impu i y esul s o SPLDA and PLDAUP in SRE10 co eex -co eex de 5
emale chopped
Ou expe imen al se up is he same as in ou p e ious speake e i ica ion expe imen , i.e. a
GMM-UBM i- ec o ex ac o ollowed by a PLDA model. Howe e , his ime sco es a e no
conside ed o 1 s1 ial decisions bu he me ic o an AHC solu ion, which de e mines he
inal labels. In o de o p o ide a be e o e iew abou he po en ial o PLDAUP, we p e e no
using any s op c i e ion, analyzing mul iple le els o he AHC dend og am. The esul s will
be measu ed in e ms o speake and clus e impu i ies (SI and CI espec i ely) and shown in
Fig. 5.4. Thick lines ep esen clus e impu i ies and dashed lines speake impu i ies. The anal-
ysis in ol es 3 di e en PLDA e sions: adi ional SPLDA is shown in blue, PLDAUP wi h
scala no maliza ion o he unce ain y ma ix is shown in ed and g een ep esen s PLDAUP
wi h no maliza ion o he unce ain y ma ix by means o an unscen ans o ma ion. Ou ep e-
sen a ion also includes an ex a line (black) indica ing he ue alue o speake s in his subse .
The esul s in Fig. 5.4 show a g ea bene i when unce ain y p opaga ion is applied, con-
i ming ou hypo hesis o imp o emen accumula ion. Fo he ange o s udy PLDAUP clus e
impu i y consis en ly unde goes an absolu e imp o emen wi hin he ange 5-10% wi h espec
o SPLDA, while no e iden deg ada ions in he speake impu i y a e no iced. Besides, his
imp o emen has also educed he bias o he Equal Impu i y (EI) poin in e ms o he numbe
o speake s. While SPLDA eaches he EI poin a 350 speake s, PLDAUP does he same a 300
87
PLDA wi h Unce ain y P opaga ion (PLDAUP)
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁
✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
▲✂✄☎ ✆✝
❙✞✂✟✠ ✆✝
▲✂✄☎ ❙✝
❙✞✂✟✠ ❙✝
◆✡☛☞✌✟ ✂
✍ ✎✏✌✑✒✌
✟✎
❈❧✉s ❡ s
■
♠
♣
✓
✔
✐
✕
②
✭
✪
✮
✖✗✘✙✚✛✜✢ ♦❢ ✣P✤❉❆ ✇✛✜ ❤ ✤♦♥❣ ✈✥ ✣❤♦✚✜ ❯✜✜✦✚ ❛ ♥❝✦✥
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁
✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
▲✂✄☎ ✆✝
❙✞✂✟✠ ✆✝
▲✂✄☎ ❙✝
❙✞✂✟✠ ❙✝
◆✡☛☞✌✟ ✂
✍ ✎✏✌✑✒✌
✟✎
❈❧✉s ❡ s
■
♠
♣
✓
✔
✐
✕
②
✭
✪
✮
✖✗✘✙✚✛✜✢ ♦❢ P✣❉❆❯P ✇✛✜❤ ✣♦♥❣ ✈✤ ✥❤♦✚✜ ❯✜✜✦✚❛♥❝✦✤
a) SPLDA b) PLDAUP
Figu e 5.5: Impu i y esul s o a) SPLDA and b) PLDAUP wi h scala no maliza ion
in SRE10 co eex -co eex de 5 emale chopped aining wi h sho u e ances
speake s. Fu he mo e, while speake e i ica ion esul s indica ed ha unscen ans o ma ions
we e be e han scala no maliza ion o he unce ain y ma ix, hose ob ained o he clus e ing
ask show ha he scala no maliza ion o e comes he pe o mance o unscen ans o ma ions
up o an absolu e 2-5%.
Du ing he p e ious expe imen s we analyzed some PLDA model ained on exce p s om
SRE04, 05, 06 and 08. This aining pool consis s o audios wi h la ge amoun s o speech
pe u e ance. Thus, aining embeddings can be conside ed eliable enough o ou s anda ds.
Howe e , his scena io may no be so ealis ic in o he domains, as dia iza ion. In ac , dia iza-
ion da a usually consis s o a combina ion o long and sho u e ances. Thus, we mus analyze
how ou wo models, SPLDA and PLDAUP, beha e when sho u e ances a e conside ed o
model aining.
Fo his pu pose, we analyze he impac o sho u e ances on he aining pool. In his
expe imen we build an al e na i e e sion o he conside ed aining pool (SRE04, SRE05,
SRE06 and SRE08) by andomly chopping he o iginal u e ances gua an eeing he speech con-
en o be wi hin he ange o 3-60 seconds. This subse will only be conside ed o he aining
o he PLDA models, bo h SPLDA and PLDAUP. Unde hese condi ions we e alua e ou subse
wi h sho u e ances wi h bo h PLDA models, SPLDA and PLDAUP wi h scala no maliza ion
o he unce ain y ma ix. In Fig. 5.5 we compa e how each model esponds depending on he
aining coho , di ing be ween SPLDA (Fig. 5.5a) and PLDAUP (Fig. 5.5b).
The illus a ed esul s in Fig. 5.5 e eal in e es ing de ails. Fi s , SPLDA seems o adap
well o sho u e ances, ou pe o ming he e sion wi h long u e ances o any ope a ional
88
Chap e 5. Unce ain y P opaga ion o Dia iza ion
poin . An explana ion is ha he aining coho su e s om shi s o he embeddings due o i s
leng h, simila o hose in he e alua ion subse . Consequen ly, hese shi s can be conside ed
as ex a in a-speake a iabili y and a e aken in o accoun in he co esponding pa ame e
(W). By con as , PLDAUP seems o sligh ly lose some o i s pe o mance. A eason o his
beha iou is ha in a-speake a iabili y lies in a subspace con olled by Ujand W. When
long eliable u e ances a e used o ain he model mos o his a iabili y is o ced o be in he
Wsubpace, ac ing UjUT
jas an addi ion du ing e alua ion. Howe e , when aining in ol es
sho u e ances bo h e ms a e ep esen a i e and con ibu ing, hus making decisions much
noisie .
5.3 Fully Bayesian P obabilis ic Linea Disc iminan Analy-
sis wi h Unce ain y P opaga ion (FBPLDAUP)
The con i ma ion o Unce ain y P opaga ion bene icial capabili ies mo i a es i s e alua ion in
b oadcas dia iza ion. Howe e , in his domain ou bes esul s so a ha e been shown in
Sec ion 4.1 by means o he FBPLDA and i s Va ia ional Bayes esegmen a ion. This eseg-
men a ion is in ac he key poin o his bes app oach, ixing some o he mis akes and hus
imp o ing he pe o mance.
Fo his pu pose, we upda e he FBPLDA model desc ibed in Sec ion 4.1.1 including
he new unce ain y p opaga ion concep . The name o his new model is Fully Bayesian
PLDA wi h Unce ain y P opaga ion (FBPLDAUP). This new model, as well as i s p ede-
cesso , explains a se o Nembeddings Φ om Idi e en speake s modeled by he se
Y={y1, ..., yi, ..., yI}, whe e yiis a la en a iable common o all u e ances om he same
i h speake . Addi ionally, i also inco po a es an ex a la en a iable xij pe u e ance, esponsi-
ble o modeling he a iabili y due o he u e ance leng h. The assignmen o each embedding
o i s esponsible speake is done in e ms o θij, a la en a iable aking he alue o 1i he
elemen jis gene a ed by he i h speake and 0 o he wise. Thus, we de ine he condi ional
dis ibu ion o Φas:
P(Φ|Y,Θ,X,µ,V,W) =
I
Y
i=1
N
Y
j=1 Nφj|µ+Vyi+Ujxj,W−1θij (5.11)
whe e Nφj|µ+Vyi+Ujxij,W−1 ep esen s he dis ibu ion o he embedding φjacco d-
ing o speake i. This modeliza ion includes a speake independen e m µ, a speake dependen
e m Vyiand he i- ec o a iabili y e m Ujxij.Vis a low ank ma ix explaining he speake
subspace and yiis he speake la en a iable. Ujs ands o he i- ec o a iabili y subspace
89
FBPLDA wi h Unce ain y P opaga ion (FBPLDAUP)
µ
W
V
φjyi
Uj
xij
θij
πθ
ε
Ni
I
Figu e 5.6: Bayesian ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opa-
ga ion
ull ank ma ix and xij i s la en a iable. Finally Wis a ull ank ma ix explaining he wi hin
speake subspace.
Due o he ac ha we a e building a Fully Bayesian solu ion, ou model pa ame e s (µ,V
and W) a e dis ibu ions a he han poin es ima es. In ac , Vhas i s own p io dis ibu ion
ε, a p oduc o gamma dis ibu ions. In addi ion o he model pa ame e s, he speake labels
Θa e also ea ed as la en a iables, modeled by means o a mul inomial dis ibu ion. This
mul inomial dis ibu ion includes a Di ichle p io πθin o de o explain he p obabili ies pe
class. The Bayesian ne wo k o he model is shown in Fig. 5.6.
Fo dia iza ion pu poses wi h he FBPLDAUP we ollow he s a is ical app oach al eady
desc ibed in Sec ion 2.6.2 maximizing he pos e io dis ibu ion P(Θ|Φ). Howe e , he com-
plexi y o he model makes he ue pos e io in ac able o op imiza ion pu poses. Hence, we
p e e applying Va ia ional Bayes o a mo e sui able solu ion. The applied simpli ica ion o he
new model is:
PY,X,Θ, πθ,˜
V,W,ε=q(Y,X)q(Θ) q(πθ), q ˜
Vq(W)q(ε)(5.12)
Fo mo e in o ma ion abou he o mula ion o he di e en p io s he o mula ion is in-
cluded in Appendix A.
Due o he ela ionship be ween FBPLDA and FBPLDAUP, bo h ollow he dia iza ion s a -
egy desc ibed in Sec ion 4.1.1, based on a Va ia ional Bayes app oxima ion. Ou new model
assumes ha dia iza ion labels Θdia should be hose which bes explain he gi en u e ances,
90
Chap e 5. Unce ain y P opaga ion o Dia iza ion
modeled by i s mean and co a iance. Thus, we mus maximize P(Θ|Φ). By Va ia ional Bayes
we app oxima e his pos e io dis ibu ion by q(Θ), whose maximum will be ou solu ion. Un-
o una ely, VB ac o s a e in e connec ed, equi ing q(Θ) he o he ac o s o be op imized as
well o a p ope solu ion, which also need q(Θ) adjus ed as well. The e o e, we mus apply an
i e a i e ee alua ion o ac o s (q(Y,X),q(Θ) and q(πθ)) in o de o each he bes alue o
ou hypo hesis.
5.4 Dia iza ion o b oadcas da a wi h FBPLDAUP
In he ollowing lines we p esen he esul s in b oadcas dia iza ion when Unce ain y P op-
aga ion is included. Fo hese expe imen s we make use o MGB 2015 acco ding o he con-
igu a ion explained in Sec ion 3.2.1. The applied sys em ollows he schema ic p esen ed in
Sec ion 4.1.2, whe e an AHC s age o es ima e pa i ion seeds o he VB esegmen a ion. In
ou new app oach, any in ol ed PLDA is subs i u ed by i s PLDAUP coun e pa , so ou AHC
s age wo ks in e ms o he PLDAUP sco e while he VB esegmen a ion ole is now played by
he FBPLDAUP. Fo compa ison easons bo h models will p esen he same dimension o he
speake subspace han in he o iginal coun e pa .
In he i s compa ison we will compa e he ini ializa ion labels. Fo his eason, we e alua e
he whole AHC ee aking in o accoun wo simila i y me ics, SPLDA and PLDAUP. Then, o
each le el on bo h ees we will e alua e he esul ing pa i ions by means o DER and calcula e
∆DER = DERPLDAUP −DERSPLDA. This compa ison is epea ed o each in ol ed show in
MGB 2015, including bo h de elopmen and es subse s. In Fig. 5.7 we illus a e a his og am
abou he ela i e a ia ions o DER depending on he conside ed model.
In Fig. 5.7 we can obse e a dis ibu ion whose mean and mode a e biased o nega i e alues,
indica ing imp o emen s when subs i u ing he SPLDA model by he PLDAUP. Besides, he
skewness o he dis ibu ion is also nega i e, showing mo e ele ance o nega i es alues o
∆DER whe e PLDAUP ou pe o ms SPLDA.
A simila s udy can be pe o med wi h clus e and speake impu i ies. Fig. 5.8 epea s
he p eceding p ocedu e, al hough now we e alua e bo h impu i ies ins ead o DER. Simila
his og ams analyzing impu i ies o he pool o pa i ions a e shown, di e en ia ing be ween
clus e (Fig. 5.8a) and speake (Fig. 5.8b) impu i ies.
The esul s in Fig. 5.8 show dis ibu ions wi h mean and mode close o ze o o bo h cases.
Di e ences a ise when conside ing highe o de momen s, such as skewness when bo h im-
pu i ies ha e opposi e beha iou (speake impu i y is nega i e while clus e impu i y is posi-
i e), and ku osis, being highe in he clus e impu i y dis ibu ion. The beha iou obse ed
91
Dia iza ion o b oadcas da a wi h FBPLDAUP
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎❉❊❘ ✭✪✮
P
♦
♣
♦
✐
♦
♥
✆
✝
✞
✟✠s✡☛✠❜✉✡✠☞✌ ☞❢ ✍✟✎✏ ❜ ❡✡✇❡❡✌ ❙✑▲✟❆ ❛✌❞ ✑▲✟❆❯✑
Figu e 5.7: His og am o DER a ia ions be ween SPLDA and PLDAUP in MGB
2015 da a.
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎ ❈❧✉s ❡ ■♠♣✉ ✐ ② ✭✪✮
P
✆
♦
✝
♦
✆
✞
✟
♦
♥
✠
✡
☛
❉☞✌✍✎☞❜✏✍☞✑✒ ✑❢ ✓ ✔✕ ❜ ✖✍✇✖✖✒ ❙✗▲❉❆ ❛✒❞ ✗▲❉❆❯✗
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎ ❙♣ ❡❛ ❦❡ ■♠♣✉ ✐ ②✭✪✮
P
✆
♦
✝
♦
✆
✞
✟
♦
♥
✠
✡
☛
❉☞s✌✍☞❜✎✌☞✏✑ ✏❢ ✒✓✔ ❜ ✕✌✇✕✕✑ ✔✖▲❉❆ ✗✑❞ ✖▲❉❆❯✖
a) Clus e Impu i y b) Speake Impu i y
Figu e 5.8: His og am o a) clus e and b) speake impu i ies a ia ions be ween
SPLDA and PLDAUP ini ializa ions in MGB 2015 da a.
92
Chap e 5. Unce ain y P opaga ion o Dia iza ion
Expe imen DER(%)
SPK Se up
40 40.61
50 39.83
75 39.72
Bes FBPLDA 41.37
ELBO Se up
ELBO 39.83
Bes FBPLDA 39.12
Table 5.3: DER (%) esul s wi h he FBPLDAUP model in MGB 2015.
in Fig. 5.8 does no ma ch wi h hose p e iously ob ained in Sec ion 5.2.2. While in hose
expe imen s bene i s we e ob ained in he clus e impu i y, emaining speake impu i y almos
unal e ed, in b oadcas da a clus e impu i y shows an a e age 1.24% absolu e ex a deg ada-
ion, wi h 70% o he pa i ions deg ading his me ic, and speake impu i ies show a 2.05%
absolu e imp o emen , common o 63% o he pa i ions.
A conclusion ex ac ed om hese esul s indica es ha PLDAUP does no longe p o ide
such imp o emen s ob ained in elephone channel expe imen s. Some causes o his deg a-
da ion lie on he e ec o sho u e ance aining. In he elephone channel expe imen s we
obse ed how sho u e ances ials we e be e e alua ed as long as he SPLDA model con-
side ed hem du ing aining. By con as , PLDAUP did no show any imp o emen bu small
deg ada ions when ained wi h sho u e ances. In ou cu en scena io we mus deal wi h e y
sho u e ances, much sho e han hose used in elephone channel expe imen s. Hence some
educ ion o he expec ed imp o emen s seems easonable.
The inal s ep is he inclusion o he new model FBPLDAUP on op o he ini ializa ion,
ca ied ou by AHC wi h a PLDAUP model. In Table 5.3 we include hose expe imen s wi h he
new model a chi ec u e o MGB 2015 e alua ion subse . Two di e en c i e ia o hypo hesis
selec ion ha e been e alua ed: p io speake numbe es ima ion and ELBO choice. The expe -
imen s include hose esul s ob ained wi h he new app oach as well as a line indica ing hose
esul s p e iously ob ained wi h he non-UP models.
Acco ding o he esul s shown in Table 5.3, he FBPLDAUP shows po en ial bene i s de-
spi e inal esul s a e o e come by adi ional FBPLDA. When compa ing se ups wi h a ixed
numbe o speake s ou esul s clea ly imp o e hose ob ained by FBPLDA. Howe e , ou new
model does no ge any bene i om he ELBO s op c i e ion while FBPLDA does. A possible
93
PLDA ee-based clus e ing
6.2 PLDA ee-based clus e ing
The PLDA ee-based clus e ing is a s a is ical gene a i e solu ion o he dia iza ion clus e ing
ask. Hence, i ollows he p inciple desc ibed in Sec ion 2.6.2, iden i ying he a ge pa i ion
Θ o he se o embeddings Φas he one maximizing P(Φ,Θ). In o de o do so i exploi s he
ee pe spec i e o he clus e ing p oblem and he pa h decoding s a egies.
The applica ion o s a is ical s a egies on op o a ee s uc u e associa es a p obabili y o
each node. Rega ding he lea es his p obabili y is P(Φ,Θm) : m= 1..BN, i.e. he quali y
me ic o each pa i ion. In o de o ob ain a simila p obabili y o he emaining nodes we
make use o he p oduc ule o p obabili y. This ule allows he decomposi ion o a gene ic
P(a1, ..., aN)as:
P(a1, ..., aN) =P(a1)P(a2|a1)Pa3|a2
1...P aN|aN−1
1
=
N
Y
j=2
Paj|aj−1
1P(a1)(6.1)
whe e aj
1 ep esen s he se o elemen s {a1, ..., aj}. This de ini ion can also be exp essed in a
ecu si e way:
Paj
1=Paj|aj−1
1Paj−1
1(6.2)
The applica ion o he p oduc ule o p obabili y o ou dia iza ion p oblem is di ec , sub-
s i u ing he j h elemen aj om he p e ious equa ion by he j h pai o a iables, consis ing
o he embedding φjand i s clus e iden i y label θj. Hence, he p obabili y o any pa i ion
P(Φ,Θ) can be decomposed as:
P(Φ,Θ) =
N
Y
j=2
Pφj, θj|φj−1
1, θj−1
1P(φ1, θ1)(6.3)
and i s al e na i e ecu si e de ini ion:
Pφj
1, θj
1=Pφj, θj|φj−1
1, θj−1
1Pφj−1
1, θj−1
1(6.4)
In consequence, each node a dep h jin he clus e ing ee Thas he p obabili y
Pφj
1, θj
1;θj
1∈Ωθj
1associa ed. Acco ding o his decomposi ion we a e assuming Φas a
sequence o o de ed embeddings o be clus e ed. Besides, hese embeddings only depend on
p e ious speake ep esen a ions o he sequence. These assump ions a e easonable in eal li e,
whe e he oice e ol es along ime, being specially no iceable in la ge segmen s o speech.
100
Chap e 6. T ee-Based Clus e ing App oaches
6.2.1 PLDA-based model
Along he p e ious lines we explo ed a new pe spec i e abou clus e ing, ep esen ing i as a ee
s uc u e o be decoded in o de o ob ain he bes pa i ion. Mo eo e , he s a is ical s a egy
associa ed a p obabili y Pφj
1, θj
1 o all nodes along he ee. Howe e , he dis ibu ion o his
p obabili y has no been speci ied ye . Conside ing speake ecogni ion s a e o he a , PLDA
amily models seem a powe ul ype o solu ion o apply.
Thus, we mus keep on ans o ming Pφj
1, θj
1 o make PLDA de ini ion applicable. As
a i s ans o ma ion, we keep on applying he p oduc ule o p obabili y, decomposing he
p obabili y a each node in o a e m depending on he embeddings, he condi ional dis ibu ion,
and a p io dis ibu ion o he labels. This decomposi ion is:
Pφj, θj|φj−1
1, θj−1
1=Pφj|θj,φj−1
1, θj−1
1Pθj|φj−1
1, θj−1
1(6.5)
This decomposi ion allows o spli Pφj
1, θj
1in o wo simple p oblems. Now, we exclu-
si ely ocus on he i s e m, he condi ional dis ibu ion o he embedding jgi en i s j h label
as well as p e ious embeddings φj−1
1and labels θj−1
1. Decisions abou he o he e m, he p io
dis ibu ion o he cu en label θjgi en p e ious embeddings and decisions will be made a e -
wa ds. Un o una ely, his condi ional e m is s ill in ac able o use PLDA due o he p esence
o label a iables. Mo eo e , PLDA only de ines dependencies among embeddings om he
same speake . The e o e, ou nex ans o ma ions seek sepa a ing embeddings om he labels.
We ake inspi a ion om Chap e 4, imposing Pφj|θj,φj−1
1, θj−1
1 o ollow a mul inomial dis-
ibu ion on he a iable θj, a one-ho sample wi h I alues (θj={θ1j, ..., θij, ..., θIj}), whe e
Iis he numbe o candida e speake s. Thus, he alue θij will ake he alue o one i he j h
embedding was gene a ed by he speake i, being ze o o he wise. Besides, we also equi e φj,
when belonging o clus e iacco ding o θj, o be exclusi ely explained by hose embeddings
al eady assigned o his clus e . This subse o embeddings p e iously assigned o clus e ia
ime jis deno ed by Φij. Unde hese o condi ions we can exp ess:
Pφj|θj,φj−1
1, θj−1
1=
I
Y
i=1
Pφj|Φijθij ;j= 1..N (6.6)
The de ini ion o he e m Pφj|Φijnow makes he applica ion o PLDA p inciples easi-
ble. Fi s , we mus assume he exis ence o a la en a iable ep esen ing he speake in o ma ion
yi, which allow us o ede ine Pφj|Φijas:
Pφj|Φij=ZPφj|yiP(yi|Φij)dyi(6.7)
101
PLDA ee-based clus e ing
Gi en his de ini ion we can now assume ha he da a we a e dealing wi h is gene a ed
by a PLDA model. Fo his pu pose, we make use o he SPLDA condi ional dis ibu ion
Pφj|yi,MSPLDA:
Pφj|yi,MSPLDA∼ N φj|µ+Vyi,W−1(6.8)
whe e µis he speake independen e m, Va low ank ma ix desc ibing he speake subspace
and Wa ull ank ma ix explaining he in a-speake a iabili y space.
Fu he mo e, he second e m P(yi|Φij), he pos e io dis ibu ion o he la en a iable
gi en all hose embeddings p e iously assigned o clus e i(Φij) is also modeled acco ding o
SPLDA. Thus, i s de ini ion is:
P(yi|Φij,MSPLDA)∼N yi|µyi(j),L−1
yi(j)(6.9)
Lyi(j) =I+VT
j−1
X
k=1
θkiWV (6.10)
µyi(j) =L−1VTW
j−1
X
k=1
θki(φk−µ)(6.11)
whe e µyi(j)and Lyi(j) ep esen he es ima es o he mean and a iance pa ame e s espec-
i ely o he la en a iable yiwhen only j−1elemen s we e obse ed. As long as jinc eases
hese es ima ions should ge close o he eal alue.
Apa om well-known de ini ions o bo h dis ibu ions, he choice o he SPLDA
model also p o ides an ex a ad an age. I s Gaussian na u e o bo h Pφj|yi,MSPLDA
and P(yi|Φij,MSPLDA)allows a closed o m solu ion o he in eg al de ining
Pφj|Φij,MSPLDA. The esul ing o mula ion o his e m is:
Pφj|Φij,MSPLDA∼N φj|µi(j),Σi(j)(6.12)
µi(j) =µ+Vµyi(j)(6.13)
Σi(j) =W−1+VL−1
yi(j)VT(6.14)
A e comple ely de ining he condi ional dis ibu ion o he embeddings, we now can pay
a en ion o he label p io dis ibu ion Pθj|φj−1
1, θj−1
1. Fi s , we assume a simpli ied p io
dis ibu ion by elimina ing he dependence wi h espec o he pas embeddings φj−1
1. The e o e,
ou p io dis ibu ion will ollow he o m Pθj|θj−1
1. Fo he esul ing dis ibu ion we ha e
op ed o he Dis ance Dependen Chinese Res au an (DDCR) p ocess [Blei and F azie , 2011],
al eady used in dia iza ion in [Zhang e al., 2019]. This model explains he occupa ion o an
102
Chap e 6. T ee-Based Clus e ing App oaches
µW V
φjyi
θj
δ
ζ
Ni
I
Figu e 6.2: PLDA ee-based clus e ing Bayesian Ne wo k
in ini e se ies o clus e s by a sequence o elemen s. Then, he assignmen o he elemen j o
any clus e exclusi ely depends on he occupa ion o clus e s up o his poin , i.e. acco ding
o all he p e ious decisions, as in ou decomposi ion. The p obabili y o assignmen o he
elemen j o any o he al eady c ea ed k= 1..K clus e s is p opo ional o i s occupa ion
a ime j, namely nk. Besides, DDCR o e s he possibili y o c ea e a new clus e K+ 1
p opo ional o γ. The ma hema ical o mula ion o DDCR is:
Pθj=k|θ(j−1)
1∝(nki k≤K
ζi k=K+ 1 (6.15)
DDCR deeply ma ches sequen ial o de ing and assignmen p oblem. Un o una ely, DDCR
conside s easonable a con inuous ansi ion among speake s. Applied o scena ios o speake
clus e ing, whe e speake eco ding can be in e lea ed, seems easonable. Howe e , in dia iza-
ion we mus conside he segmen a ion s age, which can di ide any long segmen in o pieces
o sho e leng h. Thus, we can add o his dis ibu ion mo e chances o emain in he speake
clus e . In ou p oposal we do so by speci ically de ining he si ua ion o emaining in he cu -
en speake clus e , wi h a p obabili y p opo ional o δ. This addi ion gene a es he ollowing
modi ica ion o he DDCR dis ibu ion:
Pθj=k|θ(j−1)
1∝
δi k=θ(j−1)
nki k6=θ(j−1) and k≤K
ζi k6=θ(j−1) and k=K+ 1
(6.16)
The model P(Φ,Θ), aking in o accoun he whole se o assump ions p e iously desc ibed,
can be ep esen ed by he Bayesian ne wo k illus a ed in Fig. 6.2
6.2.2 M-algo i hm op imiza ion
Once he PLDA-based model is de ined, now i is ime o ind he way o ob ain hose labels
Θdia ha bes explain he se o embeddings Φ. Taking in o accoun ha he Vi e bi algo i hm
103
PLDA ee-based clus e ing
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.3: M-algo i hm example o a clus e ing ee o dep h 4. 2 pa hs ali e each
he dep h 2 h ough he ee (g een).
canno be applied, subop imal app oaches conside ing us wo hy pa hs, as he M algo i hm
[Jelinek and Ande son, 1971] a e s ill applicable.
The M algo i hm is an i e a i e solu ion s a egy. Gi en a scena io wi h a decision ee o
dep h N, he M algo i hm acks a subse o Msu i ing pa hs, i.e. hose pa hs mo e likely o be
he solu ion (in ou case hose wi h highe log-likelihood). Besides, all pa hs mus ha e eached
dep h jwi hin he ee s uc u e. Thus, he goal is he iden i ica ion o hose bes ansi ions
aking he M acked pa hs om dep h j o dep h j+ 1. In Fig. 6.3 we illus a e an example,
whe e a clus e ing ee o dep h 4 is analyzed by he M algo i hm wi h M= 2. Su i ing pa h
(g een lines) ha e eached dep h 2 wi hin he ee.
The M algo i hm i e a i e p ocedu e is di ided in o wo s eps, es ima ion and maximiza ion.
The es ima ion s ep s udies how he Msu i ing pa hs in le el je ol e deepe h ough he
ee, p edic ing i s pe o mance in a u u e scena io and making decisions in consequence. Fo
his pu pose we ca y ou a b u e- o ce app oach, analyzing any possible ansi ion om he M
su i ing pa hs a dep h jup o a ce ain ex a dep h d. The pa ame e dis a design choice and
esponsible o a adeo be ween accu acy and compu a ional cos s. The highe d, he wide
104
Chap e 6. T ee-Based Clus e ing App oaches
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.4: Es ima ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4.
2 pa hs ali e (g een) eaching dep h 2 a e p opaga ed o all possible nodes a dep h 3
(blue)
is he explo a ion o he ee o unseen da a and hence highe accu acy migh be expec ed, bu
inc easing in an exponen ial manne he compu a ional cos s. Hence, many sys ems es ic d
o be equal o 1. In Fig. 6.4 we ep esen he es ima ion s ep applied o ou p e ious example
scena io in Fig. 6.3. Each su i ing pa h (g een line) is p opaga ed d(d= 1) le els ahead (blue
lines), e alua ing o each con igu a ion he pe o mance a his dep h.
The esul s o he es ima ion s ep p o ide an o e iew abou how he ee beha es in u u e
s eps, wi hou comp omising any decision. This choice is made du ing he maximiza ion s ep.
In his s ep all candida e p opaga ions a e anked, only keeping hose Mwi h be e sco e.
These new Mpa hs now eaching dep h j+ 1 a e ou mos p omising candida es so a , and
hose conside ed o he nex i e a ion o he algo i hm. This s ep is ep esen ed in Fig 6.5.
6.3 Expe imen s
Fo he e alua ion o he new clus e ing app oach, we will make use o Albayzín 2018, as de-
sc ibed in Sec ion 3.2.2. Fo his pu pose, we conside an i- ec o PLDA dia iza ion sys em
105
Expe imen s
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.5: Maximiza ion s ep in a M-algo i hm example o a clus e ing ee o dep h
4. The wo su i ing pa hs a e shown in g een.
106
Chap e 6. T ee-Based Clus e ing App oaches
Expe imen DER(%)
De . Subse E al. Subse
AHC 18.88 26.36
AHC + FBPLDA 13.90 17,79
PLDA TREE-BASED CLUSTERING 13.12 17.60
Table 6.1: DER (%) esul s o he PLDA ee-based clus e ing in Albayzín 2018.
Resul s compa ed wi h hose ob ained by means o AHC wi h and wi hou FBPLDA
esegmen a ion.
whose se up is: A 256 Gaussian GMM-UBM ollowed by a 100-dimension To al Va iabili y
ma ix a e esponsible o he i- ec o ex ac ion. The ob ained embeddings unde go cen e -
ing, whi ening and leng h no maliza ion p io o clus e ing, wi hou dimensionali y educ ion.
Finally, he new clus e ing app oach, he PLDA ee-based clus e ing, uses a 100-dimension
SPLDA. This se up i s in e ms o dimensions wi h he dia iza ion sys em using he FBPLDA
eclus e ing in Chap e 4 o expe imen s wi h Albayzín 2018.
In ou i s expe imen we compa e he pe o mance o FBPLDA eclus e ing, ob ained in
Chap e 4, and ou new clus e ing app oach. As a i s app oxima ion we assume he se o
embeddings Φ o be a anged in empo al o de . We es ic hype pa ame e d o be equal o 1
o compu a ional easons. In his expe imen we conside e alua ion condi ions, i.e. we only
p esen he pe o mance o he bes hype pa ame e con igu a ion (δ,ζand M) acco ding o
Albayzín 2018 de elopmen subse . The ob ained esul s a e shown in Table 6.1.
Acco ding o he ob ained esul s, he new clus e ing app oach, wo king wi h i- ec o s, p o-
ides e y li le imp o emen wi h espec o he FBPLDA coun e pa . Howe e , hese esul s
show bene i s in he e alua ion o bo h de elopmen and es subse s despi e con aining inde-
penden shows. The e o e, we can alk abou limi ed ye consis en imp o emen s due o ou
new clus e ing app oach.
Apa om he o e all sco e o bo h de elopmen and es subse s, a mo e de ailed analysis
o esul s can also be done. Fo his pu pose, we s udy he pe o mance pe show o in e es
wi h he h ee clus e ing app oaches conside ed along his hesis: AHC, FBPLDA and PLDA
ee-based clus e ing. Fo his pu pose, we analyze wo di e en me ics: On he one hand we
p opose he analysis o ∆I=IORACLE−IHYP, he di e ence in he numbe o speake s be ween
ou hypo hesis labels and he e e ence. On he o he hand, we analyze he DER pe o mance.
Bo h analyses a e shown in Fig. 6.6, including all shows in Albayzín 2018. The in ol ed shows
om he de elopmen subse a e millenium and La Noche en 24 Ho as (LN24H). Rega ding he
es subse , he shows España en Comunidad (EC), La inoamé ica en 24 Ho as (LA24H), La
107
Expe imen s
♠✁✁✂✄✄☎♠ ▲✆✝✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌
✍✝
✷
✍
✶✷
✷
✶✷
✝
✷
✸✷
☞❚❊❊
❋✎✏▲✑✡
✡✟✠
*
*
*
◆❛✒❡ ♦❢ ❤ ❡ ❙❤♦✇
✓
■
❘✔❧✕✖✐✈✔ s♣ ✔✕❦✔ s ✗✘ ♣ ✔ s✙✚✛
♠✁✁✂✄✄☎♠ ▲✆✝ ✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌
✶✍
✝✍
✸✍
✞✍
✺✍
✻✍
☞❚❊❊
❋✎✏▲✑✡
✡✟✠
*
*
*
◆❛✒❡ ♦❢ ❤❡ ❙❤♦✇
❉
✓
❘
✭
✪
✮
✔✕✖✗✘✙ ♣✚ s✛✜✢
a) ∆Ib) DER(%)
Figu e 6.6: Analysis pe show o a) ∆Iand b) DER(%) o AHC, FBPLDA and
PLDA ee-based clus e ing. Analysis ca ied ou on shows om Albayzín 2018, in-
cluding de elopmen and es subse s.
Mañana (LM) and La Ta de en 24 Ho as Te ulia (LT24HTe ) a e also included. Resul s e lec
he in e qua ile ange o each show.
Those esul s illus a ed in Fig. 6.6 show a simila beha iou o he h ee ypes o clus e -
ing pe show. Thus, hose mo e ha m ul shows a e common o all sys ems. Howe e , ou
new clus e ing app oach shows a mino in e qua ile ange pe show compa ed o AHC and
specially FBPLDA. This educ ion a ec s bo h he es ima ion abou he numbe o speake s
and DER. Hence he pe o mance o he sys em seems mo e consis en pe indi idual show o
domain, al hough small deg ada ions migh occu . This beha iou can also be ex apola ed o
he whole da ase , specially conside ing he show La Mañana (LM). While AHC and FBPLDA
pe o mances o his show a e a leas 100% wo se han any o he show in e ms o DER, he
PLDA ee-based clus e ing achie es o beha e as bad as he second wo s show. This imp o e-
men is also obse ed in he es ima ion o he speake numbe , wi h a ela i e 25% deg ada ion
educ ion.
Apa om a speci ic se up, we can also do an analysis s udying he in luence o each o
he model hype pa ame e s δ,ζand M. Fo his analysis we will conside he ob ained sco es
o any possible se up. Fig. 6.7 is ou chosen g aphical ep esen a ion o e eal he impac o
he di e en hype pa ame e s. I is composed o wo pa s, Fig. 6.7a whe e we ep esen he
ela ionship be ween ζand DER o di e en alues o M, and Fig. 6.7b , whe e we ep esen
he ela ionship be ween δ, and DER o he di e en alues o M. In o de o include all
hype pa ame e s in each subimage, Fig. 6.7a includes some a iabili y pe measu e, illus a ing
he in e qua ile ange esul s in e ms o he missing hype pa ame e , δ. Simila ly, measu es in
108
Chap e 6. T ee-Based Clus e ing App oaches
✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁
✂
✶
✂
✶✁
✂
✝
✂
✝✁
✂✞
✂✞✁
✄✵
✄✵✁
✄✂
✄✂✁
✄✄
▼✂
▼✄
▼✆
▼✂✵
▼✄✵
▼✆✵
▼✂✵✵
✏
❉
❊
❘
✭
✪
✮
✟✠✡☛☞✌ ✐♥ ❡ ♠s ♦❢ ✍
✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁
✂
✶
✂
✶✁
✂
✝
✂
✝✁
✂✞
✂✞✁
✄✵
✄✵✁
✄✂
✄✂✁
✄✄
▼✂
▼✄
▼✆
▼✂✵
▼✄✵
▼✆✵
▼✂✵✵
✍
❉
❊
❘
✭
✪
✮
✟✠✡☛☞✌ ✐♥ ❡ ♠s ♦❢ ✎
a) ζp obabili y b) δp obabili y
Figu e 6.7: DER (%) esul s o he PLDA ee-based clus e ing wi h M-algo i hm in
Albayzín 2018 in e ms o δ,ζand M
Fig. 6.7b include some a iabili y ma gins indica ing he i s and hi d qua ile esul s in e ms
o ζ.
The in o ma ion included in Fig. 6.7 e eals many impo an cha ac e is ics abou he model.
Fi s , he esul s e idence he impo ance o M. 20% ela i e imp o emen s may be ob ained
as long as mo e and mo e simul aneous pa hs a e e alua ed. Howe e , his imp o emen is no
uni o m, being any inc ease o Mmo e signi ican o lowe alues. Fo highe alues o M,
imp o emen s a e e y sca ce and implying la ge inc emen s in he compu a ional cos s. O he
de ail o bea in mind is ha , excep o Mhype pa ame e , he in luence o he emaining
adjus able alues (δand ζ) is in gene al educed (wi h he excep ion o ζ o M= 2). a ia ions
may be a ound 5% ela i e imp o emen /deg ada ions, i.e. 1% absolu e DER a ia ions.
Up o his poin we ha e only men ioned h ee exis ing hype pa ame e s, M δ and ζ. Ne e -
heless, all he esul s we e ob ained by se ing he embeddings Φin o a sequen ial o de . I he
analysis o he clus e ing ee was comple e, i.e. analyzing each one o he lea es, he impac o
his o de ing would be null. Ne e heless, by pa ially explo ing he clus e ing ee acco ding
o limi ed da a makes his a angemen an ex a ac o o ake in o accoun . The e o e, while in
ou p e ious examples we exclusi ely applied empo al o de , i.e. we can also apply di e en
a angemen s
One o he key ac o s when using his ee-based app oach is he sequence o de , specially
aking in o accoun ha we a e exploi ing he ela ionships be ween an embeddings and i s
p edecesso s in he sequence. Thus, we mus explo e how he o de ing a ec s he esul s. While
in ou p e ious expe imen s we simply made use o he empo al o de , his a angemen is no
109
Sho u e ances as occluded u e ances
bu ions, i.e. he dia iza ion ask, usually wo ks wi h e en sho e segmen s (1s-3s) o accu a ely
deal wi h speake bounda ies. Hence, imp o emen s in his scena io a e becoming mo e and
mo e needed.
7.2 Sho u e ances as occluded u e ances
The sho u e ance p oblem is widely known wi hin he speake ecogni ion communi y
[Podda e al., 2017]. The e alua ion o ials by means o sho u e ances in ol es a se e e
deg ada ion o pe o mance. Howe e , he e is no s anda d de ini ion o sho u e ance in he
li e a u e. While some wo ks ha e epo ed losses o pe o mance wi h audios con aining less
han 30 seconds o speech, a mo e se e e deg ada ion is ob ained conside ing sho e u e ances
(less han 10 seconds) [Mandasa i e al., 2011,Kanagasunda am e al., 2011]. This sho u e -
ance p oblem has also been analyzed in he Speake Recogni ion E alua ions (SRE), p oposed
by NIST. Despi e adi ionally conside ing u e ances wi h mo e han 2 minu es o audio, some
o he e alua ions [NIST, 2008,NIST, 2010] also include a condi ion in which u e ances con-
ain less han 10 seconds.
This loss o pe o mance is a consequence o a highe in a-speake a iabili y in he
es ima ions wi h sho u e ances. In he li e a u e mul iple con ibu ions ha e been p o-
posed o he di e en s eps o he speake e i ica ion pipeline, aiming o educe he unde-
si ed a iabili y. The ea u e ex ac ion s ep has been s udied in di e en ways, a emp ing
o p o ide an al e na i e o adi ional MFCCs. In [Li e al., 2015] a mul i esolu ion ime-
equency ea u e ex ac ion was p oposed, ca ying ou a mul i-scaled Disc e e Cosine T ans-
o m (DCT) on he spec og am, combining he in o ma ion a e wa ds. Al e na i e wo ks
like [Alam e al., 2015] use di e en ea u es based on he ampli ude and phase o he spec-
um. O he con ibu ions a e ocused on he modelling s age. Fac o Analysis app oaches
we e conside ed in [Vog e al., 2008] o de elop subspace models o be e wo k wi h he
sho u e ances. When conside ing i- ec o ep esen a ions, compensa ion echniques such
as [Kanagasunda am e al., 2013,Kanagasunda am e al., 2014] p ojec he ob ained ep esen-
a ions in o subspaces wi h low a iabili y due o sho u e ances. In [Sa ka e al., 2012] i is
shown ha sys ems ained on sho u e ances should compensa e he unce ain y due o lim-
i ed audio, imp o ing he e alua ion o sho audios. Howe e , when sys ems mus deal wi h
audios wi h un es ic ed leng h, sys ems should be ained on long u e ances o a be e pe -
o mance. The balance o he Baum Welch s a is ics, equi ed o he ex ac ion o i- ec o s, is
also p oposed in [Hau amäki e al., 2013]. Besides, DNNs ha e also mapped sho -u e ance i-
ec o s wi h espec o hei long-u e ance coun e pa s [Guo e al., 2017]. O he con ibu ions
116
Chap e 7. S udy o embeddings o sho u e ances
ha e also wo ked on he backend, specially PLDA. Ano he echnique, o iginally p oposed in
[Cumani e al., 2013b,Kenny e al., 2013] and analyzed in Chap e 5, makes he PLDA model
include an ex a e m o compensa e he unce ain y o he i- ec o , which depends on he u e -
ance leng h. Finally, o he s a egies compensa e he ob ained sco e acco ding o eliabili y me -
ics o he in ol ed u e ances [Hasan e al., 2013,Mandasa i e al., 2013], specially i s du a ion.
This idea is ex ended in [Viñals e al., 2018b], whe e he Quali y Measu e Func ion (QMF) e m
s udies he in e ac ion be ween en ollmen and es u e ances. In [Vog e al., 2010] in e als o
con idence a e es ima ed, leading owa ds conside able accu acy.
Some wo ks such as [Ajili e al., 2016] ha e s udied he impac o he di e en phone ic con-
en in he embedding ep esen a ions. Acco ding o hei esul s, owels and nasal phonemes
a e help ul o disc imina ion ma e s. By con as , o he ypes o phonemes, such as ica i es
o plosi es, can be misleading du ing e alua ion. Ou hypo hesis o wo k applies his idea o
phonemes o sho u e ances. The p esence o ce ain acous ic uni s boos s he pe o mance o
speake ecogni ion sys ems. Howe e , hese boos ing phonemes mus be in bo h en oll and es
u e ances o be e ec i e. This ma ch in he phone ic in o ma ion goes beyond he p esence o
ce ain phonemes, also equi ing a ma ch in he phone ic dis ibu ion along he u e ance.
In o de o explain ou pe spec i e le ’s make an analogy o he sho u e ance p oblem wi h
a simila p oblem, ace ecogni ion wi h occlusions. In he bes scena io, bo h p oblems con ain
all possible in o ma ion. Wo king wi h aces we ha e a comple e iew o he pe son o in e es ,
including all he ace elemen s ( wo eyes, he nose, he mou h, e c.). In speake ecogni ion we
ha e comple e in o ma ion in an u e ance ha con ains aces o any possible phoneme and
i s coa icula ion. As long as he u e ance ge s longe and longe he comple e in o ma ion
condi ion is mo e likely o be achie ed. In his scena io pe o mance has imp o ed mo e and
mo e as long as echnologies ha e e ol ed.
Now we ocus on sho u e ances. These con ain much less speech, e en less han a second.
A simple "Yes/No" eply o a ques ion can cons i u e an u e ance. Hence, sho u e ances a e
e y likely o lack o phonemes. In ace ecogni ion he equi alen scena io is he ecogni ion o
pa ial in o ma ion, whe e some pa s such as he mou h and nose a e no isible. In bo h cases
he missing in o ma ion exis s, bu i is una ailable. Faces always ha e a mou h and a nose
al hough some imes hey can be occluded, e.g. by a sca . Rega ding speake cha ac e iza ion,
speake s p onounce all he phonemes o a language while alking, al hough ew o hem can be
missing in a speci ic u e ance.
In ou hypo hesis we also conside he in luence o p opo ion. Acco ding o ou analogy
o ace ecogni ion, aces p esen a ixed se o elemen s (ea s, nose, mou h, e c.) wi h a con-
s ained size, and loca ed in he ace in speci ic a eas. These es ic ions a e always he same,
117
Fo mula ion o he embedding ex ac ion wi h sho u e ances
ega dless o he pe son no any occlusion. In speake cha ac e iza ion he si ua ion is sligh ly
di e en . When u e ances ge long enough he language imposes es ic ions in he phoneme
dis ibu ion. These es ic ions lead o a e e ence phoneme dis ibu ion. The longe he u -
e ance he mo e i s phoneme dis ibu ion ends o he e e ence dis ibu ion. Howe e , sho
u e ances con ain a much sho e message, and hus i s phoneme dis ibu ion can be se e ely
dis o ed. In his dis o ion we mus ake in o accoun bo h he missing phonemes and hose
p esen bu condi ioned o he message in he u e ance. This dis o ion may lead o u e ances
om he same speake wi h di e en dominan phonemes, hence complica ing he e alua ion.
Consequen ly, he sho u e ance p oblem can be in e p e ed as an occlusion om a com-
ple e in o ma ion scena io. This occlusion may be comple e, whe e long u e ances lack om
ce ain phonemes, o pa ial, in which u e ances ha e hei phonemes seen in e y di e en
p opo ions wi h espec o hei coun e pa s. The a ailable in o ma ion abou he occlusion is
impo an o be awa e o . Du ing e alua ion we compa e how he wo speake s p onounce all
he phonemes, a ailable o no , so unbalanced in o ma ion can lead o an un ai compa ison.
7.3 Fo mula ion o he embedding ex ac ion wi h sho u -
e ances
Cu en s a e-o - he-a speake e i ica ion, as desc ibed in Sec ion 2.5, elies on he pipeline
embedding-backend. U e ances a e i s con e ed in o compac ep esen a ions, he embed-
dings, which eed he decision backend o ob ain he sco e. Among all a ailable ep esen a-
ions, wo o he mos popula ones a e i- ec o s and x- ec o s. Bo h ha e been widely es ed
in speake e i ica ion ob aining g ea esul s. Fi s , we will y o unde s and how we s o e he
speake in o ma ion in hese embeddings and hen s udy i s d awbacks o sho u e ances.
7.3.1 Gene al case
The me hod o compac a a iable leng h u e ance in o a ixed-leng h ep esen a ion is simila
o mos embedding ex ac ion echniques. Gi en he u e ance O, an o de ed se o Nacous ic
ea u es O={o1, ..., on, ..., oN}, we ans o m hem by unc ion F(·), ob aining he o de ed
sequence F(O) = { 1, ..., n, ..., N}. This unc ion maps he o iginal ea u e ec o onin o
he speake cha ac e is ics subspace as he p ojec ions n. Depending on he embedding, p o-
jec ion nin ol es he ans o ma ion o he ea u e ec o onas well as a small con ex a ound
(app oxima ely 0.15 seconds). By means o his mapping we a emp o highligh he speake
pa icula i ies in he ea u es applying linea (e.g. i- ec o s) o non-linea ans o ma ions (as
118
Chap e 7. S udy o embeddings o sho u e ances
in DNNs). The unc ion F(·)is lea n om a la ge da a pool by da a analysis, e.g. by Max-
imum Likelihood algo i hms o i- ec o s o Back-P opaga ion [Rumelha e al., 1986] wi h
DNNs. Due o he ac ha each one o hese p ojec ions nonly co e s a small pe iod o
ime, hey only ha e in o ma ion abou ew acous ic uni s. The comple e cha ac e iza ion o a
speake equi es he s udy o his/he pa icula i ies o all he phonemes. These acous ic uni s
a e widesp ead along he u e ance, hus we mus combine he e ec o all hese p ojec ions n.
The usual me hod o combine he p ojec ions is i s empo al a e age. The esul is he compac
ep esen a ion G(O), de ined as:
G(O) = 1
N
N
X
n=1
n(7.1)
This embedding G(O)keeps ack o he phone ic con en in he u e ance O. Howe e ,
we can also ea each acous ic uni independen ly. Many s a e-o - he-a embeddings, such as
i- ec o s, can be in e p e ed as he sum o C ep esen a ions Gc(O), one pe acous ic uni , each
one es ima ed acco ding o Ncp ojec ions n. Acco ding o his easoning we can exp ess he
embedding as:
G(O) =
C
X
c=1
αcGc(O)(7.2)
The ob ained exp ession desc ibes embeddings as a weighed sum o Ces ima ions Gc(O),
each one ep esen ing he es ima ed pa icula i ies o he speake in a single acous ic uni . Gc(O)
can also be in e p e ed as he esul ing embedding only aking in o accoun he da a ela ed o
he phoneme c. All he con ibu ions a e weigh ed by he e m αc, he p opo ion o his acous ic
uni in he u e ance.
The e o e, embeddings a e condi ioned o wo main pa s: On he one hand he s abili y
o he dis ibu ion o weigh s α={α1, ..., αc, ..., αC}. On he o he hand he es ima ions
Gc(O), he pa icula i ies pe phoneme. Bo h bene i om la ge u e ances. E e y language has
i s own e e ence phone ic dis ibu ion. Hence he longe he u e ance he mo e i s phone ic
dis ibu ion becomes like his e e ence. Conce ning he es ima ions Gc(O), he mo e a ailable
da a, he less unce ain is he es ima ion.
The a e age s age is he las s ep in which we keep ack o he phoneme dis ibu ion. As a
consequence, we canno dis inguish be ween speake and phone ic a iabili y a e wa ds. Fu -
he s eps in he embedding pos -p ocessing o he backend may ans o m he embedding, bu
all phonemes a e equally ea ed.
119
Fo mula ion o he embedding ex ac ion wi h sho u e ances
7.3.2 i- ec o embeddings
The p e iously desc ibed o mula ion also ma ches wi h he adi ional i- ec o s. The i- ec o
modeling pa adigm, al eady desc ibed in Sec ion 2.5, explains he u e ance Oas he esul o
sampling om a Gaussian Mix u e Model (GMM), speci ic o he u e ance wi h pa ame e s
λO. This model λOis he esul o he adap a ion om a Uni e sal Backg ound Model (UBM), a
la ge GMM ha e lec s all possible acous ic condi ions. This adap a ion p ocess is es ic ed o
only he UBM Gaussian means. Besides, he shi o he GMM Gaussians is ied, and explained
by means o a hidden a iable wO, loca ed in he To al Va iabili y subspace, desc ibed by ma ix
T. Ma hema ically:
µO=µUBM +TwO(7.3)
whe e µO ep esen s he supe ec o mean, he conca ena ion o he GMM componen means,
om he a ge λO.µUBM is he supe ec o mean om he Uni e sal Backg ound Model
(UBM), he e e ence model ep esen ing he a e age beha iou . wOis he la en a iable o
he u e ance O, wi h a s anda d no mal p io dis ibu ion and Tis a low ank ma ix de ining
he o al a iabili y subspace.
The i- ec o es ima ion looks o he bes alue o he la en a iable wOso as o explain
he gi en u e ance by means o he adap ed model. Fo his pu pose, we es ima e he pos e io
dis ibu ion o he la en a iable wOgi en he u e ance O. The i- ec o ep esen a ion w
co esponds o he mean o his pos e io dis ibu ion. De ined in [Dehak e al., 2011], he i-
ec o is o mula ed as:
w= C
X
c=1
TT
cΣ−1
cNc(O)Tc+I!−1C
X
c=1
TT
cΣ−1
c˜
Fc(O)(7.4)
=1
N(O) C
X
c=1
TT
cΣ−1
c
Nc(O)
N(O)Tc+1
N(O)I!−1C
X
c=1
TT
cΣ−1
cNc(O)˜
Fc(O)(7.5)
= C
X
c=1
TT
cΣ−1
cαcTc+1
N(O)I!−1C
X
c=1
αcTT
cΣ−1
c˜
Fc(O)(7.6)
=Ψ−1(O, α)
C
X
c=1
αcΓc(O) =
C
X
c=1
αcΨ−1(O, α)Γc(O) =
C
X
c=1
αcGc(O)(7.7)
whe e Tc ep esen s he po ion o he ma ix Ta ec ing he c h componen o he UBM. Σc
symbolizes he co a iance ma ix o he c h componen o he UBM. Nc(O)and ˜
Fc(O)a e he
120
Chap e 7. S udy o embeddings o sho u e ances
ze o h and cen e ed i s o de Baum Welch s a is ics o u e ance O. These s a is ics ep esen
he numbe o samples om componen cand he accumula ed de ia ion wi h espec o he
mean o he same componen espec i ely. N(O)symbolizes he o al numbe o ames in he
u e ance O. Finally, he e m ˜
Fc(O)is he a e age de ia ion pe sample o he u e ance o he
componen co he UBM.
The o mula ion o i- ec o s o e s special cha ac e is ics. Fi s , he alue o C, he numbe
o aced acous ic uni s o disc imina e, is ixed in he UBM. I s alue is equal o he numbe
o Gaussian componen s in he UBM. The e o e, Gc(O) ep esen s he con ibu ion pe sample
o he i- ec o om componen c, and he weigh αcis he p opo ion o ames assumed o be
sampled om same c h componen . Fu he mo e, i- ec o s ha e no speake awa eness in hei
o mula ion. They simply s o e he a ia ions in he acous ic uni s wi hin an embedding. These
de ia ions om he a e age beha iou , p ope ly ea ed by he backend, a e esponsible o he
pe o mance in speake iden i ica ion sys ems.
7.3.3 Sho u e ances
Now we conside he sho u e ance scena io. Acco ding o he p e ious analysis, embeddings
wo k well i he dis ibu ion o acous ic uni s αis simila o he e e ence dis ibu ion and he
pa icula con ibu ions Gc(O)a e es ima ed wi h low unce ain y. These wo equi emen s a e
eassu ed as long as he u e ance con ains mo e and mo e da a. Conce ning sho u e ances,
hei low amoun o da a makes hem likely o ha e hei dis ibu ion o acous ic uni s α a om
hei e e ence. Fo he same eason sho u e ances may also su e om la ge unce ain y in
hei phoneme es ima ions Gc(O). Hence, deg ada ion in sho u e ances can be explained by
he ollowing easons:
•E o s in he con ibu ion o phonemes. Some con ibu ions Gc(O)we e es ima ed
wi h e y li le in o ma ion. Then he unce ain y o hei es ima ion inc eases. Mul iple
alues wi hin his unce ain y ange as Gc(O)′can be es ima ed ins ead, commi ing he
e o E=Gc(O)′−Gc(O).
•Misma ch in he phoneme dis ibu ion. The dis ibu ion o he weigh s αdoes no
ma ch he e e ence α, de ined by language cha ac e is ics. This deg ada ion causes he
e o E=PC
c=1(αc−αc)Gc(O). The ex eme case happens when some acous ic uni s
a e no p esen in he u e ance, i.e. hey a e missing. In his si ua ion hei weigh αca e
equal o ze o, also o cing he missing es ima ions Gc(O) o be se o ze o, as i hey we e
occluded. The deg ada ion due o he misma ch in he phoneme dis ibu ion is compa ible
wi h he e o s in he con ibu ion o phonemes.
121
E ec s o he sho u e ances in i- ec o s
T adi ionally e o s ha e been a ibu ed o he con ibu ions pe acous ic uni . This is
specially ue when adi ional embeddings, e.g. i- ec o s, include an unce ain y e m in
i s calcula ions. Fo his eason, his so o e o was he i s a emp ed o deal wi h, e.g.
[Kenny e al., 2013]. Howe e , o he bes o ou knowledge no p e ious wo k has co e ed he
deg ada ion due o he phoneme dis ibu ion, which can cause simila le els o deg ada ion.
7.4 E ec s o he sho u e ances in i- ec o s
The phone ic dis ibu ion in an u e ance has impo an implica ions du ing he embedding ex-
ac ion. Embedding shi s due o inco ec con ibu ions Gc(O)a e complemen a y o hose
c ea ed by he misma ch in he phone ic dis ibu ion. In his sec ion we illus a e hei impac
wi h i- ec o s. This choice o well-known embeddings makes he s udy o bo h p oblems mo e
illus a i e in a simple way.
Fo his pu pose, we p opose a small dimension i- ec o expe imen o es he e ec s o
sho u e ances in some a i icial con olled da a. Gi en an e alua ion UBM i- ec o pipeline,
we compa e he i- ec o s ob ained om an o iginal u e ance and hose ob ained om he same
u e ance a e unde going con olled sho -u e ance modi ica ions. These modi ica ions a ec
bo h he acous ic uni dis ibu ion αand hei con ibu ions Gc(O). We make use o he ol-
lowing expe imen al se up: We i s sample a la ge a i icial da a pool om a UBM i- ec o
pipeline. This da a pool consis s o mo e han en housand independen u e ances, wi h one
hund ed wo-dimension samples each. The UBM is a 4-Gaussian GMM whose componen s a e
loca ed in (0,0),(0,10),(10,0) and (10,10), all o hem wi h he iden i y ma ix as co a iance.
The gene a i e i- ec o ex ac o has a 3-dimension hidden a iable subspace. Wi h en hou-
sand o hese u e ances we ain ou e alua ion pipeline, an al e na i e UBM i- ec o sys em.
Fo simplici y we sha e he gene a i e UBM. Rega ding he i- ec o ex ac o , we ain a model
wi h only a wo-dimension la en subspace. This dimension educ ion be ween gene a ion and
e alua ion has been conside ed o imi a e eal li e, whe e he gene a ion o da a is a oo complex
p ocess ha we only can app oxima e.
F om he emaining da a pool we choose wo ex a u e ances, unseen du ing he model
aining, o e alua ion pu poses. Because hese wo u e ances a e independen , we assume
hem o ep esen wo di e en speake s. In Fig. 7.1 we ep esen hem, ed and blue espec-
i ely. The ep esen a ion includes h ee pa s: In he i s pa we show he o iginal ea u e
domain, i.e. he u e ance se o ea u e ec o s. Each ellipse in he igu e ep esen s he dis i-
bu ion o each Gaussian in hei GMMs. The image also includes in g een he ep esen a ion
o he UBM model. The second pa in Fig. 7.1 ep esen s he same ed and blue u e ances in
122
Chap e 7. S udy o embeddings o sho u e ances
he la en space by means o he pos e io dis ibu ion o he la en a iable w. The hi d pa o
Fig. 7.1 illus a es he loca ion o he pa icula es ima ions pe componen Gc(O) o he wo
u e ances in he la en space. Reddish es ima ions co espond o he ed speake while bluish
ellipses ep esen he phonemes o he blue speake .
✲ ✲✁ ✵ ✁ ✻ ✽ ✶✵ ✶✁ ✶
✲
✲✁
✵
✁
✻
✽
✶✵
✶✁
✶
▼❋❈❈ ✂s ❉✐♠❡♥s✐♦♥
✄
☎
✆
✆
✷
✝
❞
✞
✟
✠
✡
✝
☛
✟
☞
✝
P ✌ ❥✍❝✎✏✌✑ ✏✑ ✒✍❛✎✉ ✍ ❙♣❛❝✍
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
a) Da a domain b) i- ec o domain c) Componen s in I . domain
Figu e 7.1: Scena io o in e es . a) U e ances ed and blue in he ea u edomain, wi h
he UBM componen s in g een. b) U e ances ed and blue in he i- ec o domain. c)
P ojec ions o he GMM componen s in he i- ec o domain o u e ances ed ( eddish
ellipses) and blue (bluish ellipses).
Following he desc ibed se up we can ca y ou an analysis o deg ada ion in sho u e ances.
Fi s , we illus a e he phoneme dependen es ima ion e o due o limi ed da a. Fo his eason
we es ima e he pos e io dis ibu ion o he embeddings o mul iple u e ances only di e ing
he numbe o samples. The dis ibu ion o phonemes α emains unal e ed. Theo e ically,
he embeddings should no su e any bias, bu i s unce ain y should ge la ge as long as he
u e ances con ain less da a. In Fig. 7.2 we compa e he o iginal u e ances o hose ob ained
wi h one i h o he da a and one en h o he da a.
Fig. 7.2 illus a es he pos e io dis ibu ion o he la en a iable o he sho u e ances
(dashed-line ed and blue ellipses) as well as he o iginal u e ances ( ed and blue ellipses wi h
con inuous line espec i ely). The loca ion o he ellipse ep esen s he mean o he pos e io
dis ibu ion while i s con ou he unce ain y. As expec ed, he o iginal e e ence u e ance and
hei sho e e sions p esen e y educed shi s among hemsel es, wi h almos concen ic
ellipses. While he blue speake su e s almos no deg ada ion, he ed speake biases a e mo e
no iceable. Besides, he illus a ion shows ha he less da a in he u e ance, he bigge he
unce ain y o he es ima ion.
Now we s udy he impac o he dis ibu ion o acous ic uni s αon he embedding. In he
e e ence u e ances his dis ibu ion was uni o m, his is, 25% o he samples came om each
componen . We now modi y his dis ibu ion o bo h u e ances, ed and blue. In Fig. 7.3 we
show he pos e io dis ibu ions o he o iginal u e ances ( ed and blue ellipses wi h con inuous
123
E ec s o he sho u e ances in i- ec o s
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
Figu e 7.2: Compa ison o pos e io dis ibu ion o he i- ec o s wi h e e ence
phoneme dis ibu ion. Con inuous line ellipse ep esen s he o iginal u e ance while
dashed-lined ellipses illus a e u e ances wi h he limi ed da a.
line) as well as he al e ed sho u e ances (dashed-line ed and blue ellipses). In he illus a ed
example hal o he ea u e ec o s a e sampled om a single componen o he GMM while
he emaining da a is e enly sampled along he o he componen s. We ha e s udied he e ec
wi h he ou componen s in he GMM.
Illus a ed esul s in Fig. 7.3 e eal he ele ance o he dis ibu ion o phonemes α o i s
p ope modelling. The modi ica ion o he dis ibu ion o weigh s makes he ed speake o
o e ou di e en ep esen a ions o he same embedding. Besides, hese ep esen a ions a e
no o e lapped among hemsel es, beyond he unce ain y egion om he o iginal u e ance.
The e o e, hese al e na i e embeddings a e likely o ail. Ne e heless, no all speake s beha e
equally. Whils ed speake is deg aded, ou blue speake has su e ed he same al e a ions
wi hou any isible shi on his/he embeddings.
The scena io wi h a dis o ed phoneme dis ibu ion can be aken o he limi . In his si ua ion
some componen s do no con ibu e o he inal embedding. This scena io is he mos ad e se,
signi ican ly modi ying he dis ibu ion o pa e ns αand some es ima ions pe phoneme Gc(O)
being se o ze o. In his expe imen we ha e dis u bed he dis ibu ion o acous ic uni s α
o cing wo o he componen s o ze o. In Fig. 7.4 we illus a e he six possible scena ios in
e ms o he non-con ibu ing componen s. The esul s a e shown o he wo es speake s
ed and blue, wi h con inuous line ellipses o he e e ence u e ances and dashed-line ellipses
o hei al e ed e sions. Acco ding o he ep esen a ions shown in Fig. 7.4, embeddings
om u e ances wi h missing componen s expe imen la ge biases wi h espec o he e e ence
124
Chap e 7. S udy o embeddings o sho u e ances
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
1s Componen 2nd Componen
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁ ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
3 d Componen 4 h Componen
Figu e 7.3: Compa ison o pos e io dis ibu ion o i- ec o s wi h modi ica ions in
he phoneme dis ibu ion α
embeddings. These shi s a e mo e signi ican han hose p e iously seen wi h less ex eme
dis o ions in he phoneme dis ibu ion α. Some o he hypo hesized embeddings a e a beyond
he unce ain y om he o iginal u e ance. The biases su e ed by he u e ances a e no he
same o bo h speake s. Again, he blue speake su e s no ele an deg ada ion. This beha iou
i s in ou hypo hesis because he missing componen s scena io is he limi case o phoneme
dis ibu ion deg ada ion.
In all ou expe imen s he ed speake has su e ed om s ong deg ada ions while he blue
speake has emained almos unal e ed. This di e en beha iou is a consequence o he loca-
ions o he phone ic es ima ions Gc(O) o each speake . On he one hand, as shown in Fig. 7.1,
ou blue speake has i s componen s e y close o each o he , p o iding obus ness agains dis-
ibu ion modi ica ions. On he o he hand ou ed speake has i s componen s much u he
125