scieee Science in your language
[en] (orig)

Repositorio Institucional de Documentos

Abstract

La motivación de esta tesis es la necesidad de soluciones robustas al problema de diarización. Estas técnicas de diarización deben proporcionar valor añadido a la creciente cantidad disponible de datos multimedia mediante la precisa discriminación de los locutores presentes en la señal de audio. Desafortunadamente, hasta tiempos recientes este tipo de tecnologías solamente era viable en condiciones restringidas, quedando por tanto lejos de una solución general. <br />Las razones detrás de las limitadas prestaciones de los sistemas de diarización son múltiples. La primera causa a tener en cuenta es la alta complejidad de la producción de la voz humana, en particular acerca de los procesos fisiológicos necesarios para incluir las características discriminativas de locutor en la señal de voz. Esta complejidad hace del proceso inverso, la estimación de dichas características a partir del audio, una tarea ineficiente por medio de las técnicas actuales del estado del arte. Consecuentemente, en su lugar deberán tenerse en cuenta aproximaciones. Los esfuerzos en la tarea de modelado han proporcionado modelos cada vez más elaborados, aunque no buscando la explicación última de naturaleza fisiológica de la señal de voz. En su lugar estos modelos aprenden relaciones entre la señales acústicas a partir de un gran conjunto de datos de entrenamiento. El desarrollo de modelos aproximados genera a su vez una segunda razón, la variabilidad de dominio. Debido al uso de relaciones aprendidas a partir de un conjunto de entrenamiento concreto, cualquier cambio de dominio que modifique las condiciones acústicas con respecto a los datos de entrenamiento condiciona las relaciones asumidas, pudiendo causar fallos consistentes en los sistemas.<br />Nuestra contribución a las tecnologías de diarización se ha centrado en el entorno de radiodifusión. Este dominio es actualmente un entorno todavía complejo para los sistemas de diarización donde ninguna simplificación de la tarea puede ser tenida en cuenta. Por tanto, se deberá desarrollar un modelado eficiente del audio para extraer la información de locutor y como inferir el etiquetado correspondiente. Además, la presencia de múltiples condiciones acústicas debido a la existencia de diferentes programas y/o géneros en el domino requiere el desarrollo de técnicas capaces de adaptar el conocimiento adquirido en un determinado escenario donde la información está disponible a aquellos entornos donde dicha información es limitada o sencillamente no disponible.<br />Para este propósito el trabajo desarrollado a lo largo de la tesis se ha centrado en tres subtareas: caracterización de locutor, agrupamiento y adaptación de modelos. La primera subtarea busca el modelado de un fragmento de audio para obtener representaciones precisas de los locutores involucrados, poniendo de manifiesto sus propiedades discriminativas. En este área se ha llevado a cabo un estudio acerca de las actuales estrategias de modelado, especialmente atendiendo a las limitaciones de las representaciones extraídas y poniendo de manifiesto el tipo de errores que pueden generar. Además, se han propuesto alternativas basadas en redes neuronales haciendo uso del conocimiento adquirido. La segunda tarea es el agrupamiento, encargado de desarrollar estrategias que busquen el etiquetado óptimo de los locutores. La investigación desarrollada durante esta tesis ha propuesto nuevas estrategias para estimar el mejor reparto de locutores basadas en técnicas de subespacios, especialmente PLDA. Finalmente, la tarea de adaptación de modelos busca transferir el conocimiento obtenido de un conjunto de entrenamiento a dominios alternativos donde no hay datos para extraerlo. Para este propósito los esfuerzos se han centrado en la extracción no supervisada de información de locutor del propio audio a diarizar, sinedo posteriormente usada en la adaptación de los modelos involucrados.<br /> <br /> Viñals Bailo, Ignacio; Ortega Giménez, Alfonso

Read accessible full text

Repositorio Institucional de Documentos

Publisher: Universidad de Zaragoza, Prensas de la Universidad
Year: 2020
Source: https://zaguan.unizar.es/record/99805/files/TESIS-2021-077.pdf
2021
77
Ignacio Viñals Bailo
Ad ances in Subspace-
based Solu ions o
Dia iza ion in he
B oadcas Domain
Di ec o /es
O ega Giménez, Al onso
© Uni e sidad de Za agoza
Se icio de Publicaciones
ISSN 2254-7606
Ignacio Viñals Bailo
ADVANCES IN SUBSPACE-BASED SOLUTIONS
FOR DIARIZATION IN THE BROADCAST DOMAIN
Di ec o /es
O ega Giménez, Al onso
Tesis Doc o al
Au o
2020
UNIVERSIDAD DE ZARAGOZA
Escuela de Doc o ado
P og ama de Doc o ado en Tecnologías de la In o mación y
Comunicaciones en Redes Mó iles
Reposi o io de la Uni e sidad de Za agoza – Zaguan h p://zaguan.uniza .es
UNIVERSIDAD DE ZARAGOZA
TESIS DOCTORAL - INGENIERÍA DE TELECOMUNICACIÓN
Ad ances in Subspace-based Solu ions o
Dia iza ion in he B oadcas Domain
Au ho :
Ignacio Viñals Bailo
Supe iso :
Al onso O ega Giménez
DEPARTAMENTO DE INGENIERÍA ELECTRÓNICA Y COMUNICACIONES
ESCUELA DE INGENIERÍA Y ARQUITECTURA
Ap il, 2020

A mis pad es
The human oice is
he mos pe ec ins umen o all.
A o Pä
The mos impo an ques ions o li e a e,
o he mos pa ,
eally only p oblems o p obabili y
Pie e-Simon Laplace
Resea ch is c ea ing new knowledge.
Neil A ms ong
The human b ain is an inc edible
pa e n-ma ching machine.
Je Bezos
4
Acknowledgemen s
I has been i e long yea s since I made he decision o s a a PhD p og amme. Along all his
ime I ha e been o una e enough o mee , collabo a e and be helped as well as suppo ed by
many people, wi hou whom his hesis would no be a ailable oday. These lines a e dedica ed
o all o hem.
Fi s and o emos , I wan o dedica e some lines o my pa en s. I wan o hank hem o
being on my side om he e y beginning. They always o e ed me hei suppo since I decided
o become a esea che . Fo he las i e yea s hey ha e chee ed me up du ing bad imes and
kep my ee on he g ound du ing hose limi ed success ul occasions. This wo k could no be
possible wi hou hei con ibu ion.
Besides, I also mus hank Al onso O ega o gi ing me he chance o g ow up p o es-
sionally and pe sonally. He ga e me he chance o disco e he esea che ca ee when I was
unde g adua e, and o e ed me he oppo uni y o keep on de eloping mysel wi h ViVoLAB
g oup. Fo he las i e yea s he has become a iend apa om a supe iso , who guided me
along his di icul lea ning p ocess. By means o ou mee ings he helped me o disco e some
o he bes ideas while o he imes he simply made me awa e ha some imes I could no see he
o es o he ees.
Apa om Al onso O ega, ViVoLAB g oup is also ull o wonde ul people who dese e
hei own men ion. Some o my kindes memo ies along my PhD yea s include Edua do Lleida.
He always ied o build a g oup based on iendship ela ionships a he han simply p o es-
sional ones. Besides, i is p aisewo hy how his ea men also included ViVoLAB alumni all
o e he wo ld. Ano he impo an collabo a o in his hesis is An onio Miguel. His aluable
expe ise in subspace models and neu al ne wo ks we e capi al o he wo k done in his hesis.
Besides, he heo e ical discussions we had abou hese opics we e also e y en iching, opening
my eyes abou unseen lines o esea ch o explo e.
I also wan o acknowledge Johns Hopkins uni e si y s a , specially Najim Dehak and his
5
Con en s
1 In oduc ion 1
1.1 Mo i a ion o he wo k . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Objec i es and Me hodology . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.3 Thesis o ganiza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
I Dia iza ion Basic Knowledge 7
2 Dia iza ion S a e o he A 9
2.1 In oduc ion .................................... 9
2.2 Main dia iza ion s a egies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.2.1 Bo om-Up dia iza ion sys ems . . . . . . . . . . . . . . . . . . . . . 12
2.3 Acous ic ea u es o dia iza ion . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.4 Audio segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.4.1 Me ic-based segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . 17
2.4.1.1 Bayesian In o ma ion C i e ion (BIC) . . . . . . . . . . . . . 18
2.4.1.2 Kullback-Leible Di e gence (KL) . . . . . . . . . . . . . . 19
2.4.1.3 Deep Neu al Ne wo ks (DNNs) . . . . . . . . . . . . . . . . 20
2.4.2 Model-based segmen a ion . . . . . . . . . . . . . . . . . . . . . . . . 20
2.5 Speake cha ac e iza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.5.1 Ea ly days . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.5.2 Model-based ep esen a ions . . . . . . . . . . . . . . . . . . . . . . . 22
2.5.2.1 Gaussian Mix u e Models (GMMs) . . . . . . . . . . . . . . 22
2.5.2.2 Suppo Vec o Machines (SVM) . . . . . . . . . . . . . . . 23
2.5.2.3 Join Fac o Analysis (JFA) . . . . . . . . . . . . . . . . . . 24
iii

CONTENTS
2.5.3 Embedded ep esen a ions . . . . . . . . . . . . . . . . . . . . . . . . 25
2.5.3.1 I- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.5.3.2 Hyb id i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . 26
2.5.3.3 DNN embeddings . . . . . . . . . . . . . . . . . . . . . . . 27
2.5.4 P obabilis ic Linea Disc iminan Analysis (PLDA) . . . . . . . . . . . 27
2.6 Clus e ing ..................................... 28
2.6.1 Hie a chical clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . 32
2.6.2 S a is ical app oaches . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.6.3 O he al e na i es . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
2.7 Pe o mance me ics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3 Analysis o Dia iza ion in B oadcas Da a 39
3.1 The dia iza ion e e ence sys em . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.2 Analysis o b oadcas da a . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2.1 Mul i-Gen e B oadcas Challenge 2015 (MGB 2015) . . . . . . . . . . 42
3.2.2 Albayzín 2018 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.2.3 Acous ic a iabili y . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.2.4 Va iabili y in he speake dis ibu ion . . . . . . . . . . . . . . . . . . 46
3.3 E alua ion o pe o mance o he dia iza ion e e ence sys em . . . . . . . . . 48
3.3.1 E alua ion o pe o mance in MGB 2015 . . . . . . . . . . . . . . . . 48
3.3.2 E alua ion o pe o mance in Albayzín 2018 . . . . . . . . . . . . . . 50
3.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
3.4.1 The clus e ing app oxima ion . . . . . . . . . . . . . . . . . . . . . . 52
3.4.2 The quali y o he embeddings . . . . . . . . . . . . . . . . . . . . . . 52
3.4.3 The domain misma ch p oblem . . . . . . . . . . . . . . . . . . . . . 53
II The Clus e ing P oblem 55
4 Clus e ing by means o Fully Bayesian PLDA 57
4.1 The Fully Bayesian PLDA clus e ing solu ion . . . . . . . . . . . . . . . . . . 57
4.1.1 The Fully Bayesian PLDA (FBPLDA) model . . . . . . . . . . . . . . 57
4.1.2 The clus e ing p ocedu e . . . . . . . . . . . . . . . . . . . . . . . . . 60
4.1.3 Dia iza ion using he FBPLDA model . . . . . . . . . . . . . . . . . . 62
4.2 Analysis o FBPLDA pe o mance . . . . . . . . . . . . . . . . . . . . . . . . 65
4.2.1 Ini ializa ion impac . . . . . . . . . . . . . . . . . . . . . . . . . . . 65
i
CONTENTS
4.2.2 In e ence o he numbe o speake s . . . . . . . . . . . . . . . . . . . 66
4.2.3 Numbe o speake s s DER . . . . . . . . . . . . . . . . . . . . . . . 69
4.2.4 Numbe o speake s s ELBO . . . . . . . . . . . . . . . . . . . . . . 71
4.3 Al e na i e ini ializa ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
4.3.1 Compu a ionally e icien ini ializa ion . . . . . . . . . . . . . . . . . 73
4.3.2 ELBO-based ini ializa ion choice c i e ion . . . . . . . . . . . . . . . 74
4.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
5 Unce ain y P opaga ion o Dia iza ion 79
5.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
5.2 PLDA wi h Unce ain y P opaga ion (PLDAUP) . . . . . . . . . . . . . . . . . 80
5.2.1 PLDAUP in speake ecogni ion . . . . . . . . . . . . . . . . . . . . . 82
5.2.2 PLDAUP in speake clus e ing . . . . . . . . . . . . . . . . . . . . . . 86
5.3 FBPLDA wi h Unce ain y P opaga ion (FBPLDAUP) . . . . . . . . . . . . . 89
5.4 Dia iza ion o b oadcas da a wi h FBPLDAUP . . . . . . . . . . . . . . . . . 91
5.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
6 T ee-Based Clus e ing App oaches 97
6.1 T ee-based poin o iew o clus e ing . . . . . . . . . . . . . . . . . . . . . . 98
6.2 PLDA ee-based clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
6.2.1 PLDA-based model . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
6.2.2 M-algo i hm op imiza ion . . . . . . . . . . . . . . . . . . . . . . . . 103
6.3 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
6.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112
III The Speake Rep esen a ion P oblem 113
7 S udy o embeddings o sho u e ances 115
7.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
7.2 Sho u e ances as occluded u e ances . . . . . . . . . . . . . . . . . . . . . 116
7.3 Fo mula ion o he embedding ex ac ion wi h sho u e ances . . . . . . . . . 118
7.3.1 Gene al case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
7.3.2 i- ec o embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
7.3.3 Sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
7.4 E ec s o he sho u e ances in i- ec o s . . . . . . . . . . . . . . . . . . . . 122
CONTENTS
7.5 Expe imen s & Resul s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5.1 Expe imen al se up . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5.2 Baseline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
7.5.3 Reduc ion o he misma ch in α: Phone ic balance . . . . . . . . . . . 128
7.5.4 En ollmen - es dis ance s log-likelihood a io (ll ) . . . . . . . . . . . 131
7.5.5 En ollmen - es dis ance s pe o mance (EER and minDCF) . . . . . . 133
7.5.6 Long-sho s Equalized Sho -Sho . . . . . . . . . . . . . . . . . . . 134
7.6 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
8 DNNs embeddings o Dia iza ion 137
8.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
8.2 Hyb id i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
8.2.1 Bo leneck Fea u es (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 139
8.2.2 Phone ic i- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
8.3 X- ec o s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
8.4 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144
8.4.1 Bo leneck Fea u es (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 145
8.4.2 Phone ic i- ec o s & x- ec o s . . . . . . . . . . . . . . . . . . . . . . 147
8.4.2.1 Speake ecogni ion . . . . . . . . . . . . . . . . . . . . . . 148
8.4.2.2 B oadcas dia iza ion . . . . . . . . . . . . . . . . . . . . . 150
8.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
IV The Model Adap a ion P oblem 155
9 Da a-E icien Domain Adap a ion o PLDA Models 157
9.1 In oduc ion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
9.2 Me hods o domain misma ch educ ion . . . . . . . . . . . . . . . . . . . . . 158
9.3 Expe imen s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
9.3.1 Independen unsupe ised adap a ion . . . . . . . . . . . . . . . . . . 162
9.3.2 Longi udinal unsupe ised adap a ion . . . . . . . . . . . . . . . . . . 163
9.3.3 Use o in-domain labeled da a and semi-supe ised adap a ion . . . . . 164
9.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166
i
CONTENTS
V Conclusions & Fu u e Wo k 167
10 Conclusions & Fu u e wo k 169
10.1 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
10.1.1 The clus e ing ask . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169
10.1.2 The speake cha ac e iza ion s age . . . . . . . . . . . . . . . . . . . . 170
10.1.3 Unsupe ised domain adap a ion esea ch . . . . . . . . . . . . . . . . 171
10.2 Scien i ic Con ibu ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.1 Book chap e s . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.2 Pape s published in jou nals included in he Jou nal Ci a ion Repo s
(JCR) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
10.2.3 Con e ence p oceedings . . . . . . . . . . . . . . . . . . . . . . . . . 173
10.3 Fu u e Wo k . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
VI Appendix 175
A Fully Bayesian PLDA wi h Unce ain y P opaga ion I
A.1 De ini ions ..................................... I
A.2 Da a ........................................ II
A.3 Da a condi ional likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . III
A.3.1 P(Φi|yi,Xi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . III
A.3.2 P(Xi|yi,Θi,Φi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . IV
A.3.3 P(yi|Φi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . . . IV
A.4 Va ia ional app oach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V
A.4.1 Join p obabili y . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V
A.4.2 Va ia ional Bayes app oxima ion . . . . . . . . . . . . . . . . . . . . . VI
A.4.3 Op imal de ini ion o q∗(Y,X). . . . . . . . . . . . . . . . . . . . . VI
A.4.4 Op imal de ini ion o q∗(Θ) . . . . . . . . . . . . . . . . . . . . . . . VII
A.4.5 Op imal de ini ion o q∗(πθ). . . . . . . . . . . . . . . . . . . . . . . VIII
A.4.6 op imal de ini ion o q∗˜
V. . . . . . . . . . . . . . . . . . . . . . . VIII
A.4.7 Op imal de ini ion o q∗(W). . . . . . . . . . . . . . . . . . . . . . . X
A.4.8 Op imal de ini ion o q∗(ε). . . . . . . . . . . . . . . . . . . . . . . XI
A.4.9 Necessa y Expec a ions . . . . . . . . . . . . . . . . . . . . . . . . . XI
A.4.10 Va ia ional Lowe Bound . . . . . . . . . . . . . . . . . . . . . . . . . XIV
A.5 Hype pa ame e op imiza ion . . . . . . . . . . . . . . . . . . . . . . . . . . . XV
ii
CONTENTS
iii

Lis o Figu es
1.1 Example o dia iza ion esul s . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.2 Concep ual map o he s udied opics in his Thesis . . . . . . . . . . . . . . . 5
2.1 Schema ic o Bo om-Up and Top-Down dia iza ion . . . . . . . . . . . . . . . 11
2.2 Gene al schema ic o a dia iza ion sys em . . . . . . . . . . . . . . . . . . . . 12
2.3 Schema ic o he MFCC ex ac ion pipeline . . . . . . . . . . . . . . . . . . . 14
2.4 Scheme o a sliding window me ic based segmen a ion . . . . . . . . . . . . 18
2.5 Schema ic o an Agglome a i e Hie a chical Clus e ing (AHC) pe o mance . 32
3.1 Schema ic o ou baseline dia iza ion sys em . . . . . . . . . . . . . . . . . . . 40
3.2 Sec ion a iabili y example. Fo 100 i s embeddings om a Sp ingwa ch
episode wi h SPLDA pai wise LLR simila i y me ic and G ound u h ela-
ionship. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
3.3 Va iabili y in he speake dis ibu ion o MGB 2015, numbe o speake s pe
show and he p opo ion o speech o he mos ac i e speake pe show. . . . . 47
3.4 Va iabili y in he speake dis ibu ion o Albayzín 2018, numbe o speake s
pe show. and p opo ion o speech o he mos ac i e speake pe show. . . . . 48
3.5 Dis ibu ion o speech pe speake o wo episodes: An episode wi h a domi-
nan speake and an episode wi h a mo e e en speech dis ibu ion . . . . . . . . 49
4.1 Bayesian ne wo k o he Fully Bayesian PLDA . . . . . . . . . . . . . . . . . 58
4.2 Clus e ing schema ic based on label ini ializa ion and FBPLDA esegmen a ion 61
4.3 Schema ic o he dia iza ion sys em based on he FBPLDA esegmen a ion . . 62
4.4 Analysis o ∆I=IORACLE −IHY P o shows in MGB 2015 wi h AHC and
FBPLDA esegmen a ion dia iza ion sys ems. . . . . . . . . . . . . . . . . . . 64
4.5 5-le el dend og am example. . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
ix
LIST OF FIGURES
4.6 Inpu /ou pu ela ionship o he numbe o speake s wi h FBPLDA esegmen-
a ion. ....................................... 67
4.7 Los speake s acco ding o he ela i e numbe o speake s ∆Iin he ini ial
pa i ion Θ0.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
4.8 DER (%) esul s o a) AHC and b) FBPLDA in e ms o he ela i e numbe o
speake s ∆I.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
4.9 Dis ibu ion o he ini ializa ion wi h bes DER in e ms o he ela i e numbe
o speake s. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
4.10 Dis ibu ion o he ini ializa ion wi h bounded DER, a) 1% and b) 3%, in e ms
o he ela i e numbe o speake s. . . . . . . . . . . . . . . . . . . . . . . . . 72
4.11 Dis ibu ion o he pa i ion wi h bes ELBO in e ms o ela i e speake s. . . . 73
4.12 Schema ic o dia iza ion based on he simul aneous e alua ion o K di e en
ini ializa ions. The inal pa i ion is selec ed by means o PELBO. . . . . . . . 75
5.1 Bayesian ne wo k o PLDA wi h Unce ain y P opaga ion (PLDAUP) . . . . . 81
5.2 DET cu es wi h SPLDA o SRE10 co ex -co eex de 5 emale wi h in ol ed
sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
5.3 DET cu es wi h PLDAUP o SRE10 co ex -co eex de 5 emale wi h in ol ed
sho u e ances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
5.4 Impu i y esul s o SPLDA and PLDAUP in SRE10 co eex -co eex de 5 e-
male chopped . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
5.5 Impu i y esul s o a) SPLDA and b) PLDAUP in SRE10 co eex -co eex de 5
emale chopped aining wi h sho u e ances . . . . . . . . . . . . . . . . . . 88
5.6 Bayesian ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opaga ion . 90
5.7 His og am o DER a ia ions be ween SPLDA and PLDAUP in MGB 2015 da a. 92
5.8 His og am o a) clus e and b) speake impu i ies a ia ions be ween SPLDA
and PLDAUP ini ializa ions in MGB 2015 da a. . . . . . . . . . . . . . . . . . 92
6.1 4-le el ee clus e ing example . . . . . . . . . . . . . . . . . . . . . . . . . . 99
6.2 PLDA ee-based clus e ing Bayesian Ne wo k . . . . . . . . . . . . . . . . . 103
6.3 M-algo i hm example o a clus e ing ee o dep h 4 . . . . . . . . . . . . . . 104
6.4 Es ima ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4 . . . 105
6.5 Maximiza ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4 . . 106
6.6 Analysis pe show o ∆Iand DER(%) o AHC, FBPLDA and PLDA ee-
based clus e ing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
x
LIST OF FIGURES
6.7 DER (%) esul s o he PLDA ee-based clus e ing wi h M-algo i hm in Al-
bayzín 2018 in e ms o δ,ζand M. . . . . . . . . . . . . . . . . . . . . . . 109
6.8 DER ela i e esul s be ween Random o de and Time o de . . . . . . . . . . 111
7.1 Scena io o in e es . a) U e ances ed and blue in he ea u e domain, wi h he
UBM componen s in g een. b) U e ances ed and blue in he i- ec o domain.
c) P ojec ions o he GMM componen s in he i- ec o domain o u e ances . 123
7.2 Compa ison o pos e io dis ibu ion o he i- ec o s wi h e e ence phoneme
dis ibu ion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
7.3 Compa ison o pos e io dis ibu ion o i- ec o s wi h modi ica ions in he phoneme
dis ibu ion α. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
7.4 Compa ison o pos e io dis ibu ion o i- ec o s when wo phonemes a e no
con ibu ing and αc= 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
7.5 DET cu es o he scena ios Long-Long (blue), Long-Sho and Sho -Sho
Random ( ed con inuous and dashed line espec i e), Long-Sho and Sho -
Sho Balanced (g een con inuous and dashed line espec i e) o SRE10 "co eex -
co eex de 5 emale" expe imen . . . . . . . . . . . . . . . . . . . . . . . . . 130
7.6 T ial sco e in e ms o KL2 dis ance o he whole da a pool. Rep esen ed he
mean and he mean plus/minus he s anda d de ia ion . . . . . . . . . . . . . . 132
7.7 E alua ion me ics, EER (a) and minDCF (b) in e ms o he KL2 dis ance. . . 133
7.8 DET cu es o he scena ios Long-Sho Random and Sho -Sho Equalized
in SRE10 "co eex -co eex de 5 emale" . . . . . . . . . . . . . . . . . . . . . 135
7.9 No malized dis ibu ion o sco es o Ta ge (blue) and Non- a ge ( ed) ials
o scena ios Long-Sho (con inuous line) and Sho -Sho Equalized (dashed
line).Expe imen ca ied ou wi h SRE10 "co eex -co eex de 5 emale". . . . . 135
8.1 Example o a Bo leneck Fea u e ex ac o DNN . . . . . . . . . . . . . . . . . 140
8.2 BNF pipeline om he o iginal MFCCs up o Baum Welch s a is ics . . . . . . 140
8.3 Bayesian ne wo k o he phone ic i- ec o . . . . . . . . . . . . . . . . . . . . 142
8.4 Phone ic i- ec o pipeline om he o iginal MFCCs up o Baum Welch s a is ics 143
8.5 X- ec o a chi ec u e schema ic . . . . . . . . . . . . . . . . . . . . . . . . . 143
8.6 DET cu es o x- ec o s in SRE10 wi h long and sho u e ances . . . . . . . 150
8.7 Dis ibu ion o he embedding i s componen in s anda d i- ec o s, phone ic
i- ec o s and x- ec o s o he aining co pus in Albayzín 2018 . . . . . . . . 153
9.1 Schema ic o he supe ised and unsupe ised adap a ion . . . . . . . . . . . 159
xi
LIST OF FIGURES
9.2 Schema ic o unsupe ised independen adap a ion o he episodes n−1,n
and n+ 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
9.3 Schema ic o unsupe ised longi udinal adap a ion o he episodes n−1,n
and n+ 1.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
9.4 Semi-supe ised adap a ion s a egy based on he unsupe ised independen
adap a ion app oach o he episodes n−1,nand n+ 1.. . . . . . . . . . . . 160
9.5 Semi-supe ised adap a ion s a egy based on he longi udinal unsupe ised
adap a ion app oach o he episodes n−1,nand n+ 1 . . . . . . . . . . . . 160
9.6 ∆DER (%) pe o mance episode by episode o he wo shows o he e alua-
ion se . De ined as ∆DER = (DERINDEP −DERLONG). AHC e e s o he
Agglome a i e clus e ing pseudo-speake labels. . . . . . . . . . . . . . . . . . 164
A.1 Bayesian Ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opaga ion . II
xii
Objec i es and Me hodology
o each domain.
1.2 Objec i es and Me hodology
The objec i es o his hesis a e he imp o emen o dia iza ion capabili ies so ha sys ems
could wi hs and he ha m ul condi ions o he b oadcas domain. These e olu ions should be
in eg a ed in a single sys em, obus enough o deal wi h any so o audio om he s udied
en i onmen . The e o e, we should analyze possible e olu ions in he p e iously desc ibed
h ee lines o esea ch.
Rega ding o he speake cha ac e iza ion p oblem, we wan o imp o e he ex ac ion o
he speake ep esen a ions, ob aining e icien and disc imina i e cha ac e iza ions o he
in ol ed speake s. Thus, we i s seek a deepe unde s anding abou he s a e-o - he-a mod-
elling echniques based on subspace p ojec ion. Once his knowledge is is acqui ed, i will le
us explo e he limi a ions o hese echnologies, as well as p opose new app oaches designed
acco dingly.
Wi h espec o he g ouping ask, ou goal is he imp o emen o he clus e ing echniques
es ima ing he dia iza ion pa i ions. Fo his pu pose, we make use o subspace-based ech-
niques, specially PLDA, explo ing di e en a chi ec u es and s a egies.
Finally, we also mus deal wi h he domain a iabili y. In his a ea we will y o p o ide
ools and s a egies capable o dec easing he deg ada ion o domain misma ch in ci cums ances
whe e in-domain da a is sca ce o una ailable. In o de o each his goal we will deal wi h
he domain adap a ion p oblem by explo ing he in e ence o unsupe isedly-c a ed pseudo-
speake labels, ob ained om he audio o dia ize. These labels should be la e used o speci i-
cally adap he ou -o -domain model o he e alua ion audio.
1.3 Thesis o ganiza ion
The ou line o his hesis is e y o ien ed o he di e en challenges we p e iously desc ibed.
Fo his eason, his wo k is di ided in i e main pa s, as shown in he concep ual map in
Fig. 1.2:
•Basic Knowledge: This pa is dedica ed o p esen he dia iza ion p oblem and an
o e iew o he al eady p oposed echniques in he s a e o he a (Chap e 2). Mo e-
o e , his pa also s a s he expe imen al ac i i y, analyzing he cha ac e is ics o he
b oadcas domain and he pe o mance o a baseline dia iza ion sys em (Chap e 3).
4

Chap e 1. In oduc ion
Dia iza ion
Basic
Knowledge
S a e o
he a
Baseline
dia iza ion
Sys em
Model
Adap a ion
Speake
Clus e ing
FBPLDA
Clus e ing
T ee-based
Clus e ing
Unce ain y
P opa-
ga ion
Conclusions
&
Fu u e Wo k
Speake
Rep esen a ion
Embeddings
in Sho
U e ances
DNN
Embeddings
Figu e 1.2: Concep ual map o he s udied opics in his Thesis
5
Thesis o ganiza ion
•Speake Clus e ing: This pa is ocused on he di e en ools o imp o e he pe o -
mance o he clus e ing s age. Fi s , we analyze he pe o mance o he Fully Bayesian
P obabilis ic Linea Disc iminan Analysis (FBPLDA) model, dealing wi h i s weak-
nesses (Chap e 4). Chap e 5upda es he FBPLDA mode including he concep o Unce -
ain y P opaga ion (FBPLDAUP). Finally, in Chap e 6we p esen a o ally independen
clus e ing solu ion by means o a ee-based app oach.
•Speake Rep esen a ion This pa o he hesis pays a en ion o he way speake in o -
ma ion is ex ac ed om an audio u e ance and compac ed in o a condensed ep esen-
a ion, he embedding. Fi s , we s udy he s anda d app oxima ion o his in o ma ion
ex ac ion, analyzing i s impac on sho u e ances (Chap e 7). La e on, we make use
o he lea n conclusions, applying hem on he ob en ion o DNN-based embeddings o
dia iza ion (Chap e 8).
•Model Adap a ion: This pa , consis ing on Chap e 9, wo ks on he unsupe ised ex-
ac ion o in-domain in o ma ion, sui able o he adap a ion o ou -domain labels. This
•Summa y: This inal pa summa izes he conclusions o all he di e en pa s o he
hesis and p oposes how his esea ch could be ollowed in he u u e (Chap e 10)
6
Pa I
Dia iza ion Basic Knowledge
7
Chap e 2
Dia iza ion S a e o he A
The objec i e o his chap e is he e ision o he s a e o he a in di-
a iza ion. Fo his pu pose, we ake in o accoun impo an e iews such as
[Angue a e al., 2012][T an e and Reynolds, 2006]. Ou i s goal is he iden i ica ion o
he main domains in which dia iza ion has been applied. This di e en ia ion helps unde s and-
ing he e olu ion o dia iza ion echnologies. This knowledge allows he in oduc ion o he
wo main app oxima ions o dia iza ion. Then, we explain in de ail he unc ional blocks o
he mos popula dia iza ion app oach. Finally, he las pa o he chap e includes a e iew
abou how o measu e dia iza ion pe o mance.
2.1 In oduc ion
The dia iza ion ask includes all he echniques and p ocedu es needed o di e en ia e he con-
ibu ions o speake s gi en an audio. In he mos gene al case, dia iza ion wo ks in an unsupe -
ised way, i.e. wi hou p io knowledge abou he in ol ed speake s no i s numbe . Howe e ,
dia iza ion can ge bene i ed by means o he knowledge o hese cha ac e is ics, usually sim-
pli ying he p oblem. His o ically, dia iza ion esea ch has ocused on h ee main domains o
in e es :
•Telephone channel domain. This en i onmen in ol es he analysis o elephone con e -
sa ions, cha ac e ized by he p esence o ew speake s, usually wo, and con e sa ional
speech wi h sho in e en ions. Mo eo e , elephone con ex usually conside s close- o-
mou h mic ophones and es ic ed a p io i known channel condi ions.
•B oadcas domain. This condi ion includes audios om mass media b oadcas e s (TV,
adio, VoD, e c.). The mos impo an ea u e in b oadcas da a is he la ge a iabili y
9

In oduc ion
o condi ions. The a iabili y in he numbe o speake s is almos un es ic ed: om
3-4 up o 100 di e en speake s pe hou o con en , depending on he show. The e is
also a iabili y in he ype o speech: while some shows con ain mo e ead speech, e.g.
he news, o he s ha dly e e include i , being mainly composed o con e sa ional speech,
such as alk-shows. This cha ac e is ic has g ea ele ance in dia iza ion due o he leng h
o he in e en ions. Whils con e sa ional speech usually consis s o sho in e en ions
in o de o main ain he con e sa ion low, ead speech can gene a e longe u ns due o
he absence o eedback. Mo eo e , b oadcas audio also p esen s a iabili y o acous ic
scena ios, such as s udio and ou doo s, each one wi h i s own acous ic cha ac e is ics.
Finally, excep o li e con en , speech signal usually main ains high Signal o Noise
Ra io (SNR), al hough e y o en speech is pa ially occluded by complemen a y acous ic
addi ions such as music, and noises like canned laugh e and applauses.
•Mee ings domain. This scena io implies eco dings om mee ing ooms, whe e an un-
de e mined numbe o people is eco ded om one o mul iple mic ophones. Hence,
eco dings om his domain mainly include con e sa ional speech. In his domain eco d-
ing condi ions a e also e y ele an . Despi e he ac ha close- o-mou h mic ophones
can be used, mo e o en omnidi ec ional mic ophone a ays a e conside ed. These a ays
can be loca ed in a single poin , e.g. on op o he con e ence able, o sp ead along he
oom. Rega dless o he mic ophone loca ions, he dis ance be ween speake and mic o-
phone canno be igno ed. This dis ance is esponsible o no iceable channel e ec s in
he speech p opaga ion up o he mic ophones, including deg ada ions as e e be a ion.
Besides, he s a iona i y o his ansmission channel canno be gua an eed, a ec ed by
he ela i e mo emen s be ween speake and mic ophone. Finally, hese channel e ec s
usually imply powe losses o he signal, making speech quali y mo e sensi i e o noises.
The his o ic e olu ion o dia iza ion o iginally s a ed in he elephone domain. Due o
i s cha ac e is ics his domain p o ided he mos es ic ed e sion o he dia iza ion p ob-
lem. Besides, he e was a g ea in e es o dia iza ion solu ions included in speake ecog-
ni ion applica ions. This is why since 1996 dia iza ion was pa o NIST SRE e alua ions
[P zybocki and Ma in, 2004]. Only a e speake ecogni ion e ol ed i s ools in e ms o accu-
acy and obus ness, dia iza ion was able o expo i s knowledge o al e na i e domains, as in
Rich T ansc ip ion (RT) e alua ions [Ga o olo e al., 2002], whe e al e na i e domains (b oad-
cas news and mee ings) complemen ed he con e sa ional elephone speech.
10
Chap e 2. Dia iza ion S a e o he A
SPK 1 SPK 2 SPK 3 SPK 4 SPK 1
COARSEST PARTITION
FINEST PARTITION BOTTOM-UP
TOP-DOWN
Figu e 2.1: Schema ic o Bo om-Up and Top-Down dia iza ion
2.2 Main dia iza ion s a egies
Along li e a u e se e al op ions ha e been p oposed o he ob en ion o he dia iza ion labels.
Howe e , mos o hese con ibu ions can be g ouped in o wo main concep ual app oaches:
Bo om-Up and Top-Down dia iza ion s a egies. Fig. 2.1 illus a es bo h dia iza ion app oaches
in o de o ob ain he same dia iza ion labels.
•Bo om-Up. The gi en audio is i s di ided in o indi idual segmen s, in which a single
speake is assumed o be p esen . Then, hese segmen s a e clus e ed so all blocks om
he same speake a e agged wi h he same label.
•Top-Down. This al e na i e conside s he opposi e s a ing poin . This app oach s a s
conside ing a single speake esponsible o all he audio. A e wa ds, he ini ial clus e is
di ided ying o ma ch each inal clus e wi h a eal speake in he audio.
In spi e o hei opposi e app oach, bo h s a egies need o sol e he same wo challenges:
De e mining whe he some pa o he audio con ains speech om a single speake and inding
he bounda ies i necessa y. Despi e he appa en simplici y o bo h asks, hei de elopmen
o eal applica ions has equi ed se e al con ibu ions in he li e a u e. Ne e heless, bo h asks
a e s ill a o being o ally sol ed.
Despi e bo h Bo om-Up and Top-Down app oaches a e equally alid, hey a e no simila ly
popula . While bo h op ions ha e been de eloped along mul iple publica ions, in ecen yea s
11
Main dia iza ion s a egies
Figu e 2.2: Gene al schema ic o a dia iza ion sys em
he Bo om-Up s a egy has gained much mo e awa eness han he Top-Down coun e pa . A
eason o his popula i y is he i among he la es imp o emen s in speake ecogni ion and
he Bo om-Up dia iza ion pipeline, making hei inclusion s aigh o wa d. Unde hese ci -
cums ances, Bo om-Up dia iza ion has aken i s pe o mance o unp eceden le els o quali y.
In consequence, his op ion has ecen ly gained popula i y becoming he s anda d dia iza ion
app oach nowadays.
2.2.1 Bo om-Up dia iza ion sys ems
The popula i y o Bo om-Up dia iza ion has inspi ed he de elopmen o a s anda d a chi ec-
u e, which we p esen in Fig. 2.2. This schema ic desc ibes he s anda d conside ed blocks o
ans o m he inpu aw audio in o he desi ed inal labels. The unc ionali y o each block is
desc ibed as ollows:
•Acous ic Fea u e Ex ac ion. Raw speech audio is a e y complex signal wi h many
so s o in o ma ion. While some o hem a e aluable depending on he applica ion
(speake , speech, language, e c.), o he s a e no o in e es (channel, noises, e c.) because
hey can al e ou es ima es. The ea u e ex ac ion s ep aims o ans o m he aw signal
in o a ai h ul bu compac ep esen a ion o he acous ic in o ma ion, simpli ying he
access o ou a ge in o ma ion and compensa ing hose ha m ul deg ada ions.
•Segmen a ion. Gene ally speaking, segmen a ion is he ask o di iding an audio in o
pieces acco ding o an a ibu e, which should emain homogeneous along he o al leng h
o each piece. Focusing on dia iza ion, he di ision a ibu e is he speake iden i y. Thus,
he goal o dia iza ion segmen a ion is he di ision o a gi en audio in o segmen s whe e
a single speake is p esen in hem. This sys em mus exploi he homogenei y o da a
in sho pe iods o ime o ind he bounda ies be ween speake s. An ideal segmen a ion
s ep should p o ide he ime ma ks o he di e en speake in e en ions in an audio.
12
Chap e 2. Dia iza ion S a e o he A
•Speake Cha ac e iza ion. Speake cha ac e iza ion is a high-le el in o ma ion ex ac-
ion which collec s he speake in o ma ion om he acous ic ea u es. In o de o p op-
e ly do so, i equi es wo king wi h audio om a single speake . Thanks o his equi e-
men , highly e ol ed echniques wo k along he gi en inpu segmen s, enhancing hei
speake disc imina i e p ope ies while compensa ing he ha m ul a iabili ies. Mo eo e ,
his p ocess usually con e s a iable-leng h segmen s in o ixed-dimension compac ep-
esen a ions, mo e sui able o pos p ocessing.
•Clus e ing. The ou pu o he segmen a ion s ep is a se o acous ic agmen s wi h a
single speake in each o hem. Howe e , he same speake may ha e p oduced mo e
han one segmen . The clus e ing s age is esponsible o g ouping all hose segmen s
om he same speake and label hem wi h a unique ag. Fo his pu pose, clus e ing
akes he segmen ep esen a ions as inpu , gene a ing he dia iza ion labels as ou pu .
•Resegmen a ion. Resegmen a ion is an op ional ex a segmen a ion s ep o e ine he
ini ial segmen a ion bounda ies. This ex a bo de uning akes ad an age o he in e ed
clus e ing ou pu , wi h an accu a e knowledge abou he e alua ion audio. Resegmen-
a ion ou pu may be conside ed as dia iza ion labels o be edback in o he sys em o
u he e ining.
2.3 Acous ic ea u es o dia iza ion
In o de o di e en ia e speake s, dia iza ion sys ems equi e a subsys em capable o p o iding
disc imina i e cha ac e is ics a each ime s ep o an audio. These cha ac e is ics, also known
as ea u es, should maximize hei classi ica ion capabili ies. Fo his eason, hey y o ep e-
sen he audio in o ma ion in a ac able manne , simpli ying he in o ma ion ga he ing while
educing ha m ul so s o a iabili y (noise, channel in o ma ion, e c.) meanwhile. Mo eo e ,
ea u e ex ac ion is he i s dia iza ion block, hence no assump ion like numbe o speake s
no hei iden i y, speake ansi ions, e c. can be done.
The mos popula ea u es so a a e hose commonly known as sho - e m acous ic ea u es.
O iginally designed o speech ecogni ion, hese ea u es ca y ou a spec al analysis o he
aw signal while inspi ed by bo h he human p oduc ion and pe cep ion sys ems. Because he
speech signal is no s a iona y, his analysis mus be pe o med in sho analysis windows.
The mos popula ea u es a e he Mel F equency Ceps al Coe icien s (MFCCs), o iginally
p esen ed in [Da is and Me mels ein, 1980]. These ea u es p opose a sho - ime analysis o he
13
Audio segmen a ion
DKL(P||Q) = X
x
P(x) log P(x)
Q(x)(2.6)
Un o una ely, i s o iginal de ini ion is no symme ic, i.e., he KL di e gence o Q wi h
espec o P (DKL(P||Q)) may no be he same as he di e gence o P wi h espec o Q
(DKL(Q||P)). In consequence a symme ized e sion, known as KL2 di e gence, is used in-
s ead. This di e gence o dis ibu ions P and Q is de ined as:
DKL2(P||Q) = DKL(P||Q) + DKL(Q||P)(2.7)
Mo ing o dia iza ion, his dis ibu ion is conside ed in segmen a ion in [Siegle e al., 1997]
[Delacou and Wellekens, 2000].
2.4.1.3 Deep Neu al Ne wo ks (DNNs)
Thanks o he e olu ion o neu al ne wo ks many o he asks p e iously ca ied ou by o he
means, such as s a is ics, a e now pe o med by his echnology. Rega ding segmen a ion, some
con ibu ions ha e a emp ed he inclusion o DNNs in his ask.
In [Gup a, 2015] DNNs a e used as classi ie s. The hypo he ical bounda y ame is s acked
along i s con ex window, eeding a monoli hic DNN consis ing o eed o wa d laye s. The
inal laye classi ies he bounda y ame as eal o no . Mo eo e , a likelihood measu e can be
ob ained in he p ocess.
By con as , DNN eg ession capabili ies can also been applied. In [H uz and Zajic, 2017]
he neu al ne wo k mus ca y ou he eg ession o he ansi ion p obabili y, so ened du ing
aining. Fo his pu pose, inpu da a is ea ed by means o s acks o con olu ional neu al
ne wo ks.
In bo h cases, DNNs wo k as s andalone sys ems. Howe e , bo h a chi ec u es i he gi en
mo e gene al de ini ion, whe e a neu al ne wo k p o ides a me ic o a ixed-leng h analysis
window and compa ed agains a h eshold.
2.4.2 Model-based segmen a ion
Despi e he ac ha me ic-based segmen a ions a e he mos popula ones, o he al e na i es
ha e also been p oposed. Conside ing model-based segmen a ions, [Li e al., 2009] conside s
Hidden Ma ko Models (HMMs) o segmen a ion. The model ep esen s each class by means
o a 64-Gaussian GMM. This concep is e ol ed in [Diez e al., 2018], whe e classes a e ep e-
sen ed wi h ied GMMs, mo e sui able o speake ep esen a ion.
20

Chap e 2. Dia iza ion S a e o he A
Finally, some sys ems [Ga cia-Rome o e al., 2017][Diez e al., 2019] wo k in e ms o a
coa se SbC app oach. They p e e wo king wi h e y sho ixed-leng h (a ound 1.5 seconds)
segmen s, no aking ca e o bounda ies. These sys ems ely on he la es speake cha ac e i-
za ion echniques, which ha e e ol ed o p o ide obus enough ep esen a ions when wo king
wi h e y sho segmen s. By doing so, hey alle ia e he compu a ional cos s while only in o-
ducing a small p opo ion o co up ed segmen s: as many deg aded segmen s as eal bound-
a ies. Besides, hese sys ems usually coun wi h esegmen a ion sys ems o elimina e he gene -
a ed dis o ions once speake models a e a ailable.
2.5 Speake cha ac e iza ion
The na u e o speech makes his in o ma ion o ha e a sequen ial na u e. Human beings con-
ca ena e mul iple sounds o ansmi he desi ed in o ma ion. Howe e , he e is no limi a-
ion in e ms o i s leng h no he message. I can ei he be a la ge speech o a sho eply
o a closed ques ion, i.e., "yes" o "no". Mo eo e , i can include all he acous ic uni s o
jus a es ic ed se . Speake ecogni ion echnologies should p o ide a ool o obus ly en-
code he iden i y o he in ol ed speake ega dless o he in a-speake a iabili y, i.e. he
a iabili y wi hin all he possible u e ances om he same speake . Some e iews such as
[Fu ui, 2004][Kinnunen and Li, 2010] p o ide a good o e iew abou he e olu ion o hese
echnologies.
2.5.1 Ea ly days
Some o he i s success ul speake ecogni ion sys ems we e based on he co ela ion o
spec og ams [P uzansky, 1963]. This idea was la e e ol ed o ake in o accoun he o -
man analysis [Dodding on, 1971]. Because hese echniques we e no powe ul enough o
deal wi h ex -independen ecogni ion, some al e na i es we e explo ed o he ollowing
decade. Some p oposals du ing hose yea s a e he ins an aneous spec a co a iance ma ix
[Li and Hughes, 1974], spec um and undamen al equency his og ams [Beek e al., 1977] o
linea p edic ion coe icien s [Sambu , 1972]. The ollowing g ea e olu ion appea ed wi h he
conside a ion o empla e models: Dynamic Time Wa ping (DTW) [Fu ui, 1981] and Vec o
Quan iza ion (VQ) [Rosenbe g and Soong, 1987] [Soong e al., 1985], which p oposes sho
ime ea u e ec o s comp essed in codebooks. This p inciple was la e e ol ed as long as
ma ix quan izie s o mul i- ame we e also p oposed [Juang, 1990].
21
Speake cha ac e iza ion
2.5.2 Model-based ep esen a ions
In he 80s, a g ea e olu ion in he cha ac e iza ion philosophy was p oposed. Ra he han con-
side ing speech as a de e minis ic p ocess whe e ea u es could be measu ed, s a e-o - he-a
con ibu ions s a ed o de ine s a is ical models as gene a o s o speech. Mo eo e , hese gen-
e a o s we e o en designed only aking in o accoun he acous ic in o ma ion, no conside ing
high-le el c a ed ea u es.
The gene a i e so o solu ion has many ad an ages. Fi s , all segmen s a e supposedly
gene a ed by a known dis ibu ion, a pa ame ic solu ion pe ec ly desc ibed by a closed se o
pa ame e s, some o hem speake dependen . Besides, his so o solu ion allows he same
model o wo k wi h a iable-leng h segmen s while p o iding a ixed-dimension speake ep-
esen a ion. Finally, s a is ical solu ions can also p o ide p o ec ion agains di e en ypes o
andomness associa ed wi h he oice (phone ic a iabili y, noises, e c.).
When choosing he dis ibu ion o be e ep esen speake s, Gaussian dis ibu ions a e usu-
ally aken in o accoun . Ve y well known among s a is icians, Gaussian dis ibu ions ha e wo -
hy p ope ies. Howe e , Gaussian dis ibu ions a e oo simple o p ope ly ep esen all he
a iabili y and condi ions in speech. Thus, combina ions o hem, Gaussian Mix u e Models
(GMMs) a e conside ed ins ead. This app oach is conside ed unde he assump ion ha a linea
combina ion o enough Gaussians should be able o ep oduce any dis ibu ion.
2.5.2.1 Gaussian Mix u e Models (GMMs)
Gaussian Mix u e Models a e gene a i e s a is ical models i s in oduced in speake ecogni-
ion in [Reynolds and Rose, 1995]. They a e composed by he weigh ed sum o CGaussian
componen s, each one wi h i s own weigh πc, mean ec o µcand co a iance ma ix Σc, be-
ing c= 1..C. Thus, he sequence O={o1, ..., on, ..., oN}gene a ed by a GMM has been
andomly d awn as:
P(O|M) =
N
Y
n=1
C
X
c
πcN(on|µc,Σc)(2.8)
The e alua ion o hese sys ems wo ked in e ms o a loglikelihood a io. Two loglikelihood
e ms we e conside ed, bo h aking in o accoun he es audio audio es bu conside ing wo
di e en models: A model o he claimed en oll speake (Men oll) and a model ep esen ing
speake s excep o ou en ollmen one (Men oll).
ll = ln P(audio es |Men oll)
P(audio es |Men oll)(2.9)
22
Chap e 2. Dia iza ion S a e o he A
While Men oll was s aigh o wa d, he de ini ion o Men oll was no so clea . Many sys ems
wo ked wi h a pool o coho models, chosen o each ial acco ding o di e en c i e ia.
Then, his idea was e ol ed in [Reynolds e al., 2000], which p oposes he GMM-UBM
pa adigm. Fi s , his con ibu ion in eg a es he coho o al e na i e speake s in o a single
model, esponsible o ep esen he o al a iabili y o he acous ic da a. This gene al model, a
la ge GMM ained wi h se e al speake s, is known as Uni e sal Backg ound Model o UBM.
Fu he mo e, ins ead o building om sc a ch indi idual en oll models Men oll, i p oposes he
op ion o hei cons uc ion as a MAP adap a ion om he UBM, speci ically an adap a ion o
he componen means. The ob ained bene i s a e a igh e coupling be ween models, and as e
sco ing echniques. In his scena io, he p oposed ll was:
ll = ln P(audio es |Men oll)
P(audio es |MUBM)(2.10)
2.5.2.2 Suppo Vec o Machines (SVM)
The GMM-UBM s a egy became a miles one in speake e i ica ion, specially ega ding he
way o model speake s. Howe e , al e na i e sco ing app oaches we e a emp ed. Wi hin
his line o esea ch Suppo Vec o Machines (SVMs) we e p oposed o speake ecogni ion
[Campbell e al., 2006a], leading o he SVM-GMM s a egy.
Suppo Vec o Machines a e bina y classi ie s ha p ojec he inpu da a in o a high-
dimensional space whe e a hype plane sepa a es he wo classes. The e alua ion in SVMs
is de ined as ollows:
(x) =
C
X
c=1
ycαcK(xc, x) + b(2.11)
whe e (x)s ands o he dis ance o he u e ance xwi h espec o he hype plane. xc,αc
and yc ep esen he suppo ec o s, weigh s and labels espec i ely, wi h he es ic ion ha
PC
c=1 ycαc= 0 and αc>0. The labels yc ake he alue +1 o one class and −1 o he o he
one. Besides b ep esen s he hype plane bias. Finally, K(·,·)s ands o he ke nel unc ion, e-
sponsible o p ojec ing he da a in o he high-dimension space and calcula ing dis ances e ms
I K(·,·)is es ic ed o sa is y he Me ce condi ion, he Ke nel condi ion can be exp essed as
an inne p oduc as:
K(x, y) =< g(x), g(y)>(2.12)
whe e g(·)is he ans o ma ion in o he highly dimensional space.
SVMs a e ained by a maximum ma gin s a egy. This ype o aining mus iden i y a
hype plane which p ope ly classi ies he aining elemen s while sa is ying he ollowing e-
23
Speake cha ac e iza ion
s ic ion: he chosen hype plane mus keep he maximum dis ance wi h espec o he aining
popula ions o bo h classes. This eques o ces he hype plane o p o ide he maximum ma gin
p o ec ion agains spu ious da a du ing e alua ion.
The inclusion o SVMs in speake ecogni ion [Campbell e al., 2006a] was ca ied ou by
p oposing a ke nel ha bounds he KL di e gence. In ou scena io, KL di e gence measu es he
dis ance be ween u e ances Oaand Ob, modeled by GMMs Maand Mb espec i ely. Bo h
GMMs a e ob ained acco ding o he GMM-UBM pa adigm, hence hey sha e he componen
weigh s πcand he componen co a iance ma ices Σc, only di e ing a he componen means
µc. The ke nel accomplishing his eques is:
K(Oa,Ob) =
C
X
c=1 √πcΣ1/2
cµa
c√πcΣ1/2
cµb
c(2.13)
In consequence, he ke nel unc ion can be in e p e ed as he inne p oduc o he wo GMM
supe ec o s, a conca ena ion o he GMM means unde going a diagonal scaling. Applied o ou
p e ious de ini ion o SVMs, he en ollmen supe ec o cons i u es he se o suppo ec o s
and he es supe ec o plays he ole o e alua ed u e ance, deciding whe he i comes om
he en ollmen speake .
The GMM-SVM pa adigm was complemen ed wi h he Nuisance A ibu e P ojec ion
(NAP) concep . This idea, in oduced in [Campbell e al., 2006b], conside ed he compensa ion
o he in a-speake a iabili y p esen in he supe ec o s. This compensa ion is pe o med by
es ima ing a low ank ma ix U, also known as eigen-channels ma ix, which de ines he in a-
speake a iabili y wi hin he supe ec o space. Once modeled, a ma ix P=I−UUTcan be
in oduced in he ke nel unc ion al eady seen in eq (2.13):
K(Oa,Ob) =
C
X
c=1 √πcΣ1/2
cµa
cP√πcΣ1/2
cµb
c(2.14)
2.5.2.3 Join Fac o Analysis (JFA)
Join Fac o Analysis (JFA) [Kenny, 2005] is an e olu ion o he GMM ep esen a ions, paying
special a en ion o wo concep s de eloped wi h he GMM-SVM app oach: supe ec o s and
subspaces o ce ain a iabili ies. Taking bo h concep s in o accoun he JFA me hodology
e ol es de he GMM-UBM pa adigm decomposing he adap ed GMM supe ec o as a sum o
e ms:
µj=µUBM +Vyi+Uxk+Dzj(2.15)
24
Chap e 2. Dia iza ion S a e o he A
whe e µjis he adap ed supe ec o mean o u e ance j.µUBM ep esen s he mean supe ec o
om he UBM model. The e m Vyiis he speake dependen e m. Vis a low ank ma ix
desc ibing he subspace o he in e -speake a iabili y while yiis a ied la en a iable, i.e.
a la en a iable whose alue is he same o all u e ances om speake i, esponsible o
he u e ance. Simila ly, we ha e he e m Uxko channel e m. Uis a low ank ma ix
desc ibing he channel a iabili y space and xkis he ied la en a iable o all u e ances wi h
he same channel k, including u e ance j. Finally, we ha e he e m Dzj, which mus explain
he emaining a iabili y. Fo his pu pose, Dis a diagonal ma ix and zja la en a iable unique
o he u e ance. All he h ee la en a iables, yi,xkand zja e S anda d No mal dis ibu ed.
2.5.3 Embedded ep esen a ions
The ollowing la ge e olu ion implied he imp o emen o he al eady p oposed models, bu
also a new me hodology. On he one hand, models including la en a iables o map he speake
in o ma ion signi ican ly imp o ed he pe o mance. On he o he hand, he app oach o he
GMM-SVM showed ha in o ma ion could be ex ac ed om he models and independen ly
ea ed. The combina ion o bo h ideas c ea ed he embedding pa adigm, de ining models ha
cons ain he speake in o ma ion in o a es ic ed space whe e a la en a iable should explain
each speake . F om hese la en a iables we could ex ac compac ep esen a ions, oicep in s
o each speake , also known as embeddings.
Embeddings o e se e al ad an ages compa ed wi h p e ious app oaches. Once embed-
dings a e ex ac ed, hey can be decoupled om he o iginal ex ac ion me hod, simpli ying
hei s o age. Mo eo e , his decoupling makes impossible he e u n o he o iginal audio, gua -
an eeing p i acy. Finally, he ob ained embeddings can be pos p ocessed by al e na i e me hods,
also known as backends. In ac , he cu en speake ecogni ion s a e o he a , om which
mos o hese echnologies a e concei ed, is domina ed by he embedding-backend pipeline.
In he ollowing lines some o he mos popula embeddings a e p esen ed, and one o he
mos popula backends, he P obabilis ic Linea Disc iminan Analysis (PLDA) is explained
a e wa ds.
2.5.3.1 I- ec o s
I- ec o s [Dehak e al., 2011] a e a di ec e olu ion o he JFA modeling. Ra he han di e -
en ia ing be ween speake and channel ac o s, i- ec o model in eg a es hem in o he o al
a iabili y subspace. This usion makes he la en a iable s o e bo h speake and channel in-
o ma ion oge he . Mo eo e , la en a iables a e no linked among u e ances anymo e, being
25

Speake cha ac e iza ion
only ied along he samples om he u e ance. Besides, his model no longe conside s a
esidual a iabili y e m. In consequence, he u e ance j, consis ing o he sequence o ames
O={o1, ..., on, ..., oN}, is now modeled by a GMM whose mean supe ec o µjis de ined as:
µj=µUBM +Twj(2.16)
whe e µUBM again desc ibes he UBM mean supe ec o . Ts ands o a low ank ma ix de-
sc ibing he o al a iabili y subspace and wjis he la en a iable depending on he u e ance.
The men ioned model s ill can be e alua ed in e ms o likelihoods as in JFA. Ne e heless,
his echnology e ol ed o become a oicep in ex ac o . The commonly used i- ec o is he
mean o he pos e io dis ibu ion o he la en wjgi en he u e ance j. This dis ibu ion is
Gaussian and de ined as:
wj∼N (wj|µw,Σw) = Nwj|L−1
wΓw,L−1
w(2.17)
Γw=
C
X
c=1
TT
cΣc
Nj
X
n=1
γnc (on−µc) =
C
X
c=1
TT
cΣcFc(2.18)
Lw=I+
C
X
c=1
TT
c
Nj
X
n=1
γncΣcTc=I+
C
X
c=1
TT
cNcΣcTc(2.19)
whe e µw ep esen s he mean o he pos e io dis ibu ion and Σwis i s co a iance. These
e ms a e cons uc ed in e ms o Tc, he subma ix om Tdesc ibing he con ibu ion o he
c h Gaussian componen , and Σc, he co a iance ma ix o he c h componen in he UBM. The
in o ma ion o he u e ance is con ained in Ncand Fc, he ze o h and cen e ed i s o de Baum
Welch s a is ics o he c h componen espec i ely. Finally, Nj ep esen s he o al amoun o
samples in he u e ance j. Bo h o hem a e ob ained in e ms o he esponsibili ies γnc, he
p obabili y o he n h sample on o be d awn om componen co he GMM-UBM.
2.5.3.2 Hyb id i- ec o s
The la es g ea e olu ion o neu al ne wo ks, a ec ing bo h so wa e and ha dwa e, has be-
come a miles one along mos a i icial in elligence asks. This e olu ion also eached speech
echnologies [Hin on e al., 2012], including speake cha ac e iza ion. In hese asks, a i s ,
his acquisi ion o he new app oaches was smoo h, complemen ing exis ing s a e-o - he-a
echnologies.
A p oposed inclusion o DNNs in i- ec o s was p esen ed in [Lei e al., 2014] as hyb id i-
ec o s. The i- ec o ex ac o p inciple is he same, i.e. i explo es how an u e ance speci ic
model di e s om a UBM due o he unique cha ac e is ics o he u e ance. Howe e , he
26
Chap e 2. Dia iza ion S a e o he A
UBM is no a GMM anymo e. Now his ole is played by a DNN, disc imina i ely ained
o disce n phoneme senones. This neu al ne wo k is now in cha ge o he esponsibili ies γnc
equi ed o compu e he Baum Welch s a is ics Ncand Fc, he unique inpu o i- ec o aining.
Howe e , due o he ac ha no GMM-UBM is in ol ed, γnc now ep esen s he p obabili y o
he ea u e ame on o con ain he he senon cins ead.
An al e na i e p oposal a e phone ic i- ec o s [Viñals e al., 2019d]. This p oposal se s an
o iginal i- ec o model in which he GMM-UBM esponsibili y depends on a p io ac i a ion,
con olled by a DNN phoneme classi ie . Unde his app oach, he se o C componen s is
decomposed in mul iple subse s, each one esponding o indi idual phonemes. By doing so,
pa icula phoneme models become mo e speci ic while educing acous ic unce ain ies.
2.5.3.3 DNN embeddings
The imp o emen s o hyb id i- ec o s we e ou s anding, ou pe o ming pas echnologies. The
esul s in [Sadjadi e al., 2016] p esen ed an unp eceden pe o mance combining DNN pos e i-
o s wi h BNFs. Howe e , echnologies we e s ill su e ing om i- ec o s laws.
The p oposed e olu ion was a cu ing-edge idea. Ra he han e ol ing he gene a i e i-
ec o model, i ains a o ally disc imina i e DNN. In [Snyde e al., 2016] x- ec o s we e
p oposed ollowing his idea: A neu al ne wo k is ain o classi y an audio among a closed se
o speake s. The inpu ea u es i s unde go mul iple ame-le el non-linea ans o ma ions
and hen hey a e pooled in o an u e ance p ojec ion. This p ojec ion goes h ough u e ance-
le el non-linea ans o ma ions be o e i s classi ica ion. The ne wo k is ained o ecognize
a la ge pool o speake s by means o c oss en opy. Gi en a ained ne wo k, he embeddings
also known as x- ec o s a e ex ac ed du ing he o wa d p opaga ion o he in o ma ion, in he
u e ance-le el ans o ma ions.
The g ea pe o mance o x- ec o s has encou aged he communi y o e ol e
o DNNs. Now mul iple al e na i es o x- ec o s a e a ailable, including LSTM
based a chi ec u es [Wang e al., 2018], Wide Residual Ne wo k based embeddings
[Villalba e al., 2019][Viñals e al., 2019d] o e en expanded x- ec o s [Villalba e al., 2019].
2.5.4 P obabilis ic Linea Disc iminan Analysis (PLDA)
PLDA is a s a is ical linea backend. De ined in [P ince and Elde , 2007] as a gene a i e model,
PLDA applies he subspace concep al eady conside ed in JFA, assuming he embedding φjas
a sum o a iabili y e ms:
φj=µ+Vyi+Uxj+ǫj(2.20)
27
Clus e ing
whe e Vyi ep esen s he speake a iabili y e m and Uxj he u e ance a iabili y coun e pa .
Bo h e ms consis o low ank ma ices (Vand U espec i ely), which de ine subspaces o he
la en a iables yiand xj espec i ely. Whe eas he speake la en a iable yiis ied along all
u e ances wi h he same speake , xjis pa icula o each embedding j. We conside hese la en
a iables, yiand xj, o be s anda d no mal dis ibu ed. Addi ionally, he model also includes
an ex a a iabili y e m ǫj o explain he esidual a iabili y in each pa icula embedding. ǫjis
modeled by means o a ze o-mean Gaussian dis ibu ion and diagonal co a iance ma ix D−1.
Finally, µis he cons an speake independen e m.
Al hough his model o e s a closed- o m solu ion, when i s ly applied on embeddings (i-
ec o s a ha ime), i s pe o mance was no signi ica i ely be e . I equi es embeddings
o be Gaussian in o de o p ope ly ob ain i s imp o emen , al hough he ex ac ed i- ec o s
we e a om his dis ibu ion. The mos popula solu ion o his issue is leng h no maliza ion
[Ga cia-Rome o and Espy-Wilson, 2011]. Embeddings, be o e eeding he PLDA model, a e
o ced o eassu e ha i s Euclidean no m is equal o one. This p ocess p ojec s he inpu
embeddings in o a hype sphe e o adius equal o 1. Be o e leng h-no maliza ion, embeddings
should be cen e ed and whi ened. By doing so, he esul ing embeddings a e sp ead along he
hype sphe e a he han being concen a ed in es ic ed egions o he hype sphe e, leading o
mo e disc imina i e capabili ies o he sys ems.
Mul iple al e na i es ha e appea ed o he o iginal PLDA model. The Simpli ied PLDA
(SPLDA) uses he channel and esidual e ms. Ano he al e na i e is he Disc imina i e PLDA
[Cumani e al., 2013a], which ains he same model in a disc imina i e manne . An impo -
an al e na i e is he Hea y-Tailed PLDA (HTPLDA) [Kenny, 2010]. This model was p o-
posed be o e leng h-no maliza ion as a way o deal wi h non-Gaussian embeddings by modi-
ying he p io dis ibu ions. Howe e , a e leng h-no maliza ion i s compu a ional complex-
i y discou aged i s usage. Ne e heless, wi h he ad en o DNN embeddings, a mo e non-
Gaussian han i- ec o s, HTPLDA p o ides small bene i s wi h espec o o he al e na i es
[B umme e al., 2018].
2.6 Clus e ing
The clus e ing s age in a Bo om-Up dia iza ion a chi ec u e is esponsible o he ga he ing o
he acous ic agmen s in e ms o hei speake . This du y can al e na i ely be conside ed as a
labeling ask. Being he audio o Nacous ic segmen s ep esen ed by he se Φ={φ1, .., φN}
o speake ep esen a ions o embeddings, clus e ing mus in e a pa i ion, a se o labels Θ =
{θ1, .., θN}, so ha hose segmen s om he same speake sha e a common label.
28
Chap e 2. Dia iza ion S a e o he A
Table 2.1: Bell numbe Bin e ms o he numbe o elemen s o clus e
Numbe o segmen s NNumbe o pa i ions B
1 1
22
3 5
415
552
6 203
... ...
10 115975
... ...
20 51724158235372
To do so, we i s equi e a measu e o de e mine how a ce ain pa i ion Θ i s he se o
embeddings Φ. This me ic may ha e mul iple na u es, e.g. s a is ical, g aphs, ke nels, e c.
Then, gi en he chosen me ic we mus ind he pa i ion wi h he bes me ic alue. Un o -
una ely, ega dless o he me ic hey all sha e he same di icul y: The bes pa i ion is only
gua an eed o be ob ained i all possible pa i ions a e analyzed, jus choosing he one wi h he
bes me ic. This op ion is usually e e ed as b u e- o ce app oach. Un o una ely, s udies such
as [B umme and de Villie s, 2010] e eal ha he numbe o pa i ions inc ease e y as as
long as he numbe o segmen s o clus e N ises. In ac , excep o e y low alues o N,
b u e- o ce app oaches a e in gene al no iable.
Gi en an audio wi h Nsegmen s, he o al numbe o possible pa i ions is desc ibed by he
Bell numbe B. This numbe , applicable o any clus e ing ask, ep esen s he o al numbe
o independen possible g ouping a angemen s, in ou case pa i ions, o a se o Nelemen s.
This numbe is de ined by means o a ecu en ela ion:
BN+1 =
N
X
n=0 N
nBn(2.21)
B0=1 (2.22)
This numbe inc eases e y as as long as he alue o N ises. In Table 2.1 alues o low
alues o Na e shown.
Acco ding o Table 2.1, e en e y low alues o Nimply huge numbe o candida e pa i-
ions. I we conside e en highe alues o N, e.g. 100 o 200 segmen s o a one-hou TV
show, he numbe o candida e hypo heses o compa e becomes in ac able. Fo una ely, no all
29
Pe o mance me ics
•MISS ERROR (MISS). Speech audio inco ec ly labeled as non-speech. This e m also
measu es he Voice Ac i i y De ec ion (VAD) pe o mance.
•FALSE ALARM ERROR (FA). Non-speech audio in which a speake is conside ed o
be p esen . Voice Ac i i y De ec ion (VAD) pe o mance is a ec ed by his e m as well.
•SPEAKER ERROR (SPK). Speech misclassi ied as gene a ed by an al e na i e speake .
•OVERLAP ERROR (OV). Pe iods o ime when mul iple speake s a e simul aneously
alking. This e o in ol es he es ima ion o he numbe o speake s (unde es ima ion
when no all speake s a e de ec ed and o e es ima ion when non-p esen speake s a e
also labeled) as well as he misclassi ica ion o he in ol ed speake s.
Rega ding he ou e o e ms, clea ly wo o hem, Miss E o and False Ala m, a e e-
la ed wi h segmen a ion, specially he VAD s ep. Wi h espec o he speake and o e lap e o s,
hese e ms mainly depend on he clus e ing s age. Howe e , hey a e ea ed di e en ly. Ac-
co ding o i s de ini ion, he o e lap e o in ol es any in e ence e o when mul iple speake s
a e alking. Thus, i measu es miss, alse ala m and speake e o s o hese pe iods o ime.
Hence o e lap is he mos challenging e o e m, wi h se e al p oposed con ibu ions abou i s
de ec ion (e.g. [O e son and Os endo , 2007][Zelenák and He nando, 2012]) bu wi hou any
unc ional solu ion ye . The e o e, in ce ain e alua ions his e m is ob ia ed o pe o mance
compa isons.
Due o he ac ha e o s a e non-o e lapped, DER can be decomposed on mul iple e ms,
each one e alua ing he deg ada ion due o each ype o e o . The al e na i e de ini ion o DER
is:
DER =LMISS +LFA +LSPK +LOV
L o al
(2.31)
=EMISS +EFA +ESPK +EOV (2.32)
whe e EMISS,EFA,ESPK and EOV a e he DER e o e ms o miss, alse ala m, speake and
o e lap causes espec i ely.
Despi e all i s bene i s ega ding simplici y and decomposi ion o e o , DER p esen s s ong
limi a ions. Ob ia ing he miss e o and alse ala m e ms, di ec ly ela ed wi h he VAD
pe o mance, no u he knowledge can be in e ed om DER abou he speake e o . A simila
sco e is ob ained i some amoun o audio is misclassi ied, ega dless o how many speake s
a e a ec ed. Besides, DER conside s all he audio uni o mly ele an , and e o s in ol ing
36

Chap e 2. Dia iza ion S a e o he A
he same amoun o audio a e equally ha m ul. Howe e , in eal li e nei he he speake s no
hei speech a e equally aluable, in alida ing his conside a ion. This is specially ele an
when speake s do no con ibu e o he audio wi h he same amoun o speech. Fo example,
some e o s may be i ele an o he mos alka i e speake s bu a mo e signi ican o hose
speake s con ibu ing wi h much less speech. The e o e, some c i ics abou DER me ic a e
a ising while he communi y is eage o inding an al e na i e sco e.
In ecen imes some al e na i e me ics ha e also been p oposed o dia iza ion asks. The
Mu ual In o ma ion (MI) me ic was de ined in DIHARD 2018, a dia iza ion e alua ion in
di icul condi ions. The idea behind MI is measu ing how much in o ma ion we ha e abou he
eal labels p o ided ou hypo hesed pa i ion. The p oposed me ic was p oposed o s udy and
was complemen ed by DER, which managed he leade boa d. The me ic is de ined as:
MI =
R
X
i=1
S
X
j=1
nij
Nlog2
nijN
isj
(2.33)
whe e R ep esen s he numbe o clus e s in he e e ence wi h idu a ion each, and Ss ands
o he numbe o hypo hesized clus e s, each one wi h du a ion sj. Besides, he e m nij is he
amoun o speech assigned o he speake iin he e e ence and o he clus e jin he hypo hesis.
Finally, Nsymbolizes he o al amoun o speech.
Ano he al e na i e is he Jacca d E o Ra e (JER). This me ic was p oposed as al e na i e
o DER in DIHARD 2019. The i s s ep in he e alua ion is a mapping among he Rclus e s in
he e e ence and Sclus e s in he hypo hesed pa i ion. This mapping is ca ied ou acco ding
o he Hunga ian algo i hm, so each clus e in he e e ence will be mapped o a mos one
clus e o he hypo hesis and ice e sa. Then o each speake in he e e ence we es ima e:
JER e =F A +MISS
T OT AL (2.34)
whe e T OT AL ep esen s is he amoun o audio p esen in bo h he e e ence clus e e and
he mapped coun e pa . I he e was no pai ed clus e , i s alue would be he o al amoun o
speech o speake e .F A s ands o he o al amoun o speech no p esen in he e e ence
clus e e bu conside ed as pa o he pai ed g ouping. I s alue is 0i no mapping o e
was ca ied ou . Finally, MISS is he amoun o speech p esen in speake e bu no included
in he mapped coun e pa . I e speake has no pai ed clus e , i s alue is equal o T OT AL.
Ha ing de ined he indi idual e ms JER e , he Jacca d e o a e o a eco ding is he
a e age o speci ic Jacca d e o a es:
JER =1
RX
e
JER e (2.35)
37
Pe o mance me ics
Rega dless o he used me ics, hey do no p o ide any clue abou he easons o he
misclassi ica ion o audio. Hence al e na i e me ics should be help ul o be e unde s and
he speake e o . This e o is mainly gene a ed in he clus e ing block, hus clus e ing
me ics, such as he complemen a y clus e ing and speake impu i ies, well desc ibed in
[ an Leeuwen, 2010] a e sui able o his ask.
The clus e impu i y ep esen s how well he clus e s om a hypo hesized pa i ion con ain
audio om a single speake . De ined in e ms o i s clus e pu i y coun e pa , clus e impu i y
is minimized as long as he ob ained clus e s con ain audio om a single speake . Howe e , i is
no obliga o y ha clus e s om he same speake sha e he same label, in alida ing he me ic
o dia iza ion.
Simila ly o he clus e impu i y concep , we can also de ine he speake impu i y. This
new concep desc ibes how well he speech om a speake is agged wi h a single label, and
i is minimized as long as mo e and mo e da a om one speake only equi es a single label.
This me ic is also in alid o dia iza ion because mul iple speake s in he same clus e do no
deg ade he inal sco e.
38
Chap e 3
Analysis o Dia iza ion in B oadcas Da a
Dia iza ion in b oadcas is a complex ask, composed o a la ge numbe o sub asks wo king
oge he , as seen in Chap e 2. Whils mos o he p e iously desc ibed echniques wo k well
in es ic ed condi ions as he elephone domain, dia iza ion in b oadcas da a equi es many
pa icula i ies o be aken in o conside a ion. In his chap e we analyze he b oadcas domain,
emphasizing i s pa icula i ies. Fo his pu pose, we i s in oduce a e e ence dia iza ion sys-
em. This sys em will se e o explo e he wide a iabili y along b oadcas da a a e wa ds. In
his analysis we co e bo h quan i a i e and quali a i e esul s, and s udy how esul s a e a -
ec ed by his unce ain y. Finally, acco ding o he analysis and ob ained esul s, we sugges
he di e en lines o esea ch, some o hem ea ed along his hesis.
3.1 The dia iza ion e e ence sys em
The e e ence sys em conside ed o his analysis is an AHC-based dia iza ion app oach, e y
common in he li e a u e as baseline sys em. This a chi ec u e ollows a Bo om-Up app oach,
i s di iding he aw signal in o segmen s, which a e la e clus e ed acco ding o hei speake
iden i y. Fig. 3.1 illus a es he basic a chi ec u e o he sys em.
In he ollowing lines we explain in de ail he se up o each elemen in ou sys em.
•Fea u e Ex ac ion Fo he audio ans o ma ion in o ea u es, we s ic ly conside
MFCCs, s anda d ea u es in he s a e o he a . Ou MFCC se up includes a 32-band
Mel il e bank, and a inal coe icien educ ion, only conside ing coe icien s C1-C20.
The ene gy in o ma ion is disca ded oo. The in e ed s eam o ea u e ec o s does no
include de i a i es, and unde goes a no maliza ion o i s mean and a iance (CMVN).
•Segmen a ion The ob ained s eam o ea u e ec o s a e he inpu o he segmen a ion
39
The dia iza ion e e ence sys em
Figu e 3.1: Schema ic o ou baseline dia iza ion sys em
s age. In he e e ence sys em he segmen a ion s ep is di ided in o wo independen
sub asks, Voice Ac i i y De ec ion (VAD) o di e en ia e speech/non-speech and Speake
Change Poin De ec ion (SCPD) o ob ain he speake u ns.
– Voice Ac i i y De ec ion (VAD) The in e ence o he VAD mask is done by means
o [Viñals e al., 2018a], using a segmen a ion-by-classi ica ion app oach, which
wo ks in e ms o DNNs. A 2-laye BLSTM DNN, wi h 256 neu ons pe laye ,
is aken in o accoun . Each elemen in he second BLSTM ou pu sequence is p o-
jec ed in o a bina y deciso . Thus, we in e one VAD label pe inpu ame. This
laye is ained and e alua ed in 3-second analysis windows. Whene e he audio
exceeds he window dimension, a sliding window analysis is ca ied ou . This anal-
ysis implies a 3-second window and 2.5 seconds o wa d s ep. Fo he 0.5-second
o e lap pe iod he in e ence wo ks as ollows. he i s 0.25 seconds a e ob ained
om he p eceding window while he emaining 0.25 seconds a e labeled wi h he
in e ence om he ollowing window. This o e lapping design choice is made o
a oid undesi ed windowing e ec s, specially nea he a i icial window bo de s.
– Speake Change Poin De ec ion (SCPD) Fo he SCPD ask we ely on well
known echniques. In his case we op o a SCPD hyb id solu ion led by a me ic-
40
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
based segmen a ion, in pa icula ∆BIC conside ing Gaussian dis ibu ions wi h ull
co a iance ma ix. We ope a e in a sliding window egime ollowing he desc ip ion
in Sec ion 2.4. We make use o an analysis window wi h a minimum leng h o h ee
seconds, and a 0.25-second accumula i e window expansion whene e he window
does no con ain any bounda y. Rega ding he hype pa ame e λ, i is adjus ed ac-
co ding o hose esul s ob ained du ing he de elopmen phase. The me ic-based
solu ion is combined wi h a silence-based s a egy, which assumes all speech/non-
speech ansi ions o be speake bo de s. In ac , hese bo de s a e used as ancho s
o he ∆BIC segmen a ion.
•Speake Cha ac e iza ion The es ima ed segmen s a e hen con e ed in o compac
ep esen a ions, each one summa izing he pa icula ea u e s eam o each segmen .
Among he mul iple op ions desc ibed in Sec ion 2.5, ou choice o he ype o ep e-
sen a ion is he i- ec o . In ou sys em i- ec o s a e in e ed by means o an ex ac o o
256 Gaussians and 100-dimension o al a iabili y ma ix T. The ex ac ed i- ec o s a e
cen e ed, whi ened and leng h-no malized be o e eeding he clus e ing s age.
•Clus e ing The clus e ing s age in ou e e ence sys em is cons uc ed a ound an AHC
app oach, using SPLDA pai wise log-likelihood a io (LLR) as me ic. Ra he han con-
side ing he o iginal AHC ha exhaus i ely e alua ing all simila i ies among clus e s,
we ollow a simpli ica ion desc ibed in [ an Leeuwen, 2010]. This simpli ica ion only
eques s he es ima ion o he ini ial pai wise simila i y among he embeddings. Then,
a each usion i e a ion we app oxima e he exhaus i e simila i ies by app oxima ions
conside ing he al eady es ima ed alues. Among he di e en op ions o ca y ou he
app oxima ion, we op o he UPGMA (unweigh ed pai g oup me hod wi h a i hme ic
mean) app oach [Sokal and Michene , 1958]. Rega ding he clus e ing s op c i e ion, i
is done by means o a h eshold expe imen ally adjus ed du ing de elopmen .
3.2 Analysis o b oadcas da a
B oadcas da a is a ype o domain specially cha ac e ized by i s wide a iabili y. Whene e no
es ic ions a e applied ega ding he shows o analysis, speech p ocessing asks mus be obus
enough o wi hs and a g ea ange o condi ions. Among he di e en asks a ec ed by his
a iabili y we mus ake in o accoun dia iza ion.
The e a e many easons o b oadcas da a o be so mu able. F om di e en eco ding loca-
ions like s udio and ou doo , o he conside ed equipmen . Fu he mo e, ex a ac o s should
41

Analysis o b oadcas da a
be conside ed, such as he acous ic addons, i.e. acous ic a i ac s like laugh e o applause, ha
co up he audio signal. Fo dia iza ion pu poses we will pay a en ion o his a iabili y along
he clus e ing s age. P e ious blocks can be in e p e ed as high-quali y ea u e ex ac o s, hus
clus e ing mus p o ide he knowledge o p ope ly g oup oge he hei ep esen a ions in o de
o ob ain he inal labels. Du ing clus e ing we mus deal wi h wo main ypes o a iabili y: he
one p esen in he acous ic ep esen a ions, he acous ic a iabili y, and he one a ailable in he
inal speake labels, he speake dis ibu ion a iabili y.
The e ec s o bo h ypes o a iabili y di e , specially due o hei in luence along he di-
a iza ion pipeline. The acous ic a iabili y is consequence o he di e en acous ic condi ions
along he di e en audios o in e es . These di e en condi ions cause embeddings om he
same speake o be less homogeneous, making hem less obus o clus e ing pu poses. In con-
sequence, his a iabili y may be esponsible o a deg ada ion o he dia iza ion pe o mance.
Rega ding he speake dis ibu ion a iabili y, we a e e e ing o he numbe o speake s in
an audio and how much speech hey con ibu e wi h. E o s in he es ima ion o he numbe o
speake s a e highly impo an because all he speech p oduced by he a ec ed o a o s will be
misclassi ied. Besides, he la ge is he ange o possible speake s o an audio, he la ge a e
he po en ial e o s in his es ima ion and mo e audio is usually in ol ed. A g ea ac o o ake
in o accoun in his es ima ion is he dis ibu ion o speech along he di e en speake s. Those
alka i e o a o s wi h se e al con ibu ions ha e enough audio o be obus ly modeled and small
misclassi ica ions ha e negligible e ec s on hem. By con as , hose speake s wi h e y ew
speech a e weakly ep esen ed and small e o s may cause hei loss.
We now p esen an analysis abou hese wo ypes o sou ces o a iabili y in b oadcas
da a. Fo his pu pose, we will ake in o conside a ion wo la ge da ase s: Mul i-Gen e B oad-
cas Challenge 2015[Bell e al., 2015] and Albayzín 2018 [O ega e al., 2018]. Bo h da ase s
include a la ge amoun o b oadcas audio con en co e ing a wide a iabili y o shows, gen es,
languages and media.
3.2.1 Mul i-Gen e B oadcas Challenge 2015 (MGB 2015)
This da ase was eleased o he Mul i-Gen e B oadcas challenge in 2015 [Bell e al., 2015].
This challenge aims a p ocessing asks in he b oadcas domain, including ASR, alignmen
and dia iza ion. The da ase consis s o app oxima ely 1600 hou s o B oadcas audio collec ed
om B i ish B oadcas ing Co po a ion (BBC) along ou o i s channels. The o al amoun o
audio in ol es a ound 1200 episodes om 500 di e en shows. All his audio is di ided in o
h ee subse s: ain, longi udinal de elopmen and e alua ion. The subse di ision ies no o
42
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
sha e di ec knowledge among he subse s, hus all episodes om a show a e placed in he same
subse . Aside he audio, he h ee subse s we e dis ibu ed wi h dia iza ion labels. Howe e , he
label accu acy is no uni o m among subse s. Whils he aining subse includes he o iginally
b oadcas sub i les e ined by a ligh ly-supe ised ASR alignmen as me ada a, de elopmen
and e alua ion labels a e manual anno a ions. Finally, MGB 2015 also con ains an ex a subse ,
namely de elopmen , eleased o he ASR e alua ion. This subse consis s o 28 hou s om 47
shows, wi h manually anno a ed VAD ma ks.
3.2.2 Albayzín 2018
Albayzín 2018 is he la es edi ion o he Albayzín e alua ions, he a emp om Red Temá ica
de Tecnologías del Habla (RTTH) o he e olu ion o speech echnologies in hose languages
spoken in he Ibe ian Peninsula. Rega ding dia iza ion, 2018 is he hi d edi ion a e hose hold
in 2010 and 2016.
Fo he 2018 edi ion he e alua ion consis s o app oxima ely 600 hou s o da a om mass
media domain, co e ing wo di e en languages (Spanish and Ca alan) and wo di e en mass
media (TV and adio). The whole da ase was composed by h ee di e en subse s, acqui ed
along he di e en edi ions: F om 2010 edi ion we ha e a ailable 84 labeled hou s o audio
om 3/24 TV channel in Ca alan. These da a a e complemen ed by 2016 da a: 23 hou s o
manually anno a ed audio om b oadcas adio signal om Co po ación A agonesa de Radio
y Tele isión (CARTV) in Spanish. Finally, 2018 edi ion also adds a ound 400 hou s om
b oadcas con en om Radio Tele isión Española (RTVE). The e alua ion di ides he pool o
da a as ollows: Fo aining and de elopmen bo h 3/24 and CARTV a e a ailable, as well as
10 hou s om RTVE wi h manual anno a ions. E alua ion da a consis s o 40 hou s om RTVE
subse .
3.2.3 Acous ic a iabili y
The ichness o audio con en makes bo h MGB 2015 and Albayzín 2018 a sui able choice
o analyze a iabili y in B oadcas da a, s udying how di e en ac o s a ec he pe o mance
o dia iza ion sys ems. Fo he audio a iabili y we will make use o he dia iza ion sys em
pa adigm, specially he SPLDA model.
Acco ding o he PLDA pa adigm, SPLDA models he a iabili y along i s aining da a by
p ojec ing i in wo subspaces, he in e -speake space de ined by he ma ix VVTand he in a-
speake space desc ibed by ma ix W−1. Mo eo e , due o he ac ha bo h ma ices ep esen
co a iances as well, hei analysis can lead o in e es ing in o ma ion abou he a iabili y om
43
Analysis o b oadcas da a
Da ase VVTSubspace (W−1)Subspace
Telephone
SRE 0.50 0.49
B oadcas
MGB 2015 0.19 0.86
Albayzín 2018 0.49 0.58
Table 3.1: T ace analysis o PLDA in e -speake (VVT) and in a-speake (W−1)
subspaces
each ype, in e -speake and in a-speake .
In Table 3.1 we ca y ou an analysis o bo h in e -speake and in a-speake subspaces in
e ms o hei co a iance ma ices. Fo his analysis we s udy he ace o bo h VVTand W−1
ma ices o ou wo b oadcas da ase s o in e es , MGB 2015 and Albayzín 2018. This s udy
app oxima es he o al a iabili y wi hin each subspace, equi alen o add he a iabili y along
each dimension o he subspace as i hey we e independen . This s udy is complemen ed by a
simila analysis o elephone channel da a, which plays he ole o baseline. This baseline anal-
ysis conside s SRE da a, cons uc ing ou models wi h exce p s om SRE04, SRE05, SRE06
and SRE08.
The esul s illus a ed in Fig. 3.1 show a g ea misma ch in e ms o condi ions be ween
elephone and b oadcas da a. While elephone channel p esen s a simila a iabili y con ained
in bo h subspaces, ou b oadcas da abases show a leas a ound 18% ela i e ex a in a-speake
a iabili y. This measu e inc eases up o a 352% ela i e ex a a iabili y in MGB da ase . Thus,
when conside ing he b oadcas domain, we mus ake in o accoun he ollowing ques ion:
Do simila embeddings sha e he same speake s o jus analogous acous ic condi ions?
Fu he mo e, his in a-speake a iabili y is no only caused by di e ences among shows.
In ac , b oadcas da a p esen s a high wi hin-episode a iabili y. This so o a iabili y co e-
sponds o he di e en condi ions in which he audio is eco ded, e.g. he eco ding loca ion
(s udio, ou doo s, e c.), he in ol ed ma e ial (mic ophones, pos p ocessing, ...), and acous-
ic addi ions (laugh e , applauses, e c.). Besides, he speech signal is highly a ec ed by he
p esence o emo ional speech, i.e. he ans o ma ion o he oice in o de o ansmi ex a
in o ma ion such as shou ing (w a h), whispe ing ( ea ) o whining (pain). Rega dless o he
na u e o he a iabili y, i is usually e y co ela ed along ime, emaining he acous ic cha -
ac e is ics s able du ing pe iods o ime ha can con ain mul iple in e en ions om di e en
44
Chap e 3. Analysis o Dia iza ion in B oadcas Da a
❆✁✂✄☎✆ ✝✆✞✆✟✠✡✆☎☛
✶☞ ✷☞ ✸☞ ✹☞ ✺ ☞ ✻ ☞ ✼☞ ✽ ☞ ✾ ☞ ✶☞☞
✶☞
✷☞
✸☞
✹☞
✺☞
✻☞
✼☞
✽☞
✾☞
✶☞☞
❙
❡
❣
♠
❡
♥
❥
✌✍✎✏✍✑✒ ✓
❙✁✂✄✁☎ ✆☎✝✞✟✠ ✡☎✞☛☞
✶✌ ✷✌ ✸✌ ✹✌ ✺ ✌ ✻ ✌ ✼✌ ✽ ✌ ✾ ✌ ✶✌✌
✶✌
✷✌
✸✌
✹✌
✺✌
✻✌
✼✌
✽✌
✾✌
✶✌✌
✍
❡
❣
♠
❡
♥
❥
✎✏✑✒✏✓✔ ✕
a) Pai wise LLR b) G ound u h mask
Figu e 3.2: Sec ion a iabili y example. Fo 100 i s embeddings om a Sp ingwa ch
episode a) SPLDA pai wise LLR simila i y me ic b) G ound u h ela ionship.
speake s. These pe iods wi h s able condi ions usually co espond o he di e en sec ions o a
show. A clea example could be he news, whe e in e en ions om he news eade s in s udio
condi ions a e in e lea ed wi h ou doo connec ions.
In o de o expose his a iabili y we again make use o ou e e ence dia iza ion sys em,
s udying he AHC simila i y ma ix cons uc ed by means o PLDA pai wise log-likelihood a-
io. This ma ix should con ain highe alues o hose elemen s compa ing embeddings om
he same speake , ega dless o he acous ic condi ions. In Fig. 3.2 we illus a e he acous ic
simila i y ma ix among he 100 i s de ec ed segmen s om an episode o he TV show Sp ing-
wa ch, om MGB 2015. In his analysis we co e an app oxima e 25% o he o al de ec ed
in e en ions in he episode, balancing he adeo be ween gene aliza ion and isualiza ion
capabili ies. The segmen s a e s udied in ch onological o de o imeline comp ehension. The
in o ma ion includes wo pa s, he acous ic simila i y and he g ound u h mask. Fo he acous-
ic simila i y we make use o he PLDA pai wise LLR, whe e he elemen ij e eals how simila
a e he embedding iand j. Ligh e colo s indica e highe speake simila i ies and da ke colo s
less p obabili y o sha e he same speake . Rega ding he g ound u h mask, he ij posi ion
in he igu e is whi e i bo h embeddings, iand j, ha e he same speake label, being black
o he wise.
The wo images shown in Fig. 3.2 e eal he capabili ies do disce n be ween speake and
acous ic condi ions a e limi ed. We i s analyze he speake labels o he las 50 embeddings.
Acco ding o he g ound u h, wo speake s a e esponsible o a sequence o in e lea ed u -
e ances, as in a dialog. Howe e , he LLR sco es did no ealize abou ha , p o iding an
45
Conclusions
3.4 Conclusions
The esul s ob ained along he p esen chap e ha e e ealed di e en ac o s o he inhe en
a iabili y in b oadcas da a. In o de o deal wi h he de ec ed unce ain ies, dia iza ion sys ems
mus wo k in he ollowing elemen s:
3.4.1 The clus e ing app oxima ion
Ou e e ence sys em bases i s dia iza ion choices acco ding o an agglome a i e a chi ec u e.
This a chi ec u e is well-known in he communi y due o i s simplici y. Ne e heless, mo e
e ol ed solu ions could ob ain be e dia iza ion esul s. The choice o an al e na i e clus e ing
p ocedu e equi es ha some conside a ions mus be aken in o accoun .
In i s place we mus ake ca e o he me ic o de e mine he quali y o he pa i ion. Resul s
in Fig. 3.2 illus a e he g ea in luence o channel e ec s in he PLDA LLR. Imp o emen s
abou he modeliza ion o he in a-speake a iabili y should lead o g ea bene i s. Mo eo e ,
we can also wo k in he pa i ion measu emen , combining he local in o ma ion conside ed
in AHC (pai wise simila i y) wi h a mo e gene al poin o iew. Thus, al e na i e clus e ing
app oaches should ake in o conside a ion he implica ions o some o he clus e ing choices in
he decision-making p ocess.
Ano he poin o ocus on is he s op c i e ion. Many o he al e na i e clus e ings simul-
aneously wo k wi h mul iple pa i ions, which con ain a wide ange o speake s. Whene e
compa ing hypo heses, biased measu emen s mus be compensa ed p io o i s compa ison in
o de o p e en signi ican deg ada ions.
3.4.2 The quali y o he embeddings
The obse ed undesi ed a iabili y shown in Fig. 3.2 may no be exclusi ely compensa ed du -
ing clus e ing, bu also du ing he embedding ex ac ion. In ac , he mo e disc imina i e is he
in o ma ion in he embeddings, he be e will be i s pe o mance du ing he clus e ing s age.
Un o una ely, embeddings include mo e a iabili y ha is inhe en o b oadcas da a. We
a e e e ing o mo e gene al a iabili y e ms, such as phone ic a iabili y and sho segmen s.
Any imp o emen in he managemen o hese wo a iabili ies would lead o a gene al imp o e-
men in all ypes o dia iza ion, as well as in speake ecogni ion.
52

Chap e 3. Analysis o Dia iza ion in B oadcas Da a
3.4.3 The domain misma ch p oblem
Finally, we mus also co e he domain misma ch. E en i clus e ing sys ems could co e he
a iabili y in di e en domains, hei pa icula cha ac e is ics should equi e some indi idual
adap a ion o an op imal pe o mance.
Fo his eason, we can make use o he domain adap a ion echniques, i.e. adap he p o-
posed solu ion o each one o he domains o in e es . Ano he solu ion could be he opposi e,
ans o ming he e alua ion audio o i he aining condi ions. Wha e e is he solu ion, we
mus ace ano he issue: b oadcas da a includes se e al shows and gen es, hus in-domain da a
o each o hem may be limi ed o jus una ailable.
53
Conclusions
54
Pa II
The Clus e ing P oblem
55
Chap e 4
Clus e ing by means o Fully Bayesian PLDA
The esul s ob ained in Sec ion 3.3 ha e shown he limi a ions o he baseline dia iza ion sys-
em, specially conce ning he clus e ing s age based on an AHC solu ion. The e o e, his poo
pe o mance mo i a es he sea ch o al e na i e clus e ing op ions. One o he main d awbacks
o AHC is he use o local decisions, i.e. decisions aking in o accoun e y li le in o ma ion, as
he pai wise loglikelihood a ios be ween wo single embeddings. Thus, we would a he p e e
a clus e ing me hod whose me ic e alua es he o e all pa i ion. Ano he eques is ha he
op imiza ion p ocess simul aneously op imizes all labels. A solu ion i ing bo h equi emen s
is he clus e ing by means o Fully Bayesian PLDA.
4.1 The Fully Bayesian PLDA clus e ing solu ion
This p oposal o clus e ing was i s p oposed in [Villalba and Lleida, 2014] as an unsupe ised
clus e ing o model adap a ion. Along he ollowing lines we will de ine he model and ex-
plain how i can be used o clus e ing asks, including dia iza ion. In his p ocess we will pay
a en ion o i s Va ia ional Bayes (VB) decomposi ion, key poin in his app oach.
4.1.1 The Fully Bayesian PLDA (FBPLDA) model
The Fully Bayesian PLDA model [Villalba and Lleida, 2014] is a gene a i e s a is ical model
which desc ibes he inpu embeddings in e ms o la en a iables, some o hem ied along all
embeddings om he same speake . Based on he Simpli ied PLDA, he FBPLDA also desc ibes
he embedding φj om he i h speake as:
φj=µ+Vyi+ǫj(4.1)
57

The Fully Bayesian PLDA clus e ing solu ion
µW V
ε
φjyi
θj
πθ
Ni
I
Figu e 4.1: Bayesian ne wo k o he Fully Bayesian PLDA
whe e µs ands o he speake independen e m. V ep esen s he low-dimension ma ix de in-
ing he speake subspace. yiis he speake la en a iable, s anda d no mal dis ibu ed and
common o all embeddings om he speake i. Finally, he emaining unexplained a iabili y
in he embedding jis included by he e m ǫj, which is modeled by means o a ze o mean
Gaussian wi h co a iance W.
The e olu ion o he Fully Bayesian PLDA is ha , in con as o SPLDA, speake assign-
men s o bo h aining and e alua ion a e unknown, using la en a iables ins ead. Thus, a
se o Nembeddings Φ={φ1, ..., φj, ..., φN}is explained by a se o Icandida e speake s,
each one modeled by a speake la en a iable yi om he se Y={y1, ..., yi, ..., yI}. In
o de o map each embedding o i s gene a o speake , we conside he se o la en a iables
Θ = {θ1, ..., θj, ..., θN}. Each one o he θjla en a iables ollows a mul inomial dis ibu ion,
which p oduces a one-ho sample wi h I alues (θj={θ1j, ..., .θij, ..., θIj}). Each one o hese
alues θij ep esen s he assignmen o he u e ance j o he i h speake . Thus, θjwill ha e i s
componen θij equal o one when he i h speake is esponsible o he embedding j, being ze o
o he wise. Taking his assignmen in o accoun , we can model Φin e ms o Yand Θas:
P(Φ|Y,Θ) =
N
Y
j=1
I
Y
i=1 Nφj|µ+Vyi,W−1θij (4.2)
Due o he Bayesian app oach, he speake labels Θa p io i ollow a mul inomial dis ibu-
ion. This dis ibu ion is complemen ed by i s own p io , πθ, which explains he mul inomial
weigh s acco ding o Di ichle dis ibu ion. Besides, he desc ibed Fully Bayesian PLDA p o-
poses an ex a e olu ion. Ins ead o conside ing poin es ima ions o he model pa ame e s (µ,
Vand W), his e olu ion assumes hem o be la en a iables as well. While he mean µand
he columns o he speake ma ix Va e ea ed wi h a Gaussian p io , Wis modeled in e ms
o a Wisha dis ibu ion. Finally, he model also includes a p io a iable ε o he a iable V.
The Bayesian ne wo k desc ibing he whole model is illus a ed in Fig. 4.1.
58
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
The aining p ocedu e o his model is no nea ly as simple as o he SPLDA. The ain-
ing o he la e model wo ks in e ms o he Expec a ion Maximiza ion (EM) algo i hm. This
algo i hm equi es he es ima ion o he pos e io dis ibu ion o each one o he la en a i-
ables in he model (E s ep), upda ing he poin -es ima ion model pa ame e s o maximize he
loglikelihood (M s ep). Howe e , in he FBPLDA model a closed- o m solu ion o each
pos e io is no possible, hus E s ep canno be pe o med. The e o e, he o iginal wo k
[Villalba and Lleida, 2014] also p oposes an al e na i e aining s a egy by means o he Va ia-
ional Bayes (VB) [A ias, 1999][Bishop, 2006].
Va ia ional Bayes is an app oxima ion me hod ha allows o mimic he EM algo i hm
by a a ia ional equi alen . Gi en a model depending on he se o la en a iables Z=
{Z1, ..., Zh, ..., ZH}, VB app oxima es he pos e io dis ibu ion P(Z|Φ)by a ac o ial dis i-
bu ion q(Z) = QH
h=1 q(Zh). Each one o he ob ained ac o s q(Zh)is an app oxima ion o he
eal pos e io dis ibu ion P(Zh|Φ). In o de o ob ain he bes app oxima ion ollowing he
ac o ial es ic ions each ac o q(Zi)mus ollow a dis ibu ion ollowing he ela ionship:
ln q(Zh) = E∀Z , 6=h[ln P(Z,Φ)] (4.3)
Un o una ely, he app oxima ion by means o a ac o ial dis ibu ion has limi a ions. De-
spi e he ac ha he ob ained ac o dis ibu ions q(Zh)only depend on one o he la en
a iables Zh, hey a e no comple ely independen . Taking in o accoun eq. (4.3), some de-
pendencies emain, being each ac o cons uc ed on op o he expec ed alues om he o he
ac o s.
A colla e al e ec o he Va ia ional Bayes app oxima ion is ha he loglikelihood o he eal
model is no longe a sui able me ic. These ype o solu ions wo ks in e ms o he E idence
Lowe Bound (ELBO o L).
L(Φ) = Zq(Z) ln P(Φ,Z)
q(Z)dZ(4.4)
Bo h ELBO L(Φ)and loglikelihood ln P(Φ)a e in e connec ed. In ac , he loglikelihood
e m is he sum o he ELBO e m plus he KL di e gence be ween he ac o ial dis ibu ion
q(Z)and he eal pos e io dis ibu ion P(Z|Φ). We can exp ess his as:
ln P(Φ) = L(Φ) + KL (q(Z)||P(Z|Φ)) (4.5)
In ac , he ELBO and KL e ms a e in e connec ed. The maximiza ion o ELBO makes KL
di e gence o be educed, be e app oxima ing P(Z|Φ)by means o q(Z). As long as his
app oxima ion is mo e accu a e, ou ELBO e m will be a mo e eliable ep esen a ion o he
loglikelihood om he o iginal model.
59
The Fully Bayesian PLDA clus e ing solu ion
Mo ing o he speci ic case o he Fully Bayesian PLDA, he p oposed decomposi ion o
ac o s is desc ibed as ollows:
P(Y,Θ, πθ,µ,V,W,ε|Φ) = q(Y)q(Θ) q(πΘ)q(µ)q(V)q(W)q(ε)(4.6)
Thus, o aining pu poses we can now pe o m an analogous al e na i e o he EM algo-
i hm, now maximizing ELBO. The a ia ional equi alen o he E s ep mus i e a i ely upda e
he di e en ac o s o ob ain he pos e io dis ibu ions, and he analogous M s ep will p oceed
o he poin es ima ion upda e. This p ocess is epea ed un il con e gence.
4.1.2 The clus e ing p ocedu e
The clus e ing echnique by means o he FBPLDA model p oposed in
[Villalba and Lleida, 2014] has a s a is ical backg ound. This app oach assumes ha di-
a iza ion labels Θdia a e hose ha bes explain he embeddings Φ. Then, he way we should
compa e how pa i ions Θexplain he da a Φis he p obabili y P(Θ|Φ), as desc ibed in
Sec ion 2.6.2.
Θdia = a g max
Θ
P(Θ|Φ) = a g max
ΘZP(Z′,Θ|Φ)dZ′(4.7)
whe e Z′ ep esen s he se o all la en a iables in he model (Y,πθ,µ,V,Wand ε) excep
o Θ.
The applica ion o his app oach o he FBPLDA model is no s aigh o wa d. The same
di icul ies du ing aining a e p esen in his app oach, hus we again mus ely on ou VB
decomposi ion. In consequence, we mus wo k in e ms o app oxima ions as ollows:
Θdia = a g max
ΘZq(Y)q(Θ) q(πθ)q(µ)q(V)q(W)q(ε)dZ′= a g max
Θ
q(Θ) (4.8)
Despi e ha ing simpli ied he clus e ing s ep o he maximiza ion o a single ac o q(Θ),
he same di icul ies as du ing aining emain. The p oposed ac o s a e s ill in e connec ed
by means o expec a ions, so he op imiza ion o he ac o q(Θ) needs o he ac o s o be
op imized as well. Un o una ely, hese o he ac o s also depend on q(Θ). Fo his eason
we wo k in e ms o an i e a i e upda e o ac o s in which, s a ing om an ini ial s a e, we
each a maximum ELBO. This i e a i e p ocess is simila o he conside ed EM p ocedu e
du ing he SPLDA aining. Howe e , in his occasion no poin es ima ion equi es upda e
( hese only a ec he model pa ame e s µ,V,Wand ε), hus we only conside he E s ep.
60
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
Figu e 4.2: Clus e ing schema ic based on label ini ializa ion and FBPLDA eseg-
men a ion
Mo eo e , because we assume he model pa ame e s µ,V,Wand ε o be pe ec ly uned, we
exclusi ely ee alua e q(Y),q(Θ) and q(πΘ). This clus e ing p ocedu e can be in e p e ed as
a wo-s ep sea ch: The se o embeddings Φis dis ibu ed along a se o clus e s Ydu ing he
upda e o he ac o q(θ). Then, he same clus e s Ya e ee alua ed in e ms o he ecen ly
es ima ed Θdu ing he es ima ion o q(Y). This i e a i e p ocess may be easily unde s ood as
ollows: A he beginning o each i e a ion, acco ding o he cu en alue o he speake labels
Θwe cha ac e ize each one o he conside ed Iclus e s. This cha ac e iza ion is done by he
ee alua ion o he speake la en a iables Ydu ing he upda e o he ac o q(Y). Once he
clus e s a e ede ined, each embedding is hen assigned o he mos likely clus e when q(Θ) is
ee alua ed again.
The main disad an age o his app oach is he need o some ini ializa ion Θ0. This ini-
ializa ion can be ob ained in se e al ways, ei he om some p io knowledge o mo e o en
elying on he same embeddings Φ. Hence he p oposed clus e ing s age ollows he schema ic
ep esen ed in Fig. 4.2.
This clus e ing s a egy can be in e p e ed as a wo-s ep clus e ing: A i s block is in cha ge
o ob aining an ini ial pa i ion Θ0, which is e ined a e wa ds by he FBPLDA clus e ing
app oach. Apa om he bene i s due o label eassignmen , his eclus e ing by means o he
FBPLDA o e s ano he ad an age: an es ima ion abou he speake numbe . q(Θ) dis ibu es
he embeddings Φalong Ia p io i candida e speake s. Ne e heless, i is no obliga o y ha
all candida es gene a e a leas one embedding. Those candida e speake s wi hou assigned
embeddings could be elimina ed as pa o he s op c i e ion.
61
Analysis o FBPLDA pe o mance
e s in he ini ial pa i ion Θ0and hose ob ained a e he esegmen a ion. Whene e Θ0con ains
less speake s han he g ound u h (posi i e ela i e speake s), he FBPLDA eclus e ing does
no disca d any single speake , only eassigning he embeddings along he di e en a ailable
candida e speake s. By con as , whene e he ini ializa ion con ains mo e speake s han he
g ound u h numbe , he algo i hm s a s disca ding ew speake s (1-3 speake s on a e age)
al hough he ejec ion o ex a speake s is no enough, signi ican ly o e es ima ing he numbe
o speake s in an audio. In consequence, a bad es ima ion o he speake numbe is di icul o
be ixed by his esegmen a ion.
Apa om he numbe o speake s, dia iza ion is a ec ed by o he ac o s. Ano he impo -
an conside a ion o ake in o accoun is he chance o losing eal speake s. In many occasions
he wo hy speake s a e no hose who speak he mos bu hose wi h ew bu ele an in e en-
ions. An example may be he alk shows, whe e he mos alka i e indi idual is he mode a o
despi e he eal aluable con ibu ions come om he emaining speake s. Rega ding he FB-
PLDA solu ion, ou expe imen al wo k has e ealed ha i is usually eluc an o conside small
se s o embeddings (some imes a single one) as an independen speake , op ing o assuming
hem as spu ious da a om a much la ge clus e . This end o using small clus e s wi h la ge
ones is mo e ele an as long as he balance o da a becomes odde . By con as , when wo
clus e s o simila size p esen audio om he same speake , he algo i hm is unlikely o use
hem oge he .
These limi a ions abou how he VB solu ion handles he esegmen a ion specially a ec s
o low- alka i e speake s. They con ibu e e y li le o he eal audio, bu depending on he
applica ion, hei loss is no a o dable. The e o e, we s udy he numbe o los speake s acco d-
ing o he ini ial pa i ion. The ob ained esul s a e shown in Fig. 4.7. Again he pa i ions a e
iden i ied in e ms o ela i e numbe o speake s ∆I. The esul s a e also shown in e ms o
he i s , second and hi d qua ile.
Acco ding o Fig. 4.7 clea subclus e ing ini ializa ions (we assume up o 20 ex a speake s)
lead o he loss o 2-3 speake s on a e age. This end seems s eady a e 12 ex a speake s in
ou ini ial pa i ion Θ0. This esul is specially in e es ing when compa ed wi h Fig. 4.6, which
shows a g owing o e es ima ion o he speake numbe , p opo ional o hose p esen in he
ini ial pa i ion. The combina ion o bo h sou ces o in o ma ion leads o he conclusion ha we
a e usually losing eal speake s, and mos o he o e es ima ion o speake s is a consequence o
he unde clus e ing o he emaining ones.
68

Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✵
✺
✶✵
✶✺
✷✵
✷✺
✸✵
✲
✷✵ ✲
✶ ✲
✶✷ ✲✁ ✲✂ ✵✂ ✁ ✶✷ ✶ ✷ ✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✄ ■
▲
♦
☎
✆
✝
✞
✟
✠
✡
✟
☛
☎
☞✌✍ ✎ ✏✑ ✒✓ ✔✒✕✍ ❢✌ ✕ ❱❇ ✍✌✖ ✉ ✎✗ ✌ ♥
Figu e 4.7: Los speake s acco ding o he ela i e numbe o speake s ∆Iin he
ini ial pa i ion Θ0. Resul s ob ained wi h Albayzín 2018.
4.2.3 Numbe o speake s s DER
Dia iza ion pe o mance does no exclusi ely depend on he in e ed numbe o speake s. Ac-
ually, a good es ima ion abou he numbe o speake s is no always ep esen a i e o a good
dia iza ion in e ms o DER. This is because DER is highly dependen on a p ope iden i ica-
ion o he la ges clus e s in he analysis audio. Thus, as long as eally alka i e speake s a e
p ope ly clus e ed, any ea men o non- alka i e speake s may be conside ed bene icial, e en
hei disca d.
Ou nex expe imen s udies he ela ionship be ween he ini ializa ion and he DER pe o -
mance measu e. Fo his expe imen we ha e e alua ed he in e ed pa i ions ob ained om
he FBPLDA eclus e ing o mul iple le els o he AHC dend og am. The ob ained esul s a e
illus a ed in Fig. 4.8, ep esen ing he ob ained DER in e ms o he ela i e numbe o speake s
in he ini ializa ion Θ0. The ep esen ed in o ma ion simul aneously analyzes he ini ializa ion
AHC sys em (Fig. 4.8a) as well as he AHC block ollowed by he eclus e ing s age (Fig. 4.8b).
The esul s in Fig. 4.8 show ha any unde es ima ion abou he numbe o speake s is e y
ha m ul o bo h dia iza ions, AHC and he FBPLDA eclus e ing. This deg ada ion is mo e
se e e as long as he unde es ima ion inc eases. These esul s a e easonable om he DER
pe spec i e, because se e e losses o speake s will de ini ely cause he misclassi ica ion o e y
69
Analysis o FBPLDA pe o mance
✵
✶✵
✷✵
✸✵
✹✵
✺✵
✻✵
✼✵
✽✵
✲✷✵✲
✶✻ ✲
✶
✷ ✲✽ ✲
✹ ✵ ✹ ✽✶✷✶✻ ✷✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✁■
❉
❊

✭
✪
✮
✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✠♦✡✉☛☞♦♥
✵
✶✵
✷✵
✸✵
✹✵
✺✵
✻✵
✼✵
✽✵
✲✷✵✲
✶✻ ✲
✶
✷ ✲✽ ✲
✹ ✵ ✹ ✽✶✷✶✻ ✷✵
❘❡❧❛ ✐ ✈❡ ❙♣ ❡❛❦❡ s ✁■
❉
❊

✭
✪
✮
✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✰ ❱❇ ✠♦✡✉☛☞♦♥
a) AHC esul s b) AHC + FBPLDA esul s
Figu e 4.8: DER (%) esul s o a) AHC and b) FBPLDA in e ms o he ela i e
numbe o speake s ∆I. Resul s ob ained om Albayzín 2018, indica ing he i s ,
second and hi d qua ile pe bin.
alka i e speake s. In e es ingly, acco ding o Fig. 4.8 when an o e es ima ion abou he num-
be o speake s should happen o occu , he end o DER is no so deg aded, wi h small loses
o pe o mance in AHC and no no iceable deg ada ion when he FBPLDA eclus e ing is ap-
plied. In eal li e applica ions his o e es ima ion scena io may be wo hy enough, specially
conside ing semi-supe ised applica ions. By means o au oma ic echniques dia iza ion labels
wi h mul iple pu e clus e s pe speake could be easily ob ained, only equi ing li le human
supe ision o ma ch hose clus e s wi h a common speake . This manual wo k is simple han
cleaning clus e s wi h mul iple speake s, i.e. he scena io in which an unde es ima ion o he
speake numbe is done.
An al e na i e analysis is he sea ch o he pa i ion ha p o ides he bes dia iza ion esul .
Fo his analysis each episode has been dia ized wi h mul iple ini ial pa i ions Θ0, all ob ained
om he AHC dend og am. The esul s, shown in Fig. 4.9, a e compa ed in e ms o he ela i e
speake o he ini ializa ion wi h espec o he g ound u h.
Acco ding o Fig. 4.9, dia iza ion esul s end o p e e an o e es ima ion o he numbe
o speake s, in e ing mo e speake s han hose p esen in he e e ence labels. Mo eo e , he
o e es ima ion can be signi ica i e, wi h many episodes (mo e han 90% o he episodes) wi h
a leas 5 ex a speake s ob aining he bes DER esul s. These esul s i wi h hose p e iously
ob ained in Fig. 4.8, which showed ha an unde es ima ion o he speake s would lead o sig-
ni ican deg ada ions.
The choice o he bes ini ializa ion is a g ea challenge. The choice o he ini ial pa i ion
leading o he minimum DER may be di icul o e en impossible. E en i his op ion was easi-
70
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✵
✷✵
✸✵
✹✵
✁✵
✻✵
✼✵
❘❛♥❦ ♦❢ ❘❡ ❧ ❛ ✐✈❡ ❙♣ ❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎ ③✑✏✎✓ ✍ ✔✓✕ ❇✖✗✏ ❉❊✘
Figu e 4.9: Dis ibu ion o he ini ializa ion wi h bes DER in e ms o he ela i e
numbe o speake s. Resul s ob ained om Albayzín 2018 and p esen ed acco ding o
5 bins.
ble, i can imply such an una o dable compu a ional cos . The e o e, we migh some imes seek
a adeo , assuming ce ain deg ada ions in pe o mance i a la ge simpli ica ion o he sys ems
is achie ed. In Fig. 4.10 we analyze he p opo ion o ini ial pa i ions whose dia iza ion ou pu
di e s om he bes esul up o a maximum bound. This igu e is composed o wo di e en
dis ibu ions, Fig. 4.10a showing he chance o a 1% DER bound and 4.10b illus a ing he
dis ibu ion o a 3% DER bound.
Fig. 4.10 illus a es ha he p obabili y o a pa i ion o be unde a DER deg ada ion bound
is e y educed. Only an app oxima e 17% o he ini ializa ions each unde he 1% DER bound,
being o e 33% wi h a highe bound (3% DER).
4.2.4 Numbe o speake s s ELBO
In hese lines we wan o explo e ma ke s o de e mine whe he we a e wo king wi h a good
pa i ion. E en i a closed se o pa i ions is p o ided, e.g. he mul iple le els o he AHC
dend og am, we need an unsupe ised op ion o compa e hem and decide which one is ou
bes op ion.
Because we a e aking in o accoun a s a is ical solu ion, a ai op ion should be he loglike-
71
Al e na i e ini ializa ions
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✁
✵
✁
✷✵
✷✁
✸✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐ ✈❡ ❙♣❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✶✙ ❇✓✚✍✛
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✁
✵
✁
✷✵
✷✁
✸✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐ ✈❡ ❙♣❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✙✚ ❇✓✛✍✜
1% Bound 3% Bound
Figu e 4.10: Dis ibu ion o he ini ializa ion wi h bounded DER, a) 1% and b) 3%,
in e ms o he ela i e numbe o speake s. Resul s ob ained om Albayzín 2018 and
p esen ed acco ding o 5 bins.
lihood. Ac ually, because ou model is sol ed by means o a Va ia ional Bayes app oxima ion,
we should conside ELBO ins ead. Conside ed as an app oxima ion o loglikelihood, ELBO is
s ill a good ep esen a i e numbe abou how well he pa i ion ep esen s he inpu da a Φ.
In he nex expe imen we s udy he eliabili y o ELBO as pa i ion selec ion c i e ion,
explo ing which ini ializa ion ob ains he bes ELBO a e FBPLDA eclus e ing is pe o med.
This esul will indica e us which esul s a e mo e s a is ically eliable. Due o ange issues, we
ep esen ou esul s in Fig. 4.11 as a his og am illus a ing he dis ibu ion o maximum ELBO
in e ms o he ela i e speake numbe o he ini ializa ion Θ0.
Acco ding o he esul s in Fig. 4.11, ELBO is a g ea indica o abou he numbe o speake s,
op ing o small de ia ions (±5 speake s) wi h espec he g ound u h in almos 50% o he
in ol ed da a. Howe e , ano he 25% o pa i ions op imize ELBO by o e clus e ing up o 15
speake s. This is specially undesi able when conside ing Fig. 4.10, which eques s he opposi e
(unde clus e ing) o a be e dia iza ion pe o mance.
4.3 Al e na i e ini ializa ions
In he p e ious lines we ha e s udied he g ea impac o he ini ializa ion on he pe o mance
in he FBPLDA eclus e ing solu ion. I s in luence ex ends o mul iple ac o s such as o e -
all quali y, he es ima ed numbe o speake s, how many speake s we may lose o he o e all
ELBO.
72
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁
✵
✵
✷✵
✸✵
✹✵
✁✵
✻✵
✼✵
❘❛♥❦ ♦❢ ❘❡❧❛ ✐✈❡ ❙♣ ❡❛❦❡ s ☎■
P
✆
✝
✞
✝
✆
✟
✠
✝
✡
✝
☛
❆
✉
❞
✠
✝
☞
✭
✪
✮
❊▲❇❖ ✌❝❝✍✎✏✑✒❣ ✓✍ ✔✒✑✓✑✌✕✑③✌✓✑✍✒
Figu e 4.11: Dis ibu ion o he pa i ion wi h bes ELBO in e ms o ela i e speake s.
Resul s ob ained wi h Albayzín 2018 and ep esen ed acco ding o 5 bins.
All his acqui ed knowledge was ex ac ed o be e deal wi h he ini ializa ion issue. We
expec o ind an al e na i e ini ializa ion app oach wi h espec o ou i s FBPLDA app oach
(Table 4.1), whe e a h eshold de e mines he le el o he dend og am in he AHC, e ined
a e wa ds by he FBPLDA.
Acco ding o he al eady seen in o ma ion, we illus a e wo di e en app oaches. The i s
one seeks an e icien adeo be ween DER imp o emen and compu a ional cos . The second
op ion ies o each he bes possible esul s despi e alling in o mo e elabo a ed s a egies wi h
a highe compu a ional cos .
4.3.1 Compu a ionally e icien ini ializa ion
The compu a ionally e icien ini ializa ion app oach was de eloped acco ding o he in o ma-
ion in Fig. 4.10. The illus a ed in o ma ion e eals ha hose ini ial pa i ions whose eseg-
men a ion di e s om he bes esul below a bound a e p one o con ain signi ican ly mo e
speake s han he uned ini ializa ion.
Ou i s al e na i e simply p oposes assuming an ini ializa ion whose numbe o speake s is
gua an eed o o e come he g ound u h, hus exploi ing his ci cums ance. By doing his, we
assume an ini ializa ion in which he AHC algo i hm is mo e unlikely o ha e made signi ican
e o s, and exploi he s op c i e ia om he FBPLDA algo i hm. The g ea bene i o his
app oach is ha is compu a ionally e icien , equi ing as much ime as ou FBPLDA baseline.
In Table 4.4 we analyze he impac o his app oach conside ing di e en alues o he
uppe numbe o speake s. The expe imen includes bo h de elopmen and es subse s om
73

Al e na i e ini ializa ions
Expe imen De . DER(%) E al. DER(%)
MGB 2015
50 Speake s 26.08 42.22
75 Speake s 26.13 41.37
100 Speake s 26.55 42.25
200 Speake s 29.68 45.93
300 Speake s 29.99 44.95
Fines pa i ion 29.24 44.76
Albayzín 2018
50 Speake s 16.48 18.36
100 Speake s 22.60 25.38
150 Speake s 25.94 28.14
Fines pa i ion 60.48 62.54
Table 4.4: DER (%) esul s om AHC ini ializa ion wi h a maximum numbe o
speake s. Included mul iple maxima and he ines pa i ion wi h one segmen pe
clus e . Resul s shown o de elopmen and es subse s om MGB 2015 and Albayzín
2018 co po a.
bo h MGB 2015 and Albayzín 2018. Apa om ixed numbe o speake s o all episodes,
ou esul s also include he case o he ines AHC pa i ion, i.e. one embedding pe candida e
speake , as limi case.
The esul s in Table 4.4 show ha a es ic ed numbe o speake s in bo h da ase s may sig-
ni ican ly imp o e he esul s, specially compa ed wi h Table 4.1. This may be a consequence o
he di e en audio cha ac e is ics among shows, making h esholds on op o simila i y me ics
inapp opia e. Howe e , i is impo an o no ice ha no all ini ializa ions wi h a highe numbe
a e equally use ul. As long as he alue becomes highe , e o a ios s a appea ing. This deg a-
da ion can be aken o he limi conside ing he ines ini ializa ion. The e o e, his app oach
needs o be highe han he g ound u h alue and no oo high o alling in o deg ading he
pe o mance. While in ou app oach we assume he same alue o all episodes, hei du a ion
is no he same. Thus, mo e elabo a ed al e na i es may adjus his alue acco ding o he audio
du a ion.
4.3.2 ELBO-based ini ializa ion choice c i e ion
The s a is ical na u e o he FBPLDA clus e ing solu ion makes easonable he use o al e na i e
ini ializa ion choices. F om a s a is ical poin o iew, he bes pa i ion ΘDIAR should be he
one ha bes explains he gi en da a Φ, as shown in eq. 2.28. The adi ional way o measu e
74
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
Figu e 4.12: Schema ic o dia iza ion based on he simul aneous e alua ion o K
di e en ini ializa ions. The inal pa i ion is selec ed by means o PELBO.
how well some model ep esen s he gi en da a is by means o he pos e io loglikelihood.
When adap ing his app oach o he FBPLDA model we mus deal wi h he VB na u e o
ou solu ion. This solu ion subs i u es he o iginal likelihood e m by he ELBO e m. Thus,
ollowing a simila app oach ou bes dia iza ion pa i ion should be he one ha maximizes he
ELBO L e m as ollows:
ΘDIAR = a g max
ΘL(Θ,Φ)(4.9)
Ne e heless, his idea is no comple e ye . When conside ing mul iple ini ializa ions, no
all o hem ini ially suppose he same numbe o speake s. Hence he highe he numbe o
speake s, he mo e likely his da a can o e i o he e alua ion da a. The e o e, as p esen ed
in [Viñals e al., 2018a], a penalized ELBO (PELBO) e m is p oposed ins ead. This app oach
is inspi ed in BIC, in whe e he likelihood o he model is penalized in e ms o he modelling
capabili ies. Then, he bes labels can be ob ained as:
ΘDIAR = a g max
Θ
PELBO(Θ,Φ) = a g max
Θ
(L(Θ,Φ)−λQ(Θ)) (4.10)
whe e Q(Θ) ep esen s he conside ed excess o modeling capabili ies due o he o al amoun
o speake s in he pa i ion Θ. This e m is mul iplied by a ine uning pa ame e λ.
This pa i ion choice me hod can be applied igh a e he AHC algo i hm, in o de o
choose a single pa i ion. Howe e , due o he g ea e ining p ope ies o he FBPLDA eclus-
e ing, we op ed o choosing among eclus e ed pa i ions. The e o e, a se o Kdi e en ini-
ializa ions a e simul aneously eclus e ed, choosing among hem he inal pa i ion a e wa ds.
This clus e ing s uc u e is ep esen ed in Fig. 4.12.
In Table 4.5 we illus a e he esul s ob ained wi h his algo i hm. The esul s include DER
ma ks o bo h MGB 2015 and Albayzín 2018 da ase s, including bo h de elopmen and es
subse s. Two di e en esul s a e included, a single esul exclusi ely in e ms o ELBO, and
he penalized ELBO as well.
75
Conclusions
Expe imen De . DER(%) E al. DER(%)
MGB 2015
ELBO 26.82 39.12
PELBO 25.95 39.88
Albayzín 2018
ELBO 14.48 17.77
PELBO 13.90 17.79
Table 4.5: DER (%) esul s o ELBO and PELBO ini ializa ion choice. Resul s
shown o de elopmen and es subse s om bo h MGB 2015 and Albayzín 2018
co po a
The ob ained esul s a e signi ican ly be e han ou baseline sys em, and also o e comes
he p e iously desc ibed e icien solu ion. Mo eo e , bo h ELBO and penalized ELBO show
e y simila esul s, illus a ing he obus ness o he app oach. Un o una ely, while penal-
ized ELBO helps o imp o e simple ELBO du ing aining, in e alua ion sligh ly deg ades in
pe o mance, pa ially illus a ing he high domain misma ch be ween shows.
4.4 Conclusions
In his chap e we ha e explo ed he addi ion o he FBPLDA eclus e ing o he baseline di-
a iza ion sys em desc ibed in Sec ion 3.1. Acco ding o he ob ained esul s, he pe o mance
in bo h b oadcas dia iza ion da ase s has been signi ican ly imp o ed due o his block. These
imp o emen s a ec bo h he dia iza ion me ic DER as well as he es ima ion o he speake
numbe .
Mo eo e , we ha e explo ed some o he FBPLDA limi a ions. Ou analysis has e ealed
a g ea dependence o pe o mance acco ding o he ini ializa ion. Thus, he be e he ini ial-
iza ion, he be e is he label e inemen by means o FBPLDA. Besides, ou s udy also e eals
ha ini ializa ions should be e o e es ima e he numbe o speake s (a ound 10 ex a speake s
compa ed o he o acle alue) so ha FBPLDA p o ided he bes labels. Howe e , despi e he
imp o emen s ob ained in his o e es ima ion scena io, we s ill mus assume deg ada ions as
he loss o low alka i e speake s. Fu he mo e, ou analysis has de e mined ha he E idence
Lowe Bound (ELBO) seems a easonable indica o o in e he numbe o eal speake s wi hin
an audio.
Finally, we ha e explo ed he join collabo a ion o he AHC ini ializa ion and he FBPLDA
eclus e ing a he han assuming hem as independen blocks. Thus, we p oposed wo success-
76
Chap e 4. Clus e ing by means o Fully Bayesian PLDA
ul app oaches whe e a limi ed numbe o le els o he AHC dend og am a e e ined by he
FBPLDA, which is also esponsible o he s op c i e ia. Whils ou i s app oach explo ed an
e icien sea ch by assuming a single ini ial pa i ion eassu ing an o e es ima ion o i s speake
numbe , ou second s a egy ca ies ou a simul aneous eclus e ing o mul iple ini ializa ions,
op ing o one o he ob ained pa i ions acco ding o ELBO. Bo h o hem ha e demons a ed
ha FBPLDA eclus e ing is a mo e powe ul s op c i e ia han AHC, ob aining signi ican im-
p o emen s. Wi h espec o he compa ison be ween hem, he ELBO s op c i e ia is able o
ou pe o m he o e es ima ion s a egy, a he cos o inc easing he compu a ional cos s.
77
PLDA wi h Unce ain y P opaga ion (PLDAUP)
Expe imen EER (%) minDCF
Long-Long 3.37 0.161
Long-Sho 5.98 0.291
Sho -Long 5.98 0.283
Sho -Sho 8.76 0.403
Table 5.1: EER (%) and minDCF esul s wi h SPLDA o SRE10 co eex -co eex
de 5 emale wi h in ol ed sho u e ances
Sho ) o each ole (en ollmen and es ). In Table 5.1 we ep esen he pe o mance o each
o hese combina ions. The pe o mance is e alua ed acco ding o wo di e en me ics, Equal
E o Ra e (EER) and minimum De ec ion Cos Func ion (minDCF).
Acco ding o he ob ained esul s, we obse e a se e e deg ada ion in pe o mance as long
as sho u e ances a e conside ed. These deg ada ions a e no iceable when sho u e ances
a e in ol ed, ega dless o hei ole. I bo h oles, en ollmen and es , a e played by sho
u e ances he deg ada ion is much mo e signi ica i e. Fu he mo e, he seen deg ada ion is no-
iceable in bo h me ics, EER and minDCF. This in o ma ion is complemen ed by DET cu es,
shown in Fig. 5.2.
DET cu es in Fig. 5.2 con i m hose esul s ob ained in Table 5.1, showing clea di e ences
in pe o mance when sho u e ances play any ole in e i ica ion. Besides, his deg ada ion
is mo e no iceable when sho u e ances simul aneously play he en ollmen and es oles.
Finally, his beha iou is consis en along he whole cu es, ega dless o he ope a ion poin .
The p e ious expe imen illus a es he impac o sho u e ances in s a e-o - he-a ech-
nologies and jus i ies he sea ch o al e na i e echniques as PLDAUP. This app oach is e alu-
a ed in ou nex expe imen , sco ing he same ials wi h long and sho u e ances. Howe e , in
his expe imen we es ic ou e alua ion o wo o he p e ious scena ios: Long-Sho (o iginal
long u e ance as en ollmen and sho u e ance as es ) and Sho -Sho (bo h en ollmen and
es a e sho u e ances). These wo condi ions a e e alua ed wi h ou new PLDAUP model,
unde going wo di e en al e na i es o mimic leng h-no maliza ion: scala no maliza ion and
unscen ans o ma ion. The ob ained esul s a e shown in Table 5.2, including bo h EER and
minDCF esul s.
Those esul s in Table 5.2 illus a e he bene i s due o he inclusion o Unce ain y P opa-
ga ion. Howe e , he ob ained imp o emen s do no a ec he pe o mance in he same way.
While EER is clea ly mo e a ec ed (a leas 12% ela i e imp o emen s), minDCF bene i s
a e mo e negligible. Besides, bene i s a e ob ained ega dless o he leng h-no maliza ion ap-
p oxima ion o ma ices, al hough minimum ex a imp o emen s a e ob ained wi h unscen
84

Chap e 5. Unce ain y P opaga ion o Dia iza ion
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠✆✝✞✟
✆✝✞✟✠✡☛✝☞✌
✡☛✝☞✌✠✆✝✞✟
✡☛✝☞✌✠✡☛✝☞✌
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✍✎ ✏❡ ❆✎ ✍✑♠ ♣✑✒✓ ✍✓✔✎ ✔ ✕✖ ✗✘✙
❉❊❚ ❝✉✚✈✛✜ ❢✢✚ ◆■❙❚ ❙❘❊✶ ✣
Figu e 5.2: DET cu es wi h SPLDA o SRE10 co ex -co eex de 5 emale wi h
in ol ed sho u e ances
Expe imen Long-Sho Sho -Sho
EER (%) minDCF EER (%) minDCF
SPLDA 5.98 0.291 8.76 0.403
PLDAUP scala 5.11 0.261 7.72 0.389
PLDAUP unscen 5.08 0.269 7.67 0.385
Table 5.2: EER (%) and minDCF esul s wi h PLDAUP o SRE10 co eex -co eex
de 5 emale wi h in ol ed sho u e ances
85
PLDA wi h Unce ain y P opaga ion (PLDAUP)
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠
✝✞✟✠✡
✞☛✝
✝✞✟✠☛☞✌✍✎☞✏☛✝
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜
❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧
✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵
✵✁
✵✥
✵✂
✁
✥
✂
✁✵
✥✵
✄✵
☎✵
✆✝✞✟✠
✝✞✟✠✡
✞☛✝
✝✞✟✠☛☞✌✍✎☞✏☛✝
▼
✐
s
s
P
♦
❜
❛
❜
✐
❧
✐
②
✭
✪
✮
❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜
❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧
a) Long-Sho expe imen b) Sho -Sho Expe imen
Figu e 5.3: DET cu es wi h PLDAUP o SRE10 co ex -co eex de 5 emale wi h
in ol ed sho u e ances
ans o ma ions. These conside a ions can also be obse ed in Fig. 5.3, whe e DET cu es a e
shown.
DET cu es explain he di e ences be ween minDCF and EER. PLDAUP seems o wo k
simila ly o SPLDA in hose egions o he DET cu e highly penalizing missing a ge ials.
By con as , in hose egions o he cu e o high alse ala m is whe e PLDAUP ob ains he
highes imp o emen s.
5.2.2 PLDAUP in speake clus e ing
The inclusion o he unce ain y p opaga ion in he PLDA o speake e i ica ion has led o
small imp o emen s, despi e no being such a e olu ion. Howe e , dia iza ion is a sligh ly
di e en ask. Ra he han making independen decisions, dia iza ion is he esul o a la ge se
o choices depending on each o he . Thus, small indi idual imp o emen s may esul in o an
accumula ion o bene i s.
This is why we wan o e alua e he unce ain y p opaga ion capabili ies in ou clus e ing
s ep, he s age whe e we can easily in eg a e ou PLDAUP model. As a i s app oxima ion we
do no wo k in dia iza ion ye bu in a speake clus e ing ask, wo king wi h sho u e ances.
Fo his eason, we emain wo king in he elephone channel domain, using he same chopped
subse p e iously used in speake e i ica ion. The o al amoun o in ol ed audios is 2740
u e ances, con aining 232 di e en speake s.
86
Chap e 5. Unce ain y P opaga ion o Dia iza ion
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁

✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
❙✂✄☎✆ ✝✞
✂✄☎✆P✄✟✂ ✝✞
✂✄☎✆✠✡☛☞✌✡✍
✟✂ ✝✞
❙✂✄☎✆ ❙✞
✂✄☎✆P✄✟✂ ❙✞
✂✄☎✆✠✡☛☞✌✡✍
✟✂ ❙✞
◆
✠✎✏
✌✑ ✒✓ ☛✔
✌✕✖✌
✑☛
❈❧ ✉s ❡ s
■
♠
♣
✗
✘
✐
✙
②
✭
✪
✮
✚✛✜✢✣✤✥✦ ♦❢ ✧★▲❉❆ ✈✩ ★▲ ❉❆❯★
Figu e 5.4: Impu i y esul s o SPLDA and PLDAUP in SRE10 co eex -co eex de 5
emale chopped
Ou expe imen al se up is he same as in ou p e ious speake e i ica ion expe imen , i.e. a
GMM-UBM i- ec o ex ac o ollowed by a PLDA model. Howe e , his ime sco es a e no
conside ed o 1 s1 ial decisions bu he me ic o an AHC solu ion, which de e mines he
inal labels. In o de o p o ide a be e o e iew abou he po en ial o PLDAUP, we p e e no
using any s op c i e ion, analyzing mul iple le els o he AHC dend og am. The esul s will
be measu ed in e ms o speake and clus e impu i ies (SI and CI espec i ely) and shown in
Fig. 5.4. Thick lines ep esen clus e impu i ies and dashed lines speake impu i ies. The anal-
ysis in ol es 3 di e en PLDA e sions: adi ional SPLDA is shown in blue, PLDAUP wi h
scala no maliza ion o he unce ain y ma ix is shown in ed and g een ep esen s PLDAUP
wi h no maliza ion o he unce ain y ma ix by means o an unscen ans o ma ion. Ou ep e-
sen a ion also includes an ex a line (black) indica ing he ue alue o speake s in his subse .
The esul s in Fig. 5.4 show a g ea bene i when unce ain y p opaga ion is applied, con-
i ming ou hypo hesis o imp o emen accumula ion. Fo he ange o s udy PLDAUP clus e
impu i y consis en ly unde goes an absolu e imp o emen wi hin he ange 5-10% wi h espec
o SPLDA, while no e iden deg ada ions in he speake impu i y a e no iced. Besides, his
imp o emen has also educed he bias o he Equal Impu i y (EI) poin in e ms o he numbe
o speake s. While SPLDA eaches he EI poin a 350 speake s, PLDAUP does he same a 300
87
PLDA wi h Unce ain y P opaga ion (PLDAUP)
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁

✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
▲✂✄☎ ✆✝
❙✞✂✟✠ ✆✝
▲✂✄☎ ❙✝
❙✞✂✟✠ ❙✝
◆✡☛☞✌✟ ✂
✍ ✎✏✌✑✒✌
✟✎
❈❧✉s ❡ s
■
♠
♣
✓
✔
✐
✕
②
✭
✪
✮
✖✗✘✙✚✛✜✢ ♦❢ ✣P✤❉❆ ✇✛✜ ❤ ✤♦♥❣ ✈✥ ✣❤♦✚✜ ❯✜✜✦✚ ❛ ♥❝✦✥
✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁
✁

✶✁
✶
✷✁
✷
✸✁
✸
✹✁
✹
✁
▲✂✄☎ ✆✝
❙✞✂✟✠ ✆✝
▲✂✄☎ ❙✝
❙✞✂✟✠ ❙✝
◆✡☛☞✌✟ ✂
✍ ✎✏✌✑✒✌
✟✎
❈❧✉s ❡ s
■
♠
♣
✓
✔
✐
✕
②
✭
✪
✮
✖✗✘✙✚✛✜✢ ♦❢ P✣❉❆❯P ✇✛✜❤ ✣♦♥❣ ✈✤ ✥❤♦✚✜ ❯✜✜✦✚❛♥❝✦✤
a) SPLDA b) PLDAUP
Figu e 5.5: Impu i y esul s o a) SPLDA and b) PLDAUP wi h scala no maliza ion
in SRE10 co eex -co eex de 5 emale chopped aining wi h sho u e ances
speake s. Fu he mo e, while speake e i ica ion esul s indica ed ha unscen ans o ma ions
we e be e han scala no maliza ion o he unce ain y ma ix, hose ob ained o he clus e ing
ask show ha he scala no maliza ion o e comes he pe o mance o unscen ans o ma ions
up o an absolu e 2-5%.
Du ing he p e ious expe imen s we analyzed some PLDA model ained on exce p s om
SRE04, 05, 06 and 08. This aining pool consis s o audios wi h la ge amoun s o speech
pe u e ance. Thus, aining embeddings can be conside ed eliable enough o ou s anda ds.
Howe e , his scena io may no be so ealis ic in o he domains, as dia iza ion. In ac , dia iza-
ion da a usually consis s o a combina ion o long and sho u e ances. Thus, we mus analyze
how ou wo models, SPLDA and PLDAUP, beha e when sho u e ances a e conside ed o
model aining.
Fo his pu pose, we analyze he impac o sho u e ances on he aining pool. In his
expe imen we build an al e na i e e sion o he conside ed aining pool (SRE04, SRE05,
SRE06 and SRE08) by andomly chopping he o iginal u e ances gua an eeing he speech con-
en o be wi hin he ange o 3-60 seconds. This subse will only be conside ed o he aining
o he PLDA models, bo h SPLDA and PLDAUP. Unde hese condi ions we e alua e ou subse
wi h sho u e ances wi h bo h PLDA models, SPLDA and PLDAUP wi h scala no maliza ion
o he unce ain y ma ix. In Fig. 5.5 we compa e how each model esponds depending on he
aining coho , di ing be ween SPLDA (Fig. 5.5a) and PLDAUP (Fig. 5.5b).
The illus a ed esul s in Fig. 5.5 e eal in e es ing de ails. Fi s , SPLDA seems o adap
well o sho u e ances, ou pe o ming he e sion wi h long u e ances o any ope a ional
88
Chap e 5. Unce ain y P opaga ion o Dia iza ion
poin . An explana ion is ha he aining coho su e s om shi s o he embeddings due o i s
leng h, simila o hose in he e alua ion subse . Consequen ly, hese shi s can be conside ed
as ex a in a-speake a iabili y and a e aken in o accoun in he co esponding pa ame e
(W). By con as , PLDAUP seems o sligh ly lose some o i s pe o mance. A eason o his
beha iou is ha in a-speake a iabili y lies in a subspace con olled by Ujand W. When
long eliable u e ances a e used o ain he model mos o his a iabili y is o ced o be in he
Wsubpace, ac ing UjUT
jas an addi ion du ing e alua ion. Howe e , when aining in ol es
sho u e ances bo h e ms a e ep esen a i e and con ibu ing, hus making decisions much
noisie .
5.3 Fully Bayesian P obabilis ic Linea Disc iminan Analy-
sis wi h Unce ain y P opaga ion (FBPLDAUP)
The con i ma ion o Unce ain y P opaga ion bene icial capabili ies mo i a es i s e alua ion in
b oadcas dia iza ion. Howe e , in his domain ou bes esul s so a ha e been shown in
Sec ion 4.1 by means o he FBPLDA and i s Va ia ional Bayes esegmen a ion. This eseg-
men a ion is in ac he key poin o his bes app oach, ixing some o he mis akes and hus
imp o ing he pe o mance.
Fo his pu pose, we upda e he FBPLDA model desc ibed in Sec ion 4.1.1 including
he new unce ain y p opaga ion concep . The name o his new model is Fully Bayesian
PLDA wi h Unce ain y P opaga ion (FBPLDAUP). This new model, as well as i s p ede-
cesso , explains a se o Nembeddings Φ om Idi e en speake s modeled by he se
Y={y1, ..., yi, ..., yI}, whe e yiis a la en a iable common o all u e ances om he same
i h speake . Addi ionally, i also inco po a es an ex a la en a iable xij pe u e ance, esponsi-
ble o modeling he a iabili y due o he u e ance leng h. The assignmen o each embedding
o i s esponsible speake is done in e ms o θij, a la en a iable aking he alue o 1i he
elemen jis gene a ed by he i h speake and 0 o he wise. Thus, we de ine he condi ional
dis ibu ion o Φas:
P(Φ|Y,Θ,X,µ,V,W) =
I
Y
i=1
N
Y
j=1 Nφj|µ+Vyi+Ujxj,W−1θij (5.11)
whe e Nφj|µ+Vyi+Ujxij,W−1 ep esen s he dis ibu ion o he embedding φjacco d-
ing o speake i. This modeliza ion includes a speake independen e m µ, a speake dependen
e m Vyiand he i- ec o a iabili y e m Ujxij.Vis a low ank ma ix explaining he speake
subspace and yiis he speake la en a iable. Ujs ands o he i- ec o a iabili y subspace
89

FBPLDA wi h Unce ain y P opaga ion (FBPLDAUP)
µ
W
V
φjyi
Uj
xij
θij
πθ
ε
Ni
I
Figu e 5.6: Bayesian ne wo k o he Fully Bayesian PLDA wi h Unce ain y P opa-
ga ion
ull ank ma ix and xij i s la en a iable. Finally Wis a ull ank ma ix explaining he wi hin
speake subspace.
Due o he ac ha we a e building a Fully Bayesian solu ion, ou model pa ame e s (µ,V
and W) a e dis ibu ions a he han poin es ima es. In ac , Vhas i s own p io dis ibu ion
ε, a p oduc o gamma dis ibu ions. In addi ion o he model pa ame e s, he speake labels
Θa e also ea ed as la en a iables, modeled by means o a mul inomial dis ibu ion. This
mul inomial dis ibu ion includes a Di ichle p io πθin o de o explain he p obabili ies pe
class. The Bayesian ne wo k o he model is shown in Fig. 5.6.
Fo dia iza ion pu poses wi h he FBPLDAUP we ollow he s a is ical app oach al eady
desc ibed in Sec ion 2.6.2 maximizing he pos e io dis ibu ion P(Θ|Φ). Howe e , he com-
plexi y o he model makes he ue pos e io in ac able o op imiza ion pu poses. Hence, we
p e e applying Va ia ional Bayes o a mo e sui able solu ion. The applied simpli ica ion o he
new model is:
PY,X,Θ, πθ,˜
V,W,ε=q(Y,X)q(Θ) q(πθ), q ˜
Vq(W)q(ε)(5.12)
Fo mo e in o ma ion abou he o mula ion o he di e en p io s he o mula ion is in-
cluded in Appendix A.
Due o he ela ionship be ween FBPLDA and FBPLDAUP, bo h ollow he dia iza ion s a -
egy desc ibed in Sec ion 4.1.1, based on a Va ia ional Bayes app oxima ion. Ou new model
assumes ha dia iza ion labels Θdia should be hose which bes explain he gi en u e ances,
90
Chap e 5. Unce ain y P opaga ion o Dia iza ion
modeled by i s mean and co a iance. Thus, we mus maximize P(Θ|Φ). By Va ia ional Bayes
we app oxima e his pos e io dis ibu ion by q(Θ), whose maximum will be ou solu ion. Un-
o una ely, VB ac o s a e in e connec ed, equi ing q(Θ) he o he ac o s o be op imized as
well o a p ope solu ion, which also need q(Θ) adjus ed as well. The e o e, we mus apply an
i e a i e ee alua ion o ac o s (q(Y,X),q(Θ) and q(πθ)) in o de o each he bes alue o
ou hypo hesis.
5.4 Dia iza ion o b oadcas da a wi h FBPLDAUP
In he ollowing lines we p esen he esul s in b oadcas dia iza ion when Unce ain y P op-
aga ion is included. Fo hese expe imen s we make use o MGB 2015 acco ding o he con-
igu a ion explained in Sec ion 3.2.1. The applied sys em ollows he schema ic p esen ed in
Sec ion 4.1.2, whe e an AHC s age o es ima e pa i ion seeds o he VB esegmen a ion. In
ou new app oach, any in ol ed PLDA is subs i u ed by i s PLDAUP coun e pa , so ou AHC
s age wo ks in e ms o he PLDAUP sco e while he VB esegmen a ion ole is now played by
he FBPLDAUP. Fo compa ison easons bo h models will p esen he same dimension o he
speake subspace han in he o iginal coun e pa .
In he i s compa ison we will compa e he ini ializa ion labels. Fo his eason, we e alua e
he whole AHC ee aking in o accoun wo simila i y me ics, SPLDA and PLDAUP. Then, o
each le el on bo h ees we will e alua e he esul ing pa i ions by means o DER and calcula e
∆DER = DERPLDAUP −DERSPLDA. This compa ison is epea ed o each in ol ed show in
MGB 2015, including bo h de elopmen and es subse s. In Fig. 5.7 we illus a e a his og am
abou he ela i e a ia ions o DER depending on he conside ed model.
In Fig. 5.7 we can obse e a dis ibu ion whose mean and mode a e biased o nega i e alues,
indica ing imp o emen s when subs i u ing he SPLDA model by he PLDAUP. Besides, he
skewness o he dis ibu ion is also nega i e, showing mo e ele ance o nega i es alues o
∆DER whe e PLDAUP ou pe o ms SPLDA.
A simila s udy can be pe o med wi h clus e and speake impu i ies. Fig. 5.8 epea s
he p eceding p ocedu e, al hough now we e alua e bo h impu i ies ins ead o DER. Simila
his og ams analyzing impu i ies o he pool o pa i ions a e shown, di e en ia ing be ween
clus e (Fig. 5.8a) and speake (Fig. 5.8b) impu i ies.
The esul s in Fig. 5.8 show dis ibu ions wi h mean and mode close o ze o o bo h cases.
Di e ences a ise when conside ing highe o de momen s, such as skewness when bo h im-
pu i ies ha e opposi e beha iou (speake impu i y is nega i e while clus e impu i y is posi-
i e), and ku osis, being highe in he clus e impu i y dis ibu ion. The beha iou obse ed
91
Dia iza ion o b oadcas da a wi h FBPLDAUP
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎❉❊❘ ✭✪✮
P
♦
♣
♦
✐
♦
♥
✆
✝
✞
✟✠s✡☛✠❜✉✡✠☞✌ ☞❢ ✍✟✎✏ ❜ ❡✡✇❡❡✌ ❙✑▲✟❆ ❛✌❞ ✑▲✟❆❯✑
Figu e 5.7: His og am o DER a ia ions be ween SPLDA and PLDAUP in MGB
2015 da a.
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎ ❈❧✉s ❡ ■♠♣✉ ✐ ② ✭✪✮
P
✆
♦
✝
♦
✆
✞
✟
♦
♥
✠
✡
☛
❉☞✌✍✎☞❜✏✍☞✑✒ ✑❢ ✓ ✔✕ ❜ ✖✍✇✖✖✒ ❙✗▲❉❆ ❛✒❞ ✗▲❉❆❯✗
✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁
✂
✁
✄✂
✄✁
✂
✁
✸
✂
✸
✁
☎ ❙♣ ❡❛ ❦❡ ■♠♣✉ ✐ ②✭✪✮
P
✆
♦
✝
♦
✆
✞
✟
♦
♥
✠
✡
☛
❉☞s✌✍☞❜✎✌☞✏✑ ✏❢ ✒✓✔ ❜ ✕✌✇✕✕✑ ✔✖▲❉❆ ✗✑❞ ✖▲❉❆❯✖
a) Clus e Impu i y b) Speake Impu i y
Figu e 5.8: His og am o a) clus e and b) speake impu i ies a ia ions be ween
SPLDA and PLDAUP ini ializa ions in MGB 2015 da a.
92
Chap e 5. Unce ain y P opaga ion o Dia iza ion
Expe imen DER(%)
SPK Se up
40 40.61
50 39.83
75 39.72
Bes FBPLDA 41.37
ELBO Se up
ELBO 39.83
Bes FBPLDA 39.12
Table 5.3: DER (%) esul s wi h he FBPLDAUP model in MGB 2015.
in Fig. 5.8 does no ma ch wi h hose p e iously ob ained in Sec ion 5.2.2. While in hose
expe imen s bene i s we e ob ained in he clus e impu i y, emaining speake impu i y almos
unal e ed, in b oadcas da a clus e impu i y shows an a e age 1.24% absolu e ex a deg ada-
ion, wi h 70% o he pa i ions deg ading his me ic, and speake impu i ies show a 2.05%
absolu e imp o emen , common o 63% o he pa i ions.
A conclusion ex ac ed om hese esul s indica es ha PLDAUP does no longe p o ide
such imp o emen s ob ained in elephone channel expe imen s. Some causes o his deg a-
da ion lie on he e ec o sho u e ance aining. In he elephone channel expe imen s we
obse ed how sho u e ances ials we e be e e alua ed as long as he SPLDA model con-
side ed hem du ing aining. By con as , PLDAUP did no show any imp o emen bu small
deg ada ions when ained wi h sho u e ances. In ou cu en scena io we mus deal wi h e y
sho u e ances, much sho e han hose used in elephone channel expe imen s. Hence some
educ ion o he expec ed imp o emen s seems easonable.
The inal s ep is he inclusion o he new model FBPLDAUP on op o he ini ializa ion,
ca ied ou by AHC wi h a PLDAUP model. In Table 5.3 we include hose expe imen s wi h he
new model a chi ec u e o MGB 2015 e alua ion subse . Two di e en c i e ia o hypo hesis
selec ion ha e been e alua ed: p io speake numbe es ima ion and ELBO choice. The expe -
imen s include hose esul s ob ained wi h he new app oach as well as a line indica ing hose
esul s p e iously ob ained wi h he non-UP models.
Acco ding o he esul s shown in Table 5.3, he FBPLDAUP shows po en ial bene i s de-
spi e inal esul s a e o e come by adi ional FBPLDA. When compa ing se ups wi h a ixed
numbe o speake s ou esul s clea ly imp o e hose ob ained by FBPLDA. Howe e , ou new
model does no ge any bene i om he ELBO s op c i e ion while FBPLDA does. A possible
93
PLDA ee-based clus e ing
6.2 PLDA ee-based clus e ing
The PLDA ee-based clus e ing is a s a is ical gene a i e solu ion o he dia iza ion clus e ing
ask. Hence, i ollows he p inciple desc ibed in Sec ion 2.6.2, iden i ying he a ge pa i ion
Θ o he se o embeddings Φas he one maximizing P(Φ,Θ). In o de o do so i exploi s he
ee pe spec i e o he clus e ing p oblem and he pa h decoding s a egies.
The applica ion o s a is ical s a egies on op o a ee s uc u e associa es a p obabili y o
each node. Rega ding he lea es his p obabili y is P(Φ,Θm) : m= 1..BN, i.e. he quali y
me ic o each pa i ion. In o de o ob ain a simila p obabili y o he emaining nodes we
make use o he p oduc ule o p obabili y. This ule allows he decomposi ion o a gene ic
P(a1, ..., aN)as:
P(a1, ..., aN) =P(a1)P(a2|a1)Pa3|a2
1...P aN|aN−1
1
=
N
Y
j=2
Paj|aj−1
1P(a1)(6.1)
whe e aj
1 ep esen s he se o elemen s {a1, ..., aj}. This de ini ion can also be exp essed in a
ecu si e way:
Paj
1=Paj|aj−1
1Paj−1
1(6.2)
The applica ion o he p oduc ule o p obabili y o ou dia iza ion p oblem is di ec , sub-
s i u ing he j h elemen aj om he p e ious equa ion by he j h pai o a iables, consis ing
o he embedding φjand i s clus e iden i y label θj. Hence, he p obabili y o any pa i ion
P(Φ,Θ) can be decomposed as:
P(Φ,Θ) =
N
Y
j=2
Pφj, θj|φj−1
1, θj−1
1P(φ1, θ1)(6.3)
and i s al e na i e ecu si e de ini ion:
Pφj
1, θj
1=Pφj, θj|φj−1
1, θj−1
1Pφj−1
1, θj−1
1(6.4)
In consequence, each node a dep h jin he clus e ing ee Thas he p obabili y
Pφj
1, θj
1;θj
1∈Ωθj
1associa ed. Acco ding o his decomposi ion we a e assuming Φas a
sequence o o de ed embeddings o be clus e ed. Besides, hese embeddings only depend on
p e ious speake ep esen a ions o he sequence. These assump ions a e easonable in eal li e,
whe e he oice e ol es along ime, being specially no iceable in la ge segmen s o speech.
100

Chap e 6. T ee-Based Clus e ing App oaches
6.2.1 PLDA-based model
Along he p e ious lines we explo ed a new pe spec i e abou clus e ing, ep esen ing i as a ee
s uc u e o be decoded in o de o ob ain he bes pa i ion. Mo eo e , he s a is ical s a egy
associa ed a p obabili y Pφj
1, θj
1 o all nodes along he ee. Howe e , he dis ibu ion o his
p obabili y has no been speci ied ye . Conside ing speake ecogni ion s a e o he a , PLDA
amily models seem a powe ul ype o solu ion o apply.
Thus, we mus keep on ans o ming Pφj
1, θj
1 o make PLDA de ini ion applicable. As
a i s ans o ma ion, we keep on applying he p oduc ule o p obabili y, decomposing he
p obabili y a each node in o a e m depending on he embeddings, he condi ional dis ibu ion,
and a p io dis ibu ion o he labels. This decomposi ion is:
Pφj, θj|φj−1
1, θj−1
1=Pφj|θj,φj−1
1, θj−1
1Pθj|φj−1
1, θj−1
1(6.5)
This decomposi ion allows o spli Pφj
1, θj
1in o wo simple p oblems. Now, we exclu-
si ely ocus on he i s e m, he condi ional dis ibu ion o he embedding jgi en i s j h label
as well as p e ious embeddings φj−1
1and labels θj−1
1. Decisions abou he o he e m, he p io
dis ibu ion o he cu en label θjgi en p e ious embeddings and decisions will be made a e -
wa ds. Un o una ely, his condi ional e m is s ill in ac able o use PLDA due o he p esence
o label a iables. Mo eo e , PLDA only de ines dependencies among embeddings om he
same speake . The e o e, ou nex ans o ma ions seek sepa a ing embeddings om he labels.
We ake inspi a ion om Chap e 4, imposing Pφj|θj,φj−1
1, θj−1
1 o ollow a mul inomial dis-
ibu ion on he a iable θj, a one-ho sample wi h I alues (θj={θ1j, ..., θij, ..., θIj}), whe e
Iis he numbe o candida e speake s. Thus, he alue θij will ake he alue o one i he j h
embedding was gene a ed by he speake i, being ze o o he wise. Besides, we also equi e φj,
when belonging o clus e iacco ding o θj, o be exclusi ely explained by hose embeddings
al eady assigned o his clus e . This subse o embeddings p e iously assigned o clus e ia
ime jis deno ed by Φij. Unde hese o condi ions we can exp ess:
Pφj|θj,φj−1
1, θj−1
1=
I
Y
i=1
Pφj|Φijθij ;j= 1..N (6.6)
The de ini ion o he e m Pφj|Φijnow makes he applica ion o PLDA p inciples easi-
ble. Fi s , we mus assume he exis ence o a la en a iable ep esen ing he speake in o ma ion
yi, which allow us o ede ine Pφj|Φijas:
Pφj|Φij=ZPφj|yiP(yi|Φij)dyi(6.7)
101
PLDA ee-based clus e ing
Gi en his de ini ion we can now assume ha he da a we a e dealing wi h is gene a ed
by a PLDA model. Fo his pu pose, we make use o he SPLDA condi ional dis ibu ion
Pφj|yi,MSPLDA:
Pφj|yi,MSPLDA∼ N φj|µ+Vyi,W−1(6.8)
whe e µis he speake independen e m, Va low ank ma ix desc ibing he speake subspace
and Wa ull ank ma ix explaining he in a-speake a iabili y space.
Fu he mo e, he second e m P(yi|Φij), he pos e io dis ibu ion o he la en a iable
gi en all hose embeddings p e iously assigned o clus e i(Φij) is also modeled acco ding o
SPLDA. Thus, i s de ini ion is:
P(yi|Φij,MSPLDA)∼N yi|µyi(j),L−1
yi(j)(6.9)
Lyi(j) =I+VT
j−1
X
k=1
θkiWV (6.10)
µyi(j) =L−1VTW
j−1
X
k=1
θki(φk−µ)(6.11)
whe e µyi(j)and Lyi(j) ep esen he es ima es o he mean and a iance pa ame e s espec-
i ely o he la en a iable yiwhen only j−1elemen s we e obse ed. As long as jinc eases
hese es ima ions should ge close o he eal alue.
Apa om well-known de ini ions o bo h dis ibu ions, he choice o he SPLDA
model also p o ides an ex a ad an age. I s Gaussian na u e o bo h Pφj|yi,MSPLDA
and P(yi|Φij,MSPLDA)allows a closed o m solu ion o he in eg al de ining
Pφj|Φij,MSPLDA. The esul ing o mula ion o his e m is:
Pφj|Φij,MSPLDA∼N φj|µi(j),Σi(j)(6.12)
µi(j) =µ+Vµyi(j)(6.13)
Σi(j) =W−1+VL−1
yi(j)VT(6.14)
A e comple ely de ining he condi ional dis ibu ion o he embeddings, we now can pay
a en ion o he label p io dis ibu ion Pθj|φj−1
1, θj−1
1. Fi s , we assume a simpli ied p io
dis ibu ion by elimina ing he dependence wi h espec o he pas embeddings φj−1
1. The e o e,
ou p io dis ibu ion will ollow he o m Pθj|θj−1
1. Fo he esul ing dis ibu ion we ha e
op ed o he Dis ance Dependen Chinese Res au an (DDCR) p ocess [Blei and F azie , 2011],
al eady used in dia iza ion in [Zhang e al., 2019]. This model explains he occupa ion o an
102
Chap e 6. T ee-Based Clus e ing App oaches
µW V
φjyi
θj
δ
ζ
Ni
I
Figu e 6.2: PLDA ee-based clus e ing Bayesian Ne wo k
in ini e se ies o clus e s by a sequence o elemen s. Then, he assignmen o he elemen j o
any clus e exclusi ely depends on he occupa ion o clus e s up o his poin , i.e. acco ding
o all he p e ious decisions, as in ou decomposi ion. The p obabili y o assignmen o he
elemen j o any o he al eady c ea ed k= 1..K clus e s is p opo ional o i s occupa ion
a ime j, namely nk. Besides, DDCR o e s he possibili y o c ea e a new clus e K+ 1
p opo ional o γ. The ma hema ical o mula ion o DDCR is:
Pθj=k|θ(j−1)
1∝(nki k≤K
ζi k=K+ 1 (6.15)
DDCR deeply ma ches sequen ial o de ing and assignmen p oblem. Un o una ely, DDCR
conside s easonable a con inuous ansi ion among speake s. Applied o scena ios o speake
clus e ing, whe e speake eco ding can be in e lea ed, seems easonable. Howe e , in dia iza-
ion we mus conside he segmen a ion s age, which can di ide any long segmen in o pieces
o sho e leng h. Thus, we can add o his dis ibu ion mo e chances o emain in he speake
clus e . In ou p oposal we do so by speci ically de ining he si ua ion o emaining in he cu -
en speake clus e , wi h a p obabili y p opo ional o δ. This addi ion gene a es he ollowing
modi ica ion o he DDCR dis ibu ion:
Pθj=k|θ(j−1)
1∝




δi k=θ(j−1)
nki k6=θ(j−1) and k≤K
ζi k6=θ(j−1) and k=K+ 1
(6.16)
The model P(Φ,Θ), aking in o accoun he whole se o assump ions p e iously desc ibed,
can be ep esen ed by he Bayesian ne wo k illus a ed in Fig. 6.2
6.2.2 M-algo i hm op imiza ion
Once he PLDA-based model is de ined, now i is ime o ind he way o ob ain hose labels
Θdia ha bes explain he se o embeddings Φ. Taking in o accoun ha he Vi e bi algo i hm
103
PLDA ee-based clus e ing
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.3: M-algo i hm example o a clus e ing ee o dep h 4. 2 pa hs ali e each
he dep h 2 h ough he ee (g een).
canno be applied, subop imal app oaches conside ing us wo hy pa hs, as he M algo i hm
[Jelinek and Ande son, 1971] a e s ill applicable.
The M algo i hm is an i e a i e solu ion s a egy. Gi en a scena io wi h a decision ee o
dep h N, he M algo i hm acks a subse o Msu i ing pa hs, i.e. hose pa hs mo e likely o be
he solu ion (in ou case hose wi h highe log-likelihood). Besides, all pa hs mus ha e eached
dep h jwi hin he ee s uc u e. Thus, he goal is he iden i ica ion o hose bes ansi ions
aking he M acked pa hs om dep h j o dep h j+ 1. In Fig. 6.3 we illus a e an example,
whe e a clus e ing ee o dep h 4 is analyzed by he M algo i hm wi h M= 2. Su i ing pa h
(g een lines) ha e eached dep h 2 wi hin he ee.
The M algo i hm i e a i e p ocedu e is di ided in o wo s eps, es ima ion and maximiza ion.
The es ima ion s ep s udies how he Msu i ing pa hs in le el je ol e deepe h ough he
ee, p edic ing i s pe o mance in a u u e scena io and making decisions in consequence. Fo
his pu pose we ca y ou a b u e- o ce app oach, analyzing any possible ansi ion om he M
su i ing pa hs a dep h jup o a ce ain ex a dep h d. The pa ame e dis a design choice and
esponsible o a adeo be ween accu acy and compu a ional cos s. The highe d, he wide
104
Chap e 6. T ee-Based Clus e ing App oaches
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.4: Es ima ion s ep in a M-algo i hm example o a clus e ing ee o dep h 4.
2 pa hs ali e (g een) eaching dep h 2 a e p opaga ed o all possible nodes a dep h 3
(blue)
is he explo a ion o he ee o unseen da a and hence highe accu acy migh be expec ed, bu
inc easing in an exponen ial manne he compu a ional cos s. Hence, many sys ems es ic d
o be equal o 1. In Fig. 6.4 we ep esen he es ima ion s ep applied o ou p e ious example
scena io in Fig. 6.3. Each su i ing pa h (g een line) is p opaga ed d(d= 1) le els ahead (blue
lines), e alua ing o each con igu a ion he pe o mance a his dep h.
The esul s o he es ima ion s ep p o ide an o e iew abou how he ee beha es in u u e
s eps, wi hou comp omising any decision. This choice is made du ing he maximiza ion s ep.
In his s ep all candida e p opaga ions a e anked, only keeping hose Mwi h be e sco e.
These new Mpa hs now eaching dep h j+ 1 a e ou mos p omising candida es so a , and
hose conside ed o he nex i e a ion o he algo i hm. This s ep is ep esen ed in Fig 6.5.
6.3 Expe imen s
Fo he e alua ion o he new clus e ing app oach, we will make use o Albayzín 2018, as de-
sc ibed in Sec ion 3.2.2. Fo his pu pose, we conside an i- ec o PLDA dia iza ion sys em
105

Expe imen s
1
1
2
1
2
1
2
3
12
1
2
3
1
2
3
1
2
3
1
2
3
4
1s elemen
2nd elemen
3 d elemen
4 h elemen
Figu e 6.5: Maximiza ion s ep in a M-algo i hm example o a clus e ing ee o dep h
4. The wo su i ing pa hs a e shown in g een.
106
Chap e 6. T ee-Based Clus e ing App oaches
Expe imen DER(%)
De . Subse E al. Subse
AHC 18.88 26.36
AHC + FBPLDA 13.90 17,79
PLDA TREE-BASED CLUSTERING 13.12 17.60
Table 6.1: DER (%) esul s o he PLDA ee-based clus e ing in Albayzín 2018.
Resul s compa ed wi h hose ob ained by means o AHC wi h and wi hou FBPLDA
esegmen a ion.
whose se up is: A 256 Gaussian GMM-UBM ollowed by a 100-dimension To al Va iabili y
ma ix a e esponsible o he i- ec o ex ac ion. The ob ained embeddings unde go cen e -
ing, whi ening and leng h no maliza ion p io o clus e ing, wi hou dimensionali y educ ion.
Finally, he new clus e ing app oach, he PLDA ee-based clus e ing, uses a 100-dimension
SPLDA. This se up i s in e ms o dimensions wi h he dia iza ion sys em using he FBPLDA
eclus e ing in Chap e 4 o expe imen s wi h Albayzín 2018.
In ou i s expe imen we compa e he pe o mance o FBPLDA eclus e ing, ob ained in
Chap e 4, and ou new clus e ing app oach. As a i s app oxima ion we assume he se o
embeddings Φ o be a anged in empo al o de . We es ic hype pa ame e d o be equal o 1
o compu a ional easons. In his expe imen we conside e alua ion condi ions, i.e. we only
p esen he pe o mance o he bes hype pa ame e con igu a ion (δ,ζand M) acco ding o
Albayzín 2018 de elopmen subse . The ob ained esul s a e shown in Table 6.1.
Acco ding o he ob ained esul s, he new clus e ing app oach, wo king wi h i- ec o s, p o-
ides e y li le imp o emen wi h espec o he FBPLDA coun e pa . Howe e , hese esul s
show bene i s in he e alua ion o bo h de elopmen and es subse s despi e con aining inde-
penden shows. The e o e, we can alk abou limi ed ye consis en imp o emen s due o ou
new clus e ing app oach.
Apa om he o e all sco e o bo h de elopmen and es subse s, a mo e de ailed analysis
o esul s can also be done. Fo his pu pose, we s udy he pe o mance pe show o in e es
wi h he h ee clus e ing app oaches conside ed along his hesis: AHC, FBPLDA and PLDA
ee-based clus e ing. Fo his pu pose, we analyze wo di e en me ics: On he one hand we
p opose he analysis o ∆I=IORACLE−IHYP, he di e ence in he numbe o speake s be ween
ou hypo hesis labels and he e e ence. On he o he hand, we analyze he DER pe o mance.
Bo h analyses a e shown in Fig. 6.6, including all shows in Albayzín 2018. The in ol ed shows
om he de elopmen subse a e millenium and La Noche en 24 Ho as (LN24H). Rega ding he
es subse , he shows España en Comunidad (EC), La inoamé ica en 24 Ho as (LA24H), La
107
Expe imen s
♠✁✁✂✄✄☎♠ ▲✆✝✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌
✍✝
✷
✍
✶✷
✷
✶✷
✝
✷
✸✷
☞❚❊❊
❋✎✏▲✑✡
✡✟✠
*
*
*
◆❛✒❡ ♦❢ ❤ ❡ ❙❤♦✇
✓
■
❘✔❧✕✖✐✈✔ s♣ ✔✕❦✔ s ✗✘ ♣ ✔ s✙✚✛
♠✁✁✂✄✄☎♠ ▲✆✝ ✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌
✶✍
✝✍
✸✍
✞✍
✺✍
✻✍
☞❚❊❊
❋✎✏▲✑✡
✡✟✠
*
*
*
◆❛✒❡ ♦❢ ❤❡ ❙❤♦✇
❉
✓
❘
✭
✪
✮
✔✕✖✗✘✙ ♣✚ s✛✜✢
a) ∆Ib) DER(%)
Figu e 6.6: Analysis pe show o a) ∆Iand b) DER(%) o AHC, FBPLDA and
PLDA ee-based clus e ing. Analysis ca ied ou on shows om Albayzín 2018, in-
cluding de elopmen and es subse s.
Mañana (LM) and La Ta de en 24 Ho as Te ulia (LT24HTe ) a e also included. Resul s e lec
he in e qua ile ange o each show.
Those esul s illus a ed in Fig. 6.6 show a simila beha iou o he h ee ypes o clus e -
ing pe show. Thus, hose mo e ha m ul shows a e common o all sys ems. Howe e , ou
new clus e ing app oach shows a mino in e qua ile ange pe show compa ed o AHC and
specially FBPLDA. This educ ion a ec s bo h he es ima ion abou he numbe o speake s
and DER. Hence he pe o mance o he sys em seems mo e consis en pe indi idual show o
domain, al hough small deg ada ions migh occu . This beha iou can also be ex apola ed o
he whole da ase , specially conside ing he show La Mañana (LM). While AHC and FBPLDA
pe o mances o his show a e a leas 100% wo se han any o he show in e ms o DER, he
PLDA ee-based clus e ing achie es o beha e as bad as he second wo s show. This imp o e-
men is also obse ed in he es ima ion o he speake numbe , wi h a ela i e 25% deg ada ion
educ ion.
Apa om a speci ic se up, we can also do an analysis s udying he in luence o each o
he model hype pa ame e s δ,ζand M. Fo his analysis we will conside he ob ained sco es
o any possible se up. Fig. 6.7 is ou chosen g aphical ep esen a ion o e eal he impac o
he di e en hype pa ame e s. I is composed o wo pa s, Fig. 6.7a whe e we ep esen he
ela ionship be ween ζand DER o di e en alues o M, and Fig. 6.7b , whe e we ep esen
he ela ionship be ween δ, and DER o he di e en alues o M. In o de o include all
hype pa ame e s in each subimage, Fig. 6.7a includes some a iabili y pe measu e, illus a ing
he in e qua ile ange esul s in e ms o he missing hype pa ame e , δ. Simila ly, measu es in
108
Chap e 6. T ee-Based Clus e ing App oaches
✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁
✂
✶
✂
✶✁
✂
✝
✂
✝✁
✂✞
✂✞✁
✄✵
✄✵✁
✄✂
✄✂✁
✄✄
▼✂
▼✄
▼✆
▼✂✵
▼✄✵
▼✆✵
▼✂✵✵
✏
❉
❊
❘
✭
✪
✮
✟✠✡☛☞✌ ✐♥ ❡ ♠s ♦❢ ✍
✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁
✂
✶
✂
✶✁
✂
✝
✂
✝✁
✂✞
✂✞✁
✄✵
✄✵✁
✄✂
✄✂✁
✄✄
▼✂
▼✄
▼✆
▼✂✵
▼✄✵
▼✆✵
▼✂✵✵
✍
❉
❊
❘
✭
✪
✮
✟✠✡☛☞✌ ✐♥ ❡ ♠s ♦❢ ✎
a) ζp obabili y b) δp obabili y
Figu e 6.7: DER (%) esul s o he PLDA ee-based clus e ing wi h M-algo i hm in
Albayzín 2018 in e ms o δ,ζand M
Fig. 6.7b include some a iabili y ma gins indica ing he i s and hi d qua ile esul s in e ms
o ζ.
The in o ma ion included in Fig. 6.7 e eals many impo an cha ac e is ics abou he model.
Fi s , he esul s e idence he impo ance o M. 20% ela i e imp o emen s may be ob ained
as long as mo e and mo e simul aneous pa hs a e e alua ed. Howe e , his imp o emen is no
uni o m, being any inc ease o Mmo e signi ican o lowe alues. Fo highe alues o M,
imp o emen s a e e y sca ce and implying la ge inc emen s in he compu a ional cos s. O he
de ail o bea in mind is ha , excep o Mhype pa ame e , he in luence o he emaining
adjus able alues (δand ζ) is in gene al educed (wi h he excep ion o ζ o M= 2). a ia ions
may be a ound 5% ela i e imp o emen /deg ada ions, i.e. 1% absolu e DER a ia ions.
Up o his poin we ha e only men ioned h ee exis ing hype pa ame e s, M δ and ζ. Ne e -
heless, all he esul s we e ob ained by se ing he embeddings Φin o a sequen ial o de . I he
analysis o he clus e ing ee was comple e, i.e. analyzing each one o he lea es, he impac o
his o de ing would be null. Ne e heless, by pa ially explo ing he clus e ing ee acco ding
o limi ed da a makes his a angemen an ex a ac o o ake in o accoun . The e o e, while in
ou p e ious examples we exclusi ely applied empo al o de , i.e. we can also apply di e en
a angemen s
One o he key ac o s when using his ee-based app oach is he sequence o de , specially
aking in o accoun ha we a e exploi ing he ela ionships be ween an embeddings and i s
p edecesso s in he sequence. Thus, we mus explo e how he o de ing a ec s he esul s. While
in ou p e ious expe imen s we simply made use o he empo al o de , his a angemen is no
109
Sho u e ances as occluded u e ances
bu ions, i.e. he dia iza ion ask, usually wo ks wi h e en sho e segmen s (1s-3s) o accu a ely
deal wi h speake bounda ies. Hence, imp o emen s in his scena io a e becoming mo e and
mo e needed.
7.2 Sho u e ances as occluded u e ances
The sho u e ance p oblem is widely known wi hin he speake ecogni ion communi y
[Podda e al., 2017]. The e alua ion o ials by means o sho u e ances in ol es a se e e
deg ada ion o pe o mance. Howe e , he e is no s anda d de ini ion o sho u e ance in he
li e a u e. While some wo ks ha e epo ed losses o pe o mance wi h audios con aining less
han 30 seconds o speech, a mo e se e e deg ada ion is ob ained conside ing sho e u e ances
(less han 10 seconds) [Mandasa i e al., 2011,Kanagasunda am e al., 2011]. This sho u e -
ance p oblem has also been analyzed in he Speake Recogni ion E alua ions (SRE), p oposed
by NIST. Despi e adi ionally conside ing u e ances wi h mo e han 2 minu es o audio, some
o he e alua ions [NIST, 2008,NIST, 2010] also include a condi ion in which u e ances con-
ain less han 10 seconds.
This loss o pe o mance is a consequence o a highe in a-speake a iabili y in he
es ima ions wi h sho u e ances. In he li e a u e mul iple con ibu ions ha e been p o-
posed o he di e en s eps o he speake e i ica ion pipeline, aiming o educe he unde-
si ed a iabili y. The ea u e ex ac ion s ep has been s udied in di e en ways, a emp ing
o p o ide an al e na i e o adi ional MFCCs. In [Li e al., 2015] a mul i esolu ion ime-
equency ea u e ex ac ion was p oposed, ca ying ou a mul i-scaled Disc e e Cosine T ans-
o m (DCT) on he spec og am, combining he in o ma ion a e wa ds. Al e na i e wo ks
like [Alam e al., 2015] use di e en ea u es based on he ampli ude and phase o he spec-
um. O he con ibu ions a e ocused on he modelling s age. Fac o Analysis app oaches
we e conside ed in [Vog e al., 2008] o de elop subspace models o be e wo k wi h he
sho u e ances. When conside ing i- ec o ep esen a ions, compensa ion echniques such
as [Kanagasunda am e al., 2013,Kanagasunda am e al., 2014] p ojec he ob ained ep esen-
a ions in o subspaces wi h low a iabili y due o sho u e ances. In [Sa ka e al., 2012] i is
shown ha sys ems ained on sho u e ances should compensa e he unce ain y due o lim-
i ed audio, imp o ing he e alua ion o sho audios. Howe e , when sys ems mus deal wi h
audios wi h un es ic ed leng h, sys ems should be ained on long u e ances o a be e pe -
o mance. The balance o he Baum Welch s a is ics, equi ed o he ex ac ion o i- ec o s, is
also p oposed in [Hau amäki e al., 2013]. Besides, DNNs ha e also mapped sho -u e ance i-
ec o s wi h espec o hei long-u e ance coun e pa s [Guo e al., 2017]. O he con ibu ions
116

Chap e 7. S udy o embeddings o sho u e ances
ha e also wo ked on he backend, specially PLDA. Ano he echnique, o iginally p oposed in
[Cumani e al., 2013b,Kenny e al., 2013] and analyzed in Chap e 5, makes he PLDA model
include an ex a e m o compensa e he unce ain y o he i- ec o , which depends on he u e -
ance leng h. Finally, o he s a egies compensa e he ob ained sco e acco ding o eliabili y me -
ics o he in ol ed u e ances [Hasan e al., 2013,Mandasa i e al., 2013], specially i s du a ion.
This idea is ex ended in [Viñals e al., 2018b], whe e he Quali y Measu e Func ion (QMF) e m
s udies he in e ac ion be ween en ollmen and es u e ances. In [Vog e al., 2010] in e als o
con idence a e es ima ed, leading owa ds conside able accu acy.
Some wo ks such as [Ajili e al., 2016] ha e s udied he impac o he di e en phone ic con-
en in he embedding ep esen a ions. Acco ding o hei esul s, owels and nasal phonemes
a e help ul o disc imina ion ma e s. By con as , o he ypes o phonemes, such as ica i es
o plosi es, can be misleading du ing e alua ion. Ou hypo hesis o wo k applies his idea o
phonemes o sho u e ances. The p esence o ce ain acous ic uni s boos s he pe o mance o
speake ecogni ion sys ems. Howe e , hese boos ing phonemes mus be in bo h en oll and es
u e ances o be e ec i e. This ma ch in he phone ic in o ma ion goes beyond he p esence o
ce ain phonemes, also equi ing a ma ch in he phone ic dis ibu ion along he u e ance.
In o de o explain ou pe spec i e le ’s make an analogy o he sho u e ance p oblem wi h
a simila p oblem, ace ecogni ion wi h occlusions. In he bes scena io, bo h p oblems con ain
all possible in o ma ion. Wo king wi h aces we ha e a comple e iew o he pe son o in e es ,
including all he ace elemen s ( wo eyes, he nose, he mou h, e c.). In speake ecogni ion we
ha e comple e in o ma ion in an u e ance ha con ains aces o any possible phoneme and
i s coa icula ion. As long as he u e ance ge s longe and longe he comple e in o ma ion
condi ion is mo e likely o be achie ed. In his scena io pe o mance has imp o ed mo e and
mo e as long as echnologies ha e e ol ed.
Now we ocus on sho u e ances. These con ain much less speech, e en less han a second.
A simple "Yes/No" eply o a ques ion can cons i u e an u e ance. Hence, sho u e ances a e
e y likely o lack o phonemes. In ace ecogni ion he equi alen scena io is he ecogni ion o
pa ial in o ma ion, whe e some pa s such as he mou h and nose a e no isible. In bo h cases
he missing in o ma ion exis s, bu i is una ailable. Faces always ha e a mou h and a nose
al hough some imes hey can be occluded, e.g. by a sca . Rega ding speake cha ac e iza ion,
speake s p onounce all he phonemes o a language while alking, al hough ew o hem can be
missing in a speci ic u e ance.
In ou hypo hesis we also conside he in luence o p opo ion. Acco ding o ou analogy
o ace ecogni ion, aces p esen a ixed se o elemen s (ea s, nose, mou h, e c.) wi h a con-
s ained size, and loca ed in he ace in speci ic a eas. These es ic ions a e always he same,
117
Fo mula ion o he embedding ex ac ion wi h sho u e ances
ega dless o he pe son no any occlusion. In speake cha ac e iza ion he si ua ion is sligh ly
di e en . When u e ances ge long enough he language imposes es ic ions in he phoneme
dis ibu ion. These es ic ions lead o a e e ence phoneme dis ibu ion. The longe he u -
e ance he mo e i s phoneme dis ibu ion ends o he e e ence dis ibu ion. Howe e , sho
u e ances con ain a much sho e message, and hus i s phoneme dis ibu ion can be se e ely
dis o ed. In his dis o ion we mus ake in o accoun bo h he missing phonemes and hose
p esen bu condi ioned o he message in he u e ance. This dis o ion may lead o u e ances
om he same speake wi h di e en dominan phonemes, hence complica ing he e alua ion.
Consequen ly, he sho u e ance p oblem can be in e p e ed as an occlusion om a com-
ple e in o ma ion scena io. This occlusion may be comple e, whe e long u e ances lack om
ce ain phonemes, o pa ial, in which u e ances ha e hei phonemes seen in e y di e en
p opo ions wi h espec o hei coun e pa s. The a ailable in o ma ion abou he occlusion is
impo an o be awa e o . Du ing e alua ion we compa e how he wo speake s p onounce all
he phonemes, a ailable o no , so unbalanced in o ma ion can lead o an un ai compa ison.
7.3 Fo mula ion o he embedding ex ac ion wi h sho u -
e ances
Cu en s a e-o - he-a speake e i ica ion, as desc ibed in Sec ion 2.5, elies on he pipeline
embedding-backend. U e ances a e i s con e ed in o compac ep esen a ions, he embed-
dings, which eed he decision backend o ob ain he sco e. Among all a ailable ep esen a-
ions, wo o he mos popula ones a e i- ec o s and x- ec o s. Bo h ha e been widely es ed
in speake e i ica ion ob aining g ea esul s. Fi s , we will y o unde s and how we s o e he
speake in o ma ion in hese embeddings and hen s udy i s d awbacks o sho u e ances.
7.3.1 Gene al case
The me hod o compac a a iable leng h u e ance in o a ixed-leng h ep esen a ion is simila
o mos embedding ex ac ion echniques. Gi en he u e ance O, an o de ed se o Nacous ic
ea u es O={o1, ..., on, ..., oN}, we ans o m hem by unc ion F(·), ob aining he o de ed
sequence F(O) = { 1, ..., n, ..., N}. This unc ion maps he o iginal ea u e ec o onin o
he speake cha ac e is ics subspace as he p ojec ions n. Depending on he embedding, p o-
jec ion nin ol es he ans o ma ion o he ea u e ec o onas well as a small con ex a ound
(app oxima ely 0.15 seconds). By means o his mapping we a emp o highligh he speake
pa icula i ies in he ea u es applying linea (e.g. i- ec o s) o non-linea ans o ma ions (as
118
Chap e 7. S udy o embeddings o sho u e ances
in DNNs). The unc ion F(·)is lea n om a la ge da a pool by da a analysis, e.g. by Max-
imum Likelihood algo i hms o i- ec o s o Back-P opaga ion [Rumelha e al., 1986] wi h
DNNs. Due o he ac ha each one o hese p ojec ions nonly co e s a small pe iod o
ime, hey only ha e in o ma ion abou ew acous ic uni s. The comple e cha ac e iza ion o a
speake equi es he s udy o his/he pa icula i ies o all he phonemes. These acous ic uni s
a e widesp ead along he u e ance, hus we mus combine he e ec o all hese p ojec ions n.
The usual me hod o combine he p ojec ions is i s empo al a e age. The esul is he compac
ep esen a ion G(O), de ined as:
G(O) = 1
N
N
X
n=1
n(7.1)
This embedding G(O)keeps ack o he phone ic con en in he u e ance O. Howe e ,
we can also ea each acous ic uni independen ly. Many s a e-o - he-a embeddings, such as
i- ec o s, can be in e p e ed as he sum o C ep esen a ions Gc(O), one pe acous ic uni , each
one es ima ed acco ding o Ncp ojec ions n. Acco ding o his easoning we can exp ess he
embedding as:
G(O) =
C
X
c=1
αcGc(O)(7.2)
The ob ained exp ession desc ibes embeddings as a weighed sum o Ces ima ions Gc(O),
each one ep esen ing he es ima ed pa icula i ies o he speake in a single acous ic uni . Gc(O)
can also be in e p e ed as he esul ing embedding only aking in o accoun he da a ela ed o
he phoneme c. All he con ibu ions a e weigh ed by he e m αc, he p opo ion o his acous ic
uni in he u e ance.
The e o e, embeddings a e condi ioned o wo main pa s: On he one hand he s abili y
o he dis ibu ion o weigh s α={α1, ..., αc, ..., αC}. On he o he hand he es ima ions
Gc(O), he pa icula i ies pe phoneme. Bo h bene i om la ge u e ances. E e y language has
i s own e e ence phone ic dis ibu ion. Hence he longe he u e ance he mo e i s phone ic
dis ibu ion becomes like his e e ence. Conce ning he es ima ions Gc(O), he mo e a ailable
da a, he less unce ain is he es ima ion.
The a e age s age is he las s ep in which we keep ack o he phoneme dis ibu ion. As a
consequence, we canno dis inguish be ween speake and phone ic a iabili y a e wa ds. Fu -
he s eps in he embedding pos -p ocessing o he backend may ans o m he embedding, bu
all phonemes a e equally ea ed.
119
Fo mula ion o he embedding ex ac ion wi h sho u e ances
7.3.2 i- ec o embeddings
The p e iously desc ibed o mula ion also ma ches wi h he adi ional i- ec o s. The i- ec o
modeling pa adigm, al eady desc ibed in Sec ion 2.5, explains he u e ance Oas he esul o
sampling om a Gaussian Mix u e Model (GMM), speci ic o he u e ance wi h pa ame e s
λO. This model λOis he esul o he adap a ion om a Uni e sal Backg ound Model (UBM), a
la ge GMM ha e lec s all possible acous ic condi ions. This adap a ion p ocess is es ic ed o
only he UBM Gaussian means. Besides, he shi o he GMM Gaussians is ied, and explained
by means o a hidden a iable wO, loca ed in he To al Va iabili y subspace, desc ibed by ma ix
T. Ma hema ically:
µO=µUBM +TwO(7.3)
whe e µO ep esen s he supe ec o mean, he conca ena ion o he GMM componen means,
om he a ge λO.µUBM is he supe ec o mean om he Uni e sal Backg ound Model
(UBM), he e e ence model ep esen ing he a e age beha iou . wOis he la en a iable o
he u e ance O, wi h a s anda d no mal p io dis ibu ion and Tis a low ank ma ix de ining
he o al a iabili y subspace.
The i- ec o es ima ion looks o he bes alue o he la en a iable wOso as o explain
he gi en u e ance by means o he adap ed model. Fo his pu pose, we es ima e he pos e io
dis ibu ion o he la en a iable wOgi en he u e ance O. The i- ec o ep esen a ion w
co esponds o he mean o his pos e io dis ibu ion. De ined in [Dehak e al., 2011], he i-
ec o is o mula ed as:
w= C
X
c=1
TT
cΣ−1
cNc(O)Tc+I!−1C
X
c=1
TT
cΣ−1
c˜
Fc(O)(7.4)
=1
N(O) C
X
c=1
TT
cΣ−1
c
Nc(O)
N(O)Tc+1
N(O)I!−1C
X
c=1
TT
cΣ−1
cNc(O)˜
Fc(O)(7.5)
= C
X
c=1
TT
cΣ−1
cαcTc+1
N(O)I!−1C
X
c=1
αcTT
cΣ−1
c˜
Fc(O)(7.6)
=Ψ−1(O, α)
C
X
c=1
αcΓc(O) =
C
X
c=1
αcΨ−1(O, α)Γc(O) =
C
X
c=1
αcGc(O)(7.7)
whe e Tc ep esen s he po ion o he ma ix Ta ec ing he c h componen o he UBM. Σc
symbolizes he co a iance ma ix o he c h componen o he UBM. Nc(O)and ˜
Fc(O)a e he
120
Chap e 7. S udy o embeddings o sho u e ances
ze o h and cen e ed i s o de Baum Welch s a is ics o u e ance O. These s a is ics ep esen
he numbe o samples om componen cand he accumula ed de ia ion wi h espec o he
mean o he same componen espec i ely. N(O)symbolizes he o al numbe o ames in he
u e ance O. Finally, he e m ˜
Fc(O)is he a e age de ia ion pe sample o he u e ance o he
componen co he UBM.
The o mula ion o i- ec o s o e s special cha ac e is ics. Fi s , he alue o C, he numbe
o aced acous ic uni s o disc imina e, is ixed in he UBM. I s alue is equal o he numbe
o Gaussian componen s in he UBM. The e o e, Gc(O) ep esen s he con ibu ion pe sample
o he i- ec o om componen c, and he weigh αcis he p opo ion o ames assumed o be
sampled om same c h componen . Fu he mo e, i- ec o s ha e no speake awa eness in hei
o mula ion. They simply s o e he a ia ions in he acous ic uni s wi hin an embedding. These
de ia ions om he a e age beha iou , p ope ly ea ed by he backend, a e esponsible o he
pe o mance in speake iden i ica ion sys ems.
7.3.3 Sho u e ances
Now we conside he sho u e ance scena io. Acco ding o he p e ious analysis, embeddings
wo k well i he dis ibu ion o acous ic uni s αis simila o he e e ence dis ibu ion and he
pa icula con ibu ions Gc(O)a e es ima ed wi h low unce ain y. These wo equi emen s a e
eassu ed as long as he u e ance con ains mo e and mo e da a. Conce ning sho u e ances,
hei low amoun o da a makes hem likely o ha e hei dis ibu ion o acous ic uni s α a om
hei e e ence. Fo he same eason sho u e ances may also su e om la ge unce ain y in
hei phoneme es ima ions Gc(O). Hence, deg ada ion in sho u e ances can be explained by
he ollowing easons:
•E o s in he con ibu ion o phonemes. Some con ibu ions Gc(O)we e es ima ed
wi h e y li le in o ma ion. Then he unce ain y o hei es ima ion inc eases. Mul iple
alues wi hin his unce ain y ange as Gc(O)′can be es ima ed ins ead, commi ing he
e o E=Gc(O)′−Gc(O).
•Misma ch in he phoneme dis ibu ion. The dis ibu ion o he weigh s αdoes no
ma ch he e e ence α, de ined by language cha ac e is ics. This deg ada ion causes he
e o E=PC
c=1(αc−αc)Gc(O). The ex eme case happens when some acous ic uni s
a e no p esen in he u e ance, i.e. hey a e missing. In his si ua ion hei weigh αca e
equal o ze o, also o cing he missing es ima ions Gc(O) o be se o ze o, as i hey we e
occluded. The deg ada ion due o he misma ch in he phoneme dis ibu ion is compa ible
wi h he e o s in he con ibu ion o phonemes.
121

E ec s o he sho u e ances in i- ec o s
T adi ionally e o s ha e been a ibu ed o he con ibu ions pe acous ic uni . This is
specially ue when adi ional embeddings, e.g. i- ec o s, include an unce ain y e m in
i s calcula ions. Fo his eason, his so o e o was he i s a emp ed o deal wi h, e.g.
[Kenny e al., 2013]. Howe e , o he bes o ou knowledge no p e ious wo k has co e ed he
deg ada ion due o he phoneme dis ibu ion, which can cause simila le els o deg ada ion.
7.4 E ec s o he sho u e ances in i- ec o s
The phone ic dis ibu ion in an u e ance has impo an implica ions du ing he embedding ex-
ac ion. Embedding shi s due o inco ec con ibu ions Gc(O)a e complemen a y o hose
c ea ed by he misma ch in he phone ic dis ibu ion. In his sec ion we illus a e hei impac
wi h i- ec o s. This choice o well-known embeddings makes he s udy o bo h p oblems mo e
illus a i e in a simple way.
Fo his pu pose, we p opose a small dimension i- ec o expe imen o es he e ec s o
sho u e ances in some a i icial con olled da a. Gi en an e alua ion UBM i- ec o pipeline,
we compa e he i- ec o s ob ained om an o iginal u e ance and hose ob ained om he same
u e ance a e unde going con olled sho -u e ance modi ica ions. These modi ica ions a ec
bo h he acous ic uni dis ibu ion αand hei con ibu ions Gc(O). We make use o he ol-
lowing expe imen al se up: We i s sample a la ge a i icial da a pool om a UBM i- ec o
pipeline. This da a pool consis s o mo e han en housand independen u e ances, wi h one
hund ed wo-dimension samples each. The UBM is a 4-Gaussian GMM whose componen s a e
loca ed in (0,0),(0,10),(10,0) and (10,10), all o hem wi h he iden i y ma ix as co a iance.
The gene a i e i- ec o ex ac o has a 3-dimension hidden a iable subspace. Wi h en hou-
sand o hese u e ances we ain ou e alua ion pipeline, an al e na i e UBM i- ec o sys em.
Fo simplici y we sha e he gene a i e UBM. Rega ding he i- ec o ex ac o , we ain a model
wi h only a wo-dimension la en subspace. This dimension educ ion be ween gene a ion and
e alua ion has been conside ed o imi a e eal li e, whe e he gene a ion o da a is a oo complex
p ocess ha we only can app oxima e.
F om he emaining da a pool we choose wo ex a u e ances, unseen du ing he model
aining, o e alua ion pu poses. Because hese wo u e ances a e independen , we assume
hem o ep esen wo di e en speake s. In Fig. 7.1 we ep esen hem, ed and blue espec-
i ely. The ep esen a ion includes h ee pa s: In he i s pa we show he o iginal ea u e
domain, i.e. he u e ance se o ea u e ec o s. Each ellipse in he igu e ep esen s he dis i-
bu ion o each Gaussian in hei GMMs. The image also includes in g een he ep esen a ion
o he UBM model. The second pa in Fig. 7.1 ep esen s he same ed and blue u e ances in
122
Chap e 7. S udy o embeddings o sho u e ances
he la en space by means o he pos e io dis ibu ion o he la en a iable w. The hi d pa o
Fig. 7.1 illus a es he loca ion o he pa icula es ima ions pe componen Gc(O) o he wo
u e ances in he la en space. Reddish es ima ions co espond o he ed speake while bluish
ellipses ep esen he phonemes o he blue speake .
✲ ✲✁ ✵ ✁  ✻ ✽ ✶✵ ✶✁ ✶
✲
✲✁
✵
✁

✻
✽
✶✵
✶✁
✶
▼❋❈❈ ✂s ❉✐♠❡♥s✐♦♥
✄
☎
✆
✆
✷
✝
❞
✞
✟
✠
✡
✝
☛
✟
☞
✝
P ✌ ❥✍❝✎✏✌✑ ✏✑ ✒✍❛✎✉ ✍ ❙♣❛❝✍
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
a) Da a domain b) i- ec o domain c) Componen s in I . domain
Figu e 7.1: Scena io o in e es . a) U e ances ed and blue in he ea u edomain, wi h
he UBM componen s in g een. b) U e ances ed and blue in he i- ec o domain. c)
P ojec ions o he GMM componen s in he i- ec o domain o u e ances ed ( eddish
ellipses) and blue (bluish ellipses).
Following he desc ibed se up we can ca y ou an analysis o deg ada ion in sho u e ances.
Fi s , we illus a e he phoneme dependen es ima ion e o due o limi ed da a. Fo his eason
we es ima e he pos e io dis ibu ion o he embeddings o mul iple u e ances only di e ing
he numbe o samples. The dis ibu ion o phonemes α emains unal e ed. Theo e ically,
he embeddings should no su e any bias, bu i s unce ain y should ge la ge as long as he
u e ances con ain less da a. In Fig. 7.2 we compa e he o iginal u e ances o hose ob ained
wi h one i h o he da a and one en h o he da a.
Fig. 7.2 illus a es he pos e io dis ibu ion o he la en a iable o he sho u e ances
(dashed-line ed and blue ellipses) as well as he o iginal u e ances ( ed and blue ellipses wi h
con inuous line espec i ely). The loca ion o he ellipse ep esen s he mean o he pos e io
dis ibu ion while i s con ou he unce ain y. As expec ed, he o iginal e e ence u e ance and
hei sho e e sions p esen e y educed shi s among hemsel es, wi h almos concen ic
ellipses. While he blue speake su e s almos no deg ada ion, he ed speake biases a e mo e
no iceable. Besides, he illus a ion shows ha he less da a in he u e ance, he bigge he
unce ain y o he es ima ion.
Now we s udy he impac o he dis ibu ion o acous ic uni s αon he embedding. In he
e e ence u e ances his dis ibu ion was uni o m, his is, 25% o he samples came om each
componen . We now modi y his dis ibu ion o bo h u e ances, ed and blue. In Fig. 7.3 we
show he pos e io dis ibu ions o he o iginal u e ances ( ed and blue ellipses wi h con inuous
123
E ec s o he sho u e ances in i- ec o s
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
Figu e 7.2: Compa ison o pos e io dis ibu ion o he i- ec o s wi h e e ence
phoneme dis ibu ion. Con inuous line ellipse ep esen s he o iginal u e ance while
dashed-lined ellipses illus a e u e ances wi h he limi ed da a.
line) as well as he al e ed sho u e ances (dashed-line ed and blue ellipses). In he illus a ed
example hal o he ea u e ec o s a e sampled om a single componen o he GMM while
he emaining da a is e enly sampled along he o he componen s. We ha e s udied he e ec
wi h he ou componen s in he GMM.
Illus a ed esul s in Fig. 7.3 e eal he ele ance o he dis ibu ion o phonemes α o i s
p ope modelling. The modi ica ion o he dis ibu ion o weigh s makes he ed speake o
o e ou di e en ep esen a ions o he same embedding. Besides, hese ep esen a ions a e
no o e lapped among hemsel es, beyond he unce ain y egion om he o iginal u e ance.
The e o e, hese al e na i e embeddings a e likely o ail. Ne e heless, no all speake s beha e
equally. Whils ed speake is deg aded, ou blue speake has su e ed he same al e a ions
wi hou any isible shi on his/he embeddings.
The scena io wi h a dis o ed phoneme dis ibu ion can be aken o he limi . In his si ua ion
some componen s do no con ibu e o he inal embedding. This scena io is he mos ad e se,
signi ican ly modi ying he dis ibu ion o pa e ns αand some es ima ions pe phoneme Gc(O)
being se o ze o. In his expe imen we ha e dis u bed he dis ibu ion o acous ic uni s α
o cing wo o he componen s o ze o. In Fig. 7.4 we illus a e he six possible scena ios in
e ms o he non-con ibu ing componen s. The esul s a e shown o he wo es speake s
ed and blue, wi h con inuous line ellipses o he e e ence u e ances and dashed-line ellipses
o hei al e ed e sions. Acco ding o he ep esen a ions shown in Fig. 7.4, embeddings
om u e ances wi h missing componen s expe imen la ge biases wi h espec o he e e ence
124
Chap e 7. S udy o embeddings o sho u e ances
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
1s Componen 2nd Componen
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
✲ ✲✁ ✵ ✁  ✻ ✽
✲✂
✲✁
✲
✄
✵
✄
✁
✂
■☎✈❡❝ ♦ ✶s ❉✐♠❡♥s✐♦♥
✆
✝
✞
✟
✠
✡
☛
☞
✷
✌
❞
✍
✎
✏
✟
✌
✑
✎
☛
✌
P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔
3 d Componen 4 h Componen
Figu e 7.3: Compa ison o pos e io dis ibu ion o i- ec o s wi h modi ica ions in
he phoneme dis ibu ion α
embeddings. These shi s a e mo e signi ican han hose p e iously seen wi h less ex eme
dis o ions in he phoneme dis ibu ion α. Some o he hypo hesized embeddings a e a beyond
he unce ain y om he o iginal u e ance. The biases su e ed by he u e ances a e no he
same o bo h speake s. Again, he blue speake su e s no ele an deg ada ion. This beha iou
i s in ou hypo hesis because he missing componen s scena io is he limi case o phoneme
dis ibu ion deg ada ion.
In all ou expe imen s he ed speake has su e ed om s ong deg ada ions while he blue
speake has emained almos unal e ed. This di e en beha iou is a consequence o he loca-
ions o he phone ic es ima ions Gc(O) o each speake . On he one hand, as shown in Fig. 7.1,
ou blue speake has i s componen s e y close o each o he , p o iding obus ness agains dis-
ibu ion modi ica ions. On he o he hand ou ed speake has i s componen s much u he
125