UPC System for the 2015 MediaEval Multimodal Person Discovery in Broadcast TV Task
Abstract
This project verses about the system that UPC developed to participate in the Multimodal Person Discovery in Broadcast TV task in MediaEval 2015. The main objective of this task is to answer the two questions: Who speaks when? and Who appears when using any sources of information in a TV Broadcast scenario.
Full text
UPC System for the 2015 MediaEval Multimodal Person Discovery in Broadcast TV Task universitat polit` ecnica de catalunya Miquel Angel India Massana supervised by Javier Hernando Pericas 2015
Abstract This project verses about the system that UPC developed to participate in the Multimodal Person Discovery in Broadcast TV task in MediaEval 2015. The main objective of this task is to answer the two questions: Who speaks when? and Who appears when? using any sources of information in a TV Broadcast scenario. UPC has created a system that combines different monomodal technologies (Speaker Diarization, Face Diarization and Optical character Recognition with Named Entity Recognition) to solve this problem. This system has been developed by the UPC Image Processing Group (GPI) and the Center for Language and Speech Technologies and Applications (TALP). However, this project will be more focused on the participation of TALP in the competition. From TALP, we have worked in two different blocks: the speaker diarization system and the multimodal system. In one hand, the speaker diarizatior system has been set-up for the challenge and two SAD (Speech activity detector) have been designed. In the other hand, we have created with the GPI a fusion approach based on an ILP system, whose constraints are set weighing the overlapping tracks of the different modalities. 2
Acknowledgments This project could not have been possible without the help and the supervision of my advisor Javier Hernando Pericas. I would want to thank him for the support given and for the chance of participating in the 2015 MediaEval. This competition allowed me to develop my project in a more competitive environment, increasing at the same time my interest on the research field. I would also want to mention all the support given by the TSC department. From TALP, I would want to thank Abraham Woubie, who helped me with all the issues related to the Speaker Diarization script, and Omid Ghahabi, who advised me about the SAD development and its setup. From the Image Processing Group, I would want to appreciate all the effort given by David Varas and Ramon Morros to improve the performance of their monomodal technologies and the multimodal algorithm. In addition, I would want to express all my gratitude to Artemi Casanovas, also for his help and support along the project. Finally, I could hardly not find anyone to be more thankful with than my family, who have always supported and advice me in any decision I took. Therefore, I would want to give all my gratitude to them. Special mention to my parents, who have inspired and encouraged me along the whole engineering career. Due to this I would want to dedicate this project and all the effort put on it to them. 3
Contents Abstract 2 Acknowledgments 3 1 Introduction 9 1.1 Context and Motivations ......................... 9 1.2 Objectives ................................. 10 1.3 Outline ................................... 12 2 State of Art 13 2.1 Person Recognition systems ....................... 13 2.1.1 Monomodal Technologies ..................... 13 2.1.2 Fusion Algorithms ........................ 15 2.2 Speaker Diarization Systems ....................... 15 2.2.1 Feature Extraction ........................ 16 2.2.2 Speech Activity Detection (SAD) ................ 18 2.2.3 Speaker Clustering ........................ 19 2.2.4 Evaluation Metric ......................... 23 3 Speaker Diarization System 25 3.1 Front-End Processing ........................... 27 3.2 Cluster Initialization ........................... 31 3.3 Cluster Merging .............................. 34 3.4 System output ............................... 34 4
Contents 5 4 The 2015 Multimodal Person Discovery in Broadcast TV task 36 4.1 Context .................................. 36 4.2 Task Description ............................. 37 4.3 Database .................................. 39 4.4 Evaluation System ............................ 40 5 UPC Person Discovery System 42 5.1 Monomodal Technologies ......................... 43 5.1.1 Speaker Diarization System ................... 43 5.1.2 Face Diarization System ..................... 44 5.1.3 OCR with Named Entity Detection ............... 45 5.2 UPC Multimodal Algorithm ....................... 46 5.2.1 Fusion Method .......................... 48 5.2.2 Fusion Strategies ......................... 51 5.2.3 Evidence Extraction ....................... 53 6 Results 55 6.1 Experimental Set-up ........................... 55 6.1.1 Speaker Diarization System ................... 55 6.1.2 Multimodal Person Discovery Task ............... 57 6.2 Result Analysis .............................. 58 6.2.1 SAD Results ............................ 58 6.2.2 MediaEval 2015 Results ..................... 60 7 Conclusions 63 Bibliography 65
List of Figures 2.1 Speaker Diarization example ...................... 16 2.2 Mel-scale Filter Bank example ...................... 17 2.3 MFCC Extraction ............................. 18 2.4 Interpretation of Hierarchical Clustering ................ 20 2.5 GMM Clustering example ........................ 21 3.1 UPC Diarization Algorithm ....................... 26 3.2 SAD output file example ......................... 29 3.3 Hybrid SAD Algorithm .......................... 30 3.4 HMM Toplogy for Speaker Clustering .................. 31 3.5 HMM State scheme ............................ 33 3.6 Speaker Diarization output file example ................. 35 4.1 Multimodal Person Discovery Task ................... 37 4.2 Evidence example ............................. 38 5.1 Face Diarization Algorithm ....................... 44 5.2 Optical Character Recognition scheme ................. 46 5.3 Temporary representation of the monomodal indexations ....... 47 5.4 Speech/Video Fusion scheme ....................... 48 5.5 ID relationship file example ....................... 49 5.6 Global Decision example ......................... 50 5.7 Video/Text with Speech/Video Approach ................ 52 5.8 Speech/Video with Video/Text Approach ................ 53 5.9 Evidence file example ........................... 54 6
List of Figures 7 6.1 Video/Text with Speech/Video graphic example ............ 62
List of Tables 4.1 Repere Corpus Distribution ....................... 39 4.2 Repere Corpus Composition ....................... 40 6.1 Hybrid SAD Evaluation Results ..................... 58 6.2 Energy-based SAD DER Evaluation ................... 59 6.3 Energy-based SAD Evaluation Repere Results ............. 60 6.4 2015 MediaEval Task Results ...................... 61 8
Chapter 1 Introduction The purpose of this initial chapter is to present the motivations that impulse me to do this project and the objectives aimed to achieve. This project was initially thought to verse about the research done in the Speaker Diarization field. However, given the opportunity of participating on the 2015 Multimodal Person Discovery in Broadcast TV task, the research was extended including the Person Recognition field. The study of this field implies a wide knowledge of different areas like Image Processing, Speech Processing etc. Due to this fact, this project only responds to the work related to Speaker Diarization Systems and Fusion Algorithms. Despite being briefly described, those areas related to Image Processing were not developed by the author . 1.1 Context and Motivations People interested in other people Nowadays, people is very interested in looking for information of other people. This has already been achieved somehow. It is easy to find information about someone only looking for it in a web searcher. However, almost all the information we get is purely text. If we try to find some evidences of someone in video databases, this task probably becomes harder. The main issue of this topic is the lack of technology to tag automatically person identities in video databases. Normally, the content of 9
Chapter 2. State of Art 16 of an audio file. Therefore given an audio signal, the speaker diarization answers the question Who speaks when? as it is shown in figure 2.1. Figure 2.1: Speaker Diarization example In this project, this speaker indexation has been used combined with other indexations (face and text) as a part of a Multimodal Person Recongintion System. However Speaker Diarization is more usually used purely in the Speech Recognition field. For example, speaker clustering can be used in Automatic Speech Recognition systems (ASR) to adapt the acoustic models to be speaker dependent and therefore increase recognition performance. Speaker Diarization Systems are usually divided in different blocks due the several algorithms that include, so we will describe separately some sections of the diarization. The following analysis of the state of art is split in four parts: Feature Extraction, Speech Activity Detection (SAD), Speaker Clustering and Evaluation Metric. 2.2.1 Feature Extraction Human voice is characterized by a high degree of variability within the same speaker known as intra-speaker variability. Any variations like the emotional stage or the environment, can make two speech signals of the same speaker become totally different. Otherwise, this voice characterization also has a high impact in terms of identifying different speakers. The variability between features of different speaker is known as inter-speaker variability. Due to this voice variability, there is a wide research scene about types of features that can properly characterize a single speaker.
Chapter 2. State of Art 17 There are several famous parametrization features for speaker diarization as Mel Frequency Cepstral Coefficients (MFCC), Linear frequency cepstral coefficients (LFCC), Perceptual Linear Predictive (PLP), Linear Predictive Coding (LPC) and others. Nowadays, the most common features used in the speech recognition field are the MFCCs, which are also the features used in the TALP Speaker Diarization System. Mel Frequency Cepstral Coefficients (MFCC) Mel Frequency Cepstral Coefficients are defined as a set of coefficients which represent the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency [12]. This mel scale of frequency, which is applied trough a filter bank (see figure 2.2), is implemented as an approximation of the auditory human system performance, where the capacity of discerning the difference between two closely spaced frequencies decreases on the highest frequencies. The mel scale attemps to map the perceived frequency of a tone, or pitch, onto a linear scale: approximately the scale is linear between 0−1000 Hz and logarithmic above 1000 Hz. Hence there is more emphasis on the lower frequencies than on the higher ones. Figure 2.2: Mel-scale Filter Bank example MFCCs can be computed as follows: •Windowing the signal into 20 −30 ms frames. Normally a Hamming window is applied with overlapping (10 −20 ms distance between windows).
Chapter 2. State of Art 18 Figure 2.3: MFCC Extraction •Compute the Discrete Fourier Transform (DFT). •Take the square-magnitude. •Warp the frequencies (Apply mel-scale by means of a filter bank). •Compute the Discrete Cosine Transform (DCT) of the log filter-bank energies. 2.2.2 Speech Activity Detection (SAD) Speech Activity Detectors are those kind of systems whose aim is to detect the presence or absence of human speech. These systems, also known as voice activity detectors (VAD), are usually used before any other algorithm in Speaker Diarization Systems. In terms of speaker clustering and speaker segmentation, the diarization performance is more accurate if the segments without speech are previously discarded. Hence is very important to firstly identify those segments which contain speech and those which does not in the signal. Depending on the background of the audio signal, the complexity of a SAD may vary. In signals with a high SNR and assuming that the non-speech segments only contain silence or noise, energy-based VADs perform efficiently and their computational cost is low. However, with a low SNR signal or with a background that contains other components like music segments, energy-based SADs perfomance is not reliable. As an alternative of energy-based SADs, model-based approaches are a good option. Despite having a higher a computational cost, model-based SADs are able to perform in the worst signal conditions. Nevertheless, that performance totally depends on the training of the models, which is better implemented with supervised systems. Both methods can also be combined as an hybrid speech/nonSpeech detector like in [3].
Chapter 2. State of Art 19 Energy-based Detection We refer to energy-based approaches those algorithms that compare short-term energy features with a predefined threshold in order to divide speech and non-speech segments. Normally these algorithms are split in two blocks: a preprocessing block and a decision block. The first block is used to improve the quality of the signal. For instance, sometimes noise reduction techniques are applied to improve the SNR. The second block extracts the energy features of each frame and group the frames in speech/non-speech segments according if their energy features are above or below the threshold. The decision block follow the next steps: 1. The energy feature is extracted from each frame. 2. This value is compared with the threshold. If the value is above the threshold the frame is considered speech. If it is lower than the threshold, that frame is considered non-speech . 3. With the tagged frames, a speech/non-speech segmentation is performed. This detection is a short-term approach, so in order to do a coherent segmentation some time restrictions are imposed to define the size of the segments. For example in [3], a state machine is proposed to apply time constraints and fixing a minimum speech turn length. 2.2.3 Speaker Clustering As one of the most usual techniques for offline systems, TALP diarization system uses an Agglomerative Hierarchical Clustering (AHC). Hierarchical Clustering is based in splitting or merging clusters iteratively until an optimum number of clusters is reached. Depending on if clusters are split or merged, two different clustering techniques are considered: Bottom-up Clustering and Top-down Clustering. As we see in 2.4, Bottom-up Clustering systems are those which start with a big number of clusters and via merging techniques converge to the optimum amount of clusters. Top-down Clustering in a opposite way, starts with a minimum number of clusters and those are divided with splitting techniques until the optimum number of clusters is reached.
Chapter 2. State of Art 20 Figure 2.4: Interpretation of Hierarchical Clustering These hierarchical approaches need some settings to be adjusted in order to perform efficiently. In either of both clustering types, some items are needed to implement the approach: •Only for the bottom-up clustering approach, it is needed to define an initial number of clusters, whose value must be above the optimum number. In fact, it needs to be so much higher because some merging techniques does not perform correctly if this value is close to the optimum number of clusters. Therefore is very important to have an expression which sets this initial parameter. Otherwise that parameter can also be obtained training the system in a development database testing a set of values. •For both clustering approaches it is needed a method to evaluate the similarity between the acoustic models and a stop criterion to determine when the optimum number or clusters is reached. For this criterion stop is usual to use the Bayesian Information Criterion (BIC). BIC provides statistic information about the likelihood of how the models fit the data. Therefore it allows to evaluate which clusters must be merged first and deciding when the cluster merging must finish. To implement both tasks TALP Speaker Diarization
Chapter 2. State of Art 21 System uses a modified BIC as in [1]. Gaussian Mixture Models (GMM) We refer to Gaussian Mixture Models (GMM) as the most common approach for speaker modelling. This approach is based on a weighted sum of gaussian component densities as the parametric probability function of a model. In this project, this model approach will be used twice. In one hand, in the Hybrid SAD (Energybased and Model-based) Speech and Non-Speech classes are modeled as GMMs, with several combinations in the number of gaussian component densities for each class. In the other hand, in the diarization system each speaker or cluster is modelled as a GMM. Figure 2.5: GMM Clustering example A Gaussian Mixture Model is defined as a weighted sum of Gaussian distributions[21]. Pr(x|µi,Σi) = M X i wig(x|µi,Σi) (2.1)
Chapter 2. State of Art 22 where xis a N-Dimensional vector that contains the data, i.e features, wiis the weight of each mixture and g(x|µi,Σi) is the component gaussian density. Each component gaussian density is defined as : g(x|µi,Σi) = 1 (2π)D 2|Σi|1 2 exp(−1 2(x−µi)TΣ−1 i(x−µi)) (2.2) with mean vector and covariance matrix µi, Σirespectively and the mixture weights wisatisfying the constraint: PM i=1 wi= 1. Therefore GMMs are represented with these three parameters as: λ={wi, µi,Σi}i= 1, ..., M (2.3) Cluster merging With the aim of deciding which clusters should be merged in an AHC, it is needed some algorithm able to evaluate the similarity between those clusters. In a parallel way to this comparison, it is also needed to set a criterion in order to stop the clustering once the optimum number of clusters is reached. The Bayesian Information Criterion (BIC) is implemented as an algorithm able to realize both tasks with the same procedure. Bayesian Information Criterion is defined as a method which evaluates the likelihood of a data given a model. Considering an acoustic segment X={x1, x2, ..., xX} with NXacoustic frames, modeled by θXwhich is an arbitrary model with a certain number of free parameters to estimate from the data, given by #(θX) which accounts for the complexity of such model. The general BIC expression of such model using the likelihood of the data is given by: BIC(θX) = log L(X, θX)−λ1 2#(θX) log(NX) (2.4) where log L(X, θX) is the log-likelihood of the data given the considered model. Besides, the other component of the expression is a penalty term which is adjustable setting the λparameter. According to this expression and given two clusters/segments to compare, the problem is to decide whether such two segments are tied to the same speaker/cluster. We define Xand Yas the segments to merge and
Chapter 2. State of Art 23 Z=X∪Yas its union. NX,NYand NZdefine the number of acoustic samples of each segment/cluster where obviously NZ=NX+NY. The modified BIC equation is defined as: ∆BIC(Z) = BIC(X, Y ) = L0−L1(2.5) where L0is defined as the log-likelihood performed by a model θZwhich takes into account the whole data Zand L1is defined by the sum of log-likelihoods performed by two independent models θXand θYon each data cluster Xand Y, respectively: L0= NX X i=1 log Pr(zi|θZ) + NZ X i=NX+1 log Pr(zi|θZ) (2.6) L1= NX X i=1 log Pr(zi|θX) + NZ X i=NX+1 log Pr(zi|θY) (2.7) where zi∈Rd,i= 1,2, ..., NZwhich are assumed to be independent vectors of acoustic features. The penalty term of the expression is avoided according to [1], where keeping the same number of features in both clusters and setting the number of features in θZas the sum of the features of θXand θY, the penalty term is equal to zero. This allows to skip the need of using a database to adjust that term. Hence given a pair of clusters, in the case that BIC(X, Y )>0, it implies than this two segments are best modeled merged with a θZmodel rather than split with θXand θY. If any pair of clusters BIC comparison is positive (BIC(X, Y )≤0) it implies that any of this pair combinations is best modeled merged. Therefore, as a stop criterion the number optimum of clusters is finally reached and the clustering algorithm ends. 2.2.4 Evaluation Metric . The main evaluation metric used for speaker diarization experiments is the Diarization Error date (DER). This metric measures the fraction of time where the speech or non-speech is incorrectly assigned. Numerically valued as a percentage,
Chapter 2. State of Art 24 DER is the sum of three kinds of error: •False-Alarm: Non-speech fraction time where speech segments are assigned. •Miss-Speech: Speech fraction time where speech is not assigned •Speaker error: Fraction of time where the speaker turn is incorrectly assigned. Despite being a metric thought to evaluate diarization results, DER is also very useful to analyze the performance of speech activity detection systems (SAD). Taking advantage of the information given by the miss-speech and false alarm errors, is possible to obtain an estimated evaluation of the SAD. For example if the miss-speech is so much lower than the false-alarm, that SAD is overdetecting speech frames. The opposite case (MS higher than the FA) indicates lack of speech detection of the SAD. Thus having a low but similar MS and FA, indicates a reliable implementation of the SAD.
Chapter 3 Speaker Diarization System The diarization system used for the Multimodal Person Discovery in Broadcast TV 2015 task is based on a Agglomerative Hierarchical Clustering (AHC) that uses a combination of Hidden Markov Models (HMM) and Gaussian Mixture Models (GMM) as model acoustic approach. This unsupervised system is also comprised of a cluster merging block based on a modified Bayesian Information Criterion (BIC) and uses Mel Frequency Cepstral Components (MFCC) as speaker features. As an algorithm based on a bottom-up clustering, a few blocks of the system are executed iteratevely. Due to this architecture of the system, this chapter will first make a brief description about the whole performance of the algorithm. Following to this introductory, each block will be described in more detail. TALP speaker diarization sytem, which algorithm is shown in figure 3.1, is split in three blocks. This three blocks also set the structure of the chapter, where the block A is explained in the Front-end Processing section, the block B corresponds to the Cluster initialization section, and the Cluster Merging section corresponds to the block C. Besides these three blocks, a last section in this chapter describes the file obtained as the result of the system. As it has already been mentioned, TALP Speaker Diarization System is based on a bottom-up clustering approach. Hence given an audio signal, this signal is segmented defining an initial number of K clusters. Those clusters will be merged until an optimum number of clusters is reached. With the last segmentation and clustering performed, a final speaker indexation from the audio signal is obtained. 25
Chapter 3. Speaker Diarization System 32 Kinit =N GinitRcc (3.2) where Nis total amount of data available per speaker cluster, Ginit is the number of Gaussian mixtures initially assigned to each speaker cluster and Rcc is a constant value that defines the number of frames per gaussian. As default values fixed in [18], Ginit is set to 5 gaussians and Rcc is set to 7 seconds per gaussian. Assuming that the Kinit value must be so much higher than the optimum number of clusters, an upper and a lower bound are set to guarantee the correct estimation of Kinit depending on the database used. For each segment, a GMM model is extracted to obtain the probability density of each cluster. The number of gaussian mixtures that comprises each model is initially assigned with Ginit as default value. Thus all the cluster GMMs are composed by the same amount of mixtures gaussians in the initialization of the model. However, for each iteration a pair of clusters will be merged. This means that the data of this pair of clusters will be pieced together in a new one. Due to this, in each iteration the number of gaussian mixtures per model is uploaded in the complexity section. This value is obtained according to the following expression, where Nj iis the number of frames belonging to the cluster iin the iteration j. Mj i=bNj i Rcc +1 2c(3.3) The training of the HMM is performed using an Expectation-Maximization (EM) algorithm applied on the GMMs. This means that for each cluster a GMM is trained. For each iteration those models are trained again with a new clustering according with the new complexity selection updated. Once the models are trained, the signal is segmented again creating new clusters trough applying Viterbi decoding. Speaker Segmenation With the aim to refine the data partition with each new HMM training, the speaker segmentation is implemented applying Viterbi decoding. Given the initial segmentation and a Kinit number of clusters, in terms of likelihood, the optimum number of clusters with its respective features is defined as :
Chapter 3. Speaker Diarization System 33 θopt, kopt = argmax θ,k {Pr(X, pbest|θ.k)}(3.4) where X is the audio data and pbest is the best sequence of states/models given a set of observations in terms of maximum likelihood. Hence the viterbi segmentation path gives us the best state/model for each observation. Applying this algorithm with the HMM/GMM topology allows to polish the initial segmentation, which may be wrong due to being uniform, i.e., a segment that contains two speakers after the initial segmentation. The HMM topology defines each state as a model composed by several sub-states, imposing a minimum duration for the speaker turn. As it is shown in figure 3.5, once entering to one state the model forces to jump for one sub-state to the next sub-state until the last sub-state is reached. Figure 3.5: HMM State scheme In this last sub-state, transition weights are differently defined. With an αtransition weight, the model stays in the same sub-state. In the other hand, this sub-state could jump to an another state with a β ktransition weight where kis the number of clusters reached in that time. According to [4], αand βare both set to 1 although then α+βis no longer equal to 1. This non-standard (but valid) HMM topology forces that once the minimum duration is reached in the model, the transition weights are only governed by acoustics. As a default value for TV Broadcast News according to [2,18], the minimum duration is set to 3 seconds.
Chapter 3. Speaker Diarization System 34 3.3 Cluster Merging Cluster Merging is the last step of an iteration in this agglomerative clustering. The aim of this subtask is to evaluate the similarity between clusters, and merge the most similar pairs. Therefore, firstly the distance between all the pair of clusters is evaluated. Following to this comparison process, a stop criterion algorithm decides if at least one pair of clusters must be merged or the clustering must stop in that iteration. Once a pair of clusters is merged, the training of the model and the segmentation are performed again. This process is iteratively done until no more clusters need to be merged. According to [1], where only one pair of clusters is merged per iteration, the criterion stop is defined with the following equation applied on each clusters pair: ∆BIC(D) = log p(D|θ)−(log p(Da|θa) + log p(Db|θb)) (3.5) This distance evaluation is a based on a modification of the original BIC [8]. Given a pair of clusters with their respective data Daand Dband their respective model features θaand θb,D=Da∪Dband θis a new model trained with features of both clusters. The number of features of this new model must be the sum of the number of features of θaand θb[1]. Thus 3.5 compares the log-likelihood of an independent model θtaken the data of both clusters against the log-likelihood of two separate models with its respective cluster data. This expression is applied with all the possible combinations of clusters. Therefore, the pair of clusters with the biggest value is considered the most similar pair. Unless this distance was negative BIC(D)≤0, this pair will be merged and the iteration will be restarted. When any pair of clusters satisfies this criterion BIC(D)≥0, the algorithm stops and the optimum number of clusters is reached. 3.4 System output Speaker diarization systems are implemented in order to perform both speaker clustering and segmentation tasks. Hence, once the stop criterion determines the ending of the clustering, the last segmentation is taken as the final result. The TALP
Chapter 3. Speaker Diarization System 35 speaker diarization sytem, based on [18], generates a .rttm file as the system output. This kind of file is a popular diarization extension that shows the time-stamps of the speaker turns with its cluster label. Figure 3.6 shows an example of a .rttm file, where each line contains the initial time of a speaker turn, its duration and its cluster assignation. Figure 3.6: Speaker Diarization output file example
Chapter 4 The 2015 Multimodal Person Discovery in Broadcast TV task The Multimodal Person Discovery in Broadcast TV news is a task from the Mediaeval Multimedia Benchmark. The challenge of this task is to understand how to use the information about people that comes from the speech and the image and combine them so as to determine who is speaking and who is present in the video. 4.1 Context TV archives maintained by national institutions such as the French INA, the Netherlands Institute for Sound & Vision, or the BBC are rapidly growing in size. The need for applications that make these archives searchable has led researchers to devote concerted effort to developing technologies that create indexes. Indexes that represent the location and identity of people in the archive are indispensable for searching archives. Human nature leads people to be very interested in other people. However, at the moment that content is created or broadcast, it is not always possible to predict which people will be the most important to find in the future. Someone who appeared in a broadcast, but was relatively unnoticed, might suddenly start generating a buzz and become a trending topic on social networks or search engines. For this reason, it is not possible to assume that prior biometric models will be 36
Chapter 4. The 2015 Multimodal Person Discovery in Broadcast TV task 37 available at indexing time. For some people such a model may not be available in advance, simply because they are not (yet) famous. In such cases, it is also possible that archivists annotating content by hand do not even know the name of the person. The goal of this task is to address the challenge of indexing people in the archive, under real-world conditions (i.e., there is no pre-set list of people to index). 4.2 Task Description The main objective of this challenge is to answer the two questions ”who speaks when?” and ”who appears when?” using any sources of information. Figure 4.1: Multimodal Person Discovery Task Given a set of raw TV broadcasts, each broadcast is presegmented into shots. These shots must be automatically tagged with the name(s) of people who can be both seen as well as heard in those segments. The list of people is not known a priori and their names must be discovered in an unsupervised way from provided text overlay or speech transcripts. Evidence For each unique person name in the test set, participants are also asked to provide the evidence justifying their assertion. In real-world conditions, this evidence would help a human annotator double-check the automatically-generated index, even for people they did not know beforehand.
Chapter 4. The 2015 Multimodal Person Discovery in Broadcast TV task 38 Figure 4.2: Evidence example Two types of evidence are allowed: •An image evidence is a shot during which a person is visible, and their name is written on screen. •An audio evidence is a shot sduring which a person is visible, and their name is pronounced at least once during a [shot sstart time - 5s, shot send time + 5s] neighborhood. For instance in Figure 4.2, shot #3 is a correct evidence for B because her name is pronounced in previous shot #2 (less than 5 seconds apart). Shot #1 is a correct evidence for A because his name appears while his face is visible. Unsupervised System As it was mentioned before, one of the main challenges of this task is not using external supervision. Although a development database is given to the participants to train their systems, the use of person-labeled external data to train biometric models is forbidden. Therefore, the only way to identify a person from the test set is by discovering their name somewhere in the test set, and finding a way to propagate it to every shot where they are speaking and visible at the same time. Apart from this restriction, use of external data is allowed if does not provide a direct link between a person’s biometric sample (voice, face, vocabulary, etc.) and his/her name. Here is an example list of things that are allowed using external data:
Chapter 4. The 2015 Multimodal Person Discovery in Broadcast TV task 39 •Using external data to train a person-independent face detector or speech activity detector is allowed (it does not necessitate actual person labels). •Using external data to train a UBM1 for speaker identification is allowed (it does not necessitate actual person labels). •Using external data to train a LDML matrix2 for face recognition is allowed (it does need to know whether any two faces are from the same person, but does not rely on the actual identity of each face). •Using a list of person names gathered from Wikipedia to filter the output of optical character recognition is allowed. •Adding this same list to the vocabulary of a speech-to-text system is allowed. 4.3 Database The development database used for this competition was the Repere Corpus [13]. This corpus consists on 60 hours of French videos annotated with visual (heads and embedded texts) and audio information. Its composition (see Table 4.1) is based on seven shows from two french channels and its data was collected to support the development of automatic systems for people multimodal recognition in videos. The set of shows given in this corpus is highly varied, however two different main formats are differentiated: TV News and TV debate. Table 4.2 gives a brief description about the content of each show and its portion of collected emissions. These shows were TV Show Channel Show duration (minutes) BFM Story BFM 50 Planete Showbiz BFM 2 Ca Vous Regarde LCP 50 Entre Les Lignes LCP 27 Pile Et Face LCP 27 LCP Info LCP 20 Top Questions LCP 15 Table 4.1: Repere Corpus Distribution
Chapter 4. The 2015 Multimodal Person Discovery in Broadcast TV task 40 selected according its gradual difficulties in audio and video contents. The selection criteria is the situation diversity so as to have the larger panel of examples. The focus is put on prepared vs. spontaneous speech, head size and orientation, camera motion and angle, lighting, etc TV Show Channel Kind of show Proportion) BFM Story BFM Information show 33.3% Planete Showbiz BFM Celebrity news show 8.3% Ca Vous Regarde LCP Political debate 8.3% Entre Les Lignes LCP Political debate 8.3% Pile Et Face LCP Political debate 8.3% LCP Info LCP Information show 16.7% Top Questions LCP Parliamentary questions 16.7% to the government sessions Table 4.2: Repere Corpus Composition The test database given to run the submissions was a set of 171 shows. The duration of each one was 30 minutes and its structure was purely TV news. Besides BFM Story and LCP Info shows, the test database were not very similar to the development data. 4.4 Evaluation System In this task two different outputs were evaluated: the shots and the evidences. Therefore, the idea is to evaluate if the identity assigned for each shot is correct and confirm if the participant can prove it with its evidence. Evidence Evaluation For each query q(e.g. nicolas sarkozy), the evidence with the higher Levenshtein ratio r between the evidence name and the query name is chosen to be the evidence for query q. The correctness C(q) of the q evidence is then computed as follows: •If groundtruth says that this is actually an evidence for the q query, then C(q) = rif r > 0.95 and C(q) = 0 otherwise.
Chapter 4. The 2015 Multimodal Person Discovery in Broadcast TV task 41 •If groundtruth says that this is not an evidence, then C(q) = 0. Evaluation of Labels For each query q, shots for which label is equal to an q evidence are selected to compute AP(q) average precision. Mean Average Precision (MAP) is then computed as the mean of AP(q) over all queries q: MAP =1 |Q|X q∈Q AP(q) (4.1) Evidence-weighted Mean Average Precision (EwMAP) In real-world conditions, evidences should help a human annotator double-check the automatically-generated index, even for people they did not know beforehand. Therefore, Mean Average Precision is modified slightly to take the correctness of evidences into account: EwMAP =1 |Q|X q∈Q C(q)·AP(q) (4.2)
Chapter 5. UPC Person Discovery System 48 weighing the overlaps between the tracks of each source. At the same time, the relation between clusters is estimated solving an ILP system whose constrains are set with these weights and with some restrictions imposed. Due to this relevance on the weighing performance, its method is designed according to the previous assumptions defined. Following to this introduction, the implementation of these approaches is described in more detail. First, the fusion method is explained with a detailed description of the weighing method and the ILP system. In the other hand, the two approaches considered for the task are presented. 5.2.1 Fusion Method The Fusion Method can be described as a three steps algorithm that given two indexation files, creates a file that contains each overlapped track with its corresponding identities tagged. Despite three possible combinations of fusions are considered (speech/video, speech/text and video/text), speech/text is discarded according to assumption 3. Speech/video and video/text fusions are both computed with the same method. However, in the video/text fusion only we keep the relation between faces and its names assigned. The only segmentation reliable for the output is the obtained in the speech/video fusion. In that segmentation, the set of tracks listed represent the temporary stamps with face and speech presence. Otherwise, the video/text segmentation only points those segments with a face and a person name shown, which is not useful because our aim is to search where people is shown and speaking. In Figure 5.4, the speech/video fusion scheme is presented. This fusion is split in three blocks: Figure 5.4: Speech/Video Fusion scheme
Chapter 5. UPC Person Discovery System 49 1. ID Relationship File Construction: The objective of this step is to create a file with a list of all the temporary overlaps between the tracks of the two sources. Besides being the input of the weighing step, this file (Figure 5.5) also defines the initial segmentation of the fusion. Once the assignation between the identities of the two sources is known, the labelled segmentation will be created as from this list. Tini Tend ID Speech ID Video 0.00 3.05 spk 4 ID2 2.00 5.05 spk 4 ID1 6.00 8.05 spk 2 ID5 . . . 0.00 3.05 spk 7 ID8 Figure 5.5: ID relationship file example 2. Global Decision: With the ID relationship file as input, this block assigns the relation between speech IDs and video IDs. Given a Nnumber of speech IDs and a Mnumber of video IDs, αand βare defined as NxMmatrixs. With this matrixs, the ILP system is defined as : max αij (X i X j αijβij) (5.1) X j αij ≤1 (5.2) βij =X∆tij #overlaps (5.3) where αij : [0,1] ∈Zis the assignment between the ispeech identity and the jvideo identity. Otherwise, βij is the weight of this assignment. Given that system, equation 5.1 defines the function to maximize and 5.2 is the main constraint, which restricts that each speaker ID must have one and only one video ID. The ILP system computation is performed using the library IBM CPLEX.
Chapter 5. UPC Person Discovery System 50 The weighing expression 5.3 has been designed according to the assumptions defined and the structure of each source. Face indexation, besides having more IDs than the speech or the text source, also contains overlapped tracks. This implies that is common to find speech or text segments with more than one face segment overlapped at the same time. Assuming that its more reliable the weight of a speech/text track overlapped with only one video ID than with more than one video ID, the weighing expression aims to make emphasis in this assumption. Therefore, each portion of overlapping between speech/text and a video ID is divided by the number of face IDs overlapped in that portion of the track. Figure 5.6: Global Decision example Figure 5.6 shows an example of how the weighing method is implemented in a speech/video fusion case. In this intermediate step, video/text fusion finishes its procedure because only matrix αis needed. Otherwise in the case of the speech/video fusion, a last step is necessary to end the fusion method and create the segmentation.
Chapter 5. UPC Person Discovery System 51 3. Annotation Reconstruction: Once matrix αis computed, the assignation between video and speech IDs is combined with the relation ship file construction. This process is based on discarding those tracks whose overlapped IDs are not the estimated in α. With this final step, the file obtained is an indexation of the tracks where someone is shown and speaking with its corresponding IDs. 5.2.2 Fusion Strategies For the MediaEval Multimodal Person Discovery task two different multimodal approaches were considered. As approaches which both use the previous two fusions mentioned (speech/video and video/text), the main difference between them resides in the order of these fusions. Video/Text with Speech/Video Approach This approach is defined as a two fusion step algorithm, that combining three different source extractions is able to perform a name indexation of the people who is shown and speaking in a video database. These two fusions (video/text and speech/video) are performed in the written order. Those video tracks with face identities that are not assigned to any text identity in the first fusion, are discarded to perform in the speech/video fusion. Combining those two fusions, this approach creates a result file with the name indexation proposed. This algorithm is computed as follows: 1. Video and text tracks are overlapped and weighed so as to obtain an assignation between its identities. Each text identity must be tagged to only one face identity. Face IDs that are not assigned to a text ID and their respective tracks are discarded for the next fusion. 2. Speaker and the previously filtered video tracks are fused. The overlapped video/speech tracks are weighed obtaining an assignation between its identities. In this case, each speaker must be also assigned to only one face ID. Those overlapped tracks without the correct identity pairs are discarded obtaining a labelled segmentation.
Chapter 5. UPC Person Discovery System 52 Figure 5.7: Video/Text with Speech/Video Approach 3. With the previous video/speech labelled segmentation and the video/text assignations obtained in 1, the segmentation is finally tagged with the person names. Hence, the result file reports the names of who is shown and speaking during the course of the video broadcast. Speech/Video with Video/Text Approach As a variation of the first approach, this alternative algorithm (see Figure 5.8) was also considered for the competition. The main difference with the other approach is computing first the speech/video fusion instead the video/text one. Therefore those video IDs without a speaker assigned are discarded for the video/text fusion. This algorithm is computed as follows: 1. Speech and video tracks are overlapped and weighed to obtain an assignation between its identities. Each speaker must be tagged to only one face ID. Face IDs that are not assigned to a speaker are discarded for the next fusion. Therefore, those overlapped tracks without the correct identity pairs are discarded obtaining a speech/face segmentation.
Chapter 5. UPC Person Discovery System 53 2. Text and the previously filtered video tracks are fused. The overlapped video/text tracks are weighed obtaining an assignation between its identities. In this case, each text ID must be also assigned to only one face identity. 3. With the video/speech tagged segmentation and the video/text assignations obtained in 2, the segmentation is finally tagged with the person names. Hence the result file reports the names of who is shown and speaking during the course of the video broadcast. Figure 5.8: Speech/Video with Video/Text Approach 5.2.3 Evidence Extraction Besides creating a name indexation of who is shown and speaking in video broadcasts, for the MediaEval task was also needed to prove the correct extraction of that information. As it is mentioned in 4.2, these evidences must show a shot where the person is visible and their name is written or where a person is visible and it names is pronounced. Due to the lack of implementation of an Automatic Speech
Chapter 5. UPC Person Discovery System 54 Recognition (ASR) technology in our system, the second option was initially discarded. Moreover, even in a Broadcast News scenario we consider that Automatic Speech Recognition data is not enough reliable to implement a system which creates evidences or perform the main task. With the aim of searching those shots where persons are visible and its names are written, our strategy is based on the third assumption. According to that, Speaker Diarization data is ignored depending only on the combination between the Name Extraction and the Face Diarization. Therefore, the evidence extraction is computed as follows: 1. All those frames where a text name is shown are extracted. In a Broadcast news scenario, frames with text associated have a lower presence in the video than the frames where faces are detected. So text is first used to discard a high amount of frames where evidence could not be found. 2. Those frames where any face is detected are automatically discarded. With this step we obtain all those frames where text and faces are detected. 3. For each text name set of frames, the first frame where only one face is associated is extracted. If in those set there are only frames with two or more faces, the frame with less faces is selected. 4. Finally, each text frame with its name is associated with its shot. alex salmond FPVDB07050409 VIS 01 000221 image charles berling FPVDB07031305 VIS 01 000100 image dominique hervieu FPVDB07011308 VIS 01 000484 image eric besson FPVDB07050308 VIS 01 000088 image fonky foued FPVDB07011308 VIS 01 000471 image jacques weber FPVDB07020505 VIS 01 000492 image jean pierre ceron FPVDB07011308 VIS 01 000421 image kristian autain FPVDB07011206 VIS 01 000318 image odon vallet FPVDB07022804 VIS 01 000182 image reda benjelloun FPVDB07041406 VIS 01 000105 image Figure 5.9: Evidence file example
Chapter 6 Results With the purpose of taking a look into the performance of the UPC Person Discovery system, this chapter verses about the experiments executed and the results obtained for the 2015 MediaEval competition. According to the structure followed in the project, this chapter is also split separating the Speaker Diarization Dystem and the Multimodal Person Discovery System. In both cases, the experimental set-up of the experiments tested is described in detail. In the Speaker Diarization System part, the parameterization of the two SAD approaches is highlighted in order to discuss which SAD is the most appropriate for the task. Otherwise, the multimodal part takes a look into the set of submissions uploaded for the competition. Given the results, those are discussed and the the overall performance of the algorithm is analyzed . 6.1 Experimental Set-up 6.1.1 Speaker Diarization System TALP Speaker Diarization System has been set-up expressly for the 2015 Multimodal Person Discovery in Broadcast TV task. Although the development database for this competition was comprised by several kind of scenarios, the test database was based totally on Broadcast TV News. Due to this fact, the set-up was defined according to [2,18], where we find a reliable configuration for this scenario. This configuration 55
Chapter 6. Results 56 includes all the values set as default defined in chapter 3. Besides these parameters, the initial number of clusters were bounded according to an assumption applied in the development database. Assuming than at least the double of the speakers in a show is needed as Kinit to perform correctly the clustering, Kinit was bounded to [80,120] considering less than 40 speakers per show. Speech Activity Detection Experiments Speech Activity Detection was the first block to set in the system. Thus despite not having the development database in that moment, the performance of both approaches was tested. This performance was firstly analyzed in the Tecnoparla Corpora. Hence given the results obtained in the Tecnoparla corpus, the best approach was selected so as to evaluate it later with the development database. The Tecnoparla corpus consists on a set of ` Agora TV shows. ` Agora is a catalan show which was aired on Monday nights in Television of Catalunya channel. It is a debate show highly moderated and with a high variation in topics and invited speakers. The total audio time of the database is about 42 hours divided in 34 shows (approximately 1 hour and 20 minute each show), each one corresponding to an airing day. Each show has been split in two halves to delete the commercials present during the airing. As a debate corpus, each show barely contains long silence segments and music only co-exists with the signal as background. The energy-based SAD is set up based on three parameters. In one hand, the first parameter αdefines the frame energy threshold. More specifically, according if the quotient of the frame average power and the average power mean is above or below this threshold, this frame is considered as speech or non-speech. In the other hand, the minimum duration of speech and non-speech turns are also taking to account as parameters to set in the approach. Otherwise, the Hybrid SAD approach is set up with a higher number of parameters compared to energy-based. Firstly, the same αparameter as the energy-based approach is set for the first block of the SAD, where the frames are split into speech/non-speech to create the models. Each of these models (speech and non-speech) are trained with a configurable number of mixtures per model, thus this number of components is set for each model. Finally, the minimum duration of speech and non-speech turns is also set.
Chapter 6. Results 57 SADs performance was analyzed evaluating the whole diarization system. Through the DER evaluation metric, the system was firstly evaluated with ` Agora shows using annotations created for Automatic Speech Recognition tasks. These transcriptions do not annotate accurately the silences (only big silence segments are annotated). However, the DER evaluation gives helpful information to evaluate the performance of a SAD as it is explained in 2.2.4 . Once the best approach was selected and given the development database, the SAD was evaluated with the best setups tested in the ` Agora corpus. 6.1.2 Multimodal Person Discovery Task For the 2015 MediaEval competition, it was allowed to submit more than one result file. Although only one submission could be defined as primary and evaluated for the competition, the evaluation results of each file submitted were shown. These submissions were evaluated twice. For the first evaluation, the result files were evaluated only in approximately the 8% of the whole test data. Otherwise, in the second evaluation more annotations were transcript and the submissions were evaluated with approximately the 11% of the corpus. Taking advantage of the possibility of submitting more than one result file, we uploaded several result files with some variations taken. Those variations were based on modifying the architecture of the fusion algorithm, replacing some blocks with the baseline system given or combining the monomodal technologies before the multimodal algorithm implementation. The evidence extraction was always performed with the same steps. The following variations were implemented in combination in the set of submissions uploaded: •Despite considering the whole audio signal as the input of the speaker diarization system, it was thought to introduce only those segments where faces are detected. The aim of this variation was to improve the speaker diarization considering that some segments that could worsen the system (for example some advertisement with music as foreground) would be filtered. •The two fusion architectures described in 5.2.2. Depending on the order of the fusion, the result of the system may vary considerably. Many identities are
Chapter 7. Conclusions 64 that combines the video and text tracks before the speech/video fusion has shown the better performance. As a first approach in this topic, the person recognition system created by GPI and TALP still needs some improvement. The track weighing system implies a strong dependence on each monomodal system. Therefore, its needed to develop a new algorithm or setting the actual with a set of restrictions that decreases that dependence. Otherwise, the improvement of the monomodal technologies would be the other option to obtain a better performance of the whole system. In terms of Speaker Diarization, it would be interesting to develop a supervised SAD system, taking advantage of the development database so as to create models. This dissertation could not finish without a briefly explanation of how this project has induced me to increase my interest on possible future research. In the Speaker Diarization area, considering the recent research on Deep Neuronal Networks in the speech recognition field, it would be interesting to implement a speaker diarization system with that architecture able to out-perform the actual HMM approaches. Otherwise in the Person Recognition area, nobody still has found a smart way to introduce Automatic Speech Recognition (ASR) systems in multimodal algorithms. Although other technologies such as speaker diarization or OCR with NED gives more reliable information, ASR might add some extra information impossible to extract from other sources. Hence, even as a complementary source, ASR could be an interesting technology to add in Multimodal Person Discovery systems.
Bibliography [1] Jitendra Ajmera and Chuck Wooters. A robust speaker clustering algorithm. In Proc. ASRU, 2003. [2] Xavier Anguera. Robust speaker diarization for meetings. 2006. [3] Xavier Anguera, Mateu Aguil´o, Chuck Wooters, Climent Nadeu, and Javier Hernando. Hybrid speech/non-speech detector applied to speaker diarization ofmeetings. 2006. [4] Xavier Anguera, Chuck Wooters, and Javier Hernando. Automatic cluster complexity and quantity selection: towards robust speaker diarization. In In Proceedings of the Third international conference on Machine Learning for Multimodal Interaction, pages 248–256, 2006. [5] Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals. Speaker diarization: A review of recent research. IEEE TASLP, March 2010. [6] Frederic Bechet, Meriem Bendris, Delphine Charlet, Geraldine Damnati, Benoit Favre, Mickael Rouvier, Remi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linares, Jean Martinet, Gregory Senay, and Pierre Tirilly. Multimodal understanding for person recognition in video broadcasts. 2014. [7] T. Cali´nski and J. Harabasz. A dendrite method for cluster analysis. Communications in Statistics-Simulation and Computation, 3(1):1–27, 1974. 65
Bibliography 66 [8] Scott Shaobing Chen and P.S. Gopalakrishnan. Speaker, environment and channel change detection and clustering via the bayesian information criterion. 1998. [9] N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In CVPR 2015, 2005. [10] Marco Dinarelli and Sophie Rosset. Models cascade for tree-structured named entity detection. In Proceedings of 5th International Joint Conference on Natural Language Processing, pages 1269–1278, Chiang Mai, Thailand, November 2011. Asian Federation of Natural Language Processing. URL http://www.aclweb. org/anthology/I11-1142. [11] B. Epshtein, E. Ofek, and Y. Wexler. Detecting text in natural scenes with stroke width transform. In Proc. of Computer Vision and Pattern Recognition CVPR2010, pages 2963–2970, 2010. [12] Sadaoki Furui. Cepstral analysis technique for automatic speaker verification. IEEE Transactions on Acoustics Speech and Signal Processing, 29:254–272, 1981. [13] Aude Giraudel, Matthieu Carr´e, Val´erie Mapelli, Juliette Kahn, Olivier Galibert, and Ludovic Quintard. The repere corpus : a multimodal corpus for person recognition. 2012. [14] M. Guillaumin, T. Mensink, J. Verbeek, and C. Schmid. Face recognition from caption-based supervision. IJCV, 96(1), 2012. [15] J.Finkel, T. Grenager, and C. Manning. Incorporating non-local information into information extraction systems by gibbs sampling. In Proc. of the 43nd Annual Meeting of the Association for Computational Linguistics, pages 363– 370, 2005. [16] Nam Le, Di Wu, Sylvain Meignier, and Jean-Marc Odobez. Eumssi team at the mediaeval person discovery challenge. In MediaEval 2015 Workshop, 2015. [17] Bruce D. Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. pages 674–679, 1981.
Bibliography 67 [18] Jordi Luque. Speaker diarization and tracking in multiple-sensor environments. 2012. [19] J. Poignant, L. Besacier, G. Quenot, and F. Thollard. From text detection in videos to person identification. In Multimedia and Expo (ICME), 2012 IEEE International Conference on, pages 854–859, July 2012. doi: 10.1109/ICME. 2012.119. [20] Johann Poignant, Laurent Besacier, and Georges Qu´enot. Unsupervised speaker identification in tv broadcast based on written names. In IEEE Transactions on Audio, Speech and Language Processing, Institute of Electrical and Electronics Engineers (IEEE), volume 23, pages 57–68, 2015. [21] Douglas A. Reynolds and Richard C. Rose. Robust text-independent speaker identification using gaussian mixture speaker models. In IEEE Transactions ASSP Magazine, volume 3, pages 72–83, 1995. [22] Peter Rousseeuw. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math., 20(1):53–65, November 1987. ISSN 0377-0427. doi: 10.1016/0377-0427(87)90125-7. URL http: //dx.doi.org/10.1016/0377-0427(87)90125-7. [23] P. Salembier and L. Garrido. Binary partition tree as an efficient representation for image processing, segmentation and information retrieval. IEEE TIP, 9(4): 561–575, April 2000. [24] Jianbo Shi and Carlo Tomasi. Good features to track. pages 593–600, 1994. [25] Ray Smith and Google Inc. An overview of the tesseract ocr engine. In Proc. 9th IEEE Intl. Conf. on Document Analysis and Recognition (ICDAR, pages 629–633, 2007. [26] Carlo Tomasi and Takeo Kanade. Detection and tracking of point features. Technical report, International Journal of Computer Vision, 1991.
Bibliography 68 [27] Michal Uricar, Vojtech Franc, and Vaclav Hlavac. Facial landmarks detector learned by the structured output svm. In Gabriela Csurka, Martin Kraus, RobertS. Laramee, Paul Richard, and Jos Braz, editors, Computer Vision, Imaging and Computer Graphics. Theory and Application, volume 359 of Communications in Computer and Information Science, pages 383–398. Springer Berlin Heidelberg, 2013. ISBN 978-3-642-382406. doi: 10.1007/978-3-642-38241-3 26. URL http://dx.doi.org/10.1007/ 978-3-642-38241-3_26.