Full text
Facial Expression Detection for video sequences using local feature extraction algorithms Kennedy Chengeta and Professor Serestina Viriri 1University of KwaZulu Natal 2School of Computer Science and Mathematics, Westville Campus, Durban, South Africa [email protected] Abstract. Facial expression image analysis can either be in the form of static image analysis or dynamic temporal 3D image or video analysis. The former involves static images taken on an individual at a specific point in time and is in 2-dimensional format. The latter involves dynamic textures extraction of video sequences extended in a temporal domain. Dynamic texture analysis involves short term facial expression movements in 3D in a temporal or spatial domain. Two feature extraction algorithms are used in 3D facial expression analysis namely holistic and local algorithms. Holistic algorithms analyze the whole face whilst the local algorithms analyze a facial image in small components namely nose, mouth, cheek and forehead. The paper uses a popular local feature extraction algorithm called LBP-TOP, dynamic image features based on video sequences in a temporal domain. Volume Local Binary Patterns combine texture, motion and appearance. VLBP and LBP-TOP outperformed other approaches by including local facial feature extraction algorithms which are resistant to gray-scale modifications and computation. It is also crucial to note that these emotions being natural reactions, recognition of feature selection and edge detection from the video sequences can increase accuracy and reduce the error rate. This can be achieved by removing unimportant information from the facial images. The results showed better percentage recognition rate by using local facial extraction algorithms like local binary patterns and local directional patterns than holistic algorithms like GLCM and Linear Discriminant Analysis. The study proposes local binary pattern variant LBP-TOP, local directional patterns and support vector machines aided by genetic algorithms for feature selection. The study was based on Facial Expressions and Emotions (FEED) and CK+ image sources. Keywords: Local binary patterns on TOP ·Volume Local Binary Patterns(VLBP) Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 27 DOI : 10.5121/sipij.2019.10103
1 Introduction Video based facial expression analysis has received prominent roles of late, in crowd analysis, security and border control among others[13, 23]. Its also used in image retrieval, clinical research centers and social communication. Facial expressions remain the most e↵ective way of emotion display[14, 15, 11]. Previous work on 2D facial expression recognition focused more into single image frame based analysis than video or image sequence analysis. The former assumes one image is representing the facial expression where as for the video image sequence, each facial image is a temporal dynamic process[14, 15, 11]. Facial action units change in dynamic appearance and dynamic texture and are more difficult to track than static images. A video is classified as a spatial texture mixture in a temporal environment and dynamic features are retrieved[1, 15, 20]. Facial expression algorithms are either holistic or localized feature extraction algorithms. Holistic algorithms cover the whole facial image whereas localized algorithms break down a facial image into smaller units. The study focused on facial motions and locating key facial components namely, nose, eyes, face and mouth. The dynamic features were extracted using key algorithms like local GLBP from three orthogonal planes (LGBP-TOP) which is a LBP variant with Gabor filtering[14, 15, 11]. The LBP-TOP feature descriptor was used for retrieving video dynamic textures for facial expression changes [14, 15, 11]. The facial expression video image sequences were then modeled as a histogram sequence which is a sum of concatenated local facial regions[1]. The experiments used the extended Cohn-Kanade (CK+) as well as the FEED databases. The facial expression sequence was modeled into a histogram by adding local regions histograms for the LGBP-TOP maps. KNN and Support Vector Machines algorithms were used to classify the datasets. Feature selection was enhanced by genetic algorithms. The extended Cohn-Kanade database (CK+) experiments demonstrated that local feature extraction, support vector machines and genetic algorithms achieved the better results compared to holistic feature extraction methods in last couple of years[22, 24]. Video-based facial expression recognition is made of face detection, tracking and recognition[16, 10, 6, 17]. The video sequences picked up depict key universal expressions (surprise, sadness, joy, disgust and anger). Each signal expression is performed by 7 di↵erent subjects beginning from the neutral expression. The paper mainly focused on the integration of spatial-temporal motion LBP with Gabor multi-orientation fusion and compared the 3 LBP histograms on three orthogonal planes to find accuracy of facial expression recognition[1, 14]. A support vector machine classifier was used for each plane and the overall LBP-TOP algorithm. The LBP-TOP was also compared against its variants like LBP-MOP [1, 14]. Experiments conducted on the extended Cohn-Kanade (CK+) database and FEED database proved that the 3D LBP TOP algorithm compared better than the 2D image texture local binary pattern variants. Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 28
2 Literature Review In recent researches use of spatio-temporal representations has been successfully used to address limitations of static image analysis[8, 3–5, 2]. Successful research has been done by fusing PCA, Gabor Wavelets, local binary patterns and local directional patterns[23]. Classification has included support vector machines(SVM), AdaBoost, k-nearest neighbor and neural networks[7, 8, 14]. The permanent features like eyes, mouth or lips’ feature vectors were generated from the facial appearances in the spatial and frequency domains. 2.1 Static Facial Expression Analysis Background Clinical facial expression research has been widely studied with 2D images either as a combination of facial expressions in 2D images or universal global facial [7, 8, 3, 5]. The Facial Action Coding System (FACS) describes facial expressions based on a mix of action units (AU)[7, 8, 3, 5]. Each facial action is a representation of muscular facial actions resulting in sudden facial expression variations. Global facial expressions represent the wholesome facial changes. The key expression changes are being happy, sad, anger as well as fear. The FACS solution improved its 2D image success to 3D to analyze video image changes based on time, content quality and valence [7, 8, 3, 5]. Challenges with static 2D images Based Methods Major clinical research in facial expression analysis includes subjective and qualitative scenarios in the 2D image family[8, 3]. The 2D static images lack temporary dynamics[14, 4, 13, 15, 16, 18]. They are also prone to subjective judgements and poor qualitative features. The 2D static images do not capture temporary dynamics and expression changes [14, 4, 13, 15, 16, 18]. Over recent times there has been need to capture and measure quantitatively facial expressions in video and motion scenarios. Various frameworks have used dynamic analysis to deduce emotions giving accurate results.[14, 4, 13, 15, 16, 18]. Successful studies have chosen frameworks that included video face detection and tracking incorporating shape variability where features are extracted and classifiers applied on the given histograms[18, 1]. 2.2 Facial Recognition with video image sequences Facial expressions are subdivided into 3 segment groups namely the beginning, apex and end[18, 14, 15, 2, 1]. Facial expressions are also denoted as magnitudes of facial motions or motion units moments[18, 1]. Hidden Markov models and naive bayes classifiers have also been successful in video sequence facial expression recognition[14, 16]. Successful video sequence in facial expressions applied a two stage approach to classify images in 3D by measuring the video intensity using optical flows [1, 16]. Several probabilistic methods like particle filtering and condensation can also track facial expression in video sequences [18, 1]. Separate manifold substances have also been applied in video based facial expression analysis. To track video sequences models like 3d wireframe models, facial Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 29
mesh models, net models and ASM models were successfully used[14, 20, 1, 10, 12]. Videos subtle changes of facial expression can be measured on video facial expression recognition than on static image analysis[14, 16, 1, 10, 12]. 3 Local Based Facial Expression Feature Extraction Facial expression analysis influences wide areas in human computer interaction. Local binary patterns and their 2D and 3D variants have been used in this field[14, 16, 1, 10, 12]. Holistic and local based feature extractors have also been used successfully. PCA feature extractors are prominent holistic algorithms and local binary patterns, Gabor filters and Gabor wavelets and local directional patterns have been successfully applied as local feature extractors[14, 16, 1, 10, 12]. 3.1 Local Binary Patterns (LBP) for static image feature extraction Local binary patterns are based on facial images being split into local sub regions. The challenges of facial occlusion and rigidness are plenty though grey scale image conversion is used to reduce illumination[8, 7, 3]. Local binary patterns are invariant to grey level images. Localized feature vectors derived are then used to form the histogram which is used by machine learning classifiers or deep learning methods. The local features are position dependent [8, 7, 3]. For local binary patterns, the facial region is divided into small blockers like mouth, eyes, ears, nose and forehead[4]. The aggregate histograms are then grouped to form one feature vector or histogram for the facial image. The popular local binary pattern variants include uniform local binary patterns, central symmetric local binary patterns, elongated local binary patterns, multi block local binary patterns, ternary local binary patterns and rotational local binary patterns[8, 7, 3, 12, 23, 20, 1]. The key parameters include the radius of the local binary pattern and the given number of neighbors[8, 9, 3, 12, 23, 20, 1]. The basic local binary pattern non center pixels use the central pixel as the threshold value taking binary values [8, 7, 3]. Uniform binary patterns are characterized by a uniformity measure corresponding to the bitwise transition changes. The local binary pattern has 256 texture patterns. The local binary LBP r,n operator is represented mathematically in the following equation where the radius is given as r and the number of neighborhoods as n. LBP(n,r)= n=0 X n1 s(pnpc)2n.(1) The neighborhood is depicted as an m-bit binary string leading to n unique values for the local binary pattern code. The grey level is represented by 2nbin distinct codes. The gray scale of the middle pixel is given as pcwhilst pn represents the neighboring pixels. Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 30
LBP Variants Various LBP variants were successfully proposed and used. These include TLBP for Ternary Local Binary Pattern as well as Central Symmetric Local Binary Patterns [8, 9, 4]. Over-Complete Local Binary Patterns (OCLBP) is another key variant that takes into account overlapping into adjacent image blocks. The rotation invariant LBP is designed to remove the e↵ect of rotation by shifting the binary structure[8, 4]. Other variants include the monogenic and central symmetric (MCS-LBP). 3.2 Volume Local Directional Binary Patterns (VLDBP) For video sequencing facial image analysis, Volume Local Directional binary patterns (VLDBP) and Local Gabor Binary Patterns from Three Orthogonal Planes have been successful compared to other [1, 14, 15].. Volume local directional binary patterns (VLDBP) are used as an extension of LBP in the dynamic texture field. Dynamic texture extends the temporal domain and is used in video image analysis. The face regions of the video sequence images are modeled with VLDBP which incorporates movement and appearance [1, 14, 15]. It uses 3 planes where the central plan includes the central pixel used to the LBP algorithm. VLBP considers co-occurrences of surrounding points from three planes and generates binary representatives [1, 14, 15].. The extraction considers local volumetric neighborhoods against the pixels. The center pixel grey values and the surrounding pixels are then compared against each other. VLBP LP R = 3P+1 X q=0 vq2q(2) 3.3 Local Gabor Binary Patterns based on Three Orthogonal Planes 3D dynamic texture recognition which concatenates three histograms from LBP on three orthogonal planes has been widely used (LBP-TOP) [1, 14, 17]. LBPTOP extracts features from the local neighborhoods over the 3 planes. The spatial-temporal information in the 3 dimensional (X, Y, T) space is given where X,Y are spatial coordinates and the T axis represents temporal time[20]. [1, 14, 15]. LBP-TOP derives local binary patterns from a central pixel by threshholding the neighboring pixels [1, 14, 17][20]. The algorithm decomposes the 3 dimensional volume into 3 orthogonal planes[1, 2]. The XY plane indicates appearances features in the spatial domain and XT, visual against time whilst the YT plane is for motion in the temporal space domain[20]. The spatial plane, XY is similar to the regular LBP in static image analysis. The vertical spatio-temporal YT plane and horizontal XT plane are the other 2 planes in the 3 dimensional space[20]. The resulting descriptor enables encoding of spatio-temporal information in video images. The performance and accuracy of the latter was also comparable to the LBP-TOP. The LBP STCLQP Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 31
or spatio-temporal completed local quantized patterns (STCLQP) was also used to consider the pixel sign, orientation and size or magnitude[20]. Local gabor binary patterns from Three Orthogonal Planes (LGBP-TOP) add gabor filtering to improve accuracy. With the added filtering algorithm rotational misalignment of consecutive facial images is mitigated[1, 16, 10]. To avoid LBP-TOP statistical instability, a re-parameterization algorithm based on another local Gaussian jet was suggested [1, 14, 17]. k=(HL,X Y,H L,XT,H L,Y T,H C,X Y,HC,XT,HC,YT) (3) (LBP/C)TOP feature is denoted in vector form where Hv, m (v= LBP/C, and m = XY, XT, YT where m represents 6 LBP sub-histograms with contrast features in 3 planes [1, 14, 17]. The LBP-TOP algorithm describes video sequence changes in both spatial and temporal domains hence captures structural information of the former domain and longitudinal data of the latter [16, 18, 1]. LBP histogram features encode spatial data in the XY plane and the histograms from XT, YT planes include the temporal and spatial data[20]. With facial actions causing local and expression changes over time, the dynamic descriptors have an edge in facial expression analysis over the static descriptors [16, 18, 1]. Hq,x=Xq,v,tfj(q,v,t)=q(4) Contrasts in 3 orthogonal planes are denoted as Cm represented as (m= XY, XT and YT ) [1] and defined as 3 sub-histograms Hx ,y (x= C and y= XY, XT , YT ) [16, 18, 1]. Image device quality of the facial expression videos also impacts frame rates and spatial resolution quality as well [16, 18, 1]. Six Intersection Points (SIP) The LBP-SIP or Local Binary Pattern— Six Interception Points (LBP-SIP) considered 6 unique points along the intersecting lines of the 3 orthogonal planes to derive the binary pattern histograms [16, 18, 1]. AB, DF, EG =LA\LB\LC(5) where AB, DF and EG are intersection points. Six neighbor points carry enough data to describe spatio-temporal textures around point C[16, 18, 1]. LBP-SIP gives a concentrated group of high dimensional features spaces where there is sparse data [16, 18, 1]. LBP-Three Mean Orthogonal Planes (MOP) LBP-MOP or mean orthogonal plane is another variant to have been successfully used by concatenating mean images from image stacks derived along the 3 orthogonal planes[16, 18, 1]. It also preserves essential image patterns and reduces redundancy which a↵ects encoded features. Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 32
4 Feature Selection and Edge Detection Feature selection and images are enhanced for classification through edge detection and selection of the fittest images by genetic algorithms. Prominent detection algorithms include, Sobel, Canny, Kirsch edge detector which forms the base for local directional patterns and Hewitt edge detector. The study uses the Kirsch edge detector. 4.1 Local directional patterns LDP includes compass mask which allows for information extraction based on prominent edges or directions. The focus is on facial image edges on prominent facial regions[9, 3]. Convolution is applied based on the base images to get the edge detected images[9, 3]. For local directional patterns or LDP a key edge detection local feature extractor, the images were divided into LDPx histograms, retrieved and then combined into one descriptor[9, 3, 5, 1, 11]. The local directionary pattern, includes edge detection using the Kirsch algorithm. The operaFig. 1. Local Directional Patterns (LDP) tor takes one kernel which is rotated in 8 directions at forty five degrees based on one kernel marsk[9, 3]. The Kirsch operator’s edge size is calculated as the maximum size for all directions and this is shown in the local directional pattern Kirsch convolutionary equation with the associated example M0: LDPx()= r=0 X K r=0 X L f(LDPq(o, u),).(6) M0=(85x3)+(32x3)+(26x5)+(10x5)+(45x5)+(38x3)+(60x3)+(53x3) = 399 (7) 4.2 Genetic Algorithm Genetic algorithms were introduced in the 1970s as a class of evolutionary algorithms[22, 24]. These are heuristic approaches that find solutions based on evolutionary biology concepts. Genetic algorithms select a subset of features by removing unimportant features[22]. In this study the GA algorithm is also used to select best SVM kernel Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 33
function. The algorithm uses techniques like mutation, crossover, and selection to regenerate the population. The algorithm starts with a randomly generated set of facial expression images and this evolves as a new generation where the fittest images are selected for the next iteration[22]. The fitness function optimizes the objective function. The population selected once its satisfies the fitness, undergoes mutation or crossover for them to be selected for next iteration[24]. The convergence is reached when certain given generations has been achieved or when a given fitness value has been achieved. In facial expression recognition the algorithm is used to also optimize computational cost and video temporal correlation[22, 24]. Fig. 2. 1. Initialization of the parent population from the image database and choosing population size. 2. Evaluating the population 3. Selecting fit facial images and use fitness function to find the criterion. [24] 4. Crossover of facial images. The individuals which have frame numbers nearer to each other have higher probability of crossover[24]. 5. The mutations are done and the fittest individual is computed. In this computation of fittest, the older generation does not participate. Though it does die outs only after reproducing 10 fit individuals[24]. Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 34
6. This process continues until there is at least 10 fix facial images within 100 frames. Genetic algorithms has proven to be a successful algorithm to derive optimal support vector machine kernels[24]. The kernel functions are crucial to support vector machines as they a↵ect classification. A Genetic classification algorithm approach was widely suggested to choose the best kernel and its parameters[24]. 5 Facial Expression Implementation Approach The implementation involves analyzing video streams to track facial feature points over time. The feature vectors are then calculated and emotions detected from the trained models. Training and classification of the models is done using the popular algorithms namely support vector machines and genetic algorithm to remove unwanted facial images and reduce the error rate. Recognition of the model and new expressions on new images is then done on the selected annotated databases which includes CK+ database and FEED database. The section describes the approach, databases selected and then classification algorithm chosen and implemented. The study’s objective was to recognize facial expressions from video sequences. The approach involved locating and tracking the faces and expressions during the video segmentation and sequential modeling phase. The video sequence detection involved landmark detection and tracking, which define the facial shapes[15]. Viola Jones OpenCV detection tools are used. The features were then extracted using various 3D video feature extraction variants of the LBP-TOP algorithm. Gabor filters [16, 18, 1, 12, 5] were then applied during preprocessing. Geometric features were normalized and they were immune from skin color and illumination changes. [16, 18, 1]. Data: Copy and preprocess video image datasets Result: Facial expression classification results for the image datasets while For each image I inside the CK+ and FEED database do 1. divide the database into training and test sets; 2. for each image inside the given datasets; 3. apply Viola Jones algorithm for extraction and preprocess the image using Principal Component Analysis; 3b Genetic algorithm and PCA algorithm are applied to remove unwanted images which have no bearing to the facial expression classification. 4. The study then finds features using LBP-TOP, LBP-XY,LBP-XT and LBP-YT algorithms; 5. extract the features using the LBP-MOP, LBP-SIP algorithm 5b Local directional patterns are also applied to remove unwanted edges from the images 6. apply Gabor Filters to get the LGBP-TOP and LGBP-MOP features 7. calculate the Euclidian distance matrix; 8. apply the classification on each with support vector machine and genetic algorithm;Use genetic algorithms to select the optimal support vector machine kernel. 9. the best classification results is then labelled the best algorithm; End For’ end Algorithm 1: Local Gabor Binary Patterns from Three Orthogonal Planes to analyze video sequences[21, 1] Signal & Image Processing: An International Journal (SIPIJ) Vol.10, No.1, February 2019 35