Full text
Expert Systems With Applications 206 (2022) 117809 Available online 13 June 2022 0957-4174/© 2022 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/bync-nd/4.0/). Contents lists available at ScienceDirect Expert Systems With Applications journal homepage: www.elsevier.com/locate/eswa Automation of cleaning and ensembles for outliers detection in questionnaire data Vojtěch Uher a,∗, Pavla Dráždilová a, Jan Platoš a, Petr Badura b aDepartment of Computer Science, Faculty of Electrical Engineering and Computer Science, VSB - Technical University of Ostrava, 17. listopadu 15/2172, Ostrava, 708 33, Czech Republic bDepartment of Recreation and Leisure Studies, Faculty of Physical Culture, Palacký University Olomouc, třída Míru 117, Olomouc, 771 11, Czech Republic ARTICLE INFO Dataset link:http://hbsc.cz/lockdown2020/ Keywords: Anomaly detection Outliers Questionnaire data Data cleaning HBSC ABSTRACT This article is focused on the automatic detection of the corrupted or inappropriate responses in questionnaire data using unsupervised outliers detection. The questionnaire surveys are often used in psychology research to collect self-report data and their preprocessing takes a lot of manual effort. Unlike with numerical data where the distance-based outliers prevail, the records in questionnaires have to be assessed from various perspectives that do not relate so much. We identify the most frequent types of errors in questionnaires. For each of them, we suggest different outliers detection methods ranking the records with the usage of normalized scores. Considering the similarity between pairs of outlier scores (some are highly uncorrelated), we propose an ensemble based on the union of outliers detected by different methods. Our outlier detection framework consists of some well-known algorithms but we also propose novel approaches addressing the typical issues of questionnaires. The selected methods are based on distance, entropy, and probability. The experimental section describes the process of assembling the methods and selecting their parameters for the final model detecting significant outliers in the real-world HBSC dataset. 1. Introduction The anomaly (or outlier) detection reveals observations that seem to be inconsistent with the rest of the data. Hawkins (1980): “An outlier is an observation which deviates so much from the other observations as to arouse suspicious that it was generated by a different mechanism”. An outlier may indicate corrupted data such as manual error, coding error, low-quality measurements, system failure, etc. On the other hand, suspicious data can also represent some unknown or infrequent type of data that a dataset does not capture correctly. Outliers need to be detected, assessed, excluded from the analysis, or fixed if possible. They can negatively affect data distribution and weaken the reliability and credibility of the statistical analysis and its results (Shao, Zheng, Gu, Hu, & Qin,2022;Wilcox,2019;Yuan & Gomer,2021). The objective of this article is to assemble and test a robust framework using an ensemble of unsupervised methods for outliers detection in raw questionnaire data. The data usually comes from questionnaire surveys (Saris & Gallhofer,2014) that are typically used in social and behavioral research, psychology, health, etc. They cover demographic information, personal characteristics, or opinions of people. Unlike the machine-acquired numerical data used in the industrial and security areas, the questionnaire data has some specific issues coming from ∗Corresponding author. E-mail addresses: [email protected] (V. Uher), [email protected] (P. Dráždilová), [email protected] (J. Platoš), [email protected] (P. Badura). the human factors and the whole methodology designing the questionnaires to find out the required information. In questionnaire data, an observation represents the answers of one person (respondent) in a survey. It typically suffers from the following issues: many incomplete questionnaires, empty and intentionally wrong answers (e.g., repetitive patterns, self-contradictory responses), fast responses, questions are discrete categorical un/ordered items with a various number of options, normality or independence of items cannot be generally secured. We specifically examine the data from the Health Behavior in School-aged Children (HBSC) 2020 study (Inchley, Currie, Cosma, & Samdal,2018; Ng, Cosma, Svacina, Boniel-Nissim, & Badura,2021) that focuses on health and health behaviors of 11-, 13-, and 15-year-olds in the context of their social environments. We do not have any training set or any prior information about outliers. Most of the existing approaches solve outlier detection in numerical data, using some proximity-based methods (Agrawal & Agrawal,2015; Chandola, Banerjee, & Kumar,2009;Zimek, Schubert, & Kriegel,2012) (e.g. 𝑘-nearest neighbors, local outlier factor). The results are usually validated using supervised learning and labeled data. We aim at unsupervised statistical and machine learning methods for multivariate data with a special focus on the questionnaire-related issues. The goal https://doi.org/10.1016/j.eswa.2022.117809 Received 2 March 2022; Received in revised form 23 May 2022; Accepted 6 June 2022
Expert Systems With Applications 206 (2022) 117809 2 V. Uher et al. is to maximally automate the data cleaning, validation, and outliers detection. These activities are still often done manually (Waure, Poscia, Virdis, Pietro, & Ricciardi,2015) and they take a lot of time. The main contributions of this article include: 1. A crucial automatic data cleaning and preprocessing procedure. 2. Selected outlier detection methods optimized for questionnaire data that compute an outlier score reducing the data space onto one final variable. (a) Widespread methods (Mahalanobis distance, 𝑘-nearest neighbors, local outlier factor). (b) Innovated methods (the entropy-based scores dealing with empty answers and repetitive patterns, the correlation-based methods examining the significant dependencies between variables and their corruption, the probability score dealing with rare and improbable answers). 3. A selection of the proper statistical method determining the range of scores representing the finally removed outliers. The scores are assessed according to the distribution of the outlier score variable. 4. An ensemble outlier detection procedure that unites the outliers detected by selected methods assessing the outliers from different perspectives. 5. The case study computed on the HBSC 2020 data evaluating the proposed pipeline and the impact of outliers to data variance and covariance of attributes (Cronbach’s alpha (Cronbach,1951)). The article is organized as follows. Section 2summarizes the stateof-the-art literature with a special focus on the categorical and questionnaire data. Section 3defines our proposed pipeline for preprocessing questionnaire data and ensemble outlier detection. In Section 4, the types of outliers are listed, and the applied outliers detection methods are described in detail. Section 5tests the defined methods with the HBSC 2020 data. The experiments prove that our framework efficiently addresses the named problems and significantly improves the distribution and the statistical properties of the HBSC 2020 data. 2. Related work The questionnaires suffer from many errors and difficulties that must be solved before any further analysis of the collected data is performed. The procedure addressing these issues consists of data cleaning, preprocessing, and outlier detection. The cleaning and preprocessing need to be conducted before the outlier detection methods, and it has been described in the literature many times (García, Luengo, & Herrera,2015;Van den Broeck, Argeseanu Cunningham, Eeckels, & Herbst,2005;Zhu, Hernandez, Mueller, Dong, & Forman,2013). We focus on this topic in Section 3. The outlier detection algorithms can be basically divided into two categories: supervised and unsupervised. The supervised algorithms are based on training a classifier using the labeled training dataset and they include e.g. Support Vector Machine (Utkin,2014), Artificial Neural Networks (Kieu, Yang, & Jensen,2018;Naseer et al.,2018), and Decision Tree (Ramachandran & Kishorebabu,2019). As it is very difficult to obtain labeled data, the unsupervised methods that do not require labeled data are often used to compute a degree of outlierness. They include e.g. statistical (Hubert & Vandervieren,2008; Tukey et al.,1977), proximity-based (Breunig, Kriegel, Ng, & Sander, 2000;Leys, Klein, Dominicy, & Ley,2018;Shao et al.,2022), and clustering-based methods (Jiang, Liu, Du, & Sui,2016;Wang, Wang, & Wilkes,2012). We focus on the unsupervised methods applicable to multivariate questionnaire data. Most of the named approaches are designed for numerical data. The questionnaire data usually contains categorical or ordinal discrete variables with a small number of categories. Categorical data is also common in areas such as network intrusion detection (Sari et al., 2015), social networks (Aggarwal, Zhao, & Philip,2011), industrial processes (Zhu, Ge, Song, & Gao,2018), sensor faults in sensor networks (Zhang, Meratnia, & Havinga,2010), credit card fraud (Malini & Pushpa,2017), etc. The outlier detection methods are based on different approaches of data processing (Akoglu, Tong, Vreeken, & Faloutsos,2012;Ienco, Pensa, & Meo,2016). A survey of various outlier detection methods for categorical data is presented by Taha and Hadi (2019). While the general categorical data has been widely investigated in the literature, articles about outliers detection in questionnaire data are rare (Sakurai et al.,2019;Zijlstra, van der Ark, & Sijtsma,2011). The selected methods applied to categorical data can be divided into several categories: •Probability-based: (Zijlstra et al.,2011;Zijlstra, Van Der Ark, & Sijtsma,2007) proposed a method detecting outliers in questionnaire data based on the frequency of answers. The answers are sorted by frequencies so that the infrequent answers have a higher index/rank. The final outlier score of one questionnaire is a sum of the ranks of the contained answers. The greater the score is the more improbable and suspicious answers are contained. This approach is very simple but it can be highly biased if a specific answer predominates. •Proximity-based: This category includes methods based on distance (e.g. 𝑘-nearest neighbors Chandola, Banerjee, & Kumar, 2007) or density (e.g. local outlier factor Breunig et al.,2000). The main problem of these methods is to determine the size of neighborhood or estimate the number of samples expected within the neighborhood depending on the distribution of data (Uher, Gajdoš, & Snášel,2018;Uher, Gajdoš, Snášel, Lai, & Radeck` y, 2019). Another issue can be the usability of the Euclidean distance in high dimensions. A common-neighbor-based distance function was developed by Li, Lee, and Lang (2007) to measure the proximity of a pair of data points in distance-based outlier detection method for highdimensional categorical data. A weighted density that takes into account the density and uncertainty of each categorical variable is defined by Zhao, Liang, and Cao (2014). Multivariate outliers are often detected by variants of the Mahalanobis distance which represents the distance from the distribution of data (Ben-Gal,2005;Cabana, Lillo, & Laniado,2019;Leys et al.,2018) and it is non-parametric. The principal component analysis (PCA) is used for outliers detection or outliers visualization (Har-Shemesh, Quax, Lansing, & Sloot,2020;Sakurai et al.,2019;Zhu et al.,2018). PCA is also one of the methods used by Jebreel et al. (2020) to sanitize survey data that relies on combining the classification outcomes of unsupervised machine learning algorithms aimed at detecting wrong answers. Deng and Wang (2018) proposed a modified kernel PCA (KPCA) method with local outlier factor (LOF) to construct multivariate statistical process monitoring methods. Several variants of LOF (Breunig et al.,2000) have been proposed to handle different data types. A variant of LOF combined with information entropy is applied by Xie, Li, Wu, and Zhang (2016) for detecting outliers in medical insurance data. Yu, Qian, Lu, and Zhou (2006) used a mutual-reinforcement-based local outlier detection 𝜅-LOF to handle categorical attributes. The variants of 𝑘-nearest neighbors are also often used for outliers detection (Chandola et al.,2007; Chen, Miao, & Zhang,2010). The outlier score can be computed as a sum of distances between the query object 𝑞and its 𝑘nearest neighbors. Some other types usually require parameters 𝑘and distance 𝜆to assess if neighborhood with radius 𝜆contains at least 𝑘objects, or the 𝜆is used as a distance threshold (Chandola et al., 2007;Eskin, Arnold, Prerau, Portnoy, & Stolfo,2002;Knorr & Ng, 1998).
Expert Systems With Applications 206 (2022) 117809 3 V. Uher et al. Fig. 1. Flow of data processing. •Entropy-based: Algorithms based on the entropy were investigated in the area of outlier mining for categorical data (He, Deng, Xu, & Huang,2006;Pacheco, Ali, & Trappenberg,2019; Yuan, Zhang, & Feng,2018). The Shannon entropy is often used to search for a set of observations that minimize the Shannon entropy of the observations remaining after the outliers elimination (He et al.,2006). An extension of Shannon information entropy in rough sets is the approximation accuracy entropy (AAE) introduced by Jiang, Zhao, Du, Xue, and Peng (2019). This approach addresses the limits of the proximity-based methods. Finally, the evaluation of the detected outliers is an important task. But it is difficult to determine the extent to which newly proposed methods are improving compared to established methods in unlabeled data. Campos et al. (2016) conduct an extensive experimental study on the performance of a representative set of methods for unsupervised outlier detection, across a wide variety of datasets prepared for this purpose. As the methods differ and have various distributions of scores, each of them returns slightly different outliers. The ensemble learning generally combines multiple weak classifiers to obtain better overall performance but it mostly aims at supervised methods on labeled data. Several papers (Kriegel, Kroger, Schubert, & Zimek,2011;Schubert, Wojdanowski, Zimek, & Kriegel,2012) analyze the rank similarities, correlation of scores, regularization and normalization of scores to make them comparable, and weights estimation for ensemble design. The class labels are helpful for performance evaluation and optimal parameters selection. Most of the named papers rely on the algorithms based on 𝑘NN and LOF. There are also unsupervised approaches that compute ensemble outlier score to detect the outliers from multiple perspectives without any knowledge of data. They include e.g. average scoring, maximum scoring, or threshold sum (Aggarwal & Sathe,2015; Kandanaarachchi,2021). The main weakness of these ensembles is that they usually lead to average results that blur the outliers specific for standalone methods. The parameters optimization is insecure because the outlier labels are unknown. We overcome this problem with the usage of separate score distribution analysis and unification of identified outliers. For detailed survey of outliers or anomaly detection methods in various application fields, we refer to articles (Agrawal & Agrawal, 2015;Chandola et al.,2007,2009;Das, Schneider, & Neill,2008; Hodge & Austin,2004;Wang, Bah, & Hammad,2019). 3. Proposed procedure The specific issues related to questionnaire data and the design of the final ensembles are discussed here. We present our proposed procedure which starts with raw data, then continues with the specific steps of cleaning and preprocessing of data, sets the requirements for outliers detection, and defines the ensemble method combining the results of multiple algorithms. The pipeline is illustrated in Fig. 1. 3.1. Cleaning and preprocessing of questionnaires A dataset generally forms a table where columns represent the asked questions (attributes) and rows represent the answers of respondents (individual questionnaires). Our algorithms aim only at the multiple-choice questions (predetermined list of options) with a single answer including rating scales, Likert scales, or matrix questions (batteries) (Harpe,2015). They can be seen as categorical variables with a small number of options (mostly up to 10). Other types of questions have to be assessed separately and are omitted before the outliers detection methods are applied. We do not set any other a priori conditions for data. The preserved data has to be validated and cleaned. There are several widely-used approaches for cleaning the questionnaire data (García et al.,2015;Van den Broeck et al.,2005;Zhu et al.,2013) that generally lead to three common issues: Unrealistic answers (e.g. wrong type, value out of range), Incomplete questionnaires (many missing values), Response time (very short times indicate untrustworthy answers). The unrealistic answers can be detected with the utilization of the questionnaire codebook defining unambiguous attribute details. The exclusion of incomplete questionnaires or the fastest respondents can be simply done automatically by setting reasonable thresholds. Our tested preprocessing procedure goes through the following steps: 1. Selection of questions with a predetermined list of options - the omission of irrelevant variables such as strings, dates, personal information (birth date, weight, height, etc.), system and browser information (response times, device information, software version, etc.) 2. ‘‘Other’’ answers — any open-ended responses in questions with predetermined options are said to be missing (empty). 3. Elimination of questions with bad distribution — the omission of dichotomous variables and strongly unilateral variables with more than 80% of answers choosing the same option. Such questions bring a strong bias and undesirably strong correlation between attributes. 4. Missing values — all values representing empty answers are replaced by zeros. 5. Renumbering of option indices — makes a continuous sequence indexing the options of a question in the original order. The indices start with zero for missing values and continue with other natural numbers. This unifies the indexing and annuls prioritization of any option. 6. Normalization — as each question may have a different number of options, their indices recorded as answers in questionnaire data are normalized by variable onto the interval ⟨0,1⟩. The zero still represents a missing answer. 7. Fast response times — we eliminate all the questionnaires with almost zero response time (if available). Some small threshold has to be chosen. 8. Questionnaires with many missing values — we eliminate all the questionnaires with more than 70% of missing values. This step usually also eliminates the records with zero time as the rapid respondents do not fill anything. The restrictions listed above do not mean that those variables, values, or questionnaires are completely wrong. They can carry some useful information but they are undesirable for automatic outlier detection methods that are used in this paper. Moreover, the questions with a predetermined list of options usually form a major part of questionnaires that takes most of the manual effort. The incomplete questionnaires often form a huge part of data and their elimination is crucial to get any reasonable perspective on outliers. We generally work with two datasets that are used for comparison: •Unfiltered data — the original dataset after the whole preprocessing except to the points (7. and 8.) •Filtered data — it is the unfiltered data after the elimination of questionnaires with zero times (7.) and more than 70% of missing values (8.)
Expert Systems With Applications 206 (2022) 117809 4 V. Uher et al. The preprocessed data can be represented by a matrix 𝐀=[𝑎𝑖𝑗 ]∈ R𝑛×𝑚=[𝐚1,…,𝐚𝑚]of questionnaires, where 𝑚is the number of questions, 𝑛is the number of respondents and 𝐚𝑗∈R𝑛×1 is a column of matrix 𝐀(single variable). For some outliers detection methods, the matrix 𝐀needs to be normalized by columns to the range ⟨0,1⟩. Let 𝑚𝑖𝑛(𝐚𝑗)and 𝑚𝑎𝑥(𝐚𝑗)be the minimum and maximum value of a column 𝐚𝑗, the matrix normalized by columns 𝐀=[𝑎𝑖𝑗 ]∈R𝑛×𝑚is computed as 𝐀=[𝐚1−𝑚𝑖𝑛(𝐚1) 𝑚𝑎𝑥(𝐚1)−𝑚𝑖𝑛(𝐚1),…,𝐚𝑚−𝑚𝑖𝑛(𝐚𝑚) 𝑚𝑎𝑥(𝐚𝑚)−𝑚𝑖𝑛(𝐚𝑚)].(1) 3.2. Unsupervised outlier scores Once the questionnaires are preprocessed, the outlier detection algorithms are applied. We identified several common issues that should be detected in questionnaire data and can increase suspicion of outlierness. The issues include predictable patterns of responses, self-contradictory responses, or inconsistency with data distribution. As these defects cannot be revealed by a single approach, multiple outlier detection algorithms are performed, namely, time score, the 𝑘th order empirical entropy, correlation-based methods, probability score, Mahalanobis distance, local outlier factor (LOF), and the 𝑘-nearest neighbors (𝑘NN). Each method assigns an outlier score 𝜔𝑖∈⟨0,1⟩to each 𝑖th questionnaire of the data matrix 𝐀such that 𝜔𝑖= 1 means the maximal outlier score and 𝜔𝑖= 0 means a maximally standard questionnaire. Despite the dimension of the problem (number of questions), each outlier score is represented by one variable. If an outlier score is within a different range of values it has to be normalized. All the methods are unsupervised and work without any a priori knowledge. They can process questionnaires with an arbitrary number of questions and a various number of available options per question. The tested methods are described in detail in Section 4. 3.3. Ensemble outliers detection Given the outlier scores computed by selected methods, the final set of outliers has to be detected by combining all of them. The standard approaches define an ensemble based on a collective outlier score such as average scoring, or maximum scoring (Aggarwal & Sathe,2015) which are also tested in the experimental section. These ensembles mostly expect a positive correlation between outlier scores which is typical for methods based on the distance and the nearest neighbors. In our case, the methods are based on different principles and scores do not always correlate positively. The distributions of scores are highly unbalanced and they have positive skewness. Despite the unified scores, it happens that a questionnaire with a very high entropy score has a very low correlation score. The entropy assesses the predictable patterns while correlation-based methods assess if the responses are consistent with detected dependence between questions. This problem will be discussed in detail in Section 5. Using any sum of scores leads to a very average distribution of ensemble score where it is difficult to detect reasonable outliers. Therefore, we avoid computing an ensemble score and we propose a method based on the union of outliers detected by individual algorithms. The crucial task is to define a procedure that decides which questionnaires will be selected as outliers. Traditional algorithms usually choose 𝐾-worst observations or those having scores over some threshold (Aggarwal & Sathe,2015;Kriegel et al.,2011;Schubert et al., 2012). Both methods are tricky in unsupervised outlier detection as the parameters are user-defined (estimated) and they do not represent the score distribution well. A standard statistical method for outliers identification is the box-plot method (Grubbs,1969;Tukey et al., 1977). It uses the interquartile range (𝐼𝑄𝑅) which is the difference between the 75th percentile (𝑄3) and the 25th percentile (𝑄1) of the outlier score. The correct data are expected between the lower and upper whiskers of the box-plot which are computed as 𝑙𝑜𝑤 =𝑄1− 1.5⋅𝐼𝑄𝑅 and 𝑢𝑝 =𝑄3+ 1.5⋅𝐼𝑄𝑅. Everything outside this range is said to be an outlier. The box-plot is applicable for univariate data with approximately normal distribution. The outlier score computed by any of our methods is a single variable; however, its distribution is hardly normal. Our experiments showed that practically all the scores tested in this article have a distribution with positive skewness. Therefore, the adjusted box-plot (ABP) method (Hubert & Vandervieren,2008) is preferred here which is the box-plot method modified for skewed distributions. It was optimized to simulate distributions such as 𝛤, 𝜒2,𝐹, Parento, or Lognormal distributions by properly parameterized exponential function based on skewness. The adjusted box-plot rule and its lower and upper whiskers ⟨𝑙𝑜𝑤, 𝑢𝑝⟩are defined as follows: 𝐢𝐟 𝑀𝐶 ≥0𝐭𝐡𝐞𝐧 ⟨𝑄1− 1.5⋅𝐼𝑄𝑅 ⋅𝑒−4𝑀𝐶 , 𝑄3+ 1.5⋅𝐼𝑄𝑅 ⋅𝑒3𝑀𝐶 ⟩ 𝐢𝐟 𝑀𝐶 < 0𝐭𝐡𝐞𝐧 ⟨𝑄1− 1.5⋅𝐼𝑄𝑅 ⋅𝑒−3𝑀𝐶 , 𝑄3+ 1.5⋅𝐼𝑄𝑅 ⋅𝑒4𝑀𝐶 ⟩ 𝐰𝐡𝐞𝐫𝐞 𝑀𝐶 =𝑚𝑒𝑑𝑥𝑖≤𝑄2≤𝑥𝑗ℎ(𝑥𝑖, 𝑥𝑗)𝐚𝐧𝐝 ℎ(𝑥𝑖, 𝑥𝑗) = (𝑥𝑗−𝑄2)−(𝑄2−𝑥𝑖) 𝑥𝑗−𝑥𝑖 . (2) The 𝑀𝐶 represents the medcouple of the distribution which is used for skewness computation. The non-negative 𝑀𝐶 means the right positive skewness and the negative one means the left negative skewness. The 𝑄2represents a median of the outlier score. The right skewness means that the distribution has the right tail longer with most of the data concentrated to the left. As the right tail represents the highest outlier scores of the most suspicious entries, we apply only the upper limit 𝑢𝑝 to cut off the worst outliers. The lower limit 𝑙𝑜𝑤 is not applied because the low outlier scores represent the standard entries. Therefore, the adjusted box-plot rule allows us to identify the outliers much more precisely in the right tail than the standard box-plot rule. This rule is applied to all outlier scores computed by selected outlier detection methods. The only exception is the LOF where we use its own methodology based on the outlier factor. It means that each method identifies its own set of outliers. To incorporate multiple methods into a robust ensemble model, a union of individual sets of outliers is computed: 𝑂= 𝑙 ⋃ 𝑚=1 {𝑖∈ {1,…, 𝑛},|𝜔𝑚,𝑖 > 𝑢𝑝𝑚},(3) where 𝑂is a set of outliers indices, 𝑙is the number of computed outlier scores (methods), 𝑛is the number of questionnaires, 𝜔𝑚,𝑖 is the score of the 𝑚th method and the 𝑖th questionnaire and 𝑢𝑝𝑚is the upper limit of the 𝑚th method for outlier detection identified by the adjusted box-plot rule. 4. Outliers detection methods There are several common perspectives on how to define an outlier in questionnaire data: (1) Predictable patterns: Some cheating respondents often fill the questionnaires using the same answer choice over and over again or they use some repetitive pattern of answers, such as ‘‘a,b,c,d’’. These answers are also worthless for further exploration. (2) Self-contradictory responses: Some questions can be related or logically dependent. Questionnaire designers even place several similar questions into questionnaire forms to check if a respondent understands the formulation of questions. Selfcontradictory answers reveal a lack of understanding or cheating. (3) Other inconsistent responses: Although the data is filtered using all the previous approaches, some inconsistent respondents can persist. These questionnaires have to be assessed by a data analyst. They can be wrong from the perspective of the dataset distribution, or some can represent unusual but accurate replies. Nevertheless, the detection of outlying observations is generally desired.
Expert Systems With Applications 206 (2022) 117809 5 V. Uher et al. (4) Response time: Although the fastest respondents are excluded at the beginning, the response time can be also used as an outlier score. (5) Graphic visualization: A graphic interpretation of numerical or categorical attributes can reveal some outlying observations that do not fit the data distribution. This is good for visual verification of the data cleaning process. The methods were selected and designed to be unsupervised, i.e. they are functional without any a priori knowledge of data. Some of them are well-known methods and we modified some others to better adapt them for questionnaire data. A method assigns an outlier score 𝜔𝑖∈⟨0,1⟩to each 𝑖th questionnaire of the data matrix 𝐀such that 𝜔𝑖= 1 means the maximal outlier score and 𝜔𝑖= 0 means a maximally standard questionnaire. If a method naturally returns a different range of values, the scores have to be normalized. The following subsections define the methods in detail. They aim to different aspects of data to provide various perspectives on outliers. We combine methods later in Section 5to achieve a robust model for searching outliers in questionnaire data. 4.1. Proposed methods addressing issues in questionnaires This section describes our unsupervised outlier detection methods for questionnaire data based on entropy, correlation of attributes, and probability of answers. 4.1.1. The 𝑘th order empirical entropy Roughly speaking, entropy is a measure of uncertainty and it is often used in text compression algorithms. It shows whether the text contains some repetitive patterns of symbols and this states the lower limit of the compressed representation of data. We propose to compute the entropy to find outliers within a set of questionnaires. As the questionnaire data mostly consists of categorical data, each questionnaire can be transformed into a sequence of answers. It is supposed that cheating respondents try to fill the questionnaires with some predictable patterns such as ‘‘a, a, a’’, ‘‘b, b, b’’, ‘‘a, b, c’’ etc. These patterns lead to very low entropy of the sequence, and thus, it is utilized to detect the suspicious answers for outliers detection. It also detects the questionnaires with many empty answers. The empirical entropy (Beirlant, Dudewicz, Györfi, & Van der Meulen,1997;Shannon,1948) was defined for a discrete variable 𝑋with possible outcomes 𝑥1,…, 𝑥𝑙representing the symbols with occurrence probabilities 𝑃(𝑥1),…, 𝑃 (𝑥𝑙). This is also called the zerothorder empirical entropy and it is based only on the distribution of symbols in a finite string 𝑠. The 𝑘th order empirical entropy (Manzini, 1999) extends this principle to higher orders. The 𝑘th order empirical entropy 𝐻𝑘(𝑠)is a conditional entropy, where the random variable represents all strings of length 𝑘and it is computed as 𝐻𝑘(𝑠)=− ∑ 𝑤∈𝛴𝑘∑ 𝑥∈𝛴 𝑃(𝑤𝑥) log 𝑃(𝑥|𝑤)=− ∑ 𝑤∈𝛴𝑘∑ 𝑥∈𝛴 𝑛𝑤𝑥 |𝑠|log 𝑛𝑤𝑥 𝑛𝑤 ,(4) where 𝑤is a word consisting of 𝑘symbols, 𝑃(𝑤𝑥)is occurrence probability of a word 𝑤𝑥 and 𝑃(𝑥|𝑤)is conditional probability that a symbol 𝑥follows a word 𝑤in a string 𝑠. The right side of Eq. (4) represents the calculation of entropy using frequencies 𝑛𝑤𝑥 of the word 𝑤𝑥 and 𝑛𝑤of the word 𝑤in the string 𝑠having length |𝑠|. In the questionnaire data, the string 𝑠𝑖is a sequence 𝑠𝑖= (𝑎𝑖1,…, 𝑎𝑖𝑚) of numbers representing the answers to questions of the 𝑖th respondent in the matrix 𝐴. Each question represented by answer 𝑎𝑖𝑗 ∈ {0,1,…, 𝑙𝑗} for 𝑗= 1,…, 𝑚 allows only a small finite set of 𝑙𝑗numerical values extended by 0 for empty answers, where it is possible that 𝑙𝑗≠𝑙𝑦for 𝑗≠𝑦and 𝑗, 𝑦 ∈⟨0, 𝑚⟩. The implementation of 𝐻𝑘(𝑠)is based on the dictionary of unique subsequences of 𝑘answers 𝑠𝑝= (𝑎𝑝,…, 𝑎𝑝+𝑘−1)for 𝑝= 1,…, 𝑚 −𝑘+ 1 and their frequencies in the sequence 𝑠𝑖. A subsequence 𝑠𝑝represents a word 𝑤in (4) with frequency 𝑛𝑤. The same approach can be used to find subsequences of length 𝑘+ 1 for a word 𝑤𝑥 and frequency 𝑛𝑤𝑥. In this article, 𝑘th order empirical entropy is used to compute the 𝑘th order empirical entropy of answers for each questionnaire. The lower the entropy is the more predictable and suspicious answers are contained. The questionnaires with low entropy should be excluded from the dataset or detected as outliers, therefore the outlier score is 𝜔𝑖= 1−𝐻𝑘(𝑠𝑖). This is a row-based method which means that the outlier score 𝜔𝑖is computed using just answers of one respondent. Partial entropy of batteries of questions A typical type of question is a battery of questions that contains a block of questions having the same number of answers. It usually consists of related questions or measures some quality (e.g. 1 — best, 5 — worst), etc. As the battery questions are consolidated together, they usually bring respondents to cheat more often than other types of questions. Therefore, we created another type of entropy-based outliers detection which uses only the batteries. As a dataset consists of many answered questionnaires, it is possible to detect the specific set of options for each question. Then the consolidated ranges of questions with the same possible answers are supposed to belong to the same battery of questions (part). Given a sequence of answers 𝑠𝑖= (𝑎𝑖1,…, 𝑎𝑖𝑚)and a set of all the 𝑡battery intervals 𝐵= {𝑏1,…, 𝑏𝑡}, where 𝑏𝑗= (𝑙𝑗, 𝑟𝑗),𝑙𝑗< 𝑟𝑗and 𝑙𝑗, 𝑟𝑗∈⟨1, 𝑚⟩, then 𝑒𝑗=𝐻𝑘((𝑎𝑖𝑙𝑗,…, 𝑎𝑖𝑟𝑗)) is the 𝑘th order entropy (4) of the battery interval 𝑏𝑗. The entropies of 𝑡battery intervals (parts) form a vector 𝐞𝑖= (𝑒𝑖1,…, 𝑒𝑖𝑡), whose Euclidean norm ‖𝐞𝑖‖sets the final outlier score 𝜔𝑖= 1 − ‖𝐞𝑖‖of the given sequence 𝑠𝑖. This score is computed for all questionnaires and used to detect the outliers. The 𝑘should be set to a reasonably small number (𝑘= 1 or 𝑘= 2) because the size of batteries is relatively small. 4.1.2. Correlations Another family of methods we propose is based on the assumption that there are some dependencies between questions. Outliers are detected by answers that do not suit the dependencies. Given the normalized matrix 𝐀(1), its correlation matrix 𝐑= 𝑐𝑜𝑟𝑟( 𝐀)is computed, where the function 𝑐𝑜𝑟𝑟 computes the correlation matrix using the Spearman’s correlation coefficient (Sahoo,2015). Note that the missing answers are omitted for the correlation computation to avoid fake dependencies between questions with many missing answers. The vacancies are filled with zeros for further computing. The matrix 𝐑captures the linear dependencies between all attributes represented by answers to questions. The correlation matrix 𝐑is utilized to detect the significant correlations that reveal some related questions. It is supposed that a strong violation of correlations is suspicious and such questionnaires should be detected as outliers. Let an 𝐑=[𝑟𝑘𝑙]∈R𝑚×𝑚be a correlation matrix of 𝐀=[𝑎𝑖𝑗 ]∈R𝑛×𝑚. Let a 𝑐𝑜𝑟𝑟𝑀𝑖𝑛 be a threshold of minimal significant correlation and 𝑑𝑖𝑠𝑡(𝑎𝑖𝑘, 𝑎𝑖𝑙)be a distance for values of two attributes. A 𝜏is a number of the greatest distances that are used to calculate a score for each questionnaire. The basic algorithm can be written as follows: (1) Set 𝑐𝑜𝑟𝑟𝑀𝑖𝑛,𝜏and define function 𝑑𝑖𝑠𝑡 ∶R2→R. (2) Find all significant correlations: 𝑆= {(𝑘, 𝑙) ; 𝑎𝑏𝑠(𝑟𝑘𝑙)≥𝑐𝑜𝑟𝑟𝑀𝑖𝑛 for 𝑘<𝑙}. (3) For each questionnaire 𝑖from 𝐀compute a set of distances between correlated attributes: 𝐷𝑖= {𝑑; ∀(𝑘, 𝑙) ∈ 𝑆∶𝑑= 𝑑𝑖𝑠𝑡(𝑎𝑖𝑘, 𝑎𝑖𝑙)𝐢𝐟 𝑎𝑖𝑘 ≠0 ∧ 𝑎𝑖𝑙 ≠0𝐞𝐥𝐬𝐞 0}. (4) Sort each set 𝐷𝑖in descending order. (5) Compute the score 𝜔𝑖of a questionnaire 𝑖as average of the first 𝜏distances: 𝜔𝑖=∑𝜏 𝑗=1 𝐷𝑖𝑗 ∕𝜏. (6) Normalize the scores 𝜔𝑖for 𝑖= 1,…, 𝑛 to the range ⟨0,1⟩if necessary. The score 𝜔𝑖of the 𝑖th questionnaire is represented by its 𝜏worst violations of the assumed dependencies of attributes. The distance function 𝑑𝑖𝑠𝑡 can be defined by different types of metrics. The metrics
Expert Systems With Applications 206 (2022) 117809 6 V. Uher et al. can be expressed as a distance between a pair of values of the correlated attributes or as a distance between the values and the distribution of the correlated attributes. We tested two basic methods: linear regression (Sahoo,2015;Witten & Frank,2002;Zaki, Meira Jr, & Meira, 2014) and difference as they are defined in the following subsections. Linear regression As the correlation reveals a linear dependency between attributes, it is supposed that its course can be expressed by linear function estimated by linear regression (Sahoo,2015;Witten & Frank,2002). The metric is computed as an orthogonal distance between point and line. Given the two correlated columns 𝐚𝑘and 𝐚𝑙of the normalized matrix 𝐀, the linear regression computes the intercept 𝑏0and the slope 𝑏1parameters for the slope-intercept form of the linear equation 𝑦=𝑏0+𝑏1𝑥which can be expressed in the implicit form as 𝑏0+𝑏1𝑥−𝑦= 0. The distance from the line for the columns 𝑘and 𝑙is computed 𝑅𝑒𝑔𝑟𝑒𝑠𝑠(𝑎𝑖𝑘, 𝑎𝑖𝑙) = |𝑏0+𝑏1𝑎𝑖𝑘 −𝑎𝑖𝑙| √𝑏2 1+ 1 .(5) Difference The simplest distance can be computed by the difference of attributes values. Given the correlation matrix 𝐑of the normalized matrix 𝐀, the distance between the columns 𝑘and 𝑙for the 𝑖th row of 𝐀is computed as 𝐷𝑖𝑓 𝑓(𝑎𝑖𝑘, 𝑎𝑖𝑙) = {𝑎𝑏𝑠(1 − 𝑎𝑖𝑘 −𝑎𝑖𝑙) ∶ 𝑟𝑘𝑙 <0 𝑎𝑏𝑠(𝑎𝑖𝑘 −𝑎𝑖𝑙) ∶ 𝑟𝑘𝑙 ≥0(6) The difference in Eq. (6) reflects both a positive and negative correlation coefficient. It is a cheaper estimate of linear regression. 4.1.3. Probability score Another approach is a question-based outlier score presented by Zijlstra et al. (2007) which counts an individual’s frequency of unpopular answers. The point is that the frequent answers are not suspicious while many improbable answers per a questionnaire are more suspicious. This method is called 𝑂+and it judges the improbable answers by each question/attribute separately. In short, the probabilities of existing options within a question are computed so that the most probable option has the lowest score 𝑂𝑖𝑗 = 0 and the least probable one has the highest score 𝑂𝑖𝑗 =𝑜𝑗− 1 for the 𝑖th respondent and the 𝑗th question having 𝑜𝑗existing options. The final outlier score of one questionnaire is a sum of the ranks of the contained answers. The greater the score is, the more improbable answers are included. Formally, given the matrix of questionnaires 𝐀and a probability function 𝑃(𝑋𝑗)for 𝑗= 1,…, 𝑚 representing the relative frequency of each option of the 𝑗th attribute, then the outlier score 𝑂𝑖𝑗 is determined using the rank number of 𝑃(𝑋𝑗=𝑎𝑖𝑗 )denoted 𝑟𝑎𝑛𝑘 [𝑃(𝑋𝑗=𝑎𝑖𝑗 )]∈ {1,…, 𝑜𝑗}, such that 𝑂𝑖𝑗 =𝑜𝑗−𝑟𝑎𝑛𝑘 [𝑃(𝑋𝑗=𝑎𝑖𝑗 )],(7) where 𝑜𝑗is the number of options of the 𝑗th question. The total outlier score 𝑂𝑖+of the 𝑖th questionnaire is 𝑂𝑖+ = 𝑚 ∑ 𝑗=1 𝑂𝑖𝑗 .(8) Zijlstra et al. (2007) assumes that all the 𝑚questions have the same number of existing options. However, this condition is not met very often. Thus, we expect various 𝑜𝑗for different 𝑗which means that the ranking has to be normalized to get the ranks to the comparable level. Otherwise, the questions with more options would have a greater weight in the total score. The sum in (8) is reformulated as 𝑂𝑖+ = 𝑚 ∑ 𝑗=1 𝑂𝑖𝑗 ∕(𝑜𝑗− 1).(9) Although there is some bias due to different 𝑜𝑗, all the 𝑚questions are contained in every questionnaire of 𝐀, and therefore, the total outlier scores 𝜔𝑖=𝑂𝑖+should be comparable. Note that the empty answers should be skipped by 𝑂+method as they can easily become the most frequent in some questions. This would make the empty answers the most probable and the rest highly suspicious. 4.2. Widespread methods The selected general outliers detection methods based on response time and distance are briefly summarized here: •Time: The simplest method is to compute a total time (if available) to reveal how much attention a respondent paid to questions. A very short time probably indicates fake or incomplete answers. The normalized time is used as an outlier score. •Mahalanobis distance (Leys et al.,2018;Mahalanobis,1936) is a non-parametric method that measures the distance of a point to the center of the dataset using the covariance structure of the data in the multidimensional space. •Local outlier factor (LOF) (Breunig et al.,2000) is a densitybased method using a single parameter 𝑀𝑖𝑛𝑃 𝑡𝑠, which is the number of the nearest neighbors used to define the local neighborhood. Unlike global methods that define a fixed neighborhood for all points, the LOF is able to judge the outliers from the local perspective. The basic idea is that the density around a point is compared with the density around its neighbors. The assumption is that the density around an outlier is considerably different than the density around its neighbors. The outlier factor around 1 means a similar density as neighborhood, a factor smaller than 1 represents an inlier, and a factor greater than 1 is an outlier. Unlike in other methods, we do not provide the outlier score as a single random variable but we apply this LOF rule. The LOF is stable for 𝑀𝑖𝑛𝑃 𝑡𝑠 between 10 and several hundred (Breunig et al., 2000) depending on the data and the size of clusters. The 𝑀𝑖𝑛𝑃 𝑡𝑠 should not be greater than the size of the smallest distinguishable cluster. As the LOF uses a spherical neighborhood based on the Euclidean distance, it is inaccurate in high dimensions (curse of dimensionality Zaki et al.,2014). Moreover, the dimensionality increases the time complexity. Therefore, we apply the Principal component analysis (PCA) (Witten & Frank,2002;Zaki et al., 2014) to reduce the dimension before the LOF application. •𝑘-Nearest Neighbors (𝑘NN) (Chandola et al.,2007;Chen et al., 2010) The 𝑘NN simply computes the distances to the 𝑘closest objects and uses their mean directly as an outlier score. Unlike in LOF, the outliers are identified globally as the scores are compared across the whole dataset. 5. Case study: HBSC data This section evaluates our proposed procedure including questionnaire data cleaning, transformation, and outlier detection tested on a real-world dataset. We bring a case study showing the impacts of the model on the HBSC 2020 data. We got the data available in uncleaned form with all the issues named in this article. Most of the questionnaire survey data is available only in cleaned form, thus we cannot sufficiently test our methods on them. However, the model can be applied to any uncleaned questionnaire data where the outlying and faulty observations are still included. 5.1. HBSC data The data for the present article comes from a project directly linked to the Health Behavior in School-aged Children (HBSC) study. The HBSC is an international World Health Organization-collaborative questionnaire-based survey, which focuses on health and health behaviors of 11-, 13-, and 15-year-olds in the context of their social environments (family, peers, and school). The design of the HBSC study
Expert Systems With Applications 206 (2022) 117809 7 V. Uher et al. Fig. 2. Frequency of distinct answers for all 113 questions in HBSC 2020 data represented by stacked bar plots. The number of options per question varies between 3 and 16. enables not only to check cross-national comparisons (Badura et al., 2021) but also to monitor trends in the health and health behaviors of young people during their transition from childhood to adulthood, as a crucial period for long-term establishment of such behaviors. The data file used in the present article was drawn from a survey conducted in the Czech Republic in June 2020, i.e. during the final stage of the first wave of COVID-19 epidemic. The aim of this data collection was to assess the impact of the lockdown, especially in terms of school closure and ban of sport or other leisure-time activities, on the adolescents’ lives, time use and health behaviors (Ng et al.,2021). Overall, 141 schools from 14 administrative regions of the Czech Republic took part of the survey which contains 7082 unique entries gathered between 1st and 30th June 2020. 5.2. Preprocessing of HBSC 2020 The effect of cleaning and preprocessing (Section 3.1) is briefly introduced here to illustrate its importance. The original HBSC 2020 dataset contains 7082 respondents and 232 variables. The variables include the questions, personal data, system/browser information, time data, and some other auxiliary variables. After the application of the cleaning procedure, the dataset was reduced to 113 multiple choice questions with a single answer. The number of possible options varies between 3 and 16 per question. Fig. 2 shows the distribution of answers per each question. Fig. 2(a) represents the unfiltered dataset with all 7082 respondents and Fig. 2(b) represents the filtered data after the elimination of questionnaires with more than 70% of missing values (3255 respondents). The plots show that 54% of the unfiltered dataset is the questionnaires with a huge amount of missing values (>70%).
Expert Systems With Applications 206 (2022) 117809 8 V. Uher et al. Table 1 Tested methods and their parameters and types of outliers they reveal. Name Description Outlier types Parameters en𝑘 𝑘th order empirical entropy Predictable patterns, missing values 𝑘∈ {1,2,3,4} pen𝑘Partial en𝑘of batteries Predictable patterns, missing values 𝑘∈ {1,2} Maha Mahalanobis distance Inconsistent responses – 𝑂+Probability score Inconsistent responses – dc𝜏Correlation (difference) Self-contradictory responses 𝑐𝑜𝑟𝑟𝑀𝑖𝑛 = 0.7,𝜏∈ {3,6,12} rc𝜏Correlation (linear regression) Self-contradictory responses 𝑐𝑜𝑟𝑟𝑀𝑖𝑛 = 0.7,𝜏∈ {3,6,12} LOF𝑑Local outlier factor Inconsistent responses PCA dims 𝑑∈ {3,10},𝑀𝑖𝑛𝑃 𝑡𝑠 = 100 kNN100 𝑘-Nearest Neighbors Inconsistent responses 𝑘= 100 The missing values are represented by option 0. The difference between Figs. 2(a) and 2(b) is enormous and it would not be possible to read any reasonable information from the data without proper cleaning and outliers elimination. The questionnaire data is heavily disrupted by missing values, human errors, and false and inconsistent answers. 5.3. Model for outliers detection and its parameters This section describes the utilized outlier detection ensemble that is experimentally verified in the next sections on HBSC 2020 data. Both the unfiltered (7082 respondents) and filtered (3255 respondents) datasets with 113 variables are tested for comparison. The tested methods and their parameters are summarized in Table 1. It contains all the methods defined previously in Section 4with parameters that were selected experimentally for the HBSC 2020 dataset. Four different orders of entropy are tested because the entropy for any higher order is too low (score is too high) on 113 variables. Only two orders of entropy are tested with the partial entropy because the batteries of questions are even smaller (say about 10 questions). The data often contains many incomplete questionnaires that were not filled to the end by respondents in time. A large block of missing answers leads to very low entropy and high suspiciousness. However, even incomplete questionnaires with reasonable answers in the opening can be valuable according to our experience. Thus, we judge by the entropy only the initial part of a questionnaire that is filled. The missing answers at the end are simply cut off and those in the middle are preserved. The Mahalanobis distance and 𝑂+are nonparametric. In the case of correlation-based methods, we set 𝑐𝑜𝑟𝑟𝑀𝑖𝑛 = 0.7according to the correlation table which produces 15 significant correlations for both filtered and unfiltered data. The number of the worst distances between correlated variables 𝜏is set to 3, 6, and 12 to test which 𝜏is the most representative. The LOF method is based on a spherical neighborhood which limits its applicability to high-dimensional data. As the data has 113 variables we reduce them to 3 and 10 dimensions by the PCA. According to the discussion in Section 4.2, the 𝑀𝑖𝑛𝑃 𝑡𝑠 is set to 100 as the data has several thousands of respondents. Small 𝑀𝑖𝑛𝑃 𝑡𝑠 leads to many outliers detected from the perspective of dense small clusters. Similarly, the 𝑘NN method is set to 𝑘= 100 which corresponds to the LOF method. After several experiments for 𝑘∈ {50,80,100} we found out that there is no significant difference, and therefore, 𝑘= 100 is good for our experiments. 5.4. Experiments The model described in the previous section is applied to the HBSC 2020 data. This section contains our experiments and the discussion about the properties of outlier detection methods and the effect of data cleaning. Of course, not every method returns the best-expected outliers. Thus, the experiments guided us to select the appropriate combination of methods and parameters for our final ensemble model detecting the outliers in HBSC 2020 data. 5.4.1. Outlier scores distribution All the variants of methods and parameters described in Table 1 were computed. First, we analyze the distributions of outlier scores and similarities between different methods. The methods with bad distribution are excluded from the model. Fig. 3 shows standard box-plots of different outlier scores for filtered and unfiltered data. Generally, box-plots with most of the values placed in the smaller half are expected. As the scores assign the highest values to the most inconsistent questionnaires, we expect to find outliers above the upper whisker. Another comparison can be seen in Fig. 4 which illustrates how the scores correlate and how this correlation differs between filtered and unfiltered data. We use Spearman’s correlation coefficient which compares the ranks of questionnaires and is more robust in noisy data. The box-plots of unfiltered data show that there are certain groups of methods with similar logic there. The correlation-based methods correspond better to the distance-based methods and the probability score and they all go against the entropy-based methods and time. This shows that the methods really detect different types of outliers and they complement each other. The same effect can be seen in the correlation matrix in Fig. 4 where there are positive correlations within the named groups and negative correlations between methods from different groups. The box-plots show that the scores computed on unfiltered data have a terrible distribution. They are strongly affected by incomplete questionnaires which leads to extremely low or extremely high outlier scores. This is because the incomplete records represent a majority of data which deforms the statistical indicators so much that the standard data looks inconsistent. Detection of outliers based on unfiltered data is a random selection of questionnaires. The distribution of outlier scores computed on the filtered data looks reasonable and it suits the expectation that most of the questionnaires are rated by low scores while the outliers have high scores. The only methods that do not suit this expectation very well are en3,en4 and time. The correlation matrix of the filtered data also confirms that the en3 and en4 behave on the contrary to other entropies. It seems that a subword of length 3 or 4 is too long to get any reasonable entropy on 113 questions. Therefore, we exclude en3 and en4 from the model. In both experiments, the time factor is included, however; the differences between responding times were so huge that most of the questionnaires lead to very high scores considering the long times in contrast to the short ones. Therefore, we exclude only the questionnaires with zero times during the preprocessing which strongly correspond to empty questionnaires. The time factor is not used in our model for outliers detection. The method with the least variance of scores is the LOF because it judges the outliers from the local perspective based on the density of the current cluster. It has its own technique for outliers identification (Section 4.2). The 𝑘NN also explores the neighborhood of each object/questionnaire but its scores are judged globally in comparison with all the other questionnaires using the ABP rule. It can be seen that the difference between correlation-based methods with different 𝜏is minimal. Therefore, we use only dc12 and rc12 to detect outliers as they have smoother distribution than the scores for lower 𝜏.
Expert Systems With Applications 206 (2022) 117809 9 V. Uher et al. Fig. 3. Comparison between box-plots for outliers detection methods using the unfiltered (left) and filtered (right) HBSC 2020 data. Fig. 4. Comparison between Spearman correlations for methods of outliers detection on the unfiltered (left) and filtered (right) HBSC 2020 data. Fig. 5. Comparison among some scoring methods for outliers detection on the unfiltered (left) and filtered (right) HBSC 2020 data. Fig. 5 presents the distribution of some selected scores on unfiltered and filtered data. The charts are sorted according to the rc12 score. The figure shows again that the entropy goes against the other scores. The rc12 gives a very low score to incomplete questionnaires as the difference between correlated attributes is almost zero. Other distancebased methods behave similarly. The LOF10 is very low for most of the questionnaires in the unfiltered data as there are too many incomplete entries which strongly affect the local density factor. However, the entropy of empty questionnaires is very low, so that, their score is almost one. The 𝑂+method is closer to rc12 but it detects almost no outliers in unfiltered data (Fig. 3) because it is based on the frequency of answers. As the empty answers are very frequent, they strongly shift the 𝑂+distribution to low numbers. Even though we skip the empty answers, their values are missing in the final sum in contrast to the properly filled questionnaires. In conclusion, the set of outliers (denoted 𝑂𝐴𝑙𝑙) consists of the outliers detected by en1, en2, pen1, pen2, Maha, 𝑂+, dc12, rc12, LOF3, LOF10 and kNN100 methods. The ensemble is computed using the