Advances on permutation multivariate analysis of variance for big data
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Bonnini, Stefano; Assegie, Getnet Melak Article Advances on permutation multivariate analysis of variance for big data Statistics in Transition new series (SiTns) Provided in Cooperation with: Polish Statistical Association Suggested Citation: Bonnini, Stefano; Assegie, Getnet Melak (2022) : Advances on permutation multivariate analysis of variance for big data, Statistics in Transition new series (SiTns), ISSN 2450-0291, Sciendo, Warsaw, Vol. 23, Iss. 2, pp. 163-183, https://doi.org/10.2478/stattrans-2022-0022 This Version is available at: https://hdl.handle.net/10419/266313 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/
STATISTICS IN TRANSITION new series, June 2022 Vol. 23 No. 2, pp. 163–183, DOI 10.2478/stattrans-2022-0022 Received – 01.08.2021; accepted – 06.04.2022 Advances on Permutation Multivariate Analysis of Variance for big data Stefano Bonnini 1 , Getnet Melak Assegie 2 ABSTRACT In many applications of the multivariate analyses of variance, the classic parametric solutions for testing hypotheses of equality in population means or multisample and multivariate location problems might not be suitable for various reasons. Multivariate multisample location problems lack a comparative study of the power behaviour of the most important combined permutation tests as the number of variables diverges. In particular, it is useful to know under which conditions each of the different tests is preferable in terms of power, how the power of each test increases when the number of variables under the alternative hypothesis diverges, and the power behaviour of each test as the function of the proportion of true alternative hypotheses. The purpose of this paper is to fill the gap in the literature about combined permutation tests, in particular for big data with a large number of variables. A Monte Carlo simulation study was carried out to investigate the power behaviour of the tests, and the application to a real case study was performed to show the utility of the method. Key words: big data, MANOVA, permutation test, multivariate analysis. 1. Introduction In many applications of the multivariate analyses of variance (MANOVA), the classic parametric solutions for testing hypotheses of equality in population means or multisample and multivariate location problems might not be suitable for various reasons. For instance, the strong and implausible assumptions of iid observations and multivariate normality are the main reasons for considering parametric methods neither flexible nor robust and consequently often unsuitable. Moreover, in the presence of big data with a high number of response variables, great attention should be paid when the number of response variables is larger than the sample sizes, because of the loss of degrees of freedom. 1 Department of Economics and Management, University of Ferrara, Italy. E-mail: b[email protected]. ORCID: https://orcid.org/0000-0002-7972-3046. 2 University of Parma, Italy. E-mail: [email protected]. ORCID: https://orcid.org/0000-0001-7288-9636. © Stefano Bonnini, Getnet Melak Assegie. Article available under the CC BY-SA 4.0 licence
164 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… Even if there is not a unique definition, in statistics, a dataset is usually classified as “big data” if it represents a collection of informative data, extensive in terms of volume, velocity and variety, such that specific analytical technologies and methods are required for the extraction of value or knowledge (Baro et al., 2015). Big data are typical of many empirical disciplines such as biomedicine, economics, biology, ICT, education and research, financial services, social media, automotive industries, etc. (Özköse et al., 2015). Frequently, the high volume of big data depends on the multivariate nature of the dataset, due to the large number of variables. In addition, the variety of big data, due to the presence of different types of variables (quantitative and qualitative) and to the variability and heterogeneity of data, makes inferential problems more complex and requires robust and valid techniques to make inferences. For instance, in studies focused on social media, text, video, audio, and image data are jointly analysed. Hence, tests of hypotheses for big data must be addressed with appropriate methods that lead to reliable decisions, in short times and taking into account the variability and heterogeneity of the information. A typical approach to variable oriented multivariate problems consists in the application of exploratory methods based on the dimensionality reduction such as principal component analysis (PCA) or factor analysis (FA) (Johnson and Wichern, 2007; Farcomeni and Greco, 2016). For two-sample multivariate testing problems, in the presence of numeric data, a typical solution is the Hotelling T-square test. These methods are based on strong assumptions such as the linearity of the relationships between variables or normality. Linearity is a very strong and often unrealistic assumption. Normality is a reasonable assumption only with large sample sizes due to asymptotic properties of the statistics. Nevertheless, even in cases where linearity and normality are reasonable assumptions, especially in inferential problems, in the presence of many variables the estimation of a large number of unknown parameters, such as covariances or correlations, is required. Moreover, when the sample size is less than the number of variables, a problem related to the degrees of freedom arises and some typical parametric methods, such as the Hotelling T-square test, are not applicable. In such problems, nonparametric methods are preferable because they do not require that the underlying probability law belongs to a given family of distributions and no parameters need to be estimated. In particular, permutation tests follow a distribution-free approach and are almost as powerful as parametric methods based on normality when this assumption is true but much more powerful when the true underlying distribution deviates from the Gaussian (Pesarin, 2001; Anderson, 2001). Solutions for multivariate tests within the family of permutation methods consider the dependence between response variables without modelling it explicitly, and consequently without the need of estimating parameters or assuming linearity
STATISTICS IN TRANSITION new series, June 2022 165 (Pesarin and Salmaso, 2010a; Bonnini et al., 2014; Arboretti et al., 2018). Permutation solutions for multivariate location problems have been proposed and studied mainly in terms of power and robustness with respect to the underlying distribution, especially comparing their performance with that of the classic parametric tests (Pillar, 2013; Anderson, 2001; Pesarin, 2001). An interesting proposal is based on the combination of the univariate permutation tests of the marginal variables (Pesarin, 2001). Pesarin and Salmaso (2010a,b) proved that the power of the most commonly used combined permutation tests, with fixed sample size and divergent number of variables under the alternative hypothesis, tends to one in the two-sample problem. According to the type of the combining function used, a different combined test is obtained. Hence a deep study with the goal of comparing different combined tests, especially for big data with a large number of variables, is important and suitable, in order to find the most powerful test under different scenarios. To the best of our knowledge, for the multivariate multisample location problem, a comparative study of the power behaviour of the most important combined permutation tests as the number of variables diverges is missing. In particular, it is useful to know under which conditions each of the different tests is preferable in terms of power, how the power of each test increases when the number of variables under the alternative hypothesis diverges and the power behaviour of each test as a function of the proportion of true alternative hypotheses. The purpose of this paper is to fill this gap in the literature about combined permutation tests. The paper is organized as follows. Section 2 is dedicated to a review of the literature on the MANOVA problem. The method of combined permutation tests is described in Section 3. In Section 4 the results of a comparative simulation study are reported and discussed. In Section 5, the application of the method to a real case study is presented. Finally, the conclusions are in Section 6. 2. Literature review The goal of several empirical studies is the comparison of two or more populations in the presence of multivariate response variables. Often, regardless of the number of factors, the problem consists in testing the significance of treatment effects or the presence of a shift in some location parameters. In what follows, the variation of population means is investigated using multivariate analysis of variance (MANOVA). To test whether there is a significant difference between group means, various parametric multivariate tests based on strong assumptions have been proposed. The most commonly used are the Hotelling T-square test (Hotelling, 1992), the test of Wilks (1932) and the proposal of Pillai (1955). The main assumptions of these tests are normality, constant variances and continuous responses. Moreover,
166 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… these methods cannot be applied for big datasets when the number of response variables is greater than the sample size. Nonparametric solutions have been proposed to overcome the limits of the tests mentioned above due to the lack of robustness with respect to the assumptions (Pesarin and Salmaso, 2010a; Bonnini et al., 2014; Pillar, 2013; Bonnini, 2016). For instance, Anderson (2001) introduced a nonparametric solution based on the permutation test for an ecological problem. The permutation test statistic was the Fisher F ratio obtained from a distance matrix, and the simulation results proved the appropriateness of the permutation test for both one-way and two-way MANOVA. Pillar studied the accuracy and power of permutation tests for MANOVA based on different test statistics. According to his study, the sum of squares between groups with the Euclidean distance was preferable to the Chord distance and the sum of Fs of univariate ANOVA. Moreover, the simulation study revealed that the permutation test was powerful also under heteroscedastic and with unbalanced samples. In the literature, several works concerning applications of permutation tests for oneway and two-way MANOVA have been published. A non-exhaustive list includes the following papers: Mantel and Valand (1970), Mielke et al. (1976), Clarke (1993), Pillar and Orlóci (1996), Legendre and Anderson (1999), Mielke and Berry (1999), McArdle and Anderson (2001), Arboretti et al. (2018), Finch (2016). However, the extension of the permutation test for two-way MANOVA requires great attention in permuting the statistical units between groups. This is because the exchangeability condition is guaranteed only within the levels of one factor by considering the second factor as a block. Thus, constrained permutations are essential (Anderson, 2001). The twosample multivariate problem has been frequently considered. See for instance Pesarin and Salmaso (2010), Polko-Zajac (2020), Bonnini and Melak Assegie (2019). Instead, the multi-sample case has been addressed by fewer authors (see Bonnini, 2016). In some cases permutation solutions for complex problems such as multiaspect tests (Polko-Zajac, 2019), directional alternatives (Bonnini et al., 2014; Arboretti and Bonnini, 2009), tests for categorical data (Arboretti and Bonnini, 2008; Bonnini, 2014) have been developed. In this paper, we focus on multi-sample location problems for numeric variables and nondirectional alternative hypotheses. 3. Methods 3.1. Multivariate permutation test The permutation test is a distribution-free test based on the assumption of exchangeability under the null hypothesis (Pesarin, 2001). To apply the permutation principle, the sample data are partitioned into groups based on the treatment levels in an experimental study and pseudogroups in an observational study. To this end,
STATISTICS IN TRANSITION new series, June 2022 167 the structure of the dataset for 𝑆2 independent samples and V-dimensional response is represented by: 𝒀𝑌 𝑖 1,2, … , 𝑛,𝑔1,2, … , 𝑆,𝑞1,2, … , 𝑉 (1) The dataset 𝒀 takes values on the 𝑉-dimensional sample space Ω for which a 𝜎-algebra 𝒜 and a nonparametric family 𝒫 of non-degenerate unknown distributions are defined, and supposed to be exchangeable. Hypothesis testing based on the permutation approach requires a clear formulation of the null hypothesis. The null hypothesis in the MANOVA problem is defined as the equality of S multivariate (unknown) distributions: 𝐻∶𝑃𝑃 ⋯𝑃 𝒀𝟏 𝒀𝟐 …𝒀𝑺. (2) Under homoscedasticity, the difference between the groups is due to a shift in location. Thus, the null hypothesis could be formulated as equality of group means for each response variable. Let 𝒀𝒈 be a 𝑉-variate numeric random variable such that 𝒀𝒈𝝁𝜹𝒈𝜺𝒈, with 𝝁 vector of 𝑉 unknown location parameters, 𝜹𝒈, 𝑔 1, … , 𝑆, vectors of 𝑉 treatment effects and 𝜺𝒈, 𝑔1, … , 𝑆, exchangeable 𝑉dimensional random vectors that follow an unknown probability distribution with equal variance-covariance matrix 𝚺 and such that 𝐸𝜺𝒈𝟎. The null hypothesis is: 𝐻∶ 𝜹𝟏 𝜹𝟐,…,𝜹𝑺𝟎 (3) A further decomposition of the null hypothesis with respect to the marginal distributions of the multivariate response can be considered. The multivariate hypothesis can be broken down into 𝑉 partial null hypotheses: 𝐻∶ ⋂𝛿 ,..,𝛿 0≡⋂𝐻 (4) where the intersection symbol means that the null hypothesis of the overall problem is true if all the 𝑉 partial null hypotheses are true. Accordingly, with a similar approach, the alternative multivariate hypothesis 𝐻 of inequality in distribution may be represented as follows: 𝐻∶⋃𝐻 (5)
168 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… where the union symbol indicates that the alternative hypothesis is true if at least one partial null hypothesis is false and 𝐻 denote the negation of the 𝑞-th partial null hypothesis. It is worth noting that directional alternatives are also possible but the purpose of this paper is to focus on two-tailed multi-sample multivariate problems. When the overall null hypothesis is true and the equality in distribution holds, the vector of 𝑉 observations concerning a generic statistical unit comes from any of the 𝑆 populations with equal probability. In other words, the exchangeability of units with respect to the populations/samples is satisfied. In order to determine the null distribution of the test statistic, all the possible assignments of the 𝑛 units to the 𝑆 samples can be considered. Without loss of generality, let us assume that the 𝑛 units of the first sample correspond to the first 𝑛 rows of the observed dataset 𝒀, the 𝑛 units of the second sample correspond to the next 𝑛 rows of the dataset, and so on, until the 𝑛 units of the 𝑆-th sample that correspond to the last 𝑛 rows of the dataset. Each possible assignment is equivalent to a permutation of the rows of the dataset or to resampling without replacement the 𝑛 units with 𝑛𝑛 𝑛⋯𝑛. For computational convenience, instead of considering the exact test, based on all the ! ∏! possible assignments of the 𝑛 units to the 𝑆 groups, a random sample of permutations is used according to the Conditional Monte Carlo method. 3.2. Partial tests The application of the method of Combined Permutation Test to the permutation MANOVA presented above consists in carrying out one univariate permutation test for each partial hypothesis and in combining the 𝑝-values of the univariate tests. The dependence between the univariate partial test statistics, according to the permutation distribution, is taken into account in the resampling strategy by permuting the rows of the observed dataset instead of permuting the elements of each column independently of the other columns. A suitable test statistic for each partial permutation test is the so-called Treatment Sum of Squares (𝑆𝑆), which depends on the deviations of the within-group sample means from the total sample mean. Hence, the 𝑞 partial test statistic or equivalently the test statistic of the 𝑞 partial test, with 𝑞1,2, … , 𝑉, is 𝑇∑𝑛𝑌 𝑌 ∙ (6) with 𝑌 ∙ ∑ ∑ ∑ , where 𝑌 represents the mean of the values of the 𝑞-th variable observed in the 𝑔-th sample.
STATISTICS IN TRANSITION new series, June 2022 169 The multivariate permutation distribution of the test statistic 𝑻𝑇 ,𝑇,…,𝑇 under the null hypothesis is obtained through the following procedure: 1) compute the vector of observed values of 𝑻 from the dataset 𝒀: 𝑻𝒐𝒃𝒔 𝑻𝒀𝑇 ,,𝑇,,…,𝑇, 2) randomly permute the rows of the dataset (or reassign statistical units to groups) and compute the values of the test statistics as a function of the permuted dataset: 𝑻𝑻𝒀 3) repeat step (2) 𝑅 times independently and compute the permutation test statistics. Let 𝑇, be the value of the 𝑞-th partial test statistic related to the 𝑟-th permutation of the dataset 𝒀𝒓 𝒑. Hence 𝑻𝒓 𝒑𝑻𝒀𝒓 𝒑𝑇 , ,𝑇, ,…,𝑇, 4) estimate the significance level function of the partial tests 𝜆 , 𝜆𝑇 , ∑, , . (7) with 𝑟1, 2, … , 𝑅, 𝑞1, 2, … , 𝑉, and 𝐼𝐸 indicator function of 𝐸, which takes value 1 if 𝐸 is true and 0 otherwise. The 𝑝-value of the 𝑞-th partial test is 𝜆 , 𝜆𝑇 , . 3.3. Combination According to the method based on the combination of dependent permutation tests, the test statistic for the overall problem is obtained by combining the p-values of the partial tests. The synthesis of the information provided by the partial tests regarding the marginal variables is provided by the application of a suitable combining function 𝜑. Hence, the test statistic useful for the overall test, the multivariate analysis of variance, is 𝑇 𝜑𝜆,𝜆,…,𝜆. The proposal of combining 𝑝-values of partial tests in order to solve multivariate, multi-aspect, multi-strata tests, or other complex testing problems that can be broken down into partial univariate tests, appeared for the first time in the literature twenty years ago in Pesarin (2001) and was later studied and developed by several authors. For extended but not exhaustive reviews, see Pesarin and Salmaso (2010a) and
170 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… Bonnini et al. (2014). Since, for the combination of the partial tests, 𝜑∙ must satisfy some simple, mild and easily attainable conditions, several different functions can be used and each of them corresponds to a different solution with specific properties within the family of combined permutation tests. A suitable combining function 𝜑:0,1→ℝ must satisfy the following properties: 1) ∀𝜆 ,𝜆 in 0,1, 𝜆 𝜆 ⇔ 𝜑…,𝜆 ,…𝜑…,𝜆 ,… ceteris paribus (non-increasing monotony) 2) ∃𝜆𝜖𝜆,𝜆,…,𝜆 s.t. 𝜆→0 ⇔ 𝜑𝜆,𝜆,…,𝜆→𝜑∞ (finite supremum) 3) ∀𝛼𝜖0,1, ∃𝑇,< 𝜑 where 𝑇, is the test critical value (finite critical value) The most popular combining functions in the literature of combined permutation tests are Fisher, Liptak and Tippett functions. The Fisher omnibus combining function is 𝑇2∑𝑙𝑜𝑔𝜆 (8) where 𝑙𝑜𝑔𝑥 denotes the natural logarythm of 𝑥. Liptak`s combining function is based on the transformation of the complement to one of the 𝑝-values through the inverse of the cumulative distribution function (or the quantile function) of the standard normal distribution: 𝑇∑Φ 1𝜆 (9) where Φ𝑥𝑃𝑋𝑥 with 𝑋~𝒩0,1. Tippett combination is based on an order statistic and considers, as observed value of the combined test statistic, the complement to one of the most significant 𝑝-value: 𝑇𝑚𝑎𝑥1𝜆 (10) Under the null distribution, if the 𝑉 partial tests are independent and continuous, the Tippett function follows the uniform distribution in 0,1. Without loss of generality, let us assume that the null hypotheses of the overall and partial problems are rejected for large values of the respective test statistics. It is trivial to show that all three combination rules defined above satisfy this condition. Given that the observed value of the combined test statistic is 𝑇, 𝜑𝜆 , ,𝜆 , ,…,𝜆 , . the 𝑝-value of the permutation MANOVA with the combined permutation test is given by 𝜆 , 𝜆𝑇 , (11)
STATISTICS IN TRANSITION new series, June 2022 177 Furthermore, unlike the parametric approach, the permutation test does not require that a specific underlying family of distributions is known or assumed. The null permutation distribution of the test statistics can be determined regardless of whether the underlying distribution of the data is continuous or not. The 79 statements are reported in Appendix 1. Let 𝑌 be the random variable that represents the response concerning the 𝑣-th statement of an employee belonging to group 𝑔, with 𝑣1,2, … ,79 and 𝑔∈𝐺 𝐹𝑈50, 𝐹𝑂50, 𝑀𝑈50, 𝐹𝑂50. The testing problem can be represented by the following hypotheses: 𝐻: 𝑌, 𝑌, 𝑌, 𝑌, vs 𝐻:∃𝑔,𝑔 ∈𝐺 s. t. 𝑌,𝑌, The significance level is 𝛼0.05. According to the simulation study, the most suitable testing method seems to be the combined permutation test based on the Tippett combining function. The application of this test provides a p-value of 0.755, much greater than 𝛼. Hence the null hypothesis cannot be rejected. At the significance level 0.05, there is no empirical evidence to reject the null hypothesis of no difference of the organizational well-being between groups in favor of the hypothesis that the organizational well-being of the groups is not the same. In other words, we cannot conclude that there is a significance effect of gender and age on the employees’ wellbeing. The analysis was carried out by the authors by creating specific R scripts for the implementation of the methodology. It is worth noting that the final p-value of the combined test is invariant with respect to the combination strategy. In other words, if we perform a two-level combination, i.e. the first within-domain combination of partial tests and the second combination with respect to the domains, the final result is the same as obtained by permuting the partial tests all together at the same time (see Pesarin, 2001). If we had significance in the overall test, it would be useful to identify the partial tests that contribute to the overall significance. This can be done with a suitable adjustment of the p-values of the partial tests for controlling the Family Wise Error rate and avoiding the inflation of the type I error of the final combined test.
178 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… In this case, an interesting two-stage combination strategy could be of interest, because the questionnaire is divided into sections corresponding to partial aspects of organizational well-being. Each aspect corresponds to a set of questions and consequently to a domain of variables (construct). In the case of significance of the overall combined test, the analysis of the adjusted p-values of the partial combined tests related to the constructs would make sense. Unfortunately, the overall null hypothesis is not rejected. This result proves that, in the University of Ferrara, the organizational well-being of the employees in terms of risks, working environment, respect, relationship with colleagues and office manager, transparency, motivation, etc. is not affected by age and gender. It could be considered as evidence of gender-age equality within the organization. 6. Conclusions The purpose of the work is to deepen the study of the power behaviour of combined permutation tests for MANOVA problems with big data. The assessment of the convergence rate of the power to one as the proportion of variables under the alternative hypothesis increases and a comparison between the three most commonly used members within this family of tests represent the main scientific added value of the paper. These nonparametric multi-sample location tests are well approximated, consistent, unbiased and powerful also for small sample sizes. The power is also an increasing function of the number of samples and of the number of variables of the dataset. The asymptotic behaviour of the tests when the number of variables diverges was studied and the simulations proved that the proportion of true partial alternative hypotheses is more important than the absolute number of variables of the dataset in explaining the increase of power. The test based on the Tippett combination represents an exception to this general rule. This test seems to be much more powerful than the others when the proportion of true partial alternative hypotheses is not large but competitive also when the proportions are close to one. This is the only condition in which the test based on Liptak combination is competitive but, for small proportions of true alternatives, this test is by far the least powerful. Definitely, it seems that, among the distribution free solutions to the multivariate analysis of variance in the family of combined permutation tests, the method based on the Tippet combination is in general preferable, especially if there are no preventive information about the possible percentage of variables (or marginal distributions) under the alternative hypothesis. Instead of the Tippett combination, the Fisher rule
STATISTICS IN TRANSITION new series, June 2022 179 can be applied when the percentage is close to 100%. The Liptak combination seems to be non-convenient in general. This methodological tool is an important and useful solution of testing problems for big data, especially when the number of variables is very large and the sample sizes are small. The usefulness and the effectiveness of the method is confirmed by the application to the case study concerning the survey on the organizational well-being at the University of Ferrara discussed in the paper. Acknowledgments The authors wish to thank the anonymous reviewers who have contributed significantly to the quality of the paper with their comments and suggestions. The work was partially supported by the University of Ferrara, which funded the FIR project “Measuring and assessing social sustainability, inclusion and accessibility at university: methods and applications”. References Anderson, M. J., (2001). A new method for non‐parametric multivariate analysis of variance. Austral ecology, 26(1), pp. 32–46. Arboretti, R., Bonnini, S., (2008). Moment-based multivariate permutation tests for ordinal categorical data. Journal of Nonparametric Statistics, 20(5), pp. 383–393. Arboretti, R., Bonnini, S., (2009). Some new results on univariate and multivariate permutation tests for ordinal categorical variables under restricted alternatives. Statistical Methods and Applications: Journal of the Italian Statistical Society, 18(2), pp. 221–236. Arboretti, R., Ceccato, R., Corain, L., Ronchi, F. and Salmaso, L., (2018). Multivariate small sample tests for two-way designs with applications to industrial statistics. Statistical Papers, 59(4), pp. 1483–1503. Baro, E., Degoul, S., Beuscart, R. and Chazard, E., (2015). Toward a literature-driven definition of big data in healthcare. BioMed research international (https://doi.org/10.1155/2015/639021). Bonnini, S., And Melak Assegie, G., (2019). Permutation multivariate tests for treatment effect: theory and recent developments. In SUSAN SSACAB 2019, pp. 30–30. The Biostatistics Research Unit of the South African Medical Research Council.
180 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… Bonnini, S., (2014). Testing for heterogeneity with categorical data: permutation solution versus bootstrap method. Communications in Statistics: Theory and Methods, 43(4), pp. 906–917. Bonnini, S., (2016). Multivariate approach for comparative evaluations of customer satisfaction with application to transport services. Communications in Statistics: Simulation and Computation, 45(5), pp. 1554–1568. Bonnini, S., Corain, L., Marozzi, M. and Salmaso, L., (2014). Nonparametric hypothesis testing: rank and permutation methods with applications in R. John Wiley & Sons. Bonnini, S., Prodi, N., Salmaso, L., Visentin, C., (2014). Permutation approaches for stochastic ordering. Communications in Statistics: Theory and Methods, 43(10-12), pp. 2227–2235. Clarke, K.R., (1993). Non‐parametric multivariate analyses of changes in community structure. Australian journal of ecology, 18(1), pp.117–143. Farcomeni, A. and Greco, L., (2016). Robust methods for data reduction. CRC press. Finch, W.H., (2016). Comparison of multivariate means across groups with ordinal dependent variables: a Monte Carlo simulation study. Frontiers in Applied Mathematics and Statistics, 2, p. 2. Hotelling, H., (1992). The generalization of Student’s ratio. In Breakthroughs in statistics, (pp. 54-65). Springer, New York, NY. Johnson, R., (1997). Wichern. D., (2007). Applied multivariate statistical analysis. Prentice-Hall: London. Legendre, P. and Anderson, M. J., (1999). Distance‐based redundancy analysis: testing multispecies responses in multifactorial ecological experiments. Ecological monographs, 69(1), pp.1–24. Mantel, N., Valand, R. S., (1970). A technique of nonparametric multivariate analysis. Biometrics, pp. 547-558. McArdle, B. H., Anderson, M. J., 2001. Fitting multivariate models to community data: a comment on distance‐based redundancy analysis. Ecology, 82(1), pp. 290– 297. Mielke Jr, P. W., Berry, K. J., (1999). Multivariate tests for correlated data in completely randomized designs. Journal of Educational and Behavioral Statistics, 24(2), pp. 109–131.
STATISTICS IN TRANSITION new series, June 2022 181 Mielke Jr, P. W., Berry, K. J., Johnson, E. S., (1976). Multi-response permutation procedures for a priori classifications. Communications in Statistics: Theory and Methods, 5(14), pp. 1409–1424. Özköse, H., Arı, E. S. and Gencer, C., (2015). Yesterday, today and tomorrow of big data. Procedia-Social and Behavioral Sciences, 195, pp. 1042–1050. Pesarin, F., (2001). Multivariate permutation tests: with applications in biostatistics, Vol. 240. Wiley: Chichester. Pesarin, F., Salmaso, L., (2010a). Permutation tests for complex data: theory, applications and software. John Wiley & Sons: Chichester. Pesarin, F., Salmaso, L., (2010b). Finite-sample consistency of combination-based permutation tests with application to repeated measures designs. Journal of Nonparaetric Statistics, 22(5), pp. 669–684. Pillai, K. S., (1955). Some new test criteria in multivariate analysis. The Annals of Mathematical Statistics, pp. 117–121. Pillar, V., (2013). How accurate and powerful are randomization tests in multivariate analysis of variance?. Community Ecology, 14(2), pp. 153–163. Pillar, V.D.P., Orlóci, L., (1996). On randomization testing in vegetation science: multifactor comparisons of relevé groups. Journal of Vegetation Science, 7(4), pp. 585–592. Polko-Zajac, D., (2019). On permutation location-scale tests. Statistics in Transition, 20(4), pp. 153-166. Polko-Zajac, D., (2020). A comparative study on the power of parametric and permutation tests for a multidimensional and two-sample location problem. Argumenta Oeconomica Cracoviensia, 2(23), pp. 69–79 Wilks, S. S., (1932). Certain generalizations in the analysis of variance. Biometrika, pp. 471–494.
182 S. Bonnini, G. Melak Assegie: Advances on Permutation Multivariate… Appendix 1 Code Statement A.01 My working place is safe A.02 I have been informed about the risks connected to my job A.03 I am satisfied about the environment of my working place A.04 I have suffered harassment A.05 My dignity has been harmed at work A.06 At work the smoking ban is respected A.07 I usually take enough breaks A.08 I can work hard A.09 I am not comfortable when I am working A.10 The colleagues are not polite with me A.11 I am allowed to take a break when I wish A.12 I don't have the chance to take enough breaks B.10 At work I have suffered bullying B.01 In the workplace I am respected in my trade union membership B.02 In the workplace I am respected in my political orientation B.03 In the workplace I am respected in my religious faith B.04 My gender identity is an obstacle to my enhancement at work B.05 In the workplace I am respected in my ethnicity and race B.06 In the workplace I am respected in relation to my mother tongue B.07 My age is an obstacle to my enhancement at work B.08 In the workplace I am respected in relation to my mother tongue C.01 The workload is assigned with equity C.02 The responsibilities are assigned with equity C.03 My salary is proportional to the commitment C.04 The pay is differentiated according to quantity and quality of work C.05 My manager makes work decisions impartially D.01 At UNIFE the path of professional development of each employee is well defined and clear D.02 At UNIFE the career opportunities depend on merit D.03 UNIFE gives the possibility to develop skills and aptitudes of individuals in relation to the requirements of the different roles D.04 My current role is appropriate to my professional profile D.05 I am satisfied with my professional path within UNIFE E.01 I know what is expected of my work E.02 I have the skills to do my job E.03 I have the resources and tools to do my job E.04 I have an adequate level of autonomy in my work E.05 My work gives me a sense of personal fulfilment E.06 I know how to do my job E.07 I understand what is expected of me at work E.08 I have freedom of choice in deciding how to do my job E.09 I have unattainable deadlines
STATISTICS IN TRANSITION new series, June 2022 183 Code Statement E.10 I have to work very hard E.11 I have a say in deciding how fast I can do my job E.12 I’m getting pressure to work overtime E.13 I have freedom of choice in deciding what to do at work E.14 I have to do my job very quickly E.15 I have deadlines impossible to meet E.16 I have a say in how to do my job E.17 My working hours can be flexible E.18 Job requests made to me by various people/offices are difficult to combine F.01 I feel part of a team F.02 I help colleagues even if it’s not my job F.03 I am esteemed and treated with respect by colleagues F.04 In my group, those who have information make it available to everyone F.05 The organization pushes to work in a group and to collaborate F.06 If the job becomes difficult, I can count on the help of my colleagues F.07 At work my colleagues show me the respect I deserve F.08 I receive support information that helps me in my work F.09 There are frictions or conflicts between colleagues F.10 My colleagues give me the help and support I need F.11 Colleagues are willing to listen to my work problems G.01 My organization invests in people, including through adequate training G.02 The rules of conduct are clearly defined G.03 Organisational tasks and roles are well defined G.04 The circulation of information within the organisation is appropriate G.05 My organisation promotes measures to reconcile working time and life time G.06 I have clear duties and responsibilities G.07 I must neglect some tasks because I have too much to do G.08 I know the goals of my department/office G.09 Staff are always consulted on changes in work G.10 I’m supported in emotionally challenging jobs G.11 Workplace relations are strained H.01 I am proud when I tell someone that I work at UNIFE H.02 I am proud when UNIFE achieves good results H.03 I am sorry if someone has a bad opinion of UNIFE H.04 Values and behaviours at UNIFE are similar to mine H.05 If possible, I would change company I.01 Relative and friends think that UNIFE is important for the collectivity I.02 Students think that UNIFE is important for the collectivity I.03 People think that UNIFE is important for the collectivity