Full text
Mathematics and Computers in Simulation 223 (2024) 588–600 Available online 23 April 2024 0378-4754/© 2024 The Author(s). Published by Elsevier B.V. on behalf of International Association for Mathematics and Computers in Simulation (IMACS). This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Contents lists available at ScienceDirect Mathematics and Computers in Simulation journal homepage: www.elsevier.com/locate/matcom Original articles Testing for proportions when data are classified into a large number of groups M.V. Alba-Fernándeza,∗, M.D. Jiménez-Gamerob, F. Jiménez-Jiménezc aDpt. Statistics and O.R. University of Jaén, Spain bDpt. Statistics and O.R. University of Sevilla, Spain cDpt. Economics. University of Jaén, Spain ARTICLE INFO Keywords: Multinomial data Goodness-of-fit Large number of subpopulations ABSTRACT When dealing with categorical data, a common concern is to check if the observed relative frequencies agree with a certain fixed vector of ideal proportions. Suppose that the population is divided into subpopulations or groups. In such a case, the ideal proportions could vary among groups and one may be interested in simultaneously testing if the observed proportions agree with those ideal proportions in all groups. A novel procedure is proposed for carrying out such a testing problem. The test statistic is shown to be asymptotically normal, avoiding the use of complicated resampling methods to get 𝑝-values. The asymptotic behavior of the test under alternatives is also studied. Here, by asymptotic, we mean when the number of groups increases; the sample sizes of the data from each group can either stay bounded or grow with the number of groups. The finite sample performance of the new test is empirically evaluated through an extensive simulation study. The usefulness of the proposal is illustrated with some data sets. 1. Introduction Suppose the population can take values in one of 𝐺distinct categories, that may be either defined in a natural way (e.g. nationality) or obtained by discretizing the outcomes of an experiment (e.g. age groups), and the user is interested in checking if the vector of population proportions of these categories, say 𝝅= (𝜋1,…, 𝜋𝐺)⊤, is equal to a certain fixed vector, say 𝝅0= (𝜋01,…, 𝜋0𝐺)⊤, based on a sample. This problem has been extensively studied in the statistical literature (see, e.g. the books by [15,19]). A classical procedure for dealing with this problem is the chi-square test, which compares the observed frequencies in the sample with their expected values under the null hypothesis 𝐻01 ∶𝝅=𝝅0. Along this paper it is assumed that the population is divided into a large number, say 𝑘, of subpopulations or groups, that independent samples are available from each group, and that the interest is centered in simultaneously testing if the vector of population proportions in each group is equal to a certain fixed vector, which may vary among groups. Specifically, assume we have 𝑘independent vectors of multinomial data, all of them with the same number of categories or classes, 𝐺(≥2), 𝑵𝑖= (𝑛𝑖1,…, 𝑛𝑖𝐺)⊤∼(𝑛𝑖.;𝝅𝑖),1≤𝑖≤𝑘, with 𝝅𝑖= (𝜋𝑖1,…, 𝜋𝑖𝐺)⊤∈𝛯≥,𝐺 = {(𝜋1,…, 𝜋𝐺)⊤,𝜋𝑔≥0,1≤𝑔≤𝐺, ∑𝐺 𝑔=1 𝜋𝑔= 1},1≤𝑖≤𝑘.𝐺is assumed to be a fixed quantity. Based on the data, we are interested in testing 𝐻0∶𝝅𝑖=𝝅0𝑖,1≤𝑖≤𝑘, ∗Corresponding author. E-mail address: [email protected] (M.V. Alba-Fernández). https://doi.org/10.1016/j.matcom.2024.04.019 Received 9 November 2023; Received in revised form 23 March 2024; Accepted 18 April 2024
Mathematics and Computers in Simulation 223 (2024) 588–600 589 M.V. Alba-Fernández et al. Table 1 Estimated type I errors for the nominal value 𝛼= 0.05 when using the BH method and 𝑝-values are calculated with the asymptotic null distribution. 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 50 10 3.2 13.9 15.7 8.6 15.8 24.4 500 10 5.3 16.0 47.7 9.8 24.4 39.1 20 5.1 7.5 12.4 7.4 9.8 13.4 20 4.3 12.6 25.8 11.4 20.3 24.5 30 4.4 7.4 9.6 6.4 8.5 10.2 30 5.5 13.1 22.1 10.1 15.7 23.3 100 10 3.9 11.9 18.1 15.3 20.6 29.7 1000 10 2.0 26.4 71.0 14.2 31.8 40.9 20 6.3 11.5 15.5 8.7 12.6 19.2 20 7.7 17.5 37.4 12.0 22.7 34.2 30 5.2 8.4 12.4 7.1 10.2 13.0 30 7.0 15.4 25.0 10.3 18.4 27.3 200 10 7.8 10.8 25.8 16.2 24.2 26.1 2000 10 4.1 45.9 84.4 16.2 36.8 53.7 20 2.2 11.9 17.3 8.7 16.0 22.9 20 7.1 25.8 56.7 16.2 29.7 43.5 30 5.8 9.3 17.5 8.2 12.3 17.0 30 7.2 15.7 31.2 13.0 23.2 32.4 Table 2 Estimated type I errors for the nominal value 𝛼= 0.05 when using the BH method with exact 𝑝-values. 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 50 10 13.3 4.1 5.2 7.9 5.1 5.8 500 10 5.7 7.2 4.2 4.5 5.8 4.8 20 5.1 4.6 4.5 5.3 4.9 4.4 20 5.8 5.7 4.8 5.2 5.5 4.3 30 5.6 5.1 4.7 4.3 4.8 4.7 30 5.1 4.5 4.6 5.3 4.9 5.0 100 10 6.6 4.8 4.9 9.2 5.0 5.1 1000 10 2.2 5.4 3.6 8.7 5.8 4.6 20 6.3 5.5 5.1 6.0 5.3 4.4 20 6.5 4.0 4.9 6.6 5.4 5.0 30 4.9 4.2 5.2 4.6 5.0 4.6 30 4.7 4.0 5.1 5.9 5.0 5.4 200 10 7.7 4.6 4.9 16.6 6.4 5.3 2000 10 3.9 3.0 4.1 5.6 5.8 5.6 20 6.6 5.1 5.2 6.7 5.0 4.8 20 5.0 4.5 3.2 4.9 4.6 4.3 30 4.6 4.5 5.0 5.1 5.0 4.4 30 6.4 4.2 6.2 6.1 4.5 5.8 𝐻1∶𝝅𝑖≠𝝅0𝑖,for some 𝑖, for some 𝝅0𝑖= (𝜋0𝑖1,…, 𝜋0𝑖𝐺)⊤∈𝛯>,𝐺 = {(𝜋1,…, 𝜋𝐺)⊤,𝜋𝑔>0,1≤𝑔≤𝐺, ∑𝐺 𝑔=1 𝜋𝑔= 1},1≤𝑖≤𝑘, for large 𝑘. As mentioned before, a classical test of 𝐻0𝑖∶𝝅𝑖=𝝅0𝑖v.s. 𝐻1𝑖∶𝝅𝑖≠𝝅0𝑖, for any fixed 𝑖, is the chi-square test. Let 𝜒2 𝑖denote the chi-square test statistic for testing 𝐻0𝑖vs. 𝐻1𝑖, defined as 𝜒2 𝑖= 𝐺 ∑ 𝑔=1 (𝑛𝑖𝑔 −𝑛𝑖.𝜋0𝑖𝑔)2 𝑛𝑖.𝜋0𝑖𝑔 .(1) Under 𝐻0𝑖,𝜒2 𝑖is asymptotically distributed as a 𝜒2 𝐺−1, a chi-square distribution with 𝐺−1 degrees of freedom, where asymptotically means that 𝑛𝑖. increases arbitrarily. The problem of testing 𝐻0can be handled by testing each hypothesis that composes it, that is, testing 𝐻0𝑖∶𝝅𝑖=𝝅0𝑖, against 𝐻1𝑖∶𝝅𝑖≠𝝅0𝑖,1≤𝑖≤𝑘, obtaining 𝑝1,…, 𝑝𝑘, the 𝑝-values for each test, and then applying some method to adjust them, as the Bonferroni method, which controls the family-wise error rate, or the Benjamini–Hochberg (BH) method (see, e.g. [5]), which controls the false discovery rate when the 𝑘tests are independent. Both procedures agree in rejecting 𝐻0, which is tantamount to reject 𝐻0𝑖for some 𝑖, if min1≤𝑖≤𝑘𝑝𝑖≤𝛼∕𝑘. We numerically checked the level when using this method in the following cases: lev1. 𝐺= 5,𝝅0𝑖= (0.2,…,0.2)⊤,1≤𝑖≤𝑘, lev2. 𝐺= 5,𝝅0𝑖= (0.1,0.1,0.2,0.3,0.3)⊤,1≤𝑖≤𝑘, lev3. 𝐺= 5,𝝅0𝑖= (0.05,0.1,0.2,0.3,0.35)⊤,1≤𝑖≤𝑘, lev4. 𝐺= 10,𝝅0𝑖= (0.1,…,0.1)⊤,1≤𝑖≤𝑘, lev5. 𝐺= 10,𝝅0𝑖= (0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.1,0.1,0.2)⊤,1≤𝑖≤𝑘, lev6. 𝐺= 10,𝝅0𝑖= (0.05,0.05,0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.3)⊤,1≤𝑖≤𝑘. For each case, 𝑘random samples of size 𝑛𝑖. obeying the null hypothesis were generated; after calculating 𝜒2 𝑖, its 𝑝-value was computed using the asymptotic null distribution, 1≤𝑖≤𝑘. The experiment was repeated 10,000 times, and the percentage of times that 𝐻0 was rejected taking 𝛼= 0.05 is displayed in Table 1 for 𝑛𝑖. = 10,20,30,1≤𝑖≤𝑘and 𝑘= 50,100,200,500,1000,2000. Looking at Table 1 it seems that larger sample sizes are required for the validity of this methodology, especially when 𝝅becomes far apart from the equiprobably case (observe that lev3 and lev6 give worse results than lev2 and lev5, respectively). Table 2 contains the results when the 𝑝-values were calculated with the exact null distribution (which can be computed by simulation). The results shown in Table 2 are much better than those in Table 1 in the sense that, in most cases, the empirical levels are not so far apart from the nominal value. Nevertheless, there are some settings (see, for example, 𝑛𝑖. = 10, for lev1 and lev4, with 𝑘≤200) where the empirical levels are not close to the nominal value. A practical disadvantage of this procedure is that its application requires the calculation of exact critical points, which may be very time-consuming when the sample sizes and the probability vectors 𝝅0𝑖vary across subpopulations. This requirement could constitute a disincentive to its use. The problem of simultaneously testing for 𝑘subpopulations, when 𝑘is large and independent samples are available from each of them, is not new. Some authors have studied several testing problems in this setting, where it is usually assumed that the sample
Mathematics and Computers in Simulation 223 (2024) 588–600 590 M.V. Alba-Fernández et al. sizes are much smaller than the number of subpopulations. Some examples are the tests developed in [10,16,22] for testing the equality of the 𝑘distributions; in [1,7,11,18] for testing the equality of the means of the 𝑘distributions; in [2,17] for testing the homogeneity of marginal distributions in the 𝑘populations; and in [9,12] for testing goodness-of-fit, to cite just a few. To the best of our knowledge, the problem of testing 𝐻0vs. 𝐻1has not been dealt with, and, as it will be shown in Section 4, it is relevant from a practical point of view. This paper proposes and studies a novel test of 𝐻0vs. 𝐻1, whose practical application does not require the calculation of critical points, as the test statistic has asymptotically a standard normal distribution. The test is introduced in Section 2. This section also derives its asymptotic null distribution, where by asymptotic we mean when 𝑘→∞. No restriction is assumed on the sample sizes, which can remain bounded or increase with 𝑘. Additionally, this section studies the behavior of the new test under alternatives, showing that, under quite weak assumptions, it is consistent so that it can detect any alternative for large 𝑘. It is shown that the proposed test can be also derived by using the union-intersection principle. The proposed test is asymptotic, in the sense that the critical point (or equivalently, the 𝑝-value) is based on the asymptotic null distribution. To check the actual level for moderate values of 𝑘, we repeated the previous simulation experiment for the new test and observed that the empirical levels are very close to the nominal value in all cases. We also carried out a further simulation experiment to study the power of our test and compare it with that of the previous procedure. In almost all cases, the proposal outperforms that procedure. Section 3describes and summarizes the results of the simulation study. Section 4illustrates the application to two real data sets. Section 5concludes. All proofs are deferred to Section 6. Throughout the paper we will make use of the following standard notation: all random variables and random elements will be defined on a sufficiently rich probability space (𝛺, ,P); the symbols Eand Vdenote expectation and variance, respectively; P0,E0 and V0denote probability, expectation and variance under the null hypothesis, respectively; →means convergence in distribution; all limits in this paper are taken as 𝑘→∞. The sample sizes 𝑛1.,…, 𝑛𝑘. are allowed to vary with 𝑘, so they should be denoted as 𝑛1.(𝑘),…, 𝑛𝑘.(𝑘), but to simplify notation, such dependence will be skipped. 2. The test 2.1. The test statistic and its asymptotic null distribution To derive a test of 𝐻0vs. 𝐻1, we first observe two facts, stated in Lemmas 1 and 2below. Recall the definition of 𝜒2 𝑖in (1). Lemma 1 gives the expected value and the variance of 𝜒2 𝑖when 𝐻0is true. Lemma 1. We have that E0(𝜒2 𝑖) = 𝐺− 1, 𝜎2 0𝑖=V0(𝜒2 𝑖) = 2 (1 − 1 𝑛𝑖. )(𝐺− 1) + 1 𝑛𝑖. (𝐺 ∑ 𝑔=1 1 𝜋0𝑖𝑔 −𝐺2),(2) 1≤𝑖≤𝑘. The next result shows that the expectation of the statistic 𝜒2 𝑖is greater under alternatives than under 𝐻0𝑖. Lemma 2. If 𝐻0𝑖is not true, then E(𝜒2 𝑖)> 𝐺 − 1,∀𝑛𝑖. ≥2. As a consequence of Lemmas 1 and 2, large values of (1∕𝑘)∑𝑘 𝑖=1 𝜒2 𝑖− (𝐺− 1) should lead us to reject 𝐻0. Because of this reason, for testing 𝐻0vs. 𝐻1we consider the following test statistic 𝑇𝑘=√𝑘(1∕𝑘)∑𝑘 𝑖=1 𝜒2 𝑖− (𝐺− 1) √(1∕𝑘)∑𝑘 𝑖=1 𝜎2 0𝑖 . From Lemma 2, we have that E(𝑇𝑘)≥0, with E(𝑇𝑘)=0if and only if 𝐻0is true. Thus, as observed before, it is reasonable to reject the null hypothesis for large values of 𝑇𝑘. Now, to determine what are ‘‘large values’’ we have to calculate its distribution under the null hypothesis, or at least an approximation. To this end, the next theorem derives the asymptotic null distribution of 𝑇𝑘. Theorem 1. Suppose that 𝐻0is true and that 𝑛𝑖. ≥2,∀𝑖, then 𝑇𝑘 →𝑍, where 𝑍∼𝑁(0,1). Let 𝛼∈ (0,1). For testing 𝐻0vs. 𝐻1, we consider the test that rejects the null when 𝑇𝑘≥𝑧1−𝛼,(3) where 𝛷(𝑧1−𝛼) = 1 − 𝛼and 𝛷stands for the cumulative distribution function of the standard normal distribution. From Theorem 1, it has asymptotic level 𝛼. The critical region (3) can be equivalently written as 𝑎𝑝 = 1 − 𝛷(𝑇𝑘,𝑜𝑏𝑠)≤𝛼, (4) where 𝑇𝑘,𝑜𝑏𝑠 is the observed value of the statistic 𝑇𝑘at the data, that is, 𝑎𝑝 is the asymptotic 𝑝-value. Remark 1. Notice that, to derive the above result, no assumptions have been made on the sample sizes (apart from 𝑛𝑖. ≥2); hence, it is valid whether the sample sizes remain bounded or increase with 𝑘at any rate.
Mathematics and Computers in Simulation 223 (2024) 588–600 591 M.V. Alba-Fernández et al. 2.2. Behavior of the test under alternatives To study the behavior of the test (4) under alternatives, we assume that 𝜋0𝑖𝑔 ≥𝛿, 1≤𝑖≤𝑘, 1≤𝑔≤𝐺, for some 𝛿 > 0. (5) First, we see that 𝑉1≤𝜍2 0𝑘= (1∕𝑘) 𝑘 ∑ 𝑖=1 𝜎2 0𝑖≤𝑉2,∀𝑘, for some 0< 𝑉1< 𝑉2<∞.(6) From (5), we have that 𝛿≤1∕𝐺, with 𝛿= 1∕𝐺if and only if 𝜋0𝑖𝑔 = 1∕𝐺,∀𝑖,∀𝑔. Therefore, 0≤1∕𝛿−𝐺 < ∞. From (2) and (5), it readily follows that 𝜎2 0𝑖≤𝑉2= 2(𝐺− 1) + 𝐺(1∕𝛿−𝐺)<∞,1≤𝑖≤𝑘. As a consequence of the next lemma, we can take 𝑉1=𝐺− 1. Lemma 3. If 𝑛𝑖. ≥2, then V0(𝜒2 𝑖)≥𝐺− 1. Next, we study the behavior of (1∕𝑘)∑𝑘 𝑖=1 𝜒2 𝑖− (𝐺− 1) under alternatives. We can write 𝜒2 𝑖= 𝐺 ∑ 𝑔=1 [𝜋𝑖𝑔(1 − 𝜋𝑖𝑔) 𝜋0𝑖𝑔 +𝑛𝑖. (𝜋𝑖𝑔 −𝜋0𝑖𝑔)2 𝜋0𝑖𝑔 ]+ 𝐺 ∑ 𝑔=1 (𝑛𝑖𝑔 −𝑛𝑖.𝜋𝑖𝑔)2−𝑛𝑖.𝜋𝑖𝑔 (1 − 𝜋𝑖𝑔) 𝑛𝑖.𝜋0𝑖𝑔 + 2 𝐺 ∑ 𝑔=1 (𝜋𝑖𝑔 −𝜋0𝑖𝑔)(𝑛𝑖𝑔 −𝑛𝑖.𝜋𝑖𝑔) 𝜋0𝑖𝑔 ∶= 𝜇𝑖+𝜒2 𝑖1+ 2𝜒2 𝑖2 and 1 𝑘 𝑘 ∑ 𝑖=1 𝜒2 𝑖=1 𝑘 𝑘 ∑ 𝑖=1 𝜇𝑖+1 𝑘 𝑘 ∑ 𝑖=1 𝜒2 𝑖1+1 𝑘 𝑘 ∑ 𝑖=1 𝜒2 𝑖2∶= 𝜇+𝜒2 .1+ 2𝜒2 .2. From Lemmas 1 and 2, it follows that 𝜇𝑖− (𝐺− 1) ≥0,∀𝑖, with 𝜇𝑖− (𝐺− 1) = 0 if and only if 𝐻0𝑖is true. In general, we have that if 𝝅𝑖is such that 𝐺 ∑ 𝑔=1 (𝜋𝑖𝑔 −𝜋0𝑖𝑔)2 𝜋0𝑖𝑔 = 𝐺 ∑ 𝑔=1 𝜋2 𝑖𝑔 𝜋0𝑖𝑔 − 1 = 𝛿𝑖,(7) then 𝜇𝑖= 𝐺 ∑ 𝑔=1 𝜋𝑖𝑔 𝜋0𝑖𝑔 − 1 + (𝑛𝑖. − 1)𝛿𝑖. As ∑𝐺 𝑔=1 𝜋𝑖𝑔∕𝜋0𝑖𝑔 ≥∑𝐺 𝑔=1 𝜋2 𝑖𝑔∕𝜋0𝑖𝑔 =𝛿𝑖+ 1, we get that 𝜇𝑖≥𝑛𝑖.𝛿𝑖. Therefore, if 𝐻0𝑖is false (which is equivalent to say that 𝛿𝑖in (7) is positive) then, for big enough 𝑛𝑖.,𝜇𝑖− (𝐺− 1) is a positive large quantity. The precise meaning of ‘‘big enough’’ depends on 𝝅𝑖and 𝝅0𝑖. Later (just after Lemma 4), we will consider a special case for 𝝅0𝑖where ‘‘big enough’’ means 𝑛𝑖. ≥2,∀𝑖, which is not restrictive at all. Next, we see that for large 𝑘the term 𝜇dominates (1∕𝑘)∑𝑘 𝑖=1 𝜒2 𝑖, that is, (1∕𝑘)∑𝑘 𝑖=1 𝜒2 𝑖− (𝐺− 1) ≃ 𝜇− (𝐺− 1). This fact, together with (6), implies that if the sample sizes, 𝝅1,…,𝝅𝑘and 𝝅01,…,𝝅0𝑘are such that 𝜇− (𝐺− 1) ≥𝜏 > 0, then 0≤𝑎𝑝 ≤1 − 𝛷{√𝑘 √𝑉2(1 𝑘 𝑘 ∑ 𝑖=1 𝜒2 𝑖− (𝐺− 1))}≃ 1 − 𝛷{√𝑘 √𝑉2[𝜇− (𝐺− 1)]}≤1 − 𝛷(√𝑘𝜏 √𝑉2)→0, and thus, the test (3) will reject 𝐻0. For example, if 𝛿𝑖> 𝜌 > 0for a proportion 𝑝(>0) of the groups, and 𝛿𝑖= 0 in the rest, then 𝜇− (𝐺− 1) ≥𝑝𝜌 > 0, and therefore 𝑎𝑝 →0. Lemma 4 below shows that, under certain no restrictive assumptions, 𝜒2 .1and 𝜒2 .2are asymptotically negligible. Lemma 4. If (5) holds, (a) then 𝜒2 .1converges in probability to 0; (b) if, additionally, the sample sizes satisfy 𝑛𝑖.∕𝑘→0,∀𝑖, then 𝜒2 .2converges in probability to 0.
Mathematics and Computers in Simulation 223 (2024) 588–600 592 M.V. Alba-Fernández et al. As said before, the precise meaning of ‘‘big enough’’ depends on the testing problem at hand, that is, on the alternatives 𝝅1,…,𝝅𝑘, and on 𝝅01,…,𝝅0𝑘. Now, we consider the case where 𝝅01 =⋯=𝝅0𝑘= (1∕𝐺, …,1∕𝐺). In this setting, if 𝝅𝑖is such that (7) holds, then 𝜇𝑖=𝐺− 1 + (𝑛𝑖. − 1)𝛿𝑖and thus, 𝜇− (𝐺− 1) = (1∕𝑘) 𝑘 ∑ 𝑖=1 (𝑛𝑖. − 1)𝛿𝑖≥(1∕𝑘) 𝑘 ∑ 𝑖=1 𝛿𝑖, because we are assuming that 𝑛𝑖. ≥2,∀𝑖. Therefore, if 𝛿𝑖> 𝜌 > 0for a proportion 𝑝(>0) of the groups, and 𝛿𝑖= 0 in the rest, then 𝜇− (𝐺− 1) ≥𝑝𝜌 > 0, and therefore 𝑎𝑝 →0. Observe that, in this setting, ‘‘big enough’’ means 𝑛𝑖. ≥2, which imposes no further restriction on the sample sizes. Remark 2. As noted in Remark 1, the result in Theorem 1 is valid for any sample sizes. Under alternatives, for 𝜒2 .2to be negligible, it has been assumed that 𝑛𝑖.∕𝑘→0,∀𝑖. It allows the sample sizes to remain bounded or to increase with 𝑘, but at a lower rate. Although this assumption restricts the sample sizes, in practice, it is not severe at all, as when there is a large number of populations it does not seem plausible to have many data coming from each population. Remark 3. So far we have assumed that, under the null, the 𝐺categories have positive probability in all populations. From the previous developments (Sections 2.1 and 2.2), it is clear that such an assumption can be relaxed to 𝝅0𝑖∈𝛯>,𝐺𝑖, with 2≤𝐺𝑖≤𝐺, for some fixed 𝐺,1≤𝑖≤𝑘. In such a case, to calculate 𝜎0𝑖one must replace 𝐺with 𝐺𝑖in (2), and 𝐺with (1∕𝑘)∑𝑘 𝑖=1 𝐺𝑖in the expression of 𝑇𝑘. 2.3. Another derivation of the test The test (3) has been built with the help of Lemmas 1 and 2and Theorem 1. This section provides another way of achieving it. Specifically, we will see that it can be derived by using the union-intersection principle and the result in Theorem 1. With this aim, we closely follow the reasoning provided in [4], which derived a similar result for testing each of the one-population hypotheses 𝐻0𝑖. As ∑𝐺 𝑔=1 𝜋𝑖𝑔 = 1 and ∑𝐺 𝑔=1 𝜋0𝑖𝑔 = 1,1≤𝑖≤𝑘,𝐻0is equivalent to 𝐻0∶𝝕𝑖=𝝕0𝑖,1≤𝑖≤𝑘, (8) where 𝝕𝑖= (𝜋𝑖1,…, 𝜋𝑖𝐺−1)⊤and 𝝕0𝑖= (𝜋0𝑖1,…, 𝜋0𝑖𝐺−1)⊤,1≤𝑖≤𝑘. On the other hand, 𝐻0in (8) is true if and only if 𝐻0,𝒂∶𝒂⊤(𝝕−𝝕0)=0 is true ∀𝒂= (𝒂⊤ 1,…,𝒂⊤ 𝑘)⊤∈R𝑘(𝐺−1) ⧵{0}, where 𝝕= (𝝕⊤ 1,…,𝝕⊤ 𝑘)⊤,𝝕0= (𝝕⊤ 01,…,𝝕⊤ 0𝑘)⊤and 𝒂𝑖∈R𝐺−1,1≤𝑖≤𝑘. Let 𝝕= ( 𝝕⊤ 1,…, 𝝕⊤ 𝑘)⊤, with 𝝕𝑖= (𝑛𝑖1,…, 𝑛𝑖𝐺−1)⊤∕𝑛𝑖.,1≤𝑖≤𝑘. As E( 𝝕) = 𝝕, we have that 𝒂⊤( 𝝕−𝝕0)is unbiased for 𝒂⊤(𝝕−𝝕0). Let 𝛴0𝑖be the (𝐺−1)×(𝐺−1)-matrix with diagonal entries 𝜋0𝑖𝑔(1−𝜋0𝑖𝑔),1≤𝑔≤𝐺−1, and off-diagonal entries −𝜋0𝑖𝑔𝜋0𝑖𝑣,1≤𝑔≠𝑣≤𝐺−1. The null variance of 𝒂⊤( 𝝕−𝝕0)is 𝜎2 0𝒂∶= V0{𝒂⊤( 𝝕−𝝕0)}= 𝑘 ∑ 𝑖=1 𝒂⊤ 𝑖𝛴0𝑖𝒂𝑖∕𝑛𝑖., and its null expectation is 0. So, it is reasonable to reject 𝐻0,𝒂for large values of 𝒂={𝒂⊤( 𝝕−𝝕0)}2∕𝜎2 0𝒂. However, we are not interested in testing 𝐻0,𝒂for a particular 𝒂, but for all 𝒂≠0. According to the union-intersection principle, the test statistic is = sup 𝒂≠0 𝒂. Let 𝛴0be the block diagonal matrix with diagonal blocks 𝛴01∕𝑛1.,…, 𝛴0𝑘∕𝑛𝑘..is the largest eigenvalue of the matrix = ( 𝝕−𝝕0)( 𝝕−𝝕0)⊤𝛴−1 0(see, e.g. [20]). As has range one, it readily follows that = ( 𝝕−𝝕0)⊤𝛴−1 0( 𝝕−𝝕0) = 𝑘 ∑ 𝑖=1 ( 𝝕𝑖−𝝕0𝑖)⊤𝛴−1 0𝑖( 𝝕𝑖−𝝕0𝑖). In [4] it is shown that ( 𝝕𝑖−𝝕0𝑖)⊤𝛴−1 0𝑖( 𝝕𝑖−𝝕0𝑖) = 𝜒2 𝑖, for each 1≤𝑖≤𝑘, where 𝜒2 𝑖is as defined in (1). Therefore, =∑𝑘 𝑖=1 𝜒2 𝑖. The null distribution of can be approximated using the result in Theorem 1. In this way one obtains the test in (3). A good point of the union-intersection methodology is that it allows building simultaneous confidence intervals. From the above developments, for large 𝑘, 1 − 𝛼≃P0(𝑇𝑘≤𝑧1−𝛼)
Mathematics and Computers in Simulation 223 (2024) 588–600 593 M.V. Alba-Fernández et al. Table 3 Estimated type I errors for the nominal value 𝛼= 0.05. 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 𝑘 𝑛𝑖. lev1 lev2 lev3 lev4 lev5 lev6 50 10 5.3 5.8 6.0 5.2 5.5 5.5 500 10 5.2 5.2 5.4 5.1 5.5 5.1 20 5.5 5.5 5.6 5.4 5.5 5.5 20 5.0 5.1 4.9 5.1 5.3 5.4 30 5.0 5.7 5.6 5.4 5.6 5.3 30 5.1 5.0 5.1 5.1 5.0 5.1 100 10 5.4 5.5 5.7 5.1 5.6 5.1 1000 10 5.4 5.0 5.1 4.8 5.1 5.5 20 5.6 5.5 5.7 5.3 5.0 5.2 20 5.0 5.3 5.2 5.3 5.4 5.2 30 5.3 5.7 5.9 5.3 5.2 5.8 30 4.8 5.0 5.5 5.4 5.2 5.0 200 10 5.3 5.6 5.3 5.0 5.4 5.3 2000 10 4.8 5.2 5.1 5.2 5.0 4.8 20 6.1 5.1 5.2 5.0 5.2 5.3 20 5.1 5.2 5.2 5.5 5.1 5.3 30 5.2 5.3 5.1 5.0 5.0 5.2 30 4.8 5.2 4.6 5.0 5.3 4.8 =P0( ≤ 𝑘(𝐺− 1) + √𝑘𝜍0𝑘𝑧1−𝛼) =P0(|||𝒂⊤( 𝝕−𝝕0)|||≤𝜎0𝒂√𝑘(𝐺− 1) + √𝑘𝜍0𝑘𝑧1−𝛼,∀𝒂∈R𝑘(𝐺−1)). Assume for simplicity that 𝑛1.=⋯=𝑛𝑘. =𝑛. As the length of the above intervals is of order √𝑘(𝐺− 1)∕𝑛, they are not useful in practical applications, because they are too wide, except for very large 𝑛. 3. Simulation results 3.1. Level The test that rejects 𝐻0according to (3) (or equivalently (4)) has asymptotic level 𝛼. To numerically check the actual level of the test for a small or moderate number of groups, we repeated the experiment in the Introduction. That is, in each case (lev1, …, lev6), 𝑘random samples of size 𝑛𝑖. obeying the null hypothesis were generated; then the test statistic 𝑇𝑘and its asymptotic 𝑝-value, 𝑎𝑝, were calculated; the whole experiment was repeated 10,000 times. Table 3 displays the fraction of asymptotic 𝑝-values less than or equal to 𝛼= 0.05, which is the estimated probability of type I error, expressed in percentages. Looking at this table, we see that, in all cases, the level is quite close to the nominal value. Moreover, the values are closer to 𝛼than those in Table 2, and with the new test, the user does not have to calculate critical points, which takes a long time and discourages its usage. 3.2. Power To study the power, we repeated the above experiment with 100𝑝%of the samples coming from an alternative distribution (𝝅𝑖≠𝝅0𝑖), and the remaining 100(1 − 𝑝)% of the samples generated from the null (𝝅𝑖=𝝅0𝑖). Specifically, we tried the following cases: pw1. 𝐺= 5,𝐻0∶𝝅0𝑖= (0.2,…,0.2)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.1,0.2,0.2,0.2,0.3)⊤, pw2. 𝐺= 5,𝐻0∶𝝅0𝑖= (0.2,…,0.2)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.1,0.1,0.2,0.2,0.4)⊤, pw3. 𝐺= 10,𝐻0∶𝝅0𝑖= (0.1,…,0.1)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.1,0.1,0.2)⊤, pw4. 𝐺= 10,𝐻0∶𝝅0𝑖= (0.1,…,0.1)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.05,0.05,0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.3)⊤, pw5. 𝐺= 5,𝐻0∶𝝅0𝑖= (0.1,0.2,0.2,0.2,0.3)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.2,…,0.2)⊤, pw6. 𝐺= 5,𝐻0∶𝝅0𝑖= (0.05,0.1,0.2,0.3,0.35)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.1,0.2,0.2,0.2,0.3)⊤, pw7. 𝐺= 10,𝐻0∶𝝅0𝑖= (0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.1,0.1,0.2)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.1,…,0.1)⊤, pw8. 𝐺= 10,𝐻0∶𝝅0𝑖= (0.05,0.05,0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.3)⊤,1≤𝑖≤𝑘, data: 100(1 − 𝑝)% of groups obey the null, and the rest 100𝑝%of groups have 𝝅= (0.05,0.05,0.05,0.1,0.1,0.1,0.1,0.1,0.15,0.2)⊤, for 𝑝= 0.1,0.2,0.3,𝑘= 50,100,200,500,1000,𝑛𝑖. = 10,20,30,1≤𝑖≤𝑘, in all cases. To facilitate the comparison between the power of our test and that of the BH method, we report power results corrected for size as proposed by [13]. Each of the size-corrected rejection percentages 𝑝𝑤 reported in Tables 4 and 5was calculated as follows: 𝑝𝑤 =𝛷(𝛷−1( 𝑝𝑤) − 𝛷−1(𝛼) + 𝛷−1(𝛼)), where 𝛼= 0.05 is the nominal size of the test, 𝛼 is the empirical size for the corresponding configuration as reported in Tables 1 and 3, and 𝑝𝑤 is the observed, uncorrected power (not reported). Tables 4 and 5display the results for the test proposed in this paper in the first line, and the results for the BH method with exact 𝑝-values in the second line. From the results shown in these tables, it can be highlighted that the estimated power of the proposed
Mathematics and Computers in Simulation 223 (2024) 588–600 594 M.V. Alba-Fernández et al. Table 4 Adjusted estimated powers for the nominal level 𝛼= 0.05. In each case, the first line displays the results for the new proposal, and the second line displays the results when using the BH method with exact 𝑝-values. 𝑘 𝑛𝑖. pw1 pw2 pw3 pw4 𝑝= 0.10.2 0.3 0.1 0.2 0.3 0.1 0.2 0.3 0.1 0.2 0.3 50 10 8.4 12.9 18.0 17.5 40.0 63.9 8.1 13.3 18.4 21.4 45.0 65.2 6.3 7.8 9.1 13.7 22.8 31.7 6.3 8.6 11.0 20.4 34.3 46.7 20 12.1 24.7 40.1 40.4 82.5 97.5 12.7 24.5 40.6 47.2 87.1 98.6 9.1 12.9 15.8 37.3 58.4 73.6 11.3 16.6 22.5 53.7 76.3 88.4 30 19.6 42.9 68.2 66.3 98.2 100 19.0 41.0 65.8 72.7 98.8 100 13.3 20.4 26.5 63.9 88.1 95.2 17.4 27.1 36.7 80.8 96.2 99.2 100 10 9.5 15.8 25.5 25.6 59.4 86.3 9.2 16.9 25.8 29.6 66.0 90.4 6.3 7.7 9.3 15.9 26.7 36.0 6.8 8.9 11.5 23.0 40.1 53.8 20 16.2 36.5 61.5 61.0 97.6 100 17.6 38.7 61.3 68.7 98.6 100 9.1 14.0 18.3 45.9 70.3 84.5 13.1 20.6 27.5 67.8 89.7 96.5 30 26.9 64.4 89.3 87.8 100 100 27.1 63.2 89.1 92.6 100 100 15.8 26.0 32.8 79.1 95.4 99.0 20.3 34.0 45.1 92.2 99.4 99.9 200 10 12.0 23.8 40.0 39.9 84.4 98.6 12.6 25.5 41.9 48.0 89.9 99.4 6.7 8.9 10.7 20.9 34.5 47.9 8.3 10.9 14.0 30.8 53.1 70.6 20 23.0 56.2 85.2 85.1 100 100 26.1 60.5 87.2 91.0 100 100 10.9 17.1 22.6 60.3 83.0 93.2 14.3 22.9 30.7 80.1 95.9 99.2 30 42.7 87.5 99.2 99.1 100 100 44.1 87.9 99.1 99.6 100 100 17.1 26.7 35.7 89.1 98.6 99.8 24.8 40.7 54.0 98.3 100 100 500 10 17.8 43.6 71.4 70.1 99.4 100 17.9 43.4 70.1 77.4 99.9 100 6.6 9.4 10.7 22.8 37.4 50.3 10.0 13.5 17.5 46.5 68.8 82.5 20 45.4 91.1 99.7 99.6 100 100 44.9 90.7 99.6 99.8 100 100 12.3 19.7 26.3 74.8 93.3 98.3 18.2 29.9 40.4 94.3 99.6 100 30 74.0 99.6 100 100 100 100 73.7 99.8 100 100 100 100 20.3 33.3 44.4 97.3 99.9 100 33.0 53.2 68.0 99.9 100 100 1000 10 27.0 65.7 92.5 92.1 100 100 28.4 68.6 92.5 96.1 100 100 7.8 9.2 12.4 27.8 43.6 55.3 10.6 15.8 20.3 58.0 82.1 92.7 20 68.9 99.4 100 100 100 100 68.3 99.3 100 100 100 100 13.7 21.9 30.3 85.2 97.9 99.7 22.5 36.2 48.8 98.6 100 100 30 94.4 100 100 100 100 100 93.3 100 100 100 100 100 24.4 40.7 51.3 99.5 100 100 39.7 63.0 77.5 100 100 100 test increases with the sample size, the proportion of groups with non-null distribution, the number of groups, and the distance between the probability vector in the null hypothesis and the alternative probability vector generating the non-null data (observe, for example, that the power in cases pw2 and pw4 is larger than in cases pw1 and pw3, respectively). Those results agree with the theoretical developments done for the power in Section 2. In almost all cases, the new method exhibits larger power than the BH method. There are just two instances (pw4, 𝑘= 50,𝑝= 0.1,𝑛𝑖. = 20,30,inTable 4) in which the BH method has larger power than the new test. 4. Real data set applications 4.1. Public opinion poll on health care The ‘‘Public opinion poll on health care’’ is an annual opinion poll that has been conducted since 1993 in Spain. It is carried out jointly by the Ministry of Health of the Spanish Government and the Centre for Sociological Research (CIS) (https://www.cis.es/ cis/opencms/ES/8_cis/quienessomos/). The annual survey is based on around 7800 interviews with people aged 18 and over in all Spanish provinces (free data at https://www.cis.es/cis/opencms/ES/2_bancodatos/). The reference period is usually from March to October or November in each year. During the Pandemic COVID-19 (2020/21) it was not possible to visit households, and in 2022 it was carried out by telephone interviews. We focus on studying whether the assessment of certain health services in Spain has been affected by the COVID-19 pandemic. The age of citizens and the province in which they live may influence the perception of the quality of the services received. Combining these two variables, the province of residence, with 50 provinces and 2 autonomous cities, and age, categorized into 6 groups (18–24, 25–34, 35–44, 45–64, 55–64, and over 65), 312 populations were taken into account. The health care services under investigation were ‘‘Emergencies in public hospitals’’, ‘‘Emergencies 061-112’’ and ‘‘Admission and assistance in public hospitals’’. These variables were categorized into 4 levels of satisfaction: 1 means very unsatisfied, 2 means unsatisfied, 3 means satisfied and 4 means very satisfied. Thus, our aim is to check if the observed percentages of each satisfaction score in the 2022 pool match a certain fixed vector of proportions in all populations, that we took based on those observed before the Pandemic. In order to justify our choice for 𝝅01,…,𝝅0𝑘, we investigated whether the ratings obtained in the pools carried out in 2019 and 2018 gave the same results in all groups, as in such a case it is reasonable to consider them as fixed when testing changes with respect to the 2022 values.
Mathematics and Computers in Simulation 223 (2024) 588–600 595 M.V. Alba-Fernández et al. Table 5 Adjusted estimated powers for the nominal level 𝛼= 0.05. In each case, the first line displays the results for the new proposal, and the second line displays the results when using the BH method with exact 𝑝-values. 𝑘 𝑛𝑖. pw5 pw6 pw7 pw8 𝑝= 0.10.2 0.3 0.1 0.2 0.3 0.1 0.2 0.3 0.1 0.2 0.3 50 10 13.0 25.1 38.4 9.7 16.8 23.9 13.2 24.9 40.3 9.9 17.7 26.0 9.3 14.7 18.6 9.4 14.1 17.7 11.1 17.0 22.5 6.9 9.5 12.2 20 20.8 45.6 69.8 13.5 25.1 37.7 17.9 40.6 64.9 12.9 26.5 44.8 19.3 32.2 42.8 13.4 20.0 25.8 16.8 27.2 36.9 10.8 15.0 19.8 30 29.5 64.9 88.2 16.5 33.6 51.7 24.8 57.1 82.7 17.8 40.5 62.7 29.8 47.8 61.1 19.2 29.3 39.1 24.1 40.7 51.0 14.7 22.9 30.4 100 10 17.3 37.3 59.1 12.1 23.1 37.6 16.0 36.8 60.5 12.7 24.9 41.6 11.9 17.8 24.0 10.0 13.8 18.7 13.2 20.1 26.6 7.5 10.1 12.7 20 30.4 67.4 90.4 16.5 36.1 57.4 29.4 62.8 88.0 18.8 42.3 68.7 23.0 39.0 51.9 14.4 24.8 33.4 20.5 32.2 43.8 11.8 17.7 24.7 30 45.3 87.5 98.6 22.0 49.3 73.7 38.9 81.8 97.7 24.8 59.3 85.8 38.5 56.5 73.7 19.5 33.1 42.9 30.7 48.5 61.6 17.2 26.9 35.9 200 10 24.1 55.7 82.9 16.9 35.9 57.2 23.9 57.6 84.1 16.9 37.9 62.7 14.3 23.1 32.6 10.4 14.9 20.3 13.4 21.7 29.3 8.1 11.2 13.5 20 48.2 90.4 99.5 25.4 56.0 82.2 42.2 86.4 98.7 27.7 64.6 90.3 28.4 46.2 59.7 17.6 28.4 38.1 23.6 38.6 50.7 12.2 19.2 26.2 30 69.3 98.9 100 34.4 74.5 93.9 60.4 97.5 100 40.7 84.4 98.7 46.0 69.1 82.2 23.6 38.9 51.7 38.0 59.5 73.9 20.8 33.4 43.6 500 10 44.2 87.5 99.2 27.4 61.9 87.4 43.7 88.4 99.5 27.8 68.2 92.5 11.7 20.7 27.8 13.4 21.2 28.0 17.8 28.3 38.0 8.6 13.0 16.0 20 79.0 99.8 100 43.9 87.3 98.9 71.5 99.6 100 49.3 93.4 99.9 34.5 56.0 71.9 21.4 35.7 48.2 30.4 48.8 63.4 15.2 25.3 32.3 30 95.3 100 100 59.7 96.7 99.9 90.3 100 100 69.3 99.6 100 59.6 82.2 92.8 36.9 56.2 70.1 49.8 73.8 86.2 23.3 38.8 49.7 1000 10 68.0 98.9 100 42.3 86.6 98.9 66.2 99.2 100 43.2 89.7 99.8 17.8 30.3 40.3 14.7 22.9 30.3 21.3 33.9 45.0 8.9 13.6 16.2 20 96.2 100 100 65.5 98.5 100 92.5 100 100 73.8 99.7 100 44.9 68.2 81.4 26.5 42.9 55.8 37.0 57.4 72.7 17.4 28.3 37.0 30 99.9 100 100 82.6 99.9 100 57.4 100 100 92.7 100 100 71.4 91.2 97.2 39.2 63.3 76.2 58.1 82.2 92.3 28.1 45.4 58.6 Table 6 Results of the tests for the equality of satisfaction levels between 2018 and 2019. Variable 𝑘2018 total sample size 2019 total sample size 𝑝-value Emergencies in public hospitals 259 7190 7146 0.3517 Emergencies 061-112 222 4916 4853 0.1930 Admission and assistance in public hospitals 251 6859 6773 0.1550 Thus, we began by testing for the equality of results between 2019 and 2018. The reader will guess that this problem is equivalent to testing a large number, say 𝑘, of multinomial two-sample problems, 𝑘being the number of subpopulations and the multinomial distributions to be compared are those derived from the 4 categories of satisfaction registered. A test for this problem has been proposed by [3]. Following their suggestions on the sample sizes to achieve the nominal level 𝛼= 0.05, all 312 populations were screened to ensure 𝑛𝑖. ≥5in both years. Those subpopulations not meeting such requirements were not included. Table 6 shows a summary of the application of the test statistic 𝑇1in [3] to test whether (or not) the Spanish population has the same opinion on the three health services under consideration in 2018 and 2019. According to the results shown in Table 6, there is no evidence of changes between the two years before the Pandemic on the considered variables. Thus, we took the observed pooled frequencies as the fixed proportions 𝝅01,…,𝝅0𝑘in the null hypothesis. In each health service, the starting point is the set of populations considered in the previous test (column 2, Table 6). However, additional filtering has been done: first, we eliminated those populations 𝑖with 𝜋0𝑖𝑔 = 0, and 𝑛𝑖𝑔 >0, for some 1≤𝑔≤𝐺, as in those cases where clearly 𝝅0𝑖≠𝝅𝑖, and second, we only took those populations 𝑖with 𝑛𝑖. ≥5in 2022. The final number of subpopulations considered is shown in column 2 of Table 7. Before applying the proposal, the empirical level under the null is evaluated for each health service. Thus, for each variable, 𝑘 random samples of size 𝑛1.,…, 𝑛𝑘. were generated from multinomial laws with probability vectors 𝝅01,…,𝝅0𝑘. After calculating 𝑇𝑘, its asymptotic 𝑝-value was obtained. This experiment was repeated 10,000 times and the fraction of asymptotic 𝑝-values less than or equal to 𝛼= 0.05 is collected (column 4, Table 7). In all cases, the empirical levels stick to the nominal level. In the same way, 𝜒2 𝑖was calculated for each population and its 𝑝-value was computed using the asymptotic null distribution. Column 5 of Table 7 shows the percentage of rejections (at level 5%) using the BH method, giving rather liberal results. Notice that, in this case, the calculation of the exact 𝑝-values for the BH procedure is computationally too expensive and not feasible.
Mathematics and Computers in Simulation 223 (2024) 588–600 596 M.V. Alba-Fernández et al. Table 7 Applications of the goodness-of-fit test for the satisfaction levels before and after the Pandemic. Variable 𝑘2022 total Estimated level BH estimated level 𝑇𝑘𝑝-value sample size (𝛼= 0.05) (𝛼= 0.05) Emergencies in public hospitals 229 6649 0.0509 0.1074 25.8871 0 Emergencies 061-112 164 3981 0.0507 0.1763 80.8428 0 Admission and assistance in public hospitals 163 5095 0.0534 0.1759 109.2739 0 The right side of Table 7 contains the values of the test statistics 𝑇𝑘and the associated 𝑝-values according to (4). Looking at this table, the subjective perceptions of people in Spain about the three health services have changed noticeably. 4.2. Risk preferences Risk and uncertainty are at the core of economic decisions. Recently, economic literature has particularly focused attention on measurement, determinants, and stability of risk preferences in the population [6]. In this application, risk preferences are assessed through a series of experiments carried out at the University of Jaén (Spain). Specifically, individual risk attitudes are measured using a paid lottery-choice method with high external validity, a slightly modified version of the instrument designed by [8]. As a result, subjects are classified into four categories: ‘‘Risk loving’’, ‘‘Risk neutral’’, ‘‘Risk averse’’ and ‘‘Inconsistent’’. The categorization of risk attitudes was first obtained in 2018 in an experiment conducted at the Experimental Economics Laboratory of the University of Jaén (UJAEEconLab, https://ujaeconlab.com/). In this dataset, the whole population of students was divided into groups according to their academic profile (Engineering, Science, Business, Law, Humanities), gender, and age. Four categories were considered for age (≤18,19 − 20,21 − 22 and ≥23). Of the 40 populations, only 23 provided non-degenerate probability vectors and, thereby, were taken as fixed 𝝅01,…,𝝅0𝑘. To control the potential impact of individual risk preferences in future experiments, it is of great interest to test whether the observed proportions of each type of risk attitudes match the given ones. In this regard, new experimental sessions were carried out from February to October of 2023 to test 𝐻0, taking the values previously stated in 2018 as the ideal ones. The sample of 2023 was divided according to the same 23 subpopulations. However, additional filtering was done: first, we eliminated those populations 𝑖with 𝜋0𝑖𝑔 = 0, and 𝑛𝑖𝑔 >0, for some 1≤𝑔≤𝐺, and second, we only took those populations with 𝑛𝑖. ≥2. The final number of populations was 11, involving a total of 80 observations. Again, the empirical level is evaluated from 10,000 samples generated under the null hypothesis as in Table 3. The BH method with asymptotic 𝑝-values was also applied as in Table 1. The estimated type I error for the nominal value 𝛼= 0.05 was 0.0629 and 0.0847, respectively. Although the level of the new test is a bit liberal, it is still reasonably close to the target value. The application of the proposed test gives a value for 𝑇𝑘= −1.4827 and a 𝑝-value equal to 0.9309. According to these results, there is no evidence to reject the equality of the probabilities of each risk category between 2023 and 2018. The BH method with asymptotic 𝑝-values was also applied. No hypothesis was rejected. 5. Conclusions This paper deals with simultaneously testing if the observed proportions agree with some fixed values when the target population is divided into a large number of subpopulations or groups. A new procedure has been proposed. It has been studied both theoretically and numerically. It is worth mentioning that the test is very easy to implement and does not require the use of complicated resampling methods to obtain the 𝑝-value. It applies whenever the sample sizes remain bounded or increase with 𝑘. These advantages of the proposal make it an attractive option to be considered from a practical point of view. 6. Proofs Before proving the results in Section 2, we first give a preliminary result. Lemma 5. We have that E(𝜒2 𝑖) = 𝐺 ∑ 𝑔=1 𝜋𝑖𝑔(1 − 𝜋𝑖𝑔) 𝜋0𝑖𝑔 +𝑛𝑖. 𝐺 ∑ 𝑔=1 (𝜋𝑖𝑔 −𝜋0𝑖𝑔)2 𝜋0𝑖𝑔 , V(𝜒2 𝑖) = 2 𝐺 ∑ 𝑔=1 𝜋2 𝑖𝑔(1 − 2𝜋𝑖𝑔) 𝜋2 0𝑖𝑔 + 2 (𝐺 ∑ 𝑔=1 𝜋2 𝑖𝑔 𝜋0𝑖𝑔 )2 +1 𝑛[𝐺 ∑ 𝑔=1 𝜋𝑖𝑔(8𝜋2 𝑖𝑔 − 6𝜋𝑖𝑔 + 1) 𝜋2 0𝑖𝑔 −6 (𝐺 ∑ 𝑔=1 𝜋2 𝑖𝑔 𝜋0𝑖𝑔 )2 + 4 (𝐺 ∑ 𝑔=1 𝜋2 𝑖𝑔 𝜋0𝑖𝑔 )(𝐺 ∑ 𝑔=1 𝜋𝑖𝑔 𝜋0𝑖𝑔 )−(𝐺 ∑ 𝑔=1 𝜋𝑖𝑔 𝜋0𝑖𝑔 )2⎤⎥⎥⎦