scieee AI-readable full text Open interactive document viewer

Unequal probability sampling from a finite population: A multicriteria approach

Carrizosa Priego, Emilio José

Abstract

In this note we address the problem of determining selection probabilities for multipurpose surveys, when the aim is the simultaneous minimization of variances for each variable under study. A characterization of the set of Pareto-optimal designs is given for designs with replacement and also for a class of designs without replacement, namely, Poisson designs. As an application, we describe a problem encountered in Auditing, where both the fraction of misstatements, and the average amount of such misstatements are of interest.

Full text

Decision Support Unequal probability sampling from a finite population: A multicriteria approach Emilio Carrizosa * Facultad de Matemáticas, Universidad de Sevilla, Avda Reina Mercedes s/n, 41012 Sevilla, Spain article info Article history: Received 6 October 2007 Accepted 22 February 2009 Available online 21 March 2009 Keywords: Multiple objective programming Auditing Hansen Hurwitz estimator abstract In this note we address the problem of determining selection probabilities for multipurpose surveys, when the aim is the simultaneous minimization of variances for each variable under study. A characterization of the set of Pareto-optimal designs is given for designs with replacement and also for a class of designs without replacement, namely, Poisson designs. As an application, we describe a problem encountered in Auditing, where both the fraction of misstatements, and the average amount of such misstatements are of interest. Ó2009 Elsevier B.V. All rights reserved. 1. The problem Let U¼fu 1 ;...;u N gbe a finite population. Associated with each u i we have an r-valued vector Y i ¼ðY i1 ;Y i2 ;...;Y ir Þ:The vector #of means, #¼ð# 1 ;...;# r Þ¼ 1 NX N i¼1 Y i1 ;1 NX N i¼1 Y i2 ;...;1 NX N i¼1 Y ir ! ð1Þ is estimated by drawing from Ua sample with replacement of size n, and considering as estimator b #the r-dimensional Hansen–Hurwitz estimator, [9],b #¼ð c # 1 ;...;b # r Þwith b # j ¼1 NX N i¼1 Y ij f i n a i ;j¼1;2;...;r:ð2Þ Here f i denotes the frequency of u i in the sample, and a i is the decision variable denoting the probability of selection of u i at each draw. Since ða 1 ;...;a N Þrepresents a probability vector, it is constrained to belong to the unit simplex D N , D N ¼ðb 1 ;...;b N Þ:X N j¼1 b j ¼1;06b j 618j¼1;2;...;N () :ð3Þ In this paper the Y ij are assumed to be mutually independent random variables, with expected value EðY ij Þ¼ l ij <þ1 and finite variance v arðY ij Þ¼EðY 2 ij Þ l 2 ij ¼ r 2 ij P0:For simplicity we assume in what follows: l 2 ij þ r 2 ij >08i;j;ð4Þ which holds when each Y ij is not degenerate to zero. Observe also that the case of a fixed positive value for Y ij is obtained by assuming that r ij ¼0: When the coefficients Y ij are given, then each b # j is unbiased for # j , E1 NX N i¼1 Y ij f i n a i ðY 1j ;...;Y Nj Þ ! ¼1 NX N i¼1 Y ij ð5Þ and the (design) variance of b # j is given by E1 NX N i¼1 Y ij 1 NX N i¼1 Y ij f i n a i ! 2 0 @ðY 1j ;...;Y Nj Þ 0 @1 A ¼X N i¼1 Y 2 ij N 2 n a i 1 N 2 nX N i¼1 Y ij ! 2 ;ð6Þ see p. 52 of [19]. For a given vector ð a 1 ;...; a N Þ2 D N of selection probabilities, let e j ð a 1 ;...; a N Þdenote the expected squared error in the j-th component if the sample is drawn with probabilities a j at each stage, and we consider as random variables (under the above-mentioned assumptions) Y ij and f i as well, i.e., the expectation is computed with respect to both the sampling design and the distribution of the variables Y ij : e j ð a 1 ;...; a N Þ¼E1 NX N i¼1 Y ij 1 NX N i¼1 Y ij f i n a i ! 2 0 @1 A:ð7Þ By (6), and, since, for fixed j, variables Y ij are assumed to be mutually independent, one has that 0377-2217/$ - see front matter Ó2009 Elsevier B.V. All rights reserved. doi:10.1016/j.ejor.2009.02.036 *Tel.: +34 954557943; fax: +34 954622800. E-mail address: [email protected] European Journal of Operational Research 201 (2010) 500–504 Contents lists available at ScienceDirect European Journal of Operational Research journal homepage: www.elsevier.com/locate/ejor e j ð a 1 ;...; a N Þ¼EE 1 NX N i¼1 Y ij 1 NX N i¼1 Y ij f i n a i ! 2 0 @1 AðY 1j ;...;Y Nj Þ 0 @1 A ¼X N i¼1 EðY 2 ij Þ N 2 n a i 1 N 2 nEX N i¼1 Y ij ! 2 0 @1 A ¼X N i¼1 r 2 ij þ l 2 ij N 2 n a i 1 N 2 nX N i¼1 r 2 ij þX N i¼1 l ij ! 2 0 @1 A: ð8Þ One seeks a probability vector ða 1 ;...;a N Þ2 D N making small all errors e j given in (7). In other words, one faces the multiple-objective problem of simultaneous minimization of all errors, when the probabilities a j are the decision variables. Observe that, by (8), those designs with some a i equal to 0 yield an infinite value for each e j , and thus will be automatically discarded for optimality. A first approach to address such multiple-objective problem would consist of reducing the problem to one single-objective optimization problem, as largely discussed in the literature of Multiple-Objective Optimization. In this sense, one error measure could be minimized imposing that the remaining errors are below specified threshold values, as suggested e.g. in [3,5,6,12,14]. This yields a problem of the form min e l ð a 1 ;...; a N Þ s:t e j ð a 1 ;...; a N Þ6b j 8j–l; ð a 1 ;...; a N Þ2 D N ð9Þ for given bounds b j , where D N is defined in (3). Alternatively, one can minimize a weighted sum of the errors. In other words, an optimization problem of the form min X r j¼1 x j e j ð a 1 ;...; a N Þ s:tð a 1 ;...; a N Þ2 D N ð10Þ for given non-negative weights x 1 ;...;x r , is considered. See e.g. [5,6]. Another option, e.g. [11], might be to minimize the highest error, i.e., to solve the optimization problem min max r j¼1 e j ð a 1 ;...; a N Þ s:tð a 1 ;...; a N Þ2 D N : ð11Þ Instead of using scalar problems as those mentioned above, we follow a multicriteria approach, and seek selection probabilities a 1 ;...; a N minimizing simultaneously the rcriteria. In other words, we consider the nonlinear multiple-objective optimization problem min e 1 ð a 1 ;...; a N Þ; e 2 ð a 1 ;...; a N Þ;...; e r ð a 1 ;...; a N ÞðÞ s:tð a 1 ;...; a N Þ2 D N ð12Þ and we seek the set Pof Pareto-optimal solutions to (12). We recall that ð a  1 ;...; a  N Þ2 D N is said to be Pareto-optimal if there exists no ð a 1 ;...; a N Þ2 D N satisfying e j ð a 1 ;...; a N Þ6 e j ð a  1 ;...; a  N Þ;j¼1;2;...;r;ð13Þ with at least one inequality strict. See e.g. [3,7,11,13,20,21] for other multiple-objective design problems and [4] for an introduction. 2. Results 2.1. Sampling with replacement Define, for each j¼1;...;N, the scalars k j ; c j as k j ¼1 N 2 n; c j ¼ 1 N 2 nX N i¼1 r 2 ij þX N i¼1 l ij ! 2 0 @1 A: ð14Þ By construction, each k j is strictly positive. Moreover, (8) implies that e j ð a 1 ;...; a N Þ¼k j X N i¼1 r 2 ij þ l 2 ij a i þ c j :ð15Þ This simple form of the objectives involved enables us to characterize P, as shown below. Proposition 1. Optimal solutions to (9)–(11) are elements of P: Proof. By (14), each e j is a strictly convex function, and thus the optimization problems (9)–(11), with objectives monotonic in the errors e j , have a unique optimal solution, which is then Pareto-optimal, [4].h We now give a characterization of Pareto-optimality. Proposition 2. Given ð a  1 ;...; a  N Þ2 D N ;the following statements are equivalent. 1. ð a  1 ;...; a  N Þ2P: 2. There exists ð x 1 ; x 2 ;...; x r Þ;P r j¼1 x j ¼1; x j P0 8 j, such that a  i ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi P r j¼1 x j r 2 ij þ l 2 ij  r P N l¼1 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi P r j¼1 x j r 2 lj þ l 2 lj  r;i¼1;2;...;N:ð16Þ Proof. Consider a scalarized version of (12) in the form (10) for some non-negative x 1 ;...; x r ,P r j¼1 x j ¼1:Problem (10) can be analytically solved. Indeed, dropping the nonnegativity constraint, we have a convex problem with one linear constraint. Necessary and sufficient optimality conditions are given by taking Lagrange multipliers, yielding the ð a  1 ;...; a  N Þin (16) as unique optimal solution, which also satisfies the non-negativity constraints. In other words, such ð a  1 ;...; a  N Þis the unique optimal solution to (10). See [16] for a similar result on a related problem. This implies in particular that, for each ð x 1 ;...; x r Þ, the optimal solution to (10) is Pareto-optimal. Conversely, the objectives e j are strictly convex, and thus any Pareto-optimal solution to (12) is optimal solution to some problem of type (10), e.g. [4]. Hence, the set Pcoincides with the set of solutions of the form (16), as asserted. h This result implies that, if the unit simplex is discretized by generating a dense enough grid of points x , and for each such x the corresponding a from (16) is constructed, then one obtains a finite set of Pareto-optimal solutions which approximates accurately P: In the bivariate case, one can plot the trade-off curve for e 1 ; e 2 ,as in Fig. 1. 2.2. Application in auditing As an application, consider an Auditing problem, in which one has a population Uof Nitems to be audited. For each item u i , the reported book value X i >0 is known, and two parameters are of main interest, [8]: Fraction of misstated items; Average amount of misstatement per item. E. Carrizosa / European Journal of Operational Research 201 (2010) 500–504 501 PPS sampling, frequently referred in this context as Dollar-Unit Sampling, amounts to choosing individuals with probabilities proportional to their book value X i , and seems to be specially suited for the second objective, whereas a simple random sampling (SRS) might be more convenient for the first criterion, unless the likelihood of misstatement can be related with the book value. Whether a PPS design, a SRS, or a different design ‘‘in between” these two extreme designs is implemented, may be left to the analyst. For instance, in an real Auditing problem the author has recently been involved on European Regional Development Funds, the Commission Regulation (EC) 1828/2006, [17], states in its Article 17 that The method used to select the sample and to draw conclusions from the results shall take account of internationally accepted audit standards and be documented. Having regard to the amount of expenditure, the number and type of operations and other relevant factors, the audit authority shall determine the appropriate statistical sampling method to apply. Hence, ample room exists to decide which sampling design should be used. We consider here the problem of describing the Pareto-optimal designs when two criteria are considered: minimization of the variance of the Hansen–Hurwitz estimators of the fraction of misstated items, and the average amount of misstatement per item. Misstatement of item u i is modeled as a Bernoulli variable Y i1 , with common success probability l i1 ¼ l , and thus common variance r 2 i1 ¼ l ð1 l Þ: Misstatement amount in a misstated unit with book value xis assumed to be a random variate, with second moment u ðxÞ:For instance, one may impose the second moment u ðxÞto be proportional to x 2 , u ðxÞ¼ s x 2 ð17Þ for a given s which, in practice, should be estimated from a test sample m 0 of misstated units, via, for instance, a ratio estimator, b s ¼P k2m 0 y 2 k P k2m 0 x 2 k ;ð18Þ where the pairs ðx k ;y k Þrepresent the book value and misstatement in each unit in m 0 : Define Y i2 as the amount of misstatement of unit u i :We have that EðY 2 i2 jY i1 ¼1Þ¼ u ðX i Þ;ð19Þ thus EðY 2 i2 Þ¼ l EðY 2 i2 jY i1 ¼1Þ¼ lu ðX i Þ:ð20Þ By Proposition 2, a vector ða  1 ;...;a  N Þis Pareto-optimal iff it is of the form (16), which in this setting amounts to saying that there exists x;06x61, such that, for each i¼1;2;...;N,a  i has the form a  i ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ð1 x Þ l þ xlu ðX i Þ p P N l¼1 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ð1 x Þ l þ xlu ðX l Þ p ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 x þ xu ðX i Þ p P N l¼1 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 x þ xu ðX l Þ p: ð21Þ The two extreme cases for the parameter x , namely x ¼0 and x ¼1, model the situations in which absolute priority is given, respectively, to the analysis of the fraction of misstated units and the average amount of misstatement per unit, and they yield, for each i, a  i ¼ 1 N and a  i /ffiffiffiffiffiffiffiffiffiffiffiffiffiffi u ðX i Þ; pi.e., a  i /ffiffiffiffiffiffiffiffi s X 2 i q/X i under (17).In other words, the extreme cases of the trade-off parameter x yield two well-known sampling schemes, namely, SRS and Dollar-Unit Sampling, which are shown to be Pareto-optimal by Proposition 2. Depending on the importance given to the first criterion against the second, one value of x should be taken, and thus one specific sampling plan would be obtained by applying formula (21). Another interesting consequence of (21) is the fact that the set of Pareto-optimal vectors is independent of the expected fraction l of items with misstatements, since each a  i is proportional to ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1 x þ xu ðX i Þ p: 2.3. Sampling without replacement. Poisson sampling Addressing multipurpose survey designs via Multiple-Objective methods is not restricted to designs with replacement, as analyzed in Sections 2.1, 2.2 above. Indeed, if a sampling scheme without replacement is used to estimate the r-dimensional parameter #given in (1), instead of the Hansen–Hurwitz estimator, one can use the so-called Horvitz–Thompson estimator [10], b # j ¼1 NX N i¼1 Y ij I i p i ;j¼1;2;...;r;ð22Þ where I i is the random variate which takes the value 1 if u i is selected and takes the value 0 otherwise, and p i denotes the probability unit u i belongs to the sample, i.e., the so-called first-order inclusion probability. The multiple-objective problem to be addressed now has the same form than (12), namely min e 1 ðdÞ; e 2 ðdÞ;...; e r ðdÞ ðÞ s:td2D;ð23Þ where Dis a class of designs without replacement, and each e j represents the expected squared error under design d2Din the jth component, if we consider as random variables Y ij and I i as well, i.e., if the expectation is computed jointly with respect to the sampling design and the distribution of the variables Y ij , e j ðdÞ¼E1 NX N i¼1 Y ij 1 NX N i¼1 Y ij I i p i ! 2 0 @1 A ¼1 N 2 X N i¼1 1 p i p i ð r 2 ij þ l 2 ij Þþ 1 N 2 X N i;k¼1 i–k p ik  p i p k p i p k l ij l kj :ð24Þ Here p ik denotes the second-order inclusion probability, i.e., the probability that units u i ;u k are simultaneously in the sample. For simplicity, we write p i and p ik i.o. p i ðdÞand p ik ðdÞ, although it Fig. 1. Trade-off curve. 502 E. Carrizosa / European Journal of Operational Research 201 (2010) 500–504 should be clear that both the first and the second order inclusion probabilities are design-dependent. For a particular case of class of designs D, a characterization similar to the one of Section 2.1 is easily obtained. Indeed, let us consider the class Dof Poisson designs, e.g. [2]: Scalars a i ;06 a i 61 8 i¼1;...;Nare given, and each unit u i is selected with probability a i , independently of the remaining units. Poisson designs have the advantage that samples are easily drawn: independent variables X 1 ;...;X N , uniformly distributed on ½0;1are generated; the sample consists of those u i with X i 6 a i :Observe that the size of the samples generated is a random variate, with expected value P N i¼1 a i , and variance P N i¼1 a i ð1 a i Þ: For a Poisson design, (24) simplifies to e j ð a 1 ;...; a N Þ¼ 1 N 2 X N i¼1 r 2 ij þ l 2 ij a i 1 N 2 X N i¼1 r 2 ij þ l 2 ij  :ð25Þ Hence, e j has the form (14), and thus the arguments in Section 2.1 can be repeated to show the following: Proposition 3. Given ð a  1 ;...; a  N Þ,P N i¼1 a  i ¼n;06 a  i 61 8 i, the following statements are equivalent. 1. A Poisson design with selection probabilities ð a  1 ;...; a  N Þis a Pareto-optimal solution within the class of Poisson designs with expected size n, i.e. ð a  1 ;...; a  N Þis Pareto-optimal for min e 1 ð a 1 ;...; a N Þ; e 2 ð a 1 ;...; a N Þ;...; e r ð a 1 ;...; a N ÞðÞ s:tX N j¼1 a j ¼n; 06 a j 61;j¼1;...;N: ð26Þ 2. There exists ð x 1 ; x 2 ;...; x r Þ;P r j¼1 x j ¼1; x j P0 8 j, such that a  i ¼nffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi P r j¼1 x j r 2 ij þ l 2 ij  r P N l¼1 ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi P r j¼1 x j r 2 lj þ l 2 lj  r;i¼1;2;...;N:ð27Þ Having random sample sizes, as happens with Poisson designs, is seen as a serious drawback both from a managerial viewpoint (the sample size, and thus the sampling cost is unknown in advance) and statistical viewpoint as well: the errors e j and its estimates may be rather large. Different variants of Poisson designs have been proposed to mitigate or avoid this undesirable effect [1,2] the most popular being the conditional Poisson design, in which samples are successively drawn from a Poisson design with probabilities vector ð a 1 ;...; a N Þuntil a sample of size nis obtained. First and second order inclusion probabilities are obtained from the vector ð a 1 ;...; a N Þusing a recursive procedure [1]. However, the error functions e j do not have the nice and tractable form (14). As a matter of fact, the functions e j are not necessarily (quasi)convex. This is shown in Fig. 2 for a population Uof N¼10, n¼2, and a variable Y 1 ¼5;Y 2 ¼¼Y N1 ¼20;Y N ¼5: The section of the function for a 2 ¼¼ a N1 ¼ðn0:5Þ=ðN2Þ is depicted for varying a 1 : Hence, one cannot guarantee that all Pareto-optimal solutions can be obtained by minimizing a weighted sum of the errors [4]. Moreover, problems of type (9)–(11) are not so tractable, since local search may not lead to global optima. 3. Discussion In this note we have addressed the problem of finding multipurpose sampling designs using Multiple-Objective Programming ideas. The main result, with implications in Auditing, is a characterization of the set of Pareto-optimal designs when the designs under consideration are with replacement. The characterization obtained is given by a closed formula. It is then costless to get a fine discrete approximation to the Pareto-optimal set, to plot trade-off curves as in Fig. 1, and to analyze how the input parameters affect the Pareto set and the errors e j . The same multiple-objective problem for designs without replacement appears to be much harder, though, at the same time, very interesting from both a theoretical and a practical viewpoint. A characterization of Pareto-optimal designs within the class of Poisson designs is given. However, for fixed-size sampling variants of Poisson sampling, such as Conditional Poisson Sampling, a similar analysis does not seem to be possible, since the error functions are not even unimodal. Characterizing the Pareto-optimal designs for other classes D, such as order sampling schemes, [18], deserves attention, but it does not seem either to be an easy task, since even the calculation of the probabilities is a hard problem [15]. Analysis of the problem using estimators different from the Hansen–Hurwitz and Horvitz–Thompson estimators, or under statistical assumptions different to those addressed in this paper, remains an open problem. Acknowledgements The author thanks the two referees for their valuable suggestions. The financial support from Spain via MICINN (MTM200803032), and Junta de Andalucı ´a (FQM-329) is appreciated. References [1] N. Aires, Algorithms to find exact inclusion probabilities for conditional Poisson sampling and Pareto PPS sampling designs, Methodology and Computing in Applied Probability 1 (1999) 457–469. [2] K. Brewer, L. Early, M. Hanif, Poisson, modified Poisson and collocated sampling, Journal of Statistical Planning and Inference 10 (1984) 15–30. [3] E. Carrizosa, D. Romero Morales, A biobjective method for sample allocation in stratified sampling, European Journal of Operational Research 177 (2007) 1074–1089. [4] V. Chankong, Y.Y. Haimes, Multiobjective Decision Making: Theory and Methodology, North-Holland, 1983. [5] M. Clyde, K. Chaloner, The equivalence of constrained and weighted designs in multiple objective design problems, Journal of the American Statistical Association 91 (1996) 1236–1244. [6] R.D. Cook, W.K. Wong, On the equivalence of constrained and compound optimal designs, Journal of the American Statistical Association 89 (1994) 687– 692. [7] J.E. Gentle, S.C. Narula, R.L. Valliant, Multicriteria optimization in sampling design, in: S. Ghosh, W. Schucany, T. Smith (Eds.), Statistics of Quality, MarcelDekker, New York, 1997, pp. 411–425. [8] D.M. Guy, D.R. Carmichael, R. Whittington, Audit Sampling. An Introduction, Wiley, 2001. [9] M.M. Hansen, W.N. Hurwitz, On the theory of sampling from finite populations, Annals of Mathematical Statistics 14 (1943) 333–362. [10] D. Horvitz, D. Thompson, A generalization of sampling without replacement from a finite universe, Journal of the American Statistical Association 47 (1952) 663–685. [11] L. Imhof, W.K. Wong, A graphical method for finding maximin efficiency designs, Biometrics 56 (2000) 113–117. [12] C.M.S. Lee, Constrained optimal designs, Journal of Statistical Planning and Inference 18 (1988) 377–389. Fig. 2. Errors in conditional Poisson sampling. E. Carrizosa / European Journal of Operational Research 201 (2010) 500–504 503 [13] D. Malec, Selecting multiple-objective fixed-cost sample designs using an admissibility criterion, Journal of Statistical Planning and Inference 48 (1995) 229–240. [14] S. Mandal, B. Torsney, K.C. Carriere, Constructing optimal designs with constraints, Journal of Statistical Planning and Inference 128 (2005) 609–621. [15] A. Matei, Y. Tillé, Computational aspects of order PPS sampling schemes, Computational Statistics and Data Analysis 51 (2007) 3703–3717. [16] J.A. Mayor, Optimal cluster selection probabilities to estimate the finite population distribution function under PPS cluster sampling, Test 11 (2002) 73–88. [17] Official Journal of the European Union, 27 of December 2006, L 371/1–L 371/ 163. [18] B. Rosén, Asymptotic theory for order sampling, Journal of Statistical Planning and Inference 62 (1997) 135–158. [19] C.-E. Särndal, B. Swensson, J. Wretman, Model Assisted Survey Sampling, Springer, 1992. [20] R. Valliant, J.E. Gentle, An application of mathematical programming to sample allocation, Computational Statistics and Data Analysis 25 (1997) 337–360. [21] W.K. Wong, Recent advances in multiple-objective design strategies, Statistica Neerlandica 53 (1999) 257–276. 504 E. Carrizosa / European Journal of Operational Research 201 (2010) 500–504