scieee AI-readable full text Open interactive document viewer

Improving small area estimation by combining surveys : new perspectives in regional statistics

Costa, Àlex; Satorra, Albert; Ventura, Eva

Abstract

A national survey designed for estimating a specific population quantity is sometimes used for estimation of this quantity also for a small area, such as a province. Budget constraints do not allow a greater sample size for the small area, and so other means of improving estimation have to be devised. We investigate such methods and assess them by a Monte Carlo study. We explore how a complementary survey can be exploited in small area estimation. We use the context of the Spanish Labour Force Survey (EPA) and the Barometer in Spain for our study.

Full text

Statistics & Operations Research Transactions SORT 30 (1) January-June 2006, 101-122 Statistics & Operations Research Transactions Improving small area estimation by combining surveys: new perspectives in regional statistics∗ c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort Alex Costa1, Albert Satorra2and Eva Ventura2 1Statistical Institute of Catalonia (IDESCAT) 2Universitat Pompeu Fabra Abstract A national survey designed for estimating a specific population quantity is sometimes used for estimation of this quantity also for a small area, such as a province. Budget constraints do not allow a greater sample size for the small area, and so other means of improving estimation have to be devised. We investigate such methods and assess them by a Monte Carlo study. We explore how a complementary survey can be exploited in small area estimation. We use the context of the Spanish Labour Force Survey (EPA) and the Barometer in Spain for our study. MSC: 62J07, C2J10, 62H12 Keywords: composite estimator, complementary survey, mean squared error, official statistics, regional statistics, small areas 1 Introduction The 1978 Spanish constitution established that the Spanish Government has the exclusive competence over the statistics that are of interest for the whole country (Article 149.1.31a). At the same time, the statutes of the autonomous communities (regions) state that the regional administrations have the exclusive competence over the statistics that *The authors are grateful to Xavier L´ opez and Maribel Garcia, statisticians at IDESCAT, for their help at several stages of this research. The comments by Nicholas T. Longford on a previous version of this paper are very much appreciated. Address for correspondence: Albert Satorra and Eva Ventura. Department of Economics and Business. Universitat Pompeu Fabra. Received: February 2006 Accepted: June 2006 102 Improving small area estimation by combining surveys: new perspectives in regional statistics are of interest to the region (e.g., Article 33, 1979, Statute of Catalonia).1These laws lead to an interesting overlap of competences, since many surveys and administrative registers are of interest to both the country and the regions. The country’s official statistics bureaux have a much longer-standing tradition and greater resources, and are usually in charge of producing survey-based statistics with a country-wide scope. However, the statistics produced at the country level are sometimes not satisfactory for the region. This may arise for three reasons: 1) an issue that is relevant for the region but not for the country is not reported by the survey; 2) data collected at the country level is not reliable at the regional level; 3) statistics collected at the country level may not provide reliable information for small areas of a region. Foranexampleofthefirst case, consider the tourism surveys conducted in Catalonia, for which information on cross-border day trips to France or Andorra is highly relevant. The general questionnaire for the Spanish surveys does not include information about these trips since for the country as a whole these trips are of little importance. An example of the second case is that despite being over-represented in the National Labour Force Survey (EPA), less populous regions, such as Murcia and Navarra, still have too small subsample sizes to make any reliable inferences about them. The third problem is that territorially disaggregated information at the county or municipality levels, or for small islands (Isla de Hierro in the Canary Islands, or Formentera in the Balearic Islands, for example), is very important for regions, but at the overall country level they have a much lower priority. A regional statistics office could conduct a similar survey, duplicated and improved for its purpose, but that would amount to wasting resources and would increase the burden on respondents. Some subjects (companies, households, and the like) would receive virtually identical survey questionnaires, so they are likely to develop an impression that the national and regional statistical offices do not coordinate their activities. In addition to inducing a negative attitude towards official statistics, duplication of the respondents’ costs might be unacceptable. As an alternative, regional statistical offices may ask the National Statistics Institute (INE) to modify its survey design to meet the regions’ needs: to expand the questionnaire to cover issues of regional interest, or to increase the sample size to achieve sufficient precision for the inferences of interest. These changes would not cause problems in sporadically conducted surveys. An example is the Survey of Time Usage conducted in Spain on a single occasion in 2004. INE agreed with some regions to increase the subsample size in their areas. This option may not be available in some ongoing (annual, or quarterly) INE surveys that do not meet the regions’ needs due to problems related to reliability or territorial disaggregation. Reasons of technical, legal, or professional nature make modifications of the design of the ongoing surveys problematic. The national offices could not cope with the myriad of requests of various kinds from the 1. Or the Article 135, 2006, of the recently approved Statute of Catalonia. Alex Costa, Albert Satorra and Eva Ventura 103 regional offices. In this paper, we investigate an analytical solution to these problems that is based on supplementing the country survey with auxiliary information available for some of the small areas of interest. The use of auxiliary information is not a new idea in small area estimation. When the direct estimator for a particular small area is not satisfactory, one may resort to an indirect estimator. The direct estimator uses only information or data from the area and the variable of interest. Direct estimators are usually unbiased, though they may have large variances. An indirect estimator uses information from the small area of interest as well as from other areas and other variables, or even from other data sources. Indirect estimators are based on implicit or explicit models that incorporate information from other sources. For example, information obtained in a survey can be combined with the one collected in a census or an administrative register. Indirect estimators are usually biased, although their variances are smaller than those of the direct (unbiased) estimators, and the trade-offof bias and variance is usually in their favour. The novelty of our approach is that we use the information of an auxiliary survey instead of census or administrative records. We combine the information of a country-wide survey, called the reference survey (RS), with the information from a complementary survey conducted by the regional statistics office and tailored to the specific needs of the small area. A complementary survey (CS) is conducted at the regional level and records variables that correlate with the variables in the RS. CS covers one or several regions of the country, or part of a region. We regard CS as a “light survey” since the data will be faster and cheaper to collect than for the RS. For example, in the case of unemployment, a subject in the CS identifies him or herself as unemployed by the response to a single question. In contrast, the RS follow certain guidelines set forth by the International Labour Organization to classify the subject as unemployed (actively searching for work, available to begin working immediately, and so forth), employed and economically inactive. CS can also simplify the process of contacting the subjects (persons, companies, households, etc.) by using telephone contact systems (Computer Assisted Telephone Interviewing, CATI) or other automated survey methods. So, CS provides results similar to those of RS at a much lower cost; however, as CS records the values of a slightly different variable than RS, its results are biased. This is the price for the less elaborate questionnaire, with looser wording. We differentiate three types of CS: 1) A general complementary survey (GCS) covers all the regions of the country at one or several points in time. With data from many areas we can remove the bias of GCS estimators relative to RS. One example of GCS is the Economically Active Population Survey (EPA) conducted by INE as RS, and the Barometer of Spain conducted by the Centre for Sociological Research (CIS) as GCS. In the Barometer, respondents are asked if they are unemployed. Information from the EPA and CIS is available for all the Spanish regions for several years. 104 Improving small area estimation by combining surveys: new perspectives in regional statistics 2) With a regional complementary survey (RCS) we can assess the bias at the regional level, but not at the small area level because there are no data to compare RS and RCS at the small area level. An example is the Survey on Information and Communication Technologies in Catalonia, conducted by the Statistical Institute of Catalonia (IDESCAT) (RCS) by means of CATI, complementing an equivalent RS conducted by INE. RS has a clustered sampling design in which the 41 counties of Catalonia are not well represented. In contrast, the design of the RCS ensures an even coverage of the counties. In this example, the bias of RCS cannot be disaggregated to counties. 3) A local complementary survey (LCS) is a survey conducted in a specific small area. The bias is unknown since RS does not produce valid results for this small area. One example is a survey similar to the EPA in a single small municipality in Catalonia which has very sparse or no representation in EPA. The bias of the survey relative to RS can be explicitly modeled in GCS but not in RCS or LCS. We investigate how information from CS can be integrated with RS for making inferences about small areas. We consider the specific context in which EPA is RS and the Barometer is CS, in this case, a GCS. The Barometer contains a few questions regarding the subject’s employment status which are at face value highly correlated with the corresponding variable in EPA. The accuracy in small area estimation can be increased by: a) increasing the sample size in the area of interest; b) borrowing strength from neighbouring areas (using indirect or composite estimators); c) borrowing strength from CS, especially when the variables in RS and CS are highly correlated. We explore all these alternatives, with emphasis on combining the options b) and c). The performance of the estimators and the contribution of the complementary information will be assessed by simulation. Parallel work on the use of CS has been conducted by Costa et al. (2006), who study the Survey on the Uses of Information and Communication Technologies in Catalan households; INE conducts a country-wide survey while IDESCAT is in charge of the RCS. The present paper is organized as follows. Section 2 reviews established small area estimators, with emphasis on estimating labour statistics. Section 3 describes the specific context of estimating rates of unemployment in Spain. Section 4 assesses the performance of the alternative small area estimators by simulations. Section 5 summarizes the main findings of the paper. 2 Estimators for small areas In this section we consider a two-stage clustered sampling design, motivated by the sampling design of EPA that is considered later in the paper. We consider a binary Alex Costa, Albert Satorra and Eva Ventura 105 variable Ythat takes the values Yij =1 if the characteristic under study is present for subject ij,andYij =0 otherwise. Here i(i=1,2,...,n)and j(j=1,2,...,mi) denote primary sample unit (PSU) and secondary sampling unit (SSU), respectively. We use the convention that capitals (X,Y) denote population values and lowercases (x,y) sample values. Their indexing is implied; that is, in Xij we use population indexing and in xij we use sample indexing. For every sample, we have a variable Wof sampling weights, with wij representing the sampling weight of subject ij. The population is divided in Ksmall areas, indexed by k=1,2,...,K.Weuse the notation Yk.ij for the values of variable Yon units of area k. For sampling data, the symbol +in the subscript denotes the weighted summation over the sample; for example, yk,i+=mi j=1wk,ijyk,ij . For population data, the symbol +indicates summation without weighting. Our target is the population ratio θk=Yk,+/Xk,+of two totals, for each area k.We consider also the overall population ratio θ=Y+/X+. It is assumed that the denominator is positive. Several estimators are considered. 2.1 Direct estimator A direct estimator of θkuses only data from area k.Itisdefined as ˆ θk=ˆ Yk ˆ Xk , where ˆ Yk=yk+and ˆ Xk=xk+. Here the summation extends only over the (say, nk) PSUs that intersect with area k(we assume nk>2). Straightforward application of the delta-method yields the following estimator of variance V(ˆ θk) ˆ V(ˆ θk)=1 ˆ X2 kˆ V(ˆ Xk)−2ˆ θk cov( ˆ Yk,ˆ Xk)+ˆ θ2 kˆ V(ˆ Xk),(1) where  cov( ˆ Yk,ˆ Xk)=nk nk−1 nk  i=1 (zk,i(y)−¯z(y))(zk,i(x)−¯z(x)),(2) zk,i(y)=yk,i+and ¯z(y)=n−1 k nh  i=1 zk,i(y), and similarly for x. We compute ˆ V(ˆ Xk)and ˆ V(ˆ Yk)as cov( ˆ Xk,ˆ Xk)and cov( ˆ Yk,ˆ Yk), respectively. In the general case of Lstrata, the sample values are yh,ij,werehdenotes strata. The direct estimator of the overall population ratio θ=Y+/X+is ˆ θ=ˆ Y/ˆ X,where ˆ Y=y+,++ 106 Improving small area estimation by combining surveys: new perspectives in regional statistics and ˆ X=x+,++ (summation over strata, PSUs within the strata, and units within the PSU). The estimator of var (ˆ θ) is like (1) and (2) with subscript ksuppressed and a summation over Lstrata added to the right hand side of (2). Information on population totals of some auxiliary variables would allow us to calculate post-stratified or ratio-estimators. This will not be pursued in this study. For more information on those estimators, the reader can consult Rao (2003), Ghosh and Rao (1994) or Mancho (2002). L´ opez (2000) considers some of these estimators in the context of small area estimation for EPA in Canary Islands. We consider ˆ θkand ˆ θas the only direct estimators in this study. 2.2 Small area estimators without auxiliary information An indirect estimator of θkuses data from outside area k. As an alternative to ˆ θkwe may adopt the overall-country direct estimator ˆ θ=ˆ Y/ˆ Xfor every area k. This is an indirect estimator. Being based on much more data than ˆ θk,ˆ θhas a much smaller variance than ˆ θk, but is biased for θk, unless the θk’s are all equal. In this case, ˆ θis much more efficient than ˆ θk. But if the θk’s vary substantially across areas, the bias of ˆ θwill be large, and so will be its mean squared error (MSE). An attractive alternative estimator to both ˆ θkand ˆ θis the composite estimator ˆ θ(c) kdefined as the convex combination ˆ θ(c) k=φkˆ θ+(1 −φk)ˆ θk(3) with 0 ≤φk≤1. The coefficient φkis chosen so as to minimize the MSE and is equal to φk=var (ˆ θ)−cov (ˆ θk,ˆ θ) (θk−θ)2+var (ˆ θk)+var (ˆ θ)−2cov(ˆ θk,ˆ θ).(4) The denominator of (4) is positive; in fact, it is equal to E{(ˆ θk−ˆ θ)2}. Clearly, φkdepends on some unknown parameters, and itself has to be estimated. Since ˆ θ=K k=1ˆ Yk ˆ X=K k=1ˆ Xkˆ θk ˆ X= K  k=1 qkˆ θk, where qk=ˆ Xk/ˆ X,cov( ˆ θk,ˆ θ)=σ2 kqk, so the optimal weight is φk=σ2 k(1 −qk) (θk−θ)2+σ2 k(1 −2qk)+σ2,(5) where σ2 kand σ2are the respective sampling variances of ˆ θkand ˆ θ(see Longford, 1999). When qkis very small and the survey is large (e.g., the number Kof small areas is Alex Costa, Albert Satorra and Eva Ventura 107 large and the sample sizes of most of them are small), we can ignore both qkand the variance σ2; then, φkis approximated by φk=σ2 k (θk−θ)2+σ2 k . We could estimate σ2 kfrom the sample data from area kand use (ˆ θk−ˆ θ)2as an estimator of the denominator in (5). Our experience, shows that this results in a very unstable estimator of φk(see Costa, Satorra and Ventura 2003, 2004). One way to overcome this difficulty is by averaging the estimators ˆσk’s among several areas (or several variables). For example, Purcell and Kish (1979) use a weight common to all areas that minimizes the between-area average of the mean squared errors. By assuming that the within-area variances of Yare equal, the pooled estimator of their common variance is ˆσ2 w=1 n−K K  k=1 (n−1) ˆσ2 k.(6) By using the following estimator of the square of the bias b2=1 K K  k=1 (ˆ θk−ˆ θ)2,(7) and ignoring both qkand the var (ˆ θ), we estimate the weight as ˆ φk=ˆσ2 w ˆσ2 w+b2.(8) Expressions (3) and (8) define the classic composite estimator (Costa, Satorra and Ventura, 2003). 2.3 Auxiliary information In the literature on small area estimation, we find many indirect estimators that incorporate auxiliary information. Usually this information consists of data from a census or an administrative register. Rao (2003) describes the regression synthetic estimator, which combines direct estimators with those obtained from census or records, encompassing a large variety of estimators, such as ratio or count estimators. We use auxiliary information that arises from a CS carried out in each of several small areas. In particular, we assume that for each of a set of small areas we have the direct estimator ˆ θkas well as an estimator ˆ δkderived from CS. For these estimators, 108 Improving small area estimation by combining surveys: new perspectives in regional statistics consider the simple regression equation ˆ θk=α+βˆ δk+k,(9) k=1,2,...K,andthefitted estimator of θkby the OLS regression, ˆ θF k=ˆα+ˆ βˆ δk. When historical data are available for the RS and CS across several areas, the fitted estimator ˆ θF kcould be based on more advanced regression than just simple OLS. If we have RS and CS at several time points (t=1,...,T) and several areas (k=1,2,...,K), we could estimate θkby an analysis of covariance model. The regression could also involve other covariates. As more variables are incorporated into the regression that links ˆ θkwith ˆ δk, the synthetic estimator ˆ θF kwill be more efficient for θk, although using too many covariates may inflate the sampling variance. For simplicity, we only consider the OLS regression (9). In the Monte Carlo set-up of Section 4, however, we also involve an estimator that is based on a covariate (fixed-effects, FE) regression model that serves as a benchmark for maximum information attainable from CS. Even though the variance of ˆ θF kmay be substantially smaller than var (ˆ θk), ˆ θF kmay be biased. We improve both ˆ θkand ˆ θF kestimators by considering the composite estimator ˆ θ(c) k(CS)=φkˆ θF+(1 −φk)ˆ θk(10) where φk=var (ˆ θk)−cov (ˆ θF k,ˆ θ) Δ2 k+var (ˆ θk)−var (ˆ θF k)(11) Here Δk=θk−E(ˆ θF k) has to be estimated. If αand βwere known, var (ˆ θF k)=β2(var ˆ δk). When the regression parameters are estimated, then var ˆ θF k=Evar ˆ θF k|ˆ Θ+var E(ˆ θF k|ˆ Θ), where ˆ Θ=(ˆα, ˆ β) stands for the vector of estimated regression coefficients. Since the expected value of (ˆ θk−ˆ θF k)2coincides with the denominator in (11), the weight of (11) is estimated as ˆ φk=ˆσ2 k−˜σ2 k (ˆ θk−ˆ θF k)2, where ˆσ2 kand ˜σ2 kare the respective estimators of the variances of ˆ θkand ˆ θF k.An alternative estimator of φkis more stable, Alex Costa, Albert Satorra and Eva Ventura 109 ˆ φk=ˆσ2 ˆσ2+b2,(12) where ˆσ2is given by (6) and b2=1 K K  k=1 (ˆ θk−ˆ θF k)2.(13) The estimator ˆ θc k(CS)defined by equations (10), (12) and (13) is called the composite complementary survey (CCS) estimator based on OLS regression (or CCS-OLS). In Section 4, we assess by Monte Carlo the efficiency of the CCS, direct, indirect and classic composite small area estimators. The efficiency of the CCS estimator will be compared against a direct estimator based solely on RS but with sample size increased by r%, with r=10,20,50,100. For the Monte Carlo study we construct an artificial population that resembles Spain in some aspects related to the labour force. The EPA and Barometer Surveys are used for this construction. 3 EPA and Barometer surveys This section describes general aspects of estimation of unemployment rates in Spain, both at country and area levels. The main source of information about unemployment in Spain is the EPA conducted by the INE, our RS. The Barometer is our CS. EPA is a quarterly survey that uses a lengthly questionnaire, panel design and face-toface interview (paper and pencil, PAPI); in contrast, the Barometer is a monthly survey that uses CATI as the interviewing system and an entirely new sample each month. EPA is designed for reliable estimation of several labour market statistics, including the unemployment rate at country level. The Barometer uses the self-perceived labour status of the interviewed individuals as a proxy for unemployment status. While all the provinces of Spain are represented in the EPA, this is not always the case for the Barometer. Even if a province is represented in the survey, its sample size may be very small. To have a more informative study, we grouped the 50 Spanish provinces (plus the two autonomous cities in the north of Africa) into 25 areas, according to their geographical proximity and similarity of their labour markets. With this grouping, each area is represented by at least 140 observations in CS. The second and the sixth column of Table 1 shows the sample sizes of EPA and CIS surveys across the 25 areas for the fourth quarter of 2003. To give the same time frame to both surveys, the quarterly estimate by Barometer is taken to be the average across three months of the monthly unemployment rates of the Barometer. While EPA follows the standard International Labour Organization methodology, the Barometer simply asks the subjects how they perceive their employment status. The unemployment rates in the two surveys differ for two reasons: different definitions of 116 Improving small area estimation by combining surveys: new perspectives in regional statistics Table 6:Estimation of unemployment rates for men. For different sample sizes, the table shows the RRMSE average (across areas) of the various estimators evaluated in the Monte Carlo study. The estimators are: RS-based estimators (direct with sample size boosted by r%; indirect, ˆ θ; composite, ˆ θ(c) k), and the CCS estimators based on fixed effects (FE) and simple (OLS) regression. All the values of RRMSE have been multiplied by 1000. r%CS-based Summary 0(0) 10 25 50(b) 100 ˆ θ(I) ˆ θ(c) k (1) OLS(2) FE(+) Average sample size 100 Average 669 638 601 546 469 404 478 487 388 Median 645 603 571 525 469 313 447 462 374 Max 915 892 852 776 685 1526 795 803 529 Average sample size 200 Average 475 450 423 385 330 385 371 374 281 Median 470 429 409 372 328 300 347 339 272 Max 663 631 590 558 473 1503 643 658 384 Average sample size 400 Average 335 316 300 270 238 373 285 287 205 Median 325 299 292 263 231 300 266 264 198 Max 492 454 438 390 344 1484 462 495 294 Average sample size 500 Average 298 283 268 247 213 372 259 263 185 Median 294 267 266 233 206 297 250 244 186 Max 411 392 380 347 299 1493 428 454 245 Average sample size 1000 Average 212 204 193 178 153 368 196 199 141 Median 206 195 185 170 150 296 194 190 144 Max 292 286 275 243 214 1488 311 331 184 (i) These superscripts show the symbols used in Figures 1 and 2 to represent the RRMSEs. Similar conclusions are arrived at by inspecting the results for men and women. The RRMSEs for women tend to be larger than for men for a fixed sample size, because their rates or unemployment are higher than for men and RRMSEs are approximately proportional to p/(1 −p), where pis the unemployment rate. The tables contain a lot of detail that is difficult to digest and do not indicate the performance of the estimators for the individual areas. Figures 1 and 2 display RRMSEs of four small area estimators: direct, marked as 0; indirect, I; composite, 1; and CCSOLS, 2. It shows also the benchmark estimator, marked as +; and the direct estimator with sample size boosted by 50%, marked as b. We regard estimators 0,1,2 and I as feasible because they use information that would normally be available. Estimators + and b are a benchmark and comparator, respectively. They use information that would not be available in practice. At the outset, I is discarded as competitor of 0, 1 and 2. Alex Costa, Albert Satorra and Eva Ventura 117 Table 7:Estimation of unemployment rates for women. For different sample sizes, the table shows the average (across areas) of RRMSE of the various estimators evaluated in the Monte Carlo study. The estimators are: RS-based estimators (direct with sample size boosted by r%; indirect, ˆ θ; composite, ˆ θ(c) k), and the CCS estimators based on fixed effects (FE) and simple (OLS) regression. All the values of RRMSE have been multiplied by 1000. r%CS-based Summary 0(0) 10 25 50(b) 100 ˆ θ(I) ˆ θ(c) k (1) OLS(2) FE(+) Average sample size 100 Average 539 514 476 437 378 315 386 365 327 Median 537 512 475 445 376 215 383 345 335 Max 757 736 716 648 540 1164 618 621 443 Average sample size 200 Average 376 361 338 309 267 304 292 273 235 Median 383 368 327 305 266 200 279 265 240 Max 556 521 493 444 377 1155 485 517 331 Average sample size 400 Average 269 259 241 218 190 298 228 210 180 Median 262 261 237 211 192 192 217 210 182 Max 419 374 374 314 269 1147 393 410 267 Average sample size 500 Average 238 230 215 197 172 296 206 194 163 Median 234 228 211 194 170 191 198 195 161 Max 352 324 314 283 250 1151 333 381 232 Average sample size 1000 Average 172 166 155 143 126 293 157 151 129 Median 168 164 152 136 123 190 153 155 130 Max 252 241 238 229 218 1153 239 284 192 (i) These superscripts show the symbols used in Figures 1 and 2 to represent the RRMSEs. The areas are ordered according to their unemployment rates. The diagrams show, for instance, that estimator I has serious weaknesses although it is the most efficient for a few areas in the middle of the range. In general, the composite estimators 1 and 2 are the most efficient among the feasible estimators. In Figure 1 we have a graphical representation of the RRMSE for the alternative estimators in the case of large sample size (1000). Small areas are on the vertical axis and different symbols represent the different estimators. RRMSE is on the horizontal axis, so that efficiency corresponds to being on the left. For this large sample case, the indirect estimator (I) is generally very inefficient: each RRMSE summary has been truncated at 0.25, except for areas whose unemployment rate is close to the national rate. The composite estimators, the RS-based (1) and CS-based (2) are the most efficient estimators, after the benchmark estimator CCS-FE (+). The composite estimators 1 and 118 Improving small area estimation by combining surveys: new perspectives in regional statistics 2 (RS and CS based, respectively) are generally as efficient as b, direct estimator with sample size boosted by 50%. 0.00 0.05 0.10 0.15 0.20 0.25 10 20 RRMSE: large sample size RRMSE Areas 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 I I I I I I I I I I I I I I I I I I I I I I I I I 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 + + + + + + + + + + + + + + + + + + + + + + + + + 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 b b b b b b b b b b b b b b b b b b b b b b b b b ARA NRI GLT MAD ACS AGU BCN BAL AST VAL CGT CNT CAN CLE ACR MUR LOP VIZ LCO CJA EXT AGR MAL SEV CHU 0 I 1 + 2 b Direct Indirect Composite CCS−FE CCS−OLS boost50 Figure 1:RRMSEs for the areas and for estimators in the case of large sample (average small area sample size is 1000). The areas have been ordered in increasing order of magnitude of their rate of unemployment. Values of RRMSE have been truncated at 0.25. Figure 2 shows the same information, using the same layout and symbols, for small sample size (100). Again, the indirect estimator (I) is the best for the areas which unemployment value is at the middle range (around the national rate), although they are very inefficient for the areas with extreme rates of unemployment. In those areas, the composite RS-based and the feasible new CS-based estimators (1 and 2, respectively) Alex Costa, Albert Satorra and Eva Ventura 119 are more efficient than the direct estimator b for most of the areas. For most of the areas, 2 is the most efficient, after the benchmark estimator +. 0.0 0.2 0.4 0.6 10 20 RRMSE: small sample size RRMSE Areas 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 I I I I I I I I I I I I I I I I I I I I I I I I I 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 + + + + + + + + + + + + + + + + + + + + + + + + + 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 b b b b b b b b b b b b b b b b b b b b b b b b b ARA NRI GLT MAD ACS AGU BCN BAL AST VAL CGT CNT CAN CLE ACR MUR LOP VIZ LCO CJA EXT AGR MAL SEV CHU 0 I 1 + 2 b Direct Indirect Composite CCS−FE CCS−OLS boost50 Figure 2:RRMSEs for the areas and for estimators in the case of small sample (average small area sample size is 100). The areas have been ordered in increasing order of magnitude of their rate of unemployment. Values of RRMSE have been truncated at 0.75. For an area and setting of the simulations (sample size), we define the pattern of RRMSEs by their order for the estimates 0, 1 and 2. For example, for the setting in Figure 1 (large sample size), the pattern for Arag´ on is 012, which means that the RRMSE for 0 is the smallest and the RRMSE for 2 is the largest of the three; see bottom of the diagram. We say that estimator 2 is the winner for an area if the pattern 120 Improving small area estimation by combining surveys: new perspectives in regional statistics of RRMSEs is 201 or 210. In Figure 2, we see that 2 is the winner in 22 areas and 1 is the winner in the remaining three areas. This insight can not be gained from Tables 5-7. With large sample size, estimator 2 wins only 17 areas, so it is still preferable, but less decisively so. We also see that 2 is more efficient than b in 22 areas for small sample, but in only two areas for large sample. Tables 5-7 and Figures 1 and 2 corroborate the prior expectation that composite estimators outperform direct estimators in almost all settings and for almost all areas, and that the indirect estimator is efficient only in areas with small sample size. We summarize our findings from the simulations as follows: 1) CCS-OLS (with sample data at one time point) is less efficient than the benchmark estimator CCS-FE. 2) Only for very large samples (1000), CCS-OLS has no gains over the direct estimator. For smaller samples, CCS-OLS is comparable with the benchmark estimator with sample size boosted by up to 50%. 3) A substantial part of the gains attained by the benchmark estimator is attained also by the CCS-OLS estimator. 4) The CCS-OLS estimator is slightly more efficient than the estimators that use information solely from RS. 5) The behaviour of the small area estimators does not change much whether we consider total, male or female unemployment rates. 6) In the context of the estimation of Spanish unemployment rates and for moderate area sample sizes (say, 200 subjects in the area), the simplest CCS-OLS estimator is comparable with an increase of sample size by up to 50%. As a concluding remark, our results show that regional statistics are not in conflict with the statistics produced at the country-wide level. Rather, a regional survey can be combined with a country survey to improve the precision of estimators for small areas, avoiding the costly solution of increasing the region’s subsample size. 6 References Costa, A, Garcia, M., Lopez, X., and Pardal, M. (2006). Estimaci´ o de les taxes de desocupaci´ o comarcal a Catalunya. Aplicaci´ o d’estimadors de petita ` area amb combinaci´ o d’enquestes, Working Document, IDESCAT, Barcelona. Costa, A., Satorra, A. and Ventura, E. (2003). An empirical evaluation of small area estimators, SORT (Statistics and Operations Research Transactions), 27 (1), 113-135. Costa, A., Satorra, A. and Ventura, E. (2004). Using composite estimators to improve both domain and total area estimation, SORT (Statistics and Operations Research Transactions), 28 (1), 69-86. Ghosh, M. and Rao, J. N. K. (1994). Small area estimation: an appraisal. Statistical Science, 9 (1), 55-93. Alex Costa, Albert Satorra and Eva Ventura 121 Longford, N. T. (2004). Missing data and small-area estimation in the UK Labour Force Survey. Journal of the Royal Statistical Society A, 167, 341-373. L´ opez, R. (2000). Estimaciones para ´ areas peque˜ nas. Estad´ıstica Espa˜nola, 42, 146, 291-338. Mancho, J. (2002). T´ ecnicas de estimaci´ on en ´ areas peque˜ nas. Cuaderno T´ecnico del Eustat. Rao, J. N. K. (2003). Small Area Estimation. Wiley Series in Survey Methodology. StataCorp. (2003) Stata Statistical Software: Release 8.0. College Station, TX: Stata Corporation.