scieee AI-readable full text Open interactive document viewer

SOEP-Core - 2020: Sampling, nonresponse, and weighting in the IAB-SOEP migration studies M7 and M8

Steinhauer, Hans Walter,Trübswetter, Parvati,Zinn, Sabine

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Steinhauer, Hans Walter; Trübswetter, Parvati; Zinn, Sabine Research Report SOEP-Core - 2020: Sampling, nonresponse, and weighting in the IAB-SOEP migration studies M7 and M8 SOEP Survey Papers, No. 1105 Provided in Cooperation with: German Institute for Economic Research (DIW Berlin) Suggested Citation: Steinhauer, Hans Walter; Trübswetter, Parvati; Zinn, Sabine (2022) : SOEP-Core - 2020: Sampling, nonresponse, and weighting in the IAB-SOEP migration studies M7 and M8, SOEP Survey Papers, No. 1105, Deutsches Institut für Wirtschaftsforschung (DIW), Berlin This Version is available at: https://hdl.handle.net/10419/261441 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/ Hans-Walter Steinhauer, Parvati Trübswetter, Sabine Zinn Series C - Data Documentations (Datendokumentationen) SOEP-Core – 2020: Sampling, Nonresponse, and Weighting in the IAB-SOEP Migration Studies M7 and M8 1105 2022 SOEP-Core – 2020: Sampling, Nonresponse, and Weighting in the IAB-SOEP Migration Studies M7 and M8 Hans Walter Steinhauer1, Parvati Trübswetter2, and Sabine Zinn1,3 1Deutsches Institut für Wirtschaftsforschung 2Institut für Arbeitsmarkt- und Berufsforschung 3Humboldt-Universität zu Berlin July 18, 2022 Abstract This paper provides details on sampling design, fieldwork, nonresponse and population adjustments for the 2020 samples M7 and M8 of the Socio-Economic Panel (SOEP). Sample M7 refreshes the SOEP core sample, especially samples M1 and M2, including households with household heads having a Bulgarian, Polish, or Romanian nationality. The sample M8 augments the SOEP core sample, sampling households of foreigners from third countries to evaluate the “Fachkräfteeinwanderungsgesetz.” Obtaining nearly 1,000 household interviews and panel consent of households for each sample was complicated by the first wave of the Corona pandemic and the first lockdown. Nevertheless, nonresponse on the household level is driven by a variety of characteristics, such as nationality or regional contexts as well as information contained in the Integrated Employment Biographies. 1 Introduction Starting with sample M1 in 2013 the Institute for Employment Research (IAB) in Nuremberg and the German Socio-Economic Panel (SOEP) at DIW Berlin began to jointly survey the migrant population in Germany (Liebig et al., 2021). In 2015, sample M2 became the follow-up sample, where both institutes continued their cooperation. The sample M7 refreshes the SOEP core sample, especially samples M1 (Kroh, Kühne, Goebel, & Preu, 2015) and M2 (Kühne & Kroh, 2017), with households including persons with Bulgarian, Polish, or Romanian nationalities. As before, we use the Integrated Employment Biographies (IEB) as a nationwide sampling frame. The IEB is spell data based on IAB’s employment history (BeH), IAB’s benefit recipient history (LeH), the participants-in-measures data (MTG), and job search data originating from the applicants pool database (BewA). Thus, the IEB include observations of unemployment benefits, job search, and participation in active labor market programs, see Oberschachtsiek, Scioch, Seysen, and Heining (2009) for details. Beyond that, it covers socio-demographic information on gender, age, and nationality as well as geographic information, including, for example, postal codes, municipality, and regional classification. Moreover, the cooperation allows us to link information from the IEB to the SOEP data and use them, for example, in the nonresponse analyses. Moreover, the Federal Employment Agency (Bundesagentur für Arbeit) provides information on third-country nationals who applied for working in Germany as professionals (“Fachkräfte”) based on the Residence Act (Zuwanderungsgesetz, ZuwG). This information is used to identify the population in the IEB data. The sample from this population is referred to as M8. It was sampled and surveyed in 2020 in order to provide a basis to evaluate the “Fachkräfteeinwanderungsgesetz” – Skilled Immigration Act – which became effective March 1st, 2020. This paper documents the sampling design and the weighting strategy for the 2020 samples M7 and M8 of the SOEP. Therefore, section 2 provides details on the population. Sampling is described in section 3. Section 4 provides information on the fieldwork and its results. Weighting adjustments are presented in section 5. Finally, section 6 gives a brief summary. 2 Target Population and Sampling Frame The target population of sample M7 consists of persons living in private households having a Bulgarian, Polish, or Romanian nationality, who immigrated to Germany between January 1, 2016, and December 31, 2018. The target population of sample M8 consists of third-country nationals living in private households who were granted a permission to work in Germany as professionals between January 1, 2019, and January 30, 2020. To sample from these two populations, we make use of the IEB data, which is official data provided by the IAB. In total the IEB contain 917,835 persons belonging to the population of M7 and 112,945 to the population of M8. We further restrict the persons to those having a valid address in Germany. For M7 (M8) this reduces the number of persons by 146,743 (29) to a target population of 771,092 (112,916) persons. In case of the M7 sample, the reduction is mostly driven by persons working in Germany but having an address in regions near the border outside of Germany, for example, in Poland. The number of persons with valid address information within Germany contained in the IEB SOEP Survey Paper 1105 1 v37 data by group for the two samples is displayed in Table 1a for M7 and in Table 1b for M8. Table 1: Number of target persons for in the IEB data by group. (a) Subsample M7 Nationality Number Percent Bulgarian 132,154 17.1 Polish 243,566 31.6 Romanian 395,372 51.3 Total 771,092 100.0 (b) Subsample M8 Application for Number Percent Qualified employment 50,430 44.7 Unqualified employment 32,114 28.4 Education 14,268 12.6 Other 16,104 14.3 Total 112,916 100.0 The table clearly shows that the majority of the M7 target population has immigrated from Romania (slightly over 50%); the smallest share originates from Bulgaria (around 17%). These immigrants scatter unequally across Germany. Throughout the 8,171 German postal code areas, there are several regions along the border and within the country, especially in the eastern part of Germany, where there are fewer than 50 immigrants. These are most likely to be regions that people commute to rather than move there, if they have a job in these regions. The majority of immigrants is located in urban and agrarian-oriented areas in the western and southern regions of Germany as well as in Berlin and its neighboring regions in Brandenburg. In contrast to the M7 population, the majority of the M8 target population is located in urban areas. Here, their number is highly correlated to the number of people forming the M7 population. 3 Sampling Design The sampling design can be summarized as a stratified multi-stage sampling design. Because the distribution of the immigrant population for M7 containing N= 770,709 individuals is unequally spread over Germany and within states, we form regional clusters of postal code areas as primary sampling units (PSU) stratified by federal states and a rural-urban-classification; strata h= 1, . . . , H.1The PSU , indexed j= 1, . . . , Mh, were constructed to cover a minimum of at least 600 and no more than 1,200 immigrants, that is, 600 ≥Njh ≤1,200. In a first step, each postal code exceeding the minimum number of 600 immigrants became a PSU itself. This size is chosen in order to guarantee the minimum number of 60 immigrants from Bulgaria, who form the smallest group, compare Table 1. In a second step, a postal code was selected at random and the neighboring postal codes were attached until the minimum number was achieved. In the third step, all remaining postal codes as well as those PSU exceeding 1,200 immigrants were again split, then the second step was repeated. In the last step, all remaining postal codes were attached to the neighboring PSU that covered the least number of immigrants. This procedure clustered a total of 8,171 postal codes into M= H ∑ h=1 = 773 PSU. In the first stage, m= 125 PSUs had to be selected with systematic probability proportional to size 1The number of the target population is reduced by another 383 persons because they were born before 2002, thus most likely not yet of legal age. SOEP Survey Paper 1105 2 v37 sampling. To balance between urban and agrarian-oriented areas, the latter were reduced in weight for sampling by the factor pjh = 0.5. The measure of size (xjh) for PSU jin stratum his xjh =pjh ·Njh and the stratum-specific measure of size xhis xh= Mh ∑ j=1 xjh. The number of PSUs to select from each stratum was allocated proportional to the measure of size per stratum, that is mh=m·xh H ∑ h=1 xh . Thus, the inclusion probability πjh for PSU jin stratum his given by πjh =mh·xjh Mh ∑ j=1 xjh =mh·pjh ·Njh Mh ∑ j=1 pjh ·Njh (1) Within each of the sampled PSU, a simple random sample of ns= 60 immigrants was drawn from each nationality s. Thus, the inclusion probability for immigrant iof nationality sis πis =min(ns, Nsjh) Nsjh (2) such that the final inclusion probability πisjh for immigrant iof nationality ssampled in PSU jin stratum his πisjh =πis ·πjh.(3) This sampling procedure yields a maximum sample of n=m·3·ns= 125 ·3·60 = 22,500 To realize the sample for M8 we make use of the high correlation in the number of individuals in the two populations. Because of this, we were able to use the same PSUs formed for sampling households for M7. A previous simulation study showed that using the same PSUs will provide a sample of sufficient size for M8, too. The final samples drawn from the two populations include 22,020 individuals for M7 and 21,552 for M8. 4 Fieldwork Results and Response Rates After sampling, the addresses were handed over to KANTAR Public, the field work agency, and were validated. During the fieldwork, a total of 19,751 addresses were validated for M7 and 12,992 for M8. This left 2,269 addresses in M7 and 8,560 in M8 unused. The validation yielded a noticeable number of invalid or old addresses that were not eligible. Of these, the largest number was untraceable and a huge number also had moved abroad. We find this very likely for the following reasons. First, many individuals from Poland, Bulgaria, and Romania come to Germany as seasonal workers. Second, both populations are likely to have moved back to their home country because of the Covid-19 pandemic. For these reasons, only 8,173 addresses were visited by interviewers for the M7 subsample and 7,804 for M8. Table 3 displays the results for the fieldwork. In total, there were 783 complete or partial interviews in M7 and 1,096 in M8 resulting in a response rate on the household-level, calculated according to American Association for Public Opinion Research (2016), of RR2M7= 0.096 for M7 and RR2M8= 0.141. The response rate at the household-level is quite low, but as expected because of the underlying populations and the SOEP Survey Paper 1105 3 v37 Table 2: Number of target persons, postal codes and PSU by Federal State. Federal Target persons Number of State M7 M8 postal codes PSU in population PSU in sample BB 14,933 1,594 215 18 2 BE 29,982 8,207 190 24 5 BW 120,514 23,434 1,194 131 20 BY 138,846 26,482 2,062 145 22 HB 7,345 ,643 40 7 1 HE 70,303 11,544 544 74 11 HH 15,488 3,058 100 13 3 MV 9,717 1,027 189 10 2 NI 86,136 6,685 796 80 14 NW 158,092 17,429 865 154 26 RP 50,027 4,591 659 48 8 SH 23,477 2,241 445 23 4 SL 5,999 ,545 69 6 1 SN 13,247 2,949 384 14 2 ST 12,687 1,159 201 12 2 TH 13,916 1,328 218 14 2 Total 770,709 112,916 8,171 773 125 Note: BW = Baden-Württemberg, BY = Bavaria, BE = Berlin, BB = Brandenburg, HB = Bremen, HH = Hamburg, HE = Hessen, MV = Mecklenburg-Vorpommern, NI = Lower Saxony, NW = North Rhine-Westphalia, RP = Rhineland-Palatinate, SL = Saarland, SN = Saxony, ST = Saxony-Anhalt, SH = Schleswig Holstein, TH = Thuringia. Covid-19 pandemic. The refusal rate (REF 1) is similar when compared to other samples / studies. For M7 the refusal rate is REF 1M7= 0.225 and for M8 it is REF 1M8= 0.213. For more detailed information on the fieldwork see Rathje and Glemser (2021). 5 Cross-sectional Weighting The computation of survey weights is usually performed in three steps (Brick & Kalton, 1996). In the first step, design weights are calculated as inverse of the inclusion probability, see Section 3. Second, these design weights are adjusted to correct for unit nonresponse. This step is referred to as sample weighting adjustment by Kalton and Kasprzyk (1986). Lastly, weights are calibrated so that estimates conform to known population parameters or to meet specific distributions. Kalton and Kasprzyk (1986) refer to this step as population weighting adjustment. For details on the general weighting strategy of the SOEP and the integration of new samples, see Kroh, Siegers, and Kühne (2015). To account for possible selectivity due to nonresponse, we model the participation decision of the households using information on participating and nonparticipating households. Because there usually is only limited information available on nonparticipating households, we use area level information as well as interviewer observations on the residential environment. Information collected by the interviewer on the residential environment include: SOEP Survey Paper 1105 4 v37 Table 3: Fieldwork results on the household-level according to American Association for Public Opinion Research (2016). Final Disposition M7 M8 Code Number Percent Number Percent 1. Interview (1.1) Complete 301 0.022 700 0.062 (1.2) Partial 482 0.035 396 0.035 Subtotal 783 0.057 1,096 0.097 2. Eligible, Non-Interview (2.11) Refusals 1,840 0.135 1,665 0.147 (2.20) Non-contact 3,960 0.290 3,712 0.328 (2.31) Dead 31 0.002 10 0.001 (2.32) Physically/mentally unable/incompetent 2 0.000 2 0.000 (2.33) Language 84 0.006 177 0.016 (2.36) Miscellaneous 403 0.029 419 0.037 Subtotal 6,320 0.463 5,985 0.529 3. Unknown eligibility, non-interview (3.11) Not attempted or worked 1,069 0.078 719 0.063 4. Not Eligible (4.0) Not Eligible 1,075 0.079 694 0.061 (4.2) Household moved abroad 2,194 0.161 1,150 0.102 (4.4) Household untraceable 2,223 0.163 1,680 0.148 Subtotal 5,492 0.402 3,524 0.311 Total 13,664 1.000 11,324 1.000 Note: Subtotals might not add up because of errors due to rounding. problems with speaking German, condition of the housing area, condition of the house, access problems by barriers, access problems by intercom system, other access problems, safety of the housing area, composition of the housing area, and type of house (according to number of residential parties). Area level information is obtained from INKAR online (Indikatoren und Karten zur Raum- und Stadtentwicklung; www.inkar.de) on the district level. INKAR provides information on (un)employment, construction and housing, education, infrastructure, population characteristics, and other regional indicators. The time reference of INKAR data is 2017. Detailed documentation of the variables in the data is provided by (INKAR, 2019). Lower level information used in the nonresponse analysis is provided by Microm, typically on the street level (www.microm.de). Microm provides information about social structure of neighborhoods in Germany on the regional and local levels. Local level covers different aggregations; for instance, eight digit postal code areas (PLZ8) covering approximately 500 households, street level, or household cells aggregating a few households. Finally, we are able to link some information from the IEB data, such as the date of a person’s first and last spell in the IEBs, the number of spells a person has in the IEBs, the persons date of birth (and the derived age), a person’s nationality, the source a person’s first, last, and most frequent spell originates from, a person’s highest educational degree, and whether or not a person has an apprenticeship spell. SOEP Survey Paper 1105 5 v37 References American Association for Public Opinion Research. (2016). Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys (9th ed.). AAPOR. Brick, J. M., & Kalton, G. (1996). Handling missing data in survey research. Statistical methods in medical research,5(3), 215–238. doi: 10.1177/096228029600500302 INKAR. (2019). Indikatorenübersicht – Indkatoren Raum- und Zeitbezüge. Retrieved from https://www.inkar.de/documents/Indikatoren%20Raum-%20und% 20Zeitbezuege.pdf Kalton, G., & Kasprzyk, D. (1986). The treatment of missing survey data. Survey methodology,12(1), 1–16. Kroh, M., Kühne, S., Goebel, J., & Preu, F. (2015). The 2013 IAB-SOEP Migration Sample (M1): Sampling Design and Weighting Adjustment (SOEP Survey Papers No. 271). Berlin: DIW/SOEP. Kroh, M., Siegers, R., & Kühne, S. (2015). Gewichtung und Integration von Auffrischungsstichproben am Beispiel des Sozio-oekonomischen Panels (SOEP). In Nonresponse bias (pp. 409–444). Springer. Kühne, S., & Kroh, M. (2017). The 2015 IAB-SOEP Migration Study M2: Sampling Design, Nonresponse, and Weighting Adjustment (SOEP Survey Papers No. 473). Berlin: DIW/SOEP. Liebig, S., Brücker, H., Goebel, J., Grabka, M. M., Schröder, C., Zinn, S., … Deutsches Institut Für Wirtschaftsforschung (DIW Berlin) (2021). IAB-SOEP Migrationsstichprobe 2019. SOEP Socio-Economic Panel Study. doi: 10.5684/SOEP.IAB-SOEP -MIG.2019 Oberschachtsiek, D., Scioch, P., Seysen, C., & Heining, J. (2009). Stichprobe der Integrierten Erwerbsbiografien IEBS (FDZ-Datenreport No. 03/2009). Nürnberg: Institut für Arbeitsmarkt- und Berufsforschung. Retrieved from http://doku.iab.de/fdz/ reporte/2009/DR_03-09.pdf R Core Team. (2020). R: A language and environment for statistical computing [Computer software manual]. Vienna, Austria. Retrieved from https://www.R-project.org/ Rathje, M., & Glemser, A. (2021). SOEP-Core – 2020: Report of Survey Methodology and Fieldwork (SOEP Survey Papers No. 1050). Berlin: DIW/SOEP. Retrieved from https://www.diw.de/documents/publikationen/73/diw_01.c.824248 .de/diw_ssp1050.pdf SOEP Survey Paper 1105 12 v37