scieee AI-readable full text Open interactive document viewer

The IAB Employment Subsample 1975-1995

Bender, Stefan,Haas, Anette,Klose, Christoph

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Bender, Stefan; Haas, Anette; Klose, Christoph Article The IAB Employment Subsample 1975-1995 Schmollers Jahrbuch – Zeitschrift für Wirtschaftsund Sozialwissenschaften. Journal of Applied Social Science Studies Provided in Cooperation with: Duncker & Humblot, Berlin Suggested Citation: Bender, Stefan; Haas, Anette; Klose, Christoph (2000) : The IAB Employment Subsample 1975-1995, Schmollers Jahrbuch – Zeitschrift für Wirtschaftsund Sozialwissenschaften. Journal of Applied Social Science Studies, ISSN 1865-5742, Duncker & Humblot, Berlin, Vol. 120, Iss. 4, pp. 649-662, https://doi.org/10.3790/schm.120.4.649 This Version is available at: https://hdl.handle.net/10419/291972 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Schmollers Jahrbuch 120 (2000), 649-662 Duncker & Humblot, Berlin European Data Watch This section will offer descriptions as well as discussions of data sources that may be of interest to social scientists engaged in empirical research or teaching courses that include empirical investigations performed by students. The purpose is to describe the information in the data source, to give examples of questions tackled with the data and to tell how to access the data for research and teaching. We will start with data from German speaking countries that allow international comparative research. While most of the data will be at the micro level (individuals, households, or firms), more aggregate data and meta data (for regions, industries, or nations) will be included, too. Suggestions for data sources to be described in future columns (or comments on past columns) should be send to: Joachim Wagner, University of Lueneburg, Institute of Economics, Campus 4.210, 21332 Lueneburg, Germany, or e-mailed to ([email protected]). The IAB Employment Subsample 1975 -1995 By Stefan Bender, Anette Haas and Christoph Klose 1. Introduction Since 1967 the Institute for Employment Research ("IAB") has been operating as the Federal Employment Services' ("Bundesanstalt fur Arbeit") research institution. The Social Code III (Sozialgesetzbuch - SGB III) is the legal basis for the IAB's research efforts, defining it as one of the obligations of the Federal Employment Services. Because of this reason, the IAB has access to nearly all data sources (e.g. employment and unemployment statistics) of the Federal Employment Services. In general, data from the employment statistics like the IAB employment subsample is subject to confidentiality under Social Code Book X. Passing this data to third parties would have made complicated approval procedures Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 650 Stefan Bender, Anette Haas and Christoph Klose necessary. To avoid such difficulties the IAB decided on the anonymisation of the data with a procedure based on the factual anonymisation under the Federal Statistics Act. The project could only be realized with financial assistance from the scientific community. Along with the anonymising of personand establishment-related cross-section information, the longitudinal information of the persons involved had to be anonymised factually, too. The procedure mainly consisted in the aggregation of the characteristics and in a shift of the complete employment history of each person on the time axis. The anonymised IAB employment subsample is now available for researchers in a third, anonymised version. Following the so-called basic file and the regional file from the IAB employment subsample, which encompassed the years 1975 to 1990, the actualized version of the basic file covers now the years 1975 to 1995 and contains for the first time information on Eastern Germany for the period 1992 to 1995. Therefore, the IAB employment subsample is equipped with data of one percent of all employees registered by the social insurance system within the given period of 21 years. It contains exact daily flow information on the employment history of 559,540 persons as recorded by the social insurance system and on periods of drawing benefits as well. It allows reproducing employment careers without typical problems of longitudinal surveys which do arise in social research (e.g. panel mortality, memory gaps). Nevertheless there are specific problems which appear as a result of the data generation process. The anonymised IAB employment subsample which is available to researchers through the Central Archive for Empirical Social Research at the University of Cologne is, one of the most important German sources for employment market research (Alba et al. 1994). First of all, we will give a short overview of the data generating process and the data characteristics. A discussion of the variables and the anonymisation procedures are given (chapter 3). A brief discussion of data access (chapter 4) will follow and with a conclusion the article ends. 2. The IAB employment subsample - data characteristics The basis of the IAB employment subsample1 is the integrated notifying procedure for health insurance, statutory pension scheme and unemployment insurance which was introduced in West-Germany with effect from 1 January 1973. Since 1. 1. 1991, the notifying procedure has also been ap1 A detailed representation of the IAB employment subsample can be found in Bender et al. (1996, 1999, 2000a). Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 The IAB Employment Subsample 1975 -1995 651 plied to Eastern Germany. The procedure requires that employers report all information of their employees registered by the social security system to the social security agencies. This data has been stored into files, the socalled employment statistics - by the Federal Employment Service (Bundesanstalt fur Arbeit) to provide an insurance account for each employee recorded by the German social insurance system. There are different data sets in the Federal Employment Services, which are summarized under the name employment statistics. It is the current file, the quarterly files (31.3, 30.6, 30.9 and 31.12), the annual period material and the historic file. The historic file is the source for the IAB employment subsample and contains information on employment covered by social insurance which has been passed from social insurance agencies to the Federal Employment Service since the notifying procedure was introduced in 1973. The history file is generated from the employment statistics by storing the notifications submitted in previous years through the notifying procedure in archive storage media at the end of a year in accordance with a defined algorithm. So there are all notifications inside since the introduction of the notifying procedure. "All that has been published so far out of this enormous body of data is in a highly aggregated form: tables of averages by industrial sectors and regions, which are published quarterly, annually, and for longer periods." (Alba et al. 1994: 66). The informative value of the employment statistics depends essentially on the respective definition of employment covered by social security. It should also be taken into consideration that the person providing the information might change during the time, e.g. because of interfirm mobility. This can also lead to implausible information in the employment history (e.g. change of training or nationality). The employment statistics include workers, salaried employees and all trainees, as long as they are not exempt from the obligation to pay social insurance contributions. The employment statistics do not include, among others, civil servants, those in marginal employment, students enrolled in higher education and family workers. For 1995, the employment statistics cover nearly 79.4% of all employed persons in Western Germany, and 86.2% of all employed persons in Eastern Germany,2 but the coverage varies over individual occupations and industries. Furthermore, in empirical analyses, it has to be noted that the definition of variables (e.g. wage) may have changed during the past years. Notifications are prescribed for the beginning and ending of employment, also apprenticeship since 1992. In addition, an annual report must be made 2 Calculated on the basis of employed persons (domestic concept) from the Statistisches Bundesamt, Wirtschaft und Statistik 12 /1997. Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 652 Stefan Bender, Anette Haas and Christoph Klose for each employee covered by social insurance who is employed on the 31 December of the year. These yearly notifications have the function of stocktaking, because the notifying procedure is conceived in such a way that employment must be confirmed in the course of time without any gaps, or concluded by means of a cancellation. This means that for each employment there is a notification, an annual report in case the employment continues over the end of the year, and a cancellation or termination of the employment, all within statutorily prescribed periods. The characteristics sex, year of birth, nationality, marital status, number of children and qualifications, are collected for each employee recorded by social insurance. Exact daily data on employment contains information on the occupational code, the occupational status, the gross earnings to the contribution assessment ceiling, an establishment number issued by the Employment Service, the industry and the size of the establishment.3 The available characteristics can be divided into two categories. On the one hand, some characteristics (insurance numbers, employment duration and earnings covered by social insurance contributions) mainly serve insurance law purposes. These must be notified with each interfirm change, and are, therefore, very accurate. On the other hand, the IAB employment subsample contains characteristics with purely statistical information. Changes to these characteristics are notified at the end of the year with the annual report (Cramer 1985). A systematic 1-% sample was taken from the historic file. This sample covers a period of 21 years for Western Germany (1. 1. 1975 to 31. 12. 1995) and 4 years for eastern Germany (1. 1. 1992 to 31. 12. 1995). The sample's total population consists of all employment cases in occupations covered by social insurance at least once between 1975 and 1995. In the annual averages, the IAB employment subsample includes about 200,000 persons in Western Germany, 52,000 persons in Eastern Germany. It contains exactly a total of 7,847,553 notifications; 559,540 persons from Western Germany, and 354,091 from Eastern Germany, accumulate 6,711,153 notifications. On the basis of the final notification in each case, the file contains 483,327 Western Germans and 76,213 Eastern Germans. This database is supplemented by data on benefits recipients and by establishment information: (1) The benefits recipients file contains person-related information on periods in which the Federal Employment Service paid benefits. These periods are periods in which persons draw substitute wage payments from the Federal Employment Service, such as unemployment benefit, unem3 Regional characteristics are also collected, but these cannot be passed on to users for reasons of data confidentiality. Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 The IAB Employment Subsample 1975-1995 653 ployment assistance or maintenance payments for advanced vocational training or retraining. This does not cover all possible periods of unemployment. Unless specific preconditions are fulfilled a claimant may draw unemployment benefit [so-called "primary" unemployment assistance ("originäre" Arbeitslosenhilfe)]. Unemployment assistance is only paid in the case of personal need. In addition, the file does not record all benefits paid by the Federal Employment Service to persons (cf. Bender et al. 1996: 26). (2) The establishment file contains information on establishments, which are notifying, or have notified, employment covered by social insurance. The Federal Employment Service keeps a file containing all establishment numbers, which the Employment Offices have ever issued to establishments providing employment. When an establishment number is issued, an attempt is made to keep to the concept of the workplace, but this is not carried out on a uniform basis. The establishment characteristic is supplemented by information on the notifying establishments which, like the historic file, is ascertained from the employment statistics collected by the Federal Employment Office. In contrast to the historic file, existing notifications are not simply filed. Additional information on the establishments is generated from the employment notifications submitted by establishments, because the establishment number provides an unambiguous key characteristic for identifying establishments. Using the employment statistics4 as a basis, aggregations of the person-related attributes on school and vocational training are carried out in accordance with the establishment characteristic, particularly establishment-related employment figures have been determined on 30 June each year since 1977. However, the proportion of employees covered by social insurance among all employees depends on the industry. For example, a very high degree of coverage can be assumed for manufacturing industry and a low degree of coverage (up to 50%) in the services industry (cf. Bender et al. 1996). The compressed file, which has 659,221,308 Bytes decompressed, and 131,831,566 Bytes in a compressed format, does not have to be decompressed if the user decides to use the statistic software TDA (Transition Data Analysis; c.f. Rohwer/Pötter 2000). TDA is available in the Internet via ftp://stat.ruhr-uni-bochum.de/pub/tda (WWW address: http://www.stat. ruhr-uni-bochum.de). 4 The notification procedure can be regarded as a continuous complete population survey of employees covered by social insurance. Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 654 Stefan Bender, Anette Haas and Christoph Klose 3. Information in the data - The anonymisation of the IAB employment subsample The data contained in the employment statistics is social data, and is, therefore, subject to confidentiality according to § 35 of the Social Code, Book I. For this reason, the IAB decided on anonymisation; it has kept to factual anonymisation under § 16 (6) of the Federal Statistics Act, and has also used to a great extent the methods applied during the anonymisation of the German microcensus (Miiller et al. 1991, 1995). There is a so called cross-sectional and a longitudinal anonymisation. The cross-sectional information of the variables will be given below, but first a description of the main concept of the longitudinal anonymisation will be given. The aim of anonymisation of the longitudinal information is to make it at least more difficult to date notifiable events exactly. The daily accuracy results directly from the tasks of the notification procedure (e.g. calculating pension rights). In order to prevent changes to the arrangement of events with one another with regard to time and to the duration of episodes, the complete employment history for each employee (all periods of employment and of drawing benefits) was shifted along the time axis by a constant amount to anonymise the longitudinal information. The individual shift constant for the complete employment biography of an individual was stipulated by drawing an independent identical normal distributed random variable with the expectation value zero. Anonymisation of the longitudinal information causes problems for crosssection analyses, which are linked to defined historical periods. However, if massive effects are linked to the key date, for example seasonal effects on unemployment in certain industries, this effect will be smudged with time through the anonymisation and will appear less concise the greater the variance of the random shift is. Coming to the cross-sectional anonymisation for persons, it can be seen that a large part of the person-related cross-section information can be passed in the original state, or with only slight modifications. In the following, the central, person-related characteristics and their modifications resulting from the rules for anonymisation will be shown: The insurance number was replaced by a system-free personal number. A person's sex is a component of the insurance number. This characteristic was not anonymised. The year of birth is also a component of the insurance number. Anonymisation was carried out at each end of the age distribution only. Because this Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 The IAB Employment Subsample 1975 -1995 655 characteristic is time-constant, the first, and the last, notification for a person was included to assess an anonymisation. Persons under 16 (on the first notification), or over 66 (on the last notification) in Western Germany, or 63 in Eastern Germany, were each included in one category. Marital status differentiates between "married" and "single", and is also found in the original status in the anonymised sample. Number of children: the input of the number of children is based on the information on the tax cards and is available as a continuous characteristic. The characteristic is shown in the original up to seven children; from eight children, the category is "eight and more children". With regard to nationality it must be stated that the definition of who is a German is found in Art. 116 (1) of the Basic Law. Coding is carried out in accordance with the personal systematics - nationality and area code of the Federal Office of Statistics. 188 nationalities are altogether found in the IAB employment subsample, nine nationalities and seven nationality groups are shown for Western Germany. Because of the extremely low numbers of non-Germans in Eastern Germany, the only differentiation possible here is between Germans and other nationalities. Education represents the achieved schooling and concluded vocational training in eight categories. The schooling categories "lower and middle secondary school" and "upper secondary school leaving certificate (university entrance level) are subdivided into the education categories "no vocational training" and "completed vocational training". In addition, there are categories for "completed non-university higher education", "university degree", "education not known" and "missing". Anonymisation procedures were not carried out. The occupational code refers directly to the job contents. Allocation in this category is carried out in accordance with the code list from the Federal Employment Service (1975) with a three-digit code. Of the 334 occupations, 234 are retained in their original classification for Western Germany; the remaining 100 were reclassified to form 41 occupational groups. In the case of Eastern Germany, 111 occupations can still be passed on in their original classification; the remaining 224 were aggregated to form 81 occupational groups. The employment status differentiates for those in full-time employment between "apprentice, trainee, student trainee", "unskilled worker", "skilled worker", "master craftsman, foreman", "salaried employee" and "outworker". In addition, a differentiation is made between two groups of part-time employees in accordance with the ratio between contractually agreed and the usual working hours in an establishment; there is no inforSchmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12 656 Stefan Bender, Anette Haas and Christoph Klose mation for these groups regarding their employment status. It was possible to adopt this characteristic without any anonymisation measures. The gross pay subject to social insurance contributions is given for each notification at the end of an employment period and in each annual report. The upper limit is the contribution assessment ceiling for pensions insurance, which is adapted annually to the developments in wages and salaries, and the lower limit is the pay limit for those in marginal employment. A daily income was calculated from the wages sum for a notification period and this is found in an exact sum in DM between the contribution assessment ceilings. In contrast to the older version of the I AB employment subsample, the pay is no longer rounded, but is shown truncated to the respective daily income. There was a change in the definition of income in 1984, so the researcher has to keep this in mind for longitudinal analysis (Steiner / Wagner 1997). The benefits recipients file is used to form the three statements 1. unemployment benefit and comparable benefits (e.g. unemployment allowances, settling-in allowances for the unemployed, severance payments), 2. unemployment assistance and comparable benefits (e.g. unemployment assistance for members of the armed forces on short-term contracts), and 3. maintenance allowances and comparable benefits (e.g. settling-in allowances on further training / retraining, or for German language courses) from over 60 different types of benefit. This large number of benefit types arises from a series of procedural changes, which have their basis in amendments to the Employment Promotion Act ("Arbeitsförderungsgesetz"). There is still a lack of usable, practical systematic knowledge regarding the anonymisation of establishments (cf. Brand et al. 1999). However, it is still assumed that the risk of re-identification, and the benefit of de-anonymisation, increases with the size of an establishment. Establishment characteristics and changes to them arising from the anonymisation rules are briefly shown below: The Employment Offices issue the establishment number to employers. It is replaced by a system-free establishment number. The industry is assigned to a three-digit code of the establishment number and is in accordance with the list of industries used for the statistics of the Federal Employment Service (1973 edition). To protect anonymity, this code is usually shown as two digits; however, the specific nature of the information depends on the size of the establishment and on the differentiation Eastern/Western Germany. Schmollers Jahrbuch 120 (2000) 4 OPEN ACCESS | Licensed under CC BY 4.0 | https://creativecommons.org/about/cclicenses/ DOI https://doi.org/10.3790/schm.120.4.649 | Generated on 2023-04-04 12:28:12