scieee AI-readable full text Open interactive document viewer

SOEPcompanion (v34)

Kara, Selin,Zimmermann, Stefan

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Kara, Selin; Zimmermann, Stefan Research Report SOEPcompanion (v34) SOEP Survey Papers, No. 588 Provided in Cooperation with: German Institute for Economic Research (DIW Berlin) Suggested Citation: Kara, Selin; Zimmermann, Stefan (2018) : SOEPcompanion (v34), SOEP Survey Papers, No. 588, Deutsches Institut für Wirtschaftsforschung (DIW), Berlin This Version is available at: https://hdl.handle.net/10419/191299 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/ SOEP Survey Papers Series G – General Issues and Teaching Materials SOEPcompanion (v34) 588 SOEP — The German Socio-Economic Panel study at DIW Berlin 2018 Selin Kara, Stefan Zimmermann, and SOEP Group Selin Kara, Stefan Zimmermann, and SOEP Group CONTENTS 1 Preface 1 2 Contents of SOEP-Core 2 2.1 Demography and Population ....................................... 3 2.2 Work and Employment .......................................... 4 2.3 Income, Taxes and Social Security .................................... 6 2.4 Family and Social Networks ....................................... 8 2.5 Health and Care ............................................. 11 2.6 Home, Amenities and Contributions of Private HH ........................... 13 2.7 Education and Qualification ....................................... 15 2.8 Attitudes, Values and Personality .................................... 17 2.9 Time Use and Environmental Behavior ................................. 19 2.10 Integration, Migration, Transnationalization ............................... 20 2.11 Survey Methodology ........................................... 21 3 Survey Design 23 3.1 SOEP Questionnaires .......................................... 23 3.1.1 Questionnaires Overview .................................... 25 3.1.2 Household Questionnaire .................................... 25 3.1.3 Individual Questionnaire .................................... 26 3.1.4 Biography Questionnaire .................................... 28 3.1.5 Mother-Child Instruments ................................... 28 3.1.6 Youth Instruments ........................................ 30 3.1.7 Additional Instruments ..................................... 33 3.2 Survey Concepts and Modes ....................................... 34 3.3 Panel Care ................................................ 36 4 Target Population and Samples 37 4.1 The SOEP Samples in Detail ....................................... 37 4.2 Eligibility and Follow-up ......................................... 40 4.3 Development of Sample Sizes ...................................... 41 5 Data Structure of SOEP-Core 44 5.1 Principles of Data Analysis ....................................... 44 5.1.1 Cross-sectional data files (CS) ................................. 44 5.1.2 Data Structure in wide-format (wide) .............................. 45 5.1.3 Data Structure in long Format (long) .............................. 45 5.1.4 Data Structure in spell format (spell) .............................. 45 5.2 Data Distribution File .......................................... 45 5.3 Data Sets SOEP-Core .......................................... 51 i 5.3.1 Data Set Identifiers ....................................... 52 5.3.2 Core Data ............................................ 53 5.3.3 Missing Conventions ...................................... 61 6 Working with SOEP Data 62 6.1 Syntax Generator on paneldata.org ................................... 62 6.2 Working with Tracking Data (PPFAD) ................................. 72 6.3 Generating a cross-section Data Set ................................... 81 6.4 Working with Migration Data (BIOIMMIG) .............................. 89 6.5 Generating a longitudinal Data Set .................................... 98 6.6 Longitudinal Data Analysis ....................................... 109 6.7 Fixed Effects Estimation ......................................... 122 6.8 Working with SOEP Regional Data ................................... 135 7 Working with SOEP Documentation 144 7.1 Variable Search with Questionnaires ................................... 144 7.2 Variable Search with paneldata.org ................................... 146 7.3 Topic Search with paneldata.org ..................................... 155 7.4 Documentation of Generated Data .................................... 162 7.5 Working with SOEPhelp ......................................... 167 8 Contact Information 172 ii CHAPTER ONE PREFACE SOEP-Core is THE centerpiece of the wide-ranging representative longitudinal study of private households located at the German Institute for Economic Research, DIW Berlin. SOEP-Core was started in 1984 and in 1990—just after German reunification—we enlarged the area covered by the SOEP study by adding a representative sample from East Germany. This feature makes the SOEP unique among other household panel surveys worldwide. Each year since 1984, around 15,000 households and about 30,000 individuals have been surveyed by the SOEP’s fieldwork organization, Kantar Public Germany. The data provide information on every member of every household taking part in the survey. Respondents include Germans living in the states of both the former East and West Germany, foreign citizens residing in Germany, recent immigrants, and a new sample of refugees added in 2016. Some of the many topics include household composition, education, occupational biographies, employment, earnings, health, and satisfaction indicators. The SOEPcompanion provides data users with the content of the currently released SOEP-Core data (v34) and will introduce into different versions of SOEP-Core data structures. It will also provide data users with a lot of retrievals in Stata as well as multiple instructions on how to use our various documentation services. This collection of information in the SOEPcompanion is intended to become a yearly updated main reference guide, and a practical companion in basic understanding and implementation of the SOEP. We know that starting to use any new dataset, it is a difficult challenge and this is especially true given the complexity of the panel data. We hope however that this introduction will help. We always welcome any feedback, or information on ways that we can improve this documentation. • Recommendation of our most recent version of a general short description of SOEP study: The German Socio- Economic Panel Study (SOEP) • To the information system for efficient working with complex data sets: paneldata.org 1 CHAPTER TWO CONTENTS OF SOEP-CORE The contents of our questionnaires can be assigned to different question modules, which in turn can be classified into 11 general SOEP topics. There are modules that are subject to permanent social changes and therefore repeat themselves annually, and there are modules that are only asked every few years. The “Replication” column of our topic tables allows a quick assessment of the type of question. In addition, the SOEP provides modules that can appear in adapted form in several questionnaires e.g. the “Big Five Personality Traits” module. From the mother-child instruments to the personal questionnaire, the personality traits are queried and modified according to the target group. SOEP Topics •Demography and Population •Work and Employment •Income, Taxes and Social Security •Family and Social Networks •Health and Care •Home, Amenities and Contributions of Private HH •Education and Qualification •Attitudes, Values and Personality •Time Use and Environmental Behavior •Integration, Migration, Transnationalization •Survey Methodology Overview of Modules across multiple Questionnaires 2 SOEPcompanion (v34), Release 2018, v.1 Individual Youth Mother- Child A Mother- Child B Mother- Child C Parents D Mother- Child E Affective Well-Being x x Big Five Personality Traits x x x x x Birth History x x x x x x Childcare x x x x x Educational Aspirations x x x Health of child x x x x Height and weight of child x x x Height and Weight x x Language Ability German/Mother Tongue x x Leisure and activities (with child) x x Life satisfaction x x Linguistic Usage x x Locus of Control x x Origin x x Parents interest in school performance x x Pocket money x x Political orientation x x Risk Aversion in General x x State of health x x Strength and Difficulties Questionnaire x x Temperament x x 2.1 Demography and Population In this topic you find various information about the birth dates, no matter if interviewer, children, siblings or parents. Furthermore, there is data on places and history of births in households. The household sizes and relationships between the different persons in a household are also listed, as are the sexes of all persons involved. 2.1. Demography and Population 3 SOEPcompanion (v34), Release 2018, v.1 Table 3 – continued from previous page Questionnaire Module Replication No. Vars Variables Relationship to other parent or child annually 1 biochild Mother-Child Questionnaire B (Age 2-3) Childcare annually 10 care1h -care8h &care12h &care19 Leisure and activities (with child) annually 11 activ1 -activ9 &tvyn & tvhrs Linguistic Usage annually 1 language Nursing annually 3 breastf &breastfm & breastfc Mother-Child Questionnaire C (Age 5-6) Childcare annually 10 care1h -care7h &care10h &care12h &care19 Leisure and activities (with child) annually 13 activ1 -activ4 &activ6 -ac- tiv8 &activ10 -activ13 & tvyn &tvhrs Parents Questionnaire D (Age 7-8) Childcare annually 13 maincare &care1h -care7h &care9h -care12h & care19 Parenting Goals annually 18 edgoal1 -edgoal18 Parenting Role annually 9 bepar1 -bepar6 &bepar8 - bepar10 Parenting Style annually 18 edbeh1 -edbeh18 Relationship to other parent or child annually ? Mother-Child Questionnaire E (Age 9-10) Childcare annually 12 maincare &care1h -care5h &care7h &care9h - care12h &care19 Eating Behaviour annually 10 Frage 3? , eatweek1 - eatweek3 &eatsat1 -eat- sat3 &eatson1 -eatson3 Frequency of leisure and activities annually 20 freqact1 -freqact20 Friends annually 2 frndchld &frndadlt Linguistic Usage annually language Parents interest in school performance annually 7 conscho1 -conscho7 Pocket money annually 3 allow &allowpw &allowpm 10 Chapter 2. Contents of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 2.5 Health and Care On the subject of health, numerous personal data such as the number of doctoral visits and habits like sport or alcohol consumption are recorded. There are also information on health insurance, health status and grip strength. However, health information from other people such as children or deceased persons are also displayed. Questionnaire Module Replication No. Vars Variables Household Questionnaire Satisfaction With Availability Of Care 2002 1 hlf0318 Individual Questionnaire Additional Private Insurance 2011-2014, 2016, 2018 8 ple0127 –ple0134 Alcoholic Beverages 2006, 2008, 2010, 2016 2 ple0090 –ple0093 & ple0177, ple0178 Assisted or Curative Care 1999-2011 1 ple0121 Chronicall Illness 1984-1989, 1991, 2009, 2010, 2012, 2014, 2016, 2018 1ple0036 Disabilities in everyday life (SF-12) 1997-2002, 2004-2018 (every two years) 2ple0004 –ple0005 Disability or severe disability annually 2 ple0040-ple0041 Health insurance annually 4 ple0097 ,ple0099 ,ple0104 ,ple0160 Health Insurance Debts 2017 1 Health Restrictions 2011, 2012, 2013, 2015, 2017 2ple0009 &ple0162 Height and Weight 2002 - 2018 (every two years) 2ple0006 –ple0007 Hospital stays annually 3 ple0053 ,ple0055 ,ple0056 Ilness 2009, 2011, 2013, 2015, 2017 14 ple0011 –ple0024 Individual Health Service 2016,2018 1 ple0186 Insurance status 2018 Nutritional Awareness 2004-2016 (every two years) 4 ple0179 – ple0182 Continued on next page 2.5. Health and Care 11 SOEPcompanion (v34), Release 2018, v.1 Table 4 – continued from previous page Questionnaire Module Replication No. Vars Variables Private supplementary care insurance 2016,2018 3 ple0183 – ple0185 Sickness notifications to employer annually 10 ple0044 ,ple0046 ,ple0048 -ple0052 ,ple0174 , ple0175 ,plb0024 Smoking 1998, 1999, 2001, 2002- 2018 (every two years) 6ple0081 &ple0084 - ple0088 &ple0176 State of health annually 1 ple0008 Stress and Exhaustion (SF- 12) 2002 - 2018 (every two years) 10 ple0026 –ple0036 Zype of Disability 2001, 2002, 2004, 2006, 2008, 2010, 2015 2ple0040 –ple0041 Visits to the doctor annually 2 ple0072 ,ple0073 Youth Questionnaire Height and weight annually 2 jl0219 ,jl0220 State of health* annually 1 jl0218 Mother-Child- Questionnaire A (Age 0-1) Health of child annually 12 medaid3mb &lstmedex & disord &disord1 -disord9 Height and weight of child annually 3 weightb &height &weight Physical and Mental Health of Mother annually 4 feeling1 -feeling4 Mother-Child Questionnaire B (Age 2-3) Health of child annually 14 hospital12m &ill2 &ill4 - ill9 &ill11 -ill14 &ill31 & illno Height and weight of child annually 3 weightb &height &weight Mother-Child Questionnaire C (Age 5-6) Health of child annually 11 medaid3m &hospital12m &ill0 &ill2 &ill4 &ill5 &ill7 -ill10 &illno Height and weight of child annually 3 weightb &height &weight Mother-Child Questionnaire E (Age 9-10) Health of child annually 11 chhealth &medaid3m & hospital12m &ill4 &ill5 & ill7 -ill10 &ill32 &illno 12 Chapter 2. Contents of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 2.6 Home, Amenities and Contributions of Private HH In this section you will find information about the household and everything that has to do with everyday life. What kind of home do you live in? Are you an owner or a tenant? Which expenses do you have on things like personal hygiene, the car or holidays? Who’s taking care of the kids? All this and much more information about living, its costs or the living environment can be seen here. Questionnaire Module Replication No. Vars Variables Household Questionnaire Amount of Books in household 2001, 2006, 2011, 2016 1 hlf0197 Apartment equipment annually 13 hlf0023 -hlf0037 & hlf0529 -hlf0531 Apartment owner annually 1 hlf0013 Apartment status annually 4 hlf0001 &hlf0006 & hlf0007 &hlf0009 & hlf0015 Caring Situation for child annually 6 ks_asc_r & kc_relaz & kc_frdn &kc_paid & kc_mindr &kc_none Change of living situation annually 3 hlf0523 &hlf0106 & hlf0107 Cleaning or household assistance annually 2 hlf0261 &hlf0262 Consumption Module 2010 119 hlf0163-hlf0172,hlf0209- hlf0252,hlf0159 ,hlf0371- hlf0434 Hereditary lease interest annually 2 hlf0597 & hlf0598 House type annually 4 hlf0154 &hlf0016 & hlf0155 & hlf0596 Continued on next page 2.6. Home, Amenities and Contributions of Private HH 13 SOEPcompanion (v34), Release 2018, v.1 Table 5 – continued from previous page Questionnaire Module Replication No. Vars Variables Household equipment since last year 1998, 1990, 1992, 1994, 1996, 1998, 2000, 2002, 2004, 2006, 2008, 2010 32 hlf0163-hlf0167,hlf0209, hlf0212,hlf0214-hlf0215, hlf0217-hlf0218,hlf0159, hlf0223,hlf0228-hlf0229, hlf0231,hlc0116-hlc0118, hlf0233,hlf0236-hlf0237, hlf0169-hlf0170,hlf0239- hlf0242,hlf0244-hlf0245, hlf0247-hlf0248 Independent Income of The Children 2016 Leisure Activities and Costs, Children 2006, 2008, 2010, 2012, 2014, 2016 19 ka06_spo,ka06_mus, ka06_art,ka06_oth, ka06_non,ka16_ssp, ka16_smu,ka16_sar, ka16_sth,ka16_sot, ka16_spo,ka16_mus, ka16_art,ka16_org, ka16_yth,ka16_ctr, ka16_non,kk_amtp, kk_cost Living Environment 1986, 1994, 1999, 2004, 2009, 2014 22 hlf0135-hlf0152,hlj0004, hld0001-hld0003 Loans, mortgages, building-society loans annually 2 hlf0087 &hlf0088 Lunch, Childcare 1997, 2002, 2005, 2007, 2011, 2013, 2015 1kd_lunch Lunch, School 1997, 2002, 2005, 2007, 2011, 2013, 2015 1ks_lunch Material deprivation 2016 24 hlf0178-hlf0181 & hlf0186-hlf0195 & hlf0613 - hlf0622 Modernization costs annually 2 hlf0599 & hlf0600 Name and birth of children annually 1 hlk0044 Neighbourhood annually 1 hlf0153 Owner burden annually 1 hlf0606 Owner costs annually 2 hlf0601 - hlf0605 & ‘hlf0090 &hlf0084 Participation (financial reasons) 2001, 2003, 2005, 2007, 2011, 2013, 2015 24 hlf0174–hlf0175 & hlf0439–hlf0444 & hlf0178–hlf195 Persons in need of care annually 22 hlf0291 &hlf0631 hlf0292 &hlf0300 -hlf0304 & hlf0315 &hlf0317 & hlf0319 -hlf0322 & hlf0331 &hlf0332 & hlf0369 &hlf0370 & hlf0446 -hlf0448 & hlf0595 Continued on next page 14 Chapter 2. Contents of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 Table 5 – continued from previous page Questionnaire Module Replication No. Vars Variables Pets 1996, 2006, 2011, 2016 7 hlf0254-hlf0259,hlf0196 Photovoltaic and solar thermal system annually 6 hlf0532 &hlf0535 - hlf0539 Price Comparison Apartment for Rent 1984-2014 1 hlf0094 Reasons for Moving and Comparison 1985-2013, 2015 27 hlf0109-hlf0132,hlf0524- hlf0526 Rely on care times 2002 1 kd_rely Rental and ancillary costs annually 9 hlf0069 &hlf0074 & hlf0078 &hlf0079 & hlf0081 &hlf0082 & hlf0607 & hlf0608 & hlf0610 School attendance for child annually 2 ks_gen & ks_spe Second Residence 2011,2016 3 hlf0156-hlf0158 Size and condition of the house annually 4 hlf0018 &hlf0019 & hlf0071 &hcf0011 Social housing/leased flat at a reduced rate annually 2 hcf0007 &hlf0073 Sponsors and Costs, Childcare 2011, 2013, 2015 7 kd_publ,kd_indep, kd_priv,kd_comp, kd_comm,kc_amtp, kc_cost Sponsors and Costs, School 1987, 1995, 1997, 2002, 2005, 2007, 2011, Mig 2013 7kd_publ,kd_indep, kd_priv,kd_comp, kd_comm,ks_amtp, ks_cost Tenant burden annually 1 hlf0611 Type of Energy used in Household 51 hlf0540-hlf0591 2.7 Education and Qualification Education is one of the cornerstones of our society today, and the information that can be obtained through the SOEP is numerous. Whether school achievement, vocational training or academic success in this section is everything about the education of people. The school history, reasons for lack of further training, educational goals and so on. Furthermore, basic skills of children can be found here to, whether they are able to speak in whole sentences or use scissors, for example. 2.7. Education and Qualification 15 SOEPcompanion (v34), Release 2018, v.1 Questionnaire Module Replication No. Vars Variables Individual Questionnaire Acquired qualification annualy 12 plg0072-plg0079 ,plg0284 ,plg0268, p_degree, p_field Advanced training annualy 3 plg0269 -plg0271 Apprenticeship annualy 1 plg0012 -plg0015 , plg0264 ,plg0265 Continuing education, initiative 1989, 1993, 2000, 2004, 2008, 2014 2plg0273 ,plg0274 Continuing education, reasons for failure 1989, 1993, 2000, 2004, 2014 5plg0277 –plg0281 Further education, course details and motives 1989, 1993, 2000, 2004, 2008 60 plg0108 -plg0122 , plg0129 -plg0149 , plg0152 ,plg0154 , plg0164 ,plg0165 , plg0169 ,plg0171 , plg0172 ,plg0174 - plg0177 ,plg0182 - plg0186 Further education, organizer 1 Further training, financing 1989, 1993, 2000, 2004, 2008, 2014, 2015, 2017 7plg0285 -plg0291 Lifelong learning 2014 1 plg0266 Youth Questionnaire Education and career plans annually 15 jl0177 -jl0187 ,jl0438 -jl0441 ,jl0197 - jl0199 ,jl0201 -jl0205 , j_isco08_jobwish , j_isco88_jobwish, j_kldb2010_jobwish, j_kldb92_jobwish Educational aspirations annually 12 jl0130 ,jl0131,jl0188 - jl0196 ,jl0504 School, attendance & homework annually 36 jl0125 -jl0127 ,jl0132 , jl0133 ,jl0137 -jl0157 , jl0162 -jl0166 ,jl0176 , jl0434 -jl0436 Parents Questionnaire D (Age 7-8) Educational Aspirations annualy 6 idegrad1 -idegrad3 &probgra1 -probgra3 School enrollment annualy 3 sclenrolm &sclenroly & sclenroln Mother-Child Questionnaire E (Age 9-10) Continued on next page 16 Chapter 2. Contents of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 Table 6 – continued from previous page Questionnaire Module Replication No. Vars Variables School and homework annualy 29 sclenrolm &sclenroly &curscol1 -curscol8 & lamark &matmark &nomark &scolcon1 -scolcon7 &hwplace_h &hwsupprt Educational Aspirations annualy 6 idegrad1 -idegrad3 &probgra1 -probgra3 2.8 Attitudes, Values and Personality The character of a person offers a variety of analysis possibilities. Information about the personality of the respondents, their political orientation, concerns, satisfaction, willingness to take risks and much more can be found in the “Attitudes, Values, and Personality” section. Questionnaire Module Replication No. Vars Variables Individual Questionnaire Affective Well-Being annually 4 plh0184-plh0187 Anomie 1992, 1993, 1995, 1996, 1997, 2008, 2013 4plh0188-plh0191 Attitudes towards refugees 2016 11 plj0433 – plj0443 Big Five Personality Traits 2005, 2009, 2013, 2017 16 plh0212 -plh0226,plh0255 Bundestag election 2014 1 plh0333 Control beliefs 2005, 2010, 2015 10 plh0247 –plh0252 , plh0245 ,plh0246 Depressive Traits 2016 4 plh0339 – plh0342 Donation of blood 2010,2015 3 plh0131 -plh0133 Donations 2010,2015 2 plh0129 ,plh0130 Donations of goods 2010 8 plj0108 -plj0115 Flourishing annually 1 plh0334 Goals in life (Kluckhohn) 1990, 1992, 1995, 2004, 2008, 2012, 2016 9plh0105 –plh0112, plh0343 Impulsivity, patience 2008,2013 3 plh0204 ,plh0253 , plh0254 Continued on next page 2.8. Attitudes, Values and Personality 17 SOEPcompanion (v34), Release 2018, v.1 Table 7 – continued from previous page Questionnaire Module Replication No. Vars Variables Income justice, general 2005 12 plh0116-plh0127 Life satisfaction annually 1 plh0182 Loneliness 2013,2017 3 plh0269 -plh0271 Lottery question 2004,2009 1 plh0203 Money and account balance 2016 3 plh0344 – plh0346 Narcissism 2018 Optimism/Pessimism 1999, 2005, 2009, 2014 1 plh0244 Organisational and community membership 1985, 1989, 1993, 1998, 2001, 2003, 2007, 2011, 2015 5plh0263 -plh0267 Policy objectives (Inglehart Index) 1984, 1985, 1986, 1996, 2006, 2016 4plh0054,plh0056,plh0058, plh0061 Political orientation annually 4 plh0007 ,plh0011 - plh0013 Political Tendency, Left- Right 2005, 2009, 2014 1 plh0004 Reciprocity 2005, 2010, 2015 6 plh0206 -plh0211 Religious Affiliation 1990, 1997, 2003, 2007, 2011, 2015 1plh0258 Risk Aversion in Different Domains 2004,2009, 2014 6 plh0197 -plh0202 Risk Aversion in General annually 1 plh0204 Satisfaction with various aspects annually 11 plh0171 -plh0181 Self Esteem 2010, 2015, 2016 1 plh0146 Social responsibility 1987, 1992, 1997, 2002, 2017 11 plh0016 –plh0026 Tendency to Forgive 2010, 2015, 2016 4 plh0142-plh0145 Trust, trustworthiness and fairness 2003, 2008, 2013 8 plh0192 -plh0196 , pld0043 -pld0045 Wage justice 2005, 2007, 2009, 2011, 2013, 2015, 2017 6plh0138-plh0141 ,plh0337 ,plh0338 Well-being aspects 1990 (only Ost), 1994, 1999 13 plh0091 -plh0103 Worries annually 13 plh0032 ,plh0033 , plh0035 -plh0038 , plh0040 ,plh0042 , plh0043 ,plh0046 , plh0047 ,plh0335 , plh0336 10000 Euro Question 2010,2017 3 plh0134-plh0136 Youth Questionnaire Affective Well-Being annually 4 jl0381 -jl0384 Attitudes and opinions annually 4 jl0329,jl0330,jl0360 , jl0364 Big Five Personality Traits annually 17 jl0365 -jl0380 Future annually 11 jl0222 -jl0232 Life satisfaction annually 1 jl0392 Continued on next page 18 Chapter 2. Contents of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 Table 7 – continued from previous page Questionnaire Module Replication No. Vars Variables Locus of control annually 10 jl0350 -jl0359 Political orientation annually 4 jl0388 -jl0391 Risk Aversion in General annually 1 jl0349 Sources of social inequality annually 12 jl0337 -jl0348 Trust annually 3 jl0361 -jl0363 Mother-Child- Questionnaire A (Age 0-1) Temperament annually 5 temp1 -temp5 Mother-Child Questionnaire B (Age 2-3) Big Five Personality Traits annually 4 char1a &char2 -char4 Temperament annually 7 temp1 -temp7 Vineland Adaptive Behaviour Scales annually 20 spch3 &spch5 -spch8 & skll1 -skll5 &mvmn1 & mvmn3 -mvmn6 &sclr2 - sclr6 Mother-Child Questionnaire C (Age 5-6) Big Five Personality Traits annually 10 char1b &char2 -char10 Strength and Difficulties Questionnaire annually 17 behav1 -behav17 Mother-Child Questionnaire E (Age 9-10) Big Five Personality Traits annually 10 char1b &char2 -char10 Strength and Difficulties Questionnaire annually 18 behav1 -behav18 2.9 Time Use and Environmental Behavior Time is a valuable resource for every human being. Information on how a person plans their time, what obligations they have at what time and how they spend their free time can be found in the “Time Use and Environmental Behavior” section. This section also provides comprehensive information on environmental awareness. Which transport infrastructure is used, which energy resources are used to what extent and what is the position on the subject of renewable energies? 2.9. Time Use and Environmental Behavior 19 SOEPcompanion (v34), Release 2018, v.1 • Social housing/leased flat at a reduced rate • Apartment owner • Rental and ancillary costs • Tenant burden • Cleaning or household assistance • Persons in need of care • Name and birth of children • School attendance for child • Caring Situation for child • Income and expenses from rental/lease • Repayments for loans • Credit burden • Inheritance, present, lottery prize • Investments • Income/expanses household • Savings where applicable: +migration specific modules for the IAB-SOEP-Migrationsample or where applicable: +refugee specific modules for the IAB-BAMF-SOEP-Sample of refugees 3.1.3 Individual Questionnaire The individual questionnaire has been a standard instrument since the beginning of the SOEP. In order to enable analyses over time, the individual questionnaire therefore has a large number of question modules which are asked every year. There are also questions that do not have to be asked every year as short-term changes are unlikely. In order to be able to react to current social changes, new topics on the individual questionnaire are added, which are not surveyed annually and are therefore not part of the standard questions of the individual questionnaire. Availability: Since 1984 Respondent: Persons over 18 years in the household The following question modules are part of the core program of the Individual Questionnaire: • Satisfaction with various aspects • Mood • Flourishing • Risk Aversion • Political orientation • Worries • Life satisfaction 26 Chapter 3. Survey Design SOEPcompanion (v34), Release 2018, v.1 • Origin • Apprenticeship • Acquired qualification • Advanced training • Family situation • Family changes • State of health • Disability or severe disability • Visits to the doctor • Hospital stays • Sickness notifications to employer • Health insurance • Employment earnings and collective wage agreements • Additional questions for employees • Additional questions for retirees/pensioners • Transfer payments • Calendar • Use of time • Secondary occupations • Income • Work, last 7 days • Maternity/ Parental leave • Care period (Pflegezeit) • Registered unemployed • Quitting a profession • Employment status • Start of the job • Change of job • Job search • Practiced profession • Current employment • Working hours • Overtime where applicable: +migration specific modules for the IAB-SOEP-Migrationsample or where applicable: 3.1. SOEP Questionnaires 27 SOEPcompanion (v34), Release 2018, v.1 +refugee specific modules for the IAB-BAMF-SOEP-Sample of refugees 3.1.4 Biography Questionnaire Availability: Since 1987 Respondent: Supplementary, one-time data on the personal questionnaire of all persons aged 18 and over in the HH. Content: • Nationality • Origin • Childhood • Parents • Life course since the age of 15 • Education • Occupation • Partnership/ Marriage • Information on children • Siblings where applicable: +migration specific modules for the IAB-SOEP-Migrationsample or where applicable: +refugee specific modules for the IAB-BAMF-SOEP-Sample of refugees 3.1.5 Mother-Child Instruments Mother-Child-Questionnaire A (Age 0-1) Mothers of newborn children primarily answer questions about the course of pregnancy, birth, breastfeeding and the health of the newborn child. It also asks to what extent the mother feels that her life circumstances have changed after the birth of the child, how the care of the child is regulated and how the temperament of the baby (as a precursor of the personality) is perceived by mothers. Availability: Since 2003 Respondent: Mother in household (child age 0-1) Content: • Course of pregnancy • Childbirth • Health screening • Well-being • Childcare • Life circumstances 28 Chapter 3. Survey Design SOEPcompanion (v34), Release 2018, v.1 Mother-Child Questionnaire B (Age 2-3) Mothers of 2-3-year-old children also answer some questions about their child’s health and how long they have been breastfeeding. In addition, the child’s care situation is asked, again the temperament as well as a short scale for recording the personality (agreeableness, extraversion, openness and conscientiousness of the Big Five; McCrae and Costa 1987). In addition, the use of language in the family and activities carried out with the children (e.g. going to the playground, reading or telling stories, visiting other families with children) are recorded. Mothers also assess their children’s adaptive behaviour in the dimensions of communication, everyday skills, social relationships and motor skills. The acquisition is based on a translated version of the Vineland Adpative Behavior Scale, which was reduced to 20 items for the SOEP. This scale thus investigates the stage of development of the infant in everyday life. Availability: Since 2005 Respondent: Mother in household (child age 2-3) Content: • Personality of the child • Well-being • Childcare • Language skills • Development • Abilities Mother-Child Questionnaire C (Age 5-6) The subsequent age-specific survey is carried out as soon as the children turn six years old in the survey year. Among the topics it resembles the surveys conducted in previous years: health, care situation, a more comprehensive battery of items on the personality (from this age neuroticisum is also collected) and activities that are carried out with the child. In addition, there is the Strength and Difficulties Questionnaire (SDQ), which is a shortened version of the German version of the SDQ to 17 items and is a very frequently used instrument for the mental health of children and young people. Availability: Since 2008 Respondent: Mother in household (child age 5-6) Content: • Personality of the child • Activities with children • Well-being • Childcare Parents Questionnaire D (Age 7-8) The questionnaire, which was developed for 7-8-year-old children, is the only age-specific instrument to be completed by both parents, as long as they live together in the same household. In this age range, questions about school attendance (time of school enrolment) and idealistic and realistic educational aspirations become relevant for the first time. However, the focus of this instrument is on the educational goals, parenting styles and the role of both parents. The educational objectives can be differentiated between conformity and autonomy. Educational styles are asked by answering 18 items, which can be divided into six scales: Emotional warmth, inconsistent education, monitoring, 3.1. SOEP Questionnaires 29 SOEPcompanion (v34), Release 2018, v.1 negative communication, psychological control, strict control. The items were taken from the pairfam study, as were the 10 items for recording the role of parents. The parental role can be divided into three scales (autonomy, hostile attributes, willingness to make sacrifices). Availability: Since 2012 Respondent: Parents in household (child age 7-8) Content: • Expectations for school achievements • Expectations of parental educational goals • Upbringing • Parental role • Childcare Mother-Child Questionnaire E (Age 9-10) In addition to the items on health and the care situation recorded in almost all age groups, 9-10-year-old children are asked for more detailed information on the school situation. Here, too, the idealistic and realistic educational aspirations of the mothers for their child are recorded, but also the last grades of the three main subjects, as well as the child’s homework supervision and school motivation. Since friends and leisure activities are gaining in importance in this age group, questions are also asked on these topics. Whether and how much pocket money the child receives will be asked for the first time in this age group. Availability: Since 2012 Respondent: Mother in household (child age 9-10) Content: • Expectations (school achievements, parental educational goals) • Education • parental commitment • Leisure activities for children • Family environment • Social behavior child • Personality Child • Health Child • Supervision • Pocket money 3.1.6 Youth Instruments Pupils Questionnaire In the year in which the children turn twelve, they answer questions about their situation for the first time. Here the focus is once again on the school situation: the start and end of school are asked differentiated according to the days of the week, the type of school attended, the number of pupils in the class and how many of them do not come from Germany, whether one feels discriminated against by the teacher and the last grades in math, German and English. 30 Chapter 3. Survey Design SOEPcompanion (v34), Release 2018, v.1 It also determines how much time the student spends on homework, where he or she does the homework and who helps him or her with the homework and learning. The children are asked about their idealistic and realistic graduation aspiration. Since friends play an important role as caregivers at this age, they and various family members are asked what role they play in the support and how often there are disputes. Also asked about the number of close friendships and how often the parents interfere in the choice of friends. The educational aspirations of the three best friends and a maximum of three older siblings (if any) are asked. The cultural capital and learning environment of the pupils are assessed on the basis of various questions (e.g. availability of literature, instruments, art at home; a desk and a room for oneself). Furthermore, the type and frequency of leisure activities is again asked. The student answers whether and how much pocket money he or she receives and for the first time gives information about his or her own personality, willingness to take risks and life satisfaction. The use of the language in the family (only German or other languages) and with whom the meals are usually taken is also asked. Availability: Since 2014 Respondent: 11-12-year-olds in the household Content: • Attitude • Personality • School (timetable, school-leaving qualification, Engagement) • Recreational activities • Social and family surroundings • Life circumstances Early Youth Questionnaire The questionnaire for early youth is largely similar to the questionnaire for pupils in order to provide an appropriate data structure for questions relevant to developmental psychology. Fewer questions are asked about homework and the learning environment, but the question is asked whether the young person is involved in the school (e.g. as class spokesperson or in a working group) and social capital is acquired in this way. The current importance of various family members and friends is asked and, in addition to their own educational aspirations, also that of the three best friends. With regard to parents, the question is asked how long the young person is allowed to travel and stay up alone before school days and what things the 14-year-old has already done without parents (e.g. holidays, going to the doctor, exchanging something in the shop, drinking alcohol, smoking cigarettes). They ask again for the pocket money and also whether the young person has the opportunity to save money. Another new topic in this age group is the interest in politics and the inclination towards a certain party. Availability: Since 2015 Respondent: 13-14-year-olds in the household Content: • self-perception • School (timetable, school-leaving qualification, Engagement) • Recreational activities • Friends • Siblings • Parents • Pocket money 3.1. SOEP Questionnaires 31 SOEPcompanion (v34), Release 2018, v.1 • Party preferences • Self-Perception • Willingness to take risks • Life satisfaction • Attitudes/Opinions • Future Youth Questionnaire In the SOEP, people who turn 17 in the corresponding survey year are considered adult respondents. Like other firsttime adult participants, you will thus receive a CV and a individual questionnaire. Since part of the adult biography (such as the employment biography or the relationship biography) does not yet apply to the young participants and other aspects such as the relationship with parents, leisure activities, the school situation or vocational training play a greater role, a youth questionnaire was developed in 2000 which replaces the CV questionnaire in this age group and has been used since then. The content of this questionnaire corresponds in many respects to the adult CV questionnaire, so that the data can be used to supplement the information on parents (if they do not live in the household; data set: BIOPAREN). Health status, personality, willingness to take risks, locus of control, trust, time preference, political preferences, knowledge of German as well as information on the living situation, work situation, training, career plans and educational aspirations are also surveyed. For the period from 2000 to 2005, the youth questionnaire was surveyed in addition to the personal questionnaire. Since 2006, only the youth questionnaire has been recorded for 17-year-olds. Since then, it has been available in a version extended by a few indicators, and instead a test has been used to assess cognitive potential. Based on the I-S-T 2000R (Amthauer et al. 2001) the components analogies, number series and matrices with 20 subtasks each were selected for the SOEP (cf. Solga et al. 2005). With the help of these tasks, the fluid cognitive abilities are to be recorded. This is a strongly biologically determined dimension of cognitive abilities that is not influenced by education and is primarily based on reasoning, processing rate and working memory capacity (Cattell 1971; Horn 1982). Although the format of the test differs from the usual questionnaires in surveys, the willingness of young people to participate is high (Schupp and Hermann 2009). Availability: Since 2000 Respondent: 16-17 year olds in the household Content: • Living • Relationships • Leisure and Sport • School (Graduation, Foreign languages, Engagement) • Pocket money • Education • Career Plans • Future • Origin • Childhood and Parental Home • Attitudes/Opinions • Self-Perception • Life satisfaction 32 Chapter 3. Survey Design SOEPcompanion (v34), Release 2018, v.1 • Party preferences where applicable: +migration specific modules for the IAB-SOEP-Migrationsample „Lust auf DJ“ (Denksport und Jugend) Questionnaire In SOEP 2006, a separate questionnaire with cognitive tests for adolescents was used for the first time: “Lust auf DJ”. In this case, “DJ” stands for “Thinking Sports and Youth (Denksport und Jugend)”, but was also specifically selected to arouse the more common association of “Disc Jockey”. For all interviewees aged 16 - 17 years, the questionnaire “Lust auf DJ” was used and created. Availability: Since 2007 Respondent: 16-17-year-olds in the household as a supplement to the youth questionnaire Content: • Assignment of word pairs • Complete incomplete equations • Assign figures 3.1.7 Additional Instruments „Lücke“ Questionnaire - Re-questioning of the Individual Questionnaire (Summary) The “Lücke” (english:gap) questionnaire relates to temporary drop outs for which significant missing data from the previous year are collected. Availability: Since 1987 Respondent: SOEP respondents who are temporarily unavailable. Content: All data refer to the previous survey year • Status of the respondent • Occupational change • Receipt of social benefits within the last year • Completion of education • Type of educational attainment • Change of family status Deceased Persons Questionnaire For the first time in the main wave of 2009, information should be collected on former SOEP participants who have died since the survey in 2008 or until the time of the survey in 2009. Through the questionnaire “The deceased person”, the SOEP curriculum vitae principle is thus consistently “completed”. The primary aim of the chosen concept is to obtain as much information as possible about the death circumstances of former SOEP participants. However, it also generates information about people who have never participated in the SOEP survey. The information collected in this 3.1. SOEP Questionnaires 33 SOEPcompanion (v34), Release 2018, v.1 way about otherwise “unknown” persons, however, can also be used for various analysis purposes on causes of death and the context of death can also be used in the socio-scientific analysis. Availability: Since 2009 Respondent: SOEP respondents who lost a loved one. Content: • Relationship to the deceased • Deceased part of the survey? • Domestic environment of the deceased person • Cause and place of death • Legacies • Health condition of the deceased • Life satisfaction of the deceased • Influence of loss on one’s own life Gripping Strength Test Availability: Since 2008 Respondent: Persons over 17 years in the household Content: This test measures the strength a person can exert when gripping. This can be important for assessing the physical condition. 3.2 Survey Concepts and Modes Measuring stability and detecting changes means to repeat (almost) identical measures over time. Furthermore, the SOEP-questions capture stability and change by varying with regard to the time dimension, asking about events in the past, the present, and the future. Conceptually, different measurements of time are used: • Questions about a point in time (present) e.g. current employment status or current levels of satisfaction • Single retrospective questions on certain events in the past e.g. how often did you change your job during the last ten years? • Retrospective life event history since the age of 15 (in the past) e.g. employment or marital history • Monthly calendar information on income and labor market participation (in the past) e.g. employment status January through December last year • Questions concerning a period of time (in the past) e.g. demographic changes since the last interview like marriage or death of spouse • Questions concerning future prospects (future) e.g. satisfaction with life five years from now, or job expectations Survey Modes The SOEP uses several different modes to collect the data. Originally, the respondent’s answers were recorded by an interviewer who filled in a paper questionnaire, the so called pen-and-paper interview or PAPI. The personal contact between interviewer and respondent is important for the success of the survey; however, before losing a respondent due to a scheduling conflict between interviewer and respondent, the SOEP allows mailing in the questionnaire starting 34 Chapter 3. Survey Design SOEPcompanion (v34), Release 2018, v.1 from the second wave of subsamples A-I. This concept does not resemble the concept of a regular mail survey, because the interviewer still keeps the personal contact with the household and schedules appointments with its respondents if possible. Starting with subsample J, only the computer assisted mode (CAPI) is allowed, and thus mailing in the questionnaires is no longer possible. While the interviewer is in the household she/he directly conducts an interview with any household member, but can also hand out a questionnaire to other household members, who fill it in with or without her/his help (self-administered questionnaires, SAQ). This is much more time efficient for the interviewer, because household members can work in parallel on their questionnaires. In 1998, interviews were conducted with computers for the first time, in computer-assisted personal interviews, or in CAPI mode. Compared to PAPI, CAPI is much more efficient in transferring the data into an electronic format, which was an important asset especially with the extensions of the panel starting in the year 2000. The CAPI mode was first conducted in parallel to the PAPI mode, meaning that interviewers and respondents were free to chose how they wanted to do the interview. This was important for the “older” sample members (respondents as well as interviewers), who were used to the PAPI concept. Only in the most recent samples (starting in subsample J), CAPI is the only mode. The figure depicts the development of modes up to 2011, showing that the CAPI mode has gained importance since its implementation. Since the questionnaires have to be identical in both modes, the CAPI implementation is relatively simple compared to what would be technically feasible. For example, the SOEP basically does not use any form of dependent interviewing (i.e. referring to respondent data from previous waves), because this cannot be easily implemented in the PAPI-mode. Also, the filtering structure is very simple in the SOEP, because any respondent must be able to follow the interview path on her/his own on paper. Still, some technical features like the control of value ranges (e.g. month of birth, year of first marriage) or the randomization of scale items are implemented in the CAPI version of the questionnaire. In the future, new modes will be introduced into the SOEP as they develop. The computer-assisted web interview (CAWI) is close to implementation, it will, however, not be used as a replacement of the current CAPI and PAPI modes, but rather as an extension the respondents may use similar to the mail-in or self-administered questionnaires. The core interview concept of the SOEP survey, the personal contact between respondent and interviewer, will not change. 3.2. Survey Concepts and Modes 35 SOEPcompanion (v34), Release 2018, v.1 Sample Year Households (net) Persons(gross) Respondents (net) Partial Unit Non-Response (percent) Children (gross) A 1984 4528 11422 9076 0.6 2290 B 1984 1393 4830 3169 0.7 1636 C 1990 2179 6131 4453 1.9 1591 D1 1994 236 733 471 2.9 248 D1/D2 1995 541 1668 1078 6.1 517 E 1998 1057 2446 1910 3.5 466 F 2000 6043 14510 10880 5.5 2991 G 2002 1224 3538 2671 6.1 693 H 2006 1506 3407 2616 6.0 623 I 2009 1495 3428 2432 13.4 620 J 2011 3136 6873 5161 9.9 1147 K 2012 1526 3286 2473 9.2 563 L1 2010 2074 7939 3770 6.7 3900 L2 2010 2500 9063 4227 5.1 4611 L3 2011 924 3645 1487 4.2 2092 M1 2013 2723 8522 4964 17.8 2481 M2 2015 1096 3048 1711 19.3 927 M3 2016 1775 4823 2351 22.0 1808 M4 2016 1779 7297 2465 27.1 3915 Cross-Sectional Development of Sample Size (Respondents) Download Stata Code to create figure 42 Chapter 4. Target Population and Samples SOEPcompanion (v34), Release 2018, v.1 This cross-sectional view is insufficient when examining the longitudinal development of the sample, which is influenced by different demographic and field-work related factors. As already shown, demographic reasons for entering the panel are birth and residential mobility. Analogously, the demographic reasons for a panel exit are death and moving abroad. Fieldwork related reasons are different, in that they relate to the interaction between the interviewer and the responding household. Respondents are either not reached for an interview (non-contact) or they decline to participate for the current year. The figure illustrates the longitudinal development of first-wave respondents in 1984, as well as their children, of samples A and B. Longitudinal Development of the 1984 Population Download Stata Code to create figure 4.3. Development of Sample Sizes 43 CHAPTER FIVE DATA STRUCTURE OF SOEP-CORE 5.1 Principles of Data Analysis The data structure for panel data consists of three dimensions. At first, the respective examination units (n) and a matrix of dependent and independent variables (y,x) are completely analogous to a cross-sectional design. Another level is the dimension of time (t), whereby a distinction is made between two data formats for panel data structures - “wide” or “long” (with wide format the variable matrix is indexed with the dimension of time and with long format the respective examination units). Regardless of the selected data format, when using panel data with several survey waves, the data matrices are often not completely provided with information due to the panel mortality of individual survey units or because data from new panel members are only collected at a later point in time. In both cases, the term “unbalanced panel data” is used. In contrast, the classical panel data structure, on the other hand, is “balanced”, i.e. as many observations of dependent and independent variables are available for all study units as there are waves of data collection. The data of social science panel data often show a data structure, which is characterized by many investigation units (large n) as well as, in relation to it, few waves and therefore measuring time (small t). When data from a panel study are available, even descriptive forms of data analysis are often of particular interest, since the identification of changes in a variable over time and the corresponding separation of interindividual and intraindividual changes can represent important social facts, particularly in the case of generalizable samples. It is of social scientific interest whether a constant 15 % proportion of people whose income is below the poverty risk level is repeatedly found in the same person over time, or whether there was a even balance of increases and decreases in poverty risks and only half of the population was permanently exposed to the risk. The choice of complex analysis methods for panel data depends first and foremost on the respective measurement level of the dependent and independent variables, but also on whether they are time-constant variables (such as gender or migration background) or time-invariant variables. The statistical analysis models of panel data range from structural equation models, various regression models, event analysis, sequence data analysis, latent growth models to causal analyses using matching methods. A particular advantage of panel data is that the chronological sequence of changes can be modelled and calculated and the problem of unobserved heterogeneity, which is often encountered in the social sciences, can be significantly reduced, at least in comparison with cross-sectional data. 5.1.1 Cross-sectional data files (CS) Cross sectional data is a type of data, which observes many subjects at the same point of time. Each person is assigned a row in the data set and is only included once in such a data set. By merging cross-sectional SOEP data across waves you receive a dataset in wide-format. Row ID wave sex income 1 1 2015 m 1500 2 2 2015 m 1000 3 6 2015 f 2000 4 8 2015 m 5500 44 SOEPcompanion (v34), Release 2018, v.1 5.1.2 Data Structure in wide-format (wide) The SOEP data is offered in different data structures. In wide format, a respondent’s repeated responses are displayed in a single row and each response in a separate column. Each column represents a variable. We provide four datasets in wide-format: ppfad, phrf, hpfad, hhrf Row ID sex income2015 income2016 income2017 1 1 m 1500 1500 2000 2 2 m 1000 1200 1200 3 6 f 2000 2000 2000 4 8 m 5500 6000 6500 5.1.3 Data Structure in long Format (long) The long format is a compressed and user-friendly data set structure for longitudinal section analysis. Here, each person has one line per survey year. This means that you do not have several data sets for the different waves, but a data set in which all survey waves are represented. A person can occur more than once in such a data set. In long format, one line describes a person-year combination. Row ID syear sex income 1 1 2010 f 1500 2 1 2011 f 1500 3 1 2012 f 2000 4 2 1999 m 1000 5 2 2000 m 1200 5.1.4 Data Structure in spell format (spell) In the strict sense of the word, spell data are about time periods with a defined start and end. When handling spell data it is necessary to take potential censoring into account. Censoring denotes that the beginning (left censored) or ending (right censored) of a spell is imprecise because of missing information or the beginning or ending of a spell is outside of the period of observation. It is quite conceivable that a person has only one spell over a given period, such as a male who is full-time employed. For a ten year period, there may be just the one spell “full-time employed”. In panel data, the same person would have 10 observations, one per year. A person may have many spells over a time period, and even have overlapping spells, like working part-time and receiving a disability pension. Spell data is useful for looking at stays in a certain state, and transitions in and out of that state. Row ID spellnr spelltype begin end censored 1 1 1 Retired 1983 2007 left and right censored 2 1 2 Housewife/husband 1983 1984 left censored 3 1 3 Housewife/husband 1994 1994 uncensored 4 1 4 Housewife/husband 1998 1998 uncensored 5 2 1 Full-Time Employment 1984 1984 left censored 6 2 2 Full-Time Employment 1985 1985 uncensored 5.2 Data Distribution File In the SOEP, each survey year is allocated to a data wave, which is abbreviated with the letters of the alphabet. The current data wave can contain several versions, which are displayed in SOEP with a “v” for version and the respective 5.2. Data Distribution File 45 SOEPcompanion (v34), Release 2018, v.1 version number. The version number represents the survey years since the beginning of the survey. The SOEP has recently published the 34th version since the survey began in 1984. Within a data wave, updates may occur over time, such as v34.1. If updates have been carried out, users are informed about them via various information channels and asked to order the data again. After ordering the data, the data will be sent to you as a zip-file. Within this zip file you will find various data sets, a “RAW” subdirectory and the “EU-SILC Clone” subdirectory. The data sets above the “RAW” subdirectory are highly compressed and an easy to analyze version of the SOEP data. Note: SOEP strongly recommends that users use the data above the “RAW” subdirectory. 46 Chapter 5. Data Structure of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 The data in SOEP-Core are no longer only provided as wave-specific individual files but rather pooled across all available years (in “long” format). In some cases, variables are harmonized to ensure that they are defined consistently over time. For example, the income information provided up to 2001 is given in euros, and categories are modified over time when versions of the questionnaire have been changed. The longitudinal nature is one of the biggest assets 5.2. Data Distribution File 47 SOEPcompanion (v34), Release 2018, v.1 of the SOEP. That‘s why we provide longitudinal data sets, such as pl or hl. The advantage of such a data set is that longitudinal analyses can be carried out without great effort. If you need more information about the long data structure visit the chapter Data Structure in long Format (long). Overview Core Data Sets The data sets above the “RAW” subdirectory: Tracking Data Original Data Survey Data Generated Data Spell Data ppfadl pl csamp pgen artkalen hpfadl hl design hgen biocouplm pbrutto biol exit bioage17 biocouply hbrutto jugendl bioagel biomarsm pbr_exit plueckel kidlong biomarsy abroad pequiv einkalen biobirth lifespell bioedu migspell bioimmig pbiospe biojob refugspell bioparen sozkalen bioresid biosib biosoc biotwin camces cogdj cognit gripstr hconsum health hwealth interviewer mihinc pflege pkal pwealth timepref trust Raw Data Sets In the “RAW” directory you will find all wave-specific data sets that were used to generate the long data sets on the previously presented level. Attention: Please note that the data sets above the RAW subdirectory are completely sufficient for your data analysis. The data sets used to generate the SOEP-Core data can be found in the RAW subdirectory. Detailed 48 Chapter 5. Data Structure of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 Information about the RAW Data Sets can be found here (only available online) Within this “RAW” Directiory each wave is identified by letters of the alphabet: the first wave in 1984 is wave “A”, 1985 is wave “B”, and so on. To simplify the notation, the “$” sign is used, when all waves of one group of datasets are referred to. For example, $H refers to all household level datasets AH to now. For each year of SOEP data there are single data files for households (e.g. $H) as well as for individual respondents (e.g. $P) and children (e.g. $KIND) based on interview information. These observations make up the “net” population, with each of these files containing as many records as interviews could be conducted. Additional data files with a limited number of variables based on 5.2. Data Distribution File 49 SOEPcompanion (v34), Release 2018, v.1 the “address log” constitute the “gross” number of households and persons, i.e. all households and their members which were eligible for an interview in any given year. Within the “RAW” directory, the data sets are stored on a wave-specific basis and are the generation basis for the majority of the long data sets described above. In addition to these wave-specific data sets, the “RAW” directory also contains additional data sets in cross-sectional format that have not yet been distributed in long format ($school, $school2, ev, exit, $pkalost and pbr_hhchch). Tracking Data Original Data Survey Data Generated Data ppfad $p phrf $pgen hpfad $p_mig hhrf $hgen $pbrutto $p_refugees pbr_hhch $kind $hbrutto $pausl $pequiv $pluecke $pkal $h $pkalost $h_refugees $post $jugend $school $school2 ev $vp EU-SILC-Clone Currently, the official German EU-SILC is provided only as a cross-sectional dataset by the German Federal Statistical Office. A panel dataset will presumably be available from the year 2020 onwards (Bundesrat, 2016). As a consequence, Germany is excluded from cross-country studies exploiting the longitudinal dimension of EU-SILC. The aim of the EU-SILC clone is to provide an EU-SILC-like panel dataset for Germany from the year 2005 onwards so that Germany can be included in cross-country studies using EU-SILC panel data. The EU-SILC clone is built on the Socio-Economic Panel (SOEP) and, therefore, includes all EU-SILC panel variables, for which the required information is recorded in the SOEP. The EU-SILC clone includes all of the four EU-SILC sub-datasets: The household register (D-File), the personal register (R-File), personal data (P-File) and household data (H-File). The clone datasets can be combined using the R-File which includes both, the current household ID and the personal ID. ID numbers in the EU-SILC clone are unique and do not vary between the four datasets. A complete documentation of the datasets can be found here: Documentation EU-SILC: 50 Chapter 5. Data Structure of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 5.3 Data Sets SOEP-Core SOEP-Core contains a multitude of different datasets. To get an overview of the data, a somewhat simplified categorization helps: There are Tracking Data and Survey Data files which describe the development of the sample, such that the user knows which person or household was part of the interviewed sample in any given year. Then there are Original Data files, which contain the data from each year’s questionnaires without any changes except for very basic consistency checks. To help the user with the data, there also are Generated Data. These contain consistently coded variables across all waves with common names, such that the users can easily use this information when combining datasets across waves. The SOEP also provides various data on the respondent’s background, called biographical data. Biography data in general can conceptually be separated into biographical data which are unchanging (such as information on parent’s education, or data from the mother-child questionnaires) and data which may be updated through changes in a respondent’s life (such as new children in the birth biography, or a job change in the job history). Some of the changing data is stored as Spell Data. For each spell there is a definition of the spell type, begin, end point and the censoring status, indicating if a given employment or income spell is censored (left and/or right) or uncensored. One of the biggest assets of the SOEP data is their longitudinal nature, i.e. repeated observations of the same unit (person or household) over time. That‘s why we provide longitudinal data sets, such as pl or hl. Finally, there are some files which cannot be easily categorized - some are one-time datasets, some provide information about the interviewers, some about respondents outside of Germany. There are two datasets which should be the building block of any analysis, as they allow to define longitudinal populations very easily: PPFADL and HPFADL. HPFADL includes all households which have been interviewed successfully at least once. Similarly, PPFADL contains all persons who have ever lived in a household that has participated in the SOEP, i.e. that has been captured in HPFADL, including non-respondents and children. Both data files contain one record per household or person, respectively, with wave-specific variables for each year’s survey status. In addition to some time-invariant information (like gender, year of birth, migrant status), these files contain all necessary identifiers to combine other files with PPFADL and HPFADL. Although they provide essential information, PPFADL and HPFADL alone are of little use for actual analyses. The most often used sources for additional information in SOEP-Core are the cross-sectional data files provided in each survey year (or “wave”) or the data sets in the long-format. The SOEP data sets can be viewed based on their content classification (Tracking Data, Original Data, Survey Data, Generated Data and Spell Data), the data structure (cross-sectional (cs), wide, long, spell) and also from the respondent’s perspective. From the respondent’s perspective, data sets can contain gross or net information. In addition, some data sets provide information only at the household level and other data sets provide information at the individual level. 5.3. Data Sets SOEP-Core 51 SOEPcompanion (v34), Release 2018, v.1 bioresid „Generated biographical information“ (CS): In 1994 questions with a focus on occupancy were introduced to the Biographical Questionnaire asking for the duration of residence in the current dwelling and any second residence. The information surveyed in the Biographical Questionnaire is stored in the file BIORESID. biosib „Generated biographical information“ (CS): BIOSIB provides information on siblings living within the SOEP households. The data set contains the person numbers of all siblings in an observed family. It includes information on their sex, their year of birth, the number of siblings, the individual’s position within the birth order, and on the relationship between the observed siblings. biosoc „Generated biographical information“ (CS): contains retrospective data on youth and socialization. Respondents of all ages describe aspects of their life at the age of 15, including their relationship with parents, grades in school, the federal state where they last attained educational qualifications, detailed information on vocational qualifications, as well as intentions to complete further education or vocational training. Questions concerning military and alternative services are also included in this data set. biotwin „Generated biographical information“ (CS): The file BIOTWIN contains all twins that were ever identified within the SOEP. To be classified as a twin, a person is required to have exactly the same age as his or her sibling (year & month of birth), have a relationship to the head of the household that indicates that he or her and a second persons are siblings, and have the same mother (as far as a pointer to the mother is available). Furthermore, it is not only twins that are recorded in the BIOTWIN data set, but also triplets or quadruple siblings. camces „Highest Educational Qualification, Migrants Sample M1 and M2“ (CS): The CAMCES-File provides information about Computer-Assisted Measurement and Coding of Educational Qualifications in Surveys. cogdj „Data on cognitive tests (Youth)“ (CS): In SOEP 2006, a separate questionnaire with cognitive tests for adolescents was used for the first time: “Lust auf DJ”. In this case, “DJ” stands for “Thinking Sports and Youth (Denksport und Jugend)”, but was also specifically selected to arouse the more common association of “Disc Jockey”. For all interviewees aged 16 - 17 years, the questionnaire “Lust auf DJ” was used and created. cognit „Data on cognitive potential“ (long): In the 2006 survey year, for the first time, short cognitive tests were carried out with a subsample of the SOEP. The goal was to employ a robust set of instruments that could be administered easily by trained interviewers within just a few minutes. Im COGNIT06 werden den Nutzern die aggregierten Summen- Scores (jeweils Gesamtwerte für drei Zeitpakete, sog. „parcels“ von 30, 60 und 90 Sekunden) zur Verfügung gestellt. gripstr „Measures grip strength (left and right hand)“ (long): The data on grip strength from the survey year 2012 is now included in the GRIPSTR dataset. hconsum „HH consume module“ (CS)“: We were faced with three methodological challenges in generating the final consumption data. Firstly, due to the design of the consumption module, inconsistent answers arose between the monthly and annual amounts spent for consumption. Secondly, we encountered the well-known phenomenon of missing data, here in particular item nonresponse. And thirdly, consumption data are usually blurred by heaping. For researchers who do not want their consumption variables to include changes from all steps of data preparation, the new data set “HCONSUM” contains not only the prepared consumption variables but also flag variables providing researchers the opportunity to select individual solutions. health „Data on health indicators“ (long): Starting in 2002 the SOEP health module in the individual questionnaire has been revised and put into a two year replication period. In the HEALTH-File users find i.e. the generated variables on height and weight with imputation flags and a user-friendly longitudinal checked generated variable of the Body Mass Index (BMI). hwealth „Wealth module“ (long): The generated SOEP wealth data is stored in two separate data files called PWEALTH for information at the individual level and HWEALTH for correspondingly aggregated data at the household level. HWEALTH contains all information on the household level; it is purely the result of aggregating the person-level information in PWEALTH. However for all persons with valid household level information that did refuse to respond to the Individual questionnaire (partial unit non-response) imputations have been carried out and the results are included in HWEALTH. interviewer „Data on the SOEP Interviewer“ (long): The SOEP does not only aim at collecting high-quality data on the living conditions and well-being of households, but –as a by-product of internal quality assurance processes– it lends 58 Chapter 5. Data Structure of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 itself increasingly as a empirical source for survey research. The purpose of the INTERVIEWER file is to provide user convenient access to all available, longitudinal information on the SOEP interviewers. mihinc „Multiple imputed data on monthly household income (long)“: The dataset MIHINC contains the complete imputation results and is separately available. To be compatible with methods for analysing multiply imputed data, MIHINC is constructed in the so called stacked or MIM Dataset Format. It contains the following variables: HHNRAKT, SVYYEAR, MJ, MI, IHINC and IMPFLAG. Since 1995 for every survey household in all survey years there are ten imputed values for the current household income. pflege „Persons needing care within the household“ (long): Since wave B (1985) the SOEP household questionnaire includes questions on household members in need of care. In order to support analyses on an individual level, this information has been restructured and stored in the cumulative file PFLEGE. pwealth „Wealth module“ (long): In the year 2002, the individual questionnaire included for the first time a special module focusing on wealth. This section included questions on seven different wealth components: Owner-occupied property (including debt), other property (including debt), financial assets, private pensions (including life insurance and building savings contracts), business assets, tangible assets and consumer credit. The generated SOEP wealth data is stored in two separate data files called PWEALTH for information at the individual level and HWEALTH for correspondingly aggregated data at the household level. Wealth-related variable names in the file PWEALTH consist of six digits. The first digit tells the user which wealth component is referred to, and the second to sixth digits provide more detailed information about possible filter information, the personal share, the gross amount, and the amount of any outstanding debt. In principle a digit is coded “1” if a given variable does indeed contain this specific piece of information and “0” otherwise. The wealth information in the SOEP questionnaire is surveyed at the individual level and thus also imputed or edited at the individual level (although checked against household information for consistency). timepref „Experiment on time preferences“ (CS): Following on the behavioral experiment on trust and trustworthiness carried out in the 2003, 2004, and 2005 SOEP surveys, the experiment “time preferences” was run in 2006. In this experiment on economic behavior, respondents were asked to decide how they would want to receive C200 in prize money: if they would want to receive it immediately by check, or if they would want to wait and receive a larger amount later—that is, with interest. trust „Experiment on trust“ (long): Data set of the economic behavior experiment on trust and trustworthiness from the survey years 2003, 2004 & 2005, which serves to measure trust, based on an investment game. This is a one-off game for two actors who relate to each other anonymously. The first player receives a credit of ten points and can overwrite any number of points of the second player. Each overwritten point is doubled. The second player also receives a credit of ten points. After receiving the (doubled) points from the first player, it decides how much of its own credit it will transfer to the first player (zero to ten points). As with the first transfer, your points at the recipient are doubled. After the decision of the second player, the game ends and the other players are paid their income (one point corresponds to one euro, the sum is sent out as a cheque a few days later). The TRUST data set thus contains the information from all three waves in which the behavioral experiment was conducted. Spell Data Spell, duration or event history data are used frequently in the social sciences. In the strict sense of the word, spell data are about time periods with a defined start and end. General information about the data structure of spell data can be found in the chapter Data Structure in spell format (spell) Working with spell data: Working with spell data (pdf): Working with spell data (do-files): How to generate spell data from data in wide format: Based on the Migration Biographies of the IAB-SOEP Migration Sample: Generating spell data: 5.3. Data Sets SOEP-Core 59 SOEPcompanion (v34), Release 2018, v.1 Dataset Label Format Identifier (ID) Additional Identifier artkalen Spell data from the activity calendar spell pid cid biocouplm Generated biographical information spell pid cid, coupid biocouply Generated biographical information spell pid cid biomarsm Generated biographical information spell pid cid biomarsy Generated biographical information spell pid cid einkalen [deprecated] Spell data on income spell pid cid lifespell Spell Information on the Pre- and Post-Survey History of SOEP-Respondents spell pid cid migspell Migration history spell pid cid pbiospe Generated biographical information spell pid cid refugspell Migration history spell pid cid sozkalen [deprecated] Spell data on social benefits spell hid, cid artkalen „Spell data from the activity calendar“ (long): The ARTKALEN contains spells (monthly) for events starting in January 1983. This is in contrast to PBIOSPE, where spells were in yearly durations, and events previous to 1983 were included. The information on activity status are collected on a monthly basis in the yearly Individual questionnaire and stored in the file ARTKALEN. biocouplm „Generated biographical information“ (long): With the BIOCOUPLM the SOEP provides consistent and continuous partnership histories for nearly all adult respondents. BIOCOUPLM is build on the prospective information at the time of each interview. The relationsship histories are collected on a monthly basis from all adult SOEP- participants since their entry into the SOEP. biocouply „Generated biographical information“ (long): With the BIOCOUPLY the SOEP provides consistent and continuous partnership histories for nearly all adult respondents. BIOCOUPLY is build on retrospective and prospective information at the time of each interview. The relationsship histories are provided on an annual basis. biomarsm „Generated biographical information“ (long): With BIOMARSM the SOEP provides consistent and continuous marital histories for nearly all adult respondents. BIOMARSM is build on the prospective information at the time of each interview. The martial histories are collected on a monthly basis from all adult SOEP-participants since their entry into the SOEP. biomarsy „Generated biographical information“ (long): With BIOMARSY the SOEP provides consistent and continuous marital histories for nearly all adult respondents. BIOMARSY is build on retrospective and prospective information at the time of each interview. The marital histories are provided on an annual basis. einkalen „[deprecated] Spell data on income“ (long) The income calendar is used to gain information about sources of income throughout the year. The respondent checks off for each month all appropriate sources of income. lifespell „Spell Information on the Pre- and Post-Survey History of SOEP-Respondents” The SOEP team regularly conducts drop-out studies to identify the whereabouts of attritors. These studies draw on official register data and allow us to determine whether a person is still living in Germany, is deceased, or has moved abroad since the last SOEP interview. The information is combined in a spell file LIFESPELL. This dataset reports all available information on the pre- and the post-survey history of all persons who have ever been a member of a SOEP household. migspell „Migration history“(long): MIGSPELL is derived from the migration biographies, which are collected from each new respondent of the IAB-SOEP migration samples M1 and M2. It contains data on the moves of foreign-born migrants as well as on the stays abroad of German-born respondents. pbiospe „Generated biographical information“ (long): The spell file PBIOSPE is based on the information on activity status over the life course, which is collected as a matrix from every respondent answering the Biography Questionnaire. The observations start at the age of 15 and end at the current age (up to age 65). To update the ongoing occupational career in PBIOSPE, information from the yearly Individual Questionnaire is also used. refugspell „Migration history“ (long): For migration biographies in the refugee samples, we created the spell data 60 Chapter 5. Data Structure of SOEP-Core SOEPcompanion (v34), Release 2018, v.1 set REFUGSPELL. The variables in MIGSPELL and REFUGSPELL are derived from different instruments and only partially overlap. The data structure allows the data set to be linked with MIGSPELL if desired. 1992-2000 sozkalen „[deprecated] Spell data on social benefits“: The file SOZKALEN provides spell data on receiving social assistance of households, defining begin, end, and censoring status of any period of receiving 3 different types of assistance. This file is set up, using information from the calendar, asked for the previous year (asked for the years 1992-2000). Thus, it contains information on a monthly basis. 5.3.3 Missing Conventions Survey variables might be missing, i.e. without a valid code or value for different reasons. In the SOEP, negative values are not valid for any variable, but are used instead to code different reasons for missing information. There are two distinctions for missing values: they may originate in the respondent’s answer or in the survey design. The respondent may refuse or not know an answer or she may report invalid values on the one hand, and the interview design may exclude respondents with certain characteristics from some questions on the other (e.g. men will never be asked if they are pregnant). The following codes are used: Code Label -1 no answer / don’t know -2 does not apply -3 implausible value -4 Inadmissable multiple response -5 Not included in this version of the questionnaire -6 Version of questionnaire with modified filtering -8 Question not part of the survey program this year1 1Only applicable for datasets in long format. A person might refuse to answer a question, which happens more often in sensitive questions (e.g. income related questions), or may just not know the answer to a question. In such a case, the missing code is “-1” for “no answer / don’t know”. Note that the SOEP does not distinguish between the refusal to answer and a true “don’t know”. Information may be missing when a question is not asked because it is not relevant for a specific person, e.g. owneroccupiers will not be asked about the amount of rent they pay. In such cases, the question “Does not apply” to this person, and the variable receives a code of “-2”. Sometimes invalid answers are encountered, when respondents fill out a PAPI interview themselves or the interviewer mistypes an answer, e.g. persons cannot work more than 168 hours a week. In such a case, multiple checks are carried out, and if the inconsistency remains, the variable is recoded “-3 Implausible value”. Some questions contain multiple answer possibilities, where the respondents are asked to pick one and only one answer. In the SOEP PAPI instruments, sometimes respondents ignore this request and provide more than one answer, e.g. they mark “very good” and “good” when asked about their current health status. In such cases, if the correct answer cannot be determined from the questionnaire itself, the code “-4 Invalid Multiple Answers” is given to this variable. With the extension of the SOEP in recent years, entirely new samples have been added to the core. In these samples, sometimes questions are left out completely, e.g. to shorten the questionnaire or because the focus of the sample is different as in some of the related studies. In such a case, the variable will be set to “-5 Not included in this version of the questionnaire” for an entire subsample. With the use of CAPI, recent developments include an “integrated” person questionnaire, i.e. the biography part and the “regular” part of the questionnaire are asked as one. Some of the questions in the biography part are repeated in the regular part. While in the PAPI mode, the respondent will answer the same question twice, the CAPI allows to filter the respondent around the question if it has already been asked. These cases are very rare - if they occur, they receive a code “-6 Version of questionnaire with modified filtering”. 5.3. Data Sets SOEP-Core 61 CHAPTER SIX WORKING WITH SOEP DATA The following exercises are taken from our SOEP Campus Workshops. A service especially for young scholars in the disciplines of sociology, economics, and psychology. Here we provide introductions to the use of the SOEP data. 6.1 Syntax Generator on paneldata.org Paneldata allows registered users to collect and save their research-relevant variables in a variable basket. These variables can be simply written into a single data set with the script generator. The script generator helps you with data management and can save valuable working time. Open Paneldata For our experienced users, we have temporarily equipped the old soepinfo with the current data, so that the variables basket function and the script generator can also be used there. Open soepinfo 62 SOEPcompanion (v34), Release 2018, v.1 Click on the “Register/ log in” to log in to paneldata.org. 6.1. Syntax Generator on paneldata.org 63 SOEPcompanion (v34), Release 2018, v.1 If you have already registered, you can login in the “User login” area. As a new user you can register at “Register here”. Once you have logged in successfully, you have access to the variable basket and the syntax generator. 64 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 To access the activated functions, click on the navigation field “My baskets”. You will be taken to your personal workspace on paneldata.org. “My baskets” displays your created variable baskets. If you click on “Create basket”, you can create a new basket. 6.1. Syntax Generator on paneldata.org 65 SOEPcompanion (v34), Release 2018, v.1 When creating the basket, first define the name of the variable basket. The name must be lower case to be accepted by Paneldata. Optionally, you can assign a label and enter a description. You can create a security key via “Security token”. Finally, you select the study that you want to use as a database for your research. Now click on “Create basket” and your newly created variable basket appears in the interface. Now search for your relevant variables on paneldata.org and add them to your individual basket. For example, you are interested in the monthly net household income. If you do not know the variable name, you can find the superordinate concept using the topic search. Click on the navigation field “paneldata.org” to get to the main page. Select the study SOEP-Core and click on the navigation field “Topics”. 66 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 Check the different topics for income-relevant concepts and select “income, taxes, and social security”. 6.1. Syntax Generator on paneldata.org 67 SOEPcompanion (v34), Release 2018, v.1 b) Were these people and their parents born in Germany? In the data set, search for a variable that describes the migration background. Display the information of the variable for persons 2102 and 19202. 1*b) Were these people and their parents born in Germany? 2list persnr migback if persnr == 2102 | persnr == 19202 c) If they have immigrated: In which year and from which country? Search the data set for a variable that describes the country of birth and the year of moving to Germany. Display the information of the variables for persons 2102 and 19202. 1*c) If they have immigrated: In which year and from which country? 2list persnr immiyear corigin if persnr == 2102 | persnr == 19202 d) Are these people from East or West Germany? Search the data set for a variable that describes east-west affiliation. Display the information of the variables for persons 2102 and 19202. 74 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 1*d) Are these people from East or West Germany? 2list persnr loc1989 psample if persnr == 2102 | persnr == 19202 e) From which sources does the information on the migration background and the year of death come? Search the data set for info variables that show you sources of information for the year of death and the migration background. Display the information of the variables for persons 2102 and 19202. 1*e) From which sources does the information on the migration background and the year ˓→of death come? 2list miginfo todinfo if persnr == 2102 | persnr == 19202 2. How many people lived in a realised private household in 2016 and answered the individual questionnaire? Remember that the wave-specific survey year in SOEP is abbreviated with letters. SOEP started in 1984 (wave a) and was in a survey wave “bg” in 2016. For more information on this topic, please refer to the DTC subchapter Label. If you are interested in the 2016 survey year, the wave name indicates that you should be interested in variables with the abbreviation “bg”. Search the data set for variables with the abbreviation “bg” that describe the population. Display the characteristics of the population variables: 1******************************************************************************** 2*** Exercise 2) *** 3*How many people lived in a realised private household in 2016 and answered the 4*personal questionnaire? 5 6******************************************************************************** 7 8*informationen from: 9*2016 -> Wave bg 10 *private household -> bgpop 11 *Individual questionnaire -> bgnetto 12 13 tab bgpop 6.2. Working with Tracking Data (PPFAD) 75 SOEPcompanion (v34), Release 2018, v.1 Values 1 and 2 are relevant to answer the question because they describe realized households. Search the data set for variables with the abbreviation “bg” that describe the survey status. Display the characteristics of the survey status: 1tab bgnetto 76 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 Respondents with survey status between 10 and 15 or survey status 19 completed the individual questionnaire. Crosstab the variables bgpop and bgnetto with an appropriate restricting condition to answer the question. 1tab bgnetto bgpop if ((bgnetto >= 10 &bgnetto <= 15)|bgnetto==19)&(bgpop==1| ˓→bgpop==2) 3. PPFAD allows you to see which populations can be viewed from a longitudinal perspective: 6.2. Working with Tracking Data (PPFAD) 77 SOEPcompanion (v34), Release 2018, v.1 a) How many people who answered the individual questionnaire in 2000 also took part in the survey in 2014? Remember that the wave-specific survey year in SOEP is abbreviated with letters. SOEP started in 1984 (wave a) and was in a survey wave “bg” in 2016. For more information on the subject, see the subchapter Label. The wave name shows that you are interested in the survey years 2000 and 2014. The survey years include the wave names “q”(2000) and “be”(2014). Search the data set for variables with the abbreviations “q” and “be” that describe the survey status. Display the characteristics of the survey status under the condition that the individual questionnaire has been answered: 1*a)How many people who answered the personal questionnaire in 2000 also took 2*part in the survey in 2014? 3 4*informationen from: 5*2000 -> wave q 6*2014 -> wave be 7*Individual questionnaire -> $netto 8 9tab qnetto benetto if qnetto>=10 & qnetto<=19 & benetto>=10 & benetto<=19 10 *or: 11 //fre qnetto benetto if qnetto>=10 & qnetto<=19 & benetto>=10 & benetto<=19 A total of 7639 respondents completed the individual questionnaire in 2000 and 2014. b) How many people answered the individual questionnaire every year from 2000 to 2014? The survey years include the wave designations from “q”(2000) to “be”(2014). View the relevant survey status codes to answer the question. Please consider all persons who have answered the individual questionnaire: 1*b) How many people answered the individual questionnaire every year from 2000 2*to 2014? 3 4/*to see all the codes */ 5lab list bgnetto 78 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 Define a variable list that shows all survey statuses ($netto) of the 15 survey waves considered in total. 1local v "netto" 2local vlist "q`v' r`v' s`v' t`v' u`v' v`v' w`v' x`v' y`v' z`v' ba`v' bb`v' bc`v' bd`v ˓→' be`v'" (continues on next page) 6.2. Working with Tracking Data (PPFAD) 79 SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 3/*--> 15 waves */ Generate a variable that shows the number of waves of completed person interviews. Note that the values 10,12,13,14,15,16,18,19 of the $netto variable mean realized interviews. 1capture drop h1 2egen h1 = anycount(`vlist'), values(10 12 13 14 15 16 18 19) Display a table with its newly generated variable. 1tab h1 if h1 == 15 A total of 6665 people completed the individual questionnaire every year over the period 2000-2014. c) How many people who turned 15 in 2011 and lived as children in a survey household took part in the survey in 2016? The survey year 2011 is represented by the wave “bb” and the survey year 2016 is represented by the wave “bg”. To answer the question, a variable must be generated that identifies people who were 15 years old in 2011. The age of the respondent can be determined with the year of birth and you can limit children using the net code. Generate a variable with people who turned 15 in 2011 and lived in a survey household as a child. 1*c) How many people who turned 15 in 2011 and lived as children in a survey 2*household took part in the survey in 2016? 3 4*informationen from: 5*2011 -> wave bb 6*Age -> 15 7*Child -> bbnetto 8*2016 -> wave bg 9*Individual Questionnaire -> bgnetto 10 11 /*People who turned 15 in 2011 and lived in a survey household as a child...*/ 12 capture drop a15kind 13 gen a15kind = 1 if 2011-gebjahr == 15 & bbnetto >= 20 & bbnetto < 30 14 In order to identify all persons who were 15 years old in 2011, lived in a survey household as a child and completed the individual questionnaire in 2016, you must use the net codes again. Create a table from the net code of 2016 to narrow down the cases appropriately. 1// fre bgnetto if a15kind == 1&bgnetto >= 10 &bgnetto <20 2*oder: 3tab bgnetto if a15kind == 1&bgnetto >= 10 &bgnetto <20 4 80 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 In 2016, a total of 309 people who were 15 years old and were part of a survey household as a child in 2011, completed a individual interview. d) The person with persnr=588010 was born in 1984 in a panel household and was still part of the sample in 2009. The person has changed households twice during this time. In which years? To identify how often and when a person has changed the household, you must display all available household numbers in ppfad for person 588010. 1*still part of the sample in 2009. The person has changed households twice during 2*this time. In which years? 3 4*Information from: 5*-> household numbers 6 7list *hhnr if persnr == 588010 8/*-> changed household 9in year d (1987) 10 in year y (2008) 11 no participation since bb (2011) 12 */ The person 588010 has participated in the survey since the wave “b” (1985) in household 58807. From wave “d” (1987) to wave “x” (2007) the person was in household 73407, from wave “y” (2008) the person was in household 132608. 6.3 Generating a cross-section Data Set This example involves generating a data set to analyze health satisfaction determinants in 2008, and you can either use the Paneldata.org syntax generator or write your own syntax file to perform this task. You can search for the variable names in Paneldata.org (or use the variables below directly). 1. Generate a cross-section dataset for the year 2008, which should contain all persons with the following characteristics: 6.3. Generating a cross-section Data Set 81 SOEPcompanion (v34), Release 2018, v.1 • Respondents in 2008 "ynetto" • Lives 2008 in private household "ypop" The data set should contain the following variables of interest. • Satisfaction with health "yp0101" • Smoking currently yes/no "yp10601" • current employment status "emplst08" • monthly household net income "hinc08" In addition, the data set should contain the following additional information for a 2008 cross-sectional analysis (these variables are automatically generated by paneldata.org): • Current cross-section weighting factor "yphrf" • Personal number "persnr" • Original household number "hhnr" • Current household number "yhhnr" • Sample affiliation "psample" • Gender "sex" • Year of birth "gebjahr" Create an exercise path with four subfolders: Example: • H:/material/exercises/do • H:/material/exercises/output • H:/material/exercises/temp • H:/material/exercises/log These are used to store commands, log files, data sets and temporary data sets. Open an empty do file and define your created paths with globals: 1*********************************************** 2*Set relative paths to the working directory 3*********************************************** 4global AVZ "H:\material\exercises" 5global MY_IN_PATH "\\hume\rdc-prod\complete\soep-core\soep.v33.2\stata_en\" 6global MY_DO_FILES "$AVZ\do\" 7global MY_LOG_OUT "$AVZ\log\" 8global MY_OUT_DATA "$AVZ\output\" 9global MY_OUT_TEMP "$AVZ\temp\" 82 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 The global „AVZ“ defines the main path. The main paths are subdivided using the globals “MY_IN_PATH”, “MY_DO_FILES”, “MY_LOG_OUT”, “MY_OUT_DATA”, “MY_OUT_TEMP”. The global “MY_IN_PATH” contains the path to your ordered data. Use ppfad as the source file together with the required variables. Keep all cases with completed interviews. In addition, your data set should only contain respondents who can make a statement on the content of the question. For example, you can use the net code to identify and remove children from your data set. 1***PFAD *** 2 3use hhnr persnr sex gebjahr psample yhhnr ynetto ypop using "${MY_IN_PATH}ppfad.dta" 4 5 6***BALANCED VS UNBALANCED *** 7 8keep if ( (ynetto >= 10 &ynetto <20) ) 9 10 11 ***PRIATVE VS ALL HOUSEHOLDS *** 12 13 keep if ( (ypop == 1|ypop == 2) ) 14 15 16 ***SORT PFAD *** 17 18 sort persnr 19 save "${MY_OUT_TEMP}ppfad.dta", replace 20 clear Save the modified data record temporarily. Now link your data set with the weights of the SOEP and save your data set as a master file. 1***HRF *** 2 3use "${MY_IN_PATH}phrf.dta" 4sort persnr 5save "${MY_OUT_TEMP}hrf.dta", replace 6clear 7 8 9***CREATE MASTER *** 10 11 use "${MY_OUT_TEMP}ppfad.dta" 12 merge 1:1persnr using "${MY_OUT_TEMP}hrf.dta" 13 drop if _merge == 2 14 drop _merge 15 sort persnr 16 save "${MY_OUT_TEMP}master.dta", replace 17 clear Now prepare the content variables. Search for the content variables you are looking for from the various data records and temporarily save the created data records. 1***READ DATA *** 2 3use hinc08 yhhnr using "${MY_IN_PATH}yhgen.dta" 4sort yhhnr 5save "${MY_OUT_TEMP}yhgen.dta", replace (continues on next page) 6.3. Generating a cross-section Data Set 83 SOEPcompanion (v34), Release 2018, v.1 In the following, we will use this record and other information from the SOEP to create a status variable that you can use to distinguish whether or not people with a migration background also have an escape background. Create an exercise path with four subfolders: Example: • H:/material/exercises/do • H:/material/exercises/output • H:/material/exercises/temp • H:/material/exercises/log These are used to store commands, log files, data sets and temporary data sets. Open an empty do file and define your created paths with globals: 1*********************************************** 2*Set relative paths to the working directory 3*********************************************** 4global AVZ "H:\material\exercises" 5global MY_IN_PATH "\\hume\rdc-prod\complete\soep-core\soep.v33.2\stata_en\" 6global MY_DO_FILES "$AVZ\do\" 7global MY_LOG_OUT "$AVZ\log\" 8global MY_OUT_DATA "$AVZ\output\" 9global MY_OUT_TEMP "$AVZ\temp\" The global „AVZ“ defines the main path. The main paths are subdivided using the globals “MY_IN_PATH”, “MY_DO_FILES”, “MY_LOG_OUT”, “MY_OUT_DATA”, “MY_OUT_TEMP”. The global “MY_IN_PATH” contains the path to your ordered data. Task 1: Preparation of BIOIMMIG a) In which variable can you find information about the status of each person when they immigrated to Germany? Open the record or browse the BIOIMMIG documentation and search for a variable describing the immigration status. The biimgrp variable from the BIOIMMIG data set is the appropriate variable. 1*** Exercise 1 ****************************************************************** 2 3/* 4a) In which variable can you find information about the status of each person ˓→when they immigrated to Germany? 5*/ 6 7*Immigration status is stored in the variable biimgrp. 8 9use $MY_IN_PATH\bioimmig.dta, clear b) Identify this variable in the BIOIMMIG data set and load it from the data set, together with the person number and the survey year. 90 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 Open your data set only with the required variables to maintain clarity in your analysis data set. 1/* 2b) Identify this variable in the BIOIMMIG data set and load it from the data ˓→set, together with the person number and the survey year. 3*/ 4 5use persnr syear biimgrp using $MY_IN_PATH\bioimmig.dta, clear c) What are the values of this variable? Familiarize yourself with your research-relevant analysis variable and check coding and case numbers. 1/* 2c) What are the values of this variable? 3*/ 4 5tab biimgrp, m //Characteristics of the variable are examined. d) On the basis of this variable, generate the variable “Escape”, which only distinguishes between three groups: • 0 = Cases where no information is available • 1 = All persons without escape background • 2 = Asylum seekers / fugitives After you have familiarized yourself with the research-relevant analysis variable, recode the variable to suit your project. Then check the case numbers of your generated variable with the source variable. 1/* 2d) On the basis of this variable, generate the variable "Escape", which only ˓→distinguishes between three groups: 30=Cases where no information is available 41=All persons without escape background 52=Asylum seekers /refugees 6*/ 7 (continues on next page) 6.4. Working with Migration Data (BIOIMMIG) 91 SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 8recode biimgrp (-5-2-1=0"No Answer") (12346=1"no Escape") (5=2"Escape ˓→"), gen(Escape) 9tab biimgrp Escape, m // biimgrp and escape are compared. e) It may happen that initially there is no information on the status of immigration, but this will change in a later year. Limit the data record to the last observation that is available for the respective person, since this way the specification with the most information content is used. 1e) It may happen that tinitially there is no information on the status of 2*immigration, but this will change in a later year.Limit the data record to 3*the last observation that is available for the respective person, since this 4*way the specification with the most information content is used. 5*/ 6 7bysort persnr: egen syear_max =max(syear) //A variable is created, which shows the ˓→last existing yearly observation 8keep if syear_max == syear //Annual observations which are not the last observation ˓→are deleted. f) Save the generated data record on your personal drive temporarily . 1f) Save the generated data record on your personal drive temporarily 2*/ 3 4save $MY_OUT_TEMP\biimgrp.dta, replace Task 2: Add basic variables from PPFAD and weights a) Load the following information from PPFAD: • Never changing Person ID "persnr" • Household number "hhnr" and the current household number "bghhnr" 92 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 • The net variable with information about the interview type "bgnetto" • The sex of the person "sex" • The year of birth "gebjahr" • Variables on the migration background "migback","germborn","corigin","immiyear" • Information about the survey status: "psample" If you want to familiarize yourself with the PPFAD data set, visit the chapter Working with Tracking Data (PPFAD). 1/* 2a) Use the following information from PPFAD: 3- Never changing Person ID „persnr“ 4- Household number "hhnr" and the current household number "bghhnr". 5- the net variable with information about the interview type "bgnetto". 6- the sex of the person "sex" 7- the year of birth "semester" 8- Variables on the migration background "migback", "germborn" "corigin" "immiyear" 9- Information about the survey status: "bgnetto" and "psample". 10 */ 11 12 use persnr hhnr bghhnr bgnetto psample sex gebjahr germborn corigin immiyear migback ˓→using $MY_IN_PATH\ppfad.dta, clear b) Merge the previously generated data record using the person number. If you don’t understand how to create your own cross-section dataset, visit the chapter Generating a cross-section Data Set. 1/* 2b) Merge the previously generated data record using the person number. 3*/ 4 5merge 1:1 persnr using $MY_OUT_TEMP\biimgrp.dta, nogen c) Add the corresponding person extrapolation factors to the data record. 1c) Add the corresponding person extrapolation factors to the data record. 2*/ 3 4merge 1:1 persnr using $MY_IN_PATH\phrf.dta, keepus(bgphrf) nogen d) Only keep respondents for whom a youth or individual questionnaire was realized in 2016. For example, to exclude children who have not provided immigration status information, use the net code from PPFAD. Only keep persons who have conducted a completed individual or youth interview. 1/* 2d) Only keep individuals for whom a youth or personal questionnaire was ˓→realized in 2016. 3*/ 4 5tab bgnetto, m //Variable values are displayed 6 7keep if inrange(bgnetto, 10,19)// People who have a code between 10 and 19 will be ˓→kept. 6.4. Working with Migration Data (BIOIMMIG) 93 SOEPcompanion (v34), Release 2018, v.1 Task 3: Generate a status variable with the following categories:. • No immigrant background • Migration 2nd generation • Immigration without information • Immigration, not flight • Immigration, Flight To generate this status variable, check the contents of the existing migration variables from PPFAD (migback germborn). 1/* 2Generate a status variable with the following categories: 3*/ 4 5tab migback 94 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 1tab germborn Use the migration variables from PPFAD (migback, germborn) and link this information with your previously generated escape variable to build the described status variable from Task 3. 1gen Status =0// All persons will first receive the missing code for "no info". 2replace Status =1if migback == 1&germborn == 1// "no migback" 3replace Status =2if migback == 3// "2nd generation" (2nd ˓→generation migrants born by definition in Germany, therefore "& germborn == 1" here ˓→unnecessary 4replace Status =3if germborn == 2&Escape == 0// "Immigrants without information" 5replace Status =4if germborn == 2&Escape == 1// "Immigrants, no escape" 6replace Status =5if germborn == 2&Escape == 2// "Immigrant, escape" 7 8label def Statuslbl 0"no info" 1"no migback" 2"2. Generation" 3"Immigrants without ˓→information" 4"Immigrants, no escape" 5"Immigrant, escape" 9label val Status Statuslbl // Values of the status veriable receive label Task 4: Content analysis: a) How many refugees (foreign-born with refugee/asylum titles) are now in your record? Look at your status variable previously generated in task 3 to answer the question 1*** Exercise 4 ****************************************************************** 2 3/* 4a) How many refugees (foreign-born with refugee/asylum titles) are now in your ˓→record? 5*/ 6 7tab Status, m //Display Generated Status Variable 6.4. Working with Migration Data (BIOIMMIG) 95 SOEPcompanion (v34), Release 2018, v.1 All 4,514 respondents who received the value 5 for the generated status variable have a direct migration background (migback==2), were not born in Germany (germborn==2) and fled their home country (flight==2 and biimgrp==5). b) How many are there if you take the person extrapolation factors into account? Interpret the results. Look at your status variable previously generated in task 3 to answer the question 1/* 2b) How many are there if you take the person extrapolation factors into ˓→account? Interpret the results. 3*/ 4 5tab Status [aw=bgphrf], m //Display generated status variable weighted with analytic ˓→weights After weighting, there are only about 675 fugitives in the data set. The weighting thus corrected the number of fugitives downwards. c) How many persons are represented by the sample taking the extrapolation factors into account? To use frequency weights in STATA, integer weights are required. Create an integer frequency weight from the weighting factor provided so that you can make representative statements. Then take a look at the new results. 1/* 2c) How many persons are represented by the sample taking the extrapolation ˓→factors into account? 3*/ 4 (continues on next page) 96 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 5gen fweight = round(bgphrf) //Frequency weights for stata require integer weight 6tab Status [fw=fweight], m //Display generated status variable weighted with ˓→frequency weights Around 1,600,000 people are represented. d) What is the proportion of people over 40 years of age among the fugitives? Since the data in this exercise come from the wave “bg”, we are currently in the survey year 2016; if you need a description of the wave designations, please refer to the chapter Label. To generate a suitable age variable, you can use the year of birth (year of birth). If we look at the survey year 2016, all persons born in 1976 or earlier were over 40 years old. Generate a suitable age variable and look at the proportion of fugitives over 40 years of age in weighted form: 1/* 2d) What is the proportion of people over 40 years of age among the fugitives? 3*/ 4 5gen ue_40 = 0 6replace ue_40 = 1 if gebjahr <= 1976 // Persons receive proficiency 1 if they were ˓→born before 1975. 7 8tab Status ue_40 [aw=bgphrf], m row nofreq The proportion of refugees over 40 years of age is about 47%. 6.4. Working with Migration Data (BIOIMMIG) 97 SOEPcompanion (v34), Release 2018, v.1 6.5 Generating a longitudinal Data Set This example is about generating a data set to analyze determinants of health satisfaction. You can either use the syntax generator of paneldata.org or write a syntax file yourself. You can search for variable names in Paneldata.org. In the previous examples you have already created an exercise path with four subfolders, as well as corresponding globals in the STATA do-file. You can use the same folders and globals for this exercise. 1.Generate an unbalanced panel dataset for the years 2006 to 2008 using paneldata.org if you wish. The data set should contain all respondents in private households: The data set should contain the following variables of interest: • Health satisfaction "wp0101" "xp0101" "yp0101" • Smoking at present yes/no "wp9301" "yp10601" • Current employment status "emplst06" "emplst07" "emplst08" • Monthly household net income "hinc06" "hinc07" "hinc08" In addition, the data set should include the following additional information for analysis from 2006 to 2008: • Cross-sectional weighting factors for all relevant years "wphrf" "xphrf" "yphrf" • Person ID "persnr" • Original household number "hhnr" • Household number for all relevant years "whhnr" "xhhnr" "yhhnr" • Sample membership "psample" • Sex "sex" • Year of birth "gebjahr" • population membership "wpop" "xpop" "ypop" If you need detailed instructions on how the script generator works in paneldata.org, you can find them in the chapter Syntax Generator on paneldata.org. If you would like to assemble your data set yourself, you can do this with the data sets you have supplied. From the previous exercise with tracking data, you may already have an idea where to get most of the variables. Since we want to have an unbalanced panel record, the $netto variable for the years 2006 to 2008 must also be used. In addition, our analysis must limit population membership, as we are only interested in household respondents. Tip: If a data set is created from several variables of different data sets, it is worth sorting the person number before saving the individual data sets in order to be able to merge the data sets more easily afterwards. 1.1. Create a Master-Files Use ppfad as the source file together with the required variables that you may have already researched in Paneldata or identified from the variable label of the data set. Note that only variables of the years to be analyzed should be used. 1 2use hhnr persnr sex gebjahr psample xhhnr xnetto xpop yhhnr ynetto ypop whhnr wnetto ˓→wpop using "${MY_PATH_IN}ppfad.dta" 3 98 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 Since we want to receive an unbalanced data set, i.e. persons who have completed a personal questionnaire at least once within the 3 years, you must restrict the variable $netto (survey status). Also, we only want to analyze private households, so we need a further restriction of the $pop (sample membership) variable. 1 2keep if ( (xnetto >= 10 &xnetto <20)|(ynetto >= 10 &ynetto <20)|(wnetto >= 10 ˓→&wnetto <20) ) 3 4 5***PRIVATE VS ALL HOUSEHOLDS *** 6 7keep if ( (xpop == 1|xpop == 2)|(ypop == 1|ypop == 2)|(wpop == 1|wpop == 2) ˓→) 8 Then we sort the persnr (personal number) of the data record and save it. 1 2sort persnr 3save "${MY_PATH_OUT}ppfad.dta", replace 4clear 5 What is still missing is the cross-section weighting factor and the variables of interest in terms of content. To apply the weighting factors to the data set, open the weighting data set for the person level phrf, sort it and save it again. 1 2use persnr wphrf xphrf yphrf using "${MY_PATH_IN}phrf.dta" 3sort persnr 4save "${MY_PATH_OUT}phrf.dta", replace 5clear 6 Now we come to the variables of content. In order not to have to click through all delivered data sets, it is recommended to enter the label of the variable of interest on paneldata.org. Use the filter to narrow your search. Select our main study SOEP-Core, the search type “variable”, the analysis unit “p” or “h” and the corresponding year. Once you have clicked on the year of interest, a variable history is displayed. You can use this to see in which years the variable was collected and what the variable is called. Example: Variable Label „Satisfaction Health“ 6.5. Generating a longitudinal Data Set 99 SOEPcompanion (v34), Release 2018, v.1 Since you have previously added the SOEP weighting factors to your analysis data set, you should use the weighting for a representative analysis. 1tab diff if year==2007 [aw=phrf] /*weighted*/ 106 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 The values less than 0 show a deterioration in health satisfaction. The value 0 means a constant health satisfaction and all values above 0 show a positive change in satisfaction with their health. With a value of 10, it can be assumed that these people were interviewed for the first time in 2007 or 2008. c. In what direction and how much has satisfaction with the health of people who quit smoking after 2006 changed from 2006 to 2008? The procedure is similar to the previous question, except that the element “smoke yes/no” is added. 1*c) In what direction and how much has satisfaction with the health of 2*people who quit smoking after 2006 changed from 2006 to 2008? 3 4gen diff2=health-health[_n-2] if persnr==persnr[_n-2] & year==year[_n-2]+2 & ˓→year==2008 5gen quit=. 6replace quit=0 if smoke==1 & smoke[_n-2]==1 & persnr==persnr[_n-2] & ˓→year==year[_n-2]+2 & year==2008 7replace quit=1 if smoke==2 & smoke[_n-2]==1 & persnr==persnr[_n-2] & ˓→year==year[_n-2]+2 & year==2008 8replace quit=2 if smoke==2 & smoke[_n-2]==2 & persnr==persnr[_n-2] & ˓→year==year[_n-2]+2 & year==2008 9replace quit=3 if smoke==1 & smoke[_n-2]==2 & persnr==persnr[_n-2] & ˓→year==year[_n-2]+2 & year==2008 (continues on next page) 6.5. Generating a longitudinal Data Set 107 SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 10 label define quit 0 "smoker" 1 "quit" 2 "non-smoker" 3 "begin" 11 label values quit quit 12 tabstat diff2, by(quit) To obtain a weighted mean value, address the analysis weight after the generated variable. 1tabstat diff2 [aw=phrf], by(quit) /*weighted*/ This illustration shows the mean of the health variable under the condition of the variable quit we generated beforehand. With a mean of -0.24 (weighted -0.35) the biggest change in health satisfaction is seen in people who quit smoking after 2006. For example, if a person smoked in 2006 and indicated a satisfaction value of 8, the person after he/she stopped smoking in 2008 indicates a satisfaction value of 7.76. So you can assume that when a person stops smoking, the state of health that a person perceives deteriorates. Now we have to test if the assumption is correct. d. Does quit smoking make your health worse? To what extent can the result of the analysis “Stop smoking” be distorted? 108 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 In order to establish a connection between health satisfaction and stopping smoking, one should use the ttest or to be more specific, the one-sample t test. It checks whether the mean value of a sample deviates significantly from a known expected value (specified in the null hypothesis). 1*d) Does quitting smoking make your health worse? To what extent can the 2*result of the analysis "Stop smoking" be distorted? 3 4*Notes: So far we have not tested whether the difference is statistically ˓→significant 5ttest diff2==0 if quit==1 H0 Hypothesis: If one stops smoking it has no effect on health. For this test we assume a 95% probability. What we want to check now is whether the H0 hypothesis can be rejected or not. If you look at the output of the test, you first see the mean value of value 1 (quit smoking) of the variable quit. The last line of the output shows the significance level. If it falls below the value 0.05, one can speak of a statistically significant result. In our example, the null hypothesis can be discarded because its value is less than 0.05 percent. So quitting smoking has a significant impact on a person’s perceived health. 6.6 Longitudinal Data Analysis Simple cross section analyses show that married people have a higher life satisfaction than singles. You want to check this on the basis of longitudinal analyses with the SOEP. Create an exercise path with four subfolders: Example: • H:/material/exercises/do 6.6. Longitudinal Data Analysis 109 SOEPcompanion (v34), Release 2018, v.1 • H:/material/exercises/output • H:/material/exercises/temp • H:/material/exercises/log These are used to store your script, log files, datasets and temporary datasets. Open an empty do file and define your created paths with globals: 1*********************************************** 2*Set some useful commands 3*********************************************** 4version 13 5clear all 6set more off 7**increase buffer size 8set scrollbufsize 2000000 9**now restart stata! 10 11 *********************************************** 12 *Set relative paths to the working directory 13 *********************************************** 14 global AVZ "H:\material\exercises" 15 global MY_IN_PATH "\\hume\rdc-prod\distribution\soep-long\soep.v33.1\stata_en\" 16 global MY_DO_FILES "$AVZ\do\" 17 global MY_LOG_OUT "$AVZ\log\" 18 global MY_OUT_DATA "$AVZ\output\" 19 global MY_OUT_TEMP "$AVZ\temp\" The global „AVZ“ defines the main path. The main paths are subdivided using the globals “MY_IN_PATH”, “MY_DO_FILES”, “MY_LOG_OUT”, “MY_OUT_DATA”, “MY_OUT_TEMP”. The global “MY_IN_PATH” contains the path to your ordered data. Create a master file that uses the important variables from ppfadl. You should always add some variables from PPFADL to your data set by default. Download the following information from PPFADL: • Person ID "pid" • Household number "pid" • Survey year "syear" • The net variable with information on the interview type "netto" • The weighting variable "phrf" • The sex of the person "sex" • The migration background "migback" 1*------------------------------------------------------------------------------- 2*** Step 1) Start with basic information from PPFADL *** 3 4use pid hid syear netto phrf migback sex using ${MY_IN_PATH}\ppfadl.dta Search for matching variables and add them to your data set To perform your analysis, you need different SOEP variables. The SOEP offers various options for a variable search: • Search the questionnaires for useful variables. (for more information visit the chapter Variable Search with Questionnaires) 110 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 • Find a suitable variable via the topic list of paneldata.org (for more information visit the chapter Topic Search with paneldata.org) • Search for a suitable variable using a search term in paneldata.org (for more information visit the chapter Variable Search with paneldata.org) • Use the documentation provided by the generated variables (for more information visit the chapter Documentation of Generated Data) In this case you need the variables "pgfamstd" (martial status) and "plh0182" (life satisfaction). 1*------------------------------------------------------------------------------- 2*** Step 2) Add the relavant variables: here: family status and life satisfaction *** 3merge 1:1 pid syear using ${MY_IN_PATH}\pgen, keepusing(pgfamstd) keep(1 3) ˓→nogen 4 5// merges family status from pgen 6// Documentation for PGEN can be found here 7// http://panel.gsoep.de/soep-docs/surveypapers/diw_ssp0307.pdf) 8 9 10 *describe using pl (directory) 11 // for checking out variable names without opening the dataset 12 13 merge 1:1 pid syear using ${MY_IN_PATH}\pl, keepusing(plh0182) keep(1 3) nogen 14 // merges life satisfaction from pl 15 16 save $MY_OUT_DATA\ppfad.dta, replace Clean and inspect the data Recode all missings into the format of a point. 1*------------------------------------------------------------------------------- 2*** Step 3) Clean and inspect the data 3mvdecode _all, mv(-8/-1) Since you are interested in individual characteristics in your analysis: Delete all measurements that are not based on successful personal interviews. 1tab netto 2drop if netto>19 6.6. Longitudinal Data Analysis 111 SOEPcompanion (v34), Release 2018, v.1 How many people contribute measurements and what is the proportion of people contributing at least 10 measurements? Define the data set as a panel data set. 1**define the data set as panel data (continues on next page) 112 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 2xtset pid syear 3xtdes 86079 respondents have contributed information within waves a (1984) - bg (2016) and 75% of the 86079 respondents have provided information for at least 10 waves How many people took part in the survey in 2010 and contributed to continuous measurements until 2014? 1xtdes if syear>=2010 &syear<=2014 6.6. Longitudinal Data Analysis 113 SOEPcompanion (v34), Release 2018, v.1 14673 respondents provided continuous information from 2010 to 2014. Univariate inspection & analysis How does the mean of life satisfaction change over time? 1*------------------------------------------------------------------------------- 2*** Step 4) univariate inspection &analysis 3table syear, content (mean plh0182) 114 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 How high is the proportion of people who will be a) married in 2014 or b) have a migration background. Compare weighted with unweighted frequency tables: Which people are overrepresented in SOEP? 1tab1 pgfamstd migback if syear==2014 2tab pgfamstd [aw=phrf] if syear==2014 3tab migback [aw=phrf] if syear==2014 6.6. Longitudinal Data Analysis 115 SOEPcompanion (v34), Release 2018, v.1 The graph shows that a positive effect on life satisfaction can be observed when the family status changes from single to married. In the following years of the existing marriage, life satisfaction decreases again and approaches the initial satisfaction before the marriage. 6.7 Fixed Effects Estimation You want to find out whether certain variables relevant to the labour market, such as work experience or education time, influence a person’s hourly wage. Other variables such as gender or marriage status should also be taken into account. You decide to use the SOEP data to set up a fixed effects estimation model. Create an exercise path with four subfolders: Example: • H:/material/exercises/do • H:/material/exercises/output 122 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 • H:/material/exercises/temp • H:/material/exercises/log These are used to store your script, log files, datasets and temporary datasets. Open an empty do file and define your created paths with globals: 1*********************************************** 2*Set relative paths to the working directory 3*********************************************** 4global AVZ "H:\material\exercises" 5global MY_IN_PATH "\\hume\rdc-prod\distribution\soep-long\soep.v33.1\stata_en\" 6global MY_DO_FILES "$AVZ\do\" 7global MY_LOG_OUT "$AVZ\log\" 8global MY_OUT_DATA "$AVZ\output\" 9global MY_OUT_TEMP "$AVZ\temp\" The global „AVZ“ defines the main path. The main paths are subdivided using the globals “MY_IN_PATH”, “MY_DO_FILES”, “MY_LOG_OUT”, “MY_OUT_DATA”, “MY_OUT_TEMP”. The global “MY_IN_PATH” contains the path to your ordered data. a) Generate your own SOEPWage.dta data set. The data set should contain information on gross monthly wage, marital status and other personal characteristics. To perform your analysis, you need different SOEP variables. The SOEP offers various options for a variable search: • Search the questionnaires for useful variables. (for more information visit the chapter Variable Search with Questionnaires) • Find a suitable variable via the topic list of paneldata.org (for more information visit the chapter Topic Search with paneldata.org) • Search for a suitable variable using a search term in paneldata.org (for more information visit the chapter Variable Search with paneldata.org) • Use the documentation provided by the generated variables (for more information visit the chapter Documentation of Generated Data) Use the various important variables of the ppfadl.dta data set as your start file. Your source file should contain the following variables: • Person ID "pid" • Survey year "syear" • Birth Year "gebjahr" • The net variable with information on the interview type "netto" • The weighting variable "phrf" • The sex of the person "sex" • Sample Membership "pop" 1use pid syear sex gebjahr netto pop phrf using "${MY_IN_PATH}/ppfadl.dta", clear Apply the necessary content variables to your starting data set. You need the following variables for your analysis: • Employment Status "plb0022" • Current Gross Labor Income in Euro "pglabgro" • Actual Work Time Per Week "pgtatzeit" 6.7. Fixed Effects Estimation 123 SOEPcompanion (v34), Release 2018, v.1 • Working Experience Full-Time Employment "pgexpft" • Amount Of Education Or Training In Years "pgbilzeit" • Marital Status In Survey Year "pgfamstd" 1merge 1:1pid syear using "${MY_IN_PATH}/pl.dta", keepus(plb0022) keep(master match) ˓→nogen 2merge 1:1pid syear using "${MY_IN_PATH}/pgen.dta", keepus(pglabgro pgtatzeit pgexpft ˓→pgbilzeit pgfamstd) keep(master match) nogen Only keep people who have completed an interview and who live in a private household. 1*Only select people with completed interviews 2keep if inrange(netto, 10,19) 3 4*Only private households 5keep if pop==1|pop==2 Since you are only interested in the period from 2012 to 2016 in your analysis, remove all survey information that does not fall within this period. To finish, save your data set. 1*Period from 2012 to 2016 2keep if syear>=2012 &syear<=2016 Exercise 1: Prepare your data set a) Load your created SOEPWage.dta data set. The data set contains information on gross monthly wage, marital status and other personal characteristics. 1*** Exercise 1: Prepare your data set 2*a) Load data set 3use "${MY_OUT_DATA}/SOEPWage.dta", clear b) Recode all missing values in Stata Missings (.) 1*b) Recode Missings 2mvdecode _all, mv(-8/-1= .) For more information about the missing codes of SOEP data visit the chapter Missing Conventions c) Generate the variables “hourly wage” (gross monthly wage/4.33*working time) for persons who have earned at least 1 Euro and have worked at least one hour, “Married vs. Unmarried” and age. 1*c) Generate Variables 2gen wage =pglabgro/(4.33*pgtatzeit) if pglabgro>=1&pgtatzeit>=1 3 4gen married =1if pgfamstd==1|pgfamstd==6|pgfamstd==7|pgfamstd==8 5replace married =0if inrange(pgfamstd, 2,5) 6 7gen age =syear -gebjahr d) Adjust the variable “hourly wage” from outlier values by setting values smaller than the 1st percentile to the same value. Set values greater than 3 times the 99th percentile to 3*99th percentile. Then generate the variable lwage = log(wage). 1*d) Adjust wage variable 2sum wage, detail 3replace wage =1/3*r(p1) if wage<1/3*r(p1) (continues on next page) 124 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 4replace wage =3*r(p99) if wage>3*r(p99) &wage<. 5 6gen lwage =log(wage) 7label variable lwage "Log hourly wage" 8 9save "${MY_OUT_DATA}/SOEPWage_temp.dta", replace Exercise 2: Descriptive statistics a) Define the data set as a panel data set. 1*** Exercise 2: Descriptive statistics 2*a) 3xtset pid syear // Declaring data as panel data b) What percentage of people participate in all five waves (xtdescribe) 1*b) 2xtdescribe, patterns(16)// -> unbalanced panel 6.7. Fixed Effects Estimation 125 SOEPcompanion (v34), Release 2018, v.1 42808 respondents have contributed information within waves bc (2012) - bg (2016) and about 40% (17069) of the 42808 respondents have provided information for all waves. c) Describe the variable “Married” with xttab and xttrans. Take a look at some individual wage (pid=30320901, pid=30932501, pid==3101602, pid==3101801) developments with xtline. 1*c) 2*Stability of the relationship status 3xttab married 126 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 You can observe 41.37 percent of person-year observations with Married==No. At least once 19717 people within the period from 2012 to 2016 have stated not to have been married. 25014 persons reported to have been married at least once during this period. Those who were not married for at least one year responded with “married==no” in 94.69% of the observations. Whereas those who have been married at least once responded in 95.88 percent of the observations with”Married==Yes”. A very stable response behaviour can therefore be observed. 1*Transition probabilities 2xttrans married, freq 96.87 percent of the person-year observations with “married==no” are also not yet married in the next period. 98.51 percent of the persons who are married indicate that they will also be married in the following period. A stable behaviour of the respondents can be seen. 1*Individual sequences of "wage" 2xtline wage if pid==30320901 |pid==30932501 |pid==3101602 |pid==3101801, overlay 6.7. Fixed Effects Estimation 127 SOEPcompanion (v34), Release 2018, v.1 The graphic shows a comparison of the hourly wage for four different respondents. Exercise 3: Pooled OLS Regression a) Execute a pooled OLS regression with “Log hourly wage” as dependent variable and “Married”, “Gender”, “Work experience” and “Training time” as independent variables. Interpret the coefficients for “married”, “gender” and “length of training”. Why are these not causal effects? 1*** Exercise 3: Pooled OLS Regression 2*a) Pooled OLS 3reg lwage married sex pgexpft pgbilzeit 128 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 The variables married, sex and pgbilzeit most likely correlate with other disregarded/unobserved variables that have an effect on the wage. For example, women work more frequently in occupations with lower wages. b) Run the regression again with the option “vce(cluster persnr)” to get clustered standard errors. How do the standard errors of the coefficients change? 1*b) Pooled OLS with cluster standard errors 2reg lwage married sex pgexpft pgbilzeit, vce(cluster pid) 6.7. Fixed Effects Estimation 129 SOEPcompanion (v34), Release 2018, v.1 The standard errors are getting bigger. Exercise 4: Fixed Effects a) Subtract the person-specific mean value from each variable of the model. Use the “egen” function. Ideally you should also use a loop. 1*** Exercise 4: Fixed Effects 2*a) Subtract person-specific averages 3 4gen sample = 1 5foreach var in lwage married sex pgexpft pgbilzeit { 6 7bysort pid: egen `var'Mean = mean(`var') 8replace `var'Mean = . if `var'==. 9gen `var'Demeaned = `var' - `var'Mean 10 replace sample = 0 if `var'==. 11 } 12 bysort pid (sample): replace sample = sample[1] b) Estimate the Fixed Effects model with the previously generated variables. Why is no coefficient estimated for “gender”? How do the coefficients change compared to the pooled OLS estimate? Is the effect of “married” now causally interpretable? 1reg lwageDemeaned marriedDemeaned sexDemeaned pgexpftDemeaned pgbilzeitDemeaned, ˓→vce(cluster pid) nocons No coefficient was estimated for sex because sex was stable over time for all observations. The coefficient of married is now significant at the 5% level! c) Now estimate the Fixed Effects model using the command “xtreg lwage married sex pgexpft pgbilzeit, fe “. What do you notice about the coefficients compared to task 4 b)? And with the standard errors? 1*c) xtreg, fe 2xtreg lwage married pgexpft pgbilzeit, fe vce(cluster pid) 130 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 The coefficients are not identical with 4 b) and the standard errors become larger, because model b) does not take into account the estimation of mean values in the standard errors. d) Now add dummy variables for the years (i.syear). What happens with the effect of “labour market experience”? 1*d) xtreg with dummy 2xtreg lwage married pgexpft pgbilzeit i.syear, fe vce(cluster pid) 6.7. Fixed Effects Estimation 131 SOEPcompanion (v34), Release 2018, v.1 • Community Class Sizes “ggk” Use the various important variables of the ppfad.dta data set as your start file. 1use hhnr persnr bghhnr sex gebjahr bgnetto bgpop using ${MY_IN_PATH}\ppfad.dta, clear Keep people who completed a questionnaire in 2016 and live in a private household. 1*Keep people who completed a questionnaire in 2016 and live in a private household 2keep if bghhnr>0 & inrange(bgnetto, 10, 29) & inlist(bgpop, 1, 2) 3keep hhnr persnr bghhnr sex gebjahr bgnetto bgpop 4merge 1:1 persnr using ${MY_IN_PATH}\phrf.dta, keep(match master) keepusing (bgphrf) ˓→nogenerate 5tempfile ppfad 6save `ppfad' Prepare the different data sets bgp, bghbrutto, regionl 1*Prepare data set bgp 2use ${MY_IN_PATH}\bgp.dta, replace 3keep persnr hhnr bghhnr bgp01*bgp143 4tempfile bgp 5save `bgp' 6 7*Prepare data set bghbrutto 8use ${MY_IN_PATH}\bghbrutto.dta, replace 9keep hhnr bghhnr bgsampreg bgbula bgregtyp 10 tempfile bghbrutto 11 save `bghbrutto' 12 13 *Prepare data set regionl 14 use ${region}\regionl_v33.dta, replace 15 keep if syear==2016 16 keep syear hhnr hhnrakt ggk 17 rename hhnrakt bghhnr 18 tempfile regionl 19 save `regionl' Merge all data sets. 1*Merge all data sets 2use `ppfad' 3merge 1:1 persnr using `bgp', keep(match master) nogenerate 4merge m:1 bghhnr hhnr using `regionl', keep(match master) nogenerate 5merge m:1 bghhnr hhnr using `bghbrutto', keep(match master) nogenerate Recode negative values into missings. 1*Recode negative values into missings 2mvdecode sex gebjahr bgp01*bgp143,mv(-5/-1) Categorize the community class sizes of the SOEP regional data set. 1*Categorize community class size 2gen ggk_cat=. 3replace ggk_cat=-1if ggk==-1 4replace ggk_cat=1if ggk==1|ggk==2 5replace ggk_cat=2if ggk==3 (continues on next page) 138 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 6replace ggk_cat=3if ggk==4|ggk==5 7replace ggk_cat=4if ggk>5&ggk<=7 8 9lab var ggk_cat "Community Size categorised" 10 lab def ggk_cat -1"No information" 1"<=5000" 2"5001 - 20000" 3"20001 - 100000" /// 11 4">100000" 12 lab val ggk_cat ggk_cat Generate an age variable. 1*Generate age variable 2gen alter=2016-gebjahr if gebjahr >0 3gen alter_cat=1if alter<=20 4replace alter_cat=2if alter>20 &alter<=30 5replace alter_cat=3if alter>30 &alter<=65 6replace alter_cat=4if alter>65 &alter<=120 7 8lab var alter "age" 9lab var alter_cat "age categorized" 10 lab def alter_cat 1"<=20" 2"21-30" 3"31-65" 4">65" 11 lab val alter_cat alter_cat Categorize federal states variable. 1*Categorize federal states 2gen bgbula_cat=. 3*Schleswig-Holstein +Hamburg 4replace bgbula_cat=1if bgbula==1|bgbula==2 5*Lower Saxony +Bremen 6replace bgbula_cat=2if bgbula==3|bgbula==4 7*Mecklenburg Western Pomerania +Brandenburg 8replace bgbula_cat=3if bgbula==13 |bgbula==12 9*Saarland +Rhineland Palatinate 10 replace bgbula_cat=4if bgbula==7|bgbula==10 11 *Northrhine-Westphalia 12 replace bgbula_cat=5if bgbula==5 13 *Hesse 14 replace bgbula_cat=6if bgbula==6 15 *Baden-Württemberg 16 replace bgbula_cat=7if bgbula==8 17 *Bavaria 18 replace bgbula_cat=8if bgbula==9 19 *Berlin 20 replace bgbula_cat=9if bgbula==11 21 *Saxony 22 replace bgbula_cat=10 if bgbula==14 23 *Saxony-Anhalt 24 replace bgbula_cat=11 if bgbula==15 25 *Thuringia 26 replace bgbula_cat=12 if bgbula==16 27 28 lab var bgbula_cat "Federal states categorized" 29 lab def bgbula_cat 1"Schleswig-Holstein/Hamburg" 2"Lower Saxony/Bremen" 3 ˓→"Mecklenburg Western Pomerania/Brandenburg" /// 30 4"Saarland/Rhineland Palatinate" 5"Northrhine-Westphalia" 6"Hesse" /// 31 7"Baden-Wuerttenberg" 8"Bavaria" 9"Berlin" 10 "Saxony" 11 "Saxony-Anhalt" 12 ˓→"Thuringia" (continues on next page) 6.8. Working with SOEP Regional Data 139 SOEPcompanion (v34), Release 2018, v.1 (continued from previous page) 32 lab val bgbula_cat bgbula_cat 33 drop bgbula 34 rename bgbula_cat bgbula Put the variables in your preferred order and save your data set. 1*Order demography and identifiers first 2order persnr hhnr bghhnr syear sex gebjahr alter alter_cat bgsampreg bgbula ggk /// 3ggk_cat bgregtyp 4 5save ${MY_OUT_DATA}\zeit_online.dta, replace b) You want to get an initial overview of regional differences in satisfaction with various aspects in Germany. Use the variable bgsampreg and cross-stabilize the variable with all satisfaction variables to identify differences between East and West Germany, display the absolute and relative frequencies. To save the tables, save them in a log file. 1******************************************************************************** 2capture log close 3log using "${MY_LOG_OUT}\satisfaction.log", replace 4 5*Life satisfaction 6 7local varlist bgp0101 bgp0102 bgp0103 bgp0104 bgp0105 bgp0106 bgp0107 bgp0108 /// 8bgp0109 bgp0110 bgp0111 bgp0112 9foreach x of local varlist { 10 tab bgsampreg `x' [aw= bgphrf] , row 11 } 140 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 To view all tables, look at your generated log file. c) Now take a closer look at satisfaction with various aspects of life with the help of SOEP regional data. Use the community size classes. Create a table showing you satisfaction with different aspects of life and revealing differences by gender, age, community size class and federal state. 1foreach x of local varlist { 2*Tabulation of satisfaction by size of community and federal state 3table `x' sex alter_cat, by(bgbula ggk_cat) contents(freq) column row stubwidth(20) ˓→cellwidth(8) csepwidth(2) nomissing 4*Tabulation of satisfaction by size of community 5table `x' sex alter_cat, by(ggk_cat) contents(freq) column row stubwidth(20) ˓→cellwidth(8) csepwidth(2) nomissing 6*Tabulation of satisfaction by federal state 7table `x' sex alter_cat, by(bgbula) contents(freq) column row stubwidth(20) cellwidth ˓→(8) csepwidth(2) nomissing 8} 6.8. Working with SOEP Regional Data 141 SOEPcompanion (v34), Release 2018, v.1 To view all tables, look at your generated log file. As you can see, SOEP regional data can be used to analyze variables at the smallest regional levels. d) Create a table that shows you the political interest differentiated by age, gender and community size class for Bavaria 1******************************************************************************** 2capture log close 3log using "${MY_LOG_OUT}\political_interest.log", replace 4 5*Political interest 6*Tabulation of political interest by size of community for Bavaria 7table bgp143 sex alter_cat if bgbula==8, by(ggk_cat) contents(freq) column row ˓→stubwidth(20) cellwidth (8) csepwidth(2) nomissing 142 Chapter 6. Working with SOEP Data SOEPcompanion (v34), Release 2018, v.1 It becomes clear that the SOEP offers a wide range of possibilities for region-related analyses. It is possible to allocate a multitude of regional indicators at the level of the federal states, the regional planning regions, the districts and the postal codes. 6.8. Working with SOEP Regional Data 143 CHAPTER SEVEN WORKING WITH SOEP DOCUMENTATION 7.1 Variable Search with Questionnaires If you come across a variable in the data set whose variable content is unclear, you should always check whether there is a suitable questionnaire for the data set. Under Original Core Data you can see whether the data sets correspond to a survey instrument. The related questionnaires can be found here: Example: During your research project you come across the variable bbh5508 with the German label “Auto: Gründe” (Car: Reasons) and the Englisch label “Reason for No Car in Household Unfortunately, it is difficult to determine the variables content from the output and also from the label designations. To understand the complete question and also possible filter instructions, you should use the questionnaires. Example Variable: bbh5508: Wave „bb“ (Survey Year 2011); household questionnaire („h“), question number 55, item 8 Open The variable “bbh5508” can be found in the questionnaires for 2011. Select the survey year 2011 and download the household questionnaire. 144 SOEPcompanion (v34), Release 2018, v.1 Search the variable “bbh5508” in the Since you are already in the correct questionnaire, you must now search for question 55. 7.1. Variable Search with Questionnaires 145 SOEPcompanion (v34), Release 2018, v.1 To understand which information the variable “bbh5508” contains, you have to deal with the question. For each answer category, respondents should indicate whether or not the shown items apply to the household. If the item does not apply, respondents must answer an additional question about the reasons. Both questions should be understood as separate variables. The variable “bbh5501” indicates whether a TV is present in the household. The reasons why there is no TV in the house can be found in the variable “bbh5502”. The variable “bbh5507” shows whether a car is present in the household and the variable “bbh5508” shows reasons why no car is present in the household. By looking into the questionnaire, the variable is now easier to understand. The variable “bbh5508” only contains people who do not have a car in their household and shows the reasons given. 7.2 Variable Search with paneldata.org With paneldata.org it is also possible to search for variables. For example, if you want to find more information about generated variables, a search with paneldata.org is indispensable. For example, the platform offers comprehensive frequency counts, the chronology of the variables searched for, a cross-study variable linkage via concepts, a syntax generator and a topic list for content search in the SOEP. Example Variable: bbh5508: Wave “bb” (Survey Year 2011); household questionnaire (“h”), question number 55, item 8 Open 146 Chapter 7. Working with SOEP Documentation SOEPcompanion (v34), Release 2018, v.1 Please select the study SOEP-Core. The SOEP-Core overview contains important general information about the study, e.g. data access, survey method, questionnaires, thematic diversity, terms for missing codes, all available data sets of the study and metadata-based questionnaires. To search for a variable, a data set or a publication, simply enter the desired search term in the search field. 7.2. Variable Search with paneldata.org 147