A literature review of methods to detect fabricated survey data
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Bredl, Sebastian; Storfinger, Nina; Menold, Natalja Working Paper A literature review of methods to detect fabricated survey data Discussion Paper, No. 56 Provided in Cooperation with: Justus Liebig University Giessen, Center for international Development and Environmental Research (ZEU) Suggested Citation: Bredl, Sebastian; Storfinger, Nina; Menold, Natalja (2011) : A literature review of methods to detect fabricated survey data, Discussion Paper, No. 56, Justus-Liebig-Universität Gießen, Zentrum für Internationale Entwicklungs- und Umweltforschung (ZEU), Giessen This Version is available at: https://hdl.handle.net/10419/74449 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc-nd/3.0/de/
Zentrum für internationale Entwicklungs- und Umweltforschung der Justus-Liebig-Universität Gießen A Literature Review of Methods to Detect Fabricated Survey Data by SEBASTIAN BREDL*, NINA STORFINGER** and NATALJA MENOLD*** No. 56 Gießen, August 2011 *Lehrstuhl für Statistik und Ökonometrie, Fachbereich Wirtschaftswissenschaften Justus–Liebig–Universität Gießen Licher Str. 64 35394 Gießen E-Mail: [email protected]. **Center for International Development and Environmental Research (ZEU), Section 3 Justus-Liebig-Universität Gießen Senckenbergstr.3 35390 Gießen Email: [email protected]essen.de ***GESIS Leibniz-Institut für Sozialwissenschaften P.O. Box 122155 68159 Mannheim E-Mail: [email protected]
DiesesWerkistimInternetunterfolgenderCreativeCommonsLizenzpubliziert: http://creativecommons.org/licenses/by‐nc‐nd/3.0/de/ SiedürfendasWerkvervielfältigen,verbreitenundöffentlichzugänglichmachen,wenndas DokumentunverändertbleibtundSiedenNamendesAutorssowiedenTitelnennen.DasWerkdarf nichtfürkommerzielleZweckeverwendetwerden.
Abstract This paper reviews literature dealing with the issue of detecting interviewers who falsify survey data. The most reliable method of detecting falsifiers is through face-to-face reinterviewing of survey participants. However, only a limited number of participants can usually be reinterviewed. A review of the present literature clearly indicates that reinterviewing is more effective if the reinterview sample is put together according to some indicators that might be based on metadata, survey data or interviewer characteristics. We examine existing literature with regard to the suitability of different types of indicators that have been used in this context. JEL: C83 C93 Keywords: Interviewer falsification, quality control of survey data, reinterview Acknowledgements Financial support through the DFG in project WI 2024/2-1 within SPP 1292 is gratefully acknowledged. Furthermore we would like to thank Andreas Diekmann, Gesine G¨ullner and Peter Winker for their valuable comments on previous versions of the paper. 2
1 Introduction In economic and social research, survey data is often the cornerstone of empirical investigations. Several factors that may impair the quality of such data during the collection stage, such as systematic non-response or interviewer effects on response behaviour, have gained attention in literature. Another important factor that has not received much attention thus far is the conscious deviation from prescribed procedures by the interviewer, which is referred to as interviewer falsification (Schreiner et al. 1988) or cheating (Schr¨apler and Wagner 2003). In relation to this the American Association for Public Opinion Research (AAPOR) defines “intentional departure from the designed interviewer guidelines and instructions, unreported by the interviewer, which could result in the contamination of data” as “interviewers’ falsifications”(AAPOR 2003: 1). There is a wide range of potential forms of cheating (cf. also Schr¨apler 2010). The most blatant of these is undoubtedly the fabrication of entire interviews without ever having contacted the target person. Another possibility is partial falsification, for example making the contact but only asking a portion of the questions contained in the questionnaire and faking the remaining data (Harrisson 1947). More subtle forms are listed by Case (1971). They include interviewing someone other than the intended person (for example another family member or a neighbour), changing interview mode (for example conducting the interview by telephone when a face-to-face interview is required), or changing the location of the interview (for example conducting it on a street corner if at-home interviews are required). This article reviews literature dealing with detecting the most blatant form of cheating, namely the partial or complete fabrication of questionnaires by the interviewer. Seen from the interviewer’s perspective, there are several reasons why data fabrication might be an attractive option. Interviewers do not usually have a strong interest in delivering high-quality data, apart from the potentially satisfying feeling of having done a good job. Interviewers are not involved in planning surveys or developing questionnaires and it is unlikely that interviewers are trained in scientific research ethics (AAPOR 2003). Furthermore, interviewers are not involved in processing data subsequent to data collection during the field work period. As Durant (1946: p. 290) puts it, “[o]ne day’s interviewing, however well done, merely serves to lead on to the next day’s interviewing”. Thus, the reward from doing good work might be very small, whereas the task itself can sometimes be quite unpleasant. Interviewers are required to ask people who they do not know to reveal personal information, which may trigger dismissive reactions (cf. Crespi 1945; Stewart and Flowerman 1951; K¨ohne-Finster and G¨ullner 2009). Additionally, interviewers are often faced with payment schemes based largely on the number of completed interviews (Kennickell 2002), which create pressure to augment “quantity” and neglect the “quality” of interviews, and may finally promote conditions leading to data fabrication (cf. Bennett 1948; Sudman 1966). So far very little research has been done on the consequences of data fabrication for subsequent statistical analyses. This might be due in part to the fact 3
that the severity of these consequences is obviously related to the prevalence of data fabrication. This parameter can be estimated only roughly, as it is likely that not all relevant cases can be detected. Studies reporting some estimates (e.g Schreiner et al. 1988; Koch 1995; Krejsa et al. 1999; Schr¨apler and Wagner 2005; Li et al. 2009) suggest that the proportion of fabricated interviews rarely exceeds 5%. However, these studies refer only to large-scale surveys. In smaller surveys, with only a handful of interviewers, one may observe much larger proportions of fabricated interviews. Harrison and Krauss (2002) report on a survey in which only two of five interviewers delivered reliable data. Bredl et al. (2008) mention a case in which the first round of a survey conducted by four interviewers consisted entirely of faked interviews. Not only is the quantity of fabricated data an important determinant in this context, but so is quality. If cheaters were able to reproduce “realistic” data, there would hardly be a problem. According to several studies (Hippler 1979; Reuband 1990; Schnell 1991; Schr¨apler and Wagner 2005), cheaters generally do quite a good job of fitting their data to marginal distributions found in real data, but they struggle to reproduce more complex relationships like those revealed by factor analysis or multivariate regression analysis. Consequently, even a small proportion of fabricated interviews, say of around five percent, might have a severe impact on the results of multivariate statistical analysis as shown by Schr¨apler and Wagner (2005). But this is not necessarily the case as demonstrated by Schnell (1991). To the best of our knowledge, no study has yet been published that investigates the impact of extremely high proportions of faked interviews, comparable to those that have been observed in some smallscale surveys. Considering that small scale surveys play an important role in the social sciences this topic merits attention. As interviewer data fabrication seems to be a non-negligible problem, one must be concerned about how to detect fraudulent interviews. Although the overall volume of literature on this issue is still modest, the variety of proposed methods and indicators is quite considerable, which clearly calls for some comparison and evaluation of different approaches. This is the issue we would like to address in this literature review. Based on our analysis we also try to formulate some recommendations on how to proceed in order to detect fabricated data, and we identify fields of research that need more attention in the future. For our literature review we systematically searched different data bases for the social and economic sciences. Thus, we analysed literature, published in English and German. Of the literature found, the majority concerned methods of detecting falsifiers (most were journal articles, but conference proceedings and working papers were also available). In our review, we considered articles on methods of detection based on empirical data. Overall, our search results show that up to now no extended research exists on the topic of falsifications. Nevertheless, we were able to find interesting results with respect to detection methods and to discuss the advantages and disadvantages of the different methods. We distinguish between two types of studies. On the one hand there are articles dealing with detection methods applied during field control procedures 4
Schr¨apler and Wagner (2005) Schr¨apler and Wagner (2005) examined the data from the German long-term panel study SOEP. In such panel studies fabrications are relatively rare since respondents are interviewed every year and consistency checks across the different waves immediately reveal fraudulent data. By means of two different ex-post analyses Schr¨apler and Wagner (2005) examined only data from the first waves of different samples of the SOEP (Schupp and Wagner 2002). The first one was based on the so called Benford’s Law (Benford 1938), which we illustrate in more detail in Subsection 3.3. The idea behind this method is to compare the distribution of the first digit of all numbers in the (metric) answers from the survey with the Benford distribution. If the numbers follow a specific monotonic declining distribution, simply spoken that the proportion of 1’s is higher than the proportion of the 9’s, one can assume that the data is Benford distributed. Schr¨apler and Wagner (2005) calculated the deviation from the Benford distribution by means of a chi-square value for every interviewer cluster rather than for every single interview. As a result they showed that about half of the known fakers could indeed be marked as cheaters through application of this detection method. Apart from this analysis Schr¨apler and Wagner (2005) created an interesting approach by incorporating several variables in order to detect known fakers in survey data. The authors called this the “variability method” because the idea behind it is that cheaters show a lower variance of specific answers across all their conducted interviews than accurate interviewers do. Schr¨apler and Wagner (2005) attributed the reduction of variance to the proportions of missing answers, of extreme answers in scale questions and of conspicuously consistent answers across specific questions in the questionnaire. Based on the observed variance of interviews Schr¨apler and Wagner (2005) calculated a plausibility value for every interviewer in order to identify cheaters. If the plausibility was too low an interviewer was considered to be a faker. In this way the authors ranked interviewers with respect to their plausibility values and noticed that almost all of the known cheaters appeared at the top of the ranking. Additionally, they noted that their results were much better than those which are based on Benford’s Law. Thus, we can infer from the results that the “variability method” is a more promising way to reveal falsifiers than Benford’s Law. R´esum´e Now we would like to summarize the findings of studies presented in this section. For ex-ante studies Table 1 shows that recontact is the most important method for detecting false data. The studies show that random (or unfocused) recontacts enable detection of only an insignificant number of falsified interviews. In all of these studies unfocused recontact procedures were then refined using supplemental information. This information was helpful in identifying additional fabricated interviews. For focused recontacts or reinterviews information about contact outcomes or the number of interviews carried out by one interviewer, referred to as metadata, were used (Table 1). Other simple data analyses (such as those for age and gender of interviewed persons or the proportion of non-minorities interviewed by one interviewer) were helpful in conducting more 11
focused recontacts or reinterviews. However, choosing one of these methods for ex-ante studies (and their success in detecting falsifications) was highly dependent on the specific circumstances of the study. More concretely, sampling procedures, the survey topic, and the sensitivity of questions are all associated with the usability of a particular detection method. Ex-post data analyses seem to be an effective method of identifying indicators in order to separate false and real data. But more research should be done here to determine the success of such methods, and it should be reiterated that appropriate methods are often bounded to the specifics of a survey (e.g. the sensitivity of questions asked or the proportion of appropriate questions for the application of Benford’s Law or the “variability method”). With the help of selected studies presented in this section we aim to show how several methods and combinations of these methods have been used in ex-post and ex-ante studies. In the next section we carefully examine the different methods and discuss their usability. We also introduce interviewer characteristics as an additional method. Once dishonest interviewers had been detected, different autors (Koch 1995; Turner et al. 2002) then analysed the extent to which interviewer characteristics differ between honest and dishonest interviewers. However, interviewer characteristics were not used as an identification method in these studies. Thus, we disregarded interviewer characteristics in Section 2. 3 Overview of different approaches As outlined in Section 2 existing literature suggests that the effectiveness of recontact procedures can be increased if they are combined with other indicators. In this chapter we first examine literature related to recontact procedures and subsequently discuss the suitability of other methods that could be applied to create focused reinterview samples. 3.1 Recontact procedures The most common method of detecting faked survey data is the recontact method. Using this method respondents are recontacted in person, by mail or by telephone after the initial interview in order to verify whether the initial interview actually took place. Below we will focus on this recontact method and the possible problems associated with it. In spite of the fact that AAPOR (2003) suggests that face-to-face recontact is the most effective method of detecting fraudulent data, the most common recontact method used in surveys involves sending postcards to interviewed persons with an appeal to them to reply. These postcards mainly ask respondents about the time, date, and critical components or topics of the interview, as well as the interviewer’s behaviour. A statement about not being interviewed or implausible time and date information may then be considered as indicators of falsifications. In general this method has some questionable factors as shown by Koch (1995) and Hauck (1969). These factors include memory problems and the willingness 12
of contacted persons to reply, and these are associated with a selectivity bias. Hauck (1969) sent postcards to interviewed and non-interviewed respondents and noticed firstly that only 50% of interviewed respondents returned the postcards. Secondly, he showed that there were race, age and education differences between persons who sent postcards back and those who did not. This implies that cooperative persons who sent the postcards back did not constitute a sample which makes reliable statements about interviewed persons in general. Thirdly, Hauck (1969) found that fourteen out of 100 non-interviewed persons actually stated that they had been interviewed. Thus, it is clear that memory problems or interviewed target persons not sending back postcards decrease the validity of control results obtained by postcards. In particular, with respect to response rates, telephone and/or personal contacts, also referred to as reinterviews, are more effective than contacts using postcards. Telephone or personal reinterviews were already mentioned in an early work by Case (1971). In market surveys he conducted controls with telephone reinterviews and revealed that about 27% of the interviews in all the studies examined were not conducted properly. Also, the U.S. Bureau of Census regularly checks a randomly selected portion of interviewers (between 2% and 10%) by reinterviewing particular target persons (cf. Bushery et al. 1999). But regarding response and memory problems the reinterview process is limited by sample size and duration (Cantwell et al. 1992). A large number of reinterviews increases costs and a high number of questions posed within a reinterview is a strain for respondents, and consequently biases results. Reinterviews are also limited by the elapsed time following the interview (a delayed survey for further control purposes (reinterview) bears the risk of memory effects). Thus, what is required is a reinterview sample which is large enough to generate significant results but is small enough to keep costs down. As we have illustrated above (cf. Section 2), ex-ante studies by Hood and Bushery (1997) and Krejsa et al. (1999) (see also Li et al. (2009)) show that a small, and most of all a non-randomly selected, reinterview sample is more effective in detecting cheating interviewers. Thus, we must infer from these empirical findings that “content based reinterviews” perform much better than random reinterviews, and most notably better than sending out postcards. In the following sections we would like to present the prevalent indicators used for the optimal creation of a focused reinterview sample: usage of metadata, Benford’s Law and other statistical analyses of survey data, as well as interviewer characteristics. 3.2 Metadata As already outlined in Section 2, the notion “metadata” comprises different types of information related to the process of data collection, rather than to the collected data itself. Metadata-based indicators used to detect falsifiers can be divided into two groups: indicators based on interviewer’s contact outcomes and indicators based on interview processing, such as date and time stamps. Contact outcomes refer to information related to how many participants re- 13
fused the interview or how many participants were ineligible for some reason. As outlined above, Turner et al. (2002) were able to detect a large number of fabricated interviews by focusing recontact efforts on interviewers who had shown a suspiciously high success rate, whereas Hood and Bushery (1997) employed the ineligible unit rate as an indicator to put together their focused reinterview sample. In this context it is recommended to control for the characteristics of the area where an interviewer conducts his/her work, as demonstrated by Hood and Bushery (1997). However, due to a lack of reference data this is often not possible. Date and time stamps can only be recorded if the interview is conducted using a mode which relies on computer assistance. If these are available, they can be used to examine interview length, or the number of interviews completed within one day or within periods in which interviews were conducted. These types of indicators are employed by Bushery et al. (1999), Krejsa et al. (1999) and Murphy et al. (2004). Given a very limited number of studies it has so far been quite difficult to evaluate how well metadata can be used to detect cheating interviewers. Krejsa et al. (1999) combined both types of metadata-based indicators, which delivered quite promising results. The results of Turner et al. (2002) are quite promising as well, although the results of Murphy et al. (2004) show less promise. In the case of Turner et al. it should also be kept in mind that general success rates were quite low, probably as a result of the high sensitivity of the questions asked in the course of interviews. Consequently, high rates for falsifiers were extremely noticeable. A large advantage of approaches relying on metadata analysis is that they can be applied to a vast range of surveys. Whenever interviewers are prescribed which persons or households they are to contact, then one can calculate indicators related to contact outcomes. Whenever interviews are conducted with computer assistance, there is the opportunity to record date and time stamps. 3.3 Benford’s Law The largest part of the scarce research regarding the usage of answer patterns to detect fraudulent interviews is related to the usage of Benford’s Law for the analysis of metric survey data. With accurate survey data the distribution of the first digit of these metric answers usually follows the so called Benford’s Law (Benford 1938), a logarithmic (Newcomb 1881) and scale invariant distribution (Hill 1995). Thus, the probability that the first digit of the numbers is 1 is higher than the probability that it is 9. In general, Benford’s Law could be adapted to data without a built-in maximum (Nigrini 1999) and to data which is not composed of assigned numbers like zip codes or bank accounts. For example, Nigrini (1996) and T¨odter (2007) have shown that business and financial data in particular follow this monotonic decreasing distribution. Below we present some comprehensive studies which analysed survey data by means of Benford’s Law, and we attempt to illustrate whether the authors were successful in identifying fraudulent data. 14
In order to reveal faked survey data by means of Benford’s Law one must ensure that accurate survey data is actually Benford distributed and that faked data is not. If we look at the literature concerning this topic we cannot assume that this is always true. This was shown by Schr¨apler and Wagner (2003), who made a more in-depth ex-post analysis concerning the raw survey data from the German SOEP. The authors showed that the proportion of the first and second digits of real metric survey data is close to the Benford distribution and that faked data is not. But these findings should be interpreted carefully since the authors also observed a high proportion of 0’s and 5’s in the accurate survey data, perhaps due to a rounding effect. Therefore, one cannot infer that real survey data is surely Benford distributed. Results from the faked SOEP survey data are in turn not generalisable because of a very low number of cases of fabricated interviews (27 cases for faked data vs. 894 cases for real data). These results are similar to the findings of Wang and Pedlow (2005), who also observed a rounding effect in accurate survey data. Hence, Schr¨apler and Wagner as well as Wang and Pedlow, tested some modifications to the analysis in order to improve the detection of cheaters. Schr¨apler and Wagner calculated a chi-square value for “interviewer clusters”, meaning that they pooled together all interviews by one interviewer into one cluster. But the clusters with faked data were not precisely revealed. Again the reasons are a very small number of cases and, most of all, the spatial homogeneity of some clusters. The latter is due to the dependent distribution of interviewers and areas in the SOEP. But even if the authors take this spatial homogeneity into account by means of a linear regression on the chi-square value, faked and real clusters cannot be clearly separated. By contrast – concerning the first digit – Wang and Pedlow calculated the distribution the data set actually follows, the “all cases distribution” (cf. Swanson et al. (2003)). By using this distribution instead of that of Benford’s Law, it now became possible for the authors to identify cheating interviewers. Further modifications to the usage of the digits, for example regarding the digit 5 as 0, only led to satisfying results if the authors used the “all cases distribution” as well. Thus, we infer from the results of Wang and Pedlow and Schr¨apler and Wagner that real survey data does not clearly follow Benford’s Law and that one must be aware of the occurrence of a rounding effect. And there is still a lack of evidence that faked survey data is not Benford distributed, although Diekmann (2007) made an interesting contribution to this topic. Diekmann inspected the first and second digits of unstandardised faked regression coefficients and noticed that they deviate from the Benford distribution concerning only the second digit. But one should notice that this has only been documented for regression coefficients and not for raw survey data. Additionally Diekmann (2010) pointed out that one must avoid stating that survey data which deviate from Benford’s Law are automatically falsified. It often occurs that the whole data set follows Benford’s Law while some subsamples do not, even if they have not been falsified. The authors supposed that the reason behind this is that the numbers available for the analysis depend on the types of questions and topics in questionnaires. If one disregards this dependent structure, the rate of 15
“false positives” increases significantly, and therefore the discriminatory power of Benford’s Law becomes too low. To summarize, the current work concerning Benford’s Law questions the validity of this method. There is little empirical evidence that real survey data are close to the Benford distribution and problems like rounding effects may lead to a deviation from the Benford distribution. Also, it has not been clearly shown that faked raw survey data are not Benford distributed, since the number of cases is too low in most analyses. Thus, further modifications to Benford’s Law, such as several combinations of digits, should be developed and tested in datasets with a higher proportion of falsifications. Additionally, there is a lack of research regarding the number of interviewers and the number of interviews per interviewer which are required for the identification of falsifications to produce precise results. Recent work by Storfinger and Winker (2011), in which Benford’s Law is used as one of four indicators, suggests that the performance somewhat worsens as the number of interviews per interviewer decreases and improves as the overall number of interviewers decreases. However more research is needed to asses how many questionnaires per interviewer are needed in order to successfully adopt Benford’s Law. Also, the number and types of variables in the questionnaire which are suitable for application to Benford’s Law have not been precisely identified (cf. Porras and English 2004). Is Benford’s Law only appropriate for metric variables in raw survey data, or only for statistical estimates like regression coefficients? As Scott and Fasli (2001) demonstrated in a synthetic way, data are more likely to conform to Benford’s Law if the dataset contains only positive numbers and is positively skewed with a modal value not equal to zero. But there is still a lack of evidence for real survey data which fit to the Benford distribution. So far Benford’s Law can be used to evaluate survey data quality (cf. Judge and Schechter 2007) but it is not efficient enough to precisely identify cheaters. 3.4 Other Statistical Analyses of Survey Data Inspection of concrete survey data also delivers encouraging results by revealing “at risk” interviewers and falsifications. Several forms of questions and questionnaires could be included in such an analysis. As a result of our literature review we differentiate between two kinds of indicators which can be applied to an analysis to compare false and real data. The first kind applies to more or less plausible answers to survey questions by falsifiers, while the second kind is associated with the answering behaviour of falsifiers, which may differ from that of real respondents. We will call the first kind of indicators “content related” and the second ”formal”. Answers to open numerical questions, which were analysed using Benfords Law (cf. Subsection 3.3), also belong to the formal criteria. Content related indicators are special substantial answer patterns which systematically differ between fraudulent and accurate collected data. In the study by Turner et al. (2002) falsifiers produced implausible data regarding the composition of households (most households in falsified interviews included only 16
one 18-45 year old adult). Furthermore, in this study target persons “interviewed” by falsifiers “reported” having remarkably more lifetime partners and “showed” substantially more sexual activity, having sex much more often in the past seven days than those who were really interviewed. Finally, falsifiers provided phone numbers of “interviewed” persons less often. Further examples – as found by Reuband (1990) – are differences in evaluating one’s own personal financial situation, which was estimated more optimistically by falsifiers than by real sampled persons. However, content related indexes have seldom been used in studies to detect falsified data. Koch (1995) and Hood and Bushery (1997) describe accordant analyses. Hood and Bushery (1997) used the provision of telephone numbers in combination with other indexes in their multiple approach in order to improve reinterview samples. Schr¨apler and Wagner (2005) compared data from different rounds of a panel study (SOEP) and identified false data in this way. In other studies identified differences between false and real data are only discussed as possible cues to identify falsifiers (Turner et al. 2002). Schnell (1991), who systematically analysed differences between real and falsified data and in particular found differences in correlations and multiple regressions, remains sceptical about the usage of obtained differences for the purpose of detection. Of course, applying content criteria, apart from questions about household composition, age and gender, is problematic since it requires asking sensitive questions (e.g. Turner et al. 2002) or questions about content not related to the survey topic. Application of content related indicators needs additional strong hypotheses and knowledge about the differences between falsified and real data, which are, as a rule, not available. Formal criteria are produced by analysing the answering behaviour of interviewed persons. An example is item non-response. As shown by Schr¨apler and Wagner (2005) and Bredl et al. (2008) this occurs less often in falsified than in real interviews, since falsifiers incorrectly assume that real respondents would answer all questions in an interview – and therefore they avoid item nonresponse. In addition, Bredl et al. (2008) showed that falsifiers differ from real survey respondents in the way they answer semi-open-ended questions when the category “others” was included. Falsifiers tend to avoid the “others” category in order to reduce the effort needed to formulate an open answer. In contrast, when using open-ended questions one has to consider that falsifiers should answer such questions more frequently than real interviewees, but they should tend to shorter answers. In addition, falsifiers are less extreme if they use rating scales for their answers. With the help of multivariate cluster analyses Bredl et al. (2008) were able to separate falsifiers from honest interviewers using (in addition to Benfords Law, see section 3.3) information on item non-response, extreme answers and answering of open and semi-open questions. Furthermore, cheating interviewers can also be identified by looking at the proportion of answers to questions that lead to a faster answering, and therefore a quicker interview (cf. Matschinger et al. 2005; Hood and Bushery 1997). One would suppose that falsifiers tend to answer these so-called “filter questions” in such a way that allows them to skip a part of the questionnaire “legally” and therefore save time. For a health survey in the US Hood and Bushery 17
(1997) reported that falsifiers selected “the shortest path through the interview” producing survey participants who live in one person families and are “white non-smoker, no health problems and no health insurance” (p. 821). Using filter questions in this way was one of the multiple indicators applied by Hood and Bushery (1997) to detect falsifiers. In summary, formal criteria have advantages over content criteria in that they can be used for different types of content in different surveys. Work by Schr¨apler and Wagner (2005) and Bredl et al. (2008) shows that using formal indicators is an encouraging approach. However, application of formal criteria may increase the complexity of the answering process (e.g. by providing numerous open and filter questions) and reduce the accuracy of respondent’s answers by promoting low motivation and superficial information processing (satisficing behaviour, cp. Krosnick and Alwin (1987)). Like content related criteria, falsifiers’ answering behaviour should be known to the detectors. Furthermore it is important to ensure that falsifiers cannot adapt their way of cheating easily if the criteria are known to them (e.g. if a falsifier is aware that interviewers are considered to be suspicious if they deliver questionnaires with a low prevalence of item non-responses he/she could simply increase the prevalence of item non-responses when fabricating the data). This problem might be tackled by combining several criteria which makes adaptation from the falsifier’s side more difficult. Altogether, more research is needed in order to find an optimal questionnaire form, which can help to detect falsifiers through statistical data analysis using formal criteria. 3.5 Interviewer characteristics In most of the studies presented above interviewer characteristics were discussed as one issue which can be used for more focused reinterviews. The authors agree that inexperienced interviewers are likely to show cheating behaviour and should consequently be controlled by extended procedures (Biemer and Stokes 1989; Wetzel 2003; Turner et al. 2002; Schreiner et al. 1988). Schreiner et al. (1988), who report the results of the Interviewer Falsification Study of the U.S. Bureau of the Census, discovered by data analysis of two national surveys that the mean duration of employment for falsifiers is significantly lower (1.72 years) than that of all interviewers (6.22 years). The authors recommended that “(...) for the newer interviewers it may be useful to reinterview some of their work more frequently” (Schreiner et al. 1988: p.496). Concerning other interviewer characteristics mixed results can be found in the literature. With respect to gender and age of interviewers Koch (1995) has shown that young interviewers with a higher level of education produce a higher rate of falsifications. In West-Germany Koch found no gender differences, but in East-Germany male interviewers fabricated interviews more often. However Schr¨apler and Wagner (2003) did not discover any age, gender or education effects in SOEP. As a result, interviewers’ length of service seems to be a sufficient cue for more focused reinterviews. However, experienced interviewers are not less likely to falsify, but are less likely to be detected by controls. Hood and Bushery 18
(1997) reported the results of a study by the US Census Bureau 1982 which analysed interviewers’ characteristics and the likelihood of falsification. The results showed (similar to other studies) that interviewers’ tenure in particular is associated with differences in falsification behaviour. But additionally the authors reported that new interviewers with less than five years of experience had a higher probability of being detected, since they falsify more of their assignments and they tend to falsify entire interviews. Experienced interviewers (with five or more years of experience) falsify a smaller proportion of their assignments and prefer to falsify only a part of the interview. As a consequence, falsifications by experienced interviewers are more difficult to detect. Additionally, as shown by Schreiner et al. (1988) experienced interviewers use more selective approaches for falsification than less experienced interviewers. In panel surveys, for example, experienced interviewers falsify more often in continuing households (and not in newly selected households), since data can be appropriately estimated from the previous round. But in particular the comparison between experienced and inexperienced interviewers is based on very small sample sizes of interviewers who were found to falsify, and thus analyses were often done without statistical tests or results were afflicted by high statisitical insurance. In summary, newly hired and relatively inexperienced interviewers are seen as requiring extensive controls, but at the same time they are more amateur falsifiers who are more likely to be detected. This topic needs more research, especially regarding methods which can help to identify partially fabricated interviews conducted by experienced interviewers. Using length of service or other demographic characteristics for generating focused interviews appears to be less efficient in light of the results discussed in this section. This is particularly due to the fact that extended controls of less experienced interviewers lead to lower rates of falsification detection for experienced interviewers, who are much more sophisticated falsifiers. 4 Discussion and outlook In spite of the scarcity of scientific publications related to detection of data fabrication by interviewers, an examination of existing literature has delivered useful insights. Below we attempt to deduce some recommendations for practitioners. We have distinguished between two types of studies: ex-ante studies describe approaches implemented to detect cheaters, whereas ex-post studies apply indicators to datasets with known cases of falsification. Ex-ante studies analysed in Section 2 clearly suggest that focused recontacts are more effective than recontacts based on random samples. Furthermore, examination of studies using different recontact procedures reveals that the reinterview is the most reliable of these procedures. Forwarding postcards is a questionable alternative as several studies suggest that using postcards leaves many falsifiers undetected. We thus consider the focused reinterview as a good strategy to detect falsifiers. Ex-post studies complement ex-ante studies in that they provide deeper in- 19
sights into the suitability of different methods when creating focused reinterview samples. We have defined four methods in this context: metadata, Benford’s Law, other analyses of data contained in questionnaires referring to formal or content-related information, and interviewer characteristics. These can be divided into methods available for most surveys and methods whose applicability is only possible for specific types of surveys. Metadata, most types of formal survey data and data on interviewer characteristics are available for a wide range of surveys. Approaches relying on metadata and formal survey data have delivered promising results when used to detect falsifiers. This is especially true in the case of contact outcome data. Analysis of answer patterns for filter questions seems to be another promising approach and is directly linked to hypotheses of falsifiers’ behaviour. Interviewer characteristics did not turn out to be a useful indicator for creating focused reinterview samples. However, literature reveals one important point: experienced interviewers falsify in ways that make their detection more difficult. Thus, indicators relying on metadata or survey data can be expected to deliver better results for inexperienced interviewers who, if they falsify data, do it in a more amateur way. Consequently, it is plausible to assume that focused reinterviews are less efficient when applied to experienced interviewers. The applicability of Benford’s Law and of content related indicators based on substantial answers by survey participants depends on the nature of the data collected. Using Benford’s Law requires a wide range of metric variables. Even if a multitude of these variables is available it still remains open whether data which is honestly collected can be supposed to conform to Benford’s Law, and whether fabricated data cannot. There have been some promising results, but Benford’s Law should be applied with caution and, if possible, in combination with other indicators. General statements about the usefulness of content-related survey data are difficult to make, as the type of data available depends on the type of survey. It can be stated that answer patterns to questions on very sensitive issues have turned out to be good predictors of interviewer cheating and should thus be used to detect cheaters when available. Whenever one employs indicators based on metadata or content-related survey data one has to keep in mind that striking indicator values are not necessarily caused by data fabrication but may also be the result of “conventional” interviewer effects or cluster related design effects (spatial homogeneity). If indicators for the creation of focused reinterviews are calculated on the interviewer level it is not possible to distinguish between data fabrication and interviewer effects. Several studies which attempt to control for spatial homogeneity have been discussed, however these approaches are not always replicable and they bear the risk of omitting decisive factors. It is possible that indicators based on formal survey data – for example the proportion of extreme answers or item non-response – are less impaired by conventional interviewer effects and spatial homogeneity. However, to the best of our knowledge no research has yet been done on this issue. In addition to the recommendations we have provided above, we would like to point out some ideas for further research. First of all, more research should be 20