Can Weight of Evidence Method Improve Poverty Targeting?
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Duong, Be Thanh; Sariyev, Orkhan; Zeller, Manfred Article — Published Version Can Weight of Evidence Method Improve Poverty Targeting? Social Indicators Research Provided in Cooperation with: Springer Nature Suggested Citation: Duong, Be Thanh; Sariyev, Orkhan; Zeller, Manfred (2025) : Can Weight of Evidence Method Improve Poverty Targeting?, Social Indicators Research, ISSN 1573-0921, Springer Netherlands, Dordrecht, Vol. 178, Iss. 1, pp. 305-333, https://doi.org/10.1007/s11205-025-03576-z This Version is available at: https://hdl.handle.net/10419/323668 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) Social Indicators Research (2025) 178:305–333 https://doi.org/10.1007/s11205-025-03576-z ORIGINAL RESEARCH Can Weight ofEvidence Method Improve Poverty Targeting? BeThanhDuong1,2 · OrkhanSariyev1 · ManfredZeller1 Accepted: 7 March 2025 / Published online: 6 April 2025 © The Author(s) 2025 Abstract This study explores the potential of adapting a WOE (Logit) method, a combination of techniques such as weight of evidence, information value, logit regression, and a score-scaling approach based on doubling the odds, for poverty targeting. To verify the effectiveness and accuracy of this method, the study applies it to develop targeting tools based on international and Vietnamese national poverty lines, and compares their accuracy to commonly used poverty-targeting tools such as the Simple Poverty Scorecard, Proxy Means Test, and Poverty Assessment Tool. The study uses coverage rate as a criterion to compare the accuracy of targeting tools in identifying the poor at household and individual levels. The results indicate that targeting tools constructed using the WOE (Logit) method outperform those developed using other methods. Depending on the poverty line, this method improves poverty identification accuracy by 1.9–5.9 percentage points for households and 1.5–3.3 percentage points for individuals compared to other methods. Therefore, this study contributes an additional method for constructing targeting tools with high accuracy alongside existing approaches. Keywords Poverty targeting tool· Proxy means test· Simple poverty scorecard· Poverty assessment tool· Credit risk scorecard· Poverty targeting 1 Introduction To identify impoverished people eligible for certain poverty reduction programs to receive cash or in-kind transfers aimed at improving their well-being, many countries prefer lowcost but effective poverty assessment and targeting tools such as Proxy Means Test (PMT), Poverty Assessment Tool (PAT), or Simple Poverty Scorecard (Scorecard) (Etang etal., 2022; Schreiner, 2023; Zeller etal., 2006). These are questionnaires with specific answers * Be Thanh Duong [email protected]; [email protected] Orkhan Sariyev o.sariye[email protected] Manfred Zeller [email protected] 1 Hans-Ruthenberg-Institute, University ofHohenheim, Stuttgart, Germany 2 Faculty ofAgriculture andRural Development, Kien Giang University, KienGiang, Vietnam
306 B.T.Duong et al. and related scores for each question, enabling authorities to quickly predict whether households are above or below a given poverty line through straightforward and verifiable questions. For brevity, in this study, we refer to these tools collectively as targeting tools. When using targeting tools for scoring, households with total scores below a certain benchmark are classified as poor, and vice versa. This transparency makes users and the public more likely to accept and trust the accuracy of targeting tools in identifying the poor (Attanasio etal., 2023; Leite & George, 2014; Schreiner etal., 2014). To create transparent targeting tools, variables must be transformed into categorical variables, as each question in a targeting tool requires specific answers with assigned scores. Therefore, transforming original variables into categorical ones is a fundamental step in the construction process of targeting tools. Values within a variable can be grouped in various ways, resulting in numerous new categorical variables from a single original variable. Categorization for variables with wide-ranging values is even more challenging, as without a specific criterion, it is difficult to determine the best category groupings. Each new categorical variable can affect the predictability of targeting tools differently. Although the impact of a single variable may be minor, ineffective groupings across many variables can prevent targeting tools from achieving optimal predictability. Additionally, the way categories are grouped for a variable can influence statistical values of prediction models, such as c-statistic values, adjusted R-squared, or Akaike Information Criterion. Typically, the best indicators for targeting tools are selected through a stepwise process, in which the fractional change in these statistical values serves as one of the key selection criteria (Ohlenburg etal., 2022; Schreiner, 2023; Zeller etal., 2005). Consequently, how categories are grouped within a variable can also affect the indicator selection for targeting tools, ultimately affecting their predictability. However, these aspects remain underexplored, as most existing studies primarily focus on finding optimal prediction algorithms to enhance the accuracy of targeting tools. These studies often compare various predictive modeling techniques, ranging from traditional statistical methods such as linear probability models, logistic regression, ordinary least squares (OLS), and quantile regression (Linh & Baulch, 2011; Schreiner, 2015; Skoufias etal., 2020; Zeller etal., 2006) to more advanced machine learning algorithms (Alsharkawi etal., 2021; Kambuya, 2017; McBride & Nichols, 2015). Nevertheless, no traditional regression model consistently outperforms others across different contexts (Etang etal., 2022; Schreiner, 2015). Furthermore, the accuracy gains from using sophisticated machine learning algorithms over simpler regression techniques are often marginal (Hand, 2006). Weight of Evidence (WOE) and Information Value (IV) techniques are often used to address these binning issues in the development of credit risk scorecards (Anderson, 2007; Siddiqi, 2016). The WOE technique helps assess the strength of the relationship between groups within a variable, known as bins, and a given dependent variable (Lund, 2016). TheIV, on the other hand, is a useful measure for evaluating and comparing the predictive power of variables, and it can serve as a benchmark to select optimal category groupings for indicators (Siddiqi, 2006). Incorporating these techniques facilitates the variable transformation process, helping create monotonic variables with high predictive power. Moreover, applying the WOE technique in variable transformations has been shown to improve the performance of predictive modeling in credit scoring (Chen etal., 2020; Pomazanov, 2023; Sharma, 2011). Despite the proven effectiveness of WOE and IV techniques in credit risk assessment, their application in poverty identification remains underexplored. Therefore, this study aims to examine the feasibility and efficacy of applying these techniques to poverty identification. In credit scoring, logistic regression and the score-scaling approach
307 Can Weight ofEvidence Method Improve Poverty Targeting? of doubling the odds are often used along with WOE and IV techniques to construct scorecards that predict the default probabilities of customers (Anderson, 2007; Siddiqi, 2016). Hence, we apply a combination of these approaches for identifying the poor. For brevity, we refer to this combination of approaches as the WOE (Logit) method in this study, which includes WOE, IV, Logit, and the score-scaling approach of doubling the odds. Targeting tools developed using this method are called WOE tools. To verify the accuracy of WOE tools, we compare them to other commonly used targeting tools such as PMT, PAT, and Scorecard. For clarity and distinction, we refer to the methods used to construct these targeting tools as the PMT (OLS) method, PMT (Quantile) method, and Scorecard (Logit) method, respectively. The coverage rate, or poverty inclusion rate, is a commonly used criterion to define the accuracy and benchmark score of targeting tools (Etang etal., 2022; IRIS Center, 2005; Schreiner, 2023). Therefore, we apply this criterion to measure and compare the accuracy of targeting tools while holding the predicted poverty rate. Based on an intensive literature review, to the best of our knowledge, this study is the first to apply the WOE (Logit) method to construct targeting tools. By doing so, it introduces another approach to poverty identification alongside existing methods. This method could help optimize input variables to develop high-accuracy targeting tools. In this study, we apply the four aforementioned methods to develop targeting tools based on two poverty lines: Vietnam’s national poverty line and the international poverty line. Our focus is exclusively on building tools to identify the poor in the ethnic minorities, which have the highest poverty rate in Vietnam. Although ethnic minorities1 comprise only 14% of Vietnam’s population, they constitute half of the country’s poor. Thus, this group is of special concern to the Vietnamese government (CEMA, 2020; CEMA etal., 2015). The accuracy of targeting tools will be compared not only at the household level but also at the individual level, as effective targeting tools should target households with more poor members rather than those with fewer members. The remaining sections of the paper are organized as follows: Sect.2 presents the data source and methodology; Sect.3 presents the study’s results and discussion; and Sect.4 summarizes the major findings. 2 Data andMethodology 2.1 Data This study uses data from Vietnam Household Living Standard Survey in 2018 (VHLSS 2018), conducted by the General Statistics Office of Vietnam. This official dataset is useful for developing targeting tools because it includes necessary information such as household demographics, homestead characteristics (e.g., type of house or toilet), durable goods, education, employment, and income (General Statistics Office, 2019). The primary focus of this study is the ethnic minority community, which has the highest poverty rate in Vietnam. Therefore, we analyze data from 7,863 ethnic minority households. The VHLSS 2018 dataset contains over 600 indicators covering various aspects of households, including demographics, housing, education, employment, and assets. 1 Vietnam has 53 ethnic minorities and one ethnic majority (Kinh).
308 B.T.Duong et al. First, we remove indicators with more than 10% missing values.2 We then select 129 straightforward and verifiableindicators that exhibit a strong correlation with household income, expenditure, or poverty status. After identifying candidate variables, we split the dataset using stratified random sampling into a calibration sample (70%) and a validation sample (30%). We use the calibration sample to construct targeting tools, and the validation sample to evaluate their accuracy. 2.2 Poverty line The study evaluates the accuracy of four methods for constructing targeting tools based on national and international poverty lines. However, the Vietnamese government has different national poverty lines for rural and urban areas. Since 90% of the ethnic minority population and 95% of poor ethnic minorities reside in rural areas (CEMA, 2020; General Statistics Office, 2019), this study uses only the rural poverty line for all ethnic minority communities (Table1). 2.3 Methodology 2.3.1 General Steps inDeveloping Poverty Targeting Tools To verify the feasibility and effectiveness of the WOE (Logit) method, we compare it with three widely used methods: Scorecard (Logit), PMT (OLS), and PMT (Quantile). Each method is used to develop two distinct targeting tools based on national and international poverty lines. To ensure a fair comparison across all methods, we limit the number of indicators in each targeting tool to ten. Table2 presents a general step-by-step framework for constructing targeting tools. Step 1: Grouping responses to indicators Each question in targeting tools must have a specific answer with a designated score. This requirement ensures transparency and helps enumerators accurately identify the poor in the field. To meet this requirement, input variables for constructing targeting tools should be transformed to categorical variables. Numeric variables can be transformed into many new categorical variables, as there are various ways to group categories. Moreover, creating categories for variables with a wide range of values is challenging, as there are countless ways to form category groupings. However, existing studies have not provided detailed guidance on how to effectively create these categorical variables. Therefore, we generate new categories for variables based on those in the existing tool of the Vietnamese government. To maintain practicality, we limit the number of categories for each variable to seven, aligning with the Vietnamese government’s tool. Next, we conduct a crosstabulation analysis between these new categorical variables and a given poverty status to assess the relationship between the categories and poverty. If the categories do not align with the given poverty status, we regroup them accordingly to ensure they make sense and represent at least 5% of the sample. For example, in Table3, we can categorize household size into six groups: “1 member”, “2 members”, “3 members”, “4 members”, “5 members”, “6 members”, and “7 members or more”. By creating a cross-tabulation between this variable and the national poverty 2 Regarding variables include missing values with less than 10%, we replace their missing values with their mean or mode values.
309 Can Weight ofEvidence Method Improve Poverty Targeting? Table 1 Poverty rates of the Vietnamese ethnic minority community in terms of income in 2018. Source: Authors’ computations based on the VHLSS 2018 data a The poverty status is defined based on income or expenditure, so the poverty rate in Table1 is different from the poverty rate published by the Vietnamese government, which is based on a multi-dimensional poverty status b This is based on the PPP conversion factor of 7,891.2 from the World Bank in 2018 to convert Vietnam Dong (VND) to USD Household poverty statusa based on income and expenditure Poor Non-poor All National poverty line (Income/person/day ≤ 2.92 USD in purchasing power parity)bN 1,293 (16.4%) 6,570 (83.6%) 7,863 (100%) International poverty line (Expenditure/person/day ≤ 1.9 USD in purchasing power parity) N 1,054 (13.4%) 6,809 (86.6%) 7,863 (100%)
310 B.T.Duong et al. Table 2 Steps for constructing poverty targeting tools by the method Step Activity Scorecard (Logit) method PMT (OLS) method PMT (Quantile) method WOE (Logit) method 1 Group responses to indicators Manual categorization based on cross-tabulation with poverty status variable WOE & IV technique 2 Indicator selection “MaxC” stepwise “MaxR” stepwise PMT’s indicators “MaxC” stepwise 3 Regression approach Logit OLS Quantile Logit 4 Generate scorecards Scorecard (Logit) method PMT (OLS) method PMT (OLS) method WOE (Logit) method 5 Estimate accuracy Coverage rate in identifying poor households and people based on national and international poverty lines Source for more details of each method Schreiner (2023) Nguyen & Tran (2018) Zeller etal (2005) Siddiqi (2016)
311 Can Weight ofEvidence Method Improve Poverty Targeting? status, we find the following poverty rates for each category: 16.2, 9.6, 8, 14.2, 19.9, 21.8, and 32.1 percentage points, respectively. Since the first category accounts for only 3.1% of the population, we merge it with the second category to form a new category: “ < 3 members”. This new category has a poverty rate of 11 percentage points. However, the poverty rates of these categories are inconsistent and contradict common sense, as the category “ < 3 members” has a higher poverty rate than the category “3 members”. To address this, we merge the new category “ < 3 members” with the third category “3 members” resulting in the category “ < 4 members” with a poverty rate of 9.4 percentage points. Ultimately, the household size variable is categorized into five groups: “ < 4 members”, “4 members”, “5 members”, “6 members”, and “ ≥ 7 members”. The poverty rates for these categories are now consistent and make sense: 9.4, 14.2, 19.9, 21.8, and 32.1 percentage points. Step 2: Selecting indicators for the prediction model After obtaining candidate variables in Step 1, this step selects the ten best indicators for prediction models. Depending on the methodology, these indicators can be selected using the “MaxC” or “MaxR” stepwise approaches. The “MaxC” approach refers to a forward stepwise method performed manually to select indicators based on both statistical and non-statistical criteria. The statistical criterion is the c-statistic, also known as the concentration index (Ravallion, 2007) or the area under the receiver operating characteristic curve (Wodon, 1997). A higher value of this criterion indicates better model performance (Baulch, 2002). Non-statistical criteria are our judgment about various aspects of the indicators, such as user acceptability, simplicity, verifiability, applicability, variety, and so on. For more details about the “MaxC” forward stepwise approach, please refer to Schreiner (2023). The “MaxR” approach is similar to “MaxC” approach, but it uses adjusted R-squared as the statistical criterion. For more details about the “MaxR” approach, please refer to Zeller etal. (2005). Step 3: Training prediction model In this step, we use selected indicators to develop prediction models and obtain weights for those indicators. Depending on the methodology, the prediction models may use Table 3 An example of manual categorization by using cross-tabulation analysis. Source: Authors’ computations based on the VHLSS 2018 data Initial categorization Category (Number of household members) 1 2 3 4 5 6 ≥ 7 Number of poor and non-poor (Poor: Non-poor) 28:145 58:548 78:893 215:1298 201:810 146:524 180:381 Poor rate of each category (%) 16.2 9.6 8.0 14.2 19.9 21.8 32.1 Second categorization Category (Number of household members) < 3 3 4 5 6 ≥ 7 Poor rate of each category (%) 11.0 8.0 14.2 19.9 21.8 32.1 Final categorization Category (Number of household members) < 4 4 56 ≥ 7 Poor rate of each category (%) 9.4 14.2 19.9 21.8 32.1
312 B.T.Duong et al. Logit, OLS, or Quantile regression algorithms. Detailed information about the regression algorithms and dependent variables used will be provided in the description of each method below. Step 4: Transforming the prediction model’s coefficients to score Based on the results of prediction models, the coefficients of indicators can be transformed into scores to form targeting tools in three simple steps. First, we switch reference categories of variables if their coefficients are negative. This ensures all coefficients are positive. Second, we multiply the coefficients of the variables by 100. Finally, we round the scores obtained. The sum of these scores (excluding the intercept’s coefficient) is the total score for predicting the poverty statusof households. This common and straightforward approach is widely used to create targeting tools and is also applied by the Vietnamese government for their existing tools (Nguyen & Lo, 2016). The Scorecard and WOE (Logit) methods use different approaches to transform coefficients into scores, which will be explained later. Step 5: Estimating the tool’s accuracy We use coverage rate as a key criterion to compare the accuracy of targeting tools. The coverage rate, also known as the inclusion of the poor or poverty accuracy, is currently used by the Vietnamese government to assess the accuracy of their targeting tools (IRIS Center, 2005; Nguyen & Lo, 2016). The coverage rate is the percentage of households correctly predicted as poor by a given targeting tool, expressed as a percentage of the total number of observed poor households. In addition, we also measure the coverage rate of targeting tools at the person level for comparison. This means we also calculate the coverage rate of targeting tools when identifying poor individuals. This is important because identifying poor households with larger family sizes is more beneficial than identifying those with fewer members. We measure the coverage rate of targeting tools when they perform in validation samples. This helps us avoid the overfitting issue because targeting tools certainly perform well in calibration samples, which are used to construct them. It also allows us to gauge the accuracy of targeting tools in identifying the poor within an unobserved population. Additionally, we apply a bootstrap technique to estimate the standard error of the coverage and targeted rates by calculating those measures and their differences for the 1,000 repetitions from the bootstrap sample (Chen & Schreiner, 2009; Efron & Tibshirani, 1994). This approach provides a reliable assessment of tools’ performance. For a fair comparison, we compare the coverage rates at a cut-off with the same predicted poverty rate, which is the rate of the targeted population in validation samples. Due to the granularity of the scores, it is challenging to obtain cut-off scores with the same predicted poverty rate across all targeting tools. Therefore, we fix the poverty rate predicted by a specific targeting tool and use linear interpolation to estimate the coverage rates of the remaining targeting tools. 2.3.2 The Scorecard (Logit) Method Scorecards are alsodeveloped based on the principles of the PMT (OLS) method. Scorecard (Logit) method uses the scores of proxy indicators to construct Scorecards for identifying the poor. Schreiner (2010) applies logistic regression to estimate the weights of proxy indicators, and uses his own method to convert the indicators’ weights to non-negative
319 Can Weight ofEvidence Method Improve Poverty Targeting? of the national WOE tool is 1.9, 5.9, and 4.3 percentage points higher than those of the nationalScorecard, PMT, and PAT, respectively. Regarding the international poverty line, Table 4 shows that the international WOE tool achieves the highest accuracy among international targeting tools, with a coverage rate of 54.3 percentage points. In comparison, the international Scorecard and PMT have lower accuracy, with coverage rates of 51.6 and 51.3 percentage points, respectively, while the international PAT has the lowest accuracy, with a coverage rate of 50.9 percentage points. This means that the coverage rate of the international WOE tool is 2.7, 3.0, and 3.4 percentage points higher than those of the international Scorecard, PMT, and PAT, respectively. To estimate standard errors for the coverage rates and predicted poverty rates of national and international targeting tools, the bootstrap sampling method requires the actual cutoff scores of these tools. However, as mentioned earlier, national targeting tools do not have exact cut-off scores that achieve the same predicted poverty rate, and neither do international targeting tools. Therefore, to estimate standard errors, we use cut-off scores that yield predicted poverty rates closest to the observed poverty rates. The results in Table5 show that the standard errors of the national WOE tool differ only slightly from those of other national targeting tools, and its p-value is identical to those of the other national tools. Similarly, for the international poverty line, Table 5 indicates that the standard errors of the international WOE tool are comparable to those of other international targeting tools, and the p-values of all four international targeting tools are the same. 3.2 The Methods’ Accuracy inIdentifying Poor People The accuracy comparison of the four methods in identifying poor individuals follows the same approach used for identifying poor households. We apply the previously constructed targeting tools to the corresponding poverty lines and calculate the coverage rate for each tool. The key difference is that the national and international poverty lines are now measured at the individual level (household member). As a result, these poverty lines are higher than those used in the household-level analysis. Table 6 presents the coverage and predicted poverty rates of the national Scorecard, which are 64.1 and 20.8 percentage points, respectively. This means that when targeting 20.8% of the population at the individual level, the national Scorecard correctly identifies 64.1% of them as poor people. For the same targeted population share, the coverage rates of the national PMT, PAT, and WOE tools are 62.2, 62.7, and 65.5 percentage Table 6 The accuracy of targeting tools in identifying poor people by the poverty line Unit: (%) Scorecard PMT PAT WOE tool National poverty line (in person) Coverage rate 64.1 62.2 62.7 65.5 Predicted poverty rate 20.8 20.8 20.8 20.8 International poverty line (in person) Coverage rate 55.5 56.9 56.2 58.8 Predicted poverty rate 17.2 17.2 17.2 17.2
320 B.T.Duong et al. Table 7 The standard error and p-value of targeting tools in identifying poor people by the poverty line National poverty line (in person) Unit (%) Std. error p-value International poverty line (in person) Unit (%) Std. error p-value Scorecard Coverage rate 64.1 1.078 0.001 Scorecard Coverage rate 55.5 1.173 0.001 Predicted poverty rate 20.8 0.408 0.001 Predicted poverty rate 17.2 0.354 0.001 PMT Coverage rate 61.2 1.119 0.001 PMT Coverage rate 56.7 1.200 0.001 Predicted poverty rate 20.3 0.407 0.001 Predicted poverty rate 17.0 0.356 0.001 PAT Coverage rate 61.2 1.108 0.001 PAT Coverage rate 55.1 1.194 0.001 Predicted poverty rate 20.1 0.407 0.001 Predicted poverty rate 16.9 0.359 0.001 WOE tool Coverage rate 64.2 1.082 0.001 WOE tool Coverage rate 58.5 1.168 0.001 Predicted poverty rate 20.2 0.391 0.001 Predicted poverty rate 17.0 0.352 0.001
321 Can Weight ofEvidence Method Improve Poverty Targeting? points, respectively, as shown in Table6. These results indicate that the national WOE tool achieves the highest accuracy in identifying poor individuals based on the national poverty line. Specifically, the coverage rate of the national WOE tool is 1.5, 3.3, and 2.8 percentage points higher than those of the nationalScorecard, PMT, and PAT, respectively. Regarding the identification of poor individuals based on the international poverty line, the international WOE tool also demonstrates the highest accuracy among all targeting tools. Specifically, its coverage rate exceeds those of the Scorecard, PMT, and PAT by 3.2, 1.9, and 2.6 percentage points, respectively. In identifying poor individuals, both national and international WOE tools exhibit only minor differences in standard errors compared to those ofother targeting tools, as shown in Table7. Additionally, the p-values of the WOE tools are the same as those of the other targeting tools. 3.3 Discussion The comparisons above show that the WOE (Logit) method outperforms the other three methods. However, due to differences among the four methods and the lack of controlled conditions, pinpointing the exact reasons for this outperformance is challenging. Nonetheless, we will highlight key differences between the methods to examine how these variations potentially influence accuracy and indicator selection in targeting tools. The first difference lies in how responses for each indicator are binned. Other methods create categories based on the results of cross-tabulation, aiming to create groups that consistently align with poverty status. In the WOE (Logit) method, binning with the WOE technique does not stop at that requirement, but aims to maximize IV, the predictive power of an indicator. This approach often results in different binning patterns for variables with a wide range of values, such as “electricity bill,” “number of poultry,” “dependent ratio,” and so on. This is because without the criterion of IV, it is difficult to determine the best category combinations for these variables. Variables with different ways of binning can lead to differences in selecting indicators because, within the same variable, different category combinations will result in different values of the c-statistic or adjusted R square. These are essential statistical criteria in the “MaxC” or “MaxR” forward stepwise approaches, which select indicators based on small improvements in these criteria. As a result, binning with different methods can lead to the selection of different indicators in targeting tools, which ultimately affects accuracy. This can be seen in comparing the WOE (Logit) and Scorecard (Logit) methods: despite both applying the same “MaxC” stepwise approach and using Logit models, their targeting tools differ in one indicator, and two to three indicators have different binning patterns. Consequently, their accuracy differs. Additionally, for the Scorecard (Logit) and WOE (Logit) methods, the c-statistic value is also an important criterion for assessing the accuracy of targeting tools. We observe that variables with higher IVs often4 tend to achieve higher c-statistic values. Since the WOE (Logit) method maximizes the IV of indicators during thebinning process, numeric variables with a wide range of values binned in this way often have higher c-statistic values compared to those binned using cross-tabulation. While this disparity is minor, it can accumulate when targeting tools contain many such indicators. As a result, this difference may contribute to the accuracy gap between Scorecards and WOE tools. 4 This trend weakens and not always correct when differences in IVs are very minimal.
322 B.T.Duong et al. Another factor that may contribute to differences in indicator selection and targeting accuracy is the nature of the variables used in the regressions. The WOE (Logit) method employs numeric values transformed using the WOE technique, whereas other methods rely on categorized variables. However, we cannot definitively conclude that using WOEtransformed variables directly impacts accuracy or indicator selection, as the indicators used in the four targeting tools differ, potentially introducing bias. An observation we can confirm is that regressions using categorized variables sometimes encounter issues with coefficient monotonicity. The Scorecard (Logit) and PMT (OLS) methods use categorical variables, so each category has its own coefficient. This can lead to situations where a categorical variable might help the model achieve the highest c-statistic or highest adjusted R-square but cannot be selected because their coefficients are non-monotonic.5 In such cases, we do not drop these variables; instead, we recategorize them to achieve monotonic coefficients and rerun the stepwise process. Whether the Fig. 1 The association between the total score and predicted outcomes at person level by international targeting tools 5 A variable with non-monotonic coefficients leads to inconsistencies in the scoring of responses within a question.
323 Can Weight ofEvidence Method Improve Poverty Targeting? re-selected indicator remains the same or not, this change often results in a model with a lower c-statistic or adjusted R-square. This issue arises when selecting indicators such as the number of dependents and number of agricultural laborers for the national Scorecard, and the highest education level for the international Scorecard. The PMT (Quantile) method faces similar challenges, so to maintain consistent coefficients across categories, we need to select quantiles that result in lower c-statistic values. In contrast, the WOE (Logit) method does not face this issue because WOE variables are numeric, and each variable has only one coefficient. The third difference between methods is the regression algorithms used. However, we do not consider this factor as a main reason for the accuracy improvement seen in the WOE (Logit) method. Existing studies show that the Logit model sometimes outperforms OLS or Quantile regressions, but at other times, Logit models are less effective (Linh & Baulch, 2011; Schreiner, 2015; Zeller etal., 2006). This pattern is also observed in this study. For instance, in identifying poor people, the national Scorecard achieves higher accuracy than the national PMT and PAT, while the international Scorecard performs worse than the international PMT and PAT. In some cases, both the Logit and OLS regressions yield similar targeting accuracy (Etang etal., 2022). Additionally, although both the WOE (Logit) and Scorecard (Logit) methods use logistic regression, WOE tools still outperform Scorecards. The final difference between the WOE (Logit) method and the other methods is the approach to transforming coefficients into scores. Figure1 illustrates that the total scores converted by the WOE (Logit) method are more consistent with the predicted poverty probabilities than those converted by the Scorecard (Logit) method. This suggests that the Scorecard (Logit) method may inadvertently lower its accuracy, as many individuals with similar or even higher predicted poverty rates at the selected cut-off score may still be excluded. The magnitude of the total score could be a contributing factor to this issue. However, the score transformation approaches of the PMT (OLS) and PMT (Quantile) methods are similar, with large total score magnitudes. While the total scores in the PMT (Quantile) method align with individuals’ expenditures, the total scores in the PMT (OLS) method do not. As a result, both the PMT (OLS) and PMT (Quantile) methods generally perform worse than the WOE (Logit) method. Nevertheless, concluding that the score transformation approach of the WOE (Logit) method is superior to others could be biased, as the targeting tools constructed by the four methods use different indicators, and the magnitude of the total scores varies across the methods. This discussion highlights the differences between methods and the potential outcomes resulting from these differences. Based on these outcomes, we can only conclude that using both the WOE binning technique and WOE variables leads to differences in binned and selected variables compared to other methods. We cannot provide precise quantitative assessments of accuracy improvements for individual techniques within the WOE (Logit) method because the necessary conditions for a fair assessment are not met, and this is beyond the scope of the study. However, this presents an interesting research gap to explore in the future. 4 Conclusion This study explores the feasibility and efficacy of using the WOE (Logit) method to develop targeting tools. To achieve this, we compare the accuracy of the WOE (Logit) method with that of other powerful and common methods, such as the Scorecard (Logit),
324 B.T.Duong et al. PMT (OLS), and PMT (Quantile) methods. For a fair comparison, we use the coverage rate to compare the accuracy between targeting tools while holding the predicted poverty rate. The results indicate that targeting tools constructed using the WOE (Logit) method achieve higher accuracy than those developed by other methods. Depending on the national or international poverty line, the WOE (Logit) method can improve accuracy in identifying poor households by approximately 1.9–5.9 percentage points of coverage rate compared to other methods, while this accuracy improvement ranges from 1.5 to 3.3 percentage points in identifying the poor at the person level. This small improvement is meaningful on a national scale. For instance, according to the updated report by the Committee for Ethnic Minority Affairs of Vietnam (CEMA, 2020), Vietnam has 745,441 poor households among ethnic minorities. If the Vietnamese government aims to target all of these poor households, applying the WOE (Logit) method could potentially increase the number of correctly identified poor households by 14,163 to 43,981 compared to other methods. This study introduces an additional approach for governments seeking to develop accurate targeting tools. While the WOE (Logit) method may seem complex in explanation, it is neither new nor difficult to implement, as it is commonly used in credit scoring. Various accessible resources, guidelines, and software packages are available for popular and powerful tools such as R, SAS, and Python. These packages allowusers to construct targeting tools using the WOE (Logit) method in just a few simple steps. The first limitation of this study is that it focuses solely on testing whether applying the WOE (Logit) method—a comprehensive approach for constructing credit risk scorecards, including techniques such as binning categories by WOE and IV, using WOE variables, applying logistic regression, and using the score-scaling approach of doubling the poverty odds—improves the accuracy of targeting tools. Therefore, the study can only conclude that the application of all these techniques together improves the accuracy of WOE tools. The individual contributions of each component technique to the accuracy improvement remain unclear, as many conditions are not controlled for a fair comparison, making any conclusion potentially biased. Nevertheless, this is a research gap that should be explored in the future if any researcher intends to apply any single technique in theWOE (Logit) method to the PMTs. A second limitation is that the WOE (Logit) method has only been applied to ethnic minority communities in Vietnam. However, targeting tools developed using this method have been effective in identifying the poor at both the household and person levels among the 53 ethnic minority groups, which vary in characteristics, living regions, poverty rates, and levels of poverty (poor, moderately poor, or extremely poor) (CEMA, 2020). Therefore, the WOE (Logit) method should also be applicable for constructing targeting tools to identify the poor among the ethnic majority, which includes only the Kinh ethnic group, in Vietnam. We recommend that the comparative analysis of the WOE (Logit) method presented in this paper be repeated for other countries and data sets. If results are similar to those presented here, the WOE (Logit) method can be applied to poverty scoring in those countries. Appendix See Tables8, 9 and 10.
325 Can Weight ofEvidence Method Improve Poverty Targeting? Table 8 National Scorecard constructed by the Scorecard (Logit) method Indicator Value Points Score 1. How much did you pay for your electricity bill last month? A. Less than 50,000 VND 0 B. 50,000 VND—99,000 VND 3 C. 100,000 VND—199,000 VND 7 D. 200,000 VND—349,000 VND 11 E. 350,000 VND or more 19 2. How many members of the household work as non-agricultural laborers (includes side jobs)? A. One or none 0 B. Two 7 C. Three 12 D. Four or more 17 3. How many members does the household have? A. Less than four 9 B. Four 5 Five 3 Six 2 Seven or more 0 4. Was the household included in the anti-poor program in the previous year (2017)? A. Yes 0 B. No 5 5. What is the total land area owned by the household? A. Less than two hectares 0 B. Two hectares or more 6 6. Is anyone in the household working away from home? A. Yes 10 B. No 0 7. How many members of the household work in the agricultural sector as their main job? A. None 12 B. One 9 A. Two, Three or Four 4 B. Five or more 0
326 B.T.Duong et al. Table 8 (continued) Indicator Value Points Score 8. How many members of the households are dependents? A. None 8 B. One or two 4 C. Three 1 D. Four or more 0 9. Does the household own a cow, buffalo, or horse? A. No 0 B. Yes 6 10. How many motorbikes does the household own? A. None 0 B. One 3 C. Two 6 D. Three or more 9 Total score
327 Can Weight ofEvidence Method Improve Poverty Targeting? Table 9 National WOE tool constructed by the WOE (Logit) method Indicator Value Points Score 1. How much did you pay for your electricity bill last month? A. Less than 40,000 VND 17 B. 40,000 VND—89,000 VND 30 C. 90,000 VND—159,000 VND 43 D. 160,000 VND—239,000 VND 57 E. 240,000 VND or more 74 2. How many members of the household work as non-agricultural laborers (includes side jobs)? A. One or none 12 B. Two 40 C. Three 57 D. Four or more 65 3. How many members does the household have? A. Less than four 44 B. Four 37 C. Five 31 D. Six 30 E. Seven or more 22 4. Was the household included in the anti-poor program in the previous year (2017)? A. Yes 24 B. No 45 5. How many poultries does the household own? A. Less than 30 32 B. 30 to 39 35 C. 40 to 59 53 D. 60 or more 74 6. Is anyone in the household working away from home? A. Yes 71 B. No 32 7. How many motorbikes does the household own? A. None 26 B. One 32 C. Two 51 D. Three or more 64
328 B.T.Duong et al. Table 9 (continued) Indicator Value Points Score 8. Does the household own a cow, buffalo, or horse? A. No 33 B. Yes 58 9. How many members of the household work in the agricultural sector as their main job? A. None 68 B. One 50 C. Two or three 31 D. Four 24 E. Five or more 15 10. How many members of the households are dependents? A. None 53 B. One 40 C. Two 34 D. Three 20 E. Four or more 12 Total score
