scieee AI-readable full text Open interactive document viewer

An overview on the landscape of R packages for open source scorecard modelling

Szepannek, Gero

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Szepannek, Gero Article An overview on the landscape of R packages for open source scorecard modelling Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Szepannek, Gero (2022) : An overview on the landscape of R packages for open source scorecard modelling, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 10, Iss. 3, pp. 1-33, https://doi.org/10.3390/risks10030067 This Version is available at: https://hdl.handle.net/10419/258377 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/   Citation: Szepannek, Gero. 2022. An Overview on the Landscape of R Packages for Open Source Scorecard Modelling. Risks 10: 67. https:// doi.org/10.3390/risks10030067 Academic Editors: Krzysztof Jajuga and Józef Dziechciarz Received: 16 February 2022 Accepted: 14 March 2022 Published: 18 March 2022 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2022 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). risks Article An Overview on the Landscape of R Packages for Open Source Scorecard Modelling Gero Szepannek Institute of Applied Computer Science, Stralsund University of Applied Sciences, Zur Schwedenschanze 15, 18435 Stralsund, Germany; [email protected] Abstract: The credit scoring industry has a long tradition of using statistical models for loan default probability prediction. Since this time methodology has strongly evolved, and most of the current research is dedicated to modern machine learning algorithms which contrasts with common practice in the finance industry where traditional regression models still denote the gold standard. In addition, strong emphasis is put on a preliminary binning of variables. Reasons for this may be not only the regulatory requirement of model comprehensiveness but also the possibility to integrate analysts’ expert knowledge in the modelling process. Although several commercial software companies offer specific solutions for modelling credit scorecards, open-source frameworks for this purpose have been missing for a long time. In recent years, this has changed, and today several R packages for credit scorecard modelling are available. This brings the potential to bridge the gap between academic research and industrial practice. The aim of this paper is to give a structured overview of these packages. It may guide users to select the appropriate functions for the desired purpose. Furthermore, this paper will hopefully contribute to future development activities. Keywords: credit scorecard development; open source; R 1. Introduction In the credit scoring industry, there is a long tradition of using statistical models for loan default probability prediction, and domain specific standards were established long before the hype of machine learning. An overview of the historical evolution of credit risk scoring can be found in Kaszynski (2020) and Anderson (2019). A comprehensive description of the corresponding methodology is given in Thomas et al. (2019) and Kaszynski et al. (2020). The different subsequent steps during the scorecard modelling process are worked out in Anderson (2007), Finlay (2012) and Siddiqi (2006) where the latter is closely related to the credit scoring solution as implemented by the SAS Enterprise Miner software 1 . The typical steps in credit risk scorecard modelling refer to the general process definition for data mining as given by KDD, CRISP-DM or SAS’s SEMMA (cf. Azevedo and Santos 2008). It turns out that strong emphasis is laid on possibilities for manual intervention after each modelling step. Therefore, functions to summarize and visualize the intermediate results of each single step are of great importance. The typical development steps are denoted by: 1. Binning and Weights of Evidence (Section 3) 2. Preselection of Variables (Section 4) 3. Multivariate Modelling (Section 5) 4. Performance Evaluation (Section 6) 5. Reject Inference (Section 7) In contrast, the typical scorecard modelling process is rarely taken into account in current academic benchmark studies (for an overview cf. Louzada et al. 2016). An exception is given in Bischl et al. (2016), where both approaches are covered. A reason for this gap between academic research and business practice may be due to the lack of open source frameworks for scorecard modelling. Risks 2022,10, 67. https://doi.org/10.3390/risks10030067 https://www.mdpi.com/journal/risks Risks 2022,10, 67 2 of 33 Although several commercial software companies, such as SAS, offer specific solutions for credit scorecard modelling (cf. footnote 1), explicit packages for this purpose in R have been missing for a long time and in the CRAN task view on Empirical Finance 2 the explicit topic of scorecard modelling is not covered. A “Guide to Credit Scoring in R” can be found among the CRAN contributed documentations (Sharma 2009) being dedicated to describing the application of different (binary) classification algorithms to credit scoring data rather than to emphasizing the common subsequent modelling stages that are typical for scorecard modelling processes. This can be a result of the circumstances: at that time, no explicit packages were available in R for undertaking this kind of task. In recent years this has changed, and several packages have been submitted to CRAN with the explicit scope of credit risk scorecard modelling, such as creditmodel (Fan 2022), scorecard (Xie 2021), scorecardModelUtils (Poddar 2019), smbinning (Jopia 2019) woeBinning (Eichenberg 2018), woe (Thoppay 2015), Information (Larsen 2016), InformationValue (Prabhakaran 2016), glmdisc (Ehrhardt and Vandewalle 2020), glmtree (Ehrhardt 2020), Rprofet (Stratman et al. 2020) and boottol (Schiltgen 2015). Figure 1gives an overview of the packages and their popularity in terms of the number of their CRAN downloads as well as their activity and existence as observable by their CRAN submission dates. It can be seen that the packages smbinning , InformationValue and Information are among the most popular, and they have been available for quite some time. Another popular toolbox is provided by the package scorecard , which has been frequently updated in the recent past as has also happened with the package creditmodel . Figure 1. CRAN release activity and download statistics (as returned by cranlogs ,Csárdi 2019) of packages available on CRAN. In addition, some packages are available on Github but not on CRAN, such as creditR (Dis 2020), riskr (Kunst 2020), and scoringTools (Ehrhardt 2018). As all of these packages have become available during the last few years, this paper is dedicated to the question of whether recent developments have made it possible to perform all steps of the entire scorecard development process within R. For this reason, the presentation of the package landscape will be guided by these steps. One section will be dedicated to each stage. In each section, the available packages will be presented together with their advantages and disadvantages. The aim of the paper is to give a structured overview of existing packages. It may guide users in selecting the appropriate functions for the desired purpose by working out pros and cons of existing functions. As an open source programming language, the R universe is extended by a large community with currently more than 19,000 contributed packages. It is impossible for a single user to know all of them which in turn leads to some redundant development activities by programming multiple solutions for the same task. Moreover, sometimes Risks 2022,10, 67 3 of 33 contributed packages for a similar purpose provide different desirable functionalities but are not compatible with each other because they rely on different kind of input objects. By working out the pros and cons of the functions provided by the aforementioned packages, this paper aims to analyze existing gaps and provide several remedies in the supplementary code (cf. corresponding footnotes). Note that this paper focuses on the open source statistical programming language R. Within the data science industry other open source frameworks such as python have increased in popularity during the last few years, which is beyond the scope of this paper. For interested readers, some useful python functions for scorecard development are mentioned in Kaszynski et al. (2020), and some websites are dedicated to this purpose 3,4 . In particular, a python implementation of the scorecard R package (Xie 2021) is available 5 which means that some of the results as worked out in this paper are directly transferable into the python world. Nonetheless as R denotes the lingua franca of statistics Ligges (2009), it provides access to a huge number of contributed packages and functionalities from the field of statistics outside the aforementioned ones. For this reason, the paper concentrates on R and investigates whether scorecard development can be improved by access to other already existing packages that have initially been designed for other purposes but can improve the analyst’s life. If available, such functionalities will also be mentioned in the corresponding sections. Note that, traditionally, logistic regression is used for credit risk scorecard modelling despite the current hype around modern machine learning methods as they are provided by frameworks such as e.g., mlr3 (Lang et al. 2019,2021) or caret (Kuhn 2008,2021). Studies have investigated potential benefits from using modern machine learning algorithms (Baesens et al. 2002;Bischl et al. 2016;Lessmann et al. 2015;Louzada et al. 2016;Szepannek 2017), but regulators and the General Data Protection Regulation (GDPR) require models to be understandable (cf. Financial Stability Board 2017;Goodman and Flaxman 2017). The latter issue can be addressed by methodologies of explainable machine learning (for an overview Bücker et al. 2021), e.g., using frameworks as provided by the packages DALEX (Biecek 2018) or iml (Molnar et al. 2018) while taking into account to what extent a model actually is explainable (Szepannek 2019). It further turned out that the use of current state-of-the-art ML algorithms is not necessarily always beneficial in the credit scoring context (Chen et al. 2018;Szepannek 2017), and they should be rather carefully analyzed in each specific situation, rather than relying on preferred preferred models (Rudin 2019). For this reason this paper focuses on the traditional way of scorecard modelling as briefly described above. 2. Data Probably the most common credit scoring data are the German Credit Data provided by Hoffmann (1994) that are contained in the UCI Machine Learning Repository (Dua and Graff 2019). The data consist of 21 variables: a binary target ( creditability ) and 13 categorical as well as seven numeric predictors, and 1000 observations in total with 300 defaults ( level == “bad” ) and 700 nondefaults ( level == “good” ). The data are provided by several R packages such as klaR (Roever et al. 2020), woeBinning , caret or scorecard . For the examples in this paper, the data from the scorecard package are used where in addition the levels of the categorical variables such as present.employment.since , other.debtors.or.guarantors , job or housing are sorted according to their expected order w.r.t. credit risk. Note that Groemping (2019) compared the data from the UCI repository to the original papers and made a corrected version of itavailable 6 (cf. also Szepannek and Lübke 2021). Other (partly simulated) example data sets (amongst others loan data of the peer-to-peer lending company Lending Club 7 ) are contained within the packages creditmodel,scoringTools and smbinning and riskr. It is common practise to use separate validation data which are not used for model training but only for validation purposes. The manual interventions between the different modelling steps do not allow for repetitive resampling strategies such as k-fold cross Risks 2022,10, 67 4 of 33 validation or bootstrapping for model validation as they are, e.g., provided by the package mlr3 (see Section 6). Instead, usually one single holdout set is used. The package scorecard has a function split_df() that splits data according to a prespecified percentage into training and validation sets. For the examples in the remainder of the paper, the following data are used: ### example 1: load data library(scorecard) data(germancredit) # transform character variable purpose into factor germancredit$purpose <- as.factor(germancredit$purpose) tv <- split_df(germancredit, y=‘‘creditability’’,ratio = c(0.7,0.3), seed = 42,no_dfs = 2,name_dfs = c(‘‘train’’,‘‘valid’’)) train <- tv$train valid <- tv$valid # several packages require the target variables to take values 0/1 train2 <- train; valid2 <- valid train2$creditability <- as.integer(train2$creditability == ‘‘good’’) valid2$creditability <- as.integer(valid2$creditability == ‘‘good’’) # the package creditmodel does not support variables of type Factor train3 <- as.data.frame(train2) valid3 <- as.data.frame(valid2) for (j in which(sapply(train3[,-21], is.factor))) { train3[,j] <- as.character(train3[,j]) valid3[,j] <- as.character(valid3[,j]) } Note that some of the packages ( smbinning , woe , creditR , riskr , glmdisc , scoringTools , scorecardModelUtils and creditmodel8 ) do require the target variable to take only values 0 and 1 as in the example’s data sets train2 and valid2 . Although this is of course easily obtained, the package scoringModelUtils contains a function fn_target() that does this job and replaces the original target variable with a new one of name Target. 3. Binning and Weights of Evidence 3.1. Overview Binning of numeric variables is often considered the most relevant step in scorecard development. An initial automatic algorithm-based binning is manually checked and— if necessary—modified by the analyst variable by variable. On the one hand, this is a very time-consuming task, but, on the other hand, this ensures the dependencies between the explanatory variables and the target in the final model to be plausible and helps detect sampling bias (Verstraeten and den Poel 2005). Furthermore, it allows modelling of nonlinear dependencies by linear logistic regression in the subsequent Multivariate Modelling step. The loss of information by aggregation turned out to be comparatively small while this kind of procedure does not take into account for interactions between several variables and the target variable (Szepannek 2017). The identification of relevant interactions typically needs a lot of business experience, and Sharma (2009) suggests using random forests to identify potential interaction candidates. 3.2. Requirements It is important to note that binning corresponds not just to exploratory data analysis, but its results have to be considered an integral part of the final model, i.e., the resulting Risks 2022,10, 67 5 of 33 preprocessing has to be applied to new data to be able use the resulting scorecard for business purposes. For this reason, important requirements on an implementation of the binning step are the possibility to: (i) store the binning results for all variables, and (ii) apply the binning to new data with some kind of predict() function. The importance of an option to: (iii) manually modify an initial automatic binning has already been emphasized. This leads to the requirement for a separate function to manipulate an object that stores the binning results. In order to support this: (iv) summary tables and (v) visualizations of the intermediate binning results are helpful. In addition, application of binning in practice has to: (vi) deal with missing data or new levels of categorical variables that did not occur in the training data as, e.g., by regulation it may be required that holding back information (and the resulting missing values) must not lead to an improvement of the final score. Both missing data and new levels should be taken into account by the implemented binning function. Often, binning is followed by subsequent assignment of numeric weights of evidence to the factor levels xof the binned variable which are given by: WoE(x) = logf(x|y=1) f(x|y=0). (1) Note that just like the bins, the WoEs, as computed on the training data are part of the model. Furthermore, an implementation of WoE computation has to account for potentially occurring bins that are empty w.r.t. the target level y= 0 (typically by adding a small constant when computing the relative frequencies f() ). By construction, WoEs are linear in the logit of the target variable and thus well suited for subsequent use of logistic regression. The use of WoEs is rather advantageous for small data sets, and directly using the bins may increase performance if enough data are available (Szepannek 2017). On the other hand, using WoEs fixes mononty between the resulting scorecard points and the default rates of the bins, such that only the sign of the monotonicity has to be checked. It is also usual to associate binned variables with an information value (IV) IV =∑ x (f(x|y=1)−f(x|y=0)) WoE(x)(2) based on the WoEs which describe the strength of a single variable to discriminate between both classes. 3.3. Available Methodology for Automatic Binning Several packages provide functions for automatic binning based on conditional inference trees (Hothorn et al. 2006) from the package partykit (Hothorn and Zeileis 2015): scorecard::woebin() , smbinning::smbinning() , scorecardModelUtils::iv_table() and riskr::superv_bin() . The implementation in the scorecardModelUtils package merges the resulting bins to ensure monotonicity in default rates w.r.t. with the original variable which might or might not be desired. For the same purpose, the package smbinning offers a separate function ( smbinning.monotonic() ). In contrast to all previously mentioned packages, the package woeBinning implements its own tree algorithm where either initial bins of similar WoE are merged ( woe.binning() ), or the set of bins is binary split ( woe.tree.binning() ) as long as the IV of the resulting variables decreases (increases) by a percentage less (more) than a prespecified percentage (argument stop.limit ) while the initial bins are created to be of minimum size ( min.perc.total ). The function creditmodel::get_breaks_all() uses classification and regression trees (Breiman et al. 1984) of the package rpart (Therneau and Atkinson 2019) 9 to create initial bins. An additional argument, best = TRUE , merges these bins subsequently according to different criteria such as the maximum number of bins, the minimum percentage of observations per bin, a threshold for the χ2 test or odds, a minimum population stability (cf. Section 4) Risks 2022,10, 67 6 of 33 or monotonicity of the default rates across the bins (all of these can be specified by the argument bins_control). In addition to tree-based binning, the scorecard package offers alternative algorithms (argument method ) for automatic binning based on either the χ2 statistic or equal width or size of numeric variables. An alternative concept for automatic binning is provided by the package glmdisc , which is explicitly designed to be used in combination with logistic regression modelling for credit scoring (Ehrhardt et al. 2019). The bins are optimized to maximize either AIC, BIC or the Gini coefficient (cf. Section 6) of a subsequent logistic regression model (using binned variables, not WoEs) on validation data (argument criterion= ). Second order interactions can also be considered (argument interact = TRUE ). Note that this approach is comparatively intense in terms of computation time and does not take variable selection into acount (cf. Section 5). Some packages do not provide their own implementations of an automatic binning but just interface to discretization functions within other packages. Rprofet::BinProfet() uses the function greedy.bin() of the package binr (Izrailev 2015). The package scoringTools contains a variety of functions ( chiM_iter() , mdlp_iter() , chi2_iter() , echi2_iter() , modchi2_iter() and topdown_iter() ) which provide interfaces to binning algorithms from the package discretization (Kim 2012). The dlookr package (Ryu 2021) , which is primarily designed for exploratory data analysis, has an implemented interface (binning_by()) to smbinning::smbinning(). 3.4. Manipulation of the Bins As outlined before, manual inspection and manipulation of the bins is considered a substantial part of the scorecard development process. Two of the aforementioned packages provide functions to support this. Scorecard::woebin() allows passing an argument breaks_list . Each element corresponds to a variable with manual binning and must be named like the corresponding variable. For numeric variables, it must be a vector of break points, and for factor variables, it must be a character vector of the desired bins given by the merged factor levels, separated by “ %,% ” (cf. output from Example 3 for variable purpose). In addition, a function scorecard::woebin_adj() allows for an interactive adjustment of bins. The package smbinning provides two functions, smbinning.custom() and smbinning.factor.custom(). Manipulation of the bins should be based on an analysis of the binning results. For this purpose, most of the packages provide result tables on a variable level. The subsequent code example illustrates the step of an initial automatic binning as created by the package scorecard: ### Example 2: automatic binning library(scorecard) bins <- woebin(train, y=‘‘creditability’’,method = ‘‘tree’’) # binning results table for variable purpose options(digits = 3) bins$purpose[,c(2,4,5,6,7,8)] # visualize bins for variable purpose woebin_plot(bins, x=‘‘purpose’’,line_value = ‘‘woe’’) Risks 2022,10, 67 7 of 33 Version March 10, 2022 submitted to Risks 7 of 34 interactive adjustment of bins. The package smbinning provides two functions smbinning.custom() 199 and smbinning.factor.custom() . 200 Manipulation of the bins should be based on an analysis of the binning results. For this purpose, 201 most of the packages provide result tables on a variable level. The subsequent code example illustrates 202 the step of an initial automatic binning as created by the package scorecard : 203 ### example 2: automatic binning library(scorecard) bins <- woebin(train, y = "creditability",method = "tree") # binning results table for variable purpose options(digits = 3) bins$purpose[,c(2,4,5,6,7,8)] # visualize bins for variable purpose woebin_plot(bins, x="purpose",line_value = "woe") ## bin count_distr neg pos posprob woe 204 ## 1: business%,%car (new) 0.3211 148 79 0.348 0.213 205 ## 2: car (used) 0.1089 67 10 0.130 -1.061 206 ## 3: domestic appliances%,%education 0.0622 24 20 0.455 0.659 207 ## 4: furniture/equipment%,%others 0.1938 90 47 0.343 0.192 208 ## 5: radio/television%,%repairs%,%retraining 0.3140 165 57 0.257 -0.222 209 The resulting table contains several key figures for each bin, such as the distribution (absolute and 210 relative frequency of the samples given the level of the target variable), default rate and the bin’s WoE. 211 The information value of the binned variable (cf. Sec. 4) is given in a column total_iv (not shown 212 here). 213 In addition to summary tables, many packages ( glmdisc , riskr , Rprofet , scorecard , smbinning , 214 woeBinning ) provide a visualization of the bins on a variable level. Fig. 2(left) shows the binning as 215 resulting from code in example 2 which is pretty similar for most packages. A mosaic plot of the bins, 216 which simultaneously visualizes default rates as well as size of the bins, is offered by the package 217 glmdisc (Fig. 2, right) while the names of the bins after automatic binning are not self-explanatory. 218 Figure 2. Visualization of the bins for the variable purpose as created by the package scorecard (left) and mosaicplot of the binning result by the package glmdisc (right). The resulting table contains several key figures for each bin such as the distribution (absolute and relative frequency of the samples given the level of the target variable), default rate and the bin’s WoE. The information value of the binned variable (cf. Section 4) is given in a column total_iv (not shown here). In addition to summary tables, many packages ( glmdisc , riskr , Rprofet , scorecard , smbinning , woeBinning ) provide a visualization of the bins on a variable level. Figure 2 (left) shows the binning resulting from code in Example 2 which is similar for most packages. A mosaic plot of the bins, which simultaneously visualizes default rates and the size of the bins, is offered by the package glmdisc (Figure 2, right) while the names of the bins after automatic binning are not self-explanatory. Figure 2. Visualization of the bins for the variable purpose as created by the package scorecard ( left ) and mosaicplot of the binning result by the package glmdisc (right). 3.5. Applying Bins to New Data It has been emphasized that the bins as they are built on training data constitute the first part of a scorecard model. For this reason, it is necessary to store the results of the binning and to have functions to apply it to a data set. Most of the packages such as scorecard ( woebin_ply() ), smbinning ( smbinning.gen() and smbinning.factor.gen() ), woeBinning ( woe.binning.deploy() ), creditmodel ( split_bins_all() ), glmdisc ( discretize() ) and scorecardModelUtils ( num_to_cat() ) provide this functionality. Example 3 illustrates the application of binning results to a data set. Via the to = “bin” argument, either bins or WoEs can be assigned: ### Example 3: apply binning to data train_bins <- scorecard::woebin_ply(train, bins, to=‘‘bin’’) valid_bins <- scorecard::woebin_ply(valid, bins, to=‘‘bin’’) For ctree -based binning (cf. above) a workaround using the partykit::predict.party() method for bin assignment can be obtained if the tree model is stored within the results object 10 . Risks 2022,10, 67 8 of 33 More generally, binned variables can be created via the function cut() for numeric variables or by using lookup tables for factor variables (cf. Zumel and Mount 2014, p. 23) 11 . It is worth mentioning that several packages ( smbinning and riskr ) implement binning only on a single variable level but not simultaneously for several selected variables or all variables of a data frame12. 3.6. Binning of Categorical Variables For categorical variables, initially, each level can be considered as a separate bin, but levels of similar default rate and/or meaning could be grouped together. As an additional challenge, there is no natural order of the levels. For these reasons, only some of the packages offer an automatic binning of categorical variables. For example, the package smbinning does not offer an automatic merging of levels for factor variables, and its function smbinning.factor() only returns the figures similar to the table resulting from Example 2. However, each original level corresponds to only one bin. The bins can be manipulated afterwards via smbinning.factor.custom() and further be applied to new data via smbinning.factor.gen() . An automatic binning of categorical variables based on conditional inference trees is supported by the packages riskr and scorecard ( method = “tree” ). Additional merging strategies are provided by the packages glmdisc and creditmodel (as described above), scorecard ( method = “chimerge” ) and woeBinning (according to similar WoEs). Generally, merging levels with a similar default rate should only be done if the level’s frequency is large enough to result in a reliable default rate estimate on the sample. By using woeBinning ’s woe.binning() function this can be ensured: Initial bins of a minimum size ( min.perc.total ) are created and smaller factor levels are initially bundled into a positive or negative ’miscellaneous’ category according to the sign of the corresponding WoE which is desirable to prevent overfitting. The package scorecardModelUtils offers a separate function cat_new_class() for this. All levels less frequent than specified by the argument threshold are merged together, and a data frame with the resulting mapping table is stored in the output element $cat_class_new13 . The package creditmodel provides a function merge_category which keeps the m most frequent categories and merges all other levels in a new category of name “other” but no function is available to apply the same mapping to new data. Similar to woeBinning ‘s woe.binning() , the functions scorecard::woebin()14 and creditmodel::get_breaks_all()15 also merge adjacent levels of similar default rates for categorical variables. An important difference between both implementations consists in how they deal with the missing natural order of the levels and thus the notion of what ’adjacent’ means: In woe.binning() the levels are sorted according to their WoE before merging. This is not the case for the other two functions where levels are merged along their natural order which is often alphabetical 16 . This might lead to an undesired binning, and as an important conclusion an analyst should think about manually changing the level order for factor variables when working with the package scorecard17. 3.7. Weights of Evidence Most of the abovementioned packages provide WoEs of the bins within their binning summary tables. To use WoEs within the further modelling steps, it needs a functionality to assign the corresponding WoE value for each bin to the original (/or binned) variables as given by scorecard::woebin_ply() (with argument to = “woe” ), woeBinning::woe.binning.deploy() (with argument add.woe.or.dum.var = “woe”) and creditmodel::woe_trans_all(). A general way of training, storing and assigning WoEs independently of the package used for binning is given by the function woe() in the klaR package, probably the first and most comprehensive implementation of WoE computation in R. WoEs for binned variables are computed on the training data and stored in an S3 object of class woe with a corresponding predict.woe() method that allows application to new data. Furthermore, Risks 2022,10, 67 15 of 33 −1 −0.8 −0.6 −0.4 −0.2 0 0.2 0.4 0.6 0.8 1 present..1 credit.h number.o age.in.y telephon installm number.o.1 other.in purpose_ status.o savings. duration credit.a job_bin property housing_ other.de present. personal present..1 credit.h number.o age.in.y telephon installm number.o.1 other.in purpose_ status.o savings. duration credit.a job_bin property housing_ other.de present. personal status.o duration credit.h purpose credit.a savings. present. installm personal other.de present..1 property age.in.y other.in housing number.o job bin number.o.1 telephon Figure 3. Reordered correlation matrix ( left ) and phylogenetic tree of the clustered variables ( right ). The package clustVarLV (Vigneau et al. 2015,2020) offers variable clustering such that the correlation between each variable and the first latent principal component of its variable cluster is maximized. The number of clusters K has to be prespecified. As it can be seen in the output from Example 9 (only cluster 1 is shown), for each variable the correlation to the cluster’s latent component as well as the correlation to the ‘closest’ next cluster are shown. ### Example 9: variable clustering using ClustVarLV (based on Example 4) library(ClustVarLV) clverg <- CLV(train_woes[,-(1:2)], method = 1) plot(clverg) summary(clverg, K=3) Version March 10, 2022 submitted to Risks 16 of 34 ### example 9: variable clustering using ClustVarLV (based on example 4) library(ClustVarLV) clverg <- CLV(train_woes[,-(1:2)], method = 1) plot(clverg) summary(clverg, K=3) ## cor in group |cor|next group 451 ## woe_savings.account.and.bonds_bin 0.73 0.05 452 ## woe_status.of.existing.checking.account_bin 0.72 0.17 453 ## woe_purpose_bin 0.51 0.11 454 Among the aforementioned packages dedicated to credit scoring, creditR contains a function 455 variable.clustering() that performs cluster ’s pam (Maechler et al. 2021) on the transposed data 456 for variable clustering. The (sparsely documented) function correlation.cluster() %data, output, 457 “variable”, “Group”)} can be used to compute average correlations between the variables of each cluster. 458 20 459 The package Rprofet provides two functions WOEClust_hclust() and WOEClust_kmeans() that 460 do perform stats::hclust() on the transformed data or ClustOfVar::kmeansvar() and return a 461 data frame with variable names and cluster index together with the IV of the variable, which may help 462 to select variables from the clusters. Unfortunately, they are only designed to work with output from 463 the package’s function WOEProfet() and require a list of a specific structure as input argument. In 464 addition to functions cor_plot() for visualzation of the correlation matrix, char_cor() to compute 465 a matrix of Cramer’s V between or set of categorical variables and get_correlation_group() for 466 detection fo groups of correlated (numeric) variables, the package creditmodel also contains a function 467 fast_high_cor_filter() for an automatic correlation based variable selection: In a group of highly 468 correlated variables the one with the highest IV is selected as it is shown in example 10: 469 ### example 10: Automatic correlation-based variable selection (based on example 4) library(creditmodel) # create list of variables sorted according to IV iv_list =feature_selector(dat_train = train_woes, dat_test = NULL, target = "creditability", filter = "IV",iv_cp = 0.02,vars_name = FALSE) iv_list # select variables fast_high_cor_filter(dat = train_woes, com_list = iv_list, p=0.15, cor_class = TRUE ,vars_name = TRUE) Similarly, the package scorecardModelUtils offers an alternative for an automatic variable 470 preselection based on Cramer’s V using the function cv_filter() : Among two (categorical) variables 471 of V> threshold the one with lower IV is automatically removed (cf. example 11). Lastly, two 472 functions iv_filter() and vif_filter() can be used for variable preselection based on IVs only 473 (w/o taking into account for correlations between the explanatory variables) and based on variance 474 inflation (cf. also Sec. 5). 475 20 Its argument data denotes the training data, output is a data frame with two variables specifiying the variable names of the training data ( character ) and the corresponding cluster index, as given e.g. by the result from variable.clustering() . Finally, its arguments variables and clusters denote the names of these two variables in the data frame from the output argument where the clustering results are stored. Among the aforementioned packages dedicated to credit scoring, creditR contains a function variable.clustering() that performs cluster ’s pam (Maechler et al. 2021) on the transposed data for variable clustering. The (sparsely documented) function correlation.cluster() %data, output, “variable”, “Group”)} can be used to compute average correlations between the variables of each cluster. 20 The package Rprofet provides two functions WOEClust_hclust() and WOEClust_kmeans() that perform stats::hclust() on the transformed data or ClustOfVar::kmeansvar() and return a data frame with variable names and cluster index together with the IV of the variable, which may help to select variables from the clusters. Unfortunately, they are only designed to work with output from the package’s function WOEProfet() and require a list of a specific structure as input argument. In addition to functions cor_plot() for visualization of the correlation matrix, char_cor() computes a matrix of Cramer’s V between or a set of categorical variables and get_correlation_group() for detection of groups of correlated (numeric) variables. The package creditmodel also contains a function fast_high_cor_filter() for an automatic correlationbased variable selection. In a group of highly correlated variables, the one with the highest IV is selected as shown in Example 10. Risks 2022,10, 67 16 of 33 ### Example 10: Automatic correlation-based variable selection (based on Example 4) library(creditmodel) # create list of variables sorted according to IV iv_list =feature_selector(dat_train = train_woes, dat_test = NULL, target = ‘‘creditability’’, filter = ‘‘IV’’,iv_cp = 0.02,vars_name = FALSE) iv_list # select variables fast_high_cor_filter(dat = train_woes, com_list = iv_list, p=0.15, cor_class = TRUE ,vars_name = TRUE) Similarly, the package scorecardModelUtils offers an alternative for an automatic variable preselection based on Cramer’s V using the function cv_filter() . Among two (categorical) variables of V>threshold , the one with lower IV is automatically removed (cf. Example 11). Finally, two functions, iv_filter() and vif_filter() can be used for variable preselection based on IVs only (w/o taking into account for correlations between the explanatory variables) and based on variance inflation (cf. also Section 5). ### Example 11: Cramer's V based variable selection (based on Example 3) library(scorecardModelUtils) # package requires 0/1 target: train_bins2 <- train_bins train_bins2$creditability <- as.integer(train_bins2$creditability==‘‘good’’) # first data frames of IVs and Cramer's V have to be computed ivtable <- iv_table(train_bins2, ‘‘creditability’’, cat_var_name = names(train_bins2)[-1]) cvtable <- cv_table(train_bins2, names(train_bins2)[-1]) selection <- cv_filter(cvtable$cv_val_tab, ivtable$iv_table,threshold = 0.3) selection 4.5. Further Useful Functions to Support Variable Preselection The package scorecard contains a function var_filter() that performs an automatic variable selection based on IV and further allows for specifying a maximum percentage of missing or identical values within a variable, but it does not account for correlations among the predictor variables. Alternatively, the package creditmodel has a function feature_selector() for automatic variable preselection based on IV, PSI, correlation and xgboost variable importance (Chen and Guestrin 2016). The package creditR has two functions to identify variables with missing values ( na_checker() ) and compute the percentage of variables with missing values ( missing_ratio() ). For imputation of numeric variables in a data set with mean or median values, a function na_filler_contvar() is available. Of course, this has to be handled with care as the mean or median value will typically not be the same on training and validation data. The package mlr (Bischl et al. 2016,2020) offers imputation that can be applied to new data. For an assignment of explicit values to missing the package scorecardModelUtils also provides a function missing_val() . This can be either a function such as “mean” , “median” or “mode” or an explicit value such as -99999 which can be meaningful before binning to assign missing values to a separate bin. Similarly, for categorical variables the assignment of a specific level such as “missing_value” can be meaningful. A function missing_elimination() removes all variables with a percentage above missing_ratio_threshold from training (but not from validation) data. The package creditmodel offers a convenient function data_cleansing() that can be used for auto- Risks 2022,10, 67 17 of 33 matic deletion of variables with low variance and a high percentage of missing values, to remove duplicated observations and reduce the number of levels of categorical variables. The package riskr provides two functions select_categorical() and select_numeric() to select all (non-/) numeric variables of a data frame. A univariate summary of all variables is given by the function univariate() of the scorecardModelUtils package. A summary for numeric variables can be computed using the function ez_summ_num() from the package riskr . A general overview of packages explicitly designed for exploratory data analysis that provide further functionalities are given in Staniak and Biecek (2019). The packages scorecard ( one_hot() and var_scale() ) and creditmodel (\texttt{one_hot_encoding(), de_one_hot_encoding() , min_max_norm() ) provide functions for one-hot-encoding of categorical and standardization of numeric variables. 5. Multivariate Modelling 5.1. Variable Selection Traditionally, credit risk scorecards are modelled using logistic regression (cf. e.g., Anderson 2019;Siddiqi 2006;Thomas et al. 2019;Wrzosek et al. 2020), which is in R performed via glm() (with family = binomial ). In addition to the manual variable preselection as described in the former section, typically, a subsequent variable selection is performed which can be completed by the step() function. Common criteria for variable selection are AIC ( k=2 ) or BIC ( k = log(nrow(data)) ). Example 12 gives an example for BIC based variable selection. ### Example 12: BIC variable selection (based on Example 4) # column 2 (variable foreign.worker_bin excluded as it has only one level) null <- glm(creditability ~ 1,data = train_woes[,-2], family = binomial) full <- glm(creditability ~ ., data = train_woes[,-2], family = binomial) bicglm <- step(null, scope=formula(full), direction=‘‘both’’, k=log(nrow(train_woes))) Note that an initial model (here: null ) and the scope for the search have to be specified. This offers another possibility for expert knowledge integration. After each step the criteria of all candidates are reported and can be used to decide among several variable candidates of similar performance for the one that is most appropriate from a business point of view. The corresponding variable can be manually added to the formula of a new initial model in a subsequent variable selection step. The function smbinning.logitrank() of package smbinning runs all possible combinations of a specified set of variables, ranks them according to AIC and returns the corresponding model formulas in the result data frame. Depending on the size of the preselected set of variables (cf. Section 4), this can be time-consuming. As an alternative to AIC and BIC, Scallan (2011) presents how variables can be selected in line with the concept of information values (cf. Section 3) using so-called marginal information values, but currently none of the presented packages offers an implementation of this strategy. It is also common to consider the variance inflation factor of the explanatory variables of a final model given by: VIF(Xi) = 1 1−R2 i (3) where R2 i is the R2 of a linear regression model with Xi as dependent variable and all other explanatory variables except Xi as regressors. Large values of VIF(Xi) denote that this variable can be explained by the other regessors and are an indication of multicollinearity. Both the packages car (Fox and Weisberg 2019;Fox et al. 2021) and scorecard offer a function vif() that can be used for this purpose as well as the functions vif.calc() and lr_vif() of the packages creditR and creditmodel (cf. Example 13). Risks 2022,10, 67 18 of 33 ### Example 13: VIF (based on Example 11) car::vif(bicglm) scorecard::vif(bicglm) creditR::vif.calc(bicglm) creditmodel::lr_vif(bicglm) Not only variable selection during the scorecard development but also the question of segmentation may arise, i.e., whether one single model or several separate models should be used for different subsets of the population. For this purpose, the package glmtree offers a function glmtree() that computes a potential segmentation scheme according to a tree of recursive binary splits where each leaf of the tree consists in a logistic regression model. The resulting segmentation optimizes AIC, BIC or alternatively the likelihood or the Gini coefficient on validation data. Note that this optimization does not account for variable selection as described above. 5.2. Turning Logistic Regression Models into Scorecard Points >From the coefficients of the logistic regression model, the historical shape of a scorecard is obtained by assigning the corresponding effect (aka points) to each bin (such that the score of a customer is the sum over all applicable bins and can easily be calculated by hand). Typically, the effects are scaled to obtain some predefined points to double the odds (pdo, cf. e.g., Siddiqi 2006) and rounded to integers. The package scorecard offers a function scorecard() that translates a glm object into scorecard points as described above and in addition returns key figures such as frequencies, default rates and WoE for all bins. A function scorecard_ply() is available that can be used to assign scores to new data. In addition to the glm object, the bins as created by scorecard ’s woebin() (cf. Section 3) have to be passed as an input argument. Further arguments do specify the ( pdo ) as well as a fixed number of points points0 that corresponds to odds of odds0 and whether the scorecard should contain an intercept or whether the intercept should be redistributed to all variables ( basepoints_eq0 ). The function requires WoEs (not just the binned factor s) and the variable names in the coef(glm) to match the convention of variable renaming as it is done by scorecard ’s woebin_ply() function (i.e., a postfix _woe)21. Alternatively, a function scorecard2() is available that directly computes a scorecard based on bins and a data frame of the original variables. Here, in addition, the name of the target variable ( y ) and a named vector ( x ) of the desired input variables have to be passed 22 . Example 14 illustrates the usage of scorecard2() and its application to new data (here represented by the validation set) as well as its output for the variable duration.in.month . ### Example 14: calculation of scores (based on Example 2) # note: variable 20 (foreign.worker) not used (cf. also Example 12) sc <- scorecard2(bins, train, y = 'creditability',x=names(train)[1:19]) # scorecard points table for the variable 'duration.in.month' sc$duration.in.month[,c(1,2,4,5,6,7,8,13)] Version March 10, 2022 submitted to Risks 19 of 34 sum over all applicable bins and can easily calculated by hand). Typically, the effects are scaled in 541 order to obtain some predefined points to double the odds (pdo, cf. e.g. Siddiqi (2006)) and rounded to 542 integers. 543 The package scorecard offers a function scorecard() that translates a glm object into scorecard 544 points as described above and in addition returns key figures such as frequencies, default rates and 545 WoE for all bins. A function scorecard_ply() is available that can be used to assign scores to new data. 546 In addition to the glm object, the bins as created by scorecard ’s woebin() (cf. Sec. 3) have to be passed 547 as an input argument. Further arguments do specify the ( pdo ) as well as a fixed number of points 548 points0 that corresponds to odds of odds0 and whether the scorecard should contain an intercept or 549 whether the intercept should be redistributed to all variables ( basepoints_eq0 ). The function requires 550 WoEs (not just the binned factor s) and the variable names in the coef(glm) to match the convention 551 of variable renaming as it is done by scorecard ’s woebin_ply() function (i.e. a postfix _woe )21. 552 Alternatively, a function scorecard2() is available which directly computes a scorecard based on 553 bins and a data frame of the original variables. Here, in addition, the name of the target variable ( y ) 554 and a named vector ( x ) of the desired input variables have to be passed 22 . Example 14 illustrates the 555 usage of scorecard2() and its application to new data (here represented by the validation set) as well 556 as its output for the variable duration.in.month : 557 ### example 14: calculation of scores (based on example 2) # note: variable 20 (foreign.worker) not used (cf. also example 12) sc <- scorecard2(bins, train, y = ' creditability ' ,x=names(train)[1:19]) # scorecard points table for the variable ' duration.in.month ' sc$duration.in.month[,c(1,2,4,5,6,7,8,13)] ## variable bin count_distr neg pos posprob woe points 558 ## 1: duration.in.month [-Inf,8) 0.08062 51 6 0.1053 -1.2988 65 559 ## 2: duration.in.month [8,16) 0.35785 194 59 0.2332 -0.3491 18 560 ## 3: duration.in.month [16,34) 0.37907 179 89 0.3321 0.1425 -7 561 ## 4: duration.in.month [34,44) 0.10467 44 30 0.4054 0.4583 -23 562 ## 5: duration.in.month [44, Inf) 0.07779 26 29 0.5273 0.9504 -48 563 train_scored <- scorecard_ply(train, sc, only_total_score = FALSE) valid_scored <- scorecard_ply(valid, sc, only_total_score = FALSE) In addition, the package further contains a function report() which takes the data, the (original) 564 names of all variables in the final scorecard model and a breaks list (cf. Sec. 3which can be obtained 565 from the bins) as input arguments and generates an excel report summary of the scorecard model. 566 Different sheets are reported with information and figures on the data, model, scorecard points, model 567 performance, as well as the binning figures for all variables of the model which can be used for model 568 development documentation in practice. 569 In order to translate a glm based on factor variables (bins instead of WoEs) into scorecard points, 570 the package scorecardModelUtils provides a function scalling() . Its output can be used to predict 571 scores for new data by a function scoring() (cf. example 15): 572 21 A remedy how it can be used in combination with WoE assignment using the package klaR as shown in example 4 is given in snippet 9 of the supplementary code. 22 Snippet 10 of the supplementary code illustrates how the vector x of the names of the input variables in the original data frame can be extracted from the bicglm model after variable selection from example 12. Risks 2022,10, 67 19 of 33 train_scored <- scorecard_ply(train, sc, only_total_score = FALSE) valid_scored <- scorecard_ply(valid, sc, only_total_score = FALSE) In addition, the package further contains a function report() that takes the data, the (original) names of all variables in the final scorecard model and a breaks list (cf. Section 3 that can be obtained from the bins) as input arguments and generates an excel report summary of the scorecard model. Different sheets are reported with information and figures on the data, model, scorecard points, model performance and the binning figures for all variables of the model which can be used for model development documentation in practice. To translate a glm based on factor variables (bins instead of WoEs) into scorecard points, the package scorecardModelUtils provides a function scalling() . Its output can be used to predict scores for new data by function scoring() (cf. Example 15). ### Example 15: score points for a model based on bins, not WoEs (based on Example 4) library(scorecardModelUtils) # create glm using factor variables -- foreign worker excluded (cf. above) full_bins <- glm(creditability~.,data = train_bins[,-21], family = binomial) # calculate scorecard points from effects sc2 <- scalling(train_bins, ‘‘creditability’’, full_bins, point = 15, factor = 2) sc2 # apply scorecard to new data scoring(valid_bins, target = ‘‘creditability’’, sc2) The package creditmodel transforms a glm object into scorecard points via a function get_score_card() , which requires a bin table created by creditmodel::get_bins_table_all() and thus is restricted to application within its own universe. In addition, if a table of scorecard points is not required, it offers a function score_transfer() that directly applies the glm object to data and scales the resulting points accordingly (cf. Example 16) and another function p_to_score to turn posterior probabilities into score points. ### Example 16: directly predict score points from a glm object (based on Example 12) library(creditmodel) train_scored_3 <- score_transfer(bicglm, train_woes, a=500,b = 20) valid_scored_3 <- score_transfer(bicglm, valid_woes, a=500,b = 20) Another implementation of calculating scorecard points from a glm object based on bins and not WoEs is given by the function smbinning.scaling() , which comes with a predict function smbinning.scoring.gen() that can be used to score new observations but that requires the binned variables have been generated with smbinning.gen() or smbinning.factor.gen() (cf. Section 3). A function smbinning.scoring.sql() is available that transforms the resulting scorecard into SQL code. The package Rprofet also contains a function ScorecardProfet() for this purpose, which calculates a glm with corresponding scorecard points but only based on binning and WoEs as calculated by functions from the package itself (cf. Section 3), and no function is available for application of the scorecard points to new data. The function scaled.score() of the package creditR transforms posterior default probabilities into scores where any increase points double the odds (of nondefault), and odds of increase correspond to ceiling_score points. In addition, the package creditR offers a function that can be used to recalibrate an existing glm on calibration data. A simple logistic regression is fit on the calibration_data with only one input variable: the predicted log odds by the current model. Risks 2022,10, 67 20 of 33 5.3. Class Imbalance In credit scorecard modelling, the class typically is highly unbalanced in the training sample. This issue has been addressed in several papers (Brown and Mues 2012;Crone and Finlay 2012;Vincotti and Hand 2002). Usual remedies are oversampling, undersampling, synthetic minority over-sampling (SMOTE, Chawla et al. 2002) or simply reweighing observations. A comprehensive benchmark study of these techniques as well as overbagging is undertaken in Bischl et al. (2016), and it turns out that logistic regression is less sensitive to class imbalance than tree-based classifiers. Furthermore, note that different from, e.g., the accuracy of the two most commonly used performance measures in credit scorecard modelling, the Gini coefficient and the KS statistic (cf. Section 6 ) do not depend on the class imbalance ratio. The package klaR allows for specifying observation weights for WoE computation (see Section 3.7). Within the mlr3 framework, imbalance correction can be performed using mlr3pipelines (Binder et al. 2021). Several resampling algorithms are implemented in the packages imbalance (Cordón et al. 2020,2018) and unbalanced (Pozzolo et al. 2015). The SMOTE algorithm is also implemented in the smotefamily package (Siriseriwan 2019). 6. Performance Evaluation 6.1. Overview In credit scoring modelling, performance evaluation is used not only for model selection but also for third-party assessments of an existing model by auditors or regulators and to drive future management decisions about whether an existing model should be kept in place or whether it should be replaced by a new one. Note that, as opposed to common practice in machine learning, hyperparameter tuning typically has no separate validation data used for model selection (cf. e.g., Bischl et al. (2012), Bischl et al. (2021)), but in credit scorecard modelling, the validation data serves for independent model validation (corresponding to test data in frameworks such as mlr3 ). While this is less critical in the case of simple models such as logistic regression, it should still be kept in mind, especially if the model is benchmarked against more flexible machine learning models such as support vector machines, random forests or gradient boosting (cf. e.g., Hastie et al. (2009)). 6.2. Discrimination The two most popular performance metrics for credit scorecards are the Gini coefficient, Gini = 2 (AUC − 0.5 ) and the Kolmogorov–Smirnov test statistic. While for the latter, R provides the function ks.test() , one of the most popular ways to compute the AUC in R is given by the package ROCR (Sing et al. 2005,2020). Nonetheless, for the purpose of credit scorecard modelling, it is referred to the package pROC at this point for the following three reasons: 1. Different from standard binary classification problems, credit scores are typically supposed to be increasing if the event (= default-) probability decreases. The function roc() of the package pROC has an argument direction that allows for specifying this. 2. In credit scoring applications, it may be given that not all observations of a data set are of equal importance, e.g., it may not be as important to distinguish which of two customers with small default probabilities has the higher score if his or her application will be accepted anyway. The package’s function auc() has an additional argument partial.auc to compute partial area under the curve (Robin et al. 2011). 3. Finally, its function ci() can be used to compute confidence intervals for the AUC using either bootstrap or the method of DeLong (DeLong et al. 1988;Sun and Xu 2014), e.g., to support the comparison of two models. Example 17 demonstrates how pROC can be used for performance analysis. ### Example 17: Gini coefficient using {pROC} (based on Example 13) library(pROC) curve <- roc(valid$creditability, valid_scored$score, Risks 2022,10, 67 21 of 33 levels = c(‘‘good’’,‘‘bad’’), direction = ‘‘>’’) # levels = c(‘‘controls’’, ‘‘cases’’), # direction = controls > cases plot(curve) auc(curve) # gini coefficient: 2* (auc(curve) - 0.5) # confidence limits for the auc: ci(auc(curve), method = ‘‘bootstrap’’) Among thepackagesenumeratedabove, creditR offersa function Kolmogorov–Smirnov() , and riskr has two functions, ks() and ks2() , for computation of the Kolmogorov–Smirnov test statistic. In addition, riskr provides a function divergence() to compute the divergence between two empirical distributions as well as gg_dists() and gg_cum() to visualize the score densities for defaults and nondefaults and their empirical cumulative distribution functions. To compute the Gini coefficient, the package riskr provides functions aucroc (AUC), gini (Gini coefficient), gg_roc() (visualization of the ROC curve), gain() (gains table for specified values on the x-axis) and gg_gain() / gg_lift() (for visualization of the gains-/lift-chart). In the package creditmodel , two functions ks_value() and auc_value() are available as well as a model_result_plot() to visualize the ROC curve, cumulative score distributions of defaults vs. nondefaults, lift chart and the default rate over equal-sized score bins. A table with respective underlying numbers can be obtained via perf_table() . The package InformationValue contains two functions, ks_stat() and ks_plot() , for Kolmogorov-Smirnov analysis and several functions: AUROC() , plot_ROC() , Concordance() and SomersD() (Gini coefficient) to support analyses with regard to the Gini coefficient. Additionally, the confusionMatrix() and derivative performance measures misClassError() , sensitivity() , specificity() , precision() , npv() , kappaCohen() and youdensIndex() (cf. e.g., Zumel and Mount (2014) chp. 5 for an overview) can be computed for a given cut off by the corresponding functions. Note that these measures are computed with respect to the nondefault target level (supposed to be coded as ‘1’ in the target variable) as well as a cut off optimization w.r.t. the misclassification error, Youden’s Index or the minimum (/maximum) score such that no misclassified defaults (/non-defaults) occur in the data (function optimalCutoff()). Similar measures (accuracy, precision, recall, sensitivity, specificity, F1) are computed by the function fn_conf_mat() of the scorecardModelUtils package. Numeric differences between the (0/1-coded) target and the model’s predictions in terms of MSE, MAE and RMSE can be computed by its fn_error() function. The package boottol contains a function boottol() to compute bootstrap confidence intervals for Gini, AUC and KS, where subsets of the data above different cut off values are also considered. It may be desirable to analyze the (cumulative) frequencies of the binned scores. A table of such frequencies is returned by the function gini_table() in the scorecardModelUtils package. Example 18 shows selected columns for a binned score using the function gains_table() from the scorecard package. ### Example 18: score bin frequencies (...for valid_scored from Example 14) library(scorecard) gt <- gains_table(valid_scored$score, valid$creditability, bin_num = 8) gt[,c(2,4,5,6,7,8,10,11)] ## bin cum_count neg cum_neg pos cum_pos posprob approval_rate ## 1: [628, Inf) 37 37 37 0 0 0.00000 0.1263 ## 2: [575,628) 76 36 73 3 3 0.07692 0.2594 ## 3: [529,575) 112 34 107 2 5 0.05556 0.3823 Risks 2022,10, 67 22 of 33 ## 4: [492,529) 148 30 137 6 11 0.16667 0.5051 ## 5: [448,492) 185 26 163 11 22 0.29730 0.6314 ## 6: [399,448) 222 21 184 16 38 0.43243 0.7577 ## 7: [353,399) 257 14 198 21 59 0.60000 0.8771 ## 8: [-Inf,353) 293 8 206 28 87 0.77778 1.0000 Note that although the Gini coefficient is generally bounded by − 1 and 1, the value it can take for a specific model strongly depends on the discriminability of the data. For this reason, it is suitable to compare performance on different models on the same data rather than comparing performance across different data sets. Consequently, for the purpose of an out-of-time monitoring of a scorecard, it is advisable to compare an existing scorecard’s performance against a recalibrated version of it rather than to compare it with its performance on the original (development) data. Drawbacks of the Gini coefficient as a performance measure for binary classification are discussed in (Hand 2009), and the H-measure is proposed as an alternative which is implemented in the package hmeasure (Anagnostopoulos and Hand 2019). The expected maximum profit measure (Verbraken et al. 2014) as implemented in the package EMP (Bravo et al. 2019) further takes into account the profitability of a model. 6.3. Performance Summary Many of the functionalities as provided by the packages for scorecard modelling in the previous subsection already exist in other packages and are thus not indispensable. In addition to these, however, some of the package provide performance summary reports of several performance measures. These functions are listed in the following table. In Example 19, computation of a scorecard performance summary is demonstrated using the package smbinning (which returns the largest number of performance measures of the four functions from Table 5) as well the function riskr::gg_perf() that can be used to produce several graphs on the scorecard’s performance (cf. Figure 4). Note that although ROC curves are one of the most popular tools for performance visualization of binary classifiers, they are hardly suited to visualize the performance difference of several competitive models. One reason for this is that large areas of the TPR-FPR plane (e.g., everything below the main diagonal) are typically of no interest given a specific data situation. For this reason, in practice, ROC curves are not very useful for model selection. ### Example 19: scorecard performance summary (based on Example 13) library(smbinning) perf_dat <- data.frame(‘‘creditability’’ =as.integer (valid$creditability == ‘‘good’’), ‘‘score’’ =valid_scored$score) smbinning.metrics(perf_dat, ‘‘score’’,‘‘creditability’’,cutoff = 450) # roc curve, ecdf, score distribution and gain chart library(riskr) gg_perf(as.integer(valid$creditability == ‘‘good’’), valid_scored$score) Risks 2022,10, 67 23 of 33 Table 5. Overview of scorecard performance summary functions. Package Riskr Riskr Riskr Scorecard Scorecard Scorecard ScorecardModelUtils ScorecardModelUtils ScorecardModelUtils Smbinning Smbinning Smbinning Function Perf() Perf() Perf() Perf_Eva() Perf_Eva() Perf_Eva() Gini_Table() Gini_Table() Gini_Table() Smbinning.Metrics() Smbinning.Metrics() Smbinning.Metrics() KS X X X X AUC X X X Gini X X X Divergence X Bin table X Confusion matrix X X Accuracy X X Good rate X Bad rate X TPR X FNR X X TNR X FPR X X PPV X FDR X FOR X NPV X ROC curve X X X X Score densities |yX ECDF X X X Gain chart X gain roc curve cumulative distributions 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 400 600 800 400 600 800 0.000 0.001 0.002 0.003 0.004 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 Legend: non target target Figure 4. Scorecard performance graphs: ECDF ( top left ); score densities ( top right ); gains ( bottom left); ROC (bottom right). Risks 2022,10, 67 24 of 33 6.4. Rating Calibration and Concentration >From a practical point of view, it is often desirable to aggregate scorecard points into classes (rating grades) of similar risk, which is once again a binning task (cf. Section 3 ). The package creditR contains a function master.scale() that takes a data frame with scores and corresponding default probabilities as input and uses the function woeBinning::woe.binning() to group scores of similar WoE (cf. Example 20). The function odds_table() of the riskr package allows setting a breaks argument with arbitrary bins. Rating classes should be appropriately calibrated in the sense that the predicted and observed default probabilities match for all rating grades. In order to check this, the package creditR contains three functions ( chisquare.test() , binomial.test() and adjusted.binomial.test() ) that provide a table with indicators for each rating grade (cf. Example 19). Another function, binomial.point() , compares the observed average predicted default probability on the data with prespecified boundaries around some desired central tendency default probability. Bootstrap confidence intervals for default probabilities of rating grades can be computed using the function vas.test() of the package boottol . A Hosmer–Lemeshow goodness-of-fit test (Hosmer and Lemeshow 2000) is, e.g., implemented by the function hoslem.test() in the resourceselection package (Lele et al. 2019). ### Example 20: rating calibration analysis (based on Example 14) library(creditR) # calculate PDs from scores odds <- 1/19 *2ˆ(-(valid_scored$score - 600)/50) pd <- odds / (1+ odds) pd.dat <- data.frame(pd=pd, creditability = as.integer(valid$creditability == ‘‘bad’’)) # aggregate scores to rating grades mscale <- creditR::master.scale(pd.dat, ‘‘creditability’’,‘‘pd’’) # transform $Bad.Rate into numeric mscale$Bad.Rate <- as.numeric(gsub(‘‘%’’,‘‘’’,mscale$Bad.Rate))/100 # test calibration of the rating grades # chisquare.test(mscale, ‘‘PD’’, ‘‘Bad.Count’’, ‘‘Total.Observations’’) bintest <- binomial.test(mscale,‘‘Total.Observations’’,‘‘PD’’,‘‘Bad.Rate’’) bintest[,c(1,3,8,9,14)] ## Final.PD.Range Total.Distr Bad.Rate PD Test_Result ## 1 <= 0.0692759267 25.9% 0.039 0.03835 Target Value Correct ## 2 <= 0.1477666759 17.4% 0.078 0.11149 Target Value Correct ## 3 <= 0.1904265313 7.2% 0.190 0.17482 Target Value Correct ## 4 <= 0.275937974 9.6% 0.321 0.23297 Target Value Correct ## 5 <= 0.3709582356 7.5% 0.364 0.32236 Target Value Correct ## 6 <= 0.4365863463 5.1% 0.467 0.41036 Target Value Correct ## 7 <= 0.4605851205 3.1% 0.333 0.45143 Target Value Correct ## 8 <= 0.5695614029 9.2% 0.630 0.51257 Target Value Correct ## 9 <= Inf 15.0% 0.727 0.72768 Target Value Correct According to regulation, ratings must avoid risk concentration (i.e., a majority of the observations being assigned to only a few grades). The Herfindahl–Hirschman index ( HHI =∑jˆ f(j)2 , with the empirical distribution ˆ f of the rating grades j ) can be considered to verify this, as e.g., implemented by creditR ’s Herfindahl.Hirschman.Index() or Adjusted.Herfindahl.Hirschman.Index() . Small values of HHI indicate low risk concentration. 6.5. Cross Validation Some of the mentioned packages also provide functions for cross-validation. As both binning and variable selection are interactive, they are not suited for cross-validation (cf. Risks 2022,10, 67 31 of 33 Cordón, Ignacio, Salvador García, Alberto Fernández, and Francisco Herrera. 2020. imbalance: Preprocessing Algorithms for Imbalanced Datasets. R Package Version 1.0.2.1. Available online: https://CRAN.R-project.org/package=imbalance (accessed on 15 February 2022). Crone, Sven F., and Steven Finlay. 2012. Instance sampling in credit scoring: An empirical study of sample size and balancing. International Journal of Forecasting 28: 224–38. [CrossRef] Crook, Jonathan, and John Banasik. 2004. Does reject inference really improve the performance of application scoring models? Journal of Banking & Finance 28: 857–74. [CrossRef] Csárdi, Gábor. 2019. cranlogs: Download Logs from the ’RStudio’ ’CRAN’ Mirror. R Package Version 2.1.1. Available online: https://CRAN.R-project.org/package=cranlogs (accessed on 15 February 2022). DeLong, Elizabeth R., David M. DeLong, and Daniel L. Clarke-Pearson. 1988. Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics 44: 837–45. [CrossRef] Dis, Ayhan. 2020. creditR: A Credit Risk Scoring and Validation Package. R Package Version 0.1.0. Available online: https://github.com/ ayhandis/creditR (accessed on 15 February 2022). Dua, Dheeru, and Casey Graff. 2019. UCI Machine Learning Repository. Irvine: University of California. Ehrhardt, Adrien. 2018. scoringTools: Credit Scoring Tools. R Package Version 0.1.2 Available online: https://CRAN.R-project.org/ package=scoringTools (accessed on 15 February 2022). Ehrhardt, Adrien. 2020. glmtree: Logistic Regression Trees. R Package Version 0.2. Available online: https://CRAN.R-project.org/ package=glmtree (accessed on 15 February 2022). Ehrhardt, Adrien, Christophe Biernacki, Vincent Vandewalle, and Philippe Heinrich. 2019. Feature quantization for parsimonious and interpretable predictive models. arXiv arXiv:1903.08920. Ehrhardt, Adrien, Christophe Biernacki, Vincent Vandewalle, Philippe Heinrich, and Sébastien Beben. 2019. Réintégration des refusés en credit scoring. arXiv arXiv:1903.10855. Ehrhardt, Adrien, and Vincent Vandewalle. 2020. glmdisc: Discretization and Grouping for Logistic Regression. R Package Version 0.6. Available online: https://CRAN.R-project.org/package=glmdisc (accessed on 15 February 2022). Eichenberg, Thilo. 2018. woeBinning: Supervised Weight of Evidence Binning of Numeric Variables and Factors. R Package Version 0.1.6. Available online: https://CRAN.R-project.org/package=woeBinning (accessed on 15 February 2022). Fan, Dongping. 2022. creditmodel Toolkit for Credit Modeling, Analysis and Visualization. R Package Version 1.3.1. Available online: https://CRAN.R-project.org/package=creditmodel (accessed on 15 February 2022). Financial Stability Board. 2017. Artificial Intelligence and Machine Learning in Financial Services—Market Developments and Financial Stability Implications. Available online: https://www.fsb.org/2017/11/artificial-intelligence-and-machine-learning-in-financialservice/ (accessed on 15 February 2022). Finlay, Steven. 2012. Credit Scoring, Response Modelling and Insurance Rating. London: Palgarve MacMillan. Fox, John, and Sanford Weisberg. 2019. An R Companion to Applied Regression, 3rd ed. Thousand Oaks: Sage. Fox, John, Sanford Weisberg, Brad Price, Daniel Adler, Douglas Bates, Gabriel Baud-Bovy, Ben Bolker, Steve Ellison, David Firth, Michael Friendly, and et al. 2021. car: Companion to Applied Regression. R Package Version 3.0-12. Available online: https: //CRAN.R-project.org/package=car (accessed on 15 February 2022). Goodman, Bryce, and Seth Flaxman. 2017. European union regulations on algorithmic decision-making and a “right to explanation”. AI Magazine 38: 50–57. [CrossRef] Greenwell, Brandon, Bradley Boehmke, Jay Cunningham, and GBM Developers. 2020. gbm: Generalized Boosted Regression Models. R Package Version 2.1.8. Available online: https://CRAN.R-project.org/package=gbm (accessed on 15 February 2022). Groemping, Ulrike. 2019. South German Credit Data: Correcting a Widely Used Data Set. Technical Report 4/2019. Berlin: Department II, Beuth University of Applied Sciences Berlin. Hand, David. 2009. Measuring classifier performance: A coherent alternative to the area under the roc curve. Machine Learning 77: 103–23. [CrossRef] Hand, David, and William Henley. 1993. Can reject inference ever work? IMA Journal of Management Matehmatics 5: 45–55. [CrossRef] Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. 2009. The Elements of Statistical Learning, 2nd ed. New York: Springer. Hoffmann, Hans. 1994. German credit data set (statlog). Available online: https://archive.ics.uci.edu/ml/datasets/statlog+(german+ credit+data) (accessed on 15 February 2022). Hosmer, David W., and Stanley Lemeshow. 2000. Applied Logistic Regression. Hoboken: Wiley. Hothorn, Thorsten, Kurt Hornik, and Achim Zeileis. 2006. Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical Statistics 15: 651–74. [CrossRef] Hothorn, Torsten, and Achim Zeileis. 2015. partykit: A modular toolkit for recursive partytioning in R. Journal of Machine Learning Research 16: 3905–9. Izrailev, Sergei. 2015. Binr: Cut Numeric Values into Evenly Distributed Groups. R Package Version 1.1. Available online: https: //CRAN.R-project.org/package=binr (accessed on 15 February 2022). Jopia, Herman. 2019. smbinning: Optimal Binning for Scoring Modeling. R Package Version 0.9. Available online: https://CRAN.Rproject.org/package=smbinning (accessed on 15 February 2022). Kaszynski, Daniel. 2020. Background of credit scoring. In Credit Scoring in Context of Interpretable Machine Learning. Edited by Daniel Kaszynski, Bogumil Kaminski and Tomasz Szapiro. Warsaw: SGH, pp. 17–26. Risks 2022,10, 67 32 of 33 Kaszynski, Daniel, Bogumil Kaminski, and Tomasz Szapiro. 2020. Credit Scoring in Context of Interpretable Machine Learning. Warsaw: SGH. Kim, HyunJi. 2012. Discretization: Data Preprocessing, Discretization for Classification. R Package Version 1.0-1. Available online: https://CRAN.R-project.org/package=discretization (accessed on 15 February 2022). Kuhn, Max. 2008. Building predictive models in r using the caret package. Journal of Statistical Software 28: 1–26. [CrossRef] Kuhn, Max. 2021. Caret: Classification and Regression Training. R Package Version 6.0-90. Available online: https://CRAN.R-project. org/package=caret (accessed on 15 February 2022). Kunst, Joshua. 2020. riskr: Functions to Facilitate the Evaluation, Monitoring and Modeling process. R Package Version 1.0. Available online: https://github.com/jbkunst/riskr (accessed on 15 February 2022). Landgraf, Andrew J. 2016. logisticPCA: Binary Dimensionality Reduction. R Package Version 0.2. Available online: https://CRAN.Rproject.org/package=logisticPCA (accessed on 15 February 2022). Landgraf, Andrew J., and Yoonkyung Lee. 2015. Dimensionality reduction for binary data through the projection of natural parameters. arXiv arXiv:1510.06112. Lang, Michel, Martin Binder, Jakob Richter, Patrick Schratz, Florian Pfisterer, Stefan Coors, Quay Au, Giuseppe Casalicchio, Lars Kotthoff, and Bernd Bischl. 2019. mlr3: A modern object-oriented machine learning framework in R. Journal of Open Source Software 4: 1903. [CrossRef] Lang, Michel, Bernd Bischl, Jakob Richter, Patrick Schratz, Giuseppe Casalicchio, Stefan Coors, Quay Au, and Martin Binder. 2021. mlr3: Machine Learning in R—Next Generation. R Package Version 0.13.0. Available online: https://CRAN.R-project.org/package=mlr3 (accessed on 15 February 2022). Larsen, Kim. 2016. Information: Data Exploration with Information Theory. R Package Version 0.0.9. Available online: https: //CRAN.R-project.org/package=Information (accessed on 15 February 2022). Lele, Subhash R., Jonah L. Keim, and Peter Solymos. 2019. ResourceSelection: Resource Selection (Probability) Functions for Use-Availability Data. R Package Version 0.3-5. Available online: https://CRAN.R-project.org/package=ResourceSelection (accessed on 15 February 2022). Lessmann, Stefan, Bart Baesens, Hsin-Vonn Seow, and Lyn Thomas. 2015. Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research 247: 124–36. [CrossRef] Liaw, Andy, and Matthew Wiener. 2002. Classification and regression by randomforest. R News 2: 18–22. Ligges, Uwe. 2009. Programmieren Mit R, 3rd ed. Heidelberg: Springer. Little, Roderick, and Donald Rubin. 2002. Statistical Analysis with Missing Data. Hoboken: Wiley. Louzada, Francisco, Anderson Ara, and Guilherme Fernandes. 2016. Classification methods applied to credit scoring: A systematic review and overall comparison. Surveys in OR and Management Science 21: 117–34. [CrossRef] Maechler, Martin, Peter Rousseeuw, Anja Struyf, Mia Hubert, and Kurt Hornik. 2021. Cluster: Cluster Analysis Basics and Extensions. R Package Version 2.1.2. Available online: https://CRAN.R-project.org/package=cluster (accessed on 15 February 2022). Molnar, Christoph, Bernd Bischl, and Giuseppe Casalicchio. 2018. iml: An R package for Interpretable Machine Learning. JOSS 3: 786. [CrossRef] Paradis, Emmanuel, Simon Blomberg, Ben Bolker, Joseph Brown, Julien Claude, Hoa Sien Cuong, Richard Desper, Gilles Didier, Benoit Durand, Julien Dutheil, and et al. 2021. ape: Analyses of Phylogenetics and Evolution. R Package Version 5.6.1. Available online: https://CRAN.R-project.org/package=ape (accessed on 15 February 2022). Paradis, Emmanuel and Klaus Schliep. 2018. ape 5.0: An environment for modern phylogenetics and evolutionary analyses in R. Bioinformatics 35: 526–28. [CrossRef] Poddar, Arya. 2019. scorecardModelUtils: Credit Scorecard Modelling Utils. R Package Version 0.0.1.0. Available online: https: //CRAN.R-project.org/package=scorecardModelUtils (accessed on 15 February 2022). Pozzolo, Andrea Dal, Olivier Caelen, and Gianluca Bontempi. 2015. unbalanced: Racing for Unbalanced Methods Selection. R Package Version 2.0. Available online: https://CRAN.R-project.org/package=unbalanced (accessed on 15 February 2022). Prabhakaran, Selva. 2016. InformationValue: Performance Analysis and Companion Functions for Binary Classification Models. R Package Version 1.2.3. Available online: https://CRAN.R-project.org/package=InformationValue (accessed on 15 February 2022). Robin, Xavier, Natacha Turck, Alexandre Hainard, Natalie Tiberti, Frederique Lisacek, Jean Sanchez, and Markus Müller. 2011. pRoc: An open-source package for r and s+ to analyze and compare roc curves. BMC Bioinformatics 12: 1–8. [CrossRef] Robin, Xavier, Natacha Turck, Alexandre Hainard, Natalia Tiberti, Frederique Lisacek, Jean-Charles Sanchez, Markus Müller, Stefan Siegert, and Matthias Doering. 2021. pROC: Display and Analyze ROC Curves. R Package Version 1.18.0. Available online: https://CRAN.R-project.org/package=pROC (accessed on 15 February 2022). Roever, Christian, Nils Raabe, Karsten Luebke, Uwe Ligges, Gero Szepannek, and Marc Zentgraf. 2020. klaR: Classification and visualization. R Package Version 0.6-15. Available online: https://CRAN.R-project.org/package=klaR (accessed on 15 February 2022). Rudin, Cynthia. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 1: 206–15. [CrossRef] Ryu, Choonghyun. 2021. dlookr: Tools for Data Diagnosis, Exploration, Transformation. R Package Version 0.5.4. Available online: https://CRAN.R-project.org/package=dlookr (accessed on 15 February 2022). Risks 2022,10, 67 33 of 33 Scallan, Gerard. 2011. Class(ic) scorecards—Selecting attributes in logistic regression. In Credit Scoring and Credit Control XIII. Available online: https://www.scoreplus.com/papers/paper (accessed on 15 February 2022). Schiltgen, Garrett. 2015. boottol: Bootstrap Tolerance Levels for Credit Scoring Validation Statistics. R Package Version 2.0. Available online: https://CRAN.R-project.org/package=boottol (accessed on 15 February 2022). Sharma, Dhuv. 2009. Guide to Credit Scoring in R. CRAN Documentation Contribution. Available online: https://cran.r-project.org/ (accessed on 15 February 2022). Siddiqi, Naeem. 2006. Credit Risk Scorecards: Developing and Implementing Intelligent Credit Scoring, 2nd ed. Hoboken: Wiley. Sing, Tobias, Oliver Sander, Niko Beerenwinkel, and Thomas Lengauer. 2005. ROCR: Visualizing classifier performance in R. Bioinformatics 21: 7881. [CrossRef] Sing, Tobias, Oliver Sander, Niko Beerenwinkel, and Thomas Lengauer. 2020. ROCR: Visualizing the Performance of Scoring Classifiers. R Package Version 1.0-11. Available online: https://CRAN.R-project.org/package=ROCR (accessed on 15 February 2022). Siriseriwan, Wacharasak. 2019. smotefamily: A Collection of Oversampling Techniques for Class Imbalance Problem Based on SMOTE. R Package Version 1.3.1. Available online: https://CRAN.R-project.org/package=smotefamily (accessed on 15 February 2022). Staniak, Mateusz, and Przemysław Biecek. 2019. The Landscape of R Packages for Automated Exploratory Data Analysis. The R Journal 11: 347–69. [CrossRef] Stratman, Eric, Riaz Khan, and Allison Lempola. 2020. Rprofet: WOE Transformation and Scorecard Builder. R Package Version 2.2.1. Available online: https://CRAN.R-project.org/package=Rprofet (accessed on 15 February 2022). Sun, Xu, and Weichao Xu. 2014. Fast implementation of delong’s algorithm for comparing the areas under correlated receiver operating characteristic curves. IEEE Signal Processing Letters 21: 1389–93. [CrossRef] Szepannek, Gero. 2017. On the practical relevance of modern machine learning algorithms for credit scoring applications. WIAS Report Series 29: 88–96. Szepannek, Gero. 2019. How much can we see? A note on quantifying explainability of machine learning models. arXiv arXiv:1910.13376. Szepannek, Gero, and Karsten Lübke. 2021. Facing the challenges of developing fair risk scoring models. Frontiers in Artificial Intelligence 4: 117. [CrossRef] [PubMed] Therneau, Terry, and Beth Atkinson. 2019. rpart: Recursive Partitioning and Regression Trees. R Package Version 4.1-15. Available online: https://CRAN.R-project.org/package=rpart (accessed on 15 February 2022). Thomas, Lynn C., Jonathan N. Crook, and David B. Edelman. 2019. Credit Scoring and its Applications, 2nd ed. Philadelphia: SIAM. Thoppay, Sudarson. 2015. woe: Computes Weight of Evidence and Information Values. R Package Version 0.2. Available online: https://CRAN.R-project.org/package=woe (accessed on 15 February 2022). Verbraken, Thomas, Christian Bravo, Weber Richard, and Baesens Baesens. 2014. Development and application of consumer credit scoring models using profit-based classification measures. European Journal of Operational Research 238: 505–13. [CrossRef] Verstraeten, Geert, and Dirk Van den Poel. 2005. The impact of sample bias on consumer credit scoring performance and profitability. JORS 56: 981–92. [CrossRef] Vigneau, Evelyne, Mingkun Chen, and Veronique Cariou. 2020. ClustVarLV: Clustering of Variables Around Latent Variables. R Package Version 2.0.1. Available online: https://CRAN.R-project.org/package=ClustVarLV (accessed on 15 February 2022). Vigneau, Evelyne, Mingkun Chen, and El Mostafa Qannari. 2015. ClustVarLV: An R Package for the Clustering of Variables Around Latent Variables. The R Journal 7: 134–48. [CrossRef] Vincotti, Veronica, and David Hand. 2002. Scorecard construction with unbalanced class sizes. Journal of The Iranian Statistical Society 2: 189–205. Wrzosek, Malgorzata, Daniel Kaszynski, Karol Przanowski, and Sebastian Zajac. 2020. Selected machine learning methods used for credit scoring. In Credit Scoring in Context of Interpretable Machine Learning. Edited by Daniel Kaszynski, Bogumil Kaminski and Tomasz Szapiro. Warsaw: SGH, pp. 83–146. Xie, Shichen. 2021. scorecard: Credit Risk Scorecard. R Package Version 0.3.6. Available online: https://CRAN.R-project.org/package= scorecard (accessed on 15 February 2022). Zumel, Nina, and John Mount. 2014. Practical Data Science with R. New York: Manning.