scieee AI-readable full text Open interactive document viewer

Modeling credit risk: A category theory perspective

Trân-cao-Son,Nicolau, Dan,Nayak, Richi,Verhoeven, Peter

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Trân-cao-Son; Nicolau, Dan; Nayak, Richi; Verhoeven, Peter Article Modeling credit risk: A category theory perspective Journal of Risk and Financial Management Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Trân-cao-Son; Nicolau, Dan; Nayak, Richi; Verhoeven, Peter (2021) : Modeling credit risk: A category theory perspective, Journal of Risk and Financial Management, ISSN 1911-8074, MDPI, Basel, Vol. 14, Iss. 7, pp. 1-21, https://doi.org/10.3390/jrfm14070298 This Version is available at: https://hdl.handle.net/10419/258402 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Journal of Risk and Financial Management Article Modeling Credit Risk: A Category Theory Perspective Cao Son Tran 1, Dan Nicolau 1,2, Richi Nayak 3and Peter Verhoeven 4,*   Citation: Tran, Cao Son, Dan Nicolau, Richi Nayak, and Peter Verhoeven. 2021. Modeling Credit Risk: A Category Theory Perspective. Journal of Risk and Financial Management 14: 298. https:// doi.org/10.3390/jrfm14070298 Academic Editor: Adrian Cantemir Calin Received: 29 March 2021 Accepted: 27 June 2021 Published: 1 July 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1Science and Engineering Faculty, Queensland University of Technology, Brisbane, QLD 4000, Australia; [email protected] (C.S.T.); [email protected] (D.N.) 2Green Templeton College, University of Oxford, Oxford OX2 6HG, UK 3Centre for Data Science, Queensland University of Technology, Brisbane, QLD 4000, Australia; [email protected] 4Faculty of Business and Law, Queensland University of Technology, Brisbane, QLD 4000, Australia *Correspondence: peter[email protected] Abstract: This paper proposes a conceptual modeling framework based on category theory that serves as a tool to study common structures underlying diverse approaches to modeling credit default that at first sight may appear to have nothing in common. The framework forms the basis for an entropybased stacking model to address issues of inconsistency and bias in classification performance. Based on the Lending Club’s peer-to-peer loans dataset and Taiwanese credit card clients dataset, relative to individual base models, the proposed entropy-based stacking model provides more consistent performance across multiple data environments and less biased performance in terms of default classification. The process itself is agnostic to the base models selected and its performance superior, regardless of the models selected. Keywords: credit default; category theory; enriched structures; entropy; stacking 1. Introduction Credit risk assessment is a critical component of a lender’s loan approval, monitoring and pricing process. It is achieved through the application of statistical models that provide estimates of the probability of default (PD) of the borrower, usually over a one-year period. Default risk is typically treated as a dichotomous classification problem, distinguishing potential defaulters (payers) from non-defaulters (non-payers) with information about default status contained within a set of features of the parties involved in the transaction. Altman (1968) provided the first formal approach towards corporate default modeling, reconciling accounting-based ratios often used by practitioners with rigorous statistical techniques championed by researchers. He applies a statistical technique called Multivariate Discriminant Analysis (MDA) to construct discriminant functions (axes) from linear combinations of the selected covariates. A major drawback of MDA is the large number of unrealistic assumptions imposed, which frequently results in biased significance tests and error rates (Joy and Tollefson 1978;Mcleay and Omar 2000). This has led many researchers to propose logistic models as the next best alternative, requiring fewer restrictive assumptions and allowing for more general usage without loss in performance (Altman and Sabato 2007;Lawrence et al. 1992;Martin 1977;Ohlson 1980). Whilst there have been several attempts to put the field of credit risk modeling on a more concrete theoretical foundation (Asquith et al. 1989;Jonkhart 1979;Santomero and Vinso 1977;Vassalou and Xing 2004), supported by advances in computing power, the literature has more recently moved to techniques employed in the field of machine learning (ML). Essentially, it consists of statistical models that require less restrictive assumptions regarding the data, providing more flexibility in model construction and usage. It has made this approach the fastest growing research area in credit risk modeling. Among supervised machine learning methods, Artificial Neural Network (ANN) has received J. Risk Financial Manag. 2021,14, 298. https://doi.org/10.3390/jrfm14070298 https://www.mdpi.com/journal/jrfm J. Risk Financial Manag. 2021,14, 298 2 of 21 most attention, offering improved prediction accuracy, adaptive capability and robustness (Dastile et al. 2020;Tam 1991). Since its inception, the number of studies using ML techniques has increased nearly exponentially, focusing primarily on benchmarking state of the art individual classifiers. Lessmann et al. (2015) is the first study to benchmark a wide range of supervised ML classifiers not investigated previously. It has become a key reference for other researchers on model comparison. Also notable is the study of Teply and Polena (2020), applying ML to peer-to-peer loan dataset provided by the Lending Club. Of particular interest has been the construction of ensembles of credit risk models (Abellán and Mantas 2014;Ala’raj and Abbod 2016b;Finlay 2011;Hsieh and Hung 2010) and meta-classifiers trained on combined outputs of groups of base models (Doumpos and Zopounidis 2007;Lessmann et al. 2015; Wang et al. 2018;Wolpert 1992;Xia et al. 2018). Despite the increasing sophistication in how individual base models are put together in an ensemble (stacking) or how the various outputs are combined to achieve final prediction, all face a critical issue. Essentially, there is a lack of a sound conceptual framework to guide the ensemble or stacking process. Each study specifies their own method of selecting base models for combination and generating combined outputs. As a result, the recommendations made have been highly sensitive to the data environments examined, making it difficult to perform sound comparative performance analysis. This explains why each study tends to conclude that their combination method is the best performer among competing models. Motivated by a lack of consistency in model selection, this paper outlines a conceptual framework concerned with the design of structures in credit risk modeling within a classification context. Based on the framework, various computational approaches are proposed that solves the above noted problem of inconsistency in results. First, category theory is introduced to help design common structures underlying seemingly unrelated credit risk models. These structures reveal deep connection between seemingly unrelated models, thus providing a powerful tool to study their relationships without being distracted by details of their implementation. Second, a stacking model is constructed to address issues of inconsistent and biased performance in model benchmarking. Typically, a model’s predictive value exhibits inconsistent performance when there are changes in data scope within an environment or changes in the environment itself, with the underlying model essentially remaining unchanged. Complicating this issue is a tendency for models to be biased in their prediction due to the subjective selection of performance criteria. It is not unusual to observe a model delivering an impressive overall performance, while failing to detect any credit default at all. In order to address this issue, two new structures—Shannon’s information entropy and enriched categories—are introduced. The focus of attention is on demonstrating the benefit of having a sound conceptual framework to enable optimal construction of models that minimise performance inconsistency and bias. The proposed modeling framework is applied to the Lending Club’s peer-to-peer loans dataset from 2007–2013 as well as to Taiwanese credit card clients dataset for 2005. The empirical results show that the proposed entropy-based stacking approach results in more consistent performance across multiple data environments as well as less biased performance. The process itself is agnostic as to which base model is selected. The conceptual framework developed provides an explanation as to why various ensemble and stacking models proposed in the literature arrive at different conclusions regarding classification performance—they are caught in an equivalence trap. Ensemble models, despite their seemingly sophisticated assembling process, fuse the outputs of base models either by majority voting or by some type of linear weighted combination. In doing so, no new instance of data structure is created; all that has been achieved is an extension of the operation to cover the output combination process. As a result, the categorical structure of the modeling approach is the same as that of any other credit risk model with equivalent performance. J. Risk Financial Manag. 2021,14, 298 3 of 21 This paper is organized as follows. Section 2describes in detail the modeling framework proposed, including key elements of a category and how the representation of current approaches to modeling credit risk can be built within the context of frames. Section 3 presents the data, whilst Section 4presents the empirical results. A discussion of the empirical findings is presented in Section 5. Finally, Section 6concludes the paper. 2. Modeling Framework 2.1. Categorial Equivalence Whilst at a first glance, the many statistical approaches to credit risk modeling may seem radically different from one another, with each model constructing its own relation between the various covariates, common features exist which can be integrated into a conceptual framework that captures the essence of the credit modeling process. This framework can be built on the concept of category theory, which is the abstract study of process first proposed by Eilenberg and MacLane (1945). Category theory concerns itself with how different modeling approaches relate to one another and the manner in which they relate to one another is related to the functions between them. Instead of focusing on a particular credit risk modeling approach Aand asking what its elements are, category theory asks what all the morphisms from Ato other modeling approaches. Arguably, this mindset could be extremely useful as it suppresses unimportant details, allowing the modeler to focus on the important structural components of credit risk assessment. The structure of the credit risk modeling process underlying current approaches is represented in Figure 1(for the key definitions in category theory see Appendix A). J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 3 of 21 This paper is organized as follows. Section 2 describes in detail the modeling framework proposed, including key elements of a category and how the representation of current approaches to modeling credit risk can be built within the context of frames. Section 3 presents the data, whilst Section 4 presents the empirical results. A discussion of the empirical findings is presented in Section 5. Finally, Section 6 concludes the paper. 2. Modeling Framework 2.1. Categorial Equivalence Whilst at a first glance, the many statistical approaches to credit risk modeling may seem radically different from one another, with each model constructing its own relation between the various covariates, common features exist which can be integrated into a conceptual framework that captures the essence of the credit modeling process. This framework can be built on the concept of category theory, which is the abstract study of process first proposed by Eilenberg and Maclane (1945). Category theory concerns itself with how different modeling approaches relate to one another and the manner in which they relate to one another is related to the functions between them. Instead of focusing on a particular credit risk modeling approach A and asking what its elements are, category theory asks what all the morphisms from A to other modeling approaches. Arguably, this mindset could be extremely useful as it suppresses unimportant details, allowing the modeler to focus on the important structural components of credit risk assessment. The structure of the credit risk modeling process underlying current approaches is represented in Figure 1 (for the key definitions in category theory see Appendix A). Figure 1. Structure of the credit risk modeling process underlying current approaches. Object D represents a data structure that forms the basis of which specific data are collected, processed, analyzed and used in both the testing and training process. Object M represents model choice with the morphism m between D and M defined by a computational process that optimally maps the specific training dataset to a unique model (Aster et al. 2018). Object C represents modeling outcomes with the morphism c between M and C defined by a two-stage process: (i) the testing dataset is applied to the model to obtain predictions of default; and (ii) these predictions are compared to the actual outcomes observed in the data and the results are mapped into a compressed structure such as a confusion matrix or vectors of PDs from which various performance metrics are constructed (Dastile et al. 2020). Object P represents performance criteria, i.e., agreement between prediction and observation. This measurement process defines the morphism p in the structure above. The morphisms m, c and p are well-defined computational processes in the sense that they are finite and generate unique results. Consequently, m, c and p are injective morphisms. At this stage, four more morphisms, denoted with d D , id M , id C and id P , are introduced into the structure, as shown in Figure 2 below. They essentially send each object to itself, thus representing the objects’ identity morphisms. For example, the replacement operator which replaces one instance of an object with another instance can be used as an identity morphism. The resulting category 𝑅, represents the process underlying current approaches to credit risk modeling. Figure 2. Category 𝑅 of the modeling process underlying current approaches. Figure 1. Structure of the credit risk modeling process underlying current approaches. Object Drepresents a data structure that forms the basis of which specific data are collected, processed, analyzed and used in both the testing and training process. Object Mrepresents model choice with the morphism mbetween Dand Mdefined by a computational process that optimally maps the specific training dataset to a unique model (Aster et al. 2018). Object Crepresents modeling outcomes with the morphism cbetween Mand Cdefined by a two-stage process: (i) the testing dataset is applied to the model to obtain predictions of default; and (ii) these predictions are compared to the actual outcomes observed in the data and the results are mapped into a compressed structure such as a confusion matrix or vectors of PDs from which various performance metrics are constructed (Dastile et al. 2020). Object Prepresents performance criteria, i.e., agreement between prediction and observation. This measurement process defines the morphism pin the structure above. The morphisms m,cand pare well-defined computational processes in the sense that they are finite and generate unique results. Consequently, m,cand pare injective morphisms. At this stage, four more morphisms, denoted with d D ,id M ,id C and id P , are introduced into the structure, as shown in Figure 2below. They essentially send each object to itself, thus representing the objects’ identity morphisms. For example, the replacement operator which replaces one instance of an object with another instance can be used as an identity morphism. The resulting category R , represents the process underlying current approaches to credit risk modeling. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 3 of 21 This paper is organized as follows. Section 2 describes in detail the modeling framework proposed, including key elements of a category and how the representation of current approaches to modeling credit risk can be built within the context of frames. Section 3 presents the data, whilst Section 4 presents the empirical results. A discussion of the empirical findings is presented in Section 5. Finally, Section 6 concludes the paper. 2. Modeling Framework 2.1. Categorial Equivalence Whilst at a first glance, the many statistical approaches to credit risk modeling may seem radically different from one another, with each model constructing its own relation between the various covariates, common features exist which can be integrated into a conceptual framework that captures the essence of the credit modeling process. This framework can be built on the concept of category theory, which is the abstract study of process first proposed by Eilenberg and Maclane (1945). Category theory concerns itself with how different modeling approaches relate to one another and the manner in which they relate to one another is related to the functions between them. Instead of focusing on a particular credit risk modeling approach A and asking what its elements are, category theory asks what all the morphisms from A to other modeling approaches. Arguably, this mindset could be extremely useful as it suppresses unimportant details, allowing the modeler to focus on the important structural components of credit risk assessment. The structure of the credit risk modeling process underlying current approaches is represented in Figure 1 (for the key definitions in category theory see Appendix A). Figure 1. Structure of the credit risk modeling process underlying current approaches. Object D represents a data structure that forms the basis of which specific data are collected, processed, analyzed and used in both the testing and training process. Object M represents model choice with the morphism m between D and M defined by a computational process that optimally maps the specific training dataset to a unique model (Aster et al. 2018). Object C represents modeling outcomes with the morphism c between M and C defined by a two-stage process: (i) the testing dataset is applied to the model to obtain predictions of default; and (ii) these predictions are compared to the actual outcomes observed in the data and the results are mapped into a compressed structure such as a confusion matrix or vectors of PDs from which various performance metrics are constructed (Dastile et al. 2020). Object P represents performance criteria, i.e., agreement between prediction and observation. This measurement process defines the morphism p in the structure above. The morphisms m, c and p are well-defined computational processes in the sense that they are finite and generate unique results. Consequently, m, c and p are injective morphisms. At this stage, four more morphisms, denoted with d D , id M , id C and id P , are introduced into the structure, as shown in Figure 2 below. They essentially send each object to itself, thus representing the objects’ identity morphisms. For example, the replacement operator which replaces one instance of an object with another instance can be used as an identity morphism. The resulting category 𝑅, represents the process underlying current approaches to credit risk modeling. Figure 2. Category 𝑅 of the modeling process underlying current approaches. Figure 2. Category Rof the modeling process underlying current approaches. J. Risk Financial Manag. 2021,14, 298 4 of 21 From this structure, a specific approach to credit risk modeling is just a C-Instance of the category R (R-instance I1 ), represented by four elements and seven morphisms as shown in Figure 3. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 4 of 21 From this structure, a specific approach to credit risk modeling is just a C-Instance of the category 𝑅 (R-instance 𝐼), represented by four elements and seven morphisms as shown in Figure 3. Figure 3. A specific credit risk model is a C-Instance of category 𝑅. 𝐷𝟏 consists of the training and testing dataset sharing the same structure 𝑓  →𝑐    , with d being the number of sample points and 𝑓  representing n features associated with one of m credit classes 𝑐   . 𝑀𝟏 is a symbolic expression of the model’s structure in the form 𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢. Here 𝛼  is one of the k parameters obtained during the training process represented by the morphism 𝑚𝟏, 𝑆  specifies the symbolic expression of 𝛼  and 𝑆 describes how 𝑆  𝛼   are structured together. 𝐶𝟏 consists of the modeling results with structure (𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ) where 𝑃 is a set of PDs obtained during the testing process, 𝑇𝑃,𝐹𝑃,𝐹𝑁,𝑇𝑁 are elements of the confusion matrix (see Table A1) that are generated by the testing process 𝑐𝟏, where 𝑇𝑁 is a true positive, FP is a false positive, FN is a false negative and TN is a true negative. 𝑃𝟏 consists of performance metrics (𝑚  ,…,𝑚  ) generated by the morphism 𝑝𝟏, which is a specific implementation of 𝑝. Since the morphisms always generate unique results, they serve as the functional mapping between the set 𝐷𝟏,𝑀𝟏, 𝐶𝟏 and 𝑃𝟏. Figure 4 summarises the setvalued functor 𝐼, performing the mapping process of the first R-instance. Figure 4. The mapping process of the first R-instance. Now suppose there is a second approach to credit risk modeling that can be represented as another R-instance 𝐼, represented by four elements and seven morphisms (Figure 5). Figure 5. A specific credit risk model is a C-Instance of category 𝑅. Figure 3. A specific credit risk model is a C-Instance of category R. D1 consists of the training and testing dataset sharing the same structure nfI1 n→cI1 mod 1 , with dbeing the number of sample points and fI1 n representing nfeatures associated with one of mcredit classes cI1 m . M1 is a symbolic expression of the model’s structure in the form SI1SI1 1αI1 1. . . SI1 kαI1 k . Here αI1 i is one of the kparameters obtained during the training process represented by the morphism m1 , SI1 i specifies the symbolic expression of αI1 i and S describes how SI1 iαI1 i are structured together. C1 consists of the modeling results with structure PI1,TPI1,FPI1,FNI1,TNI1 where P is a set of PDs obtained during the testing process, TP , FP , FN , TN are elements of the confusion matrix (see Table A1 ) that are generated by the testing process c1 , where TN is a true positive, FP is a false positive, FN is a false negative and TN is a true negative. P1 consists of performance metrics mI1, . . . , mIl generated by the morphism p1 , which is a specific implementation of p . Since the morphisms always generate unique results, they serve as the functional mapping between the set D1 , M1 , C1 and P1 . Figure 4summarises the set-valued functor I1, performing the mapping process of the first R-instance. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 4 of 21 From this structure, a specific approach to credit risk modeling is just a C-Instance of the category 𝑅 (R-instance 𝐼), represented by four elements and seven morphisms as shown in Figure 3. Figure 3. A specific credit risk model is a C-Instance of category 𝑅. 𝐷𝟏 consists of the training and testing dataset sharing the same structure 𝑓  →𝑐    , with d being the number of sample points and 𝑓  representing n features associated with one of m credit classes 𝑐   . 𝑀𝟏 is a symbolic expression of the model’s structure in the form 𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢. Here 𝛼  is one of the k parameters obtained during the training process represented by the morphism 𝑚𝟏, 𝑆  specifies the symbolic expression of 𝛼  and 𝑆 describes how 𝑆  𝛼   are structured together. 𝐶𝟏 consists of the modeling results with structure (𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ) where 𝑃 is a set of PDs obtained during the testing process, 𝑇𝑃,𝐹𝑃,𝐹𝑁,𝑇𝑁 are elements of the confusion matrix (see Table A1) that are generated by the testing process 𝑐𝟏, where 𝑇𝑁 is a true positive, FP is a false positive, FN is a false negative and TN is a true negative. 𝑃𝟏 consists of performance metrics (𝑚  ,…,𝑚  ) generated by the morphism 𝑝𝟏, which is a specific implementation of 𝑝. Since the morphisms always generate unique results, they serve as the functional mapping between the set 𝐷𝟏,𝑀𝟏, 𝐶𝟏 and 𝑃𝟏. Figure 4 summarises the setvalued functor 𝐼, performing the mapping process of the first R-instance. Figure 4. The mapping process of the first R-instance. Now suppose there is a second approach to credit risk modeling that can be represented as another R-instance 𝐼, represented by four elements and seven morphisms (Figure 5). Figure 5. A specific credit risk model is a C-Instance of category 𝑅. Figure 4. The mapping process of the first R-instance. Now suppose there is a second approach to credit risk modeling that can be represented as another R-instance I2, represented by four elements and seven morphisms (Figure 5). J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 4 of 21 From this structure, a specific approach to credit risk modeling is just a C-Instance of the category 𝑅 (R-instance 𝐼), represented by four elements and seven morphisms as shown in Figure 3. Figure 3. A specific credit risk model is a C-Instance of category 𝑅. 𝐷𝟏 consists of the training and testing dataset sharing the same structure 𝑓  →𝑐    , with d being the number of sample points and 𝑓  representing n features associated with one of m credit classes 𝑐   . 𝑀𝟏 is a symbolic expression of the model’s structure in the form 𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢. Here 𝛼  is one of the k parameters obtained during the training process represented by the morphism 𝑚𝟏, 𝑆  specifies the symbolic expression of 𝛼  and 𝑆 describes how 𝑆  𝛼   are structured together. 𝐶𝟏 consists of the modeling results with structure (𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ) where 𝑃 is a set of PDs obtained during the testing process, 𝑇𝑃,𝐹𝑃,𝐹𝑁,𝑇𝑁 are elements of the confusion matrix (see Table A1) that are generated by the testing process 𝑐𝟏, where 𝑇𝑁 is a true positive, FP is a false positive, FN is a false negative and TN is a true negative. 𝑃𝟏 consists of performance metrics (𝑚  ,…,𝑚  ) generated by the morphism 𝑝𝟏, which is a specific implementation of 𝑝. Since the morphisms always generate unique results, they serve as the functional mapping between the set 𝐷𝟏,𝑀𝟏, 𝐶𝟏 and 𝑃𝟏. Figure 4 summarises the setvalued functor 𝐼, performing the mapping process of the first R-instance. Figure 4. The mapping process of the first R-instance. Now suppose there is a second approach to credit risk modeling that can be represented as another R-instance 𝐼, represented by four elements and seven morphisms (Figure 5). Figure 5. A specific credit risk model is a C-Instance of category 𝑅. Figure 5. A specific credit risk model is a C-Instance of category R. Assume the set-valued functor I2performs the mapping as set out of Figure 6. J. Risk Financial Manag. 2021,14, 298 5 of 21 J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 5 of 21 Assume the set-valued functor I 2 performs the mapping as set out of Figure 6. Figure 6. The mapping process of the second R-instance. Since both R-instances have unique objects and morphisms that share the same exact structure, it follows that there is a natural transformation between them. Essentially, this natural transformation can be constructed as a term rewriting operation 𝛵 that replaces specific elements of one object in 𝐼 with a corresponding object in 𝐼 that satisfies the following  𝑓   →𝑐     → 𝑓   →𝑐    , 𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢 →𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢, (𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ) →(𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ), (𝑚  ,…,𝑚  ) →(𝑚  ,…,𝑚  ). (1) The existence of 𝑇 is warranted by the fact that any modeling approach would result in the same structures of their corresponding category and with the uniqueness of 𝑚𝟏, 𝑐𝟏 and 𝑝𝟏, while the operation 𝑇 ensures that the naturality condition holds for both 𝐼 and 𝐼. More specifically, there is a natural isomorphism between the two instances 𝐼 and 𝐼 (Figure 7). Figure 7. Isomorphism between two R-instances. The beauty of category theory thus comes from its design-as-proof feature. That is, given a proposition regarding relations between objects, as soon as a structure is properly constructed, the structure itself becomes a proof. The power lays in its capability to construct simple representations that captures the essence of credit risk modeling in a single concrete formalization (category), which may yield powerful insights into credit risk modeling that are difficult to identify using traditional comparative analysis of individual models usually seen in the literature. That is, different models and their underlying processes are just instances of the same modeling structures represented by a category. As a Figure 6. The mapping process of the second R-instance. Since both R-instances have unique objects and morphisms that share the same exact structure, it follows that there is a natural transformation between them. Essentially, this natural transformation can be constructed as a term rewriting operation T that replaces specific elements of one object in I1 with a corresponding object in I2 that satisfies the following nfI1 n→cI1 mod 1 T →nfI2 n→cI2 mod 1, SI1SI1 1αI1 1. . . SI1 kαI1 k T →SI2SI2 1αI2 1. . . SI2 pαI2 p, PI1,TPI1,FPI1,FNI1,TNI1T →PI2,TPI2,FPI2,FNI2,TNI2, mI1, . . . , mIlT →mI2, . . . , mI2. (1) The existence of T is warranted by the fact that any modeling approach would result in the same structures of their corresponding category and with the uniqueness of m1 , c1 and p1 , while the operation T ensures that the naturality condition holds for both I1 and I2 . More specifically, there is a natural isomorphism between the two instances I1 and I2 (Figure 7). J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 5 of 21 Assume the set-valued functor I 2 performs the mapping as set out of Figure 6. Figure 6. The mapping process of the second R-instance. Since both R-instances have unique objects and morphisms that share the same exact structure, it follows that there is a natural transformation between them. Essentially, this natural transformation can be constructed as a term rewriting operation 𝛵 that replaces specific elements of one object in 𝐼 with a corresponding object in 𝐼 that satisfies the following  𝑓   →𝑐     → 𝑓   →𝑐    , 𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢 →𝑆  󰇡𝑆  𝛼   …𝑆   𝛼   󰇢, (𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ) →(𝑃  ,𝑇𝑃  ,𝐹𝑃  ,𝐹𝑁  ,𝑇𝑁  ), (𝑚  ,…,𝑚  ) →(𝑚  ,…,𝑚  ). (1) The existence of 𝑇 is warranted by the fact that any modeling approach would result in the same structures of their corresponding category and with the uniqueness of 𝑚𝟏, 𝑐𝟏 and 𝑝𝟏, while the operation 𝑇 ensures that the naturality condition holds for both 𝐼 and 𝐼. More specifically, there is a natural isomorphism between the two instances 𝐼 and 𝐼 (Figure 7). Figure 7. Isomorphism between two R-instances. The beauty of category theory thus comes from its design-as-proof feature. That is, given a proposition regarding relations between objects, as soon as a structure is properly constructed, the structure itself becomes a proof. The power lays in its capability to construct simple representations that captures the essence of credit risk modeling in a single concrete formalization (category), which may yield powerful insights into credit risk modeling that are difficult to identify using traditional comparative analysis of individual models usually seen in the literature. That is, different models and their underlying processes are just instances of the same modeling structures represented by a category. As a Figure 7. Isomorphism between two R-instances. The beauty of category theory thus comes from its design-as-proof feature. That is, given a proposition regarding relations between objects, as soon as a structure is properly constructed, the structure itself becomes a proof. The power lays in its capability to construct simple representations that captures the essence of credit risk modeling in a single concrete formalization (category), which may yield powerful insights into credit risk modeling that are difficult to identify using traditional comparative analysis of individual models usually seen in the literature. That is, different models and their underlying processes are just instances of the same modeling structures represented by a category. As a result, there is an equivalence between the various modeling processes that creates a performance boundary: Generalization power has meaning only within the categorical frame representing the modeling process. Consequently, two different credit risk models having the same categorical structure will on average deliver the same result if tested over all possible instances of the category. In practice, this process could go on indefinitely as new datasets would create new instances. Thus, representing credit risk modeling as a J. Risk Financial Manag. 2021,14, 298 6 of 21 category yields a compact method to arrive at the equivalence concept without the burden of going through all possible empirical verifications. 2.2. Model Combination A natural consequence of categorial equivalence is that combining different types of models can result in better and more consistent forecasting performance. Empirically, this has been observed in the literature (Dastile et al. 2020). Conceptually, for model combination to be effective, two conditions must be satisfied. First, since an instance of Ddetermines C, the combination process must generate a new data instance having a structure different from the data initially used in the combination process. Second, the classification method adopted in the combination process must have a categorical structure different from the modeling process without combination. In this category, M is decoupled from C. Instead, it is mapped to D twice with the first morphism m describing the usual process of individual model construction, as shown in Figure 8. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 6 of 21 result, there is an equivalence between the various modeling processes that creates a performance boundary: Generalization power has meaning only within the categorical frame representing the modeling process. Consequently, two different credit risk models having the same categorical structure will on average deliver the same result if tested over all possible instances of the category. In practice, this process could go on indefinitely as new datasets would create new instances. Thus, representing credit risk modeling as a category yields a compact method to arrive at the equivalence concept without the burden of going through all possible empirical verifications. 2.2. Model Combination A natural consequence of categorial equivalence is that combining different types of models can result in better and more consistent forecasting performance. Empirically, this has been observed in the literature (Dastile et al. 2020). Conceptually, for model combination to be effective, two conditions must be satisfied. First, since an instance of D determines C, the combination process must generate a new data instance having a structure different from the data initially used in the combination process. Second, the classification method adopted in the combination process must have a categorical structure different from the modeling process without combination. In this category, 𝑀 is decoupled from C. Instead, it is mapped to 𝐷 twice with the first morphism 𝑚 describing the usual process of individual model construction, as shown in Figure 8. Figure 8. A category of model combination based on a stacking process that avoids the equivalent trap. The second morphism, 𝑑, represents the process of generating a new data structure by using model combination. Performance is measured by applying a new morphism c, which is essentially a computational process, that maps the new data structure in 𝐷 to 𝐶 without going through any specific model. The morphism d does not necessarily generate a unique instance of 𝐷 since its construction depends on how the output of the individual models are combined, thus reducing the likelihood of categorical equivalence. From a practical point of view, the main purpose of combining models based on the categorical framework is to address inconsistency and bias in classification performance. Inconsistency arises when models are sensitive to changes in the data structure, with their performance being valid only within specific contexts shaped by the structure and scope of the data. Bias is a result of the credit risk models used being sensitive to imbalance in default classes in the data. More specifically, models tend to be biased towards non-de- fault prediction, generating performance that at first glance seems to be satisfactory overall but are poor in terms of capturing actual default outcomes. Bias is also a result of the tendency of modelers to focus on good overall prediction outcomes, with more attention paid to non-default outcomes and less attention to stability in performance (Abdou and Pointon 2011; Dastile et al. 2020; Lessmann et al. 2015). Unfortunately, it is common to find models showing high accuracy while failing to capture actual default outcomes. The conceptual framework based on category theory provides an explanation as to why various ensemble (stacking) models proposed in the literature arrive at different conclusions regarding classification performance. Essentially, these models are caught in an equivalence trap. Ensemble models, despite their seemingly sophisticated assembling process, fuse the outputs of the base models either by majority voting or some type of Figure 8. A category of model combination based on a stacking process that avoids the equivalent trap. The second morphism, d , represents the process of generating a new data structure by using model combination. Performance is measured by applying a new morphism c, which is essentially a computational process, that maps the new data structure in D to C without going through any specific model. The morphism ddoes not necessarily generate a unique instance of D since its construction depends on how the output of the individual models are combined, thus reducing the likelihood of categorical equivalence. From a practical point of view, the main purpose of combining models based on the categorical framework is to address inconsistency and bias in classification performance. Inconsistency arises when models are sensitive to changes in the data structure, with their performance being valid only within specific contexts shaped by the structure and scope of the data. Bias is a result of the credit risk models used being sensitive to imbalance in default classes in the data. More specifically, models tend to be biased towards non-default prediction, generating performance that at first glance seems to be satisfactory overall but are poor in terms of capturing actual default outcomes. Bias is also a result of the tendency of modelers to focus on good overall prediction outcomes, with more attention paid to non-default outcomes and less attention to stability in performance (Abdou and Pointon 2011;Dastile et al. 2020;Lessmann et al. 2015). Unfortunately, it is common to find models showing high accuracy while failing to capture actual default outcomes. The conceptual framework based on category theory provides an explanation as to why various ensemble (stacking) models proposed in the literature arrive at different conclusions regarding classification performance. Essentially, these models are caught in an equivalence trap. Ensemble models, despite their seemingly sophisticated assembling process, fuse the outputs of the base models either by majority voting or some type of linear weighted combination. In doing so, no new instance of the data structure Dis created; all that has been achieved is an extension of the operation of the morphism cto cover the output combination process. As a result, the categorical structure remains the same as that of any other credit risk model with equivalent performance. In contrast, the stacking model proposed in this paper creates a new data structure Dand at the same time a new instance of model choice Mas a meta-classifier. It is the creation of Mthat effectively provides stacking models with a categorical structure that is identical to that of the typical credit J. Risk Financial Manag. 2021,14, 298 7 of 21 risk model. However, the concept of an equivalence trap also applies in the situation, as shown in Figure 9. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 7 of 21 linear weighted combination. In doing so, no new instance of the data structure D is created; all that has been achieved is an extension of the operation of the morphism c to cover the output combination process. As a result, the categorical structure remains the same as that of any other credit risk model with equivalent performance. In contrast, the stacking model proposed in this paper creates a new data structure D and at the same time a new instance of model choice M as a meta-classifier. It is the creation of M that effectively provides stacking models with a categorical structure that is identical to that of the typical credit risk model. However, the concept of an equivalence trap also applies in the situation, as shown in Figure 9. Figure 9. A category representing the equivalence trap. It is a category representing the equivalence trap often observed in typical stacking models. Essentially, the new data instance created by d can be used to train a new metaclassifier 𝑀, which in turn brings the combination process back to the original structure of the modeling process. The combination process proposed addresses this issue by considering two key issues. First, combining models, as the theoretical framework suggests, should first transform the initial feature space into a new data instance D with a structure different from the initial dataset, whilst still capturing information representing outcomes in the initial modeling phase. Second, the new data instance D should be transformed into PDs in a coherent and transparent manner without creating any new classifiers that puts the process into an equivalence trap. These considerations are supported by two conceptual constructs: Shannon’s information entropy and enriched categories, which are discussed next. 2.3. Shannon’s Information Entropy Shannon (1948) proposed a concept called entropy to measure the amount of information created by an ergodic source and transmitted over a noisy communication channel. Noises here reflect uncertainty in how signals arrive at the destination and, for finite discrete signals, they are represented by a set of probabilities 𝑝,𝑝,…,𝑝. Entropy H is defined as follows. 𝐻(𝑝,𝑝,…,𝑝)=−∑𝑝log𝑝   (2) Judged by its construction, Shannon’s information entropy captures uncertainty in the communication as it deals with noise. Shannon (1948) considered this uncertainty to be the amount of information contained in the signals, thus conceptually establishing a link between uncertainty and information. Essentially, the entropy value tells us how much uncertainty must be removed by some process to obtain information regarding which signals arrive at the destination. Thus, it can said that the amount of information received from progressing through the process, results from the removal of the uncertainty that existed before the modeling process begun. The notion of the communication channel can be generalized to a finite event space that consists of n mutually exclusive and exhaustive events with their probabilities. The connection between entropy and information enables the creation of structures that effectively capture the information contained in the modeling process, a feature that will be exploited in the stacking model as a new data structure D used to enhance prediction. Other studies that similarly exploit the concept of entropy in risk assessment are Gradojevic and Caric (2016); Lupu et al. (2020) and Pichler and Schlotter (2020). Figure 9. A category representing the equivalence trap. It is a category representing the equivalence trap often observed in typical stacking models. Essentially, the new data instance created by dcan be used to train a new metaclassifier MS , which in turn brings the combination process back to the original structure of the modeling process. The combination process proposed addresses this issue by considering two key issues. First, combining models, as the theoretical framework suggests, should first transform the initial feature space into a new data instance Dwith a structure different from the initial dataset, whilst still capturing information representing outcomes in the initial modeling phase. Second, the new data instance Dshould be transformed into PDs in a coherent and transparent manner without creating any new classifiers that puts the process into an equivalence trap. These considerations are supported by two conceptual constructs: Shannon’s information entropy and enriched categories, which are discussed next. 2.3. Shannon’s Information Entropy Shannon (1948) proposed a concept called entropy to measure the amount of information created by an ergodic source and transmitted over a noisy communication channel. Noises here reflect uncertainty in how signals arrive at the destination and, for finite discrete signals, they are represented by a set of probabilities p1 , p2 , . . . , pn . Entropy His defined as follows. H(p1,p2, . . . , pn)=− n ∑ i=1 pilog2pi(2) Judged by its construction, Shannon’s information entropy captures uncertainty in the communication as it deals with noise. Shannon (1948) considered this uncertainty to be the amount of information contained in the signals, thus conceptually establishing a link between uncertainty and information. Essentially, the entropy value tells us how much uncertainty must be removed by some process to obtain information regarding which signals arrive at the destination. Thus, it can said that the amount of information received from progressing through the process, results from the removal of the uncertainty that existed before the modeling process begun. The notion of the communication channel can be generalized to a finite event space that consists of nmutually exclusive and exhaustive events with their probabilities. The connection between entropy and information enables the creation of structures that effectively capture the information contained in the modeling process, a feature that will be exploited in the stacking model as a new data structure D used to enhance prediction. Other studies that similarly exploit the concept of entropy in risk assessment are Gradojevic and Caric (2016); Lupu et al. (2020) and Pichler and Schlotter (2020). 2.4. Enriched Categories Another important construct used in the post-stacking classification process is the concept of enriched categories (Kelly [1982] 2005). Enriched categories replace the category of sets and mappings, which play a crucial role in ordinary category theory, by a more general symmetric monoidal closed category, allowing the results of category theory to be translated into a more general setting. Enriched categories are potentially an important analytical tool for classifying default outcomes. Essentially, the paired data of entropy J. Risk Financial Manag. 2021,14, 298 8 of 21 value Hand prediction output for the training data can be separated into two groups according to the classification class associated with each output. Each group can be viewed as a set of objects that is enriched in (B,≤, true, f) , with their hom-object values, defined by whether they belong in the same group or not. A computational process is then constructed to obtain a borrower’s PD employing the following formula: PD =Ld Ld+Lnd , (3) where Ld is the likelihood that the applicant belongs to the default group and Lnd is the likelihood that the applicant belongs to the non-default group. Both Ld and Lnd are computed using the Hamming and Manhattan distance. Thus, the combination model not only provides a new classification process but also a new method of estimating PD. 2.5. The Stacking Process With the concepts of information entropy and enriched categories defined above, the stacking model is constructed as follows (see Figure A1 for a flow chart). First, several of the nine classifiers, consisting of a logistic regression and eight of the most popular supervised ML methods, are selected as the base models. Second, during the training and testing phases, the estimated PDs are used to compute the classifiers’ Shannon information entropy (H). The entropy value and the default classifications generated will then be paired to form a new (restructured) training dataset (D 2 ). Next, employing the concept of enriched categories, final predictions are formed by assigning new testing samples into either the default group or the non-default group just constructed. Finally, the performance results of the stacking model are subjected to location tests to check for consistency and biasedness. Several considerations differentiate the entropy-based stacking model proposed in this paper from the stacking models proposed by others (Doumpos and Zopounidis 2007; Wang et al. 2018). First, instead of selecting and processing the datasets carefully before training and testing a model only once on the dataset, as is usually done by others, the performance of the proposed entropy-based stacking model is assessed repeatedly on small randomly chosen non-overlapping subsets of the original dataset. Inherent class imbalance is utilized to make model comparison more realistic (Lessmann et al. 2015), enabling the construction of different data environments, and thus tests of performance inconsistency and bias. The data process ensures that each subsample will have a different structure regarding class ratio (default/non-default) and feature availability, especially categorical features. Further, performing many simulations allows for significance testing, which is preferred over making ad hoc judgements about average performance outcomes over limited rounds of tests. Thus, significance tests are a necessary complement to the usual average performance results reported by others. Statistical analyses of model performance are also proposed in Lessmann et al. (2015), but their non-parametric tests are performed on a sample of just 10. The second consideration concerns model selection. Typically, the combination models proposed in the literature carefully select base models according to their performance on some testing data. Some combination of these models will then be benchmarked against all other models. The fact that their selection greatly determines the combination model’s overall performance suggests that the base model selection process is more critical than the combination process itself. In contrast, the entropy-based stacking model proposed in this paper seeks to prove that the combination process likely offers more consistent and less biased performance results, regardless of which base models are selected. In order to achieve this goal, the simulation process is carried out over 100 different data environments, with a different number of base models used in each simulation. Moreover, in each scenario, each sample is trained and tested on a different set of base models. Thus, the only element that remains invariant in each simulation is the reasoning process underlying the stacking model. J. Risk Financial Manag. 2021,14, 298 15 of 21 Distinctly, the proposed stacking model delivers better performance in default prediction, relative to the individual base models, and for both data sets. The mean MCC is always higher for the stacking model that for the individual base models, with significance tests strongly supporting this conclusion. Most notably, the stacking model achieves consistently better performance across the various data environments as indicated by the low standard deviation of MCC. In contrast, the performance of the individual base models is highly inconsistent, as indicated by the high standard deviation of MCC. Amongst the nine individual base models, Naïve Bayes provides the best average prediction performance ( MCC = 0.14 ) for the Lending Club peer-to-peer loans dataset, whilst Random Forest provides the best average performance (MCC = 0.33) for the Taiwanese credit card clients dataset. Compared to the individual base models, the stacking model provides the best overall performance, with the mean MCC value exceeding that of any of the individual base models selected, with an overall agreement between prediction and observation twice as high for the Taiwanese credit card clients dataset compared to the Lending Club’s peer-to-peer loans dataset. While in a few cases the performance of the stacking model appears similar to the base model selected (as indicated by the mean MCC value), the individual base models always experience high Extreme Bias. For example, for the Taiwanese credit card clients dataset, while the mean MCC (about 0.32) for the Random Forest model is similar to that of the proposed stacking model, the Random Forest model experiences high Extreme Bias (4–8%), with the prediction of the base model no better than random. Again, in terms of Accuracy, the stacking model delivers highly and consistent performance across all data environments. Mean Accuracy of the stacking model tends to fluctuate close to 0.79 across all data environments. In contrast, for the individual base models, mean Accuracy fluctuates significantly between 0.66 to 0.84. None of the individual base models show consistency in performance across the data environments. Whilst the stacking model does not provide the highest mean Accuracy in all cases, in all cases it experiences the lowest Extreme Bias. This renders the Accuracy measure somewhat inapt in terms of judging prediction performance. At best, Accuracy should be used as a complement to MCC, with its usefulness viewed in terms of satisfactory consistency. That is, a good model should deliver relatively stable Accuracy. 5. Discussion The computational effort in this paper has been in running a large number of simulations to capture different data environments. The results of the simulations presented in the previous section support the proposed stacking model in terms of providing more consistent performance across data environments and less biased performance in terms of default classification. Unlike previous studies, which have been unable to settle which base model exhibits superior default classification performance across multiple data environments (Ala’raj and Abbod 2016a;Lessmann et al. 2015;Li et al. 2018;Xia et al. 2018), this paper shows that careful selection of base models is not necessary. The performance of the proposed stacking model remains high and consistent despite changes in the number and type of base model used or the data used to train the model on. In other words, the reasoning process itself is somewhat agnostic as to which base model is selected, thus enabling replication of the stacking method in a wide range of situations, allowing meaningful comparative analysis across multiple data environments. In essence, the power of the conceptual construct based on category theory lies in its capability to construct simple representations that captures the essence of credit risk modeling in a single concrete formalization (a category). It yields powerful insights into credit risk modeling that are difficult to identify using traditional comparative analysis of individual base models frequently adopted in the literature. That is, different models and their underlying processes are just instances of the same modeling structures represented by a category. As a result, there is an equivalence (trap) between the various modeling processes, creating a performance boundary. That is, generalization power has meaning J. Risk Financial Manag. 2021,14, 298 16 of 21 only within the categorical frame representing the modeling process. Consequently, two seemingly different credit risk models that have the same categorical structure will on average produce identical results if tested over all possible instances of the category. In practice, this process could continue indefinitely as new datasets create new instances. This has been clearly demonstrated by the empirical results, showing poor performance persistence of the base models selected across different data environments. It follows that representing credit risk modeling as a category yields a compact method to arrive at the equivalence concept without the burden of having to go through all possible empirical verifications, as revealed by the literature. 6. Conclusions Two motivations underly the use of category theory to credit risk modeling. First, it serves as a powerful tool to construct an inward view of our own reasoning processes in credit risk modeling. By using this view, invariant structures emerge and form a basis on which construction of the relationship between seemingly unrelated models can be created. Furthermore, category theory enables these structures to form relationships with new conceptual constructs in fields unrelated to credit risk modeling. This unique capability enlarges the space of potential modeling solutions, resulting in improved default prediction performance. Second, categorical constructs result in new perspective on the meaning of risk beyond PDs. From this perspective, credit risk is not just a quantification of specific features but also a property emerging out of a network of relationships between various modeling processes represented by enriched categories. Thus, credit risk assessment is no longer an endeavor carried out with an isolated model; it has become as a network phenomenon. Creating the theoretical framework is, therefore, a novel contribution to the current body of literature. By focusing on credit risk through these structures, the equivalence implication was better understood and a stacking model was introduced with two new structures, enriched categories and information entropy. The empirical results showed that the stacking framework’s performance remained robust despite changes in data environments and selection of the base models, thus enabling more objective replication. The conceptual structures, seemingly disconnected, turned out to be perfect companions in the stacking model. That said, there are some limitations to the paper. The first issue relates to substantial computational overhead associated with implementing the proposed stacking model. Whilst there is no doubt that keeping the per-unit processing cost low is an important concern to credit providers, advances in supercomputing are likely to push computational costs down considerably soon. The second issue relates to the performance of the stacking model which could be tested more extensively by application to more datasets and by comparing with a larger number of base models, including deep learning and unsupervised learning. This could not only create a more dynamic testing environment but also provide more transparency for replication purposes. A unified stacking and dynamic model selection framework would enable more extensive statistical tests of performance, an objective that has so far been absent from the literature but could be a fruitful avenue for further research. A final issue of concern is that the focus on constructing classification models has value only at the time of application. The focus of risk managers is undoubtedly on the development of credit risk models that provide lenders with on-going predictive diagnosis of clients’ credit risk status. However, this would require a richer dataset. While the approach embraced in this paper is essentially exploratory in its nature, it is likely to raise more questions than provide answers on sound credit risk modeling. Author Contributions: Conceptualization, C.S.T.; methodology, C.S.T. and R.N.; software, C.S.T.; analysis, C.S.T.; writing—original draft preparation, P.V.; writing—review and editing, C.S.T., D.N., R.N. and P.V. All authors have read and agreed to the published version of the manuscript. Funding: This research received no external funding. Institutional Review Board Statement: Not applicable. J. Risk Financial Manag. 2021,14, 298 17 of 21 Informed Consent Statement: Not applicable. Data Availability Statement: Not applicable. Conflicts of Interest: The authors declare no conflict of interest. Appendix A. Key Definitions in Category Theory Definition A1. A category Chas the following elements: A collection of objects denoted as Ob(C); For every two objects c and d, there is a set C(c,d) that consists of morhphims from c to d or f:c→d; For every object c∈Ob(C) , there is a morphism Idc∈ C(c,c) , called the identity morphism on c. For convenience, c∈ C is used instead of c∈Ob(C); For every three objects c , d , e∈Ob(C) and morphisms f∈ C(c,d) and g∈ C(d,e) , there is a morphism f ◦g∈ C(c,e), called the composite of f and g. These elements are required to satisfy the following conditions: For any morphism f:c→d , with idc◦f=f and f◦idd=f , which is called the unitality condition; For any three morphisms f:c0→c1 , g:c1→c2 and h:c2→c3 , the following are equal: (f◦g)◦h=f◦(g◦h). This is called the associativity condition. Definition A2. The category Set is defined as follows: Ob(Set) is the collection of all sets; If S and T are sets, then Set(X,Y) = {f:X→Y}, where f is a function; For each set S, the identity function idx:X→Y is given by idx(s):=x for each x ∈X; Given f :X→Y and g :Y→Z , their composite function is (f◦g)(x)◦g(f(x)). Since these elements satisfy the unitality and associativity conditions, Set is indeed a category. Definition A3. A functor between two categories C and D , denoted F:C → D , is defined as follows: For every object c ∈Ob(C), there is an object F(c)∈Ob(D); For every morphism f :c0→c1in C, there is a morphism F(f):F(c0)→F(c1)in D. These elements are required to satisfy the following conditions: For every object c ∈Ob(C),F(idc) = idF(c); For any three objects c0 , c1 and c2∈ C and two morphisms, f:c0→c1 , and g:c1→c2 , the equation F(f◦g)=F(f)◦F(g)holds in D. Definition A4. AC-instance of the category Cis functor I:C → Set. Definition A5. Let C and D be categories and F,G:C → D be functors. A natural transformation α:F→Gis defined as follows: For each object c∈Ob(C) , there is a morphism αc:F(c)→G(c) in D , called the c-component of α, that satisfies the following naturality condition; For every morphism f :c→d in C, the following equation holds. F(f)◦αd=αc◦G(f). A natural transformation α:F→G is called a natural isomorphism if each component αc is an isomorphism in D. The naturality condition can be represented as follows. J. Risk Financial Manag. 2021,14, 298 18 of 21 J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 17 of 21 For every morphism 𝑓:𝑐→𝑑 in 𝒞, the following equation holds. 𝑭( 𝑓 )°𝛼  =𝛼  °𝑮( 𝑓 ). A natural transformation 𝛼:𝑭→𝑮 is called a natural isomorphism if each component 𝛼 is an isomorphism in 𝒟. The naturality condition can be represented as follows. The concept of natural transformation plays an important role in understanding relations between two categories. It describes how the two functors 𝑭 and 𝑮 can be used to as two representations of category 𝒞 inside 𝒟 with the natural transformation connecting these two representations using the morphisms in 𝒟. In order to arrive at enriched categories, the following definitions apply. Definition A6. Let 𝑋 and 𝑌 be sets. A relation between 𝑋 and 𝑌 is a subset 𝑅⊆𝑋×𝑌. A binary relation on 𝑋 is a relation between 𝑋 and 𝑋, i.e. a subset of 𝑅⊆𝑋×𝑋. Definition A7. A preorder relation on a set 𝑋 is binary relation on 𝑋, denoted as ≤, that satisfies the following two properties: Reflexivity: 𝑥≤𝑥; and Transitivity: If 𝑥≤𝑦 and 𝑦≤𝑧, then 𝑥≤𝑧. The preorder can be denoted as (𝑋,≤). Definition A8. A symmetric monoidal structure on a preorder (𝑋,≤) has the following two elements: An element 𝐼∈𝑋, called the monoidal unit; A function ⊗: 𝑋×𝑋→𝑋, called the monoidal product. These elements must satisfy the following four properties: Monotocity: for all 𝑥  ,𝑥  ,𝑦  ,𝑦  ∈𝑋,𝑖𝑓 𝑥  ≤𝑦  𝑎𝑛𝑑 𝑥  ≤𝑦  ,𝑡ℎ𝑒𝑛 𝑥  ⊗𝑥  ≤𝑦  ⊗𝑦  ; Unitality: for all 𝑥∈𝑋, the equations 𝐼⊗𝑥=𝑥 𝑎𝑛𝑑 𝑥⊗𝐼=𝑥 hold; Associativity: for all 𝑥,𝑦,𝑧∈𝑋,𝑡ℎ𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛 (𝑥⊗𝑦)⊗𝑧=𝑥⊗(𝑦⊗𝑧) holds; Aymmetry: for all 𝑥,𝑦∈𝑋,𝑡ℎ𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛 𝑥⊗𝑦=𝑦⊗ 𝑥 holds. This structure is called a symmetric monoidal preorder and is denoted as (𝑋,≤,𝐼,⊗). Definition A9. A symmetric monoidal structure on a preorder (𝑋,≤) has the following two elements: An element 𝐼∈𝑋, called the monoidal unit; A function ⊗: 𝑋×𝑋→𝑋, called the monoidal product. These elements must satisfy the following properties: Monotocity: for all 𝑥  ,𝑥  ,𝑦  ,𝑦  ∈𝑋,𝑖𝑓 𝑥  ≤𝑦  𝑎𝑛𝑑 𝑥  ≤𝑦  ,𝑡ℎ𝑒𝑛 𝑥  ⊗𝑥  ≤𝑦  ⊗𝑦  ; Unitality: for all 𝑥∈𝑋, the equations 𝐼⊗𝑥=𝑥 𝑎𝑛𝑑 𝑥⊗𝐼=𝑥 hold; Associativity: for all 𝑥,𝑦,𝑧∈𝑋,𝑡ℎ𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛 (𝑥⊗𝑦)⊗𝑧=𝑥⊗(𝑦⊗𝑧) holds; Symmetry: for all 𝑥,𝑦∈𝑋,𝑡ℎ𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛 𝑥⊗𝑦=𝑦⊗ 𝑥 holds. The concept of natural transformation plays an important role in understanding relations between two categories. It describes how the two functors F and G can be used to as two representations of category C inside D with the natural transformation connecting these two representations using the morphisms in D. In order to arrive at enriched categories, the following definitions apply. Definition A6. Let X and Y be sets. A relation between X and Y is a subset R⊆X×Y . A binary relation on X is a relation between X and X, i.e. a subset of R ⊆X×X. Definition A7. A preorder relation on a set X is binary relation on X , denoted as ≤ , that satisfies the following two properties: Reflexivity: x ≤x; and Transitivity: If x ≤y and y ≤z, then x ≤z. The preorder can be denoted as (X,≤). Definition A8. A symmetric monoidal structure on a preorder (X,≤) has the following two elements: An element I ∈X, called the monoidal unit; A function ⊗:X×X→X , called the monoidal product. These elements must satisfy the following four properties: Monotocity: for all x1,x2,y1,y2∈X,i f x1≤y1and x2≤y2,then x1⊗x2≤y1⊗y2; Unitality: for all x ∈X, the equations I ⊗x=x and x ⊗I=x hold; Associativity: for all x,y,z∈X,the equation (x⊗y)⊗z=x⊗(y⊗z)holds; Aymmetry: for all x,y∈X,the equation x ⊗y=y⊗x holds. This structure is called a symmetric monoidal preorder and is denoted as (X,≤,I,⊗). Definition A9. A symmetric monoidal structure on a preorder (X,≤) has the following two elements: An element I ∈X, called the monoidal unit; A function ⊗:X×X→X , called the monoidal product. These elements must satisfy the following properties: Monotocity: for all x1,x2,y1,y2∈X,i f x1≤y1and x2≤y2,then x1⊗x2≤y1⊗y2; Unitality: for all x ∈X, the equations I ⊗x=x and x ⊗I=x hold; Associativity: for all x,y,z∈X,the equation (x⊗y)⊗z=x⊗(y⊗z)holds; Symmetry: for all x,y∈X,the equation x ⊗y=y⊗x holds. This structure is called a symmetric monoidal preorder denoted as (X,≤,I,⊗) . Let B=(f alse,true)and f alse ≤true , the structure (B,≤,true,f) cam be developed with f representing the AND operation defined in the following matrix. ffalse true false false false true false True It is trivial to show that this structure forms a symmetric monoidal structure. J. Risk Financial Manag. 2021,14, 298 19 of 21 Definition A10. Let V=(V,≤,I,⊗) be a symmetric monoidal preorder. A V -category X has the following two elements: A set Ob(X), elements of which are called objects; For every two objects x,y∈Ob(X), there is an element X(x,y)∈ V, called the hom-object. These elements must satisfy the following two properties: For every object x ∈Ob(X), I ≤ X (x,x); For every three objects x,y,z∈Ob(X), all X(x,y)⊗ X (y,z)≤ X (x,z). Hence, it can be said that Xis enriched in V. Table A1. Confusion Matrix. Prediction Default Non-Default Actual Default TP FN Non-Default FP TN Notes: “Positive (P)” is the term used to describe a prediction of default and “Negative (N)” for a prediction of non-default outcome. “True (T)” means the actual data agrees with the prediction, whilst “False (F)” means the data does not agree with the prediction. J. Risk Financial Manag. 2021, 14, x FOR PEER REVIEW 19 of 21 Figure A1. Flow Diagram of the Proposed Entropy-based Stacking Model. Figure A1. Flow Diagram of the Proposed Entropy-based Stacking Model. J. Risk Financial Manag. 2021,14, 298 20 of 21 References Abdou, Hussein A., and John Pointon. 2011. Credit scoring, statistical techniques and evaluation criteria: A review of the literature. Intelligent Systems in Accounting, Finance and Management 18: 59–88. [CrossRef] Abellán, Joaquin, and Carlos J. Mantas. 2014. Improving experimental studies about ensembles of classifiers for bankruptcy prediction and credit scoring. Expert Systems with Applications 41: 3825–30. [CrossRef] Ala’raj, Maher, and Maysam F. Abbod. 2016a. Classifier’s consensus system approach for credit scoring. Knowledge-Based Systems 104: 89–105. [CrossRef] Ala’raj, Maher, and Maysam F. Abbod. 2016b. A new hybrid ensemble credit scoring model based on classifiers consensus system approach. Expert Systems with Applications 64: 36–55. [CrossRef] Altman, Edward I. 1968. Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. The Journal of Finance 23: 589–609. [CrossRef] Altman, Edward I., and Gabriele Sabato. 2007. Modeling credit risk for SMEs: Evidence from the U.S. Market. Abacus 43: 332–57. [CrossRef] Asquith, Paul, David W. Mullins, and Eric D. Wolff. 1989. Original issue high yield bonds: Aging analyses of defaults, exchanges, and calls. The Journal of Finance 44: 923–52. [CrossRef] Aster, Richard C., Brian Borchers, and Clifford H. Thurber. 2018. Parameter Estimation and Inverse Problems, 3rd ed. Amsterdam: Elsevier Publishing Company. Barboza, Flavio, Herbert Kimura, and Edward Altman. 2017. Machine learning models and bankruptcy prediction. Expert Systems with Applications 83: 405–17. [CrossRef] Breiman, Leo. 2001. Random forests. Machine Learning 45: 5–32. [CrossRef] Brown, Peter F., Vincent J. Della Pietra, Peter V. Desouza, Jenifer C. Lai, and Robert L. Mercer. 1992. Class-based n-gram models of natural language. Computational Linguistics 18: 467–80. Chang, Shunpo, Simon D-O Kim, and Genki Kondo. 2015. Predicting default risk of lending club loans. CS229: Machine Learning, 1–5. Available online: http://cs229.stanford.edu/proj2018/report/69.pdf (accessed on 24 January 2021). Chen, Tianqi, and Carlos Guestrin. 2016. XGBoost: A scalable tree boosting system. Paper presented at the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13–17; New York, NY, USA: Association for Computing Machinery, pp. 785–794. [CrossRef] Cortes, Corinna, and Vladimir Vapnik. 1995. Support-vector networks. Machine Learning 20: 273–97. [CrossRef] Dastile, Xolani, Turgay Celik, and Moshe Potsane. 2020. Statistical and machine learning models in credit scoring: A systematic literature survey. Applied Soft Computing 91: 106263. [CrossRef] Doumpos, Michael, and Constantin Zopounidis. 2007. Model combination for credit risk assessment: A stacked generalization approach. Annals of Operations Research 151: 289–306. [CrossRef] Eilenberg, Samuel, and Saunders MacLane. 1945. General theory of natural equivalences. Transactions of the American Mathematical Society 58: 231–94. [CrossRef] Finlay, Steven. 2011. Multiple classifier architectures and their application to credit risk assessment. European Journal of Operational Research 210: 368–78. [CrossRef] Gradojevic, Nikola, and Marko Caric. 2016. Predicting systemic risk with entropic indicators. Journal of Forecasting 36: 16–25. [CrossRef] Henley, William, and David J. Hand. 1996. A k-nearest-neighbour classifier for assessing consumer credit risk. Journal of the Royal Statistical Society: Series D (The Statistician) 45: 77–95. [CrossRef] Hsieh, Nan-Chen, and Lun-Ping Hung. 2010. A data driven ensemble classifier for credit scoring analysis. Expert Systems with Applications 37: 534–45. [CrossRef] Jonkhart, Marius J. L. 1979. On the term structure of interest rates and the risk of default: An analytical approach. Journal of Banking & Finance 3: 253–62. [CrossRef] Joy, Maurice O., and John O. Tollefson. 1978. Some clarifying comments on discriminant analysis. Journal of Financial and Quantitative Analysis 13: 197–200. [CrossRef] Kelly, Max G. 2005. Basic Concepts of Enriched Category Theory. London Mathematical Society Lecture Note Series 64; Cambridge: Cambridge University Press. Reprinted as Reprints in Theory and Applications of Categories 10. First published 1982. Lawrence, Edward C., Douglas L. Smith, and Malcolm Rhoades. 1992. An analysis of default risk in mobile home credit. Journal of Banking & Finance 16: 299–312. [CrossRef] Lending Club. 2020. Peer-to-Peer Loans Data. Available online: https://www.kaggle.com/wordsforthewise/lending-club (accessed on 24 November 2020). Lessmann, Stefan, Bart Baesens, Hsin-Vonn Seow, and Lyn C. Thomas. 2015. Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research 247: 124–36. [CrossRef] Li, Wei, Shuai Ding, Yi Chen, and Shanlin Yang. 2018. Heterogeneous ensemble for default prediction of peer-to-peer lending in China. IEEE Access 6: 54396–406. [CrossRef] Lupu, Radu, Adrian C. Călin, Cristina G. Zeldea, and Iulia Lupu. 2020. A bayesian entropy approach to sectoral systemic risk modeling. Entropy 22: 1371. [CrossRef] Malekipirbazari, Milad, and Vural Aksakalli. 2015. Risk assessment in social lending via random forests. Expert Systems with Applications 42: 4621–31. [CrossRef] J. Risk Financial Manag. 2021,14, 298 21 of 21 Martin, Daniel. 1977. Early warning of bank failure: A logit regression approach. Journal of Banking & Finance 1: 249–76. [CrossRef] Matthews, Ben W. 1975. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochimica et Biophysica Acta (BBA)-Protein Structure 405: 442–51. [CrossRef] Mcleay, Stuart, and Azmi Omar. 2000. The sensitivity of prediction models to the non-normality of bounded and unbounded financial ratios. British Accounting Review 32: 213–30. [CrossRef] Ohlson, James A. 1980. Financial ratios and the probabilistic prediction of bankruptcy. Journal of Accounting Research 18: 109–31. [CrossRef] Pichler, Alois, and Ruben Schlotter. 2020. Entropy based risk measures. European Journal of Operational Research 285: 223–36. [CrossRef] Rish, Irina. 2001. An empirical study of the naive Bayes classifier. Paper presented at the IJCAI 2001 Workshop on Empirical Methods in Artificial Intelligence, Seattle, WA, USA, August 4–6; vol. 3, pp. 41–46. Safavian, Stephen R., and David Landgrebe. 1991. A survey of decision tree classifier methodology. IEEE Transactions on Systems, Man, and Cybernetics 21: 660–74. [CrossRef] Santomero, Anthony. M., and Joseph D. Vinso. 1977. Estimating the probability of failure for commercial banks and the banking system. Journal of Banking & Finance 1: 185–205. [CrossRef] Shannon, Claude E. 1948. A mathematical theory of communication. The Bell System Technical Journal 27: 379–423. [CrossRef] Tam, Kar Yan. 1991. Neural network models and the prediction of bank bankruptcy. Omega 19: 429–45. [CrossRef] Teply, Petr, and Michal Polena. 2020. Best classification algorithms in peer-to-peer lending. North American Journal of Economics and Finance 51: 100904. [CrossRef] Tsai, Ming-Chun, Shu-Ping Lin, Ching-Chan Cheng, and Yen-Ping Lin. 2009. The consumer loan default predicting model. An application of DEA–DA and neural network. Expert Systems with Applications 36: 11682–90. [CrossRef] Vassalou, Maria, and Yuhang Xing. 2004. Default Risk in Equity Returns. The Journal of Finance 59: 831–68. [CrossRef] Wang, Maoguang, Jiayu Yu, and Zijian Ji. 2018. Personal credit risk assessment based on stacking ensemble model. Paper presented at the 10th International Conference on Intelligent Information Processing (IIP), Nanning, China, October 19–22. Wolpert, David H. 1992. Stacked generalization. Neural Networks 5: 241–59. [CrossRef] Xia, Yufei, Chuanzhe Liu, Bowen Da, and Fangming Xie. 2018. A novel heterogeneous ensemble credit scoring model based on stacking approach. Expert Systems with Applications 93: 182–99. [CrossRef] Yeh, I-Cheng. 2006. Default of Credit Card Clients Data Set. Department of Information Management, Chung Hua University, Taiwan. Available online: https://archive.ics.uci.edu/ml/datasets/default+of+credit+card+clients (accessed on 24 November 2020). Yeh, I-Cheng, and Che-hui Lien. 2009. The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients. Expert Systems with Applications 36: 2473–80. [CrossRef]