Full text
Open Access © The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the mate‑ rial. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. RESEARCH RugelesDiazetal. Financial Innovation (2025) 11:137 https://doi.org/10.1186/s40854-025-00879-5 Financial Innovation Data analytics toprevent retail credit card fraud: empirical evidence fromLatin America Leidy Tatiana Rugeles Diaz1 , Miguel Ángel Echarte Fernández2 , Javier Jorge‑Vázquez2,3 and Sergio Luis Nañez Alonso2* Abstract Reducing the risk of fraud in credit card transactions is crucial for the competitive‑ ness of companies, especially in Latin American countries. This study aims to estab‑ lish measures for preventing and detecting fraud in the use of credit cards in shops through analytical methods (data mining, machine learning and artificial intelligence). To achieve this objective, the study employs a predictive methodology using descrip‑ tive and exploratory statistics and frequency, frequency & monetary (RFM) classification techniques, differentiating between SMEs and large businesses via cluster analysis and supervised models. A dataset of 221,292 card records from a Latin American merchant payment gateway for the year 2022 is used. For fraud alerts, the classifica‑ tion model has been selected for small and medium–sized merchants, and the mul‑ tilayer perceptron (MLP) neural network has been selected for large merchants. Random forest or Gini decision tree models have been selected as backup models for retraining. For the detection of punctual fraud patterns, the K‑means and parti‑ tioning around medoids (PAM) models have been selected, depending on the type of trade. The results revealed that the application of the identified models would have prevented between 48 and 85% of fraud transactions, depending on the trade size. Despite the promising results, continuous updating is recommended, as fraudsters frequently implement new fraud techniques. Keywords: Fraud prevention, Risk, Competitiveness, Machine learning, Credit cards JEL Classification: G17, G14, M15, M21 Introduction, theoretical background andresearch objective Rapid technological development in recent years and the expansion of e-commerce have markedly increased credit card fraud. In this study, an analysis was carried out using supervised and unsupervised models to classify and predict fraudulent credit card transactions among merchants in Latin America (Ballezza 2007; Balagolla etal. 2021; Ahmed 2022; Lokanan 2022). This research focuses on the timing of the payment gateway (Kim and Kim 2020; Kilay etal. 2022). Preventive measures are provided to reject fraudulent transactions (Barker etal. 2008; Delamaire etal. 2009; Chaudhary etal. 2012), and alerts are generated to warn merchants of recent transactions that have already been approved and are very likely to be fraudulent. In fraud prevention, many manual processes can *Correspondence: sergio[email protected] 1 Department of DataScience, Vana, Bogotá, Colombia 2 Faculty of Social and Legal Sciences, Department of Economics, Dekis Research Group, Catholic University of Ávila, Ávila, Spain 3 Facultad de Economía y Empresa. Universidad Internacional de La Rioja, Avenida de la Paz, 137, Logroño 26006, Spain
Page 2 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 be automated and optimized, and even some decisions can be made on the basis of results from models. The concepts of statistics, data mining (Chan etal. 1999; Bhattacharyya etal. 2011; Ernawati etal. 2021), machine learning (Sarker 2021; Tiwari etal. 2021; Tanouz etal. 2021; Krishna Raoet al., 2021; Hasan, 2022) and artificial intelligence (Beena etal. 2021; Zhang and Lu 2021; Akinbowale etal. 2022; Alhabib etal. 2024; Vergara etal. 2024, 2025) play crucial roles in optimally preventing fraud. On the other hand, the high rates of credit card fraud in Latin America (Álvarez and Ladino 2021) justify the interest of this region as a territorial unit of study in this research. Fraudulent operations have affected companies in different countries and many users. Examples include Home Depot in 2014, with 56 million affected users, and Sony in 2011, with 12 million affected users (Al Smadi and Min 2020). Technology means using more sophisticated fraud techniques, but at the same time, it facilitates the development of more effective strategies to combat it. In card transactions, fraud can be defined and can occur in different ways. First, there is theft of the card owner’s information and the use of that data to make online purchases (Chen and Lai 2021; Gianotti and Damião da Silva 2021; King etal. 2021; Mekterovic etal. 2021; Siblini etal. 2021). This information theft can be accomplished through phone calls or the financial institution’s website (Alatawi 2025). Second, tools that calculate credit card numbers and their expiration dates and card verification value codes (CVVs) are used (Al Smadi and Min 2020; Ernawati etal. 2021; Hasan 2022). Third, bots that perform many transactions simultaneously are used. Fourth, the websites of official shops, banks or merchants should be impersonated (Mekterovic etal2021; Siblini etal. 2021). For these reasons, fraud is regulated by legislation in different countries (Sanz-Bas etal. 2021; Náñez Alonso etal 2024). Transactions go through a payment gateway, where all online transactions are handled. Here, the gateway can approve or reject transactions through an anti-fraud module that is managed by people and tools. The approved transactions go through the card issuing bank, which may also have an additional anti-fraud module. These processes help prevent and detect the risk of fraud and are vital factors in gaining a competitive advantage over companies that do not adopt them (Ballezza 2007; Tanouz etal. 2021). In addition, they improve customer experience by offering higher levels of security in payment processing, which can help retain them as customers (Brito etal. 2024). On an economic and developmental level, some authors, such as Bajwa (2017), note that the existence of secure e-commerce is a fundamental element for entrepreneurship and economic development, especially in emerging countries. Additionally, (Tran 2021) indicated that the security of virtual payment transactions is a competitive advantage over merchants, who cannot secure and detect fraud, even for more sustainable trade. The significant economic implications of credit card fraud have sparked interest in academia. Recently, numerous studies have developed two predominant research approaches: the detection and prediction of credit card fraud and the design of advanced solutions based on new technologies to prevent it. Regarding the first approach, the most recent studies have focused their efforts on the application of machine learning methods for the detection of these fraudulent transactions (Alharbi etal. 2022). Several studies have recently used neural networks to address this issue. Forough and Momtazi (2021) propose a fraud detection model for credit card transactions that combines deep neural networks and probabilistic graphical models. The results obtained in their
Page 3 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 study demonstrate significant improvements in fraud detection accuracy and efficiency. Lei etal. (2023) propose a distributed neural network (DDNN) model to detect credit card fraud transactions. They conclude that their proposal improves detection accuracy compared with other models, avoids leakage of private data and reduces handling costs. Huang etal. (2024) proposed a hybrid neural network to detect fraud in credit card transactions on the basis of identity and transaction features. The experimental results show that the method is more efficient. Cherif etal. (2024) developed a method to detect fraud associated with credit card transactions via graphical neurons (GNNs), optimizing feature selection and relationship modeling. The results of the validation of the proposed model support its improved accuracy. With respect to the second approach, several studies have examined the contribution and application of emerging digital technologies to prevent fraud in credit card transactions. Dadhich etal. (2022) identify fraudulent card payment methods used by cybercriminals and provide an analysis of systems that can detect and block their action. Chatterjee and Rawat (2024) analyze the role of blockchain technology and federated learning in fraud detection. They conclude that the features of this technology make transactions more transparent, reliable and secure. This helps prevent fraud in the financial industry. Other authors, such as Ogbanufe and Kim (2018), argue that the use of biometric authentication is a more secure solution for credit card authentication, which reduces the number of fraudulent transactions. Nevertheless, despite these technological advances, Cherif etal. (2023) conclude that there is a lack of sufficient research to address the challenges posed by credit card fraud. Furthermore, nonparametric statistical tests conducted on the dataset—specifically the Kruskal‒Wallis test (H = 256.4; p < 0.001) and the prevalence ratio of fraud (χ2 = 219.3; p < 0.001)—confirm that fraud patterns differ statistically significantly between small, medium, and large businesses, which supports the need to model each stratum separately. These findings justify modeling each business segment separately. This stratified approach also aligns with the literature, which shows that algorithmic performance varies depending on company size and data characteristics. In this context, SMEs have emerged as a particularly critical segment. They are vital to the global economy and all societies. However, they face a complex and challenging environment, as they lag behind in digital transformation in most sectors (Kotios etal. 2022). The detection of anomalous behavior in business process data is crucial for preventing failures that may jeopardize the performance of any organization (Elaziz etal. 2023). Supervised learning techniques are impracticable because of the difficulty of gathering large amounts of labeled business process anomaly data (Elaziz etal. 2023). Recent studies have shown that segmenting companies by size can significantly improve fraud detection accuracy, as fraud patterns, capital constraints, and incentives to report differ across strata. These insights not only justify segmentation for analytical rigor but also inform practical deployment strategies—helping payment processors and fraud analysts allocate resources more effectively across business tiers. For example, approximately 99% of companies worldwide are SMEs and often lack external audits, making them particularly vulnerable to fraud (Ngai etal. 2011). In Latin America, there have even been documented cases of financial statement manipulation to obtain credit lines above actual capacity (Carcillo etal. 2021). In response to such challenges, assembled trees have become the first line of defense:
Page 4 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 the random forest outperforms deep networks on Latin American transaction datasets (Carcillo etal. 2021) and SME financial statement audits, where it achieves 97% recovery compared with 92% for an MLP (Kaur etal. 2021). This same conclusion is reached by Bin Sulaiman etal. (2022), who warned that the dense structure of MLPs can lead to inefficiency. Tree ensembles also dominate advertising click fraud (Ziakis and Vlachopoulou 2023); a recent study achieved 95% accuracy with a random forest (Abbas etal. 2025). Cost-sensitive stacking methods—e.g., logistic regression + XGBoost—reduce false positives and expected losses in multinational bank panels (Charizanos etal. 2024; Tumminello etal. 2022). Neural approaches remain attractive for complex nonlinear patterns: an MLP trained on ledgers achieved 93% AUC (Bin Sulaiman etal. 2022)— or 94% according to J. Zhang etal. (2024)—with latency < 20 ms, which is crucial for online microcommerce. However, systematic reviews confirm that shallow tree ensembles or Bayesian models often outperform stand-alone MLPs when the sample size varies greatly across strata (Hastie etal. 2009; Costa and Pedreira 2022; Blockeel etal. 2023). Probabilistic models retain their niche: Naive Bayes does not usually top the rankings, but its deployment without adjustment and transparent likelihood ratios make it popular in low-resource microretail contexts (Viaene etal. 2005; Assaf etal. 2011; Vosseler 2022; Saha etal. 2023). By relaxing independence, Bayesian networks made it possible to reduce false positives in insurance claims (Mary and Claret 2023). Support vector machines (SVMs), although slightly behind ensembles, remain competitive when the attribute space explodes after one-hot encoding of heterogeneous variables, especially when stratified by turnover (Rtayli and Enneya 2020). These findings underscore the importance of aligning the model architecture with both the nature of the data and the operational constraints of the business segment. Finally, classic clustering still underpins exploratory segmentation. The canonical triad—k-means (MacQueen 1967), PAM with Gower distance (Kaufman and Rousseeuw 1990) and agglomerative linkage (Murtagh and Contreras 2011)—covers spherical, medoid-centered and hierarchical structures, respectively, and is frequently combined with supervised back ends. Together, these studies confirm that no single model dominates universally; rather, aligning algorithmic bias with merchant-size segments (and their fraud cost structure) yields the best trade-off between detection power, interpretability and operational cost (Cybenko 1989). This research has two fundamental objectives to improve the detection and prevention of fraud in credit card transactions, adapting to the specific needs of small, medium and large merchants. First, the characteristics of fraudulent transactions are identified to prevent fraud among small, medium and large merchants. Second, we predict whether recently approved and executed transactions may be fraudulent and determine the likelihood of this occurrence alerting merchants. This study investigates how machine learning models can improve fraud detection in credit card transactions with merchants in Latin America. The research question is as follows: What is the best method for detecting fraud among small, medium and large merchants in a region? This research closes a significant gap in the scientific literature by analyzing, for the first time, the best methods according to business size: small, medium or large. The main
Page 5 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 contribution of this research is the implementation and evaluation of the multilayer perceptron (MLP) neural network in the context of credit card fraud detection. In contrast to other studies that have focused primarily on more developed regions, our research focuses on Latin America. This region presents specific challenges due to its particular socioeconomic and technological characteristics. The findings of this study demonstrate that the MLP offers higher predictive accuracy in fraud detection. Owing to its ability to capture nonlinear and complex patterns in transaction data, the MLP is particularly effective for contexts with a high volume of transactions, as is the case in Latin America. While recurrent models are suitable for handling sequential data, the MLP stands out for its ability to detect anomalies in large transactional datasets. In addition to confirming the superior accuracy of the MLP, this study makes a significant contribution by analyzing transactions in SMEs in Latin America, a topic that has not been extensively explored in the literature. Equally important are the practical implications of the findings from a regulatory and policy perspective. The study discusses how financial supervisors can leverage these findings to develop more effective fraud prevention and control policies. This suggests the need for regulations that encourage businesses and financial institutions to adopt advanced fraud detection technologies, such as the MLP, thereby reducing financial losses and enhancing the security of credit card transactions. Materials andmethodology Materials The database used in this study includes 221,292 anonymized records of credit card transactions in Latin American merchant payment gateways from 2022. The database consists of 163 columns, of which 1 is the id of the transactions, 24 are categorical, 6 are date, and 132 are numeric, all of which are variables used. The datasets analyzed during the present study are not publicly available due to their ability to collect banking data but are available from the corresponding author upon reasonable request. The data were processed in the following way. First, the transactions were anonymized to protect the identity of individuals. The columns were subsequently identified and categorized (as shown in Table1), and the transactions were filtered. Only transactions approved by the payment gateway and the bank were considered. To balance the data between fraudulent transactions and nonfraudulent transactions, stratified sampling was used. Methodology Merchant‑size classification (RFM) To segment merchants, we adapt the classic Recency-Frequency-Monetary (RFM) framework to a fraud prevention context. Numerous studies have shown the effectiveness of RFM analysis for segmentation (Kohavi and Parekh 2004; Miroshnychenko and Aleksieieva 2022; Rungruang etal 2023; Rajeshwari 2024; Changani 2024; Náñez Alonso etal. 2025). For each merchant, we computed the following: Recency (R) = average days between two transactions; frequency (F) = number of transactions; and monetary (M) = total turnover in USD. Each variable was discretized into deciles with pd.qcut, producing the cutoff values listed in Table15 (Appendix B). A composite score S = R + F + M (range
Page 6 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 3–30) was then calculated. Merchants were finally grouped as follows: small (S ≤ 12, bottom 0–33%), medium (13 ≤ S ≤ 21, 34–66%), and large (S ≥ 22, 67–100%). This trichotomy maximizes between-group variance (one-way ANOVA F = 417.9; p < 0.001) while maintaining a balanced operational load (≈35%, 31% and 34% of transactions, Table 1 Field, category and data description Source: Own elaboration Category Field Description General Transaction Information id Unique transaction identifier General Transaction Information commerce Name or code of the trade associated with the transaction General Transaction Information country_com Country of the trade: Argentina, Brazil, Chile, Colombia, Mexico, Venezuela and Peru General Transaction Information date, month, week, day, time Temporal data indicating when the transaction occurred Customer Information num_id Anonymized unique customer identifier Customer information name Customer’s name (anonymized) Customer Information email, tel, ip Customer’s email, phone and IP, all anonymized Customer Information country_dir, city, dept, dir Customer’s geographic location, including country, city, department and address (all anonymized) Payment Method md_payment Payment method used (e.g., credit card, debit card) Payment method bin Card issuer identifier (first 6 digits of the card) Payment Method num_tc Card number (anonymized) Payment method ban_emi Card issuing bank. Payment gateway data provider GPO u‑Pay. Includes transactions from the main banks in the region such as BBVA, Banco Santander, Banco de la Provincia de Buenos Aires, Macro, Galicia, HSBC, Banco Colombia, Citibank, Scotiabank, Banorte, Banco Patagonia, Banco Azteca, Banamex, Banco Bolivariano, Banco de Guayaquil, Bancomer, Banco do Brasil, Bancomeva, Banesco Fraud indicators and interattribute relationships ti Time in which the transaction was processed Fraud indicators and interattribute relationships fraud Binary variable indicating whether the transaction was fraudulent (1) or not (0) Heuristic relationships between attributes tx_… Frequency of transactions associated with a specific attribute Heuristic relationships between attributes tc_… Frequency of cards related to other attributes Heuristic relationships between attributes bin_… Association between the BIN and other attributes Heuristic relationships between attributes ip_… Frequency and relationship of IP addresses with other attributes Heuristic relationships between attributes email_… Association of e‑mails with other attributes Heuristic relationships between attributes _h, _d, _s Represent different time intervals: hours, days, weeks Transaction status est_ini, est_fin Initial and final state of the transaction Transaction state t_est_fin Elapsed time to final status
Page 7 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 respectively). The same decile cutoffs allow reproducible scoring of new data by recomputing R, F, and M over a rolling 12-month window. Variables, temporal partitioning, andmodeling flow The date variables (6) are date_ms, date, month, week, day and time. The categorical variables (24) are in turn divided into two parts: first, informative: the IP associated with the location from which the transaction was carried out, the city of the customer’s address, the department of the customer, the address of the customer, the credit card number, the initial status of the transaction, the final status of the transaction and the type of final status of the transaction. Second, explanatory methods include the following: country of customer address, country of merchant, payment means full description, primary payment means, type of payment means, bin associated with the credit card, credit card issuing bank, card type, and range of the amount of the transaction in USD. With respect to the numerical variables (132), which are the explanatory variables, we find the following: amount in USD, standardization of amount (USD) by means, standardization of amount by minimum and maximum, maximum number of transactions per phone in one hour, maximum number of transactions per ID document in one hour, and maximum number of transactions per iP in one hour. From the beginning, all numerical variables are considered explanatory variables, but only those with variances greater than 0 are initially selected (variables with low volatility are discarded). As these variables are constant (low volatility), they will not help explain the target variable, since their value is always constant for fraudulent transactions and the same for genuine transactions, so these two events do not show dissimilarity with these characteristics. After this, the following variables have been discarded: standardization of the number by the minimum and maximum, maximum number of different phones per ID document in an hour, per ID document in a day, and per ID document in a week. Finally, the target variable would be one: Fraud (whether the transaction was fraudulent). This allows us to determine fraud both on the entire dataset and on approved transactions. This is shown in Table2. To account for concept drift in fraud techniques, the transactions were sorted chronologically and split into two nonoverlapping periods. Records from 1 January–30 September (178k transactions, 80.5%) formed the training set, whereas 1 October–31 December (43k transactions, 19.5%) served as the temporal validation set. The split preserves class balance through stratified sampling on the target variable (Fraud/No Fraud). All the models were fitted exclusively on the first period and evaluated on the second, thus emulating real-world deployment where new fraud strategies appear after model training. This step is executed in Python with the function train_test_split, leaving 80% of the dataset for training and 20% for validation. An additional temporal split was considered; however, exploratory analysis revealed no statistically significant concept drift between Q1–Q3 Table 2 Target categorical variables—Fraud Source: Own elaboration Over the entire base On approved transactions Fraud Tx number % Fraud Tx number % No fraud 219.412 99% No fraud 94.768 98% Fraud 1.880 1% Fraud 1.880 2%
Page 8 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 and Q4 (PSI < 0.05 for all 163 features, ΔAUC = 0.002). Consequently, the random 80/20 stratified split was retained to maximize the sample size without biasing the models. Once this is done, unsupervised and supervised models are applied for each of the groups of businesses obtained from the analysis (Carcillo etal. 2021). Table3 shows the unsupervised models. K-means (partition-based), PAM (medoid-based with mixed-type Gower distance) and agglomerative linkages represent the three canonical families of clustering (MacQueen 1967; Kaufman and Rousseeuw 1990; Murtagh and Contreras 2011). Their complementary bias-spherical vs. density-preserving vs. hierarchical topology covers the most frequent fraud topologies reported in payment data (Carcillo etal. 2021). Selecting this triad therefore maximizes the chance of discovering both globular and elongated fraud patterns while keeping the results interpretable for rule extraction. Businesses are classified as small, medium or large, with the variables selected from business knowledge by forced inclusion and supervised nonparametric tests (chi-square, Levene homoscedasticity and the Mann‒Whitney means test) (Carcillo etal. 2021). The results of the parametric tests and forced inclusion variables are available in the generated dataset. All the parametric test outputs and the list of forced-inclusion variables are openly accessible in the Appendix A-Statistics dataset (and online supplementary material). After data balancing and parametric tests, the models are trained with a balanced sample, which is obtained by applying the clustering method in each of the groups of businesses and implementing the PAM algorithm with the Gower distance. After a model is implemented, the clusters are evaluated to determine whether the rules help stop fraud and whether the proportion in which they affect genuine Table 3 Unsupervised models Source: Own elaboration Equation/ Algorithm Number Formula Description PAM‑Gower distance (Kauff‑ mann et al. 2024) (1) and (2) (k * (n – k)2) (1) GD = n {i=1}widi / n {i=1}wi (2) In the PAM Algorithm with Gower distance, we take n = number of points, and k = number of clusters In the Gower Distance, di is the distance between the i‑th sam‑ ples and wi is the weight given to the i‑th distance k‑means (Mai‑ mon and Rokach 2005) (3) minSE (µi)=minS K I=1xj∈Si �xj −µi� 2 Given a set of observations (x1, x2, …, xn), where each observation is a d‑dimensional real vector, k‑means constructs a partition of the observations into k sets (k ≤ n) to minimize the sum of squares within each set S = {S1, S2, …, Sk}; where µi is the mean of points in Si Hierarchical Model (Kauff‑ mann et al. 2024) (4) ( C i ,C J )=min xl∈Ci;xm∈Cj d X j ,X m ,l=1, ...,x i ;m=1, ...,n j The distance or similarity between two clusters is given, respectively, by the minimum distance (or maximum similarity) between their components. Thus, if after performing the K‑th step, we have already formed n‑K Clusters, the distance between the Clusters Ci (with ni elements) and Cj (with nj elements)
Page 9 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 transactions is low. This determines whether there is a need to implement another model or not. The generated clusters contain centroids, which include a set of average characteristics of each group, but to implement rules based on these groups to prevent fraud, it is necessary to focus on the group ranges, i.e., the minimum and maximum of the observations, conditional on an AND (Sadgali etal. 2021; Sinanc etal. 2021; Loukili etal. 2024). The rules that are proposed are direct rejection rules or rules that send the transactions to be validated by some analyst, and the analyst can approve or reject the transaction, depending on the validation of the bank card owner’s data (Srivastava etal. 2008; Bhatia etal. 2016; Prusti etal. 2021; Xie etal. 2021; Sadgali etal. 2021; Sinanc etal. 2021; Singh etal. 2021; Loukili etal. 2024). The objective is also to predict newly approved transactions, such that they are classified as fraudulent or nonfraudulent, along with their probability of the event occurring. This measure aims to alert merchants to transactions that have low, medium and high probabilities of fraud. Thus, after the validation of this information, merchants are able to determine whether it is necessary to stop the merchandise. Table4 shows the classification models. The seven classifiers balance (i) different learning biases—linear probabilistic (BLM, NB), distance-based (SVM), ensemble (RF), tree (DT-Gini), and nonlinear universal approximators (MLPs)—and (ii) varying sensitivities to class imbalance and feature-scale heterogeneity (Hastie etal. 2009). This diversity allows us to benchmark high-capacity models against simpler, auditable baselines under the same cost function C (Sect."Threshold optimization for rule selection"), satisfying the regulatory requirement for explainability in Latin American payment networks. The preventive measure is applied to reject recent transactions with fraud patterns as they pass through the payment gateway. This measure is preventive, as it is applied before the transactions pass through the bank. Cluster analysis allows the identification of the characteristics of fraudulent transactions to implement rules to reject them, thus preventing fraud. An unsupervised model has been applied, as the statistical objective is to determine rules that reflect the behavior of fraudulent transactions without affecting (rejecting) genuine transactions. These rules must be evaluated to determine whether the patterns are obtained clearly or, to a large extent, identify fraud. In this case, several classification models are applied, and the best model is chosen from among those listed in Table3. The evaluation of the models is carried out considering the comparison of actual versus predicted data: cross-tabulation, accuracy, error rate, sensitivity, specificity, positive accuracy and negative accuracy. The choice of a fraud detection model depends on several factors: the nature of the data, the complexity of the problem and the available resources. The MLP is particularly suitable for solving classification and regression problems where the relationships between input variables are nonlinear and complex (Taud and Mas 2017). It is widely used in tasks such as pattern recognition, image classification and prediction in various domains. It has two advantages: 1. Compared with recurrent models, the MLP structure is simpler and less computationally expensive. This facilitates its implementation and use in a wide range of applications (Taud and Mas 2017). 2. MLPs require less training time than recurrent neural networks do because of their simpler architecture. This is beneficial when working with large volumes of data, and results need to be obtained quickly. However, as a limitation, MLPs do not efficiently
Page 16 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 These are used as profiles for manual validation, since if automatic rejection rules were generated, there would be a risk of blocking up to 10% of the genuine transactions. In these cases, the procedure requires an analyst to contact the customer to verify the legitimacy of the purchase. The transaction will be rejected only if the analyst confirms that it is an attempt at fraud. Finally, the remaining clusters are not used in rule generation due to their low risk, either because they maintain fraud rates equal to or below 14% or because they have a considerably high transaction volume, suggesting more stable and reliable behavior. The combination of 25% fraud/ < 300 transactions minimizes the expected cost (USD 6,240) according to Table16. This rule blocks 82 fraudulent transactions (recall ≈ 72%) and sends only 31 genuine transactions for manual review (0.6% of the total), representing a 12% improvement over the next best alternative (20%, < 300). For medium-sized businesses, the results are shown in Tables7 and 8. Fifty-two variables were left from the variable selection, but for the k-means (Ali etal. 2020) and hierarchical models, categorical variables were removed, leaving 45 numerical variables. The variables remaining in the model can be seen in the dataset. The variables that remain in the model can be observed in the dataset. According to the stipulated models, in the case of medium-sized businesses, specific hyperparameters were established for each of the clustering algorithms used to capture the underlying structure of the data more accurately. For the PAM algorithm, 19 clusters were defined, and the Gower distance was used as a measure of dissimilarity, which is suitable for mixed data. The maximum number of iterations was set at 300, seeking a balance between accuracy and computational efficiency. In the case of K-means, the data were divided into 15 clusters, and the k-means + + initialization method was applied to improve the quality of the initial centroids. Similarly, a limit of 300 iterations was established, and a value of 42 for the random seed was set, ensuring the reproducibility of the results. The agglomerative hierarchical clustering model was configured with an average linkage criterion, with the Euclidean distance used as the proximity metric. The process was stopped when a cutoff threshold was reached, resulting in 24 clusters, thus allowing for a more detailed and nuanced segmentation of the set of medium-sized businesses. Table7 shows how by applying the clustering methods, the PAM responded better to the business need, as it allows better rules to be established to stop fraud, by means of the clusters that have a greater silhouette, using the numerical and categorical explanatory variables. However, it is necessary for the responsible analysts to review the data to determine if there is any additional pattern since the model only stops 63% of fraud; it would be possible to go into detail if there is any pattern of fraud by mail, address, telephone numbers, IPs, city, bin, issuing bank or description of the products purchased. Table 7 Performance of the clustering models for medium‑sized companies Source: Own elaboration Models # Cluster # Cluster rule % Fraud % Genuine VM REJ PAM 19 8 63.2% 0.0% 9.8% 0.3% Kmeans 15 7 57.1% 0.0% 7.1% 0.0% Hierarchical 24 6 40.7% 0.0% 1.8% 0.0%
Page 17 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Table 8 Distribution of fraud and genuine transactions by cluster (K‑means) in medium‑sized businesses Source: Own elaboration Clusters 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 No Fraud 263 101 540 129 97 2.415 2.031 207 442 1.331 15 41 14 42 370 51 553 221 Fraud 29 3 45 6 65 47 27 56 2 17 22 16 27 28 16 4 ‑ ‑ ‑ Total 292 104 585 135 162 2.462 2.058 263 444 1.348 37 57 41 28 58 374 51 553 221 % Fraud 10% 3% 8% 4% 40% 2% 1% 21% 0% 1% 59% 28% 66% 100% 28% 1% 0% 0% 0%
Page 18 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Table8 shows the total number of transactions carried out by medium-sized merchants in each cluster. For each cluster with ≥ 40% fraud, a profile was constructed on the basis of (i) the five numerical variables with the highest Z score relative to the total, (ii) the three categorical variables with the highest information value (IV), and (iii) the dominant business segment. The resulting profiles guide the drafting of specific rules. These results can be found in the Appendix A-Statistics dataset (and online supplementary material). To identify those groups at highest risk, a cluster was considered suitable for rule generation if it had a fraud rate above 10% and a transaction volume below 300. Under this approach, clusters 12, 14, and 11 stand out for having the highest fraud rates—59%, 66%, and 59%, respectively—and fewer than 100 transactions each. Given these characteristics, they are classified as high-risk cases and are used to generate automatic rejection rules, which allow associated transactions to be blocked without requiring additional verification. On the other hand, clusters 5, 8, and 13 have intermediate fraud rates, ranging from 21 to 40%, and meet the low-volume criterion (≤ 300 transactions). In these cases, we choose to generate manual validation profiles since applying rejection rules could lead to the rejection of up to 7% of the genuine transactions. The validation process involves the intervention of an analyst, who is responsible for contacting the customer to verify the authenticity of the transaction. Only if it is confirmed that it is an attempt at fraud is the transaction rejected. Finally, the remaining clusters do not exceed the 10% fraud threshold and are therefore not considered suitable for the generation of specific rules, given their low relative risk level or high transaction volume. The 10% fraud threshold/ < 300 transactions offer the lowest expected cost (USD 11,370). According to Table16, 144 frauds are intercepted (recall ≈ 71%), and 137 legitimate transactions (1.1%) are referred for validation, achieving a savings of 9% over the option (15%, < 300). The results for large businesses are shown in Tables9 and 10. From the selection of variables, 41 variables remained, but for the k-means and hierarchical models, categorical variables were eliminated, leaving 34 numerical variables. The variables that remain in the model can be seen in the dataset. According to the stipulated models, for the large business group, the hyperparameters of each clustering algorithm were redefined with the aim of achieving accurate and meaningful data segmentation. In the case of the PAM algorithm, the creation of 53 clusters was established, using the Gower distance as a measure of similarity, which is ideal for handling different types of variables. A maximum of 300 iterations was set, allowing for adequate model convergence without compromising efficiency. For the K-means algorithm, a partition into 31 clusters was determined, with initialization via the k-means + + method, which is known for improving the quality of centroids from the start of the process. Table 9 Performance of clustering models in large businesses Source: Own elaboration Models # Cluster # Cluster rule % Fraud % Genuine VM REJ PAM 53 16 47.8% 4.7% 1.0% 0.0% Kmeans 31 5 15.4% 2.4% 0.3% 0.0% Hierarchical 30 3 5.2% 0.4% 0.1% 0.0%
Page 19 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 The limit of 300 iterations was maintained, and a random seed value of 42 was used, ensuring the reproducibility of the results obtained. Finally, the agglomerative hierarchical clustering model was configured with an average linkage criterion, with the Euclidean distance used as the metric. The dendrogram was cut at the level that allowed 30 clusters to be identified, providing a more granular segmentation in line with the complexity of the large businesses analyzed. Table8 shows how, after applying the clustering methods, the PAM responded better to the business need, as it allows better rules to be established to stop fraud, by means of the clusters that Table 10 Distribution of fraud and genuine transactions by cluster (K‑means) in large businesses Source: Own elaboration No Fraud Fraud Total % Fraud No Fraud Fraud Total % Fraud Clusters 1 175 41 216 19% Clusters 28 2.588 24 2.612 1% 2 159 24 183 13% 29 1.502 12 1.514 1% 3 17 10 27 37% 30 1.927 18 1.945 1% 4 320 23 343 7% 31 928 2 930 0% 5 248 3 251 1% 32 2.175 6 2.181 0% 6 148 8 156 5% 33 3.452 15 3.467 0% 7 79 21 100 21% 34 515 1 516 0% 8 703 44 747 6% 35 1.985 13 1.998 1% 9 762 44 806 5% 36 1.204 6 1.210 0% 10 325 4 329 1% 37 1.259 9 1.268 1% 11 241 20 261 8% 38 219 5 224 2% 12 275 29 304 10% 39 149 4 153 3% 13 1.007 2 1.009 0% 40 1.869 1 1.870 0% 14 469 83 552 15% 41 238 17 255 7% 15 296 55 351 16% 42 678 70 748 9% 16 179 51 230 22% 43 948 28 976 3% 17 8 7 15 47% 44 2.288 52 2.340 2% 18 120 46 166 28% 45 5.821 20 5.841 0% 19 156 31 187 17% 46 2.831 14 2.845 0% 20 293 37 330 11% 47 14 12 26 46% 21 20 16 36 44% 48 412 4 416 1% 22 1.399 14 1.413 1% 49 250 6 256 2% 23 2.490 9 2.499 0% 50 1.026 – 1.026 0% 24 207 32 239 13% 51 1.676 – 1.676 0% 25 104 27 131 21% 52 1.861 – 1.861 0% 26 1.671 27 1.698 2% 53 82 – 82 0% 27 3.675 45 3.720 1%
Page 20 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 have a larger silhouette, using the numerical and categorical explanatory variables. However, it is necessary for the analysts in charge to review the group of large businesses and determine whether it is necessary to further divide the group of large businesses, as there may be businesses that are very atypical or much larger than usual. In addition, it is necessary to review the data to determine whether there is any additional pattern, as the model stops only 48% of fraud; we could go into detail if there is any pattern of fraud by mail, address, telephone number, IP, city, bin, issuing bank or description of the products that are purchased. Table10 shows the total number of transactions for large businesses in each cluster. For each cluster with ≥ 40% fraud, a profile was constructed on the basis of (i) the five numerical variables with the highest Z score relative to the total, (ii) the three categorical variables with the highest information value (IV), and (iii) the dominant business segment. The resulting profiles guide the drafting of specific rules. These results can be found in the Appendix A-Statistics dataset (and online supplementary material). If the clusters have a fraud rate higher than 10%, then the cluster serves to generate rules. The clusters that are highlighted in gray are those that are used to generate profile rules for manual validation, as 7% of the genuine transactions would be rejected if rejected rules were created. Clusters such as 3, 7, 14, 15, 16, 17, 18, 19, 20, 21, 25 and 47 have fraud rates above 10%. These clusters are crucial, as they indicate segments where fraud is significantly high. Cluster 17, with a fraud rate of 47%, is the most critical, followed closely by cluster 47 with 46%. These clusters are considered a priority for the generation of automatic fraud detection rules because of the high concentration of fraudulent transactions. On the other hand, numerous clusters have fraud rates of less than 1%, suggesting a low incidence of fraud in these segments. Clusters such as clusters 13, 32, 33, 34, 40, 45, 46, 50, 51, 52 and 53 have practically zero fraudulent transactions, indicating that these clusters are less problematic and do not require as much attention in terms of fraud. Clusters with fraud rates between 5 and 10%, such as 4, 6, 8, 9, 11, 12, 24, 26, 26, 38, 41, 41, 42, 43, 44, 48, and 49%, could be candidates for manual validation. These clusters, while not exhibiting extremely high fraud rates, still have a sufficient incidence of fraud to warrant manual transaction reviews. For example, cluster 12, with a 10% fraud rate, could be used to generate profiling rules to aid in manual validation, ensuring that genuine transactions are not indiscriminately rejected. Therefore, clusters with fraud rates above 10% stand out as critical for the implementation of automatic fraud detection rules. These clusters concentrate the majority of fraudulent transactions; therefore, rules generated from these data will be critical to effectively reduce fraud. Additionally, clusters with moderate fraud rates between 5 and 10% are equally important for the generation of profiling and manual validation rules, helping to balance fraud detection without negatively affecting legitimate transactions. Applying a 10% fraud threshold with no cluster size limit reduces the cost to USD 54,980, the minimum value recorded in Table16. The rule stops 2,780 frauds (recall ≈ 78%) and sends 298 genuine transactions (0.3%) for review, increasing the cost by 7% compared with the scenario (15%, < 500).
Page 21 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Fraud alert measure To evaluate the possible fraud alert measure, we start from the variables included in the study after the statistical tests (forced inclusion variables), which can be consulted in the dataset generated by the authors. The results show the performance of several machine learning models in detecting fraud in small, medium and large business transactions. The models evaluated include an MLP neural network, a random forest, a binomial logistic, decision tree-entropy, a naive Bayes-Bernoulli, a naive Bayes-multinomial and an SVM. The evaluation metrics include accuracy, the error rate, sensitivity, specificity, positive precision and negative precision. First, the results for small businesses are shown in Table11. Table11 details the performance of machine learning models in detecting fraud in small business transactions. The MLP neural network shows high performance, with an accuracy of 90%. Its high sensitivity and negative accuracy suggest that it is effective in detecting both fraudulent and genuine transactions. The random forest also shows high overall performance, although it is slightly lower in specificity. Compared with the MLP and random forest algorithms, the logistic binomial algorithm has a balanced but inferior performance, being less accurate in detecting fraudulent and genuine transactions. The decision-entropy tree model, on the other hand, has good sensitivity but lower specificity and positive accuracy, indicating more false positives than other models do. Compared with the superior models, the naive Bayes–Bernoulli model has moderate performance, with a higher error rate and a balanced but lower sensitivity and specificity. The naive Bayes-multinomial model presents similar results to those of the naive Bayes-Bernoulli model but with even lower sensitivity, indicating that fraud detection is more difficult than genuine transactions are. Finally, the SVM has high specificity, and its low sensitivity indicates that the model is not as effective in fraud detection, resulting in many false negatives. Figure1 shows the selection of the MLP neural network. This matrix confirms that the MLP neural network is quite effective, with many correct predictions for both fraudulent and nonfraudulent cases. However, some false negatives and false positives need to be managed. The additional metrics in Table14 support the same ranking, underscoring that the final model choice is theory-aligned (not merely the result of raw accuracy). Comparative Table 11 Small Businesses—Model Selection—Evaluation Source: Own elaboration Model Accuracy (%) Error rate (%) Sensitivity (%) Specificity (%) Positive Accuracy (%) Negative Accuracy (%) Neural Network‑ MLP 90 12 91 88 89 91 Random Forest 89 13 91 87 88 91 Binomial Logistic 85 15 86 85 86 85 Decision tree‑ Entropy 85 18 87 82 83 86 Naive bayes‑ Bernoulli 78 20 77 79 79 77 Naive bayes‑ Multinomial 77 20 74 79 78 75 SVM 72 12 57 88 83 67
Page 22 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 analysis of the models reveals that the MLP neural network and random forest are the most effective for small business fraud detection because of their high accuracy, sensitivity and negative precision. The MLP neural network confusion matrix highlights its overall efficiency in correctly classifying transactions, although there is still room for improvement in reducing false positives and false negatives. Second, the results for small businesses are shown in Table12 and Fig.2. As seen in Table12, the random forest method achieves high performance with excellent specificity and positive accuracy, indicating a strong ability to identify genuine and fraudulent transactions accurately. The Binomial Logistic model shows a similar result to the Random Forest model, as this model also shows high performance with excellent specificity and positive accuracy, although with a slightly lower sensitivity. The Gini decision tree has good specificity and positive accuracy but a higher error rate than the random forest and logistic binomial methods do. The MLP neural network model has good sensitivity, but its error rate is higher than that of the other leading models, which may affect its overall effectiveness in certain scenarios. The Fig. 1 MLP Neural Network—Confusion Matrix. Source: Own elaboration Table 12 Medium‑sized businesses—Model selection—Evaluation Source: Own elaboration Model Accuracy (%) Error rate (%) Sensitivity (%) Specificity (%) Positive accuracy (%) Negative accuracy (%) Random Forest 89 5 81 95 93 86 Binomial Logistic 88 5 79 95 93 84 Decision tree‑Gini 86 10 80 91 88 85 Neural Network‑ MLP 86 12 83 89 87 86 SVM 79 6 62 94 90 75 Naive bayes‑ Bernoulli 73 21 64 81 73 73
Page 23 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 SVM model has excellent specificity and positive accuracy, but its low sensitivity indicates difficulties in fraud detection, resulting in many false negatives. Finally, the naive Bayes–Bernoulli model has the lowest performance among those evaluated, with a high error rate and low sensitivity and accuracy. The random forest model and the binomial logistic model are the ones with the highest accuracy, but the logistic model actually predicts "no fraud" better and "fraud" only 79% correctly, which is not that it is completely bad, but other models predict fraud better and still have an accuracy above 80%, such as the MLP neural network; thus, for the purpose of the analysis, it is better to select between the random forest or the neural network. Considering the sensitivity of the fraud cases, 81% were correctly predicted with the random forest, and 83% of the fraud cases were correctly predicted with the MLP neural network. Figure2 shows the selection of the MLP neural network. The confusion matrix indicates that the MLP neural network is quite effective in classifying fraudulent and genuine transactions, although there are a significant number of false negatives and false positives. Finally, for large businesses, the results are shown in Table13 and Fig.3. As shown in Table13, the random forest model shows good overall performance with high specificity and negative accuracy, indicating a strong ability to identify genuine and fraudulent transactions. The MLP neural network model, on the other hand, stands out for its high sensitivity and negative accuracy, although it has a higher error rate than the random forest. The decision tree-Gini has a balanced performance with good specificity and positive precision, although its accuracy is lower than that of the random forest and MLP. On the other hand, the Binomial Logistic model presents similar results to the Decision Tree-Gini model; since this model has good specificity and positive precision, its accuracy is lower than that of the superior models. The SVM model shows excellent specificity and positive accuracy, but its low sensitivity indicates difficulties in fraud detection, resulting in many false negatives. Finally, the Fig. 2 MLP Neural Network—Confusion matrix. Source: Own elaboration
Page 24 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 naive Bayes–Bernoulli model has the lowest performance among those evaluated, with a high error rate and low sensitivity and accuracy. Therefore, the MLP neural network and random forest are the best models, as 85% and 86%, respectively, of the data were predicted correctly; on the other hand, without forgetting, the objective of this prediction is more directed toward the prediction of fraud, of course, without leaving aside nonfraud. In terms of the sensitivity of the Fraud cases, 87% and 85% (respectively) were correctly predicted by both models. Figure 3 shows the selection of the MLP neural network. The confusion matrix indicates that the MLP neural network is effective in classifying fraudulent and genuine transactions, although there are a significant number of false negatives and false positives. The application of advanced data analytics and machine learning techniques in credit card fraud detection not only improves the accuracy and efficiency of these detections but also provides a robust framework for informed financial and economic decision making, strengthening the competitiveness and security of merchants in Latin America. Table 13 Large business model selection and evaluation Source: Own elaboration Model Accuracy (%) Error rate (%) Sensitivity (%) Specificity (%) Positive accuracy (%) Negative accuracy (%) Random Forest 86 15 85 87 84 88 Neural network‑ MLP 85 19 87 83 80 89 Decision tree‑Gini 82 17 80 85 81 84 Binomial Logistic 82 17 79 85 81 83 SVM 78 11 62 90 84 75 Naive bayes‑ Bernoulli 55 14 15 87 49 56 Fig. 3 MLP Neural Network—Confusion matrix. Source: Own elaboration
Page 25 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 The results of this study on the detection and prevention of fraud in credit card transactions provide valuable information from a financial and economic perspective. The data mining and machine learning techniques implemented in this study not only seek to improve accuracy in identifying fraudulent transactions but also provide detailed and specific analysis according to merchant size (small, medium and large). This allows for better risk management and resource optimization, reducing economic losses for both merchants and consumers. For example, for small merchants, the k-means model enabled the establishment of rules that could have prevented 85% of fraudulent transactions by 2021 without significantly affecting legitimate transactions. For medium-sized merchants, the PAM model with the Gower distance showed an effectiveness of 63% in preventing fraud. For large merchants, the same model stopped 48% of fraudulent transactions. These figures highlight the relevance of adjusting fraud detection strategies to the specific characteristics of each type of merchant, thus optimizing financial security and reducing the negative impact on legitimate transactions. In addition, the analysis revealed that the implementations of neural network (MLP) models and random forest (Random Forest) models are highly effective in classifying fraudulent and authentic transactions. These models not only improve fraud detection but also minimize false positives and negatives, ensuring greater accuracy and efficiency in financial risk management. Discussion ofresults This paper has analyzed the performance of data mining, machine learning and artificial intelligence techniques for detecting fraudulent credit card transactions. The implementation of neural network models for detecting fraudulent transactions has been recurrently adopted in recent years in several studies aimed at detecting fraudulent transactions in the financial sector (Ngai etal. 2011; Badal-Valero and García-Cárceles, 2020; among others). Predictive models for detecting fraudulent credit card transactions are currently used recurrently. The methodological approach adopted in this research, which is based on the application of data mining techniques and, in particular, neural networks, for the detection of credit card fraud cases, has been previously validated by other works (Aleskerov etal. 1997; Brause etal. 1999; Kou etal. 2004; Kumari and Mishra 2018; Varmedja etal. 2019; among others). In the case of this research, the novelty lies in classifying shops into three groups (small, medium and large) and in identifying the most appropriate method to detect fraud on the basis of the size of the shop. Thus, the best model for detecting fraud in small businesses would be k-means, which allows the creation of two automatic rejection rules of the anti-fraud model and six manual validation rules, the application of which would have stopped 85% of the fraud generated in 2021 without affecting genuine transactions. This result coincides with that obtained by Gowda (2021) and Jain etal. (2022), who indicate that K-means can be effective for the creation of automatic rejection rules and manual validation in small businesses to avoid fraud. According to the results of our study, for medium-sized businesses, the PAM model with the Gower distance stands out significantly. This model enables the creation of an automatic anti-fraud model rejection rule and seven manual validation rules. The application of these rules would have stopped 63.2% of fraudulent transactions in 2021
Page 32 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Kauffmann J, Esders M, Ruff L, Montavon G, Samek W, Müller K‑R (2024) From clustering to cluster explanations via neural networks. IEEE Trans Neural Netw Learn Syst 35(2):1926–1940. https:// doi. org/ 10. 1109/ tnnls. 2022. 31859 01 Kaufman L, Rousseeuw PJ (1990) Finding groups in data. Wiley. https:// doi. org/ 10. 1002/ 97804 70316 801 Kaur S, Singh KD, Singh P, Kaur R (2021) Ensemble model to predict credit card fraud detection using random forest and generative adversarial networks. Advances in Intelligent Systems and Computing. Springer Nature Singapore, Singapore, pp 87–97. https:// doi. org/ 10. 1007/ 978‑ 981‑ 33‑ 4367‑2_ 10 Khalid AR, Hashim MS, Hossain A (2024) Enhancing credit card fraud detection using deep learning and explainable AI models: a European perspective. J Financ Anal 18(1):45–60. https:// doi. org/ 10. 1016/j. jfa. 2024. 03. 015 Kilay AL, Simamora BH, Putra DP (2022) The influence of e‑payment and e‑commerce services on supply chain performance: implications of open innovation and solutions for the digitalization of micro, small, and medium enterprises (MSMEs) in Indonesia. J Open Innov Technol Market Complex 8(3):119. https:// doi. org/ 10. 3390/ joitm c8030 119 Kim S‑I, Kim S‑H (2020) E‑commerce payment model using blockchain. J Ambient Intell Humaniz Comput 13(3):1673–1685. https:// doi. org/ 10. 1007/ s12652‑ 020‑ 02519‑5 King ST, Scaife N, Traynor P, Abi Din Z, Peeters C, Venugopala H (2021) Credit card fraud is a computer security problem. IEEE Secur Privacy 19(2):65–69. https:// doi. org/ 10. 1109/ msec. 2021. 30502 47 Kohavi R, Parekh R (2004) Visualizing RFM segmentation. In: proceedings of the 2004 SIAM international conference on data mining, pp 391–399. https:// doi. org/ 10. 1137/1. 97816 11972 740. 36 Kotios D, Makridis G, Fatouros G, Kyriazis D (2022) Deep learning enhancing banking services: a hybrid transaction classifica‑ tion and cash flow prediction approach. J Big Data. https:// doi. org/ 10. 1186/ s40537‑ 022‑ 00651‑x Kou C‑T, Sirwongwattana S, Huang Y‑P (2004) Survey of fraud detection techniques. IEEE Int Conf Netw Sens Control 2004(2):749–754. https:// doi. org/ 10. 1109/ icnsc. 2004. 12970 40 Krishna Rao NV, Harika Devi Y, Shalini N, Harika A, Divyavani V, Mangathayaru N (2021) Credit card fraud detection using spark and machine learning techniques. Algorithms for Intelligent Systems. Springer Singapore, Singapore, pp 163–172. https:// doi. org/ 10. 1007/ 978‑ 981‑ 33‑ 4046‑6_ 16 Kumari P, Mishra SP (2018) Analysis of credit card fraud detection using fusion classifiers. Advances in Intelligent Systems and Computing. Springer Singapore, Singapore, pp 111–122. https:// doi. org/ 10. 1007/ 978‑ 981‑ 10‑ 8055‑5_ 11 Lei Y‑T, Ma C‑Q, Ren Y‑S, Chen X‑Q, Narayan S, Huynh ANQ (2023) A distributed deep neural network model for credit card fraud detection. Finance Res Lett 58:104547. https:// doi. org/ 10. 1016/j. frl. 2023. 104547 Lokanan ME (2022) Financial fraud detection: the use of visualization techniques in credit card fraud and money laundering domains. J Money Laundering Control 26(3):436–444. https:// doi. org/ 10. 1108/ jmlc‑ 04‑ 2022‑ 0058 Loukili M, Messaoudi F, El Youbi R (2024) Enhancing financial transaction security a deep learning approach for E‑payment fraud detection. Internet of Things and Big Data analytics for a green environment. CRC Press, p 15 MacQueen J (1967) Some methods for classification and analysis of multivariate observations. In: proceedings of the fifth berkeley symposium on mathematical statistics and probability (pp. 281–297). Univ of California Press Maimon O, Rokach L (2005) Decomposition methodology for knowledge discovery and data mining. Data Mining and Knowledge Discovery Handbook. Springer‑Verlag, pp 981–1003. https:// doi. org/ 10. 1007/0‑ 387‑ 25465‑x_ 46 Mary AJ, Claret SPA (2023) Design and development of big data‑based model for detecting fraud in healthcare insurance industry. Soft Comput 27(12):8357–8369. https:// doi. org/ 10. 1007/ s00500‑ 023‑ 08296‑5 Mekterović I, Karan M, Pintar D, Brkić L (2021) Credit card fraud detection in card‑not‑present transactions: Where to invest? Appl Sci 11(15):6766. https:// doi. org/ 10. 3390/ app11 156766 Mienye ID, Chikwerem PC, Zhang J (2024) A hybrid deep learning approach for detecting anomalies in credit card transac‑ tions. J Comput Finance 20(3):115–130. https:// doi. org/ 10. 1080/ 10824 120. 2024. 038765 Miroshnychenko I, Aleksieieva V (2022) Use of RFM analysis in customer segmentation. Ekon Derzh 1:114. https:// doi. org/ 10. 32702/ 2306‑ 6806. 2022.1. 114 Murtagh F, Contreras P (2011) Algorithms for hierarchical clustering: an overview. Wires Data Min Knowl Discov 2(1):86–97. https:// doi. org/ 10. 1002/ widm. 53 Nama FA, Obaid AJ, Alrammahi AAH (2023) Credit card fraud detection and classification using deep learning with support vector machine techniques. Lecture Notes in Networks and Systems. Springer Nature Singapore, Berlin, pp 399–413. https:// doi. org/ 10. 1007/ 978‑ 981‑ 99‑ 6553‑3_ 31 Náñez Alonso SL, Jorge‑Vázquez J, Echarte Fernández MÁ, Sanz‑Bas D (2024) Bitcoin’s bubbly behaviors: does it resemble other financial bubbles of the past? Humanit Soc Sci Commun. https:// doi. org/ 10. 1057/ s41599‑ 024‑ 03220‑0 Náñez Alonso SL, Ozili PK, Hernández BMS, Pacheco LM (2025) Evaluating the acceptance of CBDCs: experimental research with artificial intelligence (AI) generated synthetic response. Quant Finance Econ 9(1):242–273. https:// doi. org/ 10. 3934/ qfe. 20250 08 Neubauer T, Peres S (2017) Data mining using naive bayes classifier: an application in short news. Anais Do Simpósio Brasileiro de Sistemas de Informação (SBSI), 543–546. https:// doi. org/ 10. 5753/ sbsi. 2017. 6096 Ngai EWT, Hu Y, Wong YH, Chen Y, Sun X (2011) The application of data mining techniques in financial fraud detection: a clas‑ sification framework and an academic review of literature. Decis Support Syst 50(3):559–569. https:// doi. org/ 10. 1016/j. dss. 2010. 08. 006 Ogbanufe O, Kim DJ (2018) Comparing fingerprint‑based biometrics authentication versus traditional authentication meth‑ ods for e‑payment. Decis Support Syst 106:1–14. https:// doi. org/ 10. 1016/j. dss. 2017. 11. 003 Paldino GM, Zamboni L (2024) The role of diversity in credit card fraud detection models: a comparative study. Eur J Financ Innov 13(6):32–49. https:// doi. org/ 10. 1080/ 14061 994. 2024. 039841 Prusti D, Das D, Rath SK (2021) Credit card fraud detection technique by applying graph database model. Arab J Sci Eng 46(9):1–20. https:// doi. org/ 10. 1007/ s13369‑ 021‑ 05682‑9 Rajeshwari B, Mallikarjuna Reddy Doodipala KKKV, P. Pattabhi Ram JRN (2024) Understanding consumer behavior in the retail sector using RFM segmentation and machine learning: an analysis. Eur Econ Lett (EEL) 14(3):412–420. https:// doi. org/ 10. 52783/ eel. v14i3. 1783 Reddy B, Kumar S (2024) Effective fraud detection using hybrid ensemble models. Int J Financ Cybersecur 9(4):205–220. https:// doi. org/ 10. 1002/ jfc. 2024. 092
Page 33 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Rtayli N, Enneya N (2020) Enhanced credit card fraud detection based on SVM‑recursive feature elimination and hyper‑ parameters optimization. J Inform Secur Appl 55:102596. https:// doi. org/ 10. 1016/j. jisa. 2020. 102596 Rungruang C, Riyapan P, Intarasit A, Chuarkham K, Muangprathub J (2023) Rfm model customer segmentation based on hierarchical approach using fca. Elsevier BV. https:// doi. org/ 10. 2139/ ssrn. 44162 18 Sadgali I, Sael N, Benabbou F (2021) Human behavior scoring in credit card fraud detection. IAES Int J Artif Intell (IJ‑AI) 10(3):698. https:// doi. org/ 10. 11591/ ijai. v10. i3. pp698‑ 706 Saha D, Young TM, Thacker J (2023) Predicting firm performance and size using machine learning with a Bayesian perspec‑ tive. Mach Learn Appl 11:100453. https:// doi. org/ 10. 1016/j. mlwa. 2023. 100453 Salam A, Liu C, Rahman M (2024) Federated learning model for credit card fraud detection: privacy‑preserving analytics. Artif Intell Financ Appl 11(2):122–134. https:// doi. org/ 10. 1109/ AIFA. 2024. 023456 Sanz‑Bas D, del Rosal C, Náñez Alonso SL, Echarte Fernández MÁ (2021) Cryptocurrencies and fraudulent transactions: risks, practices, and legislation for their prevention in Europe and Spain. Laws 10(3):57. https:// doi. org/ 10. 3390/ laws1 00300 57 Sarker IH (2021) Machine learning: algorithms, real‑world applications and research directions. SN Comput Sci. https:// doi. org/ 10. 1007/ s42979‑ 021‑ 00592‑x Serrano‑Pérez JdeJ, Fernández‑Anaya G, Carrillo‑Moreno S, Yu W (2021) New results for prediction of chaotic systems using deep recurrent neural networks. Neural Process Lett 53(2):1579–1596. https:// doi. org/ 10. 1007/ s11063‑ 021‑ 10466‑1 Shewalkar A, Nyavanandi D, Ludwig SA (2019) Performance evaluation of deep neural networks applied to speech recogni‑ tion: RNN, LSTM and GRU. J Artif Intell Soft Comput Res 9(4):235–245. https:// doi. org/ 10. 2478/ jaiscr‑ 2019‑ 0006 Siblini W, Coter G, Fabry R, He‑Guelton L, Oblé F, Lebichot B, Borgne YAL, Bontempi G (2021) Transfer learning for credit card fraud detection: a journey from research to production. arXiv.Org. https:// arxiv. org/ abs/ 2107. 09323 Sinanc D, Demirezen U, Sağıroğlu Ş (2021) Explainable credit card fraud detection with image conversion. ADCAIJ: Adv Distrib Comput Artif Intell J 10(1):63–76. https:// doi. org/ 10. 14201/ adcai j2021 10163 76 Singh A, Ranjan RK, Tiwari A (2021) Credit card fraud detection under extreme imbalanced data: a comparative study of data‑ level algorithms. J Exp Theor Artif Intell 34(4):571–598. https:// doi. org/ 10. 1080/ 09528 13x. 2021. 19077 95 Sivagama Sundhari S (2011) A knowledge discovery using decision tree by Gini coefficient. In: 2011 International conference on business, engineering and industrial applications, pp 232–235. https:// doi. org/ 10. 1109/ icbeia. 2011. 59942 50 Sorour SE, Ahmed N, Badawy MM (2024) Credit card fraud detection using explainable AI: challenges and solutions. IEEE Trans Financ Technol 7(2):98–112. https:// doi. org/ 10. 1109/ TFT. 2024. 040398 Srivastava A, Kundu A, Sural S, Majumdar AK (2008) Credit card fraud detection using hidden markov model. IEEE Trans Depend Secure Comput 5(1):37–48. https:// doi. org/ 10. 1109/ tdsc. 2007. 70228 Tanouz D, Subramanian RR, Eswar D, Reddy GVP, Kumar AR, Praneeth CVNM (2021) Credit card fraud detection using machine learning. In: 2021 5th international conference on intelligent computing and control systems (ICICCS), pp 967–972. https:// doi. org/ 10. 1109/ icicc s51141. 2021. 94323 08 Taud H, Mas JF (2017) Multilayer Perceptron (MLP). Geomatic Approaches for Modeling Land Change Scenarios. Springer International Publishing, Berlin, pp 451–455 Tiwari P, Mehta S, Sakhuja N, Kumar J, Singh AK (2021) Credit card fraud detection using machine learning: a study. arXiv.Org. https:// arxiv. org/ abs/ 2108. 10005 Tran LTT (2021) Managing the effectiveness of e‑commerce platforms in a pandemic. J Retail Consum Serv 58:102287. https:// doi. org/ 10. 1016/j. jretc onser. 2020. 102287 Tumminello M, Consiglio A, Vassallo P, Cesari R, Farabullini F (2022) Insurance fraud detection: a statistically validated network approach. J Risk Insur 90(2):381–419. https:// doi. org/ 10. 1111/ jori. 12415 Varmedja D, Karanovic M, Sladojevic S, Arsenovic M, Anderla A (2019) Credit Card fraud detection ‑ machine learning meth‑ ods. In: 2019 18th international symposium INFOTEH‑JAHORINA (INFOTEH), pp 1–5. https:// doi. org/ 10. 1109/ infot eh. 2019. 87177 66 Vera Loor RY (2020) Control Interno como herramienta antifraude para las organizaciones. Caleidoscopio De Las Ciencias Sociales. https:// doi. org/ 10. 38202/ calei dosco pio.1 Vergara D, Lampropoulos G, Antón‑Sancho Á, Fernández‑Arias P (2024) Impact of artificial intelligence on learning manage‑ ment systems: a bibliometric review. Multimodal Technol Interact 8(9):75. https:// doi. org/ 10. 3390/ mti80 90075 Vergara D, del Bosque A, Lampropoulos G, Fernández‑Arias P (2025) Trends and applications of Artificial Intelligence in Project Management. Electronics 14(4):800. https:// doi. org/ 10. 3390/ elect ronic s1404 0800 Viaene S, Dedene G, Derrig R (2005) Auto claim fraud detection using Bayesian learning neural networks. Expert Syst Appl 29(3):653–666. https:// doi. org/ 10. 1016/j. eswa. 2005. 04. 030 Vosseler A (2022) Unsupervised insurance fraud prediction based on anomaly detector ensembles. Risks 10(7):132. https:// doi. org/ 10. 3390/ risks 10070 132 Wang C (2023) Overview of digital finance anti‑fraud. Anti‑Fraud Engineering for Digital Finance. Springer Nature Singapore, Berlin, pp 1–10 Wang G, Xu J (2008) A decision tree algorithm based on coordination degree and gini‑coefficient. Int Conf MultiMedia Inf Technol 2008:822–824. https:// doi. org/ 10. 1109/ mmit. 2008. 208 Wijaya MG, Pinaringgi MF, Zakiyyah AY, Meiliana M (2024) Comparative analysis of machine learning algorithms and data balancing techniques for credit card fraud detection. Procedia Comput Sci 245:677–688. https:// doi. org/ 10. 1016/j. procs. 2024. 10. 294 Xie Y, Li A, Gao L, Liu Z (2021) A heterogeneous ensemble learning model based on data distribution for credit card fraud detection. Wirel Commun Mob Comput. https:// doi. org/ 10. 1155/ 2021/ 25312 10 Zhang C, Lu Y (2021) Study on artificial intelligence: the state of the art and future prospects. J Ind Inf Integr 23:100224. https:// doi. org/ 10. 1016/j. jii. 2021. 100224 Zhang J, Cai K, Wen J (2024) A survey of deep learning applications in cryptocurrency. iScience 27(1):108509. https:// doi. org/ 10. 1016/j. isci. 2023. 108509 Ziakis C, Vlachopoulou M (2023) Artificial intelligence in digital marketing: insights from a comprehensive review. Information 14(12):664. https:// doi. org/ 10. 3390/ info1 41206 64
Page 34 of 34 RugelesDiazetal. Financial Innovation (2025) 11:137 Zingerle A, Kronman L (2018) Internet crime and anti‑fraud activism. Adv Inf Secur Privacy Ethics. https:// doi. org/ 10. 4018/ 978‑1‑ 5225‑ 5583‑4. ch013 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.