scieee AI-readable full text Open interactive document viewer

Data mining approach in detecting inaccurate financial statements in government-owned enterprises

Gadžo, Amra,Suljić, Mirza,Jusufović, Adisa,Filipović, Slađana,Suljić, Erna

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Gadžo, Amra; Suljić, Mirza; Jusufović, Adisa; Filipović, Slađana; Suljić, Erna Article — Published Version Data mining approach in detecting inaccurate financial statements in government-owned enterprises Croatian operational research review Suggested Citation: Gadžo, Amra; Suljić, Mirza; Jusufović, Adisa; Filipović, Slađana; Suljić, Erna (2025) : Data mining approach in detecting inaccurate financial statements in government-owned enterprises, Croatian operational research review, ISSN 1848-9931, Croatian Society for Operations Research, Zagreb, Vol. 16, Iss. 1, pp. 1-15, https://doi.org/10.17535/crorr.2025.0001 This Version is available at: https://hdl.handle.net/10419/324921 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Croatian Operational Research Review 1 CRORR 16:1(2025), 1–15 Data mining approach in detecting inaccurate financial statements in government-owned enterprises Amra Gadžo1,∗, Mirza Suljić2, Adisa Jusufović1, Slađana Filipović1and Erna Suljić3 1Faculty of Economics, University of Tuzla, Urfeta Vejzagića 8, 75000 Tuzla, BiH E-mail: ⟨{amra.gadzo, adisa.jusufovic, sladjana.filipovic}@untz.ba⟩ 2Center for Quality Assurance and Internal Evaluation, University of Tuzla, Armije RBiH bb, 75000 Tuzla, BiH E-mail: ⟨[email protected]a⟩ 3Public Elementary School "Simin Han", Sarajac 4, 75207 Tuzla, BiH E-mail: ⟨erna.c[email protected]om⟩ Abstract. The study aims to assess the capability of various data mining techniques in detecting inaccurate financial statements of government-owned enterprises operating in the Federation of Bosnia and Herzegovina (FBiH). Inaccurate financial statements indicate potential financial fraud. Prediction models of four classification algorithms (J48, KNN, MLP, and BayesNet) were examined using a dataset comprising 200 audited financial statements from government-owned enterprises under the supervision of the Audit Office of the Institutions in the Federation of Bosnia and Herzegovina. The results obtained through data mining analysis reveal that a dataset encompassing seven balance sheet items provides the most comprehensive depiction of financial statement quality. These seven attributes are: opening entry of accounts receivable, profit (loss) at the end of the period, operating assets at the end of the period, accounts receivable at the end of the period, opening entry of operating assets, short term financial investments at the end of the period, and opening entry of short-term financial investments. By employing these seven attributes, the MLP algorithm was implemented to construct the most precise predictive model, achieving a 76% accurate classification rate for financial statements. Leveraging the identified attributes, a mathematical model could potentially be formulated to effectively predict financial statements of government-owned enterprises in FBiH. This, in turn, could considerably facilitate the process of selecting GOEs for inclusion in the annual work plan of state auditors. Presently, due to resource constraints, government-owned enterprises in FBiH do not undergo regular annual scrutiny by state auditors, with only 10 to 15 such enterprises being subject to audits each year. The results of this research can also be beneficial to both the public and the Financial Intelligence Agency in the FBiH. The paper contributes to filling the gap in the literature regarding the applied methodology, particularly in the part concerning the attributes used in the research. Keywords: data mining, financial statement frauds, government-owned enterprises, prediction of financial statements accuracy Received: April 20, 2024; accepted: October 4, 2024; available online: February 4, 2025 DOI: 10.17535/crorr.2025.0001 Original scientific paper. ∗Corresponding author. This is an open access article under the CC BY-NC-ND 4.0 license 1 http://www.hdoi.hr/crorr-journal ©2025 Copyright of this article is retained by the author(s). CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... 1. Introduction Transitional countries, including Bosnia and Herzegovina, often face problems such as an inefficient public sector, a high corruption index, and the employment of politically affiliated personnel in the management structure and supervisory boards of government-owned enterprises. The resources of these enterprises are often used to support political campaigns and favor certain suppliers in the procurement of goods and services, under the guise of the Public Procurement Law. The challenges faced by transitional countries like Bosnia and Herzegovina, including an inefficient public sector, high corruption index, and politically affiliated personnel in government-owned enterprises, have been widely documented [3]. These issues have eroded public trust in state institutions and led to financial embezzlement in government-owned enterprises. Anti-corruption policies and good governance practices have been implemented to address these challenges, but their effectiveness in rebuilding trust in public institutions remains a subject of debate [3]. The restructuring of state-owned enterprises, with a focus on corporate governance and the role of the state, has been proposed as a potential solution. However, the complex relationship between tax evasion, state capacity, and trust in transitional countries, as well as the prevalence of public procurement corruption in Bosnia and Herzegovina, further complicate the situation. Government-owned enterprises that are owned by the state or lower levels of government are frequently accused of financial embezzlement and do not enjoy public trust in the accuracy of their financial statements. Financial statement fraud involves the intentional concealment or omission of critical information resulting from a deliberate failure to report financial data in line with generally accepted accounting principles. Financial fraud is a serious problem worldwide and is particularly pronounced in companies that are state-owned in countries such as Bosnia and Herzegovina. Lalić, Jovičić & Bošnjaković [20] highlights the inflow of money from abroad and the presentation of operating losses as common examples of financial fraud in the region. Buljubašić Musanović & Halilbegović [2] further underscores the manipulation of financial statements in failing small and medium-sized enterprises, with significant differences in accruals, asset quality, leverage, profitability, and liquidity between failing and non-failing SMEs. Isaković-Kaplan et al. [13] explores the application of Benford’s Law in detecting potential earnings manipulation in income statements of economic entities, emphasizing the need for additional forensic investigations. Yadiati, Rezwiandhari & Ramdany [35] highlight a broader perspective by identifying factors such as financial stability, external pressure, industry nature, director changes, and collaboration with government projects as possible indicators of fraudulent financial reporting in state-owned enterprises. Another issue also arises from the fact that financial statements of government-owned enterprises are not subject to annual audits by state auditors. Due to the extensive number of institutions under state audit supervision (more than 2,000 institutions), only a few of government-owned enterprises are incorporated into the annual audit plan. The Audit Office of the Institutions in the Federation of Bosnia and Herzegovina does not utilize modern data mining tools to aid in the detection of inaccurate financial statements. In the process of planning which government-owned enterprises will be included in the audit plan for the year, state auditors are guided by media reports, anonymous complaints, and internal information from previous financial audits. The fundamental research question in this study is which balance sheet items and data mining technique provide the best probability of predicting inaccurate financial statements in government-owned enterprises? Therefore, the objectives of this paper are to identify the balance sheet items that best indicate inaccurate financial statements and to determine which data mining technique achieves the best results in predicting inaccuracies in financial statements for government-owned enterprises. Data mining techniques will be employed to detect balance sheet items from financial position reports of enterprises that offer the most accurate predictions of the state auditors’ assessments of financial statement quality. Furthermore, various data mining techniques will be utilized to evaluate their predictive efficacy through comparative analysis of forecasting re2 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... sults. This paper aims to contribute to addressing the literature gap concerning the attributes used to identify balance sheet items that best predict inaccuracies in financial statements of government-owned enterprises in FBiH. These objectives will be achieved through a systematic literature review, identifying gaps in the literature. Following this, a detailed explanation of the research methodology and the application of data mining techniques will follow. Finally, the research findings will be presented, accompanied by their explanations. 2. Literature review A range of studies have explored the use of various models and techniques to predict inaccurate financial statements based on audit opinion. Gadžo et al. [5] found that the Beneish M-score model, particularly its partial indicators, can accurately predict the quality of financial statements in public enterprises. Similarly, Sánchez-Serrano et al. [27] developed a model for predicting audit opinion in consolidated financial statements, achieving high accuracy. Wu & Li [33] further improved on this by using a BP neural network with Adam optimizer to predict audit opinions in listed companies, with a high accuracy rate. Yue, Shen & Chu [37] also identified specific financial ratios, such as net assets per share and earnings per share, as strong indicators of false financial affairs, on the basis of the model of logistic regression analysis. Research consistently shows that financial ratios are a key factor in predicting an auditor’s qualified opinion on financial statements [6]. These ratios, such as retained earnings to total assets, equity to total liabilities, and net income to total assets, are used in various models to accurately classify qualified and unqualified opinions. Rudkhani & Jabbari [6] found that only two financial ratios, "earnings per share" and "fixed asset turnover," were needed for an accuracy rate of 64.1% . Similarly, Gadžo [5] achieved a high accuracy rate of 98-100% using eight partial indicators from the Beneish M-score model. Sormunen [28] further noted that the classification ability of certain financial ratios may diminish over time, suggesting the need for ongoing evaluation. In our literature review, we did not identify studies that predict inaccurate financial statements based on actual values of balance sheet items at the beginning and end of the period (without using ratio indicators). That is the research gap that this scientific paper aims to fill. The basis for determining inaccurate financial statements in government-owned enterprises is the analysis of auditors’ findings and criticisms, as well as the written grounds for issuing opinions on the financial statements. According to our own research [5], the most common causes of irregularities in financial reporting of government-owned enterprises in FBiH include inadequate accounting estimates of accounts receivable from customers, inadequate valuation of fixed assets, inventory, and provisions. A range of methodologies have been proposed for predicting the auditor’s opinion on financial statements. Sánchez-Serrano et al. [27] and Stanišić, Radojević & Stanić [30] both highlight the use of artificial neural networks and machine learning algorithms. Data mining can be an extremely useful tool for detecting irregularities within a large volume of data in financial statements. Large amounts of data often contain hidden patterns and trends that may indicate potential fraudulent activities. Therefore, data mining is used to extract knowledge from vast data sets in order to identify behavioral patterns that may indicate fraud. The study conducted by evaluated the capability of various Data Mining classification methods to identify companies that released fraudulent financial statements (FFS), with a particular emphasis on recognizing the key factors linked to such fraud. The study explored the application of Decision Trees, Artificial Neural Networks (ANN), and Bayesian Belief Networks as tools for detecting fraudulent financial reporting. The results demonstrated that the Bayesian Belief Network model achieved the best performance, successfully classifying 90.3% of the validation sample in a 10-fold cross-validation procedure. According to the research conducted by [26], utilized data mining techniques, including Multilayer Feed Forward Neural Network (MLFF), Support Vector Machines (SVM), Genetic Programming (GP), Group Method of Data Handling (GMDH), Logistic Regression (LR), and Probabilistic Neural Network (PNN), to identify 3 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... companies involved in financial statement fraud. These techniques were tested on a dataset of 202 Chinese companies and compared with and without feature selection. Without feature selection, PNN outperformed all other techniques, while with feature selection, both GP and PNN achieved nearly equal accuracy, outperforming the others. According to the research conducted by [21] on financial statement fraud, utilizing data mining methods including logistic regression, decision trees (CARTC4.5. algoritam), and artificial neural networks (ANN), the results indicate that artificial neural networks and decision trees resulted in much more accurate classification compared to logistic regression. According to the research conducted by [29], the methods used in the fraud detection process were Linear Regression, Artificial Neural Networks (ANN), k-Nearest Neighbors (KNN), Support Vector Machines (SVM), Decision Stump, M5P Tree, Random Forest, and J48. The results from the experiments indicated that data mining methods were able to detect the fraud factors between the financial statements and the e-ledger. In this study, the Decision Stump Algorithm exhibited the best performance. In addition to the mentioned studies, numerous other research studies have focused on the effectiveness of different data mining techniques in detecting fraud in financial statements [17]. Tatusch et al. [31] introduced a modified version of DBSCAN, a density-based clustering algorithm, which outperformed prior methods in detecting restated financial statements. Ravisankar et al. [26] and Gill & Gupta [7] both found that probabilistic neural network (PNN) and neural network techniques were effective in identifying companies resorting to financial statement fraud. These studies collectively suggest that a combination of clustering and classification techniques, particularly those that incorporate temporal variation and financial ratios, can be effective in predicting inaccurate financial statements. All of these authors have investigated fraud detection in the profit sector. However, We were unable to identify significant scientific studies on detecting inaccurate financial statements in government-owned enterprises. Numerous authors have demonstrated the effectiveness of utilizing the Beneish M-Score model in practice, although they did not employ data mining techniques but rather relied solely on the mathematical formula of the model [10]. 3. Research elaboration 3.1. Methodology This paper examines the data extracted from the financial statements of Government-Owned Enterprises (GOEs) in the FBiH and the corresponding audit reports prepared by the Audit Office of the Institutions in FBiH. The study encompasses a time span of 16 years, from 2004 to 2019. It comprises a total of 200 financial statements and their associated audit reports, all of which were available at the time of conducting the study. Financial statements serve as the fundamental basis and starting point for analyzing business operations and assessing the condition of a company. They can also be regarded as confidential records generated by organizations that contain their financial transactions, including expenses, realized profits, income from loans, etc. [8]. Financial statements provide an indication of the organization’s financial reality and also include management notes on business performance and projected future trends. Furthermore, inaccurate financial statements deceive users of financial reports by creating the impression that organizations are performing favorably. The task of auditing is to protect the interests of capital owners and provide a reliable information foundation for rational decision-making and management of state-owned companies. In accordance with the International Standard on Auditing (ISA) 240 (sections 2 and 3), misstatements in financial statements can occur due to either fraud or error (the International Standards of Supreme Audit Institutions-ISSAI does not specifically define inaccurate financial statements like ISA 240 does). The key differentiating factor between fraud and error lies in whether the underlying action that leads to the misstatement in financial statements is deliberate or unintentional. 4 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... While fraud is a broad legal concept, auditors, under the ISAs, focus on fraud that causes significant misstatements in financial statements. There are two types of intentional misstatements that are relevant to auditors: misstatements arising from fraudulent financial reporting and misstatements resulting from misappropriation of assets. Although auditors may suspect or, in rare cases, identify instances of fraud, they do not make legal determinations regarding the occurrence of fraud. Consequently, the term "misstatement" will be used going forward without explicitly specifying whether an error is intentional or not . Pursuant to International Standard on Auditing (ISA) 240 (sections 2 and 3), misstatements in the financial statements can arise from either fraud or error (International Standards of Supreme Audit Institutions-ISSAI does not specifically define inaccurate financial statements like ISA 240 does). The key difference between fraud and error lies in the intent behind the action that leads to inaccuracies in the financial statements—whether it is intentional or unintentional. Although the concept of fraud is broad in legal terms, under the International Standards on Auditing, the auditor focuses on frauds that cause material misstatements in the financial statements. There are two types of intentional misstatements relevant to the auditor: those arising from fraudulent financial reporting and those resulting from the misappropriation of assets [11]. Although the auditor may suspect or, in rare cases, identify the occurrence of fraud, the auditor does not make legal determinations of whether fraud has actually occurred. Therefore, the term misstatement shall be used onwards without specifying if error is intentional or not. International Standard on Auditing (ISA) 705 (paragraph 2) defines three types of modified opinions: a qualified opinion, an adverse opinion, and a disclaimer of opinion. The decision regarding which type of modified opinion is appropriate depends on the nature of the matter causing the modification, i.e., whether the financial statements are materially misstated or, in cases where sufficient and appropriate audit evidence cannot be obtained, whether they could be materially misstated; and the auditor’s judgment regarding the pervasiveness of the effects or possible effects of the issue on the financial statements [12]. Our model uses balance sheet items from financial statements as predictive attributes and the type of opinion on the financial statements as the target variable. As the input set of data for the model, we used 24 balance sheet positions: Opening and closing balances of accounts receivable (attribute code: A3, A1), Sales income for the current and previous accounting period (A4, A2), Operating expenses of the current and previous accounting periods (A6, A5), Opening and closing balances of operating assets (A11, A7), Opening and closing balances of Property, Plant, and Equipment (A12, A8), Opening and closing balances of short-term financial investments (A13, A9), Opening and closing balances of business assets (A14, A10), Opening and closing balances of Depreciation (A15, A16), Administrative expenses of the current and previous accounting period (A17, A18), Figure 1: Audit reports distribution for GOEs in FBiH. 5 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... Opening and closing balances of short-term liabilities (A21, A19), Opening and closing balances of long-term liabilities (A22, A20), Business profit/loss for the current accounting period (A23) and Net cash flow from operating activities for the current accounting period (A24). These balance sheet items were selected based on the fact that auditors have predominantly identified irregularities in the valuation of these balance sheet positions as grounds for issuing qualified opinions. Auditors currently highlight non-compliance with the provisions of IFRS 9, IFRS 15, IAS 2, IAS 16, IAS 36, IAS 37, IAS 38, as well as cash flows from operating activities and business results in the financial statements of government-owned enterprises in FBiH. The output variable was the quality of financial statements measured relative to the audit report of state auditors. The results of the final audit reports for GOEs subject to the analysis are given in Figure 1. Number of enterprises is presented on the x-axis. The output variable – audit report for GOEs in FBiH can be grouped as follows: •As four categories or classes, so that audit report is a class, as presented in Table 1, •As two classes coded as YES category – unqualified and qualified opinion, and NO category – adverse opinion and disclaimer of opinion, as presented in Table 2. Class Audit report Sample number Percentage 1 Unqualified opinion 21 10.50% 2 Qualified opinion 106 53.00% 3 Disclaimer of opinion 6 3.00% 4 Adverse opinion 67 33.50% SUM 200 100.00% Table 1: Four classes, according to the final audit report. Class Assessment Sample number Percentage 1 No 73 36.50% 2 Yes 127 63.50% SUM 200 100.00% Table 2: Two classes, according to the final audit report. It is evident that prediction error in the first case would be much higher due to different distribution of the final audit report by classes. Hence, this research gave advantage to the second case. The enterprises are divided into two basic groups: •The first group includes the enterprises whose financial statements were given unqualified or qualified opinion (127 enterprises). •The second group includes the enterprises whose financial statements were given adverse of disclaimer of opinion (73 enterprises) Such formulation of the output variable categorizes the problem as classification problem, where the aim of the model is to learn how to recognize the proper classification of the final audit report. The primary goal of prediction is to develop a model that derives insights about a specific characteristic of the dependent variable by utilizing a combination of independent variables. Predictive modeling involves determining the output variable for a constrained dataset, where the symbols represent the values of the output variable in particular instances. The choice of variables from the available dataset significantly influences the precision and accuracy of the resulting predictive models. 6 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... 3.2. Data mining Data mining, a field of knowledge discovery in databases [1], can be utilized for discovering financial frauds. Different types of data mining techniques can be employed for this purpose. Classification is a method used to evaluate and find a function that assigns items from a data set to predetermined classes of the output variable, based on the input variables’ values [18]. Through classification, models can be created to classify unknown datasets into specific categories or classes [18]. The classification process typically involves the following steps: •Selecting classifiers for implementing the classification algorithm. •Choosing the class attribute (output variable). •Dividing data sets into two: training data and test data. •Training the classifiers on the training data set with known values of the class attribute. •Testing the classifiers on the test data set with hidden values of the class attribute. In the classification process, existing techniques are applied to evaluate the proposed prediction model using collected instances. In the case of selecting a classifier for datasets obtained from the financial statements of government-owned enterprises, the situation is quite clear. This dataset is very small; the procedure of conducting the experimentation phase is more difficult due to the fact that the data is dynamically changeable. Financial statements data are mostly of a numerical and categorical type, and since they are manually extracted from databases of financial statements, they require less cleaning in the preprocessing phase. Some applications of different classifiers on such data are described in the papers mentioned in the listed references. In the paper [15], the authors employed DT model for financial fraud detection in companies. Jan [14] used decision trees in combination with other data mining techniques to achieve a high accuracy rate of 90.83% in detecting financial statements fraud. Kirkos, Spathis & Manolopoulos [15] also explored the effectiveness of decision trees in detecting fraudulent financial statements, comparing their performance with other data mining techniques. These studies collectively highlight the potential of decision trees in detecting fraud in financial statements of government-owned enterprises. However, some recent approaches have been introduced for financial fraud detection [22]. Kırda & Özçelik [16] found KNN to be a highly effective classifier, achieving an accuracy rate of 91.73% in detecting financial statement fraud. Yao et. al. [36] also highlighted the effectiveness of KNN, particularly when combined with support vector machine (SVM) and stepwise regression, in detecting fraudulent financial statements. These studies collectively underscore the potential of KNN in detecting fraud in government-owned enterprises . A range of studies have demonstrated the effectiveness of Multilayer Perceptron (MLP) in detecting fraud in financial statements. Trigueiros & Sam [32] and Mubarek & Adali [24] both found that MLP, when used in conjunction with other machine learning techniques, outperformed traditional methods in fraud detection. Ravisankar et al. [26] and Kwon & Feroz [19] further support these findings, with Ravisankar [26] noting the superior performance of MLP in identifying companies engaged in financial statement fraud, and Kwon [19] reporting an 88% accuracy rate in predicting SEC investigation targets using MLP. These studies collectively highlight the potential of MLP in detecting fraud in financial statements, particularly in government-owned enterprises. For example, some studies have employed Bayesian networks to detect inaccurate financial statements. Deng [4] found the algorithm to be effective in this context, noting its potential for proactive fraud detection. Handoko, Wiyardi & Handoko [9] also found success in using the Beneish M-Score method, which includes variables related to financial statement manipulation, in detecting fraud in Indonesian government-owned enterprises. In this work, and in accordance with the above mention, the following algorithms were chosen: Decision Trees (J48) [14,15], K-nearest neighbors [16,36], Neural Network (MLP) 7 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... [8,19,24,26,32], and Bayesian network (Naive Bayes) [4,9] for application on the training dataset. The following subsection offers an detailed overview of the machine learning techniques applied in this research to detect financial fraudulent activities. Decision Trees (DT) are one of the most well-known classification techniques often used to model data in the form of a tree structure. These algorithms establish relationships between input features and outputs using a tree-like structure, named for its resemblance to an inverted tree. The name "decision tree" derives from the fact that it resembles an inverted tree. The most commonly used and widely recognized decision tree algorithm is C4.5, and its implementation in the Weka software tool is known as the J48 algorithm. The J48 (C4.5) algorithm [25] presents an extension of Professor Ross Quinlan’s earlier ID3 algorithm, and is known for its exceptionally high accuracy. The advantage of the J48 algorithm is the ability to work with numerical and categorized data [34], in addition, it is easy to implement and effectively deals with noise and missing values [23], and it has the ability to display results graphically. Basic construction of J48 algorithm uses a method known as divide and conquer to display output [34,23]: •Select the dataset as input for the rule-making process. •Calculate the normalized information gain for each attribute. •Choose the attribute with the maximum information gain as the best attribute. This attribute becomes the root node and corresponding to the best predictor. •Repeat the above step until a stopping criterion is met, calculating the information gain for each attribute and adding that attribute as a child node. Just as a tree starts from the root, branches into individual branches, and ends in leaves, decision trees use branches to represent decision paths, with the final outcome represented by the leaves. The final result is a tree with decision nodes and leaves. Each decision node has two or more divisions, while each leaf represents an outcome or decision. K-nearest neighbors (KNN) algorithm represents a straightforward and easily understandable supervised learning method commonly applied in both classification and regression problems . It belongs to a group of algorithms known as instance-based learning, sometimes referred to as ‘lazy learning methods’ because the processing of the training data is delayed until a test instance needs to be classified. KNN operates on the principle that objects with similar characteristics tend to belong to the same class or have similar output values. The main goal of KNN is to group n objects into k groups (classes) based on their attributes or features. When a testing example is considered, it is placed in an n-dimensional metric space of attribute values. The algorithm then determines the distance between the testing example and all training examples in this space. For classification, the most popular classification among the k nearest training examples is the target estimation for the classifier. To define ‘nearest,’ various metrics can be used, including the standard Euclidean distance and Hamming distance, with Euclidean distance being commonly used. If k is greater than 1, those closer to the testing example will have a greater weight in the classification. The procedure starts with an initial division of the set items into a selected number of groups. The distance between every object and every group is determined [22], and objects are located into groups closest to them based on given characteristics. After joining an object to a group, the centroid of the group is recalculated. The distance of every object from the group centroid is recalculated, and the distribution of objects among groups continues until the selected function of the criterion suggests otherwise. Because of all the above, the KNN algorithm is easy to implement and understand, and is more suitable for small data sets. However, for large datasets, it can be computationally demanding because it requires calculating the distances to all instances in the training set for each new instance. 8 CRORR 16:1 (2025), 1–15 Gadžo et al.: Data mining approach in detecting inaccurate financial statements... [33] Wu, H.-P. and Li, L. (2021). The BP neural network with adam optimizer for predicting audit opinions of listed companies. IAENG International Journal of Computer Science, 48(2). url: https://www.iaeng.org/IJCS/issues_v48/issue_2/IJCS_48_2_16.pdf [Accessed 29/6/2024] [34] Witten, I. H. and Frank, E. (2005). Data Mining: Practical Machine Learning Tools and Techniques. 2nd Edition. San Francisco: Morgan Kaufmann Publishers. url: http://academia.dk/BiologiskAntropologi/Epidemiologi/DataMining/Witten_and_Frank_Data Mining_Weka_2nd_Ed_2005.pdf [Accessed 29/6/2024] [35] Yadiati, W. and Rezwiandhari, A. (2023). Detecting Fraudulent Financial Reporting In StateOwned Company: Hexagon Theory Approach. JAK (Jurnal Akuntansi) Kajian Ilmiah Akuntansi, 10(1), 128-147. doi: 10.30656/jak.v10i1.5676 [36] Yao, J., Pan, Y., Yang, S., Chen, Y. and Li, Y. (2019). Detecting fraudulent financial statements for the sustainable development of the socio-economy in China: a multi-analytic approach. Sustainability, 11(6), 1579. doi: 10.3390/su11061579 [37] Yue, D., Wu, X., Shen, N. and Chu, C.-H. (2009). Logistic regression for detecting fraudulent financial statement of listed companies in China. In: 2009 International Conference on Artificial Intelligence and Computational Intelligence. doi: 10.1109/AICI.2009.421 15