scieee AI-readable full text Open interactive document viewer

Statistical Modelling and Analysis of Factors Influencing Employee Attrition

Khangar, Nutan V.; Amrutkar, Kalpesh P.

Abstract

Employee attrition refers to the voluntary or involuntary departure of employees from an organization, and it poses significant challenges for businesses across industries. High attrition can disrupt workflow, increase recruitment and training costs, and impact overall organizational performance. The present paper aims to analyze and understand the factors contributing to employee attrition within an organization. Through an extensive literature review and systematic data analysis, the paper examines the relationships among various demographic, organizational, and job-related factors, thereby identifying the key drivers influencing employee departures. To explore the underlying structure of the data, Principal Component Analysis (PCA) and Multiple Correspondence Analysis (MCA) are employed to visualize patterns and identify correlated factors associated with attrition. Additionally, Factor Analysis of Mixed Data (FAMD) is performed to jointly analyze numerical and categorical variables, enabling a comprehensive assessment of their combined influence on employee turnover. Numerical results, insights, and conclusions derived from these analyses are presented in the final section of the study.

Full text

International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 913 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 Statistical Modelling and Analysis of Factors Influencing Employee Attrition Nutan V. Khangar Assistant Professor, Department of Statistics MVP’s Arts, Science and Commerce College, Ozar (Mig), Nashik [email protected] Kalpesh P. Amrutkar Assistant Professor, Department of Statistics MVP’s KRT Arts, BH Commerce and AM Science College, Nashik [email protected] Abstract Employee attrition refers to the voluntary or involuntary departure of employees from an organization, and it poses significant challenges for businesses across industries. High attrition can disrupt workflow, increase recruitment and training costs, and impact overall organizational performance. The present paper aims to analyze and understand the factors contributing to employee attrition within an organization. Through an extensive literature review and systematic data analysis, the paper examines the relationships among various demographic, organizational, and job-related factors, thereby identifying the key drivers influencing employee departures. To explore the underlying structure of the data, Principal Component Analysis (PCA) and Multiple Correspondence Analysis (MCA) are employed to visualize patterns and identify correlated factors associated with attrition. Additionally, Factor Analysis of Mixed Data (FAMD) is performed to jointly analyze numerical and categorical variables, enabling a comprehensive assessment of their combined influence on employee turnover. Numerical results, insights, and conclusions derived from these analyses are presented in the final section of the study. Keywords: Employee attrition, PCA, MCA, FAMD, etc. 1. INTRODUCTION Employee attrition is a natural aspect of organizational life. Over time, employees leave an organization for a wide range of personal and professional reasons. While a certain amount of turnover is expected, attrition becomes problematic when it rises beyond a manageable level. In today’s dynamic work environment, it is uncommon for individuals to remain with a single organization throughout their careers. Many move on voluntarily for better opportunities or International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 914 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 personal circumstances, while others exit involuntarily through layoffs, restructuring, or performance-related decisions. Broadly, the term attrition refers to a gradual reduction in workforce size when departing employees are not replaced. According to the Cambridge Dictionary, attrition involves “the process of gradually weakening or destroying something,” which, in an organizational context, reflects a shrinking employee base when staff departures outpace new hiring. When this occurs, the workforce slowly diminishes in size, sometimes intentionally and sometimes as a result of external pressures. For organizations, monitoring attrition rates is essential for understanding the effectiveness of talent retention strategies. A high attrition rate signals frequent employee departures, whereas a low rate indicates relative workforce stability. Attrition may also serve as a strategic tool. For example, organizations facing budget constraints may implement a hiring freeze, choosing not to replace employees who retire or resign, thereby reducing labour costs without formal layoffs. Employee attrition generally falls into two major categories: Voluntary attrition occurs when employees choose to leave the organization. This may include resignations due to new job opportunities, relocation, personal commitments, dissatisfaction with working conditions, or health-related issues. Even when employees feel compelled to leave because of a negative work environment, the departure is classified as voluntary if initiated by the employee. Involuntary attrition arises when the organization makes the decision to terminate employment. This can result from job redundancies, restructuring, poor performance, disciplinary reasons, or job abandonment. In such cases, the organization typically elects not to refill the vacated position. To understand the underlying causes of attrition and identify the factors influencing employee turnover, various statistical and analytical methods are applied. These methods help uncover patterns, relationships, and predictors of attrition, providing valuable insights for improving workforce stability and organizational performance. To develop a clear understanding of the analytical methods used in this study, an extensive review of relevant research was conducted. A brief synthesis of key contributions from the literature is presented below. Ljubičić et al. (2021) applied Principal Component Analysis (PCA) to clinical and biochemical variables of children diagnosed with congenital adrenal hyperplasia. Their work is noteworthy because they constructed PCA models based on sexand age-adjusted standard deviation scores for hormone parameters, enabling clear visualization of combined endocrine profiles in both classical and non-classical CAH patients. International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 915 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 Khangar and Kamalja (2017) provided a comprehensive overview of Multiple Correspondence Analysis (MCA), discussing its methodological foundations, the role of biplots in correspondence analysis, and theoretical issues associated with various MCA approaches. They emphasized the usefulness of MCA based on separate singular value decompositions (SVDs) and illustrated its application through examples, including the analysis of mother–child behavioral patterns. Husson and Josse (2014) offered a detailed explanation of Correspondence Analysis (CA) using both indicator and Burt matrices. Their work effectively demonstrates how to interpret the cloud of individuals and categories, and they highlight the use of biplots and correlation circles for visualizing relationships among variables. Extensions and methodological developments in CA and MCA were further explored by Hwang et al. (2009), who proposed an advanced version of MCA capable of capturing cluster-level heterogeneity among respondents. The practical utility of MCA has been demonstrated in various applied fields. For example, Magagula et al. (2018) used MCA to examine associations within the 2018 Mozambique Malaria Indicator Survey dataset. Their study underscored the technique’s value in epidemiological research, particularly in revealing complex multivariate patterns that aid in effective communication of results. Similarly, Katarina and Šraj (2019) implemented MCA to assess how characteristics of rainfall events influence interception by urban tree species such as birch and pine. Brunette et al. (2017) employed MCA to investigate forest adaptation strategies in response to climate-related risk factors, highlighting its relevance in climate-change-oriented ecological studies. In the context of mixed-type data analysis, Sayadi et al. (2021) discussed the application of Factor Analysis of Mixed Data (FAMD) within a distributed computing environment. Their work, motivated by the development of KITAPP—a personalized medicine application—introduced a new distributed FAMD framework suitable for multi-site analysis. They further demonstrated how individualized reference data can support decision-making while maintaining privacy and datausage control in medical settings. Recent research in the area of employee attrition demonstrates a significant shift from traditional descriptive approaches to predictive HR analytics, particularly through the application of machine learning techniques to identify employees at risk of resignation and support proactive retention planning. Studies such as Raza (2022) highlight that ensemble models like the Extra Trees Classifier achieve high predictive accuracy, illustrating the effectiveness of advanced learning algorithms for attrition forecasting. Further developments, including the 2024 study Employee Attrition: Analysis of Data-Driven Models, show that hybrid and ensemble methods outperform single-model International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 916 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 approaches, capturing complex nonlinear attrition patterns. Govindarajan et al. (2025) and other recent investigations reinforce that modern classification models such as Random Forest, k-Nearest Neighbours, Naïve Bayes, and Logistic Regression enhance intervention strategies by identifying key determinants of turnover. Research also stresses that the value of predictive HR analytics lies not only in forecasting exits but also in transforming HR practices from reactive to strategic, enabling targeted initiatives in compensation, leadership development and employee engagement (Shingh et al. (2025)). Beyond machine learning approaches, scholars have explored theoretical frameworks linking job satisfaction, organizational commitment, and leadership relationships to classical turnover constructs. Recent bibliometric analyses emphasize an increasing reliance on multivariate exploratory techniques, particularly Correspondence Analysis (CA) and Multiple Correspondence Analysis (MCA), to visually uncover associations among categorical HR variables and identify employee profiles associated with turnover behaviour. MCA has been used effectively across domains such as talent acquisition, workplace perceptions, and hospitality sector retention research, demonstrating its ability to interpret complex categorical relationships. Additionally, Factor Analysis of Mixed Data (FAMD) has gained attention as a method well-suited to datasets that combine numerical variables (e.g., age, salary, tenure) and qualitative variables (e.g., education, department, job role). Recent studies in education analytics, public health, and customer churn analysis show that FAMD provides a more holistic structural view of mixed data and improves modelling performance when used before predictive tasks. Although current FAMD applications are more common outside HR, they establish a strong methodological foundation for its use in attrition analytics, revealing low-dimensional patterns not captured by PCA or MCA alone. Collectively, this emerging body of research shows that integrating machine learning, multivariate exploratory methods, and mixed-data analysis offers deeper insights into attrition dynamics and supports evidence-based strategies for reducing turnover and improving employee retention. Collectively, this study highlight the versatility and effectiveness of PCA, MCA, and FAMD for exploring complex datasets, identifying associations, and generating interpretable multivariate structures across a range of disciplines. 2. METHODOLOGY In this Paper, a combination of multivariate statistical techniques PCA, MCA, and FAMD was employed to comprehensively analyze the factors contributing to employee attrition. PCA was first applied to the numerical variables to reduce dimensionality and identify key continuous attributes that contribute most to variance in the dataset, enabling clearer interpretation of International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 917 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 underlying patterns. However, because the dataset also includes several categorical variables, PCA alone was insufficient for capturing their structure. Therefore, Multiple Correspondence Analysis (MCA) was utilized to explore relationships among categorical attributes and to visually represent clusters and associations across employee categories, such as job role, department, education field, and satisfaction levels. While PCA and MCA separately offer valuable insights into quantitative and qualitative variables, employee attrition datasets are inherently mixed in nature. To obtain an integrated understanding of both variable types, Factor Analysis of Mixed Data (FAMD) was applied, providing a unified framework that simultaneously incorporates numerical and categorical measures and explains more variation than PCA or MCA individually. Together, these complementary techniques enable a deeper and more holistic understanding of employee attrition patterns, revealing the complex interactions that influence turnover behavior. 2.1 PCA PCA is widely recognized as one of the most effective techniques for exploring and interpreting high-dimensional datasets. It enhances the interpretability of data by transforming a large set of correlated variables into a smaller set of uncorrelated components, while retaining as much of the original information as possible. This makes PCA especially valuable when working with datasets that contain many features per observation, where visualizing and understanding multidimensional structure is otherwise difficult. In practice, the first two principal components are often used to project the data onto a two-dimensional plane, enabling analysts to visually detect clusters, relationships, or underlying patterns. From a statistical perspective, PCA reduces dimensionality through a linear transformation that reexpresses the original data in a new coordinate system. The first principal component is defined as the linear combination of the original variables that captures the maximum possible variance. The second principal component accounts for the greatest remaining variance after adjusting for the first component, and the process continues until all meaningful variation in the data is explained. PCA is particularly useful when the original variables exhibit strong correlations, as it provides a compact set of components that summarize the essential structure of the dataset. Historically, PCA was first introduced by Karl Pearson in 1901 as an extension of the principal axis theorem in mechanics, and was later independently formalized and named by Harold Hotelling in the 1930s. Over time, the method has been rediscovered and adapted across multiple disciplines, resulting in a variety of alternative names and interpretations. In signal processing, it is known as International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 918 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 the Karhunen–Loève Transform (KLT); in quality control, it is referred to as the Hotelling transform; and in mechanical engineering, it appears as Proper Orthogonal Decomposition (POD). Mathematically, PCA can be derived through the singular value decomposition (SVD) of the data matrix or from the eigenvalue decomposition (EVD) of the covariance matrix 𝑋𝑇𝑋. In meteorology, it is commonly used under the term Empirical Orthogonal Functions (EOFs), while in structural dynamics it is closely related to empirical modal analysis. Variants of PCA have also been linked to theoretical frameworks such as the Eckart–Young theorem and empirical eigenfunction decomposition. 2.2 CA CA is a multivariate exploratory technique used to graphically display and interpret associations among categorical variables. It provides a low-dimensional visual representation of the relationships between the levels of variables, making it a powerful tool for analyzing contingency tables. Simple Correspondence Analysis (SCA) is applied to two-way contingency tables, allowing the study of associations between two categorical variables. When the dataset involves more than two categorical variables, MCA serves as a natural extension, enabling the analysis of multiway categorical data. Biplots play a central role in both CA and MCA, as they allow simultaneous representation of individuals (observations) and categories (variable levels) in the same geometric space. The distances and proximities in these plots help researchers interpret patterns of similarity, association, and clustering. Because of this intuitive graphical output, CA and MCA have become widely used in behavioural sciences, social research, marketing analytics, demography, and epidemiology, where categorical variables dominate. MCA is specifically designed to uncover underlying structures in datasets composed entirely of categorical variables. It transforms the original qualitative information into a set of dimensions that summarize the main patterns of variability, making it ideal for analyzing large-scale surveys, questionnaires, and socio-demographic datasets. MCA can be performed using either the indicator matrix (also called the complete disjunctive table) or the Burt matrix. • The indicator matrix is an individuals × categories binary matrix in which each column corresponds to a category of a variable, and each row corresponds to an individual. Applying CA to this matrix allows the direct representation of individuals in the factor space. International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 919 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 • The Burt matrix is a symmetric compilation of all pairwise cross-tabulations among the variables. MCA applied to the Burt table is conceptually closer to extending simple CA to multiple variables. Both approaches yield equivalent results under appropriate scaling, although the indicator matrix is often preferred when the focus is on interpreting individual-level behaviour, whereas the Burt matrix is suited for understanding relationships among variable categories. Importantly, MCA not only identifies the dominant variables contributing to the overall variation but also allows the inclusion of supplementary (passive) variables and individuals; elements that do not influence the construction of dimensions but can be projected onto the factor map for interpretive insights. This flexibility makes MCA an effective method for exploratory data analysis, data reduction, cluster identification, and hypothesis generation in categorical data contexts. By linking SCA and MCA conceptually, it becomes evident that both techniques share the same mathematical foundation; based on the 𝜒2 metric, decomposition of inertia, and singular value decomposition (SVD). MCA essentially extends the logic of SCA to higher-dimensional categorical structures, offering a unified geometric framework for understanding associations across simple and complex categorical datasets. 2.3 FAMD FAMD is a multivariate exploratory technique specifically developed for datasets that contain a combination of quantitative and qualitative variables. Conceptually, FAMD can be viewed as a hybrid approach that integrates the strengths of PCA which is ideal for continuous variables and MCA which is suited for categorical variables. By harmonizing these two frameworks, FAMD provides a unified method that simultaneously retains the information from both types of data, ensuring that no part of the dataset is neglected. Historically, the foundations of FAMD were established by Brigitte Escofier and Gilbert Saporta, whose work was later extended and formalized by Jérôme Pagès in 2002. Pagès’ contributions, particularly in his comprehensive book on the subject, form the most complete and widely referenced exposition of FAMD in the English language. The need for FAMD arises from the limitations of applying PCA or MCA alone on mixed datasets. PCA is designed for quantitative variables and treats categorical data as noise, resulting in the loss of qualitative information. Conversely, MCA focuses exclusively on categorical variables and cannot appropriately incorporate quantitative features, causing numerical variability to be ignored. International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 920 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 FAMD resolves this dilemma by balancing the influence of both types of variables through a specific normalization scheme that ensures equal contribution in the construction of dimensions. FAMD works by transforming categorical variables into sets of binary indicators, similar to the approach used in MCA, while simultaneously preserving the numerical structure of quantitative variables as done in PCA. The method assigns appropriate weights so that each variable regardless of its type contributes proportionally to the overall analysis. This produces a factor space where individuals and variables can be visualized together, enabling interpretation of patterns that involve both quantitative and qualitative dimensions. One of the key advantages of FAMD is its ability to capture interactions between variable types. For example, it can reveal how categories of a qualitative variable (such as job role, education level, or region) relate to quantitative measures (such as salary, age, or performance scores). It also allows for the inclusion of supplementary individuals and variables, making it useful for prediction, profiling, clustering, and exploratory modelling. By connecting PCA, MCA, and FAMD within a single analytical framework, FAMD provides a powerful tool for modern datasets such as surveys, socio-economic data, employee attrition data, healthcare studies, and customer satisfaction research where mixed data types are common. The result is a comprehensive, interpretable geometric representation that preserves the richness of the dataset while uncovering its underlying structure. 3 DATA STRUCTURE AND PRESENTATION The dataset comprises 35 variables capturing a wide range of employee characteristics and workplace attributes. These include demographic factors such as age, gender, marital status, and frequency of business travel, along with compensation related information like daily rate of pay. Organizational and positional details are also represented, including employees’ distance from the home office and their level of education. In addition, several job-specific parameters are recorded, such as job involvement, job level relative to other roles within the organization, and the specific job function or role assigned to the employee. Measures of work effort such as total hours worked per week, month, or year, including both standard and overtime hours are also included, providing a comprehensive view of the professional environment and workload experienced by employees. The dataset is available on https://www.kaggle.com/datasets/thedevastator/employee-attritionand-factors . International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 921 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 3.1 Graphical Presentation Fig 1. Department-wise Attrition Fig 2. Gender-wise Attrition Fig. 3 Age-wise Attrition Fig. 4 Attrition by Education Field and level Fig. 5 Attrition by Education Field Fig. 6 Attrition by Job Level Fig. 7 Attrition by Distance from Home Fig. 8 Attrition by Marital Status International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 928 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 Table 3: ANOVA Table Interaction Inertia DF 𝒑-value AB 3.02189 9 0.963421 AC 17.3983 3 0.000585197* BC 5.36232 3 0.147109 ABC 12.9959 9 0.162792 The interaction between Job Satisfaction (A) and Relationship Satisfaction (B) is not statistically significant (p = 0.963), indicating that these two variables do not show a meaningful combined effect on the response pattern. The interaction between Job Satisfaction (A) and Attrition (C), however, is statistically significant (p = 0.000585), suggesting that attrition levels differ significantly across the categories of job satisfaction. In contrast, the interaction between Relationship Satisfaction (B) and Attrition (C) is not significant (p = 0.147), implying that relationship satisfaction alone does not significantly influence attrition when considered in combination. The three-way interaction among Job Satisfaction (A), Relationship Satisfaction (B), and Attrition (C) is also not significant (p = 0.163). This means that the combined effect of job satisfaction and relationship satisfaction on attrition is not strong enough to produce a statistically meaningful three-way association. Overall, even though the three-way interaction (A × B × C) is insignificant, the significant twoway interaction between Job Satisfaction and Attrition (A × C) indicates that job satisfaction plays an important role in influencing attrition independently of relationship satisfaction. Fig. 12 Biplot -0.25 -0.2 -0.15 -0.1 -0.05 0 0.05 0.1 -0.15 -0.1 -0.05 0 0.05 0.1 Dim 1: 63.29% Dim 2: 20.46% B1C1 B2C1 B3C1 B4C1 A1C1 A2C1 A3C1 A4C1 B1C2 B2C2 B3C2 B4C2 A1C2 A2C2 A3C2 A4C2 Association between B and A across the levels of C A and B across C1 A and B across C2 International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 929 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 From Figure 12, it is observed that individuals with very high relationship satisfaction and high or medium job satisfaction tend to exhibit a negative attitude toward attrition. Conversely, individuals with low job satisfaction and medium relationship satisfaction show a positive attitude toward attrition. Furthermore, respondents with medium relationship satisfaction and very high job satisfaction also demonstrate a negative attitude toward attrition. Fig. 13. Biplot for FAMD From Figure 13, it appears that Department and Education Field are highly associated with each other, and both show a similar level of association with Job Role. Age and Job Level also exhibit a strong mutual association. The qualitative variables Job Role, Department, Education Field, and Job Level—demonstrate high discrimination power, indicating strong interrelationships among them. Among the quantitative variables, Years in Current Role, Total Working Years, Years Since Last Promotion, Years with Current Manager, and Years at Company show strong associations with one another. These quantitative variables, together with the categorical variable Age, form a cluster that reflects a high level of internal association. All remaining numerical and categorical variables lie closer to the origin, indicating weaker discriminatory power and relatively lower contribution to the overall association structure. From Figure 14, we observe that the Human Resources category within Education Field, along with Department and Job Role, has high discrimination power, indicating a strong association among these variables. A clear association is also seen between Research Director, Job Level International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 930 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 5, and the 46–60 age group. Similarly, Manager, Job Level 4, and the 46–50 age group show strong interrelationships. Further, Job Level 2, Sales Department, Sales Executive (Job Role), and Marketing form a highly associated cluster. Job Level 1 is closely associated with the 18–30 age group, and the Research & Development (Education Field) category shows strong association with both the 18–30 age group and Job Level 1, suggesting that individuals who enter research-oriented fields often begin developing interest early in their careers. In contrast, Research and Development and the 46–60 age group show high discrepancy. Additionally, Education Level 5 and Job Level 3 exhibit a notable association. All remaining variables and their categories cluster near the origin, indicating weaker discriminatory power and strong internal association but limited contribution to the overall structure. Fig.14. Graph of Categories of Qualitative Variables International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 931 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 Fig. 15. Graph of Quantitative Variable From Figure 15, we observe that all numerical variables are closely clustered, indicating a high degree of association among them. No numerical variable appears as an outlier in the data. Variables positioned closer together exhibit stronger correlations, whereas those located farther apart show weaker or no correlation. A particularly strong association is observed between Total Working Years and Monthly Income. Similarly, Years Since Last Promotion and Years at Company show a strong linear relationship. The variables Years with Current Manager and Years in Current Role are also highly correlated with each other. Variables positioned on opposite sides of the plot, separated by a large distance, are likely to have a negative correlation. For instance, as Distance From Home increases, the values of Total Working Years, Years at Company, and Years Since Last Promotion tend to decrease, indicating negative correlations. Likewise, as Training Time Last Year increases, Total Working Years, Years at Company, and Years Since Last Promotion tend to decrease, reflecting another set of negative correlations. The combined use of PCA, MCA, and FAMD offers a comprehensive understanding of the multidimensional factors influencing employee attrition. PCA helped identify the most influential numerical variables such as monthly income, total working years, and tenure-related measures revealing that these continuous attributes play a crucial role in shaping attrition International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 932 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 outcomes. MCA complemented this by uncovering association patterns among categorical factors, illustrating how specific employee groups (based on job role, department, education field, and satisfaction levels) differ in attrition tendencies. However, since real HR datasets typically include a mixture of both quantitative and qualitative variables, neither PCA nor MCA alone could fully capture the complexity of interactions present. Therefore, FAMD was employed as an integrative approach, allowing simultaneous assessment of both variable types and providing a more holistic view of the underlying structure. The insights derived from FAMD highlight clusters of characteristics strongly associated with attrition risk, supporting more informed decision-making in employee retention strategies. Overall, the combination of these techniques enhances interpretability, strengthens analytical depth, and provides a robust foundation for identifying actionable determinants of employee turnover. 4 CONCLUSIONS High-dimensional datasets can be effectively visualized by projecting them into a lowerdimensional space using eigenvalues obtained from dimensionality-reduction techniques. Based on the PCA results, the numerical variables Monthly Income, Total Working Years, Years at Company, Years in Current Role, and Years Since Last Promotion emerge as the most influential contributors, indicating their substantial role in understanding patterns related to employee attrition. The MCA results further reveal that categorical variables such as Job Role, Department, and Education Field possess the highest discrimination power and are strongly associated with each other. Additionally, variables including Marital Status, Age, Work–Life Balance, Job Level, Environment Satisfaction, Job Satisfaction, Education, Relationship Satisfaction, Overtime, Business Travel, Job Involvement, and Gender cluster closely together, suggesting that these characteristics collectively affect attrition outcomes. The FAMD analysis offers a more integrated perspective by combining both quantitative and qualitative data, explaining a greater proportion of variability than PCA or MCA due to its higher eigenvalues. The FAMD factor map highlights several strong associations such as Department–Education Field, Age–Job Level, Research Director–Job Level 5–Age Group 46–60, Manager–Job Level 4–Age Group 46–50, and Job Level 1–Age Group 18–30. Moreover, Education Level 5shows a strong link with Job Level 3. The Human Resources International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 933 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 category within Education Field, Department, and Job Rolealso demonstrates strong discrimination power and close association in the multivariate space. Among quantitative variables, Years in Current Role, Total Working Years, Years Since Last Promotion, Years with Current Manager, and Years at Company display strong internal correlations and form an interrelated cluster along with Age. Furthermore, high-level managerial roles such as Research Director and Manager are associated with Job Levels 4 and 5 and the 46–60 age group, highlighting their influence on attrition dynamics and retention decisions. To complement these findings, Chi-square partitioning was applied to selected pairs of categorical variables, revealing several statistically significant two-way interactions. Although not all interactions were significant, those that were provide meaningful insights into the factors contributing to employee attrition. These results emphasize the importance of focusing on key significant variables to help organizations design evidence-based strategies aimed at reducing attrition and improving employee retention. REFERENCES 1. Brunette, M., Bourke, R., Hanewinkel, M., and Yousefpour, R. (2018). Adaptation to climate change in forestry: A multiple correspondence analysis (MCA). Forests, 9(1), 20. 2. Ehsani, F., & Hosseini, M. (2025). Customer churn analysis using feature optimization methods and tree-based classifiers. Journal of Services Marketing, 39(1), 20-35. 3. Govindarajan, R., Kumar, K., Reddy, S., Pravallika, S., Dhatri, E. and Kumar, P. (2025). Predicting Employee Attrition: A Comparative Analysis of Machine Learning Models Using the IBM Human Resource Analytics Dataset, International Conference on Machine Learning and Data Engineering, 258, 4084-4093. 4. Husson, F., and Josse, J. (2014). Multiple Correspondence Analysis. Visualization and verbalization of data, 165-184. 5. Hwang, H., Tomiuk, M. A., and Takane, Y. (2009). Correspondence analysis, multiple correspondence analysis and recent developments. Handbook of quantitative methods in psychology, 243-263. 6. Khangar, N. V. and Kamalja, K. K. (2017). Multiple Correspondence Analysis and its applications. Electronic Journal of Applied Statistical Analysis, 10(2), 432-462. International Journal of Research (IJR) e-ISSN: 2348-6848 p-ISSN: 2348-795X Vol. 12 Issue 03 March 2025 Received: 8 March 2025 934 Revised: 21 March 2025 Accepted: 30 March 2025 Copyright  authors 2025 DOI: HTTPS://DOI.ORG/10.5281/ZENODO.17768811 7. Magagula, M. E., Ramroop, S., and Habyarimana, F. (2023). Application of Multiple Correspondence Analysis to Understand Relationships in Categorical Variables in an Epidemiological Survey Data: Analysis of the DHS-2018 Mozambique Malaria Indicator Survey Dataset, SSRN. 8. Raza, M. (2022). Machine-learning framework for analyzing organizational factors influencing employee attrition. MDPI. 9. Sayadi, S., Geffard, E., Südholt, M., Vince, N., and Gourraud, P. A. (2021). Secure distribution of Factor Analysis of Mixed Data (FAMD) and its application to personalized medicine of transplanted patients. In Advanced Information Networking and Applications: Proceedings of the 35th International Conference on Advanced Information Networking and Applications (AINA-2021), Volume 1 35 (pp. 507-518). Springer International Publishing. 10. Shafie, S., et al. (2024). Cluster-based HR analytics framework integrating artificial neural networks for employee turnover prediction. ScienceDirect. 11. Singh D., Punwatkar S., Guha S., Verma S., Sahu K., Yadav, T. and Mandpe, S. (2025). The Strategic Role of Predictive HR Analytics in Forecasting Employee Retention, Journal of Information Systems Engineering and Management, 10 (49s), 2468-4376.