scieee AI-readable full text Open interactive document viewer

Comparative study of customer segmentation strategies based on business analytics

Aimar, Davide

Abstract

This thesis explores a comparative analysis of customer segmentation strategies supported by advanced analytical methodologies. It focuses on two foundational frameworks: Recency, Frequency, Monetary (RFM) and Customer Lifetime Value (CLV), which respectively capture short-term transactional behaviors and long-term economic contributions. These metrics are subsequently analyzed through five clustering algorithms: K-Means, Hierarchical Clustering, DBSCAN, Gaussian Mixture Models (GMM), and Fuzzy C-Means. The study utilizes the UK E-Commerce data set from the UCI repository, which undergoes meticulous preprocessing and normalization to ensure robust and consistent input for the clustering models. The evaluation framework leverages two internal validation metrics—the Silhouette Score and the Calinski–Harabasz Index—to provide complementary perspectives on local density separation and global variance partitioning. Experimental results reveal that DBSCAN consistently outperforms other methods in identifying dense microclusters, often representing high-value or niche customers. In contrast, K-Means and Hierarchical Clustering exhibit stronger performance in generating broader global partitions. While Fuzzy C-Means achieves moderate results by accommodating overlapping segment boundaries through soft membership, GMM struggles with the non-Gaussian characteristics of the RFM and CLV datasets. The findings underscore that no single approach universally outperforms the others. Instead, the selection of metrics and clustering algorithms should be strategically aligned with business goals, such as identifying anomalies or performing large-scale segmentation. This study provides actionable insights for businesses aiming to enhance marketing strategies, optimize resource allocation, and strengthen customer relationship management (CRM) through data-driven segmentation approaches.

Full text

Comparative Study of Customer Segmentation Strategies Based on Business Analytics Document: Report Author: Davide Aimar Director/Co-director: Vicen¸c Fernandez Alarcon Degree:Master’s Degree in Technology and Engineering Management Examination session: Autum Comparative Study of Customer Segmentation Strategies Based on Business Analytics Abstract This thesis explores a comparative analysis of customer segmentation strategies supported by advanced analytical methodologies. It focuses on two foundational frameworks: Recency, Frequency, Monetary (RFM) and Customer Lifetime Value (CLV), which respectively capture short-term transactional behaviors and long-term economic contributions. These metrics are subsequently analyzed through five clustering algorithms: K-Means,Hierarchical Clustering,DBSCAN,Gaussian Mixture Models (GMM), and Fuzzy C-Means. The study utilizes the UK E-Commerce data set from the UCI repository, which undergoes meticulous preprocessing and normalization to ensure robust and consistent input for the clustering models. The evaluation framework leverages two internal validation metrics—the Silhouette Score and the Calinski–Harabasz Index—to provide complementary perspectives on local density separation and global variance partitioning. Experimental results reveal that DBSCAN consistently outperforms other methods in identifying dense microclusters, often representing high-value or niche customers. In contrast, K-Means and Hierarchical Clustering exhibit stronger performance in generating broader global partitions. While Fuzzy C-Means achieves moderate results by accommodating overlapping segment boundaries through soft membership, GMM struggles with the non-Gaussian characteristics of the RFM and CLV datasets. The findings underscore that no single approach universally outperforms the others. Instead, the selection of metrics and clustering algorithms should be strategically aligned with business goals, such as identifying anomalies or performing large-scale segmentation. This study provides actionable insights for businesses aiming to enhance marketing strategies, optimize resource allocation, and strengthen customer relationship management (CRM) through data-driven segmentation approaches. I Comparative Study of Customer Segmentation Strategies Based on Business Analytics Resumen Esta tesis explora un an´alisis comparativo de estrategias de segmentaci´on de clientes respaldado por metodolog´ıas anal´ıticas avanzadas. Se centra en dos marcos fundamentales: Recencia, Frecuencia, Valor Monetario (RFM) yValor de Vida del Cliente (CLV), que capturan respectivamente los comportamientos transaccionales a corto plazo y las contribuciones econ´omicas a largo plazo. Estos indicadores se analizan posteriormente mediante cinco algoritmos de agrupamiento: K-Means,Clustering Jer´arquico, DBSCAN,Modelos de Mezcla Gaussiana (GMM) yFuzzy C-Means. El estudio utiliza el conjunto de datos de comercio electr´onico del Reino Unido disponible en el repositorio UCI, que se somete a un meticuloso preprocesamiento y normalizaci´on para garantizar entradas robustas y consistentes para los modelos de agrupamiento. El marco de evaluaci´on emplea dos m´etricas de validaci´on interna—Silhouette Score y el ´ Indice de Calinski–Harabasz—que ofrecen perspectivas complementarias sobre la separaci´on de densidades locales y la partici´on de varianza global. Los resultados experimentales revelan que DBSCAN supera constantemente a otros m´etodos al identificar microclusters densos, que a menudo representan clientes de alto valor o nichos espec´ıficos. Por el contrario, K-Means y el Clustering Jer´arquico muestran un mejor desempe˜no en la generaci´on de particiones globales m´as amplias. Mientras que Fuzzy C-Means logra resultados moderados al acomodar l´ımites de segmentos superpuestos mediante asignaciones de membres´ıa difusa, GMM tiene dificultades para manejar las caracter´ısticas no gaussianas de los conjuntos de datos RFM y CLV. Los hallazgos destacan que ning´un enfoque es universalmente superior. En cambio, la elecci´on de m´etricas y algoritmos de agrupamiento debe alinearse estrat´egicamente con los objetivos comerciales, como la detecci´on de anomal´ıas o la segmentaci´on a gran escala. Este estudio proporciona ideas pr´acticas para empresas que buscan mejorar sus estrategias de marketing, optimizar la asignaci´on de recursos y fortalecer la gesti´on de relaciones con clientes (CRM) a trav´es de enfoques de segmentaci´on basados en datos. II Comparative Study of Customer Segmentation Strategies Based on Business Analytics Contents 1 Introduction 1 2 Literature Review 3 2.1 RFM as a Baseline Segmentation Tool . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2.2 Incorporating Customer Lifetime Value (CLV) . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.3 Clustering Algorithms for Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.3.1 K-MeansClustering ..................................... 5 2.3.2 HierarchicalClustering.................................... 5 2.3.3 DBSCAN ........................................... 6 2.3.4 Gaussian Mixture Models (GMM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.3.5 FuzzyC-Means........................................ 6 2.4 Leveraging CLV and RFM: Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.5 Conclusion .............................................. 7 3 Methodology 8 3.1 DataCollection............................................ 9 3.1.1 Preprocessing and Data Integrity: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.2 SegmentationMethodology ..................................... 10 3.2.1 RFMFramework....................................... 10 3.2.2 RFM Score Calculation and Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.2.3 Customer Lifetime Value (CLV) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.3 ClusteringAlgorithms ........................................ 14 3.3.1 K-meansClustering ..................................... 14 3.3.2 Clustering Algorithms: Hierarchical Clustering . . . . . . . . . . . . . . . . . . . . . . 15 3.3.3 DBSCANClustering..................................... 16 3.3.4 Gaussian Mixture Models (GMM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 3.3.5 FuzzyC-meansClustering.................................. 19 3.4 Analytical Tools and Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 3.5 ModelEvaluation........................................... 22 3.5.1 SilhouetteScore ....................................... 22 III CONTENTS CONTENTS 3.5.2 Calinski–HarabaszIndex................................... 23 4 Results 25 4.1 RFM Segmentation & Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.1.1 RFM Segmentation Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.1.2 Clustering Results on RFM Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.2 CLV Segmentation & Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 4.2.1 CLVSegmentation...................................... 37 4.2.2 Clustering Results on CLV Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . 38 4.3 ModelsEvalutation.......................................... 51 4.3.1 EvaluationonRFMData .................................. 51 4.3.2 EvaluationonCLVData .................................. 52 5 Conclusions 54 5.1 RFM analysis, clustering code and evalutation . . . . . . . . . . . . . . . . . . . . . . . . . . 64 5.1.1 RFM3Dvisualization.................................... 73 5.2 CLV segmentation,clustering and model evaluation . . . . . . . . . . . . . . . . . . . . . . . . 75 IV Comparative Study of Customer Segmentation Strategies Based on Business Analytics List of Figures 3.1 The structure of the methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 4.1 Distribution of RFM Segments Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.2 Logarithmic boxplots of RFM indicators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.3 Elbowmethodgraph......................................... 28 4.4 KMeans2Dgraph.......................................... 29 4.5 k-NN Distance Plot for DBSCAN Parameter Selection (eps = 0.45). ............. 30 4.6 DBSCANClustering(RFM)..................................... 31 4.7 Hierarchicaldendrogram....................................... 32 4.8 Cluster 5,6,7 zoomed in (left side of the previous dendogram) . . . . . . . . . . . . . . . . . . 33 4.9 Hierarchical Clustering on RFM Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 4.10 GMM clustering results (RFM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 4.11 Heatmap of Fuzzy Memberships (Sampled Data) . . . . . . . . . . . . . . . . . . . . . . . . . 36 4.12 Fuzzy C-Means Clustering (Hard Assignments) . . . . . . . . . . . . . . . . . . . . . . . . . . 37 4.13 Elbow Method for K-Means (CLV Data) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 4.14 K-Means Clustering on CLV Data (2D Projection) . . . . . . . . . . . . . . . . . . . . . . . . 39 4.15K-MeansRadarChart........................................ 40 4.16 Hierarchical Clustering Dendrogram (CLV Data) . . . . . . . . . . . . . . . . . . . . . . . . . 41 4.17 Hierarchical Clustering on CLV Data (2D Projection) . . . . . . . . . . . . . . . . . . . . . . 42 4.18hCRadarCharts........................................... 43 4.19 k-NN Distance Plot for DBSCAN Parameters Selection (eps = 0.6).............. 44 4.20 DBSCAN Clustering on CLV Data (2D Projection) . . . . . . . . . . . . . . . . . . . . . . . . 45 4.21DBSCANRadarCharts ....................................... 46 4.22 GMM Clustering on CLV Data (2D Projection) . . . . . . . . . . . . . . . . . . . . . . . . . . 47 4.23GMMRadarChar .......................................... 48 4.24 Fuzzy C-Means Clustering on CLV Data (Hard Assignments) . . . . . . . . . . . . . . . . . . 50 4.25fCRadarCharts ........................................... 50 5.1 RFM K-Means Clustering (LOG) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 5.2 RFM DBSCAN Clustering (LOG) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 V LIST OF FIGURES LIST OF FIGURES 5.3 RFMGMMClustering(LOG) ................................... 74 5.4 RFM Hierarchical Clustering (LOG) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 5.5 RFMFuzzyClustering(LOG) ................................... 75 VI Comparative Study of Customer Segmentation Strategies Based on Business Analytics List of Tables 3.1 Customer Segmentation Based on RFM Scores . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.1 Segment Statistics for RFM Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.2 Cluster Statistics for K-Means Clustering. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4.3 Cluster Statistics for DBSCAN Clustering. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 4.4 Cluster Statistics for Hierarchical Clustering. . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.5 Cluster Statistics for GMM Clustering. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 4.6 Cluster Statistics for Fuzzy C-Means Clustering. . . . . . . . . . . . . . . . . . . . . . . . . . 36 4.7 K-Means Cluster Statistics for CLV-Based Analysis . . . . . . . . . . . . . . . . . . . . . . . . 39 4.8 Hierarchical Clustering (CLV) – Mean Values. . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 4.9 Hierarchical Clustering (CLV) – Standard Deviations. . . . . . . . . . . . . . . . . . . . . . . 42 4.10 DBSCAN Clustering (CLV) – Mean Values. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 4.11 DBSCAN Clustering (CLV) – Standard Deviations. . . . . . . . . . . . . . . . . . . . . . . . . 45 4.12 GMM Clustering (CLV) – Mean Values. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 4.13 GMM Clustering (CLV) – Standard Deviations. . . . . . . . . . . . . . . . . . . . . . . . . . . 47 4.14 Fuzzy C-Means Clustering (CLV) – Mean Values. . . . . . . . . . . . . . . . . . . . . . . . . . 49 4.15 Fuzzy C-Means Clustering (CLV) – Standard Deviations. . . . . . . . . . . . . . . . . . . . . 49 4.16 Silhouette Scores for RFM Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 4.17 Calinski–Harabasz Indices for RFM Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . 51 4.18 Silhouette Scores for CLV Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 4.19 Calinski–Harabasz Indices for CLV Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 VII 2.3. CLUSTERING ALGORITHMS FOR SEGMENTATION CHAPTER 2. LITERATURE REVIEW greater accuracy. By leveraging these tools, companies can move beyond traditional CLV calculations to incorporate variables such as social influence, channel preferences, and product affinity. 2.3 Clustering Algorithms for Segmentation While metrics like RFM and CLV define the dimensions of segmentation, clustering algorithms group customers based on these metrics. Over the past decades, unsupervised learning has become a practical solution for revealing hidden patterns in data without pre-labeled categories. 2.3.1 K-Means Clustering K-Means is among the most widely used clustering algorithms. It partitions datasets into clusters by minimizing intra-cluster variance. M. K. Pakhira et al.’s paper ”Validity Index for Crisp and Fuzzy Clusters” demonstrates its computational efficiency and applicability to large datasets [9]. However, K-Means assumes spherical cluster shapes and requires predefining the number of clusters (k), which may lead to suboptimal performance if the true data structure is non-spherical or unknown [10]. D. T. Pham et al. in ”Selection of K in K-Means Clustering” introduced methods for determining the optimal number of clusters, mitigating one of K-Means’ significant limitations [10]. The algorithm’s simplicity makes it ideal for initial exploratory analysis in segmentation studies. K-Means is particularly effective for high-volume retail data, where computational efficiency is paramount. However, it struggles with datasets containing noise or clusters of varying density. To address these limitations, hybrid approaches that combine K-Means with density-based methods like DBSCAN are increasingly being explored. 2.3.2 Hierarchical Clustering Hierarchical clustering builds a tree-like hierarchy of clusters and is particularly suited for datasets requiring nested groupings. A. D. Fallis’s article, ”Hierarchical Clustering Approaches for Large Scale Data,” explores its utility in uncovering macroand micro-segmentation [11]. By employing linkage methods such as Ward’s method, researchers minimize within-cluster variance at each step [12]. Although computationally intensive, hierarchical clustering provides visual insights through dendrograms, aiding in determining optimal cut points. One of the primary advantages of hierarchical clustering is its ability to reveal multi-level segment structures. For instance, in a retail dataset, it can identify broad customer categories such as ”frequent buyers” and ”occasional shoppers” while also uncovering subgroups within each category. However, its computational complexity limits its scalability for very large datasets. 5 2.3. CLUSTERING ALGORITHMS FOR SEGMENTATION CHAPTER 2. LITERATURE REVIEW 2.3.3 DBSCAN DBSCAN (Density-Based Spatial Clustering of Applications with Noise) excels at identifying arbitrarily shaped clusters and isolating outliers. M. Ester et al.’s foundational paper ”A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise” introduced this method, highlighting its adaptability [13]. DBSCAN does not require predefining the number of clusters, relying instead on two parameters: epsilon ( ϵ ) and the minimum number of points ( extitMinPts) [14]. Schubert et al. revisited DBSCAN in their paper ”DBSCAN Revisited, Revisited,” providing guidelines for parameter optimization and addressing challenges such as sparse data distributions [14]. DBSCAN’s strength lies in its ability to handle noise and outliers, making it ideal for retail datasets with irregular purchasing behaviors. For example, it can isolate high-value outliers such as corporate clients or bulk buyers, which might be misclassified in centroid-based methods. However, its performance depends heavily on parameter tuning, which requires domain expertise. 2.3.4 Gaussian Mixture Models (GMM) Gaussian Mixture Models assume that data arises from a mixture of Gaussian distributions. J. Han et al. in ”Data Mining: Concepts and Techniques” detailed GMM’s flexibility, which allows clusters to take on various shapes [15]. GMM uses the Expectation-Maximization algorithm to iteratively refine cluster parameters, making it suitable for datasets with overlapping or non-spherical distributions. However, its reliance on Gaussian assumptions may limit its effectiveness for highly skewed or multi-modal data. GMM is particularly suited for datasets where clusters exhibit significant overlap, such as customer groups with similar spending patterns but different product preferences. Its probabilistic framework allows for soft clustering, assigning each customer a likelihood of belonging to multiple clusters. 2.3.5 Fuzzy C-Means Fuzzy C-Means extends traditional clustering by assigning membership degrees to each cluster, accommodating blurry segment boundaries. J. C. Bezdek’s book ”Pattern Recognition with Fuzzy Objective Function Algorithms” introduced this concept, emphasizing its relevance for datasets where customers exhibit overlapping behaviors [16]. While advantageous for capturing subtle differences, Fuzzy C-Means requires careful tuning of parameters like the fuzzifier, which can complicate its application. The flexibility of Fuzzy C-Means makes it particularly valuable for customer segmentation in industries with diverse product offerings. For instance, in e-commerce, customers might exhibit characteristics of both ”bargain hunters” and ”premium buyers.” By allowing partial membership, Fuzzy C-Means provides insights into hybrid customer profiles, enabling more personalized marketing strategies. 6 2.4. LEVERAGING CLV AND RFM: CLUSTERING CHAPTER 2. LITERATURE REVIEW 2.4 Leveraging CLV and RFM: Clustering Empirical studies advocate combining CLV metrics with advanced clustering algorithms for a multidimensional understanding of customer behavior. Khajvand et al.’s study demonstrated how CLV-weighted RFM scores enhance segmentation by identifying high-lifetime-value customers with sporadic activity [4]. This aligns with Gupta et al.’s concept of persistence models, where potential future value guides marketing strategies [6]. DBSCAN’s ability to isolate outliers has proven effective for identifying “sleepers” or occasional high spenders who might otherwise be overlooked by centroid-based methods [5]. Real-time clustering approaches are gaining traction, adapting dynamically to updated transactions, returns, or seasonal variations [15]. Integrating CLV and RFM with clustering algorithms provides a holistic view of customer behavior. For example, a company might use Fuzzy C-Means to identify hybrid profiles, combining this insight with CLV metrics to prioritize high-value segments. Similarly, DBSCAN can uncover hidden patterns in noisy datasets, while GMM offers probabilistic insights into overlapping customer behaviors. 2.5 Conclusion The literature illustrates a clear progression in segmentation strategies, from simple RFM constructs to complex CLV-enhanced models and advanced clustering algorithms. While RFM retains its appeal for simplicity, its integration with CLV enables a deeper differentiation of customers, identifying not only frequent spenders but also those with significant future value. The choice of clustering algorithm is context-dependent: K-Means for scalability, Hierarchical for multi-level grouping, DBSCAN for outlier detection, GMM for overlapping clusters, and Fuzzy C-Means for accommodating blurred segment boundaries. By strategically combining these methodologies, practitioners can derive segments that are both actionable for immediate use and predictive of future behavior. 7 Comparative Study of Customer Segmentation Strategies Based on Business Analytics Chapter 3 Methodology The methodology of the thesis follows a structured, sequential approach to process and analyze the dataset for customer segmentation using advanced clustering techniques [1]. Below is an outline of the methodology as illustrated in the figure 3.1 •Dataset: – The analysis begins with a dataset sourced from the UCI Machine Learning Repository, containing transactional data from an online retailer. •Preprocessing: –The dataset undergoes preprocessing to clean and refine the data. •Segmentation: –The refined data is segmented using two approaches: ∗ RFM (Recency, Frequency, Monetary) segmentation: Assigning scores to customers based on their purchase behavior [3]. Figure 3.1: The structure of the methodology 8 3.1. DATA COLLECTION CHAPTER 3. METHODOLOGY ∗ CLV (Customer Lifetime Value) calculation: Estimating the long-term value of each customer [4]. •Clustering Models: –Both RFM and CLV segmentations are independently subjected to clustering using five models 1. K-means 2. Hierarchical clustering 3. DBSCAN 4. Gaussian Mixture Models (GMM) 5. Fuzzy C-means •Model Evaluation: –Each clustering solution is evaluated using two metrics: ∗Silhouette Score: Measures cluster cohesion and separation. ∗ Calinski-Harabasz Index: Assesses the ratio of between-cluster to within-cluster variance [9]. – The evaluation compares the performance of clustering models for both RFM and CLV segmentations. 3.1 Data Collection The dataset used in this thesis was sourced from the UCI Machine Learning Repository. It consists of transactional records from a UK-based, non-store online retailer, spanning from December 1, 2010, to December 9, 2011. This period includes the critical holiday shopping season, providing valuable insights into consumer behavior during peak retail activity. Dataset Description: The dataset includes 541,909 transactions, detailed across eight attributes that offer both multivariate and time-series insights into retail operations and customer demographics: •InvoiceNo: Transaction identifier; if prefixed with ’C’, indicates a cancellation. •StockCode: Unique product code. •Description: Text description of the product. •Quantity: Number of units purchased in each transaction. •InvoiceDate: Timestamp of each transaction (e.g., ”12/1/2010 8:26”). •UnitPrice: Price per unit of the product. •CustomerID: Unique identifier for each customer. 9 3.2. SEGMENTATION METHODOLOGY CHAPTER 3. METHODOLOGY •Country: Country of the customer. 3.1.1 Preprocessing and Data Integrity: Critical preprocessing steps were undertaken to ensure the data’s reliability for subsequent analysis: • Data Cleaning: Duplicate records were removed to ensure data uniqueness. Transactions identified as cancellations, through InvoiceNo prefixes, were segregated to prevent distortion of purchasing patterns. • Date Parsing: The ’InvoiceDate’ field was converted from string format into a date-time object to facilitate time-series analyses. • Error Handling: Entries without CustomerID or with unrealistic transaction values (e.g., negative prices or quantities) were scrutinized and handled appropriately. Finally, a column for the ”TotalPrice” was added to the dataset. This was calculated by multiplying the quantity of each item by its unit price. These step was important to obtain the Monetary Value of each transaction. Post-preprocessing, the dataset was streamlined to 392692 records distributed across nine columns (the original eight plus the new one of ”TotalPrice”). This refined dataset was free from the common data issues that could undermine the validity of the analysis, thus setting a strong foundation for the subsequent segmentation. 3.2 Segmentation Methodology 3.2.1 RFM Framework The Recency, Frequency, Monetary (RFM) model is a customer segmentation technique widely used in database marketing and retail analytics. This model evaluates customers by assigning a score based on three specific criteria: • Recency (R): Recency measures how recently a customer made a purchase. A lower recency value indicates that the customer bought from the store or business more recently, which suggests higher engagement and likelihood of repeat purchases. Recency is calculated by subtracting the date of a customer’s last purchase from the current date, often expressed in days: R= Current Date −Last Purchase Date • Frequency (F): Frequency indicates how often a customer makes a purchase within a given timeframe. Businesses often track the total number of transactions each customer has completed during a specific period to determine loyalty and engagement levels. Frequent interactions are typically a sign of a 10 3.2. SEGMENTATION METHODOLOGY CHAPTER 3. METHODOLOGY customer’s trust and satisfaction with the brand: F= Total number of transactions • Monetary (M): This metric assesses the total amount of money a customer has spent over time. Higher monetary values are an indicator of higher customer value to the organization. Companies often use this dimension to identify their most valuable customers or ’high spenders’ who are likely to contribute a significant portion of revenue: M= TotalPrice = Unit Price ×Quantity Applying the RFM model allows businesses to develop differentiated marketing strategies and personalize communications. This approach not only enhances customer satisfaction but also optimizes marketing efforts by focusing on the most lucrative segments. 3.2.2 RFM Score Calculation and Segmentation To segment the customer dataset effectively, each RFM metric was quantified and scored based on quintiles. Each customer received a score from 1 to 5 for each metric, where a score of 5 represents the top 20% of behavior (e.g., most recent purchases, highest frequency, and highest spending). • Customers were ranked based on each metric, and the ranks were then divided into five equal groups or quintiles. The quintiles helped in assigning a scaled score for each RFM parameter. • Recency scores were inverted, where customers who purchased more recently scored higher (i.e., a customer with the most recent purchase gets a score of 5). •Frequency and Monetary values were scored normally, where higher values received higher scores. This method allows for the identification of distinct groups based on their transactional behavior, enabling targeted marketing strategies. Defining Customer Segments Based on the calculated RFM scores, customers were classified into six segments. This classification helps in tailoring marketing efforts according to the specific characteristics of each segment: So based on the previous table we got the following segment • Whales: These are top-tier customers with recent purchases, high transaction frequency, and significant spending, necessitating focused retention strategies and exclusive promotions. • Active Stars: Similar to Whales in behavior but slightly less intense in their interactions, they are pivotal due to their considerable expenditure and frequent purchases. 11 3.2. SEGMENTATION METHODOLOGY CHAPTER 3. METHODOLOGY Table 3.1: Customer Segmentation Based on RFM Scores R score F score M score Segment 5 5 5 Whales ≥4≥3≥4 Active Stars ≥4≥4≤3 Loyal Regulars ≤3≥3≥4 Sleeping Giants ≥4≤3≤5 New & Occasional Buyers Other cases Lost Clients • Loyal Regulars: Regular and recent purchasers with moderate spending, forming a stable revenue base and can be targeted with promotions to increase spending. • Sleeping Giants: Previously high-value customers now inactive, identified as potential sources of revenue if re-engaged effectively. • New & Occasional Buyers: Customers with recent but infrequent purchases, holding potential for development into regular buyers through effective retention strategies. • Lost Clients: The least engaged and historically low-spending, often not prioritized unless specific re-engagement strategies are feasible. 3.2.3 Customer Lifetime Value (CLV) The concept of Customer Lifetime Value (CLV) is another way of segmenting customers to understand the long-term value of a customer to a business. It helps in determining how much a company should invest in maintaining relationships with existing customers and in acquiring new ones [6]. CLV represents the total revenue that a company can expect from a customer throughout their business relationship. The value is calculated based on the profit margin of the transactions, factoring in the retention rate and discount rate to account for the time value of money. Understanding the CLV helps businesses focus their marketing efforts on customers who are likely to deliver the highest lifetime value. Methodology for Calculating CLV The calculation of CLV in this study follows a defined series of steps, using the cleaned and preprocessed dataset described in previous chapters. Here, we elaborate on each step involved in the calculation: Calculation of Key Metrics Three primary metrics are calculated for each customer: • Total Revenue: This is the sum of the TotalPrice for all transactions per customer, giving a cumulative value of how much the customer has spent. It is a direct multiplication of the Quantity and UnitPrice for each item purchased, summed across all transactions. 12 3.2. SEGMENTATION METHODOLOGY CHAPTER 3. METHODOLOGY • Frequency: This metric measures the number of distinct invoices per customer, indicating how often the customer engages in transactions. It serves as an index of customer loyalty and purchasing frequency. • Average Transaction Value: Calculated as the average TotalPrice per transaction for each customer. This metric provides insight into the spending behavior of the customer per transaction. Lifetime Calculation The lifespan of the customer relationship is computed by determining the number of days between the first and last purchase dates. This metric provides a temporal dimension to the monetary and frequency values, offering a more comprehensive view of the customer’s engagement over time. Computation of CLV The CLV is then calculated using the formula: CLV = Frequency ×Average Transaction Value ×Expected Lifespan Where: •Frequency is the number of transactions (as calculated earlier). •Average Transaction Value is the mean spending per transaction. • Expected Lifespan is an estimated duration of the customer relationship in months. This duration can be adjusted based on historical data or industry averages. This formula integrates both behavioral (Frequency, Average Transaction Value) and temporal (Expected Lifespan) aspects to provide a comprehensive estimate of the customer’s total potential value to the business. Role of CLV in Customer Segmentation After calculating the CLV, this metric becomes a cornerstone for further segmentation analysis. It allows businesses to categorize customers into segments based on their value, enabling more targeted and effective marketing strategies. High-CLV customers can be identified for premium services, while strategies for improving the CLV of lower-scoring customers can be developed. Integrating CLV with the previously discussed RFM segmentation provides a robust framework for understanding customer behavior in a multidimensional space, where each metric offers unique insights into customer loyalty, spending, and engagement. The calculated CLV not only informs about the past and present value of the customers but also helps predict future interactions and profitability, guiding strategic decisions in customer relationship management and marketing. This systematic approach to calculating and utilizing CLV ensures that the segmentation and targeting are grounded in economic reality, aiming to maximize the return on investment in customer relationships. 13 3.3. CLUSTERING ALGORITHMS CHAPTER 3. METHODOLOGY 3.3 Clustering Algorithms 3.3.1 K-means Clustering Clustering of K-means is an influential unsupervised machine learning technique that is used to group similar data points into a predetermined number of clusters, denoted as k. This method is particularly effective in customer segmentation, aiding businesses in discerning the structure within their customer base and enabling targeted marketing strategies [10]. Algorithm Overview The K-means algorithm organizes n observations into k clusters where each observation belongs to the cluster with the closest mean, serving as the cluster’s prototype. Initially, k centroids are selected randomly from the dataset. Subsequently, each data point is assigned to the nearest centroid based on the Euclidean distance, and the centroids are recalculated as the mean of all points in the cluster. This process iterates; points are reassigned and centroids updated until the centroids stabilize and exhibit minimal or no movement, indicating convergence. Mathematical Formulation The goal of K-means is to minimize the within-cluster sum of squares (WCSS), which represents the sum of squared distances between each point and its respective centroid, mathematically expressed as: Minimize WCSS = k X i=1 X x∈Si ∥x−µi∥2 where µidenotes the mean of points in Si, and kis the number of clusters. Choosing the Number of Clusters The ’Elbow Method’ is commonly employed to determine the optimal number of clusters [10]. It involves executing the K-means algorithm across a range of kvalues and plotting the WCSS for each. The optimal k typically corresponds to the point where the WCSS curve levels off, creating an elbow shape. This method is useful because it provides a visual cue to the point at which increasing the number of clusters ceases to result in significantly lower within-cluster variation. The Elbow Method is particularly effective when the decrease in WCSS becomes negligible, suggesting that adding more clusters might lead to overfitting without substantial gains in capturing distinct groupings. Algorithm Steps 1. Initialization: Select kinitial centroids randomly from the data points. 2. Assignment: Assign each data point to the nearest centroid, determined by the Euclidean distance. 3. Update: Recalculate centroids as the mean of the points in each cluster. 14 3.4. ANALYTICAL TOOLS AND ENVIRONMENT CHAPTER 3. METHODOLOGY out. In our context, we chose m = 2 because it offers a reasonable trade-off: we wanted to allow for partial membership in multiple segments but still retain relatively distinct cluster centers for clearer interpretation. Interpreting Results In fuzzy clustering, each customer (or data point) is assigned a vector of membership values, indicating its degree of belonging to each cluster. Practical approaches include: • Using Fuzzy Memberships Directly: Analyze the soft memberships to explore nuanced or hybrid customer profiles. This can reveal subtle overlaps where customers exhibit characteristics of multiple segments. • Hard Clustering Conversion: For simpler comparison with other clustering approaches (e.g., k - means), one may convert fuzzy memberships to a single cluster label by assigning each point to the cluster with the highest membership. While this loses some granularity, it simplifies subsequent analyses and visualizations. By accommodating partial memberships, FCM enriches the understanding of customer behaviors, identifying overlapping tendencies that might not be captured by strictly partitioning methods. This can be particularly beneficial when performing RFM segmentation, where spending and purchase frequency often straddle boundaries, suggesting that customers naturally fit into more than one archetypal cluster. Implementation in R The e1071 package in R implements Fuzzy C-means clustering through the cmeans() function, allowing users to specify parameters such as the fuzzification exponent ( m ), the number of cluster centers ( centers ), maximum iterations ( iter.max ), and convergence thresholds. This function automatically computes both the centroids and the membership matrix, where each data point retains partial membership across multiple clusters. 3.4 Analytical Tools and Environment The analysis presented in this thesis was conducted using R, a language and environment for statistical computing and graphics. The specific version used was R version 4.3.1 (2023-06-16 ucrt), running on a Windows 11 platform with an x86 64-w64-mingw32/x64 (64-bit) architecture. The analysis was performed using RStudio, an integrated development environment for R, version 2024.09.1+394 ”Cranberry Hibiscus”. Libraries and Packages Several additional packages were employed to support the analysis, each chosen for its specific features that aid in data manipulation, visualization, and clustering analysis. The URLs to their Comprehensive R Archive Network (CRAN) pages. 21 3.5. MODEL EVALUATION CHAPTER 3. METHODOLOGY • dplyr (version 1.1.3): A grammar of data manipulation, providing a consistent set of verbs for tackling common data-handling challenges. URL: https://CRAN.R-project.org/package=dplyr • lubridate (version 1.9.2): Simplifies date and time parsing and manipulation in R, offering intuitive handling of time-series data. URL: https://CRAN.R-project.org/package=lubridate • ggplot2 (version 3.5.1): A system for declaratively creating graphics, based on The Grammar of Graphics, offering flexible and layered visualizations. URL: https://CRAN.R-project.org/package= ggplot2 • plotly (version 4.10.4): An interface to the Plotly JavaScript graphing library, enabling interactive, web-based data visualizations. URL: https://CRAN.R-project.org/package=plotly • cluster (version 2.1.4): Implements methods for cluster analysis, including agglomerative hierarchical clustering, essential for customer segmentation. URL: https://CRAN.R-project.org/package= cluster • dbscan (version 1.2-0): A fast reimplementation of the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm. URL: https://CRAN.R-project.org/package=dbscan • mclust (version 6.1.1): Provides model-based clustering using finite Gaussian mixture models, with automatic model selection based on BIC. URL: https://CRAN.R-project.org/package=mclust • e1071 (version 1.7-13): Includes fuzzy clustering ( cmeans ) and additional machine learning tools such as Support Vector Machines. URL: https://CRAN.R-project.org/package=e1071 • reshape2 (version 1.4.4): Facilitates reshaping and melting of data frames, aiding in data transformation prior to analysis or visualization. URL: https://CRAN.R-project.org/package=reshape2 • fpc (version 2.2-13): Contains clustering methods and validation tools, complementing DBSCAN with diagnostic and plotting functions. URL: https://CRAN.R-project.org/package=fpc These tools and libraries were integral to the execution of the analysis, enabling data manipulation, clustering, and visualization. Their explicit inclusion, alongside version numbers and URLs, ensures the transparency and reproducibility of the research. 3.5 Model Evaluation The clustering solutions for both the RFM and CLV datasets were evaluated using two internal validation metrics: the Silhouette Score and the Calinski–Harabasz Index (CH). These metrics were chosen for their ability to measure both cluster compactness and separation, providing a robust evaluation framework. 3.5.1 Silhouette Score The Silhouette Score evaluates the quality of a clustering solution by assessing how similar each data point is to points in its own cluster (cohesion) relative to points in the nearest other cluster (separation). This 22 3.5. MODEL EVALUATION CHAPTER 3. METHODOLOGY metric is particularly useful for understanding the internal structure of the clusters and identifying potential misclassifications. [15]. Mathematical Definition: For a given point i, the Silhouette Score S(i) is defined as: S(i) = b(i)−a(i) max(a(i), b(i)), where: •a(i): The average distance between iand all other points in the same cluster (intra-cluster distance). •b ( i ): The average distance between i and all points in the nearest neighboring cluster (inter-cluster distance). The overall Silhouette Score is computed as the mean of S(i) for all points: S=1 n n X i=1 S(i), where nis the total number of points. Interpretation: •S(i)≈1: The point is well-clustered. •S(i)≈0: The point lies on the boundary between clusters. •S(i)<0: The point may be misclassified. Special Handling: • DBSCAN: Points labeled as noise ( − 1) were excluded from the calculation, as they do not belong to any valid cluster. • Fuzzy C-means: The fuzzy memberships were converted to hard cluster assignments by assigning each point to the cluster where its membership value was highest (arg max). 3.5.2 Calinski–Harabasz Index The Calinski–Harabasz Index evaluates clustering quality by measuring the ratio of the between-cluster variance to the within-cluster variance. It rewards clustering solutions with well-separated and compact clusters. Mathematical Definition: The CH index is defined as: CH =Between-cluster variance/(k−1) Within-cluster variance/(n−k), where: 23 3.5. MODEL EVALUATION CHAPTER 3. METHODOLOGY •k: The number of clusters. •n: The total number of points. •Between-cluster variance: B= k X j=1 nj∥µj−µ∥2, where nj is the size of cluster j , µj is the centroid of cluster j , and µ is the overall mean of the dataset. •Within-cluster variance: W= k X j=1 X x∈Cj ∥x−µj∥2, where Cjdenotes the points in cluster j. Interpretation: • Higher CH values indicate better clustering solutions, with more compact clusters and greater separation between clusters. Special Handling: • DBSCAN: If the algorithm identifies only one cluster or assigns all points as noise, the CH index is undefined. • Fuzzy C-means: Similar to the Silhouette Score, the fuzzy memberships were converted to hard assignments before calculating the CH index. 24 Comparative Study of Customer Segmentation Strategies Based on Business Analytics Chapter 4 Results 4.1 RFM Segmentation & Clustering This subsection presents the findings of the RFM segmentation and subsequent clustering. 4.1.1 RFM Segmentation Results Segment Distribution The segmentation process categorized the customer base into six distinct groups: Whales, Active Stars, Loyal Regulars, New & Occasional Buyers, Sleeping Giants, and Lost Clients. These segments were defined based on thresholds for recency, frequency, and monetary value derived from the RFM scores. The percentage distribution of customers across these segments is summarized in figure 4.1. The results reveal significant disparities in customer behavior: • Lost Clients form the largest segment, comprising 47.12% of the total customer base. This group demonstrates the longest recency and the lowest frequency and monetary metrics, indicating minimal recent engagement and low revenue contributions. • Active Stars represent 15.65% of the customers, showing moderate recency and frequency coupled with higher monetary value, making them a valuable segment for engagement. • Whales comprise 8.02% of customers but are distinguished by their high frequency (18.2 transactions) and monetary value (€11,222 on average), representing the most valuable segment. • Smaller groups like Loyal Regulars (4.03%) and New & Occasional Buyers (12.26%) highlight stable or emerging purchasing patterns. • Sleeping Giants (12.91%) are characterized by moderate monetary contributions but longer periods of inactivity. 25 4.1. RFM SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.1: Distribution of RFM Segments Analysis Segment Statistics The key statistics for each segment are provided in Table 4.1. The analysis of recency, frequency, and monetary metrics offers deeper insights into customer behavior: • Whales exhibit the shortest RecencyMean (5.93 days), highlighting their recent activity and strong brand loyalty. • In contrast, Lost Clients show the longest RecencyMean (161 days) and minimal spending, signaling the need for re-engagement strategies or exclusion from targeted campaigns. • Active Stars have relatively high MonetaryMean (€3,156) and moderate frequency (6.52 transactions), making them a stable yet growing segment. • Sleeping Giants demonstrate potential with moderate FrequencyMean (5.15 transactions) and significant MonetaryMean (€2,656), suggesting the need for tailored campaigns to reactivate their purchasing habits. Segment R Mean R Median F Mean F Median M Mean M Median Size Active Stars 17.10 17.80 6.52 5 3156 2500 679 Lost Clients 161.00 138.00 1.60 1 495 450 2044 Loyal Regulars 15.90 16.00 4.06 4 671 520 175 New & Occasional 17.60 18.50 1.66 2 464 390 532 Sleeping Giants 88.20 68.00 5.15 4 2656 2400 560 Whales 5.93 4.94 18.20 13 11222 9500 348 Table 4.1: Segment Statistics for RFM Analysis 26 4.1. RFM SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS (a) RFM log R (b) RFM log M (c) RFM log F Figure 4.2: Logarithmic boxplots of RFM indicators 4.1.2 Clustering Results on RFM Segmentation The clustering analysis on the RFM data was performed using five distinct algorithms: K-Means, DBSCAN, Hierarchical Clustering, Gaussian Mixture Model (GMM), and Fuzzy C-Means (FCM). Each method was selected based on its ability to uncover patterns in the data and provide meaningful customer segmentation. The clustering parameters were determined based on specific evaluation techniques (e.g., the Elbow Method, k-NN Distance Plot, BIC). Below, we detail the clustering results, including statistical summaries, visualizations, and the rationale for parameter selection. K-Means Clustering Results The K-Means clustering algorithm was applied to the normalized RFM data to segment the customer base into distinct groups. The optimal number of clusters (k) was determined using the Elbow Method. 27 4.1. RFM SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.3: Elbow method graph Elbow Method for Optimal k :The Elbow Method graph, depicted in Figure 4.3, identifies k = 4 as the optimal number of clusters. At this value, the Within Sum of Squares (WSS) shows a significant reduction compared to k = 3, but the rate of decrease diminishes beyond k = 4. This choice strikes a balance between minimizing WSS and avoiding excessive model complexity. So the K-Means algorithm divided the customers into four clusters. The detailed cluster statistics are shown in Table 4.2. Cluster R Mean R Median F Mean F Median M Mean M Median Size 1 16.0 5.97 22.1 19 12510 7931 209 2 7.66 2.03 82.5 63 127338 117380 13 3 249.0 244.0 1.55 1 478 310 1061 4 44.5 33.0 3.66 3 1353 826 3055 Table 4.2: Cluster Statistics for K-Means Clustering. Visualization: The distribution of customers across clusters is visualized in Figure 4.4, showing the Recency and Frequency dimensions. The Monetary value is represented by the size of the points, highlighting the high value of Cluster 2. 28 4.1. RFM SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.4: K Means 2D graph Insights: • Cluster 2 consists of high-frequency customers with recent transactions and high monetary contributions, making it a prime target for retention strategies. The total monetary value across all clusters was calculated as approximately 8,910,557€. Of this, Cluster 2 contributed 1,655,394€, representing 18.58% of the total. This significant contribution from only 13 individuals highlights the exceptional spending behavior of this small segment. • Cluster 1 and Cluster 4 represent moderate-frequency and monetary customers with differing recency profiles. • Cluster 3, the largest, includes infrequent and low-value customers, emphasizing the need for targeted engagement strategies to convert them into higher-value segments. DBSCAN Clustering Results The Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm was applied to the normalized RFM data to identify clusters of varying shapes and densities. Unlike K-Means, DBSCAN does not require the specification of the number of clusters beforehand and is particularly effective at identifying noise points. However, it is sensitive to the parameters eps (epsilon) and MinPts. Parameter Selection for DBSCAN: The k-NN distance plot, shown in Figure 4.5, was used to determine an appropriate value for eps . A value of eps = 0.45 was selected, as this corresponds to the point where the k-NN distances begin to rise sharply. MinPts was set to 5 (as already explained in the Methodology). Cluster Statistics: DBSCAN identified three groups, including one noise cluster ( Cluster 0 ). Detailed statistics for each cluster are provided in Table 4.3. 29 4.1. RFM SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.5: k-NN Distance Plot for DBSCAN Parameter Selection (eps = 0.45). Cluster R Mean R Median F Mean F Median M Mean M Median Size 0(Noise) 51.1 8.98 39.5 31 48452 30358 62 1 93.8 52.0 3.73 2 1376 659 4272 2 3.27 3.05 38 38 7413 7065 4 Table 4.3: Cluster Statistics for DBSCAN Clustering. Visualization: The figure 4.6 visualizes the DBSCAN clustering results. Insights: • Cluster 0 (Noise): Surprisingly, the noise cluster contains 62 customers with exceptionally high monetary values (Monetary Mean = 48452€). These customers likely represent outliers or high-value clients who do not conform to the dense regions of the RFM space identified by DBSCAN. This raises concerns about the algorithm’s ability to handle such critical data points effectively, as classifying important customers as noise undermines the segmentation’s strategic value. • Cluster 1: This cluster contains the majority of customers (Group Size = 4272) with low frequency and monetary values. It represents low-value, infrequent customers, aligning with expectations for this portion of the dataset. 30 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS !h Figure 4.12: Fuzzy C-Means Clustering (Hard Assignments) Insights: • Cluster 1: encompasses the largest group of customers, with relatively average Recency and low Frequency and Monetary values, indicating moderately engaged customers. • Cluster 2: contains less frequent and less recent buyers with low Monetary contributions, potentially representing disengaged or dormant customers. • Cluster 3: includes high-value customers with frequent transactions and low Recency, aligning with characteristics of highly engaged and valuable clients. • Cluster 4: represents the least engaged customers with very low Frequency and Monetary values, typically corresponding to long-time inactive clients. This marks the conclusion of the RFM segmentation and the associated clustering processes. The next section will address CLV-based segmentation, where the same clustering methods will be applied. Since the methodology for obtaining the results has been thoroughly detailed in this section, the focus will shift to the interpretation and commentary of the outcomes. 4.2 CLV Segmentation & Clustering 4.2.1 CLV Segmentation In this section, we shift our focus from the RFM-based approach to a Customer Lifetime Value (CLV) perspective. By incorporating variables such as TotalRevenue,Frequency,Average Transaction Value,Lifespan, and the derived CLV metric, we aim to capture a more comprehensive view of each customer’s long-term worth. The following subsections detail the clustering analysis performed on these CLV-driven features, providing insights into customer segments that can guide strategic marketing, retention, and acquisition 37 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS efforts. 4.2.2 Clustering Results on CLV Segmentation K-Means Clustering Results The K-Means algorithm was applied to the normalized dataset containing TotalRevenue,Frequency,Average Transaction Value,Lifespan, and CLV. As in the RFM analysis, the optimal number of clusters k was determined using the Elbow Method. Elbow Method for Optimal k :Figure 4.13 illustrates the Elbow Method, where the Within Sum of Squares (WSS) was plotted against the number of clusters. A marked bend in the plot at k = 4 indicates that increasing the number of clusters beyond four yields marginal gains in variance reduction. Consequently, k= 4 was selected as the optimal solution. Figure 4.13: Elbow Method for K-Means (CLV Data) Cluster Statistics: The K-Means algorithm thus partitions the customer base into four clusters, each exhibiting distinct behaviors in terms of revenue generation, purchase frequency, and lifetime value. Table 4.7 provides an overview of the key metrics for each cluster. 38 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS kC TR M F M CLV M AT M LS M Size 1 2915 7.08 70,932 33.3 274 1753 2 122,828 1.50 11,512,288 66,671 102 2 3 84,711 72.30 3,396,779 192 363 23 4 628 1.74 2,369 39.2 30.9 2560 Table 4.7: K-Means Cluster Statistics for CLV-Based Analysis Legend: kC K-Means Clustering TR Total Revenue (€) FFrequency CLV Customer Lifetime Value AT Average Transaction Value LS Lifespan (days) MMean Visualization: Figure 4.14 presents a 2D scatter plot of the clusters in terms of Frequency and TotalRevenue, with point size reflecting the CLV magnitude. Additionally, Figure 4.15 shows the radar chart of K-Means. A radar chart is a graphical method used to visualize multivariate data. Each variable is represented by an axis emanating from the center, and the data points are plotted on these axes to form a polygon. In this context, radar charts display the normalized averages of key metrics (e.g., Total Revenue, Frequency, CLV) for each cluster, offering a compact view of their unique characteristics. Figure 4.14: K-Means Clustering on CLV Data (2D Projection) 39 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.15: K-Means Radar Chart Insights: • Cluster 2 (only 2 customers) exhibits extremely high TotalRevenue and AvgTransactionValue, resulting in the largest CLV by far. These represent “ultra-premium” clients, where personalized retention and upselling strategies can be highly impactful. • Cluster 3 stands out for its elevated Frequency (over 70 purchases on average) and a high CLV, suggesting a loyal and active customer group. Fostering loyalty programs and subscription models could further cement these relationships. • Cluster 1 contains moderately active customers with a notable lifetime value, though significantly lower than Cluster 2 or 3. Targeted campaigns can aim to increase their purchase frequency or upsell to boost their AvgTransactionValue. • Cluster 4 holds the largest volume of customers (2,560), but with low TotalRevenue,CLV, and Frequency. It could encompass one-time or infrequent buyers. Re-engagement or cross-selling strategies may help convert a portion of this large group into higher-value segments. Having outlined the CLV-based K-Means clustering, the following subsections will compare these findings with alternative clustering methods (Hierarchical, DBSCAN, GMM, and Fuzzy C-Means) to further validate or refine the segmentation strategy. 40 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Hierarchical Clustering Results The Hierarchical Clustering approach was applied to the same CLV-based dataset using Ward’s minimum variance method ( ward.D2 ). This algorithm recursively merges clusters to minimize the total within-cluster variance, producing a dendrogram that illustrates the nested structure of the data. Determination of the Number of Clusters: A visual inspection of the dendrogram (Figure 4.16) was used to select the cutoff height, resulting in k= 4 clusters. Rectangular boundaries drawn on the dendrogram confirm this selection, balancing interpretability against cluster granularity. Figure 4.16: Hierarchical Clustering Dendrogram (CLV Data) Cluster Statistics (Means): Table 4.8 provides an overview of each cluster’s mean values for the key variables. The abbreviations are explained in the legend below the table. hC TR M F M CLV M AT M LS M Size 1 122,828 1.50 11,512,288 66,671 102 2 2 2,586 6.63 58,597 29.7 261 1,941 3 571 1.63 1,637 34.5 21.6 2,364 4 74,050 58.2 2,933,135 772 329 31 Table 4.8: Hierarchical Clustering (CLV) – Mean Values. Legend: hC Hierarchical Cluster 41 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Cluster Statistics (Standard Deviations): To further assess variability within each cluster, Table 4.9 shows the corresponding standard deviations for the same metrics. hC TR SD F SD CLV SD AT SD LS SD 1 64,551 0.71 16,280,833 14,868 145 2 3,304 5.80 137,274 75.4 72.7 3 747 1.07 7,167 129 36.4 4 65,367 48.8 3,529,243 2,458 91.2 Table 4.9: Hierarchical Clustering (CLV) – Standard Deviations. Visualization: Figure 4.17 shows a 2D projection of the four hierarchical clusters (color-coded by cluster membership), plotted against Frequency and TotalRevenue, with the point size reflecting the CLV. Additionally, Figure 4.18 shows the radar charts of the Hierarchical Clustering Figure 4.17: Hierarchical Clustering on CLV Data (2D Projection) 42 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.18: hC Radar Charts Insights: •hC 1: With only 2customers, it shows extremely high TotalRevenue and CLV, similar to the “ultrapremium” cluster identified in K-Means. Retention and bespoke marketing can amplify the value of these elite clients. • hC 2: The largest cluster (1,941 customers) with moderate revenue and frequency. Although their CLV is relatively modest, an upsell or cross-sell approach could unlock further potential. • hC 3: A large group of low-value and infrequent buyers. This segment may be harder to convert, but re-engagement campaigns or targeted promotions might reactivate some portion. • hC 4: A niche but very active cluster (F Mean ≈ 58.2). Despite generating substantial revenue ( ≈ 74 , 050), they remain far behind the top-tier cluster in terms of CLV. Strengthening loyalty programs could increase their average transaction value. These findings mirror certain patterns seen in the K-Means segmentation, albeit with variations in cluster size and boundaries. The manual selection of four clusters introduces a degree of subjectivity, yet the hierarchical approach offers a more intuitive view of segment cohesion and separation via the dendrogram. DBSCAN Clustering Results The DBSCAN algorithm takes a density-based approach, identifying dense regions in the CLV feature space while designating low-density points as noise. Unlike centroid-based methods, DBSCAN automatically determines the number of clusters based on the parameters eps and MinPts. Parameter Selection: Ak-NN distance plot was generated to guide the choice of eps . After observing in the figure 4.19 a sharp increase in the distance values near eps = 0.6 , this threshold was adopted, with 43 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.19: k-NN Distance Plot for DBSCAN Parameters Selection (eps = 0.6) MinPts set to 5. The algorithm yielded 2main clusters in the context of CLV data. Cluster Statistics (Means): Table 4.10 provides the mean values for each cluster, employing the same notation used previously. dbC TR M F M CLV M AT M LS M Size 0 39,506 35.2 1,746,547 2,184 296 81 1 1,342 3.68 21,162 28.1 128 4,257 Table 4.10: DBSCAN Clustering (CLV) – Mean Values. Legend: dbC DBSCAN Cluster Cluster Statistics (Standard Deviations): Table 4.11 reports the standard deviations for each metric, indicating the dispersion within each cluster. 44 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS dbC TR SD F SD CLV SD AT SD LS SD 0 52,305 36.9 3,448,736 10,585 113 1 1,928 4.03 56,020 53.4 131 Table 4.11: DBSCAN Clustering (CLV) – Standard Deviations. Visualization: The 2D plot in Figure 4.20 illustrates the clusters in terms of Frequency ( x -axis) and TotalRevenue (y-axis), with the size of each point corresponding to its CLV. Figure 4.20: DBSCAN Clustering on CLV Data (2D Projection) Insights: • Cluster 0: Comprises only 81 customers but shows significantly higher Frequency,TotalRevenue, and CLV on average. The standard deviations (TR SD and CLV SD in particular) are notably large, indicating a broad range of spending patterns within this high-value segment. • Cluster 1: The vast majority of customers (over 4,000), marked by modest revenue and low frequency. Although their CLV remains comparatively small, targeted marketing campaigns could potentially lift a subset of these customers into higher-value brackets. • Noise Points: DBSCAN typically designates isolated or less dense regions as noise. In this dataset, however, most customers fall into one of the two clusters, suggesting that the chosen parameters effectively captured the primary data structure. 45 4.2. CLV SEGMENTATION & CLUSTERING CHAPTER 4. RESULTS Figure 4.21: DBSCAN Radar Charts In summary, DBSCAN identifies a small but highly valuable group of buyers (Cluster 0) alongside a large, lower-value segment (Cluster 1). The elevated SD metrics in Cluster 0 indicate diverse spending habits, meriting a closer look at sub-segmentation or personalized offers for high-spending individuals. Gaussian Mixture Model (GMM) Clustering Results The Gaussian Mixture Model (GMM) employs a probabilistic framework, assuming that data points originate from a mixture of Gaussian distributions. This flexibility enables GMM to capture clusters of various shapes and densities, often yielding more nuanced segmentations compared to strictly distance-based approaches. Model Fitting: In this analysis, the Mclust package was utilized to automatically determine the optimal number of components through the Bayesian Information Criterion (BIC). The resulting best-fit model identified 6clusters in the CLV-based feature space. Cluster Statistics (Means): Table 4.12 presents mean values for each of the six GMM clusters, using the abbreviated notation described in earlier sections. gC TR M F M CLV M AT M LS M Size 1 304 1.00 0 19.6 0 1,393 2 35,999 26.1 1,750,038 2,513 208 75 3 643 2.00 3,721 17.6 104 730 4 5,705 10.5 156,058 102 217 333 5 1,211 4.60 12,095 13.4 204 1,083 6 2,911 7.23 44,922 27.1 252 724 Table 4.12: GMM Clustering (CLV) – Mean Values. Legend: gC GMM Cluster Cluster Statistics (Standard Deviations): Table 4.13 displays the standard deviations, indicating how dispersed each metric is within every cluster. Notably, gC 2 has a high CLV SD, reflecting a broad range of 46 4.3. MODELS EVALUTATION CHAPTER 4. RESULTS • Hierarchical slightly outperforms K-Means on the CH index (994 . 80 vs. 989 . 23), suggesting that it offers marginally better global separation for these five-dimensional features (TotalRevenue, Frequency, AverageTransactionValue, Lifespan, CLV ). • GMM again shows lower performance on both metrics (0 . 2012 Silhouette, 260 . 48 CH), implying that a purely Gaussian mixture approach may not capture the irregular patterns of CLV distributions. • Fuzzy C-Means and K-Means demonstrate comparably moderate Silhouette scores (around 0 . 62), but with K-Means performing better on the CH index. This dynamic again underscores how certain algorithms may excel in local cohesion yet differ in global variance partitioning. Summary of Model Evaluation • DBSCAN consistently yields high Silhouette scores in both RFM and CLV datasets, indicating well-defined clusters in terms of local density. However, its CH scores, though decent, fall short of those of K-Means or Hierarchical in some cases. • K-Means and Hierarchical often dominate the Calinski–Harabasz index, suggesting they produce broader inter-cluster separation relative to intra-cluster variance. Hierarchical clustering shows a slight edge over K-Means for the CLV data’s CH index. • GMM lags behind other methods across both metrics, for reasons likely tied to the non-Gaussian distribution of the chosen features (RFM or CLV). • Fuzzy C-Means performs moderately in both metrics, especially for CLV. It provides flexibility through soft memberships but may not match the cluster compactness or separation of DBSCAN, K-Means, or Hierarchical methods in this context. In conclusion, the choice of clustering method depends heavily on whether the primary goal is local cohesion vs. global separation. For local, density-driven segmentation (as indicated by Silhouette), DBSCAN frequently emerges as the best candidate. For maximizing between-cluster variance (as indicated by the Calinski–Harabasz index), K-Means or Hierarchical may be preferred. The next chapter discusses these findings in detail and outlines potential use-cases and future directions for applying the various clustering methodologies. 53 Comparative Study of Customer Segmentation Strategies Based on Business Analytics Chapter 5 Conclusions Overall, this thesis has comparatively analyzed customer segmentation strategies based on the RFM (Recency, Frequency, Monetary) and CLV (Customer Lifetime Value) models, integrating them with five distinct clustering algorithms: K-Means,Hierarchical Clustering,DBSCAN,Gaussian Mixture Models (GMM) and Fuzzy C-Means (FCM). The aim was to thoroughly investigate how different segmentation techniques can provide useful indications on both a tactical and strategic basis, highlighting their respective strengths and weaknesses. The results obtained provide significant considerations both from a methodological and a practical-applicative point of view. 1. Summary of Objectives and Context The thesis is placed in the broader panorama of Customer Relationship Management (CRM) and data-driven marketing, where customer segmentation plays a central role in optimizing resources and maximizing the return on investment of campaigns. In particular, the use of the RFM model allows to interpret purchasing behaviors in terms of temporal proximity (Recency), transaction frequency (Frequency) and monetary value (Monetary). Although extremely widespread for its interpretative simplicity, the RFM model has some limitations when it comes to evaluating the future economic potential of a customer. For this reason, the concept of CLV has also been included, which considers estimated future purchases, the duration of the relationship with the customer (lifespan) and other parameters that can give a long-term view on the economic value that can be generated. Algorithmically, each of the five clustering methods offers a different perspective: • K-Means and Fuzzy C-Means use a centroid-based approach, useful for obtaining compact groups, and differ in membership (hard vs. fuzzy). • Hierarchical Clustering (particularly with Ward.D2) builds a hierarchy of clusters, which is useful when exploring the nested structure of data. • DBSCAN identifies regions of high density by separating them from areas of lower density (classified as 54 CHAPTER 5. CONCLUSIONS noise), without the need to fix the number of clusters a priori. • GMM adds a probabilistic approach, assuming that the data comes from a mixture of Gaussian distributions, each with its own mean and variance. The starting data set, coming from the UCI Machine Learning Repository, was first cleaned and pre-treated (including removal of duplicates, management of outliers and date conversion). For the RFM part, three fundamental indicators were calculated (Recency,Frequency,Monetary), while for the CLV part, TotalRevenue, Average Transaction Value,Frequency,Lifespan and the estimate of CLV itself were introduced. Once these sets of variables were obtained, we proceeded with the application and comparison of the clustering algorithms, also evaluated through the internal metrics Silhouette Score and Calinski–Harabasz Index. 2. Key Findings in RFM Analysis 2.1 Default RFM Segments A first analysis divided customers into six segments (Whales, Active Stars, Loyal Regulars, New & Occasional Buyers, Sleeping Giants, Lost Clients) based on the RFM scores calculated with the quintile method. This ”manual” categorization indicated that almost half of the customers (over 47%) fall into the Lost Clients, i.e. customers who do not show recent or frequent purchases, with an overall modest spending value. On the other hand, a small group of Whales (just over 8%) generates a spending volume that is enormously higher than the average, placing itself as a top priority segment for retention or cross-selling strategies. 2.2 Comparison of Clustering Algorithms (RFM) • K-Means: It highlighted 4 clusters (determined by Elbow Method). One of these (Cluster 2) includes very few customers (13) with high monetary contribution ( ∼ 18.5% of the total), confirming the existence of an elite group with very high economic value. • DBSCAN: It identified 3 clusters, one of which is classified as noise (Cluster 0) and contains customers with an even more out of scale spending profile. This highlights a potential limit of DBSCAN, which tends to isolate the “extreme” points and classify them as noise if the ϵ and MinPts parameters are not calibrated very carefully. • Hierarchical Clustering: With a cut to 7 clusters, it allowed to identify in a granular way segments of customers with high frequency and monetary value, including some segments of minimum size but very high value ( M≈ 164 , 658 in one case). The hierarchical approach is very transparent, allowing to visualize how the groupings are formed through a dendrogram, although the choice of the cut point remains subjective. • GMM: It revealed 9 clusters for the RFM dataset, a rather high number. Here too, a small group of customers (Cluster 9) with an extremely high average monetary value (over 46,000 €) and a series of 55 CHAPTER 5. CONCLUSIONS ”intermediate” clusters with various compositions emerge. This shows the ability of GMM to capture nuances, but raises doubts about the interpretability of segmentations that are too fragmented. • Fuzzy C-Means (FCM): By setting 4 clusters (in line with K-Means), it distributed the customers in a less clear-cut way. Thanks to the fuzzy membership, some customers are positioned between multiple groups, reflecting that in real situations the behavioral boundaries are not always clear. However, the interpretation of the ”partial memberships” requires a greater analytical effort. 2.3 Application Interpretations The RFM analysis shows that ”high-spending” customers are often few but decisive for the turnover. This suggests that the RFM segmentation—with its immediacy of calculation—is excellent for tactical snapshots, such as planning seasonal campaigns or identifying inactive clusters (Lost Clients) to recover. However, if the decision horizon extends to long-term considerations (e.g. estimate of future revenues or the probability of customer ”survival”), RFM risks losing its effectiveness because it does not incorporate the prospective temporal dynamics. 3. Main Evidence in CLV Analysis 3.1 Various Definitions of Long-Term Value The second part of the thesis focused on the CLV, calculated by integrating total spent,frequency,average value per transaction and lifespan of the customer, to obtain a metric that summarizes the potential return over a prolonged time period. This perspective better intercepts marketing strategies oriented to the balance between maintaining the ”best customers” and acquiring new high-potential ones. 3.2 Comparison of Clustering Algorithms (CLV) • K-Means: With k = 4, it was highlighted how a very small cluster (only 2 customers) has an extraordinary average CLV, higher than 11 million euros. A second cluster (23 customers) shows very high frequency and robust CLV, while the majority falls into low or medium value clusters. • Hierarchical Clustering: It provided 4 clusters, one of which is again occupied by a few ultra-premium customers, and another by customers with relatively high frequency. The hierarchical approach confirms the presence of large masses of low-value customers and elite minorities with extreme parameters. • DBSCAN: Detected only 2 main clusters, separating a group of 81 ”top spenders” (Cluster 0) from the rest (Cluster 1, over 4000 low CLV customers). This clear separation, while easy to interpret, neglects the presence of possible substructures within the high-value group, as suggested by GMM or the dendrogram. • GMM: Estimates 6 clusters via BIC-based mclust. Different gradations of high-value customers emerge, including those who make a single, very expensive purchase and those who have very high frequencies, 56 CHAPTER 5. CONCLUSIONS with varied behaviors. The risk, as always, is over-segmentation, which could be too granular for lean marketing plans. • Fuzzy C-Means: Set on 4 clusters, it reiterated the existence of an exceptional micro-segment (fC 3 with just 7 customers but average CLV > 10 million) and a medium-high segment (fC 2) with about 333 customers. Most of the population remained in low-value clusters (fC 1 and 4), highlighting a large portion of customers that, in an upselling perspective, could be encouraged to grow. 4. Comparative Considerations: RFM vs. CLV A central element of this work consists in the comparison between the segmentation based on RFM and that based on CLV. Both models identify few customers with very high economic value and many customers with low spending. However: 1. Time horizon: • RFM favors the current behavior. If a customer with high past purchases stops buying for a few months, in the RFM Recency will get worse, and that individual could move from a ”high-spending” cluster to a less desirable one. • CLV instead estimates the future propensity, possibly recognizing high spending margins if historically the frequency has been high and the analysis time window suggests a probability of repurchase. 2. Strategic: • RFM lends itself to short-medium term direct marketing (for example, how to launch a Christmas promotion or a retargeting action on recently inactive customers). • CLV is more connected to strategic decisions: defining acquisition budgets, predictively evaluating the effectiveness of investments in retention, justifying extreme customization for the ”top tiers”. 3. Computation Complexity: •RFM is simpler and easily adoptable by companies with basic IT infrastructures. • CLV calculation requires forecasts or hypotheses on future behavior and requires models or assumptions (e.g. discount rates, estimated retention rate), making the procedure more complex. 4. Segmentation Stability: • RFM scores can fluctuate rapidly when the time dimension varies (a customer considered recent today may not be so in a few weeks). • CLV offers a more stable picture, at least as long as the calculation assumptions remain valid. However, if purchasing patterns change dramatically, CLV estimates should also be reevaluated. 57 CHAPTER 5. CONCLUSIONS 5. Clustering Algorithm Selection No algorithm has proven to be unquestionably “best” overall: its quality depends on the business goals, the shape of the data, and whether or not it is necessary to have well-separated clusters vs. fluid clusters. • K-Means: Ideal if you want a fast segmentation that can be interpreted in terms of centroids, provided you define the number of clusters well and work on data without strong outliers. • Hierarchical: Excellent for exploratory analysis and for visualizing the data structure using dendrograms; determining the cluster cut remains subjective, however. • DBSCAN: Useful for finding dense micro-clusters and for highlighting outliers, but it can marginalize valuable customers by labeling them as “noise” if ϵand MinPts are not calibrated correctly. • GMM: Offers a high degree of flexibility and can accommodate complex distribution shapes; however, it risks “over-splitting” the sample, especially if the data does not follow true Gaussian shapes. • Fuzzy C-Means: It is advantageous when customers are expected to have “multiple affinities” to multiple clusters (e.g., a customer who buys both “premium” and “standard” products). However, interpreting fuzzy memberships requires an additional step compared to the more common hard clustering. 6. Managerial Implications The distinctions between RFM and CLV, coupled with the varying characteristics of clustering algorithms, provide numerous insights for developing operational strategies: 1. Customer Portfolio Management • Identify Whales or High-CLV segments as priorities for exclusive campaigns, such as personalized discounts, early access to products, and dedicated customer service. • For larger low-value clusters (Lost Clients in RFM or clusters of “one-timers” in CLV), it is advisable to implement cost-effective actions. Strategies may include mass actions (e.g., generic email marketing) or selective recovery efforts (e.g., targeted discounts, loyalty packages), as investing in expensive strategies for these segments does not yield significant added value. 2. Resource Optimization • Allocate marketing resources (budget, time, contacts) proportionally to the potential value of each cluster. Specifically, CLV can justify greater investments in retaining top customers, where a high return is anticipated in the long term. 3. Loyalty and Cross-Selling Programs • Encourage segments with high Frequency but moderate unit spending to purchase higher-margin products through cross-selling (offering related or complementary products) and up-selling (en58 CHAPTER 5. CONCLUSIONS couraging the purchase of more expensive items) initiatives. • Target segments with low Recency but a history of substantial spending (Sleeping Giants) with personalized “win-back campaigns” (strategies aimed at re-engaging inactive customers). 4. Competition and Offer Analysis • If a company observes that customers within a specific cluster are migrating to competitors, corrective measures can be adopted, such as improving service quality, reducing delivery times, or launching targeted promotions. • RFM and CLV metrics can serve as internal benchmarks to monitor how customer distributions across segments evolve over time. 5. Data-Driven Culture • Implement internal dashboards that allow managers to explore RFM and CLV metrics in real time (or near real time). • The synergy between the two perspectives (shortto medium-term RFM and long-term CLV) reinforces a corporate culture based on data and predictive analysis, rather than solely on instinctive decision-making. 7. Limitations and Future Perspectives Although the research has provided important insights, there are some aspects to consider for possible future developments: 1. Data Quality and Updating • The analyzed dataset, although robust, reflects a specific period of time (about a year of transactions). In real contexts, the data should be continuously updated, and the clustering models “recursively” recalibrated. • The possible presence of seasonality (Christmas period, Black Friday, etc.) could alter the values of Recency and Frequency significantly. 2. Relevance of Additional Variables • Both RFM and CLV ignore dimensions such as product category, net profitability (which would require considering costs and margins), demographics (age, location) or behavioral variables (feedback, reviews, return rate). Integrating additional data sources could refine the segmentation, but make the calculation more complex. 3. Clustering Parameters • DBSCAN, for example, is extremely sensitive to the choice of ϵ . The same is true for GMM, which can return a variable number of clusters based on the BIC. Greater methodological robustness could 59 CHAPTER 5. CONCLUSIONS include a more sophisticated model selection (cross-evaluation of multiple metrics and comparisons with simulated data). 4. Advanced Machine Learning Applications • If segmentation is combined with predictive models (e.g. churn forecasting or propensity scoring), even more targeted results can be obtained. Methods such as deep clustering could be tested in contexts with large amounts of unstructured data (clickstream, navigation logs). 5. External Validation • In this work, internal validation metrics were used (Silhouette,Calinski–Harabasz). It would be useful to integrate an external validation, measuring the actual impact of each cluster on real business metrics (e.g. redemption rate of campaigns, upgrade rate to premium plans, etc.). 8. Overall Conclusion Ultimately, this thesis has shown how, in a marketing and CRM context, the choice between RFM and CLV and the selection of a particular clustering algorithm cannot be reduced to universal rules valid for all. On the contrary, it is necessary to consider: • The nature of the data: A dataset with many outliers and a highly skewed distribution can benefit from the most flexible clustering models (DBSCAN, GMM), as long as parameters that distort its interpretation are avoided (for example, erroneously defining valuable customers as “noise”). • The time horizon and strategy: If marketing actions aim to recover inactive customers of the last few months, RFM is an immediate indicator. If, on the other hand, it is a question of investing in long-term loyalty programs, CLV becomes central. • The desired granularity: K-Means and FCM allow a fixed number of clusters, while DBSCAN, GMM and Hierarchical can potentially create more or less segments, proving useful or dispersive depending on the case. • The ease of interpretation: The adoption of a certain method should always take into account the possibility of communicating the results clearly to business decision-makers. An excessive fragmentation into poorly distinguishable clusters risks confusing managers rather than helping them. This work contributes to the literature on customer analytics by demonstrating how different clustering methods can lead to partly divergent interpretations, especially in the presence of highly heterogeneous data and with few individuals generating the majority of the revenue. However, the “complementary” nature of RFM and CLV metrics suggests that the ideal choice often consists in combine both perspectives. For example, a company could define primary clusters using RFM (quick to update and interpret), and then prioritize customers with the highest CLVs within each cluster. From an operational perspective, the results obtained provide a practical framework for those within the organization who want to identify, describe and retain the best customers, without neglecting conversion 60 CHAPTER 5. CONCLUSIONS opportunities among the ”“average” groups or recovery of dormant groups. The road to truly personalized marketing passes through the continuous evolution of these models: iterating the segmentation with updated data, inserting new variables that provide additional levels of depth (net profitability, purchase preferences, channels used, social media interactions) and comparing the hypotheses with tangible economic results. In summary, the thesis reinforces the idea that there is no ”“one size fits all” for customer segmentation. Each business context and each marketing objective require tailor-made analyses, both in terms of the definition of metrics (RFM vs. CLV, or a mix of both) and the choice of the clustering algorithm. In the long run, a hybrid and iterative approach appears to be the most solid solution, where the results of one tool (e.g. RFM) are enriched and validated by the perspectives of another (CLV), and where multiple clustering methods are compared to capture the patterns that best reflect the market structure and business needs. This approach allows to acquire an integrated vision and to draw increasingly effective data-driven decisions in the current competitive context. 61 Comparative Study of Customer Segmentation Strategies Based on Business Analytics Bibliography 1. SMITH, Wendell R. Product Differentiation and Market Segmentation as Alternative Marketing Strategies. Journal of Marketing. 1956, vol. 21, no. 1, pp. 3–8. 2. TEICHERT, Thorsten. Customer Segmentation Revisited: The Case of the Airline Industry. Journal of Air Transport Management. 2008, vol. 14, no. 6, pp. 329–336. 3. CHUNG, K. H.; CHEN, M. RFM Analysis: A Balancing Act Between Business Intelligence and Marketing Intelligence. Industrial Management & Data Systems. 2016, vol. 116, no. 2, pp. 20–33. 4. KHAJVAND, Mahboubeh; ZOLFAGHAR, Kiyana; ASHRAFI, M.; ALIZADEH, S. Estimating Customer Lifetime Value Based on RFM Analysis of Customer Purchase Behavior: Case Study. Procedia Computer Science. 2011, vol. 3, pp. 57–63. 5. CHRISTY, T.; B., Noor Raihani; D., Gunawan D. Enhancing RFM Analysis with Weighting. Proceedings of the International Conference on Computing and Informatics. 2018, vol. 7, pp. 84–91. 6. GUPTA, S.; LEHMANN, D. R.; STUART, J. A. Valuing Customers. Journal of Marketing Research. 2006, vol. 41, no. 1, pp. 7–18. 7. VILLANUEVA, J.; HANSSENS, D. M. Customer Equity: Measurement, Management and Research Opportunities. Foundations and Trends in Marketing. 2007, vol. 1, no. 1, pp. 1–95. 8. LEMMENS, A.; CROUX, C. Bagging and Boosting Classification Trees to Predict Churn. Journal of Marketing Research. 2006, vol. 43, no. 2, pp. 276–286. 9. PAKHIRA, M. K.; BANDYOPADHYAY, S.; MAULIK, U. Validity Index for Crisp and Fuzzy Clusters. Pattern Recognition. 2004, vol. 37, no. 3, pp. 487–501. 10. PHAM, D. T.; DIMOV, S. S.; NGUYEN, C. D. Selection of K in K-Means Clustering. Proceedings of the Institution of Mechanical Engineers, Part C: Journal of Mechanical Engineering Science. 2005, vol. 219, no. 1, pp. 103–119. 11. FALLIS, A. D. Hierarchical Clustering Approaches for Large Scale Data. Journal of Big Data Analysis. 2013, vol. 2, no. 1, pp. 35–41. 12. EVERITT, B. S.; LANDAU, S.; LEESE, M.; STAHL, D. Cluster Analysis. 5th. Wiley, 2011. 13. ESTER, M.; KRIEGEL, H.-P.; SANDER, J.; XU, X. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. In: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining. AAAI Press, 1996, pp. 226–231. 62 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY 231 xaxis = list(title =’recency’, type = ’ log ’), 232 yaxis = list(title =’frequency ’, type = ’ log ’), 233 zaxis = list(title =’monetary ’, type = ’ log ’) 234 ) 235 ) 236 237 gmm _stats <- cluster_statistics (RFM , "gmm_cluster") 238 print (gmm_stats ) 239 240 # ######### hierarchical ########## 241 dist_matrix <- dist (RFM_normalized ) 242 hc_ model <- hclust ( dist _matrix , method = " ward . D2") 243 RFM $hierarchical_cluster <- as.factor( cutree ( hc _model , k = 7)) 244 245 plot(hc_ model ,labels = FALSE , main = " hierarchical clustering dendrogram ", 246 xlab = "" ,sub =" ward method ") 247 rect. hclust ( hc _model , k = 7, border = "red ") 248 249 dendro <- as. dendrogram ( hc_ model ) 250 dendro_data <- dendro_data( dendro ) 251 252 k<- 7 253 cluster_assignments <- cutree(hc_ model , k) 254 labels _ df <- data .frame ( label = rownames ( RFM _normalized ) , cluster = cluster _assignments , stringsAsFactors = FALSE) 255 labels _ df <- merge (labels_df, dendro_data$labels ,by =" label ") 256 257 rect _ data <- aggregate (x ~cluster , data =labels _df , FUN = function (x) c(min (x), max (x))) 258 rect_data <- do.call(rbind,lapply(1: nrow(rect_data), function (i) { 259 data.frame ( 260 cluster = rect_data$cluster[i], 261 xmin = rect_data$x[i, 1], 262 xmax = rect_data$x[i, 2], 263 ymin = 0, 264 ymax = max(dendro_data$segments $yend) /2 265 ) 266 })) 267 268 ggplot() + 269 geom_segment(data = dendro_data$segments , 270 aes (x = x, y = y, xend = xend , yend = yend )) + 271 geom_rect(data =rect _data , 272 aes (xmin = xmin , xmax = xmax , ymin = ymin , ymax = ymax , fill = as.factor(cluster )), 273 alpha = 0.3) + 274 scale _ fill_manual ( name = "cluster", 275 values = c("# FF6666 " ,"#FFCC66"," #66 CC66 ","#66 CCCC "," #6666 FF " ,"# CC66FF","#FF66CC")) + 276 labs(title =" hierarchical clustering dendrogram ", x = "clients", y = "height") + 69 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY 277 theme _minimal () + 278 theme ( axis.text.x = element_blank () , axis. ticks . x = element _blank () , legend. position = " right ") 279 280 # zoom on some clusters 281 zoom_xmin <- min (rect_data$xmin[rect_data$cluster %in% c(5 , 6, 7) ]) 282 zoom_xmax <- max (rect_data$xmax[rect_data$cluster %in% c(5 , 6, 7) ]) 283 zoom_ymax <- max (dendro_data$segments $yend) 284 285 ggplot() + 286 geom_segment(data = dendro_data$segments , 287 aes (x = x, y = y, xend = xend , yend = yend )) + 288 geom_rect(data =rect_data[rect_data$cluster %in% c(5 , 6, 7) , ], 289 aes (xmin = xmin , xmax = xmax , ymin = ymin , ymax = ymax , fill = as.factor(cluster )), 290 alpha = 0.3) + 291 scale _ fill_manual ( name = "cluster", 292 values = c("# FFCC66 " ,"#66 CC66","#FF6666")) + 293 coord _cartesian ( xlim = c(zoom_xmin , zoom_xmax), ylim = c(0 , zoom_ymax /3) ) + 294 labs(title =" zoom on clusters 5, 6, and 7" ,x="clients", y = "height") + 295 theme _minimal () + 296 theme ( axis.text.x = element_blank () , axis. ticks . x = element _blank () , legend. position = " right ") 297 298 ggplot (RFM , aes (x = Recency , y = Frequency , size = Monetary , color = hierarchical _cluster)) + 299 geom_point ( alpha = 0.7) + 300 labs(title =" hierarchical clustering ( rfm )" ,x="recency", y = " frequency ") + 301 theme _minimal() 302 303 plot_ly( 304 data = RFM , 305 x = ~Recency , 306 y = ~Frequency , 307 z = ~Monetary, 308 color = ~hierarchical_cluster , 309 type = " scatter3d ", 310 mode ="markers", 311 marker = list( size = 4) 312 ) %>% 313 layout( 314 title =" hierarchical clustering ( log scale )" , 315 scene = list( 316 xaxis = list(title =’recency’, type = ’ log ’), 317 yaxis = list(title =’frequency ’, type = ’ log ’), 318 zaxis = list(title =’monetary ’, type = ’ log ’) 319 ) 320 ) 321 70 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY 322 hierarchical_stats <- cluster_statistics ( RFM , "hierarchical_cluster") 323 print (hierarchical_stats ) 324 325 # ######### fuzzy cmeans ########## 326 set .seed (333) 327 fcm _result <- cmeans ( RFM _normalized , centers = k_optimal , m = 2, iter .max = 100 , verbose = FALSE ) 328 RFM $fuzzy _cluster <- apply ( fcm _result$membership , 1, which .max) 329 330 ggplot (RFM , aes (x = Recency , y = Frequency , size = Monetary , color = factor( fuzzy _cluster))) + 331 geom_point ( alpha = 0.7) + 332 labs(title =" fuzzy c - means clustering ( rfm )" , x = "recency",y=" frequency ") + 333 theme _minimal() 334 335 sample_indices <- sample (1:nrow(RFM _normalized ), 20) 336 sample_membership <- fcm _result$membership [ sample_indices , ] 337 membership _df <- as.data.frame (sample_membership ) 338 membership _df$Client <- paste ("client", 1: nrow( membership _df)) 339 membership _long <- melt ( membership _df , id. vars = "Client", 340 variable . name = "Cluster", value . name = " Membership ") 341 342 ggplot ( membership _long , aes(x = Cluster , y = Client , fill = Membership )) + 343 geom_tile ( color = " white " ) + 344 scale _ fill_gradient ( low = " lightblue ", high = " darkblue ", name = " membership ") + 345 labs(title =" fuzzy cmeans membership ( sample )", x = "cluster", y = " client ") + 346 theme _minimal () + 347 theme ( axis.text.x = element_text( size = 10) , 348 axis.text.y = element _text( size = 8, angle = 45, hjust = 1) , 349 axis. ticks = element _blank ()) 350 351 plot_ly( 352 data = RFM , 353 x = ~Recency , 354 y = ~Frequency , 355 z = ~Monetary, 356 color = ~factor( fuzzy _cluster), 357 type = " scatter3d ", 358 mode ="markers" 359 ) %>% 360 layout( 361 title =" fuzzy c - means clustering ( log scale ) ", 362 scene = list( 363 xaxis = list(title =’recency’, type = ’ log ’), 364 yaxis = list(title =’frequency ’, type = ’ log ’), 365 zaxis = list(title =’monetary ’, type = ’ log ’) 366 ) 367 ) 368 71 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY 369 fuzzy _stats <- cluster_statistics (RFM , " fuzzy _cluster") 370 print ( fuzzy _stats ) 371 372 # ######### evaluation ( silhouette &ch index ) ########## 373 dist_matrix <- dist (RFM_normalized ) 374 375 silhouette _kmeans <- silhouette ( as.numeric( RFM $kmeans_cluster ) , dist _matrix) 376 silhouette _dbscan <- silhouette ( as.numeric( RFM $dbscan_cluster ) , dist _matrix) 377 silhouette _hierarchical <- silhouette ( as.numeric( RFM $hierarchical_cluster ) , dist _matrix) 378 silhouette _gmm <- silhouette ( as.numeric( RFM $gmm _cluster ) , dist _matrix) 379 silhouette _fuzzy <- silhouette ( as.numeric( RFM $fuzzy _cluster ) , dist _matrix) 380 381 cat (" silhouette scores :\ n") 382 cat ("k - means : " ,mean ( silhouette _kmeans [ , 3]) , "\n") 383 cat (" dbscan : " ,mean( silhouette _dbscan [, 3] , na.rm = TRUE ), "\n") 384 cat ("hierarchical: ",mean( silhouette _hierarchical[, 3]), "\n") 385 cat ("gmm : ",mean( silhouette _gmm [, 3]) , "\n") 386 cat (" fuzzy c - means : " ,mean( silhouette _fuzzy [ , 3]) , "\n") 387 388 calinski _harabasz <- function (data, cluster_vector) { 389 k<- length(unique(cluster_vector)) 390 if (k < 2) return(NA ) 391 data _ mat <- as .matrix(data) 392 n<- nrow (data _ mat ) 393 overall_mean <- colMeans(data_ mat ) 394 W<- 0 395 B<- 0 396 for ( cl in unique(cluster_vector)) { 397 cl_indices <- which (cluster_vector == cl) 398 cl_ data <- data _ mat [cl_indices , , drop = FALSE ] 399 cl_mean <- colMeans ( cl_data) 400 W<- W + sum ( rowSums (( cl _data - cl_mean) ^2) ) 401 B<- B + nrow(cl_data)* sum (( cl _mean - overall_ mean ) ^2) 402 } 403 if (W == 0 || k == 1 || (n - k) == 0) return( NA) 404 (B /(k - 1)) /(W /(n - k)) 405 } 406 407 ch_kmeans <- calinski _harabasz ( RFM _normalized , RFM $kmeans_cluster) 408 ch_dbscan <- calinski _harabasz ( RFM _normalized , RFM $dbscan_cluster) 409 ch_hierarchical <- calinski _harabasz ( RFM _normalized , RFM $hierarchical_cluster) 410 ch_gmm <- calinski _harabasz ( RFM _normalized , RFM $gmm _cluster) 411 ch_fuzzy <- calinski _harabasz ( RFM _normalized , RFM $fuzzy _cluster) 412 413 ch_results <- data.frame ( 414 method = c("k - means " ,"dbscan","hierarchical","gmm"," fuzzy c - means " ), 415 calinski _harabasz = c(ch_kmeans , ch_dbscan , ch_hierarchical , ch_gmm , ch_fuzzy ) 416 ) 417 print (ch_results) 72 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY 5.1.1 RFM 3D visualization Figure 5.1: RFM K-Means Clustering (LOG) Figure 5.2: RFM DBSCAN Clustering (LOG) 73 5.1. RFM ANALYSIS, CLUSTERING CODE AND EVALUTATION BIBLIOGRAPHY Figure 5.3: RFM GMM Clustering (LOG) Figure 5.4: RFM Hierarchical Clustering (LOG) 74 5.2. CLV SEGMENTATION,CLUSTERING AND MODEL EVALUATION BIBLIOGRAPHY Figure 5.5: RFM Fuzzy Clustering (LOG) 5.2 CLV segmentation,clustering and model evaluation 1# ######### load libraries ########## 2library( e1071 ) 3library( dplyr ) 4library( lubridate ) 5library(ggplot2) 6library(cluster) 7library(dbscan) 8library(mclust) 9library( reshape2 ) 10 library( ggdendro ) 11 library( fmsb ) 12 # ######### load and preprocess data ## 13 data <- read.csv ("data/data .csv") 14 data <- data %>% 15 mutate ( InvoiceDate = as . Date ( InvoiceDate , format ="%m/%d/%Y")) % >% 16 mutate ( TotalPrice = Quantity *UnitPrice ) % >% 17 filter(!is.na( CustomerID ) , Quantity > 0, UnitPrice > 0) %>% 18 group _by( CustomerID ) % >% 19 summarise ( 20 TotalRevenue = sum( TotalPrice ) , 21 Frequency = n_distinct ( InvoiceNo ) , 22 AvgTransactionValue = mean( TotalPrice ) , 23 Lifespan = as.numeric( difftime ( max( InvoiceDate ) , min ( InvoiceDate ) , units = " days ")) 24 ) %>% 75 5.2. CLV SEGMENTATION,CLUSTERING AND MODEL EVALUATION BIBLIOGRAPHY 25 ungroup() 26 27 # ########## calculate clv ########## 28 data <- data %>% 29 mutate ( CLV = Frequency *AvgTransactionValue *Lifespan ) 30 31 # ######### normalize data ########## 32 data_normalized <- data %>% 33 select ( TotalRevenue , Frequency , AvgTransactionValue , Lifespan , CLV ) % >% 34 scale () % >% 35 as.data.frame () 36 rownames (data_normalized ) <- data$CustomerID 37 38 # ######### define statistics functions ########## 39 calculate_statistics <- function (data , cluster_column ) { 40 data %>% 41 group _by(!!sym ( cluster _column ) ) % >% 42 summarise ( 43 TotRevenueMean = mean(TotalRevenue), 44 FMean = mean( Frequency ) , 45 CLVMean = mean(CLV), 46 AvgTMean = mean (AvgTransactionValue), 47 LifespanMean = mean( Lifespan ) , 48 TotSD = sd(TotalRevenue), 49 FSD = sd ( Frequency ), 50 CLVSD = sd(CLV), 51 AvgTSD = sd(AvgTransactionValue), 52 LifespanSD = sd ( Lifespan ) , 53 Size = n() 54 ) 55 } 56 57 ##### elbow method ######## 58 set .seed (123) 59 wss <- sapply(1:10 , function (k) { 60 kmeans(data_normalized , centers = k, nstart = 10) $tot . withinss 61 }) 62 ggplot(data.frame (k = 1:10 , wss = wss ) , aes (x = k, y = wss )) + 63 geom_line( size = 1, color = " blue ") + 64 scale _x_continuous ( breaks = 1:10) + 65 labs(title =" elbow method " , x = " number of clusters " , y = " within sum of squares ") + 66 theme _minimal () + 67 theme ( 68 panel . border = element _rect( color = " black " , fill = NA , size = 1) 69 ) 70 71 k_values <- 1:10 72 73 # print elbow values 76 5.2. CLV SEGMENTATION,CLUSTERING AND MODEL EVALUATION BIBLIOGRAPHY 74 elbow _values <- data .frame (k = k_values, wss = wss) 75 print ( elbow _values) 76 77 # ######### k - means ########## 78 k_optimal <- 4 79 80 kmeans_result <- kmeans(data_normalized , centers = k_optimal , nstart = 10) 81 data$kmeans_cluster <- as.factor(kmeans_result$cluster) 82 83 # ######### visualize kmeans ########## 84 ggplot(data , aes (x = Frequency , y = TotalRevenue , color = kmeans _cluster , size = CLV )) + 85 geom_point ( alpha = 0.7) + 86 labs(title ="k - means clustering ( clv )",x=" frequency ", y = " total revenue " , size = " clv ") + 87 theme _minimal() 88 89 90 # ######### k - means statistics ########## 91 kmeans_stats <- calculate _statistics ( data,"kmeans_cluster") 92 print (kmeans_stats ) 93 94 # ######### hierarchical ########## 95 dist_matrix <- dist(data_normalized ) 96 hc_ model <- hclust ( dist _matrix , method = " ward . D2") 97 data$hierarchical_cluster <- cutree(hc_ model , k = k_optimal) 98 99 # ######### dendrogram ########## 100 plot(hc_ model , main = " hierarchical clustering dendrogram ", 101 xlab = " index " , ylab = " distance ") 102 rect. hclust ( hc _model ,k=k_optimal , border = " red ") 103 104 # ######### enhanced dendrogram ########## 105 dendro <- as. dendrogram ( hc_ model ) 106 dendro_data <- dendro_data( dendro ) 107 108 cluster_assignments <- cutree(hc_ model , k = k_optimal) 109 labels _ df <- data .frame ( label = rownames (data_normalized ) , 110 cluster = cluster_assignments , 111 stringsAsFactors = FALSE) 112 labels _ df <- merge (labels_df, dendro_data$labels ,by =" label ") 113 114 rect _ data <- aggregate (x ~cluster , data =labels _df , 115 FUN = function(x) c(min(x), max (x))) 116 117 rect_data <- do.call(rbind,lapply(1: nrow(rect_data), function (i) { 118 data.frame ( 119 cluster = rect_data$cluster[i], 120 xmin = rect_data$x[i, 1], 121 xmax = rect_data$x[i, 2], 77 5.2. CLV SEGMENTATION,CLUSTERING AND MODEL EVALUATION BIBLIOGRAPHY 122 ymin = 0, 123 ymax = max(dendro_data$segments $yend) /2 124 ) 125 })) 126 127 ggplot() + 128 geom_segment(data = dendro_data$segments , 129 aes (x = x, y = y, xend = xend , yend = yend )) + 130 geom_rect(data =rect _data , 131 aes (xmin = xmin , xmax = xmax , ymin = ymin , ymax = ymax , fill = as.factor(cluster )), 132 alpha = 0.3) + 133 scale _ fill_manual ( name = "cluster", 134 values = c("# FF6666 " ,"#FFCC66"," #66 CC66 ","#66 CCCC ")) + 135 labs(title =" hierarchical clustering dendrogram ", x = "clients", y = "height") + 136 theme _minimal () + 137 theme ( axis.text.x = element_blank () , axis. ticks . x = element _blank () , legend. position = " right ") 138 139 # ######### visualize hierarchical ########## 140 ggplot(data , aes (x = Frequency , y = TotalRevenue , color = as.factor(hierarchical_cluster), size = CLV )) + 141 geom_point ( alpha = 0.7) + 142 labs(title =" hierarchical clustering ( clv )" ,x=" frequency " , y = " total revenue " , size = "clv", color = "cluster") + 143 theme _minimal() 144 145 146 # ######### hierarchical statistics ########## 147 hierarchical_stats <- calculate_statistics ( data ,"hierarchical_cluster") 148 print (hierarchical_stats ) 149 150 # ##### dbscan #### 151 kNNdistplot <- function (data, k) { 152 distances <- kNNdist(data, k = k) 153 distances <- sort( distances ) 154 plot( distances , type = "l", main = "kNN distance plot ", 155 xlab = " points sorted by distance " , ylab = "k -NN distance ") 156 } 157 158 kNNdistplot ( data_normalized , k = 5) 159 abline( h = 0.6 , col =" red ", lwd = 2) 160 dbscan_eps <- 0.6 161 dbscan_result <- dbscan(data_normalized , eps = dbscan _eps , MinPts = 5) 162 data$dbscan_cluster <- as.factor(dbscan_result$cluster) 163 164 # ######### visualize dbscan ########## 165 ggplot(data , aes (x = Frequency , y = TotalRevenue , color = dbscan _cluster , size = CLV )) + 166 geom_point ( alpha = 0.7) + 78