Full text
Interpretable Factor Models with Machine Learning Olamide Ayodele1, Paul Laux2 1Institute of Financial Services Analytics, University of Delaware, 591 Collaboration Way, Newark, Delaware, 19713 2Institute of Financial Services Analytics, University of Delaware, 591 Collaboration Way, Newark, Delaware, 19713 Abstract Understanding the sources of cross-sectional variation in asset returns remains central to financial economics. This paper introduces aCluster-Based Factor Model (CBFM) that integrates hierarchical clustering and locally linear embedding (LLE) to construct parsimonious, data-driven factors from characteristics-managed returns. The clustering stage identifies groups of highly correlated characteristics managed returns, while the subsequent LLE stage extracts low-dimensional nonlinear factors within each cluster. Factor loadings are then estimated through a neural network framework, enabling flexible, time-varying relationships between latent factors and asset returns. Empirical evaluation using FamaβFrench 25 value-weighted portfolios shows that the CBFM achieves nearly twice the explanatory and predictive π
2 of traditional economic and categorical groupings. Compared with benchmark models such as RPβPCA and standard PCA, the CBFM yields smaller and statistically insignificant alphas, demonstrating superior pricing accuracy and reduced overfitting. These findings highlight the efficacy of data-driven clustering for uncovering the true structure of systematic risk. Key Words: Factors, Characteristics Managed Returns (CMR), Local Linear Embedding, Hierarchical Clustering 1. Introduction Understanding why different assets earn different average returns remains one of the most enduring questions in financial economics. Despite decades of research progress, there is still no consensus on the minimal set of systematic factors that can fully explain the cross-sectional variation in expected returns. While early models provided elegant theoretical foundations for linking risk and return, empirical evidence continues to show that returns are influenced by multiple sources of risk that vary across assets and over time. The growing body of evidence on firm-level characteristicsβsuch as size, value, profitability, and momentumβsuggests that observable variables contain valuable information about risk exposures. Yet, the rapid expansion of proposed characteristics has led to an overabundance of potential factors, creating uncertainty about which ones truly represent distinct and priced sources of systematic risk. This proliferation underscores the need for models that can synthesize large sets of characteristics into a coherent, interpretable structure without imposing overly restrictive assumptions. At the same time, purely statistical methods that infer latent factors from return data have proven powerful for dimensionality reduction but often fall short in economic interpretability. Methods such as principal component analysis and its modern extensions can efficiently summarize return covariances, yet the resulting factors are frequently difficult to reconcile with known economic mechanisms. Moreover, most of these approaches assume globally linear relationships between factors and firm characteristics, overlooking local heterogeneity that may arise when assets differ structurally in their sources of risk. For example, assets in different industries, size categories, or growth regimes may share distinct local risk patterns that global factor models fail to capture. As a result, there remains a persistent tension between flexibility and interpretabilityβbetween models that fit the data well and those that provide meaningful economic insight. Our paper addresses this gap by proposing the Cluster-Based Factor Model (CBFM), a hybrid framework that combines clustering and factor modeling to identify both global and local structures of systematic risk. The CBFM partitions characteristics managed returns (CMR) in statistically meaningful clusters based on their correlation score and then extracts latent factors within each cluster to capture dominant within-group variations in returns. This approach allows for heterogeneity in factor formation across different types of assets while preserving the parsimony of the overall factor structure. Conceptually, the model generalizes the idea of characteristic-sorted factors by endogenizing the grouping process, letting the data determine similar risk exposures rather than relying on arbitrary sorting thresholds. Methodologically, it integrates machine-learning techniquesβ hierarchical clustering and locally linear embeddingβwith traditional econometric tools for factor estimation, providing a flexible but interpretable framework 1
for identifying the building blocks of systematic risk. Empirically, the CBFM is evaluated using a broad panel of CMR. The analysis compares the modelβs ability to explain and predict the cross-section of returns relative to established benchmarks, including Fama-French factors, principal componentβbased models, and deep-learningβdriven factor approaches. By examining both explanatory power and pricing errors, the study assesses whether the cluster-based factors capture incremental dimensions of risk beyond those explained by existing models. The results demonstrate that the proposed framework enhances both interpretability and empirical performance, offering a unified approach that bridges the gap between traditional factor-based asset pricing and modern machine-learning methods. In doing so, the CBFM contributes to the ongoing effort to build asset-pricing models that are data-driven yet grounded in clear economic logic. 2. Related Work Factors play a central role in asset-pricing models, representing systematic sources of risk that help explain the cross-sectional variation in expected returns. The economic rationale is that investors require compensation for bearing pervasive risks that cannot be diversified away. The construction and representation of these factors vary considerably depending on the modeling approach. Broadly, two main paradigms dominate the literature. In one, factors are pre-specified and constructed from observable firm characteristics believed to proxy for underlying sources of risk. The FamaβFrench three-factor model of Fama and French [11] , later expanded to a five-factor model by Fama and French [12] , exemplifies this approach by forming portfolio-return factors based on size, value, profitability, and investment characteristics. These models interpret firm-level variables as economically meaningful measures of systematic risk exposures. However, the characteristic-based framework has evolved into what Harvey et al. [19] termed the βfactor zoo,β with hundreds of variables proposed as return predictors. Hou et al. [20] documented the limited robustness of many of these characteristics, underscoring the difficulty of identifying truly independent and persistent risk factors. To address this, several recent studies, such as Kozak et al. [25] , Fama and French [13] , and Feng et al. [15] , have developed new methods of factor construction that move beyond traditional portfolio-sorting approaches. These range from rank-based weighting schemes to deep learning architectures designed to construct economically guided factors that minimize pricing errors, offering more flexible yet still interpretable alternatives to conventional models. The second paradigm treats factors as latent variables, inferred statistically from large panels of returns rather than specified a priori. This tradition builds on the factor-analytic framework pioneered by Connor and Korajczyk [9] and formalized by Bai and Ng [2] , Bai [1] , and Lam and Yao [26] , who established consistent estimation procedures for both the number of factors and their loadings in high-dimensional settings. Principal Component Analysis (PCA) remains the most widely used method under this framework, assuming that a few common shocks explain most of the covariation in returns. Subsequent studies expanded this approach by linking factors to observable firm characteristics. For instance, Kelly et al. [22] proposed Instrumented PCA (IPCA), where characteristics serve as instruments to estimate time-varying factor loadings using an alternating least-squares algorithm. Gu et al. [18] further extended this logic using deep autoencoders to model nonlinear mappings between firm characteristics and latent factors, while Lettau and Pelger [29] and Fan et al. [14] introduced risk-premium and projected PCA variants that improve cross-sectional fit and estimation efficiency. Yet, despite these advances, a persistent challenge remains: many of these latent factors, though statistically powerful, lack economic interpretability. Kozak et al. [24] and Kozak and Nagel [23] emphasize that parsimony and interpretability should remain central objectives, arguing that a small number of dominant factors often suffice to capture the main sources of systematic risk in returns. Beyond factor construction, an equally important dimension in the literature concerns the estimation of factor loadingsβthe sensitivities of asset returns to systematic risks. Traditional methods assume static loadings, implying that the relationship between factors and returns remains stable over time. However, empirical evidence suggests that loadings evolve with business-cycle fluctuations, firm characteristics, and market conditions (Stock and Watson [33] ; Shanken [32] ). This realization has prompted both parametric and non-parametric innovations. Kelly et al. [22] modeled loadings as linear functions of observable characteristics, estimated iteratively through alternating least squares, while Connor et al. [8] and Su and Wang [34] proposed flexible, data-driven frameworks where loadings vary smoothly over time and across assets. More recent work, including Gu et al. [18] and Bakalli et al. [3] , integrates machine-learning architectures to capture nonlinear and time-varying risk exposures. These advances align with the growing conditional asset-pricing literature (Ferson and Harvey [16] ; Lettau and Ludvigson [28] ) and highlight the importance of accounting for dynamic relationships between firm characteristics and systematic risk factors. Building on this literature, the proposed Cluster-Based Factor Model (CBFM) aims to integrate the strengths of both research streamsβconstructing factors from clusters of similar CMR while estimating time-varying loadings that reflect evolving risk exposuresβto achieve a flexible yet interpretable framework for explaining the cross-section of asset returns. 2
3. METHODOLOGY 3.1 Motivation for Clustering Clustering, though not new to asset-pricing research, provides a systematic framework for organizing high-dimensional data into groups that capture latent structures and sub-features of the original information set. Ludvigson and Ng [30] grouped a large panel of macroeconomic variables into eight categories based on prior economic intuition, arguing that factors estimated from unrestricted panels often lack interpretability. By structuring data into economically coherent blocks, they facilitated the identification and naming of factors derived from each group. Similarly, Greengard et al. [17] applied a modified t-distributed stochastic neighbor embedding (t-SNE) approach to reveal fine-grained cluster structures, identifying six distinct groups of assets that correspond closely to known anomalies such as value, momentum, investment, profitability, and volatility. Their findings emphasize how clustering enhances interpretability by aligning statistical patterns with economic meaning. More recently, Jiao et al. [21] demonstrated that models which do not employ clustering tend to over-identify the number of relevant factors, as distinguishing highly correlated characteristic-managed returns becomes increasingly difficult in high dimensions. Beyond interpretability, clustering offers a pathway to factor specificity. Dimensionality-reduction techniques such as Principal Component Analysis (PCA), when applied directly to large panels of returns or characteristics-managed portfolios, often produce factors that are global averages of multiple underlying drivers. For instance, applying PCA to characteristics managed returns constructed from value, momentum, investment, and size characteristics typically yields principal components that combine all four sources of variation into a single factor, reducing their individual interpretability. Clustering mitigates this issue by separating characteristic groups and preserving their distinctiveness prior to factor extraction. In doing so, it enhances the granularity and precision of factor construction, enabling factors that are both more interpretable and more economically meaningful. Each cluster captures CMR with similar statistical behavior or economic profiles, resulting in localized factors that better represent specific dimensions of risk exposure. This separation improves predictive utility. The choice between economic and statistical clustering approaches is also important. Brown and Goetzmann [5] , for example, applied π -means clustering to mutual funds based on historical returns and sensitivities to stochastic variables, showing that the resulting groups outperformed traditional style classifications in explaining cross-sectional return variation. While our primary focus is on statistical clusteringβwhich relies purely on the data structure rather than pre-defined economic categoriesβwe also explore the implications of economic clustering as a robustness check. Both methods seek coherence within clusters, ensuring that the resulting groups reflect either shared economic rationale or statistical similarity. Motivated by the interpretability, specificity, and flexibility of clustering, we adopt clustering as a foundational component of our framework. In the next subsection, we outline the rationale for selecting hierarchical clustering as our clustering algorithm of choice and explain how its structure aligns with the objectives of factor construction in high-dimensional financial data. 3.2 Motivation for Hierarchical Clustering Hierarchical clustering (HC) offers distinct advantages for constructing latent factors in asset pricing. Unlike partitional clustering algorithms that require pre-specifying the number of clusters, HC generates a hierarchy of nested partitions, beginning either with singleton clusters and successively merging them (agglomerative) or with one large cluster and recursively splitting it (divisive). We employ the agglomerative approach for computational tractability and interpretability. This data-driven flexibility is particularly useful in settings where the optimal number of clustersβand hence latent factorsβis unknown ex ante, a challenge emphasized by Jiao et al. [21]. A notable strength of hierarchical clustering is its ability to produce a dendrogram, a tree-like diagram that visually represents how clusters merge at varying levels of dissimilarity (Figure 1). In this structure, each leaf node (green dot) represents a characteristics-managed return, and branches indicate how similar assets or portfolios combine into clusters. The vertical height at which two branches merge reflects their degree of dissimilarity, and a horizontal cut determines the final number of clusters. The resulting clusters are both data-driven and interpretable, offering a clear view of the latent structure underlying asset characteristics. Historically, hierarchical clustering has been widely used in disciplines requiring precise and interpretable groupings, such as biology and genomics, where taxonomies are built to reflect hierarchical relationships among entities. Its robustness, transparency, and ability to reveal multi-level structures make it particularly appealing for financial applications. In the context of factor modeling, hierarchical clustering ensures that the systematic risks captured by each cluster remain relatively independent, aligning with the fundamental assumption of factor orthogonality. By leveraging this approach, we aim to identify clusters that both minimize within-group dissimilarity and preserve between-group 3
independence, yielding latent factors that are distinct, interpretable, and consistent with the statistical properties of standard factor models. Figure 1: The green dots at the bottom represent unique characteristics-managed returns, while the ovals highlight the distinct clusters formed. The dashed line denotes the cutoff height determining the optimal number of clusters. 3.3 Methodological Framework for the Cluster-Based Factor Model (CBFM) 3.3.1 Hierarchical Clustering of Characteristics-Managed Returns This subsection presents our framework on hierarchical clustering characteristics managed returns based on how correlated they are. We denote {π₯π,π‘ }π π‘=1 as the π -length time series of the π -th characteristics-managed return (CMR), for π=1, . . . , π characteristics. The stacked characteristics managed returns as xπ=(π₯π,1, . . . , π₯π,π )β²βRπ and π=[x1, . . . , xπ] β RπΓπ. A key input to the hierarchical clustering algorithm is the dissimilarity matrix, which quantifies the pairwise dissimilarity between data points. While distance metrics are widely used as dissimilarity matrix, we derive our dissimilarity matrix based on correlation. We rely on a number of studies that regard correlation as a robust measure of similarity for returns. Also, Lance and Williams [27] posits that if one is interested in interrelationships between attributes rather than their absolute values, the correlation coefficient is appropriate. As a result, we define the dissimilarity between two characteristics-managed returns, xiand xjas one minus their historical correlation: π(π₯π, π₯ π)=1βππ₯ππ₯π, π· =[π(π₯π, π₯ π)]π π, π=1.(1) and our dissimilarity matrix of size πΓπ , with each element [π(π₯π, π₯ π)]π π, π=1β [0,2] . The higher the dissimilarity score between two characteristics managed returns, the more uncorrelated they are. We define the historical correlation as ππ₯ππ₯π=Γπ π‘=1(π₯π,π‘ βΒ―π₯π)(π₯π,π‘ βΒ―π₯π) βοΈΓπ π‘=1(π₯π,π‘ βΒ―π₯π)2βοΈΓπ π‘=1(π₯π,π‘ βΒ―π₯π)2 ,Β―π₯π=1 π π βοΈ π‘=1 π₯π,π‘ . 4
Following the dissimilarity matrix, we are next faced with forming clusters, where each of the cluster contains highly correlated or similar characteristics managed returns. We make use of the Agglomerative hierarchical clustering (HC) and LanceβWilliams update. First, we begin with singleton clusters, where each cluster contain single characteristic managed return, i.e C={πΆ1, . . . , πΆπ} with πΆπ={π₯π} , and then we merge clusters upward until more merging no longer improve the result. During merging, we update our dissimilarity matrix using the Lance-Williams dissimilarity update formula (Lance and Williams [27]); π(πΆπβͺπ, πΆπ)=πΌππ(πΆπ, πΆπ) + πΌππ(πΆπ, πΆπ) + π½π(πΆπ, πΆπ) + πΎξξπ(πΆπ, πΆπ) β π(πΆπ, πΆπ)ξξ,(2) where the two clusters being merged are denoted as πΆπ and πΆπ , while πΆπ represents any other cluster not involved in the merge. The updated distance between the newly formed cluster πΆπβͺπ and πΆπ is given by π(πΆπβͺπ, πΆπ) , which is calculated using equation 2. The distances between clusters πΆπ and πΆπ , as well as between πΆπ and πΆπ , are represented by π(πΆπ, πΆπ) and π(πΆπ, πΆπ) , respectively. The distance between the two clusters before merging is π(πΆπ, πΆπ) . The coefficients πΌπ , πΌπ , π½, and πΎare determined based on choice of linkage method. We choose the complete linkage method, as it defines the distance between two clusters as the maximum distance between any two characteristics-managed returns variables in different clusters. This method is particularly effective in scenarios where tightly-knit clusters with minimal within-cluster variation are desired. Unlike single linkage, which can result in the chaining effect and produce elongated, loosely connected clusters, complete linkage emphasizes cohesion within groups, ensuring the formation of compact and distinct clusters. This characteristic makes it a natural choice for our research, where an intensive grouping strategy is essential to maintain the integrity and interpretability of the resulting clusters (Lance and Williams [27] ). However, our methodology remains flexible and accommodates alternative linkage methods as needed. For a comprehensive review of other linkage methods, we refer the reader to Lance and Williams [27]. For complete linkage, the coefficients πΌπ=1 2 , πΌπ=1 2 , π½=0 , and πΎ=β1 2 in eqn 2, and the LanceβWilliams dissimilarity update formula reduces to the maximum distance criterion: π(πΆπβͺπ, πΆπ)=max{π(πΆπ, πΆπ), π(πΆπ, πΆπ)}.(3) Algorithm 1 Complete Linkage Hierarchical Clustering 1: procedure COMPLETE LINKAGE(π·)β² π· is a given dissimilarity matrix 2: Initialization: C β {πΆ1, πΆ2, . . . , πΆπ} , where each πΆπ={π₯π} is a singleton cluster of distinct characteristic managed return. 3: while |C| >1do 4: Find the pair of clusters (πΆπ, πΆπ)such that: (πΆπ, πΆπ)=arg min πΆπ,πΆπβ C, πΆπβ πΆπ π·(π₯π, π₯ π) 5: Merge the clusters πΆπand πΆπto form a new cluster πΆπβͺπ 6: Update the dissimilarity matrix by computing π·(πΆπβͺπ, πΆπ)for all remaining clusters πΆπβ C \ {πΆπ, πΆπ}: π·(πΆπβͺπ, πΆπ)=max π₯πβπΆπβͺπ, π₯πβπΆπ π·(π₯π, π₯ π) 7: Replace πΆπand πΆπin Cwith πΆπβͺπ 8: end while 9: Output: The dendrogram representing the hierarchical structure with a cut producing πΎ clusters via elbow or silhouette score 10: end procedure The process of forming clusters from the characteristics-managed returns is summarized in Algorithm 1. Starting with characteristics managed returns of dimensions πΓπ , we first calculate the time-weighted dissimilarity matrix using equation 1. This matrix serves as the input for the complete linkage hierarchical clustering algorithm. Initially, each CMR is treated as a distinct cluster. The algorithm then identifies the pair of CMR with the smallest dissimilarity 5
(i.e., the pair exhibiting the highest correlation) and merges them into the same cluster. Subsequently, the dissimilarity matrix is updated using the Lance-Williams update rule in equation 2, with the coefficients determined by the complete linkage method. Choosing the optimal number of clusters, πΎ , is critical in hierarchical clustering, as it affects both interpretability and model robustness. Choosing too few clusters may merge unrelated economic themes, while too many can fragment highly related variables. To determine πΎ , we apply two complementary approaches: the Elbow Method and the Silhouette Method. The Elbow Method identifies a natural breakpoint, or βelbow,β in the sequence of dissimilarity values from the clustering process, marking where additional clusters provide diminishing improvement in within-cluster homogeneity. If the true number of clusters is πΎβ , then for any πΎ < πΎβ , the resulting clusters remain subsets of the underlying structure. We detect the elbow by computing the second-order differences of the dissimilarity sequence: Ξ2π·π=π·π+1β2π·π+π·πβ1,(4) and selecting the point where Ξ2π·πsharply declines. To validate cluster quality, we also use the Silhouette Method, which quantifies both intra-cluster compactness and inter-cluster separation. For each characteristic π, the silhouette score is: π (π)=π(π) β π(π) max{π(π), π(π)} ,(5) where π(π) denotes the average dissimilarity of π within its cluster, and π(π) is the minimum average dissimilarity of π to the nearest neighboring cluster. The mean silhouette score for a configuration with πΎclusters is π(πΎ)=1 π π βοΈ π=1 π (π),(6) with higher values indicating better-defined clusters. Finally, we distinguish among three clustering modes used in this study: (i) Statistical Clustering, employing hierarchical clustering as described above; (ii) Economic Clustering, which groups characteristics by economic themes from prior literature; and (iii) Categorical Clustering, which classifies variables based on data sources following [6]1. 3.3.2 Local Linear Embedding Next, we present the second step of our cluster-based factors: locally Linear Embedding (LLE), a non-linear dimension reduction method. Dimension reduction is crucial as it helps address the problem of high dimensionality, a concern raised by Cochrane [7]. Locally Linear Embedding (LLE), introduced by Roweis and Saul [31] , is a non-linear dimensionality reduction technique that focuses on preserving the local geometric structure of data in the embedding space. This fits into our goal as our key objective is to obtain a faithful representation of the clustered data of characteristics managed returns (CMR) in a lower-dimensional space. Faithful representation, in this context, implies that the relationships in the original data are preserved i.e similar characteristics managed returns in high-dimensional space are appropriately weighted in the reduced space. This ensures that factor representations respect the underlying relationships within each cluster. By emphasizing local geometry, LLE captures the most relevant structures that influence assets returns without being overly influenced by global distortions. Another significant motivation of using LLE is the simplicity in hyperparameter tuning. LLE only requires two hyperparameters: the number of nearest neighbors and the target dimensionality. In our application, one of these parametersβthe target dimensionality per clusterβis predefined to be 1 component (factor) per cluster based on the fact that we have clustered highly correlated characteristics managed returns into distinct clusters and expect that one factor should represent all the characteristics in the specific cluster. This leaves only the nearest neighbor parameter to tune, making LLE a computationally efficient and easy-to-implement choice while still flexible. This simplicity ensures consistency and reduces the risk of overfitting or suboptimal performance due to improper parameter settings. Building on the output of clusters of characteristics managed returns. Suppose there are πΎ optimal clusters from our hierarchical clustering denoted as πΆ1, πΆ2, . . . πΆπΎ with each cluster containing ππ CMR grouped together, our factor representation using locally linear embedding proceeds as follow: 1A full mapping of characteristics and categories is available at www.openassetpricing.com. We also show it in Table 4 6
Algorithm 2 Dimensionality Reduction with LLE for Clustered Characteristics Require: πΎ clusters of characteristic-managed returns π , each cluster of dimension πΓππ , where ππ is the number of characteristics in cluster π Ensure: Reduced dimensional factor ππof dimension πΓ1 in each cluster 1: Step 1: Apply LLE to Each Cluster 2: for each cluster πΆπβ {πΆ1, πΆ2, . . . , πΆπΎ} β RπΓππdo 3: Step 2a: Compute Nearest Neighbors to each characteristic: 4: for each characteristic xiβCkdo 5: Find the πβ²nearest neighbors of π§π‘based on Euclidean distance in the ππ-dimensional space. 6: end for 7: Step 2b: Compute Reconstruction weights 8: for each time sample π§π‘βππdo 9: Compute weights π€π π that best reconstruct π₯πusing its πβ²neighbors: Minimize Γππ π=1ξ ξ ξxiβΓπβ² π=1Λπ€π π π₯π π ξ ξ ξ 2 2subject to Γπβ²βN(π₯π)Λπ€π π =1 10: end for 11: Step 2c: Compute Low-Dimensional Embedding 12: Solve the eigenvector problem to minimize the embedding cost: Minimize Γππ π=1ξ ξπ¦πβΓπβ²βN(π₯π)π€π π π¦πξ ξ2 2 13: Subject to the constraints: 1 πΓππ π=1π¦π=0 1 πΓππ π=1π¦2 π=1 14: Obtain the eigenvector corresponding to the second smallest non-zero eigenvalue as the low-dimensional embedding ππof dimension πΓ1. 15: end for 16: Step 3: Collect the Factors 17: Combine the factors {π1, π2, . . . , ππΎ} , each of dimension πΓ1 , representing each cluster, into the final set of factors. 18: Return the set of reduced dimensional factors π={π1, π2, . . . , ππΎ}, each of dimension πΓ1 1. Nearest Neighbors and Reconstruction of Weights: Fix a cluster πΆπ with ππ=|πΆπ| denoting the number of CMR in each cluster and matrix ππβRπΓππ (columns standardized). Let π½be the number of nearest neighbors (NN). Step 1: NN graph (column-wise). For each column xπ of ππ , we find CMR whose index set N (π) β {1, . . . , ππ} \ {π} with |N (π)| =π½using Euclidean distance in Rπ. Step 2: Reconstruction weights. For each CMR π , we form neighbor matrix ππ=[xπ]πβN(π)βRπΓπ½ and compute local covariance πΊπ=(xπ1β² π½βππ)β²(xπ1β² π½βππ) β Rπ½Γπ½. With a small Tikhonov regularizer π > 0 (π=10β6) to ensure invertibility for near-collinear neighbors, set Λ wπ=(πΊπ+ππΌπ½)β11π½ 1β² π½(πΊπ+ππΌπ½)β11π½ βRπ½,βοΈ πβN(π) Λπ€π π =1.(7) Assemble the sparse weight matrix πππβRππΓππ with rows π having nonzeros (πππ)π π =Λπ€π π for πβ N (π) and 0 otherwise. Step 3: Embedding (eigenproblem). Next, we solve the eigendecomposition problem ππYπ=YπΞπ,where ππ=(πΌππβπππ)β²(πΌππβπππ),(8) 7
where columns of YπβRππΓππ are the eigenvectors associated with the ππ smallest nontrivial eigenvalues Ξπ (the trivial eigenvector 1 with eigenvalue 0 is discarded). For our factor construction, we set ππ=1 (one factor per cluster). Let yπβRππbe that eigenvector (normalized to β₯yπβ₯2=1). Step 4: Cluster factor time series. Finally, we project the CMR panel in cluster π onto the embedding weights to obtain a πΓ1 factor: fπ=ππyπβRπ.(9) Collect F=[f1, . . . , fπΎ] β RπΓπΎ. 3.3.3 Factor Loadings Estimation We modify the assumption that asset returns Rπ‘βRπΓ1 follow a πΎ -factor structure in a linear form to a non linear structure : Rπ‘=πΊ(πΆπ‘β1,π·β² π‘β1|fπ‘) + πΊπ‘,Eπ‘[πΊπ‘]=0,Eπ‘[πΊπ‘fβ² π‘]=0.(10) Here, fπ‘βRπΎΓ1 are the latent factors constructed via hierarchical clustering and locally linear embedding (Algorithms 1 and 2); π·π‘β1 represents time-varying factor loadings, and πΆπ‘β1 captures pricing errors. When πΆπ‘β1=0 , fπ‘ spans all systematic risk in the cross-section of returns. We first estimate π·π‘β1 non-linearly using a feedforward neural network that maps latent factors fπ‘ to predicted returns Rπ‘. The recursive transformation is: h(0)=fπ‘, h(π)=πξB(π)h(πβ1)+πΆ(π)ξ, π =1, . . . , πΏ β1, b Rπ‘=B(πΏ)h(πΏβ1)+πΆ(πΏ), (11) where B(π)βRππΓππβ1 and πΆ(π)βRππ denote layer weights and biases, and π(Β·) is an activation function (ReLU or tanh). The output weights B(πΏ)βRπΓπΎ represent estimated factor loadings π·π‘β1 , while πΆ(πΏ) represents idiosyncratic intercepts. Figure 2: Neural-network architecture for estimating time-varying loadings. The input layer contains πΎ latent factors; two hidden layers with π1and π2neurons capture non-linear interactions; the output layer produces πpredicted returns. 8
The full mapping from latent factors to predicted returns is: Rπ‘=B(πΏ)πξB(πΏβ1)Β· Β· Β· πξB(1)fπ‘+πΆ(1)ξ+ Β· Β· Β· + πΆ(πΏβ1)ξ+πΆ(πΏ).(12) This setup allows non-linear and time-varying dependence between returns and latent factors, improving flexibility relative to linear regression models. In summary, we estimate factor loadings using non-linear conditioning on latent factors fπ‘ which enhances the explanatory power of our cluster-based factor model (CBFM) relative to linear alternatives. 4. Data and Results 4.1 Data description We make use of characteristics managed portfolio returns provided by Chen and Zimmermann [6]2 . We exclude characteristics managed returns with missing values during our sample period, which spans from January 1964 to December 2023 so as not to have imbalanced panel of CMR. This approach eliminates the need for imputation and allows us to focus on 130 out of the 212 characteristics originally provided in the dataset. The choice of this sample period is driven by data availability, as comprehensive accounting information is sparsely recorded before 1964, and because key developments in asset pricing research emerged around that time. The selected characteristics are continuous firm-specific attributes classified into six financial categories: Price, Trading, Accounting, Event, 13F and other, as defined by Chen and Zimmermann [6] . Following Brandt et al. [4] and DeMiguel et al. [10] , each predictor is standardized to have a cross-sectional mean of zero and a standard deviation of one. Table 4 in the Appendix lists these characteristics along with their respective financial and economic classifications, as specified by Chen and Zimmermann [6] . Each characteristic is associated with managed portfolio returns, constructed based on the methodologies outlined in the original research papers3. 4.2 Factors from Cluster Results In this section, we now present some of our findings of our cluster-based factor when applied to 130 characteristicsmanaged returns. First, we distinguish between clusters that are based on economic intuition, statistical correlation, and category of the characteristics. The 130 characteristics we focus on in this work are economically grouped into 27 themes, which are valuation, accruals, investment, leverage, risk, other, liquidity, profitability, profitability alt, sales growth, investment alt, investment growth, external financing, payout indicator, volume, lead lag, earnings growth, info proxy, momentum, volatility, long term reversal, asset composition, R&D, short sale constraints, short-term reversal, size, and cash flow risk. We take each of these themes as a cluster, which implies that the first stageβhierarchical clustering of our cluster-based factorβis skipped. We also treat the source of the characteristics as clusters, in what we term the categorical clustering. Categorically, our characteristics are grouped into 6 themes, which are Accounting, Price, Trading, Event, Other, and 13F. Statistically, we cluster CMR based on hierarchical clustering highlighted in earlier section. First, hierarchical clustering groups the CMR into several buckets such that CMR in each bucket are highly correlated. We address the question of how many clusters the CMR should be grouped into. To determine the optimal number of clusters, we use the elbow method and the Silhouette score. Figure 3 plots the difference score, the first-order difference, and the second-order difference as the number of clusters increases. We identify the point where the drop in the difference score becomes negligible for an increasing number of clusters and consider the largest drop in the second-order difference. Both criteria suggest an optimal number of clusters at πΎ=3. To verify the robustness of this clustering choice, we calculate the average Silhouette score (Table 1) for all characteristics within each cluster. With Silhouette score range from -1 to 1, where a high value indicates that characteristic is well matched to its own cluster and poorly matched to neighboring clusters, we find that the optimal silhouette score of 0.79 occur at πΎ=3 . This combination of methodsβdifference scores and Silhouette scoresβprovides a strong justification for selecting three clusters for our statistical clustering. Having established our findings of 3 clusters as the optimal number, we visualize how the characteristics are grouped using the dendrogram in Figure 4 4 . The dashed red horizontal line represents a cut-off height implicitly defined by the optimal clusters identified using the Silhouette score and Elbow method. We observe that cluster 1 primarily includes variables associated with financial leverage and capital structure, such as Net Debt to Price and CAPM Beta. 2 We extend our gratitude to Chen and Zimmermann [6] for making their dataset publicly available at https://www.openassetpricing.com/ data/ 3See Chen and Zimmermann [6] for references to the original papers that documented each characteristic. 4A clearer version of Fig 4 on how CMR cluster is in Fig 6 9
Appendix Data: Characteristics composition Table 4: Characteristics, Data Categories and Economic Categories Acronym LongDescription Category Economic AM1 Total assets to market Accounting valuation Acc2 Accruals Accounting accruals Ass3 Asset growth Accounting investment BM4 Book to market, original (Stattman 1980) Accounting valuation BMd5 Book to market using December ME Accounting valuation BPE6 Leverage component of BM Accounting leverage Bet7 Tail risk beta Price risk Bet8 CAPM beta Price risk Bet9 Frazzini-Pedersen Beta Price other Bid10 Bid-ask spread Trading liquidity Boo11 Book leverage (annual) Accounting leverage CBO12 Cash-based operating profitability Accounting profitability CF13 Cash flow to market Accounting valuation Cas14 Cash Productivity Accounting profitability alt ChA15 Change in Asset Turnover Accounting sales growth ChE16 Growth in book equity Accounting investment ChI17 Inventory Growth Accounting investment alt ChI18 Change in capital inv (ind adj) Accounting investment growth ChN19 Change in Net Noncurrent Op Assets Accounting investment alt ChN20 Change in Net Working Capital Accounting investment alt ChT21 Change in Taxes Accounting other Com22 Composite equity issuance Accounting external financing Com23 Composite debt issuance Accounting external financing Con24 Convertible debt indicator Accounting external financing Cos25 Coskewness using daily returns Price risk Cos26 Coskewness Price risk Del27 Change in current operating assets Accounting investment alt Del28 Change in current operating liabilities Accounting external financing Del29 Change in equity to assets Accounting investment Del30 Change in financial liabilities Accounting external financing Del31 Change in long-term investment Accounting investment Del32 Change in net financial assets Accounting investment alt Div33 Dividend Initiation Event payout indicator Div34 Dividend seasonality Event payout indicator Div35 Predicted div yield next month Accounting valuation Dol36 Past trading volume Trading volume EBM37 Enterprise component of BM Accounting valuation EP38 Earnings-to-Price Ratio Accounting valuation Ear39 Earnings consistency Accounting earnings growth Ear40 Earnings Surprise Accounting earnings growth Ear41 Earnings surprise of big firms Accounting lead lag Ent42 Enterprise Multiple Accounting valuation Equ43 Equity Duration Accounting valuation Exc44 Exchange Switch Event other Fir45 Firm age based on CRSP Other info proxy Continued on next page 16
Acronym LongDescription Category Economic Fro46 Efficient frontier index Accounting valuation GP47 gross profits / total assets Accounting profitability GrL48 Growth in long term operating assets Accounting investment GrS49 Sales growth over inventory growth Accounting sales growth GrS50 Sales growth over overhead growth Accounting sales growth Her51 Industry concentration (sales) Other other Her52 Industry concentration (equity) Other other Her53 Industry concentration (assets) Other other Hig54 52 week high Price momentum Idi55 Idiosyncratic risk (3 factor) Price volatility Idi56 Idiosyncratic risk (AHT) Price volatility Ill57 Amihudβs illiquidity Trading liquidity Ind58 Industry Momentum Price momentum Ind59 Industry return of big firms Price lead lag Int60 Intangible return using CFtoP Accounting long term reversal Int61 Intangible return using EP Accounting long term reversal Int62 Intangible return using Sale2P Accounting long term reversal Int63 Intermediate Momentum Price momentum Inv64 Investment to revenue Accounting investment Inv65 change in ppe and inv/assets Accounting investment Inv66 Inventory Growth Accounting profitability LRr67 Long-run reversal Price long term reversal Lev68 Market leverage Accounting leverage MRr69 Medium-run reversal Price long term reversal Max70 Maximum return over month Price volatility Mea71 Revenue Growth Rank Accounting sales growth Mom72 Momentum (12 month) Price momentum Mom73 Momentum without the seasonal part Price other Mom74 Momentum (6 month) Price momentum Mom75 Off season long-term reversal Price other Mom76 Off season reversal years 6 to 10 Price other Mom77 Off season reversal years 16 to 20 Price other Mom78 Momentum and LT Reversal Price momentum Mom79 Return seasonality years 2 to 5 Price other Mom80 Return seasonality years 6 to 10 Price other Mom81 Return seasonality years 11 to 15 Price other Mom82 Return seasonality years 16 to 20 Price other Mom83 Return seasonality last year Price other Mom84 Momentum in high volume stocks Price momentum Mom85 Off season reversal years 11 to 15 Price other NOA86 Net Operating Assets Accounting asset composition Net87 Net debt to price Accounting leverage Net88 Net Payout Yield Accounting valuation Num89 Earnings streak length Accounting earnings growth OPL90 Operating leverage Accounting other Ope91 Operating profits / book equity Accounting profitability Ope92 Operating profitability R&D adjusted Accounting profitability Org93 Organizational capital Accounting R&D Pay94 Payout Yield Accounting valuation Pri95 Price Price other Continued on next page 17
Acronym LongDescription Category Economic RD96 R&D over market cap Accounting R&D RDA97 R&D ability Accounting other RIO98 Inst Own and Market to Book 13F short sale constraints RIO99 Inst Own and Turnover 13F short sale constraints RIO100 Inst Own and Idio Vol 13F short sale constraints Rea101 Realized (Total) Volatility Price volatility Res102 Momentum based on FF3 residuals Price momentum Ret103 Return skewness Price risk Ret104 Idiosyncratic skewness (3F model) Price risk Rev105 Revenue Surprise Accounting sales growth RoE106 net income / book equity Accounting profitability SP107 Sales-to-price Accounting valuation STr108 Short term reversal Price short-term reversal Sha109 Share issuance (1 year) Accounting external financing Sha110 Share issuance (5 year) Accounting external financing Sha111 Share Volume Trading volume Siz112 Size Price size Spi113 Spinoffs Event other Sur114 Unexpected R&D increase Accounting R&D Tax115 Taxable income to income Accounting other Tot116 Total accruals Accounting investment alt Tre117 Trend Factor Price momentum Var118 Cash-flow to price variance Accounting cash flow risk Vol119 Volume Variance Trading liquidity Vol120 Volume to market equity Trading volume Vol121 Volume Trend Trading volume dNo122 Change in net operating assets Accounting investment grc123 Change in capex (two years) Accounting investment growth grc124 Change in capex (three years) Accounting investment growth sin125 Sin Stock (selection criteria) Other other std126 Share turnover volatility Trading liquidity tan127 Tangibility Accounting asset composition zer128 Days with zero trades Trading liquidity zer129 Days with zero trades Trading liquidity zer130 Days with zero trades Trading liquidity Clusters of Characteristics 18
Figure 6: Visualization of the three clusters of CMR. Each cluster is visually separated and labeled along the horizontal axis, with random vertical spacing used to distribute the points for clarity. This layout deals with the issue of overlapping labels in the original dendrogram (Fig 4), allowing for a clearer view of the characteristics that belong to each group. Cluster 1 includes variables primarily associated with financial leverage and capital structure, such as Net Debt to Price and CAPM Beta. Cluster 2 captures momentum and earnings-related themes, with characteristics such as 12-Month Momentum and Earnings Streak Length. Cluster 3, the most diverse group, encompasses factors like Size, Book-to-Market, and Asset Growth, representing firm valuation and investment activities. 19