Understanding Matrix Function Normalizations in Covariance Pooling from the Lens of Riemannian Geometry
Full text
Published as a conference paper at ICLR 2025 UNDERSTANDING MATRIX FUNCTION NORMALIZATIONS IN COVARIANCE POOLING THROUGH THE LENS OF RIEMANNIAN GEOMETRY Ziheng Chen1, Yue Song1∗ , Xiao-Jun Wu2, Gaowen Liu3& Nicu Sebe1 1University of Trento, 2Jiangnan University, 3Cisco Systems ziheng [email protected], [email protected] ABSTRACT Global Covariance Pooling (GCP) has been demonstrated to improve the performance of Deep Neural Networks (DNNs) by exploiting second-order statistics of high-level representations. GCP typically performs classification of the covariance matrices by applying matrix function normalization, such as matrix logarithm or power, followed by a Euclidean classifier. However, covariance matrices inherently lie in a Riemannian manifold, known as the Symmetric Positive Definite (SPD) manifold. The current literature does not provide a satisfactory explanation of why Euclidean classifiers can be applied directly to Riemannian features after the normalization of the matrix power. To mitigate this gap, this paper provides a comprehensive and unified understanding of the matrix logarithm and power from a Riemannian geometry perspective. The underlying mechanism of matrix functions in GCP is interpreted from two perspectives: one based on tangent classifiers (Euclidean classifiers on the tangent space) and the other based on Riemannian classifiers. Via theoretical analysis and empirical validation through extensive experiments on fine-grained and large-scale visual classification datasets, we conclude that the working mechanism of the matrix functions should be attributed to the Riemannian classifiers they implicitly respect. The code is available at https://github.com/GitZH-Chen/RiemGCP.git. 1 INTRODUCTION Global Covariance Pooling (GCP), a method used as a replacement for Global Average Pooling (GAP) in aggregating the final activations of Deep Neural Networks (DNNs), has demonstrated exceptional performance improvements across various applications (Lin et al.,2015;Ionescu et al., 2015;Li et al.,2017;Wang et al.,2017;Koniusz et al.,2017;Li et al.,2018;Wang et al.,2020a; Rahman et al.,2020;Zhu et al.,2024). The research line of existing GCP methods mainly focuses on improving performance by adopting different normalization methods (Ionescu et al.,2015;Li et al., 2017;Wang et al.,2020a), exploiting richer statistics (Cui et al.,2017;Wang et al.,2017;Koniusz et al.,2021;Rahman et al.,2023;Chen et al.,2023a), improving covariance conditioning (Song et al.,2022d;a), and obtaining compact representations (Gao et al.,2016;Yu & Salzmann,2018; Lin et al.,2018;Wang et al.,2022a). Generally, a GCP meta-layer computes the covariance matrix of the activations as the global representation, and then performs normalization either by matrix logarithm (Ionescu et al.,2015) or matrix power (Li et al.,2017;2018;Wang et al.,2020a). Finally, the normalized matrices are fed into a Euclidean classifier. The square root has emerged as the most effective normalization scheme, outperforming the logarithm counterpart by a large margin (Li et al., 2017;Wang et al.,2020a;Song et al.,2021). Although the research community has provided some theoretical support for the matrix logarithm, there are no intrinsic explanations for the matrix power. The covariance matrices naturally lie in a Riemannian manifold, known as Symmetric Positive Definite (SPD) manifolds (Pennec et al.,2006). For matrix logarithm, it maps SPD matrices into the Euclidean space of the tangent space at the identity matrix. Euclidean classifiers can, therefore, be applied after the matrix logarithm. However, the co-domain of the matrix power is still the SPD ∗Corresponding author 1
Published as a conference paper at ICLR 2025 manifold, rendering the application of Euclidean classifiers following matrix power less mathematically supported. Several works have attempted to explain the matrix power. The initial motivation of matrix power in GCP (Li et al.,2017) is that the distance induced by the matrix square root approximates the geodesic distance under Log-Euclidean Metric (LEM) (Dryden et al.,2010). Nevertheless, Fig. 1shows that the gap between these two distances is still noticeable. Furthermore, Song et al. (2021) empirically explored the benefits of approximate matrix square root over its accurate counterpart, while Wang et al. (2020b) studied the merits of GCP from an optimization perspective. However, none of them touch upon the fundamental reason why Euclidean classifiers can be directly employed in the non-Euclidean co-domain of the matrix power. There appears to be a discrepancy between theoretical principles and practical applications of matrix power and logarithm. 0 200 400 600 800 1000 Matrix Pair Index 0 100 200 300 400 Distance Distances under LEM and PEM LEM PEM Figure 1: The metric induced by the matrix power is the Power Euclidean Metric (PEM). Although PEM approaches LEM as the power approaches 0, the distances under PEM (θ= 0.5) and LEM still differ largely. We visualize these two distances for 1000 random pairs of 256×256 SPD matrices. The average difference is 335.84 ± 1.61. This indicates that PEM is not proximate to LEM under the widely used θ= 0.5. This study aims to offer a comprehensive theoretical understanding of the matrix logarithm/power in GCP and reconcile the discrepancy between theory and practice. Without loss of generality, we refer to matrix logarithm and power collectively as matrix functions. Given that the matrix logarithm is a Riemannian logarithmic map, mapping SPD data into the tangent space, we first systematically study Riemannian logarithmic maps on SPD manifolds under seven families of metrics, resulting in three types of Riemannian logarithmic maps, the ones based on the matrix logarithm, matrix power, and Log-Cholesky Metric (LCM), respectively. Consequently, the matrix logarithm in GCP establishes a tangent classifier (Euclidean classifiers on the tangent space) for covariance classification. Also, by applying a simple affine transformation, the matrix power in GCP constructs a tangent classifier. This indicates that we might unify both matrix logarithm and power as tangent classifiers. However, our experiments suggest that this tangent classifier explanation fails to account for the efficacy of matrix power. As the tangent space distorts the intrinsic geometry of manifolds, we conjecture that tangent classifiers might not be the underlying mechanisms. To delve further, we move on to a more intrinsic explanation based on the recently developed SPD Multinomial Logistics Regression (MLR) (Nguyen & Yang,2023;Chen et al.,2024a;d), which extends the Euclidean MLR (FC + softmax) into manifolds. Based on previous work (Chen et al., 2024a), we find that matrix logarithm in GCP implicitly constructs the SPD MLR under LEM. Furthermore, we theoretically demonstrate that the matrix power in GCP implicitly respects the SPD MLR under PEM. These findings suggest that matrix functions in GCP can be uniformly interpreted as Riemann classifiers. Therefore, the observed performance gap between the matrix power and logarithm can be attributed to the characteristics of the underlying Riemannian metrics. To validate this postulation, we conduct experiments on the ImageNet-1k (Deng et al.,2009) and three FineGrained Visual Categorization (FGVC) datasets, namely Caltech University Birds (Birds) (Welinder et al.,2010), Stanford Cars (Cars) (Krause et al.,2013), and FGVC Aircrafts (Aircrafts) (Maji et al.,2013). The results confirm that the Riemannian classifier rather than the tangent classifier contributes to the efficacy of matrix functions in GCP. We expect our work to pave the way for a deeper theoretical understanding of GCP from a Riemannian perspective and inspire more research to explore the rich SPD geometries for more effective GCP applications. We present a teaser table in Tab. 1. Due to page limits, we put the related work in App. Band all the proofs in the appendix. Besides, tables of notations and abbreviations are presented in App. Cfor better readability. In summary, our main contributions are two-fold. (a). First intrinsic explanation for matrix normalization. We explain the working mechanism of matrix functions in GCP from the perspectives of tangent and Riemannian classifiers, and finally claim that the rationality of matrix functions should be attributed to the Riemannian classifiers they implicitly respect. To the best of our knowledge, this is the first Riemannian interpretation of the matrix functions in GCP. (b). Empirical validation by extensive experiments. We validate our theoretical argument on large-scale and FGVC datasets. 2
Published as a conference paper at ICLR 2025 Table 1: Main results: The working mechanisms of matrix functions in GCP are attributed to Riemannian classifiers they implicitly respect. Matrix function Intrinsic explanation Used in GCP Reference Logarithm LEM-induced Riemannian Classifier Log-EMLR (Eq. (4)) (Chen et al.,2024a, Prop. 5.1) Power PEM-induced Riemannian Classifier Pow-EMLR (Eq. (5)) Thm. 2 (B1) Refuted by experiments 1 Matrix power/logarithm 😊 Tangent Classifiers (A1) Power approximates a tangent classifier 😞 Riemannian Classifiers (B2) Validated empirically and theoretically Similar distances under matrix power and logarithm (𝜃 → 0) (A0) Previously (A2) They implicitly constructs Riemannian classifiers (respecting different metrics) (B0) Refuted by experiments (𝜃=0.5) Log respects a tangent classifier Figure 2: Illustration on the main postulations (A0 to A2) and empirical validations (B0 to B2) of our investigation, where θis the power in the matrix power. Postulation A0 is adopted by Li et al. (2017) and is refuted by our experiments in Fig. 1for the specific θ= 0.5. Postulation A1 is indicated by Tab. 2and is refuted by our experiments in Sec. 6. Postulation A2 is supported by Thm. 2and is validated by our experiments in Sec. 6. Main theoretical results: Tab. 2presents a complete list of Riemannian logarithmic maps on SPD manifolds under different metrics. It indicates that the matrix power, with a simple affine transformation, can serve as a Riemannian logarithmic map. Since the matrix logarithm has been widely recognized as a building component of tangent classifiers, we also expect that the matrix power function can be explained by tangent classifiers. However, the preliminary experiments presented in Tab. 3refute this conjecture, suggesting the existence of more fundamental mechanisms. Therefore, we delve into this mystery in Sec. 5by leveraging the recently developed Riemannian classifiers. Thm. 2indicates that matrix power in GCP implicitly establishes a Riemannian classifier for covariance matrix classification. Similar results also hold for the matrix logarithm (Chen et al.,2024a). This implies that the matrix logarithm and power can be unifiedly interpreted as essential components of Riemannian classifiers. Tab. 4summarizes all our theoretical findings. Sec. 6further validate our theoretical explanations by extensive experiments. The reasoning behind our analysis is illustrated in Fig. 2. 2 THE GEOMETRY OF SPD MANIFOLDS Let Sn ++ be the set of n×nSPD matrices. As shown by Arsigny et al. (2005), Sn ++ is an open submanifold of the Euclidean space Snof symmetric matrices. There are five popular Riemannian metrics on SPD manifolds: Affine-Invariant Metric (AIM) (Pennec et al.,2006), Log-Euclidean Metric (LEM) (Arsigny et al.,2005), Power-Euclidean Metric (PEM) (Dryden et al.,2010), LogCholesky Metric (LCM) (Lin,2019), and Bures-Wasserstein Metric (BWM) (Bhatia et al.,2019). Note that when power equals 1, PEM reduces to the Euclidean Metric (EM). All of the above five standard metrics have been generalized into parameterized families of metrics. Thanwerdas & Pennec (2023) generalized AIM, LEM, and EM into two-parameter metrics by the O(n)-invariant Euclidean inner product on Sn: ⟨V, W⟩(α,β)=α⟨V, W⟩+βtr(V) tr(W),(1) where α > 0and β > −α /n,V, W ∈ Sn, and ⟨·,·⟩ is the standard matrix inner product. These generalized metrics are denoted as (α, β)-AIM, (α, β)-LEM, and (α, β)-EM, respectively. Besides, Thanwerdas & Pennec (2022) defined the power-deformed metric ˜gof a metric gon Sn ++ as ˜gP(V, W) = 1 θ2gPθ((Powθ)∗,P (V),(Powθ)∗,P (W)) ,∀P∈ Sn ++, V, W ∈TPSn ++,(2) where Powθ(P) = Pθdenotes matrix power, (Powθ)∗,P is the differential map of Powθat P, and TPSn ++ is the tangent space at P. Following Eq. (2), (α, β)-AIM, (α, β)-LEM, (α, β)-EM, LCM, 3
Published as a conference paper at ICLR 2025 and BWM are generalized into power-deformed families of metrics, denoted as (θ, α, β)-AIM, (θ, α, β)-LEM, (θ, α, β)-EM, θ-LCM, and 2θ-BWM, respectively (Thanwerdas & Pennec,2022; Chen et al.,2024d). Chen et al. (2024d) further shows that (θ, α, β)-LEM equals (α, β)-LEM. Besides, as shown by Thanwerdas & Pennec (2022); Chen et al. (2024d), θserves as a deformation from a LEM-like metric. For instance, (θ, α, β)-AIM becomes (α, β)-AIM when θ= 1 and approaching (α, β)-LEM with θ→0. On the other hand, PEM was generalized into Mixed Power Euclidean Metric (MPEM) (Thanwerdas & Pennec,2019) by two power factors θ1and θ2, denoted as (θ1, θ2)-EM. When θ1=θ2, MPEM is reduced to PEM. Han et al. (2023) extended BWM into Generalized Bures-Wasserstein Metric (GBWM) by an SPD parameter M, denoted as M-BWM. When M=I, GBWM is reduced to BWM. We further generalize GBWM into (2θ, M)-BWM by power deformation Eq. (2). In total, three parameters (θ, α, β)are involved in the metrics on the SPD manifold. The power deformation θcharacterizes deformation (Chen et al.,2024d;Thanwerdas & Pennec,2022), while (α, β)characterizes the O(n)-invariance, a powerful property in modeling covariance (Thanwerdas & Pennec,2023). The above metrics have shown success in various applications, due to their closedform expressions for Riemannian operators, such as the Riemannian logarithmic and exponential maps. We summarize the involved Riemannian operators in App. D.2. 3 GLOBAL COVARIANCE POOLING REVISITED GCP captures the second-order statistics of the features in the last layer of the deep network. The standard GCP procedure comprises calculating the covariance matrix, normalization with a matrix function, vectorization, dimensionality reduction by an FC layer, and ultimately applying a Euclidean classifier. The sequence of these operations can be represented as follows: XCov −−→ ΣfM −−→ ˜ Σfvec −−→ xfFC −−→ ˜xfEC −−→ ˆy, (3) where fM,fvec,fFC and fEC denote the matrix function, vectorization, FC layer, and Euclidean classifier, respectively. Typical candidates of matrix functions are matrix power and logarithm. As softmax is the most widely used classifier, fEC always denotes the softmax in this paper. However, our discussions can also apply to other classifiers used in GCP, such as SVM (Li et al.,2017;Wang et al.,2020a), as other classifiers receive the FC features as their inputs. FC + softmax is known as Euclidean Multinomial Logistics Regression (EMLR). When the matrix function is the matrix power, we call the process fEC ◦fFC ◦fvec ◦fMas the Pow-EMLR, while the counterpart of matrix logarithm is referred to as Log-EMLR. Especially, setting power as 1 /2 in GCP normally reaches the optimal performance (Li et al.,2017, Fig. 3). The Pow-EMLR and Log-EMLR can be formally expressed as Log-EMLR: softmax (F(fvec (mlog(S)) ; A, b)) ,(4) Pow-EMLR: softmax Ffvec Sθ;A, b,(5) where F(·;A, b)denotes the FC layer with the transformation matrix Aand biasing vector b. 4 MATRIX FUNCTIONS AND TANGENT CLASSIFIERS The matrix logarithm is the Riemannian logarithmic map at the identity matrix I, mapping SPD matrices into the tangent space TISn ++ ∼ =Sn. As tangent spaces are Euclidean spaces, it is natural to exploit FC and Euclidean classifiers on TISn ++ directly. We refer to the Euclidean classifiers over the tangent space at the identity matrix, TISn ++, as tangent classifiers. This section systematically studies all Riemannian logarithmic maps between Sn ++ and TISn ++, under seven families of metrics. 4.1 RIEMANNIAN LOGARITHMS UNDER SEVEN DEFORMED METRICS The matrix logarithm is generally characterized as the Riemannian logarithm LogIat Iunder the standard LEM and AIM. Inspired by this, we systematically investigate Riemannian logarithms on SPD manifolds. Let Pdenote an SPD matrix and ˜ Lrepresent the Cholesky factor of Pθ. Tab. 2 4
Published as a conference paper at ICLR 2025 Table 2: LogIunder seven families of metrics. θ0=θ1+θ2 2for (θ1, θ2)-EM, θ0=θfor (θ, α, β)-EM and 2θ-BWM, and (2θ, ϕ2θ(P))-BWM. Metric LogIPMetric LogIP (α, β)-LEM mlog(P)(θ, α, β)-EM 1 θ0(Pθ0−I) (θ, α, β)-AIM (θ1, θ2)-EM θ-LCM 1 θh⌊˜ L⌋+⌊˜ L⌋⊤+ 2 Dlog(D(˜ L))i2θ-BWM (2θ, P2θ)-BWM presents the Riemannian logarithms at Iunder all seven metrics, where ⌊·⌋ is the strictly lower triangular part of a square matrix, D(·)is a diagonal matrix, and Dlog(·)is the diagonal logarithm. We leave technical details in App. E. Remark 1.Let us elaborate further on the parameter of GBWM in Tab. 2. Given an SPD point P∈ Sn ++,P-BWM coincides with the standard AIM in the neighborhood of P(Han et al.,2021). This local property could be beneficial (Han et al.,2023). Similarly, (2θ, P 2θ)-BWM is locally (2θ, 1,0)-AIM, the deformed metric of the standard AIM. Please refer to App. Ffor technical details. Tab. 2implies that there are three types of LogI: Matrix-logarithm-based: mlog(P),(6) Matrix-power-based: 1 θ(Pθ−I),(7) LCM-based: 1 θh⌊˜ L⌋+⌊˜ L⌋⊤+ 2 Dlog(D(˜ L))i,(8) We denote the tangent MLR induced by Eq. (6), i.e.,Eq. (6) + vectorization + FC + softmax, as LogTMLR, while the counterparts of Eq. (7) and Eq. (8) is referred to as Pow-TMLR and Cho-TMLR, respectively. Obviously, Log-TMLR is the exact Log-EMLR (Eq. (4)) used in GCP. Table 3: Results of GCP on the ImageNet-1k and Cars datasets with Pow-TMLR or Pow-EMLR under the architecture of ResNet-18. Method ImageNet-1k Cars Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Pow-TMLR 71.62 89.73 51.14 74.29 Pow-EMLR 73 90.91 80.43 94.15 4.2 POW-TMLR VERSUS POW-EMLR Pow-EMLR applies Euclidean MLR directly on the non-Euclidean SPD manifold, while PowTMLR applies Euclidean MLR on the Euclidean space of TISn ++.In this sense, Pow-TMLR should be more theoretically advantageous than Pow-EMLR. Moreover, the difference between Pow-EMLR and Pow-TMLR seems to be minor. Pow-EMLR differs from Pow-TMLR only in a simple affine transformation fθ(X) = 1 θ(X−I). Note that the composition of affine transformations remains affine, and the FC layer is also an affine transformation. Therefore, Pow-EMLR might be viewed as the approximation of Pow-TMLR. Based on this discussion, we hypothesize that the tangent classifier serves as the underlying mechanism of matrix functions in GCP. If this hypothesis holds, Pow-EMLR should perform worse or at least similarly to Pow-TMLR. To validate this postulation, we conduct experiments on the ImageNet-1k (Deng et al.,2009) and Stanford Cars (Cars) (Krause et al.,2013) datasets. We use the architecture of ResNet-18 and ResNet-50 (He et al.,2016) on the ImageNet and Cars datasets, respectively. Following the classic iSQRT-COV (Li et al.,2018), we set power=1 /2and use Newton-Schulz iteration to calculate the matrix square root. Note that Pow-EMLR under Newton-Schulz iteration is exactly the original implementation of iSQRT-COV. As shown in Tab. 3, opposite to our hypothesis, Pow-TMLR is inferior to Pow-EMLR for classifying covariance matrices in GCP. Similar trends are also observed in additional experiments conducted on FGVC datasets, as will be presented in Sec. 6.These findings suggest that instead of tangent classifiers, there should exist other more fundamental mechanisms for underpinning matrix functions in GCP. 5
Published as a conference paper at ICLR 2025 5 MATRIX FUNCTIONS AND RIEMANNIAN CLASSIFIERS Recently, Riemannian classifiers on the SPD manifold, which can more faithfully respect the innate geometry, have exhibited more promising performance than tangent classifiers (Nguyen & Yang, 2023;Chen et al.,2024a;d). This section will demonstrate that matrix functions in GCP implicitly respect Riemannian classifiers, which offers a unified theoretical explanation of the working mechanism of matrix functions. We start with reviewing the Riemannian SPD classifiers and then present our theoretical analysis in detail. 5.1 SPD MULTINOMIAL LOGISTICS REGRESSION REVISITED Inspired by Lebanon & Lafferty (2004); Ganea et al. (2018), some recent works (Nguyen & Yang, 2023;Chen et al.,2024a;d) extended the Euclidean MLR into SPD manifolds. We first revisit the reformulation of the Euclidean MLR, and then move on to the SPD MLRs introduced by Chen et al. (2024d), especially the ones induced by (θ, α, β)-EM and (α, β)-LEM. The Euclidean MLR calculates the probability of each class by ∀k∈ {1, . . . , C}, p(y=k|x)∝exp (⟨ak, x⟩ − bk),(9) where x∈Rnis an input vector, Cis the number of classes, bk∈R, and ak∈Rn\{0}. Eq. (9) can be further rewritten as p(y=k|x)∝exp (⟨ak, x −pk⟩),(10) where pksatisfies ⟨ak, pk⟩=bk. As shown in the previous literature (Lebanon & Lafferty,2004; Ganea et al.,2018), Eq. (10) can be further reformulated by the margin distance to the hyperplane: p(y=k|x)∝exp(sign(⟨ak, x −pk⟩)∥ak∥d(x, Hak,pk)),(11) where pk∈Rnsatisfying ⟨ak, pk⟩=bk, and the hyperplane Hak,pkis defined as: Hak,pk={x∈Rn:⟨ak, x −pk⟩= 0}.(12) Chen et al. (2024d) generalized Eqs. (11) and (12) into general manifolds and proposed the SPD MLRs under five families of metrics. The SPD MLRs under (α, β)-LEM and (θ, α, β)-EM are (α, β)-LEM-based: p(y=k|S)∝exp h⟨log(S)−log(Pk), Ak⟩(α,β)i,(13) (θ, α, β)-EM-based: p(y=k|S)∝exp 1 θ⟨Sθ−Pθ k, Ak⟩(α,β),(14) where α > 0,β > −α /n, and Sis an input SPD feature. Here, Pk∈ Sn ++ and Ak∈TISn ++ ∼ =Sn are parameters for each class k. In Eqs. (13) and (14), the formula within exp(·)can be viewed as the counterpart of the Euclidean FC layer in SPD manifolds, extracting features to calculate softmax probabilities. 5.2 MATRIX FUNCTIONS AS SPD MULTINOMIAL LOGISTICS REGRESSION Under the standard LEM ((1,0)-LEM) and PEM ((θ, 1,0)-EM), Eqs. (13) and (14) become LEM-based: exp [⟨log(S)−log(Pk), Ak⟩],(15) PEM-based: exp 1 θ⟨Sθ−Pθ k, Ak⟩,(16) Eqs. (15) and (16) appear to be far away from the Log-/Pow-EMLR in GCP, as the SPD parameters {P1...C}requires Riemannian optimization. However, Chen et al. (2024a, Prop. 5.1) show that under the LEM-based Riemannian Stochastic Gradient Descent (RSGD) for each Pkand Euclidean SGD for each Ak, Eq. (15) is equivalent to a Euclidean MLR optimized by the Euclidean SGD in the co-domain of the matrix logarithm. Similar to LEM, we have the following proposition w.r.t. PEM. 6
Published as a conference paper at ICLR 2025 Table 4: Intrinsic explanations of some classifiers for GCP. For Cho-TMLR, ˜ L= Chol(Sθ). For Pow-TMLR, θ0=θ1+θ2 2for (θ1, θ2)-EM, θ0=θfor (θ, α, β)-EM, θ0= 2θfor 2θ-BWM and (2θ, ϕ2θ(S))-BWM. Here, fs(·)denotes the softmax, F(·)denotes the FC layer, and ˜ V= 1 θh⌊˜ L⌋+⌊˜ L⌋⊤+ 2 Dlog(D(˜ L))iwith Sθ=˜ L˜ L⊤as the Cholesky decomposition. Log-EMLR Pow-EMLR ScalePow-EMLR Pow-TMLR Cho-TMLR Expression fs(F(fvec (mlog(S)))) fsFfvec Sθ (θ > 0) fsFfvec 1 θSθ (θ > 0)fsFfvec 1 θ0(Sθ0−I) fsFfvec ˜ V Explanation SPD MLR SPD MLR SPD MLR Tangent Classifier Tangent Classifier Metrics LEM (θ, 1,0)-EM (θ, 1,0)-EM (θ, α, β)-EM, (θ1, θ2)-EM, 2θ-BWM, (2θ, ϕ2θ(S))-BWM θ-LCM Used in GCP ✓(Eq. (4)) ✓ (θ= 0.5in Eq. (5)) ✗ ✗ ✗ Reference (Chen et al.,2024a, Prop. 5.1) Thm. 2Thm. 2Tab. 2Tab. 2 Theorem 2. [↓]Under PEM with θ > 0, optimizing each SPD parameter Pkin Eq. (16)by PEMbased RSGD and Euclidean parameter Akby Euclidean SGD, the PEM-based SPD MLR is equivalent to a Euclidean MLR illustrated in Eq. (10)in the co-domain of ϕθ(·) : Sn ++ → Sn ++, defined as ϕθ(S) = 1 θSθ, θ > 0,∀S∈ Sn ++.(17) We define ScalePow-EMLR as softmax Ffvec 1 θSθ;A, b.Then, ScalePow-EMLR respects the SPD MLR under the standard PEM. The only difference between ScalePow-EMLR and PowEMLR (Eq. (5)) is the scalar product before vectorization, which is expected to have minor effects on DNNs. Obviously, we have Ffvec 1 θSθ;A, b=Ffvec Sθ;˜ A, b.(18) where ˜ A=1 θA. Therefore, from a forward perspective, ScalePow-EMLR is equivalent to the original Pow-EMLR. Besides, by scaled initialization and learning rate of A, ScalePow-EMLR could be completely equivalent to Pow-EMLR during network training. Note that this analysis cannot be transferred into Pow-TMLR. Please refer to App. Gfor more details. Therefore, the Pow-EMLR in GCP is implicitly an SPD MLR induced by (θ, 1,0)-EM. For the widely used matrix square root normalization, it respects the SPD MLR induced by (1 /2,1,0)-EM. We summarize all the findings in Tab. 4. Besides, Thm. 2can be easily extended into the case of θ < 0. In this case, our work can also offer theoretical insights for the inverse of covariance (θ=−1) proposed by Rahman et al. (2023). More details are presented in App. J. 5.3 THEORETICAL INSIGHTS ON THE MATRIX POWER AND LOGARITHM Previous studies on GCP (Li et al.,2017;Wang et al.,2020a;Song et al.,2021) have empirically demonstrated a clear advantage of the matrix power (particularly matrix square root) over matrix logarithm. This subsection offers novel theoretical insights to disentangle the different performance between the matrix logarithm and power in GCP. As shown by Tab. 4, both matrix logarithm and matrix power implicitly build SPD MLRs. However, the Riemannian metrics they respect are different. Matrix power respects (θ, 1,0)-EM, while matrix logarithm respects LEM. Both (θ, 1,0)-EM and LEM share O(n)-invariance (Chen et al.,2024d), a powerful property in characterizing covariance matrices. Besides, (θ, 1,0)-EM is a deformed metric of LEM, interpolating between the standard EM (θ= 1) and LEM (θ→0) (Thanwerdas & Pennec,2022). The standard EM might suffer from a swelling effect for characterizing SPD matrices (Arsigny et al.,2005), while LEM might over-stretch the eigenvalues of SPD matrices due to the computation of matrix logarithm (Song et al.,2021). Consequently, (θ, 1,0)-EM represents balanced alternatives between the standard LEM and EM. In addition, as shown by Chen et al. (2024d, Tab. 4), (θ, 1,0)-EM could perform better than LEM regarding SPD MLR. Therefore, the empirical advantages of matrix power over matrix logarithm in the GCP could be attributed to the characteristics of the underlying Riemannian metrics. 7
Published as a conference paper at ICLR 2025 Table 5: Results of iSQRT-COV on four datasets with different covariance matrix classifiers. The backbone network on ImageNet is ResNet-18, while the one on the other three FGVC datasets is ResNet-50. Power is set to be 1 /2for Pow-TMLR, ScalePow-EMLR and Pow-EMLR. Classifier ImageNet-1k Aircrafts Birds Cars Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Cho-TMLR N/A N/A 78.97 91.81 48.07 72.59 51.06 74.33 Pow-TMLR 71.62 89.73 69.58 88.68 52.97 77.80 51.14 74.29 ScalePow-EMLR 72.43 90.44 71.05 89.86 63.48 84.69 80.31 94.07 Pow-EMLR 73 90.91 72.07 89.83 63.29 84.66 80.43 94.15 Figure 3: The validation top-5 accuracy on the three FGVC datasets for iSQRT-COV with different classifiers using the ResNet-50 backbone. 6 EXPERIMENTS In this section, we validate the following hypothesis based on our previous theoretical analysis. (A1) As Pow-EMLR approximates the tangent classifier Pow-TMLR, the working mechanism of Pow-EMLR is attributed to the tangent classifier; (A2) As both Pow-EMLR and Log-EMLR in GCP are equivalent to Riemannian classifiers, the mechanism of matrix normalization should be attributed to Riemannian classifiers. We implement different classifiers for covariance matrix classification, including the original PowEMLR, the tangent classifiers Pow-TMLR and Cho-TMLR, and the intrinsic ScalePow-EMLR. We use the Caltech University Birds (Birds) (Welinder et al.,2010), FGVC Aircrafts (Aircrafts) (Maji et al.,2013), Stanford Cars (Cars) (Krause et al.,2013), and ImageNet-1k (Deng et al.,2009) datasets. As the matrix square root is the most effective matrix function in GCP, we set power = 1 /2. In all experiments, we train the network from scratch. More implementation details are in App. H. 6.1 MAIN RESULTS Notably, although ScalePow-EMLR is equivalent to Pow-EMLR under scaled settings, we implement them under the same network settings for a complete comparison. The results on four datasets are shown in Tab. 5. Our main empirical observations are as follows: (1). Pow-EMLR>Pow-TMLR. Pow-EMLR generally outperforms Pow-TMLR, especially on Cars and Birds datasets. Recalling in Tab. 4, the expression of Pow-EMLR differs from Pow-TMLR only in an affine transformation. However, across all four datasets, Pow-EMLR consistently surpasses Pow-TMLR. On the Birds and Cars datasets, Pow-EMLR outperforms Pow-TMLR by a large margin. For example, on the Birds dataset, the top-5 accuracy of Pow-EMLR and Pow-TMLR is 84.66% and 77.80%, respectively, whereas, on the Cars dataset, it is 94.15% and 74.29%. (2). Pow-EMLR≈ScalePow-EMLR. Pow-EMLR shows comparable performance to ScalePowEMLR. Recalling in Tab. 4, the only difference between Pow-EMLR and ScalePow-EMLR is a scalar product. Moreover, as discussed in Sec. 5.2 this minor difference can be further solved by scaled initialization of the FC layer. Although we use the same initialization for a fair comparison, Pow-EMLR and ScalePow-EMLR show similar performance. (3). Pow-EMLR≫Cho-TMLR. While Cho-TMLR demonstrates the best performance on the Aircrafts datasets, it exhibits the worst performance on the other two FGVC datasets. On the Cars and Birds datasets, Pow-EMLR surpasses Cho-TMLR by a large margin. The unstable performance of 8
Published as a conference paper at ICLR 2025 Cho-TMLR might be attributed to the diagonal logarithm, which might overly stretch the diagonal elements of the Cholesky factor. Based on the above empirical findings, we can reach the following conclusion. (A1) is refuted by (1). The inferior performance of Pow-TMLR against Pow-EMLR in (1) indicates that Pow-EMLR can not be simply viewed as equivalent to the tangent classifier Pow-TMLR. (A2) is validated by (2). (2) validates our theoretical postulation that the effectiveness of matrix power should be attributed to the Riemannian classifier it implicitly constructs. Other findings. In the first and last observations, tangent classifiers are less effective than the Riemannian classifier. Tangent classifiers can distort the innate geometry of the manifold, as the tangent space is only a local linear approximation of the manifold. In contrast, the Riemannian classifier can faithfully respect the geometry of the manifold. Besides, although Log-EMLR coincides with both tangent and Riemannian classifiers, the real underlying mechanism of matrix logarithm should also be attributed to the Riemannian classifier instead of the tangent classifier. Table 6: Ablations of Pow-EMLR, ScalePow-EMLR, and Pow-TMLR under different settings. (a) Results of different powers under the ResNet-50. Classifier Aircrafts Cars Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Pow-TMLR-0.25 65.41 86.71 41.47 66.66 ScalePow-EMLR-0.25 72.76 90.31 61.78 84.04 Pow-EMLR-0.25 71.47 90.04 62.88 84.14 Pow-TMLR-0.5 67.9 88.75 55.01 77.95 ScalePow-EMLR-0.5 74.29 91.12 62.42 84.82 Pow-EMLR-0.5 74.17 91.21 62.83 84.85 Pow-TMLR-0.7 65.92 87.49 50.68 74.12 ScalePow-EMLR-0.7 74.26 91.15 64.22 83.67 Pow-EMLR-0.7 74.17 90.49 61.41 82.39 (b) Results under the AlexNet. Dataset Result Pow-TMLR Pow-EMLR Aircrafts Top-1 Acc (%) 38.01 65.02 Top-5 Acc (%) 74.4 87.79 Cars Top-1 Acc (%) 28.57 59.13 Top-5 Acc (%) 59.51 82.04 6.2 TRAINING DYNAMICS AND ABLATIONS Training dynamics. Fig. 3presents the top-5 validation accuracy curves on three FGVC datasets. Pow-EMLR exhibits comparable performance to ScalePow-EMLR throughout the training. Moreover, Pow-EMLR consistently outperforms Pow-TMLR, particularly on the Cars and Birds datasets. This again suggests that the effectiveness of Pow-EMLR should be attributed to the Riemannian MLR rather than the tangent classifier. Furthermore, we note that the decreasing learning rate plays a crucial role in Cho-TMLR. On the Aircrafts dataset, before the 50th epoch, Cho-TMLR exhibits the worst performance among all classifiers. However, after the 50th epoch, when the learning rate reduces, Cho-TMLR surpasses all the other classifiers. Nonetheless, on the remaining two datasets, Cho-TMLR remains inferior throughout the training. This discrepancy may be attributed to the logarithm operation in Cho-TMLR. Recalling Eq. (8), there is a diagonal logarithm for the Cholesky factor. Similar to the matrix logarithm, Eq. (8) will also over-stretch the diagonal elements of the Cholesky factor, compromising the overall performance of Cho-TMLR. Ablations. To further validate our postulation, we compare Pow-EMLR, SaclePow-EMLR, and Pow-TMLR with different powers under the ResNet-50 architecture, i.e.,,θ= 0.25,0.5,0.7. We also compare Pow-EMLR against Pow-TMLR under the AlexNet architecture. The ablations are conducted on the Aircrafts and Car datasets. The results discussed below confirm again our findings that the mechanism of matrix functions in GCP should be attributed to Riemannian classifiers. Impact of matrix power. Following Song et al. (2021), we use accurate SVD to calculate the matrix power and Pad´ e approximant for backpropagation. The results are reported in Tab. 6a. Since we use SVD for the matrix power here, the results in Tab. 6a under θ= 0.5are slightly different from Tab. 5. Nevertheless, Pow-EMLR consistently shows similar performance to ScalePow-EMLR and outperforms Pow-TMLR under different powers. Impact of architectures. We also use the vanilla AlexNet (Krizhevsky et al.,2012) as an alternative backbone. Tab. 6b presents the comparison results under the AlexNet architecture. Consistent with our previous observation, Pow-EMLR still outperforms Pow-TMLR. 9
Published as a conference paper at ICLR 2025 APPENDIX CONTENTS A Future work 17 B Related work 17 C Notations and abbreviations 18 D Additional preliminaries 18 D.1 Pullback metrics .................................... 18 D.2 Riemannian operators on the SPD manifold ..................... 19 E Technical details on Riemannian Logarithm 19 F Power-deformed GBWM as local power-AIM 21 G Additional discussions on Pow-TMLR, Pow-EMLR, and ScalePow-EMLR 22 G.1 The equivalence of Pow-EMLR and ScalePow-EMLR ............... 22 G.2 The in-equivalence of Pow-EMLR and Pow-TMLR ................. 23 G.3 A Riemannian perspective of Pow-TMLR vs. Pow-EMLR ............. 23 H Additional experimental details 23 H.1 Datasets ........................................ 23 H.2 Implementation details ................................ 23 H.3 Experiments on the second-order Transformer .................... 24 H.4 Ablations on the bias reformulation ......................... 25 I Proof of Thm. 2 25 J Additional discussions on Thm. 2 27 16
Published as a conference paper at ICLR 2025 A FUTURE WORK While Chen et al. (2024d) also explored Riemannian MLRs induced by other metrics, these MLRs involve computationally expensive Riemannian computations, rendering them unsuitable for largescale datasets. As a future avenue, we aim to simplify the Riemannian computations in these alternative Riemannian classifiers and apply them to GCP for improved covariance matrix classification. B RELATED WORK Global covariance pooling. GCP aims to leverage the second-order statistics of deep features to enhance the learning competence of DNNs. DeepO2P(Ionescu et al.,2015), acknowledged as the first end-to-end global covariance pooling network, employs matrix logarithm for the classification of covariance matrices. This method also offers matrix backpropagation to differentiate the gradient w.r.t the decomposition-based matrix functions. Following this pioneering work, B-CNN (Lin et al., 2015) employs the outer product of global features and carries out element-wise power normalization. However, there exist three limitations of the above two methods. Firstly, the high dimensional covariance feature considerably escalates the parameters of the FC layer, thereby introducing the risk of overfitting. Secondly, the matrix logarithm could over-stretch the small eigenvalues, undermining the effectiveness of GCP. Thirdly, the matrix logarithm is based on matrix decomposition, which is computationally expensive. The subsequent research primarily focuses on four aspects: (a) adopting richer statistical representation (Wang et al.,2017;Zheng et al.,2019;Nguyen,2021); (b) reducing the dimensionality of the covariance feature (Gao et al.,2016;Kong & Fowlkes,2017;Cui et al.,2017;Acharya et al.,2018;Rahman et al.,2020;Wang et al.,2022a); (c) investigating effective and efficient matrix normalization (Li et al.,2018;Zheng et al.,2019;Lin & Maji,2017;Yu et al.,2020;Song et al.,2022c;b); (d) improving covariance conditioning for better generalization ability (Song et al.,2022d;a). In this work, we do not aim to achieve state-of-the-art performance over the existing GCP-based methods but rather to unravel the underlying theoretical mechanism of GCP matrix functions. Interpretations of global covariance pooling. Along with the progress of GCP, several works began to study its mechanism. Wang et al. (2020b) investigated the effect of GCP on deep Convolutional Neural Networks (CNNs) from an optimization perspective, including accelerated convergence, stronger robustness, and good generalization ability. Wang et al. (2023) further broadened these investigations, substantiating the merits of GCP in other networks, such as vision transformers (Touvron et al.,2021;Yuan et al.,2021;Liu et al.,2021) and differentiable Neural Architecture Search (NAS) (Liu et al.,2019). Song et al. (2021) empirically studied the advantage of approximate matrix square root over the accurate one. Wang et al. (2022a) considered the matrix power as decorrelating representations and developed a channel-adaptive dropout to produce lower-dimensional covariance matrices. Nevertheless, existing literature does not fully address the fundamental question of why Euclidean classifiers operate effectively in the non-Euclidean space generated by the matrix power. Our research fills in this theoretical gap, offering intrinsic explanations regarding the role of the matrix functions in GCP. Riemannian classifiers on SPD manifolds. Since the matrix logarithm is a diffeomorphism between the SPD manifold and its tangent space at the identity (Arsigny et al.,2005), the most widely used classifier on SPD manifolds is composed of the matrix logarithm and a Euclidean classifier (Wang et al.,2021;Chen et al.,2023b;Wang et al.,2022b;Nguyen,2022a;b;Wang et al.,2022c; Chen et al.,2024b;Wang et al.,2024b). However, this tangent classifier might distort the intrinsic geometry of SPD manifolds. Similar issues also arise in other manifolds (Huang et al.,2017;Wang et al.,2024a;Chen et al.,2025) Inspired by HNNs (Ganea et al.,2018), recent studies have developed intrinsic classifiers directly on SPD manifolds. Nguyen & Yang (2023) introduced three gyro structures on SPD manifolds induced by AIM, LEM, and LCM, respectively. Based on these gyro structures, the authors generalize the Euclidean Multinomial Logistics Regression (MLR). Concurrently, Chen et al. (2024a) proposed a formula for SPD MLR under Riemannian metrics pulled back from the Euclidean space. However, both works require specific Riemannian properties and focus on certain metrics on SPD manifolds. Chen et al. (2024d) presented a general framework for designing Riemannian MLRs on general geometries and showcased their framework under various metrics on SPD manifolds, covering the SPD MLRs introduced by Chen et al. (2024a); Nguyen & Yang (2023). Based on this framework, Chen et al. (2024c) showcased the SPD MLR under their 17
Published as a conference paper at ICLR 2025 Table 7: Summary of notations. Notation Explanation Sn ++ The SPD manifold SnThe Euclidean space of symmetric matrices LnThe Euclidean space of n×nlower triangular matrices TPSn ++ The tangent space at P∈ Sn ++ gP(·,·)or ⟨·,·⟩PThe Riemannian metric at P∈ Sn ++ ⟨·,·⟩ or ·:·The standard Frobenius inner product LogPThe Riemannian logarithm at P Ha,p The Euclidean hyperplane f∗,P The differential map of fat P∈ Sn ++ f∗gThe pullback metric by ffrom g ad(·)The adjoint operator of linear maps ST ST ={(α, β)∈R2|min(α, α +nβ)>0} ⟨·,·⟩(α,β)The O(n)-invariant Euclidean inner product g(α,β)-LE The Riemannian metric of (α, β)-LEM g(α,β)-AI The Riemannian metric of (α, β)-AIM g(θ,α,β)-AI The Riemannian metric of (θ, α, β)-AIM g(α,β)-E The Riemannian metric of (α, β)-EM g(θ,α,β)-E The Riemannian metric of (θ, α, β)-EM g(θ1,θ2)-E The Riemannian metric of (θ1, θ2)-EM gBW The Riemannian metric of BWM gM-BW The Riemannian metric of M-BWM g(2θ,M)-BW The Riemannian metric of (2θ, M)-BWM gLC The Riemannian metric of LCM gθ-LC The Riemannian metric of θ-LCM fFC or F(·;A, b)The FC layer Powθor (·)θThe matrix power fvec The vectorization fEC A Euclidean classifier mlog The matrix logarithm LP[·]The Lyapunov operator Chol The Cholesky decomposition LP,M [·]The generalized Lyapunov operator Dlog(·)The diagonal element-wise logarithm fMThe matrix function of matrix power or logarithm ⌊·⌋ The strictly lower triangular part of a square matrix D(·)A diagonal matrix with diagonal elements from a square matrix proposed product metrics. In this paper, based on the Riemannian classifiers developed by Chen et al. (2024d), we present an intrinsic explanation for matrix functions in GCP. C NOTATIONS AND ABBREVIATIONS For better clarity, we summarize all the notations in Tab. 7and all the abbreviations in Tab. 8. D ADDITIONAL PRELIMINARIES D.1 PULLBACK METRICS The power-deformed metrics on the SPD manifold are special cases of pullback metrics. Pullback metrics are common techniques in Riemannian geometry, connecting different Riemannian metrics. 18
Published as a conference paper at ICLR 2025 Table 8: Summary of Abbreviations. Abbreviation Explanation SPD Symmetric Positive Definite GCP Global covariance pooling GAP Global Average Pooling LEM Log-Euclidean Metric AIM Affine-Invariant Metric EM Euclidean Metric PEM Power Euclidean Metric MPEM Mixed Power Euclidean Metric BWM Bures-Wasserstein Metric GBWM Generalized Bures-Wasserstein Metric FGVC Fine-Grained Visual Categorization MLR Multinomial Logistics Regression EMLR Euclidean Multinomial Logistics Regression RMLR Riemannian Multinomial Logistics Regression SPD MLR RMLR on SPD manifolds Log-EMLR Eq. (4) Pow-EMLR Eq. (5) Pow-TMLR EMLR in the tangent space generated by Eq. (7) ScalePow-EMLR ScalePow-EMLR in Tab. 4 Cho-TMLR EMLR in the tangent space generated by Eq. (8) Definition 3 (Pullback Metrics).Suppose M,Nare smooth manifolds, gis a Riemannian metric on N, and f:M→Nis smooth. Then the pullback of gby fis defined point-wisely, (f∗g)p(V1, V2) = gf(p)(f∗,p(V1), f∗,p(V2)),(19) where p∈ M,f∗,p(·)is the differential map of fat p, and Vi∈TpM. If f∗gis positive definite, it is a Riemannian metric on M, which is called the pullback metric defined by f. D.2 RIEMANNIAN OPERATORS ON THE SPD MANIFOLD The O(n)-invariant Euclidean inner product on Sn(Thanwerdas & Pennec,2023) is defined as ⟨V, W⟩(α,β)=α⟨V, W⟩+βtr(V) tr(W),(20) where (α, β)∈ST with ST ={(α, β)∈R2|min(α, α +nβ)>0},V, W ∈ Sn, and ⟨·,·⟩ is the standard matrix inner product. We summarize deformed SPD metrics and associated Riemannian operators in Tab. 9with the following notations. Specifically, P, Q, M ∈ Sn ++ are SPD matrices, and V, W are tangent vectors in the tangent space at P,i.e.,TPSn ++. We denote gP(·,·)as the Riemannian metric at P, and LogP(·) as the Riemannian logarithm at P, respectively. Also, Chol and mlog represent the Cholesky decomposition and matrix logarithm, with their differential maps at Pdenoted as Chol∗,P and mlog∗,P , respectively. We denote ˜ V= Chol∗,P (V),˜ W= Chol∗,P (W),L= Chol(P), and K= Chol(Q).⌊·⌋ is the strictly lower part of a square matrix, D(·)is a diagonal matrix with diagonal elements of a square matrix, and Dlog(·)is a diagonal matrix consisting of the logarithm of the diagonal entries of a square matrix. We denote LP,M [V]as the generalized Lyapunov operator, i.e.,the solution to the matrix linear system MLP,M [V]P+PLP,M [V]M=V. When M=I,LP,I [V]is reduced to the Lyapunov operator, denoted as LP[V]. E TECHNICAL DETAILS ON RIEMANNIAN LOGARITHM We first review a well-known result for the pullback metric (Thanwerdas & Pennec,2022, Tab. 2). 19
Published as a conference paper at ICLR 2025 Table 9: Riemannian operators and deformed metrics of seven basic metrics on SPD manifolds. Note that for MPEM, Pand Qmust be commuting matrices when computing the Riemannian logarithm. Name Riemannian Metric gP(V, W )Riemannian Logarithm LogPQDeformation (θ= 0) (α, β)-LEM (Thanwerdas & Pennec,2023)⟨mlog∗,P (V),mlog∗,P (W)⟩(α,β)(mlog∗,P )−1[mlog(Q)−mlog(P)] 1 θ2Pow∗ θg(α,β)-LE (α, β)-AIM (Thanwerdas & Pennec,2023)⟨P−1V, W P−1⟩(α,β)P1/2mlog P−1/2QP−1/2P1/21 θ2Pow∗ θg(α,β)-AI (α, β)-EM (Thanwerdas & Pennec,2023)⟨V, W ⟩(α,β)Q−P1 θ2Pow∗ θg(α,β)-E (θ1, θ2)-EM (Thanwerdas & Pennec,2022) 1 θ1θ2⟨Powθ1∗,P (V),Powθ2∗,P (W)⟩(Powθ∗,P )−1(Qθ−Pθ), with θ= (θ1+θ2)/2N/A LCM (Lin,2019)Pi>j ˜ Vij ˜ Wij +Pn j=1 ˜ Vjj ˜ WjjL−2 jj (Chol−1)∗,L ⌊K⌋−⌊L⌋+D(L) Dlog(D(L)−1D(K))1 θ2Pow∗ θgLC BWM (Bhatia et al.,2019)1 2⟨LP[V], W⟩(PQ)1/2+ (QP )1/2−2P1 4θ2Pow∗ 2θgBW GBWM (Han et al.,2023)1 2⟨LP,M [V], W ⟩MM−1PM−1Q1/2+QM−1PM−11/2M−2P1 4θ2Pow∗ 2θgM-BW Lemma 4. Given a Riemannian metric gon the SPD manifold Sn ++ and a diffeomorphism f: Sn ++ → Sn ++, the Riemannian logarithm ˜ LogPunder the pullback metric ˜g=f∗gis ˜ LogPQ= (f∗,P )−1Logf(P)f(Q),(21) where f∗,P is the differential map at P, and Log is the Riemannian logarithm under g. Next, we show a lemma about the scaling of a Riemannian metric. Lemma 5. Supposing Sn ++ is endowed with a Riemannian metric gand a > 0is a positive real scalar, the scaling metric ag shares the same Riemannian logarithm map with g. Proof. Since the Christoffel symbols of ag are identical to those of g, the geodesic functions under both ag and gremain unchanged. This implies that the Riemannian exponential maps are the same for ag and g. As the inverse of the Riemannian exponential maps, the Riemannian logarithm maps under ag and gare also identical. By the above lemmas, we can readily prove Tab. 2. Proof. By Lem. 5, for the power-deformed metric of a metric gin Sn ++, the Riemannian logarithm at Iis the same as the counterpart under Pow∗ θg. Therefore, in the following, without loss of generality, we compute LogIunder Pow∗ θg. We further denote the Riemannian logarithm under g as ¯ Log. In the following, we denote Pas an SPD matrix, 0as the n×nzero matrix, and Vas a tangent vector in TISn ++. Besides, we note that Powθ∗,I (V) = θV. (22) We first deal with (α, β)-LEM and θ-LCM, as both of them are pullback metrics from the Euclidean space. Then, we proceed to deal with other metrics (α, β)-LEM: As shown by Thanwerdas & Pennec (2023), the Riemannian logarithm at Iis LogI(P) = mlog−1 ∗,I (mlog(P)−mlog(I)) = mlog(P).(23) θ-LCM: We define a map as f=ψ◦Chol ◦Powθ,(24) where ψ(L) = ⌊L⌋+ Dlog(D(L)) for the lower triangular matrix L.Chen et al. (2024e) shows that LCM is the pullback metric by ψ◦Chol from the Euclidean space Lnof lower triangular matrices. 20
Published as a conference paper at ICLR 2025 Therefore, Pow∗ θgLC is the pullback metric from Lnby f. Besides, we have the following: f(P) = ⌊˜ L⌋+ Dlog(D(˜ L)),(25) f(I) = 0,(26) f∗,I(V) = θ⌊V⌋+1 2D(L),(27) where ˜ L= Chol(Pθ). We have LogI(P)=(f∗,P )−1(f(P)−f(I)) =1 θh⌊˜ L⌋+⌊˜ L⌋⊤+ 2 Dlog(D(˜ L))i.(28) For (θ, α, β)-EM, (θ1, θ2)-EM, (θ, α, β)-AIM, 2θ-BWM, and (2θ, P 2θ)-BWM, we denote LogIas their logarithm at I, while ¯ LogIas the logarithm under the metric before deformation. The results can be directly obtained by Eq. (22), Lem. 4, Lem. 5, and Tab. 9. (θ, α, β)-EM: LogI(P) = 1 θ¯ LogI(Pθ) =1 θPθ−I. (29) (θ1, θ2)-EM: The LogIcan be directly obtained by Tab. 9and Eq. (22). (θ, α, β)-AIM: LogI(P) = 1 θ¯ LogI(Pθ) =1 θmlog(Pθ) = mlog(P). (30) 2θ-BWM: LogI(P) = 1 2θ¯ LogI(P2θ) =1 θ(Pθ−I). (31) (2θ, P2θ)-BWM: Under M-BWM, we have LogI(M) = 2(M1 2−I).(32) Therefore, for (2θ, P2θ)-BWM, we have LogI(P) = 1 2θ¯ LogI(P2θ) =1 θ(Pθ−I). (33) F POWER-DEFORMED GBWM AS LOCAL POWER-AIM Let us first formalize this property. Proposition 6. For any P∈ Sn ++ and V, W ∈TPSn ++, we have the following: g(2θ,P 2θ)-BW P(V, W) = 1 4g(2θ,1,0)-AI P(V, W).(34) 21
Published as a conference paper at ICLR 2025 Proof. As shown by Bhatia (2009), the Riemannian metric of the standard AIM ((1,1,0)-AIM) is gAI P(V, W) = vec(V)⊤(P⊗P)−1vec(W),(35) where vec(V)is the column vectorization of V,⊗is the Kronecker product. For the (2θ, P 2θ)-BWM, we have the following: g(2θ,P 2θ)-BW P(V, W) = 1 4θ2gϕ2θ(P)-BW ˜ P(˜ V , ˜ W) =1 4·1 4θ2vec( ˜ V)⊤(˜ P⊗˜ P)−1vec( ˜ W) =1 4·1 4θ2gAI ˜ P(˜ V , ˜ W) =1 4g(2θ,1,0)-AI P(V, W), (36) where ˜ V= Pow2θ∗,P (V),˜ W= Pow2θ∗,P (W),˜ P=P2θ, and Eq. (36) can be obtain by Han et al. (2023, Eq. 3) G ADDITIONAL DISCUSSIONS ON POW-TMLR, POW-EMLR, AND SCALEPOW-EMLR G.1 THE EQUIVALENCE OF POW-EMLR AND SCALEPOW-EMLR It can be proven that Pow-EMLR is equivalent to ScalePow-EMLR under scaled initial weight and learning rate in the FC layer. We denote the network as x0∈Rd0g(·;Θ) −→ x∈RdfFC −→ y∈Rc→L∈R,(37) where x0,g(·; Θ),fFC, and Lare the input feature, feature extraction with parameter Θ, FC layer, and loss, respectively. The FC layers in Pow-EMLR and ScalePow-EMLR are denoted as y=Ax+b and ¯y=1 θ¯ A¯x+¯ b. We set the initial values and learning rates of Aand ¯ Asatisfying A0=1 θ¯ A0and ¯γ=θ2γ, and maintain all the other settings the same. Then, we have the following for the gradient at A=A0(or ¯ A=¯ A0): ∂L ∂¯ A=1 θ ∂L ∂¯y¯x⊤=1 θ ∂L ∂y x⊤=1 θ ∂L ∂A, ∂L ∂¯x=1 θ¯ A⊤∂L ∂¯y=A⊤∂L ∂y =∂L ∂x . (38) Under SGD, the updated values of ¯ Asatisfying the following: 1 θ¯ A1=1 θ(¯ A0+ ¯γ∂L ∂¯ A) =1 θ(¯ A0+ ¯γ1 θ ∂L ∂A) =1 θ¯ A0+ ¯γ1 θ2 ∂L ∂A =A0+γ∂L ∂A =A1. (39) Therefore, the updated values of Aand ¯ Astill satisfy A1=1 θ¯ A1. In addition, the gradients of Pow-EMLR w.r.t. xand bare identical to ScalePow-EMLR w.r.t. ¯xand ¯ b. Therefore, Pow-EMLR is equivalent to ScalePow-EMLR under scaled settings. 22
Published as a conference paper at ICLR 2025 G.2 THE IN-EQUIVALENCE OF POW-EMLR AND POW-TMLR We denote X=Sθ. Then for Pow-TMLR, we have the following y=Ffvec 1 θ(X+I);A, b =Ffvec (X+I) ; ˜ A, b =Ffvec (X) ; ˜ A,˜ b, (40) where ˜ A=1 θAand ˜ b=1 θAfvec(I). As Aappears in ˜ b, the gradient of Ais composed of two parts, one w.r.t. yand the other one w.r.t. ˜ b. In contrast, in the standard FC layer y=F(X;A, b), the gradient of Ais independent of b. Therefore, Pow-TMLR cannot be simply viewed as equivalent to Pow-EMLR with transformed initialization. G.3 A RIEMANNIAN PERSPECTIVE OF POW-TMLR VS. POW-EMLR Although the numerical expressions of Pow-TMLR and Pow-EMLR differ by a constant transformation, they differ fundamentally in theory: Pow-TMLR is a tangent classifier, whereas Pow-EMLR is a Riemannian classifier. 1. Tangent Classifier: The tangent classifier treats the entire manifold as a single tangent space at the identity matrix. When mapping data into this tangent space, critical structural information, such as distances and angles, cannot be preserved. This distortion undermines classification performance. In contrast, Riemannian MLR is constructed based on Riemannian geometry, fully respecting the manifold’s geometric structure. 2. Tangent as a Special Case of Riemannian Classifier. The tangent classifier can be seen as a reduced case of the Riemannian classifier. For example, let us take Eq. (16) as an example. When all SPD parameters Pkare set to the fixed identity matrix, Eq. (16) exactly corresponds to Pow-TMLR. In summary, the Riemannian classifier enjoys significant theoretical advantages over the tangent classifier while incorporating the tangent classifier as a special case. H ADDITIONAL EXPERIMENTAL DETAILS H.1 DATASETS The Caltech University Birds (Birds) (Welinder et al.,2010) dataset is composed of 11, 788 images distributed over 200 different bird species. The FGVC Aircrafts (Aircrafts) (Maji et al.,2013) dataset comprises 10, 000 images of 100 classes of airplanes, while the Stanford Cars (Cars) (Krause et al., 2013) dataset consists of 16, 185 images representing 196 classes of cars. In addition to these widely used FGVC datasets, we also evaluate our proposed theory on the large-scale ImageNet-1k (Deng et al.,2009) dataset, which contains 1.28M training images, 50K validation images and 100K testing images distributed across 1K classes. H.2 IMPLEMENTATION DETAILS We follow the official Pytorch code of iSQRT-COV1(Li et al.,2018) to reimplement GCP. Following Wang et al. (2020a); Song et al. (2022a), we use ResNet-18 as our backbone network on the ImageNet dataset, and ResNet-50 on the other three FGVC datasets. On Both the ImageNet-1k and FGVC datasets, the ResNet-18 and ResNet-50 are trained from scratch with the GCP layer. 1https://github.com/jiangtaoxie/fast-MPN-COV 23
Published as a conference paper at ICLR 2025 As the matrix square root is the most effective matrix function in GCP, we set power = 1 /2for matrix power normalization. Following Song et al. (2022a), we reduce the channels of the final convolutional features from 2048 to 256 for compact representation of covariance matrices, producing 256×256 spatial covariance matrices. We train the network from scratch with an SGD optimizer on all datasets. For a fair comparison, the learning settings are identical for Pow-EMLR and ScalePowEMLR, while we fine-tune Pow-TMLR w.r.t. learning rate, classifier factor, and batch size on three FGVC datasets. The learning rate of the FC layer is set to be ktimes larger than the convolutional layers, where kis called the classifier factor. We use a step-wise learning scheduler, dividing the learning rate by 5 at n-th epoch. Tab. 10 summarizes the hyperparameters in our main experiments in Tab. 5. For Cho-TMLR, the learning rate is set as 1e−2,5e−3, and 3e−3for the convolutional layers on the Aircrafts, Birds, and Cars dataset. The batch size on the Cars dataset is 4. On the Aircrafts dataset, the training lasts 60 epochs with the learning rate divided by 5 at epoch 50. On the Cars and Birds datasets, the training lasts 120 epochs with a learning rate reduction by a divisor of 10 at epoch 100. The experiments on ImageNet use a workstation with 32-core AMD EPYC 7302 CPU and an NVIDIA RTX A6000, while other experiments use a workstation with 16-core AMD EPYC 7302 CPU and an NVIDIA GeForce RTX 2080 Ti GPU. Due to the heavy computational burden of Cholesky decomposition, we do not implement Cho-TMLR on the ImageNet. Table 10: Summary of hyperparameters. Backbone ResNet18 ResNet50 Dataset ImageNet Aircrafts Birds Cars Classifiers Pow-EMLR ScalePow-EMLR Pow-TMLR Pow-EMLR ScalePow-EMLR Pow-TMLR Pow-EMLR ScalePow-EMLR Pow-TMLR Pow-EMLR ScalePow-EMLR Pow-TMLR Learning Rate 1e−1.15e−35e−25e−35e−25e−3 Weight Decay 1e−41e−41e−41e−41e−41e−4 Classifier Factor 1 5 10 1 10 1 LR Scheduler [30,45,60] [21,50] [50,100] [50,100] [50,100] [50,100] Batch Size 256 8 10 6 10 10 Epoch 60 50 100 100 100 100 H.3 EXPERIMENTS ON THE SECOND-ORDER TRANSFORMER Table 11: Comparison of Pow-EMLR, ScalePow-EMLR and Pow-TMLR under the SoT-7 backbone on the ImageNet-1k dataset. Classifier Top-1 Acc (%) Top-5 Acc (%) Pow-TMLR 75.79 92.91 ScalePow-EMLR 76.14 93.18 Pow-EMLR 76.11 93.05 To further validate our findings, we follow Song et al. (2022b) to conduct experiments using the Second-order Transformer (SoT) (Xie et al.,2021) on the ImageNet-1k dataset. Specifically, we use the 7-layer SoT (SoT-7) architecture as the backbone network and train the model up to 250 epochs with a batch size of 512, keeping the other settings the same as Song et al. (2022b). As shown in Tab. 11, Pow-EMLR still achieves similar performance to ScalePow-EMLR, but outperforms Pow-TMLR. These results further support our claim that tangent classifiers cannot adequately explain the matrix functions used in GCP, while the underlying mechanism can be better explained by our Riemannian perspective. 24
Published as a conference paper at ICLR 2025 H.4 ABLATIONS ON THE BIAS REFORMULATION Recalling the vanilla Pow-EMLR Ak, Sθ−bk, it can be rewritten as Ak, Sθ−Pkaccording to Eq. (10), where ⟨Pk, Ak⟩=bk. This reformulation is the key step to extending Euclidean MLR. Although this reformulation has shown success in different Riemannian MLRs (Ganea et al.,2018; Shimizu et al.,2020;Chen et al.,2024a;d;Nguyen & Yang,2023), we conduct ablations on this reformulation. For simplicity, we set Pkidentical for all k; otherwise, it will bring a [k, n, n]intermediate tensor for each [n, n]covariance. We compare the following to MLR: Pow-EMLR: Ak, Sθ−bk,(41) Pow-EMLR’: Ak, Sθ−P,(42) where Ak, P ∈ Sn. Tab. 12 shows the results on all three fine-grained datasets. We observe that Pow-EMLR performs similarly to Pow-EMLR’. Table 12: Comparison of Pow-TMLR, Pow-EMLR, and Pow-EMLR’ on all three fine-grained datasets. Classifier Air Birds Car Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Top-1 Acc (%) Top-5 Acc (%) Pow-TMLR 69.58 88.68 52.97 77.8 51.14 74.29 Pow-EMLR’ 73.03 90.4 63.96 85.02 80.06 94.02 Pow-EMLR 72.07 89.83 63.29 84.66 80.43 94.15 I PROOF OF THM.2 This proposition is mainly inspired by Chen et al. (2024a, Thm. 5). However, all the results by Chen et al. (2024a) require the metric to be a pullback metric from a standard Euclidean space, while the metric in our Thm. 2is a pullback metric from the SPD manifold. Nevertheless, we still can reach similar theoretical results. We first recap RSGD and then begin to present our proof. RSGD (Bonnabel,2013) is formulated as ¯ W= ExpW(−γΠW(∇Wf)) (43) where ExpWis the Riemannian exponential map at W, and ΠWmaps the Euclidean gradient ∇Wf to the Riemannian gradient, and γdenotes learning rate. We denote (1,0)-EM as EM, and the metric tensor of it as gE. Instead of providing an ad hoc proof exclusively for PEM, we present the following two more general lemmas. Lemma 7. Given a diffeomorphism ϕ:Sn ++ → Sn ++,ϕinduces a pullback metrics on Sn ++ from {Sn ++, gE}, denoted as gϕ-E. The gϕ-E-induced SPD MLR is p(y=k|S)∝exp [⟨ϕ(S)−ϕ(Pk), ϕ∗,I(Ak)⟩],(44) where S∈ Sn ++ is an input feature, Pk∈ Sn ++ and Ak∈ Snare parameters for each class k. Proof. According to Chen et al. (2024d, Thm. 3.3), the Riemannian MLR based on gϕ-E is given as p(y=k|S)∝exp hgϕ-E Pk(LogPkS, PTI→PkAk)i = exp [⟨ϕ(S)−ϕ(Pk), ϕ∗,I (Ak)⟩],(45) where Eq. (45) can be obtained by the properties of deformed metrics (Thanwerdas & Pennec,2022, Tab. 2) and EM (Thanwerdas & Pennec,2023, Tab. 3). Following the notations in Lem. 7, we have the following lemma. Lemma 8. Supposing ϕ∗,I is the identity map and each SPD parameter Pk(Euclidean parameter Ak) in Eq. (44)is optimized by gϕ-E-based RSGD (Euclidean SGD), the gϕ-E-based SPD MLR is equivalent to a Euclidean MLR illustrated in Eq. (10)in the co-domain of ϕ. 25