scieee AI-readable full text Open interactive document viewer

Gyrogroup Batch Normalization

Chen, Ziheng; Song, Yue; Wu, Xiao-Jun; Sebe, Niculae

Full text

Published as a conference paper at ICLR 2025 GYROGROUP BATCH NORMALIZATION Ziheng Chen1∗ , Yue Song1, Xiao-Jun Wu2& Nicu Sebe1 1University of Trento, 2Jiangnan University ABSTRACT Several Riemannian manifolds in machine learning, such as Symmetric Positive Definite (SPD), Grassmann, spherical, and hyperbolic manifolds, have been proven to admit gyro structures, thus enabling a principled and effective extension of Euclidean Deep Neural Networks (DNNs) to manifolds. Inspired by this, this study introduces a general Riemannian Batch Normalization (RBN) framework on gyrogroups, termed GyroBN. We identify the least requirements to guarantee GyroBN with theoretical control over sample statistics, referred to as pseudoreduction and gyroisometric gyrations, which are satisfied by all the existing gyrogroups in machine learning. Besides, our GyroBN incorporates several existing normalization methods, including the one on general Lie groups and different types of RBN on the non-group SPD geometry. Lastly, we instantiate our GyroBN on the Grassmannian and hyperbolic spaces. Experiments on the Grassmannian and hyperbolic networks demonstrate the effectiveness of our GyroBN. The code is available at https://github.com/GitZH-Chen/GyroBN.git. 1 INTRODUCTION Deep Neural Networks (DNNs) on Riemannian manifolds have gained increasing interest in various machine learning applications, such as computer vision (Huang et al.,2017;Huang & Van Gool, 2017;Huang et al.,2018;Skopek et al.,2019;Wang et al.,2022b;a;Chen et al.,2023c;Gao et al., 2023;Wang et al.,2024b;Chen et al.,2025), natural language processing (Ganea et al.,2018;Shimizu et al.,2020), drone classification (Brooks et al.,2019;Chen et al.,2024a), human neuroimaging (Pan et al.,2022;Kobler et al.,2022a;Ju et al.,2024;Wang et al.,2024a), medical imaging (Huang et al., 2019;Chakraborty et al.,2020), node and graph classification (Chami et al.,2019;Dai et al.,2021; Zhao et al.,2023;Chen et al.,2023b;Nguyen et al.,2024;Chen et al.,2024c). As core techniques in DNNs, normalization techniques (Ioffe & Szegedy,2015;Ba et al.,2016;Ulyanov et al.,2016;Wu & He,2018;Chen et al.,2023a) have also been extended into different geometries. However, most existing Riemannian normalization methods are designed for a selected few geometries or fail to normalize the sample statistics. Brooks et al. (2019); Kobler et al. (2022b;a) introduced Riemannian Batch Normalization (RBN) on SPD manifolds under the specific Affine-Invariant Metric (AIM). Chakraborty (2020) generalized this idea and proposed a Riemannian normalization framework over homogeneous spaces. However, this approach cannot generally normalize the sample statistics. Similar formulation and issue can also be found in Lou et al. (2020, Alg. 2) and Bdeir et al. (2024, Sec. 4.2). Besides, Chakraborty (2020) also developed a Riemannian normalization for matrix Lie groups. Although this approach can control firstand second-order statistics, it is limited to a specific type of distance (Chakraborty,2020, Sec. 3.2). Chen et al. (2024b) took one step further and developed RBN over the general Lie group, referred to as LieBN. Although LieBN can effectively normalize sample statistics, many geometries do not admit a group structure. In summary, the existing RBN methods fail to normalize manifold-valued samples in a principled manner. Recently, building Riemannian networks based on gyro structures has shown notable success across various geometries, including Symmetric Positive Definite (SPD) (Nguyen,2022a;b), Grassmannian (Nguyen,2022b), hyperbolic (Ganea et al.,2018), and spherical (Skopek et al.,2019) manifolds. Gyro structures, natural extensions of vector structures, offer powerful mathematical tools for building Riemannian neural networks. Moreover, gyrogroups naturally encompass Lie groups and extend to non-group geometries. For instance, AIM on the SPD manifold, as well as Grassmannian, hyperbolic, and spherical manifolds, do not form Lie groups but instead gyrogroups. ∗Correspondence to [email protected] 1 Published as a conference paper at ICLR 2025 Figure 1: Illustration of GyroBN on the Grassmannian and hyperbolic spaces. As Gr(1,3) is homeomorphic to the real projective space RP2 , we illustrate Gr(1,3) as the unit hemisphere with antipodal points identified. We set the bias as I1,3= (1,0,0)⊤ and the scaling scalar as 0.2 for better illustration. For the hyperbolic space, we visualize the GyroBN on the Poincaré ball model P3 −1 , which is the interior of the unit sphere in R3 . We set bias and shift as zero vector and 0.7, respectively. Table 1: Comparison of previous RBN methods with our GyroBN, where M and V denote the sample mean and variance. Compared with the existing RBN methods, our GyroBN can normalize sample statistics in a principled manner. Besides, several previous RBN methods with theoretical control over sample statistics are special cases of our GyroBN. Method Controllable Statistics Applied Geometries Incorporated by GyroBN SPDBN (Brooks et al.,2019)M SPD manifolds under AIM ✓ SPDBN (Kobler et al.,2022b)M+V SPD manifolds under AIM ✓ SPDDSMBN (Kobler et al.,2022a)M+V SPD manifolds under AIM ✓ ManifoldNorm (Chakraborty,2020, Algs. 1-2) N/A Riemannian homogeneous space ✗ ManifoldNorm (Chakraborty,2020, Algs. 3-4) M+V Matrix Lie groups under the distance d(X, Y ) = mlog X−1Y✓ RBN (Lou et al.,2020, Alg. 2) N/A Geodesically complete manifolds ✗ LieBN (Chen et al.,2024b)M+V General Lie groups ✓ GyroBN M+V Pseudo-reductive gyrogroups with gyro isometric gyrations N/A Based on the above analysis, this paper introduces GyroBN, an RBN framework for general gyrogroups. We use the gyro addition, subtraction, and scalar product to extend the centering (vector subtraction), biasing (vector addition), and scaling (vector scalar product) in the Euclidean BN into manifolds in a principled manner. For broader applicability and in-depth theoretical analysis, we adapt the existing gyrogroup into a more relaxed structure, termed pseudo-reductive gyrogroup. Our theoretical analysis shows that pseudo-reductive gyrogroups with gyroisometric gyrations can enable GyroBN with theoretical control over sample statistics. More importantly, these requirements are satisfied by all the existing gyrogroups in machine learning. Therefore, compared with the existing RBN methods, our GyroBN can normalize sample statistics in a principled manner. Besides, several existing RBN methods are incorporated by our GyroBN as special cases, including the LieBN on Lie groups, such as three SPD Lie groups and rotation matrices, and different types of RBN on the non-group SPD geometry. We provide a detailed comparison in Tab. 1. Empirically, we instantiate our GyroBN on the Grassmannian and hyperbolic spaces, as illustrated in Fig. 1. To the best of our knowledge, our Grassmannian GyroBN is the first implementation of Grassmannian RBN. Experiments on the Grassmannian and hyperbolic networks validate the effectiveness of our GyroBN. Our main theoretical contributions are summarized as follows: (1) We propose the pseudo-reductive gyrogroup, a relaxed structure of the gyrogroup, and present relevant theoretical analyses; (2) We identify the requirements that guarantee our GyroBN with theoretical control over sample statistics, i.e.,pseudo-reduction and gyroisometric gyrations; (3) We propose a GyroBN framework for RBN over general gyrogroups, which can be manifested in various geometries in a plug-and-player manner. (4) We implement our GyroBN on the Grassmannian and hyperbolic spaces. Extensive experiments on popular Grassmannian and hyperbolic networks validate the effectiveness of our framework. 2 Published as a conference paper at ICLR 2025 Main theoretical results: Def. 3.1 relaxes the existing gyrogroup into the pseudo-reductive gyrogroup. Prop. 3.2 reveals that the non-reductive Grassmannian gyrogroup is, in fact, pseudo-reductive. Thms. 3.3 and 3.5 highlight that the invariance of gyronorm under gyrations is crucial for enabling several operators to act as gyroisometries. Prop. 3.6 confirms that all gyrogroups listed in Tab. 2 are pseudo-reductive and their gyrations are gyroisometries. These two properties are essential for enabling GyroBN in Alg. 1to normalize sample statistics, which are formalized in Thm. 4.1. Sec. 5 discusses how several existing RBN methods are special cases of our GyroBN. Lastly, Sec. 6manifest our GyroBN on the Grassmannian and hyperbolic spaces, where Prop. 6.1 discuss the efficient implementation on the Grassmannian. Due to page limits, all the proofs are presented in App. G. 2 PRELIMINARIES This section recaps gyrogroups (Ungar,2009) and several concrete gyrogroups in machine learning. Definition 2.1 (Gyrogroups (Ungar,2009)).Given a nonempty set G with a binary operation ⊕:G×G→G , {G, ⊕} forms a gyrogroup if its binary operation satisfies the following axioms for any a, b, c ∈G: (G1) There is at least one element e∈G called a left identity (or neutral element) such that e⊕a=a . (G2) There is an element ⊖a∈Gcalled a left inverse of asuch that ⊖a⊕a=e. (G3) There is an automorphism gyr[a, b] : G→Gfor each a, b ∈Gsuch that a⊕(b⊕c) = (a⊕b)⊕gyr[a, b]c(Left Gyroassociative Law). (1) The automorphism gyr[a, b] is called the gyroautomorphism, or the gyration of G generated by a, b . (G4) Left reduction law: gyr[a, b] = gyr[a⊕b, b]. Definition 2.2 (Gyrocommutative Gyrogroups (Ungar,2009)).A gyrogroup {G, ⊕} is gyrocommutative if it satisfies a⊕b= gyr[a, b](b⊕a)(Gyrocommutative Law). (2) Definition 2.3 (Nonreductive Gyrogroups (Nguyen,2022a)).A groupoid {G, ⊕} is a nonreductive gyrogroup if it satisfies axioms (G1), (G2), and (G3). Intuitively, gyrogroups are natural generalizations of groups. Unlike groups, gyrogroups are nonassociative but have gyroassociativity characterized by gyrations. Since all gyrations in any (Lie) group are the identity map, every (Lie) group is automatically a gyrogroup. As shown by Nguyen & Yang (2023), given P , Q and R in a manifold M and t∈R , the gyro structures can be defined as: Gyro addition: P⊕Q= ExpP(PTE→P(LogE(Q))) ,(3) Gyro scalar product: t⊙P= ExpE(tLogE(P)) ,(4) Gyro inverse: ⊖P=−1⊙P= ExpE(−LogE(P)) ,(5) Gyration: gyr[P, Q]R= (⊖(P⊕Q)) ⊕(P⊕(Q⊕R)),(6) Gyro inner product: ⟨P, Q⟩gr =⟨LogE(P),LogE(Q)⟩E,(7) Gyro norm: ∥P∥gr =⟨P, P⟩gr ,(8) Gyrodistance: dgry(P, Q) = ∥⊖P⊕Q∥gr ,(9) where E is the gyro identity element, and LogE and ⟨·,·⟩E is the Riemannian logarithm and metric at E. A bijection ω:G→Gis called gyroisometry, if it preserves the gyrodistance dgry(ω(P), ω(Q)) = dgry(P, Q).(10) Note that the gyro scalar product ⊙ is required for a gyrogroup to form a gyrovector space (Nguyen, 2022b). Although this paper only involves gyrogroups, we also recap gyrovector spaces in App. C. Several geometries in machine learning admit a gyro structure defined in Eq. (3)-Eq. (9) and form a (nonreductive) gyrogroup, such as Affine-Invariant Metric (AIM) (Pennec et al.,2006), Log-Euclidean Metric (LEM) (Arsigny et al.,2005), and Log-Cholesky Metric (LCM) (Lin,2019) on the SPD manifold Sn ++ (Nguyen,2022a), Orthonormal Basis (ONB) perspective Gr(p, n) (Bendokat et al., 2024) and projector perspective f Gr(p, n) (Bendokat et al.,2024) for the Grassmannian (Nguyen, 3 Published as a conference paper at ICLR 2025 2022a;Nguyen & Yang,2023), Poincaré ball Pn K for the hyperboloid (Ungar,2009;Ganea et al., 2018), and projected hypersphere Dn K for the hypersphere (Skopek et al.,2019). Besides, the gyrogroups proposed by Nguyen (2022b) on the SPD manifold under the LEM and LCM coincide with the Lie groups proposed by Arsigny et al. (2005); Lin (2019). We denote MK as Pn K(K < 0) , Dn K(K > 0) and Rn(K= 0) , respectively. MK is known as the Constant Curvature Space (CCS) (Do Carmo & Flaherty Francis,1992). We summarize all the necessary gyro properties in Tab. 2. Table 2: Gyrogroup properties on several geometries. Related notations are defined in App. C.3.2. Geometry Symbol P⊕Qor x⊕y E ⊖Por ⊖xLie group Gyrogroup References AIM Sn ++ ⊕AI P1 2QP 1 2InP−1✗✓(Nguyen,2022b) LEM Sn ++ ⊕LE mexp(mlog(P) + mlog(Q)) InP−1✓ ✓ (Arsigny et al.,2005) (Nguyen,2022b) LCM Sn ++ ⊕LC ψ−1 LC(ψLC(P) + ψLC(Q)) InψLC(−ψLC(P)) ✓ ✓ (Lin,2019) (Nguyen,2022b) (Chen et al.,2024e) f Gr(p, n)e ⊕Gr mexp(Ω)Qmexp(−Ω) e Ip,n mexp(−Ω)e Ip,n mexp(Ω) ✗Non-reductive (Nguyen,2022a) Gr(p, n)⊕Gr mexp(Ω)V Ip,n mexp(−Ω)Ip,n (Nguyen & Yang,2023) MK⊕K(1−2K⟨x,y⟩−K∥y∥2)x+(1+K∥x∥2)y 1−2K⟨x,y⟩+K2∥x∥2∥y∥20−x✗(✓for K=0)✓ (Ungar,2009) (Ganea et al.,2018) (Skopek et al.,2019) 3 PSEUDO-REDUCTIVE GYROGROUPS As shown in Tab. 2, the Grassmannian does not satisfy left reduction (G4) in Def. 2.1. However, we find that a relaxed version of (G4) is necessary to guarantee the sample normalization over gyrogroups. Therefore, this section proposes an intermediate between the gyrogroup and the nonreductive one, called pseudo-reductive gyrogroups. Unless specifically emphasized, the gyro structure in this paper is defined as Eq. (3)-Eq. (9). Given a gyrogroup {G, ⊕}, the left gyrotranslation by P∈Gis defined as LP:G→G, with LP(Q) = P⊕Q, ∀Q∈G. (11) If any left gyrotranslation is a gyroisometry, we can use gyrotranslation to center samples. Nguyen & Yang (2023) show that any left gyrotranslation on the SPD and Grassmannian manifolds is a gyroisometry. However, the proof relies on the left cancellation law of gyrogroups, which does not hold for nonreductive gyrogroups. Therefore, the proof is questionable for the Grassmannian. This subsection proposes an intermediate structure, referred to as pseudo-reductive gyrogroups, which can support left cancellation law in general and, therefore, gyroisometry of left gyrotranslation. We illustrate the derivation logic in Fig. 2. Invariance of gyronorm under any gyration Left cancellation law Left gyrotranslation law Axiom (G1-3) Pseudo-reduction Gyroisometries of any gyration and left gyrotranslation Gyroisometry of the gyroinverse Gyrocommutativity Axiom (G1-3) Left reduction (G4) Ours: Previous: Or Figure 2: The conceptual comparison of derivation logic of gyroisometries of our work against previous work (Nguyen & Yang,2023), where the left gyrotranslation law is presented in Lem. G.1. The previous work proves the results on the SPD and Grassmannian manifolds in a case-by-case manner. In contrast, we relax the left reduction into pseudo-reduction and give a general theoretical framework. Our framework also corrects the proof for the Grassmannian cases. Definition 3.1 (Pseudo-reductive Gyrogroups).A groupoid {G, ⊕} is a pseudo-reductive gyrogroup if it satisfies axioms (G1), (G2), (G3) and the following pseudo-reductive law: gyr[X, P] = 1 ,for any left inverse Xof Pin G, (12) 4 Published as a conference paper at ICLR 2025 where 1 is the identity map. Eq. (12) can be intuitively viewed as the intermediate between reduction and non-reduction. For gyrogroups, Eq. (12) can be directly obtained from the left gyroassociativity (G3) and reduction (G4) (Ungar,2009, p. 12). However, there is no theoretical guarantee that Eq. (12) holds for general non-reductive gyrogroups. Therefore, we name Eq. (12) as pseudo-reduction. Nevertheless, for the specific non-reductive Grassmannian, it is indeed pseudo-reductive. Proposition 3.2. [ ↓ ] Gr(p, n) and f Gr(p, n) form pseudo-reductive and gyrocommutative gyrogroups. Our pseudo-reductive gyrogroup naturally generalizes the vanilla gyrogroup, as it shares most of the basic properties of gyrogroups (Ungar,2009, Thms. 1.13 - 1. 14), which are detiled in Thm. D.1. The most related property in Thm. D.1 is the left cancellation law, one of the key prerequisites for gyro translation as gyroisometry. Note that the left cancellation on the gyrogroup comes from left gyro associativity and Eq. (12) (Ungar,2009, p. 12). Therefore, left cancellation does not generally hold for non-reductive gyrogroups, but exists in pseudo-reductive gyrogroups. Next, we point out an iff statement about gyroisometry, which will be useful in the following. Theorem 3.3. [ ↓ ]Given a pseudo-reductive gyrogroup {G, ⊕} , gyr[P, Q] preserves gyronorm for any P, Q ∈G, iff gyr[P, Q]is a gyroisometry for any P, Q ∈G. We find that the isometry of any gyration is a key prerequisite for other gyro operators as isometries. Definition 3.4 (Gyro Left-invariance).The gyrodistance or gyrogroup is gyro left-invariant if any left gyrotranslation is a gyroisometry. Theorem 3.5 (Gyroisometries).[ ↓ ]Given a pseudo-reductive gyrogroup {G, ⊕} with any gyr[·,·] as a gyroisometry, then we have the following: 1. The gyrodistance (Eq. (9)) is gyro left-invariant; 2. Symmetry of the gyrodistance: ∀P, Q ∈G, dgry(P, Q)=dgry(Q, P); 3. If {G, ⊕}is gyrocommutative, then the gyroinverse is a gyroisometry; Proposition 3.6. [ ↓ ]For every (pseudo-reductive) gyrogroup in Tab. 2, the gyrodistance is identical to the geodesic distance (therefore symmetric). The gyroinverse, any gyration and any left gyrotranslation are gyroisometries. Credit of the proof. For the Grassmannian and SPD manifolds, the isometries of gyration, gyroinverse, and left gyrotranslation have been proven by Nguyen & Yang (2023, Thms. 2. 12 - 2. 14 and 2.16 - 2. 18). Nevertheless, in the proof of Thms. 2. 12 - 2. 14, the authors view the non-reductive Grassmannian as a gyrogroup, as they use the left cancellation property of the gyrogroup. Fortunately, our Prop. 3.2 and Thm. D.1 shows that the Grassmannian still enjoys left cancellation. Therefore, all the results on isometries in their Thms. 2.12-2.14 are correct. Nevertheless, these results on the SPD and Grassmannian manifolds can be directly obtained by our Thms. 3.3 and 3.5, as these gyrogroups are pseudo-reductive and the gyrations preserve the gyronorm. Besides, we show the expressions for the gyrodistance of all gyrogroups in Tab. 2, and the associated gyroisometries on MK , both of which is none-trivial. The detailed proof is presented in App. G.4. 4 GYROBN ON GENERAL PSEUDO-REDUCTIVE GYROGROUPS Prop. 3.6 shows that several geometries enjoy isometric gyrotranslation, offering a theoretical foundation for normalizing samples over gyrogroups in a principled manner. Inspired by this, this section develops Riemannian Batch Normalization (RBN) for general pseudo-reductive gyrogroups, referred to as GyroBN. In the following, {M,⊕}is assumed as a pseudo-reductive gyrogroup with the gyro structure defined as Eq. (3)-Eq. (9)1. 4.1 EUCLIDEAN BATCH NORMALIZATION REVISITED As the core operations of different Euclidean normalization variants (Ioffe & Szegedy,2015;Ba et al., 2016;Ulyanov et al.,2016;Wu & He,2018) are similar, this paper focuses on BN. Given a batch of activations {xi...N }, the core operations of BN can be expressed as: ∀i≤N, xi←γxi−µ √v2+ϵ+β(13) 1In our GyroBN, ⊙is not required to comply with the axioms of the gyrovector space (Def. C.4). 5 Published as a conference paper at ICLR 2025 where µ , v2 , γ , and β are the sample mean, sample variance, scaling parameter, and biasing parameter, respectively. 4.2 GYROBN To generalize the Euclidean BN into gyrogroups, we first define sample mean, sample variance, centering, biasing, and scaling over gyrogroups. Then, we introduce our GyroBN framework with a theoretical analysis of the ability to normalize sample statistics. We define the gyromean as the Fréchet mean under the gyrodistance: M= FM({Pi}) = argmin Q∈M 1 NXN i=1 d2 gry (Pi, Q)(14) The gyrovariance is the Fréchet variance, i.e.,the minimization of the rightest hand side of Eq. (14). When the distance in Eq. (14) is the geodesic distance, the Fréchet mean and variance are known as the Riemannian mean and variance. Although the gyromean and gyrovariance are not necessarily the same as the Riemannian ones, Prop. 3.6 indicates the equivalence for the gyrogroups in Tab. 2. Easy computation shows that the centering and biasing in the Euclidean BN (Eq. (13)) can be viewed as gyro addition (Eq. (3)) in Rn , and the scaling can be viewed as gyro scalar product (Eq. (4)), scaling in the tangent space at the identity element. Inspired by this, we define the normalization over gyrogroups by gyro addition and gyro scalar product. Given a batch of activations {Pi...N ∈ M} , we defined the core operations of GyroBN as ∀i≤N, ˜ Pi= Biasing z}|{ B⊕    Scaling z }| { s √v2+ϵ⊙  Centering z }| { ⊖M⊕Pi    ,(15) where M∈ M and v2 are gyromean and gyrovariance, B∈ M is the biasing parameter, s∈R is the scaling parameter, and ϵis a small value for numerical stability. Theorem 4.1 (Homogeneity).[ ↓ ]Supposing {M,⊕} is a pseudo-reductive gyrogroup with any gyration gyr[·,·] as a gyroisometry, for N samples {Pi...N ∈ M} , we have the following properties: Homogeneity of gyromean: FM({B⊕Pi}) = B⊕FM({Pi}),∀B∈ M,(16) Homogeneity of dispersion from E:1 NXN i=1 d2 gry(t⊙Pi, E) = t2 NXN i=1 d2 gry(Pi, E),(17) The most important property of the Euclidean BN (Eq. (13)) lies in its ability to normalize data distribution by the control over sample mean and variance. Similarly, our formulation in Eq. (15) can also normalize gyromean and gyrovariance. Specifically, given a pseudo-reductive gyrogroup with isometric gyrations, Eq. (16) indicates that the centering and biasing can transfer the gyromean, while Eq. (17) can scale the sample variance, since after centering, the resulting gyromean is the identity element E. To finalize our GyroBN, we define the running mean updates over gyrogroups as the binary barycenter based on gyrodistance: Barη(P1, P2) = argminQ∈M ηd2 gry (P1, Q) + (1 −η)d2 gry (P2, Q),with η∈[0,1].(18) Notably, when the gyrodistance is identical to the geodesic distance, the binary barycenter can be calculated by geodesic. With all the above ingredients, the general framework for our GyroBN is presented in Alg. 1. Thm. 4.1 indicates that given a pseudo-reductive gyrogroup with any gyration as a gyroisometry, our GyroBN enjoys a theoretical guarantee of control over the gyro statistics. Specifically, for the gyrogroups in Tab. 2, our GyroBN can control the gyromean and gyrovariance. Besides, as the gyromean and gyrovariance are identical to the Riemannian counterparts, the GyroBNs on these gyrogroups also normalize the Riemannian statistics. Especially, simple computation shows that our GyroBN recovers the standard Euclidean BN (Ioffe & Szegedy,2015) when M=Rn. 6 Published as a conference paper at ICLR 2025 Algorithm 1: Gyrogroup Batch Normalization (GyroBN) Require : batch of activations {P1...N ∈ M}, small positive constant ϵ, and momentum η∈[0,1], running mean Mr, running variance v2 r, biasing parameter B∈ M, scaling parameter s∈R. Return :normalized batch {˜ P1...N ∈ M} 1if training then 2Compute batch mean Mband variance v2 bof {P1...N }; 3Update running statistics Mr= Barγ(Mb, Mr),v2 r=γv2 b+ (1 −γ)v2 r; 4end 5(M, v2)=(Mb, v2 b)if training else (Mr, v2 r) 6∀i≤N, ˜ Pi=B⊕s √v2+ϵ⊙(⊖M⊕Pi) 5 AGYRO PERSPECTIVE FOR THE EXISTING RIEMANNIAN NORMALIZATIONS Several existing Riemannian normalization methods on different geometries enjoy theoretical control of sample mean and variance, including LieBN on general Lie groups (Chen et al.,2024b), and SPDBNs based on the specific AIM geometry (Brooks et al.,2019;Kobler et al.,2022b;a). This subsection further reveals that they are concrete implementations of our GyroBN. 5.1 LIEBN AS A SPECIAL CASE OF GYROBN Chakraborty (2020, Algs. 3-4) first proposed Riemannian normalization on matrix Lie groups under a specific distance. Chen et al. (2024b) extended their framework into general Lie groups, referred to as LieBN, with a theoretical control over Riemannian mean and variance. This subsection shows that LieBN is indeed a special case of our GyroBN. LieBN is established under a left-invariant metric on the Lie group. The centering and biasing are defined by left group translation, and scaling is defined by the scaling on the tangent space at the identity element (Chen et al.,2024b, Eq.13-15). As every Lie group is automatically a gyrogroup, the gyrotranslation is the exact group translation. Therefore, the centering, biasing, and scaling are the same under GyroBN and LieBN. However, the mean, variance, and running mean updates on LieBN are defined based on the geodesic distance. In contrast, the counterparts on the GyroBN are based on gyrodistance. Nevertheless, the following proposition demonstrates the equivalence of these operators under the GyroBN with LieBN. Proposition 5.1. [ ↓ ]Given a Lie group with a left-invariant metric, the gyrodistance and geodesic distance are identical. The GyroBN is, therefore, identical to the LieBN (Chen et al.,2024b, Alg. 1). Chen et al. (2024b) implemented LieBN on four left-invariant geometries, including SPD manifold with AIM 2 , LEM and LCM, and rotation matrices. According to Prop. 5.1, these implementations are immediately the special cases of our GyroBN. 5.2 AIM-BASED SPDBNS AS SPECIAL CASES OF GYROBN Several RBNs on the SPD manifold were developed based on AIM (Brooks et al.,2019;Kobler et al., 2022b;a). The core operations of these approaches can be expressed as the following: Normalization: ∀i≤N, ˜ Pi=B1 2M−1 2PiM−1 2s √v2+ϵB1 2(19) where M and v2 are the Riemannian mean and variance, i.e.,the Fréchet mean and variance under the geodesic distance. The running mean is updated by binary barycenter under the geodesic distance. Prop. 3.6 indicates that, under the AIM geometry, the gyrodistance is identical to the geodesic distance. Therefore, the gyromean and gyrovariance are identical to the Riemannian mean and variance. The running mean updates are also identical under the gyrodistance and geodesic distance. Besides, simple computations show that Eq. (19) is exactly the specific implementation of Eq. (15) under {Sn ++,⊕AI, gAI} , where gAI denotes AIM. Therefore, the SPDBNs developed by Brooks et al. (2019); Kobler et al. (2022b;a) are also special cases of our GryoBN. 2 AIM is left-invariant w.r.t. the Lie group operation P⊕LieAI Q=LQL⊤ with L as the Cholesky factor of P=LL⊤(Thanwerdas & Pennec,2022). This group structure differs from the one presented in Tab. 2. 7 Published as a conference paper at ICLR 2025 Remark 5.2.Brooks et al. (2019) only consider centering and biasing. Kobler et al. (2022b) use running mean for centering during the training. Kobler et al. (2022a) use different momentum to update running statistics for training and testing and multi-channel mechanisms for domain adaptation. Nevertheless, all of them are based on Eq. (19). Therefore, tricks such as multi-channel and separate momentum can also be applied to our GyroBN. This is what we mean by claiming that our GyroBN incorporates their approaches. 6 GYROBNS ON GRASSMANNIAN AND HYPERBOLIC SPACES As indicated by Prop. 3.6 and Thm. 4.1, our GyroBN can be implemented on the Grassmannian and hyperbolic spaces, with the ability to normalize sample statistics. Our Alg. 1allows us to implement GyroBN in a plug-and-play manner. This section clarifies additional technical details regarding its application in these two spaces. As Prop. 3.6 has demonstrated the equivalence of the gyrodistance and geodesic distance on these spaces, we use the terms "gyromean" and "gyrovariance" interchangeably with their Riemannian counterparts. 6.1 GRASSMANNIAN GYROBN We focus on the ONB perspective. To the best of our knowledge, this is the first RBN on the Grassmannian. Given a batch of activations {U1···N} over {Gr(p, n),⊕Gr} , Eq. (15) can be expressed as Centering to the identity Ip,n:U1 i= mexp −[MM⊤,e Ip,n]Ui,(20) Scaling the dispersion from Ip,n:U2 i= mexp s √v2+ϵ[U1 i(U1 i)⊤,e Ip,n]Ip,n,(21) Biasing towards parameter B∈ M:U3 i= mexp [B, e Ip,n]U2 i(22) where (·) = g Loge Ip,n (·) with g Log as the Riemannian logarithm under the projector perspective, and M is the Riemannian mean of {U1···N} , and e Ip,n =Ip,nI⊤ p,n is the identity element under the projector perspective. The Riemannian mean can be calculated by the Karcher flow (Karcher,1977). We use Alg. 5.3 by Bendokat et al. (2024) for a stable and efficient computation of the Riemannian logarithm required in the Karcher flow. For [MM⊤,e Ip,n] and [U1 i(U1 i)⊤,e Ip,n] , inspired by Bendokat et al. (2024, Alg. 5.3) and Nguyen et al. (2024, Prop. 3.12), we have the following for fast computation. Proposition 6.1. [ ↓ ]Given U= (U⊤ 1, U⊤ 2)⊤∈Gr(p, n) with U1∈Rp×p and U2∈R(n−p)×p , then we have the following [UU⊤,e Ip,n] = 0−e UT 2 e U20,(23) where e U2=U2Qarcsin( ˆ S) ˆ SR⊤ and U⊤ 1 SVD := QSR⊤ . Here S is in ascending order, Q and R are column-wisely flipped accordingly, and ˆ S=√1−S2. Remark 6.2.Although we focus on the GyroBN under the ONB perspective, the GyroBN under the projector perspective can be calculated via the ONB perspective by the following process: (1) mapping data into the ONB perspective by π−1:f Gr(p, n)→Gr(p, n) ; (2) normalizing data by the GyroBN under Gr(p, n) ; (3) mapping normalized data back to f Gr(p, n) by π . Technical details are presented in App. E. 6.2 HYPERBOLIC GYROBN We focus on the Poincaré ball model over the hyperbolic space, i.e., Pn K . The specific manifestation can be conducted in a plug-in manner. We simply need to plug the required operators from Tab. 2and Tab. 8into Alg. 1. Besides, the Poincaré Fréchet mean can be calculated by Lou et al. (2020, Alg. 1)]. Remark 6.3.As shown by Cannon et al. (1997), there are five isometric models that one can work for the hyperbolic spaces. Although we focus on the Poincaré ball, the GyroBN under other isometric metrics can also be easily constructed via the Poincaré GyroBN. The overall process is similar to Rmk. 6.2. For more detail, please refer to Lem. E.1 and Thm. E.2. 8 Published as a conference paper at ICLR 2025 7 EXPERIMENTS Our GyroBN layers are model-agnostic and can be seamlessly integrated into any network operating over the gyrospaces listed in Tab. 2. This section evaluates the effectiveness of our GyroBN on Grassmannian and hyperbolic neural networks. 7.1 EXPERIMENTS ON THE GRASSMANNIAN Implementation. We focus on a newly developed Grassmannian network, GyroGr (Nguyen & Yang,2023), which replaces the non-intrinsic transformation block (FRMap + ReOrth layers) in the classic GrNet (Huang et al.,2018) with Grassmannian left gyrotranslation. This modification has demonstrated improved numerical performance and stability (Nguyen & Yang,2023). GyroGr is constructed over the ONB Grassmannian and consists of three basic blocks: left gyrotranslation, pooling (Huang et al.,2018), and the projection map (ProjMap) (Huang et al.,2018). The projection map functions as an activation layer that maps data into symmetric matrices for classification. Following Nguyen & Yang (2023), we evaluate our method on skeleton-based action recognition tasks, including the HDM05 (Müller et al.,2007), NTU60 (Shahroudy et al.,2016), and NTU120 (Liu et al.,2019) datasets, focusing on mutual actions for NTU60 and NTU120. For a fair comparison, we also extend ManifoldNorm (Chakraborty,2020, Alg. 1-2) and RBN (Lou et al.,2020, Alg. 2) to the Grassmannian. Although these two BN methods were not originally implemented over the Grassmannian, they can be adapted by leveraging Riemannian operators such as geodesics, exponential and logarithmic maps, and parallel transport. The key difference is that our GyroBN can normalize data distributions across different geometries, whereas the other two methods cannot. Further details on datasets, implementation, and training efficiency are provided in App. F.2. Table 3: Comparison of GyroBN against other Grassmannian BNs under GyroGr backbone. BN None ManifoldNorm-Gr RBN-Gr GyroBN-Gr Acc. Mean±std Max Mean±std Max Mean±std Max Mean±std Max HDM05 48.97±0.24 49.23 49.67±0.76 50.41 48.64±0.77 49.49 51.89±0.37 52.43 NTU60 70.13±0.16 70.32 68.56±0.43 69.14 67.77±0.52 68.35 72.60±0.04 72.65 NTU120 53.76±0.18 53.96 51.41±0.38 51.92 50.56±0.22 50.82 55.47±0.10 55.59 Table 4: Ablation of Grassmannian GyroBN under various network architectures. HDM05 NTU60 NTU120 Architecture 1Block 2Block 3Block 4Block 1Block 2Block 3Block 4Block 1Block 2Block 3Block 4Block GyroGr 49.23 49.09 47.02 27.36 70.32 70.14 70.23 65.03 53.96 54.1 54.59 47.59 GyroGrBN 52.43 50.62 51.56 30.29 72.65 71.93 72.25 66.67 55.59 56.15 54.63 48.9 Main results. We compare our GyroBN with previous BN methods under the 1Block GyroGr backbone, which consists of one block of gyrotranslation and pooling layers followed by a ProjMap layer. We apply the BN after the pooling layer. The 5-fold results are presented in Tab. 3. GyroBN consistently delivers improved performance, enhancing the average performance of the vanilla GyroGr by 2.92%, 2.47%, and 1.71%. In contrast, Grassmannian ManifoldNorm and RBN could degrade the vanilla GyroGr network, particularly on the NTU60 and NTU120 datasets. This is primarily due to GyroBN’s theoretical guarantee of normalizing sample statistics, a capability lacking in previous methods such as ManifoldNorm and RBN, as summarized in Tab. 1. Additionally, we observe that GyroBN mitigates the performance gap between training and testing, indicating its ability to enhance the model’s generalization. Detailed discussions are provided in App. F.5. Overall, the above findings highlight the effectiveness of our GyroBN in facilitating network training. Ablations on the architecture. We further validate our GyroBN across different architectures within the GyroGr baseline, which includes up to four blocks of gyrotranslation and pooling. GyroBN is applied after the first pooling layer, and we denote GyroGr with GyroBN as GyroGrBN. As implied by Tab. 3, both GyroGr and GyroGrBN exhibit relatively small variances, allowing us to conduct ablations using a single trial. Tab. 4reports the results across all three datasets. We observe that GyroBN consistently improves the performance of the vanilla GyroGr baseline, highlighting the effectiveness of the GyroBN framework. Notably, as the network depth increases, the performance of the GyroGr backbone, with or without GyroBN, degrades. This is because the dimensionality of the 9 Published as a conference paper at ICLR 2025 G Proofs 30 G.1 Proof of Prop. 3.2 ................................... 30 G.2 Proof of Thm. 3.3 ................................... 31 G.3 Proof of Thm. 3.5 ................................... 32 G.4 Proof of Prop. 3.6 ................................... 32 G.5 Proof of Thm. 4.1 ................................... 35 G.6 Proof of Prop. 5.1 ................................... 36 G.7 Proof of Prop. 6.1 ................................... 37 16 Published as a conference paper at ICLR 2025 A LIMITATION As shown in Prop. 3.6, several geometries, including SPD, Grassmannian, hyperbolic, and hyperpherical manifolds, have gyro structures that can enable GyroBN with a theoretical control over sample statistics. However, some geometries do not have gyro structures, or their gyro structures are still untouched. Therefore, GyroBN cannot be established on these manifolds. We will explore other techniques to establish an RBN framework for these non-gyro geometries. B NOTATIONS Tab. 6summarizes all the notations in the main paper. Table 6: Summary of notations. Notation Explanation {G, ⊕}or abbreviated as GA gyrogroup {M,⊕, g}or abbreviated as MA Riemannian manifold {M, g}with a gyrogroup structure induced by g 1 identity map TPMTangent space at P∈ M gp(·,·)or ⟨·,·⟩PRiemannian metric at P∈ M ∥·∥PThe norm induced by ⟨·,·⟩Pon TPM dgeo(·,·)Geodesic distance LogPRiemannian logarithm at P ExpPRiemannian exponentiation at P f∗,P The differential map of the smooth map fat P∈ M PTP→QParallel transportation along the geodesic connecting Pand Q EGyro identity of {M,⊕} ⊖PGroup inverse of P∈ M ⊙Gyro scalar product gyr[·,·]Gyration ⟨·,·⟩gr Gyro inner product ∥·∥gr Gyronorm dgry(·,·)Gyrodistance FM Fréchet mean under a gyrodistance Barη(·,·)Binary barycenter based on a gyrodistance ⟨·,·⟩ The standard Frobenius inner product ∥·∥ Norm induced by the standard Frobenius inner product mlog Matrix logarithm mexp Matrix exponentiation LCholesky decomposition Dlog The diagonal element-wise logarithm ψLC Dlog ◦L Sn ++ The SPD manifold SnThe Euclidean space of symmetric matrices Pn Kn-dimensional Poincaré ball with curvature K < 0 Rnn-dimensional Euclidean space Dn K,n-dimensional projected hypersphere with curvature K > 0 MKConstant Curvature Spaces (CCS) Gr(p, n)The Grassmannian under the ONB perspective f Gr(p, n)The Grassmannian under the projector perspective Ip,n The Grassmannian identity under the ONB perspective e Ip,n The Grassmannian identity under the projector perspective πThe Riemannian isometry from Gr(p, n)onto f Gr(p, n) [·,·]Matrix commutator [·]An element in Gr(p, n), which is a equivalent class Inn×nidentity matrix PθMatrix power for SPD matrix P ⊕AI,⊕LE and ⊕LC Gyro additions on the SPD manifold under AIM, LEM and LCM e ⊕Gr and ⊕Gr Gyro additions on the Grassmannian under the ONB and projector perspectives ⊕KGyro additions on the CCS (·) (·) = g Loge Ip,n (·)with g Log as the Riemannian logarithm on f Gr(p, n) gAI,gLE and gLC AIM, LEM, LCM on the SPD manifold gGr and egGr Riemannian metrics on the Grassmannian under the ONB and projector perspectives tanKtanK= tan if k > 0, elif K > 0,tanK= tanh 17 Published as a conference paper at ICLR 2025 C PRELIMINARIES C.1 RIEMANNIAN GEOMETRY Manifolds can be intuitively understood as locally Euclidean spaces. Differentials generalize the concept of classical derivatives. For a detailed introduction to smooth manifolds, we refer readers to Tu (2011); Lee (2013). A Riemannian manifold is a manifold equipped with a Riemannian metric, which can be intuitively interpreted as a point-wise inner product. This metric allows for the adaptation of various Euclidean operators to the manifold setting. For an in-depth discussion on Riemannian manifolds, see Do Carmo & Flaherty Francis (1992); Lee (2018). Definition C.1 (Riemannian Manifolds (Do Carmo & Flaherty Francis,1992)).A Riemannian metric on M is a smooth symmetric covariant 2-tensor field on M , which is positive definite at every point. A Riemannian manifold is a pair {M, g} , where M is a smooth manifold and g is a Riemannian metric. The isometries generalize the bijection in the set theory into the Riemannian geometry. If two manifolds are isometric, they can be viewed as equivalent. The Riemannian operators in these two manifolds are also closely related (Chen et al.,2024d, App. A. 2). The following defines the Riemannian isometry. Definition C.2 (Isometries (Lee,2018)).If {M, g} and {f M,eg} are both Riemannian manifolds, a smooth map f:M → f Mis called a (Riemannian) isometry if it is a diffeomorphism that satisfies gp(V, W) = egf(p)(f∗,p(V), f∗,p(W)),(24) where f∗,p(·) : TpM → Tf(p)f M is the differential map of f at p∈ M , and V, W ∈TpM are two tangent vectors. The exponential & logarithmic maps and parallel transportation are also crucial for Riemannian approaches in machine learning. To bypass the notation burdens caused by their definitions, we review the geometric reinterpretations of these operators (Pennec et al.,2006;Do Carmo & Flaherty Francis, 1992). In detail, in a manifold M , geodesics correspond to straight lines in Euclidean space. A tangent vector −→ xy ∈TxM can be locally identified to a point y on the manifold by geodesic starting at x with an initial velocity of −→ xy , i.e. y= Expx(−→ xy) . On the other hand, the logarithmic map is the inverse of the exponential map, generating the initial velocity of the geodesic connecting x and y , i.e., −→ xy = Logx(y) . These two operators generalize the idea of addition and subtraction in Euclidean space. For the parallel transportation PTx→y(V) , it is a generalization of parallelly moving a vector along a curve in Euclidean space. we summarize the reinterpretation in Tab. 7. Table 7: The geometric reinterpretations of Riemannian operators. Operations Euclidean spaces Riemannian manifolds Straight line Straight line Geodesic Subtraction −→ xy =y−x−→ xy = Logx(y) Addition y=x+−→ xy y = Expx(−→ xy) Parallelly moving V→VPTx→y(V) A Lie group is a manifold with a smooth group structure. It is a combination of algebra and geometry. Definition C.3 (Lie Groups).A manifold is a Lie group, if it forms a group with a group operation ⊙ such that m(x, y)7→ x⊙yand i(x)7→ x−1 ⊙are both smooth, where x−1 ⊙is the group inverse of x. The following are some naive examples of the Lie group: 1. The set of real numbers R, whose group operation is the addition. 2. The set of n×n invertible matrices GL(n) , whose group operation is the matrix product. This group is known as the general linear group. 3. The set of n×n orthogonal matrices O(n) , whose group operation is the matrix product. This group is a subgroup of GL(n), known as the orthogonal group. 18 Published as a conference paper at ICLR 2025 C.2 GYROVECTOR SPACES Gyrogroups in Def. 2.1 generalize groups to non-associative algebraic systems by gyrations. Similarly, the gyrovector space generalizes the vector space, which has shown impressive success in hyperbolic geometry (Ungar,2005;2009;2012;2014). Definition C.4 (Gyrovector Spaces (Nguyen,2022a)).A gyrocommutative gyrogroup {G, ⊕} equipped with a scalar multiplication ⊙:R×G→G is called a gyrovector space if it satisfies the following axioms for s, t ∈Rand a, b, c ∈G: (V1) 1⊙a=a, 0⊙a=t⊙e=e, and (−1) ⊙a=⊖a. (V2) (s+t)⊙a=s⊙a⊕t⊙a. (V3) (st)⊙a=s⊙(t⊙a). (V4) gyr[a, b](t⊙c) = t⊙gyr[a, b]c. (V5) gyr[s⊙a, t ⊙a] = 1 , where 1 is the identity map. Gyrovector spaces generalize vector spaces to curved geometries, such as the SPD and Grassmannian manifolds (Nguyen,2022a). While retaining familiar properties like distributivity (V2) and associativity (V3), they incorporate the complexities of gyrations. C.3 MATRIX AND VECTOR MANIFOLDS C.3.1 DEFINITIONS SPD: The set Sn ++ of n×n SPD matrices form a manifold, named the SPD manifold (Pennec et al., 2006). We focus on three popular Riemannian metrics on the SPD manifold: Affine-Invariant Metric (AIM) (Pennec et al.,2006), Log-Euclidean Metric (LEM) (Arsigny et al.,2005), and Log-Cholesky Metric (LCM) (Lin,2019). Grassmannian: The Grassmannian is the set of p -dimensional subspace of n -dimensional vector space (Tu,2011), which has two matrix representations (Bendokat et al.,2024): Projector perspective: f Gr(p, n) = {P∈ Sn:P2=P, rank(P) = p}, ONB perspective: Gr(p, n) = {[U]:[U] := {e U∈St(p, n)|e U=UR, R ∈O(p)}},(25) where Sn is the Euclidean space of symmetric matrices, St(p, n) is the Stiefel manifold, and O(p) is the orthogonal group. By abuse of notations, we use [U] and U interchangeably for the element of Gr(p, n) . As shown by Helmke & Moore (2012), the ONB perspective Gr(p, n) is diffeomorphism to f Gr(p, n)by π(U) = UU⊤,∀U∈Gr(p, n),(26) where the n×p column-wise orthonormal matrix U should be more precisely understood as a representative of an equivalence class (Bendokat et al.,2024). CCS: The Poincaré ball Pn K for the hyperbolic space (Ungar,2009;Ganea et al.,2018), projected hypersphere Dn K for the hypersphere (Skopek et al.,2019), and standard Euclidean space Rn (Zorich & Paniagua,2016) are more generally called the Constant Curvature Space (CCS) (Do Carmo & Flaherty Francis,1992), as they have constant sectional curvature K . The n -dimensional Poincaré ball and projected hypersphere are represented as Pn K=x∈Rn:⟨x, x⟩<−1 K with K < 0 and Dn K=Rnwith K > 0. When K= 0, the CCS becomes the standard Euclidean space Rn. 19 Published as a conference paper at ICLR 2025 C.3.2 GYRO AND RIEMANNIAN STRUCTURES For a matrix manifold Mand a CCS MK, we make the following notations: {M, g}=             {Sn ++, gAI}(The SPD manifold under AIM) {Sn ++, gLE}(The SPD manifold under LEM) {Sn ++, gLC}(The SPD manifold under LEM) {Gr(p, n), gGr}(The Grassmannian under the ONB perspective) {f Gr(p, n),egGr}(The Grassmannian under the projector perspective) (27) MK=   Pn K,for K < 0 Rn,for K= 0 Dn K,for K > 0 (28) tanK=tan if K > 0 tanh if K < 0(29) Notations in Tab. 2:We summarize all the necessary group operations in Tab. 2with the following notations. Given any P, Q ∈ M with M as Sn ++ or f Gr(p, n) , and x, y ∈ MK with MK as Pn K(K < 0) , Dn K(K > 0) or Rn(K= 0) , we make the following notations. For the Grassmannian, U=π−1(P) and V=π−1(Q) are the ONB representations. We denote matrix exponential, matrix logarithm, and Cholesky decomposition as mexp , mlog , and L , respectively. We denote ψLC = Dlog ◦L , where ψLC is the diagonal logarithm. As shown by Chen et al. (2024e), LCM and the associated Lie group on Sn ++ are pulled back by ψLC for the Euclidean space of lower triangular matrices. In is the n×n identity matrix, Ip,n = (Ip,0)⊤∈Rn×p , and e Ip,n =π(Ip,n) . For the Grassmannian f Gr(n, p) , Ω=[P, e Ip,n] , where P= Loge Ip,n (P) and [·,·] is the matrix commutator. We denote ⟨·,·⟩ and ∥·∥as the standard (matrix & vector) inner product and norm. We further make the following notations for the related Riemannian operators in Tab. 8. Given P∈ M ( x∈ MK ), we denote the tangent vector as V∈TPM ( v∈TxMK ). We denote the geodesic distance, Riemannian logarithm, and Riemannian exponential as dgeo , Log and Exp , respectively. Table 8: Riemannian geometries of several matrix and vector manifolds. For CCS MK, we present the operators for K= 0 , as when K= 0 , all the Riemannian operators are reduced to the familiar vector operators. Manifolds dgeo(P, Q)or dgeo(x, y) LogPQor LogxyExpPVor ExpxvReferences {Sn ++, gLE} ∥mlog(P)−mlog(Q)∥(mlog∗,P )−1(mlog(Q)−mlog(P)) mexp mlog(P) + mlog∗,P (V)(Arsigny et al.,2005) {Sn ++, gAI}mlog Q−1 2PQ−1 2P1 2mlog P−1 2QP−1 2P1 2P1 2mexp P−1 2V P−1 2P1 2(Pennec et al.,2006) {Sn ++, gLC} ∥ψLC(P)−ψLC(Q)∥(L−1)∗,L (ψLC(Q)−ψLC(P))) ψ−1 LC (ψLC(P)+(ψLC)∗,P (V)) (Lin,2019) (Chen et al.,2024e) Gr(p, n)∥arccos(Σ)∥ P⊤QSVD := OΣR⊤ Oarctan(Σ)R⊤ (In−PP ⊤)Q(P⊤Q)−1SVD := OΣR⊤PR cos(Σ)RT+Osin(Σ)RT VSVD := OΣR⊤ (Edelman et al.,1998) (Bendokat et al.,2024) f Gr(p, n)1 2∥mlog ((In−2Q) (In−2P))∥1 2[mlog ((In−2Q) (In−2P)) , P ] mexp([V, P ])Pmexp(−[V, P ]) (Bendokat et al.,2024) (Batzies et al.,2015) MK2 √|K|tan−1 Kp|K|∥−x⊕Ky∥2 √|K|λK x tan−1 Kp|K|∥−x⊕Ky∥−x⊕Ky ∥−x⊕Ky∥x⊕KtanKp|K|λK x∥v∥ 2v √|K|∥v∥(Do Carmo & Flaherty Francis,1992) (Petersen,2006) (Skopek et al.,2019) Remark C.5 (The Grassmannian and cut locus).Due to the cut locus of the Grassmannian, the logarithm map does not exist globally (Bendokat et al.,2024). In this paper, when we use LogP(Q) on the Grassmannian, we implicitly assume P and Q are not in each other’s cut locus. Besides, more precisely, the gyro addition and scalar multiplication on Grassmannian are also not globally defined (Nguyen,2022a), due to the cut locus. Following Nguyen (2022a) and Nguyen & Yang (2023), we implicitly assume the gyro operations are well-defined on the Grassmannian. C.4 INTUITIVE EXPLANATIONS OF GYROGROUPS This subsection intuitively explains gyrogroups by contrast with the trivial R . Gyrogroups are proposed to generalize the concept of addition in Ror Rnto non-Euclidean spaces. Gyrogroups in Def. 2.1 extend the concept of groups to non-associative algebraic systems, i.e.,a⊕ (b⊕c)= (a⊕b)⊕c . The gyrogroups over the manifolds in Tab. 2are defined by Eqs. (3), (5) and (6). 20 Published as a conference paper at ICLR 2025 More importantly, these definitions are natural generalizations of Euclidean vector operations. Take the gyro addition Eq. (3) as an example. When the manifold M is the Euclidean space Rn , the gyro addition defined by Eq. (3) becomes the familiar vector addition. Gyration & gyroassociativity: The gyration is an automorphism gyr[a, b] : G→G for each a, b ∈G. It is called an automorphism as it can preserve the gyro addition: gyr[a, b](c⊕d) = gyr[a, b](c)⊕gyr[a, b](d),∀c, d ∈G. (30) The gyration is used to generalize the associativity: a+ (b+c)=(a+b) + c, ∀a, b, c ∈R.(31) Take the Grassmannian gyrogroup {Gr(p, n),⊕Gr} as an example. For any U, V, R ∈ ⊕Gr , we have Non-associativity: V⊕Gr (U⊕Gr R)= (V⊕Gr U)⊕Gr R, (32) Left gyroassociativity: V⊕Gr (U⊕Gr R) = (V⊕Gr U)⊕Gr gyrGr[U, V ](R).(33) Eq. (33) differs with Eq. (31) only in a gyration. The concrete expression of the gyration over the Grassmannian is presented by Nguyen (2022a, Def. 3.18). Left reduction is defined as gyr[a, b] = gyr[a⊕b, b] for any a and b in the gyrogroup G . It is called a "left reduction" because bin a⊕bcan be canceled out. Gyrocommutativity in Def. 2.2 generalizes the commutative property ( a+b=b+a ) by the gyration operator gyr[a, b]. Nonreductive gyrogroups in Def. 2.3 are relaxed gyrogroups, allowing structures where the left reduction law (G4) does not hold. Its prototype comes from the gyrogroup of the Grassmannian (Nguyen,2022a, Thm. 3.20), where (G4) does not hold. Relation between left reduction, non-reduction, & pseudo-reduction: The left reduction can induce many basic properties of gyrogroups ( see Thm. D.1 and Ungar (2009, Thms. 1.13 and 1.14)). However, non-reduction does not guarantee several basic properties. This drawback will undermine the rationality of the non-reductive gyrogroups. Therefore, as an intermediate, we propose pseudo-reduction. The associated pseudo-reductive gyrogroups maintain most of the basic properties of gyrogroups. Please refer to Thm. D.1 and its remark for more details. D FIRST PSEUDO-REDUCTIVE GYROGROUPS PROPERTIES Theorem D.1 (First Pseudo-reductive Gyrogroups Properties).Let {G, ⊕} be a pseudo-reductive gyrogroup. For any elements P, Q, R, X ∈G, we have: 1. If P⊕Q=P⊕R, then Q=R(General Left Cancellation law; see (9) below). 2. gyr[E, P] = 1 for any left identity Ein G. 3. gyr[X, P] = 1 for any left inverse Xof Pin G. 4. There is P left identity which is P right identity. 5. There is only one left identity. 6. Every left inverse is P right inverse. 7. There is only one left inverse, ⊖P, of P, and ⊖(⊖P) = P. 8. The left cancellation law: ⊖P⊕(P⊕Q) = Q. 9. The gyrator identity: gyr[P, Q]X=⊖(P⊕Q)⊕{P⊕(Q⊕X)}. 10. gyr[P, Q]E=E. 11. gyr[P, Q](⊖X) = ⊖gyr[P, Q]X. 12. gyr[P, E] = 1 . 13. The gyrosum inversion law: ⊖(P⊕Q) = gyr[P, Q](⊖Q⊖P) 21 Published as a conference paper at ICLR 2025 Credit of the proof. Most of the proof is borrowed from Ungar (2009, Thms. 1.13 and 1.14), which presents the corresponding properties for gyrogroups. After re-analyzing these two theorems, we find that most of the proof can be readily extended into pseudo-reductive gyrogroups. Proof. This theorem follows Ungar (2009, Thms. 1.13 and 1.14), which presents some useful properties for gyrogroups. We argue that all the properties except gyr[a, a] = 1 are independent of the left reduction law (G4), and are therefore satisfied on pseudo-reductive gyrogroups. All the properties can be proven in the same way as the ones for Thms. 1.13 and 1. 14 by Ungar (2009). We summarize the logic in the following: • left gyroassociativity ⇒1 • left gyroassociativity + 1⇒2 • definition ⇒3 • left gyroassociativity + 1+3⇒4 • definition ⇒5 • left gyroassociativity + (G2) + 1+3+4+5⇒6 •1+6⇒7 • left gyroassociativity +3⇒8 • left gyroassociativity + left cancellation in 8⇒9 • gyro identity in 9⇒10 •10 ⇒11 • left cancellation in 8+ gyro identity in 9⇒12 • left cancellation in 8+ gyro identity in 9⇒13 Remark D.2.For non-reductive gyrogroups, 2and 3are agnostic. Therefore, every property based on 2or 3is not guaranteed, such as 4, and from 6to 10. The missing of these basic properties will undermine the rationality of non-reductive gyrogroups. In contrast, most basic properties of gyrogroups are preserved in our pseudo-reductive gyrogroups. E GRASSMANNIAN GYROBN UNDER THE PROJECTOR PERSPECTIVE Given two isometric manifolds {M1, g1} and {M2, g2} , the induced gyro structures in Eq. (3)- Eq. (9) have the following relations. We first present a useful lemma. Lemma E.1. Given manifolds {M1, g1} and {M2, g2} and a Riemannian isometry f:M1→ M2 , we have the following: 1. The groupoid {M1,⊕1} induced by g1 is pseudo-reductive (left-invariant), iff the groupoid {M2,⊕2}induced by g2is pseudo-reductive (left-invariant); 2. Any gyration in {M1, g1} preserves gyronorm iff any gyration in {M2, g2} preserves gyronorm; 3. fpreserves the gyrodistance. Proof. First, we review some facts about gyrogroups under the Riemannian isometry. As shown by Nguyen & Yang (2023, Thm. 2.5), {M1,⊕1} satisfies (G1-G3) in Def. 2.1 iff {M2,⊕2} satisfies (G1-G3). Besides, f−1is an isomorphism satisfying f−1(P⊕2Q) = f−1(P)⊕1f−1(Q),(34) f−1(t⊙2P) = t⊕1f−1(P),(35) f−1(⊖2P) = ⊖1f−1(P),(36) gyr2[P, Q](R) = fgyr1[f−1(P), f−1(Q)](f−1(R)).(37) 22 Published as a conference paper at ICLR 2025 where P, Q, R ∈ M2 are arbitrary points, t∈R is a real scalar, ⊖i , gyri are the gyro inverses and gyrations on Mifor i= 1,2.fis also an isomorphism with similar properties. We only need to prove one direction for the iff condition for the pseudo-reduction or norm invariance. We focus on ⇒and follow the above notations in the following. Pseudo-reduction: gyr2[⊖P, P] = f◦gyr1[f−1(⊖2P), f−1(P)] ◦f−1(Eq. (37)) =f◦gyr1[⊖1f−1(P), f−1(P)] ◦f−1(Eq. (36)) = 1 . (38) Gyronorm invariance under gyrations: For simplicity, we denote the gyronorm, identity element, and Riemannian logarithm on Mi as ∥∥iEi , and Logi , respectively. First, the following demonstrates that the Riemannian isometry fpreserves gyronorm: ∥P∥2=Log2 E2(P)E2 (1) =Log1 f−1(E2)f−1(P)f−1(E2) (2) =LogE1f−1(P)E1 =f−1(P)1 (39) (1) As f:M1→ M2is a Riemannian isometry, we have the following equations: Log2 PQ=f∗,f−1(P)Log1 f−1(P)f−1(Q),∀P, Q ∈ M2,(40) g2 P(V, W) = g1 f−1(P)((f−1)∗,P (V),(f−1)∗,P (W)),∀V, W ∈TPM2,(41) where (·)∗is the differential map; (2) E2=f(E1). Then we have the following: ∥gyr2[P, Q](R)∥2 (1) =fgyr1[f−1(P), f−1(Q)](f−1(R))2 (2) =gyr1[f−1(P), f−1(Q)](f−1(R))1 (3) =f−1(R)1 (4) =∥R∥2 (42) The above derivation comes from the following. (1) Eq. (37); (2) Eq. (39); (3) Any gyration on M1can preserve the gyronorm; (4) Eq. (39). Invariance of gyrodistance under f :Denoting the gyrodistance on Mi as di , for any U, V ∈ M1 , we have the following: d1(U, V ) = ∥⊖1U⊕1V∥1 (1) =∥f(⊖1P⊕1Q)∥2 (2) =∥⊖1f(P)⊕2f(Q)∥2 = d2(f(P), f(Q))2 (43) The above derivation comes from the following. (1) fpreserves gyronorm; 23 Published as a conference paper at ICLR 2025 (2) fis an isomorphism. Given a batch of activations {P1···N}on a gyrogroup {M,⊕}, we denote the GyroBN as GyroBN({Pi};B, s, ϵ, η),(44) where B∈ M and s are biasing and scaling parameters, ϵ is a small positive value, and η is the momentum. Theorem E.2. Given manifolds {M1, g1} and {M2, g2} and a Riemannian isometry f:M1→ M2 , for a batch of activation {P1...N } in M1 , GyroBN1(Pi;B, s, ϵ, γ) in M1 can be calculated as GyroBN1(Pi;B, s, ϵ, γ) = f−1(GyroBN2(f(Pi); f(B), s, ϵ, γ)) ,(45) where GyroBN2is the GyroBN in M2. Proof. This theorem is inspired by Chen et al. (2024b, thm. 5.3), which characterizes the LieBNs under isometric manifolds. As gyrogroups are natural generalizations of Lie groups, our GyroBN is expected to have similar results. The following proof follows a similar logic to the one by Chen et al. (2024b, thm. 5.3), except that all operations are gyro operations. For i= 1,2 , we denote Eq. (15) and binary gyro barycenter on Mi as ξi(·|M, v2, B, s) and Bari η(·,·) . Let B={P1...N }and f(B) = {f(P1...N )}. We only need to show the following: M2=f(M1), v1=v2,(46) ξ1(Pi|M1, v2, B, s) = f−1(ξ2(f(Pi)|M2, v2, f(B), s)),(47) Bar1 η(P, Q) = f−1Bar2 η(f(P), f(Q)),∀P, Q ∈ M1.(48) where Mi and vi are the batch Fréchet mean and variance over Mi for i= 1,2 . Eqs. (46) and (48) can be directly obtained by the invariance of gyrodistance under f (Lem. E.1). We only need to show Eq. (47). We have the following: f−1(ξ2(f(Pi)|M2, v2, f(B), s)) = f−1(f(B)⊕2(t⊙2(⊖2M2⊕2f(Pi))) (1) =f−1◦f(B⊕1(t⊙(⊖1M1⊕1Pi))) =B⊕1(t⊙(⊖1M1⊕1Pi)) =ξ1(Pi|M, v2, B, s), (49) where t=s √v2+ϵ. The above derivation comes from the following. (1) fis an isomorphism preserving gyro operations. As π−1:f Gr(p, n)→Gr(p, n) is a Riemannian isometry, Thm. E.2 indicates that the GyroBN under the projector perspective can be calculated by the ONB perspective by the following process: 1. mapping data into the ONB perspective by π−1:f Gr(p, n)→Gr(p, n); 2. normalizing data by the GyroBN under Gr(p, n); 3. mapping normalized data back to f Gr(p, n)by π. Besides, both Lem. E.1 or Prop. 3.6 can guarantee theoretical control over the gyromean and gyrovariance under the projector perspective. F EXPERIMENTAL DETAILS AND ADDITIONAL ABLATIONS F.1 SUMMARY OF OPERATORS IN GRASSMANNIAN AND HYPERBOLIC GYROBNS The specific implementation of Alg. 1on a gyrogroup can be carried out in a plug-in manner. This involves simply substituting the required operators from Tabs. 2and 8into Alg. 1. To streamline this process, we summarize the discussion in Sec. 6in Tab. 9, where we present all the required operators for the Grassmannian and hyperbolic GyroBNs. 24 Published as a conference paper at ICLR 2025 Table 9: Key operators in calculating GyroBN on the Grassmannian and hyperbolic manifolds. Here P, Q ∈Gr(p, n) are two ONB Grassmannian points, while x, y ∈Pn K are two Poincaré vectors. Other notations follow from Tabs. 2and 8. Operator Gr(p, n)Pn K Identity element Ip,n 0∈Rn P⊕Gr Qor x⊕Kymexp(Ω)V(1−2K⟨x,y⟩−K∥y∥2)x+(1+K∥x∥2)y 1−2K⟨x,y⟩+K2∥x∥2∥y∥2 ⊖GrPor ⊖Kxmexp(−Ω)Ip,n −x t⊙Gr Por t⊙Kxmexp(tΩ)Ip,n 1 √|K|tanh ttanh−1(p|K|∥x∥)x ∥x∥ BarGr γ(Q, P)or BarK γ(y, x) ExpGr P(γLogGr P(Q)) x⊕K(−x⊕Ky)⊙Kt Fréchet Mean Karcher Flow (Karcher,1977) (Lou et al.,2020, Alg. 1) 0 50 100 150 Epoch 50 60 70 80 90 Acc. NTU60 Training GyroGr - Train GyroGrBN - Train 0 50 100 150 Epoch 50 55 60 65 70 75 Acc. NTU60 Testing GyroGr - Test GyroGrBN - Test 0 50 100 150 Epoch 50 60 70 80 90 Acc. NTU120 Training GyroGr - Train GyroGrBN - Train 0 50 100 150 Epoch 40 45 50 55 Acc. NTU120 Testing GyroGr - Test GyroGrBN - Test Figure 4: Training and testing performance on the NTU datasets of 1Block GyroGr. Our BN improves the generalization abilities of GyroGr. F.2 DETAILS ON THE GRASSMANNIAN EXPERIMENTS F.2.1 DATASETS AND PREPROCESSING HDM05 3 (Müller et al.,2007). It consists of 2,273 skeleton-based motion capture sequences executed by different actors. Each frame consists of 3D coordinates of 31 joints. We remove the under-represented clips, trimming the dataset down to 2086 instances scattered throughout 117 classes. Following Nguyen & Yang (2023), we model each sequence as a 93 ×10 Grassmannian matrix. NTU60 (Shahroudy et al.,2016). It has 56,880 sequences of 3D skeleton data classified into 60 classes, where each frame contains the 3D coordinates of 25 or 50 body joints. We use mutual actions and follow the cross-view protocol (Shahroudy et al.,2016). Following Nguyen & Yang (2023), we model each sequence as a 150 ×10 Grassmannian matrix. NTU120 4 (Liu et al.,2019). This dataset contains 114,480 sequences in 120 action classes. We use mutual actions and follow the cross-setup protocol (Liu et al.,2019). Following Nguyen & Yang (2023), we model each sequence as a 150 ×10 Grassmannian matrix. 3https://resources.mpi-inf.mpg.de/HDM05/ 4https://github.com/shahroudy/NTURGB-D 25 Published as a conference paper at ICLR 2025 G.3 PROOF OF THM.3.5 Proof of Thm. 3.5.Given any P, Q, R ∈G, we make the following proof. Gyroisometry of the left gyrotranslation: This property generalizes Thms. 2.12 and 2.16 by Nguyen & Yang (2023), which deal with the gyrotranslations in the SPD and Grassmannian manifolds. We have the following: dgry(LP(Q), LP(R)) = dgry(P⊕Q, P ⊕R), =∥⊖(P⊕Q)⊕(P⊕R)∥gr , =∥gyr[P, Q] (⊖Q⊕R)∥gr (left gyrotranslation), =∥⊖Q⊕R∥gr (gyroisometry of the automorphism), = dgry(Q, R). (63) Symmetry of the gyrodistance: dgry(P, Q) = ∥⊖P⊕Q∥gr , =∥⊖(⊖P⊕Q)∥gr (Eq. (5) and Eq. (7)), =∥gyr[⊖P, Q](⊖Q⊕P)∥gr (gyrosum inversion law), =∥⊖Q⊕P∥gr (gyroisometry of the automorphism), = dgry(Q, P). (64) Gyroisometry of the gyroinverse: dgry(⊖P, ⊖Q) =∥P⊖Q∥gr , =∥⊖Q⊕P∥gr ( gyrocommutativity and gyroisometry of the automorphism), = dgry(Q, P), = dgry(P, Q)(Symmetry of the gyrodistance). (65) G.4 PROOF OF PROP.3.6 Proof of Prop. 3.6. We first show the expressions of the gyrodistance on M . Then we proceed to show the gyrodistance and gyroisometries on MK . We follow all the notations in Tab. 2and denote the geodesic distance under a specific geometry as dgeo. Expressions of gyrodistance on M: For {Sn ++,⊕AI}, we have the following: dgry(P, Q) = P−1 2QP−1 2gr , =mlog(P−1 2QP−1 2)(LogI= mlog), = dgeo(P, Q). (66) For {Sn ++,⊕LE}, we have the following: dgry(P, Q) = ∥mexp(−mlog(P) + mlog(Q))∥gr , =mlog−1 ∗,I(−mlog(P) + mlog(Q))I,(LogI(P) = mlog−1 ∗,I(mlog(P))), =∥mlog(P)−mlog(Q)∥,(the pullback of LEM (Chen et al.,2024e)) = dgeo(P, Q), (67) where mlog−1 ∗,I is the inverse of the differential map of mlog , and ∥·∥I is the norm induced by the LEM at I. 32 Published as a conference paper at ICLR 2025 As LCM is also a pullback metric (Chen et al.,2024e), {Sn ++,⊕LC}follow the same logic: dgry(P, Q) = ψ−1 LC(−ψLC(P) + ψLC(Q))gr , =(ψLC)−1 ∗,I(−ψLC(P) + ψLC(Q))I(LogI(P)=(ψLC)−1 ∗,IψLC(P)), =∥ψLC(P)−ψLC(Q)∥,(the pullback of LCM (Chen et al.,2024e)) = dgeo(P, Q). (68) For {f Gr(p, n),e ⊕Gr}, we have the following: dgry(P, Q) = mexp(−[P , e Ip,n])Qmexp([P, e Ip,n])gr (⊖P=−P), =e P⊤Qe Pgr (e P= mexp([P, e Ip,n]) ∈SO(n)), =1 2Loge Ip,n e P⊤Qe P(∥·∥I=1 2∥·∥), =1 2[Ω,e Ip,n], (69) where [Ω,e Ip,n] = Loge Ip,n e P⊤Qe P. For Ω, we have the following: Ω = 1 2mlog In−2e P⊤Qe PIn−2e Ip,n((Bendokat et al.,2024, Prop. 5.6)), =1 2mlog e P⊤(In−2Q)e Pe P⊤In−2e Pe Ip,n e P⊤e P =1 2mlog e P⊤(In−2Q)In−2e Pe Ip,n e P⊤e P, (1) =e P⊤1 2mlog (In−2Q)In−2e Pe Ip,n e P⊤e P, (2) =e P⊤1 2mlog ((In−2Q) (In−2P))e P, (3) =e P⊤Ωe P, (70) Eq. (70) follows from the following facts: (1) When mlog(B) is well-defined and A is non-singular, A−1(log B)A= log A−1BA (Horn & Johnson,2012). (2) P= mexp([P, e Ip,n])e Ip,n mexp(−[P, e Ip,n]) (Nguyen,2022a, Eq. (36)). (3) Let Ω = 1 2mlog ((In−2Q) (In−2P)) Combining Eqs. (69) and (70), we have dgry(P, Q) = 1 2[Ω,e Ip,n], =1 2[e P⊤Ωe P, e Ip,n], (1) =1 2e P⊤[Ω,e Pe Ip,n e P⊤]e P, (2) =1 2[Ω,e Pe Ip,n e P⊤], (3) =1 2[Ω, P], (4) =1 2∥LogP(Q)∥, (5) =∥LogP(Q)∥P, (6) = dgeo(P, Q). (71) 33 Published as a conference paper at ICLR 2025 Eq. (71) follows from the following facts: (1) [e P⊤Ωe P, e Ip,n] = e P⊤Ωe Pe Ip,n −e Ip,n e P⊤Ωe P, =e P⊤Ωe Pe Ip,n e P⊤−e Pe Ip,n e P⊤Ωe P, =e P⊤[Ω,e Pe Ip,n e P⊤]e P. (72) (2) Euclidean norm is invariant under the action A7→ OAO⊤,∀A∈Rn×n, O ∈O(n). (3) P= mexp([P, e Ip,n])e Ip,n mexp(−[P, e Ip,n]). (4) LogP(Q) = [Ω, P ]. (5) ∥·∥P=1 2∥·∥,∀P∈f Gr(p, n). (6) dgeo(P, Q) = ∥LogP(Q)∥Pfor any P, Q ∈f Gr(p, n)not in each other’s cut locus. For {Gr(p, n),⊕Gr} , we first make the following notations: We denote the geodesic distance, gyrodistance, Riemannian logarithm at Ip,n , and Riemannian metric at Ip,n on Gr(p, n) as dgeo , dgry , LogIp,n , and ∥·∥ONB Ip,n . The counterparts on f Gr(p, n) are g dgeo , g dgry , g Loge Ip,n , and ∥·∥PP e Ip,n . As shown by Nguyen et al. (2024, App. N), π: Gr(p, n)→f Gr(p, n) is a Riemannian isometry. Then, for any U, V ∈Gr(p, n)with π(U) = Pand π(V) = Q, we have the following: dgry(U, V ) = LogIp,n ⊖GrU⊕Gr V ONB Ip,n (1) =LogIp,n π−1(e ⊖GrPe ⊕GrQ) ONB Ip,n (2) =π−1 ∗,Ip,n g Loge Ip,n e ⊖GrPe ⊕GrQ ONB Ip,n (3) =g Loge Ip,n e ⊖GrPe ⊕GrQ PP e Ip,n =g dgry(P, Q) =g dgeo(P, Q) (4) = dgeo(U, V ) (73) The above comes from the following. (1) Gyro additions under the Riemannian isometry (Nguyen & Yang,2023, Lem. 2.1). (2,3,4) Riemannian operators under the Riemannian isometry (Gallier & Quaintance,2020). Constant Curvature Spaces: When MK=Rn(K= 0) , the gyro structures defined in Eq. (3)-Eq. (9) are reduced to the vector structures. The claim can be directly proved. In the following, we present the proof for K= 0 . We first show the expression for the gyrodistance under MK , then the isometry of gyration, and finally the results on the gyroisometry of gyrotranslation and gyroinverse. In the following, a, b, x, y are arbitrary points in MK. Gyrodistance and geodesic distance: The Riemannian metric, logarithm, and geodesic distance on the CCS (Skopek et al.,2019) are gK x=λK x2gE,(74) LogK x(y) = 2 p|K|λK x tan−1 Kp|K|∥−x⊕Ky∥−x⊕Ky ∥−x⊕Ky∥,(75) dgeo(x, y) = 2 p|K|tan−1 Kp|K|∥−x⊕y∥,(76) 34 Published as a conference paper at ICLR 2025 where λK x=2 1+K∥x∥2 , gE is the standard Euclidean inner product, and ⊕K is the gyro addition in Tab. 2. Especially, when x= 0, we have gK 0= 22gE,(77) LogK 0(y) = 1 p|K|tan−1 Kp|K|∥y∥y ∥y∥(78) For the gyrodistance, we have the following: dgry(x, y) = ∥−x⊕y∥gr , =∥Log0(−x⊕y)∥0, = 2  1 p|K|tan−1 Kp|K|∥−x⊕y∥−x⊕y ∥−x⊕y∥, =2 p|K|tan−1 Kp|K|∥−x⊕y∥(∀s > 0,tan−1 K(s)>0), (79) where ∥·∥0is the norm induced by gK 0. Norm invariance under gyration: As MK forms a real inner product gyrovector spaces (Ungar, 2009, Def. 3.2), any gyration on CCS preserves the norm induced by standard inner product: ∥gyr[a, b](x)∥=∥x∥,∀x. (80) For the gyronorm, we have the following: ∥gyr[a, b]x∥gr = 2∥Log0(gyr[a, b]x)∥0, =2 p|K|tan−1 Kp|K|∥gyr[a, b]x∥, =2 p|K|tan−1 Kp|K|∥x∥, =∥x∥gr . (81) Isometry of left gyrotranslation and gyroinverse: Note that MK forms a gyrocommutative gyrogroup. According to Thm. 3.5, we can obtain the results. G.5 PROOF OF THM.4.1 Proof of Thm. 4.1. According to Thm. 3.3 and Thm. 3.5, any left gyrotranslation is a gyroisometry. Therefore, for any Q∈ M, we have the following: dgry (B⊕Pi, Q)(1) = dgry (⊖B⊕(B⊕Pi),⊖B⊕Q) (2) = dgry ((⊖B⊕B)⊕gyr[⊖B, B](Pi)),⊖B⊕Q) (3) = dgry (Pi,⊖B⊕Q). (82) The above comes from the following. (1) Any left gyrotranslation is a gyroisometry. (2) Left gyroassociative law. (3) ⊖B⊕B=Eand pseudo-reduction. Denoting the gyromean of {Pi}and {B⊕Pi}as Mand f M, we have the following: B⊕M(1) =B⊕(⊖B⊕f M) (2) = gyr[B, ⊖B](f M) (3) =f M. (83) The above comes from the following. 35 Published as a conference paper at ICLR 2025 (1) Eq. (82) indicates that M=⊖B⊕f M. (2) Left gyroassociative law. (3) Pseudo-reduction. Now, we proceed to deal with the second property. We have the following: dgry(t⊙Pi, E)(1) =∥⊖E⊕(t⊙Pi)∥gr (2) =∥t⊙Pi∥gr =∥tLogEPi∥E =|t|∥LogEPi∥E =|t|∥Pi∥gr (3) =|t|∥⊖E⊕Pi∥gr =|t|dgry(E, Pi) (4) =|t|dgry(Pi, E) (84) The above follows from the following. (1) Symmetry of gyrodistance (Thm. 3.5). (2) ⊖E=E. (3) Pi=⊖E⊕Pi. (4) Symmetry of gyrodistance. The last equation in Eq. (84) indicates the homogeneity of dispersion from E. G.6 PROOF OF PROP.5.1 Proof of Prop. 5.1. We only need to prove the equivalence of gyrodistance and geodesic distance. Note that every Lie group is automatically a gyrogroup with each gyration as the identity map. We denote {M,⊕, g} as a Lie group with left-invariant metric g . For any P and Q in M , we have the following: dgry(P, Q)(1) =∥⊖P⊕Q∥gr (2) =∥LogE(⊖P⊕Q)∥E = dgeo(E, ⊖P⊕Q) (3) = dgeo(P, P ⊕(⊖P⊕Q)) (4) = dgeo(P, Q). (85) The above derivation comes from the following. 1. Definition of gyrodistance Eq. (9). 2. Definition of gyronorm Eq. (8). 3. Under a left-invariant metric, any left Lie group translation is a Riemannian isometry. 4. P⊕(⊖P⊕Q)=(P⊖P)⊕Q( the associative of group addition) =E⊕Q =Q. (86) Therefore, the gyromean and gyrovariance are exactly the Fréchet mean and variance under the geodesic distance, while the running mean updates are also identical under gyrodistance and geodesic distance. 36 Published as a conference paper at ICLR 2025 G.7 PROOF OF PROP.6.1 We first review a fast and stable algorithm for the ONB Grassmannian logarithm (Bendokat et al., 2024, Alg. 5.3), and the calculation of Grassmannian logarithm under the projector perspective by the ONB Grassmannian logarithm (Nguyen et al.,2024, Prop. 3.12). Algorithm 2: Grassmann logarithm under the ONB perspective (Bendokat et al.,2024, Alg. 5.3) Input: U, Y ∈Gr(p, n)are Stiefle representatives under ONB perspective. 1QSRTSVD := YTUwith Sin ascending order, and Qand Rcolumn-wisely flipped accordingly; 2ˆ S=√In−S2; 3∆=(In−UU⊤)Y Qarcsin( ˆ S) ˆ SRT; Output: LogU(Y)=∆ Alg. 2reviews a fast and stable algorithm for the Grassmannian Riemannian logarithm under the ONB perspective Gr(p, n) . The vanilla Riemannian logarithm in Tab. 8requires an n×p SVD and a p×p matrix inverse, while Alg. 2only requires an p×p SVD. Therefore, Alg. 2is more efficient than the vanilla logarithm. Besides, Alg. 2can also return a unique tangent vector when Y is in the cut locus of U. For more details, please refer to Bendokat et al. (2024, Sec. 5.2). As the projector perspective is isometric to the ONB perspective, the Grassmannian logarithm under the projector perspective can be calculated by the ONB Grassmannian logarithm (Nguyen et al.,2024, Prop. 3.12). Proposition G.2 ((Nguyen et al.,2024)).Given any P, Q ∈f Gr(p, n) with U=π−1(P) and V=π−1(Q), the Riemannian logarithm g LogP(Q)on f Gr(p, n)is given as g LogP(Q) = π∗,U (LogUV),(87) where Log is the Riemannian logarithm under the ONB perspective, π∗,U :TUGr(p, n)→ TPf Gr(p, n)is the differential map of πat U, which is defined as π∗,U (∆) = ∆U⊤+U∆⊤,∀∆∈TUGr(p, n).(88) Now, we begin to present the proof. Proof of Prop. 6.1.We first show the expression for LogIp,n and g Loge Ip,n . First note the following: (In−Ip,nI⊤ p,n) = 0 0 0In−p,(89) U⊤Ip,n =U⊤ 1, U⊤ 2Ip 0 =U⊤ 1, (90) By the above two equations, the ONB Grassmannian logarithm at Ip,n is LogIp,n (U) = 0 0 0In−pU1 U2Qarcsin( ˆ S) ˆ SRT(Alg. 2) = 0 U2Qarcsin( ˆ S) ˆ SRT! =0 e U2, (91) where QSRTSVD := U⊤ 1 with S in ascending order, and Q and R column-wisely flipped accordingly, and ˆ S=√In−S2. 37 Published as a conference paper at ICLR 2025 For g Loge Ip,n , we have g Loge Ip,n (UU⊤)(1) =π∗,Ip,n LogIp,n (U) (2) =π∗,Ip,n  0 e U2 (3) =0e U⊤ 2 e U20. (92) The above derivation comes from the following. (1) Prop. G.2 (2) Eq. (91) (3) For any ∆ = (∆⊤ 1,∆⊤ 2)⊤∈TIp,n Gr(p, n), where ∆1is p×p, we have the following π∗,Ip,n ∆1 ∆2=∆1 ∆2(Ip,0) + Ip 0∆⊤ 1,∆⊤ 2 =∆10 ∆20+∆⊤ 1∆⊤ 2 0 0  =∆1+ ∆⊤ 1∆⊤ 2 ∆20. (93) Combining all the above results together, we have the following: [UU⊤,e Ip,n] = hg Loge Ip,n (UU⊤),e Ip,ni = 0e U⊤ 2 e U20,e Ip,n =0e U⊤ 2 e U20Ip0 0 0 −Ip0 0 0  0e U⊤ 2 e U20 =0−e UT 2 e UT 20. (94) 38