scieee AI-readable full text Open interactive document viewer

Isotropic Deep Learning: You Should Consider Your (Foundational) Biases

Bird, George

Abstract

This work was originally written per the requirements of the 2025 NeurIPS Position Paper track, which included bold meta-level arguments for what the field is doing right and wrong. Therefore, the paper's rhetoric aimed to present a strong position and a title which reflects this. Abstract: This position paper explores an alternative mathematical formulation, `Isotropic Deep Learning', by analysing the implications of current functional forms in deep learning. Modern networks almost universally rely on foundational forms respecting discrete permutation symmetry. However, this is an underappreciated choice in form, argued to introduce unrecognised biases without suitable alternatives. Initially, this discrete symmetry observation is promoted to a continuous rotation defined framework, then broadened to primitive sets defined by various other symmetries. This constitutes a new symmetry-led design-axis: rather than enforcing it through model design, which transfers symmetry through the structure, it studies how foundational form symmetries inherently act on and interact within general architectures --- one objective is a systematic approach to the consequences of network symmetry breaking in addition to symmetry making and emerging from the primitive-level. In addition, determining whether non-trivial expressibility is contingent on which function symmetries are preserved moreover broken. The goal is to expose and leverage unintended biases by deducing principles applicable in broader contexts for beneficial computation. Proposed is a systematic reformulation of all foundational primitives into classes that respect particular groups, and to determine the resultant implications. This constitutes an inverted ontology framework where general symmetries are situated definitionally prior to neurons, rather than a permutation symmetry being deduced from them. This design axis motivates reselection of compositions upwards, as they underpin current constructions and may enable new models contingent on alternative foundations. Hence, the paper advocates for a distinctly bottom-up reformulation aiming to deduce general principles for broad leverage. This is motivated by prior work demonstrating that current functional forms influence activation distributions: discrete symmetries in functions induce similar discrete structure in embedded representations through training. Thus, geometric artefacts can arise in learned representations solely due to human-imposed design choices rather than task-driven necessity. Therefore, the prevailing choice is shown to carry unappreciated and unintended task-agnostic biases. Moreover, there appears to be no compelling a priori justification for why such representations or functional forms are universally desirable; this paper hypothesises three testable pathologies of the current formulation with significant connections to mechanistic interpretability. Hence, this motivates the construction and analysis of alternative foundational primitives, aiming to alter geometric constraints on representations and improve performance. The underlying inductive biases of the isotropic approach may constitute a preferable default which could be adopted if a wide array of suitable and well-performing functions are developed. A variety of preliminary functions are proposed, including new activation functions, normalisers, and operations, and an audit is provided across various primitives in use. The symmetry-principled construction is then generalised, enabling a broad class of group-defined reformulations across primitives, positing a new foundational design axis with distinct inductive biases. Thus, Isotropic Deep Learning becomes just one case study among such parallel implementations for all models. This initial group-theoretic generalisation of primitives is systematically extended upwards to encompass their hierarchical compositions, motivating its applicability across all scales in architectures. This yields an initial three generations of symmetry strength for categorisation in the framework. This extension recovers Geometric Deep Learning as the strongest generation when composing functions, ensuring model-scale compliance with the symmetry constraint derived from data for specialist applications. This substantially contrasts with this paper's enquiry, diverging through a bottom-up philosophy starting from primitives rather than working recursively top-down from model constraints. This establishes a further role for symmetry emerging within deep learning. This taxonomic formalism also encompasses the Parameter Symmetry approach as a distinct compositional case, studying the consequences of computational equivalences under reparameterisations deduced from current permutation-like primitives. In contrast, this work redefines primitives through a symmetry-led design-axis and investigating the ramifications more broadly, not just restricted to discrete parameter degeneracies. Hence, this ``Taxonomic Deep Learning'' approach reveals all three to be distinct special cases, characteristic of various compositional scales and strengths --- a unification of contemporary approaches to symmetry in an intuitive, hierarchical, and complementary formalism. This may facilitate a better, comprehensive comparison and exploration of their interplay, while clarifying further regimes that may remain to be considered. Encouraged is a systematic audit into the influence of symmetry generally, but particularly the reformulation and comparison of various group-defined primitive sets. From this, the study of downstream phenomena can proceed after a primitive algebra is fixed. This can span from determining representation biases, reassessing theorems contingent on prior primitives, optimisation, performance, and diverse new model architectures. Zenodo will be used to maintain a continuously updated copy of this work as it evolves. Please make sure that any share link used connects to this general page, rather than a specific version of the paper. Changelog: Changed abstract and aspects of introduction. This is now considered the finalised version of the paper's content. Formatting/grammatical errors may be corrected but the content will be stable moving forwards.

Full text

Isotropic Deep Learning: You Should Consider Your (Foundational) Biases George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] May 1, 2025 Abstract This position paper explores an alternative mathematical formulation, ‘Isotropic Deep Learning’, by analysing the implications of current functional forms in deep learning. Modern networks almost universally rely on foundational forms respecting discrete permutation symmetry. However, this is an underappreciated choice in form, argued to introduce unrecognised biases without suitable alternatives. Initially, this discrete symmetry observation is promoted to a continuous rotation defined framework, then broadened to primitive sets defined by various other symmetries. This constitutes a new symmetry-led design-axis: rather than enforcing it through model design, which transfers symmetry through the structure, it studies how foundational form symmetries inherently act on and interact within general architectures — one objective is a systematic approach to the consequences of network symmetry breaking in addition to symmetry making and emerging from the primitive-level. In addition, determining whether non-trivial expressibility is contingent on which function symmetries are preserved moreover broken. The goal is to expose and leverage unintended biases by deducing principles applicable in broader contexts for beneficial computation. Proposed is a systematic reformulation of all foundational primitives into classes that respect particular groups, and to determine the resultant implications. This constitutes an inverted ontology framework where general symmetries are situated definitionally prior to neurons, rather than a permutation symmetry being deduced from them. This design axis motivates reselection of compositions upwards, as they underpin current constructions and may enable new models contingent on alternative foundations. Hence, the paper advocates for a distinctly bottom-up reformulation aiming to deduce general principles for broad leverage. This is motivated by prior work demonstrating that current functional forms influence activation distributions: discrete symmetries in functions induce similar discrete structure in embedded representations through training. Thus, geometric artefacts can arise in learned representations solely due to human-imposed design choices rather than task-driven necessity. Therefore, the prevailing choice is shown to carry unappreciated and unintended task-agnostic biases. Moreover, there appears to be no compelling a priori justification for why such representations or functional forms are universally desirable; this paper hypothesises three testable pathologies of the current formulation with significant connections to mechanistic interpretability. Hence, this motivates the construction and analysis of alternative foundational primitives, aiming to alter geometric constraints on representations and improve performance. The underlying inductive biases of the isotropic approach may constitute a preferable default which could be adopted if a wide array of suitable and well-performing functions are developed. A variety of preliminary functions are proposed, including new activation functions, normalisers, and operations, and an audit is provided across various primitives in use. The symmetry-principled construction is then generalised, enabling a broad class of group-defined reformulations across primitives, positing a new foundational design axis with distinct inductive biases. Thus, Isotropic Deep Learning becomes just one case study among such parallel implementations for all models. This initial group-theoretic generalisation of primitives is systematically extended upwards to encompass their hierarchical compositions, motivating its applicability across all scales in architectures. This yields an initial three generations of symmetry strength for categorisation in the framework. This extension recovers Geometric Deep Learning as the strongest generation when composing functions, ensuring model-scale compliance with the symmetry constraint derived from data for specialist applications. This substantially contrasts with this paper’s enquiry, diverging through a bottom-up philosophy starting from primitives rather than working recursively top-down from model constraints. This establishes a further role for symmetry emerging within deep learning. This taxonomic formalism also encompasses the Parameter Symmetry approach as a distinct compositional case, studying the consequences of computational equivalences under reparameterisations deduced from current permutation-like primitives. In contrast, this work redefines primitives through a symmetry-led design-axis and investigating the ramifications more broadly, not just restricted to discrete parameter degeneracies. Hence, this “Taxonomic Deep Learning” approach reveals all three to be distinct special cases, characteristic of various compositional scales and strengths — a unification of contemporary approaches to symmetry in an intuitive, hierarchical, and complementary formalism. This may facilitate a better, comprehensive comparison and exploration of their interplay, while clarifying further regimes that may remain to be considered. Encouraged is a systematic audit into the influence of symmetry generally, but particularly the reformulation and comparison of various group-defined primitive sets. From this, the study of downstream phenomena can proceed after a primitive algebra is fixed. This can span from determining representation biases, reassessing theorems contingent on prior primitives, optimisation, performance, and diverse new model architectures. 1 1 Introduction Elementwise functional forms singularly dominate contemporary deep learning [ 1 , 2 , 3 , 4 , 5 , 6 , 7 , 8 , 9 , 10 ]. This is particularly evident in, but not limited to, activation functions sometimes referred to as ‘ridge’ [ 11 ] activation functions. Activation functions are often displayed in univariate form [ 12 , 13 , 14 ], generally characterised by the form shown in Eqn. 1, with σ being a placeholder activation function, e.g., ReLU (f(x) = max (0, x)) [15], Tanh (f(x) = tanh (x)), etc. f:R→R, x 7→ f(x) = σ(x)(1) However, this display choice obfuscates a crucial (standard) basis dependence. This dependence is made explicit in Eqn. 2, which displays the multivariate functional form of a given activation function. This should be considered a more implementation-faithful form 1 . This reveals the functional form’s usually hidden ˆei basis dependence. The multivariate form is depicted for an n -neuron layer, with activation vector x ∈Rn . This standard basis dependence is arbitrary and appears to be largely a historical precedent, rather than a problem-aligned, intentional inductive bias. This is discussed further in App. F. f:Rn→Rn, x 7→ f(x) = n X i=1 σ(x ·ˆei) ˆei(2) Due to this basis dependence, non-linear transformations differ angularly in effect [ 16 , 17 ]. Therefore, this will be termed an anisotropic function, indicating this rotational asymmetry. Particularly, it could be termed a standard-anisotropic function, indicating its dependence on the standard basis. Due to the pervasive use of these functional forms, including activation functions, normalisers, initialisers, regularisers, optimisers, architectures, operations, and gradient clipping, amongst others, contemporary deep learning as a paradigm may consequently be termed a form of ‘anisotropic deep learning’. Despite its implications and prevalence, this choice of basis-dependent anisotropic form appears underappreciated and incidental in the development of most contemporary models. Anisotropic forms have largely become an unquestioned default, approaching an axiomatic-like definition for the general field rather than a considered choice. Hence, re-evaluating their impact and then systematically reformulating this foundational aspect of modern deep learning, with potentially wide-reaching consequences, is suggested to constitute distinct approaches, such as ‘Isotropic Deep Learning’ — effectively based upon differing foundational, axiomatic-like, primitive definitions. This is emphasised by such choices underlying all downstream compositions, including those in models. It should be determined whether their respective phenomena, theoretical results, and various consequences are contingent upon these foundational choices. It is arguably the general composition of such primitives, including parameterised maps, which defines the current ontology of deep learning — contrasting it with other machine learning approaches. The asymmetry in current non-linear transforms is usually about the standard (Kronecker) basis vectors, and frequently their negative, {+ˆei,−ˆei} ∀i∈[0,1,··· , n] for n width layers, and these are equal in privilege. This is often due to their elementwise application. Therefore, it can be said to distinguish the standard basis — a ‘distinguished basis’ 2 . This basis-dependence is often overlooked in consequence, and this work argues that it acts as an implicit inductive bias for representational geometry; therefore, it must be evaluated. For example, the standard basis’ activation space distortions are visible in Fig. 1 showing the mapping of elementwise-tanh on a variety of test shapes. Crucially, one can consider this choice of functional form to break a continuous rotational-like symmetry, and reduce it to a discrete rotational-like symmetry (alongside some specific mirrors). The latter is specifically referred to as a permutation ( Sn ) symmetry of the standard basis, and the former is an orthogonal ( O(n) ) symmetry about the origin. In effect, if a function is treated in its multivariate form, in the current formulation of deep learning, it is equivariant to a permutation of the components of its vector decomposed in the standard basis — this also defines the notion and individuality of neurons. For (a representation of) an element of the permutation group, notated in shorthand by P∈ Sn , the following equivariance relation holds: f(Px) = Pf(x) . However, permutation symmetry is a subgroup of the orthogonal symmetry proposed: Sn⊂O (n) — therefore, this discrete permutation symmetry could be considered a broken continuous orthogonal symmetry 3 . Further details on the categorisation and nuance of such symmetries are discussed further in Sec. 5.2. Non-linearities are usually pivotal to the network’s ability to achieve a desired computation, as seen through the universal approximation theorem’s [ 18 ] explicit dependence on the form of the activation function [ 19 ]. Non-linearities produce differing local transformations, such as stretching, compressing, and generally reshaping a manifold — displayed elegantly in Olah [20]4 . Consequently, the network may be expected to adapt by moving representations to geometries about these distinguished directions, using specific localised mappings to achieve the desired computation, discussed further in Sec. 2.2. Hence, an anisotropy about distinguished vectors may be expected to be induced into the activation distribution, through optimisation in general networks. This anisotropic inductive bias on representations appears to be reinforced directly through most functional forms, indirectly through many optimisers, and is inherent in the connectivities of many architectures. Hence, it seems largely systematic to contemporary deep learning through each aspect of this triad — all of which typically share the same underlying and characterising permutation symmetry. Each is hypothesised to contribute to such representational structures. 1Softmax has an extra denominator term, but still displays the basis-dependent nature of elementwise forms. 2 This is suggested as a generalisation from a ‘privileged basis’ discussed in Elhage et al. [16] .‘Distinguished directions’ may better reflect how representations can be encouraged or deterred in alignment about these directions; whereas ‘privileged basis’ would suggest a greater preference for alignment. The term ’basis’ will often be retained despite the set of ‘distinguished vectors’ potentially being under-/over complete for spanning the activation space, as demonstrated by Bird [17]. 3 The mirror transform in the group O (n) appear to result in no change in (single-argument) functional forms from those defined through the pure rotation group SO (n). Hence, O (n)is used since Isotropic deep learning automatically respects this larger group. 4Olah also stated disillusionment with the elementwise form for other reasons in this article. 2 Figure 1: Left shows a 2-dimensional plane, R2 , populated with various shapes: black concentric circles, green lines through the origin, red parallel lines and in faint black the standard (cartesian) coordinate axes ˆe1 and ˆe2 (which remain untransformed). If this space is then imaged through elementwisetanh , the individual pointwise coordinates making up the shapes are passed through the standard tanh activation. The resultant shapes are shown in the centre plot. The rightmost plot is similar to the centre plot, but for the so-called isotropictanh presented in Sec: 3.1. One can see that the objects in the centre plot are distorted around the basis directions, whilst in the right-most plot, they are not distorted due to the basis directions. For example, the green lines are significantly curved towards the corners of the boundary. An interactive demonstration of these functions is available here. The direct effect has been empirically demonstrated in activation functions [ 17 ]: training results in the discrete symmetry of the functional forms, inducing a broken symmetry in the activations which transforms with transformations to the distinguished directions of the form. Since these non-linear zones are centred around the distinguished directions, the embedded representations are expected to adopt advantageous angular arrangements with respect to the arbitrarily imposed geometry of the distinguished basis. For example, they appear to move towards the non-linearities’ extremums, aligned, anti-aligned, or other geometries [ 16 ], through training [ 17 ]. This may correspond to a local, dense, sparse coding [ 21 ] or superposition [ 16 ], respectively. This indicates that such directions mark out an absolute reference structure about which representations are observably shaped. Therefore, the network has adapted its representations through optimisation due to the properties of the foundational functional form choices present. These general representational biases are entirely distinct considerations from enforcing a specific end-to-end symmetry in a network. In such cases, the form is leveraged to preserve a data structure through the network for a targeted application, where representations predictably transform or remain unchanged with respect to the group. In these circumstances, the symmetries do not ‘act’ on the network internally as a bias in the general manner being suggested in this work — it is these latter considerations that are argued to be unappreciated and unintentional, in universal settings. Thus, a differing approach to symmetry’s role in deep learning: a task-driven versus a function-driven approach. Due to this separate motivation, this formalism is a tangent line of enquiry, and the group-theoretic reasoning emerged independently in response. It is argued to be an important consideration, regardless of the underlying data structure, and hence its applicability is considered general. Nevertheless, such approaches both draw on group-theoretic roots and can be unified under a formalism presented in Sec. 5.2, which may enable beneficial cross-considerations and sharing of tooling at times. Moreover, this function-driven causal hypothesis aids in explaining the observed tendency of distinguished-direction alignment. This is the hypothesis underlying the encouraged position: functional forms should be deliberate and carefully considered design choices, with a suitably optimal and minimally harmful default since they can induce a representational structure not required by the task. Currently, anisotropic primitives appear to induce a human-imposed representational collapse onto the distinguished directions. Hence, this was shown to frequently not be a task-necessitated collapse, but instead a task-agnostic structure induced by function primitives. There appears to be little justification for why this particular form and induced structure is universally desirable, with several key negative implications predicted in Sec. 2. Without a priori justification, this inductive bias may be detrimental to computation; therefore, unconstraining the activation is argued to be generally preferable. In addition, the added structure into functional forms, which produces these distinguished directions, may be considered a needless additional assumption for some applications applying deep-learning models. Throughout the rest of this paper, it is argued that a departure from this anisotropic functional form paradigm towards the isotropic reformulation may be generally preferable as a universal inductive bias, unless supported by task-aligned justification for differing primitive algebra. This paper encourages consideration of these choices when designing a model generally, alongside the usual architectural toolkit. In particular, isotropic choices, initially led by a basis independence principle, are argued to unconstrain the representations into more optimal arrangements for general tasks and architectures, free from imposed discrete structure. Some instances where isotropy may be particularly beneficial, such as the amendments discussed for self-attention, are discussed in App. D. However, the development of functions, then models, which suitably leverage isotropy may require substantial time to parallel the existing approach to deep learning in empirical results, since they are based upon a fundamentally differing foundational set of primitives. These will then require downstream verification of the argued optimality. Additionally, many contemporary architectures may be a product of selection over anisotropic primitives and anisotropic benchmarks, whose intrinsic anisotropies may be indicative of. Such outcomes of selection may not pertain 3 to isotropically defined primitives or benchmarks (or other taxonomies). Therefore, reselection of architectures from a clean-slate foundational approach may be necessitated, but generally productive ideas could be analogised. Hence, due to these considerations, such reformulations may require considerable development before they mature into practical implementations. The understanding of such phenomena, and the potential resultant impact of this, is argued to be a worthwhile exploratory avenue. The most fundamental addition of this work is that these inductive bias considerations motivate a broader symmetryunifying construction for functional forms, discussed in Sec. 5. This produces a taxonomic class for deep learning extending bottom-up from sets of primitives defined through their group structures. Isotropic and Anisotropic deep learning represent just two among the array of possible groups capable of generating functional form branches using the tools presented — a seemingly rich and unexplored proposal. Exploring such alternatives may offer better optimised functional forms beyond those discussed in this paper. Thus, the specific case study of Isotropic deep learning should not detract from the wider scope of primitives defined over the broader symmetry taxonomy and general groups. This taxonomic approach extends a group-theoretic formalism for considering and imposing symmetry constraints on functional forms, thereby organising them under distinct branches, each generating a complete set of primitives. This taxonomy is organised by group criteria across three degree-tiers/generations and three flavours, discussed further in Sec. 5.2. Through these group-defined sets, one can choose which primitive class to implement. This is argued to approach an axiom-like choice in ramifications, since it is these primitives which are later composed into all downstream models. Hence, the primitive selected occupies an underivable base choice preceding any model design, and its group-defined form is usually assumed instead of a leveraged design-axis. Therefore, reformulation requires a broad reevaluation, extending from the reselection of models following from primitive changes, investigating each’s resultant emergent phenomena, as well as the reanalysis of theorems predicated on the current form for primitive, among numerous further implications. Comparison between choices may be advantageous in yielding more fundamental insights into the innate properties of deep learning. Such sets of primitives can be constructed a priori for general applications; yet, the taxonomy also indicates how arbitrary graphs can inherently break such symmetries. This symmetry-breaking is argued to be both useful, if influences are leveraged correctly, or perhaps detrimental, if it occurs haphazardly; therefore, establishing such a link is critical, and this paper advocates for it being a careful design choice. This connection is achieved through evaluating automorphisms of an arbitrary graph structure. Typically, one would then set the functional form symmetry constraints as a subset of the available automorphisms to fix a primitive class from those available 5 . Hence, generally, forms are applied which remain unbroken by the inherent connectivity. This, in turn, would specify how to typically apply the primitive form, such as channel-wise for convolution, since these connectivities are not inherently symmetry broken by the graph’s connectivities. This motivates the expanded set of generation strengths to which group-theoretic constraints can apply. ‘Closures’ are the broadest functional class, which would typically be chosen and derived from automorphisms of an arbitrary graph and then can be selectively elevated to stronger generations in design for narrower classes. Geometric deep learning’s considerations arise when a single group is elevated to the strongest level network-wide, and hence provides a restriction on the subset of architectures and functions which can provide these initial closures. However, it is argued that awareness of the generalised group-theoretic choices and their consequences may be beneficial in universal settings. Due to the argued pervasive applicability of such a symmetry-formalism approach across foundational primitives, the terms "branch" or "fork" are used, e.g. "Isotropic-fork". These are felt to be appropriate descriptors for the groups of distinct forms produced, as well as downstream architectures, theorems, and phenomena contingent upon them. This indicates that the use of graph-based computation, from the continued use of linear algebra, is preserved; however, all intermediate functions have parallel implementations that respect their new chosen symmetries. These alternative classes of primitives would typically diverge substantially from the form of contemporary functions, and likely their respective models and consequences — warranting a differing subclassification system. Encouraged are differing semi-autonomous subdisciplines of exploration to determine the implications and leveragability of each, a systematic analysis revealing how they may incur different biases through their various reformulations. However, the consistent use of linear algebra and a primitive in a layered structure makes it appropriate to continue grouping them under the "deep learning" heading, rather than a distinct machine learning approach. Nevertheless, it is argued that in most other meaningful ways, the forks may be largely distinct, likely preferring differing architectures, applications, and resultant phenomena (such as interpretability consequences). This highlights a potentially broader ontology for what may be constitute a deep learning system, where these taxonomically-organised and underpinning choices offer axiomatic-like branching within the field. This axiomatic-like nature is underscored by each symmetry functional form requiring a respective Universal Approximation Theorem due to the current theorems [ 18 , 22 , 19 ] being contingent on the contemporary activation function primitive. This should be undertaken, for each new class of primitives, providing existence proofs for dense networks as standard. It is also hypothesised that theoretical efforts may be able to extend this through the group structure, enabling the determination of which symmetry families may yield useful functional forms a priori. This speculation would constitute a more overarching Universal Approximation Theorem, and would likely be a desirable long-term objective in any case. This could be termed a ‘Group Universal Approximation Theorem’ GU(A)T for discussion purposes. Additionally, the symmetry automorphisms are derivable from arbitrary graphs, including intrinsic privileged directions. This extended UAT approach could perhaps be further developed to derive bounds for a given symmetry on any given graph structure (generalising the typical dense network assumptions) — which could be referred to as a ‘Group Universal Bound Theory’ GU(B)T for discussion purposes. Both remain conjectures and may serve as long-term aspirational objectives, representing a beneficial theoretical direction to be 5Then one would choose to implement a beneficial instantiation from a selection of functions which abide by the chosen functional form. 4 pursued, particularly in terms of the proposed symmetry formalism. This may aid in further narrowing down which groups are suitable for deep learning and enable a better directed search, alongside inductive bias considerations. Overall, the approach outlined, which also utilises symmetry and a new taxonomic organisation, stems from a grouptheoretic formalism as a definitional tool for all primitives. Then the implications of differing sets of primitive-level algebras are extended upwards in general compositions, considering their respective inductive biases, generalised resultant model architectures, scale-interplays, theorems, and characteristic phenomena. This is distinct from both Geometric Deep Learning’s Invariant/Equivariant networks [ 23 , 24 , 25 , 26 , 27 ] as well as recent observations, and leveraging of Parameter Symmetries [ 28 , 29 ]. The former is a task-driven end-to-end respect of a particular symmetry, such that the entire model transforms predictably under its action. Hence, it is a model-level consideration that can extend down to achieve this, such as into layer maps for group convolution and further. The latter is a compositional consideration that concerns the computational equivalences under reparameterisation of surrounding affine layers, deduced from contemporary permutation-like activation function algebras. All can be similarly united under the overarching formalism of Sec. 5.2. Hence, the taxonomic system, with varying generation strengths, flavours, and scales, can be shown to recover the new and prior approaches to symmetry in deep learning as particular regimes/philosophies of an overarching group-theoretic perspective. This unification also enables novel findings when considering differing compositional regimes and layerwise constructions that have not been researched thus far. In conclusion, this paper argues that the existence of such a definitional choice for functional forms, and their consequences extending upwards, has remained a substantially underappreciated approach and should be investigated thoroughly with an aim to leverage findings generally. Such choices and their effects are typically obfuscated, neglected, and seldom [ 30 ] questioned in general model-design. A basis dependence has resulted in an internal absolute frame that appears to have become ubiquitous throughout most primitives in nearly every model. This may be partly a result of accidental notational oversimplification, suppressing basis factors, enabling the consequences of that form to remain obscure and unquestioned. There is also a decades-long history of successful and practical precedent behind it, which has become entrenched in even hardware alignment, having formed around and potentially having also shaped the wider practice. Additionally, its practicality has so far surfaced minimal apparent tensions in observations 6 . Hence, it is argued that it has become largely an unintended default, as there is a lack of suitable alternative forms, much less primitive sets, in wide circulation and unrecognised consequences arising from the current form. However, a causal link between the current arbitrary basis’s transforms and internal representations has recently been empirically demonstrated as significant [ 17 ]. An influence on models’ internal representations in turn will alter their behaviour and likely downstream performance, where it is hypothesised to display some pathological consequences. These motivate the need for a reconsideration. Hence, alternative choices and establishing their implications can now be systematically explored and developed. This includes a reselection of instantiations for each form to leverage their symmetries better, which has not been undertaken even when alternative forms have sometimes surfaced. Well-justified and understood decisions can then be drawn from a range of choices. A suitable, minimally detrimental default can also be selected, and specialist choices can be made for particular applications. This culminates in an extended formalism which provides a unifying perspective of several naturally emerging group-theoretic approaches. This has had the effect of demarcating complementary but distinct regimes and scales to consider. Pursuing this may eventually yield cross-disciplinary findings if these disparate approaches are bridged, while other scales and compositions may yield further insights beyond what is currently established. This is argued to be a good motivation for considering this broad and unifying approach. The following section discusses the hypothesised pathological consequences, which initially motivated such reconsiderations. 2 Predicted Detriments of Anisotropy This section outlines a non-exhaustive set of predicted pathologies that anisotropic functional forms may introduce. These mainly centre on the role of the activation functions, since this is the area that has been primarily explored thus far. However, similar considerations may be equally applicable to other primitives (particularly quantisation). To the author’s knowledge, some of these failure modes are newly characterised phenomena, such as the so-called ‘neural refractive problem’. It may indicate that if Isotropic Deep Learning is substantially mature, it may form a better default inductive bias unless an alternative is task-necessitated. A further intuition is expressed in App. F.1. 2.1 The Neural Refractive Problem The ‘neural refractive problem’ describes how linear and origin-intersecting trajectories of activations may converge or diverge from their initial path after an activation function is applied. This is analogous to a light ray refracting through optically varying media or boundaries. This phenomenon appears to occur in all anisotropic activation functions examined to date. The ‘refraction effect’ typically occurs more significantly at larger magnitudes — potentially producing a failure mode under network extrapolation. Neural refraction is demonstrated by curvature of previously straight origin-intersecting lines (in green) in the centre plot of Fig. 1, but is absent in the rightmost plot of the same figure. This refraction is a mathematical consequence of the form, but its impact on representations and pathological nature requires validation. Mathematically, this phenomenon has several representations, a magnitude-varying ‘dynamic refraction’ shown in Eqn. 3 or differentially in Eqn. 4. Also defined is a ‘static refraction’ definition shown in Eqn. 5. Geodesic-based constructions may also be defined. These are intended only as provisional formalisms of the phenomenon. These current formalisms are 6Except, perhaps, in observations of interpretability phenomena that may emerge from the form’s structure. 5 described for a multivariate activation function f and vector x =αˆx where ˆx is a unit vector. This relation may be satisfied for a single direction, a subset of the space or all directions ˆx∈ X ⊆ Sn . The relations generally show how the activation function alters the direction of its input vector in an anisotropic manner. ∃ˆx∈ Sn−1,∃α1=α2>0 : f(α1ˆx) ∥f(α1ˆx)∥=f(α2ˆx) ∥f(α2ˆx)∥(3) ∃ˆx∈ Sn−1,∃α0:∂ ∂α f(αˆx) ∥f(αˆx)∥α0= 0(4) ∃ˆx∈ Sn−1,∃α:f(αˆx) ∥f(αˆx)∥= ˆx(5) It can be seen that along a straight-line trajectory in direction ˆx , the result of the activation function is a curved line if dynamically refracted. Therefore, if the linear feature hypothesis is followed, then every linear feature, in refracted directions, becomes curved following the activation function. The network may exploit some of this curvature to construct new linear features in the subsequent layers; however, there may be many instances where this curvature is detrimental to established semantics. The network may lose semantic separability, produce magnitude-based semantic inconsistency or produce compensatory maladaptations in later layers. However, due to the non-linear nature of refractions in generalised directions, which continues to be compounded over subsequent layers, the network may struggle to mitigate the effect. Hence, these maladaptations may fail disproportionately for out-of-distribution samples. This may hinder the generalisation performance of the network and indicates a mode which may make representations more susceptible to adversarial attacks. Isotropic choices would resolve the refraction, potentially resulting in fewer such adaptations. An illustrative example of neural refraction is shown in Fig. 2. 0-1 1 -1 1 Identity 0-1 1 -1 1 Leaky-ReLU 0-1 1 -1 1 Standard-Tanh 0-1 1 -1 1 Isotropic Tanh Softmax 0-1 1 -1 1 Figure 2: Displays the f:R2→R2 maps, for the identity map (leftmost), standard Leaky-ReLU (centre-left), standard Tanh (centre), Softmax (centre-right), and isotropic-Tanh (rightmost). These maps transform various objects within the space, including two lines with zero-intercept, shown in red and green, as well as sets of horizontal and vertical lines in pale grey. The three centre plots demonstrate ‘neural refraction’ in its static form for Leaky-ReLU and its dynamic form for standard Tanh, as well as a more general case for Softmax. The identity plot and isotropic plot do not cause such refractions to these objects. Postulated to be especially detrimental, in both refraction cases, is the loss of semantic separability. If two distinct trajectories, representing different semantics, are transformed into curves which intersect or converge, then the separability of these concepts is lost or misrepresented. For example, suppose one direction is a linear feature for the presence of a dog in an image, whilst the other is for a horse. In that case, if these activations are of sufficient magnitude where the activation function causes convergence, the identity of the activation’s meaning can be misconstrued. This is in addition to the aforementioned deflection of linear trajectories, which may reduce the effectiveness of following linear transforms to effectively separate representations. The convergence may be particularly consequential for functions such as Sigmoid and Tanh, since large magnitude inputs end up at particular limit points (discussed as trivial representational alignments in Bird [17] ). For example, Tanh produces the limit points shown in Eqn. 6 when ˆx·ˆei= 0 for all i . If there exists an i such that ˆx·ˆei= 0 , then the transformed vector has a 0 in the corresponding index. Therefore, all vectors end up at limit points with sufficient magnitude when using elementwise-Tanh or Sigmoid. A fully-connected layer would typically only effectively separate two such converging directions at a time, which are then further curved by a subsequent activation function. lim ∀i,x·ˆei=0 α→∞ f(αˆx) = N X i=1 tanh (αˆx·ˆei) ˆei≈ N X i=1 ±ˆei= (±1,··· ,±1)T(6) Consequently, semantic separability is lost for large magnitude representations except for 3n discrete limit points for Tanh and Sigmoid. Therefore, embedded activations may be expected to align with these limit points. This explains some results empirically observed by Bird [17] . Similarly, ReLU has one distinct limit point,  0 , but otherwise an orthant unaffected by neural refraction. It is speculated that this is an additional reason for the success of ReLU, as only a subset of directions experiences the neural refraction phenomenon. Furthermore, this suggests an advantage of Leaky-ReLU: despite featuring static refraction, directions do not become overlapped, so semantic separability is retained. The network may otherwise ‘expend’ training time on producing robust semantic separability, having a potentially discretising effect on representations. This would be a needless compensatory adaptation, which may lower representational capacity and extend training as a result of inefficiency. 6 More generally, the dynamic deflection of trajectories may cause semantic ambiguity for the network, where only samples interpolable from training samples are reliably semantically identifiable. Particularly, the more significant the deflection, the greater the semantic ambiguity may be expected due to the resultant position of representations becoming unpredictable. Therefore, a magnitude-dependent semantic inconsistency may arise due to such deflections. A deflection function can be a trivial diagnostic measure, defined by Eqn. 7 for a particular activation function. θ(α; ˆx, f) = arccos f(αˆx)·ˆx ∥f(αˆx)∥(7) This may result in an additional mode of degraded performance for a network, especially on out-of-training-distribution samples. For example, suppose a linear feature roughly represents the quantity of cows in a field. In that case, the network may fail to extrapolate its function when an anomalous amount of cows are present. This would be due to a considerably larger magnitude of the linear feature, which is typically deflected significantly. Therefore, the deflection is unprecedented and becomes uninterpretable. The activation function would result in a loss of semantic consistency. Consequently, a network seeking to preserve linear features may constrain activation magnitudes through training to regions where the non-linear response is approximately predictable and stable, thereby avoiding the damaging consequences of neural refractions. Moreover, the network may move representations towards locally linear positions, limiting the beneficial transformative properties of the non-linearity. Current angular anisotropies fundamentally cause the refraction phenomenon. If compression and rarefaction occur in certain angular regions, linear features will be deflected in various ways. A fix for this is to introduce isotropy (or norm-based forms of quasi-isotropy). This is the initial motivation for developing the approach. Isotropy does not prevent compression and rarefaction of activation distributions in general, as a bias can be added to reintroduce these useful phenomena predictably. It is argued that these issues only arise when they affect linear features, not affine ones, in a potentially unpredictable and thus semantically uninterpretable manner. Applying this to all affine features would be restrictive enough to return linear approximations only; whilst shifting the origin of linear features could be considered through a symmetry-broken transformation. The phenomenon is eliminated from networks by rearranging Eqn. 5 shown in Eqn. 8, then applying the simplification ∥f(αˆx)∥=σ(α)in Eqn. 9. f(αˆx) = ∥f(αˆx)∥ˆx′(8) f(αˆx) = σ(α) ˆx′(9) Finally choosing ˆx′=Rˆxfor isotropy and Rˆx= Inˆx= ˆxfor simplicity, shown in Eqn. 10. f(αˆx) = σ(α) ˆx(10) In standard notation, Eqn. 10 can be rewritten into the functional form for isotropic activation functions shown in Eqn. 11. This should be a piecewise function, defined using the identity at x = 0 , but this is suppressed for simplicity. Alongside an appropriate smoothness condition on the Jacobian, this ensures the apparent ‘singularity’ at x = 0 is only a coordinate singularity present only due to how the functional form is denoted. Future work involves establishing a universal approximation theorem for this functional form, which is currently an ongoing area of research for the author. This form can be generalised to other functional forms in App. A, but is discussed briefly below using symmetry equivariance. f:Rn→Rn, x 7→ f(x) = σ(∥x∥) ˆx(11) The form of Eqn. 11 is O(n) time for Rn activation vectors, and only computes the non-linear term once, unlike n -computations for the non-linear term in current activation functions. In addition, radial basis functions suffered from a O(nm) cost. This bilinear scaling arguably impeded the widespread adoption of this functional form in ever-larger models, displayed in Eqn. 12 and Tab. 3.1. f:Rn→Rm, x 7→ f(x) = m X i=1 σ(∥x −ci∥) ˆei(12) Isotropy can be generalised to a result of rotational equivariance of the function, expressed as a condition in Eqn. 13. This uses a commutator bracket for convenience, with ∀R∈O (n) . This bracket can be used to similarly define the current anisotropic discrete rotational (permutation) paradigm, by using the transform ∀P∈ Sn instead of the rotation. Connecting forms of machine learning through symmetry is further elaborated on in Sec. 5, including the apparant functional form indifference between O (n) and SO (n) . Hence, these constraints constitute an effective definitional tool for generating and categorising all primitives across various taxonomies. The equivariance relation may be recognised as superficially similar to equivariant neural networks, due to an analogous equivariance relation; however, the differences in both implementation and motivations are substantial, and discussed further in App. E.1. [R,f]=(Rf −Rf) =  0(13) The relation may be more familiar as f(Rx) = Rf (x) . This relation only applies to single-argument functions and requires generalising to more circumstances, shown in App. A. A similar condition suffices: f(Rx1,··· ,RxN) = Rf (x1,··· , xN) for f:NNRn→Rn. 7 The ‘neural refractive problem’ outlines how semantic meanings may become intertwined or ambiguous due to current functional forms skewing linear features in undesirable ways. This is predicted to be especially detrimental for out-ofdistribution activations, which are likely to be most deflected and hence most semantically corrupted. Thus, the network’s generalisation may then fail in such circumstances. It may be expected that the network produces compensatory adaptations for the phenomenon, which may be narrow in the scope of their corrections. Since neural refraction is a non-linear and anisotropic phenomenon, it cannot be inverted by a single subsequent layer, potentially incurring unnecessary training overhead on producing corrections due to unintended refraction. 2.2 Quantised Representations, Emergence of Linear Features and Semantic Interpolatability Symmetry-broken functional forms have been shown to induce symmetry-broken representations which transform with the basis [ 17 ], which indicates a dependency on the anisotropy and offers an explanation why approximately discrete embedding directions are tended towards [ 31 , 16 , 17 ]. In this section, that conjecture of dependence will be made clear. Additionally, it can be hypothesised that because embedded activations are often discretised and meaningful directions may be expected to align with these embeddings, then semantic directions also become quantised. This generally appears to be the case in observations [ 31 , 32 , 17 ]. Reversing this proposed causality would indicate that a continuous rotational symmetry may enable a continuous embedding. Functional forms would not directly induce arbitrary direction-based symmetry breaking in their embeddings through training; such a structure would only emerge from task necessity. One can start with the prediction of form-induced representational collapse. In this context, representational collapse is the following heuristic: The induced discretisation of what would otherwise be an approximately smooth continuum of representations as samples drawn over a dataset. Where Discretisation is the increasing concentration of representations of clusters through training, until they eventually approach a nearly discrete-like cluster in representation space. Hence, this is also described as a quantisation, to differentiate it from other representational collapses. Quantisation would be the induced discretisation of an otherwise continuous quantity. At this early stage, until the nature of this predicted phenomenon is suitably understood, this heuristic may be more appropriate than a premature, rigid mathematical definition. The following discussion provides an informal motivation for the prediction of quantisation, followed by a more principled discussion. One may expect that the angular unevenness of various anisotropic primitives will result in some form of general effect on optimisation. Particularly, such unevenness would likely result in slight preferred directions for embeddings and slightly discouraged directions for embeddings organised around the anisotropic distinguished directions. It is the degree to which this effect may occur which is of interest. For example, in extreme cases, this may result in the absence of representations over discouraged directions and a clustering of representations over encouraged directions. This is the suggested form-induced quantisation into discrete-like clusters. Such a collapse results in informative representational degrees of freedom being suppressed in otherwise approximately connected data. This is suggested to be pathological when arbitrarily imposed through task-agnostic functional form inductive biases. This may be beneficial where redundancy can be suppressed, as discussed in Sec. 4; however, the form-induced arbitrary structure may remain detrimental. This extreme discretisation would be very distinct and may aid in detection — which is suggested to have observably already appeared [ 16 , 17 ]. However, extreme discretisation itself may not be ubiquitous; it is the more general production of task-agnostic ’structure’ about these directions which constitutes the general inductive bias, and these are suggested to be indicative of the algebraic symmetries of the forms. Without such initial unevenness, preferential angular regions would not exist, and representations may distribute more ’naturally’, perhaps smoothly or be indicative of structure in the dataset, rather than task-agnostic structure due to the choice of primitives. Practically, this may materialise in numerous modes through optimisation, depending on the function’s particular analytical qualities; however, these are suggested to all result from the underlying group structure of the forms. Particularly in primitives defined through discrete group algebras. This may arise regarding both forward and backwards pass considerations. For example, current anisotropic primitives respect at least a standard-basis permutation symmetry ( Sn ), which can result in a discrete orthant partitioning of the space in two or more dimensions. A functional form can then be described piecewise through this partitioning. Given a univariate function, f , which is applied elementwise, it can be represented piecewise as two differing functions f< and f≥ for the negative and positive semi-definite domains, respectively. When applied elementwise, the representation space’s orthants have various combinations of these two functions acting on elements, dependent upon the particular orthant. Several of these orthants are hence analytically equivalent, but rotated, in function. Generally, there are n+ 1 distinct orthants for n -width layers, with a m n degeneracy for m∈ {0,1,··· , n} . This is indicative of the underlying Sn standard-basis permutation symmetry arising from the elementwise application. Effects on optimisation may then result, where representations shift over different orthants to leverage the differing localised maps for computation. Hence, structure is expected to arise as a consequence of this symmetry partitioning. Hence, all Sn functions are expected to be influenced in this generalised symmetric manner, with specific modes contingent upon each orthant’s particular map, yet remaining tied through this underlying construction. Other permutation-based symmetries in functional forms can be considered. For example, hyperoctahedral Bn , including the standard-basis permutation with sign-flip symmetry, makes all orthants analytically degenerate when correcting for rotation; hence, its effect on representations through optimisation may indicate this. Similar for even-sign flips, which produce two sets of analytically degenerate orthants, if constructed piecewise using 0≤Qn i=1 sign (x ·ˆei)(These are denoted Dnbut are not to be confused with the dihedral group). Other discrete symmetries can result in other partitionings to be considered, such as the simplex-based symmetries in Bird [17] . Such a partitioning does not occur in the continuous-symmetry definition for isotropic primitives under O (n) , so such an aligned structure emerging through optimisation is not expected. This is illustrated in Fig. 3 for 3D multivariate maps under symmetry, while Fig. 2 demonstrates the phenomenon in 2D for Leaky-ReLU and 8 standard Tanh, where the Sn and Bn respective symmetries produce quadrant partitioning of their maps (isotropic-Tanh does not partition in this manner). Identity (1n) Permutation (Sn)Hyperoctahedral (Bn) Even-Sign Flips & Permutation (Dn)Orthogonal (O(n)) ⊂ ⊂ ⊂ ⊂ Figure 3: Illustrates the effect of the various symmetries in 3D about the standard basis. The standard bases are shown as red, green and blue arrows, with the various octants (orthants in nD ) demonstrated for discrete symmetries. Left-to-right shows, the identity symmetry of elementwise functions In , the permutation symmetry Sn , the even-sign permutation symmetry Dn , the hyperoctahedral symmetry Bn , and the continuous orthogonal symmetry O(n) producing an angularly continuous depiction. The colour-shading of the various octants demonstrates which octants are analytically identical under a rotation/permutation of their map. This intuitively shows how different octant regions relate in their maps and may influence the representation space. Particularly, the discrete maps effectively incur an absolute frame for the internal representation spaces, whilst orthogonal maps only incur an absolute origin. Additionally, there may be a hierarchical interplay of influences on representations from various functional forms. These may interact non-trivially, potentially privileging differing bases, with an overall privilege which may evolve through training. Potentially, accumulation may occur, as suggested, up to a point where an alternative basis becomes privileged and begins to disperse the existing structure; this may result in interesting dynamics, additional phase-changes and steady-state equilibrium behaviour. Whether this occurs is speculation, but it could be explored. One may also consider the consequences on the associated semantics being represented through these embeddings. Many real-life semantics are continuums: colours, positions and poses of objects, broad morphology, even within a single species or objects. Induced representational collapse onto a single discrete semantic may lose vital nuance and meaningful degrees of freedom. Discretised representations encouraged by functional forms appear to be a poor default inductive bias under these considerations. Without spurious structure added to representations from functions, the quantising bias would vanish, potentially enabling more continuous representation for the task. In this manner, isotropic functions would not prevent discrete semantics, which can be clustered through bias parameters; however, they do not promote discretisation either. Hence, Isotropic deep learning would be well-positioned to enable networks to acquire more naturally distributed representations, driven by the task and free from structure. Therefore, moving towards isotropy is hypothesised to encourage embeddings to be more smoothly distributed and better representative of the task and data. In addition, this is expected to better enable interpolatable semantics for intermediate representations between typically discrete linear features. This may substantially enhance the expressivity and representational capacity of networks — only limited by concept interferences. This may position an Isotropic approach as producing more optimal representations. Moreover, in such a case, the discrete concept of ‘representation capacity’ may become inapplicable. Each layer may express different continuous arrangements, where differing concepts are angularly suppressed and expressed in analogy to the linear features hypothesis [ 33 ]. Instead, the ‘magnitude-direction hypothesis’ is proposed as a continuous extension: with magnitudes indicating the amount of stimulus present, direction indicating the concept. Activations then populate this more continuous manifold, which is argued to enable more meaningful interpolations. This continuous semanticity may also produce a better-organised semantic map at each network layer, since intermediate representations may now relate otherwise discrete features. The lack of discretising bias may allow semantics to be brought continuously into proximity (which ‘weight locking’ discussed in Sec. 2.3 may typically prevent). A manifold without form-induced structure may aid researchers in the emerging field of representational alignment, discussed further in App. D.4. Therefore, in terms of representations, the inductive bias of isotropy appears more appropriate as a default, due to many real-world semantics being continuous and not being quantised into discrete bins through functional form induced structure. However, anisotropy may also be a good inductive bias if universal discretisation of concepts at all scales and abstraction levels is expected along the standard basis. Isotropy can be thought of as introducing an inductive bias that enables continuous and interpolatable semantics while retaining discrete semantics when task-necessitated, as opposed to being design-imposed structure. Hence, it generalises the discrete linear features paradigm into a more continuous setting. 2.3 Weight Locking, Optimisation Barriers and Disconnected Basins ‘Weight locking’ is a term to describe how, particularly, the weight parameter may suffer from being stuck in local minima found further into loss valleys, encountered only after a sufficient amount of training. Similar locking of biases near  0 may also occur. This optimisation artefact is predicted to occur through two modes — both a result of the anisotropic functional 9 form. Similarly, if used definitionally for generating new functional forms, they are constrained by this symmetry maximally. For example, if a function is left invariant to both Snand O (n), denote only the latter, since Sn⊂O (n). Additionally, the three primary categories also form a hierarchy, with probabilistic ensuring symmetry closure and algebraic ensuring the other two of their respective subdivisions. This is indicated by Eqn. 30. Algebraic ⇒Probabilistic ⇒Closure (30) Each category will be discussed, followed by examples that motivate it, and then a table will outline the symmetries of several functions. Symmetry-closure indicates that any element of a functional class can be transformed under a symmetry, and the result is also a member of the class. This is displayed in Eqns. 31, 32 and 33 for left-invariant, right-invariant and equivariant respectivly. The class Fwould be a chosen subset of all maps. ∀f∈ F,∀g∈ G (f◦g)∈ F (31) ∀f∈ F,∀g∈ G (g◦f)∈ F (32) Equivariant closure is a differing requirement, involving the pairing of its group-inverse (but can be generalised): ∀f∈ F,∀g∈ G g−1◦f◦g∈ F (33) These statements indicate that any function in a class which is transformed under symmetry is still a member of the class. Using representation theoretic generalisation it can be denoted ∀f∈ F , ∀g∈ G , ρ(1)(g−1)◦f◦ρ(2)(g)∈ F for two representations ρ(1) and ρ(2) . Closure conditions do not say if these two instances of the class are equally likely to be initialised, which is a stronger condition. This latter case is the probabilistic condition and is given in Eqns. 34, 35 and 36 for left-invariant, right-invariant and equivariant respectivly. P gives the probability of the member of the class F to be initialised. These could be termed a ‘weak’ accordance with a symmetry group. ∀f∈ F,∀g∈ G P(f◦g) = P(f)(34) ∀f∈ F,∀g∈ G P(g◦f) = P(f)(35) Again, probabilistic equivariance is a differing requirement: ∀f∈ F,∀g∈ G Pg−1◦f◦g=P(f)(36) The probabilistic condition can be specified as time-like, initialisation-like, data-like, or any combination of these. This depends on how the probability is considered. Time-like would be, for example, over subsequent iterations of forward passes in the network, discussed in App. C. Initialisation-like indicates the distributions of parameters which are spontaneously symmetry broken on initialisation. Data-like can consider the probability defined over samples of the dataset. Other situational subcategories for probabilistic conditions may exist and require extension to the formalism. Using representation theoretic generalisation it can be denoted ∀f∈ F,∀g∈ G,P(ρ(1)(g−1)◦f◦ρ(2)(g)) = P(f)for two representations ρ(1) and ρ(2). Finally, there are algebraic symmetry relations, or a ‘strong’ accordance with a symmetry. These indicate that every instance of a function in the functional class respects a symmetry which leaves the computation unchanged. They are defined by Eqns. 37, 38 and 39 for left-invariant, right-invariant and equivariant, respectively. These are the familiar bracket relations. ∀f∈ F,∀g∈ G f◦g=f(37) ∀f∈ F,∀g∈ G g◦f=f(38) Again, algebraic equivariance would be a differing requirement: ∀f∈ F,∀g∈ G f◦g=g◦f(39) Using representation theoretic generalisation it can be denoted ∀f∈ F , ∀g∈ G , ρ(1)(g)◦f=f◦ρ(2)(g) for two representations ρ(1) and ρ(2) . One can also consider if the function has multiple arguments or concatenated output spaces. These generalised domains and codomains can have these conditions applied in various ways, e.g. direct sums or tensor products. Hence, one can extend the above definitions to functions with multiple arguments, including differing relations applied to any combination of arguments. Additionally, weight sharing can be considered another extension to the model. The following discussion concerns several applications of the formalism. To begin with, App. B.1, discusses a functional form which introduces anisotropy but in an isotropically initialised manner. This motivated the construction of this formalism. This enabled such nuance in classifications beyond algebraic constraints. For example, f(x;W) = Wx is considered anisotropic since its algebraic equivariance is the identity, but it could still be weakly isotropic. Then considering f(x) = PN i=1 f(x ·ˆei) ˆei , which has algebraic permutation equivariance, one may consider their 16 composition: f(x) = PN i=1 f((Wx)·ˆei) ˆei . This latter class is not strongly isotropic but could be weakly isotropic under suitable initialisation — this difference is significant and motivated by classifying forms through this meta-framework. This produced the original distinction between weak and strong symmetry accordance, which was extended with closure. This is an example of how the composition of layers can then undergo spontaneous symmetry breaking. Deep learning models are hypothesised to leverage symmetry-breaking phenomena, which are further adapted through training, to achieve practical computation. Therefore, defining a alternative approach which prevents any such phenomenon would be impractical to purpose. This is one of the primary motivations of this paper: elucidating the role of symmetry breaking in networks systematically by exploring this taxonomy and altered primitives. The approach to defining each fork generally considers algebraic symmetries in primitives, while allowing parameterised maps to be closed under the general linear group, in whichever relevant flavour. However, a pure branch would also initialise such parameters under probabilistic constraints to the symmetry of the branch. Hence, this would result in primitives respecting the overall intended symmetry before symmetry-breaking initialisations and learning. As stated, the specific group used in such constraints is chosen from a selection that is considered derivable from arbitrary directed graph automorphisms, due to their connectivity, and then chosen to be applied to associated primitives. This can then result in the production of functional classes which have a closure under these respective automorphism symmetries. These maps, with closure under general automorphisms, can then be chosen to be upgraded probabilistically or up to algebraic constraints for specific groups they are closed under. Such considerations can also be applied at all scales: each individual primitive functional form and any possible composition of these through the arbitrary graph. In effect, this indicates which symmetries are in principle available before specific downstream choices of functional forms for primitives are chosen or freshly formulated. This interplay between architectures and primitive-constraints outlines the family of deep learning approaches and is suggested to be axiomatic-like. As stated, from the closures available, one can choose a specific automorphism to form a stronger constraint to. This can be a subgroup of, or equal to, the whole automorphism group produced by the graph. For ‘pure’ branches, this choice results in parameterised mappings being upgraded to probabilistic, whilst functional forms over nodes pick up algebraic constraints. Functions from this functional class, which maximally abide by such restrictions, are then used to produce particular primitives for the specific network. For example, this would typically be a restriction to a permutation family over the standard basis in contemporary deep learning. This constraint is then applied either probabilistically (such as affine maps fx;W, b=Wx + b , which can have left and right probabilistic invariance) or algebraically (such as in activation functional forms f(x) = Pn i=1 f(x ·ˆei) ˆei , which is algebraically-equivariant). A different choice can yield Isotropic probabilistic and algebraic forms, or any other symmetry fork of primitives. Some orthogonal probabilistic initialisers are used; this can now be termed a hybridisation. Overall, the suggestion is that these are a considered choice in design. Isotropic networks are the statement that, at minimum, the primitives should be probabilistically in/equivariant maximally to orthogonal family actions, but strive for algebraic isotropy wherever feasible. Yet this is not dogmatic: a linear map would be too restrictive for computation if algebraically-orthogonal in/equivariant, so in such cases, only probabilistic-orthogonal in/equivariance, as spontaneous symmetry breaking, would be desirable. It is the consideration of such choices which is important, and knowledge of their induced biases, not the restriction to them. Hybridisations between differing symmetry group definitions in a single model can be explored and appear to have justification already; these would occupy intermediate positions compared to the more purely defined branches proposed. Therefore, although the pure branches are rigid in their form constraints, explorations of hybridisations are also encouraged, analysis of which is enabled through this formalism. To the best of our knowledge, these are believed to be distinct research directions compared to previous symmetry-based literature. As stated, this formalism is effective at distinguishing several approaches, including the considerations of this paper from Geometric Deep Learning’s end-to-end symmetry-defined networks, such as equivariant networks, and Parameter Symmetries. The latter two regard a different scale to which these relations apply. Geometric deep learning’s equivariant networks can be recovered from the framework by considering the function class of the entire model and restricting the model class to those which are algebraically equivariant to the intended symmetry group, observed in the underlying data distribution. A similar applies to invariant flavoured constraints on models. Additionally, one can consider group convolution as the next compositional scale down to which these apply: functional class blocks over which these constraints are applied. Such components do occupy smaller-scale compositions, though they are in furtherance of an end-to-end algebraic symmetry of the full model. This is an entirely different objective from those presented in this paper, which are intended for general application in arbitrary architectures, allowing and shaping symmetry breaking through considering differing taxonomic generations composed in models. This represents a significant distinction in approach. These taxonomic considerations consider the implicit inductive biases which act on and interplay within the network’s dynamics; it is not solely focused on constructing models about a desired equivariance or invariance to the group necessary to preserve the data structure. Hence, one can consider this proposal as the consequences of symmetry breaking in general networks. Nevertheless, this overall formalism can encompass both approaches as a special case of symmetries applied over functional classes at differing scales and philosophies. Hence, this typically differs in scale from most of the considerations in this paper’s approach, which focuses on the functional form primitive’s relation and its increasing interacting compositions, as opposed to the model as a whole and the restricted functional classes required downwards to achieve such model-scale results under composition. These are complementary but differing approaches, with independent objectives and philosophies; yet, some interplay can likely be 17 established. Such interplay may be considered through this broader formalism. However, at present, the approaches between these primitive foundation reformulations and equivariant networks appear to often be mutually exclusive in many GDL models due to the latter’s frequent dependence on elementwise primitives, but this is not always the case. However, this may be reconciled over time under alternative implementations and is discussed further in App. E.1. This significantly differentiates the typical scales at which Geometric Deep Learning’s equivariant models consider symmetry from those at which this paper considers them. An algebraic equivariance over the entire model has not been the primary concern of this paper. Moreover, other differences arise, such as how the reformulation of primitive pertains to symmetry families due to the different dimensionalities of their respective maps, whereas models adhere to a single group preserved throughout. Additionally, this formalism appears to suitably distinguish recent observations and consequences due to Parameter Symmetries, alongside their parameter-space degeneracies. This is the avenue investigating how specific symmetry actions can leave aspects of existing networks unchanged in functionality, resulting in parameter degeneracies. This phenomenon can be united into the framework as a compositional consequence. Three such relations make this evident: a right-closure of a general linear layer, an algebraic-permutation equivariance of activation functions and a general linear left-closure of a linear layer. Using this formalism, one can now extend the considerations to other primitive compositions for analysis as well. For example, a linear layer f1(x) = W1x+ b1 ( Rm→Rn ) has a right-closure to n×n general linear transforms, since the functional class can take on differing parameter values. Similarly for f3(x) = W3x + b3 ( Rn→Rp ) which has a left-closure to n×n general linear transforms. Informally, this means these layers have the capacity to ‘absorb’ any general-linear, or subset of, transforms whilst remaining in the class. Finally, the activation function, say f(x) = Pn i=1 tanh (x ·ˆei) ˆei has a signed-permutation algebraic equivariance, meaning Eqn. 40 follows from the algebraic-equivariance to P∈Bngroup. f(x) = f(Inx) = fP−1Px=P−1f(Px)(40) Following this, the composition of f3◦f◦f1 means that the left and right general-linear closures can ‘absorb’ these P and P−1 transforms whilst remaining in the class. This is because f3◦P∈ F3 and P−1◦f1∈ F1 . This combination of properties reproduces the parameter symmetries under composition f3◦f◦f1. This recasting of parameter symmetries under the above symmetry formalism may aid generalised considerations and comparisons. This approach reveals a large number of degeneracies, per such sandwhiched construction of an algebraicequivariant maps Rk→Rk and associated closed linear layers, the network acquires a multiplicative (k!)2 factor degeneracy in its parameter space if the activation function is Sn algebraically equivariant. Independently, this was considered a pathology between anisotropy and isotropy within this work, and can be connected to the discussion in Sec. 2.3. Additionally, this also highlights that the emergence of a Bn or Sn symmetry, in particular, is not due to the parameters (which are general linear invariant closures); instead, it is the function they sandwich, in this case the activation function, which is algebraically equivariant to a transform. This aligns with Godfrey et al. [28] , which identified and explored permutationrelated symmetries over existing activation functions. Their intertwiner groups in activation functions correspond to those of parameters, and can be directly mapped to the above formalism discussed8. Extensions to this can be considered under this generalised formalism, such as making clear the consequences if the affine maps are not left/right closed under group G , when sandwiched with an algebraic-equivariant G . In such cases, parameter symmetries are not applicable and the network can become unique. One could also consider the promotion or prevention of such closure symmetries up to probabilistic conditions in a similar manner. Prevention may cause the initialiser always to favour particular arrangements of the network under symmetry, potentially aiding in alignment efforts. Furthermore, one can consider the compositional consequences of adding noise under a probabilistic constraint. This is a suggested regulariser discussed in App. C. It may also have consequences for generative efforts, which could be explored. In addition, more careful treatment of symmetries in Radial Basis Functions [ 30 ] ( Rn→Rm ) can be explored. These appear to feature a permutation right-closure ( Sm ) and orthogonal left-closure ( O (n) ), which, when combined with linear layers, can form similar parameter symmetries to those discussed. One can also consider other compositions, such as the maps which are composed to form residuals f(x) = x +g(x) , where one can consider how facquires the symmetry of g. This is because the function is the summation of an identity map with a general map g . The identity commutes with any algebraic symmetry, so the resulting symmetry of their composition is only defined through g’s subset symmetry. Furthermore, it appears that there is a tradeoff between the level of the maximal constraint on the functional form and the result on representations. For example, algebraic Sn equivariance is less of a constraint on functional forms than algebraic O(n) permutation equivariance. Yet, the latter appears to produce fewer constraints on representations by removing absolute reference directions, the distinguished directions. This tradeoff between algebraic constraints and resultant representational constraints presents an interesting avenue for exploration. Overall, this classification construction seems capable of both demarcating and unifying multiple differing approaches to symmetries within deep learning. Hence, this broader symmetry formalism may be highly advantageous to explore further. The present focus of each’s approach is pictorally demonstrated in Fig. 4. It also indicates how naturally the hierarchy of constructing models may require analysis of smaller compositions, such as discrete convolutions for equivariant networks. Still, it is in furtherance of the model-scale regime to which the symmetry constraints are applied. Due to its lowest hierarchical positioning, if a reformulation of the primitive foundation occurs, then it has consequences for all layers, all compositions, and all models in all applications, extending upwards. It is the study of the interplay and emergent phenomena between symmetry 8 They also demonstrate a connection to representations. This is further supporting evidence for this paper’s hypothesis that activation functions produce an inductive bias on representations. 18 and networks these bring, as well as using it as a definitional tool across all primitive to generate group-defined classes, from which reselection of specific instantiations can occur. Flavours Generations Left Invariant Right Invariant Equivariant ClosureProbabilistic Algebraic Taxonomic Table Example Groups: Foundational Primitives Layers Composed Layers Models Network Construction Heirarchy Primitive-First Programme Parameter Symmetries Equivariant NetworksGroup-Convolutions Approaches: Depends On Built From Figure 4: Pictorally demonstrates the various approaches to symmetry in contemporary deep learning, through their generation and flavour regimes as well as their typical scales. Left demonstrates the hierarchical dependencies in the construction of deep learning systems. Dark orange indicates a top-down approach philosophy, which seeks model-scale symmetries derived for purpose-built, targeted applications, and consequently constrains algebraically downwards to ensure this. In contrast, the mint colour represents a bottom-up, causal-effect, and group-theoretic philosophy, where new compositions are generated from and are contingent upon smaller-scale constructions and may occupy more general generations within the taxonomy. Centre-top provides several group taxonomies which may be considered for implementation. The centre-bottom specifies several approaches to symmetry and is colour-coded to identify regimes in the leftmost and rightmost diagrams. The rightmost depicts the taxonomic organisation put forward. Parameter symmetries are the composition of left/right-invariant closures with contemporary activation function equivariance to permutation-like groups, which is indicated by the fused triangle in the taxonomy. This also raises the problem of defining the notions of layers and primitives, which is addressed in an upcoming paper. As a consequence, end-to-end models can contain instances where specific primitives can be reformulated such that the model as a whole respects the symmetry; however, they cannot encompass the primitive-first paradigm as a whole, as this is a superset due to forming the foundational base and its argued consequences for all general models. This indicates the present approach’s universal, axiomatic-like importance for consideration, as any consequences further interact with compositions contingent upon them. This is not to imply that one philosophy is superior to another; they target differing objectives yet may be complementary. One is already well-established and growing, with several state-of-the-art results that already indicate success in achieving its intended goals. The other is attempting to determine how symmetries from functional forms may interact with networks as unintended inductive biases, and then reformulating primitives for beneficial purposes in general architectures. Distinctions are drawn to avoid confusion between their objectives and considerations, as they independently share a group-theoretic root as their natural expression, which may appear similar at first due to its relative infrequency in deep learning. This is similar to how parameter symmetries also share a group-theoretic root and are again distinct. The objective of this formalism is to draw on this shared overlap of group-theoretic considerations to construct an overarching framework that provides a more comprehensive and high-level perspective. Hence, this group-theoretic approach provides unifying terms to compare these within the context of deep learning. The approach of this paper is to be primarily conscious of such decisions regarding functional classes, at all scales, in their introduction of inductive biases and interactions. This process begins with generating numerous new families of foundational primitive implementations through careful searches and selection within these functional classes, followed by building these upwards in compositions for novel architectures and potentially improved, generally applicable models. Tab. 5.2 roughly indicates the typical symmetry properties of conventional forms. For example, a linear layer can be chosen to be initialised isotropically, but itself does not display the associated algebraic symmetry. Each such choice restricts the functional class and incurs specific inductive biases to consider. Question marks on the equivariant network indicate the specific model’s chosen initialisations. Comments Function Closure Probabilistic Algebraic L R E L R E L R E Affine Layer Wx + bGL (n) GL (n) GL (n) O (n) O (n) O (n) InInIn Standard-Tanh PN i=1 f(x ·ˆei) ˆeiInInBnInInBnInInBn Composed PN i=1 f((Wx)·ˆei) ˆeiGL (n)SnSnO (n)SnSnInInIn Equivariant-Nets fmodels G?G?G G?G?GInInG CE Loss L:Rn→RSnInInSnInInSnInIn Isotropic-Tanh σ(∥x∥) ˆxInInO (n) InInO (n) InInO (n) One can extend this formalism in numerous ways. For example, one can upgrade the group-theoretic considerations to 19 gauge-theoretic, if an application suitably justifies such a generalisation of the approach. Additionally, one can consider more general symmetries, for example O (1,3) , where a metric-tensor can be inserted into the functional form to produce a pseudo-norm: f(x) = fxβgβγxγˆxα. This has some interesting consequences. Particularly, one can stack the metric of such a function, similar to the manner in App. D.3, producing f(x) = f(xβgβγ hxγ)ˆxα . If the contravariant and covariant indicating indices are dropped, whilst allowing non-symmetric metric for later pairwise comparisons then the following equation can be considered: f(x) = f(xβgh βγxγ)ˆxα . Moreover, the metric gh βγ can then be expressed generally as the product of two matrices Wh k and Wh q , returning to standard matrix notation: f(x) = f(xT(Wh q)TWh kx)ˆx and considering K=Wkx and Q=Wqx , then a self-attention-like structure, f(x) = f(QT hKh)ˆx , is closely recoverable, and could be generalised for the value matrix, pairwise inner-products and normalisation factor 9 . Changing to a non-softmax activation function, which doesn’t depend on other elements or components, reinforces this generalised symmetry consideration and partially motivates the discussion App. D.1. Furthermore, one can make the apparent metric position dependent, having representational similarity follow from metrics over a non-linearly contracting and expanding space. This could be a meaningful avenue to explore, clustering regions of representation space or dispersing others. Hence, this symmetry formalism enables a recontextualisation which may have consequences for different comparisons between representations and potentially improved expressibility in attention. Therefore, this symmetry approach also allows the reinterpretation of self-attention-like operations within this symmetry formalism. Overall, this formalism is highly versatile in categorising symmetry techniques within deep learning, potentially enabling improved cross-communications whilst also making clear alternative avenues to explore. The intention was to enable further comparisons and generalisation to more examples, using this approach to both audit existing functions into a categorised taxonomy and to use it as a definitional and generative method for producing functions and models collated by group structure. This can occur in parallel to understanding the ramifications of such symmetry definitions through representational geometry and mechanistic interpretability, optimisation and performance, parameter degeneracies and more. 5.2.1 Note on Representations The following three paragraphs briefly discuss representation-theoretic additional considerations that may be important for taxonomisation, but are nuances that may complicate the utility of the overall taxonomy for general practice. Moreover, these flavours can all be reformulated in terms of representations, which encompass and extend the equivariance/invariance formulae detailed below, and could be organised using a highest weight approach, such as indexing isotropy with Casimir operators. This can be used to label these representations and provide a more principled, primitive, and compositional framework. Hence, using a representation theoretic approach, for corresponding left and right actions, can add further nuance to the primary flavour categorisations notated, and likely representation theory may better organise foundational biases and their interactions. This should be undertaken, but is notationally suppressed for the sake of approachability, and is assumed as an implicit and vital part of the taxonomisation. Furthermore, considering the countless groups possible, and all the various possible representations, it is encouraged that general foundational bias principles are distilled down over particular families of groups defining primitives, e.g. orthogonal as opposed to particular instances O (5) , O (8) , O (100) etc. This generalised collation of biases may often offer better practical benefit and leveragability, rather than more niche instances of specific groups, so this is encouraged foremost. For example, it was shown in Bird [17] that the axis-anti/alignment generally persisted independent of the network width, which would indicate differing specific instances of the permutation family: S24 , S32 , etc. — this is the primary objective of this programme. Therefore, although a more fine-grained categorisation of primitives is possible, it may be advantageous in general to categorise them only by group and dissimilar representations. Additionally, representations connected under a conjugated transform ρ′(g) = A−1ρ(g)A may have meaningfully differing inductive biases depending on A , particularly whether there exists an element g′∈ G for which A=ρ(g′) is a representation. If this is not the case, then the resultant bias may be non-trivial and could be organised by what group A is a representation of. An example is that if a weight-decay regularisation is used, and two differing equivalent representations are chosen for a primitive composed with affine layers, then if A has column-wise or row-wise vectors which do not normalise to one, there may be meaningful compositional biases between L2 regulariser-affine-activation function interactions. A similar argument applies if A is not in the representation of the permutation defining an anisotropic activation function; then it will interact meaningfully with an L1 regulariser. This suggests that more than just irreducible representations may be relevant as foundational biases, and further investigation is required. This will also inflate the considerations of the foundational bias scheme; however, it again is likely beneficial to consolidate these into general principles for wider adoption, although specific instances could still be applied if desirable. An example of this is the permutations used in Bird [17] for rotated and non-standard-basis representations, where it was found that no observable differences in biases arose. Furthermore, there is an exponential growth in considerations when considering compositions; the desire is to distil general principles over constructions of only group-defined primitives. Overall, it is suggested that despite finer distinctions being possible through representation-theoretic principles, the practicality of the taxonomy may remain primarily within group-theoretic definitions of primitives and general principles for their induced biases. This adds to why the group-theoretic approach is primarily showcased, with fewer primary flavours, alongside the notational approachability of this primitive-first programme. 9Such a normalisation factor may be similarly applicable to standard isotropic activation functions too. 20 6 Conclusion This paper focuses on a novel case study of an isotropic functional form as a hypothesised better default inductive bias for deep learning. Current forms have been demonstrated in previous literature to produce task-unmotivated representational artefacts [ 17 ], which this work hypothesised may limit the networks’ semantic expressibility. It is further argued that the current anisotropic functional forms may have detrimental effects on performance and learning through the predicted ‘neural refraction’, ‘discrete semantics’, and ‘weight locking’ phenomena. Removing such constraints from the model is also argued to unconstrain the representations from any particular basis. Hence, it is expected to produce a more natural activation representation based upon task necessities, rather than a structure induced by human-imposed functional forms. This may improve semantic structure and produce high-capacity embeddings, particularly important for applications discussed in App. D. In isotropic networks, functional forms are promoted from the existing discrete permutation symmetry to a continuous orthogonal symmetry. This has substantial consequences for the form of almost every function in modern-day deep learning. This paper and appendices also outline a framework for connecting future work in these directions. Several preliminary activation functions, normalisers, optimisers, regularisers and operations are also described as a starting point. The tenets of reviewing such functional form choices are encouraged. Through comparison studies, this philosophy should, at the very least, reveal which characteristics of functional forms most significantly contribute to performance. This includes the role of symmetry breaking, which is one of the foremost concerns when developing this taxonomy, as it determines how various compositions may induce hierarchical interactions that aid or detract from learning. A change to isotropic deep learning is argued to be generally advantageous, but may need substantial time for development as a mature alternative. New models and benchmarks may also require development to determine the practicality of this alternative approach. Additionally, better-optimised implementations that suitably leverage isotropy are preferable, as the ones described remain illustrative placeholders. The functions proposed so far are analogous to existing functions, which may not be inherently optimal for an isotropic network, even if they share superficial similarities. Therefore, empirical work on these placeholder functions will be presented in future papers, to not distract from the primary motivation for this shift to isotropic deep learning and the wider group-theoretic alternatives. The proposed ideas aim to stimulate the community’s interest in conducting a directed search for better isotropic functions and determining whether this approach should be adopted in wider applications. The predicted pathologies may also offer a suitable falsifiable mode to validate the principles at this early stage. Finally, other approaches to symmetry in deep learning are shown to be distinct from the approach proposed in this paper; however, an overarching symmetry formalism is also introduced to unify these disparate approaches and make clear other avenues to explore in a similar regard. It is proposed that the breadth of the reformulations may constitute an alternative direction for deep learning: Isotropic deep learning, with the taxonomy of Sec. 5 extending this further. Connecting this to graph automorphisms yields axiomatic-like choices, forming distinct ‘branches’ of primitives and downstream models to consider. In general, it situates symmetry prior to neurons, rather than symmetry deduced from the neuron-defined computation graph — an ontological inversion of what is defined or deduced from a neural network. Extending this further may provide a complete axiomatic construction for deep learning. Additionally, new categories of primitives may be highly distinct from all that have come before, particularly with potential consequences on representations and learning. This may be leveraged to offer new approaches for deep learning in both practice and theory. This may require the development of various subdisciplines for the study of respective branches and any interrelatedness, such as the conjectured group universal approximation and bound theorems. Reanalysis of existing phenomena and results may also be undertaken to determine if they are predicated upon specific primitive choices. In general, this taxonomic group classification may be an effective approach in categorising new functional forms and judging their interactions, foundational biases and symmetry-breaking phenomena. However, it is not a replacement for other analytical factors that may be considered in setting the functional form or after a group-defined functional form is fixed. For example, the identity group can generate such a broad repertoire of functional forms that is insufficiently distinguished by group theory alone. Hence, other analytical tools, besides group-theoretic considerations, remain crucial to distinguish their foundational biases. Yet, a group-defined approach does enable a principled way to generalise over sets of primitives and is argued to be important in considering the useful computational maps that networks may leverage — such as in the orthant setup. Overall, it is argued that group theory provides a good foundation for defining initial primitive forms and an initial framework for categorising interactions. However, the categorisation of foundational biases is likely to outgrow the limits of group and representation-theoretic taxonomisation, alongside symmetry breaking. It may practically require an extension of the taxonomy beyond this, and group theory should be employed to extend the exploration, but not limit it. Overall, this generates a considerable and novel design axis for the field of deep learning, which may be explored with the aim of producing better-performing models, gaining a better mechanistic understanding of their functions, discovering fundamental phenomena and results, and hopefully generalising to more applications. 6.1 Final Note on Philosophical Implications Philosophically, this paper advocates for a shift in perspective on symmetry in relation to deep learning and an ontological shift in what could be considered a deep learning model, as well as the emergent consequences that stem from this. This paper concerns the emergence of symmetry within deep learning itself and how it may, inherently and crucially, task-agnostically, act on a model’s computation. This approach covers exploring the implications of this generally on models for all applications, contingent upon various choices of functional form definitions. This is a markedly different assumption about the relationship between symmetry and deep learning compared to end-toend model approaches, which leverage the established symmetries of the natural world using observations and arguments 21 frequently emerging externally to the discipline, and extending these into deep learning for models to adhere to. These are extensions of a known external physical approach to symmetry and transferring into models. However, this paper proceeds from a drastically differing assumption regarding the emergence of symmetry, emphasising that it is internal to deep learning as well and has importance in its own right. It is making the argument and assumption that group-theoretic considerations are natively important characterising tools for the field, reinforced by it being already unintentionally set within the contemporary defining choices of primitives. Hence, it is not a perspective on emulating natural symmetries into a system, but considering that the functional form structure already within the system carries its own symmetry biases naturally attributable to and categorisable through symmetry — an externalist top-down versus internalist bottom-up perspective on symmetry for deep learning. Hence, this constitutes a philosophical reorientation of symmetry, not just a tool for aligning models with the external world, but also as an internal, function-driven influence acting task-agnostically on representations, optimisation, interpretability, and more generally — the primary and generalised perspective shift this paper advocates for. Primarily, it is argued that these considerations may naturally arise internally, existing as important factors within the field, in addition to, but not requiring, motivation from the external world. Hence, it is contingent on an assumption that symmetry’s importance as a categorising principle is not only imported but also native to deep learning. Perhaps one of the most crucial aspects is the new ontology of what constitutes a deep learning approach, which this reframing provokes. It highlights that it may no longer be contingent on a computational system constructed upon a neuronwise interconnectivity approach, but generalised across various group-defined primitive sets. This even generalises the notion of a neuron as an object generalised for higher-dimensional considerations, which is a downstream consequence after a symmetry definition is set. Definitions reverse from neurons preceding deducted symmetries, to symmetries defining neurons — and if symmetries are derived from automorphisms, then graphs precede these. It also contends that this is likely not limited to foundational biases stemming from parameter degeneracies, as a consequence contingent on current formalisms, such as affine layers plus a non-linearity; it also generalises the implications beyond this compositional structure, suggesting that even in such cases they do not need to be emergent phenomena solely predicated and formulated on degeneracies, but general function-driven consequences attributable atomically, such as neural refraction, and more in general compositions. Hence, this differs from parameter symmetries, which identify computational equivalences under reparameterisations deduced when assuming the existing fixed set of primitives and the consequences therein. Instead, this work also changes the assumptions upon which those deductions are predicated by broadening the definition of primitives to a plethora of group-defined relations and considering the implications of these more broadly within the new ontology. Moreover, these function-driven implications are argued to exist for a model even where parameter-induced computational degeneracies may not. Hence, this is being considered as an exploratory approach to phenomena which may extend beyond those which are predicated on computational equivalences through parameter degeneracies, and therefore is not contingent on these to exist. For example, identity-defined functions may have similar foundational biases. Overall, the emphasis extends to more general ramifications, from all primitive definitions, both atomically and general compositional cases for networks. A trivial counterexample of how these extend past parameter degeneracies could be considering the diagonal basis-dependent and permutation-defined approximation of Hessian-like scalings for adaptive optimisers. This is a more trivial example, but similar constructions could be considered for activation functions and many other primitives, which may not depend on the specific parameter-symmetry composition structure for influence — such as in cases where a restricted linear map may not have the required closures, but still have phenomena dependent on the definition of the primitives used, for a single example: consequences of refraction can exist independent of degeneracy. This also reframes several interpretability approaches, which may also be dependent on the existing set of primitives. Determining the semantics of representations may be mediated by these axiomatic-like choices, where artefactual structure may arise from primitive algebra alone. By broadening over differing group-defined sets, the geometry of representations may also alter. This may also affect the knowledge and deductions a model can make — broadening its epistemic horizon. This interplay between imposed geometry of primitives and emergent consequences on representations and optimisations may constitute a ’no-free-geometry’ consideration — reframing that observation of representation structure may be as much conditioned on the primitives as the data. Hence, care should be taken not to use it as evidence in support of the primitives circularly, but perhaps ascertain more fundamental and shared organisations for representations as insight into semantic relations and learning. This also motivates a shift in perspective on representations, from solely copying the external structure of reality to also being strongly structured by the non-derivable choices of functional forms inherent to the model. Overall, this is a notably differing philosophy for symmetry in deep learning. It assumes that it already pre-exists and argues that it may be a constitutive and intrinsic property of deep learning, forming a leveragable, native, and foundational design axis. Philosophically, this is suggesting a pluralism redefinition for what constitutes deep learning by generalising it across various group-theoretic components constituting a model and determining and categorising their implications, such that they can be beneficially leveraged. 22 References [1] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012. URL https://proceedings.neurips.cc/ paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf. [2] Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56):1929–1958, 2014. URL http://jmlr.org/papers/v15/srivastava14a.html. [3] Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions, 2014. URL https://arxiv.org/abs/ 1409.4842. [4] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition, 2015. URL https://arxiv.org/abs/1409.1556. [5] Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015. URL https://arxiv.org/abs/1502.03167. [6] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015. URL https://arxiv.org/abs/1512.03385. [7] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017. URL https://arxiv. org/abs/1412.6980. [8] Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks, 2018. URL https://arxiv.org/abs/1608.06993. [9] Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks, 2020. URL https://arxiv.org/abs/1905.11946. [10] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. URL https://arxiv.org/abs/1706.03762. [11] Benjamin F Logan and Larry A Shepp. Optimal reconstruction of a function from its projections. 1975. [12] Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10), pages 807–814, 2010. [13] Prajit Ramachandran, Barret Zoph, and Quoc V. Le. Searching for activation functions, 2017. URL https://arxiv. org/abs/1710.05941. [14] Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus), 2023. URL https://arxiv.org/abs/ 1606.08415. [15] Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al. Rectifier nonlinearities improve neural network acoustic models. In Proc. icml, volume 30, page 3. Atlanta, GA, 2013. [16] Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy models of superposition, 2022. URL https://arxiv.org/abs/2209. 10652. [17] George Bird. The spotlight resonance method: Resolving the alignment of embedded activations. In Second Workshop on Representational Alignment at ICLR 2025, 2025. URL https://openreview.net/forum?id=alxPpqVRzX. [18] George Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems, 2(4):303–314, 1989. [19] Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural Networks, 4(2):251–257, 1991. ISSN 0893-6080. doi: https://doi.org/10.1016/0893-6080(91)90009-T. URL https://www.sciencedirect. com/science/article/pii/089360809190009T. [20] Chris Olah. Neural networks, manifolds, and topology — colah.github.io. https://colah.github.io/posts/ 2014-03-NN-Manifolds-Topology/, April 2014. [Accessed 15-05-2025]. [21] Peter Foldiak and Dominik Endres. Sparse coding, Jan 2008. URL http://www.scholarpedia.org/article/ Sparse_coding#:~:text=Sparse%20coding%20is%20the%20representation,subset%20of% 20all%20available%20neurons. 23 [22] Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks, 2(5):359–366, 1989. [23] Taco S. Cohen and Max Welling. Group equivariant convolutional networks, 2016. URL https://arxiv.org/ abs/1602.07576. [24] Taco S. Cohen and Max Welling. Steerable cnns, 2016. URL https://arxiv.org/abs/1612.08498. [25] Daniel E. Worrall, Stephan J. Garbin, Daniyar Turmukhambetov, and Gabriel J. Brostow. Harmonic networks: Deep translation and rotation equivariance, 2017. URL https://arxiv.org/abs/1612.04642. [26] Taco S. Cohen, Mario Geiger, Jonas Koehler, and Max Welling. Spherical cnns, 2018. URL https://arxiv.org/ abs/1801.10130. [27] Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇ ckovi´ c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges, 2021. URL https://arxiv.org/abs/2104.13478. [28] Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge. On the symmetries of deep learning models and their internal representations, 2023. URL https://arxiv.org/abs/2205.14258. [29] Derek Lim, Theo Moe Putterman, Robin Walters, Haggai Maron, and Stefanie Jegelka. The empirical impact of neural parameter symmetries, or lack thereof, 2024. URL https://arxiv.org/abs/2405.20231. [30] David Lowe and D Broomhead. Multivariable functional interpolation and adaptive networks. Complex systems, 2(3): 321–355, 1988. [31] Chris Olah, Alexander Mordvintsev, and Ludwig Schubert. Feature visualization, Nov 2017. URL https://distill. pub/2017/feature-visualization/. [32] David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: Quantifying interpretability of deep visual representations, 2017. URL https://arxiv.org/abs/1704.05796. [33] Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. Linear algebraic structure of word senses, with applications to polysemy. Transactions of the Association for Computational Linguistics, 6:483–495, 2018. [34] Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2019. URL https://arxiv.org/abs/1803.03635. [35] Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton. Dynamic routing between capsules, 2017. URL https: //arxiv.org/abs/1710.09829. [36] R. Hadsell, S. Chopra, and Y. LeCun. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, pages 1735–1742, 2006. doi: 10.1109/CVPR.2006.100. [37] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations, 2020. URL https://arxiv.org/abs/2002.05709. [38] Vardan Papyan, X. Y. Han, and David L. Donoho. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40):24652–24663, September 2020. ISSN 1091-6490. doi: 10.1073/pnas.2015509117. URL http://dx.doi.org/10.1073/pnas.2015509117. [39] Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. How does batch normalization help optimization?, 2019. URL https://arxiv.org/abs/1805.11604. [40] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization, 2016. URL https://arxiv.org/ abs/1607.06450. [41] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization, 2017. URL https://arxiv.org/abs/1607.08022. [42] Yuxin Wu and Kaiming He. Group normalization, 2018. URL https://arxiv.org/abs/1803.08494. [43] Pascal Mettes, Elise van der Pol, and Cees G. M. Snoek. Hyperspherical prototype networks, 2019. URL https: //arxiv.org/abs/1901.10514. [44] Sergey Ioffe. Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017. URL https://arxiv.org/abs/1702.03275. [45] Andrew M Saxe, James L McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv preprint arXiv:1312.6120, 2013. 24 [46] Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Yee Whye Teh and Mike Titterington, editors, Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9 of Proceedings of Machine Learning Research, pages 249–256, Chia Laguna Resort, Sardinia, Italy, 13–15 May 2010. PMLR. URL https://proceedings.mlr.press/v9/glorot10a.html. [47] John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12(61):2121–2159, 2011. URL http://jmlr.org/papers/v12/ duchi11a.html. [48] Matthew D. Zeiler. Adadelta: An adaptive learning rate method, 2012. URL https://arxiv.org/abs/1212. 5701. [49] Aston Zhang, Zachary C Lipton, Mu Li, and Alexander J Smola. Dive into deep learning. Cambridge University Press, 2023. [50] Jiaxuan Wang and Jenna Wiens. Adasgd: Bridging the gap between sgd and adam, 2020. URL https://arxiv. org/abs/2006.16541. [51] Charles George Broyden. The convergence of a class of double-rank minimization algorithms 1. general considerations. IMA Journal of Applied Mathematics, 6(1):76–90, 1970. [52] Roger Fletcher. A new approach to variable metric algorithms. The computer journal, 13(3):317–322, 1970. [53] Donald Goldfarb. A family of variable-metric methods derived by variational means. Mathematics of computation, 24 (109):23–26, 1970. [54] David F Shanno. Conditioning of quasi-newton methods for function minimization. Mathematics of computation, 24 (111):647–656, 1970. [55] Jorge Nocedal. Updating quasi-newton matrices with limited storage. Mathematics of computation, 35(151):773–782, 1980. [56] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9:1735–80, 12 1997. doi: 10.1162/neco.1997.9.8.1735. [57] Ward Cheney and David Kincaid. Linear algebra: Theory and applications. The Australian Mathematical Society, 110: 544–550, 2009. [58] G. W. Stewart. The efficient generation of random orthogonal matrices with an application to condition estimators. SIAM Journal on Numerical Analysis, 17(3):403–409, 1980. ISSN 00361429. URL http://www.jstor.org/stable/ 2156882. [59] Francesco Mezzadri. How to generate random matrices from the classical compact groups, 2007. URL https: //arxiv.org/abs/math-ph/0609050. [60] Kai Hu and Barnabas Poczos. Rotationout as a regularization method for neural network, 2020. URL https: //openreview.net/forum?id=r1e7M6VYwH. [61] Wallace Givens. Computation of plain unitary rotations transforming a general matrix to triangular form. Journal of the Society for Industrial and Applied Mathematics, 6(1):26–50, 1958. doi: 10.1137/0106004. URL https: //doi.org/10.1137/0106004. [62] Adelaide P Yiu, Valentina Mercaldo, Chen Yan, Blake Richards, Asim J Rashid, Hwa-Lin Liz Hsiang, Jessica Pressey, Vivek Mahadevan, Matthew M Tran, Steven A Kushner, Melanie A Woodin, Paul W Frankland, and Sheena A Josselyn. Neurons are recruited to a memory trace based on relative neuronal excitability immediately before training. Neuron, 83 (3):722–735, August 2014. [63] Lingxuan Chen, Kirstie A Cummings, William Mau, Yosif Zaki, Zhe Dong, Sima Rabinowitz, Roger L Clem, Tristan Shuman, and Denise J Cai. The role of intrinsic excitability in the evolution of memory: Significance in memory allocation, consolidation, and updating. Neurobiol. Learn. Mem., 173(107266):107266, September 2020. [64] John Bridle. Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. In D. Touretzky, editor, Advances in Neural Information Processing Systems, volume 2. Morgan-Kaufmann, 1989. URL https://proceedings.neurips.cc/paper_files/paper/ 1989/file/0336dcbab05b9d5ad24f4333c7658a0e-Paper.pdf. [65] Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Christopher J. Cueva, Erin Grant, Iris Groen, Jascha Achterberg, Joshua B. Tenenbaum, Katherine M. Collins, Katherine L. Hermann, Kerem Oktar, Klaus Greff, Martin N. Hebart, Nathan Cloos, Nikolaus Kriegeskorte, Nori Jacoby, Qiuyi Zhang, Raja Marjieh, Robert Geirhos, Sherol Chen, Simon Kornblith, Sunayana Rane, Talia Konkle, Thomas P. O’Connell, Thomas Unterthiner, Andrew K. Lampinen, Klaus-Robert Müller, Mariya Toneva, and Thomas L. Griffiths. Getting aligned on representational alignment, 2024. URL https://arxiv.org/abs/2310.13018. 25 However, in Wang and Wiens [50] their optimiser is ADAM-like with isotropic definitions. This AdamSGD algorithm is displayed in Eqn. 51. This may be a starting point for a more optimal isotropic adaptive optimiser. θt+1 =θt−ηtmt ηt=ηs1−βt 2 vt/dim θ vt=β2vt−1+ (1 −β2)∥gt∥2 2, v0= 0 mt=β1mt−1+ (1 −β1)gt, m0= 0 (51) Alternatively, a return to an inverse-Hessian approximating quasi-Newton algorithms such as BFGS [ 51 , 52 , 53 , 54 ] or L-BFGS [ 55 ], may be another approach. Gradient clipping is another operation that requires consideration, as its anisotropy can produce a representational bias. The author is actively researching both directions. A.4 Operations: Several new operations can also be defined within the Isotropic framework; these include min-like and max-like functions, as well as a new multiplication operation. A link demonstrating these is available at https://www.desmos.com/ calculator/ttmvis7av4. The standard minimum function and maximum function are displayed in Eqns. 52 and 53 respectively, for a function f:RN×RN→RN. f(x,n) = min (x,n) = N X i=1 min (x ·ˆei, n ·ˆei) ˆei(52) f(x,n) = max (x,n) = N X i=1 max (x ·ˆei, n ·ˆei) ˆei(53) These functions are applied elementwise to components of x and n as indicated by the sum over standard basis directions, ˆei . This basis dependency is an example of an inductive bias which affects representations. These functions have two arguments, which require generalisation of the framework to multi-argument cases. One way to generalise such an isotropic equivariance is by applying the group action to both arguments in a modifed equivariance relation: Rf (x,n) = f(Rx, Rn) , for all R∈O (n) . Accordingly, one can define Eqns. 54 and 55 as functional forms which follow this relation. In the case where n ·x = 0, the operation should be defined as the identity map. f(x,n) = min sign (x ·n), n ·n x ·nsign (x ·n)x (54) f(x,n) = max sign (x ·n), n ·n x ·nsign (x ·n)x (55) These operations are a suggestion for such an isotropic alternative, but may not be optimal. They clip vectors which cross a hyperplane boundary defined by the choice of n ∈RN ; the hyperplane equation is: x ·n =∥n∥2 2 . The minimum function preserves the coordinates of all samples on the origin side of the hyperplane and projects all other coordinates onto a hyperplane. This projection is carefully chosen so that the projected point remains on a line passing through the origin and the original coordinates. Consequently, neural refraction does not occur on the plane boundary for origin-intersecting lines, like it does in the standard functions. The maximum function computes the opposing case, where points are kept constant on the far side of the hyperplane, which does not include the origin. In this case, points within the other sector are similarly projected onto the hyperplane. Similarly, Hadamard multiplication, also termed elementwise multiplication, is inherently basis-dependent. This can be seen through its basis dependence, through ˆei , in Eqn. 56. This is used in settings such as the LSTM gates mechanism [ 56 ], also one can reinterpret the dropout-mask in such a form [2], alongside many more applications. f(x,n) = x ⊗n = N X i=1 (x ·ˆei) (n ·ˆei) ˆei(56) One interpretation of Eqn. 56 is that it scales each component of x by each component of n , such that each axis in the standard basis is rescaled. This is equivalent to premultiplying x with a diagonal matrix W= diag (n)∈RN×N , where diag decomposes n along the standard basis and puts these components along the diagonal of a square matrix. This axis-wise rescaling is anisotropic, particularly a two-component algebraic permutation equivariance similar to before. To reproduce a similar behaviour, one could use isotropic-multiplication shown in Eqn. 57. This rescales the component of x which lies in the n direction by an amount determined by ∥n∥. f(x,n) = x + ((∥n∥−1) x ·ˆn) ˆn(57) These are just examples and may not be a simple drop-in replacement for existing operations, due to the differences in how the operations act. Initial anisotropic operations rescale in multiple axes/hyperplanes at once, whereas the isotropic operations 32 only act in single directions. This may make them suboptimal as gate mechanisms, such as in LSTMS, since they can only collapse the representations in one direction at a time: f:RN×RN→RN−1,→RN . These single-direction gates may be limited in how they reshape the representations and have more significant computational cost; therefore, further development is undoubtedly needed. These specific implementations may require a custom isotropic operation that fully replicates their desired effect in an isotropic and basis-independent manner. Finally, there may be some instances where representations must be clipped into a hypercube shape. This is not an isotropic operation, nor basis independent; however, this formulation does reduce neural refraction, so it is included for convenience. Using Eqn. 58 results in substantial neural refraction along the boundary of the hypercube: lines passing through the origin are projected across the boundary, substantially changing direction. f(x) = max (min (x, 1) ,−1) = N X i=1 max (min (x ·ˆei,1) ,−1) ˆei(58) Using Eqn. 59 achieves the same result without neural refraction at the boundary. Any coordinate which is outside the boundary is projected back onto the boundary, along a line between the origin and the original coordinate. As a result, direction is preserved. An affine transform can recenter this box at arbitrary coordinates and reshape the box due to the linear transform. f(x) = min 1 maxi(x ·ˆei),1x (59) 33 B Quasi-Isotropic Functional Forms A middle ground, following from Sec. 4, balancing the predicted problems of anisotropy whilst enabling representation compression, not just through the bias, is to relax the hard isotropy condition and introduce slight symmetry breaking in many directions. This can be achieved using many small perturbations to the direction unit vector only, producing a softer symmetry breaking. Then the network has many distinguished vectors, a subset with which it may align its representations in a task-dependent manner. Therefore, it does not favour a particular basis, but still introduces some desirable consequences of anisotropy, for problems such as classification. If one further restricts the functions from featuring dynamic refraction, it limits detrimental anisotropic effects. This recovers an arguably more principled and slightly increased basis-independent form of anisotropy. One method is to apply a non-linearity based on rounding the vector’s directions. This is shown in Eqn. 60, where [·] indicates the rounding operation and ϕ(x)=x . In fact, the anisotropic perturbation may be implemented as simply as: ϕ(x) = βx for β= 1 . The overall angular term is approximately unit-normed, but can be trivially modified to be exactly norm-1. Φ (ˆx;α) = [αˆx] α+ϕˆx−[αˆx] α≈ˆx(60) This produces a quasi-isotropic functional form shown in Eqn. 61, with an isotropy-breaking parameter α . Slight anisotropic refraction is added, independent of magnitude, so it is predictable and thus extrapolatable to the network. Due to the angular rarefaction and compression by the proposed non-linearity, representation overand underdensities may then occur, where semanticity may begin to be assigned. However, for α→ ∞ , isotropy is continuously reintroduced and could be an optimisable parameter. Such an approach may be beneficial to contrastive learning methods [36, 37]. f(x) = σ(∥x∥) Φ (ˆx)(61) Similarly, one could construct Eqn. 62 as another form of quasi-isotropic activation functional forms. In this case, a finite set of unit-vectors ˆ bi is distributed over a lattice across Sn−1 in Rn and a discrete rotation group transforms between various of these unit-vectors. This method may be less computationally efficient than the aforementioned method due to the numerous dot-product evaluations; however, it remains faithful to the Ψn hierarchical group structure. Moreover, one can recover a Bn symmetry without neural-refraction with this method, which may be desirable. f:Rn→Rn, x 7→ f(x) = σ(∥x∥) ˆx max ˆ biˆ bi·ˆxi(62) It is likely that many other functional forms could be considered. B.1 Parameterised Probabalistic-Isotropy Following from Eqn. 62, an alternative activation functional form can be constructed which appears similar but uses trainable ˆ bivectors, which notationally can be stacked into a matrix ˆ Zij ∈Rm×n. Following an appropriate initialisation of trainable parameters ˆ Zij , as discussed in App. A, the functional form can become weakly-isotropic. Yet it can undergo spontaneous symmetry breaking to reproduce desirable behaviours of anisotropy whilst preventing phenomena like neural refraction. Whether the denominator terms ˆx and ˆ Z should be unit-normalised can be evaluated, as well as the specific use of a max function. f:Rn→Rn, x 7→ f(x) = σ(∥x∥) ˆx max iˆ Zij ˆxj(63) 34 C Stochastic Isotropy — Producing Immediate Anisotropic Analogues One method to approximate isotropy with current functions is to stochastically choose a basis on which the anisotropic function operates. This enables anisotropic functions to be used in an isotropic network, without inducing a representational alignment to an arbitrary basis. The example is a training-time probabilistic condition; each batch exhibits spontaneous symmetry breaking, which ‘averages out’ over multiple batches. This approach can be generalised, and its overall utility assessed. For example, current (anisotropic) dropout by Srivastava et al. [2] appears to privilege the basis anti-aligned with the standard basis, thereby maximally preserving information when a direction of the standard basis is collapsed. So it is expected to incur an arbitrary basis dependence in the representations. However, anisotropic dropout can be applied to a stochastically chosen basis. This randomness is hypothesised to prevent a representational-anisotropy induced by an arbitrarily chosen basis. This can be achieved by producing a basis, uniformly drawn from the layer’s orthogonal symmetry: B∼SO (n)10 . There are several methods to produce a uniform random matrix, each with varying computational costs. These include the exponentiation of a Lie generator scaled by an appropriately drawn random variable, the Gram-Schmidt procedure [ 57 ], and many others [ 58 , 59 ]. The below matrix-multiplication procedure is computationally cumbersome; simpler formulations may be desirable. Hence, the following proposed forms are only a starting point for converting anisotropic functions directly to probabilistically-isotropic forms. In practice, isotropic functions should be constructed from the ground up rather than merely analogous functions converted from existing anisotropic ones. Therefore, this remains a placeholder, but with some interesting extensions. For the example of standard dropout, shown in Eqn. 64, it can be made stochastically-isotropic by including the basistransform shown in Eqn. 65. Where x is the activation vector, with normalisation factor Sa and Ssi , dropout-mask Mi= M·ˆei and standard-basis vectors ˆei. x′=Sa N X i=1 Mi(x ·ˆei) ˆei(64) x′=Ssi N X i=1 B M·ˆei(x ·ˆei) ˆei(65) Similar formulations, such as RotationOut by Hu and Poczos [60] , generally improved performance when the basis is stochastically rotated. However, this method remains stochastically anisotropic as the rotations are generated through Given’s rotations [ 61 ], which is not uniform over the space of orthogonal matrices — a necessity for full stochastic-isotropy. Nevertheless, the implementation by Hu and Poczos [60] is somewhat encouraging. This procedure can be generalised and applied to any existing anisotropic function. Yet, as stated, it is generally preferable to construct an isotropic function from first principles rather than relying on stochastic isotropy, which may be computationally costly. C.1 Considering Correlating the Stochastic-Isotropy Correlating the random bases in time would be a curious extension, particularly to stochastically isotropic dropout. This may produce a time-like structure in the embedded activation distribution of a network. This also constitutes an initialisation-based symmetry breaking, since a random walk may not evenly cover the space in practice and its starting point may bias training. Nevertheless, this is a suggested extension that is considered interesting and informative. If one imagines a random walk in the Lie-algebra space for rotation matrices (no mirror action): SO (n)∋R(t+dt)= R(t)δR , with δR=er·n , with r being the corresponding (normalised) anti-symmetric generators for rotations and n ∼ N( 0, σIn)with 0< σ ≪1. This procedure results in a random walk of the rotation matrix at each time step. Following this, a time-correlated Bernoulli distribution can be defined. Beginning with  D(0) , divide up the layer of neurons into two sets: inactive In=ni| D(n−1) i= 0o and active An=ni| D(n−1) i= 1o . Then we have two hyper-parameters: the standard dropout probability λ and an overlap probability Γ , such that |An|q+|In|Γ = (|An|+|In|)λ — where q is not a free parameter. If |An|= 0 or |In|= 0 , then temporarily define q= Γ = λ . If not, one must prevent unnormalised probabilities as shown in Eqn. 66. Γ = max 0,max λ+|An| |In|(λ−1) ,min 1,min λ+|An| |In|λ, Γ (66) Leading to q=λ+|In| |An|(λ−Γ) . Then use one Bernoulli function across all active neurons using R|An|∋ D(n) (A)∼ BernoulliDist.|An|(q) likewise for inactive neurons R|In|∋ D(n) (I)∼BernoulliDist.|In|(Γ) . Therefore, correlating the inactive ‘neurons’ 11 across the time steps, whilst still introducing a degree of random dropout. Thus, the ‘basis of dropout’ undergoes a random walk at every time step, and ‘neurons’ are randomly chosen to be dropped from the network, with a differing likelihood if they were just previously dropped. The coherence time can be adjusted through Γ for the specific time-dependent task needed. 10 Generating SO (n) may be computationally simpler than generating O (n) due to the former’s connected nature. Moreover, in either case, there is no effect for dropout. 11 These ‘neurons’ do have a stochastic drift in their definition due to the random walk. In general, individual ‘neurons’ are an ambiguous concept in an isotropic network. 35 This creates a link between the stimulus’s presentation time to the network and the ‘neurons’ it alters, such that stimuli presented in a smaller time window perturb a similar subset of the network’s neurons. This may produce an encoding qualitatively similar to that found in human cognition, where neurons are thought to go through excitability cycles of slightly differing frequencies and phases. When the excitability is higher, information (engrams) preferentially encodes upon those neurons [ 62 , 63 ]. As groups of ‘neurons’ begin to decohere, some overlap remains, such that memories are interlaced if they occur within a temporal window of coherence. This potentially gives neural networks using isotropic dropout an advantage in time-series data. However, this is not suggested as a model of such neurological processes, only a similar behaviour in deep learning, which is made possible through isotropic choices. Similar correlations could be considered for anisotropic dropout, likely to have a similar resultant effect. 36 D Potential Applications Besides the proposed general applicability of the isotropic modifications, the following are some areas where they may yield significant performance benefits or enable desirable network behaviours. D.1 Isotropy In Transformers It is argued that isotropic deep learning may be a more appropriate inductive bias for deep learning. However, there may also be some architectures which are particularly enhanced by its inclusion. One of these is the self-attention step of transformers [10], where isotropic-tanh may be of particular benefit, in replacing the softmax operation [64]. Softmax is defined through elements being bounded between zero and one, f(x)·ˆei∈[0,1] and summing to one. Consequently, it is non-negative, and there are regimes where this may be a limiting factor. It forms a n−1 -simplex embedded in the Rnspace, normal to  1, therefore a degree-of-freedom is also lost: f:Rn→ △n−1,→Rn. It has been shown that representations can exist in an antipodal superposition [ 16 ], particularly when stimuli do not tend to coexist in samples; thus, antipodal arrangements can exist with minimal interference. Such a stimulus may be a continuous quantity, but its two extremes are mutually exclusive. Many of these semantics are present in the real world, such as daytime-to-nighttime, motion towards or away, and smiling versus frowning. These could be represented through a zero-to-one scale; however, a [−1,1] scale may be a better representation, with zero as a better neutral middle point. This is because, in the linear features hypothesis, the magnitude often indicates the strength of the stimulus’s presence. In this case, the negative of a semantic direction may be equally meaningful and present in varying amounts. It may be expected that enabling this behaviour within the self-attention step is favourable. Moreover, the sum-to-one case may not always be desirable: it always encourages a change to the semantics when considering the residual-step-modification. This may encourage a semantic correction to an activation in transformers, even when it is inappropriate, or optimisation may force the existence of a near-zero value vector to prevent corrections. The residual step only encourages an identity transform independent of the activation, in the linear layer, rather than the identity dependent on a particular activation. The self-attention step compares the pairwise similarities between several vectors grouped into the so-called ‘keys’ and ‘queries’. The degree of similarity then affects how much of another semantic is expressed: the ‘values’. However, the softmax layer is basis-dependent and prevents a negative expression of these value semantics. A more suitable choice may be isotropictanh . In analogy of its sum-to-one constraint, its vector-magnitude is at maximum one, 0≤ ∥f(x)∥ ≤ 1 , whilst elementwise its values are −1≤f(x)·ˆei≤1 . Hence, it can express a negative of the value semantic, or any scaling of it between −1 and 1 . This suggests that isotropictanh may be an appealing drop-in replacement for softmax in the attention step, at least conceptually. Its continuous rotational symmetry may also offer an advantage, since the underlying pairwise similarity of self-attention QKT=xTWT QWKxˆ=xTW′ kqx is also basis-independent in x for isotropic initialisations of W , which somewhat aligns with the principles of isotropic deep learning. Hence, removing further bases may enable a more even interpolation between, and perturbation to, the value vectors. Hence, an isotropic adaptation to a self-attention may appear as shown in Eqn. 67, which will be explored in future work. Attention (Q, K, V ) = Isotropic-Tanh QKT √dkV(67) However, this does not make transformers ‘isotropic’ as a whole, since there are many further anisotropic steps present. Nevertheless, single-layer isotropic adaptations remain compatible with a larger anisotropic network or radial-basis network. Therefore, one may hybridise these approaches if appropriate. In general, isotropy may not be applicable to current transformers, due to their likely selection upon anisotropic primitives. Hence, a ground-up approach may be generally required, whilst borrowing concepts from the established transformers. D.2 Real-Time Dynamical Network Topology An appealing feature of isotropic deep learning is the relation displayed in Eqn. 14, showing that due to rotational equivariance, a rotation to one weight matrix can be counteracted with the inverse-rotation of another, preserving the network’s function. Consequently, a gauge freedom is formed due to functional forms commuting with symmetries. With gauge freedom, a particular gauge that expresses the weights beneficially without affecting network functionality can be chosen. One such gauge expresses the parameters in a basis with a magnitude ordering of the singular values for the matrix rows/columns. One could then set a threshold for the singular value, and bias, to determine if each corresponding direction in such a matrix has a meaningful contribution to the overall functionality. If it is deemed to have negligible value, it can be pruned with little adverse effect on the network. Moreover, ζ latent neurons can be included, with zero-initialised singular values fully connected to existing neurons. These do not impact performance, but enlarge the activation space and parameters available. Since the Jacobians of the isotropic activation functions are not strictly diagonal, these latent neurons may be rapidly trained if required. Therefore, the otherwise static, fully connected network is now dynamic, growing and shrinking in response to task-necessitated demand, with minimal impact on performance. This is enabled through an isotropic functional form. It poses an interesting research direction, enabled by the continuous rotational symmetry available. Transfer learning and task-swapping may become more straightforward. The network may grow to accommodate new tasks, or prune to optimise the model and stabilise computation. For example, it may stabilise by removing near-negligible, but sometimes significantly non-zero effects, which may be maladaptive due to their infrequent 37 effect. Output and input neurons could also be appended and removed in such a way, allowing for real-time changes to a dataset, or even training on multiple datasets. Such a procedure could be extended to convolutional networks, allowing a dynamic number of kernels. Similar possibilities may exist for other architectures. It appears that it may sidestep the Lottery Ticket Hypothesis [ 34 ] in choosing the optimal network size before training. Due to the computational cost, this does not need to be computed at every step; instead, it can be performed periodically and layerwise. This could offer substantial insight into how parameters may be shared between tasks in real-time. For example, the author postulates that if a new dataset is introduced partway through training on a different dataset, there might be a short-term increase in parameters until the network’s parameter-sharing begins, followed by a pruning phase until a more compact architecture is reached. Questions such as these could motivate the development of a subfield offering insight into these research avenues. These network dynamics may be incredibly insightful. Overall, this enables task-driven, real-time neural plasticity in deep learning. Other continuous symmetry-primitives may enable similar approaches. D.3 Multi-Headed Layers Although not limited to isotropic deep learning, producing ‘multi-head’ feed-forward layers may be desirable, enabling perturbative-like corrections to activations at each layer. This could be achieved by summing over a series of activation functions in each layer. An example of this is shown in Eqn. 68 or a more general construction in Eqn. 69, for weight matrices Wjand W′j, biases  bjand  b′jand an activation function f:Rn→Rn. xl+1 =X j fWjxl+ bj(68) xl+1 =X j W′jfWjxl+ bj+ b′j(69) This is effectively a sum over several feed-forward layers, which could increase the expressibility of a layer. It may be especially beneficial for isotropic functions, such as isotropictanh , where each sum produces a further perturbative ‘correction’ to an output vector. A scaling could be enforced if an ordering of such perturbative corrections was desired. Furthermore, one could consider an alternative, but similar, structure based upon a fibre-bundle symmetry extension over these various heads. Each layer could constitute a base space, and the multiple heads could be considered a fibre, or vice versa. This can enable a definitional form for matrix-valued functions. Additionally, in such a case, one may also consider the heads as a vector representation over each base neuron, and restructure the parameter maps accordingly for computation between various vector-valued-like neurons. This generalisation may find applicability in certain specialisms, whilst the general structure above may have broader applicability. D.4 Isotropic Representations May Aid Semantic Alignment An emerging interdisciplinary field of semantic alignment [ 65 ] are trying to produce comparable representations between deep learning and the brain. The author believes it is worthwhile to investigate how the representations generated by the Isotropic Deep Learning approach can aid in achieving this objective. This is because anisotropic functional forms have been shown to create representational structure due to functional forms; this structure is not a naturally emerging consequence of the data [ 17 ]. These artificial structures may be detrimental to representational alignment objectives. Removing anisotropy may help with alignment methods that use continuous rotation-like, such as in the work of Williams et al. [66] , since the inductive bias of isotropy is equivariance to continuous rotation. However, this connection is largely speculative, but is included as a point of discussion and potential research avenue for isotropic deep learning. For isotropy, this may provide a testable route for the hypothesis that isotropic deep learning forms more ‘natural’ semantic structure in representations. However, this is not to suggest that the brain is likely to be isotropic, especially since the approach produces delocalised functional forms. In isotropic networks, neurons instead act as a collective and are arbitrarily decomposable into any set of individual neurons, due to the gauge invariance. Consequently, there is likely no identifiable and agreeable definition of a neuron in an isotropic network. This is fundamentally incompatible with the brain’s structures. On the other hand, observations of distortion in deep learning due to anisotropy do not imply that the brain also produces anisotropic and discrete features when time-averaging its neuron firings. Therefore, despite the incompatibility of functional forms with biological neuron behaviour, their representations may have substantial similarity, and it may be possible to produce better alignment between their respective activation distributions once anisotropic-incurred structures are removed from deep learning. A similar approach may be extendable to inferring meaning from languages where a large corpus is available. If isotropy produces a continuous representation, free from basis distortions, then one may expect a more structured and interpolatable semantic structure. One may speculate whether an approximately language-agnostic structure may develop in a similar analogy to representational alignment — the latter field is arguably assuming (and often encouraging) an approximately model-agnostic representation structure. Hence, there may be a chance that representations free from artificial anisotropies may aid in deducing an alignment between known and unknown vocabulary. This may be an interdisciplinary application of Isotropic Deep Learning. Smaller-scale empirical approaches could begin with a hybrid of the two methods to establish whether embedded activation representations can be aligned between other simpler biological systems’ neural activity or vocalisations, and similar domained deep learning models. This may give insight into the approximate semanticity of some of these. Similar can be tried for various models within deep learning, determining if convergence onto one, or several, universal-like representations occurs within a similar domain. 38 Despite this, the success of ensemble models may limit this alignment. Ensemble models use diverse individual models to collectively approximate a task solution. The model diversity would suggest that there are more minima than those connected through a permutation or continuous symmetry, which would not yield diverse individual solutions, as they are functionally identical. Therefore, there are likely functionally diverse models with substantially different internal representations. This may challenge assumptions of a universal and comparable representational structure for semantics. Despite this, they may all produce different approximations to a universal representation. Substantial testing will elucidate this, and the basis-undistorted representations produced by isotropic networks may be particularly beneficial in such an endeavour. 39 E Comparisons With Geometric Deep Learning Approaches E.1 Distinction from Equivariant Networks Geometric Deep Learning and this Primitive-First approach both consider deep learning through the use of symmetries and group-theoretic lenses. Naturally, due to this, several implementation, formulaic and terminological convergences occur, and such interdisciplinary considerations and understanding of their consequences may be practically advantageous in general. For example, similarities include the (algebraic) invariant and equivariant relation in their definition, which extends to other group-theoretic usage, representation theory, Lie Groups, Gauges, etc., which can make them appear similar due to a incommon formalism. Yet, these perspectives on symmetry in deep learning remain distinct in their intended purpose, the resultant consequences for applications, their broadly different regimes of consideration, their independent motivation for such symmetry principles, and their differing potential impacts and implications for the field. Several places in this work already briefly outline overlaps and dichotomies, particularly regarding taxonomic unification. This section aims to provide an extended outline of the parallels and divergences between these philosophies, establishing a holistic picture of these complementary considerations of symmetries in deep learning. Primarily, the guiding philosophy of Geometric Deep Learning is to ascertain which symmetries (specifically the algebraic symmetries in the taxonomic terminology) are inherent to a given dataset and in the intended application. Using this, one then constructs models which are guaranteed to respect this data-derived structure. A clear example of this approach is the broad array of equivariant networks that have been developed, and several of these in particular will be discussed further below to compare similarities and differences. These have required the construction of individual maps to further this model-scale equivariance when composed together. This does create some overlap, specifically within algebraic considerations, but it is primarily a top-down approach — it starts with model-scale constraints and recursively applies them downwards to ensure it is retained over all the functions and their compositions. For primitive-first reformulations, it starts with determining the implications of primitive functions as foundational biases and their compositional biases, which are expected to be hierarchical in nature. This is its initial line of enquiry before leveraging such findings generally. This is not just pertinent to activation functions but extends over primitives more generally: optimisers, normalisers, operations, initialisations and many more foundational maps. These can be collated into sets defined by particular symmetries through a taxonomised system of three generations and three flavours to indicate the strength/degree and type of symmetry categorising each map — this merges into the broader taxonomic approach. In this, Geometric Deep Learning occupies the algebraic sector in furtherance of model-scale symmetries. This is not the sole constraint nor scale-remit of the primitive-first approach, for which the whole taxonomy is relevant at all scales. This approach is the suggestion that these primitive characterising symmetries may have important implications internally for networks. Their form may have direct and indirect implications for the internal dynamics of networks, through representations and learning dynamics, which, once suitably understood, can be leveraged in more general applications. It is hypothesised that such symmetries may be important and generalising analytical quality of many primitives for each set 12 . Hence, the development and investigative audit of all such primitives and their consequences, followed by the rebuilding of general architectures for applications, is the primary guiding principle for this work. Hence, it constitutes a bottom-up approach which is less predicated on specific data-driven structures. To extend this brief comparative analysis, this section will continue with a discussion and outline of the several key approaches within Geometric Deep Learning, namely Equivariant Group-Convolutions [ 23 ] and discussion of Steerable-CNNs [ 24 ], Harmonic networks [ 25 ] and Spherical-CNNs [ 26 ]. Following this is a summary of the similarities between these methods and the primitive-first approach — namely, one can produce some alignments when considering only algebraic symmetries of the taxonomy. Finally, crucial differences will be highlighted to demonstrate the distinctiveness of these approaches in general, particularly in terms of symmetry in deep learning. Overall, this indicates distinct but potentially complementary group-theoretic approaches. An outline of Cohen and Welling [23] ’s Group Equivariant Networks is to utilise a specific symmetry, particularly the one expressed in the underlying task’s data-structure domain, and ensure the network as a whole respects the task-relevant symmetry through use of a modified convolution operation: Group Convolutional Neural Networks (G-CNNs). This is generalising the traditional translation equivariance of the convolution operation (ignoring edge effects) to instead be equivariant to a general discrete group G . This symmetry group is chosen a priori by considering the given dataset, so the approach is to leverage the known symmetries of the task as a strong and constraining inductive bias to ensure accurate desired solutions. Consequently, this symmetry-respecting constraint is applied end-to-end over the whole network architecture. This can have a wealth of benefits, including increased weight-sharing efficiency, physically accurate modelling, and a resultant increased expressive capacity. A summary of this design procedure is: identify if the specific task has its data distributed over a particular linear base space and if it is expected that applications upon this data are expected to follow a symmetry of that space. If so, then the data can be ‘lifted’ onto a symmetry group acting on that space, and an associated equivariant model can then be used to achieve the intended application. The connection to symmetry can be denoted for data and general activations by the following map: f:G → Rn — every element within the group is assigned an n -dimensional vector by f , such that the resultant vectors are interrelated through the intended group action. For Cohen and Welling [23] , their convolution operation then preserves group-structure in its convolution map for discrete symmetries, producing new features that are still structured over the group. These can then be stacked to form a model which respects the end-to-end symmetry. 12 Although it is recognised that other analytical qualities specific to individual implementations also contribute, so to some extent could be considered an ‘effective theory’. 40 The group-convolution is implemented as a modification of the classic discrete convolution operation: by applying the filter over the group, as shown in Eqn. 70. In Eqn. 70, k indexes the filter ψ and notably the sum is over the base space X in the first layer: h∈ X , not h∈ G . Then, the resultant equivariant group convolution respects the symmetries of the task at every layer, given by the aforementioned data structure f:G → Rn. A naive implementation results in augmenting the number of filters to accommodate every action of the group; however, crucially, these are shown to be related through permutation, so can be achieved more efficiently in practice by an indexing that exploits the group’s structure. For more details of precise implementation, see Cohen and Welling [23]. [f ⋆ ψ] (g) = X h∈G X k fk(h)ψkg−1h(70) Overall, this adapts the convolutional operation and padding to respect the symmetries in the underlying data structure. It is also shown that several existing primitives commute with these group actions. Particularly, the existing elementwise non-linearities commute with the considered group actions, and hence must be retained in their current functional form to ensure end-to-end adherence to the symmetry. In other cases, only small modifications to other primitives, such as those specified for normalisations, need to be made, which still enables them to retain their current functional form. Hence, in this approach, primitive reformulations are potentially detrimental to breaking the end-to-end equivariance of the construction — in this case, retaining the current elementwise form is important for commuting with the group action. In short, the current primitives are retained. In the subsequent works of Cohen and Welling [24] , Worrall et al. [25] , and Cohen et al. [26] , considerable progress is made in developing models capable of a greater range of equivariance symmetries through extending the architectures and tools. In Cohen and Welling [24] , the authors build upon earlier work by creating steerable capsules, in which vectors transform under irreducible representations of the discrete group. These steerable filters are constructed as linear combinations of base filters, resulting in a more parameter-efficient design. In [ 26 ], they generalise these concepts for images over a spherical shell, S2 , and lift it to an SO (3) continuous symmetry equivariance, using a Fourier transform-like method. Through these and others, networks are made algebraically equivariant to discrete group transformations and extended to specific continuous group transformations. This subfield is highly active and rich with many other successful discoveries along the same vein of data-driven symmetry approaches for the model scale. The key differences can be clarified with the stated examples. Returning to Worrall et al. [25] , the authors use the steerable filters to construct equivariance to continuous patch rotation using finite filters constrained onto circular harmonic functions exhibiting the desirable rotational equivariance. This develops into complex-valued activations and maps, where the initial input data becomes the real part in the complexification. To maintain the rotational equivariance in the harmonic network, an activation function is introduced which must act upon the complex-valued activations but is also constrained to ensure rotational equivariance. The result is an activation function which acts on the absolute value of the complex number elementwise. This instance of an activation function which applies over the absolute value, can be considered to constitute a single activation function of the algebraic Sn×U (1) branch, and if expressed multivariately could be manipulated into the functional form such as Eqn 29. This indicates that application-driven instances of primitives do occur and could be retrospectively classified through the taxonomic considerations of this paper, and particularly strictly as algebraic generations. However, these also emerged in the context of a primitive that must further the broader model’s constraints, so consideration of their foundational biases is not discussed in relation to its wider impact on general networks outside of model-scale equivariance, which contributes further divergences in philosophy. Building upon the Harmonic network, Thomas et al. [67] utilise a more generalised instance intended for geometric tensors which enables equivariance to local rotation, translation and permutation, extending the prior work of Harmonic networks for various geometric tensors. As a necessity, the non-linearity must act on the tensors in a manner which is not changed under the specified group actions, for rank-1 tensors this is manifestly a form of norm-based activation function instance to ensure equivariance, drawing some parallels with instances of isotropic activation functions but only for tensor objects, meanwhile for rank-0 (scalars) tensors this reduces to the standard elementwise activation function of the typical anisotropic forms and higher rank tensors are similarly demonstrated. These choices ensure that the non-linearity does not break the equivariance to these actions. Overall, the shared language of group theory emerged naturally in both approaches as a result of their respective objectives. One uses it in ensuring model-scale adherence to a specified algebraic symmetry group lifted from the data, whilst the other considers sets of primitives with respective function-driven foundational and compositional biases. These do sometimes constitute overlapping considerations and potential for shared tooling, mostly limited to instances of activation functions for specific algebraic symmetries. Yet, they remain tangent in both their motivating purpose, particular versus general applications and the consequences of symmetry in a network. Although the primitive reformulation of deep learning is of geometrical and deep learning construction, it does not appear to sit cleanly into the current field of geometrical deep learning [ 27 ]. Instead, it is constructed around the geometry of embedded representations, altering the internal symmetries of general networks rather than a network-wide externally applied symmetry instilled by a predominantly task-driven inductive bias. In many ways, the primitive first approach could be considered the consequences and leveraging of symmetry breaking in many circumstances, where a network is not enforced to be end-to-end symmetric and instead these actions can be thought to act on the network and influence its behaviour in unintentional and peculiar ways. These are then conjectured to reemerge as a plethora of scattered phenomena discussed. A further discussion of the differences is provided below. There are several more distinct differences in the approach, particularly concerning where symmetries arise, the distinct inductive biases considerations, and how these may influence vector spaces. These are further detailed below. 41 Local coding, as discussed in App. F, is characterised by a one-to-one correspondence between a neuron and a semantic. Therefore, each neuron’s activation is associated with the prevalence of a distinct semantic concept. In contrast, distributed codes have semantic meaning that is dispersed across a population of neurons. Consequently, semantics no longer tend to align with individual neurons. There are conveniences to local coding; its simplicity makes networks very interpretable, a benefit for both AI safety and diagnosing network pathologies. It is therefore often given as an intuitive first-order approximation to the action of deep learning models. However, this first-order heuristic may inadvertently suggest that such inductive biases are benign — a position this paper challenges. Modulation of distinct semantics is a desirable action. Operations such as bounding the strength, rectifying the signal, and distorting aspects of its stimulus to neuron-response curves may all be beneficial semantic actions for a network, altering its internal expressions and enabling better interactions between concepts. These actions can all be achieved using the application of an activation function: Tanh and Sigmoid, ReLU [ 15 ], or Leaky-ReLU [ 15 ], Swish [ 13 ] and SiLU [ 89 ], respectively. Alongside comparative and logical actions between semantics, such as those achieved by Softmax or gate mechanisms [ 56 ], respectively. If each neuron encodes a single semantic feature, then under this assumption, it is appropriate to apply such operations elementwise, independently acting on each neuron to scale their semantic representation. Under a local coding assumption, elementwise operations suffice in modulating and enabling interactions between entire semantic features, since each neuron corresponds to an independent semantic. An additional influence may have been the expert system approach, where fixed if-then logic produces a condition on each meaningful quantity to produce the desired output. This methodology may also have influenced the on-off activation function approach, such as the Heaviside step function. This approach aligns conceptually with local coding. A consensus is emerging that whilst networks often tend to this local coding [ 31 ], the representations are often more nuanced in practice [ 31 , 88 , 16 ]. It is argued that the network may balance local coding with interference and representational capacity needs [ 16 ]. This paper and Bird [17] ’s work have further explored the issue by implicating functional forms as responsible for the formation of this coding tendency. Under these more generalised distributed codes, single activation actions and comparisons are insufficient for interacting with whole semantics. Such operations may only distort the fraction of the semantic, represented through a single neuron. This underpins the neural refractive problem. Instead, a distributed logic is required. Instead of assigning operations upon numbers as components of arrays, it is suggested to consider a full vector space treatment (inner-product space), with magnitude and directions as foundational quantities rather than components. It is argued that operations should act on these, more fundamental and basis-free quantities. This is reinforced by the changing affine transforms — due to parameters adapting during training or random symmetrybreaking initialisations. Therefore, vector directions may be unpredictably distributed and change rapidly during training, causing activations to move around the representation space. As activations evolve around the representation space through training, so too do the semantics they individually and collectively represent. Therefore, activation functions as semantic moderators should not be chosen to only act upon single neurons, as semantics are not wholly expressed by individual neurons. Especially, since the tendency towards symmetry-broken local coding only emerges after considerable training [ 17 ]. Moreover, this is compounded by lasting polysemanticity in specific neurons [31, 16] even after training. Since semantics are often found to be distributed in the representation space, and adapt through learning, without a predictive theory of their trajectories, a sensible inductive bias is to apply the modulation isotropically. Hence, ensuring that any linear feature, including those off-axis, are modulated consistently by isotropic forms. This respects both distributed coding schemes and polysemantic neurons. Furthermore, the application of functions isotropically still correctly modulates locally coded semantics since their functionality can be identical along the standard basis. Thus, applying non-linearities exclusively along the standard basis is inconsistent with the goal of manipulating internal semantic meaning in distributed codes. Removing such an inductive bias is expected to remove this local coding bias, enabling more distributed representations with increased representation capacity whilst balancing concept interferences. In conclusion, the intuitive local coding approach may have had some encouragement on the elementwise functional form used in contemporary deep learning; however, relaxing this inductive prior to a distributed neural code suggests that isotropic activation functions may be considerably more appropriate. Had distributed coding been prevalent during the early developments of deep learning, then isotropic functional forms may have emerged as the default paradigm: modulating vector magnitudes instead of standard components decomposed on an arbitrary basis. This could have been seen as a more biologically inspired computational model. While exact Isotropic forms are likely biologically implausible, due to constraints of individual neurons with localised responses, such limitations do not constrain artificial systems. Hence, Isotropy may be considered a better inductive bias to approximate coding intuitions. 48