Full text
On The Ontology of Individuated Neurons: Symmetry-Reformulations Enable Dynamic Topologies George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] October 10, 2025 Abstract This paper primarily leverages a new symmetry-principled class of primitives to demonstrate architectures that can grow and shrink in real-time with task demand. This is made possible by structural changes, connected by symmetry, which leave the computation invariant under neurogenesis and nearly identical under neurodegenesis. Using such ‘isotropic’ primitives, the notion of an individuated neuron is lost, allowing freedom in how layers are represented and interpreted. Permitted from this a formal procedure for layer-wise diagonalisation, where typically interconnected layers, such as dense layers, convolutional kernels and more, can be reexpressed such that neurons have a one-to-one ordered connectivity. This indicates which one-to-one neuron communications are strongly impactful on functionality and those which are not. Inconsequential neurons can be removed, and new ‘scaffold’ neurons added, whilst remaining analytically functionally invariant. A new tunable model parameter, ‘intrinsic length’, is also introduced to aid in this analytical invariance. This approach makes connectivity pruning equivalent to neurodegenesis. This diagonalisation also offers new possibilities for interpretability into such networks, and it is demonstrated that isotropic dense networks can asymptotically reach a sparsity factor of 50% whilst retaining exact network functionality. The work includes investigations into integration and separation within networks that respond to both differing and joint tasks. Finally, the construction is generalised, demonstrating a nested functional class for this form of isotropic primitive architectures. 1 Introduction Neural plasticity in the quantity and connectivity of neurons is biologically advantageous [ 1 , 2 ] and routinely occurs in the brains of mammals [ 3 , 4 ]. It can improve neural efficiency through pruning [], whilst enabling the accumulation of knowledge [], robustness [] and function through growth []. It is therefore hypothesised that analogous behaviour within artificial neural networks may be similarly beneficial. This may range from improved efficiency to reduced anomalous computations to increased capacity — overall, a real-time restructuring for knowledge acquisition and stabilisation. The practical consequences of a network that can adapt its architecture and connectivity in real-time would likely be transformative. Proposed is a novel methodology that leverages an alternative formulation of primitives, known as “isotropic primitives”, to achieve real-time restructuring of networks with minimal degradation in functionality. This is achieved by exploiting a symmetry-structural functional invariance of a network. This both indicates the flexible use cases of these alternative formulations and enables the exploration of model adaptations to new information. Typically, one may define a neural network system as a computation predicated upon individual units intercommunicating to yield an overall desirable learnt function. This situates neurons, with their interconnections and computations, as ontologically fundamental building blocks to this computation system. In other words, this model is constituted by and defines such delineated artificial neurons, at least thus far. These then communicate and adapt to reproduce the desired function through collective computation. Currently, the design core of this approach is the agreeable separability, and hence individuality, of those neurons, which serve as the atomic units for forming ever-larger networks. These units interplay to produce the broad functional class of artificial neural networks we employ today. Mathematically, this individuality is conferred by the contemporary choices of network operations, where the individuality of neurons is assumed into implicit constraints on the possible functional forms. This affects most primitives in widespread use, including, but not limited to, the majority of: activation functions, normalisers, optimisers, several regularisers, and gradient clippings — all of which collectively reinforce this neuron-wise perspective. The question may arise: why do we often choose to apply our activation functions elementwise? 1
Extending from this current construction, one can work deductively, determining the computational equivalences [] inherent to most networks. Exhibited by such models are parameter-space degeneracies, usually centred around the exchange of neurons — formally termed a permutation symmetry, Sn , of the neurons in the network. Under appropriate reparameterisations, swapping neurons and similarly exchanging corresponding weights and internal biases, can be considered a computational invariance of the network 1 . In other words, the network functions identically before and after the exchange by permuting the actions. Hence, a permutation symmetry is deduced from this current construction. It fundamentally arises from the implicit constraint of a neuron-wise functional form for network operations, stemming from the notion of what constitutes a neural network model. However, this work argues that such a perspective on individuated neurons and their resultant constraints on functional forms 2 is limiting. The functional class of all artificial neural networks is confined based on this assumed fundamental individuality. Where these contemporary primitives’ forms have potentially stemmed from historical naturalistic ideals, reinforced later through hardware alignment, positive performance precedent, and some cognitive entrenchment [ 6 ] — often appealing to their biological counterparts for intuition and inspiration, where the individuality of cells is a physical limitation. Yet, by abandoning this implicit constraint on functional forms, this broadens the possible functional class. It will be shown that the behaviours and flexibility of neural network models are notably enriched, in addition to recently established representation changes [7, 8]. Fundamentally, this is the generalised approach surrounding ‘isotropic deep learning’ and associated primitivesymmetry redefinitions [ 6 ]. These proposed redefinition approaches effectively constitute an ontological reversal of the neuron’s priors. Instead of current neurons as the fundamental constituents, which implicitly define this computational approach and from which permutation-like symmetries emerge and are deduced, the foundational base will be inverted. Hence, instead of working deductively from neurons leading to symmetry, this work proposes to reconsider the observed relation in reverse: symmetry leads to neurons. Current networks then emerge from one symmetry: the permutation group 3 , producing a conventional notion of a neuron. Yet this reversal allows other symmetries to take permutations’s place, generating new forms for primitives and, hence, the induced notion of neurons. Philosophically, this situates primitive symmetries as foundational and hence ontologically prior. Neurons, as a notion, are then derived/emergent from them, a qualitative concept resulting from the mathematically altered functional forms. Each primitive symmetry can then be definitional for each branch of a new approach. A new set of primitives for each foundational symmetry and linked through them. Comparable in intuition to the substantively similar physical ontological inversion which occurred between individuated particles and symmetry-defined fields. The ramifications of this redefinition are argued to be broad: from a generalised notion of an artificial neural network computational system that extends beyond its biologically-inspired premise, to new sets of primitives and consequential, leveragable behavioural and representational changes for networks. The functional class of artificial neural networks is broadened through such a shifted ontology. Symmetries are no longer derivative or confined to functional invariances under reparameterisations, or implicated as an externalist inductive bias on a model’s construction, but become internally foundational to the computational system. Moreover, the notion of individuated neurons becomes an emergent property of permutation definitions in the primitives currently used. However, using this alternative perspective, one can instead consider neurons as a unit extending from the permutation definition, rather than the other way around. As stated, a natural next step is to substitute alternative symmetries, yielding new sets of primitives and, consequently, new notions of what a neuron even is. This work leverages one such alternative symmetry-defined branch in particular. This work explores the ‘isotropic’ redefinition, an approach predicated on the continuous orthogonal superset, O (n) , of the preexisting discrete permutation group, Sn⊂O (n) . From this, primitive sets emerge that are orthogonally equivariant to standard representations of these actions [6]. From this ontological inversion, one can work deductively to rediscover emergent neurons and leverage new behaviours. For isotropic networks, the notion of a neuron is more general. Typically, we may consider a layer construction as n copies of 1 -dimensional neuron-like objects concatenated, due to the permutation symmetry; however, under the orthogonal symmetry redefinition, this shifts to considering a layer as 1 copy of an n -dimensional neuron. The fundamental qualitative ‘neuron’ is higher-dimensional and can be decomposed into any linear combination, provided it is normalised, to yield equally valid alternative individuated bases. Crucially, in this approach, there is no longer a clean, individuated and agreeable definition of a neuron. There effectively exists a (global) gauge freedom to choose whatever decomposition of individual neurons one likes; there is no distinguishing mathematical boundary delineating each unit — shifting the modular object from a neuron to a layer. Any choice of individual neurons is no more fundamental than any other; they may shift arbitrarily in definition due to their continuous orthogonal nature. Hence, varied decompositions do not alter the computation when interspersed between affine maps, yielding familiar computational invariances under affine composition. Moreover, this freedom to change the basis of the representation 1 This is also connected to the symmetries of the discrete directed graphs underlying the communication pathways between models predicated on discrete, individuated neurons. These will later be generalised continuously in an upcoming manuscript [5] requiring completion. 2Which likely cyclically reinforces this perspective 3And associated permutation-like groups, such as hyperoctahedral Bn=Zn 2⋊Snamongst others 2
space extends to circumstances in which one can project into a higher-dimensional space without loss of functionality (a neurogenesis ability). Similarly, under more restricted conditions 4 , one can down-project to subspaces with minimal functional consequences or catastrophic failure, providing the network with neurodegenerative/pruning ability — all of which are enabled through this symmetry-principled approach. What this amounts to is a new form of computational invariance to leverage with isotropic redefinitions — one in which the architectural structure becomes plastic and adaptable to tasks through reparameterisations. These are structural automorphisms for networks. This contrasts with the contemporary permutation definition, which is limited to exchange symmetries which do not alter the structure. Primitive symmetries, neural architecture and functional invariance can become desirably intertwined. Hence, this new definition of primitives enables networks to expand and contract dynamically in response to demand. This is considered a beneficial consequence directly enabled by the orthogonal reformulation of the network’s basic constituents. This paper explores such reparameterisations, allowing neurons within a layer to compensate for alterations to the structure while maintaining functionality. One can also generalise this beyond isotropic functions to any continuous group-defined primitive sets, unlike the traditional discrete-defined permutation primitives. Although briefly featured, the approach to be discussed does not primarily reside within the approach of sparsifying a preexisting architecture [ 9 ]. Rather, it can best be described through altering both neuron number and connectivity, which are analytically ensured to reproduce the function under certain reparameterisations. Hence, it aligns thematically with other works where the central objective is to truly alter network topology whilst preserving functionality [ 10 , 11 ], rather than masking effectively preexisting graph connections. However, except for the unifying objective, it is distinct both mathematically and in implementation from prior approaches, such as diverging through diagonalisation and a thresholded singular-value-based approach. Hence, it contrasts with a static-nodal underlying directed graph method, allowing a dynamic nodal graph. A more thorough comparative analysis between prior works is given in App. A 2 Theory This section outlines how a network can be diagonalised to enable growth and pruning under various reparameterisations and conditions. Moreover, considering such an approach allows for a network to be sparsified asymptotically to 50% of its original number of weights, without any loss in functionality. First, a general introduction to isotropic activation functions is given. These considerations also apply if isotropic normalisations are used. Following this, a layer-wise basis change to a diagonalised configuration via singular-value decomposition will be described, along with a more efficient partial diagonalisation procedure for training-time. Concluding the theory section, it will be discussed how the network can undergo growth and pruning for each layer, and a brief discussion of applications. 2.1 Isotropic Activation Function Overview Isotropy is most generally an orthogonal group family constraint applied to form a set of primitives. This generalised constraint can form activation functions, normalisers, initialisers, and optimisers with the same underlying symmetry properties [ 6 ] and argued network-wide implications, particularly on representations [ 7 , 8 ]. These primitives then transform under a particular, typically standard, representation of the orthogonal group specific to the width of the layer. In this section, activation functions and normalisers are the primary concern5. For the primitives in this scenario, and for an n -width network, the standard representation is given by ρ: O (n)→ GLn(R) and each element will be notated as a matrix R=ρ(g)∈GLn(R) where g is an element of the orthogonal group, g∈O (n) . This is just the standard Rn×n representation for orthogonal (rotation-like) matrices. The representation theoretic notation will be suppressed moving forward for approachability to R∈O (n). For the isotropic activation functions functional class of study, f∈ F , this action is desired to commute with the isotropic activation functions, f:Rn→Rn , as shown in Eqn. 1. Additionally, the functions should maximally abide by this symmetry. ∀g∈O (n) [g, f] = g◦f−f◦g= 0 (1) More straightforwardly, it can be denoted as shown in Eqn. 2. f(Rx) = Rf (x)(2) Then, from this constraint, the activation function class can be constructed for this symmetry-defined primitive. A functional form which is within this functional class is displayed in Eqn. 3, where ˆx is the unit-normalised vector given the 4One that weight-decay regularisations may further moderate. 5 Within this framework, the definitional distinctions between activation functions and normalisers become less clear; for example, Layer-Normalisation could be considered a parameterised activation function with a specific symmetry constraint [ 6 ]. Therefore, a proposed distinguishing difference defining normalisers is proposed to be the presence of probabilistic constraints, time-like or data-like, and whether representative degrees-of-freedom are lost as a normaliser or preserved as an activation function. 3
2-norm6. f(x) = f(∥x∥2) ˆx(3) Many functions, which are non-linear in f , maximally satisfy this relation. The particularities of which function is used do not matter for dynamic network topologies. Several placeholders are suggested [ 6 ]; however, these superficially replicate existing activation functions, and there is no reason to assume their continued optimality in this altered function form. Therefore, the general function will just be denoted as f(x) = σ(∥x∥) ˆx . This equivariance is demonstrated in Eqns. 4 to 7, using x′=Rx for orthogonal matrix Rand that ˆx′=Rˆxremains unit-normalised. f(x′) = σ(∥x′∥) ˆx′=σ(∥Rx∥)Rˆx=f(Rx)(4) σ(∥Rx∥)Rˆx=σ√xTRTRxRˆx(5) RσpxTInxˆx=Rσ(∥x∥) ˆx(6) ∴f(Rx) = Rf (x)(7) 2.2 Single-Sided Continuous Reparameterisations This paper employs diagonalisation of layers for its dynamic topology method. This can most clearly be seen when two such orthogonal functions, or more sequential orthogonal primitives, are composed with three affine layers. This is a distinct compositional consideration from parameter symmetry work, which typically focuses on a single discrete permutation activation function sandwiched by two affine layers and is discussed briefly in this subsection. In effect, the full diagonalisation considers the general and continuous orthogonal reparameterisations twice, concurrently acting on the left and right of the middle affine layer. This reparameterisation is possible due to the functional class of affine layers exhibiting left and right closure under general linear actions. Since the orthogonal group is a subset of the general linear group, the affine layer is closed under orthogonal actions. It can therefore be reparameterised to produce an equivalence class of functionally identical models. Combining this with orthogonal equivariance enables a reparameterisation where the network’s functionality remains exactly invariant — everything about the network is preserved. This is derivation is demonstrated in the transformations between Eqns. 8 through 12, where fAff.1 (x) = W1x + b1 and fAff.2 (x) = W2x + b2 and fIso. =σ(∥x∥) ˆx , and using the relation In=R−1R=RTR for orthogonal matrices. This is for the standard two-affine with one non-linearity reparameterisation. fAff.2 ◦fIso. ◦fAff.1 =W2fW1x + b1+ b2(8) =W2fRTRW1x + b1+ b2(9) =W2RT | {z } W′ 2 f RW1 |{z} W′ 1 x +R b1 |{z} b′ 1 + b2(10) =W′ 2fW′ 1x + b′ 1+ b2(11) =W′ 2fW′ 1x + b′ 1+ b′ 2=f′ Aff.2 ◦fIso. ◦f′ Aff.1 (12) When considering the full three affine and two non-linearity compositional scenarios, a layer can be fully diagonalised by applying such transformations to the left and right of the intermediate affine layer, isolating a sparse and simplyconnected diagonalised matrix within the central affine layer. This diagonalisation is achieved through a singular value decomposition allowed through this two-sided gauge freedom. The diagonalised weights can then be ordered by singular value magnitude, enabling a perturbative consideration of the layer’s action — especially when incorporating normalisation. This diagonalisation procedure is discussed in the following subsection. 6 One may notice the 2 -norm present in this description and wonder whether this approach can be generalised to l -norm. Although appropriate primitive redefinition can be achieved using this to provide functional classes within the hyperoctahedral Bn primitive set, these considerations do not generalise well to dynamic network topologies. One may be tempted to adapt the singular value decomposition (SVD) diagonalisation procedure such that orthogonal group matrices are defined with a unit-norm defined by ∥a∥n ; however, there is a lack of an inner-product inducing the norm. This means that the l -norm does not define a generalised orthogonality, so one has to use the standard inner product. This leaves only the discrete hyperoctahedral group to which the map is equivariant, and it lacks the continuous group needed for this approach to project down to smaller subspaces. Although if all incoming weights to a neuron tend to zero, then it can still be pruned. Growth remains possible regardless. 4
2.3 Layer Diagonalisation First, a procedure for fully diagonalising a layer will be demonstrated. This makes clear that one layer at a time can be reparameterised to express a one-to-one connectivity between the current layer’s neurons, individuated in a particular basis, and the preceding layer’s neurons. For a full diagonalisation, as stated, this requires three affine layers interspaced around two isotropic primitives. However, growth and pruning of the network do not necessarily require a full diagonalisation. For efficiency, several matrices can be contracted in an abridged approach to the method. The latter partial diagonalisation using singular value decomposition will be discussed following the full diagonalisation. Similar to before, one can consider two isotropic activation functions, or generally non-linearity, interspaced with three affine maps: • First Affine Layer: fAff.1 :Rl→Rm . A functional class with form fAff.1 (x) = W1x + b1 . Has both left and right general-linear closure (functional class) symmetries; the left closure will be utilised. • First Non-Linearity: fIso.1 :Rm→Rm . A functional class with form fIso.1 (x) = σ1(∥x∥) ˆx . Exhibits an standard orthogonal algebraic equivariance symmetry by definition, this will be leveraged. • Second Affine Layer: fAff.2 :Rm→Rn . A functional class with form fAff.2 (x) = W2x + b2 . Has both left and right general-linear closure (functional class) symmetries; both closures will be utilised. • Second Non-Linearity: fIso.2 :Rn→Rn . A functional class with form fIso.2 (x) = σ2(∥x∥) ˆx . Exhibits an standard orthogonal algebraic equivariance symmetry by definition, this will be leveraged. • Third Affine Layer: fAff.3 :Rn→Ro. A functional class with form fAff.3 (x) = W3x + b3. Has both left and right general-linear closure (functional class) symmetries; the right closure will be utilised. With their composition specified by: fAff.3 ◦fIso.2 ◦fAff.2 ◦fIso.1 ◦fAff.1 :Rl→Ro . Then fAff.2 can be reexpressed in a basis where it becomes diagonalised and ordered by singular values. The diagonalisation transformations are displayed in Eqns. 13 to 17, which act on the intermediate affine layer with a double-sided transformation similar to the previous section’s reparameterisation. In particular, the weights W2 can be represented with a singular value decomposition W2=UΣ2VT , where U and V are orthogonal so can commute with the non-linearity orthogonal equivariance symmetry by construction, and Σ2is a diagonal matrix. W3fIso.2 W2fIso.1 W1x + b1+ b2+ b3(13) =W3fIso.2 UΣ2VTfIso.1 W1x + b1+UUT |{z} In b2 + b3(14) =W3UfIso.2 Σ2fIso.1 VTW1x + b1+UT b2+ b3(15) =W3U |{z} W′ 3 fIso.2 Σ2fIso.1 VTW1 | {z } W′ 1 x +VT b1 |{z} b′ 1 +UT b2 |{z} b′ 2 + b3(16) =W′ 3fIso.2 Σ2fIso.1 (W′ 1x +b′ 1) + b′ 2+ b3(17) This reparameterisation corresponds to the following transforms for the modified affine layers. •W′ 1=VTW1 •W′ 2=UTW2V=UTUΣ2VTV=Σ2 •W′ 3=W3U • b′ 1=VT b1 • b′ 2=UT b2 • b′ 3= b3 5
Before Diagonalisation After Diagonalisation Figure 1: This illustration depicts the qualitative effects on a network from full diagonalisation — the chosen layer has a double-sided basis change to the connectivity, reducing the map to a one-to-one correspondence between neurons. This drastically simplifies the interrelations between the chosen layer, allowing the application of the dynamic network implementation. In general, for sequential layers, only one layer can be diagonalised at a time, as diagonalising one layer often destroys the diagonalised state of the immediately preceding and following layers in the process. Layers which are interspaced by other affine transforms can be concurrently diagonalised. Where W′ 2 is now fully diagonalised, in effect the ‘neurons’ from this perspective only communicate one-to-one in this layer, as depicted in Fig. 1. This greatly simplifies the construction and may aid in isotropic network interpretability. This is because the generalised neuron is non-individuated, expressing a gauge freedom for reparameterisation in many differing and equivalent bases, hence results in an analytically identical computation. In other words, the choice of individual neurons is arbitrary, as there is no map which individuates them so can be chosen in a convenient expression. Additionally, this decomposition can be computed whilst ordering the singular values. This can give an indication of the importance of each weight and will be leveraged throughout the following methodology. Regularisation could be applied to these singular values as a form of weight decay if desirable. One may be concerned that the one-to-one connectivity of layers infers a suboptimality of these layers; however, this is not necessarily so, as one can still analyse classical anisotropic networks with their parameters reexpressed as a diagonalised spectrum of one-to-one connection, it is just not possible to reparameterise the model in such a way as it does not commute with the elementwise operations. Therefore, this does not necessarily imply lack of expressibility. 2.3.1 Comment on Sparsity of Diagonalisation This full layer diagonalisation can be leveraged as is to reparameterise the existing network into a sparser basis, whilst preserving the function exactly. The approach above enables every other layer of a multilayer perceptron, and potentially generalised across other architectures, to be reexpressed in a diagonalised parameterisation whilst retaining all performance. If one considers a dense network with all layers of n -width, this would ordinarily consist of N(N+1) parameters per affine mapping — N2 for weights and N for biases 7 . However, if that layer is able to be diagonalised, then the affine transform can be equally described by just 2N parameters. Incidentally, to the objective of this paper, this can be considered a vast improvement in sparsity of the dense network without any performance loss. The map is now equivalent to an elementwise product and summation. This grows linearly with neuron size as opposed to quadratically (typically bilinearly). Naturally, this must be associated with two additional affine layers to enable such a reparameterisations. Nevertheless, this sparsity is a drastic computational improvement, particularly for a static inference-time network, unlike the naive parameterisation. Moreover, the full diagonalisation can be computed for every alternative layer surrounded by nondiagonalised affine layers. For a (odd-length) dense network, with n -width layers and 2D+ 1 sequential affine mappings, then Eqn. 18 gives the sparsity factor Sp . This is the proportion of remaining parameters compared to the original count. Even-length networks have a similar relations with slightly higher sparsity. Sp=2DN +N(D+ 1) (N+ 1) N(2D+ 1) (N+ 1) =D+ 1 2D+ 1 +2D (2D+ 1) (N+ 1) (18) Thus the network can be reexpressed in a sparsified basis, with a proportion of parameters that is equal to a constant plus a term which drops reciprocally with layerwidth as one may expect. Assymptopically, as D→ ∞ and N→ ∞ ,the network preserves exact function with only 50% of the parameters. Moreover, this may be further improved by jointly optimising over all of the layerwise basis changes, such as to maximise the overall sparsity across all parameters instead of every alternative layer. Therefore, Eqn. 18 provides an upper-bound for odd length dense isotropic networks. This may be achieved through working with the direct sum of parameterised orthogonal groups, such as through their Lie algebras and optimising for elementwise sparsity to leverage sparse-matrix computations. Both of these approaches yield exact analytical sparsification, this is a symmetry-leveraged structural not statistical sparsity mask. However, from these parameterisations, the already symmetry-sparsified network can serve as a base from which to then perform classical sparsifying of the parameters. Hence, allowing some approximation of function, this sparsity factor may be drastically improved using conventional methodologies of pruning — moreover, the singular values in already diagonalised layers can aid in an analysis of the function approximation (particularly when using a form of 7Additional parameters for isotropic normalisations or other operations may be present. 6
normalisation to ensure magnitude compensation does not occur in the subsequent layer). This is felt to be a worthwhile and exciting direction for isotropic networks, which already show some preliminary improved performance on some tasks [8]. 2.3.2 Partial Layer Diagonalisation Despite being illustrative and perhaps more broadly practical and interpretable, the full diagonalisation procedure is not necessitated for dynamic topologies. The proposed algorithm for dynamic topologies can proceed with a simplified form, which may be more practical. This reparameterisation follows more closely the standard two affine layers surrounding an (isotropic) non-linearity. This partial diagonalisation can occur in two forms: left-sided and right-sided, corresponding to the following parameterised maps: mixing followed by scaling or scaling followed by mixing, respectively. The right-sided transform is explicitly given in Eqns. 19 through 23. This also influences which direction the pruning and growth acts, forward or backwards, which are both permissible applications under this construction. However, the forward approach exhibits the most versatility. W2fIso.1 W1x + b1+ b2(19) =W2fIso.1 UΣ1VTx +UUT |{z} In b1 + b2(20) =W2UfIso.1 Σ1VTx +UT b1+ b2(21) =W2U |{z} W′ 2 fIso.1 Σ1VT | {z } W′ 1 x +UT b1 |{z} b′ 1 + b2(22) =W′ 2fIso.1 W′ 1x + b′ 1+ b′ 2(23) This right partial diagonalisation is given by the following reparameterisations, which correspond to a modified affine layers: •W′ 1=UTW1=UTUΣ1VT=Σ1VT •W′ 2=W2U • b′ 1=UT b1 • b′ 2= b2 This partial diagonalisation is sufficient to prune and grow neurons defined by the layer to which fIso.1 applies over. It does require that the diagonalised matrix Σ1 be retained for the thresholding. Overall, this partial diagonalisation is sufficient for the methodology, but not as illustrative as the full diagonalisations, where the one-to-one mapping is explicit, sparsified and better interpretable. Both methods remain applicable for the procedures presented, with the latter partial diagonalisation enabling dynamic topologies additionally in the first layer, since no preceding layer need ‘absorb’ the V orthogonal matrix. Hence, full-diagonalisation is not possible on this first layer; however, prior work suggests that pruning of the first layer is disproportionatly unfavourable to performance anyway [ 9 ] — this would require reestablishing such an empirical finding again for isotropic networks due to their foundational and representational distinctions. Additionally, partial diagonalisation may offer short-term improvement in efficiency due to fewer matrix multiplications, though fulldiagonalised sparsified reparameterisation may remain preferable at inference time. For clarity, the full diagonalisation will be referenced moving forward; however, code implementations use the partial form to additionally enable growth and pruning of the first hidden layer. 2.4 Dynamic Pruning Considering the intermediate affine map fAff.2 , which has been expressed in a diagonalised basis with weights ordered as singular values, W′ 2= Σ ∈Rn×m , Σii ≥Σ(i+1)(i+1) . As this diagonalised weight tends to zero, the ‘neuron’ becomes entirely independent of the preceding layer, Σii →0 . This enables the pruning of the entire neuron, not just a single connection, since the neuron has only a single one-to-one connectivity due to diagonalisation so the connectivity and neuron become one of the same. This pruning results in minimal and measurable degradation of network functionality 8 . 8 It is suggested to use some form of isotropic normaliser to ensure magnitude-compensations do not occur in subsequent layers. Moreover, this should not be one which supresses radial degrees-of-freedom in representation, as discussed in App. B, instead the previously proposed Chi-normaliser [6] may be well-suited. 7
Before Diagonalisation Left-Partial Diagonalisation Right-Partial Diagonalisation Figure 2: This illustration depicts the qualitative effects on a network from partial diagonalisation. This time, the chosen layer has a single-sided basis change to the connectivity. Depending on the transform side, leftor right-sided, an initial or later mixing of connectivities occurs, followed by scaling by the singular values. This setup may be more convenient to implement than a full diagonalisation. Some interpretability and explanatory convenience are lost due to the remaining connectivity mixing. Additionally, this is a singular-value based thresholding for neural pruning as opposed to top −k or thresholded-magnitude based connectivity pruning as commonly implemented9. However, in the current construction, a bias parameter remains and cannot be simply ‘gauged away’ under this reparameterisation. This requires the addition of a new tunable parameter, the "intrinsic length parameter", denoted o , and there are numerous ways to interpret such a parameter, and its behaviour is largely unique to isotropic networks. This additional parameter has become a crucial additional insight into isotropic networks which enables this unique structural reparameterisation. To gauge away residual bias into the intrinsic length requires reexpressing the isotropic function, as detailed in Eqns. 24. f(x) = f(∥x∥) ˆx=f(∥x∥)x ∥x∥=f(∥x∥) ∥x∥x =g(∥x∥)x (24) Additionally expressing the affine layer as Σx + band implementing into Eqn. 24 to yield Eqn. 25. f(x) = g Σx + b Σx + b(25) One can then expand the norm-term as shown in Eqn. 26 where Σii →0 and bi indicate the ith components decomposed in the standard basis. Σx + b 2 2= min(n,m) X j=0 Σjjxj+ bj2=Σiixi+ bi2+ min(n,m) X i=j=0 Σjjxj+ bj2= b2 i+ min(n,m) X i=j=0 Σjjxj+ bj2 (26) As Σii →0 , the network’s norm term becomes asymptotically functionally identical to a network of a smaller ‘pruned’ size, except for the residual bias bi . The residual bi in the linear part of the isotropic activation function can be forward projected to update the biases of the following affine layer, such that it identically cancels; however, the bi within the norm cannot. This requires a new tunable parameter: ‘intrinsic length’, which can ‘absorb’ the bias, making the network asymptotically exhibit full invariance to the action of pruning. This intrinsic length is assumed to be positive definite o= exp (λ) such that isotropic activation functions such as isotropic-tanh [ 6 ] remain defined as intended; however, this positive-definite property is not a strict requirement if the isotropic non-linearity is designed to permit it. Therefore, this intrinsic length can be generalised but will be assumed to be positive in the following discussion. Crucially, it can also be made trainable as a novel optimisable parameter specific for isotropic networks. Geometrically, it acts much like a bias that is orthogonal to the linear space. An intuition is that it is embedding the existing subspace away from the origin, e.g. Σx + bT o⊥= 0 and oT ⊥o⊥=o= exp (λ) — where λ can be instead optimised to ensure positive definiteness. If interpreted in such a geometric picture, the associated offset vector does exhibit a rotational degeneracy in the orthogonal complement space. Except for the following activation function, this complement space is projected out before the following affine layer. This returns the vector space to its prior dimensionality. 9 However, it is stressed that any of the previous approaches to decide connection-pruning may be similarly applicable, and can be explored empirically. This singular-value thresholding is presented merely as a novel alternative and may not be optimal; however, it will be assumed moving forwards. In practice, the methodology is to be kept open to any alternative thresholding schemes, such as additionally using gradients which may be preferable in many circumstances. Similarly, one may combine approaches and consider the gradient with respect to the diagonalised singular-value as a threshold for pruning — this is also effectivly a batch-statistic. However, naturally, pruning should always begin with the smallest singular value or loss-gradient magnitude to avoid function degredation. 8
Overall, the intrinsic-length offers a novel and beneficial trainable parameter for an isotropic network analogous to a perpendicular bias. If allowed to be negative, it may allow direction flipping which may increase network support for manipulating representations into antipodal arrangements where desirable. It could also be reinterpreted as a parameterised activation function instead of an additional orthogonal offset parameter for the affine map. In Eqn. 27, the norm is generalised with this new parameter. Σx + b+o⊥ 2 2=oT ⊥o⊥+ min(n,m) X j=0 Σjjxj+ bj2=o+ b2 i |{z} =o′=o′T ⊥o′ ⊥ + min(n,m) X j=0 j=iΣjjxj+ bj2= Σ′x + b′+o′ ⊥ 2 2 (27) It can also be shown that this remains invariant to orthogonal group actions, preserving the overall isotropic activation function’s equivariance, fIso. (Rx) = RfIso. (x). Overall, this additional parameter enables one to ‘gauge away’ residual bias using parameter degrees-of-freedom, further minimising function degradation during whole neuron pruning. This accounts for the normalisation term, but the linear aspect can be similarly accounted for through reparameterisation of the subsequent layer accomodating the pruned bias into the subsequent biases. However, functional degradation does remain when the pruned diagonalised weight is not identically zero, but close to: Σii =ϵ with 0< ϵ ≪1 . As stated, this should be mitigate using an isotropic normalisation composed with the activation function. Typically, one may attempt the standard approaches, layer-normalisation and batch-normalisation; however, isotropic alternatives are instead required. The latter is particularly discussed in App. B, which had the additional consequence of revealing an isotropic architecure which inherently displays a depthwise nested functional class due to the isotropic terms. This is felt to be rather significant as it reproduces the nested functional class structure in an alternative way to the typical Residual Network construction, and in a way which doesn’t constrain layer dimensionality. Instead, it is a natural consequence of the symmetry-principled reformulations of primitives. In summary, generalising layer-normalisation to isotropy requires one to not normalise by the mean, as this results in a change in intrinsic geometry: Rn→Rn−1,→Rn [ 6 ], while dividing by the standard-deviation projects to a hyper-spherical shell. In App. B this is shown to result in affine expressibility of the network, rendering it limited in its application. Therefore, a batch-normaliser such as the Chi-normaliser [ 6 ] may be employed to reduce function degredation under pruning. Additionally, this forward correction for the bias in the linear term ( x of g(∥x∥)x ) is similarly applicable to any remaining Σii =ϵ term producing x . If one considers the linear term, with original diagonalised weights Σ and its pruned form Σ′ alongside subsequent weight matrix before and after pruning, W(2) and Y(2) respectivly, then we desire the following map to be closely preserved: Y(2)Σ′x ≈W(2)Σx . Therefore, one can find a suitable adjusted weight Y(2) through a moore-penrose pseudo-inverse, where rank permits, such as to minimise a least-squares difference in these maps. This is displayed in Eqn. 28, when an inverse can be determined (if not one may continue to use the original W(2) with the associated column pruned.). If it is the smallest singular value which is deleted, as intended, then this reduces to a corresponding column deletion of Y(2) as would expected — since removing the smallest singular value and associated parameters minimally changes the mapping. Hence, this can be considered to derive the subsequent layer’s columnwise deletion. This will be further detailed in Sec. 2.6 which outlines the overall implementation. Y(2) =W(2)ΣΣ ′TΣ′Σ ′T−1 (28) 2.5 Dynamic Growth Dynamic growth is comparativly trivial. If acting forwards, it amounts to embedding the last layer’s space into a higher dimensional vector space. This may be interpreted in two ways for its implications on neurons. The first, is that if neurons are chosen to be individuated in some arbitrary basis, it corresponds to appending additional neurons to this layer which are connected in such a manenr that the model is functionally independent of them — these could be considered ‘scaffold neurons’ in this conventional individuated picture. Alternativly, it is comparable to treating the higher-dimensional generalised notion of a neuron to have further increased in dimensionality. Due to the isotropic activation function’s jacobian acting to distribute learning gradients, alongside standard connectivity mixing, these apparant individuated neurons may be rapidly trained despite being functionally independent of the original model. This is because they are a symmetrical reparameterisation with respect to the forward pass, but non-trivially break the structural symmetry of the backward pass. This can occur even when connecting parameters are zeroed. This can be seen compartivly for a typical anisotropic activation function’s diagonal jacobian compared to an isotropic non-diagonal jacobian in Eqns. 29 and Eqn. 30 respectivly for a n -‘neuron’ layer, f:Rn→Rn and standard basis-vectors ˆei . The singularity at x = 0 can become a non-troublesome coordinate singularity under suitable choices of σ , f( 0) = 0 ; however, requires a custom implementation to prevent autodiff issues which is available in the code repository at . 9
x(3) =W3g x(2) x(2) + b3(39) x(3) =W3g x(2) g x(1) W′ 2x(0) +g x(1) b′ 1+ b2+ b3(40) x(3) =g x(2) g x(1) W′ 3x(0) +g x(2) g x(1) b′′ 1+g x(2) b′ 2+ b3(41) Generalising this to yield the full isotropic dense network recursive relation. x(l)= l−1 Y i=1 g x(i) !W◦ lx(0) + l−1 X i=1 b◦ i i Y j=1 g x(j) + b◦(42) The angular dependence may appear trivial due to the W◦ lx(0) ; however, this is misleading due to it also appearing in the argument of map, g , the output angular distribution non-trivially depends on the input. However, interspacing other maps besides isotropic and affine primitives may also be explored from this view. Moreover, this can be rewritten under suitable parameter changes as Eqn. 43. Relaxing the reparameterisation constraint for Eqn. 43 and enabling parameters to take any values also broadens the notion of a dense neural network in this context, showing that isotropy generalised networks can act as a purturbative sum of decreasing sized networks. Therefore, this generalised construction enables a nested functional class and hence expressivity inclusion, much like the nested functional classes of Residual networks yet instead emerging from symmetry. Furthermore, the intermediate layers can vary in dimensionality whilst still preserving this nested functional class structure, this symmetry emergence of nested functional classes stems from primitive constraints not architectural constraints; therefore it does not constrain the layer dimensionality like Residual Networks classically do. x(l)= l−1 X i=1 W◦ ix(0) + b◦ i i Y j=1 g x(j) (43) To make clear, this is not a classical dense network but is a larger encompassing but derivative architecture including it as a enclosed functional class — hence, it generalises classical dense networks. It displays a natural nested functional class, and a purtubative-like expansion which can recursivly includes shallower networks. If g:R+→[0, α] , then this also acts like a purturbative expansion, allowing convergence if α≤1 . Similarly normalisation can be used to implement a series convergence. This approach may also allow a new route to recursive reimplementation of backward and forward propagation methods. It is applicable for any non-linearity which can be expressed as f(x) = g(x)x for g:Rn→R or similarly for complex networks. This includes any generalised norm-based term, such as hyperoctahedral primitives defined using an l -norm for l= 2 , although isotropy’s orthogonal equivariance ensures this. Generalising further, one could substitute independent networks for each gargument producing an ensemble network-like effect. The initial intention was to show the application of normalisation to this system and how this may mitigate the input-error under pruning if placed carefully. One may consider how the residual input error term from pruning neurons, after residual bias has been gauged away, may continue to negativly impact performance. One could consider preor post-composing activation functions with a hyper-spherical normalisation such that each layer produces a vector with a fixed norm. This can be seen to be an undesirable placement, despite mitigating the input-error under pruning, since all the non-linearities become constant and the layer becomes linear and equivilant under reparameterisation to Eqn. 44 — which is guaranteed not to have a universal approximation theorem due to the constraint to affine maps. This suggests this form of hyperspherical shell normalisation collapses expressibility to affine maps. x(l)=Wˆx(0) + b(44) Therefore, normalising to a hyperspherical shell must not happen occur within an isotropic layer, as this guarantees linear expressibility by the network. Instead, either generalised normalisation, which does not project to the hyperspherical shell such as the proposed Chi-Normalisation [ 6 ], which can be preor post-composed with isotropic activations which is likely a preferable form of normalisation. These constrain the magnitude of xi in the product Σiixi , to ensure that as Σii →0 it cannot be compensated by a growth in xiwhich may otherwise adversily affect the network’s functionality. Generally, one should also consider the preor postcomposure of normalisations that can also impact vanishing or exploding gradients within these generalised isotropic architectures. 16
C Reparameterisation Non-Equivilance Through Gradient Coupling Although function equivilance is retained through the various reparameterisations, these reparameterisations are nonequivilant when considering gradient updates. The most trivial case, which appears to occur uniquely in isotropic-enabled orthogonal reparameterisations is coupling to adaptive optimisers. In effect, although a diagonalised network performs equally to its non-diagonalised counterpart, since adaptive optimisers often use an elementwise scaling, then the result trajectories are not equiviariant to the orthogonal actions — these networks will learning substantially differently. Thus, they diverge in function equivilance through training. This phenomena should not be underappreciated, since this gradient trajectory divergence does not typically happen in anisotropic reparameterisations. Ordinarily this would not be relevant under parameter symmetries; however, the isotropic parameter symmetries do cause these adaptive optimiser (elementwise) accumulations to diverge, meaning that isotropic parameter symmetries are only valid to first order in this case. This is a new consideration for the nature of parameter symmetries in networks with generalised primitives. The second trivial case of gradient coupling is products, AB =W . If diagonalisation methods do not recontract terms, then gradient trajectories will also diverge. Considering ym=Wmnxn and AB =W , produces two differing update equations, assuming standard gradient descent, as shown in Eqns. 45 and 46. Other optimisers yield analogous results and the author is unaware of an optimiser, or an existance possibility, which nativly reconciles these. Notating gp=∂L ∂yp y′ m=W′ mnxn= (Wmn −ηgmxn)xn(45) y′ m= (A′ mkB′ kn)xn= (Amk −ηgmBkixi) (Bkn −ηgjAjkxn)xn(46) Subtracting these equations to determine the divergence with ϵm= (W′ mn −A′ mkB′ kn)xn . These are displayed in Eqns. 47 and 48. ϵm= (Wmn −ηgmxn)xn−AmkBkn −ηgmBknBkixi−ηgjAjkAmkxn+η2gjgmAjkBkixnxixn(47) ϵm=η(gmBknxnBkixi+gjAjkAmkxnxn−gmxnxn(1 + ηgjAjkBkixi)) (48) This result establishes that leaving a layer in a partially diagonalised parameterisation may cause divergences if ΣVT are not contracted back to W . Moreover, various reparameterisations W=AB =CD may cause differing trajectories, which remains pertinant to the linear aspects of isotropic functions, including in the expansion given in App. B. However, what is more significant is the null-like reparameterisations appearing under Σii = 0 as previously discussed. These leave the network functionally identical; yet reparameterisations can be couple to the optimisation in non-trivial ways. This may allow a new mode to steer such networks, acting like a new choice situated somewhere between a reparameterisation and initialisation. This can be discusd both as a foward and backward process, but the forward partial diagonalisation is displayed in Eqn. 49 for consideration. yi=Wij fIso.1 ΣjkVlk | {z } Yjl xl+ bj j + di(49) Differentiating y with respect to Y , W and b , indicates their interdependency, and how initialisation even for rows/columns corresponding to zero-singular values has a non-trivial gradient coupling, producing a symmetry breaking consideration. These are given in Eqns. 50, 51 and 52, respectivly using notation zj=ΣjkVlkxl+ bj. ∂yi ∂Ymn =g(∥z∥) (Wimxn) + Wijzj ∂ ∂Ymn g(∥z∥)(50) ∂yi ∂Wmn =δmig(∥z∥)Ynlxl+ bn(51) ∂yi ∂ bm =g(∥z∥)Wim +Wijzj ∂ ∂ bm g(∥z∥)(52) One can see that in all these equations, the choices of reparameterisation for entries premultiplied by Σii = 0 still have a meaningful coupling to optimisation and therefore gradient trajectories. Therefore, these null-like reparameterisations only preserve forward pass network functionality, whilst altering the backward pass gradient step. Therefore, these parameter symmetries are preserved upto zeroth order but not first order. One can observe that Eqn. 51 will become relevant after two gradient steps when the update makes the original zero singular value become non-zero. 17
This has ramifications for the initialisation choices under neurogenesis. Various choices for initialisations can result in spontaneous symmetry breaking if zero initialised, or one can bias gradient trajectories using a non-zero initialisation which performs an explicit symmetry break into the backward pass. These considerations may be important aspects of dynamical networks moving forward and the potential for parameter symmetries to only maintain the symmetry upto zeroth order should be studied and findings factored into such implementations. These may steer learning without immediatly altering the network. These appear to generally manifest when a reparameterisation is available which preserves network functionality, but that action is outside the group actions defining the primitives. Overall this is a function invariance under parameter symmetries, whilst gradient coupling breaks the optimisation symmetry and hence the networks are inequivilant under these reparameterisations. D Backward Adaption 18