On Dynamic Architectures For Isotropic Networks
Full text
Dynamic Topologies of Isotropic Networks George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] August 1, 2025 1 Introduction: Neural plasticity in the quantity and connectivity of neurons is biologically advantageous and routinely occurs in the brains of animals. It can improve neural efficiency through pruning, whilst enabling the accumulation of knowledge, robustness and function through growth. It is therefore hypothesised that analogous behaviour within artificial neural networks may be similarly beneficial. Ranging from improved efficiency to reduced anomalous computations to increased capacity, the practical consequences of networks which can adapt their architecture and connectivities in real-time could be transformative. Proposed is a novel methodology which leverages an alternative formulation of primitives, known as “isotropic primitives”, to achieve real-time restructuring of networks with minimal degradation in functionality. This is achieved by exploiting a symmetry-structural functional invariance of a network. This both indicates the flexible use cases of these alternative formulations and enables the exploration of model adaptations to new information. Typically, a neural network is defined by computation predicated upon individual units intercommunicating to yield a desirable learnt function. This situates neurons, with their interconnections and computations, as ontologically fundamental building blocks to this computation system, so far. Resulting from this is an agreeable separability, and hence individuality, to neurons. This is implicitly bestowed definitionally by this contemporary perspective. They become the constructing atomic units for forming networks. The individuality of neurons is underscored by the primitives in widespread adoption: activation functions, normalisers, optimisers, all reinforce this. They collectivly continue an observable permutation symmetry to our modern networks as consequence to neuron-wise priors. Working deductivly from this current construction, one can then determine the well-known computational equivalences, exhibited as parameter-space degeneracies, which exist under exchange of neurons — formally termed a permutation symmetry, Sn , of neurons in the network. Under appropriate reparameterisations, swapping of neurons, and similarly exchanging corresponding weights and internal biases, can be considered a computational invariance of the network. In other words, the network functions identically before-and-after exchange by permuting actions. Hence, a permutation symmetry is deduced from this construction, which has emerged particularly through the contemporary constraints on primitives implicitly predicated on this perspective of individual neurons forming a network. Fundamentally, the generalised approach surrounding ‘isotropic deep learning’, reverses this ontological prior of neurons. Instead of neurons as the fundamental constituents implicitly defining this computational approach, they are emergent from symmetry. Hence, instead of deductively working from neurons leading to symmetry, this work proposes to reconsider the observed relation in reverse: symmetry leads to neurons. Philosophically, this situates foundational primitive symmetries as ontologically prior, and neurons are derived. Each primitive symmetries can then be definitional to each branch of appraoch. The ramifications of this redefinition are broad: from a broader notion of an artificial neural network system, to new sets of primitives and consequential, leveragable, behavioural changes for the network. Moreover, the notion of individuated neurons becomes an emergent property of permutation definitions in the primitives currently used. However, using this alternative perspective, one can instead consider neurons as a unit extending from the permutation definition, rather than the other way around. As a natural result, one can substitute alternative symmetries to yield new sets of primitives and consequently new notions of what a neuron is. This present work leverages one such alternative symmetry defined branch. In particular, this work explores one of these: the ‘isotropic’ redefinition predicated on the continuous orthogonal superset, O (n) , of the preexisting discrete permutation group, Sn⊂O (n) .From this, primitives emerge that are orthogonally equivariant to standard representations. From this ontological inversion, one can work deductivly to rediscover emergent neurons. For isotropic networks, the notion of a neuron is more general. Typically we may consider a layer construction as n copies of 1 -dimensional neuron-like objects concatenated, due to the permutation symmetry; however, due to the orthogonal symmetry redefinition this shifts now to considering a layer as 1copy of an n-dimensional neuron. In this approach, there is no longer a clean, individuated and agreeable definition of a neuron. There exists a gauge freedom to choose whatever decomposition to individual neurons one likes. Moreover, varied decompositions do not alter 1
the computation — the fundamental neuron is higher-dimensional and can be decomposed in any linearly combined, but normalised, way to yield equally valid alterantive bases. Moreover, this freedom to alter the basis for the represenation space extends to circumstances where one can project up to a higher-dimensional space without loss of functionality, and similarly, under more restricted conditions, one can down-project to subspaces without functionality consequences. What this amounts to is a new form of computational invariance to leverage with isotropic redefinitions — one in which the architectural structure becomes plastic and adaptable to task through reparameterisations. These are structural automorphisms for networks. This contrasts to the contemporary permutation definition which is limited to exchange symmetries which do not alter the structure. This new definition of primitives allows networks to dynamically grow and shrink to demand. This is felt to be a useful consequence directly enabled by orthogonal reformulation of network’s basic constituants. Reparameterisations allow neurons within a layer to compensate for alterations to the structure, whilst maintaining functionality. 1.1 Related Work 2 Theory This section outlines how a network can grow and prune under various reparameterisations and conditions. First, a general introduction to isotropic activation functions — these considerations similarly apply if also using normalisation, etc. Following this, a layer-wise basis change into a diagonalised configuration using scalar-valued decomposition will be described. Concluding the theory section, it will be discussed how the network can have growth and pruning for each layer, and a brief discussion of applications. 2.1 Isotropic Activation Function Overview Isotropy is most generally an orthogonal group family constraint applied to form a set of primitives. These transform under a particular, typically standard, representation of the orthogonal group specific to the width of the layer. For the primitives in this scenario, and for an n -width network, the standard representation is given by ρ: O (n)→ GLn(R) and each element will be notated as a matrix R=ρ(g)∈GLn(R) where g is an element of the orthogonal group, g∈O (n) . This is just the standard Rn×n representation for orthogonal (rotation-like) matrices. The representation theoretic notation will be suppressed moving forward for approachability to R∈O (n). For the isotropic activation functions functional class of study, f∈ F , this action is desired to commute with the isotropic activation functions, f:Rn→Rn , as shown in Eqn. 1. Additionally, the functions should maximally abide by this symmetry. ∀g∈O (n) [g, f] = g◦f−f◦g= 0 (1) More straightforwardly, it can be denoted as shown in Eqn. 2. f(Rx) = Rf (x)(2) Then, from this constraint, the activation function class can be constructed for this symmetry-defined primitive. A functional form which is within this functional class is displayed in Eqn. 3, where ˆx is the unit-normalised vector given the 2-norm1. f(x) = f(∥x∥2) ˆx(3) Many functions, which are non-linear in f , maximally satisfy this relation. The particularities of which function is used do not matter for dynamic network topologies. Therefore, these will just denoted as general, f(x) = σ(∥x∥) ˆx . This equivariance is demonstrated in Eqns. 4 to 7, using x′=Rx for orthogonal matrix R and that ˆx′=Rˆx remains unit-normalised. f(x′) = σ(∥x′∥) ˆx′=σ(∥Rx∥)Rˆx=f(Rx)(4) σ(∥Rx∥)Rˆx=σ√xTRTRxRˆx(5) RσpxTInxˆx=Rσ(∥x∥) ˆx(6) 1 One may notice the 2 -norm present in this description and wonder whether this approach can be generalised to l -norm. Although appropriate primitive redefinition can be achieved using this, providing functional classes within the hyperoctahedral Bn primitive set, these considerations do not generalise well to dynamic network topologies. One may be tempted to adapt the SVD diagonalisation procedure such that orthogonal group matrices are defined with a unit-norm defined by ∥a∥n ; however, there is a lack of an inner-product inducing the norm meaning that the l -norm does not define a generalised orthogonality, so one has to use the standard inner-product. This leaves only the discrete hyperoctahedral group to which the map is equivariant, and lacking the continuous group needed for this approach to projection down to smaller subspaces. Although if all incoming weights to a neuron tend to zero, then it can still be pruned. Growth remains possible regardless. 2
∴f(Rx) = Rf (x)(7) This paper looks at the consequences under composition when two such orthogonal functions, or more sequential orthogonal primitives, are composed with three affine layers. This is a distinct compositional consideration from parameter symmetry work, which typically focuses on a single discrete permutation activation function sandwiched by two affine layers. In effect, the diagonalisation considers the general and continuous orthogonal reparameterisations twice, concurrently acting on the left and right of the middle affine layer. This reparameterisation is possible due to affine layers exhibiting a left and right closure to general linear actions. Since the orthogonal group is a subset of the general linear group, then the affine layer is closed under its action and therefore can be reparameterised. Combining this with the orthogonal equivariance enables a reparameterisation where the network’s functionality is invariant. This is derivation is demonstrated in the transformations between Eqns. 8 through 12, where fAff.1 (x) = W1x + b1 and fAff.2 (x) = W2x + b2 and fIso. =σ(∥x∥) ˆx , and using the relation In=R−1R=RTR for orthogonal matrices. This is for the standard two affine with one non-linearity reparameterisation. fAff.2 ◦fIso. ◦fAff.1 =W2fW1x + b1+ b2(8) =W2fRTRW1x + b1+ b2(9) =W2RT | {z } W′ 2 f RW1 |{z} W′ 1 x +R b1 |{z} b′ 1 + b2(10) =W′ 2fW′ 1x + b′ 1+ b2(11) =W′ 2fW′ 1x + b′ 1+ b′ 2=f′ Aff.2 ◦fIso. ◦f′ Aff.1 (12) When considering the full three affine and two non-linearity compositional scenarios, a layer can be diagonalised by applying such transformations to the left and right of the intermediate affine layer. This diagonalisation is achieved through a singular value decomposition allowed through this two-sided gauge freedom. The diagonalised weights can then be ordered by singular value magnitude, enabling a purturbative consideration of the layer’s action. This diagonalisation procedure is discussed in the following subsection. 2.2 Layer Diagonalisation Similar to before, one can consider two isotropic activation functions, or generally non-linearity, interspaced with three affine maps: • First Affine Layer: fAff.1 :Rl→Rm. A functional class with form fAff.1 (x) = W1x + b1. • First Non-Linearity: fIso.1 :Rm→Rm. A functional class with form fIso.1 (x) = σ1(∥x∥) ˆx • Second Affine Layer: fAff.2 :Rm→Rn. A functional class with form fAff.2 (x) = W2x + b2. • Second Non-Linearity: fIso.2 :Rn→Rn. A functional class with form fIso.2 (x) = σ2(∥x∥) ˆx • Third Affine Layer: fAff.3 :Rn→Ro. A functional class with form fAff.3 (x) = W3x + b3. With their composition specified by: fAff.3 ◦fIso.2 ◦fAff.2 ◦fIso.1 ◦fAff.1 :Rl→Ro . Then fAff.2 can be reexpressed in a basis where it becomes diagonalised and ordered by singular values. The diagonalisation transformations are displayed in Eqns. 13 to 17, which acts on the intermediate affine layer with a double-sided transformation similar to the previous’ section reparameterisation. In particular, the weights W2 can be represented with a singular value decomposition W2=UΣ2VT, where Uand Vare orthogonal so can commute with the non-linearity, and Σ2is a diagonal matrix. W3fIso.2 W2fIso.1 W1x + b1+ b2+ b3(13) =W3fIso.2 UΣ2VTfIso.1 W1x + b1+UUT |{z} In b2 + b3(14) =W3UfIso.2 Σ2fIso.1 VTW1x + b1+UT b2+ b3(15) 3
=W3U | {z } W′ 3 fIso.2 Σ2fIso.1 VTW1 | {z } W′ 1 x +VT b1 |{z} b′ 1 +UT b2 |{z} b′ 2 + b3(16) =W′ 3fIso.2 Σ2fIso.1 (W′ 1x +b′ 1) + b′ 2+ b3(17) This reparameterisation corresponds to the following transforms for the modified affine layers. •W′ 1=VTW1 •W′ 2=UTW2V=UTUΣ2VTV=Σ2 •W′ 3=W3U • b′ 1=VT b1 • b′ 2=UT b2 • b′ 3= b3 Where W′ 2 is now fully diagonalised, in effect the ‘neurons’ from this perspective only communicate one-to-one in this layer. This is because the generalised neuron has a gauge freedom to be expressed in many differing bases which result in an equivilant computation, as the choice of individual neurons is arbitrary as there is no map which individuates them. Additionally, this decomposition can be computed whilst ordering the singular values. This can give an indication into the importance of each weight and will be leveraged throughout the following methodology. Regularisation could be applied to these singular values as a form of weight decay if desirable. Moreover, for improved computational efficiency using sparsity, one can jointly optimise the layerwise bases changes such as to maximise the overall sparsity across all parameters. This could be achieved through working with the direct-sum of parameterised orthogonal groups, such as through its Lie algebras. 2.3 Approximate Dynamic Pruning: Let’s consider the intermediate affine map A2, which has been expressed in a diagonalised basis with weights ordered as singular values. We can observe that as one tends to zero, the neuron becomes independent of the previous neuron. We can then prune this neuron without affecting functionality. However, a bias parameter remains and cannot be ‘gauged away’ in this manner. This requires the addition of a new tunable parameter, the "intrinsic length parameter", o , and there are numerous ways to interpret such a parameter, and its action is largely unique to isotropic networks. (It can also be added in a general norm/pseudonorm) Consider the following activation function form changes: f(x) = f(∥x∥) ˆx(18) f(x) = f(∥x∥)x ∥x∥(19) f(x) = f(∥x∥) ∥x∥x (20) f(x) = g(∥x∥)x (21) Reintroducing the affine layer (in diagonalised form Wx + b=˜ W⊗x + b f(x) = g ˜ W⊗x + b ˜ W⊗x + b(22) The aspect which requires consideration is the generalised norm ˜ W⊗x + b . Considering a toy example of three neurons (w1x1+b1)2+ (w2x2+b2)2+ (w3x3+b3)2 . If w3→0, then: (w1x1+b1)2+ (w2x2+b2)2+ (w3x3+b3)2 = (w1x1+b1)2+ (w2x2+b2)2+b2 3 (23) As w3→0 , the network becomes functionally identical to a network of a smaller ‘pruned’ size, except for b3 . The b3 in the linear part of the activation function can be forward projected to update the biases of the following affine layer, such that it cancels, yet b3within the norm cannot. 4
This requires a new tunable parameter: ‘intrinsic length’, which can ‘absorb’ the bias, making the network exhibit full invariance to the action of pruning. Whether the intrinsic length is positive definite or generally real can be experimented with, but crucially, it can also be made trainable as a novel optimisable parameter specific for isotropic networks. (w1x1+b1)2+ (w2x2+b2)2+ (w3x3+b3)2+o = (w1x1+b1)2+ (w2x2+b2)2+b2 3+o (24) (w1x1+b1)2+ (w2x2+b2)2+b2 3+o = (w1x1+b1)2+ (w2x2+b2)2+o′ (25) The parameter itself can be interpreted in multiple ways. For example, a parameterised activation function, modifying all lengths, or perhaps more interestingly, an orthogonal embedding offset in the affine map. Let’s consider it positive definite and we can consider ∥o⊥∥=o and o⊥ is in a differing orthogonal dimension to the original affine map’s span. It acts as a trainable parameterised embedding vector (it could be complexified for real-space values). Wx + b+o⊥(26) Of course, o⊥ is rotationally degenerate in such circumstances, only defined by its length, so practically it should not be computed in the affine layer, yet this provides an interesting intuition and allows for full equivalence as diagonalised wi→0 . When this is propagated to the next layer we project out this o⊥ degree of freedom, returning the network to the standard mdimensions. Generally, this intrinsic length may be a novel and beneficial trainable parameter in general settings. 2.4 Perfect Dynamic Pruning 2.5 Dynamic Growth: SCAFFOLDING NEURONS Dynamic growth is practically trivial; one can add zero-initialised weights and biases with no effect. These will become trained through backpropagation and may activate. This can also be undertaken in anisotropic networks. One can define a threshold and require two neurons below that singular value threshold at any time. Then, if the threshold is exceeded, perform pruning; if insufficient neurons meet the threshold, perform growth. These can be computed every epoch of less for computational tractability. 2.6 Implementation Considerations 2.7 Notes: One can consider these actions can be undertaken leftwards or rightwards, going from Rn→Rm to Rn−1→Rm or Rn→Rm−1, but is displayed in rightward fashion 5