On Dynamic Architectures For Isotropic Networks
Full text
Isotropic Plasticity: Symmetry-Defined Primitives Enable Dynamic Topologies George Bird Department of Computer Science & Department of Physics and Astronomy University of Manchester [email protected] September 29, 2025 Abstract A novel method is described and implemented for dynamic network topologies by leveraging isotropic primitives. Using such primitives, the notion of an individuated neurons is lost allowing a freedom in how layers are represented and interpreted. By considering a layer in a diagonalised representation one can modify the structure of the network, appending and removing neurons, in response to task neccesity. 1 Introduction Neural plasticity in the quantity and connectivity of neurons is biologically advantageous and routinely occurs in the brains of animals. It can improve neural efficiency through pruning, whilst enabling the accumulation of knowledge, robustness and function through growth. It is therefore hypothesised that analogous behaviour within artificial neural networks may be similarly beneficial. Ranging from improved efficiency to reduced anomalous computations to increased capacity, the practical consequences of networks that can adapt their architecture and connectivities in real-time could be transformative. Proposed is a novel methodology that leverages an alternative formulation of primitives, known as “isotropic primitives”, to achieve real-time restructuring of networks with minimal degradation in functionality. This is achieved by exploiting a symmetry-structural functional invariance of a network. This both indicates the flexible use cases of these alternative formulations and enables the exploration of model adaptations to new information. Typically, a neural network is defined by computation predicated upon individual units intercommunicating to yield a desirable learnt function. This situates neurons, with their interconnections and computations, as ontologically fundamental building blocks to this computation system, so far. Resulting from this is an agreeable separability, and hence individuality, to neurons. This is implicitly bestowed definitionally by this contemporary perspective. They become the constructing atomic units for forming networks. The individuality of neurons is underscored by the primitives in widespread adoption: activation functions, normalisers, optimisers, and many other implementations, which collectively reinforce this. They collectively continue an observable permutation symmetry to our modern networks as a consequence of neuron-wise priors. Working deductively from this current construction, one can then determine the well-known computational equivalences, exhibited as parameter-space degeneracies, which exist under the exchange of neurons — formally termed a permutation symmetry, Sn , of neurons in the network. Under appropriate reparameterisations, swapping of neurons, and similarly exchanging corresponding weights and internal biases, can be considered a computational invariance of the network. In other words, the network functions identically before and after exchange by permuting actions. Hence, a permutation symmetry is deduced from this construction, which has emerged particularly through the contemporary constraints on primitives implicitly predicated on this perspective of individual neurons forming a network. Fundamentally, the generalised approach surrounding ‘isotropic deep learning’, reverses this ontological prior of neurons. Instead of neurons as the fundamental constituents implicitly defining this computational approach, they are emergent from symmetry. Hence, instead of deductively working from neurons leading to symmetry, this work proposes to reconsider the observed relation in reverse: symmetry leads to neurons. Philosophically, this situates foundational, primitive symmetries as ontologically prior, and neurons are derived from them. Each primitive symmetry can then be definitional for each branch of approach. As an intuition, this is a similar to the physical ontological inversion between particles and symmetry-defined fields. The ramifications of this redefinition are broad: from a broader notion of an artificial neural network system, to new sets of primitives and consequential, leveragable, behavioural changes for the network. Moreover, the notion of individuated 1
neurons becomes an emergent property of permutation definitions in the primitives currently used. However, using this alternative perspective, one can instead consider neurons as a unit extending from the permutation definition, rather than the other way around. As a natural result, one can substitute alternative symmetries to yield new sets of primitives and consequently new notions of what a neuron is. This present work leverages one such alternative symmetry-defined branch. In particular, this work explores one of these: the ‘isotropic’ redefinition predicated on the continuous orthogonal superset, O (n) , of the preexisting discrete permutation group, Sn⊂O (n) . From this, primitives emerge that are orthogonally equivariant to standard representations. From this ontological inversion, one can work deductively to rediscover emergent neurons. For isotropic networks, the notion of a neuron is more general. Typically, we may consider a layer construction as n copies of 1 -dimensional neuron-like objects concatenated, due to the permutation symmetry; however, due to the orthogonal symmetry redefinition, this shifts now to considering a layer as 1copy of an n-dimensional neuron. In this approach, there is no longer a clean, individuated and agreeable definition of a neuron. There exists effectivly a (global) gauge freedom to choose whatever decomposition of individual neurons one likes. Moreover, varied decompositions do not alter the computation — the fundamental neuron is higher-dimensional and can be decomposed in any linearly combined, but normalised, way to yield equally valid alternative bases. Moreover, this freedom to alter the basis for the representation space extends to circumstances where one can project up to a higher-dimensional space without loss of functionality, and similarly, under more restricted conditions, one can down-project to subspaces without functionality consequences. What this amounts to is a new form of computational invariance to leverage with isotropic redefinitions — one in which the architectural structure becomes plastic and adaptable to tasks through reparameterisations. These are structural automorphisms for networks. This contrasts with the contemporary permutation definition, which is limited to exchange symmetries which do not alter the structure. This new definition of primitives enables networks to expand and contract in response to demand dynamically. This is considered a beneficial consequence directly enabled by the orthogonal reformulation of the network’s basic constituents. Reparameterisations allow neurons within a layer to compensate for alterations to the structure while maintaining functionality. 1.1 Related Work 2 Theory This section outlines how a network can be diagonalised to enable growth and pruning under various reparameterisations and conditions. First, a general introduction to isotropic activation functions is given. These considerations similarly apply if also using isotropic normalisations. Following this, a layer-wise basis change into a diagonalised configuration using scalar-valued decomposition will be described alongside a partial diagonalisation procedure which is more efficient. Concluding the theory section, it will be discussed how the network can have growth and pruning for each layer, and a brief discussion of applications. 2.1 Isotropic Activation Function Overview Isotropy is most generally an orthogonal group family constraint applied to form a set of primitives. These primitives then transform under a particular, typically standard, representation of the orthogonal group specific to the width of the layer. For the primitives in this scenario, and for an n -width network, the standard representation is given by ρ: O (n)→ GLn(R) and each element will be notated as a matrix R=ρ(g)∈GLn(R) where g is an element of the orthogonal group, g∈O (n) . This is just the standard Rn×n representation for orthogonal (rotation-like) matrices. The representation theoretic notation will be suppressed moving forward for approachability to R∈O (n). For the isotropic activation functions functional class of study, f∈ F , this action is desired to commute with the isotropic activation functions, f:Rn→Rn , as shown in Eqn. 1. Additionally, the functions should maximally abide by this symmetry. ∀g∈O (n) [g, f] = g◦f−f◦g= 0 (1) More straightforwardly, it can be denoted as shown in Eqn. 2. f(Rx) = Rf (x)(2) Then, from this constraint, the activation function class can be constructed for this symmetry-defined primitive. A functional form which is within this functional class is displayed in Eqn. 3, where ˆx is the unit-normalised vector given the 2-norm1. 1 One may notice the 2 -norm present in this description and wonder whether this approach can be generalised to l -norm. Although appropriate primitive redefinition can be achieved using this, providing functional classes within the hyperoctahedral Bn primitive set, these considerations do not generalise well to dynamic network topologies. One may be tempted to adapt the SVD diagonalisation procedure such that orthogonal group matrices are defined with a unit-norm defined by ∥a∥n ; however, there is a lack of an inner-product inducing the norm meaning that the l -norm does not define a generalised orthogonality, so one has to use the standard inner-product. This leaves only the discrete hyperoctahedral group to which the map is 2
f(x) = f(∥x∥2) ˆx(3) Many functions, which are non-linear in f , maximally satisfy this relation. The particularities of which function is used do not matter for dynamic network topologies. Therefore, these will just denoted as general, f(x) = σ(∥x∥) ˆx . This equivariance is demonstrated in Eqns. 4 to 7, using x′=Rx for orthogonal matrix R and that ˆx′=Rˆx remains unit-normalised. f(x′) = σ(∥x′∥) ˆx′=σ(∥Rx∥)Rˆx=f(Rx)(4) σ(∥Rx∥)Rˆx=σ√xTRTRxRˆx(5) RσpxTInxˆx=Rσ(∥x∥) ˆx(6) ∴f(Rx) = Rf (x)(7) This paper employs diagonalisation of layers for its dynamic topology method. This can most clearly be seen when two such orthogonal functions, or more sequential orthogonal primitives, are composed with three affine layers. This is a distinct compositional consideration from parameter symmetry work, which typically focuses on a single discrete permutation activation function sandwiched by two affine layers. In effect, the diagonalisation considers the general and continuous orthogonal reparameterisations twice, concurrently acting on the left and right of the middle affine layer. This reparameterisation is possible due to the functional class of affine layers exhibiting left and right closure under general linear actions. Since the orthogonal group is a subset of the general linear group, the affine layer is closed under orthogonal actions. It can therefore be reparameterised to produce an equivalence class of functionally identical models. Combining this with orthogonal equivariance enables a reparameterisation where the network’s functionality remains invariant. This is derivation is demonstrated in the transformations between Eqns. 8 through 12, where fAff.1 (x) = W1x+ b1 and fAff.2 (x) = W2x+ b2 and fIso. =σ(∥x∥) ˆx , and using the relation In=R−1R=RTR for orthogonal matrices. This is for the standard two-affine with one non-linearity reparameterisation. fAff.2 ◦fIso. ◦fAff.1 =W2fW1x + b1+ b2(8) =W2fRTRW1x + b1+ b2(9) =W2RT | {z } W′ 2 f RW1 |{z} W′ 1 x +R b1 |{z} b′ 1 + b2(10) =W′ 2fW′ 1x + b′ 1+ b2(11) =W′ 2fW′ 1x + b′ 1+ b′ 2=f′ Aff.2 ◦fIso. ◦f′ Aff.1 (12) When considering the full three affine and two non-linearity compositional scenarios, a layer can be fully diagonalised by applying such transformations to the left and right of the intermediate affine layer. This diagonalisation is achieved through a singular value decomposition allowed through this two-sided gauge freedom. The diagonalised weights can then be ordered by singular value magnitude, enabling a perturbative consideration of the layer’s action — especially when incorporating normalisation. This diagonalisation procedure is discussed in the following subsection. 2.2 Layer Diagonalisation First, a procedure for fully diagonalising a layer will be demonstrated. This makes clear that one layer at a time can be reparameterised to express a one-to-one connectivity between the current layer’s neurons, individuated in a particular basis, and the preceding layer’s neurons. For a full diagonalisation, this requires three affine layers interspaced around two isotropic primitives. However, growth and pruning of the network do not necessarily require a full diagonalisation. For efficiency, several matrices can be contracted in an abridged approach to the method. This latter partial diagonalisation will be discussed following the full diagonalisation. Similar to before, one can consider two isotropic activation functions, or generally non-linearity, interspaced with three affine maps: equivariant, and lacking the continuous group needed for this approach to projection down to smaller subspaces. Although if all incoming weights to a neuron tend to zero, then it can still be pruned. Growth remains possible regardless. 3
• First Affine Layer: fAff.1 :Rl→Rm. A functional class with form fAff.1 (x) = W1x + b1. • First Non-Linearity: fIso.1 :Rm→Rm. A functional class with form fIso.1 (x) = σ1(∥x∥) ˆx • Second Affine Layer: fAff.2 :Rm→Rn. A functional class with form fAff.2 (x) = W2x + b2. • Second Non-Linearity: fIso.2 :Rn→Rn. A functional class with form fIso.2 (x) = σ2(∥x∥) ˆx • Third Affine Layer: fAff.3 :Rn→Ro. A functional class with form fAff.3 (x) = W3x + b3. With their composition specified by: fAff.3 ◦fIso.2 ◦fAff.2 ◦fIso.1 ◦fAff.1 :Rl→Ro . Then fAff.2 can be reexpressed in a basis where it becomes diagonalised and ordered by singular values. The diagonalisation transformations are displayed in Eqns. 13 to 17, which act on the intermediate affine layer with a double-sided transformation similar to the previous section’s reparameterisation. In particular, the weights W2 can be represented with a singular value decomposition W2=UΣ2VT, where Uand Vare orthogonal so can commute with the non-linearity, and Σ2is a diagonal matrix. W3fIso.2 W2fIso.1 W1x + b1+ b2+ b3(13) =W3fIso.2 UΣ2VTfIso.1 W1x + b1+UUT |{z} In b2 + b3(14) =W3UfIso.2 Σ2fIso.1 VTW1x + b1+UT b2+ b3(15) =W3U |{z} W′ 3 fIso.2 Σ2fIso.1 VTW1 | {z } W′ 1 x +VT b1 |{z} b′ 1 +UT b2 |{z} b′ 2 + b3(16) =W′ 3fIso.2 Σ2fIso.1 (W′ 1x +b′ 1) + b′ 2+ b3(17) This reparameterisation corresponds to the following transforms for the modified affine layers. •W′ 1=VTW1 •W′ 2=UTW2V=UTUΣ2VTV=Σ2 •W′ 3=W3U • b′ 1=VT b1 • b′ 2=UT b2 • b′ 3= b3 Where W′ 2 is now fully diagonalised, in effect the ‘neurons’ from this perspective only communicate one-to-one in this layer, as depicted in Fig. 1. This is because the generalised neuron has a gauge freedom to be expressed in many differing bases, which results in an equivalent computation, as the choice of individual neurons is arbitrary, as there is no map which individuates them. Before Diagonalisation After Diagonalisation Figure 1: This illustration depicts the qualitative effects on a network from full diagonalisation — the chosen layer has a double-sided basis change to the connectivity, reducing the map to a one-to-one correspondence between neurons. This drastically simplifies the interrelations between the chosen layer, allowing the application of the dynamic network implementation. In general, for sequential layers, only one layer can be diagonalised at a time, as diagonalising one layer often destroys the diagonalised state of the immediately preceding and following layers in the process. Layers which are interspaced by other affine transforms can be concurrently diagonalised. 4
Additionally, this decomposition can be computed whilst ordering the singular values. This can give an indication of the importance of each weight and will be leveraged throughout the following methodology. Regularisation could be applied to these singular values as a form of weight decay if desirable. Moreover, for improved computational efficiency using sparsity, one can jointly optimise the layerwise basis changes, such as to maximise the overall sparsity across all parameters. This could be achieved through working with the direct sum of parameterised orthogonal groups, such as through their Lie algebras. 2.2.1 Partial Layer Diagonalisation Despite being illustrative and perhaps more broadly practical and interpretable, the full diagonalisation procedure is not, and dynamic topologies can proceed with a simplified form, which may be more practical. This reparameterisation follows more closely the standard two affine layers surrounding an (isotropic) non-linearity. This partial diagonalisation can occur in two forms: left-sided and right-sided, corresponding to the following parameterised maps: mixing followed by scaling or scaling followed by mixing, respectively. The right-sided transform is explicitly given in Eqns. 18 through 22. This also influences which direction the pruning and growth acts, forward or backwards, which are both permissible applications under this construction. W2fIso.1 W1x + b1+ b2(18) =W2fIso.1 UΣ1VTx +UUT |{z} In b1 + b2(19) =W2UfIso.1 Σ1VTx +UT b1+ b2(20) =W2U |{z} W′ 2 fIso.1 Σ1VT | {z } W′ 1 x +UT b1 |{z} b′ 1 + b2(21) =W′ 2fIso.1 W′ 1x + b′ 1+ b′ 2(22) This right partial diagonalisation is given by the following reparameterisations, which correspond to a modified affine layers: •W′ 1=UTW1=UTUΣ1VT=Σ1VT •W′ 2=W2U • b′ 1=UT b1 • b′ 2= b2 This partial diagonalisation is sufficient to prune and grow neurons defined by the layer to which fIso.1 applies over. It does require that the diagonalised matrix Σ1 be retained for the thresholding. Overall, this partial diagonalisation is sufficient for the methodology, but not as illustrative as the full diagonalisations, where the one-to-one mapping is explicit and better interpretable. Both methods remain applicable for the procedures presented, with the latter partial diagonalisation enabling dynamic topologies additionally in the first layer, since no preceding layer can ‘absorb’ the V orthogonal matrix, so full-diagonalisation is not possible and some improvement in efficiency due to fewer matrix multiplications. For clarity, the full diagonalisation will be referenced moving forward; however, implementations use the partial form to additionally enable growth and pruning of the first hidden layer. 2.3 Dynamic Pruning Considering the intermediate affine map fAff.2 , which has been expressed in a diagonalised basis with weights ordered as singular values, W′ 2= Σ ∈Rn×m . As this diagonalised weight tends to zero, the ‘neuron’ becomes entirely independent of the preceding layer, Σii →0 . This enables the pruning of the entire neuron, not just a connection, since the neuron has only a single one-to-one connectivity due to diagonalisation. This pruning results in minimal and measurable degradation of network functionality. However, in the current construction, a bias parameter remains and cannot be ‘gauged away’. This requires the addition of a new tunable parameter, the "intrinsic length parameter", denoted o , and there are numerous ways to interpret such a parameter, and its behaviour is largely unique to isotropic networks. To gauge away residual bias into the intrinsic length requires reexpressing the isotropic function, as detailed in Eqns. 23. f(x) = f(∥x∥) ˆx=f(∥x∥)x ∥x∥=f(∥x∥) ∥x∥x =g(∥x∥)x (23) 5
Before Diagonalisation Left-Partial Diagonalisation Right-Partial Diagonalisation Figure 2: This illustration depicts the qualitative effects on a network from partial diagonalisation. This time, the chosen layer has a single-sided basis change to the connectivity. Depending on the transform side, leftor right-sided, an initial or later mixing of connectivities occurs, followed by scaling by the singular values. This setup may be more convenient to implement than a full diagonalisation. Some interpretability and explanatory convenience are lost due to the remaining connectivity mixing. Additionally expressing the affine layer as Σx + band implementing into Eqn. 23 to yield Eqn. 24. f(x) = g Σx + b Σx + b(24) One can then expand the norm-term as shown in Eqn. 25 where Σii →0 and bi indicate the ith components decomposed in the standard basis. Σx + b 2 2= min(n,m) X j=0 Σjjxj+ bj2=Σiixi+ bi2+ min(n,m) X i=j=0 Σjjxj+ bj2= b2 i+ min(n,m) X i=j=0 Σjjxj+ bj2 (25) As Σii →0 , the network’s norm term becomes asymptotically functionally identical to a network of a smaller ‘pruned’ size, except for the residual bias bi . The residual bi in the linear part of the isotropic activation function can be forward projected to update the biases of the following affine layer, such that it identically cancels; however, the bi within the norm cannot. This requires a new tunable parameter: ‘intrinsic length’, which can ‘absorb’ the bias, making the network asymptotically exhibit full invariance to the action of pruning. This intrinsic length is assumed to be positive definite o= exp (λ) such that isotropic activation functions such as isotropic-tanh [] remain defined as intended; however, this positive-definite property is not a strict requirement if the isotropic non-linearity is designed to permit it. Therefore, this intrinsic length can be generalised but will be assumed to be positive in the following discussion. Crucially, it can also be made trainable as a novel optimisable parameter specific for isotropic networks. Geometrically, it acts much like a bias that is orthogonal to the linear space, an intuition is that it is embedding the existing subspace away from the origin, e.g. Σx + bT o⊥= 0 and oT ⊥o⊥=o= exp (λ) — where λ can be optimised for positive definiteness. This does have a rotational degeneracy in the orthogonal complement space, if interpreted in such a manner. Except for the following activation function, this complement space is projected out before the following affine layer, returning the vector space to its prior dimensionality. Overall, it may offer a novel and beneficial trainable parameter for an isotropic network, and if negative, may allow direction flipping, which may increase network support for manipulating representations into antipodal arrangements where desirable. It could also be reinterpreted as a parameterised activation function instead of an additional orthogonal offset parameter for the affine map. In Eqn. 26, the norm is generalised with this new parameter. Σx + b+o⊥ 2 2=oT ⊥o⊥+ min(n,m) X j=0 Σjjxj+ bj2=o+ b2 i |{z} =o′=o′T ⊥o′ ⊥ + min(n,m) X j=0 j=iΣjjxj+ bj2= Σ′x + b′+o′ ⊥ 2 2 (26) It can also be shown that this remains invariant to orthogonal group actions, preserving the overall isotropic activation function’s equivariance, fIso. (Rx) = RfIso. (x). Overall, this additional parameter enables one to gauge away residual bias using parameter degrees-of-freedom, further minimising function degradation during whole neuron pruning. This accounts for the normalisation term, but the linear aspect can be similarly accounted for through reparameterisation of the subsequent layer accomodating the pruned bias into the subsequent biases. 6
However, functional degradation does remain when the pruned diagonalised weight is not identically zero, but close to: Σii =ϵ with 0< ϵ ≪1 . This can be mitigate using an isotropic normalisation composed with the activation function. There are two standard approaches which can be used to achieve this: layer-normalisation and batch-normalisation. These are discussed in App. A, which had the additional consequence of revealing an isotropic architecure which inherently displays a depthwise nested functional class due to the isotropic terms — this is significant as reproduces the nested functional class structure in an alternative way to the typical Residual Network construction and in a way which doesn’t constrain layer dimensionality. Additionally, this forward correction for the bias in the linear term is similarly applicable to any remaining Σii =ϵ term. If one considers the linear term, with original diagonalised weights Σ and its pruned form Σ′ alongside subsequent weight matrix before and after pruning, W(2) and Y(2) respectivly, then we desire the following map to be closely preserved: Y(2)Σ′x ≈W(2)Σx . Therefore, one can find a suitable adjusted weight Y(2) through a pseudo-inverse, where applicable, such as to minimise a least-squares difference in these maps. This is displayed in Eqn. 27, if an inverse can be determined (if not one may continue to use the original W(2) with the associated column pruned.). If it is the smallest singular value which is deleted, as intended, then this reduces to a corresponding column deletion of Y(2). This will be further detailed in Sec. 2.5 which outlines the overall implementation. Y(2) =W(2)ΣΣ ′TΣ′Σ ′T−1 (27) In summary, generalising layer-normalisation to isotropy requires one to not normalise by the mean, as this results in a change in intrinsic geometry: Rn→Rn−1,→Rn [], while dividing by the standard-deviation projects to a hyper-spherical shell. In App. A this is shown to result in affine expressibility of the network, rendering it limited in its application. Therefore, a batch-normaliser such as the Chi-normaliser [] may be employed to reduce function degredation under pruning. Similarly, one can consider the gradient with respect to the diagonalised weight as a threshold for pruning — this is also effectivly a batch-statistic. Pruning should always begin with the smallest singular value or loss-gradient magnitude. 2.4 Dynamic Growth Dynamic growth is comparativly trivial. If acting forwards, it amounts to embedding the last layer’s space into a higher dimensional vector space. This may be interpreted in two ways for its implications on neurons. The first, is that if neurons are chosen to be individuated in some arbitrary basis, it corresponds to appending additional neurons to a layer which are connected, although the model is functionally independent of them — these could be considered ‘scaffold neurons ’in this conventional individuated picture. Alternativly, it is comparable to treating the higher-dimensional generalised notion of a neuron to have further increased in dimensionality. Due to the isotropic activation function’s jacobian acting to distribute learning gradients, alongside connectivity mixing, these apparant individuated neurons may be rapidly trained despite being functionally independent of the original model. This can occur even when connecting parameters are zeroed. This can be seen compartivly for a typical anisotropic activation function’s diagonal jacobian compared to an isotropic non-diagonal jacobian in Eqns. 28 and Eqn. 29 respectivly for a n -‘neuron’ layer, f:Rn→Rn and standard basis-vectors ˆei . The singularity at x = 0 can become a non-troublesome coordinate singularity under suitable choices of σ , f( 0) = 0 ; however, requires a custom implementation to prevent autodiff issues which is available in the code repository at . f(x) = N X i=1 σ(x ·ˆei) ˆei⇒∂(f·ˆei) ∂(x ·ˆej)=σ′(x ·ˆei)δij (28) f(x) = σ(∥x∥) ˆx⇒∂(f·ˆei) ∂(x ·ˆej)=σ(∥x∥) ∥x∥δij +σ′(∥x∥)−σ(∥x∥) ∥x∥(x ·ˆei) (x ·ˆej) ∥x∥2(29) Below, Eqn. 30, shows a reforumlated isotropic activation funciton and its jacobian. f(x) = g(∥x∥)x ⇒∂(f·ˆei) ∂(x ·ˆej)=g(∥x∥)δij +g′(∥x∥)xixj ∥x∥(30) One can see, that Eqn. 29 and Eqn. 30, for the isotropic activation function’s jacobian, contains non-diagonal terms in addition to the normal diagonalised derivative — this may increase gradient flow through these scaffold neurons hastening their operationalisation. The way in which such neurons are initialised depends on the diagonalisation picture being considered. This is in practicality non-trivial and choices can range in effect by explicit or spontaneous symmetry breaking. This arises since the reparameterisations are only functionally identical on the forward pass, they interact and are non-equivilant with gradient descent algorithms during updates. This is a potentially insightful direction for future study. Moreover, after the diagonalisation procedure, the two occurances of either U on the following layer or amending the bias term, similarly for V on the backwards approach, can become decoupled, especially evident under neuronal growth. These considerations are discussed in App. B. 7
Both the weights and biases associated with that scaffold neuron may be initialised in a variety of ways, with implications for learning and resulting function. In either case, the new singular values should be set to zero, yet the surrounding affine maps must also changes dimensionality to accomodate this new ‘neuron’. This may require a new column or row depending on which direction the growth acts in. How these are initialised does not affect the network’s functionality; however, for convention their SVD decomposition can have a new row or column added that is orthogonal to the existing span — this can be achieved through the Gram-Schmidt procedure. 2.5 Overview of Implementation In addition to growth and pruning considerations, it must be decided when to apply them. What will be discussed is a functional buffer of neurons, consisting of two primary hyperparameters: the number of scaffold neurons, Ξ∈Z+ , and singular-value threshold,ϑ∈R+. The threshold value indicates at which point the singular value connections are considered significant to performance, below this value their inclusion is considered neglible or perhaps contributes rarily-used, non-robust adaptations. It would typically be chosen to be a small positive number. The number of scaffold neurons indicates the buffer of neurons which is maintained below this threshold. If it is set to Ξ=2 , then two scaffold neurons are maintained in the network at any time, i.e. two diagonalised neurons below the singular-value threshold. If more neurons are below the threshold then neuronal pruning can occur; if fewer, neuronal growth occurs. These can be computed every epoch of less for computational tractability. In the following two subsections, a step by step methodology is provided on how to prune or grow neurons in dense networks. The extension to convolution growth and pruning of kernels is exactly analogous, treating the kernels channelwise dense networks. 2.5.1 Forward Adaption This subsection outlines the full forward adaption process such that it can be easily implemented, the backward adaption is given in App. C. The forward process is when considering an affine map Rn→Rm , then the following layer characterised by m -dimensionality can be increased or decreased. Several of these steps and reparameterisations can be contracted for computational efficiency, whether or not these contractions are implemented affects gradient flow as discussed in App. B. First perform a left-partial diagonalisation of the weight parameters using singular value decompisition, W1=UΣVT . This reparameterises from Eqn. 31 to 32, as previously demonstrated2. W2fIso.1 W1x + b1;o+ b2(31) =W2UfIso.1 Σ1VTx +UT b1;o+ b2(32) The dimensionality of these matrices is given by: U∈Rm×m , VT∈Rn×n , Σ∈Rm×n ≥0 , W2∈Rp×m , b1∈Rm , b2∈Rpand o∈R≥0, . The next step is to threshold the singular values to determine how many are below the threshold, ϑ . These are the entries along the diagonal. If the number of scaffold neurons falls below the desired count, |{Σii|Σii < ϑ}| <Ξ then iterativly add neurons until the condition is equalised. This is iterative growth (neurogenesis) of the network. This is described first below with Ξ increasing by one3. In this case proactive contractions of the matrices can be considered, which also highlights that the two instances of U can be decoupled as shown in Eqn. 33, where dashes indicate the new parameterisations in contrast with the prior. =W′ 2fIso.1 Σ′ 1V ′Tx + b′ 1;o′+ b′ 2(33) From this equation, one implements neurogenesis through the following steps. 1. For matrix, Σ∈Rm×n ≥0 , a row of zeros should be appended to the bottom of the matrix to form a matrix Σ′∈ R(m+1)×n ≥0. This is given by Σ′=Σ 0T 2. For vector, b1∈Rm this is modified with a new last entry, b∗ to form b′ 1∈R(m+1) . Additionally, b∗ must be chosen such that b2 ∗+o′=o, forming the non-negativity restriction b2 ∗< o. The new vector is of form b′ 1= b1 b∗. 2 One does not need to undo the diagonalisation after this procedure the basis change leaves the network functionally equivilant. Both training and diagonalising other surrounding layers will destroy the diagonalisation of the specified layer, there is no purpose to do this manually. Although the basis change will couple differently to dynamic optimisers and certain initialisations for grown layers will prevent reversing diagonalisation exactly. 3 Again this threshold can be redefined to be the magnitude with respect to loss as an alterantive loss-based approach as effectivly a mean importance statistic over a batch. A logarithm before thresholding could be taken to make order-of-magnitude scales comparable. Overall, this is less ideal as it doesn’t indicate the consequence on loss if the entire neuron is grown or pruned, only the local gradient. 8
3. For scalar, o∈R≥0, an adjustment follows accordingly from the prior step: o′=o−b2 ∗. 4. For matrix, VT∈Rn×n, no change need be made. 5. For matrix, W2∈Rp×m , this matrix requires a new column W2∈Rp×(m+1) . The exact initialisation of this column, U∗ , does not matter for the forward pass since it does not materially alter the functionality due to the premultiplication by a zeroed singular value and if b∗= 0 . However, in terms of a gradient coupling this has subtle yet significant implications. Setting the column to the zero-vector can result in a spontaneous symmetry breaking upon the next optimisation step, whilst specifying a non-zero vector acts as an explicit symmetry breaking which will bias learning accordingly. There are a range of such choices, such as choosing to form a semi-orthogonal matrix by using a Gram-Schmidt procedure to produce a column that ensure semi-orthogonality of the matrix. Alternativly initialising to be equal to an existing column or linear combination of, may be desirable to increase fidelity around existing high singular values. A convention may be selected such as unit-normalised vector forming a semi-orthogonal matrix, a range equally valid degenerate options is possible and these may couple with training dynamics unexpectedly. This will take the form W′ 2=W2 U∗ . This is highly interesting, as although these parameter symmetries all retain functionality in forward pass as expected, they can diverge in terms of resultant gradients which may be unexpected, coupling differently to optimisers and resulting in later functionality divergences — a first-order broken symmetry which is preserved only upto zeroth-order, while adaptive methods break them to second-order. 6. For vector, b2∈Rp , an adjustment arising from the prior bias term can be considered: b′ 2+g(·) U∗b∗= b2 ; therefore, b′ 2= b2−g(·) U∗b∗ . Since the norm remains constant when the network grows, no additional discrepancy arises due to the non-linear norm scaling in this reparameterisation within the batch. However, choices can then be made for U∗ and b∗ to ensure this is identically zero, by intialising either to the zeros. This means a spontaneous symmetry must occur in either or both variables, if set to zero; and one may be explicitly symmetry broken by being non-zero on initialisation. Approximatly correct solutions can also be made for b′ 2 with respect to the batch, the discrepancy arises from g(·) yet can be minimised with an appropriate normalisation such that a expected scaling is known. Then both may be explicitly symmetry broken. However, conventionally, U∗ , can satisfy the semi-orthogonality and b∗= 0. Moving onto neuralpruning when the number of scaffold neurons exceed the desired count, |{Σii|Σii < ϑ}| >Ξ then iterativly prune neurons until the condition is equalised. This is iterative pruning of the network. This is described first below with Ξdecreasing by one. 1. For matrix, Σ∈Rm×n ≥0 , the row closest approximating the zero vector, such as the smallest singular value row, may be deleted to form a matrix Σ′∈R(m−1)×n ≥0. This is given by Σ=Σ′ ΣT≈ 0T 2. For vector, b1∈Rmthis is modified by deleting the last entry, b∗to form b′ 1∈R(m−1), b1= b′ 1 b∗. 3. For scalar, o∈R≥0, an adjustment follows accordingly from the prior step: o′=o+b2 ∗. 4. For matrix, VT∈Rn×n, no change need be made. 5. For matrix, W2∈Rp×m , this matrix requires deletion of the corresponding column W2∈Rp×(m−1) . One could similarly implement Eqn. 27 if the full-rank invertibility is satisfied. Generally this forms a new matrix, W′ 2, from W2=W′ 2 U∗. 6. For vector, b2∈Rp , an adjustment arising from the prior bias term can be considered: b′ 2= b2+g(·) U∗b∗ , which preserves approximate functionality, accurate upto the current batch determining the non-linear gterm. 2.5.2 On the Nature of Bottlenecks and Embeddings 3 Experiments 3.1 Singular Value Distribution 3.2 Single Task Adaptions 3.2.1 Implications on Performance 3.3 Multi-Task Adaptions 3.3.1 Dual Task versus Two Single Tasks: Task Sharing 3.3.2 Real Time Dynamic Task Adaptions 4 Conclusion 9