Disentangling neural network structure from the weights space
Abstract
Deep Neural Networks have been used to tackle a wide variety of tasks achieving great performance. However, there is still a lack of knowledge of how the training of these models converge and how weights relate to their properties. In this thesis we investigate the structure of the weight space and try to disentangle its properties. Attention mechanisms are introduced to capture relations among neurons' weights that help in weight reconstruction, hyper-parameter classification and accuracy prediction. Our approach further has the potential to work with variable input size allowing different network width, depth or even architecture types.
Full text
Disentangling neural network structure from the weights space Master Thesis submitted to the Faculty of the Escola T`ecnica d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by Pol Caselles Rico In partial fulfillment of the requirements for the master in Advanced Telecommunication Technologies Advisors: Konstantin Sch¨urholt (ICS-HSG), Damian Borth (ICS-HSG), Xavier Gir´o-i-Nieto (UPC) Barcelona, Date 20/01/2021
2
Contents List of Figures 4 List of Tables 5 1 Introduction 8 1.1 Ganttdiagram ................................. 11 2 State of the art and related work 12 2.1 Representation learning in NN weight space . . . . . . . . . . . . . . . . . 12 2.2 Attention architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.3 Visualization and dimensionality reduction . . . . . . . . . . . . . . . . . . 16 3 Methodology 17 3.1 Attention auto-encoder I (AttnAE I) . . . . . . . . . . . . . . . . . . . . . 17 3.2 Attention auto-encoder II (AttnAE II) . . . . . . . . . . . . . . . . . . . . 20 3.3 Dataaugmentation............................... 22 3.4 Manifold Visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 4 Experiments and results 23 4.1 Metrics...................................... 23 4.1.1 Signal-to-noise ratio (SNR) . . . . . . . . . . . . . . . . . . . . . . 23 4.1.2 Coefficient of determination score (R2) ................ 24 4.2 Datasetexploration............................... 24 4.2.1 Tetris-Seed dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.2.2 Tetris-hyper dataset . . . . . . . . . . . . . . . . . . . . . . . . . . 29 4.2.3 Googledataset ............................. 31 4.3 Tetris-seed and tetris-hyper datasets experiments . . . . . . . . . . . . . . 35 4.3.1 Model weights reconstrucction . . . . . . . . . . . . . . . . . . . . . 36 4.3.2 Downstream tasks: accuracy and hyper-parameters prediction . . . 40 4.3.3 Attention score maps interpretability . . . . . . . . . . . . . . . . . 44 4.4 Google dataset experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.4.1 Model exploration and code validation . . . . . . . . . . . . . . . . 50 4.4.2 Model weights reconstrucction . . . . . . . . . . . . . . . . . . . . . 52 4.4.3 Downstream tasks: accuracy and hyper-parameter prediction . . . . 53 5 Environment impact 59 6 Conclusions 60 6.1 Futurework................................... 61 References 62 3
List of Figures 1 Mainsetup ................................... 9 2 Ganttdiagram ................................. 11 3 Meta-classification performance maps . . . . . . . . . . . . . . . . . . . . . 13 4 Tetris model architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 5 AttnAE I encoder architecture . . . . . . . . . . . . . . . . . . . . . . . . . 19 6 AttnAE I decoder architecture . . . . . . . . . . . . . . . . . . . . . . . . . 19 7 AttnAE II encoder architecture . . . . . . . . . . . . . . . . . . . . . . . . 20 8 Tetrispiecesdataset .............................. 24 9 UMAP: weight space and statistics of the tetris dataset . . . . . . . . . . . 26 10 Visualization of basic statisics of the tetris dataset . . . . . . . . . . . . . . 27 11 UMAP: comparison of data encoding of the tetris dataset . . . . . . . . . . 28 12 UMAP: weights and statistics space on tetris-hyper dataset . . . . . . . . . 30 13 UMAP: weights on google dataset . . . . . . . . . . . . . . . . . . . . . . . 33 14 UMAP: statistics on google dataset . . . . . . . . . . . . . . . . . . . . . . 34 15 Reconstruction R2score over epochs and representation learning progress . 38 16 AttnAE attention score maps . . . . . . . . . . . . . . . . . . . . . . . . . 45 17 Correlation matrix of attention score maps . . . . . . . . . . . . . . . . . . 46 18 AttnAE architecture noise robustness over different SNR thresholds . . . . 47 19 AttnAE architecture noise robustness over epochs . . . . . . . . . . . . . . 48 20 Recall curves for hyper-parameter classification . . . . . . . . . . . . . . . 56 21 Accuracy distribution over epochs . . . . . . . . . . . . . . . . . . . . . . . 57 22 Epoch prediction in the Google dataset . . . . . . . . . . . . . . . . . . . . 58 4
List of Tables 1 Datasetproperties ............................... 25 2 Correlation between accuracy and statistics on the tetris-seed dataset . . . 28 3 Correlation between accuracy and statistics on the google dataset . . . . . 32 4 Accuracy prediction results on tetris dataset . . . . . . . . . . . . . . . . . 36 5 Accuracy prediction on tetris-seed and tetris-hyper datasets . . . . . . . . 41 6 Hyper-paramaters prediction on the tetris-hyper dataset . . . . . . . . . . 42 7 Accuracy prediction from the AE embeddings . . . . . . . . . . . . . . . . 43 8 Accuracy prediction baselines comparsion on google dataset . . . . . . . . 51 9 Accuracy prediction results on google dataset . . . . . . . . . . . . . . . . 53 10 Hyper-parameters classification results on google dataset . . . . . . . . . . 55 5
Abstract Deep Neural Networks have been used to tackle a wide variety of tasks achieving great performance. However, there is still a lack of knowledge of how the training of these models converge and how weights relate to their properties. In this thesis we investigate the structure of the weight space and try to disentangle its properties. Attention mechanisms are introduced to capture relations among neurons’ weights that help in weight reconstruction, hyper-parameter classification and accuracy prediction. Our approach further has the potential to work with variable input size allowing different network width, depth or even architecture types. 6
Acknowledgments I would like to thank my tutor, Xavier Giro-i-Nieto, for granting me the opportunity to do my master’s internship at ICS-HSG in Sant Gallen, Switzerland. I want to acknowledge Konstantin Sch¨urholt for his huge help during the development of the thesis. I would also thank to Damian Borth for the kind hospitality I received at all times and for having accepted me in the ICS-HSG lab. 7
1 Introduction In recent years, Deep Neural Networks (DNN) have been used to tackle a wide variety of tasks, achieving incredible results. Looking at the amount of papers where these methodologies have been used, it has not stopped growing. DNN have been applied successfully to a wide different domains such as language, 3D reconstruction and analysis, image and video classification, creation and understanding, speech analysis and synthesis and so much more. DNN are increasingly employed to real-world use-cases and have become the state of the art of a lot of approaches. However, their structures and operations are still poorly understood. There is a lot of information and intuition but more research into this direction is needed to fully describe and define them in depth. Mathematical fundamentals and a lot of techniques are know to train DNN, but due to the high dimension of the space in which these technologies work (in some cases up to 100B trainable parameters) it is difficult to explore it with the computation power that there is nowadays. There are still countless open questions about the solutions reached and how and why they really work: why do some networks generalize, while others do not? Why do some networks learn a bias, while others do not? Looking directly at the weight space of the networks, are different networks trained on the same task learning the same features or are different solutions? How unique are the solutions in the weight space? Can we map or group these unique solutions to some latent space that we can then take advantage to be applied into any other downstream task? How do different architectures trained on the same task relate to each other? Do different domains and tasks guide how the model ends up to different weights? Nowadays it is difficult to answer all these questions without uncertainty, and more explanations are still needed. Gaining insight into the model weight structure may change the way new models are trained. Having a better understanding can lead us to potentially develop new strategies during the model design processes. This knowledge can be also used in the field of intellectual property to define metrics to certificate and patent them. Versioning and diagnostics can also be potentiated by the influence of models’ knowledge providing explanations of their behaviour and how the appropriate structure should be. Having full knowledge of models’ weights trajectories and how they relate to accuracy we could save time and computing power consumption. This knowledge can be used to predict which models are or are not going to flourish. Being able to compress model weights can also be useful in weight deployment transmission and storage savings. Another interesting topic that can take advantage in the field of the DNN is called interpretability. Fully understanding them 8
can lead us to discover new properties of the neural networks. Nowadays, in order to fine tune our models we usually retrain them by applying small changes in a wide variety of hyper-parameters. There is a field called meta-learning that studies the evolution of the models during training in order to learn how to learn. The explainable artificial intelligence field is concerned with explaining how models work beyond the theoretical part. It attempts to answer questions of a moral nature or of interpretation of their functioning and to explain the reason why some decisions are made or others not. Given that the field of artificial intelligence is the mathematics that defines the criteria of decision, investigating the question can help to understand the solutions obtained. The data for this thesis are the models trained on different datasets, tasks, architectures, properties of training and the randomness of the iterative procedures (split of the dataset, batch composition, dropout and so on). Our goal is to generate model-level embeddings that are rich enough to define the entire architecture in time (including current state in training, e.g. epoch), dataset on which the model was trained on and properties defined during training as well as the structure of the architecture. Achieving this goal would mean that it will be able to understand the underlying structure. Due to the dimensionality of the solution space, finding analytical solutions is generally infeasible. Therefore, we have opted for the use of deep learning methodologies as they are designed to work with large volumes of data. The main idea is to focus on a subset of models with a limited number of tunable parameters and try to disentangle the structure from the weight perspective. There is a wide range of architectures and possibilities during training. For this reason we have worked on vanilla convolutional and feed forward networks, of model sizes between 100 to 5000 weights for each sample. Figure 1: The main setup: define a bidirectional mapping between model architecture and its properties and a model-level embeddings that are rich enough to define it. 9
the concept of tokenising model weights to generate a sequence of embeddings. We also incorporate global tokens at the beginning to generate global representations. 2.3 Visualization and dimensionality reduction In domains where each of the data available has a large number of features, it is usually necessary to perform one or more methods to reduce their size. This process is used in order to eliminate information that is not relevant or does not provide anything new. In many occasions it is necessary to reduce the size of the data to be able to train models or also to be able to visualize the data in 2D representations. The main techniques of dimensionality reduction can be summarised in the methods based on feature selection, those based on matrix factorization, those based on manifold learning or the auto-encoder methods. Moreover, these can be classified in two main groups, those that make a linear reduction and those that are non-linear. In order to be able to represent large dimensional data, the principal component analysis (PCA) method and the non-linear tSNE [16] and the uniform manifold approximation and projection for dimension reduction (UMAP) [17] methods have become very popular. In general, PCA is used for fast reduction and elimination of information that is not relevant. In contrast, tSNE and PCA are used for large compression up to two or three dimensions. They are mainly used to visualise large dimensional data. tSNE methodology is based on t-distributed stochastic neighbour embedding with nonlinear scaling to represent changes at different levels, it preserves local structure in the data. On the other hand, UMAP claims to preserve both local and most of the global structure in the data. Unlike PCA which is deterministic, these technique does not expect the relationship to be linear. Both algorithms are highly stochastic and very much dependent on choice of hyperparameters (t-SNE more than UMAP) and can yield very different results in different runs, so the plots might obfuscate an information in the data that a subsequent run might reveal. 16
3 Methodology In this thesis I have focused on the representation learning from the weights space. I have tackled different tasks such as accuracy prediction, hyper-parameter classification, weights’ compression and noise robustness using in all cases attention mechanisms. In this chapter, we present the auto-encoder architecture attention auto-encoder I (AttnAE I) and its evolution (AttnAE II) in order to be able to deal with variable input sizes. I also present the insights of the Unterthiner et al.[1] Google’s dataset and ours tetrisseed and tetri-hyper datasets. I finish off with some definitions of the metrics and data augmentation techniques that have been used. A model based on neural networks uses a series of mathematical operations to modify its weights during training. This process establishes structure in the weight space. Due to the dimensionality of the environment, the space is going to be sparse. Therefore, it may be difficult to establish relationships between models and to understand the interactions. Attention-based architectures have the potential to learn the relationships that exist among weights. Multi-head attention and positional embeddings both provide information about the relationship between different embeddings. Besides, the transformer encoder architecture does not suffer from long dependency issues and it is suitable for domains with a large amount of training data. Data augmentation techniques such as applying Gaussian noise, erasing input sequences and applying permutations (they do not modify the functioning of the model) have been studied. 3.1 Attention auto-encoder I (AttnAE I) Attention mechanisms have been used in a wide variety of tasks. In recent years the wellknown Transformer architecture has been used and adapted to different fields due to its great capacity to work well. For this reason we decided to use the encoder of the complete model as the main block of our architecture to compute the attention. The main block does not reduce the dimensionality, and it also needs all input embeddings Zito be the same size. For this reason, the difficulty to use it lies in the methodology used to adapt or translate the model weight information into a series of embeddings. In previous works [1] [2], the input was defined as a vector of a concatenation of all the weights of the model. In our case, a criteria has been defined to separate the weights to generate a sequence of embeddings that represents our model. Depending on how the weights have been transformed into these embeddings, they are described as follows: 17
•Neuron: Each weight is passed through a Multi Layer Perceptron (MLP) to obtain an embedding of size dmodel. •NeuronGroup: All the wheights that compose each neuron are mapped into an embedding of size dmodel. In case the model is a convolutional neural network (CNN) then all the weights of each kernel are flattened and are used as a group. Biases are included in its corresponding group/kernel weights added at the end. •Layer: All the weights of all neuron for each layer are mapped into an embedding of size dmodel. In case the model is a convolutional neural network (CNN), then all the weights of all kernels for each layer are used as a group. Figure 4: Model architecture used in the tetris dataset. The weights of this model at different checkpoints is considered one sample of the tetris-seed dataset. The colours shown refer to the NeuronGroup’s embedding system (an embedding will be generated from all the weights corresponding to each of the neurons). Colors refer to the embeddings in figure 5. Each group of weights for a model is encoded with a different MLP to a fixed size dmodel (see Figure 5). After the transformer encoder, all the dmodel embeddings are concatenated. The function fseq2neck is defined as a MLP of one layer from the size length sequence∗d model to the size of the bottleneck. With this approach we have to be aware of how it scales as the length of the sequence and size of the embeddings increases. The reason is that with this structure the dimension of all the embeddings must be reduced to the size of the bottleneck. For large models this could be an issue becasue the dimensionality of the function fseq2neck should be increased. What it would be doing then is using a usual autoecoder MLP, where attention mechanism has been used to generate such a vector. In short, attention would be simply added before performing a regular dimensionality reduction. During all the Nblocks applied in the transformer encoder it does neither reduce the dimensionality nor telling the model to summarize the key features of each embedding. The best results were obtained by forcing the bottleneck to have values between -1 and 1 by applying tanh function on it. 18
Figure 5: Attention encoder used to map weight’s vector to a smaller latent space embedding. The number of input embedding into the transformer encoder remains exactly the same at the output of it. In order to recover again the input vector from the bottleneck, a different MLP is defined from the neck size to dmodel size. We have then n(where nis the length input embedding sequence) linear layers to generate the input embeddings to the decoder transformer encoder block. Afterwards, each of the embedding Z0 iis mapped to the correspondent group of weights that it belongs to. For example, if neuronGroup is being used, each embedding Z0 iwould contain the necessary information to be able to recover the weights corresponding to that embedding. For each of the embeddings a MLP is used to transform it back to the weights. For example, the Z0 1would be converted back to the first 16 values. With this approach the model is forced to learn attention in order to be able to capture information of the other neurons in the reconstructed vector. Figure 6: Attention decoder used to map bottleneck to the reconstructed weight’s vector. This architecture forces the model to use attention in order to capture information of other groups. 19
3.2 Attention auto-encoder II (AttnAE II) In the case that the datasets contain architectures of different lengths, the previous architecture cannot be used, because it would be necessary to train a specific embedder for those new weights. In addition, if the models are very large, the function fseq2neck will be very complex, and the difficulty will be focused on a standard auto-encoder. So the previous architecture does not scale and is not flexible enough to deal with variable input size models. Expanding the idea from vision transformer (ViT) [7] and BERT [8] (where special token is applied to represent the meaning of the entire model), we have introduced some changes in the input embedder and the way we compute the bottleneck. Figure 7: Attention encoder II. It uses unique embedders for each group of weights. It implements sine/cosine position encoding as well as Lktokens to generate global embeddings of the models. Depending on the Encoder type, the embedders will end up slightly different. In the case of the NeuronGroup, a single MLP is defiend for each type of neuron. In the tetris-seed dataset (described in section 4.2)there are two types, the first layer (where each neuron has 16 weights) and the second layer (where each neuron has 5 weights). In this particular case, two linear mappings will be learnt from sizes 16 and 5 to dmodel embedding size, called MLP1and MLP2. In order to make this step as scalable as possible being able to use different model sizes, sine/cosine position embedding have been used. Taking into account that in an autoencoder setup the size of the bottleneck needs to be reduced with respect to the input vector, with a transformer encoder the size is not actually being reduced, because the input and the output are exactly the same sequence size. Therefore, the interpretation of the output of the autoencoder has been changed: 20
•Lklearneable embeddings have been defined which are concatenated at the input sequence. We cannot take any other embedding from the input sequence because its output is the token model’s representation. We add a token which has no other purpose than being a model-level representation. After N blocks of the transformer encoder, we are going to take as output only the corresponding embeddings to the ones we introduced. In essence, we can understand these learnebale embeddings as a way of telling the model that those embeddings are not actually the information of the input model, but to encode the entire model during each block of the transformer. •In order to generate the bottleneck, the function seq2neck is a linear layer from the concatenated Xkembeddings to the size of the neck. This setup is able to deal with different layer sizes without changing the architecture. The size of the compression can be changed in the transformer encoder by varying the number of Lktokens. Due to the domain field, it will be very possible to have to deal with very large sequences. In this case the implementation of Reformer [9] for fast attention computation can be done. They replace dot-product attention by one that uses localitysensitive hashing, changing its complexity from O(L2) to O(LlogL), where L is the length of the sequence. In case it is needed to increase the number of transformer encoder blocks N, reversible residual layers are used instead of the standard residuals, which allows storing activations only once in the training process instead of N times. We have not implemented it in the proposed methodology. For the decoder, a similar approach may apply, a MLP may be mapped from the bottleneck to embedding of d model size plus sine/cosine position embedding. In this case the learneable embeddings have not been used due to the fact that the information is not being compressed. The length of the sequence is exactly the same used to encode the model. From the output of the transformer encoder block a MLP is used for each type of neuron that is going to be reconstructed. In our example, each Z0nis directly the compressed information representation embedding to reconstruct each GroupNeuron. Afterwards, all the output vectors are concatenated to reconstruct the input vector. 21
3.3 Data augmentation Generally to be able to train models with many parameters, we usually need very large datasets. One way to be able to get new samples is by using data augmentation techniques applied at the existing dataset. In the domain of neural networks, one method of being able to generate new samples without modifying their behaviour is by performing permutations among the neurons within a same layer. With Lithe number of neurons at each iof Nlayers N∈Nin a deep fully connected network, it can be generated up to QN−1 i=1 Li!−1, forN > 1 permutations without affecting its operation. Same procedure will be applied when using convolutional neural networks (CNN) on the order of the kernels. When applying permutations layer-wise, the order of the weights of each neuron of the following layer must change accordingly to the permutation applied before. 3.4 Manifold Generation and Visualization UMAP vs. t-SNE The main difference between t-SNE and UMAP is the interpretation of the distance among clusters. t-SNE preserves local structure in the data. UMAP claims to preserve both local and most of the global structure in the data. In t-SNE the distance does not mean anything, close proximity is highly informative, distant proximity is not very interesting and cannot rationalise distances, or add in more data. On the other hand UMAP is faster to compute than tSNE. It can preserve more global structure than tSNE, it can run on raw data without PCA preprocessing, an allow new data to be added to an existing projection. Instead of the single perplexity value in tSNE, UMAP defines nearest neighbours as the number of expected nearest neighbours and minimum distance as how tightly UMAP packs points which are close together. For these reasons, we used UMAP as it can cope with non-linear scaling and can reduce to 2D well. Due to the dimensionality of the domain the difference in time of the two methods is of great importance. In all the representations of the models in 2D visualizations they correspond in applying UMAP reduction up to two components. Where the first component is the x-axis and the second component is the y-axis. 22
4 Experiments and results Training neural networks is generally assumed to impress a pattern in the parameter space, determined by dataset, task and training regime. The main goal is to understand the relationships that exists among neurons. From this knowledge we expect that it helps in a wide variety of downstream tasks. To investigate the weight space for patterns, we apply representation learning on three datasets (tetris-seed, tetris-hyper, and google). On each dataset, we present results of weight reconstruction, accuracy prediction and hyperparameter classification. On the tetris-dataset we exapand the downstream tasks to be computed on the compressed embeddings generated by the auto-encoder setup. The experiments and results obtained in the different sections will be presented subsequently. They have been separated into two main groups, the experiments performed in the tetris datasets and the experiments performed in the Google dataset. 4.1 Metrics 4.1.1 Signal-to-noise ratio (SNR) Signal-to-noise ratio is defined as the ratio of the power of the input vector to the power of the distortion injected noise. Logarithmic decibel scale has been used due to very wide dynamic range of the input. SNRdB = 10 log10 Pinput Pnoise ;Px=E[X2] = 1 N N X i=0 x2 i(1) In order to perturb different signals in our setup, we decided to use Gaussian noise Xnoise ∼ N(0, σ2) . Taking it into account and using the expression described in 1, we can define sigma as follows: PXnoise =E[X2 noise] = σ2 noise where σnoise(Pinput, snr) = qPinput ·10− snrdB 10 (2) Where σnoise is computed for a given input power and the desired signal to noise ratio to be applied. 23
4.1.2 Coefficient of determination score (R2) R2is a statistic that provides information about the goodness of fit of a model. In regression, the R2coefficient of determination is a statistical measure of how well the regression predictions approximate the real data points. Non-positive values indicate that we are not doing better than fitting a constant predictor (horizontal hyperplane), and values close to 1 indicate that the regression predictions perfectly fit the data. This implies that 100% of the variability of the dependent variable has been accounted for. R2= 1 −MSEres MSEtot = 1 −Pi(yi−byi)2 Pi(yi−¯y)2(3) R2score is scale invariant and multiplying the outputs by a constant will not change the metric. In essence, the coefficient of determination is a relative measure which compares the MSE of the Model to the MSE of a constant prediction. 4.2 Dataset exploration In this chapter, we will start off with the three datasets used in the thesis (Google , tetris-seed and tetris-hyper datasets), continue with its explanation and describing its properties. Public datasets already exists for neural networks trained on popular vision datasets as described in Chapter 2, although, it is for its complexity that only the google dataset [1] has been picked and we have worked on a smaller custom generated datasets, either to validate code as well as to better understand the results due to the reduction in size. Figure 8: Tetris pieces used to generate tetris-seed and tetris-hyper datasets on the classification task. This representation has been adapted into a color space images. The pictures used in training are composed of one channel images where the background is white and the pieces are black. 24
The custom tetris dataset is designed to be much more tractable in terms of complexity and size. It is composed of models of fixed size (two feed forward layers, five and four neurons each, respectively). The images on which these models have been trained were also generated. It is composed of four tetris pieces in a 4 by 4 black-and-white pixels setting. The models were trained to classify which piece each image had. The models have been trained for 75 epochs storing each checkpoint of a vectorized version of size 100 (model’s weights have been flattened). Dataset Model type Samples Parameters Datsets trained Model diferences Google [1] CNN ∼1,1M 4970 Mnist Fashion mnist Cifar10 Svhn Optimizer Init method Activation Tetris-seed FCN ∼75K 100 Tetris shapes Init seed Tetris-hyper FCN ∼105K 100 Tetris shapes Init method Activation Table 1: Dataset properties. Comparison between Google and tetris datasets. The google dataset consists of a 120k models trained on Mnist, Fashion-mnist, cifar-10 and svhn equally distributed. For each model trained there have been stored 9 checkpoints at epochs 0, 1, 2, 3, 4, 20, 40, 60, 80 and 86. The difference between trainings are the initialization function (Random Normal, Truncated Normal, glorot normal, he normal, orthogonal) and the activation function (relu, tanh). Same seed have been used for all the trainings. In the following model representations Uniform Manifold Approximation and Projection for Dimension Reduction (UMAP) has been used to generate two-dimensinal plots. In all representations, the x-axis and y-axis refers to the first and second components. In the following graphs I do not differentiate between different trainings, so all models and their different checkpoints will be shown as points. All plots have been generated using the same hyperparameters with minimum distance set to 0.1, number of neighbours set to 50 and using euclidean distance function. 25
Correlation Linear regression Pearson Spearman MSE R2 Mean -0.26 -0.20 0.11 0.07 Variance 0.18 0.72 0.11 0.03 Percentile 0 -0.40 -0.73 0.099 0.16 Percentile 25 0.36 -0.56 0.10 0.13 Percentile 50 -0.05 -0.09 0.12 0.002 Percentile 75 0.30 0.52 0.11 0.09 Percentile 100 0.40 0.73 0.099 0.156 Stack All - - 0.08 0.30 Table 3: Pearson and Spearman correlation between the statistics of the model weights and the accuracy prediction of the model on the google dataset. Linear regression applied for each of the statistics and also stacking all as seven ordered variables. As explained in the tetris datasets, Pearson and Spearman correlation have been applied between basic statistics and accuracy (see Table 3). Correlation values between those are really high. Same pattern can be seen in previous smaller datasets. In this dataset the model trained has a difference on the architecture (in this case is a convolutional neural network). Nonetheless, we are able to predict accuracy up to 0.3R2score with linear regression from the concatenation of the statistics. Comparing results with the tetrisdataset, the correlation values between statsitics are different, meaning that I cannot generalize which statistics better define the overall accuracy score. Overall, these results suggests that having information of the distribution of the model’s weights give us a lot of information of the model’s behaviour. This shortcut can lead us to fail when defining and designing architectures. If computing mean and variance is enough for achieving good performance, the model will end up learning this transformation instead of learning the relation between weights. Besides, permutations as a data augmentation will not help because this data augmentation is invariant to the mean and variance transformations. 32
Figure 13: UMAP embedding applied directly from the input vector of dimension 4970 to two dimensions. Top left image show the accuracy of each sample (low accuracy in blue, high accuracy in yellow). Top right show the checkpoint during training (checkpoints: 0-magenta, 1,2-green, 3-blue, 40,60,80-red, 86-yellow). Bottom left image show the activation function (relu-magenta, tanh-green). Bottom right image show the inisialitation function (RandomNormal-red, TruncatedNormal-blue, glorotNormal-green, heNormal-magenta, orthogonal-yellow) 33
Figure 14: UMAP embedding of statistics from the input vector of dimension 7 to two dimensions. Top left image show the accuracy of each sample (low accuracy in blue, high accuracy in yellow). Top right show the checkpoint during training (checkpoints: 0-magenta, 1,2-green, 3-blue, 40,60,80-red, 86-yellow). Bottom left image show the activation function (relu-magenta, tanh-green). Bottom right image show the inisialitation function (RandomNormal-red, TruncatedNormal-blue, glorotNormal-green, heNormal-magenta, orthogonal-yellow) 34
4.3 Tetris-seed and tetris-hyper datasets experiments We begin our experiments on the tetris-seed and tetris-hyper datasetes, as they are the smallest dataset. On both datasets, we have performed compression and accuracy prediction tasks with AttnAE I and AttnAE II architectures as well as vanilla DNN and PCA based auto-encoders. In a further step, we also present information regarding the interpretability of the attention score maps as well as the robustness to noise when performing compression tasks. To establish the existence of patterns and in an attempt to learn useful embeddings, we learn incomplete auto-encoders with a compression ratio of 5:1 in all experiments (from input vector of size 100 to a vector size of 20). The dataset was divided into two groups of 50% of the samples for each group. I have also ensured that each model with all its 75 checkpoints are to the same corresponding split. All the values presented are computed on the test split. The number of models for each checkpoint are equally distributed on both splits. We investigate whether auto-encoder for weight reconstruction yields embedding spaces, that are useful for general downstream tasks and preserve information on the samples embedded. To evaluate the amount of information that is preserved, accuracy prediction was performed using linear regression and vanilla multi-layer perceptron (MLP). 35
4.3.1 Model weights reconstrucction For the AE reconstruction task, we compare several architectures. As baseline, we consider PCA and a MLP-based Auto-encoder. Further, we apply the transformer-based architectures we propose in the previous section, the first version (AttnAE I) and the second update (AttnAE II) using the general context tokens concepts used in Bert and ViT. To evaluate the quality of the reconstruction we have used the metrics R2between the input vector and the reconstructed vector of all the samples of the test set. Each sample of the dataset is represented by a vector where all the weights are concatenated. The vanilla auto-encoder uses this vector directly. However, AttnAE I/II are defined to work with embeddings which represents the models. The neuronGroup approach (where all the weights associated with each neuron are used to generate an embedding) was used to encode the input data. This decision was taken due to its better performance. It was considered a reasonable form to spread the information without generating too many embeddings (an embedding for each weight) or losing information among different neurons (generating one embedding for all the weights for each layer). Deep neural networks are able to modulate nonlinear functions, therefore they should obtain at least the same performance as PCA, which relies on a linear auto-encoder. All experiments were trained for 8000 epochs (except PCA). The experiments where data augmentation was applied only permutations on the weights of each model were used as permutation. As a reminder, this augmentation generates equivalent samples without changing the performance. 120 possible permutations were generated and applied randomly during training. Architecture Embed Size Parameters (N, dmodel)Model Size Data Augmentation R2Test Reconstruction PCA 20 - - No 0,352 10 layer MLP 20 - ∼90k Permutations 0,756 AttnAE I 20 1, 20 ∼16k No 0,643 AttnAE I 20 1, 128 ∼80k Permutations 0,830 AttnAE I 20 2, 128 ∼180k Permutations 0,843 AttnAE II 20 4, 128 ∼100k Permutations 0,865 Table 4: Model performance in weight reconstruction task. PCA computed with a linear kernel. 36
The results of our experiments are presented in Table 3. All non-linear models have achieved better performance than linear auto-encoder (Principal Component Analysis for reconstruction in a setup 100-20-100). Comparatively, the two attention-based versions (AttnAE I/II) achieved better scores than regular multi-layer perceptron auto-encoder (10 layer MLP for the encoder and symmetric decoder). The AttnAE II model achieved the best reconstruction score. In both cases it was necessary to use permutations as data augmentation techniques in order to improve the results, otherwise the model suffers from huge over-fitting and performance is considerably lower. In the case of MLP, without permutations, values around 0 R2were obtained. Without data augmentation the variance could not be predicted and only the mean value was forecasted. The results shown in Table 4 clearly show how transformer-based systems achieve superior results in reconstruction. It has also been proven that they have a higher convergence speed than the regular vanilla MLP auto-encoder as seen in Figure 15. In the case of AttnAE II, it was able to work with models that had not seen before with a different number of weights. However, in this comparison the whole dataset is composed of a fixed architecture. The fact that AttnAE II uses the same encoder for each type of neuron greatly reduces the size of the model and scales better as the input has more neurons. For the model to converge it is necessary to increase the number of transformer encoder blocks so that the model-level representation token is rich enough to perform the compression. This behaviour may be due to the reduction in the number of trainable parameters used in the embedder to generate the tokens that describe the input models. That is the reason for increasing the complexity in the transformer encoder block where the compression is generated. This combination enables to reduce the complexity of the encoding of the input with the disadvantage of increasing the number of transformer encoder blocks needed to make the model converge. The main hypothesis is that as the models are being trained, their weights will get structured depending on different parameters, such as which data have been used for training, which architecture has been chosen or which hyper-parameters have been selected. Therefore, based on the good functioning of AttnAE I/II architecture that uses attention, the hypothesis is that it is possible to better identify and capture the relationship between the embeddings. In these experiements, we used neuronGroup to generate an embedding from the weights that composes each neuron. We expect that the reconstruction performance increases as the weights gain more structure. 37
Figure 15: Left figure: Reconstruction R2score performance over epochs. Each dot represents the score computed only for models that belongs to the selected checkpoint in the test split. The model has been trained with models from all epochs without any distinction. Right Figure: Model reconstruction R2score performance on the first 1k epochs training. Results given in the test split. To get a better understanding where structure exist, we evaluate the reconstruction performance over the epoch of the samples (see Figure 15). The AttnAE II model with 100K parameters was chosen to evaluate the reconstruction R2score. The test split consists of models belonging to different epochs. The results presented in Table 4 were computed on all the samples without differentiating in which group they belonged. The model used was trained with samples belonging to all groups without distinction. The reconstruction R2in models belonging to the firsts epochs is considerably lower than the models that have already converged (the models are expected to be converged on the last checkpoints). The weights of the models in the first epochs are samples from a known distribution (they have not gotten structures from their properties). This implies that we are able to predict the mean of the population, but it is difficult to predict the exact realisation from the given input. Even knowing the distribution we are sampling from, the best it an be done is to sample from the same distribution so the reconstruction predictions will match the mean and will have the same variance. This behaviour can be attributed to the difference in structure that exists between the models with random weights and those that have acquired structure during their training. This results validate the idea that at the beginning there is no structure among weights (just samples from distribution) that is why we cannot understand and capture the relation because there have not been place any back-propagation update yet. However, as the training progresses, the weights are updated in a particular manner depending on wide 38
variety of parameters (such us the dataset to train it or the hyper parameters used) giving them structure. Our hypotesis is that more structure in the weights space allows the attention mechanism to capture this interaction and therefore it is able to reconstruct the weights (because these values are not random samples from a given distribution). 39
4.3.2 Downstream tasks: accuracy and hyper-parameters prediction We want to explode the interaction between neurons to better define properties of the models and be meaningful in downstream tasks. In this section, we explore the accuracy prediction and hyperparameters classification (activation function and initialization method) on the tetris-seed and tetris-hyper datasets. To evaluate what properties are encoded in the embeddings learned for reconstruction (for each input vector of size 100 an embedding of size 20 from the bottleneck is generated), we apply linear probing and vanilla MLP from the embeddings to accuracy. Performance is measured using the R2 score metric for the entire test set. Mainly two approaches were used to construct the samples in the two datasets. On the one hand, a linear regressor was used to predict the accuracy in two modalities: Directly from the weights’ vector and on a vector generated with the concatenatenation of basic statistics. In the latter group, the statistics are compared by applying them directly to all the weights in each layer and to the neuronGroup approaches. On the other hand, the AttnEnc architecture was used (adapted from the AttnAE I/II) and the three different methods of grouping the weights have been compared to generate the embeddings (layerwise, neuron and neuronGroup). In order to tackle accuracy prediction and hyperparameters classification tasks, the architecture AttnAE was adapted. To adapt the architecture the encoder part was used (a neural network that transforms the embeddings that describe the model into a reduced fixed size embedding). In order to be able to make predictions, we have adapted the last layer from the bottleneck to fit the requirements: •Accuracy prediction: From the bottleneck vector (generated by the encoder block from AttnAE) a linear layer and sigmoid function is applied. The model generates for each sample a prediction in the range [0,1] as accuracy is defined in this range. •Hyper-parameters classification: From the bottleneck vector (generated by the encoder block from AttnAE) a linear layer of size the number of possible classes to be predicted for each group is used. Afterwards, a softmax function is applied to force the output to be a probability distribution. The class chosen is the one that has achieved greater output value or probability. 40
Architecture Input type Data Augmentation Accuracy R2 tetris-seed Accuracy R2 tetris-hyper Linear regression weights No 0.48 0,23 Linear regression weights statistics No 0.86 0,67 Linear regression weights statistics layer-wise No 0,87 0,74 Linear regression weights statsitics neuronGroup No 0,89 0,74 MLP (4 layers) weights No 0,87 0,71 Attn Enc weights layer-wise No 0,87 0,72 Attn Enc weights neuronGroup No 0,92 0,87 Attn Enc weights neuronGroup Permutations 0,95 0,89 Table 5: Accuracy prediction on tetris-seed and tetris-hyper datasets. Comparison of models using directly the weights and using the basic statistics transformation. As shown in Table 4, applying a linear regressor on the weights’ vector to predict accuracy has obtained 0,48 R2and 0.23 R2for tetris-seed and tetris-hyper datasets. The usage of the statistical transformation with neuronGroup encoding has increased the results up to 0,89 R2and 0,74 R2respectively. Results reported on previous works concluded that pre-processing the weights by computing its statistics layer-wise helps on predicting the accuracy. In our datasets, computing the statistics layer-wise did not give any advantage with respect to the global statistics. We extend these results by applying our attention encoder architecture, which outperforms the linear regression results in all experiments. This results were expected due to the fact that our model is able to map non linear relations and we suppose the interactions among neurons are highly non linear. In our approach we expect the model to learn in an unsupervised manner the interconnections by learning from the neuron’s relations. We have not used neuron encoder because of the poor scalability that this approach offers when dealing with bigger models. We need some criteria to reduce the size of the weight vector. From Table 5 we can state that treating all the weight for each neuron gives us the best performance. Using permutations in training helped to make the model more accurate and gain in terms of generalization. 41
When noise is applied to the latent sapace (see Figure 18), two main regions can be differentiated. The first one is found for low SNR values (the noise is high compared to the signal) where the model that has not been trained with noise in any of the presented methods obtains worse scores. The best model in this aspect is the one that has been trained specifically for this task. Secondly, when the SNR is high (signal is considerably higher than noise) the model that is able to obtain better results is the one that has not been trained with noise. On the other hand, when noise is applied directly to the vector of weights that define the models, a slightly different behaviour is observed. In this case the model able to obtain better results for low SNR values is again the one that has been entered with noise in the input vector. However, when the SNR is high this last model obtains results practically similar to the model that has not used noise in the training. Figure 19: Comparison of attention auto-encoder models for weight reconstruction. Left figure: Reconstruction R2score performance over different Signal to noise ratio (SNR) when injecting Gaussian noise on the latent embedding in test time. Right figure: Reconstruction R2score performance over different Signal to noise ratio (SNR) when injecting Gaussian noise on the input vector in test time. 48
The results show the reconstruction model is very sensitive to noise. Even the models that have been trained to take noise into account in their multiple configurations have a low robustness to it. However, the model that was trained applying noise in the input vector presents a middle point, improving when there is more interference and obtaining practically the same results when there is no noise. The main conclusions are that the weights are very sensitive for reconstruction tasks, where the AttnAE model is not able to correctly capture the interactions. The results presented previously refer to all the samples from all the checkpoints. However, as the values for each of the groups are computed (Figure 20) it is possible to see a difference in the results. The model that has not been trained with noise obtains a better robustness in those models that obtained a low accruacy prediction (in general they correspond to the groups of the initial checkpoints where statistically the weights have not yet converged). However, when the model is evaluated for the groups that in general have converged, the best model is the one that was trained to deal with noise in the latent space, followed by the one that was trained with noise applied in the input vector. In conclusion, the models that were trained with noise obtain a lower score for reconstruction mainly due to a significantly poor performance in the lower groups. However, if only the converged models are considered (only the converged dataset samples, corresponding to the highest checkpoints) and therefore have a higher structure, the noise-trained model applied in the latent space is the one with the best capacity to capture the relationships among neurons. 49
4.4 Google dataset experiments As the last and most complex dataset, we apply my architecture on the published dataset from Google [1] in their four splits of models trained on mnist, fashion-mnist, cifar10 and svhn. Results presented in the tetris-seed and tetris-hyper datasets showed the capability of the attention architectures to successfully work in this domain. The training times have been significantly increased by working with a significantly larger and a more complex dataset. In their work they did not neither present tasks of reconstruction nor hyperparameter prediction. In our case, we have tackled all these downstream tasks as well as the accuracy prediction. 4.4.1 Model exploration and code validation We begin with reproducing the results from the paper Predicting neural network accuracy from weights [1] from the google paper described in the state of the art. I decided to implement their approach to validate that the metrics and setup is working properly. In [1] two technologies were used (deep neural networks and gradient boosting machines) to predict accuracy. They also use two methodologies to perform the computation, using the vector with all weights of each model concatenated and using a basic statistical vector. The results presented by the authors show that both technologies work better using the statistics. In our case a vanilla MLP was implemented similar to the DNN presented in the paper in order to predict accuracy of the model. In addition, we implemented different well known basic architectures to see their behaviour in comparison. Since in our work we have focused on the understanding of weights, we have only considered to use directly the weights’ vector of the models. Except for my vanilla DNN, the other models (have been designed to be able to deal with embeddings that represents the entire model) have used chunks of the the weights’ vector. In this first approach the weights’ vector were splited into equal size chunks. The results presented is using four chunks in total to represent each sample. In all the models the accuracy prediction was achieved using one last neuron plus a sigmoid function. All trainings used MSE as a cost function: •DNN: Vanilla MLP of ten layers. Dropout set to 0.1 and no norm layers. •GRU+Attention: Vanilla bidirectional GRU sequence to sequence. From the resulting sequence attention layer has been used to generate a 128 size embedding. 50
•Transformer Encoder + MLP: After the transformer encoder seq2seq layer, all the embeddings have been concatenated and a vanilla MLP has been applied. •Tranformer Encoder* + MLP: Same approach used in the regular Transformer encoder but in this case using sine/cosine position embeddings. For each of the chunks a different MLP has been used to generate embeddings to feed the transformer encoder. •MLPLayer + SUM: Different MLP have been used to generate each embedding. Afterwards all the embeddings are equally added. Architecture Model Size Input Type MNIST R2 SVHN R2 Baseline GBM - weights 0,988 0,971 Baseline GBM -weights statistics 0,993 0,986 Google Paper [1] Baseline DNN ∼1.5-3.6M weights 0,980 0,931 DNN ∼1.8M weights 0,983 0,946 GRU+Attn ∼2M weights 0,985 0,948 Transformer Encoder ∼2M weights 0,989 0,920 Transformer Encoder* ∼54M weights 0,993 0,975 Our implementation MlpLayer + Sum ∼800K weights 0,990 0,975 Table 8: Comparsion between results from Google Paper and our vanilla implementation in the MNIST and SVHN splits. We have either not use cross validation nor different realisations. Just for validation purposes. The results in Table 7 indicate that, the reproduced DNN architecture obtained similar results to those reported by the authors. Our custom vanilla DNN is in the range of model size with respect to the baseline. Without entering into too many details, the different architectures used obtain a very similar result. However, their functioning is different compared to the vanilla DNN. The main property is that they work with tokens that describe the models. This allows us to develop models capable of working with variable inputs in size. Table 8 it can be seen that regardless of the number of parameters of the models, similar results are achieved. We cannot state that those models work the best for this domain 51
because we did not extensively fine-tune them. This scores make us think that with this dataset (google dataset) it is really easy to predict the accuracy. Taking into account that there is a wide variety of architectures, it is necessary an approach that is able to work with variable input size. Moreover, since according to the first results in Table 8 the models based on attention mechanisms work well in this field, we have decided to explore into this direction. Since models can be very large, we have discouraged the use of LSTM or GRU based systems due to the issue with vanishing gradients. 4.4.2 Model weights reconstrucction In the dataset tetris-seed the reconstruction task has been performed using neuronGroup to generate the tokens that describe each of the models. For this reason, it was decided to use the same approach for the Google dataset. PCA (with linear kernel) was used as the linear method and AttnAE II as the non-linear method. Permutations were also used in the training split as data augmentation technique. The results were evaluated by computing the R2score between the reconstructed weights and the initial weights for all the samples coming from all the checkpoints in the test set. For the transformer-based model, it was necessary to modify the training method. Its size was increased to 2.5M parameters, using 4 blocks N= 4, and four heads for each one h= 4. The learning rate was decreased lr = 1e−5and it was reduced every 150 epochs with a ratio 10:1. It was trained for 600 epochs. The size of the latent space was set at 128. Using the linear system for a bottleneck measurement of 128, a R2of 0.225 was achieved in test on the MNIST partition. In the case of the non-linear model a score of 0.214 R2has been achieved. The use of permutations has been crucial to make the model converge, however the results have not been as expected. Different configurations of the hyperparameters, modification of the bottleneck size and different schedulers were used to modify the learning rate, however in no case it has been possible to exceed the baseline imposed by the PCA. The results are similar for all the dataset partitions (Fashion-Mnist, svhn, cifar10 and mnist). In addition, all models present a great instability between train and test partitions being very sensitive to small modifications. Due to the limited time available to complete the thesis, it was decided not to continue in this direction. 52
4.4.3 Downstream tasks: accuracy and hyper-parameter prediction In [1] they predicted the accuracy from the weights and from the basic statistics. In order to tackle this task, an adaptation of the AttnAE II model was used to predict accuracy. The same approach described in the tetris-seed dataset to predict the accuracy of the models (the encoder part of the attention auto-encoder with neuronGroup embeddings) was used. As explained in the tetris-seed downstream section, a last layer plus sigmoid function has been applied from the bottleneck vector to predict accuracy values in the range of [0,1]. The metric used is R2score computed between the predictions and the real value of the accuracys for all the samples from the test set. In this dataset the first layers of each sample in the collection of models are convolutional layers. For this reason, the neuronGroup embedding was applied using all the weights that compose each kernel for each of the dimensions. In this case, the weights were flattened and 2-dimensional position information has not been included. In addition, the value of the bias corresponding to each kernel was added at the end of each of these vectors. The number of embeddings necessary to represent the model is composed by 58 tokens (3x16 corresponding to the three initial convolutional layers and 1x10 for the last dense layer). Finally, the model was composed of a two global information context tokens plus the 58 representative tokens of the model. Architecture Input type MNIST SVHN CIFAR10 FASHION DNN Google Paper [1] weights 0,980 0,931 0,954 0,980 AttnPred neuronGroup (ours) weights 0,992 0,972 0,975 0,990 GBM Google Paper [1] weights statistics 0,993 0,986 0,984 0,993 Table 9: Comparsion accuracy prediction results between our approach (AttnPred) and the best experiments obtained in the google paper for their respective input type. Values are expressed in R2score values. DNN from [1] is composed of 1.5-3.6M parameters. AttnPred is composed of 600K parameters. The results show in Table 9 better accuracy prediction performance in all four splits of the dataset compared to the models that were trained using the weights vector. However, the performance is not as good as the best approach the authors performed, using GBM applied on the statistics vector layer-wise. However, the main characteristic of our model 53
with respect to the proposed ones is that it is designed to be able to predict models of variable length and it has 600K prameters instead of 1.5-2.6M of the Gools DNN [1]. The structure generated in the weights space during training tends to be highly correlated with accuracy. The basic statistics show good behaviour in capturing the information needed for prediction. It is possible that the attention-based model is learning to generate these statistics for prediction rather than modelling the connections between the weights. This could explain the similarity of the results without surpassing them. This correlation between weights structure and accuracy may be due to the fact that the generated models have been trained only on one dataset. Furthermore, we are assuming that the partitions split between train and test come from the same source. Once the model is trained, the accuracy obtained will also depend on the source of the image used in the test. The same model with defined weights will have a different accuracy depending on the test dataset. The results presented show that the statistics work well in predicting the performance of the model using the same dataset, however, not understanding the behaviour of the model is probably not enough to predict its performance on other datasets. In [1] there is no data regarding the predictions made on the hyper-parameters that identify each of the models, however in the dataset the authors provide all the information concerning their training and architecture. For each sample they provide the type of activation, initialization and optimizer function used as well as the epoch where the model becomes. In particular they provide models in just nine different groups (0, 2, 3, 4, 20, 40, 60, 80, 86). To perform the hyperparameter prediction task it was approached as a classification problem. However, in the epoch prediction task, we decided to approach it as a regression because there is a difference between classifying an epoch closer to or further away from the true value. Two predictors were used, a log-linear classifier (using the statistic’s vector) and a non-linear architecture (using directly the weights). The same architecture used for accuracy prediction was adapted to obtain the number of neurons corresponding to the number of classes to predict. We apply softmax on the logits and assign labels on the highest prediction value. To evaluate the predictor the F1 score was used (to measure the accuracy and the recall in the same metric). For each of the tasks a specific classifier was trained, minimising the binary cross entropy function as a cost function. 54
Classification Labels Architecture MNIST SVHN CIFAR10 FASHION Linear (weights s) 0,822 0,823 0,807 0,839 Activation (2 classes) AttnPred (neuronGroup) 0,903 0,924 0,894 0,912 Linear (weights s) 0,735 0,749 0,709 0,773 Init method (5 classes) AttnPred (neuronGroup) 0,946 0,929 0,941 0,957 Linear (weights s) 0,674 0,700 0,679 0,696 Optimizer (3 classes) AttnPred (neuronGroup) 0,786 0,740 0,742 0,806 Linear (weights s) 0,280 0,290 0,268 0,287 Epoch (9 classes) AttnPred (neuronGroup) 0,324 0,319 0,313 0,312 Table 10: Hyper-parameters classification results using an adaption of AttnAE II named AttnPred. The metric used to evaluate the performance is F1 score. The results presented are the average of five different splits between train and test (scores obtained from the test split with checkpoints group balanced). The variance in all fields is less than 0.001. Using attention-based neural networks it is possible to identify the different properties that compose each of the models only from the information of the weights as shown in Table 9. In all cases the classes are equally balanced. In all the hyper-parameters under study, scores higher than random guessing were obtained. We found the model with N= 4 and heads = 4 ( 190k parameters) to be the best fit to obtain the best performance found on the classification task. In all hyper-parameters and all datasets the non-linear attention-based model has obtained better results than the linear baseline. However, we consider that the difference in performance is not too large considering the computational complexity realised by the DNN-based model. For epoch prediction the result obtained by the log-linear prediction was entered for classification since as a regressor the results were equal to random guessing. In the case of the AttnPred the results presented were trained as a regressor to obtain better results than as a classification. For each of the predictors we have evaluated their performance by plotting the recall score for the predictions of the initialisation method and the activation and optimiser functions for each of the groups composed by different epochs. The figure 21 shows the 55
recall obtained for each of the classes during training. For the epoch prediction evaluation, the confusion matrix including the distribution of predictions for each group of models coming from the same epoch is shown in Figure 22. Figure 20: Recall curves for hyper-parameter classification on the Google dataset. Right column are results for the log-linear predictor. Left column are results for the attention based architecture predictor, AttnPred. The Google dataset [1] is composed of models initialised by five different methodologies using the same seed. For this reason, the group corresponding to epoch zero has only five possible configurations of weights. From this point on, depending on the configuration used for the hyperparameters of each model, it progresses in a different manner. Each of these configurations share a set of rules that correlate the weights with its properties 56
of the models that are configured the same way. Figure 20 shows how the non-linear model is able to perfectly classify the first group. However, the linear regression-based model does not manage to group them correctly. In the two predictors the two worst predicted initialisation functions coincide, being truncatedNormal and random Normal. This may make sense given that these distributions are very similar to each other. In general, the linear predictor performs worse in the epochs 20, 40, 60 where there is more variety of models that are at stages further away from convergence and some that are already converged (see Figure 21). The non-linear predictor predict better the classes in all hyper-parameter classification tasks. Figure 21: Accuracy distribution at different epoch groups of the google dataset [1] In the case of epoch prediction, the regression-based implementation for the AttnPred model was chosen instead of classification because of its better performance. In general the results compared to the other hyper-parameters classification tasks are much worse. It is more difficult to predict in which epoch the model comes from just by looking at the weights space. The amount of update weights may not be proportional to the convergence of such models. One model may obtain the same solution trained for only two epochs while another model may use a longer path requiring more updates to arrive at the same point. Conversely, a model can be trained indefinitely and never converge to a good solution. At the distribution of predictions (see figure 22) for each epoch it can be seen that the model tends to predict values close to the correct ones. The variance of the predictions is higher in the central range, coinciding in the region with more diversity of converged and non-converged models. 57