xSTAE: Explaining Classifier Decisions through EEG Signal Style Transfer Autoencoding Natalia Koliou1,*,†,Panagiotis Zazos2,†,Christoforos Romesis1,†,Cristian Bosch3, Stasinos Konstantopoulos1and Panagiotis Trakadas2 1Institute of Informatics and Telecommunications, NCSR ‘Demokritos’, Ag. Paraskevi, Greece 2Four Dot Infinity, Athens, Greece 3CeADAR, University College Dublin, Dublin, Irelard Abstract Style transfer methods are a powerful visualization tool that can be used to generate counterfactual explanations, plausible alternatives to the original input that leads to a different classification. In this paper we present xSTAE, a system that restyles a misclassified example into the correct class, in order to help the expert understand what patterns the classifier was looking for to assign the correct class, and failed to see in the instance. The system is based on an Autoencoder trained on a loss function that balances between identity loss (similarity with the original instance) and a classification loss derived from a pre-trained classifier, allowing xSTAE to remain completely agnostic with respect to the internals of the classifier it interprets. We present promising experimental results on sleep-stage classification decisions over EEG data, which validate the core of the idea and show future research directions. Keywords EEG, Style Transfer, Autoencoders, Sleep Stage Classification, Generative AI, Explainable AI 1. Introduction Explainable AI (XAI) has gained significant attention in the last few years, mainly due to the wide application of deep learning models in high-stake domains, prominently including healthcare where understanding model decisions can directly impact patient needs, treatment and overall well-being. In practice, XAI helps users not only determine whether to trust individual predictions, but also compare different models and identify areas for improvement when a model performs poorly [1]. While most XAI research has focused on image and tabular data, time-series remain relatively under-explored. One possible reason regarding this gap is that, unlike images, the semantics of time-series cannot be easily visualized [ 2 ] and their interpretation requires combining domain expertise with an understanding of temporal dependencies. This gap has considerable impact in the healthcare domain, where bio-signals are widely used in clinical practice. Electroencephalography (EEG) signals for instance—which are the focus of this paper—are widely used in clinical practice to monitor brain activity and support the diagnosis of neurological conditions [ 3 ], sleep and mental disorders [4,5], and cognitive impairments [6]. EXPLIMED 2025 - Second Workshop on Explainable Artificial Intelligence for the Medical Domain - 25-30 October 2025, Bologna, Italy *Corresponding author. †These authors contributed equally. $nataliak[email protected] (N. Koliou);
[email protected] (P. Zazos); c[email protected] (C. Romesis); cristian.bosc[email protected] (C. Bosch); konstan[email protected] (S. Konstantopoulos);
[email protected] (P. Trakadas) https://www.linkedin.com/in/natalia-koliou-b37b01197/ (N. Koliou); https://www.linkedin.com/in/panagiotis-zazos-ba1a02188/ (P. Zazos); https://www.linkedin.com/in/cristian-bosch/ (C. Bosch) 0009-0004-3920-9992 (N. Koliou); 0009-0004-2127-9714 (P. Zazos); 0009-0001-6485-5548 (C. Romesis); 0000-0002-2962-4226 (C. Bosch); 0000-0002-2586-1726 (S. Konstantopoulos); 0000-0002-5146-5954 (P. Trakadas) ©2025 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0
In this paper we propose a method for generating instance-based interpretations of sleep-stage classification decisions over EEG data. Our method is based on training a generative AI model to interpret a pre-trained classifier by giving counterfactual examples of what a misclassified instance should have looked like to be correctly classified. Observing such examples can help the operator observe the patterns that lead to misclassifications and create a basis for a more focused collection and annotation of further training data. In the remainder of this paper we first discuss related work (Section 2) and formally state the problem (Section 3). We then present our method (Section 4) and proceed to provide and discuss experimental results (Section 5) and conclude (Section 6). 2. Related Work While deep learning is used for a variety of tasks, classification has been a primary focus within the XAI research community due to its broad applicability and conceptual simplicity. Relevant research has primarily focused on image and tabular data, where model interpretability is more intuitive and visual explanations are easier to generate [7,8,9]. As attention has only recently turned to the, more challenging, problem of explaining timeseries classifiers, the literature on this topic remains relatively limited [ 10 , 2 ]. In many cases, the need to interpret decisions drives the development of the classifier itself, with LAXCAT and PatchX being characteristic examples. LAXCAT [ 11 ] is a deep neural network architecture that simultaneously identifies the time intervals and the variables of a multi-variate time-series that contribute to classification decisions. LAXCAT uses convolutional layers to extract features from the time series, and two attention modules to identify which variables and time intervals are most important for the classification. PatchX [ 12 ] divides time series into smaller patches and performs fine-grained classification on each patch using deep neural networks. These patch-level results are then combined with a traditional classifier to generate the final prediction, and explanations are provided by interpreting the importance of each patch in the overall decision. A different line of research couples the training of the classifier with the training of its interpreter. Gee et al. [ 13 ] propose a prototype-based approach for explaining deep time-series classifiers by learning diverse and representative latent prototypes that highlight class-discriminative patterns across multiple modalities, including ECG, respiration, and audio waveforms. They achieve this by integrating a prototype learning mechanism directly into the training process, encouraging the model to associate input instances with learned prototypes that reflect meaningful and diverse latent representations for each class. XTF-CNN [ 14 ] is a dual-channel convolutional neural network that learns representations of microseismic waveforms from both the time and frequency domains to improve classification of rock fracturing events. XTF-CNN is integrated with EUG-CAM that generates fine-grained, gradient-based activation maps over input waveforms to illustrate which parts of the signal influence the model’s decisions. Finally, DeepVix [ 15 ] is a visual analytics system designed to explain Long Short-Term Memory (LSTM) networks applied to high-dimensional multivariate time series data. DeepVix provides interactive visualizations of the LSTM architecture, including node activations and gate weights across layers and time steps, enabling users to investigate intermediate computations and trace how different input variables contribute to the model’s predictions. Naturally, the lines of research above require a tight coupling between the classifier and its interpretation module, restricting the range of possible networks that can be used to those for which an appropriate interpretation module has been devised. On the other hand, timeXplain [ 16 ] is a post-hoc, model-agnostic explanation framework for time series classifiers based on SHAP. timeXplain uses domain-specific perturbation strategies organized by time, frequency, and statistical mappings. By observing the effect of perturbing the inputs to the results, timeXplain evaluates feature importance. timeXplain demonstrates that certain mappings (e.g., time slices with noise replacements) can produce explanations that are more faithful than some
model-specific methods. Focusing in the EEG domain, Apicella et al. [ 17 ] evaluated several established XAI methods to explain machine learning models trained on EEG data for emotion recognition, aiming to address the dataset shift problem, a common issue in Brain-Computer Interfaces where EEG signal characteristics vary across recording sessions causing models trained on one session to perform poorly on others. They applied these XAI techniques to identify which EEG signal components the models rely on and tested how consistently these important features appear within and across different recording sessions of the same subjects. Their results showed that many relevant features detected by XAI remain stable across sessions, indicating that these explanations can be used to improve the generalization and robustness of EEG-based classification systems. Zanola et al. [18] developed xEEGNet, a compact and fully interpretable neural network for EEG-based dementia classification that transforms a traditional “black box” model (ShallowNet) into a “white box” by progressively modifying its architecture to highlight explainable components. They achieve interpretability by designing the network to learn EEG band-specific filters and spatial topographies, which correspond to meaningful brain signal features clinicians can relate to dementia pathology. This approach not only reduces the number of parameters by over 200 times—helping to resist overfitting—but also allows direct inspection of the learned kernels and weights to provide clear, medically relevant explanations for the model’s decisions. Hussain et al. [ 19 ] developed machine learning models to classify human activities—resting, motor, and cognitive—using EEG spectral features collected from healthy individuals. They applied the model-agnostic explainability technique LIME to interpret which EEG features most influenced the classification decisions, providing clinically relevant insights into brain activity during these tasks. Their results showed strong classification performance and meaningful explanations that could support improved patient monitoring and rehabilitation. 3. Problem Statement According to the taxonomy introduced by Theissler et al. [ 2 ], counterfactual explanations are considered a promising instance-based approach for interpreting time-series classifiers. A counterfactual is a plausible alternative to the original input that leads to a different classification, while remaining as similar as possible to the original. By comparing the original time series with its counterfactual, one can infer which parts of the signal had the greatest impact on the model’s decision. Let 𝐷 = {𝑥(𝑖)}𝑁 𝑖=1 be an EEG dataset of 𝑁 sequences, where each sequence 𝑥(𝑖)∈ 𝒳 ⊆ R𝑑 is represented as a 𝑑 -dimensional vector, and 𝐶 : 𝒳 → 𝒴 a trained classifier that maps each input 𝑥∈ 𝒳 to one of 𝑛 discrete class labels in 𝒴 = { 1 , 2 , . . . , 𝑛} . Given an input 𝑥 , the classifier outputs a label 𝑦 = 𝐶 ( 𝑥 ). Our goal is to provide an explanation for why the classifier assigned that specific label to the input. This becomes particularly insightful in cases of misclassification, where understanding what made the classifier predict an incorrect label can reveal class-specific patterns that were strong enough to override the correct label. Identifying the distinguishing characteristics that the model associates with the incorrect class provides a means to explain its decision-making process. Formally, we define the problem as follows: Given that 𝐶 is an imperfect classifier, we focus on the subset of inputs 𝑥𝑚∈ 𝒳 for which the predicted label 𝐶 ( 𝑥𝑚 )differs from the true label 𝑦* ( 𝑥𝑚 ). For each such misclassified instance, we aim to identify the minimal modification 𝑥′ 𝑚 of the input such that the classifier’s prediction aligns with the ground truth, i.e., 𝐶 ( 𝑥′ 𝑚 ) = 𝑦* ( 𝑥𝑚 ), while ensuring that the modification is as small as possible according to a chosen distance metric 𝑑(·,·). This can be expressed as the optimization problem: 𝑥′ 𝑚= arg min 𝑧∈𝒳 𝑑(𝑥𝑚, 𝑧)so that 𝐶(𝑧) = 𝑦*(𝑥𝑚)(1)
Here, 𝑥′ 𝑚 serves as a counterfactual that reveals how the original input 𝑥𝑚 would need to change to be correctly classified. By analyzing the difference Δ 𝑥𝑚 = 𝑥′ 𝑚−𝑥𝑚 , we can identify dominant patterns associated with class 𝐶 ( 𝑥𝑚 )and gain insight into the model’s decision boundaries. 4. Proposed Methodology To generate meaningful counterfactuals, we propose a generative framework based on a set of class-conditional autoencoders, where each autoencoder is trained to reconstruct inputs from any class, while restyling them toward a specific target class 𝑦tgt ∈ 𝒴. For every target class, we train a separate autoencoder 𝐸tgt : 𝒳 → 𝒳tgt . Given an input 𝑥 and a target label 𝑦tgt = 𝐶 ( 𝑥 ), the corresponding autoencoder 𝐸tgt generates a counterfactual 𝑥′ = 𝐸tgt ( 𝑥 )that closely resembles 𝑥 , but is modified just enough for 𝐶 to classify it as 𝑦tgt , i.e., 𝐶 ( 𝑥′ ) = 𝑦tgt . By comparing the original input 𝑥 with its counterfactual 𝑥′ , one can identify the patterns in 𝑥 that were responsible for the classifier’s original decision, and thus detect the characteristics of class 𝐶(𝑥)that were most prominent. Training each of these autoencoders requires incorporating two forms of feedback: 1. The first evaluates how well the generated output approximates the original input. This can be quantified using a similarity function 𝑑 ( 𝑥, 𝑥′ ), where 𝑥∈ 𝒳 is the original input and 𝑥′ is the reconstructed output. The goal is to keep 𝑑 ( 𝑥, 𝑥′ )sufficiently small to ensure the generated output remains close to the original. 2. The second evaluates how well the output aligns with the desired target class 𝑦tgt ∈ 𝒴 . This feedback is provided by the classifier 𝐶 that we aim to explain. During training, the generated output 𝑥′ is passed through 𝐶 , and its prediction 𝐶 ( 𝑥′ )is compared against the target label 𝑦tgt. This dual feedback guides the autoencoders towards producing counterfactuals that are both similar to the input and representative of the target class. 4.1. EEG Data Our methodology is specifically designed for time-series EEG data. Let 𝑒∈R𝑛𝑠×𝑛𝑐 denote an EEG segment in the time domain (or epoch), where 𝑛𝑠 is the number of time samples and 𝑛𝑐 is the number of recording channels. Each epoch corresponds to a fixed-duration window of 𝑤 seconds of EEG signal recording. To reduce the complexity of the input data, we transform raw EEG time-series from the time domain to the frequency domain. Given an input signal 𝑒∈R𝑛𝑠×𝑛𝑐 , we first apply the Fast Fourier Transform (FFT) to obtain its spectral representation. To retain only domain-relevant information, we filter out frequencies outside a predefined range of interest (e.g., those not associated with meaningful brain activity). The remaining frequency band is then divided into 𝑛𝑠′ non-overlapping segments, each represented by three features: the midpoint frequency, the phase at that frequency, and the average amplitude across the segment. This results in a matrix 𝑒𝑓∈R𝑛𝑠′×𝑛𝑓, where each row represents a frequency-region waveform using 𝑛𝑓features. 4.2. Classifier The classifier is a two-stage convolutional neural network designed to process EEG data in the time domain. The architecture is based on the one proposed by Youness [ 20 ] and Esparza-Iaizzo et al. [ 21 ]. This architecture captures short-term temporal dependencies by analyzing sequences of consecutive epochs. To predict the label for some epoch 𝑒𝑓 𝑡∼𝑒𝑡 (for simplicity), the model considers a sequence of 𝑘epochs, including the current epoch and the previous 𝑘−1epochs:
Table 1 Summary of the classifier architecture. Component Description Base Feature Extractor Processes each epoch independently: 2×Conv1D(16, k=5) + ReLU + MaxPool(s=2) + Dropout 2×Conv1D(64, k=5) + ReLU + MaxPool + Dropout 2×Conv1D(128, k=3) + ReLU AdaptiveMaxPool + Flatten + Dense(128) + ReLU + Dropout Output: 128-dim feature vector Temporal Sequence Model Takes 𝑘stacked epoch features (sequence length = 𝑘): Conv1D(64, k=3) + ReLU + Dropout Conv1D(64, k=3) + ReLU Conv1D(n, k=3) + AdaptiveMaxPool Output: n-length logits vector 𝑋𝑓 𝑡= [𝑒𝑡−𝑘+1, 𝑒𝑡−𝑘+2, . . . , 𝑒𝑡−1, 𝑒𝑡]∈R𝑘×𝑛𝑠′×𝑛𝑓→^𝑦𝑡∈ {1,2, . . . , 𝑛}(2) The architecture consists of two primary components as illustrated in Table 1. 4.3. Autoencoder The autoencoder is a hybrid architecture combining self-attention mechanisms and convolutional operations. The input to 𝐸is a sequence of 𝑛𝑠′vectors, each of dimensionality 𝑛𝑓+ 1: 𝑋𝑓 𝑡=𝑒𝑡⊕𝑝𝑜𝑠 ∈R𝑛𝑠′×(𝑛𝑓+1) →𝑋𝑓 𝑡,𝑡𝑔𝑡 (3) The first 𝑛𝑓 dimensions correspond to the extracted features, while the last dimension is a positional encoding added to inject a sense of temporal order. Since Transformers inherently treat their input as a set of unordered tokens, this positional encoding, implemented as a simple increasing counter (e.g., 1, 2, 3, . . . , 𝑛𝑠′ ), allows the model to capture temporal dependencies across the samples in the epoch. A detailed overview of the autoencoder architecture is provided in Table 2. 5. Experiments We apply our methodology to sleep stage classification using EEG data from the Bitbrain Open Access Sleep (BOAS) dataset [ 22 ]. This dataset consists of 128 full-night recordings collected from healthy volunteers wearing a two-channel EEG headband ( 𝑛𝑐 = 2). The EEG signals were sampled at 256 Hz and segmented into non-overlapping 30-second epochs, each containing 7,680 samples per channel ( 𝑛𝑠 = 7680). Every epoch was independently scored by three certified sleep experts following the American Academy of Sleep Medicine (AASM) guidelines [ 23 ]. These guidelines define five standard sleep stages: Wake (W), N1 (light sleep), N2 (intermediate sleep), N3 (deep sleep), and REM (rapid eye movement sleep). Since our focus is on sleep stage classification, we exclude the Wake class and restrict our task to the four sleep-related stages: N1, N2, N3, and REM ( 𝑛 = 4). To address typical inter-scorer variability ( ∼ 85% agreement [ 24 ][ 25 ]), a fourth expert reviewed the annotations and produced a consensus label for each epoch. These consensus sleep-stage labels were then aligned with the EEG segments to provide a reliable ground truth for each 30-second window. 5.1. Setup According to our methodology, the first step involves transforming the Bitbrain EEG data from the time domain to the frequency domain. To achieve this, we pre-process each epoch
Table 2 Summary of the autoencoder architecture. Component Description Attention Mechanism Learns meaningful temporal dependencies 3×MultiHeadAttention(num_heads = 1) Output: Refined input vector 𝑥𝐴∈R𝑛𝑠×𝑛𝑓 Encoder Downsamples refined input into a latent representation ConvINReLU(16, k=7, s=1, p=3) ConvINReLU(32, k=4, s=2, p=1) ConvINReLU(64, k=4, s=2, p=1) ConvINReLU(128, k=3, s=2, p=1) ConvINReLU(256, k=3, s=1, p=1) Output: Latent vector 𝑧∈R𝑛𝑧×256 Bottleneck Applies deep transformations in latent space 𝑁×ConvINReLU(256, 256, k=3, s=1, p=1) Output: Transformed latent vector 𝑧∈R𝑛𝑧×256 Decoder Upsamples the latent representation to reconstruct the original input DeconvINReLU(256, 128, k=3, s=1, p=1) DeconvINReLU(128, 64, k=3, s=2, p=1) DeconvINReLU(64, 32, k=4, s=2, p=1) DeconvINReLU(32, 16, k=4, s=2, p=1) ConvTranspose1D(16, F, k=7, s=1, p=3) Output: Reconstructed vector ^𝑥∈R𝑛𝑠×𝑛𝑓 (per channel) using the pipeline shown in Figure 1. The raw time-domain signal 𝑒∈R𝑛𝑠×𝑛𝑐 (Figure 1a) is first converted to the frequency domain using a Fast Fourier Transform (FFT) (Figure 1b). We then filter out all frequency components outside the 0.4-30 Hz range (Figure 1c), which includes the primary EEG waveforms relevant to sleep staging: Delta (0.5-4 Hz), Theta (4-8 Hz), Alpha (8-13 Hz), and Beta (13-30 Hz) [26]. Next, we split the retained frequency range into 𝑛𝑠′ equal, non-overlapping segments. From each segment, we extract a representative waveform defined by its midpoint frequency and phase, along with the average amplitude across the segment (Figure 1d). This produces a compact and structured representation of the original time-series in the frequency domain: 𝑒𝑓∈R𝑛𝑠′×𝑛𝑓 , where 𝑛𝑓 = 6, corresponding to three features (frequency, phase, amplitude) for each of the two EEG channels. In our experiments, we set 𝑛𝑠′ = 300 to balance expressiveness with computational efficiency. We split the dataset subject-wise into training, validation, and test sets to ensure that all epochs from a given participant remained within a single split. This way, we prevent subject leakage and preserve the validity of evaluation results. Approximately 80% of the nights (80 participants) were assigned to the training set, 10% (12 participants) to the validation set, and 20% (25 participants) to the test set. To standardize the data, we apply z-score normalization using the mean and standard deviation computed from the training set. These statistics are then used to normalize the validation and test sets accordingly. ˜𝑥=𝑥−𝜇train 𝜎train (4) 5.1.1. Classifier The sleep stage classifier described in Section 4.2 is the model we aim to explain. It takes as input a sequence of 𝑘 = 5 consecutive EEG epochs in the frequency domain and produces a prediction for the final epoch 𝑒5(Equation 5). 𝑋𝑓 𝑡= [𝑒1, 𝑒2, 𝑒3, 𝑒4, 𝑒5]∈R5×300×6→^𝑦𝑡∈ {1,2,3,4}(5)
Figure 1: EEG signal preprocessing: (a) Time-domain EEG signal 𝑒 ,(b) Full-spectrum FFT, (c) Band-limited FFT (0.4-30 Hz), (d) Compressed frequency-domain spectrum 𝑒𝑓. A critical part of training this model is choosing a loss function that matches the data distribution and learning objectives. The Bitbrain dataset has a highly imbalanced class distribution: N2 accounts for around 50% of the epochs, while N1 and N3 are less common (about 5% and 20%, respectively), and REM makes up the remaining 25%. This skewed class distribution can bias the model toward over-predicting the majority class, resulting in poor performance on minority classes. To address this pronounced class imbalance, we adopt the sparse categorical focal loss, which modifies the standard cross-entropy loss to emphasize learning
Table 3 Classifier’s Hyperparameter Search Space and Optimal Values. Parameter Search Space Optimal Value base_filters {16, 32, 64} 64 filter_1 Equal to base_filters 64 filter_2 base_filters ×4 256 filter_3 base_filters ×8 512 seq_mult {2, 4, 6} 6 seq_filter base_filters ×seq_mult 384 kernel_1 {3, 5, 7} 3 kernel_2 {3, 5, 7} 7 kernel_3 {3, 5, 7} 5 seq_kernel {3, 5} 5 dropout_conv (0.1, 0.5) 0.134 dropout_seq (0.1, 0.5) 0.205 dense_units (64, 256) with steps of 32 128 optimizer {Adam, AdamW} AdamW learning rate (lr) 5×10−5to 5×10−45.92 ×10−5 focal loss alpha (𝛼) (0.0, 1.0) 0.50 focal loss gamma (𝛾) (0.0, 5.0) 2.95 batch_size {16, 32, 64} 16 from hard-to-classify examples. The focal loss is defined as: ℒfocal =− 𝑛 ∑︁ 𝑐=1 𝛼𝑐(1 −𝑝𝑐)𝛾log(𝑝𝑐),(6) where 𝑝𝑐 is the predicted probability for the true class 𝑐 , 𝛼𝑐 is a weighting factor that balances the importance of each class, and 𝛾≥ 0is a focusing parameter that reduces the contribution of well-classified examples. By tuning 𝛼𝑐 and 𝛾 , the focal loss places greater importance on difficult or misclassified instances, enhancing performance on underrepresented classes. Although the loss is defined as a sum over all classes, in practice only the term corresponding to the true class contributes to the loss for each training sample. This is due to the one-hot encoding of the target labels: the correct class is represented with a value of 1, while all other classes are 0. As a result, all terms involving incorrect classes are multiplied by zero and do not affect the outcome. To identify the optimal configuration for the classifier, we conducted an automated hyperparameter search using the Optuna framework. The goal was to maximize the macro-F1 score on the validation set, which better reflects performance across all sleep stages. We performed 50 trials, with each trial running for up to 20 training epochs. Early stopping was applied if the validation macro-F1 did not improve for 5 consecutive epochs. The search space included a variety of architectural and training parameters, presented in the middle column of Table 3. These ranged from convolutional filter sizes and kernel widths to dropout rates, dense layer units, learning rate, and optimizer choice. The focal loss parameters 𝛼 and 𝛾 were also included to fine-tune the loss function’s sensitivity to class imbalance. The optimal hyperparameters identified by Optuna’s best trial are presented in the rightmost column of the same table. Figure 2reveals that the most influential hyperparameters were base_filters and dense_units . The focal loss parameters 𝛼and 𝛾also had great impact. During evaluation, our classifier achieved an overall accuracy of 0.87. Table 4provides a detailed per-class performance breakdown, including precision, recall, F1-score, and support. To benchmark our approach, we compare against the work of Esparza-Iaizzo et al. [ 21 ], who performed non-causal, single-channel sleep stage decoding directly on raw EEG signals, including the Wake stage. Comparing the confusion matrices in Figure 3, we note that Light Sleep (N1) remains the most challenging stage to classify, with our model achieving a recall of only 27%.
Figure 2: Optuna’s Hyperparameter Importance Chart. Table 4 Classification report including precision, recall, F1-score, and support per class on the test set. Class Precision Recall F1-score Support N1 0.65 0.27 0.38 1267 N2 0.94 0.91 0.92 17357 N3 0.56 0.76 0.64 725 REM 0.75 0.93 0.83 3870 Accuracy 0.87 23219 Macro avg 0.72 0.72 0.69 23219 Weighted avg 0.88 0.88 0.87 23219 However, by excluding Wake, our pipeline avoids masking N1 errors within dominant Wake bins, as seen in previous works. Instead, misclassified N1 epochs are primarily confused with N2 (52%) and REM (21%), reflecting the transitional nature of light sleep. Mid-stage non-REM sleep (N2) detection shows considerable improvement, with recall increasing from 85% in the baseline to 91% in our model. Performance on deep slow-wave sleep (N3) remains comparable, with both methods effectively capturing delta-band features. The largest boost is observed in REM sleep detection, where recall rises from 75% to 93%. This suggests that spectral representation combined with sequential modeling better isolates the characteristic EEG patterns of REM. Together, these results validate that converting raw headband signals into compact, twochannel spectral slices and modeling their temporal evolution over consecutive epochs yields a more discriminative representation for the predominant sleep stages (N2 and REM) without sacrificing accuracy in deep sleep. Light sleep (N1) remains inherently difficult due to its low prevalence (∼5%) and high intra-stage variability in healthy adults [27]. 5.1.2. Autoencoders To perform style transfer across the 𝑛 = 4 sleep stages, we train a separate autoencoder 𝐸tgt for each target class 𝑦tgt ∈ { 1 , 2 , 3 , 4 } . Each model learns to restyle input epochs from any class to
Figure 8: The original signal (top) and its decomposition into the four brainwave bands. This signal is labeled as N3 but was misclassified as N2. approach on open data and publish the complete experimental setup as open source.1 Although the empirical validation was promising, there are some further steps before the system can be validated in trials with experts. Specifically, we plan to further explore different ways to define the identity loss. For one, visual observation has shown that mean square error biases the model towards making small changes everywhere, which makes it harder for the expert to identify what changes have been effected. Before trials, we need to define (and validate the convergence and low classification loss) of alternative identity loss definitions, e.g. preferring bigger local changes or preferring changes that do not affect all four brainwave bands. The trials can then be used to establish with notion of ‘identity’ makes it easier to spot the changes affected in order to achieve the intended re-labelling. A further, more ambitious, step is to link the insights extracted from interpreting misclassifications to possible actions for alleviating them. Since xSTAE is specifically designed to be model-agnostic and can be matched to any pre-trained classifier, such actions also need to operate at the same level of abstraction to maintain the generality of the system. In other words, the outcome of an expert’s understanding of the classifier’s pain-points should operate at the level of re-balancing data or of post-hoc establishing the confidence of specific classification decisions; as opposed to recommending hyper-parameter or architectural changes or similar 1 The data is the BOAS dataset [ 28 ] and the experimental setup is available at https://doi.org/10.5281/zenodo. 17085776.
Figure 9: The signal from Fig. 8after restyling into N3; both the full signal (top) and its decomposition into the four brainwave bands are shown. classifier-specific actions. Acknowledgments This research was co-funded by the European Union under GA no. 101135782 (MANOLO project). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or CNECT. Neither the European Union nor CNECT can be held responsible for them. Declaration on Generative AI As the subject of this article is generative AI, the example outputs in Figures 7and 9are generated by AI. No generative AI was used to prepare any of the remaining content, either textual or graphical. References [1] M. T. Ribeiro, S. Singh, C. Guestrin, “Why should I trust you?”: Explaining the predictions of any classifier, arXiv:1602.04938 [cs.LG], 2016. URL: https://arxiv.org/abs/1602.04938.
[2] A. Theissler, F. Spinnato, U. Schlegel, R. Guidotti, Explainable AI for time series classification: A review, taxonomy and research directions, IEEE Access PP (2022) 1–1. doi:10.1109/ACCESS.2022.3207765. [3] M. Allahbakhshi, A. Sadri, S. O. Shahdi, Diagnosis of Parkinson’s disease using EEG signals and machine learning techniques: A comprehensive study, arXiv:2405.00741 [eess.SP], 2024. URL: https://arxiv.org/abs/2405.00741. [4] H.-N. Jo, Y.-S. Kweon, S.-H. Lee, EEG spectral analysis in gray zone between healthy and insomnia, 2024. doi:10.48550/arXiv.2411.09875. [5] M. Jafari, D. Sadeghi, A. Shoeibi, H. Alinejad-Rokny, A. Beheshti, D. L. García, Z. Chen, U. R. Acharya, J. M. Gorriz, Empowering precision medicine: AI-driven schizophrenia diagnosis via EEG signals: A comprehensive review from 2002-2023, arXiv:2309.12202 [eess.SP], 2023. URL: https://arxiv.org/abs/2309.12202. [6] A. Ortiz, F. J. Martínez-Murcia, M. A. Formoso, J. L. Luque, A. Sánchez, Dyslexia Detection from EEG Signals Using SSA Component Correlation and Convolutional Neural Networks, Springer International Publishing, 2020, pp. 655–664. URL: http://dx.doi.org/10.1007/ 978-3-030-61705-9_54. doi:10.1007/978-3-030-61705-9_54. [7] M. M. Karim, Y. Li, R. Qin, Towards explainable artificial intelligence (XAI) for early anticipation of traffic accidents, arXiv:2108.00273 [cs.CV], 2022. URL: https://arxiv.org/ abs/2108.00273. [8] E. Kadar, G. Gilboa, DXAI: Explaining classification by image decomposition, arXiv:2401.00320 [cs.CV], 2024. URL: https://arxiv.org/abs/2401.00320. [9] T. Vermeire, D. Martens, Explainable image classification with evidence counterfactual, arXiv:2004.07511 [cs.LG], 2020. URL: https://arxiv.org/abs/2004.07511. [10] T. Rojat, R. Puget, D. Filliat, J. D. Ser, R. Gelin, N. Díaz-Rodríguez, Explainable artificial intelligence (XAI) on timeseries data: A survey, arXiv:2104.00950 [cs.LG], 2021. URL: https://arxiv.org/abs/2104.00950. [11] T.-Y. Hsieh, S. Wang, Y. Sun, V. Honavar, Explainable multivariate time series classification: A deep neural network which learns to attend to important variables as well as informative time intervals, arXiv, 2020. URL: https://arxiv.org/abs/2011.11631.arXiv:2011.11631. [12] D. Mercier, A. Dengel, S. Ahmed, Patchx: Explaining deep models by intelligible pattern patches for time-series classification, in: 2021 International Joint Conference on Neural Networks (IJCNN), IEEE, 2021, pp. 1–8. URL: http://dx.doi.org/10.1109/IJCNN52387. 2021.9533293. doi:10.1109/ijcnn52387.2021.9533293. [13] A. H. Gee, D. Garcia-Olano, J. Ghosh, D. Paydarfar, Explaining deep classification of time-series data with learned prototypes, arXiv:1904.08935 [cs.LG], 2019. URL: https: //arxiv.org/abs/1904.08935. [14] X. Bi, Z. Chao, Y. He, X. Zhao, Y. Sun, Y. Ma, Explainable time-frequency convolutional neural network for microseismic waveform classification, Information Sciences 546 (2021) 883–896. doi:10.1016/j.ins.2020.08.109. [15] T. Dang, H. Van, H. Nguyen, P. Vung, R. Hewett, DeepVix: Explaining long short-term memory network with high dimensional time series data, in: Proceedings of the 11th International Conference on Advances in Information Technology (IAIT ’20), 2020, pp. 1–10. doi:10.1145/3406601.3406643. [16] F. Mujkanovic, V. Doskoč, M. Schirneck, P. Schäfer, T. Friedrich, timexplain – a framework for explaining the predictions of time series classifiers, arXiv:2007.07606 [cs.LG], 2023. URL: https://arxiv.org/abs/2007.07606. [17] A. Apicella, F. Isgrò, A. Pollastro, R. Prevete, Toward the application of XAI methods in EEG-based systems, arXiv:2210.06554 [cs.LG], 2024. URL: https://arxiv.org/abs/2210. 06554. [18] A. Zanola, L. F. Tshimanga, F. D. Pup, M. Baiesi, M. Atzori, xEEGNet: Towards explainable AI in EEG dementia classification, arXiv:2504.21457 [cs.LG], 2025. URL: https://arxiv.org/ abs/2504.21457.
[19] I. Hussain, R. Jany, R. Boyer, A. Azad, S. A. Alyami, S. J. Park, M. M. Hasan, M. A. Hossain, An explainable EEG-based human activity recognition model using machinelearning approach and lime, Sensors 23 (2023). doi:10.3390/s23177452. [20] M. Youness, CVxTz/EEG_classification: v1.0, Zenodo, 2020. doi: 10.5281/zenodo. 4060151. [21] M. Esparza-Iaizzo, M. Sierra-Torralba, J. G. Klinzing, J. Minguez, L. Montesano, E. LópezLarraz, Automatic sleep scoring for real-time monitoring and stimulation in individuals with and without sleep apnea, bioRxiv, 2024. doi:10.1101/2024.06.12.597764. [22] E. López-Larraz, M. Sierra-Torralba, S. Clemente, G. Fierro, D. Oriol, J. Minguez, L. Montesano, J. G. Klinzing, Bitbrain open access sleep dataset, OpenNeuro, 2024. doi:10.18112/openneuro.ds005555.v1.0.0. [23] R. B. Berry, R. Brooks, C. Gamaldo, S. M. Harding, R. M. Lloyd, S. F. Quan, M. T. Troester, B. V. Vaughn, AASM Scoring Manual Updates for 2017 (Version 2.4), Journal of Clinical Sleep Medicine 13 (2017) 665–666. doi:10.5664/jcsm.6576. [24] H. Danker-Hopfe, D. Kunz, G. Gruber, G. Klösch, J. L. Lorenzo, S.-L. Himanen, et al., Interrater reliability between scorers from eight european sleep laboratories in subjects with different sleep disorders, Journal of Sleep Research 13 (2004) 63–69. doi: 10.1046/j. 1365-2869.2003.00375.x. [25] R. S. Rosenberg, S. Van Hout, The American Academy of Sleep Medicine inter-scorer reliability program: Sleep stage scoring, Journal of Clinical Sleep Medicine 9 (2013) 81–87. doi:10.5664/jcsm.2350. [26] L. Cao, Fourier-based spectral analysis of eeg signals for sleep stage classification, Theoretical and Natural Science 109 (2025) 130–140. doi:10.54254/2753-8818/2025.GL24090. [27] A. Patel, V. Reddy, K. Shumway, et al., Physiology, Sleep Stages, StatPearls Publishing, Treasure Island (FL), 2024. URL: https://www.ncbi.nlm.nih.gov/books/NBK526132/, [Updated 2024 Jan 26]. [28] E. López-Larraz, M. Sierra-Torralba, S. Clemente, G. Fierro, D. Oriol, J. Minguez, L. Montesano, J. G. Klinzing, The Bitbrain open access sleep (BOAS) dataset, OpenNeuro, 2025. doi:10.18112/openneuro.ds005555.v1.1.1.