scieee AI-readable full text Open interactive document viewer

Novel lossy compression method of noisy time series data with anomalies: Application to partial discharge monitoring in overhead power lines

Klein, Lukáš

Abstract

In overhead power transmission lines, particularly in regions like natural parks where establishing a safe zone is difficult, the adoption of cross-linked polyethylene insulated covered conductors (CCs) helps prevent outages due to vegetation contact. However, these CCs are susceptible to partial discharge (PD) activity, which can degrade insulation and lead to system failures. Detecting and analyzing PD are essential for maintaining power system reliability and safety. A key challenge in PD monitoring is transmitting the large volumes of PD signal data over unreliable 2G networks, as existing compression methods either compromise on data integrity or are ineffective. This paper introduces a novel lossy compression technique utilizing an autoencoder with skip connections and correction data to address this issue. Unlike previous algorithms that struggle with noisy time series data and fail to preserve crucial anomaly information, our method reconstructs the signal without anomalies, which are subsequently restored using correction data. Achieving a compression factor of about 25 (reducing data to 4.1% of its original size), this approach maintains essential PD signal features for analysis. The effectiveness of our method is validated by three classification algorithms, showing promise for future fault detection, diagnosis, and memory space reduction. This innovative compression solution marks a significant advancement in PD data processing, offering a balanced trade-off between compression efficiency and data fidelity, and paving the way for enhanced remote monitoring in power transmission systems.

Full text

Engineering Applications of Artificial Intelligence 133 (2024) 108267 Available online 16 March 2024 0952-1976/© 2024 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/bync/4.0/). Contents lists available at ScienceDirect Engineering Applications of Artificial Intelligence journal homepage: www.elsevier.com/locate/engappai Research paper Novel lossy compression method of noisy time series data with anomalies: Application to partial discharge monitoring in overhead power lines Lukáš Kleina,b,∗, Jiří Dvorskýb,c, David Seidla,b, Lukáš Prokopa aENET centre - CEET, VSB - Technical University of Ostrava, 17. listopadu, Ostrava, 708 00, Czech Republic bDepartment of Computer Science, VSB - Technical University of Ostrava, 17. listopadu, Ostrava, 708 00, Czech Republic cIT4Innovations National Supercomputing Center, Studentská 6231/1B, Ostrava, 708 00, Czech Republic ARTICLE INFO Keywords: Data compression Partial discharge Autoencoder Deep learning Anomalies ABSTRACT In overhead power transmission lines, particularly in regions like natural parks where establishing a safe zone is difficult, the adoption of cross-linked polyethylene insulated covered conductors (CCs) helps prevent outages due to vegetation contact. However, these CCs are susceptible to partial discharge (PD) activity, which can degrade insulation and lead to system failures. Detecting and analyzing PD are essential for maintaining power system reliability and safety. A key challenge in PD monitoring is transmitting the large volumes of PD signal data over unreliable 2G networks, as existing compression methods either compromise on data integrity or are ineffective. This paper introduces a novel lossy compression technique utilizing an autoencoder with skip connections and correction data to address this issue. Unlike previous algorithms that struggle with noisy time series data and fail to preserve crucial anomaly information, our method reconstructs the signal without anomalies, which are subsequently restored using correction data. Achieving a compression factor of about 25 (reducing data to 4.1% of its original size), this approach maintains essential PD signal features for analysis. The effectiveness of our method is validated by three classification algorithms, showing promise for future fault detection, diagnosis, and memory space reduction. This innovative compression solution marks a significant advancement in PD data processing, offering a balanced trade-off between compression efficiency and data fidelity, and paving the way for enhanced remote monitoring in power transmission systems. 1. Introduction Partial discharge (PD), a localized dielectric breakdown within a segment of electrical insulation under high voltage stress, signals insulation system degradation (Bartnikas,2002). While traditionally associated with exceeding electric field thresholds, PD’s initiation is multifaceted, encompassing factors such as insulation material imperfections, environmental stresses, aging, and contamination. This broader understanding underscores PD’s role as a crucial indicator of insulation health, essential for developing advanced diagnostic and mitigation strategies to ensure power system reliability (Bartnikas, 2002;Stone,2005;Orellana et al.,2023). Therefore, PD detection and analysis are important in ensuring power system reliability and safety. PD can be detected by various methods, such as acoustic, optical, electrical, or electromagnetic measurements (Lu et al.,2020). Among these methods, electromagnetic measurements have an advantage because they are non-invasive, sensitive, and easy to implement (Boggs,1982;Kaziz et al.,2023). One of the applications of PD monitoring using electromagnetic measurements is in covered conductors (CCs), which are overhead ∗Corresponding author at: ENET centre - CEET, VSB - Technical University of Ostrava, 17. listopadu, Ostrava, 708 00, Czech Republic. E-mail address: [email protected] (L. Klein). power transmission lines with cross-linked polyethylene (XLPE) insulation (Voldhaug and Robertson,1995). CCs are used in areas where a safe zone cannot be established, such as natural parks, to prevent outages caused by vegetation contact (Leskinen,2004). However, CCs are susceptible to PD activity, especially when contact is prolonged (e.g., by a fallen tree) (Hamacek,2012). Rather than employing the galvanic contact method for signal capture (Misak and Pokorny,2015; Misak et al.,2016), PD in CCs generates an electromagnetic signal. This signal emanates from the PD area and can be intercepted by an antenna. Characterized by their transient and non-stationary nature, these signals possess exceedingly brief durations. In conditions of minimal ambient noise, the PD-induced signal manifests as distinctive peaks within the signals recorded via ADC in the time domain (Kabot et al., 2020). Such signals propagate along the conductor and subsequently radiate into the adjacent space (Boggs,1982;Misak et al.,2017). These signals can be collected by antennas placed near the CCs and analyzed to identify the PD sources and their characteristics (Fulnecek and Misak,2021;Martinovic and Fulnecek,2021). https://doi.org/10.1016/j.engappai.2024.108267 Received 31 October 2023; Received in revised form 13 February 2024; Accepted 11 March 2024 Engineering Applications of Artificial Intelligence 133 (2024) 108267 2 L. Klein et al. However, a challenge in the monitoring of PD using antennas is the large size of the PD signals, making it difficult to transmit them over unreliable 2G networks for further analysis (Martinovic and Fulnecek, 2021). The authors of an article (Klein et al.,2023b) attempted to address the issue by employing classification through edge computing, i.e., at the location of the detector. However, for more effective training and precise classification, it is still necessary to transmit the data to a remote server. The PD signals collected by antennas are sampled at high frequencies (e.g. 500 MHz) in order to capture PDs, resulting in large data files (e.g. 800 kB per sample) (Misak et al.,2017). Transmitting such large files over 2G networks can incur a high cost and time, as well as data loss and corruption. Therefore, there is a need for compression methods that can reduce the size of PD signals while preserving their information and features. Existing compression methods for PD signals are inadequate or distortive. Some methods use lossless compression techniques, such as the Huffman coding (Huffman,1952;Knuth,1985) or the Lempel–Ziv– Welch algorithm (Kinsner and Greenfield,1991), which can preserve the original signal without any distortion. However, these methods have low compression factor (for example from 2 to 4), which are not sufficient to significantly reduce transmission cost and time. Other methods use lossy compression techniques, such as wavelet transform (Chaovalit et al.,2011) or principal component analysis (Asahi et al.,2021), which can achieve higher compression factor (for example from 10 to 100), but at the expense of distorting the signal and losing some information and features. Moreover, most of these methods fail to retain the anomalies that are essential for PD analysis. Anomalies are abnormal or irregular patterns in the signal that indicate the presence and nature of PD sources (Fulnecek and Misak,2021). They are usually characterized by high-amplitude peaks or impulses that deviate from the normal signal behavior. Most lossy compression methods tend to smooth out or remove these anomalies during the compression process, which can affect the accuracy and reliability of PD analysis. Moreover, there are few algorithms, such as SZ (Liang et al.,2023) and its variants (Zhao et al.,2021;Liang et al.,2018;Zhao et al.,2020) or ZPF (Lindstrom and Isenburg,2006a) and its variants (Hammerling et al.,2019;Diffenderfer et al.,2019;Hoang et al.,2020), that can compress noisy time series data, and none that we have found that can handle these specific data. Noisy time series data are data that contain random fluctuations or variations that obscure the underlying signal patterns or trends (Chiarot and Silvestri,2023). They are common in real-world applications, such as sensor networks, financial markets, or biomedical signals (Gupta and Gupta,2019). Noisy time series data pose challenges to compression methods because they require more bits to represent them accurately and make it harder to identify and extract relevant signal features (Liang et al.,2018). The PD signals collected by antennas are examples of noisy time series data, as they are affected by various sources of noise, such as environmental interference, measurement errors, or background noise. Given the limitations of existing methods for our problem, we introduce a novel compression method that takes advantage of deep learning to approximate the normal signal and detect anomalies. To tackle the challenge of compressing our data for efficient storage and transmission to our remote server, we developed a new lossy compression algorithm for PD signals in CCs based on a deep autoencoder with anomaly correction data. The deep autoencoder is a neural network that learns a low-dimensional representation of the input signal and reconstructs it with minimal distortion. The anomaly correction data are supplementary data that store outliers that are not captured by the autoencoder, as the autoencoder was specifically not trained to reconstruct anomalous signals. The proposed deep autoencoder also has skip connections that enable information to bypass some layers of the network and facilitate the training process. Our method achieved a compression factor of approximately 25 and preserved the information and characteristics of the PD signal, as assessed by three different classification algorithms. The compressed data can be utilized for fault detection and diagnosis, as well as saving significant memory space and transmission time from places with slow and unreliable connection. In this paper, we used a unique public dataset of PD signals collected by antennas from different stations under various conditions. The dataset contains more than 20k samples of 20 ms each, with a sampling rate of 500 MHz. The dataset also provides the ground truth labels of the PD types. We compared our compression method with other state-of-the-art methods, such as SZ, ZFP, bzip2, and zip, in terms of compression ratio, reconstruction error, and preservation of PD features. We also evaluated the performance of three classifiers on the original and reconstructed signals to show that our method did not affect the fault detection and diagnosis accuracy. These algorithms were of three different types. One was used as time series classification, the second classified a spectogram of a signal, and the last one was experimental, using only a picture of a plot to classify a signal. Our results demonstrated that our method achieved a minimal reduction in performance after compression and reconstruction using our proposed method. Our contributions can be summarized as follows: •We propose a novel approach that combines an autoencoder with data corrections using a specially designed loss function. •We employ aggressive down-sampling in the initial layers of the autoencoder to ensure it can fit into standard memory and utilize a state-of-the-art ResNet-based autoencoder architecture. •Our method achieves the best compression performance for PD signal data, while preserving anomalies within the compressed representation. •Furthermore, our approach effectively handles very large input data (800 kB), compressing them to a representation that is 25 times smaller. The advantages of the proposed algorithm over alternative methods are readily apparent. This algorithm is capable of performing lossy compression on such data types without substantial loss of significant features. In contrast, other leading-edge, publicly available lossy compression techniques have not been able to compress this data efficiently. An alternative to these are lossless compression methods; however, they provide only a limited compression ratio. Furthermore, our approach capitalizes on the use of GPUs, which facilitate efficient and adaptable deep learning algorithms, as these data are characterized by high levels of noise and the useful signals are sparse. Traditional compression methods fall short in effectively handling this type of data. In the development of our study’s methodology, we integrate a ResNet-based autoencoder architecture (Asahi et al.,2021), leveraging the distinctive and underutilized capabilities of deep learning in the realm of hardware acceleration. The decision to employ an autoencoder is rooted in its documented proficiency in signal compression tasks, a critical consideration given the unique attributes and high noise levels inherent in PD signals. Such signals exhibit characteristics markedly different from other types (Misak et al.,2017;Martinovic and Fulnecek, 2021), where conventional leading-edge algorithms have often fallen short in effective processing. The incorporation of ResNet blocks within the autoencoder framework is strategically aimed at enhancing the downsampling process for the extensive PD signal data. ResNet’s proven success (Wickramasinghe et al.,2021) in various fields makes it an ideal candidate for this architecture, ensuring an optimized compression process that meticulously maintains signal integrity through deliberate large steps in the downsampling phase. This approach not only addresses the challenges posed by the specific nature of PD signals but also sets a new precedent in the efficient handling of similarly complex data. The rest of the paper is organized as follows. Section 2reveals a background and reviews related work on signal compression methods. Section 3describes the proposed algorithm and its components. Section 4presents the results and discusses the implications and limitations of our method. Section 5concludes the paper and suggests future work. Engineering Applications of Artificial Intelligence 133 (2024) 108267 3 L. Klein et al. 2. Background This section provides a concise overview of the relevant literature on PD detection and its applications. Furthermore, it surveys various techniques for compressing signals that contain PD or other scientific data, and discusses their advantages and limitations. 2.1. Detection of PD One of the challenges in power engineering is the detection of PD, which are electrical phenomena that can cause insulation degradation and failure of power equipment. PD are localized dielectric breakdowns of insulation defects or degradation of power equipment. PD detection is an important technique for condition monitoring and fault diagnosis of power equipment (Yaacob et al.,2014). Various methods have been proposed and applied for PD detection, such as electrical, chemical, acoustic, and optical methods, Yaacob et al. (2014) and Refaat and Shams (2018). Many of these methods also employ machine learning techniques to improve the accuracy and reliability of PD detection (Lu et al.,2020). UHF methods have been extensively researched, demonstrating a significant focus within the academic community. Furthermore, there are studies specifically concerned with Partial Discharges employing UHF techniques, as evidenced by works such as Orellana et al. (2023) and Roslizan et al. (2020), which contribute valuable insights into this area. In particular, for overhead transmission power lines with CC, most of the existing algorithms use a contact galvanic method (Xu et al.,2023;Thomas and K.V.,2023; Xi et al.,2023), which relies on a large public dataset provided by the ENET Centre for a kaggle competition (Kaggle,2019). However, a contact galvanic method has some drawbacks, such as high cost and difficulty of installation. Recently, a new dataset for PD detection in overhead power lines with CC using an antenna, which is a contactless method, was published. This method has the advantages of being cheaper and easier to install than a contact galvanic method (Martinovic and Fulnecek,2021;Fulnecek and Misak,2021), as it does not require any high voltage component or power outage during the installation process. Both methods can be seen in Fig. 1. In the development of this device, a concerted effort was made to align with principles of cost-efficiency, thereby deviating from the normative specifications of Ultra-High Frequency systems, which characteristically exceed 500 MHz. The discourse surrounding the feasibility of incorporating an ADC converter, characterized by elevated vertical resolution and a proximal sampling rate of 5 GHz, necessitates a comprehensive evaluation. Although technically plausible, the implementation of such a high-frequency AD converter would invariably lead to a surge in manufacturing costs, thus detracting from the device’s cost-effectiveness. This consideration is of paramount importance in the engineering of systems where the equilibrium between technical performance and economic viability is sought. The critique posited herein serves to illuminate a critical domain of research and development within the field, accentuating the ongoing challenge of reconciling technological advancements with cost constraints. 2.2. Algorithms for PD designed for antenna method Several algorithms have been proposed for the detection of PD using antennas in the literature. One of the first was presented by Fulnecek and Misak (2021), who applied a three-step approach consisting of noise and interference filtering, PD pulse clustering, and PD event identification based on signal steepness. However, this algorithm showed rather low performance and was not suitable for real environment applications. Martinovic and Fulnecek (2021) developed a fast algorithm for contactless PD detection on a remote gateway device, which was designed for low computational cost and high speed. Their algorithm used a boosting classifier and peak clustering to differentiate PD signals from Fig. 1. Data acquisition HW. noise and interference, and also employed signal steepness as a feature for PD detection. A more recent algorithm by Klein et al. (2023b) used a stacking ensemble of deep learning algorithms and focused mainly on signal peaks to detect faults. They implemented a custom preprocessing denoising algorithm to extract peaks from the signals, and then used a 1D-CNN and autoencoders networks as base learners, and a multilayer perceptron (MLP) as a meta learner to classify the peaks into PD and non-PD classes. They showed that their algorithm outperformed other methods on real-world data. The last algorithm that we will review is also proposed by Klein et al. (2023c). They compared different algorithms on hardware accelerators used for edge computing and proposed a custom classification of the spectrogram approach, which was compatible with hardware accelerators. They converted the signals into spectrogram images and used a custom ensemble of convolutional neural network (CNN) to classify them into PD and non-PD categories. They showed that their ensemble was more efficient in terms of power efficiency than a ResNet neural network (He et al.,2016) and other architectures and approaches. These algorithms illustrate the potential of antenna as a tool for PD detection and classification, as it can capture both the spectral and temporal features of PD signals. However, they also face some challenges, such as noise and interference suppression, signal compression and transmission, and computational complexity and efficiency. Therefore, further research is warranted to enhance the performance and reliability of antenna-based PD detection methods. To improve results, more training data should be used. 2.3. Compression method When we speak of a compression method, we refer to two algorithms. There is the compression algorithm that takes an input 𝛼and generates a representation 𝛽that requires fewer bits. And there is the decompression algorithm that operates on the compressed representation 𝛽to generate the output 𝛾(see Fig. 2). Based on the requirements of decompression, data compression algorithms can be divided into two broad classes: lossless compression schemes, in which 𝛾is identical to 𝛼, and lossy compression, which generally provide much better compression than lossless compression, but allows 𝛾to be different from 𝛼, in other words, 𝛾is ‘‘reasonable’’ approximation of 𝛼. Engineering Applications of Artificial Intelligence 133 (2024) 108267 4 L. Klein et al. Fig. 2. Compression method – conceptual schema. 2.4. Compression methods for scientific data The rapid growth of data (Cappello et al.,2019;Wen et al.,2018; Nalbantoglu et al.,2009) and the emergence of big data and Internet of Things (IoT) technologies (Azar et al.,2020;Ukil et al.,2015) have triggered a interest in lossy compression of scientific data, where data are rather noisy and slight loss of information is acceptable. Consequently, many lossy compression methods have been proposed and developed in recent years for general data compression (Kouznetsov,2021;Blalock et al.,2018). A common challenge for general scientific data is to compress them efficiently without losing important information. Two major families of data compression algorithms are zfp (Lindstrom and Isenburg,2006b) and SZ (Di and Cappello,2016), which have different approaches and characteristics. Zfp uses fixed-point representation conversion, near orthogonal block transformations, and embedded coding to encode each bit-plane. This algorithm has been widely analyzed and applied in various contexts (Hammerling et al.,2019;Diffenderfer et al.,2019;Hoang et al.,2020,2021). For example, article (Hammerling et al.,2019) used the zfp algorithm to compress climate data, or for particle data (Hoang et al.,2021). SZ employs a block-wise prediction-based compression model with three main steps: data point predictions, linear-scale quantization to convert the data to integer codes, and compression of the integer codes using lossless compression. This algorithm has also been extensively studied and applied (Zhao et al.,2022;Yu et al.,2022;Liu et al.,2022;Wang et al.,2022). SZ was applied for three-dimensional adaptive mesh refinement simulations or for molecular dynamics (Zhao et al.,2022). Moreover, some work has been done on error correcting codes for these algorithms (Fulp et al.,2021;Diffenderfer et al.,2019). Both of these frameworks were primarily used for floating point data compression, which was not that suitable for our data consisting of single bytes of signal. Machine learning algorithms (Said,2018;Chen et al.,2014) and deep learning techniques have been widely applied to various data types, such as image compression (Jamil et al.,2023). For example, a deep autoencoder was used for phonocardiography signals (Chien et al.,2020), and a hybrid model of LSTM and XGBoost (Yan et al., 2021) was proposed as a general method for data compression. Another application of the autoencoder was Das et al. (2020), where it was used to compress the smart grid data. They achieved good results with transfer learning. Researchers (Borova et al.,2022) conducted the performance analysis of different types of compression (wawelet-based, redundant points, and an autoencoder), but found no conclusive result. However, the neural network showed good performance even in their case where they had insufficient data for training. A similar approach of using CNN and autoencoder was proposed for the compression of long-span suspension bridge data. A very small reconstruction error was obtained even with a very high compression rate (Ni et al.,2019). Quantized autoencoder was also proposed for industrial data compression (Zhang et al.,2023) where the dual path structure of the autoencoder and the swish activation function was used. A multi-hierarchical network compression method was developed for rolling bearing data (Sun et al.,2022). Furthermore, Azar et al. (2020) introduced a deep learning compression algorithm that integrated the wavelet SZ (Di and Cappello,2016) for compression and then reconstructed the data using a neural network. The most recent approach, which was in some ways similar to our approach, was a work by Wu et al. (2022). They proposed splitting the data into anomalies and normal data for spectrum data compression. Normal data was compressed using an attention-based asymmetric convolutional neural network, and a nonlinear outlier compression algorithm performs outlier data compression. The design of their neural network was pretty heavy and can handle only limited size of input data. In addition, their approach was focused on compression of floating points with a wide range of values. 2.5. Compression methods for PD A few methods have been developed for PD compression, but they differ in their data acquisition methods and have less data to compress with a higher amount of noise. One of the methods developed for the compression of PD is based on Gaussian process regression (Mashimo et al.,2022). This method could reduce the size of the PD data by at least 10 %. They used a contact method and their data were much smaller (only a few kilobytes) and the waveform to compress was much cleaner than in our article, which has many sources of background noise (Fulnecek and Misak,2021) (broadcast, corona, sunset, internal parts of measuring device, etc.). One of the studies that is relevant to our work is the one conducted by Zhao et al. (2023), who propose a compression reconstruction method called TSR-DRRT. It aims to sparsely represent and accurately reconstruct noisy signals. It decomposes signals to obtain PD pulses and trains a transfer SR dictionary with different signals. It compresses signals by enhancing the match between dictionary atoms and PD pulses. This technique reconstructs signals by adaptively adjusting the iteration criteria based on the difference between the correlation of the dictionary and the signal. The method uses a similar signal size as ours, but with much less noise and more sparsity, which facilitate compression, as their data come from simulations and not from production data from real environment containing many noise sources. Also, their data volume is compressed to only 20 % of the original signal. The rest of the methods were used for the classification or detection of PD using compressed data (Yunpeng et al.,2002;Majidi et al.,2015; Zhang et al.,2015;Yang et al.,2017). These algorithms cannot be used to compress a signal, and reconstruction is also not possible. An interesting example is an article by Vantuch et al. (2019a), which used a text compression-based feature extraction for classification. 3. Methodology This section describes the methodology used in this study to compress and reconstruct PD signals using a deep auto-encoder. First, the preprocessing of the dataset, and the characteristics of the PD data are introduced. Next, the design of the deep autoencoder is explained, including the parameters and hyperparameters of the network. Then, the loss function used to backpropagate the reconstruction error is defined. After that, the data correction method is proposed to keep anomalies in the reconstructed signal. Furthermore, the training and testing procedures of the autoencoder are described. Finally, the performance of the proposed method is compared to other compression algorithms in terms of compression ratio and signal quality. Additionally, the classification accuracy of using reconstructed and original signals for PD classification is evaluated. 3.1. Dataset and brief description of data We employed a publicly available dataset (Klein et al.,2023a) that was obtained from a partial discharge detection method using an antenna. The antenna type was Bonywhip and the analog-to-digital converter (ADC) had a high sampling rate of 500 MHz. The data was collected from 8 stations that recorded the signal from 22 kV overhead power transmission lines with coupling capacitors. Each sample captured a 20 ms signal that corresponded to a single cycle of the utility Engineering Applications of Artificial Intelligence 133 (2024) 108267 5 L. Klein et al. Table 1 Samples per station. Station ID Without PDs With PDs 52007 3231 269 52008 2502 998 52009 3336 164 52010 3437 63 52011 3491 9 52012 3495 5 52013 3145 353 52014 3420 80 frequency. The sample size was 800 kB with values ranging from −128 to 127. The data contained significant noise from various sources (Fulnecek and Misak,2021). For example, two clusters of signals in most samples originated from the internal power supply. PDs, which were the target of the data acquisition, appeared mainly as single peaks (as PDs are small sparks in the insulation) that deviated from normal data. However, there were also many other peaks that could indicate PD, but could have been caused by other phenomena such as corona discharge, rim effect, and others. As the dataset is pretty large (151,434 samples ≈ 120 GB), we sampled from every station signal which has annotation that contains PD and also randomly up to a total of 3500 samples without PD. In total 3500 samples for each station. The total number of samples was 31,500. The number of samples per station can be seen in Table 1. 3.1.1. Noise in the data Signal data are primarily affected by two types of noise (Fulnecek and Misak,2021): Discrete Spectral Interference (DSI) and Random Pulse Interference (RPI). DSI typically stems from narrowband sources, like radio broadcast stations. It overlays PD patterns with a DSI noise signal, which can significantly decrease the visibility of these patterns by lowering their amplitude ratio in relation to the noise. In cases of high amplitude DSI noise, PD patterns may be completely masked, making detection based solely on amplitude impractical (Martinovic and Fulnecek,2021). The amplitude of radio wave noise, across a frequency spectrum up to 10 MHz, exhibits fluctuations due to various factors such as seasonal shifts, diurnal cycles, and meteorological conditions. These fluctuations are exacerbated by the dual modes of wave propagation: groundwave and skywave. Solar activities notably impact DSI noise levels, which typically escalate during nighttime hours (Govindarajan et al.,2019; Fulnecek and Misak,2021). On the other hand, RPI is characterized by intermittent pulse occurrences within the signal that are not associated with CCs or PDs. These pulses can originate from natural events like lightning, electrical activities such as switching, or electrical discharges. RPI, identified as broadband interference, can create signal peaks that resemble CCinduced pulses. A specific occurrence of RPI can be seen when disconnectors are in an open state, where the PD detected is a result of floating voltage electrodes rather than CC. This form of PD is recognized by a sequence of pulses, or a ‘‘pulse train’’, with nearly consistent intervals and amplitudes, creating a pattern of repetitive pulse sequences (Fulnecek and Misak,2021;Chaudhuri et al.,2023). 3.2. Motivation for proposed compression algorithm We opted to utilize deep learning techniques primarily due to the inherent limitations observed in traditional lossy compression algorithms concerning their ability to preserve the integrity of signal patterns. It is crucial to emphasize that our primary focus resided in the compression and transmission of signals, with no intention of addressing noise reduction. De-noising did not constitute an objective within the scope of our research. Instead, our aim was to safeguard the distinctive patterns present in PD data, which might otherwise be erroneously interpreted as noise. Furthermore, deep learning offers the advantage of GPU acceleration, leading to a substantial enhancement in processing efficiency. This approach remains relatively unexplored in this domain, further motivating our decision to investigate its potential advantages. It is essential to clarify that the motivation behind the proposed method arises from the inherent constraints associated with neural networks in compressing anomalies (Qian et al.,2022). These anomalies, which can be pivotal in specific scenarios, tend to be less effectively represented by neural networks due to their primary focus on learning conventional patterns as opposed to exceptions or outliers. Moreover, we delve into the context of PD as an illustrative example where anomalies play a critical role. Our data correction algorithm excels at capturing such anomalies with precision, in contrast to a neural network, which might learn to represent the general signal while overlooking these crucial irregularities (Qu and Qi,2018), including high-amplitude transient noise. To address this challenge, we employ autoencoders, a well-establi shed approach commonly used for denoising and anomaly detection (Chen et al.,2018). Autoencoders prove particularly effective in this context due to their inherent capability to learn standard data features (Majumdar,2018), enabling the identification and correction of anomalies that deviate from these established norms. 3.3. Overview of the proposed compression method The complete compression method is visually represented in Fig. 3. Also the algorithm can be found on Github.1The compression algorithm involves several steps: 1. The algorithm begins by collecting input data, e.g. on the Edge device. The algorithm preprocesses the data (see Section 3.3.1). 2. Subsequently, these data are processed using a 1D ResNet-based deep neural network encoder (see Section 3.3.2 and further). 3. Using the output of the autoencoder in conjunction with the original signal, correction data are generated (see Section 3.3.6). 4. The encoder’s output, latent space data, and the correction data are subsequently compressed and transmitted to a remote server (see Section 3.3.7). When the compressed data reaches the remote server, it can be decompressed or left intact in compressed form. Decompression algorithm consists of following steps: 1. Using a 1D ResNet-based deep neural network decoder, the signal is decoded. 2. Then data correction procedures are applied. 3. The final reconstructed signal is then restored and made available for further use. 3.3.1. Preprocessing To preprocess the data, we normalized the signal to facilitate the compression of periodic patterns. We did this by finding the window with the highest average absolute value summed over a window and rolling the signal by the position of that window. This way, a signal containing clusters had the clusters always in the same place (for example, internal components create two clusters of signal each period). This preprocessing did not corrupt the data since the signal is continuous and captures only a single period of the oscillations of alternating current. To standardize each sample individually, we applied Z-Score Normalization. This reduced the demands on the autoencoder, which only 1https://github.com/Lukykl1/DeepPDCompressor Engineering Applications of Artificial Intelligence 133 (2024) 108267 6 L. Klein et al. Fig. 3. Overview of the compression and transmission process. The cylindrical nodes represent data, while the rectangular nodes represent data processing. Fig. 4. Multiple samples without any PDs. needs to encode the differences between data points and not the varying scale of the signal. The mean and standard deviation are then transmitted with the encoded data to be used for reconstruction. We observed that this normalization improved the accuracy and speed of the autoencoder training. The example of a signals can be seen in a Fig. 4, where they do not contain partial discharges and in Fig. 5 with PDs. 3.3.2. Design of autoencoder We used similar architecture to a convolutional ResNet autoencoder (Wickramasinghe et al.,2021), but we modified it to use it for a large 1D signal. We used 1D convolution and instead of maxpooling in a first layer, we used a stride and other small modifications like activation function after skip connection and number of convolutions and activations in block, where our approach should be more inline with original ResNet architecture. Also, the architecture and downsampling and upsampling were modified to support very large input. A 1D CNN deep autoencoder is a type of neural network that consists of two parts: an encoder and a decoder. The encoder takes the input data and transforms them into a lower-dimensional latent representation, while the decoder reconstructs the original data from the latent representation. The goal of the autoencoder is to minimize the reconstruction error between the input and output, while learning a compact and meaningful representation of the data. To improve the performance of the autoencoder, we used residual connections between the encoder and the decoder. Residual connections are shortcuts that allow information to flow directly from one layer to another without passing through intermediate layers. They help Engineering Applications of Artificial Intelligence 133 (2024) 108267 7 L. Klein et al. Fig. 5. Multiple samples with visible PDs. Fig. 6. Conceptualized design of proposed autoencoder. to avoid the problem of vanishing gradients and degradation problem, and enable deeper networks to be trained. The residual connections were implemented using ResNet blocks, which consisted of two 1D convolutional layers with Mish activation function (Misra,2019) and the skip connection (with a convolution with kernel of size one and interpolation) is added to the result, and another Mish activation function is used. We used a Mish activation function, which should perform a little bit better than other state-of-the-art activation functions. Batch normalization was also used and applied before activation functions. Our 1D CNN deep autoencoder was symmetrical and has 8 ResNet blocks for the decoder and 8 blocks for the encoder, where each block consisted of two 1D convolutional layers with the same number of filters and kernel size, followed by a skip connection that adds the input of the block to its output. The first layer and the last layer were not ResNet blocks, and they were only combination of convolution, batch normalization, and non-linearity (Mish activation function). We utilized large kernels in the first layers, which should increase the effective perception field (Ding et al.,2022), which was needed when we used aggressive down scaling in the first layers in order to decrease the input size. However, the goal of the autoencoder was also to preserve the shape of the signal. The proposed part of the autoencoder architecture can be seen in Fig. 6 (the scale does not correspond as the image would be incomprehensible). The layer parameters are shown in Table 2 for encoder and for decoder in Table 2, where the size of kernel, filter count, and scale of down/up scaling are described and also if the layer consists of a ResNet block. The architecture of our network, particularly exemplified through each ResNet block, incorporates a strategic downsampling of the signal. This process is pivotal to the design’s efficacy, as the omission of numerous ResNet blocks would precipitate a significant augmentation in memory consumption. Consequently, ResNet blocks are not merely critical for downsampling; they also mandate a deep network architecture. This requirement stems from the imperative to mitigate memory constraints, a consideration of paramount importance given that our input data approximates 800kB. In our experimental endeavors, we explored the modification of the step size to facilitate a less profound network depth. Nevertheless, this modification proved inadequate in capturing the quintessential details, resulting in outputs that lacked critical features, often omitting essential aspects of the signal’s morphology, and broadly failing to facilitate effective learning. Although further fine-tuning may engender improvements, it remains crucial to strike an optimal balance between the network’s depth and the fidelity of feature extraction, ensuring that essential details are preserved while adhering to the limitations imposed by memory capacity. 3.3.3. Downsampling and upsampling As the input data is pretty large, we used aggressive scaling down to make it possible to use a deep autoencoder. For the first layer of encoder, we used a stride and large kernel to decrease an input size. Then we applied for rest of layers max pooling as it made a network faster and less computationally demanding. For the first layer, we downscaled by a factor of 8, then we used maxpooling with size of 4 and then two maxpooling of size 2. Then we changed only a number of filters. For the decoder we applied analogical steps. For the last four layers, we applied stride in transposed convolution in order to upsample the result. The last layer had a stride of 8, layer before 4 and two before stride of 2. 3.3.4. Proposed 1D ResNet block A ResNet-like block for the proposed algorithm is illustrated in Fig. 7. This block is adapted from C-RAE (Wickramasinghe et al., Engineering Applications of Artificial Intelligence 133 (2024) 108267 8 L. Klein et al. Table 2 Parameters of proposed autoencoder. (a) encoder Kernel Filters Scale ResNet Size Block 17 112 8 False 13 256 4 True 7 256 2 True 5 128 2 True 3 128 1 True 3 128 1 True 3 64 1 True 3 32 1 True 3 32 1 True 3 1 1 False (b) decoder Kernel Filters Scale ResNet Size Block 3 32 1 False 3 32 1 True 3 64 1 True 3 128 1 True 3 128 1 True 3 128 1 True 5 256 2 True 7 256 2 True 13 112 4 True 17 1 8 False 2021), a compression algorithm that uses a convolutional recurrent autoencoder (C-RAE) for 1D signals. To the best of our knowledge, no other compression algorithm combined a ResNet-like architecture and an autoencoder for 1D signals. The block consists of the following components: a 1D convolution layer, a batch normalization layer, and an activation function. If necessary, a maxpooling layer is applied after the activation function. Then, another 1D convolution layer and a batch normalization layer are added. A skip connection is also added to this output. The skip connection is composed of a 1D convolution layer with a kernel size of 1 and a batch normalization layer. The output of the skip connection is resized to match the size of the previous output (if necessary) by interpolation. Finally, another activation function is applied after adding the skip connection to the previous output. 3.3.5. Loss function In this paper, we introduce an adaptive loss function for autoencoders that aims to mitigate the impact of anomalies in the input and output vectors. We define anomalies as the largest mean squared errors (MSE) between the vectors, as well as the windowed maxima and minima in both vectors, which may indicate outliers or noise. By progressively discarding these anomalies during training, we expect the autoencoder to learn the regular patterns in the correction data more effectively. The development of our adaptive loss function was predicated on the initial acquisition of the signal’s morphology, with a subsequent phase of attenuating the influence of anomalies as training advances. Such anomalies are mitigated through the deployment of our data correction algorithm, thereby diminishing the neural network’s necessity to assimilate these irregularities. It is acknowledged that the underpinning rationale of our adaptive loss function leans more towards an empirical basis rather than a theoretical one. Nonetheless, empirical evidence from our experiments indicates that the strategy of omitting non-learnable data points exclusively during the early phases of training, as opposed to a continuous exclusion throughout the training process, adversely affects the convergence efficiency of the neural network. We designed a loss function that incrementally eliminates the most anomalous errors from the computation of MSE, our baseline loss Fig. 7. ResNet-like block. function. The elimination process is driven by the parameter 𝑔. We began with 𝑔= 0, that is, no elimination, and increased the number of eliminated errors by 𝛥𝑔 = 300 every two training epochs, up to a maximum of 𝑔= 4000 errors for each signal. We selected MSE as our baseline loss function based on our preliminary experiments that demonstrated its superior performance over mean absolute error (MAE) loss or a hybrid of spectral and MSE loss. The pseudocode of our proposed loss function is shown in Algorithm 1. Furthermore, our loss function disregarded the 𝑙= 10 highest points and the 𝑙lowest points in a window of 𝑤= 2 × 104points to diminish the effect of anomalies. An example of a function of this adaptive loss can be seen in Fig. 8(b), where the autoencoder does not learn to reconstruct peaks and anomalies from the original signal, Fig. 8(a), and these peaks are added later using a data correction procedure, Fig. 8(c). In a more rigorous and scholarly framework, we articulate the concept as follows. Consider 𝑢 = [𝑢1, 𝑢2,…, 𝑢𝑛]to represent the output vector derived from the autoencoder, and 𝑣 = [𝑣1, 𝑣2,…, 𝑣𝑛]to denote the original input vector. Here, 𝑛symbolizes the dimensionality of these vectors, 𝑤denotes the window size for local analysis, 𝑔represents the global exclusion threshold, and 𝑙signifies the count of exclusions per window on a local scale. The adaptive loss function, pivotal in evaluating the discrepancy between 𝑢 and 𝑣, is mathematically formulated as: Adaptive Loss(𝑢, 𝑣) = ∑𝑖∈𝑆(𝑢𝑖−𝑣𝑖)2 |𝑆| In this expression, 𝑆embodies the set of indices eligible for inclusion, meticulously curated by negating the indices disqualified through either local or global exclusion criteria, described by: 𝑆={𝑖∈ {1,2,…, 𝑛} ∣ 𝑖∉𝐸and 𝑖∉𝐺} Wherein, 𝐸is delineated as the aggregation of indices precluded based on local window analysis—specifically, those indices corresponding to the 𝑙minimal and 𝑙maximal values within each window of size 𝑤. Concurrently, 𝐺encapsulates the indices of the top 𝑔values pursuant to the global criterion, which entails arranging the squared discrepancies (𝑢𝑖−𝑣𝑖)2in a descending sequence and selecting the foremost 𝑔entries. Engineering Applications of Artificial Intelligence 133 (2024) 108267 9 L. Klein et al. Fig. 8. Signal processing examples. To enrich the mathematical discourse, let us introduce formulations for the sets 𝐸and 𝐺more explicitly. The set 𝐸, pertaining to local exclusions, can be mathematically defined as: 𝐸= ⌈𝑛 𝑤⌉ ⋃ 𝑘=1 {Indices of 𝑙smallest and largest (𝑢𝑖−𝑣𝑖)2within window 𝑘} Here, the operation ⋃denotes the union of indices across all windows, and ⌈⋅⌉symbolizes the ceiling function, ensuring an integer count of windows. Furthermore, the set 𝐺, which delineates global exclusions, can be characterized after sorting the squared differences (𝑢𝑖−𝑣𝑖)2globally as: 𝐺={Indices of the top 𝑔values of (𝑢𝑖−𝑣𝑖)2} These expansions not only extend the mathematical framework but also facilitate a deeper understanding of the adaptive mechanism underlying the loss function, optimizing the autoencoder’s learning process by judiciously selecting the discrepancies that are most informative for model adjustment. 3.3.6. Data correction We propose a three-step data correction procedure based on the squared error between the original and reconstructed signals. The reconstructed signal was obtained by applying an autoencoder to the original signal. In the first step, we selected the 𝑁𝑥points that had the largest squared error in the reconstructed signal. In the second step, we identified the 𝑁𝑦points that have a high squared error in a non overlapping sliding window of size 𝑤over the reconstructed signal. In the third step, we detect the 𝑁𝑧points that are outliers in the original signal in a sliding window of size 𝑤over the reconstructed signal and that are not corrected by the previous two steps. Correction points are characterized by their value and position in the signal. The parameters for data correction can be seen in Table 3. Also a simple version of this algorithm can be seen in Algorithm 2. 3.3.7. Compressed data The output from compression algorithm consists of two types of compressed representation of the original signal: the latent space data generated by the encoder and the correction data. We applied compression techniques to both in our proposed algorithm. Algorithm 1: Adaptive Loss Function Input : 𝑢: output from autoencoder, 𝑢 ∈𝑅𝑛;𝑣: original input vector, 𝑣 ∈𝑅𝑛;𝑤: size of non-overlapping windows, 𝑤∈𝑁,1≤𝑤≤𝑛∧𝑤|𝑛;𝑔: global exclusion count (top 𝑔global values), 𝑔∈𝑁,𝑔 < 𝑛;𝑙: local exclusion count for given window (𝑙lowest and 𝑙highest values in a window), 𝑙∈𝑁,2𝑙 < 𝑤 Output: Value of adaptive loss function for given vectors 𝑢 and 𝑣 1𝐼←∅; 2for 𝑖←0to 𝑛 𝑤do 3𝐿←𝑖𝑤 + 1; 4𝑅←(𝑖+ 1)𝑤; 5𝐼←𝐼∪set of indices of 𝑙lowest values from 𝑣𝐿…𝑣𝑅; 6𝐼←𝐼∪set of indices of 𝑙highest values from 𝑣𝐿…𝑣𝑅; 7end 8𝐿←empty list; 9for 𝑖←1to 𝑛do 10 if 𝑖∉𝐼then 11 𝐴𝑝𝑝𝑒𝑛𝑑 (𝐿, (𝑢𝑖−𝑣𝑖)2); 12 end 13 end 14 Sort 𝐿in ascending order; 15 Remove last 𝑔elements from 𝐿; 16 return Mean value of 𝐿; We employed a bfloat16 (Brain Floating Point) format to compress the latent space data of our model. This format, which occupies 16 bits in computer memory, represents a wide dynamic range of numeric values by using a floating radix point. It is a truncated version of the 32-bit IEEE 754 single-precision floating-point format (binary32), but it preserves the approximate dynamic range of 32-bit floating-point numbers by retaining 8 exponent bits and reducing the precision to 8 bits. We found that this format was adequate for our latent space representation and could also potentially improved the training speed if our GPU supported mixed precision. For correction data, we used a lossless fastpfor256 algorithm (Trotman and Lin,2016), which was suitable for this kind of data and has a Engineering Applications of Artificial Intelligence 133 (2024) 108267 16 L. Klein et al. Table 16 Metrics results for optimizers. Metric Adadelta RMSprop NAdam Adam MSE 45.2 12.29 11.59 12.10 MAE 4.74 2.53 2.28 2.32 Spectral Distortion 1203.1 573.51 519.5 524.2 PSNR 34.1 38.2 39.1 39.0 FFT MAE 1765 977 884 758 SNR 6.151 6.151 6.151 6.151 SNR Reconstructed 7.413 6.845 6.747 6.602 PSNR Original 25.36 25.36 25.36 25.36 PSNR Reconstructed 26.72 25.95 26.03 27.88 rapidly. The choice of optimizer is also a critical component of hyperparameter tuning, warranting further investigation. Additionally, the potential to apply newer, more advanced optimizers exists, presenting avenues for enhanced performance. The analysis of metrics (see Table 16) across four optimization algorithms — Adadelta, RMSprop, NAdam, and Adam — highlights distinct performance characteristics. Adam and NAdam optimizers outperform others in accuracy, as evidenced by their lower MSE and MAE. Conversely, Adadelta exhibits the highest error rates, indicating lesser precision in reconstruction. Adam stands out for its minimal spectral distortion and highest PSNR, suggesting superior signal quality preservation. Adadelta, with higher spectral distortion and the lowest PSNR, faces challenges in maintaining signal fidelity. FFT analyses further demonstrate Adam and Nadam’s superior performance in the frequency domain, crucial for spectral analysis applications. Despite all optimizers maintaining consistent SNR, Adam slightly leads in SNR reconstructed, indicating better signal reconstruction. In summary, Adam emerges as the most effective optimizer, with Nadam closely behind. Adadelta, despite its limitations, shows potential in signal reconstruction. NAdam offers quicker training time. 4.5. Limitations A possible limitation of this study is the insufficient exploration of hyperparameters to optimize the compression method. Compression performance might be improved by selecting more appropriate parameters. Another potential drawback is the incomplete coverage of all features of the reconstructed signal. The signal might lack some features that are not captured by the three different algorithms used in this study. Moreover, the correction data might be more sophisticated and effective than the simple anomaly capture. This algorithm is designed for this specific type of signal and field, and it might not perform well for other types of data. However, it can provide a direction for future research and applications. The compression method is also not compatible with hardware accelerators (Edge TPU, Neural compute stick, etc.), but it is suitable for places without a reliable connection (for example, unstable 2G) but with enough power to support larger devices such as devices from the NVIDIA Jetson family (our case). The proposed compression also offers significant space savings for long-term storage. The architectural design of the neural network has not been exhaustively optimized, indicating a substantial opportunity for advancements in both efficiency and the compactness of the neural network’s structure. This domain merits additional investigation to fully realize the potential improvements. Contrastingly, when compared to conventional galvanic contact methodologies, antenna-based PD detection exhibits certain drawbacks, including a reduced detection range and the incapability to conduct Phase-Resolved Partial Discharge (PRPD) analysis, attributable to the lack of a carrier signal at 50 Hz. Nonetheless, it is imperative to acknowledge that this technique provides various advantages, such as cost-effectiveness, enhanced fault localization capabilities, increased feasibility for widespread industry adoption, scalability for extensive deployments, and the elimination of the need for power interruption during installation. Current endeavors are aimed at mitigating these drawbacks to bolster the reliability and applicability of antenna-based PD detection systems. In conclusion, the progression of contactless antenna detection technology is an active area of research, necessitating continued exploration to attain a level of reliability and preparedness for industrial application. Such efforts are instrumental in enriching the training dataset, thereby augmenting both the deployment efficiency and the validation of authenticity. Moreover, this facilitates the categorization of data on remote servers, enabling the transmission and classification of a larger dataset. This, in turn, enhances the accuracy of the detection methodology. 5. Conclusion and future work We proposed a novel method for lossy compression of signals with anomalies, which has two main components: a deep autoencoder with 1D CNN ResNet blocks and a correction data generator. The deep autoencoder compresses the shape of the signal by encoding it into a latent space and then decoding it into a reconstructed signal. The correction data generator calculates the difference between the original and reconstructed signals and stores the values and positions of the anomalies. This way, the signal preserves all the important features and can be compressed to a size 25× smaller than the original (4.1% of original size). The method is suitable for scenarios where there is an unreliable connection or a need to save storage space. 5.1. Practical application We have demonstrated that our methodology is apt for deployment on edge devices, where energy concerns are secondary. A practical example of this application involves utilizing a device from the NVIDIA Jetson series, which was the target for our algorithm’s design (limiting memory usage on the GPU, though it remains substantial and may not be suitable for smaller, more efficient accelerators like the Edge TPU). This approach facilitates the effective compression of signals to be transmitted. In numerous instances, energy concerns are negligible, especially when devices are installed directly beneath power distribution lines. However, the reliability of transmission and connectivity becomes paramount due to the prevalent use of unreliable 2G GSM connections in these settings. Based on our experience, transmitting a single uncompressed signal can take up to 45 min. Our proposed algorithm could reduce this duration to merely two minutes, significantly enhancing the frequency of sample transmission. Furthermore, by potentially reducing the number of samples required for transmission (solely with PDs), we could achieve highly efficient transmission and detection. 5.2. Future research directions As a possible direction for future research, we propose to optimize the proposed compression algorithm by applying hyperparameter tuning techniques and experimenting with different architectures for the deep autoencoder, such as multipath stacked autoencoder with dilated convolutions. We also aim to enhance the secondary compression of the latent space and the correction data by employing more efficient methods for selecting and encoding the anomalies. Furthermore, we plan to explore how to achieve error-bound compression by adjusting the size of the latent space according to the desired level of accuracy. Another interesting direction is to incorporate attention mechanisms into the algorithm, but this poses challenges in terms of performance and compatibility with simpler GPUs. In addition, we intend to use the compressed data for direct detection of PDs and subsequent faults (such as contacts with vegetation, etc.). Engineering Applications of Artificial Intelligence 133 (2024) 108267 17 L. Klein et al. In this paper, we proposed a compression method with novel aspects that leverages neural networks to achieve effective lossy compression for partial discharge detection signals obtained from a noisy antenna method. Our algorithm outperformed other lossless and lossy algorithms in reducing the data size to 4 % of the original size without compromising the accuracy in data classification. The reconstructed signals preserved the most important features of the data, which enables reliable PD identification and localization. Our method also facilitates simpler data acquisition and transmission from detection stations with higher frequency and lower bandwidth. Furthermore, our algorithm can reduce the workload and storage requirements for data analysis and management. CRediT authorship contribution statement Lukáš Klein: Conceptualization, Data curation, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. Jiří Dvorský: Data curation, Investigation, Methodology, Supervision, Visualization, Writing – original draft, Writing – review & editing. David Seidl: Conceptualization, Formal analysis, Funding acquisition, Project administration, Supervision, Validation. Lukáš Prokop: Data curation, Funding acquisition, Project administration, Supervision, Validation. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Data availability Data are public and published dataset. Acknowledgments This article has been produced with the financial support of the European Union under the REFRESH – Research Excellence For Region Sustainability and High-tech Industries project number CZ.10.03.01/00/22_003/0000048 via the Operational Programme Just Transition and TN02000025 National Centre for Energy II. This work was supported by SGS, VŠB – Technical University of Ostrava, Czech Republic, under the grant No. SP2024/006 ‘‘Parallel processing of Big Data XI’’. This work was supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90254). This work was partially supported by SGS, VŠB – Technical University of Ostrava, Czech Republic, under the grant No. SP2024/043. References Asahi, Y., Fujii, K., Heim, D.M., Maeyama, S., Garbet, X., Grandgirard, V., Sarazin, Y., Dif-Pradalier, G., Idomura, Y., Yagi, M., 2021. Compressing the time series of five dimensional distribution function data from gyrokinetic simulation using principal component analysis. Phys. Plasmas 28 (1), 012304. Azar, J., Makhoul, A., Couturier, R., Demerjian, J., 2020. Robust IoT time series classification with data compression and deep learning. Neurocomputing 398, 222–234. http://dx.doi.org/10.1016/j.neucom.2020.02.097, URL https://www.sciencedirect. com/science/article/pii/S0925231220302939. Bartnikas, R., 2002. Partial discharges. Their mechanism, detection and measurement. IEEE Trans. Dielectr. Electr. Insul. 9 (5), 763–808. http://dx.doi.org/10.1109/TDEI. 2002.1038663. Blalock, D., Madden, S., Guttag, J., 2018. Sprintz: Time series compression for the internet of things. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2 (3), http://dx.doi.org/10.1145/3264903. Boggs, S., 1982. Electromagnetic techniques for fault and partial discharge location in gas-insulated cables and substations. IEEE Trans. Power Appar. Syst. (7), 1935–1941. Borova, M., Prauzek, M., Konecny, J., Gaiova, K., 2022. A performance analysis of edge computing compression methods for environmental monitoring nodes with LoRaWAN communications. IFAC-PapersOnLine 55 (4), 387–392. http://dx.doi.org/ 10.1016/j.ifacol.2022.06.064, URL https://www.sciencedirect.com/science/article/ pii/S2405896322003792. Cappello, F., Di, S., Li, S., Liang, X., Gok, A.M., Tao, D., Yoon, C.H., Wu, X.-C., Alexeev, Y., Chong, F.T., 2019. Use cases of lossy compression for floating-point data in scientific data sets. Int. J. High Perform. Comput. Appl. 33 (6), 1201–1220. http://dx.doi.org/10.1177/1094342019853336. Chaovalit, P., Gangopadhyay, A., Karabatis, G., Chen, Z., 2011. Discrete wavelet transform-based time series analysis and mining. ACM Comput. Surv. 43 (2), 1–37. Chaudhuri, S., Ghosh, S., Dey, D., Munshi, S., Chatterjee, B., Dalai, S., 2023. Denoising of partial discharge signal using a hybrid framework of total variation denoising-autoencoder. Measurement 223, 113674. http://dx.doi.org/10.1016/j. measurement.2023.113674. Chen, Z., Son, S.W., Hendrix, W., Agrawal, A., Liao, W.-K., Choudhary, A., 2014. NUMARCK: Machine learning algorithm for resiliency and checkpointing. In: SC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. pp. 733–744. http://dx.doi.org/10.1109/SC. 2014.65. Chen, Z., Yeo, C.K., Lee, B.S., Lau, C.T., 2018. Autoencoder-based network anomaly detection. In: 2018 Wireless Telecommunications Symposium. WTS, IEEE, pp. 1–5. Chiarot, G., Silvestri, C., 2023. Time series compression survey. ACM Comput. Surveys 55 (10), 1–32. http://dx.doi.org/10.1145/3560814. Chien, Y.-R., Hsu, K.-C., Tsao, H.-W., 2020. Phonocardiography signals compression with deep convolutional autoencoder for telecare applications. Appl. Sci. 10 (17), http://dx.doi.org/10.3390/app10175842, URL https://www.mdpi.com/2076-3417/ 10/17/5842. Das, L., Garg, D., Srinivasan, B., 2020. NeuralCompression: A machine learning approach to compress high frequency measurements in smart grid. Appl. Energy 257, 113966. http://dx.doi.org/10.1016/j.apenergy.2019.113966, URL https: //www.sciencedirect.com/science/article/pii/S0306261919316538. Dempster, A., Schmidt, D.F., Webb, G.I., 2021. MiniRocket: A very fast (almost) deterministic transform for time series classification. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. KDD ’21, Association for Computing Machinery, New York, NY, USA, pp. 248–257. http: //dx.doi.org/10.1145/3447548.3467231. Di, S., Cappello, F., 2016. Fast error-bounded lossy HPC data compression with SZ. In: 2016 IEEE International Parallel and Distributed Processing Symposium. IPDPS, pp. 730–739. http://dx.doi.org/10.1109/IPDPS.2016.11. Diffenderfer, J., Fox, A., Hittinger, J., Sanders, G., Lindstrom, P., 2019. Error analysis of ZFP compression for floating-point data. SIAM J. Sci. Comput. 41, A1867–A1898. http://dx.doi.org/10.1137/18M1168832. Ding, X., Zhang, X., Han, J., Ding, G., 2022. Scaling up your kernels to 31 ×31: Revisiting large kernel design in CNNs. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. CVPR, pp. 11953–11965. http://dx.doi.org/10. 1109/CVPR52688.2022.01166. Dozat, T., 2016. Incorporating Nesterov Momentum into Adam. In: Proceedings of the 4th International Conference on Learning Representations. pp. 1–4. Fulnecek, J., Misak, S., 2021. A simple method for tree fall detection on medium voltage overhead lines with covered conductors. IEEE Trans. Power Delivery 36 (3), 1411–1417. http://dx.doi.org/10.1109/tpwrd.2020.3008482. Fulp, D., Poulos, A., Underwood, R., Calhoun, J.C., 2021. ARC: An automated approach to resiliency for lossy compressed data via error correcting codes. In: Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing. Govindarajan, S., Subbaiah, J., Cavallini, A., Krithivasan, K., Jayakumar, J., 2019. Development of Hankel-SVD hybrid technique for multiple noise removal from PD signature. IET Sci. Measur. Technol. 13 (8), 1075–1084. http://dx.doi.org/10.1049/ iet-smt.2018.5679. Graves, A., 2013. Generating sequences with recurrent neural networks. http://dx.doi. org/10.48550/ARXIV.1308.0850, URL https://arxiv.org/abs/1308.0850. Gupta, S., Gupta, A., 2019. Dealing with noise problem in machine learning data-sets: A systematic review. Procedia Comput. Sci. 161, 466–474. http://dx.doi.org/10. 1016/j.procs.2019.11.146, URL https://www.sciencedirect.com/science/article/pii/ S1877050919318575. Hamacek, S., 2012. Problems of covered conductors running. URL https://dspace.vsb. cz/handle/10084/90358?show=full. Hammerling, D., Baker, A., Pinard, A., Lindstrom, P., 2019. A collaborative effort to improve lossy compression methods for climate data. http://dx.doi.org/10.1109/ DRBSD-549595.2019.00008. He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition. CVPR, pp. 770–778. http://dx.doi.org/10.1109/CVPR.2016.90. Hoang, D., Bhatia, H., Lindstrom, P., Pascucci, V., 2021. High-quality and low-memoryfootprint progressive decoding of large-scale particle data. In: 2021 IEEE 11th Symposium on Large Data Analysis and Visualization. LDAV, IEEE, http://dx.doi. org/10.1109/ldav53230.2021.00011. Engineering Applications of Artificial Intelligence 133 (2024) 108267 18 L. Klein et al. Hoang, D., Summa, B., Bhatia, H., Lindstrom, P., Klacansky, P., Usher, W., Bremer, P.- T., Pascucci, V., 2020. Efficient and flexible hierarchical data layouts for a unified encoding of scalar field precision and resolution. IEEE Trans. Vis. Comput. Graphics 27, 603–613. http://dx.doi.org/10.1109/TVCG.2020.3030381. Huffman, D.A., 1952. A method for the construction of minimum-redundancy codes. Proc. IRE 40 (9), 1098–1101. http://dx.doi.org/10.1109/JRPROC.1952.273898. Jamil, S., Piran, M.J., Rahman, M., Kwon, O.-J., 2023. Learning-driven lossy image compression: A comprehensive survey. Eng. Appl. Artif. Intell. 123, 106361. http:// dx.doi.org/10.1016/j.engappai.2023.106361, URL https://www.sciencedirect.com/ science/article/pii/S0952197623005456. Kabot, O., Fulneček, J., Mišák, S., Prokop, L., Vaculík, J., 2020. Partial discharges pattern analysis of various covered conductors. In: 2020 21st International Scientific Conference on Electric Power Engineering. EPE, pp. 1–5. http://dx.doi.org/10. 1109/EPE51172.2020.9269171. Kaggle, 2019. VSB power line fault detection. Kaggle URL https://www.kaggle.com/c/ vsb-power-line-fault-detection. Kaziz, S., Said, M.H., Imburgia, A., Maamer, B., Flandre, D., Romano, P., Tounsi, F., 2023. Radiometric partial discharge detection: A review. Energies 16 (4), 1978. Kingma, D.P., Ba, J., 2014. Adam: A method for stochastic optimization. http://dx.doi. org/10.48550/ARXIV.1412.6980, URL https://arxiv.org/abs/1412.6980. Kinsner, W., Greenfield, R., 1991. The lempel-ziv-welch (LZW) data compression algorithm for packet radio. In: [Proceedings] WESCANEX’91. IEEE, pp. 225–229. Klein, L., Fulneček, J., Seidl, D., Prokop, L., Mišák, S., Dvorský, J., Piecha, M., 2023a. A data set of signals from an antenna for detection of partial discharges in overhead insulated power line. Sci. Data 10 (1), 544. http://dx.doi.org/10.1038/s41597-02302451-1. Klein, L., Seidl, D., Fulneček, J., Prokop, L., Mišák, S., Dvorský, J., 2023b. Antenna contactless partial discharges detection in covered conductors using ensemble stacking neural networks. Expert Syst. Appl. 213, 118910. http://dx.doi.org/ 10.1016/j.eswa.2022.118910, URL https://www.sciencedirect.com/science/article/ pii/S0957417422019285. Klein, L., Žmij, P., Krömer, P., 2023c. Partial discharge detection by edge computing. IEEE Access 11, 44192–44204. http://dx.doi.org/10.1109/ACCESS.2023.3268763. Knuth, D.E., 1985. Dynamic huffman coding. J. Algorithms 6 (2), 163–180. Kouznetsov, R., 2021. A note on precision-preserving compression of scientific data. Geosci. Model Dev. 14 (1), 377–389. http://dx.doi.org/10.5194/gmd-14-377-2021, URL https://gmd.copernicus.org/articles/14/377/2021/. Leskinen, T., 2004. Finnish and slovene experience of covered conductor overhead lines. In: CIGRÉ 2004. Liang, X., Di, S., Tao, D., Li, S., Li, S., Guo, H., Chen, Z., Cappello, F., 2018. Errorcontrolled lossy compression optimized for high compression ratios of scientific datasets. In: 2018 IEEE International Conference on Big Data (Big Data). pp. 438–447. http://dx.doi.org/10.1109/BigData.2018.8622520. Liang, X., Zhao, K., Di, S., Li, S., Underwood, R., Gok, A.M., Tian, J., Deng, J., Calhoun, J.C., Tao, D., Chen, Z., Cappello, F., 2023. SZ3: A modular framework for composing prediction-based error-bounded lossy compressors. IEEE Trans. Big Data 9 (2), 485–498. http://dx.doi.org/10.1109/TBDATA.2022.3201176. Lindstrom, P., Isenburg, M., 2006a. Fast and efficient compression of floating-point data. IEEE Trans. Visual. Comput. Graph. 12, 1245–1250. http://dx.doi.org/10. 1109/TVCG.2006.143. Lindstrom, P., Isenburg, M., 2006b. Fast and efficient compression of floating-point data. IEEE Trans. Visual. Comput. Graph. 12 (5), 1245–1250. http://dx.doi.org/10. 1109/tvcg.2006.143. Liu, Y., Di, S., Zhao, K., Jin, S., Wang, C., Chard, K., Tao, D., Foster, I., Cappello, F., 2022. Optimizing error-bounded lossy compression for scientific data with diverse constraints. IEEE Trans. Parallel Distrib. Syst. 33 (12), 4440–4457. http://dx.doi. org/10.1109/TPDS.2022.3194695. Löning, M., Király, F., Bagnall, T., Middlehurst, M., Ganesh, S., Oastler, G., Lines, J., Walter, M., ViktorKaz, Mentel, L., Chrisholder, Tsaprounis, L., RNKuhns, Parker, M., Owoseni, T., Rockenschaub, P., Danbartl, Jesellier, Eenticott-Shell, Gilbert, C., Guzal Bulatova, Lovkush, Schäfer, P., Khrapov, S., Buchhorn, K., Take, K., Subramanian, S., Meyer, S.M., AidenRushbrooke, Rice, B., 2022. sktime/sktime: v0.13.4. Zenodo, http://dx.doi.org/10.5281/ZENODO.7117735, URL https://zenodo.org/ record/7117735. Lu, S., Chai, H., Sahoo, A., Phung, B.T., 2020. Condition monitoring based on partial discharge diagnostics using machine learning methods: A comprehensive stateof-the-art review. IEEE Trans. Dielectr. Electr. Insul. 27 (6), 1861–1888. http: //dx.doi.org/10.1109/TDEI.2020.009070. Majidi, M., Fadali, M.S., Etezadi-Amoli, M., Oskuoee, M., 2015. Partial discharge pattern recognition via sparse representation and ANN. IEEE Trans. Dielectr. Electr. Insul. 22 (2), 1061–1070. http://dx.doi.org/10.1109/TDEI.2015.7076807. Majumdar, A., 2018. Blind denoising autoencoder. IEEE Trans. Neural Netw. Learning Syst. 30 (1), 312–317. Martinovic, T., Fulnecek, J., 2021. Fast algorithm for contactless partial discharge detection on remote gateway device. IEEE Trans. Power Delivery 1. http://dx.doi. org/10.1109/tpwrd.2021.3104746. Mashimo, T., Takahashi, T., Ishino, R., Hori, Y., 2022. Development of data compression method of partial discharge waveform for remote insulation diagnosis in manhole for power transmission cable. In: 2022 9th International Conference on Condition Monitoring and Diagnosis. CMD, pp. 92–97. http://dx.doi.org/10.23919/ CMD54214.2022.9991387. Matthews, B., 1975. Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochim. Biophys. Acta (BBA) - Protein Struct. 405 (2), 442–451. http://dx.doi.org/10.1016/0005-2795(75)90109-9. Misak, S., Fulnecek, J., Jezowicz, T., Vantuch, T., Burianek, T., 2017. Usage of antenna for detection of tree falls on overhead lines with covered conductors. Adv. Electr. Electron. Eng. 15 (1), http://dx.doi.org/10.15598/aeee.v15i1.1894. Misak, S., Kratky, M., Prokop, L., 2016. A novel method for detection and classification of covered conductor faults. Adv. Electr. Electron. Eng. 14 (5), http://dx.doi.org/ 10.15598/aeee.v14i5.1733. Misak, S., Pokorny, V., 2015. Testing of a covered conductor’s fault detectors. IEEE Trans.Power Delivery 30 (3), 1096–1103. http://dx.doi.org/10.1109/tpwrd.2014. 2357072. Misra, D., 2019. Mish: A self regularized non-monotonic activation function. http://dx. doi.org/10.48550/ARXIV.1908.08681, URL https://arxiv.org/abs/1908.08681. Nalbantoglu, Ö., Russell, D., Sayood, K., 2009. Data compression concepts and algorithms and their applications to bioinformatics. Entropy 12 (1), 34–52. http: //dx.doi.org/10.3390/e12010034. Ni, F., Zhang, J., Noori, M.N., 2019. Deep learning for data anomaly detection and data compression of a long-span suspension bridge. Comput.-Aided Civ. Infrastruct. Eng. 35 (7), 685–700. http://dx.doi.org/10.1111/mice.12528. Orellana, L., Ardila-Rey, J., Avaria, G., Davis, S., 2023. Danger assessment of the partial discharges temporal evolution on a polluted insulator using UHF measurement and deep learning. Eng. Appl. Artif. Intell. 124, 106573. http://dx.doi.org/10.1016/ j.engappai.2023.106573, URL https://www.sciencedirect.com/science/article/pii/ S0952197623007571. Qian, J., Song, Z., Yao, Y., Zhu, Z., Zhang, X., 2022. A review on autoencoder based representation learning for fault detection and diagnosis in industrial processes. Chemometr. Intell. Lab. Syst. 104711. Qu, Y., Qi, H., 2018. uDAS: An untied denoising autoencoder with sparsity for spectral unmixing. IEEE Trans. Geosci. Remote Sens. 57 (3), 1698–1712. Refaat, S.S., Shams, M.A., 2018. A review of partial discharge detection, diagnosis techniques in high voltage power cables. In: 2018 IEEE 12th International Conference on Compatibility, Power Electronics and Power Engineering. CPE-POWERENG 2018, pp. 1–5. http://dx.doi.org/10.1109/CPE.2018.8372608. Roslizan, N.D., Rohani, M.N.K.H., Wooi, C.L., Isa, M., Ismail, B., Rosmi, A.S., Mustafa, W.A., 2020. A review: Partial discharge detection using UHF sensor on high voltage equipment. J. Phys. Conf. Ser. 1432 (1), 012003. http://dx.doi.org/ 10.1088/1742-6596/1432/1/012003, URL http://dx.doi.org/10.1088/1742-6596/ 1432/1/012003. Ruder, S., 2016. An overview of gradient descent optimization algorithms. http://dx. doi.org/10.48550/ARXIV.1609.04747, URL https://arxiv.org/abs/1609.04747. Said, A., 2018. Machine learning for media compression: challenges and opportunities. APSIPA Trans. Signal Inf. Process. 7, e8. http://dx.doi.org/10.1017/ATSIP.2018.12. Salomon, D., 2007. Data compression: The complete reference, fourth ed. Springer-Verlag, London. Stone, G., 2005. Partial discharge diagnostics and electrical equipment insulation condition assessment. IEEE Trans. Dielectr. Electr. Insul. 12 (5), 891–903. http: //dx.doi.org/10.1109/tdei.2005.1522184. Sun, J., Liu, Z., Wen, J., Fu, R., 2022. Multiple hierarchical compression for deep neural network toward intelligent bearing fault diagnosis. Eng. Appl. Artif. Intell. 116, 105498. http://dx.doi.org/10.1016/j.engappai.2022.105498, URL https: //www.sciencedirect.com/science/article/pii/S0952197622004882. Thomas, J.B., K.V., S., 2023. Neural architecture search algorithm to optimize deep transformer model for fault detection in electrical power distribution systems. Eng. Appl. Artif. Intell. 120, 105890. http://dx.doi.org/10.1016/ j.engappai.2023.105890, URL https://www.sciencedirect.com/science/article/pii/ S095219762300074X. TorchVision maintainers and contributors, 2016. TorchVision: Pytorch’s computer vision library. https://github.com/pytorch/vision. Trotman, A., Lin, J., 2016. In vacuo and in situ evaluation of SIMD codecs. In: Proceedings of the 21st Australasian Document Computing Symposium. ADCS ’16, Association for Computing Machinery, New York, NY, USA, pp. 1–8. http: //dx.doi.org/10.1145/3015022.3015023. Ukil, A., Bandyopadhyay, S., Pal, A., 2015. IoT data compression: Sensor-agnostic approach. In: 2015 Data Compression Conference. pp. 303–312. http://dx.doi.org/ 10.1109/DCC.2015.66. Vantuch, T., Prílepok, M., Fulneček, J., Hrbáč, R., Mišák, S., 2019a. Towards the text compression based feature extraction in high impedance fault detection. Energies 12 (11), http://dx.doi.org/10.3390/en12112148, URL https://www.mdpi.com/19961073/12/11/2148. Vantuch, T., Prílepok, M., Fulneček, J., Hrbáč, R., Mišák, S., 2019b. Towards the text compression based feature extraction in high impedance fault detection. Energies 12 (11), 2148. http://dx.doi.org/10.3390/en12112148. Voldhaug, L., Robertson, C., 1995. MV overhead lines using XLPE covered conductors. Scandinavian experience and NORWEB developments. In: Second International Conference on the Reliability of Transmission and Distribution Equipment, 1995.. pp. 52–60. http://dx.doi.org/10.1049/cp:19950218. Wang, D., Pulido, J., Grosset, P., Jin, S., Tian, J., Ahrens, J., Tao, D., 2022. TAC. In: Proceedings of the 31st International Symposium on High-Performance Parallel and Distributed Computing. ACM, http://dx.doi.org/10.1145/3502181.3531458. Engineering Applications of Artificial Intelligence 133 (2024) 108267 19 L. Klein et al. Wen, L., Zhou, K., Yang, S., Li, L., 2018. Compression of smart meter big data: A survey. Renew. Sustain. Energy Rev. 91, 59–69. http://dx.doi.org/10.1016/j.rser.2018.03. 088, URL https://www.sciencedirect.com/science/article/pii/S1364032118301849. Wickramasinghe, C.S., Marino, D.L., Manic, M., 2021. ResNet autoencoders for unsupervised feature learning from high-dimensional data: Deep models resistant to performance degradation. IEEE Access 9, 40511–40520. http://dx.doi.org/10.1109/ ACCESS.2021.3064819. Wu, G., Zhou, F., Ding, G., Wu, Q., Li, X.-Y., 2022. An efficient heterogeneous edge-cloud learning framework for spectrum data compression. IEEE Trans. Mob. Comput. 1. http://dx.doi.org/10.1109/TMC.2022.3153049. Xi, Y., Zhou, F., Zhang, W., 2023. Partial discharge detection and recognition in insulated overhead conductor based on bi-LSTM with attention mechanism. Electronics 12 (11), http://dx.doi.org/10.3390/electronics12112373, URL https://www.mdpi. com/2079-9292/12/11/2373. Xu, N., Wang, W., Fulneček, J., Kabot, O., Mišák, S., Wang, L., Zheng, Y., Gooi, H.B., 2023. TBMF framework: A transformer-based multilevel filtering framework for PD detection. IEEE Trans. Indu. Electron. 1–10. http://dx.doi.org/10.1109/tie.2023. 3274881. Yaacob, M.M., Alsaedi, M.A., Rashed, J.R., Dakhil, A.M., Atyah, S.F., 2014. Review on partial discharge detection techniques related to high voltage power equipment using different sensors. Photonic Sens. 4 (4), 325–337. http://dx.doi.org/10.1007/ s13320-014-0146-7. Yan, Z., Wang, J., Sheng, L., Yang, Z., 2021. An effective compression algorithm for real-time transmission data using predictive coding with mixed models of LSTM and xgboost. Neurocomputing 462, 247–259. http://dx.doi.org/10. 1016/j.neucom.2021.07.071, URL https://www.sciencedirect.com/science/article/ pii/S0925231221011474. Yang, F., Sheng, G., Xu, Y., Hou, H., Qian, Y., Jiang, X., 2017. Partial discharge pattern recognition of XLPE cables at DC voltage based on the compressed sensing theory. IEEE Trans. Dielectr. Electr. Insul. 24 (5), 2977–2985. http://dx.doi.org/10.1109/ TDEI.2017.006553. Yu, X., Di, S., Zhao, K., Tian, J., Tao, D., Liang, X., Cappello, F., 2022. SZx: an ultra-fast error-bounded lossy compressor for scientific datasets. arXiv:2201.13020. Yunpeng, L., Fangcheng, L., Zhiye, C., Yanqing, L., 2002. Data compression and pattern recognition for partial discharge ultrasonic signal based on fractal theory. In: Proceedings. International Conference on Power System Technology. 2, pp. 958–961. http://dx.doi.org/10.1109/ICPST.2002.1047541. Zeiler, M.D., 2012. ADADELTA: An adaptive learning rate method. http://dx.doi.org/ 10.48550/ARXIV.1212.5701, URL https://arxiv.org/abs/1212.5701. Zhang, Y., Upton, D., Jaber, A., Ahmed, H., Khan, U., Saeed, B., Mather, P., Lazaridis, P., Atkinson, R., Vieira, M.Q., et al., 2015. Multiple source localization for partial discharge monitoring in electrical substation. In: 2015 Loughborough Antennas & Propagation Conference. LAPC, IEEE, pp. 1–4. Zhang, M., Zhang, H., Zhang, C., Yuan, D., 2023. Communication-efficient quantized deep compressed sensing for edge-cloud collaborative industrial IoT networks. IEEE Trans. Ind. Inform. 19 (5), 6613–6623. http://dx.doi.org/10.1109/TII.2022. 3202203. Zhao, K., Di, S., Dmitriev, M., Tonellot, T.-L.D., Chen, Z., Cappello, F., 2021. Optimizing error-bounded lossy compression for scientific data by dynamic spline interpolation. In: 2021 IEEE 37th International Conference on Data Engineering. ICDE, pp. 1643–1654. http://dx.doi.org/10.1109/ICDE51399.2021.00145. Zhao, K., Di, S., Liang, X., Li, S., Tao, D., Chen, Z., Cappello, F., 2020. Significantly improving lossy compression for HPC datasets with second-order prediction and parameter optimization. In: Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing. HPDC ’20, Association for Computing Machinery, New York, NY, USA, pp. 89–100. http://dx.doi.org/10. 1145/3369583.3392688. Zhao, K., Di, S., Perez, D., Liang, X., Chen, Z., Cappello, F., 2022. MDZ: An efficient error-bounded lossy compressor for molecular dynamics. In: 2022 IEEE 38th International Conference on Data Engineering. ICDE, pp. 27–40. http://dx.doi.org/ 10.1109/ICDE53745.2022.00007. Zhao, S., Zhao, H., Libo, M., Yuehan, Q., Hui, R., 2023. Partial discharge signal compression reconstruction method based on transfer sparse representation and dual residual ratio threshold. IET Science, Measurement & Technology http://dx. doi.org/10.1049/smt2.12148.