Full text
6428 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 ResBaGAN: A Residual Balancing GAN with Data Augmentation for Forest Mapping Álvaro G. Dieste , Francisco Argüello , and Dora B. Heras , Member, IEEE Abstract—Although deep learning techniques are known to achieve outstanding classification accuracies, remote sensing datasets often present limited labeled data and class imbalances, two challenges to attaining high levels of accuracy. In recent years, the GAN architecture has achieved great success as a data augmentation method, driving research toward furtherenhancements. This work presents ResBaGAN, a GAN-based method for the classification of remote sensing images, designed to overcome the challenges of data scarcity and class imbalances by constructing an advanced data augmentation framework. This framework builds upon a GAN architecture enhanced with an autoencoder initialization and class balancing properties, a superpixel-based sample extraction procedure with traditional augmentation techniques, and an improved residual network as classifier. Experiments were conducted on large, very high-resolution multispectral images of riparian forests in Galicia, Spain, with limited training data and strong class imbalances, comparing ResBaGAN to other machine learning methods such as simpler GANs. ResBaGAN achieved higher overall classification accuracies, particularly improving the accuracy of minority classes with F1-score enhancements reaching up to 22%. Index Terms—BAGAN, classification, data augmentation, multispectral, residual network, superpixels. ABBREVIATIONS AA Average accuracy. ACGAN Auxiliary classifier generative adversarial network. Adam Adaptive moment estimation. BAGAN Balancing generative adversarial network. CGAN Conditional generative adversarial network. Manuscript received 16 December 2022; revised 17 April 2023 and 23 May 2023; accepted 25 May 2023. Date of publication 1 June 2023; date of current version 21 July 2023. This work was supported in part by the Agencia Estatal de Investigación, Government of Spain under Grant PID2019-104834GB-I00 and Grant TED2021-130367B-I00, in part by the Consellería de Cultura, Educación, Formación Profesional e Universidades, Xunta de Galicia under Grant ED431G 2019/04 and Reference Competitive Group accreditation under Grant ED431C 2022/16, and in part by the Junta de Castilla y León under Grant Project VA226P20 (PROPHET–II). All are cofounded by the European Regional DevelopmentFund(ERDF).Theworkof ÁlvaroG.Diestewasalsosupportedby the Consellería de Cultura, Educación, Formación Profesional e Universidades, Xunta de Galicia under Grant ED481A 2022/257. (Corresponding author: Álvaro G. Dieste.) Álvaro G. Dieste and Dora B. Heras are with the Centro Singular de Investigación en Tecnoloxías Intelixentes, Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain (e-mail: alvaro.goldar[email protected]; [email protected]). Francisco Argüello is with the Departamento de Electrónica e Computación, Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain (e-mail: [email protected]). The code developed for this work will be available at https://github.com/ alvrogd/ResBaGAN. Digital Object Identifier 10.1109/JSTARS.2023.3281892 CNN Convolutional neural network. ELU Exponential linear unit. ETPS Extended topology preserving segmentation. F1 F1-score. FID Fréchet inception distance. FV Fisher vectors. GAN Generative adversarial network. GCN Graph convolutional network. κCohen’s kappa. KELM Kernel extreme learning machine. LeakyReLU Leaky rectified linear unit. OA Overall accuracy. PA Producer’s accuracy. PReLU Parametric rectified linear unit. ReLU Rectified linear unit. ResBaGAN Residual balancing generative adversarial network. ResNet Residual network. RBM Restricted Boltzmann machines. SEEDS Superpixels extracted via energy-driven sampling. SLIC Simple linear iterative clustering. UA User’s accuracy. UAV Unmanned aerial vehicle. VAE Variational autoencoder. WP WaterPixels. I. INTRODUCTION REMOTE sensing is essential for monitoring the Earth’s surface, aiding in tasks, such as tracking human-made constructions [1], detecting cropland changes [2], and studying ecosystems [3]. Accurate image classification methods are crucial for many of these applications, such as identifying areas invaded by non-native plant species for environmental monitoring. Unfortunately, remote sensing datasets often present limited labeled data and class imbalances (i.e., certain classes have significantly more samples than others), making it challenging to develop accurate classification methods [4]. These challenges become even more pronounced when processing data from multispectral sensors, given their limited spectral resolution [5]. It is crucial to develop advanced methods that optimally use the available data to overcome these challenges. Neural networks have become increasingly prevalent in remote sensing applications owing to their exceptional performance. For instance, Zhu et al. [6] and Yuan et al. [7] This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6429 reviewed the application of different deep learning techniques in agriculture and environmental monitoring. The CNN stands out for its particularly well-suited image processing capabilities, which has led to its widespread application. For example, in Morales et al. [8], DeeplabV3+ CNNs were used to monitor the deforestation of the Mauritia flexuosa palm, a dominant species in the Amazon rainforest ecosystem, through high-resolution aerial RGB images acquired by UAV. In Hamdi et al. [9],a modified U-Net CNN was implemented to automatically detect and map the damaged areas in a forested area in Bavaria, Germany, by using aerial photographs with four spectral bands. The U-Net architecture was also used in Isaienkov et al. [10] for deforestation detection in the forest-steppe zone in the Kharkiv region of Ukraine, by using multispectral images of Sentinel-2. Finally, a region-based Mask CNN was used in Chiang et al. [11] for an automated forest health diagnosis in RGB aerial images from the Wood of Cree in Scotland, aiding in the early detection of dead trees. Over the past few years, remote sensing research has developed more sophisticated deep learning architectures to better utilize the available training data. The CNN design has progressed from shallow architectures with few layers [12] to deeper, more powerful ones with tens or hundreds of layers, commonly known as residual networks [13],[14],[15],[16], which can be recognized by their innovative integration of skip-connections between layers. Recent advances have also introduced new components to CNNs. For example, Li et al. [17]includedattentionmechanisms into a simple CNN to extract richerspatialandspectralinformation.AnotherexampleisHong et al. [18], which combined a CNN with a GCN to model relationships between training samples, enhancing the classification performance. Researchers have also explored deep learning architectures beyond CNNs. For instance, He et al. [19] leveraged the transformer architecture [20] to assimilate high volumes of data, advancing the state-of-the-art in the multimodal semantic segmentation of images. Another promising research direction isusing multimodaldata fromvarioussensors for complexscene analysis. For instance, Hong et al. [21],[22] studied the use of dual-branch CNN architectures to combine features from two different modalities (e.g., hyperspectral and LiDAR). Furthermore,designingself-supervisedarchitecturesthatcanlearnfrom both unlabeled and labeled data is another promising area. For instance, Sun et al. [23] developed a general-purpose model by analyzing millions of unlabeled scenes with a transformerbacked autoencoder architecture, acquiring extensive remote sensing knowledge that can be fine-tuned for a specific task, outperforming various state-of-the-art architectures. Data augmentation provides another promising solution to address data scarcity and class imbalance constraints. These techniques can assist advanced deep learning classification architectures by synthetically generating new training samples to enrich the learning data. Numerous data augmentation techniques exist [24] and have been applied to remote sensing. The traditional approaches involve transformations that convert a sample into a new one and techniques that combine different samples of the same class to create new ones. For instance, Haut et al. [25] showed the effectiveness of augmentation through the random deletion of input patch segments, while Acción et al. [26] subdivided each patch and applied independent transformations to each segment. In Nalepa et al. [27],new samples were generated from the first principal component of the dataset or by calculating the mean value of each band. More novel augmentation approaches use generative techniques that synthesize samples from scratch after estimating the data distribution, as opposed to transforming existing samples into new ones. Currently, the most widely adopted data generative approach is the GAN [28]. This deep learning-based augmentation approach has shown great potential compared to traditional techniques [29], owing to its remarkable capacity to generate highly realistic and diverse data from scratch, outperforming earlier generativearchitectures,suchastheRBM[30]andtheVAE[31]. FurtheradvancementsintheGANarchitecturehavealsoenabled itsuseasaclassifiernetwork,establishingitasavaluabletoolfor remote sensing applications like environmental monitoring. For instance, in Shashank et al. [32], a GAN was used to identify the target epiphyte Werauhia kupperiana in RGB images acquired by UAV in Costa Rican forests. Other interesting scenarios of GANs for environmental monitoring are described in [6] and [33].Unfortunately,successfully training thesenetworksischallengingunlesslargedatasetsareavailable[34].Second,classimbalanceshinder the learning of minorityclasses, evenpreventing it entirely. In such cases, GANs tend to synthesize identical samples for these classes or even fail to capture their data distribution,resulting in pure noise [35]. However, a recent development in GANs, the BAGAN [35], provides a promising solution to enable the successful application of GANs when facing these two challenges. This work presents ResBaGAN, a novel GAN-based classification method for remote sensing images applied to environmental monitoring. The ResBaGAN architecture overcomes the challenges of data scarcity and class imbalances by constructing an advanced data augmentation framework to support the classifier. The framework builds upon the cutting-edge innovations of BAGAN to provide effective GAN-based data augmentation. In addition, the framework leverages a superpixel segmentation-based sample extraction process with traditional augmentation techniques, and a ResNet-based classifier [13] improved to further enhance its capabilities. Experiments were conducted to demonstrate how the combined synergistic effects of all components significantly improved classification performance, particularly in the minority classes, compared to simpler methods like a stand-alone BAGAN. More specifically: 1) The proposed method integrates a GAN architecture inspired by BAGAN to enable effective deep learning-based data augmentation for scarce datasets and their minority classes. 2) The sample extraction process is guided by a superpixel segmentation and includes traditional augmentation techniques, thereby facilitating the more accurate modeling of the different classes. The segmentation ensures that each
6430 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 sample predominantly contains pixels from a single class, while the additional augmentation techniques increase the diversity of learning data. 3) The classifier features a ResNet-based design improved to integrate data from different levels of abstraction for better generalization. This architecture enhances both the quality of the GAN-based augmentation and the final classification accuracy of ResBaGAN. 4) In terms of evaluation, the FID score [36], which is the current standard metric for GAN-based architectures on RGB data, was adapted to multispectral images. This modification allowed for an accurate assessment of the proposed deep learning-based data augmentation. The rest of this article is organized as follows. Section II outlines the GAN architecture and its evolutions. Then, Section III describes ResBaGAN, detailing its different components. Thereafter, Section IV presents the experiments for evaluating ResBaGAN in terms of classification performance and robustness. Next, Section Vcarries out the discussion of this work. Finally, Section VI concludes this article. II. RELATED WORK A GAN [28] is a generative deep learning architecture that consists of the two networks represented in Fig. 1(a): a generator Gand a discriminator D. These networks are trained in an adversarial fashion: Glearns to synthesize samples as realistically as possible to deceive Dinto believing that they are real, while Dis trained to distinguish between real and fake samples. As a result of this opposition of objectives, the improvement of one network encourages the other to perform better. Both Dand Guse CNN architectures. Dis a standard CNN, while Greplaces conventional convolutions with transposed convolutions, which function inversely. Dneeds to transform a sample into a unidimensional feature vector of length Zfor the final classification. In contrast, Gaims to generate data akin to the input of D. To achieve this, it starts with a random vector of Zelements, often referred to as a latent vector, reshaped into a low-resolution sample. The size of this sample is gradually increased by applying transposed convolutions to produce a fake sample. Building upon the initial GAN architecture in Fig. 1(a), numerous optimizations have been proposed in recent years to furtherextenditscapabilities[37].Thefollowingareparticularly relevant to this work: 1) CGAN [38].In the original GAN, Gcannot synthesize samples for a specific class on demand. The CGAN architecture, shown in Fig. 1(b), addressed this limitation by incorporating information corresponding to the desired class in the latent vector. 2) ACGAN [39].A natural extension of CGAN is to enable Dtoassigneachsampletothebest-fittingclass,inaddition to discerning between real and fake samples. To this end, Din ACGAN has two outputs as shown in Fig. 1(c):a (a) (b) (c) (d) Fig. 1. Comparison between the initial GAN architecture and various evolutions. Note that BAGAN also introduced an autoencoder module to assist the main GAN module in parameter initialization; this autoencoder is not represented for brevity but can be seen in Fig. 2. (a) GAN. (b) CGAN. (c) ACGAN. (d) BAGAN. binary output for detecting fake samples and an output with as many elements as classes. 3) BAGAN [35].This architecture introduced several modifications to stabilize training with small datasets and effectively synthesize the minority classes. First, the networks’ parametersareinitializedbyanautoencoderthatprocesses
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6431 the dataset in advance, instead of from scratch. As a result, both Dand Gstart learning from an approximation of the data distribution, significantly stabilizing their training on small datasets. In addition, a more sophisticated sampling strategy for the latent vectors is introduced by modeling class-specific distributions in the latent space provided by the autoencoder, which further stabilizes the training of G. Moreover, the dual outputs of Dare combined into a single output, as shown in Fig. 1(d). By extension, the lossfunctionused inthetraining process becomessimpler, making it easier for Gto learn to synthesize the minority classes. As discussed in Section I, the evolving capabilities of GANs have established them as valuable tools for data augmentation in remote sensing applications. In particular, BAGAN has shown promise in different computer vision tasks to address the challenges of scarce and imbalanced datasets, two common limitations in remote sensing. However, to the best of our knowledge, the use of BAGAN in remote sensing has been limited to object detection, as demonstrated in Zhang et al. [40]. The method proposed in this article, ResBaGAN, is a GAN-based classification approach inspired by innovations in BAGAN, among other features. We not only aim to expand its usage in remote sensing, particularly for classification applied to environmental monitoring, butalso enhance its performance further. This is achieved by incorporating traditional augmentation techniques, superpixel segmentation-based sample extraction, and an improved ResNet-based classifier. III. PROPOSED METHOD ResBaGAN, the proposed method for remote sensing classification applied to environmental monitoring, is illustrated in Figs. 2and 3. This section describes its different components as follows. First, the GAN architecture designed for deep learning-based data augmentation is discussed in Section III-A. Next, the sample extraction procedure that combines superpixel segmentation and traditional augmentation techniques is introduced in Section III-B. Then, following these explanations, a complete step-by-step explanation of how ResBaGAN operates is provided. Finally, the design of the neural architectures in ResBaGAN is detailed in Section III-C. A. Network Architecture The core of ResBaGAN is a BAGAN-based architecture that includes an improved ResNet-based classifier [13], resulting in a combination that enhances both the quality of the GAN- based augmentation and the final classification accuracy. The integration of shortcuts in residual networks allows building much deeper architectures compared to stacking convolutional layers alone, circumventing training issues like gradient vanishing. This makes ResNet more suitable than shallow CNNs for classification with limited data, as it facilitates the extraction of moredetailedfeatures from the availablesamples.Moreover,the generalization capabilities of the designed ResNet architecture have been improved by fusing features from every convolutional stage into the final classification layers, instead of using features just from the last stage. Furthermore, the designed GAN module leverages two BAGAN features to assist the generator network in learning: autoencoder initialization and an improved loss function. This approach leads to the synthesis of more realistic samples when dealing with scarce and imbalanced datasets, thus providing an effective data augmentation for such situations. As shown in Fig. 2, three main elements can then be identified in the network architecture of ResBaGAN: 1) The input data are patches from the different classes, as displayed on the left side of the figure. The specialized sample extraction procedure provides this augmented training data and will be detailed in Section III-B. 2) ResBaGAN leverages an autoencoder module [41] to stabilize the subsequent training of the GAN on limited data. The autoencoder is depicted in the upper part of the figure. Autoencoders are easier to train than GANs under data scarcity, but they do not assign classes to the acquired knowledge, making them unsuitable for classification. However, they can help in initializing the weights of the GAN. The autoencoder must learn to compress and reconstruct samples as accurately as possible, encouraging it to acquire an initial understanding of the dataset. The autoencoderconsistsofanencoder(Einthefigure)andadecoder (Δ). As the encoding and decoding steps resemble the tasks of the discriminator (D) and the generator (G), respectively, weight-sharing by using the same network topologiesallowstransferringtheknowledgeofthetrained autoencoder to the uninitialized GAN, thus complementing each other’s strengths and weaknesses. 3) The main networks of ResBaGAN are Dand G, collectively referred to as the GAN module in the figure. Once initialized from the autoencoder, these networks are trained in an adversarial manner, so that Dlearns on both the real samples and the samples that Gsynthesizes. To facilitate Gthe learning of minority classes, Dcombines its real/fake prediction and class prediction in a singleoutput.This output contains N+1elements,where Ndenotes the number of possible classes, and samples suspected as fake are assigned the additional class. The specific network topologies of the Dand Gnetworks (and thus Eand Δ) in ResBaGAN will be detailed in Section III-C. Therefore, the ResBaGAN training process involves three steps: 1) Learning of the autoencoder module: The autoencoder trains on all of the available samples without considering their labels, to produce the initial understanding of the data distribution. L2 loss, shown as the dotted line at the top of the autoencoder, guides the training process. This loss is calculated as the mean squared difference between the input and the reconstructed samples, which the autoencoder aims to minimize. Given a reconstructed sample ˆyand its target y, the L2 loss is as follows: LL2=(ˆyi−yi)2.(1)
6432 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 Fig. 2. Neural network architecture of ResBaGAN. Its main features are a BAGAN-based design that incorporates an improved ResNet-based classifier (D). The input augmented training samples are obtained by using the extraction procedure presented in Fig. 3. 2) GAN initialization: All of the learned parameters are transferred from the autoencoder to the GAN networks, enabling them to acquire the knowledge of the autoencoder. This is shown in the rightmost line of the figure. 3) Learning of the GAN module: Owing to the use of the transferred weights as values, the training starts from a more stable point than randomly initialized parameters, as the GAN refines the initial understanding of data. Categoricalcross-entropyloss,representedbythedottedlineatthe bottom of the Fig. 2, guides this learning process. This loss is calculated as the difference between the predicted and the expected class-probability distributions for a sample. Given a predicted distribution ˆy, obtained by applying the softmaxfunctiontoD’srawoutput,andaone-hotencoded class target y, the cross-entropy loss is as follows: LCE =− N+1 c=1 yclog(ˆyc)(2) where ycdenotes the target probability of the sample belonging to class c.BothDand Gstrive to minimize their respective losses, LDand LG, which are based on the cross-entropy loss. As Dwants to accurately classify the real samples and assign the fake label to the samples from G, its loss can be split into a real part and a fake part: LD=Lreal D+Lfake D.(3) The real part is obtained by applying the cross-entropy loss between the predicted class probabilities of the real samples and their reference labels, and the fake part of the loss is obtained by applying the cross-entropy loss between the predicted probabilities of synthetic samples and the fake label. In contrast, G’s objective is to ensure that Dassigns its generated fake samples to the class that they intend to represent. To do so, the cross-entropy loss is applied between the predicted probabilities of the synthetic samples and their intended labels. Training proceeds by alternately optimizing the discriminator and the generator, which can be summarized in a minimax manner as follows: min Gmax DL(D,G) =Ex∼pdata(x),y ∼pdata(y)[log D(y|x)] +Ez∼pz(z),y ∼pdata(y)[log(1−D(N+1|G(z,y)))]. (4) The first term corresponds to the probability of all real samplesx, drawnfromthe pdata(x)data distribution,being correctly classified according to their labels y. The second term represents the probability of all synthetic samples being classified as the fake class N+1;zis a latent vector drawn from the latent space distribution pz(z), which is conditioned by the intended class yin G(z,y)to produce a synthetic sample. Once all of the neural networks in ResBaGAN have been trained, the fake output of Dis disabled, yielding the final classifier. B. Sample Extraction via Superpixel Segmentation and Traditional Augmentation To further improve ResBaGAN in handling scarce and imbalanced datasets, a specialized procedure for extracting samples from the dataset is introduced. This procedure leverages two extensively studied techniques in remote sensing: superpixel segmentationandtraditionaldataaugmentation.Fig.3illustrates
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6433 Fig. 3. Procedure for extracting training samples from the dataset in ResBa- GAN. A superpixel segmentation guides the process, and traditional augmentation techniques enrich the learning data that enters the network architecture in Fig. 2. The augmentation is not performed on samples destined for the validation or test sets. this procedure step by step, yielding the training samples fed to networks described in Section III-A. A superpixel segmentation groups similar pixels in an image into homogeneous and contiguous regions called superpixels. These regions have a relatively constant size but not a specific shape, as they adapt to the objects in the scene [42]. Thus, superpixels can be used as larger pixels for image processing and have been widely used in remote sensing classification [3], [26],[43],[44]. Instead of extracting a patch for each pixel in a sliding-window fashion [45], one patch per superpixel can be extracted, assigning the predicted class to all of its pixels. If the patch size is adjusted to fit within the average size of the superpixels, patches centered on the superpixels predominantly contain pixels of a single class. This approach helps avoid the Algorithm 1: Superpixel-Guided Sample Extraction. function EXTRACTPATCHES(dataset, N) Segment dataset into superpixels Initialize patches ←∅ for all superpixels with reference data Measure minimum enclosure Locate central point Extract a patch of N×Npixels Assign class via majority voting Add patch to patches end for return patches end function Algorithm 2: Traditional Augmentation on Training Samples. function AUGMENT(sample) Select random rotation from {0◦,90◦,180◦,270◦} Rotate sample if random() <0.5then Flip sample horizontally end if if random() <0.5then Flip sample vertically end if return sample end function function GETTRAININGBATCH(patches, batch_size) Initialize batch ←∅ while |batch|<batch_size do Select random patch from patches augmented_sample ←AU GMENT (patch) Add augmented_sample to batch end while Initialize fake_samples ←∅ Initialize samples_per_class ←batch_size classes while |fake_samples|< samples_per_class do Generate synthetic_sample with G Add synthetic_sample to fake_samples end while Add fake_samples to batch return batch end function noise introduced in learning data by patches containing multiple classes, particularly at the object edges, thus facilitating the modeling of the different classes for neural networks. On this basis, ResBaGAN extracts samples from the dataset byusing a superpixelsegmentationas guidance, as shownin Fig. 3and detailed in Algorithm 1. Moreover, traditional data augmentation techniques are applied to all of the training samples, as described in Algorithm 2. Regarding the superpixel segmentation algorithm, numerous options exist in the literature, such as SLIC [46], ETPS [47], and
6434 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 Algorithm 3: ResBaGAN Classification Method. function RESBAGAN(target_dataset) patches ←SAMPLEEXTRACTION(target_dataset) TRAINING(patches) CLASSIFICATION(patches) end function function SAMPLEEXTRACTION(target_dataset) Segment target_dataset using WP Extract a sample per superpixel with reference data Split samples into train, validation, and test end function function TRAINING(patches) Train autoencoder on augmented train samples to learn data distribution Transfer autoencoder parameters to GAN module Train GAN on augmented train samples and synthetic samples from G end function function CLASSIFICATION(patches) Disable additional fake class in D’s output Map target dataset using Das the classifier end function SEEDS [48]. Typically, these algorithms are based on gradient descent or graphs [42], with the latter being more computationally expensive and lacking control of the segment size or regularity [46].WP[49] is a popular gradient descent algorithm thatprovidesgood-qualitysegmentationswithmoderatecompu- tationalconsumption[42],[46],andallowsforthecustomization of the segment size and regularity. Therefore, WP is chosen as the superpixel segmentation algorithm for ResBaGAN. As all ResBaGAN components are explained, its step-by-step operation can be summarized as Algorithm 3. C. Network Topologies As explained in Section III-A, the autoencoder and the GAN modules share network topologies. First, the topology of the residual discriminator (Din Fig. 2) will be discussed, which is also applicable to the encoder (E). Then, the design of the generator(G) will follow, which is alsoapplicable to the decoder (Δ). The ResBaGAN classifier (D) features a ResNet-based topology, with the novelty of integrating data from different levels of abstraction to enhance generalization, as it will be detailed in the following paragraph. This architecture is depicted in Fig. 4. When compared to shallow CNNs, a residual approach allows for more efficient use of the available learning data and pushes the capabilities of Gfurther owing to the adversarial nature of GANs. More specifically, the classifier consists of three residual stages (colored differently in the figure), each containing three residualblocks that use the same number ofconvolutionalfilters. Fig. 4. Diagram of the classifier (D) for ResBaGAN. The output dimensions correspond to an input sample with dimensions H×W×B=32×32 ×B. A different color is assigned to each residual stage. For compactness, the figure does not elaborate either on the convolutional layers that adapt the input of each stage to its dimensionality, or on the layers that adapt the output of the first two stages to the final feature fusion.
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6435 TABLE I DETAILS OF THE LAYERS IN THE CLASSIFIER (D)FOR RESBAGAN The growing filters allow for extracting increasingly complex visual features. Additional convolutional layers precede each stage to ensure appropriate input dimensionality, and the first convolutional layer in each stage has a stride of 2 for learned dimensionality reduction. In contrast to the original ResNet architecture, this improved design fuses the features extracted from the last stage with those derived from all preceding stages, by adding the corresponding feature maps. This approach merges data from the low-, mid-, and high-level abstractions into the input for the final layers, leading to better generalization. Two additional convolutional layers map the dimensionality of the outputs from Stage 1 and Stage 2 to Stage 3 to perform the addition. A global average-pooling layer transforms the fused feature maps into a vector of features processed by a fully connected layer for the final classification. The details of the classifier’s layers are presented in Table I. For brevity, the table omits the convolutional layers adapting the input dimensionality between stages and those adapting the outputfromthefirsttwostagesforfeaturefusion;theyhave3×3 filters and use the same activation function as the other layers. As usual in residual networks, a dropout layer follows each convolutional layer to stabilize training and address problems such as overfitting [50]. The selected dropout probability and activation function will be explained in Section IV-A4. Spectral normalization [51] is also applied to all convolutional layers, as it commonly improves GAN networks. In the design of ResBaGAN, achieving an equilibrium between Gand Dis crucial for stable learning and thus, sustained adversarial behavior. This balance enables both networks to progress together, rather than allowing one to outperform the other significantly, disrupting the adversarial learning and impeding their ability to compete. Such disruption would prevent Gfrom learning to perform a deep learning-based augmentation that is useful to D. Thus, a standard GAN generator has been found to be an ideal counterpart for the residual classifier in ResBaGAN. AsdepictedinFig.5,thegeneratorconsistsofaseriesoftransposed convolutional layers. The initial embedding layer [52] transforms the desired class into a Z-element vector, combined Fig. 5. Diagram of the generator (G) for ResBaGAN. The output dimensions correspond to an input latent vector with Zelements. with the latent vector via a dot product operation. The result, interpreted as a patch with the dimensions N×N×B=1× 1×Z, is supplied to the transposed convolutions to generate the synthetic sample by gradually increasing its resolution. The details of the generator’s layers are presented in Table II. Notably, the choice of Zcan significantly impact the synthesis performance. Short latent vectors may struggle to create realistic samples, whereas large vectors could be too difficult to handle effectively. Hence, the impact of Zwill be examined in Section IV-A4. The activation function also follows the rationale in Section IV-A4. The only exception is the activation of the last layer, which uses a tangent function, as all the dataset is scaled to the [−1,1] range before being input to ResBaGAN. Spectral normalization is also applied to all of the convolutional layers.
6436 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 TABLE II DETAILS OF THE LAYERS IN THE GENERATOR (G)FOR RESBAGAN TABLE III DESCRIPTIONS OF THE DATASETS USED IN THIS WORK IV. EXPERIMENTS In this section, ResBaGAN is evaluated in terms of classification performance to assess both its general effectiveness and the contributions of its individual components, such as the synthesis quality of the GAN-based augmentation. ResBaGAN’s capabilities are compared with simpler classification approaches, such as standalone CNNs and other GAN-based approaches like ACGAN. The section is organized as follows. First, the datasets, metrics, and experimental environment used for the assessment are outlined in Section IV-A. A range of design choices that were left open in Section III are also addressed, such as the selection of activation functions, the optimizer, and the number of training epochs, among other factors. Then, the experimental results are presented and analyzed in Section IV-B. A. Experimental Setup 1) Datasets: Eight large, very high-resolution multispectral images of natural regions with dense vegetation were used [3]. These images were captured in 2018, 2019, and 2020, by flying a UAV at a 120 m altitude over various river basins in Galicia, Spain, resulting in a spatial resolution of 10 cm/px. The UAV carried a MicaSense RedEdge-MX multispectral camera, capturing five spectral bands corresponding to wavelengths of 475 nm (blue), 560 nm (green), 668 nm (red), 717 nm (red edge), and 842 nm (near-infrared). Table III details the specific locations and dimensions of the scenes. The composite color image and the reference data for each dataset can be found in Fig. 6. Table IV enumerates the ten identifiable classes in the reference data, detailing the number of samples in each dataset. These classes range from native vegetation to human-made structures such as roads or buildings. It is important to highlight the strong imbalances between classes in all datasets, with minority classes having multiple orders of magnitude fewer samples at times. This imbalance introduces a bias toward majority classes, potentially preventing a balanced classification accuracy across the different classes. All of the datasets were segmented using the WP algorithm by choosing an average size of 400 px/superpixel, allowing a minimum size of 100 px/superpixel, and employing a compactness factor of 0.5 points, following the approach in [3].The extracted patches had spatial dimensions of N×N=32 ×32 px. In addition, all the data were normalized to the [−1,1] range. For a given dataset, all associated experiments used the same subset of 15% training samples, and 5% validation samples to monitor the training progress by identifying potential issues like overfitting. This limited availability of learning data is enforced
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6443 TABLE VIII RECORDED PER-CLASS CLASSIFICATION ACCURACIES FOR ALL THE CLASSES USING THE CLASSIFIERS DEVELOPED IN THIS WORK TABLE IX AGGREGATION OF ALL RECORDED CONFUSION MATRICES FOR RESBAGAN ACROSS ALL TEN EXPERIMENTS FOR EACH DATASET In contrast, the enhancements in BAGAN provided an effective GAN-based augmentation for the CNN, significantly increasing all metrics to match the performance of the improved ResNet. BAGAN and the ResNet showed varied performance, with one method slightly outperforming the other at times, as summarized by the F1 score. Interestingly, the improved ResNet tended to improve PA the most, achieving better values than BAGAN in six out of ten classes, while BAGAN was more effective in increasing UA, with better values in six out of ten classes. Finally, ResBaGAN further improved performance across all classes, achieving the best PA, UA, and F1 scores in almost every case. These enhancements were particularly noticeable in the minority classes. For example, concrete is consistently one of the scarcest classes, and ResBaGAN boosted its F1 by 22%. The solid performance of ResBaGAN when facing scarce and imbalanced data is further exemplified by the confidence intervals for F1 scores. Even at its lowest potential, ResBaGAN consistently obtains competitive results compared to the improved ResNet and BAGAN. This behavior is further reiterated in the aggregated confusion matrix for ResBaGAN across all ten experiments for each dataset, depicted in Table IX. The strong diagonal dominance indicates high classification accuracy across all classes.
6444 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 Fig. 13. Recorded F1 scores for all the classes using the classification approaches developed in this work, averaging results across datasets. The bars represent the mean values along their confidence intervals. Higher values are better. 5) Ablation Study: To assess the importance of the various components of ResBaGAN to its final performance, further experiments were conducted by removing one component at a time and comparing the resulting classification accuracy to the baseline ResBaGAN. These additional experiments were run on the Eiras Dam dataset, with reported OA, AA, and κ. Specifically, the following four alterations were examined: rRemoving the BAGAN innovations, namely the autoencoder and the improved loss function, transforming Res- BaGAN into an ACGAN-derived architecture. rReplacingtheresidualtopologyinDwiththeshallowCNN described in Table V. rRemoving the traditional augmentation techniques applied to training samples. rRemoving the superpixel segmentation-guidance from the sample extraction procedure, to instead extract one patch per pixel in a sliding window manner. The training used the same number of samples as with superpixels. The results of the ablation experiments are represented in Fig. 14. It is evident that removing any of ResBaGAN’s components significantly impacted its performance, as the potential maximum accuracies are always lower than, or comparable at most, to the minimum expected accuracies in the baseline ResBaGAN. Fig. 14. Recorded OA, AA, and κmetrics for the Eiras Dam dataset in the ablation study of ResBaGAN. The bars represent the mean values along their confidence intervals. Higher values are better. TABLE X COMPUTATIONAL COSTOFTHEDEEP LEARNING METHODS USED IN THIS WORK.THE RESULTS ARE PRESENTED AS AVERAGED SPEEDUP VALUES FOR TRAINING AND CLASSIFICATION ACROSS ALL EXPERIMENTS,RELATIVE TO THE SHALLOW CNN. HIGHER VALUES ARE BETTER Interestingly, AA was the most affected accuracy metric. Given that it assigns equal weights to minority and majority classes, this observation highlights the particular importance of all ResBaGAN components in developing an effective data augmentation framework to adequately assist the classifier in learning minority classes with limited data. 6) Computational Cost: Finally, the computational cost of the various deep learning methods tested in this work is compared. The cost is presented in Table Xas averaged speedup values for the training and classification phases across all experiments, using the shallow CNN as a baseline. First, the improved ResNet required approximately twice as much computation time as the shallow CNN during both phases. This is expected due to the much larger neural topology. In contrast, ACGAN also required twice as much training time compared to the CNN, as it uses this network as Dand a companion Gnetwork of similar complexity. However, the classification time for ACGAN was only slightly higher than the CNN, as Gdid not play a role in this phase. We attribute the minor slowdown to the PyTorch overhead of having both networks on the GPU. BAGAN further increased the training time by requiring the learning of the autoencoder before the GAN networks themselves. However, as the autoencoder did not play a role during classification, the inference time remained roughly the same as that of ACGAN. Finally, ResBaGAN exhibits increased complexity due to combining features from both the improved ResNet and the BAGAN, in addition to the traditional augmentation techniques, resulting in the longest training times, and slightly slower classification compared to the improved ResNet. We attribute this slight difference again to the PyTorch overhead of having multiple networks on the GPU.
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6445 V. DISCUSSION Enhancing the accuracy of remote sensing image classification is crucial for numerous monitoring tasks of the Earth’s surface. Although deep learning techniques are known for their excellent classification performance, they usually require large amounts of data to unleash their full potential. This can be especially challenging for remote sensing applications, as datasets often have limited labeled data and class imbalances [4].Data augmentation techniques, particularly those based on GAN architectures [28], present a promising and powerful solution for generating rich additional learning data [29], as demonstrated in this work. BAGAN [35] is a particularly interesting GAN architecture, as it is specifically designed to address the aforementioned limitations. Despite its potential, to the best of our knowledge, the application of BAGANs in remote sensing has been limited to object detection [40] until now. This study shows that BAGAN-inspired architectures outperform more common GAN approaches like ACGAN [39] in providing effective deep learning-based data augmentation, particularly when working with limited and imbalanced multispectral datasets. Consequently, BAGAN-inspired architectures are proposed as a new baseline for researchers working with GANs in remote sensing. Another crucial finding of this study is that integrating additional techniques can further improve the quality of the GAN-based data augmentation and the final classificationaccuracy, as demonstrated bythe proposedmethod ResBaGAN, which improves upon BAGAN. ResBaGAN’sadaptability presentsasignificant advantagefor thefuture,asitcanbeeasilyintegratedwithclassifiersfromother works to increase their accuracy. To achieve this, the classifier only needs to be integrated as the Dnetwork within ResBaGAN, with adjustments made to the design of the Gnetwork to match the capabilities of both modules, as illustrated throughout this work. The primary limitation of ResBaGAN is its long training time compared to other deep learning approaches. However, it presents a good tradeoff between cost and performance. As shown in the experiments in this work, none of the other classification methods achieved 80% for the F1 score across all classes. It is also worth noting that the classification cost of Res- BaGANmainly depends on the computationalcost of the chosen classifier, so it could be reduced if the classifier is replaced by one with lower complexity. One additional advantage of ResBaGAN is that applying the superpixel segmentation to the image greatly reduces the computational cost of classification compared to a pixel-focused approach. This features makes ResBaGAN well-suited for rapidly mapping extensive terrain areasinrealapplicationsinvolvingremotesensingclassification. Another significant contribution of this work is the successful adaptation of the FID score metric to assess the quality of GAN-based synthesis while using the full spectral resolution of remote sensing data. This modification is necessary as there are currently no standard models for evaluating FID on remote sensing data, similar to the Inception v3 network used for the RGB data. This adaption allows for more accurate comparisons among GAN-based networks and can be easily included in other research works. Nonetheless, it would be valuable to explore the development of standard models for evaluating the FID score on remote sensing data. Various aspects of ResBaGAN open up future workdirections to improve its performance. For instance, the BAGAN-based designs in this work do not feature the sophisticated latent vector sampling strategy, which could further ease working with limited data. Another possibility for improvement could involve integrating state-of-the-art deep learning architectures like GCN [18] and transformers [19] as classifiers in ResBa- GAN. Moreover, other traditional augmentation techniques not considered in this analysis, such as those described in [25] or [26], could be integrated into the proposed solution. Regarding the field of application, this study focused on applyingResBaGANtomappingimagesofforestareas.Itwould alsobe interestingtoevaluateResBaGANonnewdatasetsto further explore its strengths and limitations, whether in vegetation terrains or in other types of images. Moreover, to evaluate the robustness of ResBaGAN, it would be valuable to analyze its performance using datasets captured under variable illumination and atmospheric conditions, or including noise produced by the sensor. As ResBaGAN extracts patches centered on superpixels and works with the whole spectral dimensionality of the image, it is expected that it could handle the intraclass spectralvariabilityproducedbythesevaryingconditionstosome extent. The solution proposed in this work operates over very highresolution images. However, ResBaGAN could be adapted to operate over remote sensing images with different spectral and spatial resolutions by adapting its stages. For instance, modifying patch and superpixel size parameters would be necessary for dealing with datasets with different spatial resolutions. For coarser grained datasets, such as satellite images, employing spectral unmixing techniques [76] might be necessary for situations where multiple, different elements are captured within the same pixel. VI. CONCLUSION This work proposes ResBaGAN, a deep-learning method for the classification of remote sensing images, designed to overcome the prevalent challenges of data scarcity and class imbalances. This is achieved by constructing an advanced data augmentation framework that combines BAGAN-inspired augmentation, sample extraction on top of traditional augmentation techniques and superpixel segmentation, and an improved ResNet-based classifier. This integrated approach maximizes the usage of spectral and spatial information from the available training data to effectively overcome the aforementioned issues. Moreover, the superpixel segmentation also makes ResBaGAN suitable for recognizing large terrain areas. To the best of the authors’ knowledge, this is the first study that applies a BAGAN-based architecture to remote sensing classification. In addition, this work also proposes an adaptation of the FID score metric for evaluating the synthesis quality of GANs in multi- and hyperspectral remote sensing images.
6446 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 16, 2023 ResBaGAN’s performance was evaluated in the context of environmental monitoring by using eight large, very highresolution multispectral images of riparian forests, with limited learningdataandstrongclassimbalances.Acomparisontoother state-of-the-art classification methods in the remote sensing field, such as ACGAN, was carried out. The results revealed the effectiveness of ResBaGAN in achieving high classification accuracies in constrained datasets with improved performance across all classes of elements, particularly in minority ones. Finally, numerous interesting research directions have been identified for future work. For instance, it would be valuable to assess the performance of ResBaGAN on new datasets featuring diverse spatial and spectral resolutions. Additionally, exploring the integration of more advanced components into ResBa- GAN could yield further improvements, such as substituting the traditional augmentation techniques with more sophisticated methods or replacing the residual classifier with cutting-edge architectures like transformers. REFERENCES [1] Z. Shao, P. Tang, Z. Wang, N. Saleem, S. Yam, and C. Sommai, “BRRNet: A fully convolutional neural network for automatic building extraction from high-resolution remote sensing images,” Remote Sens., vol. 12, no. 6, 2020, Art. no. 1050. [2] B. Lu, P. D. Dao, J. Liu, Y. He, and J. Shang, “Recent advances of hyperspectralimaging technologyandapplications inagriculture,” Remote Sens., vol. 12, no. 16, 2020, Art. no. 2659. [3] F.Argüello,D.B.Heras,A.S.Garea,andP.Quesada-Barriuso,“Watershed monitoring in Galicia from UAV multispectral imagery using advanced texture methods,” Remote Sens., vol. 13, no. 14, 2021, Art. no. 2687. [4] A. E. Maxwell, T. A. Warner, and F. Fang, “Implementation of machinelearningclassificationinremotesensing:Anappliedreview,”Int.J.Remote Sens., vol. 39, no. 9, pp. 2784–2817, 2018. [5] B. A. Bradley, “Remote detection of invasive plants: A review of spectral, textural and phenological approaches,” Biol. Invasions, vol. 16, pp. 1411–1425, Jul. 2014. [6] N. Y. Zhu et al., “Deep learning for smart agriculture: Concepts, tools, applications, and opportunities,” Int. J. Agricultural Biol. Eng., vol. 11, no. 4, pp. 32–44, 2018. [7] Q. Yuan et al., “Deep learning in environmental remote sensing: Achievements and challenges,” Remote Sens. Environ., vol. 241, 2020, Art. no. 111716. [8] G. Morales, G. Kemper, G. Sevillano, D. Arteaga, I. Ortega, and J. Telles, “Automatic segmentation of mauritia flexuosa in unmanned aerial vehicle (UAV) imagery using deep learning,” Forests, vol. 9, no. 12, 2018, Art. no. 736. [9] Z. M. Hamdi, M. Brandmeier, and C. Straub, “Forest damage assessment using deep learning on high resolution remote sensing data,” Remote Sens., vol. 11, no. 17, 2019, Art. no. 1976. [10] K. Isaienkov, M. Yushchuk, V. Khramtsov, and O. Seliverstov, “Deep learning for regular change detection in Ukrainian forest ecosystem with Sentinel-2,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 14, pp. 364–376, 2021. [11] C.-Y. Chiang, C. Barnes, P. Angelov, and R. Jiang, “Deep learning-based automatedforest healthdiagnosis fromaerialimages,”IEEEAccess,vol.8, pp. 144064–144076, 2020. [12] Y. Chen, H. Jiang, C. Li, X. Jia, and P. Ghamisi, “Deep feature extraction and classification of hyperspectral images based on convolutional neural networks,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 10, pp. 6232–6251, Oct. 2016. [13] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 770–778. [14] H. Lee and H. Kwon, “Going deeper with contextual CNN for hyperspectral image classification,” IEEE Trans. Image Process., vol. 26, no. 10, pp. 4843–4855, Oct. 2017. [15] Z. Zhong, J. Li, Z. Luo, and M. Chapman, “Spectral–spatial residual networkforhyperspectralimageclassification:A 3-Ddeeplearningframe- work,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 2, pp. 847–858, Feb. 2018. [16] M.E.Paoletti,J.M.Haut,R.Fernandez-Beltran,J.Plaza, A.J.Plaza,andF. Pla, “Deep pyramidal residual networks for spectral–spatial hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 57, no. 2, pp. 740–754, Feb. 2019. [17] R. Li, S. Zheng, C. Duan, Y. Yang, and X. Wang, “Classification of hyperspectral image based on double-branch dualattention mechanism network,” Remote Sens., vol. 12, no. 3, 2020, Art. no. 582. [18] D.Hong, L.Gao,J. Yao,B. Zhang,A.Plaza,andJ. Chanussot,“Graphconvolutional networks for hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 59, no. 7, pp. 5966–5978, Jul. 2021. [19] Q. He, X. Sun, W. Diao, Z. Yan, D. Yin, and K. Fu, “Transformer-induced graphreasoning formultimodalsemanticsegmentationinremotesensing,” ISPRS J. Photogrammetry Remote Sens., vol. 193, pp. 90–103, 2022. [20] A. Vaswani et al., “Attention is all you need,” in Proc. 31st Int. Conf. Neural Inf. Process. Syst., 2017, pp. 6000–6010. [21] D. Hong et al., “More diverse means better: Multimodal deep learning meets remote-sensing imagery classification,” IEEE Trans. Geosci. Remote Sens., vol. 59, no. 5, pp. 4340–4354, May 2021. [22] X. Wu, D. Hong, and J. Chanussot, “Convolutional neural networks for multimodal remote sensing data classification,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022, Art. no. 5517010. [23] X. Sun et al., “Ringmo: A remote sensing foundation model with masked image modeling,” IEEE Trans. Geosci. Remote Sens., 2022, to be published, doi: 10.1109/TGRS.2022.3194732. [24] C.Shortenand T.M. Khoshgoftaar,“Asurveyonimagedata augmentation for deep learning,” J. Big Data, vol. 6, Jul. 2019, Art. no. 60. [25] J. M. Haut, M. E. Paoletti, J. Plaza, A. Plaza, and J. Li, “Hyperspectral image classification using random occlusion data augmentation,” IEEE Geosci. Remote Sens. Lett., vol. 16, no. 11, pp. 1751–1755, Nov. 2019. [26] A. Acción, F. Argüello, and D. B. Heras, “Dual-window superpixel data augmentation for hyperspectral image classification,” Appl. Sci., vol. 10, no. 24, 2020, Art. no. 8833. [27] J. Nalepa, M. Myller, and M. Kawulok, “Training- and test-time data augmentation for hyperspectral image segmentation,” IEEE Geosci. Remote Sens. Lett., vol. 17, no. 2, pp. 292–296, Feb. 2020. [28] I. J. Goodfellow et al., “Generative adversarial networks,” CoRR, vol. abs/1406.2661, 2014. [Online]. Available: http://arxiv.org/abs/1406. 2661 [29] M. Frid-Adar, I. Diamant, E. Klang, M. Amitai, J. Goldberger, and H. Greenspan, “GAN-based synthetic medical image augmentation for increased CNN performance in liver lesion classification,” Neurocomputing, vol. 321, pp. 321–331, 2018. [30] G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of datawithneuralnetworks,”Science,vol.313,no.5786,pp. 504–507,2006. [31] D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Proc. 2ndInt.Conf.Learn.Representations,Y.BengioandY.LeCun,Eds.Banff, AB, Canada, Apr. 14-16, 2014. [Online]. Available: http://arxiv.org/abs/ 1312.6114 [32] A. Shashank, V. V. Sajithvariyar, V. Sowmya, K. P. Soman, R. Sivanpillai, and G. K. Brown, “Identifying epiphytes in drones photos with a conditional generative adversarial network (C-GAN),” Int. Arch. Photogrammetry, Remote Sens. Spatial Inf. Sci., vol. 99, pp. 99–104, 2020. [33] L.Zhu,Y. Chen, P. Ghamisi,andJ. A. Benediktsson,“Generativeadversarial networks for hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 9, pp. 5046–5063, Sep. 2018. [34] M. Lucic, M. Tschannen, M. Ritter, X. Zhai, O. Bachem, and S. Gelly, “High-fidelity image generation with fewer labels,” in Proc. Int. Conf. Mach. Learn., 2019, pp. 4183–4192. [35] G.Mariani,F.Scheidegger,R. Istrate,C.Bekas,andC.Malossi,“BAGAN: Data augmentation with Balancing GAN,” 2018, arXiv:1803.09655. [36] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local Nash equilibrium,” in Proc. Adv. Neural Inf. Process. Syst., vol. 30, 2017, pp. 6629–6640. [37] A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, “Generative adversarial networks: An overview,” IEEE Signal Process. Mag., vol. 35, no. 1, pp. 53–65, Jan. 2018. [38] M.MirzaandS.Osindero,“Conditionalgenerativeadversarialnets,”2014, arXiv:1411.1784.
DIESTE et al.: RESBAGAN: A RESIDUAL BALANCING GAN WITH DATA AUGMENTATION FOR FOREST MAPPING 6447 [39] A. Odena, C. Olah, and J. Shlens, “Conditional image synthesis with auxiliary classifier GANs,” in Proc. Int. Conf. Mach. Learn., 2017, pp. 2642–2651. [40] Y. Zhang, X. Liu, S. Wa, S. Chen, and Q. Ma, “GANsformer: A detection network for aerial images with high performance combining convolutional network and transformer,” Remote Sens., vol. 14, no. 4, 2022, Art. no. 923. [41] M. A. Kramer, “Nonlinear principal component analysis using autoassociative neural networks,” AIChE J., vol. 37, no. 2, pp. 233–243, 1991. [42] D. Stutz, A. Hermans, and B. Leibe, “Superpixels: An evaluation of the state-of-the-art,” Comput. Vis. Image Understanding, vol. 166, pp. 1–27, 2018. [43] P. G. Bascoy, A. S. Garea, D. B. Heras, F. Argüello, and A. Ordóñez, “Texture-based analysis of hydrographical basins with multispectral imagery,” in Proc. SPIE, 2019, vol. 11149, pp. 225–234. [44] S. R. Blanco, D. B. Heras, and F. Argüello, “Texture extraction techniques for the classification of vegetation species in hyperspectral imagery: Bag of words approach based on superpixels,” Remote Sens., vol. 12, no. 16, 2020, Art. no. 2633. [45] L. Zhang, L. Zhang, and B. Du, “Deep learning for remote sensing data: A technical tutorial on the state of the art,” IEEE Geosci. Remote Sens. Mag., vol. 4, no. 2, pp. 22–40, Jun. 2016. [46] R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Süsstrunk, “SLIC superpixels compared to state-of-the-art superpixel methods,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 11, pp. 2274–2282, Nov. 2012. [47] J. Yao, M. Boben, S. Fidler, and R. Urtasun, “Real-time coarse-to-fine topologically preserving segmentation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2015, pp. 2947–2955. [48] M. Van den Bergh, X. Boix, G. Roig, B. de Capitani, and L. Van Gool, “SEEDS: Superpixels extracted via energy-driven sampling,” in Proc. Eur. Conf. Comput. Vis., 2012, pp. 13–26. [49] V. Machairas, M. Faessel, D. Cárdenas-Peña, T. Chabardes, T. Walter, and E. Decencière, “Waterpixels,” IEEE Trans. Image Process., vol. 24, no. 11, pp. 3707–3716, Nov. 2015. [50] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res., vol. 15, no. 56, pp. 1929–1958, 2014. [51] T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, “Spectral normalization for generative adversarial networks,” 2018, arXiv:1802.05957. [52] D. Jurafsky and J. H. Martin, Speech and Language Processing–An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, 2nd Edition, Prentice Hall, Pearson Education International, 2009. [Online]. Available: https://www.worldcat.org/oclc/ 315913020 [53] R. G. Congalton, “A review of assessing the accuracy of classifications of remotely sensed data,” Remote Sens. Environ., vol. 37, no. 1, pp. 35–46, 1991. [54] Y. Xu, W. Yu, P. Ghamisi, M. Kopp, and S. Hochreiter, “Txt2Img–MHN: Remote sensing image generation from text using modern Hopfield networks,” 2022, arXiv:2208.04441. [55] P. Shamsolmoali, M. Zareapoor, H. Zhou, R. Wang, and J. Yang, “Road segmentation for remote sensing images using adversarial spatial pyramid networks,” IEEE Trans. Geosci. Remote Sens., vol. 59, no. 6, pp. 4673–4688, Jun. 2021. [56] AnandTech, “Intel Core i7–11700 K review: Blasting off with Rocket Lake,” Accessed on: Jun. 11, 2022. [Online]. Available: https://www.anandtech.com/show/16535/intel-core-i7-11700k-review- blasting-off-with-rocket-lake [57] TechPowerUp, “NVIDIA GeForce RTX 3080 Ti specs—TechPowerUp GPU database,” Accessed on: Jun. 11, 2022. [Online]. Available: https: //www.techpowerup.com/gpu-specs/geforce-rtx-3080-ti.c3735 [58] Canonical Ltd., “Enterprise open source and Linux—ubuntu,” Accessed: Jun. 11, 2022. [Online]. Available: https://ubuntu.com/ [59] Python software foundation, “Welcome to Python.org,” Accessed on: Jun. 11, 2022. [Online]. Available: https://www.python.org/ [60] A. Paszke et al., “PyTorch: An imperative style, high-performance deep learning library,” in Proc. Int. Conf. Adv. Neural Inf. Process. Syst., 2019, vol. 32, pp. 8024–8035. [61] NVIDIA Corporation, “CUDA zone–library of resources—NVIDIA developer,” Accessed on: Jun. 11, 2022. [Online]. Available: https:// developer.nvidia.com/cuda-zone [62] S. Chetlur et al., “cuDNN: Efficient primitives for deep learning,” CoRR, vol. abs/1410.0759, 2014. [Online]. Available: http://arxiv.org/abs/1410. 0759 [63] JTC1/SC22/WG14,“ISO/IEC 9899,”Standard,ISO/IEC, 2018.Accessed: Jun. 11, 2022. [Online]. Available: https://www.open-std.org/jtc1/sc22/ wg14/ [64] Standard C++ Foundation, “Standard C++,” Accessed on: Jun. 11, 2022. [Online]. Available: https://isocpp.org/ [65] Free Software Foundation, Inc., “GCC, the GNU compiler collection– GNU project,” Accessed on: Jun. 11, 2022. [Online]. Available: https: //gcc.gnu.org/ [66] TensorHub, Inc., “GuildAI–Experiment tracking, ML developer tools,” Accessed: Jun. 13, 2022. [Online]. Available: https://guild.ai/ [67] D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (ELUs),” 2015, arXiv:1511.07289. [68] A. L. Maas, “Rectifier nonlinearities improve neural network acoustic models,” in Proc. 30th Int. Conf. Mach. Learn., vol. 28, no. 1, p.3, 2013. [69] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassinghuman-levelperformanceon ImageNetclassification,” inProc. IEEE Int. Conf. Comput. Vis., 2015, pp. 1026–1034. [70] V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmannmachines,”in Proc. 27thInt. Conf.Mach.Learn.,2010, pp.807–814. [71] D. P. Kingma and J. Ba, “ADAM: A method for stochastic optimization,” in Proc. 3rd Int. Conf. Learn. Representations, San Diego, CA, USA, May 7-9, 2015. [Online]. Available: http://arxiv.org/abs/1412.6980 [72] A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” 2015, arXiv:1511.06434. [73] X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. 13th Int. Conf. Artif. Intell. Statist., 2010, vol. 9, pp. 249–256. [74] J. Sánchez, F. Perronnin, T. Mensink, and J. Verbeek, “Image classification with the fisher vector: Theory and practice,” Int. J. Comput. Vis., vol. 105, pp. 222–245, Dec. 2013. [75] G.-B. Huang, “An insight into extreme learning machines: Random neurons, random features and kernels,” Cogn. Comput., vol. 6, pp. 376–390, Sep. 2014. [76] D. Hong, N. Yokoya, J. Chanussot, and X. X. Zhu, “An augmented linear mixing model to address spectral variability for hyperspectral unmixing,” IEEE Trans. Image Process., vol. 28, no. 4, pp. 1923–1938, Apr. 2019. Álvaro G. Dieste received the B.S. degree in computer engineering and the M.S. degree in high performance computing in 2021 and 2022, respectively, from the University of Santiago de Compostela, Santiago, Spain, where he is currently working toward the Ph.D. degree in computer science. He is currently an Assistant Researcher with the Centro Singular de Investigación en Tecnoloxías Intelixentes, University of Santiago de Compostela. His research interests focus on computer vision tasks, whilebringingtogetherknowledgefromdiversecomputing areas, such as artificial intelligence, high performance computing, and big data. Francisco Argüello received the B.S. and Ph.D. degrees in physics from the University of Santiago de Compostela, Santiago, Spain, in 1988 and 1992, respectively. Heiscurrently an AssociateProfessorwith the Department of Electronics and Computer Engineering, University of Santiago de Compostela. His research interests include signal and image processing, computer graphics, parallel and distributed computing, and quantum computing. Dora B. Heras (Member, IEEE) received the M.S. and Ph.D. degrees in physics from the University of SantiagodeCompostela,Santiago,Spain,in1994and 2000, respectively. She is currently an Associate Professor with the Department of Electronics and Computer Engineering, University of Santiago de Compostela. Her research interests cover a range of topics in the combined fields of image processing, remote sensing, machine learning, and high performance computing. In particular, she has published papers on high performance computing, registration, classification, domain adaptation, and change detection applied to remotely sensed images.