scieee AI-readable full text Open interactive document viewer

EffBaGAN: an efficient balancing GAN for Earth observation in data scarcity scenarios

Vilela Pérez, Nicolás; Blanco Heras, Dora; Argüello Pedreira, Francisco

Abstract

Generative Adversarial Networks (GAN) can be used as a data augmentation technique in scenarios with limited labeled information and class imbalances, common issues in remote sensing datasets. The EfficientNet architecture has gained attention for achieving high accuracy with moderate computational cost. This work introduces EffBaGAN, a generative network specifically designed for the classification of multispectral remote sensing images based on EfficientNet, addressing data scarcity and class imbalances while minimizing network complexity. EffBaGAN is built upon a BAGAN architecture, incorporating a custom EfficientNet-based discriminator and generator. In particular, for the discriminator we propose RedEffDis, a reduced version of EfficientNet-B0 adapted to multispectral imagery. The generator, ResEffGen, includes a residual EfficientNet-based path, which enhances the quality of the generated synthetic samples. Additionally, a superpixel-based sample extraction procedure is used to further reduce the computational cost of the method. Experiments were conducted on large, very high-resolution multispectral images of vegetation, demonstrating that EffBaGAN achieves higher accuracy than other advanced classification methods, including vision transformers and residual BAGAN, while maintaining a significantly lower computational cost. In fact, EffBaGAN is more than twice as fast as the residual BAGAN, making it an efficient solution for remote sensing image classification in data-scarce environments.

Full text

IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 2477 EffBaGAN: An Efficient Balancing GAN for Earth Observation in Data Scarcity Scenarios Nicolás Vilela-Pérez ,DoraB.Heras , Member, IEEE, and Francisco Argüello Abstract—Generative adversarial networks (GANs) can be used as a data augmentation technique in scenarios with limited labeled information and class imbalances, common issues in remote sensing datasets. The EfficientNet architecture has gained attention for achieving high accuracy with moderate computational cost. This work introduces efficient balancing generative adversarial network (EffBaGAN), a generative network specifically designed for the classification of multispectral remote sensing images based on EfficientNet, addressing data scarcity and class imbalances while minimizing network complexity. EffBaGAN is built upon a balancing generative adversarial network (BAGAN) architecture, incorporating a custom EfficientNet-based discriminator and generator. In particular, for the discriminator we propose reduced EfficientNet discriminator, a reduced version of EfficientNet-B0 adapted to multispectral imagery. The generator, residual EfficientNet generator, includes a residual EfficientNet-based path, which enhances the quality of the generated synthetic samples. In addition, a superpixel-based sample extraction procedure is used to further reduce the computational cost of the method. Experiments were conducted on large, very high-resolution multispectral images of vegetation, demonstrating that EffBaGAN achieves higher accuracy than other advanced classification methods, including vision transformers and residual BAGAN, while maintaining a significantly lower computational cost. In fact, EffBaGAN is more than twice as fast as the residual BAGAN, making it an efficient solution for remote sensing image classification in data-scarce environments. Index Terms—Balancing generative adversarial network (BAGAN), classification, data augmentation, EfficientNet, multispectral, residual generator, transformer, vegetation. NOMENCLATURE AA Average accuracy. ACGAN Auxiliary classifier generative adversarial network. Received 5 August 2024; revised 22 October 2024; accepted 1 December 2024. Date of publication 4 December 2024; date of current version 3 January 2025. This work was supported in part by the Agencia Estatal de Investigación, Government of Spain, Contract PID2022–141623NB–I00, and Contract TED2021–130367B–I00 funded by MCIN/AEI/10.13039/501100011033 and by the European Union NextGenerationEU/PRTR funds, in part by the ConselleríadeCultura,Educación,FormaciónProfesionaleUniversidades,Xuntade Galicia, through the aid of accreditation of Galician Research Center 2024-2027 ED431G-2023/04, and accreditation of competitive Research Group ED431C 2022/16; all of them are cofounded by the European Regional Development Fund (ERDF). (Corresponding author: Nicolás Vilela-Pérez.) Nicolás Vilela-Pérez and Dora B. Heras are with the Singular Research Center onIntelligent Technologies (CiTIUS),Universidadede Santiagode Compostela, 15782 Santiago de Compostela, Spain (e-mail: [email protected]; [email protected]). Francisco Argüello is with the Department of Electronics and Computing, Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain (e-mail: [email protected]). Digital Object Identifier 10.1109/JSTARS.2024.3510859 Adam Adaptive moment estimation. BAGAN Balancing generative adversarial network. CGAN Conditional generative adversarial network. CNN Convolutional neural network. ConViT Convolutional-like vision transformer. DWConv Depthwise convolution. DWSConv Depthwise separable convolution. EffBaGAN Efficient balancing generative adversarial network. ELU Exponential linear unit. FLOPs Floating-point operations. GAN Generative adversarial network. κCohen’s kappa coefficient. LeakyReLU Leaky rectified linear unit. MBConv Mobile inverted bottleneck. OA Overall accuracy. PReLU Parametric rectified linear unit. PWConv Pointwise convolution. RedEffDis Reduced EfficientNet discriminator. ResBaGAN Residual balancing generative adversarial network. ResEffGen Residual EfficientNet generator. ResNet Residual network. SE Squeeze-and-excitation. SEEDS Superpixels extracted via energy-driven sampling. SiLU Sigmoid linear unit. SLIC Simple linear iterative clustering. TrSp Training speedup. TTPE Training time per epoch. ViT Vision transformer. WP Waterpixels. I. INTRODUCTION REMOTE sensing multiand hyperspectral images are valuable tools for classifying elements in a scene [1]. Applicationsofsuchclassificationincludeforestmapping,monitoring the evolution of invasive species in watersheds [2], and estimating crop yield. For example, crop yield estimation can be achieved by analyzing the electrical conductivity of the soil from various spectral bands of the images [3]. Various machine learning techniques have been employed for image classification tasks in remote sensing [4]. In recent years, the focus has shifted primarily toward deep learningbased techniques, which have proven more effective for multi- © 2024 The Authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. For more information, see https://creativecommons.org/licenses/by-nc-nd/4.0/ 2478 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 and hyperspectral image classification compared to traditional approaches [5],[6]. Among these techniques, CNN [7], such as ResNet [8], are particularly prominent and are commonly chosen for this purpose [9],[10],[11]. The integration of skip connections in ResNets facilitates the training of deep networks by mitigating the vanishing gradient problem, which can prevent the learning process in very deep networks [12]. However,thesetechniques arecharacterizedby ahighcomputationalcost[13].Thus,inrecentyears,researchhasincreasingly focused on developing fast and effective models that require fewer computational cost, thus minimizing training and testing times [14]. This is particularly relevant for applications on mobile devices with limited computing capacity, such as those utilizing MobileNet networks [15], or in scenarios requiring real-time or near real-time processing. One of the techniques that are applied to reduce the computational cost in deep learning are network compression techniques, such as pruning and quantization [16]. Pruning involves removing parameters that do not significantly impact the network’s performance, thereby enhancing computational efficiency. Quantization, on the other hand, reduces the number of operations by lowering the precision of the data type used for weights or activations in the networks. For instance, Hernández et al. [17] demonstrated the application of quantization in models for IoT devices within a facial recognition system, highlighting how this technique allows model training on IoT devices by reducing computational cost without significantly compromising accuracy. Another alternative to reduce the computational cost is to develop network models focused on this purpose, being able to highlight the EfficientNet architecture [18], proposed by Google in 2019. The main novelty introduced in this architecture is that different networks can be designed by uniformly varying the three dimensions that define them: depth, width, and resolution. This is done through a compound coefficient, which is defined as an exponent affecting the three dimensions mentioned above. It allows an adaptation to the computational resources of the device used and/or to the computational time requirements without affecting the resulting accuracy. EfficientNet achieves comparable and often superior image classification accuracies compared to other CNN-based architectures, such as ResNet. There are previous works in the literature where EfficientNets havebeenused fortheclassification ofimagesinremote sensing, but not for the classification of the different elements present in watersheds. One problem with deep learning networks is that they require a large amount of data to train properly [19]. Added to this is the factthatdatascarcityandclassimbalancesareprevalentinmultiand hyperspectral remote sensing datasets, leading to biased classifiers that do not generalize well. Data augmentation helps alleviate this issue. Although the most common way to perform data augmentation is by applying simple transformations to existing samples [20], it has also recently been performed by using a type of neural architecture known as GAN [21]. These architectures allow the creation of completely new synthetic samples from an estimate of the reference data distribution [22], [23]. GAN architectures that generate synthetic images from random noise are known as noise-to-image, but some architectures generate synthetic images from input images, known as image-to-image. This work is focused on noise-to-image architectures. Different GAN designs are found in the literature [24]. rCGAN [25]: incorporates the corresponding information to synthesize a sample of a specific class on demand. rACGAN [26]: a natural extension of CGAN that allows the discriminator to assign each sample the most probable class, as well as distinguish between whether it is a real or synthetic sample. rBAGAN [27]: when dealing with imbalanced datasets among the different classes present in them, GANs may not have enough information from the minority classes to train on the most relevant features of these classes. Furthermore, in these cases, GANs synthesize identical samplesfor each class, with no varietyto enrich the dataset, sometimes even failing to synthesize noise. To solve these two problems, this architecture was developed. It introduces multiple modifications to stabilize the training of small and imbalanced datasets. These improvements are the introduction of an autoencoder [28] and the combination of both discriminator outputs (the predicted class and whether it is a real or synthetic sample). In summary, the main problem with deep learning methods for classification of remote sensing images is their high computational cost. Thus, in this work, we have designed a proposal to maintain very good accuracy metrics in this type of classification while reducing the computational cost compared to other state-of-the-art methods. This work presents an alternative to perform data augmentation by taking advantage of the computational efficiency of EfficientNetnetworks,proposingEffBaGANas a result.EffBaGAN is a novel EfficientNet-based augmentation and classification method for multispectral remote sensing imaging for environmental monitoring. The EffBaGAN architecture includes a customdiscriminator andgenerator. Inparticular, thecomputational cost reduction is achieved through the RedEffDis, our simplified proposal of EfficientNet. The proposed residual EfficientNet generator (ResEffGen) includes an EfficientNet-based residual pathtoallowthesynthesisofhigher-qualitysamples.Itpreserves the features of the initial latent vector while creating new ones and synthesizing these samples in two steps: one for the spatial expansion and one for the spectral information. The challenges of data scarcity and class imbalances are addressed through a data augmentation technique that combines traditional transformations with sample synthesis by the proposed EffBaGAN. Thus, the main contributions of this work are the following. 1) The proposed method incorporates an EfficientNet-based architecture and a sample extraction process based on superpixel segmentation and traditional augmentation techniques combined with sample synthesis by BAGAN. This achieves high accuracies due to the enrichment obtained through the sample extraction procedure and both data augmentation techniques, as well as reduced computational cost due to the sophisticated building blocks used in EfficientNet-based architectures: the MBConv blocks. VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2479 The combination of these technical innovations is crucial in scarce and imbalanced datasets, where the available labeled samples should be used to their full potential. 2) The discriminator RedEffDis allows achieving very good accuracy metrics while significantly reducing computational cost compared to other classifiers based on residual architectures such as ResNet. This discriminator is characterized by having a reduced number of blocks, adapting to the samples extracted from the datasets used. Each of these blocks extracts spatial features from each channel independently, thus avoiding interaction between them and being faster and more computationally efficient. It does this through a DWSConv, which has two steps: first the independent feature extraction mentioned above, and then the combination of these through a one-point convolution. 3) The generator ResEffGen includes an EfficientNet-based residual path, which allows for the combination of the most relevant features of the initial noise and the new ones generated through the main path. As a result, richer and more complete samples are generated through a two-step process, being one for the spatial resolution and the other for the spectral information. This proposed residual path first expands the initial noise to have the spatial resolution of the desired samples, and then synthesizes the values of the different bands of the synthesized sample. This is done through a transposed DWSConv, which performs the two steps of this operation but with the transposed behavior. The rest of this article is organized as follows. Section II shows the EfficientNet architecture and those of GAN used for the proposed method. Then, Section III describes EffBaGAN in detail. Thereafter, Section IV presents the different experiments for evaluating EffBaGAN in terms of classification performance and computational efficiency. Next, Section Vcarries out the discussion of our proposal. Finally, Section VI concludes this article. II. RELATED WORK As discussed above, the basis of this work is the EfficientNet to reduce computational cost and GANs to address the problem of scarcity and imbalance in remote sensing images. A. EfficientNet Its low computational cost makes this type of network a very interesting alternative for the classification task. The EfficientNet [18] is characterized by a uniform scaling of the three dimensions of this type of network through a compound coefficient, thus adapting to the computational or temporal constraints imposed while trying to obtain a scheme as accurately as possible, as already explained in the introduction. The main component of the EfficientNet architecture is MBConv,firstintroducedinMobileNetV2networks[29].Thisblock consists of the following two phases. rExpansion through a PWConv: the block starts with a convolutional operation with 1×1size filter called PWConv. This convolution aims to increase the input dimensionality, projecting it into a high-dimensional space to capture more complex representations. Thechannel expansion isdefined by the expansion factor e. rDWSConv: the result of the previous phase is then passed through a DWSConv, significantly reducing the computational cost compared to a traditional convolution by avoidingintensiveinteractionsbetween channels. This operation is further divided into the following two parts. 1) DWConv: applies a convolution to each input channel separately, thus identifying spatial patterns independently to each channel, preserving the number of channels but reducing the spatial resolution. 2) Reduction through a PWConv: is the last operation in the block, and is applied to learn more complex representations by combining the spectral information from the output of the previous operation. In turn, this block reduces the number of output channels to adapt it to the processing of the subsequent stages. It should be noted that, in EfficientNets, MBConv blocks additionally include an SE block [30], which consists of two phases: a first phase that obtains a representative value of each feature map/channel as a global summary of everything present in it, and a second phase that consists of learning through fully connected layers the weights that will be applied to each feature. The percentage of channels processed through this block is defined by the reduction rate se. Thesameauthorsas EfficientNethavelater proposed a second version of it: EfficientNet V2 [31]. The main difference between this network and the first version is the modification of the MBConv block architecture, becoming a Fused MBConv block, introduced in [32]. The Fused MBConv block differs from the standard by replacing the first two operations (PWConv and DWConv) by a standard convolution. The MBConv block used in the EfficientNet architecture, as well as the Fused MBConv block used in the EfficientNet V2 architecture [31], are depicted in Fig. 1, where the differences between them can be seen. By using the MBConv block, CNNs will achieve higher computationalefficiency,as this blockreducesthe numberofcomputations compared to traditional convolutions. Furthermore, the addition of the SE block allows CNNs to focus on the most relevant features of the input sample and attenuate those that are less relevant with almost no added computational cost [33].On the other hand, the standard convolution applied in the Fused MBConv block is faster than PWConv + DWConv at runtime due to the exploitation of the spatial locality of the processed data. As mentioned before, the main building block of EfficientNets, MBConv, was previously introduced in MobileNet networks, also with the same objective of reducing the computational cost. MobileNet-based networks demonstrated, on RGB remote sensing images, to obtain very competitive accuracies while the computational operations used by such models are lower than for other standard networks [34],[35]. Similarly, EfficientNet-basednetworksalsoobtainverycompetitiveresults 2480 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 Fig. 1. Structure of (a) MBConv and (b) Fused MBConv blocks. The input sample has Cin channels with height Hand width W. The accumulated output dimensions of each operation are shown on its left. The blocks MBConv and Fused MBConv apply expansion factor eon the first operation, stride sH×sWon the second operation in MBConv and on the first in Fused MBConv, and reduction rate se ∈[0,1] on the first red operation. The number of output channels is Cout, selected in the last operation of both blocks. Fig. 2. Architecture of a GAN. It can be observed how the generated latent vectors pass through Gto generate synthetic samples. in terms of classification accuracy for remote sensing problems, also with a lower computational cost [36],[37],[38],[39]. B. Generative Adversarial Networks We chose GANs as a data augmentation technique on scarce and imbalanced datasets. A GAN consists of two neural networks, the generator G, and the discriminator D, which compete with each other in an adversarial learning process. Gcreates syntheticdatasamplesfromarandomlatentinputspace,whileD evaluates the authenticity of these samples, trying to distinguish them from the real ones. As the networks are trained together, Gimproves its ability to produce increasingly realistic samples, while Dimproves its ability to discern between real and synthetic samples. This process of competition and feedback allows GANs to generate synthetic samples that are indistinguishable from real ones, making them especially useful as a data augmentation technique, with great potential compared to traditional techniques [40]. The GAN architecture is depicted in Fig. 2. Both Gand Duse convolutional neural architectures. While Dremains a common architecture, Guses transposed convolutionsasitsmainoperation,workinginversely.SothenDneedsto transform an input sample into a unidimensional feature vector of length Z, while Ghas as input a vector of length Zand will transform it into a sample whose sizes are those needed for D’s input. This unidimensional vector of length Zis known as latent space, being Zthe size of said space. This space is nothing more than the abstract representation of the characteristics and attributes of the samples synthesized by G. BAGAN’s features make it a very good choice for applying data augmentation technique on datasets with significant class imbalances, such as remote sensing datasets [27]. This architecture introduces various improvements to stabilize training, especially in situations of data scarcity and class imbalances. First, the parameters of the two networks are initialized by an autoencoder[28], thus initiatingthetraining fromagoodstarting point and subsequently learning how to represent the different classes in the latent space. In addition, thanks to the initialization of Gwith the autoencoder’s decoder, the generator can learn an accurate class-conditioning in the latent space. Various proposals for using GANs for data augmentation in the classification of images with EfficientNet have been presented in the bibliography. In [41], different methods to improve the classification of an EfficientNet-B0 network on a COVID-19 chest X-ray image set are proposed, one of them being a GAN-based augmentation one, which did not obtain the bestresults, butinstead obtained thebest results by balancingthe dataset. Kwak and Kim [42] used a GAN-based augmentation method to further classify very high-resolution multispectral remote sensing images with an EfficientNet-based classifier, obtaining better results than the alternatives without augmentation. Regarding this method, it is worth mentioning that they used a CycleGAN [43], being this an image-to-image augmentation method, and not a noise-to-image one. Finally, Abady et al. [44] usedGAN-based architecturestoaugment multispectralsatellite images from different regions through season transfer with these VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2481 Fig. 3. Procedure for extracting training samples from the dataset in EffBaGAN. It starts with a segmentation in superpixels to subsequently obtain the patch of each superpixel and enrich the data through traditional augmentation, introducing the output of this procedure to the network architecture of Fig. 4. The augmentation is not performed on the validation and test sets. The multispectral image has height H, width W,andBbands. architectures. Then, they performed detection and localization of these areas with EfficientNet-B4. The results show that in scenarios where the training and test sets were generated with the same GAN architecture, it is very accurate, but in those where this was not the case the results can be improved. All of the abovementioned proposals have been the starting point for our thinking that a method combining the low computational cost of EfficientNet and data augmentation through GAN would be promising. Transversely to the previous proposals, it is also worth mentioning that Feng et al. [45] demonstrated, on hyperspectral remote sensing images, how a residual generator obtains better accuracies in GAN-based methods on noise-to-image than one that does not have these residual features. This last approach was the starting idea for our later proposal for a residual generator with EfficientNet features (which in turn has residual features), presented in more detail in Section III-C. Shortcuts in the residual generator allow the creation of synthetic samples preserving the most important features of the initial latent vector while generating new ones, and the EfficientNet features present in it allow the synthesis of the samples in two steps: one for the spatial expansion and one for the spectral information. III. PROPOSED METHOD EffBaGAN, the proposed method for remote sensing classification applied to data scarcity scenarios, is illustrated in Figs. 3 and4.This section describes its differentcomponentsasfollows. First, the sample extraction procedure that combines superpixel segmentation and traditional augmentation techniques is introduced in Section III-A. Next, the designed network architecture is discussed globally and jointly in Section III-B. Finally, the design of the neural architectures in EffBaGAN is detailed in Section III-C. A. Sample Extraction Via Superpixel Segmentation and Traditional Augmentation Aiming to improve EffBaGAN for scarce and imbalanced datasets, a procedure for extracting samples from these datasets has been used [46]. This procedure combines two techniques widely used in remote sensing: superpixel segmentation and traditional augmentation techniques. Fig. 3shows this procedure, step by step, the result of which are input samples to the network architecture described in Section III-B. Fig. 4. Neural network architecture of EffBaGAN. Its main features are an EfficientNet-based BAGAN design that incorporates an EfficientNet-based residual generator (ResEffGen) and an EfficientNet-based classifier (RedEffDis). The input real training samples are obtained using the extraction procedure presented in Section III-A. Superpixel segmentation groups contiguous pixels with similar characteristics, resulting in a homogeneous region. These regions do not have to be of a specific and/or regular shape, but have to be adapted to the image to group pixels with similar characteristics. Since these are large and very high-resolution images, a representative patch of the superpixel will be selected for classification to classify the central pixel of it, and then the class resulting from this classification will be propagated to the whole superpixel. In this way, the number of samples to be processed by the classification scheme is significantly reduced, making 2482 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 the computational cost of the scheme much lower than if it were done at the level of individual pixels. This representative patch is obtained from the already segmented image, determining the minimumquadrilateral that containsthesuperpixelwithin it,and setting the center point as the superpixel’s central pixel. Once we have this central pixel, we establish a search range of size N×Npixels centered on this central pixel, forming the patch ofthissuperpixel.Itisworthmentioningthattheclassassociated to the superpixel will be the class of the aforementioned central pixel. The patches have a spatial resolution of 32 ×32. To perform the superpixel segmentation, different algorithms in the literature, such as SEEDS [47] or SLIC [48], could be used. In this work, WP [49] is used to facilitate comparison with other classification techniques already published as well as for its low computational time and good performance, providing homogeneous, compact regions with high adherence to the edges [2]. Allextractedpatches are further subjected to a traditional augmentation technique, which consists of, first, a random rotation of 0◦,90 ◦, 180◦, or 270◦, and then a possible horizontal and/or vertical flip, with a 50% probability of occurrence each. B. Network Architecture The EffBaGAN is an EfficientNet-based BAGAN architecture. It has been designed to include a new residual generator ResEffGen with EfficientNet features, allowing the synthesis of higherqualitysamples.ThediscriminatorRedEffDis,also based on EfficientNet, aims at reducing the computational cost. Since the high-resolution datasets for the classification applied to forest mapping have very limited labeled samples, an autoencoder module [28] is inserted. It stabilizes the subsequent training of the GAN, thus achieving a more realistic sample synthesis even with scarce and imbalanced datasets. The EffBaGAN training process consists of three steps, as shown in Fig. 4. 1) First, the autoencoder is trained with all the available samples without considering their label, to produce an initial understanding of the data distribution. To guide the unsupervised training, mean squared error loss is used, with the autoencoder aiming to minimize it as much as possible. This autoencoder module consists of an encoder and a decoder, which must have the same topology as the discriminator(RedEffDis)and the generator(ResEffGen), respectively, to allowknowledgetransfer through the sharing of weights of the autoencoder with the uninitialized GAN. 2) Once the autoencoder has been trained, all the learned parameters are transferred to the GAN, thus making it acquire the knowledge of the autoencoder. The objective of the autoencoder is to learn how to compress and reconstruct the samples as accurately as possible. 3) Finally, the last step is the training of the GAN module itself, where due to the transfer of the previous step, it starts from a more stable point than if the parameters were randomly initialized. To guide the training of this step, categorical cross-entropy loss is used in both the ResEffGen and the RedEffDis. It should be noted that the total loss of the discriminator results from the sum of the categorical cross-entropy loss of the real samples and that of the synthetic samples. The first part results from the calculation of the loss function between the predicted class probabilities of the real samples and their labels. On the other hand, the second part results from the calculation of the loss function between the predicted class probabilities of the synthetic samples and their labels. Once the EffBaGAN networks have been trained, the synthetic output of the discriminator is disabled, so that the final classifier only makes predictions about the classes in the dataset. C. Network Topologies As mentioned in Section III-B, the autoencoder and the GAN modules share topology. First, the topology of the EfficientNetbased discriminator RedEffDis, also applicable to the encoder, will be discussed. Subsequently, the topology of the EfficientNet-based residual generator ResEffGen, also applicable to the decoder, will be discussed. Itshouldbenotedthatthetopologieshavebeenselectedaftera study that will be detailed in Section IV-B1 for the discriminator, and in Section IV-B2 for the generator. For the discriminator topology, we have proposed the RedEffDis architecture, intending to obtain a better adaptation to the BAGAN and to the multispectral remote sensing datasets. We can highlight the following features. 1) RedEffDis is based on EfficientNet-B0: EfficientNet has multiple versions, ranging from EfficientNet-B0 to EfficientNet-B7 uniformly increasing the three dimensions of the network. These three dimensions are depth, width and resolution, so that the EfficientNet-B7 will have more layers, and a larger number of channels and sample sizes than the EfficientNet-B0. All of these versions are designed for much larger datasets than those used in this work (such as ImageNet [50]), so we have focused on the simplest architecture of this type of network to avoid possible overfitting to the training data, which is EfficientNet-B0. 2) RedEffDis avoids the excessive reduction of spatial resolution provided by EfficientNet-B0 by eliminating the first 3 stages with MBConv of this network. In other case, the last stages of the network will not be able to extract the spatial features correctly even having the necessary operations to do so, since these stages have an input spatial resolution of 1 ×1. The proposed elimination decreases the reduction of spatial resolution across the network, so that the final stages are able to extract spatial features by having an input spatial resolution greater than 1 ×1. 3) Finally, RedEffDis includes several modifications focused on MBConv block stages. It should be noted that the same proportion among the different stages of the network architecture than for EfficientNet-B0 is respected. The modifications focused on the following aspects. VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2483 Fig. 5. Diagram of the RedEffDis for EffBaGAN. The output dimensions correspond to an input sample with dimensions H×W×B=32×32 ×B, being Bthe number of spectral bands. The values of the modifiable parameters (r1,r 2,e,se)will be those of the corresponding configuration. rNumber of repetitions for each MBConv stage, r1 and r2: these values are selected complying with the constraint r1≤r2. This is shown in Fig. 5, where each stage with MBConv blocks is repeated the number of times indicated on its right. rExpansion factor e: the MBConv block of each stage performs an expansion of the number of input channels in the first operation, as shown in Fig. 1.The number of channels resulting from the expansion is the number of input channels Cin multiplied by e.The values of these factors are chosen to be the same for all stages with MBConv, thus respecting the proportions of EfficientNet-B0. rReduction rate se: the MBConv block of each stage includes an SE module that performs the SE operations. The reduction rate se indicates the percentage of channels that are selected for these operations. It is a decimal number in the range [0,1], and, similar to the expansion factor e,se is the same for all stages with MBConv. It is worth mentioning that the modifications proposed for the MBConv block can also be applied to the Fused MBConv one. Following the approach of EfficientNet V2 [31], two additional versionsweretested:onewiththefirsthalfbeingFusedMBConv blocks, and another with all of them. We call these versions RedEffDis-FusedHalf and RedEffDis-FusedAll, respectively. TABLE I DETAILS OF THE LAYERS IN THE REDEFFDIS FOR EFFBAGAN Fig. 6. Diagram of the ResEffGen for EffBaGAN. The size of the input latent vector is Z. The output dimensions of the generator are 32 ×32 ×B, being B the number of spectral bands. These output dimensions must be the same as the real samples and as the input dimensions of the RedEffDis. The proposed RedEffDis topology is depicted in Fig. 5and detailed in Table I. As a result of the discriminator study in terms of accuracy and training time, we have proposed two different configurations of EffBaGAN depending on the values of (r1,r 2,e,se)defined for RedEffDis. EffBaGAN-Base is defined by values (3,4,6,0.5), which means that it has (2 ·3+4+1)=11 MBConv blocks with expansion factor e=6and reduction rate se =0.5. EffBaGAN-Small is defined by (1,1,6,0), thus having only (2 ·1+1+1)=4 MBConv blocks with expansion factor e=6and reduction rate se =0. These two configurations EffBaGAN-Base and EffBaGANSmall will be evaluated in Section IV. RegardingthegeneratorResEffGen,itisdepictedinFig.6and detailed in Table II. We have proposed to add an EfficientNetbased residual path next to a main path that includes two transposed convolutions. The EfficientNet-based residual path performs the transposition of the DWSConv operation of the EfficientNet(seethetwotransposedconvolutionsinthegreenpath, 2484 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 TABLE II DETAILS OF THE LAYERS IN THE RESEFFGEN FOR EFFBAGAN the residual one, in Fig. 6). The transposed convolutions of the main path increase the spatial resolution of the synthetic sample by two steps to finally obtain a sample of the desired spatial and spectral resolution. For its part, the residual path consists of a transposed DWSConv which, as previously explained, is formed by a DWConv followed by a PWConv, and in this case we have adapted them to have this transposition behavior. This combination of the RedEffDis and the ResEffGen is effective for the following two main reasons. rThe EfficientNet-based discriminator trains faster thanks to the optimized MBConv block. Its design includes the DWSConv operation, which has a lower computational cost than a standard convolution as it avoids interaction between channels. rThe EfficientNet-based residual generator produces an enhanced-accuracy network thanks to the increased quality of the synthetic samples. This is achieved by complementing the information contained in the main path with the synthetic samples produced by the residual path with transposed DWConv and transposed PWConv. IV. EXPERIMENTS In this section, EffBaGAN is evaluated in terms of classificationperformancetoassessitsoveralleffectiveness.EffBaGAN’s capabilities are compared with other classification methods, such as a CNN, a ResNet, a BAGAN-based method, a ViT, and other methods designed for low computational cost such as MobileNet or EfficientNet-B0. The section is organized as follows. First, the datasets, metrics, and execution environment used for the experiments are outlined in Section IV-A. The optimizer, weight initialization, and the hyperparameter optimization procedure, are also detailed here. Then, the EffBaGAN’s discriminator and generator topology selection, as well as the experimental results, are presented and discussed in Section IV-B. A. Experimental Setup 1) Datasets: Eight large, very high-spatial resolution multispectral images of natural regions with dense vegetation, which were used in [2], were considered. These images have been captured in 2018, 2019, and 2020 flying an autonomous aerial vehicle at 120 m altitude over several river basins in Galicia (Spain), resulting in a spatial resolution of 10 cm/px. The vehicle carried a MicaSense RedEdge-MX multispectral camera, capturing five bands corresponding to wavelengths of 475 nm (blue), 560 nm (green), 668 nm (red), 717 nm (red-edge), and 840 nm (near-infrared). Table III details the specific locations and dimensions of the scenes. Fig. 7shows the composite color images corresponding to the eight river images, together with their reference information. Table IV lists the ten classes identifiable in the reference information, detailing the number of samples in each dataset. These classes range from native vegetation to human-made structures such as roads or buildings. It is important to note the data scarcity and large imbalances between classes in all datasets. These imbalances introduce a bias toward the majority classes, which could prevent a balanced classification accuracy across classes. All datasets were segmented into superpixels using the WP algorithm choosing an average size of 400 px/superpixel, with a minimum of 100 px/superpixel, and utilizing a compactness factor of 0.5 points, following the approach of [2]. The extracted patches have a spatial dimension of N×N=32×32 px. In addition, all data were normalized to the range [−1,1] . For all datasets, a training set of 15% of the samples and a validation set of 5% of the samples will be used to monitor training progress to identify potential problems such as overfitting. This leaves the remaining 80% of the samples for the test set. 2) Metrics: The classification performance of EffBaGAN is determined by class prediction of each labeled sample and comparison of results with reference information. For this, three standard pixel-level metrics in remote sensing classification [51] will be used, excluding for their calculation only the central pixelsofthesuperpixelsofthetrainingset:OA,AA,andCohen’s kappa coefficient (κ). The computational cost of the EffBaGAN will be evaluated throughdifferentmetrics.Executiontimeisevaluated intermsof TTPE, in seconds. Based on this metric, the TrSp evaluates how much faster a classification method trains per epoch compared to a pre-established one. Three metrics related to the size and complexity of the network will also be obtained: network size in memory (in MiB), number of trainable parameters (floatingpoint variables) of the network, and number of operations, in particularFLOPs, ofthenetworkforoneforwardpasswithbatch size one. 3) Execution Environment: All experiments have been performed on a computing cluster. The used node has two AMD EPYC 7543 CPUs with 32 cores each, 256 GB of RAM, and two NVIDIA Ampere A100 GPUs with 40 GB of VRAM each, but only one core and one GPU have been used in the experiments carried out. Regarding the software, the code has been executed within a Conda environment, with Python 3.8.13, and CUDA 11.3 with cuDNN 8.3.2. The programming language to be used will be Python, on which the following main packages should be highlighted: PyTorch 1.12.0, for the creation, training, and testing of the networks; NumPy 1.24.3, for the manipulation of the datasets; scikit-learn 1.3.2, for the preprocessing of the datasets and obtaining κ; fvcore 0.1.5.post20221221, for obtaining the FLOPs; torchinfo 1.7.2, for obtaining the detailed information on the architecture of the networks; and Guild AI 0.8.1, for the registration of the experiments. VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2485 Fig. 7. Composite color images (left) and reference information (right) of the datasets used in this work. All representations follow the same size scale. The class corresponding to each color in the reference information is described in Table IV, while black means that there is no reference information about those pixels. 2492 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 Fig. 11. Composite color image (a), reference information (b), ViT classification map (c), and EffBaGAN-Small classification map (d) of a region of the Eiras Dam dataset where pines and rocks are present. The class and color identification is consistent with the details provided in Table IV. (a) Composite color. (b) Reference information. (c) ViT. (d) EffBaGAN-Small. Fig. 12. Composite color image (a), reference information (b), ViT classification map (c), and EffBaGAN-Small classification map (d) of a region of the Eiras Dam dataset where eucalyptus are present. The class and color identification is consistent with the details provided in Table IV. (a) Composite color. (b) Reference information. (c) ViT. (d) EffBaGAN-Small. TABLE IX COMPUTATIONAL RESOURCES AND AVERAGE RECORDED ACCURACY METRICS FOR EIRAS DAM DATASET USING DIFFERENT EXPANSION FACTORS FOR THE EFFBAGAN-SMALL DISCRIMINATOR VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2493 Fig. 13. Accuracy metrics in the ablation study of the two data augmentation techniques used in EffBaGAN on the Eiras Dam dataset. The bars represent the mean values along their confidence intervals. Higher values are better. TABLE X RECORDED ACCURACY METRICS FOR THREE WIDELY USED REMOTE SENSING DATASETS USING EFFBAGAN-SMALL 610 ×340 px), making pixel-by-pixel processing computationally feasible. Each run for these datasets was repeated five times. Table Xshows that EffBaGAN-Small achieved OA and AA values above 99% for two out of the three datasets. For Indian Pines,withhasamuchlowernumberoflabeledsamples,specific parameter and architecture configurations would be required for maximum performance. V. DISCUSSION Deep learning-based techniques for the classification of remote sensing images present high classification performance, but also high computational cost and, typically, require large amounts of data to be adequately trained. Scenarios with data scarcity, such as the forest mapping one studied in this work, pose a challenge to these techniques. In general, in remote sensing for Earth observation, the scarcity of labeled data, and the class imbalances [4] are common issues. In this context, data augmentation techniques, especially those based on GANs [21] such as BAGAN [27], play a significant role. Regarding the high computational cost of the deep learning classification techniques, several new efficient methods, such as EfficientNet [18], have been developed to specifically reduce it [14]. In this work, we focus on EfficientNets due to their structure based on the compound coefficient that modifies the depth,width andresolutionof thenetworktoadapt tothespecific computational device used without significantly affecting classification accuracy. EffBaGAN, a combination of EfficientNet, an efficient classification method, and BAGAN, a data augmentation technique, has been shown to be effective in this work. The analyzed results show lower computational cost than other approaches specifically designed with similar objectives, such as ResBaGAN. It is important to mention that the proposed method in this work is adapted to the spatial and spectral resolution of the considered datasets. If it is desired to use this method with input samples with different spatial resolution or from remote sensing images with different number of bands than those shown here, the patch size should be changed. In our case the size is 32 ×32 ×5. The spatial resolution in the network is reduced by the consecutive convolutional layers, so the number of layers and the stride should be adapted to avoid an excessive reduction of the spatial resolution. If the number of bands is very different from that of the images used here, it will also be appropriate to change the number of feature maps that are obtained throughout the network, in both discriminator and generator. If maximum performance is required, some additional changes can be applied. First, in the discriminator, the sizes of the kernels used in the convolution operations for the extraction of features should be selected according to the resolution of the input patches; and, second, in the generator, the size of the latent space should be adapted to the desired resolution for the synthetic samples. Overfitting should also be avoided as it was explained in Section IV-B3. In addition, as future work to improve the accuracy of the proposed method, more sophisticated architectures for the discriminator such as transformers could be proposed. Although GANs are a very good alternative for sample synthesis and, consequently, dataset enrichment, they have some limitations. The main limitation of GANs lies in their instability during the training process. Due to the nature of their adversarial architecture, training can be affected by a difficult to achieve equilibrium. On many occasions, the discriminator tends to improve much faster than the generator, which can result in the latter eventually generating poor quality samples. Another limitation of GANs is the high computational cost involved in training, making their applicability difficult in scenarios with limited computational resources or strong temporal constraints. In addition, this type of network is difficult to adjust in terms of the selection of hyperparameters and the architecture itself, and this adjustment is done empirically, being laborious and unsystematic.Forallthesereasons,thedesignofthearchitecture of the proposed method in this work has been done following a careful process. Another method that shows good results in our experiments is ViT, which also presents very high accuracy metrics but lower in a 1% on average than the proposed method EffBaGAN over the studied datasets. It is important to note that EffBaGAN and ViT are two deep learning architectures with different approaches and purposes. EffBaGAN represents an advancement of GANbasednetworks,therefore,employingadversarialtraining,while ViT is based on the transformer architecture that captures complex relations among the input samples. Also, as discussed in the experimentation, ViT distinguishes less well between the different vegetation types in the datasets used, which is the critical point of the classification in this particular domain. The computational cost of EffBaGAN opens an interesting line for future research in the use of GAN-based architectures in real-time applications. While EffBaGAN represents a significant step toward near real-time processing and suitability for 2494 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 hardware architectures with limited capabilities in data-scarcity scenarios, it still presents some limitations. Although it achieves a good balance between performance and computational cost, the training time remains higher compared to other methods. Future work on this network should focus on optimizing its architecture to further reduce computational costs without compromising accuracy. These enhancements could enable more efficient execution on resource-constrained platforms, making it evenmorepracticalforreal-worldapplications. Also concerning potential future improvements for increasing performance, techniques such as quantization and pruning could reduce computational costs even further [16]. These techniques would allow for moreefficientinferencebyminimizingthenumberofoperations and memory requirements. Finally, the adaptation to other computing platforms could also be an interesting line for future research. Although the current experiments were conducted using a single core and one high-performance GPU, aiming to minimize the amount of computing resources used, evaluating the performance on alternative computing infrastructures such as distributed computing infrastructures could prove beneficial. This flexibility would allow the model to scale efficiently for problems involving bigger datasets and, therefore, requiring more substantial computational resources. VI. CONCLUSION This work introduces EffBaGAN, a deep learning method tailored for multispectral remote sensing image classification, specifically addressing the challenges posed by data scarcity while maintaining computational efficiency. The method integrates an EfficientNet-based residual generator and an EfficientNet-based discriminator within a BAGAN-based data augmentation framework. In addition, the use of superpixelbased sample extraction helps to reduce computational costs. This work represents the first application of a fully EfficientNetbased architecture with BAGAN for remote sensing classification. To optimize performance, various configurations of the EffBaGANdiscriminatorwere tested, modifying factorssuch as the repetition of each stage, the expansion factor, and the reduction rate in the SE module. Generator topologies were also explored, with variations in the number of transposed convolutions in the main path and the type of residual paths used, including a simple transposed convolution, upsampling with convolution, or transposed DWSConv. As a result, a more effective and computationally efficient overall network topology was achieved. The method was evaluated under a data scarcity scenario, focusing on classifying eight high-resolution multispectral images of forests in Galicia (Spain) with limited training data and pronounced class imbalances. Compared to other classification methods, including some based on transformers, EffBaGAN demonstrated high overall and average classification accuracy, with an average TTPE of 8.90 s, being more than twice as fast as ResBaGAN, another method designed for data scarcity scenarios. The best-performing configuration, EffBaGAN-Small, achieved an OA of 96.16% and an AA of 87.15%. Finally, this work opens several research directions. One potential direction is replacing the EffBaGAN discriminator withmoreadvancedarchitecturesliketransformers,whichcould enhanceaccuracy.Theproperconfigurationtomanagecomputational costs for such models would need thorough investigation. In addition, reducing the computational burden of EffBaGAN through techniques like quantization and pruning, as well as experimentation on different computing platforms, could improve its efficiency. REFERENCES [1] D. Chutia, D. K. Bhattacharyya, K. K. Sarma, R. Kalita, and S. Sudhakar, “Hyperspectral remote sensing classifications: A perspective survey,” Trans. GIS, vol. 20, no. 4, pp. 463–490, 2016. [2] F.Argüello,D.B.Heras,A.S.Garea,andP.Quesada-Barriuso,“Watershed monitoring in Galicia from UAV multispectral imagery using advanced texture methods,” Remote Sens., vol. 13, no. 14, 2021, Art. no. 2687. [3] M. Teke, H. S. Deveci, O. Halilo˘glu, S. Z. Gürbüz, and U. Sakarya, “A short survey of hyperspectral remote sensing applications in agriculture,” in Proc. 6th Int. Conf. Recent Adv. Space Technol., 2013, pp. 171–176. [4] T. A. W. Aaron, E. Maxwell, and F. Fang, “Implementation of machine-learning classification in remote sensing: An applied review,” Int. J. Remote Sens., vol. 39, no. 9, pp. 2784–2817, 2018, doi: 10.1080/01431161.2018.1433343. [5] M. Paoletti, J. Haut, J. Plaza, and A. Plaza, “Deep learning classifiers for hyperspectral imaging: A review,” ISPRS J. Photogrammetry Remote Sens., vol. 158, pp. 279–317, 2019. [6] S.Li,W.Song,L.Fang,Y.Chen,P.Ghamisi,andJ.A.Benediktsson,“Deep learningfor hyperspectralimageclassification:An overview,” IEEE Trans. Geosci. Remote Sens., vol. 57, no. 9, pp. 6690–6709, Sep. 2019. [7] Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998. [8] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 770–778. [9] W. Rawat and Z. Wang, “Deep convolutional neural networks for image classification: A comprehensive review,” Neural Comput., vol. 29, no. 9, pp. 2352–2449, 2017. [10] S. Yu, S. Jia, and C. Xu, “Convolutional neural networks for hyperspectral image classification,” Neurocomputing, vol. 219, pp. 88–98, 2017. [Online]. Available: https://www.sciencedirect.com/science/article/ pii/S0925231216310104 [11] Z. Zhong, J. Li, L. Ma, H. Jiang, and H. Zhao, “Deep residual networks for hyperspectral image classification,” in Proc. IEEE Int. Geosci. Remote Sens. Symp., 2017, pp. 1824–1827. [12] S.Basodi,C.Ji,H.Zhang,andY.Pan,“Gradientamplification:Anefficient way to train deep neural networks,” Big Data Mining Analytics,vol.3, no. 3, pp. 196–207, 2020. [13] N. Thompson, K. Greenewald, K. Lee, and G. F. Manso, “The computational limits of deep learning,” in Proc. 9th Comput. Within Limits, Jun. 2023. [Online]. Available: https://limits.pubpub.org/pub/wm1lwjce [14] B. R. Bartoldson, B. Kailkhura, and D. Blalock, “Compute-efficient deep learning: Algorithmic trends and opportunities,” J. Mach. Learn. Res., vol. 24, no. 122, pp. 1–77, 2023. [Online]. Available: http://jmlr.org/ papers/v24/22-1208.html [15] A. G. Howard et al., “MobileNets: Efficient convolutional neural networks for mobile vision applications,” 2017. [Online]. Available: https://arxiv. org/abs/1704.04861 [16] T. Liang, J. Glossner, L. Wang, S. Shi, and X. Zhang, “Pruning and quantization for deep neural network acceleration: A survey,” Neurocomputing, vol. 461, pp. 370–403, 2021. [Online]. Available: https://www. sciencedirect.com/science/article/pii/S0925231221010894 [17] N. Hernández, F. Almeida, and V. Blanco, “Performance and energy efficiency: Quantization of models for IoT devices,” Res. Square, 2023, doi: 10.21203/rs.3.rs-3405705/v1. [18] M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. 36th Int. Conf. Mach. Learn., Jun. 2019, pp. 6105–6114. [Online]. Available: https://proceedings.mlr.press/v97/ tan19a.html VILELA-PÉREZ et al.: EFFBAGAN: AN EFFICIENT BALANCING GAN FOR EARTH OBSERVATION IN DATA SCARCITY SCENARIOS 2495 [19] B. Zohuri and M. Moghaddam, “Deep learning limitations and flaws,” Modern Approaches Mater. Sci., vol. 2, pp. 241–250, 2020. [20] S. Yang, W. Xiao, M. Zhang, S. Guo, J. Zhao, and F. Shen, “Image data augmentation for deep learning: A survey,” 2023. [Online]. Available: https://arxiv.org/abs/2204.08610 [21] I. Goodfellow et al., “Generative adversarial nets,” in Proc. Adv. Neural Inf. Process. Syst.., 2014, pp. 2672–2680. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2014/file/ 5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf [22] Q. Su, H. N. A. Hamed, M. A. Isa, X. Hao, and X. Dai, “A GAN-based data augmentation method for imbalanced multi-class skin lesion classification,” IEEE Access, vol. 12, pp. 16498–16513, 2024. [23] E. Strelcenia and S. Prakoonwit, “A survey on GAN techniques for data augmentation to address the imbalanced data issues in credit card fraud detection,” Mach. Learn. Knowl. Extraction, vol. 5, no. 1, pp. 304–329, 2023. [Online]. Available: https://www.mdpi.com/2504-4990/5/1/19 [24] A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, “Generative adversarial networks: An overview,” IEEE Signal Process. Mag., vol. 35, no. 1, pp. 53–65, Jan. 2018. [25] M.MirzaandS. Osindero,“Conditionalgenerativeadversarialnets,”2014. [Online]. Available: https://arxiv.org/abs/1411.1784 [26] A. Odena, C. Olah, and J. Shlens, “Conditional image synthesis with auxiliary classifier GANs,” in Proc. 34th Int. Conf. Mach. Learn., Aug. 2017, pp. 2642–2651. [Online]. Available: https://proceedings.mlr.press/v70/ odena17a.html [27] G.Mariani,F. Scheidegger,R.Istrate, C.Bekas,andC.Malossi,“BAGAN: Dataaugmentationwith balancingGAN,”inProc. Int. Conf. Mach. Learn., 2018. [Online]. Available: https://research.ibm.com/publications/bagandata-augmentation-with-balancing-gan [28] M. A. Kramer, “Nonlinear principal component analysis using autoassociative neural networks,” AIChE J., vol. 37, no. 2, pp. 233–243, 1991. [Online]. Available: https://aiche.onlinelibrary.wiley.com/doi/abs/ 10.1002/aic.690370209 [29] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2018, pp. 4510–4520. [30] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2018, pp. 7132–7141. [31] M. Tan and Q. Le, “EfficientNetV2: Smaller models and faster training,” in Proc. 38th Int. Conf. Mach. Learn., Jul. 2021, pp. 10096–10106. [Online]. Available: https://proceedings.mlr.press/v139/tan21a.html [32] S. Gupta and M. Tan, “EfficientNet-EdgeTPU: Creating acceleratoroptimized neural networks with AutoML,” 2019. [Online]. Available: https://ai.googleblog.com/2019/08/efficientnet-edgetpu-creating.html [33] V.-T. Hoang and K.-H. Jo, “Practical analysis on architecture of EfficientNet,” in Proc. 14th Int. Conf. Hum. Syst. Interact., 2021, pp. 1–4. [34] L. Cao, “A MobileNetV2 model of transfer learning is employed for remote sensing image classification,” Adv. Eng. Technol. Res., vol. 10, no. 1, pp. 596–596, 2024. [35] S. Du, J. Li, and M. Noto, “Comparison and analysis of three MobileNetbased models for wildfire detection,” J. Adv. Inf. Technol., vol. 15, no. 4, pp. 511–518, 2024. [36] P. Charoenchittang, P. Boonserm, K. Kobayashi, and N. Cooharojananone, “Airport buildings classification through remote sensing images using EfficientNet,” in Proc. 18th Int. Conf. Elect. Eng./Electron., Comput., Telecommun. Inf. Technol., 2021, pp. 127–130. [37] H. Alhichri, A. S. Alswayed, Y. Bazi, N. Ammour, and N. A. Alajlan, “Classification of remote sensing images using EfficientNet-B3 CNN model with attention,” IEEE Access, vol. 9, pp. 14078–14094, 2021. [38] R. D. I. Puspitasari, F. Q. Annisa, and D. Ariyanto, “Flooded area segmentation on remote sensing image from unmanned aerial vehicles (UAV) using DeepLabV3 and EfficientNet-B4 model,” in Proc. Int. Conf. Comput., Control, Informat. Appl., 2023, pp. 216–220. [39] R. Wang, Z. Yang, H. Qiu, X. Liu, and D. Wu, “Spatial and channel exchange based on EfficientNet for detecting changes of remote sensing images,” in Proc. 26th Int. Conf. Comput. Supported Cooperative Work Des., 2023, pp. 1595–1600. [40] M. Frid-Adar, I. Diamant, E. Klang, M. Amitai, J. Goldberger, and H. Greenspan, “GAN-based synthetic medical image augmentation for increased CNN performance in liver lesion classification,” Neurocomputing, vol. 321, pp. 321–331, 2018. [Online]. Available: https://www. sciencedirect.com/science/article/pii/S0925231218310749 [41] O.Fedoruk,K. Klimaszewski,A.Ogonowski,andR.Mo˙zd˙zonek,“Performance of GAN-based augmentation for deep learning COVID-19 image classification,” AIP Conf. Proc., vol. 3061, no. 1, 2024, Art. no. 030001, doi: 10.1063/5.0203379. [42] T. Kwak and Y. Kim, “Semi-supervised land cover classification of remote sensing imagery using CycleGAN and EfficientNet,” KSCE J. Civil Eng., vol. 27, no. 4, pp. 1760–1773, Apr. 2023, doi: 10.1007/s12205-023-2285-0. [43] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired Image-to-Image translation using cycle-consistent adversarial networks,” in Proc. IEEE Int. Conf. Comput. Vis., 2017, pp. 2242–2251. [44] L. Abady et al., “Detection and localization of GAN manipulated multispectral satellite images,” in Proc. Euro. Symp. Artif. Neural Netw., 2022, pp. 339–344. [45] B.Feng,Y.Liu,H. Chi,andX.Chen, “Hyperspectralremotesensingimage classification based on residual generative adversarial neural networks,” Signal Process., vol.213,2023,Art.no. 109202.[Online].Available:https: //www.sciencedirect.com/science/article/pii/S0165168423002761 [46] Á. G. Dieste, F. Argüello, and D. B. Heras, “ResBaGAN: A residual balancing GAN with data augmentation for forest mapping,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 16, pp. 6428–6447, 2023. [47] M. Van den Bergh, X. Boix, G. Roig, and L. Van Gool, “SEEDS: Superpixels extracted via energy-driven sampling,” Int. J. Comput. Vis., vol. 111, no. 3, pp. 298–314, Feb. 2015, doi: 10.1007/s11263-014-0744-2. [48] R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Süsstrunk, “SLIC superpixels compared to state-of-the-art superpixel methods,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 11, pp. 2274–2282, Nov. 2012. [49] V. Machairas et al., “Waterpixels,” IEEE Trans. Image Process., vol. 24, no. 11, pp. 3707–3716, Nov. 2015. [50] O. Russakovsky et al., “ImageNet large scale visual recognition challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015, doi: 10.1007/s11263-015-0816-y. [51] R. G. Congalton, “A review of assessing the accuracy of classifications of remotely sensed data,” Remote Sens. Environ., vol. 37, no. 1, pp. 35–46, 1991. [Online]. Available: https://www.sciencedirect.com/science/article/ pii/003442579190048B [52] D. P. Kingma and J. L. Ba, “Adam: A method for stochastic gradient descent,” in Proc. ICLR: Int. Conf. Learn. Representations, 2015, pp. 1–15. [53] X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. 13th Int. Conf. Artif. Intell. Statist., Ser. Proc. Mach. Learn. Res.,May 2010, pp. 249–256.[Online]. Available: https://proceedings.mlr.press/v9/glorot10a.html [54] D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (ELUs),” Under Rev. ICLR, 2015. [Online]. Available: https://www.researchgate.net/ publication/284579051_Fast_and_Accurate_Deep_Network_Learning_ by_Exponential_Linear_Units_ELUs [55] A. L. Maas et al., “Rectifier nonlinearities improve neural network acoustic models,” in Proc. ICML, vol. 30, no. 1, 2013, p. 3. [Online]. Available: https://robotics.stanford.edu/∼amaas/papers/relu_ hybrid_icml2013_final.pdf [56] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassinghuman-levelperformanceon ImageNetclassification,”inProc. IEEE Int. Conf. Comput. Vis., 2015, pp. 1026–1034. [57] D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),” 2023. [Online]. Available: https://arxiv.org/abs/1606.08415 [58] Y. Choi, “PyTorch tutorial,” 2017. Accessed: Oct. 21, 2024. [Online]. Available: https://github.com/yunjey/pytorch-tutorial/tree/master [59] A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Representations, 2021. [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy [60] S. d’Ascoli, H. Touvron, M. L. Leavitt, A. S. Morcos, G. Biroli, and L. Sagun, “ConVit: Improving vision transformers with soft convolutional inductive biases,” in Proc. Int. Conf. Mach. Learn., 2021, pp. 2286–2296. [61] Y. Chen et al., “Mobile-Former: Bridging MobileNet and transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 5270–5279. [62] W. Han et al., “A survey of machine learning and deep learning in remote sensing of geological environment: Challenges, advances, and opportunities,” ISPRS J. Photogrammetry Remote Sens., vol. 202, pp. 87–113, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/ pii/S0924271623001582 [63] ROSIS, “Hyperspectral remote sensing scenes,” 2013, Accessed: Oct. 21, 2024. [Online]. Available: http://www.ehu.eus/ccwintco/index.php/ Hyperspectral_Remote_Sensing_Scenes 2496 IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, VOL. 18, 2025 Nicolás Vilela-Pérez received the B.S. degree in computer engineering and the M.Sc. degree in big data in 2023 and 2024, respectively, from the Universidade de Santiago de Compostela, Santiago de Compostela, Spain, where he is currently working toward the Ph.D. degree in computer science. He is currently a Collaborative Researcher with the Singular Research Center on Intelligent Technologies (CiTIUS), Universidade de Santiago de Compostela. His research interests focus on computer vision tasks, whilebringingtogetherknowledgefrom diversecomputing areas, such as artificial intelligence, high-performance computing, and big data. Dora B. Heras (Member, IEEE) received the M.Sc. degree in physics and the Ph.D. degree in physics from the Universidade de Santiago de Compostela, Santiago de Compostela, Spain, in 1995 and 2000, respectively. Since 2023, she is a Vice-Chair of the International Parallel Computing conference (Euro-Par). Since 2020, she has been the Chair for the HighPerformance and Disruptive Computing in Remote Sensing (HDCRS) Working Group under the IEEE GRSS Earth Science Informatics Technical Committee (ESI TC). She is currently a Full Professor with the Department of Electronics and Computing, Universidade de Santiago de Compostela. Her research interests cover a range of topics in the combined fields of image processing, remote sensing, machine learning, and high-performance computing applied to Earth observation. In particular, she has published papers on registration, classification, domain adaptation, and change detection applied to multispectral and hyperspectral remotely sensed images. Francisco Argüello received the B.S. and Ph.D. degrees in physics from the Universidade de Santiago de Compostela, Santiago de Compostela, Spain, in 1988 and 1992, respectively. He is currently a Full Professor with the Department of Electronics and Computing, Universidade de Santiago de Compostela. His research interests includesignalandimageprocessing,computergraphics, parallel and distributed computing, and quantum computing.