Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing : Replication of Cognitive Models
Full text
This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY-NC-ND 4.0 https://creativecommons.org/licenses/by-nc-nd/4.0/ Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing : Replication of Cognitive Models © 2024 the Authors Published version Terziyan, Vagan; Tiihonen, Timo Terziyan, V., & Tiihonen, T. (2024). Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing : Replication of Cognitive Models. Procedia Computer Science, 232, 890- 902. https://doi.org/10.1016/j.procs.2024.01.089 2024
ScienceDirect Available online at www.sciencedirect.com Procedia Computer Science 232 (2024) 890–902 1877-0509 © 2024 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing 10.1016/j.procs.2024.01.089 10.1016/j.procs.2024.01.089 1877-0509 © 2024 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing Available online at www.sciencedirect.com ScienceDirect Procedia Computer Science 00 (2023) 000–000 www.elsevier.com/locate/procedia 1877-0509 © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing 5th International Conference on Industry 4.0 and Smart Manufacturing Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing: Replication of Cognitive Models Vagan Terziyan a,*, Timo Tiihonen a a Faculty of Information Technology, University of Jyväskylä, 40014, Jyväskylä, Finland Abstract As Industry 4.0 and 5.0 evolve to be highly automated but human-centric, there is a need for process modeling based on digital replicas of physical objects including humans. Knowledge distillation and cognitive cloning offer a way to train operational copies of decision-making black boxes, or donors, without requiring additional data. In this paper, we propose an architecture and analytics for a generative adversarial network, called Cloning-GAN, which enables donor-clone knowledge transfer, including the donor’s individual biases. The architecture involves generating challenging samples to be labeled by the donor and used as training data for the clone. We consider several multicriteria requirements for the generated data, including closeness to the decision boundary, uniform distribution in the decision space, maximal confusion for the donor, and challenge for the clone. We present various strategies to balance these conflicting criteria forcing the clone learning quickly the hidden cognitive skills and biases of the donor. © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing Keywords: Smart Manufacturing; digital twins; cognitive clones; knowledge transfer; knowledge distillation; Generative Adversarial Network; adversarial distillation 1. Introduction Modelling and simulation enhanced with emergent technologies is a key for efficient Industry 4.0 and beyond, including smart manufacturing [1]. In addition to digital twins [2] simulating a physical entity or operation, we can observe a strong trend towards human-centric cyber-physical production systems [3] driven by mental models and smart operators [4] aiming at human values [5], including digital cognitive clones of humans [6] and groups [7]. * Corresponding author. E-mail address: vagan.terziy[email protected] Available online at www.sciencedirect.com ScienceDirect Procedia Computer Science 00 (2023) 000–000 www.elsevier.com/locate/procedia 1877-0509 © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing 5th International Conference on Industry 4.0 and Smart Manufacturing Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing: Replication of Cognitive Models Vagan Terziyan a,*, Timo Tiihonen a a Faculty of Information Technology, University of Jyväskylä, 40014, Jyväskylä, Finland Abstract As Industry 4.0 and 5.0 evolve to be highly automated but human-centric, there is a need for process modeling based on digital replicas of physical objects including humans. Knowledge distillation and cognitive cloning offer a way to train operational copies of decision-making black boxes, or donors, without requiring additional data. In this paper, we propose an architecture and analytics for a generative adversarial network, called Cloning-GAN, which enables donor-clone knowledge transfer, including the donor’s individual biases. The architecture involves generating challenging samples to be labeled by the donor and used as training data for the clone. We consider several multicriteria requirements for the generated data, including closeness to the decision boundary, uniform distribution in the decision space, maximal confusion for the donor, and challenge for the clone. We present various strategies to balance these conflicting criteria forcing the clone learning quickly the hidden cognitive skills and biases of the donor. © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing Keywords: Smart Manufacturing; digital twins; cognitive clones; knowledge transfer; knowledge distillation; Generative Adversarial Network; adversarial distillation 1. Introduction Modelling and simulation enhanced with emergent technologies is a key for efficient Industry 4.0 and beyond, including smart manufacturing [1]. In addition to digital twins [2] simulating a physical entity or operation, we can observe a strong trend towards human-centric cyber-physical production systems [3] driven by mental models and smart operators [4] aiming at human values [5], including digital cognitive clones of humans [6] and groups [7]. * Corresponding author. E-mail address: [email protected] 2 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 Emergent development of such smart assets results in significant benefits for smart manufacturing, making it more efficient, safer, and capable of producing high-quality products. Assisting technologies to enable digital replica of industrial reality (e.g., in the form of smart models, twins, clones, etc.) include machine (deep) learning [8], knowledge transfer [9], knowledge distillation [10], particularly adversarial distillation [11] among others. In this paper, we introduce an adversarial cloning architecture capable of making digital replicas of hidden cognitive assets with respect to their individual biases. The architecture is supposed to be a kind of adversarial distillation driven by a generative adversarial network with a special configuration and objectives regarding the generator and discriminator networks. It facilitates supervised machine learning of the clone model due to a smart way to generate challenging inputs for the donor (the target smart and black-box asset for the clone), who is supposed to act as a supervisor in such an adversarial learning process. The multicriteria-driven generator (adversarial facilitator) is the key cloning enabler and the major innovation in the suggested architecture. The following text of the paper is organized as follows: in Section 2, we present the main contribution of the paper, i.e., the intended cloning architecture; in Section 3, we present the related work; in Section 4 we discuss the added value of this study compared to our previous studies regarding cloning architecture; and we conclude in Section 5. 2. Introduction to Cloning-GAN In this section, we are going to introduce a new architecture capable of unlocking and replication of hidden and smart industrial assets, including human decision-makers and other digital, cognitive and data-driven (e.g., machine learning based) models. 2.1. DONOR, cloning, and CLONE Assume we have a cognitive model aka black box (abstract, physical, cyber, social, etc.) named as DONOR and represented by a hidden “labeling function” 𝓕𝓕:[0,1]𝑛𝑛⟼∆𝑐𝑐 defined as follows: (a) The domain of 𝓕𝓕 is defined by an Euclidean space ℝ𝑛𝑛 bounded by a 𝒏𝒏-dimensional unit hypercube: [0,1]𝑛𝑛, each 𝑖𝑖-th point of which is an 𝒏𝒏-dimensional vector 𝑥𝑥𝑖𝑖 of rational values, aka coordinates of some object of potential labeling: [0,1]𝑛𝑛={∀𝑥𝑥𝑖𝑖:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛)∈ℝ𝑛𝑛|𝑥𝑥1∈[0,1],𝑥𝑥2∈[0,1],…,𝑥𝑥𝑛𝑛∈[0,1]}; (b) The range of 𝓕𝓕 is defined by the unit 𝒄𝒄-simplex ∆𝑐𝑐 (i.e., the 𝒄𝒄-dimensional probability simplex), each 𝑖𝑖-th point of which is a 𝒄𝒄-dimensional vector 𝑝𝑝𝑖𝑖 of probabilities [regarding each of 𝒄𝒄 possible class labels from the set 𝕃𝕃={𝑙𝑙1,𝑙𝑙2,…,𝑙𝑙𝑐𝑐} assigned to the corresponding object represented by point 𝑥𝑥𝑖𝑖 within the domain space]: ∆𝑐𝑐={∀𝑝𝑝𝑖𝑖:(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)∈[0,1]𝑐𝑐|∑ 𝑝𝑝𝑘𝑘 𝑐𝑐 𝑘𝑘=1 =1}; c) ∀𝑥𝑥𝑖𝑖:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛)∈[0,1]𝑛𝑛,∃!𝑝𝑝𝑖𝑖:(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)𝑖𝑖∈∆𝑐𝑐, 𝑝𝑝𝑖𝑖=𝓕𝓕(𝑥𝑥𝑖𝑖). Simply speaking, a DONOR behaves as a labeling (classification) function (like a neural network), which takes a vector of numeric values (normalized to [0,1]) of some target object features as an input and outputs the probability distribution of that object having a particular label (i.e., belonging to a particular class) among several possible ones. Cloning is a supervised learning process, in which the DONOR is a supervisor (i.e., labels the training set samples), and the task is to train the neural network model as a CLONE of the DONOR, i.e., to learn the hidden labeling function 𝓕𝓕 defined above, as precisely as possible. For similar tasks, the traditionally used term is “knowledge distillation”. However, we are using the “cloning” term to follow consistently the terminology from our former articles where the ultimate objective was formulated as designing digital cognitive clones of humans (as decision-makers) and groups of collective hybrid intelligence (see, e.g. [12], [7], [13], [6], and [14]). Cloning-GAN architecture (see Fig. 1), is a special kind of Generative Adversarial Network (GAN) architecture (see review on GANs in [15]), and it is suggested to facilitate cloning process by smart generation of special (challenging or adversarial) samples-as-queries for the DONOR for labeling and then synchronously and incrementally train the CLONE on the basis of generated and labeled samples.
Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 891 Available online at www.sciencedirect.com ScienceDirect Procedia Computer Science 00 (2023) 000–000 www.elsevier.com/locate/procedia 1877-0509 © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing 5th International Conference on Industry 4.0 and Smart Manufacturing Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing: Replication of Cognitive Models Vagan Terziyan a,*, Timo Tiihonen a a Faculty of Information Technology, University of Jyväskylä, 40014, Jyväskylä, Finland Abstract As Industry 4.0 and 5.0 evolve to be highly automated but human-centric, there is a need for process modeling based on digital replicas of physical objects including humans. Knowledge distillation and cognitive cloning offer a way to train operational copies of decision-making black boxes, or donors, without requiring additional data. In this paper, we propose an architecture and analytics for a generative adversarial network, called Cloning-GAN, which enables donor-clone knowledge transfer, including the donor’s individual biases. The architecture involves generating challenging samples to be labeled by the donor and used as training data for the clone. We consider several multicriteria requirements for the generated data, including closeness to the decision boundary, uniform distribution in the decision space, maximal confusion for the donor, and challenge for the clone. We present various strategies to balance these conflicting criteria forcing the clone learning quickly the hidden cognitive skills and biases of the donor. © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing Keywords: Smart Manufacturing; digital twins; cognitive clones; knowledge transfer; knowledge distillation; Generative Adversarial Network; adversarial distillation 1. Introduction Modelling and simulation enhanced with emergent technologies is a key for efficient Industry 4.0 and beyond, including smart manufacturing [1]. In addition to digital twins [2] simulating a physical entity or operation, we can observe a strong trend towards human-centric cyber-physical production systems [3] driven by mental models and smart operators [4] aiming at human values [5], including digital cognitive clones of humans [6] and groups [7]. * Corresponding author. E-mail address: vagan.terziy[email protected] Available online at www.sciencedirect.com ScienceDirect Procedia Computer Science 00 (2023) 000–000 www.elsevier.com/locate/procedia 1877-0509 © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing 5th International Conference on Industry 4.0 and Smart Manufacturing Using Cloning-GAN Architecture to Unlock the Secrets of Smart Manufacturing: Replication of Cognitive Models Vagan Terziyan a,*, Timo Tiihonen a a Faculty of Information Technology, University of Jyväskylä, 40014, Jyväskylä, Finland Abstract As Industry 4.0 and 5.0 evolve to be highly automated but human-centric, there is a need for process modeling based on digital replicas of physical objects including humans. Knowledge distillation and cognitive cloning offer a way to train operational copies of decision-making black boxes, or donors, without requiring additional data. In this paper, we propose an architecture and analytics for a generative adversarial network, called Cloning-GAN, which enables donor-clone knowledge transfer, including the donor’s individual biases. The architecture involves generating challenging samples to be labeled by the donor and used as training data for the clone. We consider several multicriteria requirements for the generated data, including closeness to the decision boundary, uniform distribution in the decision space, maximal confusion for the donor, and challenge for the clone. We present various strategies to balance these conflicting criteria forcing the clone learning quickly the hidden cognitive skills and biases of the donor. © 2023 The Authors. Published by ELSEVIER B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the 5th International Conference on Industry 4.0 and Smart Manufacturing Keywords: Smart Manufacturing; digital twins; cognitive clones; knowledge transfer; knowledge distillation; Generative Adversarial Network; adversarial distillation 1. Introduction Modelling and simulation enhanced with emergent technologies is a key for efficient Industry 4.0 and beyond, including smart manufacturing [1]. In addition to digital twins [2] simulating a physical entity or operation, we can observe a strong trend towards human-centric cyber-physical production systems [3] driven by mental models and smart operators [4] aiming at human values [5], including digital cognitive clones of humans [6] and groups [7]. * Corresponding author. E-mail address: [email protected] 2 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 Emergent development of such smart assets results in significant benefits for smart manufacturing, making it more efficient, safer, and capable of producing high-quality products. Assisting technologies to enable digital replica of industrial reality (e.g., in the form of smart models, twins, clones, etc.) include machine (deep) learning [8], knowledge transfer [9], knowledge distillation [10], particularly adversarial distillation [11] among others. In this paper, we introduce an adversarial cloning architecture capable of making digital replicas of hidden cognitive assets with respect to their individual biases. The architecture is supposed to be a kind of adversarial distillation driven by a generative adversarial network with a special configuration and objectives regarding the generator and discriminator networks. It facilitates supervised machine learning of the clone model due to a smart way to generate challenging inputs for the donor (the target smart and black-box asset for the clone), who is supposed to act as a supervisor in such an adversarial learning process. The multicriteria-driven generator (adversarial facilitator) is the key cloning enabler and the major innovation in the suggested architecture. The following text of the paper is organized as follows: in Section 2, we present the main contribution of the paper, i.e., the intended cloning architecture; in Section 3, we present the related work; in Section 4 we discuss the added value of this study compared to our previous studies regarding cloning architecture; and we conclude in Section 5. 2. Introduction to Cloning-GAN In this section, we are going to introduce a new architecture capable of unlocking and replication of hidden and smart industrial assets, including human decision-makers and other digital, cognitive and data-driven (e.g., machine learning based) models. 2.1. DONOR, cloning, and CLONE Assume we have a cognitive model aka black box (abstract, physical, cyber, social, etc.) named as DONOR and represented by a hidden “labeling function” 𝓕𝓕:[0,1]𝑛𝑛⟼∆𝑐𝑐 defined as follows: (a) The domain of 𝓕𝓕 is defined by an Euclidean space ℝ𝑛𝑛 bounded by a 𝒏𝒏-dimensional unit hypercube: [0,1]𝑛𝑛, each 𝑖𝑖-th point of which is an 𝒏𝒏-dimensional vector 𝑥𝑥𝑖𝑖 of rational values, aka coordinates of some object of potential labeling: [0,1]𝑛𝑛={∀𝑥𝑥𝑖𝑖:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛)∈ℝ𝑛𝑛|𝑥𝑥1∈[0,1],𝑥𝑥2∈[0,1],…,𝑥𝑥𝑛𝑛∈[0,1]}; (b) The range of 𝓕𝓕 is defined by the unit 𝒄𝒄-simplex ∆𝑐𝑐 (i.e., the 𝒄𝒄-dimensional probability simplex), each 𝑖𝑖-th point of which is a 𝒄𝒄-dimensional vector 𝑝𝑝𝑖𝑖 of probabilities [regarding each of 𝒄𝒄 possible class labels from the set 𝕃𝕃={𝑙𝑙1,𝑙𝑙2,…,𝑙𝑙𝑐𝑐} assigned to the corresponding object represented by point 𝑥𝑥𝑖𝑖 within the domain space]: ∆𝑐𝑐={∀𝑝𝑝𝑖𝑖:(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)∈[0,1]𝑐𝑐|∑ 𝑝𝑝𝑘𝑘 𝑐𝑐 𝑘𝑘=1 =1}; c) ∀𝑥𝑥𝑖𝑖:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛)∈[0,1]𝑛𝑛,∃!𝑝𝑝𝑖𝑖:(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)𝑖𝑖∈∆𝑐𝑐, 𝑝𝑝𝑖𝑖=𝓕𝓕(𝑥𝑥𝑖𝑖). Simply speaking, a DONOR behaves as a labeling (classification) function (like a neural network), which takes a vector of numeric values (normalized to [0,1]) of some target object features as an input and outputs the probability distribution of that object having a particular label (i.e., belonging to a particular class) among several possible ones. Cloning is a supervised learning process, in which the DONOR is a supervisor (i.e., labels the training set samples), and the task is to train the neural network model as a CLONE of the DONOR, i.e., to learn the hidden labeling function 𝓕𝓕 defined above, as precisely as possible. For similar tasks, the traditionally used term is “knowledge distillation”. However, we are using the “cloning” term to follow consistently the terminology from our former articles where the ultimate objective was formulated as designing digital cognitive clones of humans (as decision-makers) and groups of collective hybrid intelligence (see, e.g. [12], [7], [13], [6], and [14]). Cloning-GAN architecture (see Fig. 1), is a special kind of Generative Adversarial Network (GAN) architecture (see review on GANs in [15]), and it is suggested to facilitate cloning process by smart generation of special (challenging or adversarial) samples-as-queries for the DONOR for labeling and then synchronously and incrementally train the CLONE on the basis of generated and labeled samples.
892 Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 3 Fig 1. Cloning-GAN architecture GENERATOR vs CLONE “game” in Cloning-GAN is driven by the conflicting objectives of the two synchronously trained adversaries. The objective of the GENERATOR training is to generate the most challenging (adversarial, puzzling) training samples for CLONE. The objective of the CLONE training is the capability to predict (guess, imitate) the labeling outcomes (especially biased ones) from the DONOR as close as possible in the challenging cases. The loss function for the CLONE is clear – it provides punishment feedback for the mismatch between its own outcomes and corresponding outcomes from the DONOR. Therefore, the most sophisticated task in the Cloning-GAN architecture is to define the loss function for the GENERATOR, which is naturally more complex one because it has to take into account several different criteria for the quality of generated samples. Let us provide more details on the loss functions regarding GENERATOR vs CLONE training. 2.2. CLONE loss in Cloning-GAN The loss of the CLONE in each sample 𝐱𝐱 𝒊𝒊 (denoted as 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)) is a normalized measure (𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)∈[0,1]) of the two probability distribution vectors’ mismatch (aka “Turing” mismatch): how CLONE addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕 (𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)) vs how DONOR addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕(𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)): 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)=Mismatch[𝓕𝓕 (x i)↔𝓕𝓕(x i)]= 1 √2∙𝒅𝒅(𝓕𝓕 (𝐱𝐱 𝒊𝒊),𝓕𝓕(𝐱𝐱 𝒊𝒊))=√(𝑝𝑝1−𝑝𝑝1)2+(𝑝𝑝2−𝑝𝑝2)2+⋯+(𝑝𝑝𝑐𝑐−𝑝𝑝𝑐𝑐)2 2, (1) where 𝒅𝒅 – Euclidian distance. Another, more computationally expensive but also more solid, option would be use of Jensen-Shannon Distance instead of Euclidean distance in formula (1). It is a metric [16] bounded to [0,1]; it is based on the concept of information entropy; and it is specifically designed to measure distance between probability distributions, while Euclidean distance is a general-purpose measure. Therefore, another option of the CLONE’s loss formula would be: 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ JSD(𝐱𝐱 𝒊𝒊)=𝐉𝐉𝐉𝐉𝐉𝐉(𝓕𝓕 (𝐱𝐱 𝒊𝒊),𝓕𝓕(𝐱𝐱 𝒊𝒊))=√1 2∙∑[𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘] 𝑐𝑐 𝑘𝑘=1 , (1*) where 𝐉𝐉𝐉𝐉𝐉𝐉 – Jensen-Shannon Distance defined as a metric in [16].
Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 893 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 3 Fig 1. Cloning-GAN architecture GENERATOR vs CLONE “game” in Cloning-GAN is driven by the conflicting objectives of the two synchronously trained adversaries. The objective of the GENERATOR training is to generate the most challenging (adversarial, puzzling) training samples for CLONE. The objective of the CLONE training is the capability to predict (guess, imitate) the labeling outcomes (especially biased ones) from the DONOR as close as possible in the challenging cases. The loss function for the CLONE is clear – it provides punishment feedback for the mismatch between its own outcomes and corresponding outcomes from the DONOR. Therefore, the most sophisticated task in the Cloning-GAN architecture is to define the loss function for the GENERATOR, which is naturally more complex one because it has to take into account several different criteria for the quality of generated samples. Let us provide more details on the loss functions regarding GENERATOR vs CLONE training. 2.2. CLONE loss in Cloning-GAN The loss of the CLONE in each sample 𝐱𝐱 𝒊𝒊 (denoted as 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)) is a normalized measure (𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)∈[0,1]) of the two probability distribution vectors’ mismatch (aka “Turing” mismatch): how CLONE addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕 (𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)) vs how DONOR addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕(𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)): 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)=Mismatch[𝓕𝓕 (x i)↔𝓕𝓕(x i)]= 1 √2∙𝒅𝒅(𝓕𝓕 (𝐱𝐱 𝒊𝒊),𝓕𝓕(𝐱𝐱 𝒊𝒊))=√(𝑝𝑝1−𝑝𝑝1)2+(𝑝𝑝2−𝑝𝑝2)2+⋯+(𝑝𝑝𝑐𝑐−𝑝𝑝𝑐𝑐)2 2, (1) where 𝒅𝒅 – Euclidian distance. Another, more computationally expensive but also more solid, option would be use of Jensen-Shannon Distance instead of Euclidean distance in formula (1). It is a metric [16] bounded to [0,1]; it is based on the concept of information entropy; and it is specifically designed to measure distance between probability distributions, while Euclidean distance is a general-purpose measure. Therefore, another option of the CLONE’s loss formula would be: 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ JSD(𝐱𝐱 𝒊𝒊)=𝐉𝐉𝐉𝐉𝐉𝐉(𝓕𝓕 (𝐱𝐱 𝒊𝒊),𝓕𝓕(𝐱𝐱 𝒊𝒊))=√1 2∙∑[𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘] 𝑐𝑐 𝑘𝑘=1 , (1*) where 𝐉𝐉𝐉𝐉𝐉𝐉 – Jensen-Shannon Distance defined as a metric in [16]. 4 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 See some examples below: 𝓕𝓕(𝐱𝐱 𝒊𝒊):(1,0,0);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0,0,1)⇒𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)=√(1−0)2+(0−0)2+(0−1)2 2=1; 𝓕𝓕(𝐱𝐱 𝒊𝒊):(1,0,0);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0,0,1)⇒𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ JSD(𝐱𝐱 𝒊𝒊)=√[1∙𝑙𝑙𝑙𝑙𝑙𝑙22∙1 1+0∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0 1]+[0∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0 0+0∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0 0]+[0∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0 1+1∙𝑙𝑙𝑙𝑙𝑙𝑙22∙1 1] 2→1; 𝓕𝓕(𝐱𝐱 𝒊𝒊):(0.5,0.3,0.2);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0.1,0.2,0.7)⇒𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)=√(0.5−0.1)2+(0.3−0.2)2+(0.2−0.7)2 2≈0.458; 𝓕𝓕(𝐱𝐱 𝒊𝒊):(0.5,0.3,0.2);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0.1,0.2,0.7)⇒𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ JSD(𝐱𝐱 𝒊𝒊)= =√[0.5∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.5 0.5+0.1+0.1∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.1 0.5+0.1]+[0.3∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.3 0.3+0.2+0.2∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.2 0.3+0.2]+[0.2∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.2 0.2+0.7+0.7∙𝑙𝑙𝑙𝑙𝑙𝑙22∙0.7 0.2+0.7] 2≈0.467. 2.3. GENERATOR overall loss and its components in Cloning-GAN The loss of the GENERATOR in each generated sample 𝐱𝐱 𝒊𝒊 (denoted as 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝔾𝔾(𝐱𝐱 𝒊𝒊)) is constructed from three loss components as follows: 𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝔾𝔾(𝐱𝐱 𝒊𝒊)=𝛌𝛌𝐓𝐓∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊)+𝛌𝛌𝐂𝐂∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝐱𝐱 𝒊𝒊)+𝛌𝛌𝐔𝐔∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊). (2) The “Turing” component of the GENERATOR’s loss on each generated sample 𝐱𝐱 𝒊𝒊 (denoted as 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊)) is a normalized measure (∈[0,1]) of the two probability distribution vectors’ match (aka “Turing” match): how CLONE addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕 (𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)) vs how DONOR addresses 𝐱𝐱 𝒊𝒊 (i.e., 𝓕𝓕(𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)): 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊)=Match[𝓕𝓕 (x i)↔𝓕𝓕(x i)]=𝟏𝟏−𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊)=1−√(𝑝𝑝1−𝑝𝑝1)2+(𝑝𝑝2−𝑝𝑝2)2+⋯+(𝑝𝑝𝑐𝑐−𝑝𝑝𝑐𝑐)2 2. (3) One may notice that “Turing” component of the GENERATOR’s loss on each sample 𝐱𝐱 𝒊𝒊 is opposite to the loss of the CLONE on the same sample, because one of the GENERATOR’s objectives is to generate such samples, which would be the most difficult ones for the CLONE in predicting the DONOR’s outcomes. For the same purpose, we can also use a modification of formula (3) if we apply the Jensen-Shannon Distance defined by formula (1*): 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓 JSD (𝐱𝐱 𝒊𝒊)=𝟏𝟏−𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ JSD(𝐱𝐱 𝒊𝒊)=1−√1 2∙∑[𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘∙log22∙𝑝𝑝𝑘𝑘 𝑝𝑝𝑘𝑘+𝑝𝑝𝑘𝑘] 𝑐𝑐 𝑘𝑘=1 . (3*) See some examples below: 𝓕𝓕(𝐱𝐱 𝒊𝒊):(1,0,0);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0,0,1)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝒙𝒙 𝒊𝒊)=𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓 JSD (𝐱𝐱 𝒊𝒊)=1−1=0; 𝓕𝓕(𝐱𝐱 𝒊𝒊):(0.3,0.4,0.3);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0.2,0.4,0.4)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝒙𝒙 𝒊𝒊)=1−0.1=0.9; 𝓕𝓕(𝐱𝐱 𝒊𝒊):(0.3,0.4,0.3);𝓕𝓕 (𝐱𝐱 𝒊𝒊):(0.2,0.4,0.4)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓 JSD (𝒙𝒙 𝒊𝒊)≈1−0.111=0.889. Further in the paper, we will use only terms 𝑳𝑳𝑳𝑳𝑳𝑳𝑳𝑳ℂ(𝐱𝐱 𝒊𝒊) to represent CLONE loss and 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊) to represent the “Turing” component of the GENERATOR’s loss, assuming that particular computing options for it, either (1) and (3) or (1*) and (3*), could be chosen depending on the task. The “Challenge” component of the GENERATOR’s loss on each generated sample 𝐱𝐱 𝒊𝒊 (denoted as 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝐱𝐱 𝒊𝒊)) is a normalized measure (∈[0,1]) of how easy it would be for the DONOR to confidently label the generated sample or, therefore, how far is the sample from being the “adversarial” one and being able to confuse the DONOR (i.e., how far is the probability distribution provided by DONOR on sample 𝐱𝐱 𝒊𝒊 from the uniform one): 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝐱𝐱 𝐓𝐓)=𝑐𝑐 √𝑐𝑐−1∙𝛔𝛔(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)=√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐−1 , (4) where 𝝈𝝈 – standard deviation.
894 Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 5 𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)= √ (𝑝𝑝1−1 𝑐𝑐) 2 +(𝑝𝑝2−1 𝑐𝑐) 2 +⋯+(𝑝𝑝𝑐𝑐−1 𝑐𝑐) 2 𝑐𝑐=⋯=√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐. 𝐌𝐌𝐌𝐌𝐌𝐌[𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)]=𝝈𝝈(1, 0,…,0)=𝝈𝝈(0,…,0,1)=𝝈𝝈(0, …,1,…,0)=√𝑐𝑐∙(0+⋯+1+⋯+0)−1 𝑐𝑐=√𝑐𝑐−1 𝑐𝑐. 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝐱𝐱 𝒊𝒊)= 𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐) 𝐌𝐌𝐌𝐌𝐌𝐌[𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)]=𝑐𝑐 √𝑐𝑐−1∙𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)=𝑐𝑐 √𝑐𝑐−1∙√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐=√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐−1 . Therefore, 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝐱𝐱 𝒊𝒊)∈[𝟎𝟎,𝟏𝟏] - aka normalized standard deviation 𝝈𝝈. The more standard deviation 𝝈𝝈 – the less confusion for the DONOR – the more loss for the GENERATOR. The “Challenge” component of the GENERATOR’s loss ensures appearance of challenging-and-adversarial (close to decision boundaries and corner cases) samples, helping the CLONE to learn faster the individual biases of the DONOR. See some examples below: 𝓕𝓕(𝒙𝒙 𝒊𝒊):(1,0,0)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙(12+02+02)−1 2=1; 𝓕𝓕(𝒙𝒙 𝒊𝒊):(0.5,0.3,0.2)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙(0.52+0.32+0.22)−1 2≈0.26; 𝓕𝓕(𝒙𝒙 𝒊𝒊):(1 3,1 3,1 3)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙3∙(1 3)2−1 2=0. It would be important to mention here that the “Challenge” component of the GENERATOR’s loss makes sense only in such “optimistic” cases when the DONOR responds with the so-called “soft targets”, i.e., not only indicates the “winning” class but provides the complete “ranking list” with the scores-as-probabilities for each of the classes involved (the universe here is a disointUnionOf 𝑐𝑐 classes). This case is a typical outcome from the DONOR, which is a hidden neural network classifier with disclosed output (e.g., softmax) layer. In “the-winner-takes-all” classification output cases, our probability vector (so called “hard target”) will be represented by one-hot encoded vector containing one “1” with the rest “0” and, therefore, the “challenge” component of the loss function will always be constant and equal to √𝑐𝑐−1 𝑐𝑐 (see above). However, such “blind” cases will be approached similarly (although less efficient), i.e., by applying the same formula (2) but without the “Challenge” component in it. The “Uniformity” component. One of the GENERATOR’s objectives is to generate adversarial samples “everywhere” (close to uniform distribution) within the decision space bounded by the unit hypercube. “Uniformity” component of the GENERATOR’s loss regarding each generated sample 𝐱𝐱 𝒊𝒊 (denoted as 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐂𝐂𝐔𝐔𝐔𝐔𝐋𝐋𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊)) is a normalized measure (∈[0,1]) of how far is the distribution of previously generated samples including the new one (𝐱𝐱 1,𝐱𝐱 2,…, 𝐱𝐱 𝑖𝑖−1, 𝐱𝐱 𝒊𝒊) from the uniform distribution. In fact, we compare the estimated average-nearest-neighbor- distance (𝑨𝑨𝒊𝒊) for 𝑖𝑖 samples uniformly distributed within an 𝑛𝑛-dimensional unit hypercube with the actual averagenearest-neighbour-distance (𝒂𝒂𝒊𝒊) for 𝐱𝐱 1,𝐱𝐱 2,…, 𝐱𝐱 𝑖𝑖−1, 𝐱𝐱 𝒊𝒊 generated samples (i.e., after 𝒊𝒊th generated sample 𝒙𝒙 𝒊𝒊 in 𝒏𝒏dimensional unit hypercube) as follows: 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐂𝐂𝐔𝐔𝐔𝐔𝐋𝐋𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊)=1−𝐌𝐌𝐌𝐌𝐌𝐌(𝐴𝐴𝑖𝑖,𝑎𝑎𝑖𝑖) 𝐌𝐌𝐌𝐌𝐌𝐌(𝐴𝐴𝑖𝑖,𝑎𝑎𝑖𝑖), (5) where 𝑨𝑨𝒊𝒊=𝒊𝒊−𝟏𝟏 𝒏𝒏 is a good heuristic approximation for the intended average nearest neighbor distance from [17]; 𝑎𝑎1=1; 𝑎𝑎𝑖𝑖=𝟏𝟏𝒊𝒊∙∑𝐔𝐔𝐔𝐔𝐂𝐂𝒌𝒌=𝟏𝟏,…𝒊𝒊;𝒌𝒌≠𝒋𝒋[𝒅𝒅(𝒙𝒙 𝒋𝒋,𝒙𝒙 𝒌𝒌)] 𝒊𝒊𝒋𝒋=𝟏𝟏 (𝒅𝒅 – Euclidian distance). One may see that this calculation regarding each generated sample takes into account information on previously generated samples. However, it does not consume much memory for it. See example in Fig. 2. It shows how the “Uniformity” loss is computed during the iterative process of ten samples’ generation. For each previously generated sample, this process keeps in its memory only the distance to its nearest neighbor (i.e., collecting the minimal distances needed to calculate 𝒂𝒂𝒊𝒊) and refines it when a new sample arrives. One may see these minimal distances, to be kept in memory, outlined within the distance matrixes shown in Fig 2.
Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 895 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 5 𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)=√(𝑝𝑝1−1 𝑐𝑐)2+(𝑝𝑝2−1 𝑐𝑐)2+⋯+(𝑝𝑝𝑐𝑐−1 𝑐𝑐)2 𝑐𝑐=⋯=√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐. 𝐌𝐌𝐌𝐌𝐌𝐌[𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)]=𝝈𝝈(1, 0,…,0)=𝝈𝝈(0,…,0,1)=𝝈𝝈(0, …,1,…,0)=√𝑐𝑐∙(0+⋯+1+⋯+0)−1 𝑐𝑐=√𝑐𝑐−1 𝑐𝑐. 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝐱𝐱 𝒊𝒊)= 𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐) 𝐌𝐌𝐌𝐌𝐌𝐌[𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)]=𝑐𝑐 √𝑐𝑐−1∙𝝈𝝈(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐)=𝑐𝑐 √𝑐𝑐−1∙√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐=√𝑐𝑐∙(𝑝𝑝1 2+𝑝𝑝2 2+⋯+𝑝𝑝𝑐𝑐2)−1 𝑐𝑐−1 . Therefore, 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝐱𝐱 𝒊𝒊)∈[𝟎𝟎,𝟏𝟏] - aka normalized standard deviation 𝝈𝝈. The more standard deviation 𝝈𝝈 – the less confusion for the DONOR – the more loss for the GENERATOR. The “Challenge” component of the GENERATOR’s loss ensures appearance of challenging-and-adversarial (close to decision boundaries and corner cases) samples, helping the CLONE to learn faster the individual biases of the DONOR. See some examples below: 𝓕𝓕(𝒙𝒙 𝒊𝒊):(1,0,0)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙(12+02+02)−1 2=1; 𝓕𝓕(𝒙𝒙 𝒊𝒊):(0.5,0.3,0.2)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙(0.52+0.32+0.22)−1 2≈0.26; 𝓕𝓕(𝒙𝒙 𝒊𝒊):(1 3,1 3,1 3)⇒𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂(𝒙𝒙 𝒊𝒊)=√3∙3∙(1 3)2−1 2=0. It would be important to mention here that the “Challenge” component of the GENERATOR’s loss makes sense only in such “optimistic” cases when the DONOR responds with the so-called “soft targets”, i.e., not only indicates the “winning” class but provides the complete “ranking list” with the scores-as-probabilities for each of the classes involved (the universe here is a disointUnionOf 𝑐𝑐 classes). This case is a typical outcome from the DONOR, which is a hidden neural network classifier with disclosed output (e.g., softmax) layer. In “the-winner-takes-all” classification output cases, our probability vector (so called “hard target”) will be represented by one-hot encoded vector containing one “1” with the rest “0” and, therefore, the “challenge” component of the loss function will always be constant and equal to √𝑐𝑐−1 𝑐𝑐 (see above). However, such “blind” cases will be approached similarly (although less efficient), i.e., by applying the same formula (2) but without the “Challenge” component in it. The “Uniformity” component. One of the GENERATOR’s objectives is to generate adversarial samples “everywhere” (close to uniform distribution) within the decision space bounded by the unit hypercube. “Uniformity” component of the GENERATOR’s loss regarding each generated sample 𝐱𝐱 𝒊𝒊 (denoted as 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐂𝐂𝐔𝐔𝐔𝐔𝐋𝐋𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊)) is a normalized measure (∈[0,1]) of how far is the distribution of previously generated samples including the new one (𝐱𝐱 1,𝐱𝐱 2,…, 𝐱𝐱 𝑖𝑖−1, 𝐱𝐱 𝒊𝒊) from the uniform distribution. In fact, we compare the estimated average-nearest-neighbor- distance (𝑨𝑨𝒊𝒊) for 𝑖𝑖 samples uniformly distributed within an 𝑛𝑛-dimensional unit hypercube with the actual averagenearest-neighbour-distance (𝒂𝒂𝒊𝒊) for 𝐱𝐱 1,𝐱𝐱 2,…, 𝐱𝐱 𝑖𝑖−1, 𝐱𝐱 𝒊𝒊 generated samples (i.e., after 𝒊𝒊th generated sample 𝒙𝒙 𝒊𝒊 in 𝒏𝒏dimensional unit hypercube) as follows: 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐂𝐂𝐔𝐔𝐔𝐔𝐋𝐋𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊)=1−𝐌𝐌𝐌𝐌𝐌𝐌(𝐴𝐴𝑖𝑖,𝑎𝑎𝑖𝑖) 𝐌𝐌𝐌𝐌𝐌𝐌(𝐴𝐴𝑖𝑖,𝑎𝑎𝑖𝑖), (5) where 𝑨𝑨𝒊𝒊=𝒊𝒊−𝟏𝟏 𝒏𝒏 is a good heuristic approximation for the intended average nearest neighbor distance from [17]; 𝑎𝑎1=1; 𝑎𝑎𝑖𝑖=𝟏𝟏𝒊𝒊∙∑𝐔𝐔𝐔𝐔𝐂𝐂𝒌𝒌=𝟏𝟏,…𝒊𝒊;𝒌𝒌≠𝒋𝒋[𝒅𝒅(𝒙𝒙 𝒋𝒋,𝒙𝒙 𝒌𝒌)] 𝒊𝒊𝒋𝒋=𝟏𝟏 (𝒅𝒅 – Euclidian distance). One may see that this calculation regarding each generated sample takes into account information on previously generated samples. However, it does not consume much memory for it. See example in Fig. 2. It shows how the “Uniformity” loss is computed during the iterative process of ten samples’ generation. For each previously generated sample, this process keeps in its memory only the distance to its nearest neighbor (i.e., collecting the minimal distances needed to calculate 𝒂𝒂𝒊𝒊) and refines it when a new sample arrives. One may see these minimal distances, to be kept in memory, outlined within the distance matrixes shown in Fig 2. 6 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 Fig 2. Example illustrating “Uniformity” component of Cloning-GAN (based on iterative computing regarding 10 generated samples) 2.4. Trade-off among the GENERATOR’s loss components Performance of the GENERATOR and, therefore, the whole Cloning-GAN will depend on the choice of the importance weights 𝝀𝝀𝑻𝑻,𝝀𝝀𝑪𝑪,𝝀𝝀𝑼𝑼, which correspond to different components of the GENERATOR’s loss from formula (2). Let us consider different approaches to balance with these weights, which will correspond to different options of the Cloning-GAN architecture. The Basic Cloning-GAN (B-C-GAN) architecture supposes manual initialization and control of weights 𝝀𝝀𝑻𝑻,𝝀𝝀𝑪𝑪,𝝀𝝀𝑼𝑼 (assuming that 𝝀𝝀𝑻𝑻+𝝀𝝀𝑪𝑪+𝝀𝝀𝑼𝑼=𝟏𝟏), which will be considered as constants until changed manually if needed. The Dynamic Cloning-GAN (D-C-GAN) architecture assumes the dependence of the weights, which correspond to different components of the GENERATOR’s loss from formula (2), on the current learning iteration (epoch) 𝒊𝒊 as follows:
896 Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 7 𝝀𝝀𝑻𝑻(𝒊𝒊)+𝝀𝝀𝑪𝑪 (𝒊𝒊)+𝝀𝝀𝑼𝑼(𝒊𝒊)=1 ; 𝝀𝝀 𝑻𝑻 (𝒊𝒊)=0.5 , i.e., constant, which takes half of the overall importance always during training; 𝝀𝝀𝑪𝑪(𝒊𝒊)=𝑖𝑖 2∙(𝑖𝑖+𝑖𝑖∗), i.e., monotonously increasing importance from 0 (at the beginning) to 0.5 (at infinity); 𝝀𝝀𝑼𝑼(𝒊𝒊)= 𝑖𝑖∗ 𝟐𝟐∙(𝒊𝒊+𝒊𝒊∗), i.e., monotonously decreasing importance from 0.5 (at the beginning) to 0 (at infinity), where 𝒊𝒊∗ is the controlling parameter (integer >1, which indicates at which iteration 𝝀𝝀𝑪𝑪=𝝀𝝀𝑼𝑼=0.25). See illustrative explanation in Fig. 3. Fig 3. Plots show a trade-off among the importance of the GENERATOR’s loss function components during training. Therefore, for D-C-GAN, the loss of the GENERATOR from formula (2) on each generated sample 𝐱𝐱 𝒊𝒊 could be updated as follows: 𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝔾𝔾(𝐱𝐱 𝒊𝒊)=0.5∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊)+𝑖𝑖 2∙(𝑖𝑖+𝑖𝑖∗)∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝐱𝐱 𝒊𝒊)+𝑖𝑖∗ 𝟐𝟐∙(𝒊𝒊+𝒊𝒊∗)∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊). (6) Training process performance regarding the GENERATOR and, therefore, the CLONE can be controlled just by one parameter 𝑖𝑖∗ depending on the cloning case specifics. We consider the “Turing” component to be equally important (50% of the overall importance) during the training process because it directly influences the performance of the CLONE to simulate the DONOR closely enough. We assume that the ability of the GENERATOR to generate “diverse” (everywhere within the decision space) rather than “difficult” (challenging, adversarial) inputs for the DONOR-CLONE couple is more important at the beginning at the training process, and the trend towards more challenging rather than different inputs will become influential later in the training process. The trend itself (i.e., how fast the capability to generate challenging inputs will become as important as the capability to generate diverse inputs) could be controlled by parameter 𝒊𝒊∗. The reasonability for dynamic weights of different loss components is yet to be checked experimentally. The intuition is that Cloning-GANs (of the D-C-GAN option) may converge better when some criterion is dominating at the beginning of the process and another one - at the end. GENERATOR, like a kind of “young journalist”, in the beginning, learns to ask “different” questions and, when mature with this skill, switches to training his own capability of asking “difficult” questions also. The “Three Musketeers” (“All for one and one for all”) Cloning-GAN (3M-C-GAN) architecture differs significantly from the previous two architectures. It supposes splitting the GENERATOR to three GENERATORs (“Turing” or 𝔾𝔾𝐓𝐓, “Challenge” or 𝔾𝔾𝐂𝐂, and “Uniformity” or 𝔾𝔾𝐔𝐔), one responsible for each loss component from formula (2). Before that, we considered only the incremental learning option for the Cloning-GAN architecture, which could be extended to the “minibatch” learning case when all the loss components are computed as aggregates over few
Vagan Terziyan et al. / Procedia Computer Science 232 (2024) 890–902 897 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 7 𝝀𝝀𝑻𝑻(𝒊𝒊)+𝝀𝝀𝑪𝑪 (𝒊𝒊)+𝝀𝝀𝑼𝑼(𝒊𝒊)=1; 𝝀𝝀𝑻𝑻(𝒊𝒊)=0.5, i.e., constant, which takes half of the overall importance always during training; 𝝀𝝀𝑪𝑪(𝒊𝒊)=𝑖𝑖 2∙(𝑖𝑖+𝑖𝑖∗), i.e., monotonously increasing importance from 0 (at the beginning) to 0.5 (at infinity); 𝝀𝝀𝑼𝑼(𝒊𝒊)= 𝑖𝑖∗ 𝟐𝟐∙(𝒊𝒊+𝒊𝒊∗), i.e., monotonously decreasing importance from 0.5 (at the beginning) to 0 (at infinity), where 𝒊𝒊∗ is the controlling parameter (integer >1, which indicates at which iteration 𝝀𝝀𝑪𝑪=𝝀𝝀𝑼𝑼=0.25). See illustrative explanation in Fig. 3. Fig 3. Plots show a trade-off among the importance of the GENERATOR’s loss function components during training. Therefore, for D-C-GAN, the loss of the GENERATOR from formula (2) on each generated sample 𝐱𝐱 𝒊𝒊 could be updated as follows: 𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝐿𝔾𝔾(𝐱𝐱 𝒊𝒊)=0.5∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝐱𝐱 𝒊𝒊)+𝑖𝑖 2∙(𝑖𝑖+𝑖𝑖∗)∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝐱𝐱 𝒊𝒊)+𝑖𝑖∗ 𝟐𝟐∙(𝒊𝒊+𝒊𝒊∗)∙𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝐱𝐱 𝒊𝒊). (6) Training process performance regarding the GENERATOR and, therefore, the CLONE can be controlled just by one parameter 𝑖𝑖∗ depending on the cloning case specifics. We consider the “Turing” component to be equally important (50% of the overall importance) during the training process because it directly influences the performance of the CLONE to simulate the DONOR closely enough. We assume that the ability of the GENERATOR to generate “diverse” (everywhere within the decision space) rather than “difficult” (challenging, adversarial) inputs for the DONOR-CLONE couple is more important at the beginning at the training process, and the trend towards more challenging rather than different inputs will become influential later in the training process. The trend itself (i.e., how fast the capability to generate challenging inputs will become as important as the capability to generate diverse inputs) could be controlled by parameter 𝒊𝒊∗. The reasonability for dynamic weights of different loss components is yet to be checked experimentally. The intuition is that Cloning-GANs (of the D-C-GAN option) may converge better when some criterion is dominating at the beginning of the process and another one - at the end. GENERATOR, like a kind of “young journalist”, in the beginning, learns to ask “different” questions and, when mature with this skill, switches to training his own capability of asking “difficult” questions also. The “Three Musketeers” (“All for one and one for all”) Cloning-GAN (3M-C-GAN) architecture differs significantly from the previous two architectures. It supposes splitting the GENERATOR to three GENERATORs (“Turing” or 𝔾𝔾𝐓𝐓, “Challenge” or 𝔾𝔾𝐂𝐂, and “Uniformity” or 𝔾𝔾𝐔𝐔), one responsible for each loss component from formula (2). Before that, we considered only the incremental learning option for the Cloning-GAN architecture, which could be extended to the “minibatch” learning case when all the loss components are computed as aggregates over few 8 Vagan Terziyan et al. / Procedia Computer Science 00 (2023) 000–000 generated samples. In particular, the 3M-C-GAN architecture assumes that, at 𝑖𝑖 -th iteration, each of three GENERATORs independently generates one sample each and, therefore, resulting in a minibatch of three samples {𝒙𝒙 𝒊𝒊𝑻𝑻,𝒙𝒙 𝒊𝒊𝑪𝑪,𝒙𝒙 𝒊𝒊𝑼𝑼} as shown in Fig. 4. This minibatch goes through the same process as in normal Cloning-GAN described above so that the component losses are aggregated and distributed among the GENERATORs as follows: 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓({𝒙𝒙 𝒊𝒊𝑻𝑻,𝒙𝒙 𝒊𝒊𝑪𝑪,𝒙𝒙 𝒊𝒊𝑼𝑼})=1 3∙(𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝒙𝒙 𝒊𝒊𝑻𝑻)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝒙𝒙 𝒊𝒊𝑪𝑪)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓(𝒙𝒙 𝒊𝒊𝑼𝑼)); (7) 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂({𝒙𝒙 𝒊𝒊𝑻𝑻,𝒙𝒙 𝒊𝒊𝑪𝑪,𝒙𝒙 𝒊𝒊𝑼𝑼})=1 3∙(𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝒙𝒙 𝒊𝒊𝑻𝑻)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝒙𝒙 𝒊𝒊𝑪𝑪)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐂𝐓𝐓𝐓𝐓𝐂𝐂(𝒙𝒙 𝒊𝒊𝑼𝑼)); (8) 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔({𝒙𝒙 𝒊𝒊𝑻𝑻,𝒙𝒙 𝒊𝒊𝑪𝑪,𝒙𝒙 𝒊𝒊𝑼𝑼})=1 3∙(𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝒙𝒙 𝒊𝒊𝑻𝑻)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝒙𝒙 𝒊𝒊𝑪𝑪)+𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐔𝐔𝐓𝐓𝐓𝐓𝐔𝐔𝐋𝐋𝐓𝐓𝐔𝐔𝐓𝐓𝐔𝐔𝐔𝐔(𝒙𝒙 𝒊𝒊𝑼𝑼)). (9) Fig 4. Architecture of the “Three Musketeers” (“All for one and one for all”) Cloning-GAN (3M-C-GAN). The main feature of 3M-C-GAN architecture is that each of the GENERATORs, after generating just one sample, will be punished for the loss created by all the minibatch (i.e., will be responsible for “colleagues’” performance also). For example, all the loss 𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝐋𝔾𝔾𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓𝐓({𝒙𝒙 𝒊𝒊𝑻𝑻,𝒙𝒙 𝒊𝒊𝑪𝑪,𝒙𝒙 𝒊𝒊𝑼𝑼}) computed according to formula (7) will be backpropagated through the “Turing” GENERATOR 𝔾𝔾𝐓𝐓, although it is actually responsible only for one sample of the three. In this way, the architecture ensures that all the three GENERATORs are responsible for each one and each one is responsible for all. Therefore, we used “Three Musketeers” (with the motto “All for one and one for all”) as a name for such an architecture. One of the advantages of this architecture is that we do not need to worry about the weights 𝝀𝝀𝑻𝑻,𝝀𝝀𝑪𝑪,𝝀𝝀𝑼𝑼 from formula (2) because each GENERATOR addresses independently only its own specific loss criteria. We can also expect better convergence of such a training process managed by multiple GENERATORs compared with one GENERATOR with complex and conflicting criteria. If necessary, the missing “D’Artagnan” from the “Three Musketeers” architecture can be enabled as an additional (fourth) generation quality objective and, therefore, an additional GENERATOR to be responsible for the so-called “hostility” of generated samples. Such a component could enable not only challenging content for the DONOR, which is the duty of the “Challenge” component, but rather to ensure that more challenging (“hostile”) areas in data space will be covered more with the diverse-by-labeling generated samples. Assume that: 𝑛𝑛-dimensional sample 𝐱𝐱 𝑁𝑁𝑁𝑁:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛), which belongs to previously generated samples (𝐱𝐱 1,𝐱𝐱 2,…, 𝐱𝐱 𝑖𝑖−1) and labeled (to 𝑐𝑐 classes) by the DONOR as 𝓕𝓕(𝐱𝐱 𝑁𝑁𝑁𝑁):(𝑝𝑝1,𝑝𝑝2,,…,𝑝𝑝𝑐𝑐), is the nearest neighbor (NN) of the newly generated sample 𝐱𝐱 𝒊𝒊:(𝑥𝑥1,𝑥𝑥2,…,𝑥𝑥𝑛𝑛) labeled by the DONOR as 𝓕𝓕(𝐱𝐱 𝒊𝒊):(𝑝𝑝1,𝑝𝑝2,…,𝑝𝑝𝑐𝑐); The “uncertainty gain” of 𝐱𝐱 𝒊𝒊 in comparison with its nearest neighbor 𝐱𝐱 𝑁𝑁𝑁𝑁 is the following asymmetric measure: