scieee AI-readable full text Open interactive document viewer

Federated Learning and Differential Privacy: Software tools analysis, the Sherpa.ai FL framework and methodological guidelines for preserving data privacy

Rodríguez Barroso, Nuria,Jiménez López, Daniel,Ruiz Millán, José Antonio,Martínez Cámara, Eugenio,Luzón García, María Victoria,Herrera Triguero, Francisco

Abstract

SHERPA Europe S.L. OTRI-4137

Full text

Federated Learning and Differential Privacy: Software tools analysis, the Sherpa.ai FL framework and methodological guidelines for preserving data privacy Nuria Rodríguez-Barrosoa, Goran Stipcichb, Daniel Jiménez-Lópeza, José Antonio Ruiz-Millána, Eugenio Martínez-Cámaraa,∗, Gerardo González-Secob, M. Victoria Luzóna, Miguel Ángel Veganzonesb, Francisco Herreraa aAndalusian Research Institute in Data Science and Computational Intelligence, University of Granada, Spain bSherpa.ai, Bilbao, Spain Abstract The high demand of artificial intelligence services at the edges that also preserve data privacy has pushed the research on novel machine learning paradigms that fit these requirements. Federated learning has the ambition to protect data privacy through distributed learning methods that keep the data in its storage silos. Likewise, differential privacy attains to improve the protection of data privacy by measuring the privacy loss in the communication among the elements of federated learning. The prospective matching of federated learning and differential privacy to the challenges of data privacy protection has caused the release of several software tools that support their functionalities, but they lack a unified vision of these techniques, and a methodological workflow that supports their usage. Hence, we present the Sherpa.ai Federated Learning framework that is built upon a holistic view of federated learning and differential privacy. It results from both the study of how to adapt the machine learning paradigm to federated learning, and the definition of methodological guidelines for developing artificial intelligence services based on federated learning and differential privacy. We show how to follow the methodological guidelines with the Sherpa.ai Federated Learning framework by means of a classification and a regression use cases. Keywords: federated learning, differential privacy, software framework, Sherpa.ai Federated Learning framework ∗Corresponding author Email addresses: [email protected] (Nuria Rodríguez-Barroso), [email protected] (Goran Stipcich), [email protected] (Daniel Jiménez-López), [email protected] (José Antonio Ruiz-Millán), [email protected] (Eugenio Martínez-Cámara), [email protected] (Gerardo González-Seco), [email protected] (M. Victoria Luzón), [email protected] (Miguel Ángel Veganzones), [email protected] (Francisco Herrera) arXiv:2007.00914v2 [cs.LG] 6 Oct 2020 1. Introduction The last advances in fundamental and applied research in artificial intelligence (AI) has aroused interest in industry and end users. This interest goes beyond the traditional centralised setting of AI, and nowadays there is a high demand of AI services at the edges. One of the main pillars of AI is data, whose larger availability has boosted the progress of AI in the last years. However, data is a sensitive element, especially when it describes users’ personal features, such as clinical or financial data. This sensitive nature of personal data has raised the awareness of end users on data privacy protection, promoting the publication of legal frames [1] and recommendations for developing AI services that preserve data privacy [2]. In this context, the progress of AI applications is based on (1) using data generated or stored at the edges, (2) working with large amounts of data from a wide range of sources, and (3) protecting data privacy in order to comply with the legal restrictions and to pay attention to end users’ concerns. Some use cases of AI with these dependencies are: • When data contains sensitive information, such as email accounts, personalised recommendations or health records, applications should employ privacy-preserving techniques to learn from a population of users whilst keeping the sensitive information on each user’s device [3]. • When information is located in data silos, for instance, healthcare industry is usually reluctant to disclose its records, keeping it as sequestered data [4]. Nevertheless, joint learning from data silos of different health institutions would allow to improve the robustness of the resulting models. • Due to data privacy legislation, banks [5] and telecom [6] companies cannot share individual records. However they would benefit from models that learn from several entities’ data. The standard machine learning paradigm does not match the previous dependencies, as it learns from a centralised data source. Likewise, distributed machine learning does not fit the preserving data privacy challenge, because data is shared among several computational elements. Moreover, distributed machine learning cannot cope with the challenges associated to decentralised data processing, such as the ability to work with a great amount of clients with non homogeneous data distributions [7]. Federated learning (FL) is a nascent machine learning paradigm where many clients, in the sense of electronic devices or entire organisations, jointly train a model under the orchestration of a central server, while keeping the training data decentralised [8]. Roughly speaking, data is not shared with the central server, indeed it is kept in the devices where it is stored or generated. Accordingly, FL addresses the challenges of developing AI services on scattered data across a large amount of clients with non homogeneous data distributions. 2 Maintaining the data in its corresponding storage silos does not completely assure privacy preservation, since several adversarial attacks can still be damaging [9]. Data obfuscation, anonymisation techniques, such as blindly trusting artificial intelligence black box models (i.e. convolutional neural networks), or randomly sampling data from the clients’ models have been proven to be inadequate to preserve privacy [10,11]. Moreover, the complete obfuscation of the data greatly reduces its value, thus a balance between privacy and utility is needed. Differential privacy (DP) is proposed as a data access technique which aims to maintain personal data privacy while maximising its utility [12]. The characteristics of FL and DP, and by extension their combination, make them candidates to address the challenges of distributed AI services that preserve data privacy. The research and progress of FL and DP need the support of software tools that ease the design of privacy-preserving AI services while not requiring development from scratch. Consequently, in recent years several software tools with FL and DP functionalities have been released with this aim. We perform a comparative analysis of the FL and DP software tools released so far, and we conclude that their lack of a holistic view of FL and DP hinders the development of unified FL and DP AI services, as well as the furtherance of addressing the challenges of AI services at the edges that preserve data privacy. Therefore, we present the Sherpa.ai Federated Learning framework,1,2an open-source unified FL and DP framework for AI. Sherpa.ai FL aims to bridge the gap between the fundamental and applied research. Moreover, it will facilitate open research and development of new solutions built upon FL and DP for the challenges posed by AI at the edges and data privacy protection. A flexible approach to a wide range of problems is assured by its modular design that takes into account all the key elements and functionalities of FL and DP, which consist of: 1. Data. Different data sets can be processed. 2. Learning model. Several core machine learning algorithms are incorporated. 3. Aggregation operator. Different operators for fusing the parameters of the clients’ learning models are embodied. 4. Clients. It is where the learning models are run. 5. Federated server. The clients can be orchestrated by different communication strategies. 6. Communication among clients and server. Different solutions are encompassed to reduce the communication iterations, to protect the learning from adversarial attacks, and to obfuscate the parameters with DP techniques. 7. DP mechanisms. The fundamental DP mechanisms, such as the Laplace mechanism, as well as the composition of DP mechanisms are incorporated. The progress of AI is not only supported by the release of software tools, but it needs fun1https://developers.sherpa.ai/privacy-technology/ 2https://github.com/sherpaai/Sherpa.ai-Federated-Learning-Framework 3 damental guidelines defining how to put together the different software tools’ attributes for reaching the intended learning goal while at same time matching the problem restrictions. Accordingly, since FL is a machine learning paradigm, we first study the principles of machine learning and how to make them fit the FL requirements. We see that most machine learning methods can be directly adapted to a FL setting, but some of them require ad-hoc amendments. As a result of this study, we define the experimental workflow of FL in terms of methodological guidelines for preserving data privacy in the development of AI services at the edges. These methodological guidelines are grounded in the machine learning workflow, and they have guided the design and development of Sherpa.ai FL, therefore they can be followed with Sherpa.ai FL. It is shown how to follow the mentioned methodological guidelines with Sherpa.ai FL through two examples encompassing a classification and a regression use cases, namely: 1. Classification. We use the EMNIST Digits dataset to describe how to conduct a classification task with Sherpa.ai FL. We also compare the federated classification with its centralised counterpart. Both approaches achieve similar results. 2. Regression. We describe how to perform a regression experiment using the California Housing dataset. We compare the FL experiment with its centralised version. In addition, it is shown how to assess and limit the privacy loss using DP. The main contributions of this paper are: 1. To analyse the most recently released FL and DP software tools, revealing the lack a unified view of FL and DP that hinders the possibility of addressing the challenges of AI at the edges with data privacy. 2. To present Sherpa.ai FL, an open-source unified FL and DP framework for AI. 3. To study the adaptation of machine learning models to the principles of FL and, accordingly, to define the methodological guidelines which can be followed with Sherpa.ai FL for developing AI services that preserve data privacy with FL and DP. The rest of the paper is organised as follows: the next section formally defines FL and DP as well as their key elements. Section 3analyses the main FL and DP frameworks’ features. Section 4introduces Sherpa.ai FL including software architecture and functionalities. Section 5explains the adaptation of the machine learning paradigm to FL, taking into account the adaptation of core algorithms and the methodological guidelines of an experimental workflow. Section 6shows some illustrative examples consisting in a classification and regression problem. Finally, the concluding remarks and future work are reported in Section 7. 4 2. Federated Learning and Differential Privacy The development of a framework for FL and DP requires a thorough understanding of what FL is and what its key elements are. Accordingly, we formally define FL in Section 2.1, and we detail each key element of a FL scheme in Section 2.2. Similarly, DP is defined in Section 2.3, and its key elements are described in Section 2.4. 2.1. The definition of Federated Learning FL is a distributed machine learning paradigm that consists of a network of nodes where we distinguish two types of nodes: (1) Data owner nodes, {C1, . . . , Cn}, that possess a collection of data, {D1, . . . , Dn}, and (2) Aggregation nodes, {G1, . . . , Gk}, aiming at learning a model from data owners. The deployment of these two types of nodes defines, at least, two kind of federated architectures according to Yang et al. [13], namely: 1. Peer-to-peer: It is the architecture in which all the nodes are both Data owner and Aggregation nodes. This scheme does not require a coordinator. The main advantages are the elevated security and data privacy while the main disadvantage is the computation cost. This FL architecture is illustrated in Figure 1. 2. Client-server: It consists of a coordinator Aggregation node named server and a set of Data owner nodes named clients.1In this architecture, the client does not share its local data ensuring its privacy. We represent the client-server scheme in Figure 2. Data Owner A Data Owner B Aggregation Node A Aggregation Node B Model update Model update Model update Model update Updated Model Updated Model Model update WB Model update WA Figure 1: Representation of peer-to-peer FL architecture. 1In the literature we find different ways to refer to the clients in a FL architecture, namely: nodes, agents or clients. In this paper, we rather prefer the term clients. 5 Server (Aggregation node) Client A (Data Owner node) Client B (Data Owner node) Client C (Data Owner node) Model update Model update Model update Updated Model Updated Model Updated Model Figure 2: Representation of client-server FL architecture. Since the peer-to-peer model is a generalisation of the client-server model, we consider the latter for the formal definition of FL. In this architecture, each of the clients Cihas a local learning model LLMirepresented by the parameters θi. FL aims at learning the global learning model GLM, represented by θ, using the scattered data across clients through an iterative learning process known as round of learning. For that purpose, in each round of learning t, each client trains the LLMiover its local training data Dt i, updating their local parameters θt i. Subsequently, the global parameters θtare computed aggregating the local parameters {θt 1, . . . , θt n}using a specific federated aggregation operator ∆: θt=∆(θt 1,θt 2, . . . , θt n)(1) After the aggregation of the parameters in the GLM, the LLMs are updated with the aggregated parameters: θt+1 i←θt,∀i∈ {1, . . . , n}(2) The communication between server and clients can be synchronous or asynchronous. In the first option, the server awaits for the clients updates, aggregates all the local parameters and sends them to each client. Nevertheless, in the second option, the server merges the local parameters with the GLM as soon as it receives them, using a weighted scheme based on the age difference among the models. 6 We repeat this iterative process for as many rounds of learning as needed. Thus, the final value of θwill sum up the clients’ underlying knowledge. In particular, the learning goal is typically to minimise the following objective function: min θF(θ), with F(θ):= n ∑ i=1 wiFi(θ)(3) where nis the number of clients, Fiis the local objective function for the i-th client which is the common objective function of the problem fitted to each client’s data, wi≥0 and ∑iwi=1. 2.2. Key elements of Federated Learning The development of a FL environment requires the right combination of a set of necessary key elements. Since FL is a specific configuration of a machine learning environment, FL shares some key elements with it, namely: (1) data and (2) the learning model. However, the particularities of FL make necessary additional key elements, such as: (1) federated aggregation operators, (2) clients, (3) federated server and (4) communication among the federated server and the clients. The adaptation of the common key elements among FL and machine learning, and the FL specific ones are described as what follows. Data.Data plays a central role in FL as in machine learning. The distribution of data becomes crucial in FL since it is distributed among the different clients. Regarding the splitting of the data among clients, there are two possibilities depending on the data distribution: •IID (Independent and Identically Distributed) scenario: when the data distribution in each client corresponds to the population data distribution. In other words, the data in each client is independent and identically distributed, as well as representative of the population data distribution. •Non-IID (non Independent and Identically Distributed) scenario: when the data distribution in each client is not independent or identically distributed from the population data distribution. In a real FL scenario, each client only stores the data generated on the client itself, ensuring the non-IID property of the global data. Hence, the non-IID scenario is the most likely one and it represents a real challenge for FL. Learning model.The learning model is the shared structure between the server and the clients, where each client trains a local model using its own data, while the global model on the server is never trained, but instead it is obtained aggregating the clients’ model parameters. Thus multiple models are trained without explicit data sharing, a configuration that is essentially different from the classical (centralised) learning paradigm. 7 Federated aggregation operators.The aggregation operator is in charge of aggregating the parameters in the server. It has to: (1) assure a proper fusion of the local learning models in order to optimise the objective function in Equation 3; (2) reduce the number of communication rounds among the clients and the federated server and (3) be robust against clients with poor data quality or malicious clients. Some of the most commonly used federated aggregation operators in the literature are: •Federated Averaging (FedAvg) [14]. It is based on keeping a shared global model that is periodically updated by averaging models that have been trained locally on clients. The training process is arranged by a central server which hosts the shared global model. However, the actual optimisation is done locally on clients. •CO-OP [15]. It proposes an asynchronous approach, which merges any received client model with the global model. Instead of directly averaging the models, the merging between a local model and the global model is carried out using a weighting scheme based on a measure of the difference in the age of the models. This is motivated by the fact that in an asynchronous framework, some clients will be trained on obsolete data while others will be trained on more up-to-date data. Clients.Each client of a federated scenario represents a node of the distributed scheme. Typical clients in FL could be smartphones, IoT devices or connected vehicles. Each client owns its specific training dataset and its local model. Their principal aim is to train local models on their own private data and share the trained model parameters with the federated server where the parameters fusion is performed. Federated server.The federated server orchestrates the iterative learning of FL, which is composed of several rounds of learning. The server participates in: (1) receiving the trained parameters of the local models, (2) aggregating the trained parameters of each client model using federated aggregation operators and (3) updating every learning model with the aggregated parameters. The learning process involved in both training the local models and updating them, is known as a round of learning. The global model, which is stored in the federated server, represents the final model after the learning process. Therefore it is used for predicting, testing or any posterior evaluation. Communication among the federated server and the clients.Communication between clients and server is the most tricky element of a FL scheme. On the one hand, an efficient communication is a crucial requirement due to the high communication times needed because of the network speed limitations and availability. For that reason, FL should minimise communications and maximise their efficiency by means of, for example, reducing the number of rounds of learning. On the other hand, the interchange of model parameters between the server and the clients constitutes a vulnerability to the federated server scheme, since the original data may be reconstructed from the model parameters through model-inversion adversarial attacks [16], resulting in a great risk of private data leakage. For this reason, DP techniques [12] are commonly used in order to share model parameters [17]. 8 2.3. The definition of Differential Privacy DP is the property of an algorithm whose input is typically a database, and whose encoded response allows to obtain relatively accurate answers to potential queries [12,18]. The motivation for DP stems from the necessity of ensuring the privacy of individuals whose sensitive details are part of a database, while at the same time being able to gain accurate knowledge about the whole population when learning from the database. DP does not imply a binary concept, i.e., the guarantee or not of an individual’s data privacy. Instead DP establishes a formal measure of privacy loss, allowing for comparison between different approaches. Thus, DP will rigorously bound the possible harm to an individual whose sensitive information belongs to the database by fixing a budget for privacy loss. The formal definition of DP requires a few preliminary notions. Namely, we define the probability simplex over a discrete set B, denoted ∆(B), as the set of real valued vectors whose |B|components sum up to one and are non-negative: ∆(B):=(x∈R|B|: |B| ∑ i=1 xi=1, xi≥0, i=0, . . . , |B|)(4) Arandomised algorithm M:A→B, with Ba discrete set, is defined as a mechanism which is associated with a mapping M:A→∆(B)such that, with input a∈A, the mechanism produces M(a) = bwith probability (M(a))b=P(b|a), for each b∈B. The probability is taken over the randomness employed by the mechanism M. In general, databases are collections of records from a universe X. It is convenient to express databases xby their histogram x∈N|X |, where each component xistands for the number of elements in the database of type i in X. This interpretation naturally leads to define the distance between databases: two databases x,yare said to be n-neighbouring if they differ by nentries as ||x−y||1=n, where || · ||1 denotes the `1norm. In particular, if the databases only differ in a single data element (n=1), the databases are simply addressed as neighbouring. At this stage, DP can be formally introduced. A randomised algorithm (mechanism) M with domain N|X | preserves e-DP for e>0 if for all neighbouring databases x,y∈N|X | and all S ⊆ Range(M)it holds that: P[M(x)∈ S]≤exp(e)P[M(y)∈ S](5) If, on the other hand, for 0 <δ<1 it holds that: P[M(x)∈ S]≤exp(e)P[M(y)∈ S] + δ(6) then the mechanism possesses the weaker property of (e,δ)-DP. The probability is taken over the randomness employed by the mechanism M. In essence, Equation 5tells us that for every run of the randomisation mechanism M(x), it is almost equally likely to observe the same output for every neighbouring database y, 9 3.1. PySyft PySyft1is a Python library for secure and private deep learning. PySyft decouples private data from model training, using FL, DP, and Encrypted Computation (like Multi-Party Computation (MPC) and Homomorphic Encryption (HE)) within the main deep learning frameworks like PyTorch and TensorFlow. Features.It is compatible with existing deep learning frameworks such as TensorFlow and PyTorch. Their low level FL implementation allows developing and debugging projects with complex communication networks in a local environment with almost no overhead. It is mainly focused on providing Secure MPC through HE, it thus allows to apply computations on ciphertext which is ideal for developing FL models while preserving privately the results of the computations to the participants. Last, they offer many Python notebooks, which greatly softens the learning curve of this framework. Shortcomings.Its low level of FL support is missing some key features: neither it includes any dataset by default nor it implements any model aggregation operators. Its low level implementation and the two drawbacks stated before make this framework quite complex to use, requiring considerable knowledge in this field to correctly assemble a FL model. While its webpage2advertises many DP mechanisms, they are nowhere to be found. As a matter of fact, in their github documentation they state the following: “Do NOT use this code to protect data (private or otherwise) - at present it is very insecure. Come back in a couple of months”. Overview.We conclude that PySyft is a low level FL framework for advanced users which is compatible with many well-known deep learning frameworks and it does provide neither any DP mechanism nor any DP algorithm. 3.2. TensorFlow TensorFlow implements DP and FL through its libraries TensorFlow Privacy and TensorFlow Federated, respectively. Features.TensorFlow Privacy3is a Python library for training machine learning models with privacy for training data. It integrates seamlessly with existing TensorFlow models and allows the developer to train its models with DP techniques. In addition they have many tutorials to quickly learn how to use it. TensorFlow Federated4is an open-source framework for machine learning and other computations on decentralised data. As TensorFlow Privacy, it integrates easily with existing TensorFlow Models. In addition, it has built-in many known training datasets. 1https://github.com/OpenMined/PySyft 2https://www.openmined.org 3https://github.com/tensorflow/privacy 4https://www.tensorflow.org/federated 16 Shortcomings.TensorFlow Privacy only focuses on differentially private optimisers and it does not provide any DP mechanisms to implement your own differentially private optimisers. It does not officially support any other deep learning library and it is still not compatible with the latest TensorFlow 2.x. In addition, it is a ”library under continual development” according to its Github documentation1, it is not thus mature enough for production usage. While TensorFlow Federated provides both low level and high level interfaces for FL settings and it has some high level interfaces to create aggregation operators, it does not provide any built-in aggregation operators. Last, it is not yet compatible with the latest TensorFlow 2.x. Overview.These TensorFlow frameworks in conjunction allow us to develop FL models, but they are tied to the TensorFlow framework, which greatly denies any portability of the generated model. They are neither compatible with the latest version of TensorFlow nor they are ready for final products. In addition, they lack DP mechanisms to implement new privacy-preserving algorithms. 3.3. FATE FATE2is an open-source project initiated by Webank’s AI Department to provide a secure computing framework to support the federated AI ecosystem. Features.It provides many interesting FL algorithms and it exposes3a high level interface driven by custom scripts. Shortcomings.Its high level interface made of scripts relies too much on command line parameters and on a poorly documented domain specific language. It is unclear how to implement a low level FL model, which makes us think this framework is designed as a black box model. Their modular architecture seems quite complex. Also, it does not feature any DP algorithm, and there are no signs of future plans for implementing them. Overview.This framework is mainly focused on FL, making one of its biggest weaknesses that it does not implement any DP algorithm, in order to improve its data protection regulation compliance. Secure computation protocols ensure that data is not eavesdropped by an adversary, but it does not ensure that individuals’ privacy, roughly speaking, is preserved. In addition, it is expected to be used as a high level interface which relies on a barely documented custom language. 1https://github.com/tensorflow/privacy 2https://fate.fedai.org/overview/ 3https://fate.readthedocs.io/en/latest/examples/federatedml-1.x-examples/README. html 17 3.4. LEAF LEAF1is a benchmarking framework for learning in federated settings, with applications including FL, multi-task learning, meta-learning, and on-device learning. Features.This framework mainly focuses on benchmarking FL settings. It provides some basic FL mechanisms such as the Federated Averaging Aggregator and given its modular design it can be adapted to work on any existing framework. Last, it has some known built-in datasets such as FEMNIST, Shakespeare and Celeba. Shortcomings.It does not provide any benchmark for preserving privacy in a FL setting, even though privacy must be taken into consideration as it is a desired property of many FL settings. Moreover, it does not offer as many official documentation or tutorials as the other frameworks discussed in this section. Overview.LEAF offers a baseline implementation for some basic FL methods but its main purpose is benchmarking FL settings. However, DP benchmarks are not provided, even though nowadays privacy is a concern in most FL settings. 3.5. PaddleFL PaddleFL2is an open source FL framework based on PaddlePaddle3. PaddlePaddle is an industrial platform with advanced technologies and rich features that cover core deep learning frameworks, basic model libraries, end-to-end development kits, tool and component as well as service platforms. Features.PaddleFL provides a high level interface to develop FL models with DP. In the FL field it implements the Federated Averaging Aggregator and its secure multi-party computation equivalent. When it comes to DP, it provides an implementation of the differentially private stochastic gradient descent. Shortcomings.This framework has little documentation. It lacks any other DP algorithm so there is great difficulty in developing alternative privacy-preserving techniques. Last, since it is based on PaddlePaddle it is not compatible with other frameworks, and there is little documentation which makes it really hard to use and understand. Overview.PaddleFL provides a high level interface for some basic and well-known FL aggregators and implements a differentially private algorithm, being one of its main drawbacks that it is little documented and it does not implement any tool to easily extend its capabilities. 1https://leaf.cmu.edu/ 2https://paddlefl.readthedocs.io/en/latest 3https://github.com/paddlepaddle/paddle 18 3.6. Frameworks analysis The discussed software tools share some shortcomings for developing distributed AI services that preserves data privacy. Among them, we stress out the following: 1. They focus on FL or DP, but they do not provide a unified approach for both of them. 2. They lack DP mechanisms and related methods from the DP area. Likewise, they do not allow to develop and integrate new DP mechanisms in the frameworks. 3. Only the most basic federated aggregation operators are implemented. They are mainly focused on deep learning models, and they do not provide support for other machine learning algorithms that may be also used in the FL setting. We summarise and compare the characteristics of the frameworks reviewed in Table 1. We conclude that a unified FL and DP framework is required, and this is the ambitious aim of Sherpa.ai FL, which we present in the following section. FL & DP features PySyft TensorFlow LEAF PaddleFL Federated Learning: •Use federated models with different datasets •Support for other libraries •Sampling environment: IID or non-IID distribution •Federated aggregation mechanisms •Federated attack simulator Differential Privacy: •Mechanisms: Exponential, Laplacian, Gaussian •Sensitivity sampler •Subsampling methods to increase privacy •Adaptive Differential Privacy Desired properties: •Documentation & tutorials •High level API •Ability to extend the framework with new properties Table 1: FL and DP features comparison among existing frameworks. Complete Partial Do not work Unknown 4. Sherpa.ai Federated Learning Framework We develop Sherpa.ai FL,1,2which is an open-research unified FL and DP framework that aims to foster the research and development of AI services at the edges and to preserve data privacy. We describe the hierarchical and modular software architecture of Sherpa.ai FL, related to the key elements of FL and DP shown in Section 4.1. Likewise, 1https://developers.sherpa.ai/privacy-technology/ 2https://github.com/sherpaai/Sherpa.ai-Federated-Learning-Framework 19 we detail the functionalities and the implementation details of Sherpa.ai FL in Section 4.2 and Section 4.3. 4.1. Software architecture The software is structured in several modules that encapsulate the specific functionality of each key element of a FL setting. The architecture of these software modules allows the extension of the framework in relation to the progress of the research on FL. Figure 3 shows the backbone of the software architecture of Sherpa.ai FL, and we describe each module as what follows: •data_base: it is in charge of reading the data according to the chosen database. It is related to the data key element. •data_distribution: it performs the federated distribution of data among the clients involved in the FL process. It is also related to the data key element and completes its functionality. •private: it includes several interfaces such as the node interface which represents the clients key element and other ones that allow to access and modify the federated data distribution. •learning_approach: it represents the whole FL scheme including the federated server model and the communication and coordination among federated server and clients. It encapsulates the federated server and the communication key elements. •federated_aggregator: it defines the software structure to develop federated aggregation operators. It is linked to the federated aggregation operator key element. Federated Goverment Data base Data distribution Private Federated aggregator Model Differential privacy implements Figure 3: Links between the different modules of Sherpa.ai FL. 20 •model: it defines the learning model using predefined models and their functionalities. This learning model could be any machine learning model that can be aggregated by its representation in parameters. It is related to the model key element, as we associate a model object with the clients and the federated server. •differential_privacy: it preserves DP of the clients by specifying the data access. It is related with the DP key elements, and also with the data, the clients and the communication FL key elements. 4.2. Software functionalities In this section we highlight the main contributions of Sherpa.ai FL, which are summarised in a wide range of functionalities, namely: • To define and customise a FL simulation with a fixed number of clients using classical data sets. • To define the previous FL simulation using high-level functionalities. • To train machine learning models among different clients. Currently, Sherpa.ai FL offers support for a Keras models (neural networks), and for several models from Scikit-Learn (linear regression, k-means clustering, logistic regression). • To aggregate the information learned from each of the clients into a global model using classical federated aggregation operators such as: FedAvg, weighted FedAvg [24] and an aggregation operator for the adaptation of the k-means algorithm to the federated setting [25]. • To apply modifications on federated data such as normalisation or reshaping. • To evaluate the FL approach in comparison with the classical centralised one. • To preserve DP of clients’ data and model parameters in the FL context. The platform currently offers support for the fundamental DP mechanisms (Randomized Response, Laplace, Exponential, Gauss), and the composition of DP mechanisms (Basic and Advanced adaptive composition using privacy filters for the maximum privacy loss). Moreover, it is possible to increase privacy by subsampling. In Table 2, we summarise the main contributions of Sherpa.ai FL in comparison with the key points analysed for each framework in the previous Section. Thanks to the hierarchical implementation of each module, the aforementioned functionalities can be extended and customised just by adding software classes that inherit from the original software classes. For example, the already available machine learning models, DP mechanisms and federated aggregation operators can be modified, or new ones can be created, simply by overwriting the corresponding methods in the classes TrainableModel,DataAccessDefinition,FederatedAggregator, respectively. 21 FL & DP features PySyft TensorFlow LEAF PaddleFL Federated Learning: •Use federated models with different datasets •Support for other libraries •Sampling environment: IID or non-IID distribution •Federated aggregation mechanisms •Federated attack simulator Differential Privacy: •Mechanisms: Exponential, Laplacian, Gaussian •Sensitivity sampler •Subsampling methods to increase privacy •Adaptive Differential Privacy Desired properties: •Documentation & tutorials •High level API •Ability to extend the framework with new properties Table 2: FL & DP features comparison between existing frameworks and Sherpa.ai FL. Complete Partial Do not work Unknown 4.3. Implementation details Sherpa.ai FL has been developed by DaSCI1Institute and Sherpa.ai.2We developed the software using Python language for the whole architecture. Furthermore, Keras3, TensorFlow4and scikit-kearn5APIs are employed for the machine learning part which ensures efficiency and compatibility. It can also be run on computing devices such as CPUs or GPUs. In order to use GPUs, the adequate versions of TensorFlow and CUDA must be installed. For detailed installation instructions, please see the installation guide.6,7 The framework is licensed under the Apache License 2.0,8a permissive license whose main conditions require preservation of copyright and license notices. 5. Machine learning matches federated learning. Methodological guidelines for preserving data privacy FL is a paradigm of machine learning, but its particularities force to adapt the machine learning settings to FL. For this reason, Sherpa.ai FL functionalities outlined in Section 1https://dasci.es/ 2https://sherpa.ai/ 3https://keras.io/ 4https://www.tensorflow.org/ 5https://scikit-learn.org/stable/ 6https://github.com/sherpaai/Sherpa.ai-Federated-Learning-Framework/blob/master/ install.md 7The implementation details described in this paper corresponds to the release 0.1.0 of Sherpa.ai FL. 8https://www.apache.org/licenses/LICENSE-2.0 22 4stem from the need to develop machine learning algorithms specialised for Federated Artificial Intelligence that protect clients’ data privacy. However, the adaptation is not only focused on the algorithms, but also on the workflow of machine learning. In this section, we first discuss the key aspects of distributed computing for machine learning in the federated setting. Following, Section 5.2 defines a rather specific paradigm for adapting machine learning models to the federated setting, followed by some remarkable exceptions to the defined adaptation in Section 5.3. Finally, we also define the adaptation of the machine learning workflow to FL in Section 5.4 as methodological guidelines for preserving data privacy with FL using Sherpa.ai FL. 5.1. Key aspects of distributed computing for Federated Machine Learning The recent introduction of FL [26,14,7] responds to the need for novel distributed machine learning algorithms in a setting that clashes with several assumptions of conventional parallel machine learning in a data centre. The differences are substantially originated by the unreliable and poor network connection of the clients, since clients are typically mobile phones. Thus reducing the number of rounds of learning is essential as communication constrains are more severe. Additionally, the data is unevenly scattered across Kclients and must be considered as non-IID, that is, the data accessible locally is not in any way representative of the overall trend.1The data is in general sparse, where the features of interest take place on a reduced number of clients or data points.2Ultimately, the number of total clients greatly exceeds the number of training points available locally on each client (Kn/K). In the federated machine learning setting, the training is decoupled from the access to the raw data. In fact, the raw data never leaves users’ mobile devices and a high-accuracy model is produced in a central server by aggregating locally computed updates. At each FL round, an update vector θ∈Rdis sent from each client to the central server to improve the global model, with dthe parameters’ dimension of the computed model. It is worth noting that the magnitude of the update θis thus independent from the amount of raw data available on the local client (e.g., θmight be a gradient vector). One of the advantages of this approach is the considerable bandwidth and time saved in data communication. Another motivation for the FL setting (but that also constitutes one of its intrinsic advantages) is the concern for privacy and security. By not transferring any raw data to the central server, the attack surface reduces to only the single client, instead of both client and server. On the other hand, the update θsent by the client might still reveal some of its private information, however the latter will be almost always dramatically reduced with respect to the raw training data. Besides, after improving the current model by the update θ, this can (and should) be deleted. 1Sherpa.ai FL allows for both IID and non-IID client data. 2Sherpa.ai FL allows for weighted aggregation, emphasising the contribution of most significant clients to the global model. 23 Additional privacy can be provided by randomised algorithms providing DP [12], as detailed in Sections 2.3 and 2.4. In particular, the centralised algorithm could be equipped with a DP layer allowing the release of the global model without compromising the privacy of the individual clients who contributed to its training (see e.g., Abadi et al. [27]). On the other hand, in the case of a malicious or compromised server, or in the case of potential eavesdropping, DP can be applied on the local clients1for protecting their privacy [28,29,30]. 5.2. The Machine Learning paradigm in a federated setting In the following, we describe the federated machine learning paradigm by recognising some relevant attributes that ease the natural adaptation of a ML model in the federated setting. Primarily, we observe that a great number of machine learning methods resemble the minimisation of an objective function with finite-sum as in Equation 3. The aforementioned problem structure encompasses both linear and logistic regressions, support vector machines, and also more elaborated techniques such as conditional random fields and neural networks [26]. Indeed, in neural networks predictions are made through a nonconvex function, yet the resulting objective function can still be expressed as Fi(θ)and the gradients can be efficiently obtained by backpropagation, thus resembling Equation 3. A variety of algorithms have been proposed to solve the minimisation problem in Equation 3in the federated setting, where, as mentioned earlier, the primary constrain is the communication efficiency for reducing the number of FL rounds in the aggregation of local models. In this context, another characteristic trait of the Federated ML paradigm is constituted by the intrinsic compatibility with baseline aggregation operators (e.g. Federated Averaging), and where no ad-hoc adaptation is required. Ultimately, several of these FL algorithms have been supplied with DP [27,29,30]. We thus identify a rather important aspect of the Federated machine learning paradigm as being prone to straightforward application of the common building components of DP. In addition, the latter feature eases the task of estimating the privacy loss in the FL rounds by the application of composition theorems for DP.2 To summarise, a machine learning method is prone to adaptation in the federated setting if it adheres to the principles of the federated machine learning paradigm described above, namely: (1) the problem structure resembling the minimisation of an objective function as in Equation 3, (2) the attribute of easy aggregation of local models’ parameters, and (3) the direct applicability of DP techniques for additional privacy. Among such machine learning models we cite neural networks [31], linear [32,33] and logistic [34] 1Sherpa.ai FL allows to apply sophisticated and customised DP mechanisms on the model’s parameters, as well as on client’s raw data (see Section 2.4). 2Sherpa.ai FL offers support for both common building components for DP, as well as for its basic and advanced composition theorems using privacy filters (see Section 2.4). 24 regressions.1 5.3. Models deviating from the federated machine learning paradigm It is worth mentioning specific machine learning models whose structure only partially fits in the federated machine learning paradigm described above. Although the problem structure can still be represented by a minimisation of an objective function as in Equation 3, their adaptation to a federated setting requires additional and ad-hoc procedures. One example is found in the k-means clustering algorithm for unsupervised learning [35,36], where the non-IID nature of the data distribution is seen as a major obstacle in a federated setting. Namely, the direct application of average aggregation is unfeasible due to potentially different number and ordering of local clusters, and more advanced algorithms need to be employed. A workable solution is to fix the number of local clusters, and apply an additional k-means clustering in the average aggregation.2Alternatively, one might try grouping the clients’ population sharing jointly trainable data distributions, as proposed by Sattler et al. [37] in the context of deep neural networks. An additional complication is constituted by the preservation of clients’ privacy. For instance, the baseline DP building components necessitate some adjustments in order to be applied. Although not in a FL context, in Zhang et al. [38] the authors adjust the Laplace noise added to each centroid based on the contour coefficients. Another notable example is represented by the federated version of matrix factorizationbased ranking algorithms for recommendation systems [39,40]. The peculiar architecture of this algorithm involves the communication of only a portion of the update vector for improving the model, thus the round of learning implies additional communications between the central server and the clients. Moreover, the multiple communication iterations necessitates further caution with privacy loss in the DP context. A viable approach is to implement a two-stage randomised response mechanism on the local data, and to allow the clients to specify their privacy level [30]. 5.4. Methodological guidelines for preserving data privacy with federated learning in Sherpa.ai FL The experimental settings of FL and machine learning are very similar because FL is a machine learning paradigm. Nonetheless, the particularities of FL force to revise the machine learning workflow, and adapt it to the FL definition. In this section, we define the workflow of FL based on the machine learning workflow, and we present it as methodological guidelines. Sherpa.ai FL comply with these methodological guidelines assuring the following of good practises in the development of AI services at the edges that preserve data privacy. 1See also notebooks on deep learning, linear and logistic regressions available in Sherpa.ai FL at https: //github.com/sherpaai/Sherpa.ai-Federated-Learning-Framework/tree/master/notebooks 2See notebook on k-means clustering in Sherpa.ai FL at https://github.com/sherpaai/Sherpa. ai-Federated-Learning-Framework/tree/master/notebooks 25 As a result, we have created federated data from the EMNIST dataset with 5 nodes and using every available data. Hence, the data collection process have finished. This data is a set of data nodes containing private data. [5]: print(type(federated_data)) print(federated_data.num_nodes()) federated_data[0].private_data <class 'shfl.private.federated_operation.FederatedData'> 5 Node private data, you can see the data for debug purposes but the data ,→remains in the node <class 'dict'> {'112883278416': <shfl.private.data.LabeledData object at 0x1a486393d0>} As we can see, private data in a node is not accesible directly but the framework provides mechanisms to use this data in a machine learning model. Once data is prepared, the next step is the definition of the neural network architecture (model selection) used along the learning process. The framework provides a class to adapt a Keras (or Tensorflow) model to the framework, so you only have to create a function that will act as model builder. [6]: import tensorflow as tf def model_builder(): model =tf.keras.models.Sequential() model.add(tf.keras.layers.Conv2D(32, kernel_size=(3,3), ,→padding='same', activation='relu', strides=1, input_shape=(28,28, ,→1))) model.add(tf.keras.layers.MaxPooling2D(pool_size=2, strides=2, ,→padding='valid')) model.add(tf.keras.layers.Dropout(0.4)) model.add(tf.keras.layers.Conv2D(32, kernel_size=(3,3), ,→padding='same', activation='relu', strides=1)) model.add(tf.keras.layers.MaxPooling2D(pool_size=2, strides=2, ,→padding='valid')) model.add(tf.keras.layers.Dropout(0.3)) model.add(tf.keras.layers.Flatten()) model.add(tf.keras.layers.Dense(128, activation='relu')) model.add(tf.keras.layers.Dropout(0.1)) model.add(tf.keras.layers.Dense(64, activation='relu')) model.add(tf.keras.layers.Dense(10, activation='softmax')) 32 model.compile(optimizer="rmsprop", ,→loss="categorical_crossentropy", metrics=["accuracy"]) return shfl.model.DeepLearningModel(model) The following step is the definition of the federated aggregation operator in order to complete the model selection in FL. The framework provides some aggregation operators that we can use immediately and the possibility to define your own operator. In this case, we use the provided FedAvg operator. [7]: aggregator =shfl.federated_aggregator.FedAvgAggregator() federated_government =shfl.federated_government. ,→FederatedGovernment(model_builder, federated_data, aggregator) The framework also provides the possibility of making data transformation for the data preprocessing step, defining federated operations using FederatedTransformation interface. We first reshape data and then normalise it using test data mean and standard deviation (std) as normalisation parameters. [8]: import numpy as np class Reshape(shfl.private.FederatedTransformation): def apply(self, labeled_data): labeled_data.data =np.reshape(labeled_data.data, ,→(labeled_data.data.shape[0], labeled_data.data.shape[1], ,→labeled_data.data.shape[2],1)) shfl.private.federated_operation. ,→apply_federated_transformation(federated_data, Reshape()) [9]: import numpy as np class Normalize(shfl.private.FederatedTransformation): def __init__(self, mean, std): self.__mean =mean self.__std =std def apply(self, labeled_data): labeled_data.data =(labeled_data.data -self.__mean)/self. ,→__std 33 mean =np.mean(train_data.data) std =np.std(train_data.data) shfl.private.federated_operation. ,→apply_federated_transformation(federated_data, Normalize(mean, std)) We are now ready to train the FL algorithm. We run 2 rounds of learning showing test accuracy and loss of each client and test accuracy and loss of the global aggregated model. [10]: test_data =np.reshape(test_data, (test_data.shape[0], test_data. ,→shape[1], test_data.shape[2],1)) federated_government.run_rounds(2, test_data, test_labels) Accuracy round 0 Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a485e2450>: [15.087034225463867, 0.9314000010490417] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x106a0ffd0>: [21.040000915527344, 0.9094250202178955] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a48639e90>: [11.712089538574219, 0.9425749778747559] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a486396d0>: [10.11756420135498, 0.9498249888420105] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a48639c50>: [24.04242706298828, 0.8968499898910522] Global model test performance : [7.954472064971924, 0.9403749704360962] Accuracy round 1 Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a485e2450>: [21.94520378112793, 0.9227499961853027] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x106a0ffd0>: [16.780630111694336, 0.9445000290870667] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode 34 object at 0x1a48639e90>: [13.413337707519531, 0.9463250041007996] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a486396d0>: [9.085938453674316, 0.9628000259399414] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x1a48639c50>: [20.926694869995117, 0.918524980545044] Global model test performance : [10.171743392944336, 0.958299994468689] If we focus our attention on test accuracy in each client, we realise that there are widely varying results.1This is because of the scattered nature of the data distribution, which causes disparity in the quality of training data between clients. 6.1.3. Comparison with a centralised convolutional neural network approach We analyse the behaviour of the FL approach in comparison with the equivalent centralised approach, which means training the neural network represented in Figure 5on the same data using centralised learning. For this experiment, we use 25 clients and 10 rounds of learning with 5 epochs in both IID and non-IID scenario, where the nodes’ data contain only a portion of all labels. For a fair comparison, in the classical approach we train for epochsFL ×roundsFL epochs. IID non-IID Centralised approach 0.9904 0.9901 Federated approach 0.9921 0.9855 Table 4: Accuracy of the FL and the classical approach, in both IID and non-IID scenarios. In the FL case, the data is distributed over 25 clients, and 10 FL rounds of learning with 5 epochs per client are employed. The high performance of the federated approach stands out in Table 4, where the accuracy for the considered scenarios is reported. In the IID scenario, it beats the centralised approach results, which shows the robustness of the approach caused by the combination of the information learned by each client. In the non-IID scenario, the federated approach attains lower results than the centralised one due to the additional challenge of non-homogeneous distribution of data across clients.2However, the results are very competitive highlighting the strength of the federated approach. 1Running more learning rounds results in better performance as in the next section. The purpose of this example is to show how it works. 2The performance of the centralised approach using non-IID data is not perfectly identical to the IID case due to the random sampling employed when generating the non-IID nodes’ data. 35 6.2. Linear regression with DP This section presents a linear regression FL simulation with DP following the methodological guidelines with Sherpa.ai FL.1The Laplace mechanism is used when the model’s sensitivity is estimated by a sampling procedure [19]. Moreover, we demonstrate the application of the advanced composition theorem for DP for not exceeding the maximum privacy loss allowed (see Section 2.4). 6.2.1. Case of study We will use the California Housing dataset,2which consists of approximately 20000 samples for median house prices in California. Although the dataset possesses eight features, in this example we will only make use of the first two, in order to reduce the variance in the prediction. The (single) target is the cost of the house. As it can be observed in the code below, we retain 2000 samples for later use with the sensitivity sampling for DP, and the rest of the data is split in train and test sets as detailed in Table 5. Train set Test set Total 14912 3728 18 640 Table 5: Distribution of the California Housing dataset. For the FL simulation we use 5 clients among which the train dataset is IID. FedAvg is chosen as the federated aggregation operator. The code of the example is detailed in the following section. 6.2.2. Description of the code Sherpa.FL allows to easily convert a generic dataset to interact with the platform: import shfl from shfl.data_base.data_base import LabeledDatabase import sklearn.datasets import numpy as np from shfl.private.reproducibility import Reproducibility # Comment to turn off reproducibility: Reproducibility(1234) all_data =sklearn.datasets.fetch_california_housing() n_features = 2 1Sherpa.ai FL offers support for the linear regression model from scikit-learn https:// scikit-learn.org/stable/index.html 2https://scikit-learn.org/stable/modules/generated/sklearn.datasets.fetch_ california_housing.html 36 data =all_data["data"][:,0:n_features] labels =all_data["target"] # Retain part for DP sensitivity sampling: size = 2000 sampling_data =data[-size:, ] sampling_labels =labels[-size:, ] # Create database: database =LabeledDatabase(data[0:-size, ], labels[0:-size]) train_data, train_labels, test_data, test_labels =database.load_data() We will simulate a FL scenario by distributing the train data over a collection of clients, assuming an IID setting: iid_distribution =shfl.data_distribution.IidDataDistribution(database) federated_data, test_data, test_labels =iid_distribution. ,→get_federated_data(num_nodes=5) At this stage, we need to define the linear regression model, and we choose the aggregation operator to be the average of the clients’ models: from shfl.model.linear_regression_model import LinearRegressionModel def model_builder(): model =LinearRegressionModel(n_features=n_features) return model aggregator =shfl.federated_aggregator.FedAvgAggregator() 6.2.3. Running the model in a Federated configuration We are now ready to run the FL model. Note that in this case, we set the number of rounds n=1 since no iterations are needed in the case of linear regression. The performance metrics used are the Root Mean Squared Error (RMSE) and the R2score. It can be observed that the performance of the Global model (i.e. the aggregated model) is in general superior with respect to the performance of each node, thus the federated learning approach proves to be beneficial: federated_government =shfl.federated_government. ,→FederatedGovernment(model_builder, federated_data, aggregator) federated_government.run_rounds(n=1, test_data=test_data, ,→test_label=test_labels) 37 Accuracy round 0 Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a08d0>: [0.8161535463006577, 0.5010049851923566] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0ac8>: [0.81637303674763, 0.5007365568636023] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a09b0>: [0.8155342443231007, 0.5017619784187599] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0be0>: [0.8158502097728687, 0.5013758352304256] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0cf8>: [0.8151607067608612, 0.5022182878756591] Global model test performance : [0.8154147770321544, 0.5019079411109164] 6.2.4. Differential Privacy: sampling the model’s sensitivity In the case of applying the Laplace privacy mechanism (see Section 2.4), the noise added has to be of the order of the sensitivity of the model’s output, i.e. the model parameters of our linear regression. In the general case, the model’s sensitivity might be difficult to compute analytically. An alternative approach is to attain random differential privacy through a sampling over the data [19]. That is, instead of computing analytically the global sensitivity ∆f, we compute an empirical estimation of it by sampling over the dataset. This approach is convenient since it allows for the sensitivity estimation of an arbitrary model or a black-box computer function. The Sherpa.FL framework provides this functionality in the class SensitivitySampler. In order to carry out this approach, we need to specify a distribution of the data to sample from. This in general requires previous knowledge and/or model assumptions. In order not make any specific assumption on the distribution of the dataset, we can choose a uniform distribution. To the end, we define our class of ProbabilityDistribution that uniformly samples over a data-frame. We use the previously retained part of the dataset for sampling: class UniformDistribution(shfl.differential_privacy. ,→ProbabilityDistribution): """ Implement Uniform sampling over the data """ 38 def __init__(self, sample_data): self._sample_data =sample_data def sample(self, sample_size): row_indices =np.random.randint(low=0, high=self._sample_data. ,→shape[0], size=sample_size, dtype='l') return self._sample_data[row_indices, :] sample_data =np.hstack((sampling_data, sampling_labels.reshape(-1,1))) The class SensitivitySampler implements the sampling given a query, i.e. the learning model itself in this case. We only need to add the method get to our model since it is required by the class SensitivitySampler. We choose the sensitivity norm to be the `1 norm and we apply the sampling. The value of the sensitivity depends on the number of samples n: the more samples we perform, the more accurate the sensitivity. Indeed, increasing the number of samples n, the sensitivity gets more accurate and typically decreases. from shfl.differential_privacy import SensitivitySampler from shfl.differential_privacy import L1SensitivityNorm class LinearRegressionSample(LinearRegressionModel): def get(self, data_array): data =data_array[:, 0:-1] labels =data_array[:, -1] train_model =self.train(data, labels) return self.get_model_params() distribution =UniformDistribution(sample_data) sampler =SensitivitySampler() n_samples = 4000 max_sensitivity, mean_sensitivity =sampler.sample_sensitivity( LinearRegressionSample(n_features=n_features, n_targets=1), L1SensitivityNorm(), distribution, n=n_samples, gamma=0.05) print("Max sensitivity from sampling: " +str(max_sensitivity)) print("Mean sensitivity from sampling: " +str(mean_sensitivity)) Max sensitivity from sampling: 0.008294354064053988 Mean sensitivity from sampling: 0.0006633612087443363 39 Unfortunately, sampling over a dataset involves the training of the model on two datasets differing in one entry [19]. Thus in general this procedure might be computationally expensive (e.g. in the case of training a deep neuronal network). 6.2.5. Running the model in a Federated configuration with Differential Privacy At this stage we are ready to add a layer of DP to our federated learning model. Specifically, we will apply the Laplace mechanism from Section 2.4, employing the sensitivity obtained from the previous sampling, namely ∆f≈0.008. The Laplace mechanism provided by the Sherpa.FL framework is then assigned as the private access type to the model’s parameters of each client in a new FederatedGovernment object. This results into an e-differentially private FL model. For example, picking the value e=0.5, we can run the FL experiment with DP: from shfl.differential_privacy import LaplaceMechanism params_access_definition = ,→LaplaceMechanism(sensitivity=max_sensitivity, epsilon=0.5) federated_governmentDP =shfl.federated_government.FederatedGovernment( model_builder, federated_data, aggregator, ,→model_params_access=params_access_definition) federated_governmentDP.run_rounds(n=1, test_data=test_data, ,→test_label=test_labels) Accuracy round 0 Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a08d0>: [0.8161535463006577, 0.5010049851923566] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0ac8>: [0.81637303674763, 0.5007365568636023] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a09b0>: [0.8155342443231007, 0.5017619784187599] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0be0>: [0.8158502097728687, 0.5013758352304256] Test performance client <shfl.private.federated_operation. ,→FederatedDataNode object at 0x7f63606a0cf8>: [0.8151607067608612, 0.5022182878756591] Global model test performance : [0.8309024800913748, 0.48280707735516126] In the above example we observed that the performance of the model has slightly deteri40 orated due to the addition of DP. In general, the privacy increases at expenses of accuracy (i.e. for smaller values of e). 6.2.6. Comparison with centralised and non private approaches It is of practical interest to assess the performance loss due to DP in the FL context. Table 6reports the performance metrics for the centralised model, for the FL non-private model, and for the FL differentially-private model. In all the cases, the models have learned on the train set, and the performance results have been computed over the test set. The data is IID over 5 clients in the FL cases. For the federated DP cases, different values of e={0.2,0.5,0.8}are used and the total privacy expense is limited at eT=4. Thus, for each case, we took the average over the total runs before the budget is expended (see discussion about advanced composition theorems for privacy filters in Section 2.4). The sensitivity is fixed, employing the value obtained from sampling above.1 Approach RMSE R2 Classical 0.81540 0.50192 Federated non-private 0.81541 0.50190 Federated DP (e=0.2, average of 20 runs) 1.05541 0.04224 Federated DP (e=0.5, average of 8 runs) 0.84501 0.46457 Federated DP (e=0.8, average of 5 runs) 0.82171 0.49414 Table 6: Federated linear regression: comparison between the classical centralised model, the non-private FL model, and the FL model with a DP layer using the Laplace mechanism. For the DP cases, the results are the average over the total runs allowed for the maximum privacy budget eT=4. Different values of e={0.2,0.5, 0.8}are considered, and the sensitivity is fixed. The data is IID distributed over 5 clients. It can be observed that the centralised model and the non-private FL model exhibit comparable performance, thus the accuracy is not degraded by applying a FL approach. The application of the Laplace mechanism guarantees e-DP, and the accuracy of the FL model can be leveraged by setting the value of e: for higher values, lower privacy is guaranteed, but the accuracy increases. 7. Concluding remarks The characteristics of FL and DP make them good candidates to support AI services at the edges and to preserve data privacy. Hence, several software tools for FL and DP have been released. After a comparative analysis, we conclude that these software tools do not provide a unified support for FL and DP, and they do not follow any particular methodological guidelines that direct the developing of AI services to preserve data privacy. 1Note that, when applying the composition theorems for privacy filters in the present example, we are assuming that the estimated sensitivity is a good enough approximation of the analytic sensitivity [22]. 41