Model Disparity in Distributed Learning
Abstract
We propose to experimentally evaluate how disparity actually impacts prediction accuracy for selected distributed or collaborative learning protocols. The results of this assessment should provide a guideline in how to define disparity limits in distributed of collaborative learning training processes.
Full text
Model Disparity in Distributed Learning Ana Nunes Alonso INESC TEC & UMinho Portugal [email protected] Jos´ e Orlando Pereira INESC TEC & UMinho Portugal [email protected] Rui Oliveira INESC TEC & UMinho Portugal rui.oli[email protected] Index Terms—federated leaning, distributed learning, collaborative learning, approximate distributed agreement Federated Learning algorithms enable distributed training of ML models [1]. Typically, the contributions of the training nodes are collected by a central note which integrates these to build/update a global model. The global model can then be made available to interested parties. However [1]: 1) This architecture makes the central node a single point of failure as results (the model) can be compromised if the central node fails. 2) Additionally, it requires interested parties to trust the central node implicitly, in that the global model that it provides corresponds to the intended (best) result. Replication is often proposed to mitigate these issues, by using a group of nodes aggregating contributions to build redundant instances of the global model instead of relying on a single node for that task. The replicated setup provides more flexibility on how nodes can assign trust in the system. An alternative to introducing a set of nodes with the specific aggregation/model building tasks is to leverage the existing training nodes to each build a replica of the global model, in what is sometimes referred to as distributed or collaborative learning [2]. However, in a setting with faults, different nodes can receive different sets of messages as messages can fail to be delivered (omission faults), nodes can fail (crash faults) or nodes can have unexpected or malicious behaviour (Byzantine faults). Left unchecked, faults could cause the global model replicas to diverge and this issue has been addressed in the research community in different ways, namely by using robust aggregation methods over training contributions to model parameters [3], which, at best, provide a bound on the disparity between model replicas. Still, the effect of the disparity in the models in prediction accuracy is rarely addressed. We propose to experimentally evaluate how disparity actually impacts prediction accuracy for selected distributed or collaborative learning protocols. The results of this assessment should provide a guideline in how to define disparity limits in distributed of collaborative learning training processes. We will then explore solutions for limiting disparity in the models. A first approach could be to use a distributed agreement protocol [4] to exactly synchronize model parameters across model replicas. However, it has been shown that exact distributed agreement in an asynchronous system cannot be guaranteed in the presence of even a single omission fault [5]. An alternative is to explore approximate distributed agreement protocol [6] implementations [7] which guarantee limited dispersion among agreed upon values. REFERENCES [1] J. Verbraeken, M. Wolting, J. Katzy, J. Kloppenburg, T. Verbelen, and J. S. Rellermeyer, “A survey on distributed machine learning,” ACM Computing Surveys (CSUR), vol. 53, no. 2, pp. 1–33, 2020. [2] R. Guerraoui, N. Gupta, and R. Pinot, “Byzantine machine learning: A primer,” ACM Comput. Surv., aug 2023, just Accepted. [Online]. Available: https://doi.org/10.1145/3616537 [3] E. El-Mhamdi, R. Guerraoui, A. Guirguis, and S. Rouault, “Garfield: System support for byzantine machine learning,” CoRR, vol. abs/2010.05888, 2020. [Online]. Available: https://arxiv.org/abs/2010.05888 [4] M. Pease, R. Shostak, and L. Lamport, “Reaching agreement in the presence of faults,” Journal of the Association for Computing Machinery, vol. 27, no. 2, pp. 228–234, April 1980. [5] M. J. Fischer, N. A. Lynch, and M. S. Paterson, “Impossibility of distributed consensus with one faulty process,” Journal of the Association for Computing Machinery, vol. 32, no. 2, pp. 374–382, April 1985. [6] D. Dolev, N. Lynch, S. Pinter, E. Stark, and W. Wheil, “Reaching approximate agreement in the presence of faults,” Journal of the Association for Computing Machinery, vol. 33, no. 3, pp. 499–516, July 1986. [7] E. L. da Conceic¸˜ ao, A. Nunes Alonso, R. C. Oliveira, and J. O. Pereira, “TADA: A Toolkit for Approximate Distributed Agreement,” in Distributed Applications and Interoperable Systems, M. Pati˜ no-Mart´ ınez and J. Paulo, Eds. Cham: Springer Nature Switzerland, 2023, pp. 3–19.