Published January 12, 2023 | Version v1

Deployment of Federated Learning Infrastructure

Description

Federated Learning (FL) is a technique to train machine learning (ML) models on decentralized data, meaning the data is distributed across multiple devices. FL aids model managers in improving their ML models by training on large distributed data. It also enables privacy since local data is kept on clients' devices, and only model updates are shared. In an FL setting, two main types of entities exist the server and the client. Usually, there is one server and multiple clients, each having its local dataset. The server holds the global model, while the clients hold the data. In an FL training round, the server sends a copy of its global model to each client. The clients then train the received models on their local data and send the model weights back to the server, which are aggregated to update the global model. The process is repeated multiple times until the model achieves convergence. Sometimes, it is necessary to change training parameters and the network architecture to help the model converge. Such changes entail restarting the training process, which induces an overhead. In this project, we are developing a tool that automates the deployment of FL on multiple devices. The tool is built on top of existing FL frameworks and configuration tools. Several comparisons have been made to choose the best FL framework and configuration tool for our project. Details about the implementation and training flow description are provided in the following sections.

Files

Mohammed-Hemdan.pdf

Files (666.1 kB)

Name Size Download all
md5:170131b02decb6b8316a221c85c0dcb3
666.1 kB Preview Download