Deployment of Federated Learning Infrastructure
Authors/Creators
Description
Federated Learning (FL) is a technique to train machine learning (ML) models on decentralized data, meaning the data is distributed across multiple devices. FL aids model managers in improving their ML models by training on large distributed data. It also enables privacy since local data is kept on clients' devices, and only model updates are shared. In an FL setting, two main types of entities exist the server and the client. Usually, there is one server and multiple clients, each having its local dataset. The server holds the global model, while the clients hold the data. In an FL training round, the server sends a copy of its global model to each client. The clients then train the received models on their local data and send the model weights back to the server, which are aggregated to update the global model. The process is repeated multiple times until the model achieves convergence. Sometimes, it is necessary to change training parameters and the network architecture to help the model converge. Such changes entail restarting the training process, which induces an overhead. In this project, we are developing a tool that automates the deployment of FL on multiple devices. The tool is built on top of existing FL frameworks and configuration tools. Several comparisons have been made to choose the best FL framework and configuration tool for our project. Details about the implementation and training flow description are provided in the following sections.
Files
Mohammed-Hemdan.pdf
Files
(666.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:170131b02decb6b8316a221c85c0dcb3
|
666.1 kB | Preview Download |