Published February 21, 2023 | Version v1

Heterogeneous Datasets for Federated Survival Analysis Simulation

  • 1. Politecnico di Milano
  • 2. Technische Universität Dresden

Description

Heterogeneous Datasets for Federated Survival Analysis Simulation

This repo contains three algorithms for constructing realistic federated datasets for survival analysis. Each algorithm starts from an existing non-federated dataset and assigns each sample to a specific client in the federation. The algorithms are:

  • uniform_split: assigns each sample to a random client with uniform probability;
  • quantity_skewed_split: assigns each sample to a random client according to the Dirichlet distribution [3, 4];
  • label_skewed_split: assigns each sample to a time bin, then assigns a set of samples from each bin to the clients according to the Dirichlet distribution [3, 4].

For more information, please take a look at our paper at https://arxiv.org/abs/2301.12166 [1].

Content

  • federated_survival_datasets.zip: the content of the repository at https://github.com/archettialberto/federated_survival_datasets
  • Heterogheneous_Datasets_for_Federated_Survival_Analysis_Simulation.pdf: the conference paper describing the work.

Installation

Federated Survival Datasets is built on top of numpy and  scikit-learn. To install those libraries you can run pip install -r requirements.txt. To import survival datasets into your project, we strongly recommend SurvSet (https://github.com/ErikinBC/SurvSet) [2], a comprehensive collection of more than 70 survival datasets.

Usage

import numpy as np
import pandas as pd

from federated_survival_datasets import label_skewed_split

# import a survival dataset and extract the input array X and the output array y
df = pd.read_csv("metabric.csv")
X = df[[f"x{i}" for i in range(9)]].to_numpy()
y = np.array([(e, t) for e, t in zip(df["event"], df["time"])], dtype=[("event", bool), ("time", float)])

# run the splitting algorithm
client_data = label_skewed_split(num_clients=8, X=X, y=y)

# check the number of samples assigned to each client
for i, (X_c, y_c) in enumerate(client_data):
    print(f"Client {i} - X: {X_c.shape}, y: {y_c.shape}")

We provide an example notebook in the zipped folder to illustrate the proposed algorithms. It requires scikit-survival, seaborn, and pandas.

References

[1] Archetti, A., Lomurno, E., Lattari, F., Martin, A., & Matteucci, M. (2023). Heterogeneous Datasets for Federated Survival Analysis Simulation. arXiv preprint arXiv:2301.12166.

[2] Drysdale, E. (2022). SurvSet: An open-source time-to-event dataset repository. arXiv preprint arXiv:2203.03094.

[3] Hsu, T. M. H., Qi, H., & Brown, M. (2019). Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335.

[4] Li, Q., Diao, Y., Chen, Q., & He, B. (2022, May). Federated learning on non-iid data silos: An experimental study. In 2022 IEEE 38th International Conference on Data Engineering (ICDE) (pp. 965-978). IEEE.

Files

federated_survival_datasets.zip

Files (833.0 kB)

Additional details

Related works

Funding

European Commission
AI-SPRINT - Artificial Intelligence in Secure PRIvacy-preserving computing coNTinuum 101016577