RE-SAMPLE Synthetic ML Training Dataset
Authors/Creators
Description
Synthetic version of the RE-SAMPLE ML training dataset. Created based on data from Fondazione Policlinico Universitario Agostino Gemelli (Gemelli), Medisch Spectrum Twente (MST) and Tartu University Hospital (TUK). The distributions of the variables in the synthetic dataset closely approximate those observed in the original datasets collected during the project. However, some discrepancies exist. Notably, the dataset provided by the partner Gemelli exhibited a significantly higher proportion of missing physical activity data (~87%), whereas the datasets from MST and TUK contained approximately 39% and 28% missing values for this variable, respectively. As a result, the distribution of physical activity in the synthetic dataset deviates more substantially from the real dataset for Gemelli compared to the other partners.