There is a newer version of the record available.

Published September 3, 2022 | Version v3

On the benefits of self-taught learning for brain decoding - Data

  • 1. Univ Rennes, Inria, CNRS, Inserm, Rennes, France
  • 2. Univ Rennes, IUF, Inria, CNRS, IRISA, Rennes, France

Description

DERIVED DATA FROM PAPER "On the benefits of self-taught learning for brain decoding"

Here are stored the data necessary to reproduce the full analysis of the paper "On the benefits of self-taught learning for brain decoding". 

We study the benefits of using a large public neuroimaging database composed of fMRI statistic maps, in a self-taught learning framework, for improving brain decoding on new tasks. First, we leverage the NeuroVault database to train, on a selection of relevant statistic maps, a convolutional autoencoder to reconstruct these maps. Then, we use this trained encoder to initialize a supervised convolutional neural network to classify tasks or cognitive processes of unseen statistic maps from large collections of the NeuroVault database. We show that such a self-taught learning process always improves the performance of the classifiers but the magnitude of the benefits strongly depends on the number of data available both for pre-training and finetuning the models and on the complexity of the targeted downstream task.

Contents overview 

1. original

The original directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset

Each subdirectory contains:
- text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ; 
- csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ;
- an `original` directory in which original statistic maps downloaded from NeuroVault will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`.

2. preprocessed

The preprocessed directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset

Each subdirectory contains:
- text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ; 
- csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ;
- several subdirectores (`resampled`, `resampled masked`...) in which preprocessed statistic maps will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`.

3. derived

The derived directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset

Each subdirectory contains subdirectories in which the parameters of models trained on the different datasets are stored. These subdirectories are named in the following way: 

{name_of_the_dataset}_maps_classification_{classification_task}_model_cnn_{model_architecture}_valid_{type_of_experiment}_retrain_{type_of_initialization}_{preprocessing_type}_epochs_{number_of_epochs}_batch_size_{batch_size}_lr_{learning_rate}

For instance, parameters for the following experiment:
- Dataset: HCP Dataset subset 50 subjects
- Classification task: contrast classification
- Model: 4 layers CNN
- Type of experiment: Performance evaluation
- Initialization: Default 
- Preprocessing type: Resampled masked normalized
- Epochs: 500
- Batch: 32
- Learning rate: 1e-04 
will be contained in the directory: 

hcp_dataset_50_maps_classification_contrast_model_cnn_4layers_valid_perf_retrain_no_resampled_masked_normalized_epochs_500_batch_size_32_lr_1e-04

 

Files

data.zip

Files (37.0 GB)

Name Size
md5:c4d9f8b086e083e468038b768799bf43
37.0 GB Preview Download