On the benefits of self-taught learning for brain decoding - Data
Authors/Creators
- 1. Univ Rennes, Inria, CNRS, Inserm, Rennes, France
- 2. Univ Rennes, IUF, Inria, CNRS, IRISA, Rennes, France
Description
DERIVED DATA FROM PAPER "On the benefits of self-taught learning for brain decoding"
Here are stored the data necessary to reproduce the full analysis of the paper "On the benefits of self-taught learning for brain decoding".
We study the benefits of using a large public neuroimaging database composed of fMRI statistic maps, in a self-taught learning framework, for improving brain decoding on new tasks. First, we leverage the NeuroVault database to train, on a selection of relevant statistic maps, a convolutional autoencoder to reconstruct these maps. Then, we use this trained encoder to initialize a supervised convolutional neural network to classify tasks or cognitive processes of unseen statistic maps from large collections of the NeuroVault database. We show that such a self-taught learning process always improves the performance of the classifiers but the magnitude of the benefits strongly depends on the number of data available both for pre-training and finetuning the models and on the complexity of the targeted downstream task.
Contents overview
1. original
The original directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset
Each subdirectory contains:
- text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ;
- csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ;
- an `original` directory in which original statistic maps downloaded from NeuroVault will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`.
2. preprocessed
The preprocessed directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset
Each subdirectory contains:
- text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ;
- csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ;
- several subdirectores (`resampled`, `resampled masked`...) in which preprocessed statistic maps will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`.
3. derived
The derived directory contains 3 subdirectories:
- NeuroVault dataset
- HCP dataset
- BrainPedia dataset
Each subdirectory contains subdirectories in which the parameters of models trained on the different datasets are stored. These subdirectories are named in the following way:
{name_of_the_dataset}_maps_classification_{classification_task}_model_cnn_{model_architecture}_valid_{type_of_experiment}_retrain_{type_of_initialization}_{preprocessing_type}_epochs_{number_of_epochs}_batch_size_{batch_size}_lr_{learning_rate}
For instance, parameters for the following experiment:
- Dataset: HCP Dataset subset 50 subjects
- Classification task: contrast classification
- Model: 4 layers CNN
- Type of experiment: Performance evaluation
- Initialization: Default
- Preprocessing type: Resampled masked normalized
- Epochs: 500
- Batch: 32
- Learning rate: 1e-04
will be contained in the directory:
hcp_dataset_50_maps_classification_contrast_model_cnn_4layers_valid_perf_retrain_no_resampled_masked_normalized_epochs_500_batch_size_32_lr_1e-04