Published December 8, 2021
| Version 1.0
Dataset
Open
scMARK an 'MNIST' like benchmark to evaluate and optimize models for unifying scRNA data
Description
Here we present a novel benchmark dataset (scMARK), that consists of 100,000 cells and 10 studies and can be used to ask how well models integrate data from different scRNA studies.
- Data is provided as aData *h5ad files that can be read with Python's library Scanpy.
- File 10k_cells_per_study.tar.bz2 contains scMARK, with one *h5ad file for each of 10 studies, downsampled to 10,000 cells per study.
- File 2k_cell_per_study_10studies.tar.bz2 contains a single *h5ad file with all 10 studies, downsampled to 2,000 cells per study. The matrix of UMI counts contains the intersection of genes in all 10 studies.