Published June 13, 2026 | Version v2

Example AnnData for label transfer

Authors/Creators

  • 1. ROR icon University College London

Description

To simulate the situation of having new data to annotate & integrate, we use here a scRNA-seq study (https://pmc.ncbi.nlm.nih.gov/articles/PMC7439502/) of lung cells from idiopathic pulmonary fibrosis (IPF), a fatal interstitial lung disease. The raw data can be found here on the GEO database (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE136831).

The .h5ad data object here was generated in the following way:

  1. .gz files were downloaded from GSE136831 (mtx of transcript counts, cell metadata, gene and cell identifiers) and assembled into an AnnData object (n_obs × n_vars = 312928 × 45947).
  2. The following cell types (from metadata: "Manuscript_Identity" column) were retained: [ "Macrophage",  "Macrophage_Alveolar",  "cMonocyte", "ncMonocyte", "Fibroblast", "Myofibroblast", "Aberrant_Basaloid"]
  3. The "Macrophage", "Macrophage_Alveolar",  "cMonocyte" and "ncMonocyte" cell types were overrepresented relative to the other cell-types retained. We therefore downsampled specifically these 4 cell types to 5000 cells each. All cells labelled as the other retained cell-types were retained.
  4. The cells from IPF samples (metadata column "Disease_Identity" == "IPF") were further subsetted to leave n = 12651 cells.

 

Files

Files (378.2 MB)

Name Size
md5:7318cbe4e8cef43ec29c15f5d332f478
378.2 MB Download