Published January 12, 2026 | Version v3

NGSTroubleFinder train dataset

Description

This dataset comprises high-throughput RNA and DNA sequencing data generated from biological samples of the 1000GP. Samples have been admixed to create files with known contamination status.

The dataset includes both non-contaminated (clean) and contaminated samples and the source code from commit "9671efea84cd11e12bbd868dbe5ac763854d8334" that has been used to train the model for contamination detection. The full github history is available at https://github.com/STALICLA-RnD/NGSTroubleFinder

 

 

Files

Files (17.7 GB)

Name Size
md5:a831d61bd4d705163545a7e2e923293c
17.7 GB Download
md5:05cb722028c087b802e0ea4eb67fe49c
8.7 MB Download

Additional details

Related works

Is supplement to
Preprint: 10.1101/2025.01.31.635690 (DOI)

Funding

European Commission
REPO4EU - Precision drug REPurpOsing For EUrope and the world 101057619

Dates

Submitted
2025-04-07

Software

Repository URL
https://github.com/STALICLA-RnD/NGSTroubleFinder
Programming language
Python , C
Development Status
Active