There is a newer version of the record available.

Published July 29, 2021 | Version v1
Dataset Open

Data and Weights for Reverse Homology


Training data, weights, and classification datasets for "Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning". 

Training data:

  • scer_idr_homologues and human_idr_homologues contain a zip file of the fasta files of IDR homologues used to train the yeast and human model, respectively. Note that these fasta files are aligned, but we strip away the alignment symbol "-" before input into our model. 


  • scer_idr_model and human_idr_model contain a zip file of the weights for the yeast and human model respectively, which can be loaded into the model files at It also contains z-scores of all of the model features across all IDRs, which are required to run the mutational scanning map code. 

Classification datasets:

  • IDR_classification_datasets contains datasets used in our benchmarks. These datasets are encoded as binary csv matrixes. cdc28_classification contains IDRs labeled as Cdc28 phosphorylation sites, mitochondrial_targeting_classification contains IDRs labeled as mitochondrial targeting signals, evosig_cluster_classification contains IDRs labeled by clusters assigned in previous computational work by Zarin et al. eLife 2019, and go_SLIM_classification contains proteins labeled by GO Slim annotations. 


Files (156.3 MB)

Name Size Download all
86.5 MB Preview Download
43.1 MB Preview Download
167.7 kB Preview Download
6.6 MB Preview Download
19.9 MB Preview Download