Published August 30, 2018 | Version v1.0.0

Processed features in support of Liebeskind et al (2018)

  • 1. University of Texas at Austin

Description

Processed feature matrices used in Liebeskind et al. (2018). Supporting code: https://github.com/marcottelab/plum

Datasets 1 - 4 correspond to those used in Figure 4:

Dataset 1: No AP-MS, yeast CF-MS, training species: Human

Dataset 2: AP-MS, yeast CF-MS, training species: Human

Dataset 3: AP-MS, yeast CF-MS, training species: Human, Yeast

Dataset 4: AP-MS, no yeast CF-MS, training species: Human, Yeast

".train_labeled.missing_annotated.csv" files are those used for training the model and include only orthogroups for which interactions are known in the training species. These known interactions come either from gold-standard test sets, such as CORUM or EMBL's training portal, or from the fact that at least of the orthogroup pairs is missing in the focal taxon.

".missing_annotated.csv" files were used for prediction, and include the entire feature matrices, plus known missing pairs. Note that there is no dataset 3 file. This is because data sets 2 and 3 differ only in the training species used, so dataset 3 predictions used dataset2_07302018.missing_annotated.csv as a feature matrix.

dataset4_prediction_07302018.csv contains the predictions for all pairs on data set 4, the best performing data set that was used for all downstream analyses.

Files

Files (10.8 GB)

Name Size
md5:e74ea436e00bd6b24528c8957e5ddc5a
2.3 GB Download
md5:cf645f4261fda5ae14b6f6ec919d1544
28.9 MB Download
md5:ce95e2470ed6cbe7beb27493f5bcd359
2.4 GB Download
md5:8a3bb1035e038160feaeb791cb3acccb
30.0 MB Download
md5:eae5edf338b308fc540d1941a8807b8f
38.7 MB Download
md5:3e6c527e7d2b10369bc994ed8ccde370
2.3 GB Download
md5:825b0d950705b78e4fb80e0354a1b662
36.6 MB Download
md5:cef3e178a46f58bf6fe835827ce30c9e
3.7 GB Download

Additional details

Related works

Is supplement to
10.5281/zenodo.1406146 (DOI)