Processed features in support of Liebeskind et al (2018)
Authors/Creators
- 1. University of Texas at Austin
Description
Processed feature matrices used in Liebeskind et al. (2018). Supporting code: https://github.com/marcottelab/plum
Datasets 1 - 4 correspond to those used in Figure 4:
Dataset 1: No AP-MS, yeast CF-MS, training species: Human
Dataset 2: AP-MS, yeast CF-MS, training species: Human
Dataset 3: AP-MS, yeast CF-MS, training species: Human, Yeast
Dataset 4: AP-MS, no yeast CF-MS, training species: Human, Yeast
".train_labeled.missing_annotated.csv" files are those used for training the model and include only orthogroups for which interactions are known in the training species. These known interactions come either from gold-standard test sets, such as CORUM or EMBL's training portal, or from the fact that at least of the orthogroup pairs is missing in the focal taxon.
".missing_annotated.csv" files were used for prediction, and include the entire feature matrices, plus known missing pairs. Note that there is no dataset 3 file. This is because data sets 2 and 3 differ only in the training species used, so dataset 3 predictions used dataset2_07302018.missing_annotated.csv as a feature matrix.
dataset4_prediction_07302018.csv contains the predictions for all pairs on data set 4, the best performing data set that was used for all downstream analyses.
Files
Files
(10.8 GB)
| Name | Size | |
|---|---|---|
|
md5:e74ea436e00bd6b24528c8957e5ddc5a
|
2.3 GB | Download |
|
md5:cf645f4261fda5ae14b6f6ec919d1544
|
28.9 MB | Download |
|
md5:ce95e2470ed6cbe7beb27493f5bcd359
|
2.4 GB | Download |
|
md5:8a3bb1035e038160feaeb791cb3acccb
|
30.0 MB | Download |
|
md5:eae5edf338b308fc540d1941a8807b8f
|
38.7 MB | Download |
|
md5:3e6c527e7d2b10369bc994ed8ccde370
|
2.3 GB | Download |
|
md5:825b0d950705b78e4fb80e0354a1b662
|
36.6 MB | Download |
|
md5:cef3e178a46f58bf6fe835827ce30c9e
|
3.7 GB | Download |
Additional details
Related works
- Is supplement to
- 10.5281/zenodo.1406146 (DOI)