Published October 19, 2024 | Version v2

Integration of protein and coding sequences enables mutual augmentation of the language model

Authors/Creators

Description

The file structure is as follows:

Project Root
├── TE_MRL
│   ├── MRL_dataset.zip
│   └── TE_dataset.zip

├── finetuned_model
│   ├── FoldP
│   ├── LocP
│   ├── SSP
│   └── SolP

├── tax_tsne
│   └── emb_3models.zip

└── training_data
    ├── FoldP.csv
    ├── LocP.csv
    ├── SolP.csv
    ├── SSP.pkl
    └── pretrain_source_GCF.txt

Files

source_data.zip

Files (449.9 MB)

Name Size
md5:6368a38719d70e1fe03f9e791b36d9cb
449.9 MB Preview Download