Published October 19, 2024
| Version v2
Dataset
Open
Integration of protein and coding sequences enables mutual augmentation of the language model
Authors/Creators
Description
The file structure is as follows:
Project Root
├── TE_MRL
│ ├── MRL_dataset.zip
│ └── TE_dataset.zip
│
├── finetuned_model
│ ├── FoldP
│ ├── LocP
│ ├── SSP
│ └── SolP
│
├── tax_tsne
│ └── emb_3models.zip
│
└── training_data
├── FoldP.csv
├── LocP.csv
├── SolP.csv
├── SSP.pkl
└── pretrain_source_GCF.txt