Clusters of Alpha- and Betaproteobacterial Genomes for in silico Prospecting of Nodulating Diazotrophs
Authors/Creators
Description
Clusters of Alpha- and Betaproteobacterial Genomes for in silico Prospecting of Nodulating Diazotrophs
File descriptions below.
File: AlphaBeta_sequences.fas
Header pattern format:
The headers contain the follow markers:
- id_#_[protein name] - where # is the number related to the protein, markerd followed by the protein name
- ProductID - WP id
- Organism - species
- AccessionID - NCBI accession number
example:
>id_5_leucine--tRNA ligase ProductID-WP_025456399N1 Organism-Neisseria gonorrhoeae FA 1090 AccessionID-NC_002946N2
File: clusters.txt
The file contains the cluster number corresponding to each protein in MULTIFASTA (AlphaBeta_sequences.fas), in the same order.
File: pointers.txt
The file contains the pointer to the beginning of each sequence in the (AlphaBeta_sequences.fas) file. That is, each number in the pointers.txt file corresponds to the start character of each sequence (in its header).
File: dbstructAB.mat
File containing cluster information in MATLAB format, for internal use by the RAFTS3G software (Nichio, et al., 2019).
References