Published June 5, 2019 | Version v1

Diamond formatted protein database for taxonomic classification

Authors/Creators

  • 1. Dept of Biochemistry and Biophysics, National Bioinformatics Infrastructure Sweden, Science for Life Laboratory, Stockholm University, Box 1031, SE-17121 Solna, Sweden

Description

This is a diamond formatted database (diamond version 0.9.22) built on December 14th 2018.

The database contains a total of 17,694,143 sequences:

  • 2,708,401 protein sequences from the taxmapper database (commit 450d337), containing 121 unique taxa.
  • 14,976,193 protein sequences from 1055 unique fungal taxa (downloaded from JGI 1000 fungi project on November 23 2018).
  • 9,549 protein sequences from the Hygrophorus russula genome obtained from Genbank (accession GCA_003314125.1) on November 28 2018.

The Hygrophorus russula  protein sequences were obtained by running Augustus (v. 3.2.3) gene caller on the genomic fasta file using the laccaria_bicolor gene model.

Taxonomic information was built into the diamond database by running:

zcat fasta.gz | diamond makedb -d diamond -p 4 --taxonmap taxonmap.gz --taxonnodes nodes.dmp

The nodes.dmp file was obtained from the taxdump.tar.gz file on December 11 2018.

Files

Files (8.0 GB)

Name Size
md5:99c5c00564b2b34fe8515cdff4f3363d
8.0 GB Download