There is a newer version of the record available.

Published November 7, 2024 | Version v1

Data for evaluation of diverse-seq algorithms

  • 1. Australian National University

Contributors

Contact person:

  • 1. Australian National University

Description

The algorithms required for phylogenetics — multiple sequence alignment and phylogeny estimation — are both compute intensive. diverse-seq implements computationally efficient alignment-free algorithms that enable efficient prototyping for phylogenetic workflows. It can accelerate parameter selection searches for sequence alignment and phylogeny estimation by identifying a subset of sequences that are representative of the diversity in a collection. diverse-seq can further boost the performance of phylogenetic estimation by providing a seed phylogeny that can be further refined by a more sophisticated algorithm.

The data sets in this archive are either HDF5 stored whole microbial genomes or multiple sequence alignments of one-to-one orthologs from mammal species. The `wol.dvseqs` HDF5 file is derived from the data used in Zhu et al Nature Communications, 10(1), 5477 with the original fasta formatted files in wol.zip. The `soil.dvseqs` HDF5 file is derived from the genomes included in REFSOIL (Choi et al The ISME Journal, 11(4), 829–834), with the original genbank formatted files included in refsoil.zip. The data in `mammal-aligned.zip` are fasta formatted multiple sequence alignments of sequences sampled from Ensembl release 112.

Files

mammals-aligned.zip

Files (22.0 GB)

Name Size
md5:08160398b72a78a5ae736a0bc118d477
1.4 MB Preview Download
md5:50e1b0d7583b6bd49bffa693af236cf5
2.6 GB Preview Download
md5:8c1a1f14d3244c98f87839597cc80d14
1.1 GB Download
md5:97b590e68d7a6015950e442fc9ae0f02
8.9 GB Download
md5:d21867e9f947408e038322d5f84cc148
9.4 GB Preview Download

Additional details

Dates

Available
2024-11

Software

Repository URL
https://github.com/HuttleyLab/DiverseSeq
Programming language
Python
Development Status
Active