Published March 26, 2025 | Version v1

Additional data for the paper "CREMSA : Compressed indexing of (ultra) large alignments"

  • 1. ROR icon Laboratoire d'Informatique de l'École Polytechnique

Description

Three datasets used in the paper “CREMSA : Compressed indexing of (ultra) large  alignments” made available here for reproducibility:

  • random_datasets_n10000_m30000.zip : An artificial dataset generated as described in the paper.
  • HIV1_ALL_2022_genome_DNA.fasta.xz : A multiple sequence alignment of 5,381 HIV1 genomes, retrieved from the Los Alamos National Laboratory on March 2025.
  • MFS_1.fasta.xz : A multiple sequence alignment of 214,283 protein sequences of the Major Facilitator Superfamily (MFS), retrieved from Pfam on March 2025.

Files

random_datasets_n10000_m30000.zip

Files (645.6 MB)

Name Size
md5:d403c3b42ccbdbc8d854d3274631a9af
4.1 MB Download
md5:bf8e686e23a7e14c6a1dcbeb74b8669a
55.3 MB Download
md5:23e8056017806655c29f732996a30a16
586.2 MB Preview Download