Published March 26, 2025
| Version v1
Dataset
Open
Additional data for the paper "CREMSA : Compressed indexing of (ultra) large alignments"
Authors/Creators
Description
Three datasets used in the paper “CREMSA : Compressed indexing of (ultra) large alignments” made available here for reproducibility:
random_datasets_n10000_m30000.zip: An artificial dataset generated as described in the paper.HIV1_ALL_2022_genome_DNA.fasta.xz: A multiple sequence alignment of 5,381 HIV1 genomes, retrieved from the Los Alamos National Laboratory on March 2025.MFS_1.fasta.xz: A multiple sequence alignment of 214,283 protein sequences of the Major Facilitator Superfamily (MFS), retrieved from Pfam on March 2025.