K-FluDB
Authors/Creators
Description
K-FluDB: A Novel K-Mer Based Database for Enhanced Genomic Surveillance of Influenza A Viruses
K-FluDB is a compressed database composed of distinct sub-sequences specific to 50 influenza A subtypes. It includes subtype-specific sequences for all 18 hemagglutinin (HA) and 11 neuraminidase (NA) subtypes. The original influenza sequences were obtained from the NCBI database on May 8, 2022, comprising a total of 895,900 Influenza A sequences.
To generate this database, sequences were first subsampled based on genomic segment and variant group, resulting in 81,262 sequences. These sequences served as input for the PanGen-InfluenzaA tool (GitHub), which constructs pangenomes by identifying both subtype-specific sequences and sequences shared across multiple subtypes.
Repository Contents
This repository contains a ZIP archive with three folders, each corresponding to pangenome datasets designed for reads of 75, 150, and 300 nucleotides in length. Each folder includes the following files:
-
Dispensable data files (
1_dispensable.fastato8_dispensable.fasta):
These files contain the dispensable genomic fragments for each of the eight segments of the Influenza A virus. -
Subtype-specific data files (
1_specific.fastato8_specific.fasta):
These files contain the subtype-specific genomic fragments for each segment. The recommended files for mapping against genomic reads are4_specific.fastaand6_specific.fasta, corresponding to segments 4 and 6, which are the targets commonly used for Influenza A subtyping.
Within each core directory (75, 150, and 300), three distinct subdirectories—specific, pangenome, and dispensable—contain the pangenome files stratified by subtype. Specifically, the files located in the specific subdirectories for segments 4 and 6 are the recommended datasets for use in subtype identification during subsequent genomic analysis.
Compression Efficiency and Classification Accuracy
K-FluDB achieves a relative compression index of 96.54% when using the complete pangenome and 99.64% when considering only subtype-specific sequences. The average precision for correctly classifying Hx and Nx subtypes using the subtype-specific sequences is 99.2% and 99.71%, respectively.
This database provides a highly efficient and accurate resource for influenza A subtype classification while significantly reducing the storage and computational requirements associated with full-genome analyses.
Acknowledgements
This work has been supported by the Universidad Nacional Autónoma de México grant number [PAPIIT-DGAPA-IN230523] granted to Blanca Taboada and Secretaría de Educación, Ciencia, Tecnología e Innovación de la Ciudad de México with grant number [SECTEI/138/2024] granted to Selene Zárate.
The first author gratefully acknowledges the scholarship provided by CONAHCYT. We also extend our sincere appreciation to the National Autonomous University of Mexico (UNAM) for grant-ing access to the MIZTLI supercomputer, supported by the Gen-eral Directorate of Computing and Information and Communica-tion Technologies (DGTIC) through project LANCAD-UNAM-DGTIC-350. Lastly, we wish to thank Jerome Verleyen, Juan Manuel Hurtado, and Roberto Bahena from UNAM’s Instituto de Biotecnología for their indispensable assistance with computation-al support.
Files
K-FluDB_sep_25v2.zip
Files
(14.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:198d544a7c81d7a615c7b55f74c9a420
|
14.1 MB | Preview Download |
Additional details
Software
- Repository URL
- https://github.com/usjunco/pangen