Published June 23, 2025 | Version v2

Complete gene content and sequences for all Hpar samples in CHUVI pangenomic analysis study

  • 1. Microbiology and Infectology Research Group, Galicia Sur Health Research Institute (IIS Galicia Sur), Vigo, 36312, Spain

Description

259 genomes of Hamophilus parainfluezae were annotated, with genes assigned to a KEGG pathway. The following dataset includes, for each sample, all genes with their protein sequence, gene description and KEGG pathway assigned. This output was generated using prokka (for genetic annotation and protein sequence generation) and KEGGREST (for KEGG ortholog and pathway assignment).

Columns of the dataset work as follows:

  • SampleID: Sample ID from the SRA database (NCBI)

Columns from prokka output:

  • locus_tag: unique identifier for each sequence generated by prokka

  • ftype: type of genomic feature (CDS = coding sequence; CRISPR; tRNA = transference RNA; tmRNA transfer-messenger RNA)

  • length_bp: length of genomic feature in base pairs

  • gene: gene name

  • EC_number: Enzyme Comission number

  • COG: Cluster of Ortholog Groups family

  • Sequence: protein aminoacid sequence (for CDS types)

Columns from KEGGREST output:

  • ortholog: KEGG Orthology database identifier

  • Pathway: KEGG Pathway database identifier

 

Files

complete_gene_content_Hpar.csv

Files (214.9 MB)

Name Size Download all
md5:d29f5f80cd791b009131884465a471fa
214.9 MB Preview Download

Additional details

Dates

Available
2025-01-06