Complete gene content and sequences for all Hpar samples in CHUVI pangenomic analysis study
Authors/Creators
- 1. Microbiology and Infectology Research Group, Galicia Sur Health Research Institute (IIS Galicia Sur), Vigo, 36312, Spain
Description
259 genomes of Hamophilus parainfluezae were annotated, with genes assigned to a KEGG pathway. The following dataset includes, for each sample, all genes with their protein sequence, gene description and KEGG pathway assigned. This output was generated using prokka (for genetic annotation and protein sequence generation) and KEGGREST (for KEGG ortholog and pathway assignment).
Columns of the dataset work as follows:
-
SampleID: Sample ID from the SRA database (NCBI)
Columns from prokka output:
-
locus_tag: unique identifier for each sequence generated by prokka
-
ftype: type of genomic feature (CDS = coding sequence; CRISPR; tRNA = transference RNA; tmRNA transfer-messenger RNA)
-
length_bp: length of genomic feature in base pairs
-
gene: gene name
-
EC_number: Enzyme Comission number
-
COG: Cluster of Ortholog Groups family
-
Sequence: protein aminoacid sequence (for CDS types)
Columns from KEGGREST output:
-
ortholog: KEGG Orthology database identifier
-
Pathway: KEGG Pathway database identifier
Files
complete_gene_content_Hpar.csv
Files
(214.9 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:d29f5f80cd791b009131884465a471fa
|
214.9 MB | Preview Download |
Additional details
Dates
- Available
-
2025-01-06