Published August 30, 2026 | Version 1.0.0

PHORAGER - mining prophages from ESKAPE pathogens

Description

ESKAPE pathogen prophage repertoire : PHORAGER output collection

This archive contains the key prophage-mining outputs generated by PHORAGER (Prophage Hunting, vOTU Retrieval, Annotation and Genomic ExploRation) applied to a large collection of publicly available ESKAPE pathogen genomes.

- Pipeline: PHORAGER v0.5.0-beta https://github.com/aponsero/PHORAGER
- Associated manuscript: Defining the ESKAPE pathogen prophage repertoire with PHORAGER (Dyball et al. bioRXiv)
- Generated by: the PHORAGER `bacterial`, `prophage`, `annotation`, and `summarize` workflows, run with default parameters.

------------------------------------------------------------------------------------------------------------------------------------------------

1. What is in this archive

The collection is split into one ZIP file per ESKAPE pathogen. Each ZIP unpacks to a single folder containing five files for that species:

File (per species) Type Content PHORAGER stage
<Species>_genomes_metadata.tsv TSV NCBI metadata for the input genomes external from NCBI
<Species>_bacterial_genome.tsv TSV Per-genome QC metrics + prophage counts `summarize` (bacterial_genome)
<Species>_filtered_phage_set.fasta FASTA (nt) Prophage sequences after annotation-based filtering, before dereplication `annotation` (post-filter)
<Species>_final_representatives.fasta FASTA (nt) vOTU cluster representative prophages `annotation` (post-clustering)
<Species>_prophage_table.tsv TSV Per-prophage table (host, cluster, quality, length, CDS) `summarize` (prophage_table)

Species included (six ZIPs):

- Enterococcus faecium
- Staphylococcus aureus
- Klebsiella pneumoniae
- Acinetobacter baumannii
- Pseudomonas aeruginosa
- Enterobacter spp.

Folder / file naming. Folders follow `<Species>_<DD.MM.YY>`, where the date is the NCBI download/build date for that species' genome set (these differ between species, e.g. most ESKAPE genomes were retrieved 10 Oct 2025, K. pneumoniae on 22 Oct 2025). Files within a folder are prefixed with the same `<Species>` token.

------------------------------------------------------------------------------------------------------------------------------------------------

2. File descriptions

2.1 `<Species>_genomes_metadata.tsv` : NCBI input metadata

Tab-separated metadata for the bacterial genomes downloaded from NCBI and used as input to PHORAGER. This file is produced by the NCBI download step, not by PHORAGER.

2.2 `<Species>_bacterial_genome.tsv` : per-genome summary

One row per genome that passed CheckM2 quality assessment. Columns (exact, as emitted
by the pipeline):

Column Description
Genome_name Genome basename (from CheckM2 `Name`); matches the input genome filename stem (e.g. the NCBI accession)
Completeness CheckM2 completeness estimate (%)
Contamination CheckM2 contamination estimate (%)
Genome_size Total assembly size (bp)
GC_content GC fraction
Contig_N50 Contig N50 (bp)
Coding_sequences Total predicted CDS count
Prophage_count Number of prophages attributed to this genome in `filtered_phage_set.fasta` (i.e. the filtered, pre-dereplication set). `0` if none passed filtering

> Note: `Prophage_count` is derived from the filtered (non-dereplicated) prophage set,
> by attributing each prophage to the genome named before the first `_ctg` in its ID.

2.3 `<Species>_filtered_phage_set.fasta` : filtered prophage sequences

Nucleotide FASTA of all prophage sequences that passed PHORAGER's annotation-based filtering, before vOTU dereplication. Each record corresponds to exactly one row in `prophage_table.tsv` (matched on the sequence ID). Sequence-ID format is described in §3.

2.4 `<Species>_final_representatives.fasta` : vOTU representatives

Nucleotide FASTA containing one representative sequence per vOTU cluster (95% ANI / 85% aligned fraction by default). These are extracted directly from `filtered_phage_set.fasta`, so every header here also appears in that file and equals a value in the `Cluster_representative` column of `prophage_table.tsv`. Use this file as the non-redundant, species-level prophage set for downstream analysis.

2.5 `<Species>_prophage_table.tsv` : per-prophage summary

One row per prophage in `filtered_phage_set.fasta`. Columns (exact, as emitted):

Column Description
Prophage_name Prophage sequence ID; matches a header in `filtered_phage_set.fasta`
Bacterial_host Host genome basename (everything before the first `_ctg` in `Prophage_name`); matches `Genome_name` in the bacterial table.
Cluster_representative The vOTU representative for this prophage's cluster. If equal to `Prophage_name`, this prophage is a representative (and appears in `final_representatives.fasta`). `Not_in_clusters` if no cluster assignment was found.
CheckV_quality CheckV quality tier (e.g. Complete / High-quality / Medium-quality). `NA` if not found in the CheckV summary.
Phage_length Sequence length (bp). Uses CheckV's trimmed `proviral_length` where available, otherwise `contig_length`.
CDS_number CDS count from the annotation step. `NA` if detailed annotation was skipped or the prophage was not matched.

------------------------------------------------------------------------------------------------------------------------------------------------

3. Prophage sequence-naming convention

Input contigs are renamed by PHORAGER to `{genome_basename}_ctg{NNN}`, where `genome_basename` is the input genome's filename stem and `NNN` is a zero-padded contig number (3–5 digits, depending on contig count per genome). Prophage records then take one of two forms:

- Coordinate-based (merged/consolidated predictions):
  `{genome_basename}_ctg{NNN}_{start}_{end}`
  Coordinates are 1-based and inclusive, relative to the renamed contig.
  Example: `GCF_000123456.1_ctg001_4521_38790`
- Whole-contig predictions:
  `{genome_basename}_ctg{NNN}_complete`
  Used when a mining tool flagged the entire contig as prophage.

CheckV may append a trailing `_1` to trimmed proviral sequences; where present, that suffix is part of the FASTA header and the `Prophage_name` value.

------------------------------------------------------------------------------------------------------------------------------------------------

4. How the files link together

The four PHORAGER files share consistent identifiers, so they can be joined directly:

- Prophage → host genome: `prophage_table.Bacterial_host` = `bacterial_genome.Genome_name`.
- Prophage row → sequence: `prophage_table.Prophage_name` = FASTA header in `filtered_phage_set.fasta`.
- Prophage → its vOTU representative: `prophage_table.Cluster_representative`; the representatives themselves are in `final_representatives.fasta`.
- Genome → NCBI metadata: `bacterial_genome.Genome_name` = the accession column in `genomes_metadata.tsv` *(confirm which column — see §2.1)*.

------------------------------------------------------------------------------------------------------------------------------------------------

5. How these outputs were generated

Genomes were processed with PHORAGER v0.5.0-beta using default parameters. In summary
(see the associated manuscript for full methods and citations):

1. Bacterial QC & dereplication : CheckM2 quality filtering (default ≥90% completeness, <5% contamination) followed by dRep dereplication at the strain level (default 99.9% ANI).
2. Prophage mining & consolidation : geNomad and VIBRANT run in parallel; overlapping predictions consolidated into non-redundant coordinate ranges; sequences extracted.
3. Annotation & filtering : CheckV quality assessment; length/quality filtering; annotation with Pharokka (and Phold); structural-gene-based filtering. The `filtered_phage_set.fasta` is the output of this stage.
4. vOTU clustering : BLAST + CheckV ANI scripts, clustered at the species level (default 95% ANI / 85% aligned fraction). `final_representatives.fasta` holds the cluster representatives.
5. Summaries : `bacterial_genome.tsv` and `prophage_table.tsv`.

Tool versions:

Tool Version
CheckM2 1.0.1
dRep 3.5.0
geNomad 1.8.1
VIBRANT 1.2.1
CheckV 1.0.3
Pharokka 1.9.1
Phold 0.2.0
BLAST+ 2.16.0

 

Files

Acinetobacter_baumannii_15.10.25.zip

Files (416.0 MB)

Name Size
md5:e208e9fc908ef2ffe98ad2969fca7dda
89.3 MB Preview Download
md5:aaddc8a6a5b05707754e24a04648b926
123.2 MB Preview Download
md5:e2976acb60d4feddb10c4535c443df0a
31.9 MB Preview Download
md5:22a8577b17e3e3d1ba6dc951b5fa49e0
94.4 MB Preview Download
md5:ee3137d335b83e576a5ade5e42dc34ba
28.6 MB Preview Download
md5:57d8c3793688d8106b88f7f13b523dba
48.5 MB Preview Download