Published October 29, 2025 | Version v1

NLP4Pheno v1.0: Bacterial Phenotype Predictions Dataset

Authors/Creators

Description

  This dataset contains the complete output from the NLP4Pheno pipeline, including Named Entity Recognition (NER) predictions, Relation Extraction (RE) predictions, and XGBoost-based phenotype predictions for 121,042 bacterial strains extracted from PubMed Central literature.

  The dataset comprises:
  - Named Entity Recognition predictions for 9 entity types (STRAIN, SPECIES, ISOLATE, COMPOUND, MEDIUM, ORGANISM, PHENOTYPE, EFFECT, DISEASE)
  - Relation Extraction predictions for 16 relationship types (primarily STRAIN-centered)
  - 1,046,299 aggregated strain-phenotype-compound relationships
  - XGBoost predictions integrating genomic features (protein domains) with text-mined relationships
  - Network representations suitable for graph analysis

  This dataset corresponds to the manuscript: "Integrating natural language processing and genome analysis enables accurate bacterial phenotype prediction"

 

Files

README.md

Files (10.4 GB)

Name Size
md5:aeb41f7738cee6e245c88caec223f4c4
7.3 GB Download
md5:33f6087fe3978a5e73e88796cde8e19e
60.7 MB Download
md5:771cc160199ea2e7e78632f31f0de84b
12.8 kB Preview Download
md5:1e1a0fd2e1c4204a413a5c7a68931edd
969.9 MB Download
md5:d3435aeb104ee8a0b1ce18f914f82c6b
59.1 MB Download
md5:d66035ef54017560854be704a181b830
2.0 GB Download