NLP4Pheno v1.0: Bacterial Phenotype Predictions Dataset
Authors/Creators
Description
This dataset contains the complete output from the NLP4Pheno pipeline, including Named Entity Recognition (NER) predictions, Relation Extraction (RE) predictions, and XGBoost-based phenotype predictions for 121,042 bacterial strains extracted from PubMed Central literature.
The dataset comprises:
- Named Entity Recognition predictions for 9 entity types (STRAIN, SPECIES, ISOLATE, COMPOUND, MEDIUM, ORGANISM, PHENOTYPE, EFFECT, DISEASE)
- Relation Extraction predictions for 16 relationship types (primarily STRAIN-centered)
- 1,046,299 aggregated strain-phenotype-compound relationships
- XGBoost predictions integrating genomic features (protein domains) with text-mined relationships
- Network representations suitable for graph analysis
This dataset corresponds to the manuscript: "Integrating natural language processing and genome analysis enables accurate bacterial phenotype prediction"
Files
README.md
Files
(10.4 GB)
| Name | Size | |
|---|---|---|
|
md5:aeb41f7738cee6e245c88caec223f4c4
|
7.3 GB | Download |
|
md5:33f6087fe3978a5e73e88796cde8e19e
|
60.7 MB | Download |
|
md5:771cc160199ea2e7e78632f31f0de84b
|
12.8 kB | Preview Download |
|
md5:1e1a0fd2e1c4204a413a5c7a68931edd
|
969.9 MB | Download |
|
md5:d3435aeb104ee8a0b1ce18f914f82c6b
|
59.1 MB | Download |
|
md5:d66035ef54017560854be704a181b830
|
2.0 GB | Download |