Structured Nanotoxicity Datasets with Physicochemical and Toxicological Attributes of Metal Oxide Nanoparticles
Description
This repository provides three curated nanotoxicity datasets — HaHa-Auto, HaHa-Manual, and Ha IIIB — comprising structured physicochemical and toxicological information on metal oxide nanoparticles. The datasets are intended for use in predictive modeling, risk assessment, and data-driven research in nanotoxicology.
HaHa-Auto: A semi-automatically generated dataset using large language models (LLMs) with a prompt-engineered data extraction pipeline. Contains 2696 entries with 15 descriptors, filtered for quality based on a physicochemical score.
HaHa-Manual: A manually curated counterpart to HaHa-Auto based on the same source articles. Offers a higher number of entries (3440) for benchmarking automated extraction performance.
Ha IIIB: A high-quality, smaller dataset (666 entries) curated from previously published training data, used to evaluate model applicability and robustness with limited but highly reliable data.
Each dataset includes detailed metadata such as:
1. Core size, hydrodynamic size, surface charge, surface area
2. Quantum-mechanical descriptors (e.g., band gap, enthalpy of formation)
3. Biological assay conditions (e.g., cell type, exposure time, dosage)
4. Toxicity endpoints (e.g., cell viability)
These datasets are suitable for use in automated machine learning (AutoML) applications, (Q)SAR model development, or validation of nanotoxicity prediction pipelines. All data entries were filtered using a scoring system to ensure quality and consistency, and are provided in CSV format with clear headers and descriptors.