Steering potentials, not bug fixes, eliminate catastrophic outliers in Boltz-2 cofold protein-ligand affinity prediction: a six-way protocol evaluation on zinc-hydroxamate-like MMP-1 active-site ligands
Authors/Creators
- 1. Genesis_Medicine Lab; HAN PREDICT, Inc.; Recover Korean Medicine Clinic
Description
Correction (new version, 2026-07-18) — integrity correction. Corrected version (2026-07-18), identity disclaimer for a fabricated ligand panel. A primary-source audit (2026-07-16) established that the calibration panel's compound names and potencies are fabricated. This version softens 'inhibitors' to 'active-site ligands' in the title/subtitle, rewrites the Section 2.1 selection sentence to describe 15 structures under nominal ChEMBL identifiers used only as structure handles (removing the single pIC50 statement), and adds an identity note. The entire scientific result — that the --use_potentials steering flag eliminates catastrophic cofold-pose outliers — is unaffected. Versioned correction, not a retraction. In silico only.
End-to-end protein-ligand structure prediction with affinity heads (Boltz-2, AlphaFold3-class cofolding) has become a routine first-pass tool for virtual screening, but the rate at which these models emit physically implausible poses — and how that rate depends on inference-time configuration — remains incompletely characterised. We evaluate six published cofold configurations on 15 ChEMBL MMP-1 zinc-hydroxamate active-site ligands (1500 cofold poses per condition; 9000 total), using GFN2-xTB single-point energies on the predicted ligand geometries as a physical-plausibility readout. Standard Boltz-2 without the --use_potentials (Boltz-2x) steering flag emits catastrophic outliers (per-ligand population σ up to 14.27 kcal mol⁻¹ on CHEMBL94487, with one pose reaching an implausible +8911 kcal mol⁻¹ relative energy). Enabling --use_potentials reduces σ on the same ligand to 3.18-4.29 kcal mol⁻¹ across three independent seeds, with zero positive-energy outliers in 4500 samples (0.022%, vs. 0.188% without the flag, 9× reduction). A recently released community fork of Boltz-2 (Volgin et al., March 2026) that fixes a silent wrong-answer bug, a metal-ion C-alpha filter, and a bfloat16 dtype path does not eliminate the outliers when run without the steering flag (σ_filtered = 6.66 kcal mol⁻¹; 2/100 catastrophic poses on the canary ligand). Adding --use_potentials to the fork removes the outliers but the within-population precision remains ~2× worse than standard Boltz-2x (σ_filt 6.98 vs. 3.18 kcal mol⁻¹). The operative factor is therefore the steering-potential flag, not the bug fixes. We recommend standard Boltz-2 + --use_potentials as the canonical cofold protocol for protein-ligand affinity prediction, especially for metalloprotein active sites; a residual ≤0.025% catastrophic-failure rate remains, so xtb-based filtering of cofold poses is still mandatory downstream.
Notes
Files
manuscript_corrected_2026-07-18.md
Files
(110.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:a32d0156451696e6e4cda2577f4994ef
|
33.0 kB | Preview Download |
|
md5:62d492919183f15668ba95b4d903bd2f
|
77.7 kB | Preview Download |