There is a newer version of the record available.

Published May 29, 2026 | Version v1.0.0

Enveda-180: generation of a large open multimodal MS/MS and ion mobility spectral library for drug-like small molecules

Description

Small-molecule characterization depends heavily on annotated MS/MS spectral libraries for reference matching and, in the case of open-source libraries, for training machine-learning-based structure elucidation. A primary bottleneck is the scarcity of such open data: current libraries such as GNPS, MSnLib, and MassBank, combined, contain spectra for approximately 150,000 unique compounds, representing a small fraction of known chemical space. These libraries are also sparse in complementary measurements such as liquid chromatography retention time (RT) and collisional cross section (CCS) from ion mobility mass spectrometry, the latter of which is seeing growing adoption in the field. To address these gaps, we present Enveda-180, an open spectral library built from high-throughput, pooled characterization of 184,330 structurally diverse small molecules, more than doubling the combined structural coverage of MS/MS open libraries, and more than 5x the combined open CCS values. Compounds were profiled at three collision energies in both positive and negative ion modes. We describe our profiling methods in detail, including the use of ion mobility to resolve multiple species of the same nominal ion structure. Enveda-180 substantially expands the chemical space covered by open MS/MS libraries and enables multimodal learning by integrating retention time, CCS, MS/MS, and metadata within a single resource.

Technical info (English)

Enveda-180 is a curated tandem-MS spectral library generated from high-throughput, pooled LC-MS/MS profiling of the Enamine HLL-200 hit-locator collection. The release comprises 1,158,293 curated MS/MS spectra for 184,330 unique compounds (distinct full InChIKeys; 183,476 distinct InChIKey-14 connectivity blocks). Each compound is represented by a median of 6 spectra (mean 6.28, maximum 48), and 99.3% of compounds carry two or more spectra.

Acquisition. Data were acquired on a Bruker timsTOF Pro 2 trapped-ion-mobility Q-TOF operating in PASEF mode, under heated electrospray ionization (300 °C, 4500 V) in both positive and negative ion modes. Full-scan range was m/z 20 to 1300 with an ion-mobility range of 0.3 to 1.45 V·s·cm⁻². Each selected precursor was fragmented at three nominal collision energies (20, 40, 60 eV). Liquid chromatography used a Waters ACQUITY Premier BEH C18 column (2.1 × 50 mm, 1.7 µm) with water and acetonitrile mobile phases containing 0.1% formic acid, a 5% to 95% acetonitrile gradient over 5.5 minutes (7 minutes total), 0.5 mL/min flow, and a 5 µL injection.

Adducts and polarity. Spectra span 14 adducts across both polarities (ESI+ 77.0%, ESI− 23.0%). [M+H]+ and [M−H]− together account for roughly 67% of all spectra. The full distribution:

Adduct Count
[M+H]+ 590,683
[2M+Na]+ 188,432
[M-H]- 185,030
[2M+H]+ 89,625
[2M-H]- 53,005
[M-H+FA]- 24,648
[M+Na]+ 16,153
[M+NH4]+ 5,913
[M+Cl]- 2,410
[M+K]+ 1,097
[2M-H+FA]- 663
[M-H+Hac]- 411
[2M-H+Hac]- 198
[M+Br]- 25

Ion mobility and CCS. 96% of spectra carry an experimental collision cross section derived from calibrated 1/K0 reference points. Enveda-180 reports multiple mobility-resolved CCS values per precursor rather than a single value: among CCS-bearing spectra, 56% carry two or more species and 31.5% carry three or more, with a median inter-species ΔCCS of 11.8 Ų (6.6% of the primary CCS). These secondary species capture protomers, isomers, conformers, and monomer/dimer pairs that share a nominal m/z and retention time.

Spectral quality. Spectra passed multi-criteria curation combining SIRIUS rank-1 fragmentation-tree scoring, explained-intensity and minimum-peak thresholds, precursor-integrity checks, and round-trip InChIKey validation. The rank-1 tree explains a median of 88.6% of total fragment intensity; 91.4% of spectra exceed the 70% explained-intensity level targeted by manual annotators, and 98.1% exceed 50%. Median normalized spectral entropy is 0.44. Spectra contain a median of 138 peaks (minimum 4). We also provide an intensity-filtered version of the spectra, with fragment peaks below 1% relative intensity removed.

Files and identifiers. The release ships MGF, MSP, JSONL and Parquet serializations of the spectra. Each spectrum has a unique, sortable identifier of the form enveda_<compound index>_<spectrum index>. 

Files

Files (8.9 GB)

Name Size
md5:9b6f5b7c1fe35bd2dcc4046f841cfc33
298.2 MB Download
md5:e9beaec694c9c86409bbdaf737a3672a
272.7 MB Download
md5:a73ad6f47c2e9cf6002aa4e531aa55d5
273.4 MB Download
md5:5a46d40040ac7aad5c879c70cf73f103
349.0 MB Download
md5:a405415d649e5b62ba64beede7f483c9
1.8 GB Download
md5:2b6a10f3688f8c7c132d04873e105d5b
1.7 GB Download
md5:03fc1ddc6f3d5d7305959569522c965e
1.7 GB Download
md5:1ba320853ca954010d8f25dbab6bdb89
2.6 GB Download

Additional details

Software

Repository URL
https://github.com/enveda/enveda-180
Programming language
Python
Development Status
Active