Published May 4, 2023 | Version v1

"cleanventory": A Data Science Approach to Identify Persistent and Mobile Substances Regulated in Global Trade Markets

  • 1. Norwegian Geotechnical Institute (NGI)
  • 2. University of Luxembourg
  • 3. Empa

Description

To enable data-driven decision making and to advance the ability to assess, manage and regulate the use of persistent and mobile substances, a sound overview of the chemical diversity regulated within global trade markets is needed. However, this information is notoriously difficult to obtain. Official documentation and lists, if available, come in a diverse range of document types and sometimes entirely different annotations and formats. With the number of regulated chemicals rising, there is need for a harmonized information system on regulated chemicals.

A pioneering effort to index the global chemical inventory, i.e., all chemicals registered on (super-)national economic trade markets, was first published by Wang et al. in 2020. This ground-breaking work has enabled a greater understanding of the regulatory chemical space and allowed for initial widescale analyses of the inventories and their chemicals.

As part of the H2020 project ZeroPM (https://zeropm.eu), a fully reproducible and open-source re-construction of the global chemical inventory – the "cleanventory" – is being developed to help identify, prioritize, and group persistent and mobile substances. A modern database infrastructure will facilitate wide-spread use of the database (also outside the context of persistent and mobile substances), strictly following FAIR principles (Findable, Accessible, Interoperable, Reproducible). The database will be publicly available and will also include features for programmatic access. Snapshots of the database will be continuously posted on ZeroPM's Zenodo community (https://zenodo.org/communities/zeropm-h2020) and all code will be made available on ZeroPM's GitHub repository (https://github.com/ZeroPM-H2020).

All data sources listed in Wang et al. are being re-evaluated and updated, and new inventories are being identified. So far, 33 individual files from 18 inventories in twelve trade markets are considered. These contain over 960,000 inventory entries with over 215,000 unique CAS Registry Numbers and over 360,000 unique chemical names (but not every inventory entry contains both identifiers). While this amount of information could be considered "big data", we put considerable efforts towards the quality of the data, i.e., also ensuring "good data". Extensive cleaning and curation of the data sources into a harmonized "cleanventory" format was necessary to ensure a high degree of compatibility.

Even so, structural information is not included in these inventories. To identify chemical structures from inventory entries, CAS Registry Numbers and chemical names are used, because they are the most common identifiers. To convert inventory identifiers to InChI strings (i.e., structural information), four freely available API services are used: PubChem (compound and substance domain), CAS Common Chemistry, NCI/CADD Chemical Identifier Resolver, and ChemSpider.

A visual example of the chemical structure identification workflow is shown in Figure 1. The workflow aggregates the structure information returned by the API services. Because of its design, this workflow can result in multiple InChI strings for every identifier; and for each API service. To identify the "most probable" chemical structure for every inventory entry (i.e., the combination of CAS Registry Number and chemical name), a weighted consensus ranking approach was developed to assign each InChI strings an identification score between 0 and 1. Furthermore, all identified InChI strings were assessed for the presence of multiple components via Open Babel.

So far, over 320,000 unique InChI strings were retrieved by the API services. After the weighted consensus ranking approach described above, 160,000 unique InChI strings are identified as being the "most probable" chemical structure for the given inventory entries. All InChI strings have full information traceability, e.g., about their presence in individual inventories and alternative ("less probable") chemical structures.

The work on the "cleanventory" continues and new data sources will be incorporated. These include important additions such as chemical substances regulated as pharmaceuticals, pesticides, and food-packaging materials. Additionally, workflows to identify corresponding SMILES strings are being investigated, along with an automated identification of "moieties of interest" for prioritization and grouping of persistent and mobile substances. Future work will also explore the integration of fully defined chemical mixtures as well as polymers and UVCBs, and explore the possibility of automatic integration of curated QSAR model predictions.

This high-quality database of chemical structures on global trade markets will be the foundation of ZeroPM's work on compiling high-quality data on uses, exposure, and the persistence, mobility and toxicity properties. This will enable effective prioritization and substance grouping strategies to help regulators, industry and the water treatment sector achive zero pollution of persistent and mobile substances.

Files

Wolf_SETAC_EU_23_V1.pdf

Files (2.9 MB)

Name Size Download all
md5:d21c834e885f765da98464d54764fa86
2.9 MB Preview Download

Additional details

Funding

European Commission
ZeroPM - ZeroPM: Zero pollution of Persistent, Mobile substances 101036756