Published September 23, 2025 | Version v1

Donor Specificity of UDP-dependent Glycosyltransferases Across Diverse Acceptor Molecules

  • 1. The Novo Nordisk Center for Biosustainability, Technical University of Denmark, 2800 Kgs. Lyngby, Denmark
  • 2. Loschmidt Laboratories, Department of Experimental Biology and RECETOX, Faculty of Science, Masaryk University, Czech Republic
  • 3. International Clinical Research Center, St. Anne's University Hospital Brno, Czech Republic

Description

This dataset describes binary donor specificity of GT1-family glycosyltransferases across a range of acceptor molecules. It was compiled through manual curation of published data from the literature, as specified in the References. The dataset comprises 218 unique GT1 enzyme sequences, 10 distinct donor sugars, and 108 acceptor compounds, yielding 1189 curated enzyme-donor-acceptor activity records.

When different literature sources reported conflicting outcomes for the same enzymatic reaction, the reaction was classified as active if at least one source reported a positive result, regardless of other conflicting reports. This approach ensures the dataset reflects the most permissive interpretation of enzyme activity.

Additionally, if the measured activity for a given donor-acceptor pair was less than 5% of the activity observed with the enzyme’s most active donor-acceptor pair, the reaction was considered inactive.

The dataset is provided in two complementary formats:

.xlsx (for visual inspection)

A human-readable spreadsheet where each row represents a unique enzyme-acceptor pair, and activity against multiple donors is given across separate columns (Glc, Gal, Glu, Rha, GlcNAc, etc.). Binary activity values (1/0) are colour-coded (green = active, red = inactive) to aid quick interpretation.

Columns:

  • Plant: The source organism (species) from which the GT1 enzyme was derived.
  • UGT: The name or identifier of the glycosyltransferase (UGT) enzyme, often following gene naming conventions.
  • Uniprot/Genbank: A unique accession ID referencing the enzyme’s sequence in UniProt or GenBank databases.
  • Glc, Gal, Glu, Rha, GlcNAc, Xyl, Ara, GalNAc, GalUA, Man: Binary activity values (0 = inactive, 1 = active) indicating whether the enzyme accepted each of these donor sugars when combined with the listed acceptor.
    • Glc: Glucose
    • Gal: Galactose
    • Glu: Glucuronic acid
    • Rha: Rhamnose
    • GlcNAc: N-acetylglucosamine
    • Xyl: Xylose
    • Ara: Arabinose
    • GalNAc: N-acetylgalactosamine
    • GalUA: Galacturonic acid
    • Man: Mannose
  • Substrate: The name of the acceptor molecule used in the activity assay.
  • SMILES: The SMILES (Simplified Molecular Input Line Entry System) string representing the chemical structure of the acceptor.
  • Protein sample: A description of how the enzyme sample was prepared (e.g., purified protein).
  • Analysis: The method used to detect activity (e.g., HPLC).
  • doi: Digital Object Identifier (DOI) of the original publication from which the data was extracted, if applicable.
  • Seq: The full amino acid sequence of the enzyme.

Note: This version does not include donor SMILES; only acceptor structures are provided.

.csv (for computational use)

A long-format version structured for machine learning and programmatic analysis, where each row corresponds to a single enzyme-donor-acceptor reaction pairing, thereby facilitating dataset reshaping, statistical modelling, and high-throughput prediction tasks.

Columns:

  • enzyme: The name or label of the GT1 enzyme involved in the reaction.
  • donor: The abbreviated name of the donor sugar used in the glycosylation reaction.
  • acceptor: The name of the acceptor compound to which the donor sugar is transferred.
  • activity: Binary indicator of enzyme activity for the enzyme-donor-acceptor pairing (1 = active, 0 = inactive).
  • enz_ID: A unique accession ID referencing the enzyme’s sequence in UniProt or GenBank databases.
  • enz_sequence: The full amino acid sequence of the enzyme.
  • donor_id: The Compound Identifier (CID) from PubChem of the donor compound.
  • donor_SMILES: The SMILES string representing the chemical structure of the donor sugar.
  • acceptor_ID: The CID from PubChem of the acceptor compound.
  • acceptor_SMILES: The SMILES string representing the chemical structure of the acceptor molecule.
  • doi: The Digital Object Identifier (DOI) of the publication from which the data was obtained.

This format provides full molecular structures for both donors and acceptors via canonical SMILES strings, along with unique identifiers for each molecule and enzyme. It should be noted that some acceptors did not have a reported CID, leaving the "acceptor_ID" blank for these molecules.

The dataset is intended to support research in glycosyltransferase specificity, donor promiscuity profiling, enzyme engineering, and data-driven modelling of enzyme-donor-substrate interactions. For a more in-depth review of the donor specificity of GT1 enzymes, please refer to:
Gharabli H, Welner DH. The sugar donor specificity of plant family 1 glycosyltransferases. Front Bioeng Biotechnol. 2024 May 2;12:1396268. doi: 10.3389/fbioe.2024.1396268. PMID: 38756413; PMCID: PMC11096472.

Files

UGT_Donor_Dataset.csv

Files (860.9 kB)

Name Size Download all
md5:8021bcca9622d508ada5f367bad83d89
740.3 kB Preview Download
md5:2eb75c181f6cfc17fb707620cf56ded8
120.6 kB Download

Additional details

Related works

Is supplement to
Publication: 10.3389/fbioe.2024.1396268. (DOI)

Funding

Novo Nordisk Foundation
NNF20CC0035580
RECETOX
LM2023069
Czech National Infrastructure for Biological Data
LM2023055

References