Published December 1, 2020 | Version 0.6.7

Enhancing Open Modification Searches via a Combined Approach Facilitated by Ursgal

Description

The identification of peptide sequences and their post-translational modifications (PTMs) is a crucial step in the analysis of bottom-up proteomics data. The recent development of open modification search (OMS) engines allows virtually all PTMs to be searched for. This not only increases the number of spectra that can be matched to peptides but also greatly advances the understanding of biological roles of PTMs through the identification, and thereby facilitated quantification, of peptidoforms (peptide sequences and their potential PTMs). While the benefits of combining results from multiple protein database search engines has been established previously, similar approaches for OMS results are missing so far. Here, we compare and combine results from three different OMS engines, demonstrating an increase in peptide spectrum matches of 8-18%. The unification of search results furthermore allows for the combined downstream processing of search results, including the mapping to potential PTMs. Finally, we test for the ability of OMS engines to identify glycosylated peptides. The implementation of these engines in the Python framework Ursgal facilitates the straightforward application of OMS with unified parameters and results files, thereby enabling yet unmatched high-throughput, large-scale data analysis.

This dataset includes all relevant results files, databases, and scripts that correspond to the accompanying journal article. Specifically, the following files are deposited:

  • Homo_sapiens_PXD004452_results.zip: result files from OMS and CS for the dataset PXD004452
  • Homo_sapiens_PXD013715_results.zip: result files from OMS and CS for the dataset PXD013715
  • Haloferax_volcanii_PXD021874_results.zip: result files from OMS and CS for the dataset PXD021874
  • Escherichia_coli_PXD000498_results.zip: result files from OMS and CS for the dataset PXD000498
  • databases.zip: target-decoy databases for Homo sapiens, Escherichia coli and Haloferax volcanii as well as a glycan database for Homo sapiens
  • scripts.zip: example scripts for all relevant steps of the analysis
  • mzml_files.zip: mzML files for all included datasets
  • ursgal.zip: current version of Ursgal (0.6.7) that has been used to generate the results (for most recent versions see https://github.com/ursgal/ursgal)

Files

databases.zip

Files (37.4 GB)

Name Size
md5:3da5697d75c4533cb5a8e93e370a5bb1
33.3 MB Preview Download
md5:781fbbfdcd58ffc418641e405a51f31f
1.9 GB Preview Download
md5:f5d593397199d99160204dea1feac854
625.8 MB Preview Download
md5:edb8991b7a47a36cd5370a379c931006
643.5 MB Preview Download
md5:f871452b36d22216c8f96a92980f6ef9
284.7 MB Preview Download
md5:c29f91c6eb7fc034f5d82bd2301f59c7
11.1 GB Preview Download
md5:8e2af347fb3d600e4849b54a37779af7
10.1 GB Preview Download
md5:ecb676d78cd92a79ba1d1023d52dfcfe
8.8 GB Preview Download
md5:20dbc650a50ef6d4a4eea4bd99f7c2ce
3.9 GB Preview Download
md5:812d653614c4b9e1ec3c617d287c841b
13.7 kB Preview Download
md5:0c9eb14864cd2d97c1d3e9e7da78fdf8
10.1 MB Preview Download