Published August 17, 2021 | Version v1

Protein Identification by Nanopore Peptide Profiling

  • 1. Groningen Biomolecular Sciences and Biotechnology Institute, University of Groningen
  • 2. Stratingh Institute for Chemistry, University of Groningen

Description

This dataset belongs to “Protein Identification by Nanopore Peptide Profiling” and describes the raw data and analysis of tryptic digested peptides translocating through a mutant Fragaceatoxin C nanopore. A jupyter notebook describing the analysis and structure is added to this dataset.

 

Data description:

Protein Identification by Nanopore Peptide Profiling.ipynb

               Jupyter notebook contained data analysis of data contained in data_0.zip and data_1.zip (Python 3.7)

python_scripts.zip

               Supplementary  scripts belonging to “Protein Identification by Nanopore Peptide Profiling.ipynb”. See explanation of custom classes in the jupyter notebook.

data_0.zip - Folder containing raw electrophysiology data and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:

               Alpha casein:                                    Tryptic digest of alpha casein

               Beta casein:                                      Tryptic digest of beta casein

               BSA:                                                  Tryptic digest of bovine serum albumin

               Control:                                              Tryptic digest of water (no protein, control measurement)

               Cytochrome c:                                   Tryptic digest of cytochrome c

               DHFR_His6:                                      Tryptic digest of dihydropholate reductase (His6 tagged)

               EFP:                                                  Tryptic digest of elongation factor P

               HMW1Act:                                         Tryptic digest of high molecular weight adhesin protein

data_1.zip - Folder containing raw electrophysiology data, comma-separated MS peptide masses, and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:

               Lysozyme:                                         Tryptic digest of lysozyme              

               PAN:                                                   Tryptic digest of proteasome-activating nucleotidase

               TbpA_Y27A:                                       Tryptic digest of periplasmic binding protein

               Trypsin:                                               Tryptic digest of bovine trypsin

               Mass_spec:                                        csv files containing measured ESI-MS peptides

               Lysozyme synthetic peptides:            Synthetic peptides:

                                                                              Lys1:       TPGSR

                                                                              Lys2alk:  C(+57.02)ELAAAMK

                                                                              Lys3:       HGLDNYR

                                                                              Lys4alk:  WWC(+57.02)NDGR

                                                                              Lys5:       GTDVQAWIR

                                                                              Lys6alk:  GYSLGNWVC(+57.02)AAK

                                                                              Lys7:       FESNFNTQATNR

The structure of the data files is registered data_1.zip in 'index.csv' (digested proteins) and  'index_peptides.csv' (synthetic peptides) contained in the data folder. In this file, we describe the protein that was measured as well as the folder location and the expected baseline / standard deviation.

Structure of ./data/index.csv

Protein (string) | Folder (string) | Baseline (pA) (float) | Baseline Error (pA) (float)

 

In each Folder, there is another 'index.csv', explaining which files are with protein and which are without (blank).

Structure of ./data/[protein]/[repeat]/index.csv

blank (boolean) | fname (string)

 

Each folder in data_0.zip and data_1.zip contains a folder for each measure protein, which contains a folder for each repeat. The repeats contain raw axon binary files (.abf), each file contains measurement conditions as follows:

    [Date of measurement]_[Pore type]_[Buffer conditions]_[added analyte(s)]_[operator initials]

    e.g: 20200312_1M_KCl_50mM_Citricacid_50mM_BTP_pH_38_FraC_G13F_neg70mV_20ul_CytC_TrypsinGold_FL_0000

    Measured on 12-03-2020, in 1M KCl buffered with Citricacid (50 mM) adjusted using bis-tris-propane to pH 3.8, using Fragaceatoxin C mutant G13F at a negatively applied potential of 70 mV. 20 µL cytochrome c was added to the cis compartment.

The total volume of the container used for all electrophysiology experiments was 400 µL, all samples were prepared at a 1 g/L concentration. A prefix “perf” before analyte description indicates that the chamber was flushed with approximately 2 mL fresh buffer prior to analysis. The buffer condition "BTP" means bis-tris-propane, which is used to titrate to the exact pH of 3.8.

Each analysed folder contains results.pkl file, containing the analysis result as provided by “Protein Identification by Nanopore Peptide Profiling.ipynb” - see the jupyter notebook

Each analysed folder contains results_analysis.xlsx, which contains sheets with excluded currents, standard deviations, dwell time and beta value for the pore without analyte added “Blank” and results from the analyte added in “Results”. Parameters used for fitting are contained in “Parameters”. The “Histograms” tab shows the raw data of the excluded current spectra.

 

mass_spec_peaks.zip – Folder containing mass spectrometry files as analysed by PEAKS Studio

The folder contains an subfolder for each protein measured using electrospray ionisation mass spectrometry (ESI-MS).

               acasein:                alpha casein protein

               b_casein:              beta casein protein

               BSA:                      bovine serum albumin

               CytC:                     cytochrome C digested

               DHFR:                   dihydropholate reductase

               HMW1_Act:           high molecular weight adhesin protein

               PAN:                      proteasome-activating nucleotidase

               ThBP:                    periplasmic thiamine binding protein

               Trypsin:                 bovine trypsin

Notes

This work is part of the research program of the Foundation for Fundamental Research on Matter (FOM), which is part of the Netherlands Organization for Scientific Research (NWO) under grant number 16SMPS05. In addition, M.T.C.W. was supported by the Dutch Organization for Scientific Research (VENI 722.016.006).

Files

data_0.zip

Files (19.7 GB)

Name Size
md5:640c0c7d03e14c116629eab3ffb269bb
12.0 GB Preview Download
md5:a1a165e99d51969402bca9b4a31f16e2
6.0 GB Preview Download
md5:a62f2eb195eabdb90a0a97ea20af1465
1.7 GB Preview Download
md5:8047c2ea9a5419c513e923b1f584f156
1.8 MB Preview Download
md5:04bd5627f11b3509dbeb28d4fd884041
94.4 kB Preview Download

Additional details

Funding

European Commission
DeE-Nano - Design and Engineering Next-Generation Nanopore Devices for Bioplymer Analysis 726151
European Commission
ROSALIND FRANKLIN - Rosalind Franklin Fellowship Cofund Programme 600211