HUBRIS - protein-protein interaction (PPI) networks
Authors/Creators
Description
This repository contains code and files related to HUBRIS, a protein-protein interaction network compiled from 8 different public databases. It also includes data about the newly identified interactors of HHIP in two lung cell lines (IMR90 and 16HBE) and RNA-Seq data from the two cell lines. The associated manuscript is entitled:
HHIP protein interactions in lung cells provide insight into COPD pathogenesis
Authors: Dávid Deritei1,*, Hiroyuki Inuzuka2,*, Peter J. Castaldi1, Jeong Hyun Yun1, Zhonghui Xu1, Wardatul Jannat Anamika1, John M. Asara3, Feng Guo4, Xiaobo Zhou1, Kimberly Glass1,†, Wenyi Wei2,†, Edwin K. Silverman1,†,‡
1 Channing Division of Network Medicine, Brigham and Women’s Hospital, Harvard Medical School, Boston, MA 02215, USA
2 Department of Pathology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA 02215, USA
3 Division of Signal Transduction, Department of Medicine, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA 02215, USA
4 Jiangsu Key Laboratory of Immunity and Metabolism, Jiangsu International Laboratory of Immunity and Metabolism, Department of Pathogen Biology and Immunology, Xuzhou Medical University, Xuzhou, Jiangsu 221004, China
* co-first authors
† co-senior authors
‡ corresponding author; Channing Division of Network Medicine Department of Medicine,
181 Longwood Avenue, Boston, MA 02115-5804
Phone: 617-525-0856 / Fax: 617-731-1541
Email: ed.silverman@channing.harvard.edu
Abstract
Chronic obstructive pulmonary disease (COPD) is the third leading cause of death worldwide. The primary causes of COPD are environmental, including cigarette smoking; however, genetic susceptibility also contributes to COPD risk. Genome-Wide Association Studies (GWASes) have revealed more than 80 genetic loci associated with COPD, leading to the identification of multiple COPD GWAS genes. However, the biological relationships between the identified COPD susceptibility genes are largely unknown. Genes associated with a complex disease are often in close network proximity, i.e. their protein products often interact directly with each other and/or similar proteins. In this study, we use affinity purification mass spectrometry (AP-MS) to identify protein interactions with HHIP, a well-established COPD GWAS gene which is part of the sonic hedgehog pathway, in two disease-relevant lung cell lines (IMR90 and 16HBE). To better understand the network neighborhood of HHIP, its proximity to the protein products of other COPD GWAS genes, and its functional role in COPD pathogenesis, we create HUBRIS, a protein-protein interaction network compiled from 8 publicly available databases. We identified both common and cell type-specific protein-protein interactors of HHIP. We find that our newly identified interactions shorten the network distance between HHIP and the protein products of several COPD GWAS genes, including DSP, MFAP2, TET2, and FBLN5. These new shorter paths include proteins that are encoded by genes involved in extracellular matrix and tissue organization. We found and validated interactions to proteins that provide new insights into COPD pathobiology, including CAVIN1 (IMR90) and TP53 (16HBE). The newly discovered HHIP interactions with CAVIN1 and TP53 implicate HHIP in response to oxidative stress.
Additional information:
Manuscript available on BioRxiv: https://www.biorxiv.org/content/10.1101/2024.04.01.586839v1
Associated GitHub repository: https://github.com/deriteidavid/hubris
RNA-Seq data (IMR90 and 16HBE cells): GSE285360
AP-MS data: ftp://massive.ucsd.edu/v07/MSV000096594/
(For more details see the documentation on the GitHub page and the Methods section of our manuscript)
Description of the files:
- biorosetta_data/ - local files of the biorosetta gene id mapping tool for consistent remapping
- db_local_files/ - local saves of the 8 constituent databases of HUBRIS (date of download 01/24/2024)
- PPI_networks_for_analysis/ - different PPI networks derived from HUBRIS, new experimental edges and RNA-Seq data
- RNASeq_lists/ - two text files (for the two cell lines) containing lists of genes that are considered expressed based on the RNA-Seq data. We use these lists to filter HUBRIS into cell type-specific versions. The files are generated by RNASeq_based_filtering_hg38.py.
- RNA_Seq_hg38/ - RNA-Seq data from IMR90 and 16HBE cells (tpm). For more details see the Methods section of our manuscript.
- all_significant_links_HHIP.xlsx - table containing the list of newly identified, significant protein-protein interactors of HHIP in 16HBE and IMR90 cells, based on analysis with SAINTExpress, with and without the CRAPome database.
- create_HUBRIS.py - main script to merge the already downloaded PPI databases (with the option to download the newest versions), translate between gene IDs, and filter the network based on the representation of edges in the different databases. See the GitHub README or the Methods section of our manuscript for more details
- databases.xlsx - configuration file for processing the database files (read by create_HUBRIS.py). For the files included in the db_local_files/ no re-configuration is needed, however, if the newest version of the databases is downloaded it may be necessary to adapt the file manually.
- generate_PPIs_for_analysis.py - Generate 2x2x2x3 = 24 different PPI network variants generated by the parameters new_edges (True, False); cell_line (16HBE, IMR90, union); filter_crapome (True, False); filter_networks_based_on_expression (True, False). The script saves the different network variations into .gml files in PPI_networks_for_analysis/
- G_merged_raw.gpickle - gpickle file of the raw HUBRIS network (no filtering)
- G_hubris.gpickle - gpickle file of filtered HUBRIS (every edge must be represented in at least 2 of the 8 databases. This parameter can be changed in create_HUBRIS.py)
- G_hubris.gml - gml file of the filtered HUBRIS
- G_hubris.txt - edgelist of the filtered HUBRIS
- G_hubris_lcc.gml - gml file of the largest connected component of HUBRIS (filtered)
- G_hubris_lcc.txt - edgelist of the largest connected component of HUBRIS (filtered)
- hgnc_mapping.tsv - gene id mapping from https://www.genenames.org/download/statistics-and-files/ (downlaoded 05/10/2022).
- hubris_functions.py - helper functions for create_HUBRIS.py
- RNASeq_based_filtering.py - script generating the lists of expressed genes (RNASeq_lists/) based on RNA-Seq data (rnaseq_PPI_GWAS/)
Methods
From the Methods section of our manuscript
Creating HUBRIS
For the creation of the HUBRIS network, we merged 8 different protein-protein interaction databases from the following sources:
-
HuRI: http://www.interactome-atlas.org/data/HuRI.tsv
-
HIPPIE: http://cbdm-01.zdv.uni-mainz.de/~mschaefer/hippie/hippie_current.txt
-
Interactome3D: https://interactome3d.irbbarcelona.org/downloadset.php?queryid=human&release=current&path=complete
-
StringDB: https://stringdb-static.org/download/protein.links.detailed.v11.5/9606.protein.links.detailed.v11.5.txt.gz
For all databases the date of download is 01/24/2024. All analyses are based on the state of the databases on this date.
All networks were used in their linked form, except for StringDB, where we performed a preliminary filtering for the experimental score being larger than zero, as we are looking for concrete physical interactions not functional associations. In addition, for databases where non-human species were included, we filtered to include only interactions found in humans.
Merging the networks: The first step of the merging required using a common gene/protein identifier as many networks used different identifiers, such as Entrez Gene Id (HumanNet, HIPPIE, BioGrid, NCBI), Uniprot Id (Interactome3D, Reactome), Ensembl Gene Id (HuRI), and Ensembl Peptide Id (StringDB). Since Entrez Gene Id was the most common, we decided to use it as the common id. We then used the BioRosetta tool (https://github.com/reemagit/biorosetta) to re-label the networks not using Entrez Gene Id.
In cases where there was no valid or unique translation, we kept the node in the network with the original identifier (as it could still be structurally significant) and appended the shortened form of the source id type to the node id with an underscore. For example, the Ensembl Peptide Id ENSP00000266991 did not have a valid translation to Entrez Gene Id in BioRosetta, and therefore the node received the new label ENSP00000266991_ensp. This helps us to manually look up relevant nodes when interpreting results. In the raw, unfiltered version of HUBRIS roughly one-third of the nodes do not have a valid translation in our protocol; however, in the final, filtered version of HUBRIS (N=20,103) there are only 244 such nodes.
After the relabeling of the individual networks, we merged them into a single network that for every edge aggregated the names of the different databases that contained them. We call the number of databases that contain a single specific edge the database support count and it is a property of every edge.
Properties of the merged HUBRIS: The raw merged network consists of N=32,384 nodes and E=3,233,850 edges. For the edges, the database support distribution is the following: 1: 2,337,191, 4: 304,861, 5: 242,908, 3: 167,029, 2: 158,215, 6: 20,493, 7: 2,946, 8: 207.
Filtering the HUBRIS network: The above statistics show that most of the edges (>72%) are from a single database source and these edges are overwhelmingly from StringDB (>89%).
Having manually examined parts of the network that are relevant for our study (especially the neighborhood of HHIP and the other COPD GWAS gene products), we decided to trim our network and only keep edges that have support from at least 2 databases. The resulting network has the following statistics: N=32,384 nodes and E=896,659 edges, where the giant component is composed of N=20,103 nodes and E= 896,648 edges.
In all analyses described in the paper and the accompanying documents when we refer to HUBRIS, we refer to the network defined above.
Cell type-specific filtering of HUBRIS based on RNA-Seq data
To better understand the cell type-specific network processes between HHIP and other proteins, it is necessary to consider that not all proteins present in the PPI network are present in a specific cell type. To account for this effect, we used RNA-Seq data from the two cell types we used in this study: IMR90 and 16HBE (GSE285360).
RNA extraction and library preparation: RNA samples from cell lines were extracted by RNAeasy kit
(Qiagen) and measured for quality and quantity by Qbit (Invitrogen). Samples with a RIN over 8 were
submitted for sequencing for gene expression.
Alignment and quantification: RNA-seq data were processed through the nf-core RNA sequencing
pipeline (https://nf-co.re/rnaseq/ ) version 3.14.0 using Nextflow 24.04.4 1 .
Stranded reads were aligned to the human GRCh38 reference genome using the STAR aligner 2.7.10a 1 in
one-pass mode and the GENCODE-45 GTF with default parameters. Aligned and unaligned reads were
output to a single BAM file. Junctional reads were assigned according to default parameters. For gene and
isoform level quantification, isoform counts were estimated with Salmon version 1.10.1 1 and default
parameters. Gene and isoform level counts were generated with the tximport Bioconductor package
(bioconductor-tximeta 1.12.0, r-base 4.1.3) 1 , and Salmon isoform TPM counts were aggregated to the gene level using tximport with the countsFromAbundance flag set to “No”and using the
assays(*)$abundance table.
For both cell lines, we used the same criteria: we excluded genes with an absolute count of 0 since the normalization can skew these values into nonzero values, and then used a cutoff of 0.25 on the log2 of the FPKM values (see Supplementary Figure S12).
The value of 0.50 was determined to strike an appropriate balance between removing nodes with low expression while still allowing us to retain all proteins measured in the AP-MS experiments (including HHIP).
Finally, we kept genes that were expressed and removed every other node from HUBRIS, and by keeping the remaining largest connected component (LCC) thus created the cell type-specific network.
Author Contributions
D.D. – conceptualization, network generation and analysis, data curation (public databases), software (HUBRIS, network methods, GSEA), writing (original draft), visualization (Figures 1-4, 7, S2-S9, S11, S12)
H.I. – experimental investigation (AP-MS, co-IP), validation, writing (review and editing), visualization (Figures 5, S10)
P.C. – formal analysis (AP-MS data), writing (review and editing), supervision
J.H.Y. – experimental investigation (scRNA-Seq), validation, formal analysis, writing (review and editing), visualization (Figure 6)
Z.X. – formal analysis (AP-MS data), validation, visualization (Figure S1), writing (review and editing
W.J.A. – experimental investigation (siRNA KO), validation, formal analysis, writing (review and editing)
J.M.A – resources (sample analyses), MS data acquisition and analysis, writing (review and editing)
F.G. – experimental investigation (RNA-Seq of cell lines), writing (review and editing)
X. Z. – supervision, writing (review and editing)
K.G. – supervision, writing (review and editing)
W.W. – conceptualization, funding acquisition, supervision, writing (review and editing)
E.K.S. – conceptualization, funding acquisition, project administration, formal analysis (siRNA KO data), supervision, writing (original draft)
Acknowledgments
The scRNA-Seq work was funded by K08HL146972 and HMS Shore Faculty Development Award (J.Y.).
The mass spectrometry work was partially funded by NIH grants 5P01CA120964 (J.M.A.) and 5P30CA006516 (J.M.A.)
K.G. is supported by R01 HL155749.
EKS is supported by R01 HL152728, R01 HL147148, R01 HL133135, and P01 HL114501.
Conflict of Interest Statement
In the past three years, Edwin K. Silverman received grant support from Bayer and Northpond Laboratories.
Jeong Yun received consulting fees from Bridge Biotherapeutics.
Other authors declare no conflict of interest.
Files
hubris_zenodo.zip
Additional details
Software
- Repository URL
- https://github.com/deriteidavid/hubris
- Programming language
- Python