Proteo-genomics-guided interpretation of somatic mutations in cancer genomes
Authors/Creators
Description
Cancer genomes harbor hundreds of mutated genes, but their exact roles in tumorigenesis often remain incompletely understood. Here, by integrating protein structures with mutation data from 139,818 patients, we generate detailed, functional maps of mutations for 173 cancer genes in 18+ tumor types. In most cancer genes, oncogenic mutations accumulate in “significantly mutated regions” aligned with key functions, including DNA-binding domains in transcription factors (e.g., FOXA1), ligand-binding sites in receptors (e.g., ERBB2), catalytic domains in chromatin remodelers (e.g., EP300), substrate-recognition sites in ubiquitin-proteasome regulators (e.g., SPOP), and degradation signals in cell-cycle proteins (e.g., CCND1). In other cancer genes, mutations align more closely with constraints in three-dimensional protein structures, such as targeting the interior core (e.g., TP53) or a specific secondary structure (e.g., PTEN). Broadly, our study provides a proteo-genomic framework to characterize the effects of mutations in tumorigenesis and integrate them into workflows for clinical interpretation in precision medicine.
In this repository, we provide the source code for identifying significantly mutated regions in cancer genes (TrainHMM.zip), along with a precompiled version, a detailed user manual, and a dataset for testing its functionality and reproducing key results of our study. We also provide details on the coordinates of the significantly mutated regions, their mutation rates, and other biological variables in our proteo-genomics model (SupplDataset.zip).
Files
SupplDataset.zip
Files
(104.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:15037ce9bc59f9579dc4b48c954a53dc
|
52.0 MB | Preview Download |
|
md5:58462ca90b5ba064a7e73ee934bafe4a
|
52.1 MB | Preview Download |