Published June 12, 2025 | Version v1

Pangolin precomputed scores

  • 1. EDMO icon Technical University of Munich

Description

This dataset contains Pangolin precomputed scores for all SNVs in protein-coding genes (hg38 genome version) computed with default parameters: window size 50 nt, scores are masked based on GENCODE splice site annotations (see below for more information, and see the original paper Zeng & Li, 2022).

Pangolin is a deep learning model that predicts the effect of a variant on the splice site usage. It computes a gain and a loss score for every position within a user-defined window around the variant that represents the increase and decrease in the usage of a potential splice site at the respective positions. Pangolin outputs the maximum gain and the maximum loss scores within the window together with the corresponding positions. It also provides an option to mask scores when a genome annotation is provided to the model, which sets those scores to zero if Pangolin predicts activation for annotated splice sites and deactivation for unannotated splice sites.

The directory contains per-gene TSV files with the following columns:

  • chrom: Chromosome
  • pos: Genomic position
  • ref: Reference allele
  • alt: Alternative allele
  • gain_score: Pangolin gain score
  • gain_pos: relative position of the gain score
  • loss_score: Pangolin loss score
  • loss_pos: relative position of the loss score

Files

Pangolin_hg38_snvs_masked.zip

Files (13.0 GB)

Name Size
md5:679ef0b50e511b6102b4b88fbf811108
13.0 GB Preview Download