Published June 15, 2020 | Version v1

Semantic text mining in early drug discovery for type 2 diabetes

Description

BACKGROUND: Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g.\ recently discovered modes of action) and identifies potential novel drug target, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs.
    
METHODS: We extracted over 7 million n-grams from PubMed and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant `semantic concepts'. To score papers, these concepts were weighted depending on co-mentioning with core T2D proteins. A protein's current T2D relevance was determined by combining the scores of the papers mentioning it in the preceeding five years. The significance of a jump in a protein's rank was assessed by comparing it to previously observed jumps.
    
RESULTS: We show that T2D relevant papers, also those not mentioning T2D explicitly, got assigned high scores by mentioning semantic concepts often used in connection with T2D, as shown by the enrichment of well known T2D proteins among the top scoring proteins. Our `high jumpers' identified important past developments in the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.

Files

README.md

Files (3.7 GB)

Name Size
md5:3f7be97fcbd10fedaf8b3991cb6b124b
384.7 kB Download
md5:6fcc2cbdda0632d2f1f57660be744eaa
1.3 MB Download
md5:c872bfd6e6bf1409a226e68d84c4680e
13.2 MB Download
md5:7cef81e5de0f6388c8e2ee0be390f8b9
2.4 GB Download
md5:ebe881d7c074fa27489ec689fa1635ee
202.7 MB Download
md5:47e9349ea285982101a94e7f1383509a
2.7 MB Download
md5:4d6991db37e6ee3a0de22aeec39316a3
24.8 MB Download
md5:07c3304877c41d0990856e25385e310a
320.2 MB Download
md5:9abd25b7c9a55b9289108441d8882c4b
698.1 MB Download
md5:29ec2e42c2bbfaf46b993aeea754b1ec
4.2 kB Preview Download

Additional details

Related works

Is supplement to
Journal article: 10.1371/journal.pone.0233956 (DOI)