Published July 26, 2021 | Version v1

Two document-concept representations of the biomedical literature

Authors/Creators

  • 1. Trinity College Dublin

Description

These two datasets represent the biomedical literature (Medline abstracts and PubMedCentral articles) in the "document-concept matrix" format produced by TDC Tools.  These datasets can be used in downstream IR applications such as Literature-Based Discovery.

Each of the two datasets corresponds to a specific data extraction method, see details here and in the paper linked below.

Important: the raw data from which this data is derived was downloaded from Medline, PubMedCentral and PubTatorCentral, provided courtesy of the U.S. National Library of Medicine (NLM). The data was extracted in January 2021 and do not reflect the most current/accurate data available from NLM. See the github repository above in order to generate similar datasets from up to date data.

Files

Files (11.3 GB)

Name Size
md5:dfc6ff81c51e3b95122cc3fde78507f2
8.2 GB Download
md5:998af2893396cd77158c8059350026d9
3.1 GB Download