There is a newer version of the record available.

Published August 20, 2024 | Version v1

PrimeKGQA, the dataset from paper: Bridging the Gap: Generating a Comprehensive Biomedical Knowledge Graph Question Answering Dataset

  • 1. Universität Hamburg

Description

Despite the plethora of resources such as large-scale corpora and manually curated Knowledge Graphs (KGs), the ability to perform reasoning with natural language inputs over biomedical graphs remains challenging due to insufficient training data. We propose a novel method for automatically constructing a Biomedical Knowledge Graph Question Answering (BioKGQA) dataset sourced from PrimeKG, the largest precision medicine-oriented KG. In total,
we create 83999 question-answer pairs along with their respective SPARQL queries. Our approach generates a diverse array of contextually relevant questions covering a wide spectrum of biomedical concepts and levels of complexity. We evaluate our method based on automatic metrics alongside manual annotations. We establish novel standards tailored for KGQA systems to highlight the linguistic correctness and semantical faithfulness of the generated questions based on extracted KG facts. The compiled dataset – PrimeKGQA – serves as a valuable benchmarking resource for advancing knowledge-driven biomedical research and evaluating KGQA system.

Files

final_test.json

Files (330.5 MB)

Name Size
md5:39301bc15ca180291b772a4ba22fe114
134.4 MB Preview Download
md5:e7816de112f8ebbcb6f5367baa70bf81
110.9 MB Preview Download
md5:e32de58b3c647162ccb289babf238c59
85.2 MB Preview Download

Additional details

Dates

Accepted
2024-08-20

Software

Repository URL
https://github.com/xixi019/primeKGQG
Programming language
Python