Published February 13, 2020 | Version v1

1000 Disease Ontology terms and their Wikidata mappings to 17 mostly Indian languages and English

Authors/Creators

  • 1. School of Data Science, University of Virginia

Description

This dataset contains the result of a SPARQL query run on the Wikidata Query Service on 13 February 2020 around 22:25 UTC. They are archived here as a means to determine progress with the coverage of disease-related terms in languages other than English, particularly in languages of India.

The SPARQL query

  • was for
    • Wikidata items for concepts that have a Disease Ontology ID (P699)

      • sorted by number of sitelinks

      • optionally with their Wikidata label in English

      • optionally with their Wikipedia article title in English

      • optionally with their Wikidata label in Hindi, Bangla and Swahili

      • optionally with their Wikidata label in Marathi, Telugu, Eastern Punjabi, Western Punjabi, Gujarathi, Maithili, Kannada, Odia, Bhojpuri, Tamil, Nepali, Urdu, Malayalam, Esperanto

  • is contained in the file SPARQL.txt,

whereas the results are available in several formats, as provided by the Wikidata Query Service:

  • query.csv
  • query.tsv
    query.html
  • query.json.txt (Zenodo produced an error upon trying to upload the file as query.json, so I renamed it, which worked fine).

A simplified version of the SPARQL query can also be fed into the TABernacle tool that represents the live data in a way that facilitates editing the missing pieces.

Files

query.csv

Files (1.4 MB)

Name Size Download all
md5:dc821b46802044fd2015644cfda6938a
199.8 kB Preview Download
md5:34f49cd5cf0fb2ded7ccc9be2df8f2ec
377.0 kB Download
md5:383bd81c0eba376293dd8520501cea32
606.8 kB Preview Download
md5:c0f42653d54a23a5f1a16603fcfe2510
228.0 kB Download
md5:c994bc93b57b3d7a02e7fef942b164df
2.5 kB Preview Download