DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases

Luis Gasco; Eulàlia Farré; Miranda-Escalada, Antonio; Salvador Lima; Martin Krallinger

doi:10.5281/zenodo.6408477

Published April 2, 2022 | Version 1.0.0

Dataset Open

DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases

1. Barcelona Supercomputing Center

DisTEMIST corpus - training set

Introduction

The DisTEMIST corpus is a collection of 1000 clinical cases with disease annotations linked with Snomed-CT concepts. All documents are released in the context of the BioASQ DisTEMIST track for CLEF 2022. For more information about the track and its schedule, please visit the website.

File structure:

The DisTEMIST corpus has been randomly divided into a training set, containing 750 clinical cases, and a test set, consisting of 250 additional cases. Participants must train their systems using the train set and submit predictions for the test set, on which they will be evaluated. The file structure of the corpus is as follows:

train_set:
- text_files: Folder with plain text files of the clinical cases
- subtrack1_entities: It contains annotations in a tab-separated file (TSV) with the following columns:
  - filename: document name
  - mark: identifier mention id
  - label: mentions type (ENFERMEDAD)
  - off0: starting position of the mention in the document
  - off1: ending position of the mention in the document
  - span: text span
- subtrack2_linking: It will be released on 15 April. It will contain the disease mentions normalized to Snomed-CT concepts.

test_set: The 250 clinical cases that will be used to evaluate the systems will be published in accordance with the schedule of the task.

Resources

Notes

Funded by the Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).

Files

distemist_corpus.zip

Files (1.2 MB)

Name	Size	Download all
distemist_corpus.zip md5:84dcea024411d3074d51c5b564725804	1.2 MB	Preview Download

	All versions	This version
Views	7,332	216
Downloads	1,708	22
Data volume	110.1 GB	30.0 MB

DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases

Authors/Creators

Description

Notes

Files

distemist_corpus.zip

Files (1.2 MB)