MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction
Authors/Creators
- 1. Barcelona Supercomputing Center
- 2. Dublin University
Description
MEDDOPLACE stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.
This repository includes the Training Set of the corpus in three different formats (brat, .json, .tsv), as well as relevant gazetteers. For more information, please check the attached README file.
** Update April 19th 2023: A new version has been uploaded, with the following changes:
- The format for subtask 3 has been changed, with the four columns for each class type being replaced with a single column;
- A few inconsistencies in the SNOMED normalization have been fixed;
- A gazetteer has been added for the SNOMED normalization
MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: https://temu.bsc.es/meddoplace.
Related Links:
- MEDDOPLACE website: https://temu.bsc.es/meddoplace
- Annotation Guidelines: https://doi.org/10.5281/zenodo.7775234
License
This work is licensed under a Creative Commons Attribution 4.0 International License.
Contact
If you have any questions or suggestions, please contact us at:
- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)
- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)
Files
meddoplace_train_set_230419.zip
Files
(5.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:cb0f58cef217fc6a5bdc45b22418d57e
|
5.2 MB | Preview Download |