SympTEMIST Corpus: Gold Standard annotations for clinical symptoms, signs and findings information extraction
Authors/Creators
- 1. Barcelona Supercomputing Center
Description
SympTEMIST stands for Symptoms TExt MIning Shared Task. It is a shared task and set of resources focused on the detection, normalization and indexing of symptoms, signs and findings in medical documents in Spanish. SympTEMIST is complementary to the DisTEMIST corpus (https://temu.bsc.es/distemist) and MedProcNER/ProcTEMIST (https://temu.bsc.es/medprocner) as they all use the same document collection.
[**UPDATE 2023-09-14**] This repository includes the Train Set for the three subtasks, a SNOMED symptoms, signs and findings gazetteer and the Multilingual Silver Standard in 9 languages. Please read the README file attached for more information on folder structure and file format.
SympTEMIST was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioCREATIVE 2023. For more information on the corpus, annotation scheme and task in general, please visit: https://temu.bsc.es/symptemist.
Related Links:
- SympTEMIST website: https://temu.bsc.es/symptemist
- SympTEMIST annotation guidelines: https://doi.org/10.5281/zenodo.8246439
License
This work is licensed under a Creative Commons Attribution 4.0 International License.
Contact
If you have any questions or suggestions, please contact us at:
- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)
- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)
Additional resources
If you are interested in SympTEMIST, you might want to check out these corpora and resources that use the same text documents:
- DisTEMIST (diseases)
- MedProcNER/ProcTEMIST (clinical procedures)
- PharmaCoNER (chemicals and proteins)
Files
symptemist-train_all_subtasks+gazetteer+multilingual_230914.zip
Files
(22.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:56e612db54113747abe21bab3f762e61
|
22.2 MB | Preview Download |