GGPONC 2.0 (Major Release incl. Gold Standard Annotations)
Creators
Description
About this Release
Version 2.0 (major release), contains gold-standard annotations described in the GGPONC 2.0 paper.
Once you have downloaded the data, the most convenient way to access them is through our Hugging Face BigBIO dataloader.
⚠️ More recent versions of GGPONC (with new / updated guidelines, but without gold-standard annotations) are published at:
https://zenodo.org/records/12520623
Project Description
The GGPONC project aims to provide a freely distributable corpus of German medical text for NLP researchers. Clinical guidelines are particularly suitable to create such corpora, as they contain no protected health information (PHI), which distinguishes them from other kinds of medical text.
The second major release of the corpus (GGPONC 2.0, 2024/03) consists of 30 German oncology guidelines with 1.87 million tokens. It has been completely manually annotated on the entity level by 7 medical students using the INCEpTION platform over a time frame of 6 months in more than 1200 hours of work. This makes GGPONC 2.0 the largest annotated, freely distributable corpus of German medical text at the moment.
Annotated entities are Findings (Diagnosis / Pathology, Other Finding), Substances (Clinical Drug, Nutrients / Body Substances, External Substances) and Procedures (Therapeutic, Diagnostic), as well as Specifications for these entities. In total, annotators have created more than 200000 entity annotations. In addition, fragment relationships have been annotated to explicitly indicate elliptical coordinated noun phrases, a common phenomenon in German text.
Files
Additional details
Related works
- Is original form of
- Dataset: 10.5281/zenodo.12520368 (DOI)
- Is published in
- Conference paper: https://aclanthology.org/2022.lrec-1.389 (URL)
- References
- Software: 10.5281/zenodo.6473122 (DOI)
Dates
- Available
-
2023-03-24
Software
- Repository URL
- https://github.com/hpi-dhc/ggponc_annotation