Published July 31, 2018 | Version v1

Manual Dependency Annotation of Three German Text Extracts from the Project hermA (Gold Standard Data)

  • 1. Universität Hamburg

Description

This dataset was created in the digital humanities project hermA (www.herma.uni-hamburg.de) and comprises annotated extracts of the following three texts:

  • Modern literature (Lit2009): novel Corpus Delicti: Ein Prozess by German author Juli Zeh, published in Frankfurt/Main in 2009.
  • Non-contemporary literature (Lit1850): Eine Frauenfahrt um die Welt ('A woman’s journey around the world') by Austrian author Ida Pfeiffer (1850). Full text available at Deutsches Textarchiv: http://www.deutschestextarchiv.de/pfeiffer_frauenfahrt01_1850/6.
  • Modern academic writing (Aca2009): Stand, Möglichkeiten und Grenzen der Telemedizin in Deutschland ('Telemedicine in Germany: status, chances and limits') by Rüdiger Klar and Ernst Pelikan, published in Bundesgesundheitsblatt ('Federal Health Gazette') in 2009. DOI 10.1007/s00103-009-0787-7.

The texts are annotated for part-of-speech and dependency syntax and are made available in CoNLL file format. We describe the annotation process and report inter-annotator agreements in:

Adelmann, Benedikt, Melanie Andresen, Wolfgang Menzel & Heike Zinsmeister. 2018. Evaluation of Out-Of Domain Dependency Parsing for its Application in a Digital Humanities Project. Proceedings of the 14th Conference on Natural Language Processing (KONVENS 2018). Vienna, Austria.

Files

Files (236.9 kB)

Name Size Download all
md5:00d7322efafa2fd30becf5b9459c8b98
87.8 kB Download
md5:e549c66a9de5911b0e4601eba9ddd0e9
77.8 kB Download
md5:7f3305071265490dd50636dbd4c86e78
71.4 kB Download

Additional details

References

  • Adelmann, Benedikt, Melanie Andresen, Wolfgang Menzel & Heike Zinsmeister. 2018. Evaluation of out-of Domain Dependency Parsing for its Application in a Digital Humanities Project. Proceedings of the 14th Conference on Natural Language Processing (KONVENS 2018). Vienna, Austria.