Published August 23, 2017 | Version v1

Textmining von Konferenzabstracts: Dokumentation eines Arbeitsprozesses in den Digitalen Geisteswissenschaften. Ein Werkstattbericht

  • 1. Max-Planck-Institut für Wissenschaftsgeschichte
  • 2. Universität Hamburg
  • 3. Humboldt-Universität zu Berlin

Description

This dataset accompanies a research of the mining of conference abstracts. The article describes the digital workflow of automatically extracting various entities from a corpus of conference abstracts and performing network analysis on the results in order to find relationships. The aim is to identify the spread and diversity of tools, the representation of institutions and federations, and the thematic and personal networks of speakers at a German-speaking DH conference. Particular emphasis is placed on the documentation of the entire process, which can be interpreted as an initiative to develop interdisciplinary standards in this field.

Included in the dataset are

  • a training corpus (Trainingskorpuserweitert.csv)
  • a configuration file (Dariah6.prop)
  • a model file gained from the Stanford parser (dariah—6ner-model2.ser.gz)

Files

Trainingskorpuserweitert.csv

Files (27.1 MB)

Name Size Download all
md5:cb3f0987f7b94636652f987746a87b6b
1.2 kB Download
md5:fb15d2f2a8c0ebd6ae53bc0e6349e6c6
26.8 MB Download
md5:502e9664c0ad02fb10827761c69ab79b
370.3 kB Preview Download