Published September 10, 2025 | Version 1

ZEFYS2025: A German Dataset for Named Entity Recognition and Entity Linking for Historical Newspapers

Description

In order to enable the training of machine learning models capable of correctly identifying named entities and linking them to wikidata entities, we provide a large corpus of 100 German-language newspaper pages published between 1837 and 1940. The machine learning task for which this dataset was collected falls into the domain of token classification and, more generally, of natural language processing.

The dataset was compiled by collaborators in the research project "Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe" at the Staatsbibliothek zu Berlin – Berlin State Library (SBB). The research project was funded by the Federal Government Commissioner for Culture and the Media (BKM), project grant no. 2522DIG002. The Minister of State for Culture and the Media is part of the German Federal Government.

Files

ZEFYS2025_ A German Dataset for Named Entity Recognition and Entity Linking for Historical Newspapers.md

Additional details

Software

Repository URL
https://github.com/qurator-spk/sbb_ner_hf
Programming language
Python
Development Status
Active