ZEFYS2025: A German Dataset for Named Entity Recognition and Entity Linking for Historical Newspapers
Authors/Creators
Description
In order to enable the training of machine learning models capable of correctly identifying named entities and linking them to wikidata entities, we provide a large corpus of 100 German-language newspaper pages published between 1837 and 1940. The machine learning task for which this dataset was collected falls into the domain of token classification and, more generally, of natural language processing.
The dataset was compiled by collaborators in the research project "Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe" at the Staatsbibliothek zu Berlin – Berlin State Library (SBB). The research project was funded by the Federal Government Commissioner for Culture and the Media (BKM), project grant no. 2522DIG002. The Minister of State for Culture and the Media is part of the German Federal Government.
Files
ZEFYS2025_ A German Dataset for Named Entity Recognition and Entity Linking for Historical Newspapers.md
Files
(2.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:76f0948086b3923ef751e1fd3a1805a1
|
2.1 MB | Preview Download |
|
md5:d49745156cd4a4e6b57c350ed3a506bd
|
23.5 kB | Preview Download |
Additional details
Software
- Repository URL
- https://github.com/qurator-spk/sbb_ner_hf
- Programming language
- Python
- Development Status
- Active