Published May 16, 2023 | Version v1
Other Open

OCR model for German prints trained from several datasets

  • 1. Universitätsbibliothek Mannheim
  • 2. ROR icon University of Mannheim

Description

german_print is a model which was trained for the recognition of printed (German and Latin) texts. The goal was to get a generic model which can be used with prints of all epochs, ranging from 15th incunabula until 20th century prints.

See https://github.com/UB-Mannheim/kraken/wiki/Training-German-Print for details on the training process.

Files

metadata.json

Files (16.4 MB)

Name Size Download all
md5:707720bff8e9f0e7848c1c859d1bc60e
16.4 MB Download
md5:9df38517cdf0310860239f8fe8d99fe7
2.9 kB Preview Download

Additional details

Funding

Deutsche Forschungsgemeinschaft
Workflow für werkspezifisches Training auf Basis generischer Modelle mit OCR-D sowie Ground-Truth-Aufwertung 460547474