Published July 8, 2026 | Version 2

Diplomatic HTR ground truth dataset for an Early New High German transcription model (15th century), Version 2

  • 1. ROR icon Berlin-Brandenburg Academy of Sciences and Humanities
  • 2. ROR icon Academy of Sciences and Literature

Description

This repository contains a set of training data for ATR models (Kraken). It contains 50 pages of ground truth as image files (jpg) and transcription files (PAGE xml).
 
The ground truth contains 50 pages including 2,177 lines with 18,626 word tokens and 110,618 characters.
 
Please refer to the README.md file for further information.

A ground truth dataset following a graphemic transcription of the same data conntained within this repository may be found here: Graphemic HTR-Ground Truth dataset .

Authors

The data in this repository was prepared and curated by Adam Juszczak (ORCiD: 0009-0000-5330-6183) of the BBAW / Regesta of Emperor Frederik III and Frederik Skidzun (ORCiD: 0009-0002-7712-4207) of the AdW Mainz / Regesta Imperii Online.

License

This dataset is made available under the CC-BY 4.0 license.

Files

images.zip

Files (153.3 MB)

Name Size Download all
md5:fd3ba6556fd9f1b89f9ea709b53d0ebf
773 Bytes Download
md5:c5b0d0e441ba84cd4b7dd77d7afc855b
152.6 MB Preview Download
md5:022586c6f84dd4805a5b99205fe933c3
19.1 kB Preview Download
md5:166569d6dca627713467df0f068f3cf1
4.2 kB Preview Download
md5:19dd3a8fbeb5b2e7dc04d0bd6ba3b4ed
6.5 kB Preview Download
md5:93b1be1cec94d2edf5ed80776518d3a3
655.5 kB Preview Download

Additional details

Related works

Is new version of
Dataset: 10.5281/zenodo.18441030 (DOI)