Published April 3, 2017 | Version v1

Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.

Authors/Creators

  • 1. Pattern Recognition and Human Language Technologies Research Center

Description

Train-B Dataset.   Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. 

This dataset is complementary to this other dataset:

https://zenodo.org/record/439807#.WOIBZ3WLSkA

More information at:

https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/

 

Files

Files (3.8 GB)

Name Size
md5:e11b9d0cb97169d64069268a23e90ef2
1.9 GB Download
md5:93ea0b7285f65c8438155e9490c691ed
1.9 GB Download

Additional details

Funding

European Commission
READ - Recognition and Enrichment of Archival Documents 674943