ICDAR 2019 Competition on Post-OCR Text Correction

Rigaud, Christophe; Doucet, Antoine; Coustaty, Mickael; Moreux, Jean-Philippe

doi:10.5281/zenodo.3459116

Published September 24, 2019 | Version v1

Conference paper Open

ICDAR 2019 Competition on Post-OCR Text Correction

1. L3i Laboratory, University of La Rochelle
2. National Library of France

This paper describes the second round of the ICDAR 2019 competition on post-OCR text correction and presents the different methods submitted by the participants. OCR has been an active research field for over the past 30 years but results are still imperfect, especially for historical documents. The purpose of this competition is to compare and evaluate automatic approaches for correcting (denoising) OCR-ed texts. The present challenge consists of two tasks: 1) error detection and 2) error correction. An original dataset of 22M OCR-ed symbols along with an aligned ground truth was provided to the participants with 80% of the dataset dedicated to training and 20% to evaluation. Different sources were aggregated and contain newspapers, historical printed documents as well as manuscripts and shopping receipts, covering 10 European languages (Bulgarian, Czech, Dutch, English, Finish, French, German, Polish, Spanish and Slovak). Five teams submitted results, the error detection scores vary from 41 to 95% and the best error correction improvement is 44%. This competition, which counted 34 registrations, illustrates the strong interest of the community to improve OCR output, which is a key issue to any digitization process involving textual data.

Dataset

In addition to the paper, you may also be interested in the datasets of the ICDAR 2019 Competition on Post-OCR Text Correction.

Files

ICDAR2019_POCR_report.pdf

Files (226.9 kB)

Name	Size	Download all
ICDAR2019_POCR_report.pdf md5:5224addcce8ebf49dcb2eb75a3d33ed4	226.9 kB	Preview Download

Additional details

European Commission
NewsEye - NewsEye: A Digital Investigator for Historical Newspapers 770299

	All versions	This version
Views	1,144	1,142
Downloads	928	926
Data volume	228.7 MB	228.2 MB

ICDAR 2019 Competition on Post-OCR Text Correction

Creators

Description

Files

ICDAR2019_POCR_report.pdf

Files (226.9 kB)

Additional details

Funding