There is a newer version of the record available.

Published 2025 | Version v2

Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations

Contributors

Data collector:

Description

A corpus of 105,738 synthetic Tamil text-line images with exact ground-truth transcriptions, rendered across 27 Unicode Tamil typefaces and intended for training and evaluating line-level OCR models — in particular for fine-tuning the LSTM recogniser used by Tesseract 4 and later.

Contents

Each instance is one rendered line of typeset Tamil: a grayscale TIFF image paired with its transcription, in tesstrain layout (<stem>.tif and <stem>.gt.txt). Transcriptions are Unicode NFC and are exact by construction, because the text is known before it is rendered.

How it was generated

  • Source texts are merged and segmented into fixed 12-word lines, normalised to NFC, and deduplicated (2,630 exact duplicate lines removed, 1.37%).
  • Lines are laid out 50 to a synthetic A4 page and rasterised at 300 dpi (2480 × 3508 px), with HarfBuzz shaping via Raqm so that conjunct formation and vowel-sign reordering are applied correctly.
  • Line crops are then recovered from the rendered page by horizontal projection profile, rather than emitted from known layout coordinates. This is deliberate: it gives training crops the same geometry a segmenter produces at inference time.

Typefaces

27 Unicode Tamil typefaces — 18 with conventional print letterforms and 9 with handwriting-style letterforms. All 27 carry GSUB tables and cover every Tamil codepoint occurring in the corpus. Sources include Google Fonts (SIL Open Font License 1.1) and the Tamilnadu Virtual Academy families.

Note: the handwriting-style faces are typefaces, not handwriting. This corpus contains no handwritten material and should not be used to train or evaluate handwriting recognisers.

Script coverage

Across 13.1M graphemes the corpus covers 227 of the 247 traditional Tamil syllabary units — all 12 uyir, all 18 mey, the aytham, and 196 of 216 uyirmey — plus 50 of 52 Grantha forms used in loanwords. Ten of the twenty absent units are the vowel series of ங, which in written Tamil occurs almost exclusively as ங் before க.

Known limitations

  • Clean rendering only. No degradation is modelled: no scanning noise, blur, skew, show-through or ink defects. Expect a gap to real document images.
  • Rare units are under-served. 37 syllabary units occur fewer than 27 times in the corpus, so they cannot appear in every typeface at any corpus size. Models trained on this data should be assumed weak on those units.
  • Register skew. Literary sources dominate; contemporary newsprint is roughly 12.7% of lines.
  • Lines only. No page context, so the corpus does not support layout analysis, tables or multi-column documents.
  • Not for held-out evaluation. A model scored on synthetic lines drawn from this same generative process will report an optimistic figure. Evaluate on real document images.

Source texts

Source Licence
Tamil Wikisource CC BY-SA 4.0
aitamilnadu/tamil_stories (AI Tamil Nadu) Apache-2.0
Theekkathir CC BY-SA 4.0
Tamil Wikinews CC BY-SA 4.0
Maattru CC BY-SA 4.0

The Apache-2.0 story collection is one-way compatible into BY-SA; its attribution and licence notices are preserved.

Generation pipeline

The complete pipeline that produced this corpus is open source at github.com/khaleeljageer/tesseract-gt-builder (GPL-3.0), so the corpus can be regenerated at a different scale, typeface set or register rather than only consumed as fixed.

File Structure

tam_new-ground-truth/
├── 00001.gt.txt
├── 00001.tiff
├── 00001.box
├── 00001.lstm
├── 00002.gt.txt
├── 00002.tiff
├── 00002.box
├── 00002.lstm
├── ...

Cite this work

@dataset{tamilocr_dataset_2025,
  author    = {Syedkhaleel Jageer},
  title     = {Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations},
  year      = {2025},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.16881612},
  url       = {https://doi.org/10.5281/zenodo.16881612}
}

 

Files

tam_new-ground-truth.zip

Files (3.6 GB)

Name Size
md5:07c1360d91fada4d129c361c47282120
3.6 GB Preview Download

Additional details

Related works

References
Dataset: 10.1109/IALP57159.2022.9961304 (DOI)

Dates

Available
2025

Software

Repository URL
https://github.com/khaleeljageer/tesseract-gt-builder
Programming language
Python
Development Status
Active