Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations
Authors/Creators
Contributors
Data collector:
Description
A corpus of 105,738 synthetic Tamil text-line images with exact ground-truth transcriptions, rendered across 27 Unicode Tamil typefaces and intended for training and evaluating line-level OCR models — in particular for fine-tuning the LSTM recogniser used by Tesseract 4 and later.
Contents
Each instance is one rendered line of typeset Tamil: a grayscale TIFF image paired with its transcription, in tesstrain layout (<stem>.tif and <stem>.gt.txt). Transcriptions are Unicode NFC and are exact by construction, because the text is known before it is rendered.
How it was generated
- Source texts are merged and segmented into fixed 12-word lines, normalised to NFC, and deduplicated (2,630 exact duplicate lines removed, 1.37%).
- Lines are laid out 50 to a synthetic A4 page and rasterised at 300 dpi (2480 × 3508 px), with HarfBuzz shaping via Raqm so that conjunct formation and vowel-sign reordering are applied correctly.
- Line crops are then recovered from the rendered page by horizontal projection profile, rather than emitted from known layout coordinates. This is deliberate: it gives training crops the same geometry a segmenter produces at inference time.
Typefaces
27 Unicode Tamil typefaces — 18 with conventional print letterforms and 9 with handwriting-style letterforms. All 27 carry GSUB tables and cover every Tamil codepoint occurring in the corpus. Sources include Google Fonts (SIL Open Font License 1.1) and the Tamilnadu Virtual Academy families.
Note: the handwriting-style faces are typefaces, not handwriting. This corpus contains no handwritten material and should not be used to train or evaluate handwriting recognisers.
Script coverage
Across 13.1M graphemes the corpus covers 227 of the 247 traditional Tamil syllabary units — all 12 uyir, all 18 mey, the aytham, and 196 of 216 uyirmey — plus 50 of 52 Grantha forms used in loanwords. Ten of the twenty absent units are the vowel series of ங, which in written Tamil occurs almost exclusively as ங் before க.
Known limitations
- Clean rendering only. No degradation is modelled: no scanning noise, blur, skew, show-through or ink defects. Expect a gap to real document images.
- Rare units are under-served. 37 syllabary units occur fewer than 27 times in the corpus, so they cannot appear in every typeface at any corpus size. Models trained on this data should be assumed weak on those units.
- Register skew. Literary sources dominate; contemporary newsprint is roughly 12.7% of lines.
- Lines only. No page context, so the corpus does not support layout analysis, tables or multi-column documents.
- Not for held-out evaluation. A model scored on synthetic lines drawn from this same generative process will report an optimistic figure. Evaluate on real document images.
Source texts
| Source | Licence |
|---|---|
| Tamil Wikisource | CC BY-SA 4.0 |
aitamilnadu/tamil_stories (AI Tamil Nadu) |
Apache-2.0 |
| Theekkathir | CC BY-SA 4.0 |
| Tamil Wikinews | CC BY-SA 4.0 |
| Maattru | CC BY-SA 4.0 |
The Apache-2.0 story collection is one-way compatible into BY-SA; its attribution and licence notices are preserved.
Generation pipeline
The complete pipeline that produced this corpus is open source at github.com/khaleeljageer/tesseract-gt-builder (GPL-3.0), so the corpus can be regenerated at a different scale, typeface set or register rather than only consumed as fixed.
File Structure
tam_new-ground-truth/
├── 00001.gt.txt
├── 00001.tiff
├── 00001.box
├── 00001.lstm
├── 00002.gt.txt
├── 00002.tiff
├── 00002.box
├── 00002.lstm
├── ...
Cite this work
@dataset{tamilocr_dataset_2025, author = {Syedkhaleel Jageer}, title = {Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations}, year = {2025}, publisher = {Zenodo}, doi = {10.5281/zenodo.16881612}, url = {https://doi.org/10.5281/zenodo.16881612}}
Files
tam_new-ground-truth.zip
Additional details
Related works
- References
- Dataset: 10.1109/IALP57159.2022.9961304 (DOI)
Dates
- Available
-
2025
Software
- Repository URL
- https://github.com/khaleeljageer/tesseract-gt-builder
- Programming language
- Python
- Development Status
- Active