There is a newer version of the record available.

Published March 17, 2025 | Version 1.0
Dataset Restricted

Kurdish Scene Text Recognition Version 1.0 (KSTRV1) Dataset

  • 1. ROR icon University of Zakho
  • 2. ROR icon Duhok Polytechnic University

Description

KSTRV1 (Kurdish STR) Version 1 

KSTRV1 is a large-scale dataset developed for Kurdish Scene Text Recognition (KSTR), addressing the scarcity of resources for non-Latin script like Kurdish. It includes 1,420 real-world scene images and 19,872 extracted word-level samples across Kurdish (Sorani and Badini dialects), Arabic, and English. To expand coverage and improve generalizability, the dataset is augmented with 20,000 synthetic text examples, crafted with diverse typography, multi-angle orientations, simulated distortions, and intricate background textures. This synthesis enhances the dataset’s capacity to handle real-world variability, supporting robust training for text recognition systems in underrepresented languages.

For more information please refere to the paper article in this link 

⚠️The KSTRV1 dataset is no longer supported and has
been superseded by KSTRV2.Please use
KSTRV2 for all
future research and development.
KSRTV2 link



Files

Restricted

The record is publicly accessible, but files are restricted. Log in to check if you have access.

Additional details

Software