Kurdish Scene Text Recognition Version 1.0 (KSTRV1) Dataset
Authors/Creators
Description
KSTRV1 (Kurdish STR) Version 1
KSTRV1 is a large-scale dataset developed for Kurdish Scene Text Recognition (KSTR), addressing the scarcity of resources for non-Latin script like Kurdish. It includes 1,420 real-world scene images and 19,872 extracted word-level samples across Kurdish (Sorani and Badini dialects), Arabic, and English. To expand coverage and improve generalizability, the dataset is augmented with 20,000 synthetic text examples, crafted with diverse typography, multi-angle orientations, simulated distortions, and intricate background textures. This synthesis enhances the dataset’s capacity to handle real-world variability, supporting robust training for text recognition systems in underrepresented languages.
For more information please refere to the paper article in this link
⚠️The KSTRV1 dataset is no longer supported and has
been superseded by KSTRV2.Please use KSTRV2 for all
future research and development.
KSRTV2 link
Files
Additional details
Software
- Repository URL
- https://github.com/Sardar14/KSTRV1