There is a newer version of the record available.

Published December 30, 2024 | Version v2
Other Open

Printed Urdu Base Model Trained on the OpenITI Corpus

Authors/Creators

  • 1. École Pratique des Hautes Études, Aoroc - CNRS PSL

Description

# Printed Urdu Base Model Trained on the OpenITI Corpus

This is a text recognition model trained on the OpenITI dataset of printed
Arabic-script text available [here](https://github.com/OpenITI/arabic_print_data.git) in its state of 2022-09-03. It encompasses Urdu (~11k lines) material in a variety of typefaces. The model has been obtained by fine-tuning the [Arabic-script base model](https://doi.org/10.5281/zenodo.7050296) on the purely Urdu subset of the corpus.

 The ground truth was lightly normalized to NFD but is otherwise untouched.

## Architecture

The default model architecture and hyperparameters of kraken 4.x where used.

## Uses

The model is trained on a variety of highly diverse typefaces it is mostly intended as a base model for fine-tuning more specific models from it. In line with this it has not been extensively verified or optimized.

## How to Get Started with the Model

Follow the instructions on installing and using kraken from the [website](https://kraken.re).

#### Metrics

CER: 4.13%

Files

README.md

Files (16.3 MB)

Name Size Download all
md5:ea2cf4baf331624b961227e60ec29023
2.6 kB Preview Download
md5:f4541c1ab4c6846d29d920d6f6825052
1.5 kB Preview Download
md5:6b9a7f3f8fc2ae68019b8dd457b0b1f3
16.3 MB Download