PP-OCRv6 (small) multilingual handwritten and printed text recognition model
Description
PP-OCRv6 (small) multilingual text recognition base model
Description
This is the small variant (~3.24M parameters) of a family of PP-OCRv6
text-line recognition models (tiny, small, medium) for
kraken. The models are trained from scratch with baseline
+ bounding polygon data with a very diverse corpus containing historical,
contemporary and born-digital document line images, handwritten and
machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian,
Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).
Architecture
PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.
The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:
- increasing line height to 128px
- uncapping line width/CTC label budget
- replacing the optimizer with Adam+Muon
Uses
This is a base model that is supposed to produce usable output across a wide
range of scripts and materials out of the box while also allowing fine-tuning
with ease. It should offer similar accuracy and generalization to VLM-based
recognizers without hallucinations and with vastly higher throughput. This
medium variant model achieves the highest scores on the test set, tiny and
small trade accuracy for inference speed and a smaller memory footprint.
Transcription guidelines, Normalization, and Transformations
No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.
Bias, Risks, and Limitations
The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.
The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.
How to Get Started with the Model
Install kraken (>= 7.1.0), download the model, and run recognition on an
input image:
shell
kraken -i image.png output.txt \
segment -bl \
ocr -m small.safetensors
For more information, refer to the documentation.
Training Details
Training Data
The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A † marks languages that were additionally augmented with synthetic training data (see below).
| Language | Script | Datasets | |----------|:------:|----------| | Ancient Greek † | Greek | EPARCHOS, HPGTR, HTR_CPgr23, Stavronikita Monastery Greek Handwritten Document Collection No. 114, Stavronikita Monastery Greek Handwritten Document Collection No. 53, Stavronikita Monastery Greek Handwritten Document Collection No. 79, 11 private datasets | | Arabic † | Arabic | Agapet, arabic_ms_data, iskandar, Muharaf: Manuscripts of Handwritten Arabic Dataset, OpenITI Arabic Print Data, RASAM dataset, TariMa | | Catalan | Latin | FONDUE-CA-PRINT-20, htromance-spa, 1 private dataset | | Church Slavonic | Cyrillic | 3 private datasets | | Classical Armenian † | Armenian | synthetic only | | Corsican | Latin | OCR Corse | | Czech † | Latin | 2024--medieval-czech-main, 2023--medieval-czech, HTR Winter School 2025 - Medieval Czech - Biblioteka Jagiellonska BJ Rkp 441 IV, Paderov Bible handwriting ground truth, ehri | | Danish | Latin | ehri | | Dutch † | Latin | 6000 ground truth of VOC and notarial deeds / HTR of VOC, WIC and notarial deeds, ARletta, Dagboek Ernest Clarysse, FONDUE-NE-MSS-17-PR | | English | Latin | FONDUE-EN-PRINT-20, IAM Handwriting Database, jcrs_train, jcrs_val, JosephHookerHTR, OCR-D gt_structure_text, sloanelab, The Revolutionary City / RevCity documentation, Memorials for Jane Lathrop Stanford, ehri, 2 private datasets | | Finnish † | Latin | NewsEye/READ OCR Finnish Newspapers | | French | Latin | Antoine Verard extracts, Copiste-d-un-jour, corpus-HTR-lignes-mixtes, dataset-celestine-doniau-danest, FONDUE-FR-MSS-19, FONDUE-FR-MSS-19-PR, FONDUE-FR-PRINT-20, FONDUE-MLT-ART, genauto-td-htr, HTR Front Justice, La Correspondance Doucet-Rene Jean, Memoire sur St Domingue par H. M. Michel, Moonshines, NewsEye READ AS French Newspapers, NuBIS-OCR, Recensement Valaisan (Valais Time Machine), Tapus Corpus, TIMEUS Corpus, TitresNobiliaires_17_18, CREMMA Manuscrits du 20e, CREMMA Wikipedia, Maxime Kovalewsky - Coutume contemporaine et loi ancienne (1893), HTRomance, Modern Roman languages corpus, Peraire Ground Truth, PARES, 1 private dataset | | Georgian † | Georgian | 15 private datasets | | German | Latin | 2024--medieval-german, 2025--Early-Modern-German, Bullinger Digital Gwalther handwriting ground truth, charlottenburger-amtsschrifttum, Chronicling Germany, dach-gt, DigiTheo Ground Truth, Dresdner Hofdiarium, Fibeln, FONDUE-DE-MSS-16-PR, FONDUE-DE-MSS-18, FONDUE-DE-MSS-19-PR, FONDUE-DE-MSS-20-PR, FONDUE-MLT-ART, FONDUE-MLT-PRINT-TEST, FoNDUE_Kunsthistorisches-UZH_Archivdatenbank, Ground truth for Neue Zurcher Zeitung black letter, Hakenkreuzbanner, inzigkofen, Klosterneuburg, Stiftsbibl., Cod. 48, koenigsfelden, mkn-kurrent-gt, NewsEye / READ OCR Austrian Newspapers, nuremberg_letterbooks, OCR-D gt_structure_text, reichsanzeiger-gt, Training Data Incunabula Reichenau, Weisthuemer, gt-fraktur, german_kurrent_handwritten_text_lines, Ground Truth (Tagebücher Edwin Hennig), ehri, Fanny loves Wilhelm, Frauen im Fokus, Graphemic Early New German, 1 private dataset | | German (shorthand) | Latin | 1 private dataset | | Geʽez † | Ethiopic | synthetic only | | Hebrew | Hebrew | 2025-hebrew, 2 private datasets | | Hungarian | Latin | ehri | | Irish † | Latin | synthetic only | | Italian | Latin | Diario del Soldato Bruno Celestino, EpiSearch HTR, FONDUE-IT-PRINT-20, FONDUE-IT-PRINT-20-PR, HTRogène Medieval Italian Manuscripts, HTRomance, Medieval Italian corpus of ground-truth for Handwritten Text Recognition, leopardi, LiDi1.0-project, LAM, 1 private dataset | | Latin | Latin | 2025--late-medieval-latin-main, burchards-dekret-digital, Caroline Minuscule ground truth, Carolingian Latin Group HTR Wien Winter School 2025, CREMMA Medii Aevi, DISTINGUO Latin ground truth, Eutyches, FONDUE-LA-MSS-16-PR, FONDUE-LA-MSS-17-PR, FONDUE-LA-MSS-MA, FONDUE-LA-PRINT-16, HTR Winter School 2023/2024 - Late Medieval Latin, ONB 3891, HTR Winter School 2024/2025 - Late Medieval Latin, ONB 4135; ONB 4680, HTRogène Medieval Latin Manuscripts, HTRomance, Medieval Latin corpus of ground-truth for Handwritten Text Recognition, notarial_charter, nubis, OCR-D gt_structure_text, Paris Bible Project, Training Data Incunabula Reichenau, Wien ONB Cod 2160 ground truth | | Latvian † | Latin | synthetic only | | Lithuanian † | Latin | synthetic only | | Malayalam | Malayalam | Ground Truth data for printed Malayalam | | Middle Dutch | Latin | data | | Middle French | Latin | Cremma Medieval, De la généalogie des dieux, Données imprimés du 16e siècle, Données HTR incunables du 15e siècle, Données HTR manuscrits du 15e siècle, Données imprimés du 18e siècle, Données imprimés gothiques du 16e siècle, Fabliaux, FONDUE-FR-AAEB-16, FONDUE-FR-AAEB-17, FONDUE-FR-MSS-18, FONDUE-FR-PRINT-16, FONDUE-FR-PRINT-17, HTR-SETAF-Jean-Michel, HTR-SETAF-LesFaictzJCH, HTR-SETAF-Pierre-de-Vingle, HTRogene French, HTRomance, Medieval French corpus of ground-truth for Handwritten Text Recognition, Imprimés 17e siècle, Liber, OCR17plus, TNAH-2021-DecameronFR, transcription-chastel | | Multilingual (mixed) | Latin | Training Data Incunabula Reichenau, TranscriboQuest25_MedVernacReligio | | Norwegian | Latin | NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian | | Occitan | Latin | HTRogène Medieval Occitan Manuscripts, 1 private dataset | | Ottoman Turkish | Arabic | mehmed_ibn_mehmed_uskubi_cukrikcizade_altiparmak_risale, OpenITI Arabic-script OCR Catalyst Project print/typeface data | | Persian † | Arabic | hafiz_divan, OpenITI Arabic-script OCR Catalyst Project print/typeface data, sadi_gulistan | | Picard | Latin | 1 private dataset | | Polish † | Latin | ehri | | Portuguese | Latin | iForal-Dataset, Portuguese Handwriting 16th-19th c. | | Romanian † | Latin | synthetic only | | Russian † | Cyrillic | 2 private datasets | | Serbian (Cyrillic) † | Cyrillic | synthetic only | | Slovak | Latin | Slovensky Supermodel P&T1, ehri | | Slovenian † | Latin | synthetic only | | Spanish | Latin | FoNDUE Spanish chapbooks 19th c. Dataset, FONDUE-ES-MSS-19-PR, FONDUE-ES-PRINT-19, HTR - Araucania manuscript XIX, HTRogène Medieval Spanish Manuscripts, HTRomance, Medieval Spain corpus of ground-truth for Handwritten Text Recognition, ohg, 3 private datasets | | Swedish | Latin | Finnish Court Records-sub500, kat57 Swedish ground truth dataset, NewsEye / READ OCR training dataset from Swedish Newspapers, riskarchiv | | Syriac † | Syriac | zenodo.18157525, winter_school_vienna, 2 private datasets | | Ukrainian | Cyrillic | 1 private dataset | | Urdu | Arabic | OpenITI Arabic-script OCR Catalyst Project print/typeface data | | Yiddish | Hebrew | 7 private datasets |
Synthetic Training Data
Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.
Training Procedure and Hyperparameters
The model was trained with kraken (feature/ppocrv6_rec branch).
| | | |---|---| | Hardware | 3 × NVIDIA H100 | | Precision | bf16-mixed | | Optimizer | AdamW + Muon (momentum 0.95, weight decay 0.01) | | Learning rate | 8.5e-4 (cosine schedule, 1500 step warmup, min 1e-6) | | Batch size | 96 | | Gradient clipping | 1.0 | | Epochs | 16 | | Normalization | NFD + whitespace | | Augmentation | enabled |
Evaluation
Testing Data
Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.
Evaluations with purely synthetic data are marked with ‡. CER and WER are
the character- and word-level error rates, computed with torchmetrics using
greedy CTC decoding and the NFD + whitespace normalization (equivalent to
ketos test -u NFD -n).
Metrics
| Language | Lines | CER (%) | WER (%) | |----------|------:|--------:|--------:| | Ancient Greek | 451 | 13.24 | 66.76 | | Arabic | 1,926 | 13.63 | 52.44 | | Catalan | 128 | 2.48 | 13.33 | | Church Slavonic | 6,599 | 11.47 | 48.99 | | Classical Armenian ‡ | 3,580 | 0.15 | 0.90 | | Corsican | 50 | 1.23 | 8.24 | | Czech | 922 | 9.61 | 45.72 | | Danish | 25 | 0.81 | 6.44 | | Dutch | 4,384 | 8.97 | 35.90 | | English | 2,005 | 8.18 | 28.45 | | Finnish | 7,907 | 0.77 | 4.73 | | French | 5,967 | 10.01 | 20.27 | | Georgian | 394 | 19.41 | 70.57 | | German | 3,007 | 2.51 | 9.86 | | German (shorthand) | 830 | 25.78 | 62.53 | | Geʽez ‡ | 2,992 | 0.32 | 1.41 | | Hebrew | 3,426 | 6.50 | 18.48 | | Hungarian | 38 | 2.82 | 18.79 | | Irish ‡ | 2,824 | 0.32 | 1.57 | | Italian | 2,611 | 3.90 | 15.32 | | Latin | 5,748 | 9.12 | 31.91 | | Latvian ‡ | 2,397 | 0.48 | 2.68 | | Lithuanian ‡ | 2,615 | 0.71 | 3.69 | | Malayalam | 59 | 36.44 | 96.04 | | Middle Dutch | 3,014 | 8.00 | 29.59 | | Middle French | 3,970 | 5.03 | 21.14 | | Multilingual (mixed) | 241 | 1.63 | 8.78 | | Norwegian | 2,335 | 7.62 | 27.35 | | Ottoman Turkish | 451 | 7.22 | 33.05 | | Persian | 990 | 5.98 | 25.40 | | Polish | 2,766 | 0.70 | 4.21 | | Portuguese | 2,763 | 14.29 | 51.72 | | Romanian ‡ | 2,412 | 0.66 | 3.27 | | Russian | 3,053 | 16.34 | 49.26 | | Serbian (Cyrillic) ‡ | 2,789 | 0.10 | 0.56 | | Slovak | 1,550 | 3.24 | 14.08 | | Slovenian ‡ | 2,741 | 1.65 | 3.88 | | Spanish | 5,465 | 5.00 | 18.82 | | Swedish | 3,555 | 5.88 | 25.89 | | Syriac | 1,801 | 6.73 | 30.25 | | Ukrainian | 1,253 | 7.57 | 29.72 | | Urdu | 1,656 | 5.70 | 25.51 | | Yiddish | 4,320 | 5.81 | 21.54 | | Aggregate (micro-average) | 108,010 | 5.43 | 20.22 | | Aggregate (macro-average) | | 6.93 | 25.33 |
License
Released under the Apache-2.0 license.
Acknowledgements
Training of this model was funded by the European Union under Grant Agreement No.~101132163 (ATRIUM) and No.~101071829 (MiDRASH). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.
This project also received funding from the BPI Scribe project.
Citation
If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.
Files
README.md
Files
(13.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:d2dfcce1310b0fe6504acdb69fa9185e
|
31.7 kB | Preview Download |
|
md5:07279cd68cd4402807542799f12d0f4a
|
13.1 MB | Download |