Published August 4, 2026 | Version v1
Other Open

PP-OCRv6 (small) multilingual handwritten and printed text recognition model

Authors/Creators

  • 1. ALMAnaCH, Inria Paris

Description

PP-OCRv6 (small) multilingual text recognition base model

Description

This is the small variant (~3.24M parameters) of a family of PP-OCRv6 text-line recognition models (tiny, small, medium) for kraken. The models are trained from scratch with baseline + bounding polygon data with a very diverse corpus containing historical, contemporary and born-digital document line images, handwritten and machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian, Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).

Architecture

PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.

The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:

  • increasing line height to 128px
  • uncapping line width/CTC label budget
  • replacing the optimizer with Adam+Muon

Uses

This is a base model that is supposed to produce usable output across a wide range of scripts and materials out of the box while also allowing fine-tuning with ease. It should offer similar accuracy and generalization to VLM-based recognizers without hallucinations and with vastly higher throughput. This medium variant model achieves the highest scores on the test set, tiny and small trade accuracy for inference speed and a smaller memory footprint.

Transcription guidelines, Normalization, and Transformations

No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.

Bias, Risks, and Limitations

The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.

The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.

How to Get Started with the Model

Install kraken (>= 7.1.0), download the model, and run recognition on an input image:

shell kraken -i image.png output.txt \ segment -bl \ ocr -m small.safetensors

For more information, refer to the documentation.

Training Details

Training Data

The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A marks languages that were additionally augmented with synthetic training data (see below).

| Language | Script | Datasets | |----------|:------:|----------| | Ancient Greek † | Greek | EPARCHOS, HPGTR, HTR_CPgr23, Stavronikita Monastery Greek Handwritten Document Collection No. 114, Stavronikita Monastery Greek Handwritten Document Collection No. 53, Stavronikita Monastery Greek Handwritten Document Collection No. 79, 11 private datasets | | Arabic † | Arabic | Agapet, arabic_ms_data, iskandar, Muharaf: Manuscripts of Handwritten Arabic Dataset, OpenITI Arabic Print Data, RASAM dataset, TariMa | | Catalan | Latin | FONDUE-CA-PRINT-20, htromance-spa, 1 private dataset | | Church Slavonic | Cyrillic | 3 private datasets | | Classical Armenian † | Armenian | synthetic only | | Corsican | Latin | OCR Corse | | Czech † | Latin | 2024--medieval-czech-main, 2023--medieval-czech, HTR Winter School 2025 - Medieval Czech - Biblioteka Jagiellonska BJ Rkp 441 IV, Paderov Bible handwriting ground truth, ehri | | Danish | Latin | ehri | | Dutch † | Latin | 6000 ground truth of VOC and notarial deeds / HTR of VOC, WIC and notarial deeds, ARletta, Dagboek Ernest Clarysse, FONDUE-NE-MSS-17-PR | | English | Latin | FONDUE-EN-PRINT-20, IAM Handwriting Database, jcrs_train, jcrs_val, JosephHookerHTR, OCR-D gt_structure_text, sloanelab, The Revolutionary City / RevCity documentation, Memorials for Jane Lathrop Stanford, ehri, 2 private datasets | | Finnish † | Latin | NewsEye/READ OCR Finnish Newspapers | | French | Latin | Antoine Verard extracts, Copiste-d-un-jour, corpus-HTR-lignes-mixtes, dataset-celestine-doniau-danest, FONDUE-FR-MSS-19, FONDUE-FR-MSS-19-PR, FONDUE-FR-PRINT-20, FONDUE-MLT-ART, genauto-td-htr, HTR Front Justice, La Correspondance Doucet-Rene Jean, Memoire sur St Domingue par H. M. Michel, Moonshines, NewsEye READ AS French Newspapers, NuBIS-OCR, Recensement Valaisan (Valais Time Machine), Tapus Corpus, TIMEUS Corpus, TitresNobiliaires_17_18, CREMMA Manuscrits du 20e, CREMMA Wikipedia, Maxime Kovalewsky - Coutume contemporaine et loi ancienne (1893), HTRomance, Modern Roman languages corpus, Peraire Ground Truth, PARES, 1 private dataset | | Georgian † | Georgian | 15 private datasets | | German | Latin | 2024--medieval-german, 2025--Early-Modern-German, Bullinger Digital Gwalther handwriting ground truth, charlottenburger-amtsschrifttum, Chronicling Germany, dach-gt, DigiTheo Ground Truth, Dresdner Hofdiarium, Fibeln, FONDUE-DE-MSS-16-PR, FONDUE-DE-MSS-18, FONDUE-DE-MSS-19-PR, FONDUE-DE-MSS-20-PR, FONDUE-MLT-ART, FONDUE-MLT-PRINT-TEST, FoNDUE_Kunsthistorisches-UZH_Archivdatenbank, Ground truth for Neue Zurcher Zeitung black letter, Hakenkreuzbanner, inzigkofen, Klosterneuburg, Stiftsbibl., Cod. 48, koenigsfelden, mkn-kurrent-gt, NewsEye / READ OCR Austrian Newspapers, nuremberg_letterbooks, OCR-D gt_structure_text, reichsanzeiger-gt, Training Data Incunabula Reichenau, Weisthuemer, gt-fraktur, german_kurrent_handwritten_text_lines, Ground Truth (Tagebücher Edwin Hennig), ehri, Fanny loves Wilhelm, Frauen im Fokus, Graphemic Early New German, 1 private dataset | | German (shorthand) | Latin | 1 private dataset | | Geʽez † | Ethiopic | synthetic only | | Hebrew | Hebrew | 2025-hebrew, 2 private datasets | | Hungarian | Latin | ehri | | Irish † | Latin | synthetic only | | Italian | Latin | Diario del Soldato Bruno Celestino, EpiSearch HTR, FONDUE-IT-PRINT-20, FONDUE-IT-PRINT-20-PR, HTRogène Medieval Italian Manuscripts, HTRomance, Medieval Italian corpus of ground-truth for Handwritten Text Recognition, leopardi, LiDi1.0-project, LAM, 1 private dataset | | Latin | Latin | 2025--late-medieval-latin-main, burchards-dekret-digital, Caroline Minuscule ground truth, Carolingian Latin Group HTR Wien Winter School 2025, CREMMA Medii Aevi, DISTINGUO Latin ground truth, Eutyches, FONDUE-LA-MSS-16-PR, FONDUE-LA-MSS-17-PR, FONDUE-LA-MSS-MA, FONDUE-LA-PRINT-16, HTR Winter School 2023/2024 - Late Medieval Latin, ONB 3891, HTR Winter School 2024/2025 - Late Medieval Latin, ONB 4135; ONB 4680, HTRogène Medieval Latin Manuscripts, HTRomance, Medieval Latin corpus of ground-truth for Handwritten Text Recognition, notarial_charter, nubis, OCR-D gt_structure_text, Paris Bible Project, Training Data Incunabula Reichenau, Wien ONB Cod 2160 ground truth | | Latvian † | Latin | synthetic only | | Lithuanian † | Latin | synthetic only | | Malayalam | Malayalam | Ground Truth data for printed Malayalam | | Middle Dutch | Latin | data | | Middle French | Latin | Cremma Medieval, De la généalogie des dieux, Données imprimés du 16e siècle, Données HTR incunables du 15e siècle, Données HTR manuscrits du 15e siècle, Données imprimés du 18e siècle, Données imprimés gothiques du 16e siècle, Fabliaux, FONDUE-FR-AAEB-16, FONDUE-FR-AAEB-17, FONDUE-FR-MSS-18, FONDUE-FR-PRINT-16, FONDUE-FR-PRINT-17, HTR-SETAF-Jean-Michel, HTR-SETAF-LesFaictzJCH, HTR-SETAF-Pierre-de-Vingle, HTRogene French, HTRomance, Medieval French corpus of ground-truth for Handwritten Text Recognition, Imprimés 17e siècle, Liber, OCR17plus, TNAH-2021-DecameronFR, transcription-chastel | | Multilingual (mixed) | Latin | Training Data Incunabula Reichenau, TranscriboQuest25_MedVernacReligio | | Norwegian | Latin | NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian | | Occitan | Latin | HTRogène Medieval Occitan Manuscripts, 1 private dataset | | Ottoman Turkish | Arabic | mehmed_ibn_mehmed_uskubi_cukrikcizade_altiparmak_risale, OpenITI Arabic-script OCR Catalyst Project print/typeface data | | Persian † | Arabic | hafiz_divan, OpenITI Arabic-script OCR Catalyst Project print/typeface data, sadi_gulistan | | Picard | Latin | 1 private dataset | | Polish † | Latin | ehri | | Portuguese | Latin | iForal-Dataset, Portuguese Handwriting 16th-19th c. | | Romanian † | Latin | synthetic only | | Russian † | Cyrillic | 2 private datasets | | Serbian (Cyrillic) † | Cyrillic | synthetic only | | Slovak | Latin | Slovensky Supermodel P&T1, ehri | | Slovenian † | Latin | synthetic only | | Spanish | Latin | FoNDUE Spanish chapbooks 19th c. Dataset, FONDUE-ES-MSS-19-PR, FONDUE-ES-PRINT-19, HTR - Araucania manuscript XIX, HTRogène Medieval Spanish Manuscripts, HTRomance, Medieval Spain corpus of ground-truth for Handwritten Text Recognition, ohg, 3 private datasets | | Swedish | Latin | Finnish Court Records-sub500, kat57 Swedish ground truth dataset, NewsEye / READ OCR training dataset from Swedish Newspapers, riskarchiv | | Syriac † | Syriac | zenodo.18157525, winter_school_vienna, 2 private datasets | | Ukrainian | Cyrillic | 1 private dataset | | Urdu | Arabic | OpenITI Arabic-script OCR Catalyst Project print/typeface data | | Yiddish | Hebrew | 7 private datasets |

Synthetic Training Data

Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.

Training Procedure and Hyperparameters

The model was trained with kraken (feature/ppocrv6_rec branch).

| | | |---|---| | Hardware | 3 × NVIDIA H100 | | Precision | bf16-mixed | | Optimizer | AdamW + Muon (momentum 0.95, weight decay 0.01) | | Learning rate | 8.5e-4 (cosine schedule, 1500 step warmup, min 1e-6) | | Batch size | 96 | | Gradient clipping | 1.0 | | Epochs | 16 | | Normalization | NFD + whitespace | | Augmentation | enabled |

Evaluation

Testing Data

Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.

Evaluations with purely synthetic data are marked with . CER and WER are the character- and word-level error rates, computed with torchmetrics using greedy CTC decoding and the NFD + whitespace normalization (equivalent to ketos test -u NFD -n).

Metrics

| Language | Lines | CER (%) | WER (%) | |----------|------:|--------:|--------:| | Ancient Greek | 451 | 13.24 | 66.76 | | Arabic | 1,926 | 13.63 | 52.44 | | Catalan | 128 | 2.48 | 13.33 | | Church Slavonic | 6,599 | 11.47 | 48.99 | | Classical Armenian ‡ | 3,580 | 0.15 | 0.90 | | Corsican | 50 | 1.23 | 8.24 | | Czech | 922 | 9.61 | 45.72 | | Danish | 25 | 0.81 | 6.44 | | Dutch | 4,384 | 8.97 | 35.90 | | English | 2,005 | 8.18 | 28.45 | | Finnish | 7,907 | 0.77 | 4.73 | | French | 5,967 | 10.01 | 20.27 | | Georgian | 394 | 19.41 | 70.57 | | German | 3,007 | 2.51 | 9.86 | | German (shorthand) | 830 | 25.78 | 62.53 | | Geʽez ‡ | 2,992 | 0.32 | 1.41 | | Hebrew | 3,426 | 6.50 | 18.48 | | Hungarian | 38 | 2.82 | 18.79 | | Irish ‡ | 2,824 | 0.32 | 1.57 | | Italian | 2,611 | 3.90 | 15.32 | | Latin | 5,748 | 9.12 | 31.91 | | Latvian ‡ | 2,397 | 0.48 | 2.68 | | Lithuanian ‡ | 2,615 | 0.71 | 3.69 | | Malayalam | 59 | 36.44 | 96.04 | | Middle Dutch | 3,014 | 8.00 | 29.59 | | Middle French | 3,970 | 5.03 | 21.14 | | Multilingual (mixed) | 241 | 1.63 | 8.78 | | Norwegian | 2,335 | 7.62 | 27.35 | | Ottoman Turkish | 451 | 7.22 | 33.05 | | Persian | 990 | 5.98 | 25.40 | | Polish | 2,766 | 0.70 | 4.21 | | Portuguese | 2,763 | 14.29 | 51.72 | | Romanian ‡ | 2,412 | 0.66 | 3.27 | | Russian | 3,053 | 16.34 | 49.26 | | Serbian (Cyrillic) ‡ | 2,789 | 0.10 | 0.56 | | Slovak | 1,550 | 3.24 | 14.08 | | Slovenian ‡ | 2,741 | 1.65 | 3.88 | | Spanish | 5,465 | 5.00 | 18.82 | | Swedish | 3,555 | 5.88 | 25.89 | | Syriac | 1,801 | 6.73 | 30.25 | | Ukrainian | 1,253 | 7.57 | 29.72 | | Urdu | 1,656 | 5.70 | 25.51 | | Yiddish | 4,320 | 5.81 | 21.54 | | Aggregate (micro-average) | 108,010 | 5.43 | 20.22 | | Aggregate (macro-average) | | 6.93 | 25.33 |

License

Released under the Apache-2.0 license.

Acknowledgements

Training of this model was funded by the European Union under Grant Agreement No.~101132163 (ATRIUM) and No.~101071829 (MiDRASH). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.

This project also received funding from the BPI Scribe project.

Citation

If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.

Files

README.md

Files (13.1 MB)

Name Size Download all
md5:d2dfcce1310b0fe6504acdb69fa9185e
31.7 kB Preview Download
md5:07279cd68cd4402807542799f12d0f4a
13.1 MB Download

Additional details