Published November 29, 2024 | Version v1

Namgyal Manuscript Collection Datasets

Description

These are the official datasets created for the Tibetan Manuscript Project Vienna (TMPV) in the years 2023 and 2024. These datasets contain:

  • OCR datasets (line image - line label pairs) created from the PageXML annotations
  • PageXML (Transkribus) annotations in Unicode and Wylie
  • PageXML Layout annotations (lines, images, captions, margins) used for image segmentation training
  • OCR models (PyTorch checkpoints and ONNX model files)

Files

OCR_Model_2024_11_29_23_8.zip

Files (1.3 GB)

Name Size
md5:060811a2063ac2d3d90ad0f59212c2f0
95.3 MB Preview Download
md5:b3fd7f1e5fe8695203ab9c1d29fcc1a9
96.0 MB Preview Download
md5:7964474eb25886da57fb5ad7fdcbfb64
55.1 MB Preview Download
md5:4b50828e4105c56030bafbf525c8fe0f
55.2 MB Preview Download
md5:ad59fb48b78f4b4d68b5423d85748a52
444.7 MB Preview Download
md5:0cfe23731bdf332ca61720646f0a93cd
269.5 MB Preview Download
md5:e96331b04d81e9a86156a4d378adb685
266.3 MB Preview Download

Additional details

Software

Repository URL
https://github.com/eric86y/Namgyal-OCR
Programming language
Python