Published April 2026 | Version v1

Jibo-labeled classical hiragana image dataset

  • 1. ROR icon University of Washington
  • 2. University of Washington, Seattle

Description

JIBO-LABELED CLASSICAL HIRAGANA IMAGE DATASET

This dataset is intended to be used to train an optical character recognition (OCR) model to recognize hiragana 平仮名 in classical Japanese manuscripts and woodblock printed books. Hiragana are Japanese phonetic symbols derived from cursivized forms of Chinese characters. Among them, so-called hentaigana 変体仮名 (lit. variant kana) are forms of hiragana that have been more-or-less obsolete since 1900. Decoding hentaigana is a non-trivial task that is essential for text extraction from images of classical Japanese manuscripts, i.e. those inscribed before 1900.

In order to train a deep-learning model that can recognize individual jibo variants, we modified the hiragana section of an existing dataset called the Kotenseki kuzushiji dataset (Version 2, 2019, http://codh.rois.ac.jp/char-shape/, consisting of 44 items, 6,151 scans, 4,328 character types, and 1,086,326 individual characters) and refined it with jibo labels. Rather than adding the jibo labels by hand, we used the Kokugoken hentaigana jikei database (consisting of 7 items) to train a small model that automatically added jibo labels, which we then corrected by human inspection. Note: as of this writing (20 April 2026), the CODH website is offline, so we provide a working link for the former dataset through the internet archive: https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/.

The Kotenseki kuzushiji dataset provides coordinate information for another dataset consisting of scans of classical Japanese manuscripts, called the Kotenseki dataset. Both datasets were jointly developed by the National Institute of Japanese Literature (NIJL) and the Center for Open Data in the Humanities (CODH). Regarding the contents of the Kotenseki dataset, please see the following page: https://codh.rois.ac.jp/char-shape/book/ (accessed 27 August 2025). Note: as of this writing (20 April 2026), the CODH website is offline, so we provide a working link through the internet archive: https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/book/.

The current Jibo-labeled classical hiragana image dataset contains the hiragana portion of the Kotenseki kuzushiji dataset with jibo labels. It was derived from the Kotenseki kuzushiji dataset, with bootstrapping of the labeling from the Kokugoken hentaigana jikei database. We made the best effort to correct auto-labeled data, but potential mistakes remain in the current version.

We are releasing this dataset in hopes that other will be able to use it to train other models or to add jibo-decoding functionalities to their existing classical Japanese OCR models. 

For more details, please see the file README.md. 

ACKNOWLEDGEMENTS

We express our deepest appreciation to: • the National Institute for Japanese Literature (Kokubungaku kenkyu shiryokan 国文学研究資料館) for creating, tagging, and sharing the images from their collections; • the Center for Open Data in the Humanities (Jinbungaku oopun deeta kyodoriyo sentaa 人文学オープンデータ共同利用センター) for processing and releasing the original dataset; • the National Institute for Japanese Language and Linguistics (Kokuritsu kokugo kenkyusho 国立国語研究所) for sharing their valuable data online.

We thank Amazon Web Services (AWS) and the eScience Institute of the University of Washington, Seattle, for providing cloud computing credits. We thank the Simpson Center for the Humanities, University of Washington, Seattle for Digital Humanities Summer Fellowships (Summer 2025).

ATTRIBUTION

If you reuse this dataset, the following citation pattern is recommended:

"Michael R. Zeng, Herman Chau, and Paul S. Atkins (2026). Jibo-labeled classical hiragana image dataset. doi:10.5281/zenodo.18765466."

LICENSE AND DISCLAIMERS

This dataset is bound by the terms of the underlying Nihon kotenseki kuzushiji deetasetto, which was publicly released under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/). The underlying dataset is provided by ROIS-DS Center for Open Data in the Humanities (CODH) (https://web.archive.org/web/20250827113242/http://codh.rois.ac.jp/). Here is the recommended citation for the original dataset:

"Nihon kotenseki kuzushiji deetasetto (from the collection of NIJL and others, processed by CODH) 『日本古典籍くずし字データセット』 (国文研ほか所蔵/CODH加工) doi:10.20676/00000340"

Corrections and minor changes were made to the original dataset, in addition to the jibo tags. The creators of the original dataset do not necessarily endorse this derivative.

No warranties are given regarding this dataset. We cannot provide technical support.

If you find the dataset useful, please let us know! We would love to hear from you.

Files

Jibo-labeled Hentaigana Image Dataset.zip

Files (1.4 GB)

Name Size
md5:fa49df49868436c6bbb2bf3996e62d8c
1.4 GB Preview Download
md5:f20275f689581d1a25c76b5901f3d9c1
11.4 kB Preview Download

Additional details

Related works

Is derived from
Dataset: https://codh.rois.ac.jp/char-shape/ (URL)

Funding

University of Washington

Dates

Issued
2026-04-22

References

  • Version 2, Nihon kotenseki kuzushiji deetasetto 日本古典籍くずし字データセット. National Institute of Japanese Literature (NIJL) and Center for Open Data in the Humanities (CODH). 11 November 2019. http://codh.rois.ac.jp/char-shape/. Accessed on 13 August 2023.
  • Kokugoken hentaigana jikei deetabeesu 国語研変体仮名字形データベース. National Institute for Japanese Language and Linguistics (NINJAL). https://cid.ninjal.ac.jp/hentaiganaDB/index.html. Accessed on 20 April 2026.
  • Herman Chau, Michael R. Zeng, and Paul S. Atkins. "Deriving Orthographic Data from Classical Japanese Texts with Machine-Learning Methods." Jinmoncom 2025 rombunshū じんもんこん2025論文集 (December 2025): 19–24. https://cir.nii.ac.jp/crid/1050306495453310464 (accessed on 20 April 2026).