Published August 19, 2019 | Version v1

A Character-Level LSTM Network Model for Tokenizing the Old Irish text of the Würzburg Glosses on the Pauline Epistles

  • 1. National University of Ireland Galway

Description

This paper examines difficulties inherent in tokenization of Early Irish texts and demonstrates that a neural-network-based approach may provide a viable solution for historical texts which contain unconventional spacing and spelling anomalies. Guidelines for tokenizing Old Irish text are presented and the creation of a character-level LSTM network is detailed, its accuracy assessed, and efforts at optimising its performance are recorded. Based on the results of this research it is expected that a character- level LSTM model may provide a viable solution for tokenization of historical texts where the use of Scriptio Continua, or alternative spacing conventions, makes the automatic separation of tokens difficult.

Files

doyle2019character.pdf

Files (280.7 kB)

Name Size Download all
md5:22636e5b8c988b0578d5417914c948d8
280.7 kB Preview Download

Additional details

Funding

European Commission
ELEXIS - European Lexicographic Infrastructure 731015