ACHILLES: Ancient and Historical Language Evaluation Set
Authors/Creators
Description
The dataset used in the SIGTYP 2024 Shared Task on Word Embedding Evaluation for Ancient and Historical Languages. The task included four problems; problems 1-3 were offered in both constrained and unconstrained tracks on CodaLab, while problem 4 was only a part of the unconstrained track.
- POS-tagging
- Lemmatisation
- Morphological feature prediction
- Mask filling
- Word-level
- Character level
For problems 1-3, data from Universal Dependencies v.2.12 was used for Ancient Greek, Ancient Hebrew, Classical Chinese, Coptic, Gothic, medieval Icelandic, Latin, Old Church Slavonic, Old East Slavic, Old French and Vedic Sanskrit. Old Hungarian texts, annotated to the same standard as UD corpora, were added to the dataset from the MGTSZ website. In Old Hungarian data, tokens which were POS-tagged PUNCT were altered so that the form matched the lemma to simplify complex punctuation marks used to approximate manuscript symbols; otherwise, no characters were changed.
As the ISO 639-3 standard does not distinguish between historical stages of Latin, as it does between other languages like Irish, but it was desirable to approximate this distinction for Latin, we further split Latin data. This resulted in two Latin datasets: Classical and Late Latin, and Medieval Latin. This split was dictated by the composition of the Perseus and PROIEL treebanks that served as a source for Latin UD treebanks.
Historical forms of Irish were only included in mask filling challenges (problem 4), as the quantity of historical Irish text data which has been tokenised and annotated to a single standard to date is insufficient for the purpose of training models to perform morphological analysis tasks. The texts were drawn from CELT, Corpas Stairiúil na Gaeilge, and digital editions of the St. Gall glosses and the Würzburg glosses. Each Irish text taken from CELT is labelled "Old", "Middle" or "Early Modern" in accordance with the language labels provided in CELT metadata. Because CELT metadata relating to language stages and text dating is reliant on information provided by a variety of different editors of earlier print editions, this metadata can be inconsistent across the corpus and on occasion inaccurate. To mitigate complications arising from this, texts drawn from CELT were included in the dataset only if they had a single Irish language label and if the dates provided in CELT metadata for the text match the expected dates for the given period in the history of the Irish language.
The upper temporal boundary was set at 1700 CE, and texts created later than this date were not included in the dataset. The choice of this date is driven by the fact that most of the historical language data used in word embedding research dates back to the 18th century CE or later, and our intention was to focus on the more challenging and yet unaddressed data. The resulting datasets for each language were then shuffled at the sentence level and split into training, validation and test subsets at the ratio of 0.8 : 0.1 : 0.1.
A detailed list of text sources for each language in the dataset, as well as other metadata and the description of data formats used for each problem, is provided on the Shared Task's GitHub. The structure of the dataset is as follows:
π morphology (data for problems 1-3) βββ π test
βββ π ref (reference data used in CodaLab competitions)
βββ π lemmatisation
βββ π morph_features
βββ π pos_tagging
βββ π src (source test data with labels) βββ π train βββ π valid
π fill_mask_word (data for problem 4a)
βββ π test
βββ π ref (reference data used in CodaLab competitions)
βββ π src (source test data with labels in 2 different formats)
βββ π json
βββ π tsv
βββ π train (train data in 2 different formats)
βββ π json
βββ π tsv
βββ π valid (validation data in 2 different formats)
βββ π json
βββ π tsv
π fill_mask_char (data for problem 4b)
βββ π test
βββ π ref (reference data used in CodaLab competitions)
βββ π src (source test data with labels in 2 different formats)
βββ π json
βββ π tsv
βββ π train (train data in 2 different formats)
βββ π json
βββ π tsv
βββ π valid (validation data in 2 different formats)
βββ π json
βββ π tsv
We would like to thank Ekaterina Melnikova for suggesting the name for the dataset.
Files
SIGTYP2024.zip
Files
(146.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:88d557e8921c92afcb1d81e8468e38bf
|
146.7 MB | Preview Download |
Additional details
Additional titles
- Subtitle (English)
- The Dataset for SIGTYP 2024 Shared Task on Word Embedding Evaluation for Ancient and Historical Languages
Related works
- Is new version of
- Dataset: https://github.com/sigtyp/ST2024/ (URL)
Funding
- Irish Research Council
- CARDAMOM β Comparative Deep Models of Language for Minority and Historical Languages IRCLA/2017/129
- Science Foundation Ireland
- Insight SFI/12/RC/2289
- Science Foundation Ireland
- Insight 2 SFI/12/RC/2289_P2
Dates
- Created
-
2023-11-05Train & validation data published on Github & CodaLab
- Updated
-
2024-01-02Test data published on Github & CodaLab
- Updated
-
2024-01-12fill_mask_word test data & Classical Chinese fill_mask_char test data updated on Github & CodaLab
- Updated
-
2024-02-14Train & validation data for fill_mask_word and fill_mask_char tasks updated for publication on Zenodo
Software
- Repository URL
- https://github.com/sigtyp/ST2024/
- Programming language
- Python
References
- Acadamh RΓoga na hΓireann. (2017). Corpas StairiΓΊil na Gaeilge 1600-1926. Text Archive. http://corpas.ria.ie/index.php?fsg_function=1
- Bernhard Bauer, Rijcklof Hofman, and PΓ‘draic Moran. (2017). St. Gall Priscian Glosses, version 2.0. http://www.stgallpriscian.ie/
- Adrian Doyle (2018). WΓΌrzburg Irish Glosses. https://wuerzburg.ie/
- HAS Research Institute for Linguistics. (2018). Old Hungarian Codices. Hungarian Generative Diachronic Syntax. http://oldhungariancorpus.nytud.hu/en-codices.html
- Daniel Zeman, Joakim Nivre, Mitchell Abrams, Elia Ackermann, NoΓ«mi Aepli, Hamid Aghaei, Ε½eljko AgiΔ, Amir Ahmadi, Lars Ahrenberg, Chika Kennedy Ajede, Salih Furkan Akkurt, GabrielΔ AleksandraviΔiΕ«tΔ, Ika Alfina, Avner Algom, Khalid Alnajjar, Chiara Alzetta, Erik Andersen, Lene Antonsen, Tatsuya Aoyama, ..., and Rayan Ziane (2023). Universal Dependencies 2.12 [dataset]. http://hdl.handle.net/11234/1-5150
- Eszter Simon. (2014). Corpus building from Old Hungarian codices. In: Katalin Γ. Kiss (ed.): The Evolution of Functional Left Peripheries in Hungarian Syntax. Oxford: Oxford University Press.
- Donnchadh Γ CorrΓ‘in, Hiram Morgan, Beatrix FΓ€rber, Benjamin Hazard, Emer Purcell, CaoimhΓn Γ DΓ³naill, Hilary Lavelle, Julianne Nyhan, and Emma McCarthy. (1997). CELT: Corpus of Electronic Texts. Retrieved: March 15, 2021.
- Dag T. T. Haug and Marius L. JΓΈhndal. (2008). Creating a parallel treebank of the Old Indo-European Bible translations. In Proceedings of the Second Workshop on Language Technology for Cultural Heritage Data (LaTeCH 2008), pages 27β34.
- Giuseppe G. A. Celano, Daniel Zeman, and Federica Gamba. (2014). The Ancient Greek and Latin Dependency Treebank 2.0. Accessed: February 08, 2024.