Annotated Corpora of Historical Catalan (HisCat) - Llibre dels Fets
Authors/Creators
- 1. University of Cambridge and University of Alacant
- 2. University of Cambridge
Description
This repository is part of the Annotated Corpora of Historical Catalan (HisCat). It contains the first POS-tagged text that is partially manually corrected and used to train Old Catalan POS taggers, described in the following paper:
Meelen, Marieke & Pujol i Campeny, Afra, (2021) 'Old Catalan Morphosyntax: developing an annotated corpus' in Journal of Open Humanities Data.
This POS-tagged text is the 13th century Llibre dels Fets, a historical chronicle. The version of the text used for this project is
Bruguera, J. (1991). El Llibre dels Fets del Rei en Jaume. Barcelona: Barcino.
as prepared for the Corpus Informatitzat del Català Antic
Torruella, J., Pérez Saldanya, M., & Martines, J. (2009). Corpus Informatitzat del Català Antic. URL: http://cica.cat/.
The subcorpus counts with 164,096 POS-annotated tokens (165,538 tokens including punctuation and folio markers), of which 60,000 have been manually corrected. This subcorpus contains a total of and 4,506 main clauses. POS tagging of this text was done with the Memory-Based Tagger by TiMBL (https://languagemachines.github.io/mbt/). The code accompanying the paper can be found on GitHub: https://github.com/lothelanor/catalancorpora). In addition to memory-based tagging, have tried neural-based tagging with TARGER (https://github.com/achernodub/targer) for which we created word embeddings that can be found on Zenodo. Results for memory-based tagging were better, however, which is why this version is uploaded here.
Notes
Files
LlibredelsFets_60000corrected.txt
Files
(2.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:21e21de064c0af190bc9c582c5aa682c
|
2.1 MB | Preview Download |
Additional details
Related works
- References
- Other: https://zenodo.org/record/5615556#.YXvZkR3TVqs (URL)