Inferring quantitative typological trends from multilingual treebanks. A case study
Authors/Creators
- 1. CNR-ILC
Description
In this paper, we propose a novel approach to address the newly defined needs of linguistic typology recently interested in fine-grained features underlying language diversity. In fact, we introduce a method to extract qualitative and quantitative information about a wide range of features from multilingual annotated corpora based on Natural Language Processing methods and techniques. We tested our method in a case study focusing on word order variation in two widely investigated constructions, VERB-SUBJ(ect) and NOUN-ADJ(ective), with a specific view to structural and functional factors underlying the preference for one or the other order, both intra- and cross-linguistically, and their interaction. Preliminary experiments have been carried out aimed at acquiring typological evidence from a selection of linguistically annotated treebanks for three different languages, namely Italian, Spanish and English.
Files
LeL_alzetta_et_al_CameraReady_final.pdf
Files
(571.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e0d8080e3691e0aac685e04a46856fa6
|
571.0 kB | Preview Download |