Published December 16, 2019 | Version v1

Inferring quantitative typological trends from multilingual treebanks. A case study

Description

In this paper, we propose a novel approach to address the newly defined needs of linguistic typology recently interested in fine-grained features underlying language diversity. In fact, we introduce a method to extract qualitative and quantitative information about a wide range of features from multilingual annotated corpora based on Natural Language Processing methods and techniques. We tested our method in a case study focusing on word order variation in two widely investigated constructions, VERB-SUBJ(ect) and NOUN-ADJ(ective), with a specific view to structural and functional factors underlying the preference for one or the other order, both intra- and cross-linguistically, and their interaction. Preliminary experiments have been carried out aimed at acquiring typological evidence from a selection of linguistically annotated treebanks for three different languages, namely Italian, Spanish and English.

Files

LeL_alzetta_et_al_CameraReady_final.pdf

Files (571.0 kB)

Name Size Download all
md5:e0d8080e3691e0aac685e04a46856fa6
571.0 kB Preview Download