There is a newer version of the record available.

Published September 8, 2022 | Version 2.0.0

Dataset of CONTAINMENT and SUPPORT in the Uralic languages of the Volga-Kama area

  • 1. University of Turku, University of Oulu
  • 2. Ludwig-Maximilians-Universität München, University of Helsinki

Description

This open access dataset contains examples of the expressions of CONTAINMENT and SUPPORT in the Uralic languages of Volga-Kama area. The exact languages and the sources of data are given in Table 1. The dataset contains data on the relational nouns (RN) and plain spatial cases expressing prototypical CONTAINMENT and SUPPORT in the languages. The RN included in the dataset are listed in Table 2, and case forms in Table 3.

 

language

corpora

Erzya (MdE)

Syatko-subcorpus of the MokshEr corpus (MokshEr 2010)

Moksha (MdM)

Subcorpora in Moksha of the MokshEr corpus (MokshEr 2010)

Meadow Mari (MaM)

Marko East (Marko [no year]), Oncyko (Oncyko 2000), Meadow Mari corpus (Arkhangelskiy 2019b), Wanca (Meadow Mari) (Helsingin yliopisto et al. 2019)

Hill Mari (MaH)

Marko West (Marko [no year]), Wanca (Hill Mari) (Helsingin yliopisto et al. 2019)

Udmurt (Udm)

Pilot version of Udmurt corpus (relational nouns; presently included into [Arkhangelskiy 2018])

Udmurt corpus (content nouns) (Arkhangelskiy 2018)

Komi Zyrian (KoZ)

Komi Zyrian Web Corpus (Arkhangelskiy 2019a), Коми корпус (Fu-Lab team 2021)

Komi Permyak (KoP)

Komi Permyak text collection from the University of Turku (Permyak 2009)

Table 1. Languages included into the dataset and the sources of the data for each language.

 

 

MdE

MdM

MaM

MaH

Udm

KoZ

KoP

containment

pot(mo)-

potmə-

kørgø, kørgə-

kørgə̈-

puʃk-

pɨt͡ʃk-

pɨt͡ʃk-

support

lang-

lang-

ymba-

βə̈(l)-

vɨl-

vɨl-/vɨv-

vɨl-/vɨv-

Table 2. RN included in the dataset.

 

 

MdE

MdM

MaM

MaH

Udm

KoZ

KoP

location

-so/-se (inessive)

-sa (inessive)

-ʃte/-ʃto/-ʃtø (inessive)

-ʃtə/-ʃtə̈ (inessive)

-ɨn (inessive)

-ɨn (inessive)

-ɨn (inessive)

source

-sto/-ste (elative)

-sta (elative)

gət͡ɕ (source postposition)

gə̈t͡s (source postposition)

-ɨɕ (elative)

-ɨɕ (elative)

-iɕ (elative)

goal

-s (illative)

-s/-t͜s (illative)

-ʃke/-ʃko/-ʃkø/-ʃ, (illative)

-ʃkə/-ʃkə̈/-ʃ, (illative)

-e/-ɨ (illative)

-ɘ (illative)

-ɘ (illative)

path

-ka/-ga/-va (prolative)

-ka/-ga/-va/-gæ (prolative)

-

-

-ti/-eti/-jeti/-ɨti (prolative)

-ɘd (prolative); -ti (transitive)

-ɘt (prolative); -ti (transitive)

Table 3. Cases that have been included into the dataset. All cases do not necessary show in every set, as for some combinations of RN and case there is no data.

 

The main purpose of the dataset is to enable the study of variation between a plain case and RN inflected in case when expressing CONTAINMENT or SUPPORT. To facilitate this each expression of relation has been given a prototypicality score 4 = most prototypical, 1 = non-prototypical, which tells if the relation between landmark and trajector expressed in the sentence is typical for the entities participating in it. The prototypicality scores are based on the pre-linguistic concepts of containment and support, which are robustly attested and therefore should be independent of any single language. The scoring is based on the authors understanding of the language external relations, and no native consultants are used to verify the results. Therefore, some caution is in order when using the dataset.

 

The dataset contains files with data of CONTAINMENT RN, SUPPORT RN, and plain case on all the included languages. The files are named according to the scheme element_languge (e. g. Containment_Erzya for the containment data on Erzya). In addition, files named element_frequencies show the number of examples divided by case and prototypicality score for each language, and element_summary shows the total number of prototypicality scores for each language. For plain case there are also summary files for the scores of CONTAINMENT and SUPPORT separately.

 

The dataset is annotated for following information:

  1. The case in which the content noun or RN is inflected.
  2. The predicate as inflected in the data.
  3. The content noun as given in the data.
  4. Translations of both (mainly in citation form, but in predicate sometimes with some grammatical information, cf. abbreviations below).
  5. The prototypicality score. In the data on plain cases the prototypicality score is given only for the clauses where the relation is either CONTAINMENT or SUPPORT (i. e. the prototypicality score indicates the prototypicality of the relation as CONTAINMENT or SUPPORT according to the type of relation expressed).
  6. In the data on plain cases, the relation expressed by the case is marked (CONT = CONTAINMENT, SUP = SUPPORT, N/A = some other relation).
  7. The original sentence context.
  8. Free translation. Some of the translations are done following the lexical meanings and syntactic structures of the languages, so the English is unidiomatic from time to time.
  9. The file name with which the original sentence can be located in the corpus.

 

The three final columns are partly lacking at the moment from the Mari and Komi languages. The translations in the data are intended only as guidelines, and anyone using the dataset should refer to the original language data in the analysis. The data in the columns is presented according to the following conventions :

  • If the predicate is in square brackets, it means that the predicate is not present in the clause with the target LM. This can be because of two reasons: 1) The predicate is given in a previous clause, and is elliptically omitted, 2) the “predicate” is copula, which is not obligatory in the present tense in the languages studied.
  • The following abbreviations are used to specify the meaning of the predicate when the English translation is ambiguous (note that the use is not checked, and the abbreviations might be lacking from some predicates):

CAUS    causative

CONT   continuative

CVB      converb

FRQ      frequentative

INCH     inchoative

INF        infinitive

ITR        intransitive

MOM     momentaneous

NEG      negative

NMLZ    nominalization

PASS    passive

PTCP    participle

REFL    reflexive

TRA      transitive

The authors of this dataset are Tomi Koivunen and Riku Erkkilä and it is published under CC-BY-NC-ND licence. If used in a publication, please refer to this publication as well as mention the original source(s):

This dataset has been used in following publications:

 

References to used corpora:

Arkhangelskiy, Timofey. 2018. Udmurt corpus. http://udmurt.web-corpora.net/index.html.

Arkhangelskiy, Timofey. 2019a. Komi-Zyrian corpus. http://komi-zyrian.web-corpora.net/index.html.

Arkhangelskiy, Timofey. 2019b. Meadow Mari corpus. http://meadow-mari.web-corpora.net/index_en.html.

Fu-Lab team. 2021. Корпус коми языка. http://komicorpora.ru/.

Helsingin yliopisto, FIN-CLARIN, H. Jauhiainen, T. Jauhiainen & K. Lindén. 2019. Wanca 2016, Korp Version. Kielipankki. http://urn.fi/urn:nbn:fi:lb-2019052401.

Marko. (no year). MARKO - Corpus of Mari language. University of Turku.

MokshEr, V.3. 2010. Mokšan ja ersän sähköinen korpus. Turun yliopisto.

Oncyko. 2000. Oncyko corpus. University of Turku.

Permyak. 2009. Turku Komi-Permyak Corpus. University of Turku.

Files

Case_Erzya.csv

Files (336.6 kB)

Name Size Download all
md5:893dc5ba1617436270a54d5afdbf8065
25.4 kB Preview Download
md5:c1a6ea8a3ecbb86c7c777015e98c1990
1.2 kB Preview Download
md5:b676a8c0cefa66b5fa0baa64f64ee01e
4.9 kB Preview Download
md5:b38945f6644bd79414d65e6a0874fbc8
6.2 kB Preview Download
md5:a957f613d9fadbe81267f9663e5a7c41
6.9 kB Preview Download
md5:a83b70d8ff0839c7fcd579248d0aa06d
16.4 kB Preview Download
md5:3472cdd54ecc110b873451ead426f695
48.5 kB Preview Download
md5:f4ffb6dbe280728dd5979342cf548bf1
246 Bytes Preview Download
md5:32774b7ad601f8c29f55715a7f6477bb
266 Bytes Preview Download
md5:33c61acd65e82386e7b091756b1c895b
254 Bytes Preview Download
md5:a3c19742203da9a99355d033d3add8d0
27.8 kB Preview Download
md5:fd61b0143c3d7f461915965bdf60d3d1
27.0 kB Preview Download
md5:764bbf8ca9e7c4014cec4b7a7780e02c
722 Bytes Preview Download
md5:1d9fc4d7f341e81567a28f5b83dd799f
4.2 kB Preview Download
md5:39f39f2eb4c18350e2b63a9d90129125
4.9 kB Preview Download
md5:a065e636c9ca751443db7be26e98ce0f
5.3 kB Preview Download
md5:f925d1f90874be90255d096c813172dd
4.0 kB Preview Download
md5:7db2ec123e29e1b23d69d5f8a1fa3eb3
26.7 kB Preview Download
md5:51fb0ef4d120dfdc0c5215423e657832
267 Bytes Preview Download
md5:fad5567c1ae2dbcee147d9be1cd85876
27.2 kB Preview Download
md5:59b6953c76a1bdc9c366a021a705b5f4
24.7 kB Preview Download
md5:4a26c33ae6c49a917aa4eb8431ec0353
739 Bytes Preview Download
md5:1353faf6ce1e66f8e31a604880d72731
4.0 kB Preview Download
md5:52aa12663269d60291b7eb16df69325c
6.6 kB Preview Download
md5:2d2d43360b290856da084efec329e9ac
5.4 kB Preview Download
md5:76ad4b49f34b4ada894d98dfb6e92a31
4.0 kB Preview Download
md5:1b5519222f099087e2a814137d562476
24.5 kB Preview Download
md5:a34f9b1a9d0d3c9c0de4d4308ef0c2ad
267 Bytes Preview Download
md5:2a859caad965d5cfc8bc65c4f11a2e00
28.1 kB Preview Download