Published May 22, 2026 | Version v1

Annotated dataset of Russian -nie nominalisations

  • 1. ROR icon University of Graz
  • 2. ROR icon University of Novi Sad

Description

Annotated dataset of Russian -nie nominalisations

Daria Seres (University of Graz)

Marko Simonović (University of Graz)

Predrag Kovačević (University of  Novi Sad)

Goal and rationale

This dataset was created as a modern Russian comparison dataset for the analysis of -nie nominalisations in Simonović, Kovačević and Milićev (accepted). Its purpose is to provide comparable modern Russian evidence against which the Slavonic-Serbian -nie nominalisation data can be contrasted.

Dataset description

Each row represents one attested token of a modern Russian -nie nominalisation drawn from the Russian National Corpus (RNC; ruscorpora.ru). The data were extracted from RNC texts classified as учебно-научная ‘academic/educational-scientific’ and публицистика ‘journalistic writing’ and restricted to texts produced after 2000. The sample was randomly selected and then manually checked and annotated.

The sample contains:

  • 1,705 tokens

  • 435 lemmas

The data were extracted from the Russian National Corpus using the following lemma query:

[lemma=".*(т|н)ие"]

This query targets lemmas ending in -ние or -тие. Since the query targets formal lemma endings, it can retrieve false positives. In this dataset, such items were not removed, but their status is explicitly marked in Departicipial and by blank or uncertain values in the subsequent annotation columns.

In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus.

In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus.

Column A — ID

ID assigns an arbitrary unique number to each example in the dataset.

Column B — Citation form

Contains the citation form of the nominalisation, for example разрешение ‘permission’, клонирование ‘cloning’, исследование ‘research’, образование ‘education’, исполнение ‘execution; fulfilment’, and создание ‘creation’.

Due to Russian inflectional morphology, the citation form may differ from the attested surface form in the Center column. For example, the citation form клонирование ‘cloning’ corresponds to attested forms such as клонирования, and разрешение ‘permission’ may correspond to разрешение or разрешения depending on case and number.

Column C — Departicipial

Marks whether the item is analysed as belonging to the target deverbal/departicipial -nie nominalisation class.

Values:

  • 1 = the item is analysed as a target deverbal/departicipial -nie nominalisation.

  • 0 = the item is not analysed as a target deverbal/departicipial -nie nominalisation.

  • ? = the status is uncertain.

Examples marked 1 include разрешение ‘permission’, клонирование ‘cloning’, использование ‘use’, создание ‘creation’, исследование ‘research’, and выполнение ‘performance; fulfilment’. Items that do not have an attested corresponding verb, but for which that verb could be constructed and its aspect could be determined were also marked 1.

Examples marked 0 include formal false positives such as тысячелетие ‘millennium’, 200-летие ‘200th anniversary’, 850-летие ‘850th anniversary’, поколение ‘generation’ and Минобразования ‘Ministry of Education’. These items match the formal extraction pattern but were not analysed as target deverbal/departicipial nominalisations.

Examples marked ? include uncertain cases where the nominal is perceived as deverbal/departicipial, but the corresponding verb could not be constructed. Examples are волнение ‘agitation’ and возникновение ‘appearance, emergence’. These were retained in the dataset but marked as uncertain rather than forced into a binary classification.

Column D — Compound stem

Compound stem marks whether the nominalisation contains a compound stem.

Values:

  • 1 = the nominalisation contains a compound stem.

  • 0 = the nominalisation does not contain a compound stem.

  • blank = not applicable or not annotated, usually because the item was not analysed as a target nominalisation or its status was uncertain.

Examples marked 1 include налогообложение ‘taxation’, правоотношение ‘legal relation’, здравоохранение ‘health care’, водоснабжение ‘water supply’, телевидение ‘television’, and мироощущение ‘worldview; sense of the world’.

Examples marked 0 include разрешение ‘permission’, клонирование ‘cloning’, исследование ‘research’, использование ‘use’, создание ‘creation’, and решение ‘decision; solution’.

Note that Slavic prefixes and the negation particle were not counted as compound material.

Column E — Perfective base

Perfective base marks whether the nominalisation is based on a perfective verb.

Values:

  • 1 = the nominalisation has a perfective base.

  • 0 = the nominalisation does not have a perfective base, or the base is imperfective, biaspectual, lexicalised, or otherwise not clearly perfective.

  • ? = the aspectual status of the base is uncertain.

  • blank = not applicable or not annotated, usually because the item was not analysed as a target nominalisation or its status was uncertain.

Examples marked 1 include разрешение ‘permission’, воссоздание ‘recreation’, создание ‘creation’, окончание ‘ending; completion’, возрождение ‘revival’, выполнение ‘performance; fulfilment’, сокращение ‘reduction’, and повышение ‘increase’.

Examples marked 0 include клонирование ‘cloning’, исследование ‘research’, использование ‘use’, образование ‘education’, страхование ‘insurance’, течение ‘course; flow’, содержание ‘content; maintenance’, and соревнование ‘competition’.

 

Column F — Unattested verb

Marks whether the base verb is an actually attested verb. For instance,   разрешение ‘permission’ has the value 1 because разрешить is an attested verb, but упражнение ‘exercise’ got the value 0 because  упражнить is not an attested verb, even though it can be constructed and its aspect can be determined.

Column G — Left context

Contains the text immediately preceding the attested target form. It functions as the left side of a concordance line.

For example, in the row with разрешение ‘permission’, the left context is:

мамонта, японской стороне предстоит получить специальное

Column H — Center

Contains the exact attested surface form of the nominalisation in the corpus. This form may be inflected, capitalised, or otherwise different from the citation form.

For example, the citation form растение ‘plant’ corresponds to attested forms such as растений, and решение ‘decision; solution’ corresponds to forms such as решение, решения, and решений.

Column I — Right context

Contains the text immediately following the attested target form. Together with Left context and Center, this column provides the concordance context for each example.

For example, in the row with разрешение ‘permission’, the right context is:

от российских властей, что предположительно будет

Column J — Title

Contains the title or bibliographic identifier of the RNC source text from which the example was extracted.

Examples include Японцы хотят клонировать якутского мамонта ‘The Japanese want to clone a Yakut mammoth’, Юрий Патрикеев: «Проиграл, потому что боролся с Гарднером слишком долго», and Учет и налогообложение операций по страхованию работников ‘Accounting and taxation of employee insurance operations’.

Column K — Author

Contains the author name recorded in the Russian National Corpus metadata, where available.

Column L — Birthday

Contains the author’s year of birth where this information is available in the corpus metadata. The field is blank when the information is not provided.

Column M — Header

Contains the document header recorded in the RNC metadata. In many cases this repeats or shortens the source title.

Column N — Created

Contains the year in which the source text was created. The dataset is restricted to texts produced after 2000.

Column O — Sphere

Contains the broad RNC domain classification. In this dataset, the relevant values are учебно-научная ‘academic/educational-scientific’ and публицистика ‘journalistic writing’.

Column P — Type

Contains the text-type classification assigned by the RNC.

Examples in the dataset include заметка ‘short article/note’, интервью ‘interview’, статья ‘article’, аннотация ‘abstract’, письмо деловое ‘business letter’, and характеристика ‘character reference/evaluation’.

Column Q — Topic

Contains the thematic classification assigned by the RNC.

Examples include наука и технологии ‘science and technology’, искусство и культура ‘art and culture’, политика и общественная жизнь ‘politics and public life’, право ‘law’, образование ‘education’, спорт ‘sport’, and бизнес, коммерция, экономика, финансы ‘business, commerce, economics, finance’.

Column R — Publication

Contains the publication source recorded in the RNC metadata.

Examples include «Известия», «Домовой», «Вечерняя Москва», «Вопросы статистики», «Физика твердого тела», and «Бухгалтерский учёт».

Column S — Publ_year

Contains the publication year recorded in the RNC metadata. This usually corresponds to the year in Created, but the two fields are kept separate because they come from distinct metadata fields.

Column T — Medium

Contains the medium classification assigned by the RNC.

Examples include газета ‘newspaper’, журнал ‘magazine/journal’, машинопись ‘typescript’, and электронный текст ‘electronic text’.

Column U — Ambiguity

Contains information on ambiguity resolution in the RNC metadata. In the present dataset, examples commonly have the value омонимия снята ‘homonymy resolved’, indicating that morphological ambiguity was resolved in the corpus annotation.

Column V — Full context

Contains the complete corpus context from which the example was extracted. This field allows the reader to verify the local interpretation of the nominalisation and its annotation.

For example, the full context for разрешение ‘permission’ is:

Для того чтобы вывезти в Японию фрагмент ткани мамонта, японской стороне предстоит получить специальное разрешение от российских властей, что предположительно будет сделано осенью.

Reference

Russian National Corpus. Available at: ruscorpora.ru.

Simonović, Marko, Predrag Kovačević and Tanja Milićev. Accepted. When -nie met -nje: Slavonic-Serbian loan deverbal nominals. To appear in Stefan Milosavljević, Daria Seres, Jelena Stojković, Marko Simonović and Jelena Živojinović (eds.), Advances in Formal Slavic Linguistics 2023. Berlin: Language Science Press.

Files

Annotated dataset of Russian -nie nominalisations.pdf

Files (586.5 kB)

Name Size Download all
md5:034133c1282363cd85c2bde301ccda01
162.9 kB Preview Download
md5:3194ef1fc2c1a810c40c40778c965840
423.6 kB Download