Published January 19, 2021 | Version v1.0.0
Dataset Restricted

Corpus of 19th century Spanish American novels for family resemblance analysis (part of data-nh)

Authors/Creators

Description

This dataset contains different formats of a corpus of 19th century Spanish American novels which were used in a family resemblance analysis as a part of the dissertation "Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)" by Ulrike Henny-Krahmer. The texts are included as plain text files, linguistically annotated files (using TreeTagger), as text files with only noun lemmas, and as chunks of 1,000 noun tokens derived from the lemmatized texts. The texts were prepared in this way to be used with topic modeling. Because 22 of the novels still are under copyright, this dataset has restricted access. The other 234 novels are in the open domain. This dataset is part of "data-nh" (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.

Files

Restricted

The record is publicly accessible, but files are restricted. Log in to check if you have access.

Request access

If you would like to request access to these files, please fill out the form below.

You need to satisfy these conditions in order for this request to be accepted:

As part of the novels in the corpus is still under copyright, access is only granted for the inspection of research results and to ensure their reproducibility, but the texts may not be published freely.

You are currently not logged in. Do you have an account? Log in here

Additional details

Related works

Is part of
Dataset: https://github.com/cligs/data-nh (URL)