Published January 19, 2021 | Version v1.0.0

Features for the classification of Spanish American 19th century novels by subgenre (part of data-nh)

Authors/Creators

  • 1. Ulrike

Description

This dataset contains feature sets that were prepared for the classification of Spanish American 19th century novels by subgenre. There are two main types of features sets: MFW-based features, and topic features. The MFW-based features include basic MFW, word n-grams, and character n-grams with different numbers of MFW and tf, tf-idf, or z-score normalization. The topic features are derived from topic models created with different parameters (number of topics, optimization intervals).

These feature sets were used in a classification analysis as a part of the dissertation "Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)" by Ulrike Henny-Krahmer. The dataset is part of "data-nh" (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.

Files

features.zip

Files (16.9 GB)

Name Size
md5:465ad14e2ab569c6fb7e60a4766d4629
16.9 GB Preview Download
md5:eface9b4d7245e73cc71c6591362b78d
252 Bytes Download

Additional details

Related works

Is part of
Dataset: https://github.com/cligs/data-nh (URL)