Skip to main content
zenodoopen

Features for the classification of Spanish American 19th century novels by subgenre (part of data-nh)

<p>This dataset contains feature sets that were prepared for the classification of Spanish American 19th century novels by subgenre. There are two main types of features sets: MFW-based features, and topic features. The MFW-based features include basic MFW, word n-grams, and character n-grams with different numbers of MFW and tf, tf-idf, or z-score normalization. The topic features are derived from topic models created with different parameters (number of topics, optimization intervals).</p> <p>These feature sets were used in a classification analysis as a part of the dissertation &quot;Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)&quot; by Ulrike Henny-Krahmer. The dataset is part of &quot;data-nh&quot; (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
8
Reuse readiness
8
Engagement
4

Topics