Features for the classification of Spanish American 19th century novels by subgenre (part of data-nh)
<p>This dataset contains feature sets that were prepared for the classification of Spanish American 19th century novels by subgenre. There are two main types of features sets: MFW-based features, and topic features. The MFW-based features include basic MFW, word n-grams, and character n-grams with different numbers of MFW and tf, tf-idf, or z-score normalization. The topic features are derived from topic models created with different parameters (number of topics, optimization intervals).</p> <p>These feature sets were used in a classification analysis as a part of the dissertation "Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)" by Ulrike Henny-Krahmer. The dataset is part of "data-nh" (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 8
- Engagement
- 4