childTale-A: A corpus of eighty fairy tales from the 7th edition by the Brothers Grimm, manually annotated for textually encoded emotions
<p>The childTale-A corpus is a collection of eighty fairy tales, a core set of the Grimms’ Children's and Household Tales as introduced in Herrmann & Lüdtke (2023). Within the CHYLSA project, textually encoded emotions were annotated in each sentence in each of the eighty fairy tales. Annotations were collected for the dimensions <em>valence</em> and <em>arousal</em>, as well as for the six basic emotions <em>anger</em>, <em>disgust</em>, <em>fear</em>, <em>joy</em>, <em>sadness</em>, and <em>surprise</em>. Each fairy tale was annotated by two persons (for details see Hermann & Lüdtke, 2023).</p> <p>In detail, this dataset contains:</p> <ul> <li>instructions and texts used for annotation (zip-files): <ul> <li>all N=80 fairy tales as txt-files (with normalised orthography)</li> <li>instructions (in German) for the valence and arousal annotation as well as for the annotation of the six basic emotions</li> </ul> </li> <li>scripts for preparing and analysing annotations for textually encoded emotions: <ul> <li>Python scripts to prepare Excel files for annotation (including a script to separate texts into individual sentences)</li> <li>R-scripts for data preparation, calculation of inter-rater reliability and smoothing of the valence annotations (discrete cosine transformation (DCT) with length normalisation) <ul> </ul> </li> </ul> </li> <li>sentence-level data (for each sentence in each of the eighty fairy tales): <ul> <li>annotations for the dimensions valence and arousal (continuous values)</li> <li>categorisation of each sentence (as negative, neutral or positive) based on the continuous valence annotation</li> <li>annotations on the occurrence of the six basic emotions anger, disgust, fear, joy, sadness, and surprise</li> </ul> </li> <li>transformed and length-normalised valence annotations as basis for the <em>Emotional Arcs</em> (DCT and length normalisation results in one hundred data points for each fairy tale, saved in a separate data file)</li> <li>text-level data (for each of the eighty fairy tales): <ul> <li>general information, for example title in German and English, number of sentences, Kinder- und Hausmärchen-ID (KMH-ID), corpus ID</li> <li>results of the analysis of the annotated data with values for: <ul> <li><em>Average Valence</em> and <em>Average Arousal</em></li> <li>inter-rater reliability index (Krippendorff's alpha coefficient) for the valence and arousal annotations</li> <li>proportions of positive, negative and neutral sentences</li> <li><em>Emotion Potential</em> (percent of both positive and negative sentences)</li> <li><em>Valence Span</em>, <em>Arousal Span</em> and range of the <em>Emotional Arc</em></li> <li><em>Emotion Profile</em> (relative frequency for each of the six basic emotions <em>anger</em>, <em>disgust</em>, <em>fear</em>, <em>joy</em>, <em>sadness</em>, and <em>surprise)</em></li> <li>inter-rater reliability indices (Krippendorff's alpha coefficient and the percentage of agreement) for each of the six basic emotions</li> </ul> </li> </ul> </li> </ul> <p>Reference: <br> Herrmann, Berenike & Lüdtke, Jana (2023). A Fairy Tale Gold Standard. Annotation and Analysis of Emotions in the Children's and Household Tales by the Brothers Grimm.<em> </em>Zeitschrift für digitale Geisteswissenschaften (ZfdG). DOI: 10.17175/2023_005.</p> <p> </p> <p>DFG Schwerpunktprogramm SPP 2207 “Computational Literary Studies“<br> Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/402743989">https://gepris.dfg.de/gepris/projekt/402743989</a></li> <li><a href="https://dfg-spp-cls.github.io/">https://dfg-spp-cls.github.io<em>/</em></a></li> </ol> <p>Teilprojekt: „CHYLSA - Children’s and Youth Literature Sentiment Analysis“</p> <p>Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/424250469">https://gepris.dfg.de/gepris/projekt/424250469</a></li> <li><a href="https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/">https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/</a></li> </ol>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4