Supplementary code and data for the paper `From stage to page: language independent bootstrap measures of distinctiveness in fictional speech`
<p>The repository provides full data and processing / analysis pipeline for the paper <strong>'From stage to page: language independent bootstrap measures of distinctiveness in fictional speech</strong>'<br> <br> Rendered notebooks are also available through Github:</p> <p>1) <a href="https://github.com/perechen/difs-character-voices/blob/master/data/all_stars_clean.ipynb">Preparation, energy distance and exploration</a> (main)</p> <p>2) <a href="https://github.com/perechen/difs-character-voices/blob/master/03_analysis.md">Keyword curves & formal modeling</a></p> <p> </p> <p>- `00_dracor_get_data.R`. Script uses <a href="https://dracor.org/">DraCor</a> dedicated API to get texts spoken by characters</p> <p>- `01_distinctiveness_energy.ipynb` does the heavy lifting of data wrangling, cleaning and preprocessing, plus implements energy distance bootstrapping and does exploratory analysis</p> <p>- `02_logodds_curves.R` calculates keyword curves for characters<br> <br> - `03_analysis_and_models.R` explores keyword curves and does Bayesian models</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0