Datasets and code for "Mapping the plague through natural language processing"
<p>This project investigates the performance of various NLP libraries and geocoding services for the semi-automated generation of quantitative datasets from narrative texts. We provide the original files, several intermediate data products as well as the final plague datasets. Please note that some of the steps in this process were done manually, thus some of the scripts cannot be run completely. </p> <p>This work is based on two plague treatises:</p> <p>- Sticker, G. 1908 <em>Abhandlungen aus der Seuchengeschichte und Seuchenlehre. Band 1: Die Pest</em>. Giessen, A. Töpelmann.</p> <p>- Biraben, J.-N. 1975 <em>Les hommes et la peste en France et dans les pays européens et méditerranéens</em>. Paris, Mouton.</p> <p>The final geocoded, plague datasets are:</p> <p><strong>- plague_sticker_v1.csv</strong></p> <p><strong>- plague_biraben_v1.csv</strong></p> <p>A data dictionary is available as</p> <p><strong>plague_datadict.xlsx</strong></p> <p>Other files:</p> <table> <tbody> <tr> <td>file name</td> <td>content</td> </tr> <tr> <td>sticker_OCR_orig.txt</td> <td>Original OCR text</td> </tr> <tr> <td>sticker_OCR.txt</td> <td>Original OCR text without parenthesis (author names)</td> </tr> <tr> <td>sticker_textprep.rds</td> <td>Original OCR text with further preparations</td> </tr> <tr> <td>sticker_goldstandard_annotated_1.tsv</td> <td>manual annotations file 1</td> </tr> <tr> <td>sticker_goldstandard_annotated_2.tsv</td> <td>manual annotations file 2</td> </tr> <tr> <td>sticker_goldstandard_annotated_consensus.tsv</td> <td>consenus annotation file</td> </tr> <tr> <td>sticker_standard_toponyms.csv</td> <td>Gold standard for toponym recognition. Contains the tokenization, the start/end character respective to the OCR text (orig and without parenthesis) and whether a token is a location or other</td> </tr> <tr> <td>sticker_comparison_NER.rds</td> <td>Comparison of NER performance</td> </tr> <tr> <td>sticker_comparison_geocoding.rds</td> <td>Comparison of Geocoding performance</td> </tr> </tbody> </table> <p> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 12
- Reuse readiness
- 8
- Engagement
- 4