Skip to main content
zenodoopen

Datasets and code for "Mapping the plague through natural language processing"

<p>This project investigates the performance of various NLP libraries and geocoding services for the semi-automated generation of quantitative datasets from narrative texts. We provide the original files, several intermediate data products as well as the final plague datasets.&nbsp;Please note that some of the steps in this process were done manually, thus some of the scripts cannot be run completely.&nbsp;</p> <p>This work is based on two plague treatises:</p> <p>- Sticker, G. 1908 <em>Abhandlungen aus der Seuchengeschichte und Seuchenlehre. Band 1: Die Pest</em>. Giessen, A. T&ouml;pelmann.</p> <p>- Biraben, J.-N. 1975 <em>Les hommes et la peste en France et dans les pays europ&eacute;ens et m&eacute;diterran&eacute;ens</em>. Paris, Mouton.</p> <p>The final geocoded, plague datasets are:</p> <p><strong>- plague_sticker_v1.csv</strong></p> <p><strong>- plague_biraben_v1.csv</strong></p> <p>A data dictionary is available as</p> <p><strong>plague_datadict.xlsx</strong></p> <p>Other files:</p> <table> <tbody> <tr> <td>file name</td> <td>content</td> </tr> <tr> <td>sticker_OCR_orig.txt</td> <td>Original OCR text</td> </tr> <tr> <td>sticker_OCR.txt</td> <td>Original OCR text without parenthesis (author names)</td> </tr> <tr> <td>sticker_textprep.rds</td> <td>Original OCR text with further preparations</td> </tr> <tr> <td>sticker_goldstandard_annotated_1.tsv</td> <td>manual annotations file 1</td> </tr> <tr> <td>sticker_goldstandard_annotated_2.tsv</td> <td>manual annotations file 2</td> </tr> <tr> <td>sticker_goldstandard_annotated_consensus.tsv</td> <td>consenus annotation file</td> </tr> <tr> <td>sticker_standard_toponyms.csv</td> <td>Gold standard for toponym recognition. Contains the tokenization, the start/end character respective to the OCR text (orig and without parenthesis) and whether a token is a location or other</td> </tr> <tr> <td>sticker_comparison_NER.rds</td> <td>Comparison of NER performance</td> </tr> <tr> <td>sticker_comparison_geocoding.rds</td> <td>Comparison of Geocoding performance</td> </tr> </tbody> </table> <p>&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
12
Reuse readiness
8
Engagement
4

Topics