Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “Middle Dutch”
Middle Dutch syllabified words
<p><strong>Specifics of the data:</strong></p> <ul> <li>Text file (<em>syllabified_crm.txt</em>) containing 43,710 syllabified Middle Dutch words, taken from the <em>Corpus Van Reenen-Mulder</em>. This corpus, created by Pieter van Reenen en Maaike Mulder at the Free University Amsterdam, contains about 2,500 Middle Dutch charters. It has about 750,000 tokens. The charters were written in the Netherlands and Flanders between 1300 and 1400.</li> <li>The 43,710 syllabified words in this list is the total amount of unique words from the <em>Corpus Van Reenen-Mulder</em>. Some tokens from this corpus were, however, excluded when assembling the data set due to the fact that they contained diacritic symbols to indicate abbreviations, clitics, or unclear parts in the original charter.</li> <li>A dash-symbol (-) is used as separator.</li> <li>Apart from the entire data set, this DOI also includes: <ul> <li>A pdf-file visualizing the data set</li> <li>The splits used for the automatic syllabification experiment by Haverals, Kestemont & Karsdorp (2018).</li> <li>A gold standard out-of-corpus sample of 1,748 Middle Dutch words, taken at random from the <em>Cd-rom Middelnederlands</em>, also used in the above-mentioned syllabification experiment</li> </ul> </li> </ul>
Dataset of Middle Dutch lexical stress patterns and syllabifications
<p>This dataset consists of <strong>48.219 Middle Dutch words</strong> taken from in total 205 rhymed texts of the <em>Cd-rom Middelnederlands </em>(1998). All of these words have been <strong>assigned a syllabification and lexical stress pattern</strong>.</p> <p>E.g.: <em>proevede</em> is syllabified as <em>proe-ve-de</em> and has a stress index set at -3, which means that – counting from the rightmost syllable – the third syllable receives stress.</p> <p>This upload contains the following files:</p> <ul> <li>The <strong>JSON-file</strong> (compressed), which was used as input data for a machine learning algorithm trained for the automatic syllabification and stress assignment of Middle Dutch polysyllabic words (for the code of this experiment, see <a href="https://github.com/WHaverals/stresser">GitHub</a>)</li> <li>An <strong>Excel-file</strong>, containing the same data as the JSON (for more convenient reference)</li> <li>A <strong>split file </strong>(compressed), used in the training proces of the above-mentioned experiment</li> <li>A pdf-file with some <strong>insightful illustrations</strong> about the contents of the dataset</li> </ul> <p>This dataset is part of the research of <a href="https://www.uantwerpen.be/en/staff/wouter-haverals/research/">Wouter Haverals</a> (FWO, University of Antwerp), carried out under the supervision of prof. Mike Kestemont and em. prof. Frank Willaert.</p>
The Middle Dutch Manuscripts Surviving from the Carthusian Monastery of Herne (14th century)
<p>This repository contains the dataset described in the following conference paper:</p><blockquote><p>Wouter Haverals & Mike Kestemont, "The Middle Dutch Manuscripts Surviving from the Carthusian Monastery of Herne (14th century): Constructing an Open Dataset of Digital Transcriptions". CHR 2023: Computational Humanities Research Conference. December 6-8, 2023, Paris, France.</p></blockquote><p>The dataset consists of (automatically created) hyper-diplomatic, digital transcriptions of 18 Middle Dutch manuscripts that survive from the carthusian monastery in Herne in nowadays Belgium (or manuscripts which have meaningful ties with the charterhouse). These manuscripts primarily date to the second half of the fourteenth century and offer exciting possibilities for the analysis of authorship, translatorship and scribal practices in the history of the Low Countries. The transcriptions have been (partly) automated through the use of handwritten text recognition (on the Transkribus platform). This dataset is licensed under a CC-BY 4.0 licence, encouraging the further re-use of this data for all purposes, provided an unambiguous scholarly reference to the paper above is given.</p><p><strong>Content</strong></p><p>Transcriptions for the following 18 manuscripts are included in various formats:</p><ul><li>Brussels, RL, 1805-1808</li><li>Brussels, RL, 2485</li><li>Brussels, RL, 2849-51</li><li>Brussels, RL, 2877-78</li><li>Brussels, RL, 2879-80</li><li>Brussels, RL, 2905-09</li><li>Brussels, RL, 2979</li><li>Brussels, RL, 3091</li><li>Brussels, RL, 3093-95</li><li>Ghent, UL, 1374</li><li>Ghent, UL, 941</li><li>Paris, Bibl. Mazarine, 920</li><li>Paris, Bibl. de l'Arsenal, 8224</li><li>Saint Petersburg, BAN, O 256</li><li>Vienna, ÖNB, SN 12.857</li><li>Vienna, ÖNB, SN 12.905</li><li>Vienna, ÖNB, Cod. 13.708</li><li>Vienna, ÖNB, SN 65</li></ul><p>The contents of the repository have been structured as follows:</p><ul><li><i>transcriptions</i>: transcriptions of the 18 manuscripts in various formats (hyper-diplomatic; i.e. without brevigraph expansion):<ul><li>pagexmls: One file per folium, encoded in the PAGEXML format as outputted by Transkribus. One zip-file per manuscript folder.</li></ul></li><li><i>spreadsheets.zip</i>: detailed metadata on various aspects of the data in spreadsheat format.<ul><li>silent_voices_summary.xlsx: summary statistics at the codex-level (cf. Table 2 in the paper)</li><li>codex_info.xlsx: folium-level metadata</li><li>manuscript_data_metadata.xlsx: text region-level metadata</li><li>manuscript_data_metadata_rich.xlsx: contains the most convenient and complete version of the dataset, including the texts with automatically expanded abbreviations and the linguistic enrichment (lemma's and part-of-speech tags).</li></ul></li><li><i>code</i>: Python notebooks (requiring Python >= 3.8).<ul><li>transduction.ipynb: the notebook for the replication of the abbreviation expansion experiments described in the paper. (See also the configuration file for there tagger norm.json.</li><li>enrich.ipynb: the notebook used for the linguistic enrichment of the expanded texts, on the basis of the PIE(-NLP) lemmatizer. (See also the PIE model file herne-norm.tar, which is used in the enrichment.)</li><li>requirements.txt: third-party dependencies for running the code in these notebooks. Note: enrich.ipynb will require you the Middle Dutch (DUM) model for nlp-pie.</li></ul></li></ul><p><strong>Related data</strong></p><ul><li>The final Transkribus model used to generate the transcriptions will be make publicly available on the platform.</li><li>The accompanying images are released in a separate, restricted access repository on Zenodo, because we were unable to clear the copyright on some of the facsimiles. We will only be able to share these images under very strict conditions.</li></ul><p><strong>Acknowledgments</strong></p><p>Thanks to Anouck Kuypers, Sam Verellen and Frans de Jonge for their work on the transcriptions. The transcription of Brussels, RL, 3093-95 was contributed by Dr. Ine Kiekens. We acknowledge the help of Renée Gabriël and Peter Boot in previous collaborations that relate to the present paper. Finally, we would like to thank Caroline Vandyck who has helped with the finalization of the dataset.</p><p><strong>Funding statement</strong></p><p>This work has been funded by the Flemish Research Agency (FWO) in the context of the project "Silent voices: A Digital Study of the Herne Charterhouse as a Textual Community (ca. 1350-1400)".</p>
The Dutch middle route
<u>Source</u>: Europeana <br><u>4DCity URL</u>: <a href="https://4dcity.org/imgupload/1652523754.2146.jpg">https://4dcity.org/imgupload/1652523754.2146.jpg</a> <br><u>Original Image URL</u>: <a href="https://api.europeana.eu/thumbnail/v2/url.json?uri=https%3A%2F%2Fwww.openbeelden.nl%2Fimages%2F676379%2FDe_Hollandse_middenroute_%25280_38%2529.png&type=VIDEO">https://api.europeana.eu/thumbnail/v2/url.json?uri=https%3A%2F%2Fwww.openbeelden.nl%2Fimages%2F676379%2FDe_Hollandse_middenroute_%25280_38%2529.png&type=VIDEO</a> <br><br><u>Image-Metadata:</u><br>Filename: 1652523754.2146.jpg<br>Image Dimensions: 360x288<br>Megapixels: 0.10 MP<br>Filesize: 95.11 KB<br>
The Middle Dutch Manuscripts Surviving from the Carthusian Monastery of Herne (14th century) [facsimile image data]
<p>(*This repository is still being completed and will be finalized by 6 Dec 2023.)</p><p>This repository contains the original images underlying the dataset described in the following conference paper:</p><blockquote><p>Wouter Haverals & Mike Kestemont, "The Middle Dutch Manuscripts Surviving from the Carthusian Monastery of Herne (14th century): Constructing an Open Dataset of Digital Transcriptions". CHR 2023: Computational Humanities Research Conference. December 6-8, 2023, Paris, France.</p></blockquote><p>High-resolution facsimiles (photographic reproductions) are included for the following 18 manuscripts:</p><ul><li>Brussels, RL, 1805-1808</li><li>Brussels, RL, 2485</li><li>Brussels, RL, 2849-51</li><li>Brussels, RL, 2877-78</li><li>Brussels, RL, 2879-80</li><li>Brussels, RL, 2905-09</li><li>Brussels, RL, 2979</li><li>Brussels, RL, 3091</li><li>Brussels, RL, 3093-95</li><li>Ghent, UL, 1374</li><li>Ghent, UL, 941</li><li>Paris, Bibl. Mazarine, 920</li><li>Paris, Bibl. de l'Arsenal, 8224</li><li>Saint Petersburg, BAN, O 256</li><li>Vienna, ÖNB, SN 12.857</li><li>Vienna, ÖNB, SN 12.905</li><li>Vienna, ÖNB, Cod. 13.708</li><li>Vienna, ÖNB, SN 65</li></ul><p>The images in this repository have been used as the raw data in Transkribus for creating the open-access transcription dataset which can be freely accessed from this repository: <a href="https://zenodo.org/doi/10.5281/zenodo.10005253">https://zenodo.org/doi/10.5281/zenodo.10005253</a>. While we own the intellectual rights over the digital transcriptions in this other repository, we have so far not been able to clear the intellectual rights over (all of) the underlying images, which is why this repository is offered under "restricted access" only. Under Belgian copyright law, the image data could only be shared under specific circumstances and this repository is therefore mainly used for long-term sustainability purposes. Please contact us via the Zenodo platform if you are interested in re-using this data.</p><p><strong>Funding statement</strong></p><p>This work has been funded by the Flemish Research Agency (FWO) in the context of the project "Silent voices: A Digital Study of the Herne Charterhouse as a Textual Community (ca. 1350-1400)".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.