Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

878

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

878 results for “corpus”

Learn how ShareScore rates datasets ↗
zenodo52/100

S1000 corpus, large-scale tagging results and other supplementary files

<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>

opencc-by-4.0Sep 2022View details →
zenodo52/100

Corpus of critical citations contexts

<p>We present here a corpus of 505 critical citation contexts, i.e. a set of sentences or propositions that contain at least one citation of a study towards which the author(s) has/have a negative opinion. Those contexts come from other existing annotated corpora, from our readings about critical citation and disagreement in science, and from contexts manually annotated by native speakers of English. We have re-annotated all those contexts in order to be sure that they match our definition of critical citations. This corpus can be helpful to train tools dedicated to the automatic retrieval of critical citations. English (2024-02-20)</p>

opencc-by-4.0Feb 2024View details →
zenodo52/100

List of TEI rolename annotations in the ISicily EpiDoc corpus

<p>This CSV file details every instance of a 'roleName' tag in the I.Sicily (sicily.classics.ox.ac.uk) EpiDoc TEI files, reporting the ID number of the file in which it appears, and the value of the @type and @subtype attributes in each case - as such it serves as an index of roleName attestations in the I.Sicily dataset (also recoverable directly from the EpiDoc files). The file will be updated in future.</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

UNIC Templates for uploading corpus metadata v1.11

<p>The UNIC platform (https://unic.dipintra.it) accepts a JSON file for uploading corpus metadata based on the template here. Alternatively, use the spreadsheet template to input the corpus metadata and convert the resulting .xlsm file to JSON using this application at https://huggingface.co/spaces/nannanliu/UNIC_metadata_conversion. When opening the Excel spreadsheet template, please enable Macros, which will automatically validate your input in the columns. Please do not change the order of the columns because they are embedded with code. To add elements and components not included by the UNIC schema, create new columns after the existing ones.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

ZooCor Corpus (fr-it)

<p>ZooCor is a bilingual (fr-it) specialized corpus, consisting of 350 texts concerning marine fauna and conservation biology. The ZooCor corpus covers a time window from 2000 to 2022. The partial metadata related to the subdomain of the threatened species of chelonids (marine turtles) are published. For further information on the entire corpus, please adress to silvia.zollo@uniparthenope.it.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Word Embedding of Amazon Product Review Corpus

<p>A word embedding of the <a href="https://www.cs.uic.edu/~liub/FBS/sentiment-analysis.html#datasets">Amazon Product Review Corpus</a> (<a href="https://www.doi.org/10.1145/1341531.1341560">Jindal and Liu, 2008</a>).</p> <p>Created using <a href="https://code.google.com/archive/p/word2vec/">Word2Vec</a> in CBOW mode, 500 dimensions and window size 5.</p> <p>Words have been lemmatised and particle verbs have been merged into a single token (e.g. <code>calm_down</code>).</p> <ul> </ul> <p>&nbsp;</p> <p><strong>Attribution</strong></p> <p>This dataset was created as part of the following publication:</p> <p>Marc Schulder,&nbsp;Michael Wiegand,&nbsp;Josef Ruppenhofer&nbsp;and&nbsp;Benjamin Roth&nbsp;(2017).&nbsp;<strong>&quot;Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features&quot;</strong>. Proceedings of the 8th International Joint Conference on Natural Language Processing (IJCNLP). Taipei, Taiwan, November 27 - December 3, 2017.&nbsp;<a href="https://doi.org/10.5281/zenodo.3365609">DOI: 10.5281/zenodo.3365609</a>.</p> <p>If you use the data in your research or work, please cite the publication.</p>

opencc-by-4.0Nov 2017View details →
zenodo52/100

BioASQ-QA: A manually curated corpus for Biomedical Question Answering

<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>

opencc-by-2.5Dec 2022View details →
zenodo52/100

Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets

<p>Murreviikko (literally &#39;Dialect week&#39;) is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko (&#39;dialect week&#39;) and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, H&auml;me, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni &#39;my book-my&#39;) are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata &gt; pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id&#39;s, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union&rsquo;s Digital Single Market directive (2019/790).</p>

opencc-by-4.0May 2023View details →
zenodo48/100

WTC1.1 (WikiTailor corpus v. 1.1)

<p>&nbsp;</p><p><strong>Content:</strong></p><p>List of the 743 domains, their term vocabularies in 10 languages, and the Wikipedia articles associated to each domain extracted by the best model described in:</p><p><i>&nbsp; Cristina España-Bonet, Alberto Barrón-Cedeño and Lluís Màrquez. "Tailoring and Evaluating the Wikipedia for in-Domain Comparable Corpora Extraction." &nbsp;</i>Knowledge and Information Systems, Volume 65, pages 1365-1397. 2023. Springer-Verlag, London Ldt. https://doi.org/10.1007/s10115-022-01767-5</p><p><i>&nbsp; https://github.com/cristinae/WikiTailor</i></p><p>&nbsp;</p><p><strong>Files Description:</strong></p><ul><li>commonCats2015.enesdefrcaareuelrooc.tsv</li></ul><p>Multilingual domains listed one per line, languages are separated by a tab in the order en, es, de, fr, ca, ar, eu, el, ro and oc. For each language we include the pair "ID categoryName" separated by a blank space.</p><ul><li>[LAN].0.tar.bz</li></ul><p>A folder per domain for language [LAN] containing the vocabulary and IDs of the extracted articles by the Wikitailor model 50-WT100.</p><ul><li>extraction[LAN]0.tar.bz</li></ul><p>A folder per domain for language [LAN] containing the text of the extracted articles. The name of the file corresponds to the IDs in [LAN].0.tar.bz.</p><p>&nbsp;</p><p>&nbsp;</p>

opencc-by-sa-4.0May 2020View details →
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

InVID Fake Video Corpus v1.0

<p>The InVID TV Fake Video Corpus is a small collection of verified fake videos. It was developed in the context of the InVID project with the aim of gaining a perspective of the types of fake video that can be encountered in the real world.</p> <p>Currently the Corpus consists of 59 videos. For each video, information is provided describing the fake, its original source, and the evidence proving it is a fake. As we do not own the videos, the dataset only provides the video URLs and metadata, in the form of a tab-separated value (TSV) file. See the README file for more information.</p>

opencc-by-4.0Jan 2017View details →
zenodo48/100

Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)

<h3>&Uuml;berblick</h3> <p>Das <strong>Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)</strong> ist einer der gr&ouml;&szlig;ten, frei verf&uuml;gbaren Datens&auml;tze von Plenarprotokollen des Deutschen Bundestages. Er ist eine Zusammenstellung aller Plenarprotokolle von der 1. Wahlperiode bis zur aktuellsten 21. Wahlperiode die im XML-Format auf dem <a href="https://www.bundestag.de/services/opendata">Open Data Portal</a> des Deutschen Bundestages und dem <a href="https://dip.bundestag.de/">Dokumentations- und Informationssystem f&uuml;r parlamentarische Materialien (DIP)</a> bis zum jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die TXT-Sammlung. Parquet-Dateien sind f&uuml;r Big Data-Anwendungen verf&uuml;gbar.</p> <p>Der CPP-BT ist der <strong>Zwillings-Korpus</strong> des<strong> </strong><a href="https://doi.org/10.5281/zenodo.4643065"><strong>Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT)</strong></a><strong>.&nbsp;</strong> Durch die Verbindung beider Korpora k&ouml;nnen Sie Plenarprotokolle und Drucksachen &mdash; und damit alle Vorg&auml;nge des Bundestages &mdash; in einheitlichen Analysen untersuchen.</p> <p>&nbsp;</p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird mehrmals pro Wahlperiode aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <h3>NEU in Version 2025-05-24</h3> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten (bis einschlie&szlig;lich aktuellste Wahlperiode)</li> <li>Neukonzeptionierung des Datensatzes als deklarative {targets} Pipeline</li> <li><strong>Wichtige &Auml;nderung: </strong>Variable "nummer_original" zu "protokoll_nr" umbenannt</li> <li><strong>Wichtige &Auml;nderung:</strong> Variable "datum" zu "sitzung_datum" umbenannt</li> <li>Neues Feature: Alle Einzelreden des Bundestages in tabellarischem Format mit vielen neuen Metadaten verf&uuml;gbar (ab 18. Wahlperiode)</li> <li>Neues Feature: Datensatz im Parquet-Format verf&uuml;gbar</li> <li>Neues Feature: Zus&auml;tzlicher Bericht zur Qualit&auml;tskontrolle</li> <li>Inhaltiche Erweiterung und Verbesserung der TXT-Variante</li> <li>Viele zus&auml;tzliche Tests zur Qualit&auml;tspr&uuml;fung</li> <li>Pipeline ruft automatisch die tagesaktuell neuesten Bundestagsprotokolle ab (API Key notwendig)</li> <li>Pipeline speichert viele Checkpoints und kann jederzeit unterbrochen und fortgesetzt werden</li> <li>Delta Updates m&ouml;glich</li> <li>Grundlegende &Uuml;berarbeitung des Codebooks</li> </ul> <p>&nbsp;</p> <h3>Features</h3> <ul> <li>Insgesamt bis zu 35 Variablen in der CSV-Variante</li> <li>Plenarprotokolle von der 1. Wahlperiode bis zur neuesten Wahlperiode am Stichtag</li> <li>Aufteilung in Einzelreden&nbsp;u.a. mit ID, Name, Fraktion und Amt der Redner:in (ab 18. Wahlperiode)</li> <li>Aufteilung in Protokollbestandteile: Inhaltsverzeichnis, Sitzungsverlauf, Anlagen, Rednerliste (ab 18. Wahlperiode)</li> <li>Fortlaufende Aktualisierung (Datensatz kann zus&auml;tzlich via Pipeline t&auml;glich aktualisiert werden)</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV, XML, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report, um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'ANALYSE')</li> <li>Diagramme liegen jeweils in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor</li> <li>Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4542661">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <h3>&nbsp;</h3> <h3>Eckdaten</h3> <p><em>Stichtag:</em> 24. Mai 2025</p> <p><em>Inhaltlicher Umfang</em>: 4566 Plenarprotokolle / ~362 Millionen Tokens</p> <p><em>Zeitlicher Umfang:</em> 1949 bis 2025</p> <p><em>Wahlperioden:</em> 1. bis 21. Wahlperiode</p> <p><em>Formate:</em><strong> </strong>CSV, TXT, XML und Parquet</p> <h3>&nbsp;</h3> <h3>Source Code und Compilation Report</h3> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Qualit&auml;tskontrollen auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Source Code f&uuml;r die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt:&nbsp;<a href="https://doi.org/10.5281/zenodo.4542661">https://doi.org/10.5281/zenodo.4542661</a></p> <h3>&nbsp;</h3> <h3>Kryptographische Signaturen</h3> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <h3>&nbsp;</h3> <h3>Kein Urheberrecht: Public Domain</h3> <p>An den Plenarprotokollen besteht gem. &sect; 5 Abs. 2 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <h3>Disclaimer</h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zum Deutschen Bundestag oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <h3>Alternativen</h3> <ul> <li>Blaette, A., &amp; Leonhardt, C. (2024). GermaParl Corpus of Plenary Protocols (v2.1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12794676">https://doi.org/10.5281/zenodo.12794676</a></li> <li>Richter, F., Koch, P., Franke, O., Kraus, J., Kuruc, F., Thiem, A., H&ouml;gerl, J., Heine, S., &amp; Sch&ouml;ps, K. (2020). Open Discourse (Version V4) [dataset]. Harvard Dataverse. <a href="https://doi.org/doi:10.7910/DVN/FIKIBO">https://doi.org/doi:10.7910/DVN/FIKIBO</a></li> <li>Rauh, Christian; Schwalbach, 2020, "The ParlSpeech V2 data set: Full-text corpora of 6.3 million parliamentary speeches in the key legislative chambers of nine representative democracies", <a href="https://doi.org/10.7910/DVN/L4OAKN">https://doi.org/10.7910/DVN/L4OAKN</a>, Harvard Dataverse, V1</li> <li>Open Knowledge Foundation, "Offenes Parlament", <a href="https://offenesparlament.de/daten/">https://offenesparlament.de/daten/</a></li> </ul> <p>&nbsp;</p> <h3>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</h3> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <h3>Kontakt</h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im Issue Tracker auf Codeberg oder kontaktieren Sie mich &uuml;ber <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>

opencc-zeroFeb 2021View details →
zenodo48/100

Corpus der Entscheidungen des Bundespatentgerichts (CE-BPatG)

<h3>&Uuml;berblick</h3> <p>Das <strong>Corpus der Entscheidungen des Bundespatentgerichts (CE-BPatG)</strong> ist der bislang gr&ouml;&szlig;te, frei verf&uuml;gbare Datensatz von Entscheidungen des Bundespatentgerichts. Er ist eine Zusammenstellung aller Entscheidungen die in der <a href="https://www.bundespatentgericht.de/">amtlichen Datenbank des Bundespatentgerichts</a> am jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die PDF-Sammlung.</p> <p>F&uuml;r Praktiker:innen stelle ich zus&auml;tzlich nach Senat sortierte PDF-Sammlungen aller <em>Leitsatzentscheidungen</em> zur Verf&uuml;gung.</p> <p>&nbsp;</p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird <em>1-2 mal im Jahr</em> aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <h3>NEU in Version 2025-07-08</h3> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten</li> <li>NEU: Datensatz im Parquet-Format</li> <li>Expliziter R Package Version Lock f&uuml;r 2024-06-13 (CRAN Date)</li> <li>&Uuml;berarbeitung des Dockerfiles</li> <li>Vereinfachung der Run-Skripte und st&auml;rkere Integration mit Docker Compose</li> <li>Vereinheitlichung der Extraktion von PDF-Dateien, der Berechnung linguistischer Kennzahlen und der Berechnung kryptographischer Hashes</li> <li>&Uuml;berarbeitung der Dokumentation zu Varianten des Datensatzes</li> <li>Entfernung von exakten Prozentzahlen in den Frequenztabellen</li> <li>Entfernung der Tesseract System Library</li> <li>Entfernung der Nummerierung der Diagramme</li> </ul> <h3>&nbsp;</h3> <h3><strong>Features</strong></h3> <ul> <li>Insgesamt bis zu 32 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'Analyse')</li> <li>Jedes Diagramm liegt in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor. Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="../doi/10.5281/zenodo.6667305">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <p>&nbsp;</p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em> 8. Juli 2025</p> <p><em>Inhaltlicher Umfang</em>: 31.203 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 2000 bis 2025</p> <p><em>Formate:</em><strong> </strong>PDF, TXT, CSV und Parquet</p> <p>&nbsp;</p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist ab Version 2022-07-12 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Robustness Checks auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Source Code f&uuml;r die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="../doi/10.5281/zenodo.6667305">https://zenodo.org/doi/10.5281/zenodo.6667305</a></p> <p>&nbsp;</p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p>&nbsp;</p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Entscheidungstexten und amtlichen Leits&auml;tzen besteht gem. &sect; 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Beh&ouml;rden, Gerichten oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <p><strong>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Kommentieren Sie gerne im Issue Tracker auf Codeberg oder kontaktieren Sie mich via <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>&nbsp;</p>

opencc-zeroJul 2020View details →
zenodo48/100

Corpus der Entscheidungen des Bundesverfassungsgerichts (CE-BVerfG)

<p><strong>&Uuml;berblick</strong></p> <p>Das <strong>Corpus der Entscheidungen des Bundesverfassungsgerichts (CE-BVerfG)</strong> ist der bislang gr&ouml;&szlig;te, frei verf&uuml;gbare Datensatz von Entscheidungen des Bundesverfassungsgerichts. Er ist eine Zusammenstellung aller Entscheidungen die auf der <a href="https://www.bundesverfassungsgericht.de">amtlichen Webseite des Bundesverfassungsgerichts</a> am jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die PDF-Sammlung.</p> <p>Alle die <em>Corona-Pandemie</em> betreffenden Entscheidungen des Bundesverfassungsgerichts finden Sie zus&auml;tzlich separat dokumentiert und analysiert im Datensatz <strong><a href="https://doi.org/10.5281/zenodo.4459405">Corona-Rechtsprechung des Bundesverfassungsgerichts (BVerfG-Corona)</a></strong>.</p> <p>Das <em>CE-BVerfG</em> sollte nicht mit dem <a href="http://doi.org/10.5281/zenodo.3831111"><strong>Corpus der amtlichen Entscheidungssammlung des Bundesverfassungsgerichts (C-BVerfGE)</strong></a> verwechselt werden. Letzterer zielt nur auf eine Abbildung der amtlichen Sammlung ab und ist deutlich kleiner.</p> <p>&nbsp;</p> <p><strong>Aktualisierung</strong></p> <p>Dieser Datensatz wird <em>1-2 mal im Jahr </em>aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <p><strong>NEU in Version 2025-09-22</strong></p> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten</li> <li>Neue Variablen: Tenor und Beschlussformel</li> <li>Neues Feature: Datensatz im Parquet-Format verf&uuml;gbar</li> <li>Variable entfernt: Kurzbeschreibung (wird vom BVerfG nicht mehr angeboten)</li> <li>Amtliche Sammlung bis inklusive Band 169 mit Name, Band und Seite versehen</li> <li>Anpassung der Sammlung von Entscheidungen an neues Datenbankformat der BVerfG-Webseite</li> <li>&Uuml;berarbeitung der Dokumentation zu den Varianten des Datensatzes</li> <li>Expliziter R Package Version Lock f&uuml;r 2024-06-13 (CRAN Date)</li> <li>Parallelisierung von Quanteda repariert</li> <li>&Uuml;berarbeitung des Dockerfiles</li> <li>Zus&auml;tzliche Tests auf Einzigartigkeit von ECLI und doc_id</li> <li>Definition von Inhalt des Source Code-Archivs via git</li> <li>/tmp in Arbeitsspeicher ausgelagert</li> <li>Fix f&uuml;r Metadaten-Extraktion</li> </ul> <p>&nbsp;</p> <p><strong>Features</strong></p> <ul> <li>Insgesamt bis zu 38 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV, HTML)</li> <li>Entscheidungsnamen und&nbsp;BVerfGE-Fundstelle (nur BVerfGE)</li> <li>Verkn&uuml;pfung mit Pr&auml;sidentIn/Vize-Pr&auml;sidentIn</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'ANALYSE').</li> <li>Jedes Diagramm liegt in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor. Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4308216">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <p>&nbsp;</p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em>&nbsp;22. September 2025</p> <p><em>Inhaltlicher Umfang</em>: 9.302 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 1998 bis 2025, plus vereinzelte Entscheidungen aus anderen Jahren</p> <p><em>Zitationsnetzwerk:&nbsp;</em>ca. 170.000 Zitate zwischen ca. 14.500 Entscheidungen</p> <p><em>Formate:</em><strong> </strong>PDF, TXT, CSV, HTML und GraphML</p> <p>&nbsp;</p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist ab Version 2021-01-08 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Robustness Checks auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt die gesamte Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="../doi/10.5281/zenodo.3902658">https://doi.org/10.5281/zenodo.4308216</a></p> <p>&nbsp;</p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p>&nbsp;</p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Entscheidungstexten und amtlichen Leits&auml;tzen besteht gem. &sect; 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht weder mit dem Bundesverfassungsgericht noch mit den Herausgebern der BVerfGE in Verbindung.</p> <p>&nbsp;</p> <p><strong>&Auml;hnliche Datens&auml;tze zum BVerfG<br></strong></p> <p>Wendel, L., &amp; M&ouml;llers, C. (2023). Korpus der Entscheidungen des Bundesverfassungsgerichts (2.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10369205">https://doi.org/10.5281/zenodo.10369205</a></p> <p><em>[Senatsentscheidungen 1972-2010] Engst, Benjamin G., Thomas Gschwend, Christoph H&ouml;nnige &amp; Caroline E. Wittig. (2020). &ldquo;The Constitutional Court Database." &nbsp;<a href="https://ccdb.eu/">https://ccdb.eu/</a></em></p> <p><em>[BVerfGE von 1951 bis 2015]</em>&nbsp; Coupette, C. (2019). Juristische Netzwerkforschung: Modellierung, Quantifizierung und Visualisierung relationaler Daten im Recht (Online-Appendix). Zenodo. <a href="https://doi.org/10.1628/978-3-16-157012-4-appendix">https://doi.org/10.1628/978-3-16-157012-4-appendix</a></p> <p>&nbsp;</p> <p><strong>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Notieren Sie diese bitte im Issue Tracker auf Codeberg oder schreiben Sie mir via <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>

opencc-zeroJun 2020View details →
zenodo48/100

The Distant Listening Corpus

<p dir="auto">This is a README file for a data repository originating from the&nbsp;<a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and serves as welcome page for both</p> <ul> <li>the GitHub repo <a href="https://github.com/DCMLab/distant_listening_corpus">https://github.com/DCMLab/distant_listening_corpus</a> and the corresponding</li> <li>documentation page <a href="https://dcmlab.github.io/distant_listening_corpus" rel="nofollow">https://dcmlab.github.io/distant_listening_corpus</a></li> </ul> <p dir="auto">For information on how to obtain and use the dataset, please refer to <a href="https://dcmlab.github.io/distant_listening_corpus/introduction" rel="nofollow">this documentation page</a>.</p> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#the-distant-listening-corpus-a-corpus-of-annotated-scores">The Distant Listening Corpus (A corpus of annotated scores)</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#getting-the-data">Getting the data</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#data-formats">Data Formats</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#opening-scores">Opening Scores</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#opening-tsv-files-in-a-spreadsheet">Opening TSV files in a spreadsheet</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#loading-tsv-files-in-python">Loading TSV files in Python</a></li> </ul> </li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#version-history">Version history</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#questions-suggestions-corrections-bug-reports">Questions, Suggestions, Corrections, Bug Reports</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#cite-as">Cite as</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#license">License</a></li> </ul> </li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#overview">Overview</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#abc">ABC</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bach_en_fr_suites">bach_en_fr_suites</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bach_solo">bach_solo</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bartok_bagatelles">bartok_bagatelles</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#beethoven_piano_sonatas">beethoven_piano_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#c_schumann_lieder">c_schumann_lieder</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#chopin_mazurkas">chopin_mazurkas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#corelli">corelli</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#couperin_clavecin">couperin_clavecin</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#couperin_concerts">couperin_concerts</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#cpe_bach_keyboard">cpe_bach_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#debussy_suite_bergamasque">debussy_suite_bergamasque</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#dvorak_silhouettes">dvorak_silhouettes</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#frescobaldi_fiori_musicali">frescobaldi_fiori_musicali</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#grieg_lyric_pieces">grieg_lyric_pieces</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#handel_keyboard">handel_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#jc_bach_sonatas">jc_bach_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#kleine_geistliche_konzerte">kleine_geistliche_konzerte</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#kozeluh_sonatas">kozeluh_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#liszt_pelerinage">liszt_pelerinage</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mahler_kindertotenlieder">mahler_kindertotenlieder</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#medtner_tales">medtner_tales</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mendelssohn_quartets">mendelssohn_quartets</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#monteverdi_madrigals">monteverdi_madrigals</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mozart_piano_sonatas">mozart_piano_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#pergolesi_stabat_mater">pergolesi_stabat_mater</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#peri_euridice">peri_euridice</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#pleyel_quartets">pleyel_quartets</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#poulenc_mouvements_perpetuels">poulenc_mouvements_perpetuels</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#rachmaninoff_piano">rachmaninoff_piano</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#ravel_piano">ravel_piano</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#scarlatti_sonatas">scarlatti_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schubert_winterreise">schubert_winterreise</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schulhoff_suite_dansante_en_jazz">schulhoff_suite_dansante_en_jazz</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schumann_kinderszenen">schumann_kinderszenen</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schumann_liederkreis">schumann_liederkreis</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#sweelinck_keyboard">sweelinck_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#tchaikovsky_seasons">tchaikovsky_seasons</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#wagner_overtures">wagner_overtures</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#wf_bach_sonatas">wf_bach_sonatas</a></li> </ul> </li> </ul> <div dir="auto"> <h1>The Distant Listening Corpus (A corpus of annotated scores)</h1> <a href="https://github.com/DCMLab/distant_listening_corpus/#the-distant-listening-corpus-a-corpus-of-annotated-scores"></a></div> <p dir="auto"><em>A modular infrastructure for the empirical study of (an)notated music</em></p> <p dir="auto">This corpus has been created within the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and employs the <a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p dir="auto">The publication covers the following public corpora (the DOI links always point at the latest version respectively):</p> <ul> <li>J.S. Bach &ndash; English and French Suites [<a href="https://doi.org/10.5281/zenodo.14996489" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bach_en_fr_suites">repo</a>][<a href="https://github.com/DCMLab/bach_en_fr_suites/archive/refs/heads/main.zip">ZIP</a>]</li> <li>J.S. Bach &ndash; Solo Pieces (A corpus of annotated scores) [<a href="https://doi.org/10.5281/zenodo.14996765" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bach_solo">repo</a>][<a href="https://github.com/DCMLab/bach_solo/archive/refs/heads/main.zip">ZIP</a>]</li> <li>B&eacute;la Bart&oacute;k &ndash; 14 Bagatelles, Op. 6 [<a href="https://doi.org/10.5281/zenodo.14996945" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bartok_bagatelles">repo</a>][<a href="https://github.com/DCMLab/bartok_bagatelles/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Fran&ccedil;ois Couperin &ndash; L'art de toucher le clavecin [<a href="https://doi.org/10.5281/zenodo.14984598" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/couperin_clavecin">repo</a>][<a href="https://github.com/DCMLab/couperin_clavecin/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Clara Schumann &ndash; Lieder [<a href="https://doi.org/10.5281/zenodo.14996952" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/c_schumann_lieder">repo</a>][<a href="https://github.com/DCMLab/c_schumann_lieder/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Fran&ccedil;ois Couperin &ndash; Concerts Royaux [<a href="https://doi.org/10.5281/zenodo.15027239" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/couperin_concerts">repo</a>][<a href="https://github.com/DCMLab/couperin_concerts/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Carl Philipp Emanuel Bach &ndash; Works for Keyboard [<a href="https://doi.org/10.5281/zenodo.14996326" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/cpe_bach_keyboard">repo</a>][<a href="https://github.com/DCMLab/cpe_bach_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Girolamo Frescobaldi (1583-1643) &ndash; Fiori Musicali, op. 12 (1635) [<a href="https://doi.org/10.5281/zenodo.14984864" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/frescobaldi_fiori_musicali">repo</a>][<a href="https://github.com/DCMLab/frescobaldi_fiori_musicali/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Georg Friedrich H&auml;ndel &ndash; Grobschmied Variations (The Harmonious Blacksmith), HWV 430 [<a href="https://doi.org/10.5281/zenodo.14996996" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/handel_keyboard">repo</a>][<a href="https://github.com/DCMLab/handel_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>J.C. Bach &ndash; Keyboard Sonatas [<a href="https://doi.org/10.5281/zenodo.14996292" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/jc_bach_sonatas">repo</a>][<a href="https://github.com/DCMLab/jc_bach_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Heinrich Sch&uuml;tz &ndash; Kleine Geistliche Konzerte [<a href="https://doi.org/10.5281/zenodo.14997003" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/kleine_geistliche_konzerte">repo</a>][<a href="https://github.com/DCMLab/kleine_geistliche_konzerte/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Leopold Koželuch &ndash; Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.14997015" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/kozeluh_sonatas">repo</a>][<a href="https://github.com/DCMLab/kozeluh_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Gustav Mahler &ndash; Kindertotenlieder [<a href="https://doi.org/10.5281/zenodo.14997022" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mahler_kindertotenlieder">repo</a>][<a href="https://github.com/DCMLab/mahler_kindertotenlieder/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Felix Mendelssohn &ndash; String Quartets [<a href="https://doi.org/10.5281/zenodo.14996150" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mendelssohn_quartets">repo</a>][<a href="https://github.com/DCMLab/mendelssohn_quartets/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Claudio Monteverdi &ndash; Madrigals [<a href="https://doi.org/10.5281/zenodo.15003026" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/monteverdi_madrigals">repo</a>][<a href="https://github.com/DCMLab/monteverdi_madrigals/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Giovanni Battista Pergolesi &ndash; Stabat Mater (1736) [<a href="https://doi.org/10.5281/zenodo.14990099" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/pergolesi_stabat_mater">repo</a>][<a href="https://github.com/DCMLab/pergolesi_stabat_mater/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Jacopo Peri &ndash; Euridice (1600) [<a href="https://doi.org/10.5281/zenodo.14996445" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/peri_euridice">repo</a>][<a href="https://github.com/DCMLab/peri_euridice/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Ignaz Pleyel &ndash; String Quartets [<a href="https://doi.org/10.5281/zenodo.14997048" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/pleyel_quartets">repo</a>][<a href="https://github.com/DCMLab/pleyel_quartets/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Francis Poulenc &ndash; Mouvements Perpetuels [<a href="https://doi.org/10.5281/zenodo.14997053" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/poulenc_mouvements_perpetuels">repo</a>][<a href="https://github.com/DCMLab/poulenc_mouvements_perpetuels/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Sergei Rachmaninoff &ndash; Piano Pieces [<a href="https://doi.org/10.5281/zenodo.14984155" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/rachmaninoff_piano">repo</a>][<a href="https://github.com/DCMLab/rachmaninoff_piano/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Maurice Ravel &ndash; Piano Pieces [<a href="https://doi.org/10.5281/zenodo.14997064" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/ravel_piano">repo</a>][<a href="https://github.com/DCMLab/ravel_piano/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Domenico Scarlatti &ndash; Keyboard Sonatas [<a href="https://doi.org/10.5281/zenodo.14992884" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/scarlatti_sonatas">repo</a>][<a href="https://github.com/DCMLab/scarlatti_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Franz Schubert &ndash; Winterreise [<a href="https://doi.org/10.5281/zenodo.14997095" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schubert_winterreise">repo</a>][<a href="https://github.com/DCMLab/schubert_winterreise/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Erwin Schulhoff &ndash; Suite dansante en jazz [<a href="https://doi.org/10.5281/zenodo.14997098" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schulhoff_suite_dansante_en_jazz">repo</a>][<a href="https://github.com/DCMLab/schulhoff_suite_dansante_en_jazz/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Robert Schumann &ndash; Liederkreis [<a href="https://doi.org/10.5281/zenodo.14997104" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schumann_liederkreis">repo</a>][<a href="https://github.com/DCMLab/schumann_liederkreis/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Jan Sweelinck &ndash; Organ Pieces [<a href="https://doi.org/10.5281/zenodo.14997111" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/sweelinck_keyboard">repo</a>][<a href="https://github.com/DCMLab/sweelinck_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Richard Wagner &ndash; Overtures [<a href="https://doi.org/10.5281/zenodo.14997120" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/wagner_overtures">repo</a>][<a href="https://github.com/DCMLab/wagner_overtures/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Wilhelm Friedemann Bach &ndash; Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.14997133" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/wf_bach_sonatas">repo</a>][<a href="https://github.com/DCMLab/wf_bach_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Rammos, Y., Neuwirth, M., Moss, F. C., &amp; Rohrmeier, M. (2024). An annotated corpus of tonal piano music from the long 19th century. Empirical Musicology Review, 18(1), 84&ndash;95. <a href="https://doi.org/10.18061/emr.v18i1.8903" rel="nofollow">https://doi.org/10.18061/emr.v18i1.8903</a></em></p> <ul> <li>Ludwig van Beethoven - Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.7473560" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/beethoven_piano_sonatas">repo</a>][<a href="https://github.com/DCMLab/beethoven_piano_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Fr&eacute;d&eacute;ric Chopin - Mazurkas [<a href="https://doi.org/10.5281/zenodo.7473566" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/chopin_mazurkas">repo</a>][<a href="https://github.com/DCMLab/chopin_mazurkas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Claude Debussy - Suite Bergamasque [<a href="https://doi.org/10.5281/zenodo.7473568" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/debussy_suite_bergamasque">repo</a>][<a href="https://github.com/DCMLab/debussy_suite_bergamasque/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Anton&iacute;n Dvoř&aacute;k - Silhouettes [<a href="https://doi.org/10.5281/zenodo.7473576" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/dvorak_silhouettes">repo</a>][<a href="https://github.com/DCMLab/dvorak_silhouettes/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Edvard Grieg - Lyric Pieces [<a href="https://doi.org/10.5281/zenodo.7473578" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/grieg_lyric_pieces">repo</a>][<a href="https://github.com/DCMLab/grieg_lyric_pieces/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Franz Liszt - Ann&eacute;es de P&egrave;lerinage [<a href="https://doi.org/10.5281/zenodo.7473580" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/liszt_pelerinage">repo</a>][<a href="https://github.com/DCMLab/liszt_pelerinage/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Nikolai Medtner - Tales [<a href="https://doi.org/10.5281/zenodo.7473528" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/medtner_tales">repo</a>][<a href="https://github.com/DCMLab/medtner_tales/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Robert Schumann - Kinderszenen [<a href="https://doi.org/10.5281/zenodo.7473582" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schumann_kinderszenen">repo</a>][<a href="https://github.com/DCMLab/schumann_kinderszenen/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Pyotr Tchaikovsky - The Seasons [<a href="https://doi.org/10.5281/zenodo.7473586" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/tchaikovsky_seasons">repo</a>][<a href="https://github.com/DCMLab/tchaikovsky_seasons/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Moss, F. C., Neuwirth, M., &amp; Rohrmeier, M. A. (2021). A semi-automated workflow paradigm for the distributed creation and curation of expert annotations. Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR, 262&ndash;269. <a href="https://doi.org/10.5281/ZENODO.5624417" rel="nofollow">https://doi.org/10.5281/ZENODO.5624417</a></em></p> <ul> <li>Arcangelo Corelli &ndash; Trio Sonatas [<a href="https://zenodo.org/doi/10.5281/zenodo.7504011" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/corelli">repo</a>][<a href="https://github.com/DCMLab/corelli/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Neuwirth, M., &amp; Rohrmeier, M. (2021). The Annotated Mozart Sonatas: Score, harmony, and cadence. Transactions of the International Society for Music Information Retrieval, 4(1), 67&ndash;80. <a href="https://doi.org/10.5334/tismir.63" rel="nofollow">https://doi.org/10.5334/tismir.63</a></em></p> <ul> <li>Wolfgang Amadeus Mozart - Piano Sonatas [<a href="https://zenodo.org/doi/10.5281/zenodo.7424962" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mozart_piano_sonatas">repo</a>][<a href="https://github.com/DCMLab/mozart_piano_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Neuwirth, M., Harasim, D., Moss, F. C., &amp; Rohrmeier, M. (2018). The Annotated Beethoven Corpus (ABC): A Dataset of Harmonic Analyses of All Beethoven String Quartets. Frontiers in Digital Humanities, 5(July), 1&ndash;5. <a href="https://doi.org/10.3389/fdigh.2018.00016" rel="nofollow">https://doi.org/10.3389/fdigh.2018.00016</a></em></p> <ul> <li>Ludwig van Beethoven - String Quartets [<a href="https://zenodo.org/doi/10.5281/zenodo.7441343" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/ABC">repo</a>][<a href="https://github.com/DCMLab/ABC/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <div dir="auto"> <h2>Getting the data</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#getting-the-data"></a></div> <ul> <li>download individual subcorpora as ZIP files using the URLs provided above</li> <li>download a <a href="https://specs.frictionlessdata.io/data-package/" rel="nofollow">Frictionless Datapackage</a> that includes concatenations of the TSV files in the four folders (<code>measures</code>, <code>notes</code>, <code>chords</code>, and <code>harmonies</code>) and a JSON descriptor: <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/releases/latest/download/distant_listening_corpus.zip">distant_listening_corpus.zip</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/releases/latest/download/distant_listening_corpus.datapackage.json">distant_listening_corpus.datapackage.json</a></li> </ul> </li> <li>clone the repo (~2.4 GB): <code>git clone --recursive -j12 https://github.com/DCMLab/distant_listening_corpus.git</code></li> </ul> <div dir="auto"> <h2>Data Formats</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#data-formats"></a></div> <p dir="auto">Each piece in this corpus is represented by five files with identical name prefixes, each in its own folder. For example, the <em>Pr&eacute;lude</em> of J.S. Bach&rsquo;s first English Suite, BWV 806, has the following files:</p> <ul> <li><code>MS3/BWV806_01_Prelude.mscx</code>: Uncompressed MuseScore 3.6.2 file including the music and annotation labels.</li> <li><code>notes/BWV806_01_Prelude.notes.tsv</code>: A table of all note heads contained in the score and their relevant features (not each of them represents an onset, some are tied together)</li> <li><code>measures/BWV806_01_Prelude.measures.tsv</code>: A table with relevant information about the measures in the score.</li> <li><code>chords/BWV806_01_Prelude.chords.tsv</code>: A table containing layer-wise unique onset positions with the musical markup (such as dynamics, articulation, lyrics, figured bass, etc.).</li> <li><code>harmonies/BWV806_01_Prelude.harmonies.tsv</code>: A table of the included harmony labels (including cadences and phrases) with their positions in the score.</li> </ul> <p dir="auto">Each TSV file comes with its own JSON descriptor that describes the meanings and datatypes of the columns ("fields") it contains, follows the <a href="https://specs.frictionlessdata.io/tabular-data-resource/" rel="nofollow">Frictionless specification</a>, and can be used to validate and correctly load the described file.</p> <div dir="auto"> <h3>Opening Scores</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#opening-scores"></a></div> <p dir="auto">After navigating to your local copy, you can open the scores in the folder <code>MS3</code> with the free and open source score editor <a href="https://musescore.org" rel="nofollow">MuseScore</a>. Please note that the scores have been edited, annotated and tested with <a href="https://github.com/musescore/MuseScore/releases/tag/v3.6.2">MuseScore 3.6.2</a>. MuseScore 4 has since been released which renders them correctly but cannot store them back in the same format.</p> <div dir="auto"> <h3>Opening TSV files in a spreadsheet</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#opening-tsv-files-in-a-spreadsheet"></a></div> <p dir="auto">Tab-separated value (TSV) files are like Comma-separated value (CSV) files and can be opened with most modern text editors. However, for correctly displaying the columns, you might want to use a spreadsheet or an addon for your favourite text editor. When you use a spreadsheet such as Excel, it might annoy you by interpreting fractions as dates. This can be circumvented by using <code>Data --&gt; From Text/CSV</code> or the free alternative <a href="https://www.libreoffice.org/download/download/" rel="nofollow">LibreOffice Calc</a>. Other than that, TSV data can be loaded with every modern programming language.</p> <div dir="auto"> <h3>Loading TSV files in Python</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#loading-tsv-files-in-python"></a></div> <p dir="auto">Since the TSV files contain null values, lists, fractions, and numbers that are to be treated as strings, you may want to use this code to load any TSV files related to this repository (provided you're doing it in Python). After a quick <code>pip install -U ms3</code> (requires Python 3.10 or later) you'll be able to load any TSV like this:</p> <div dir="auto"> <pre><span>import</span> <span>ms3</span> <span>labels</span> <span>=</span> <span>ms3</span>.<span>load_tsv</span>(<span>"harmonies/BWV806_01_Prelude.harmonies.tsv"</span>) <span>notes</span> <span>=</span> <span>ms3</span>.<span>load_tsv</span>(<span>"notes/BWV806_01_Prelude.notes.tsv"</span>)</pre> <div>&nbsp;</div> </div> <div dir="auto"> <h2>Version history</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#version-history"></a></div> <p dir="auto">See the <a href="https://github.com/DCMLab/distant_listening_corpus/releases">GitHub releases</a>.</p> <div dir="auto"> <h2>Questions, Suggestions, Corrections, Bug Reports</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#questions-suggestions-corrections-bug-reports"></a></div> <p>Please <a href="https://github.com/DCMLab/distant_listening_corpus/issues">create an issue</a> and/or feel free to fork and submit pull requests.</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Données supplémentaires: Repérage automatisé de l'hyponymie dans des corpus spécialisés en français à l'aide de Sketch Engine

<p>Ces figures sont des donn&eacute;es suppl&eacute;mentaires de l&#39;article suivant :<br> San Mart&iacute;n A., Trekker C., et Le&oacute;n-Ara&uacute;z P. 2022. Rep&eacute;rage automatis&eacute; de l&rsquo;hyponymie dans des corpus sp&eacute;cialis&eacute;s en fran&ccedil;ais &agrave; l&rsquo;aide de Sketch Engine. <em>Terminology</em>. doi:&nbsp;10.1075/term.20044.san</p> <p>Les figures suivantes repr&eacute;sentent le r&eacute;sultat complet de l&rsquo;&eacute;valuation des WS. La premi&egrave;re colonne repr&eacute;sente le terme &eacute;valu&eacute; (c&rsquo;est-&agrave;-dire les termes de recherche) et les trois colonnes suivantes, les trois premiers r&eacute;sultats. Enfin, les colonnes suivantes repr&eacute;sentent visuellement la pr&eacute;cision de chaque paire, le chiffre &agrave; gauche &eacute;tant le nombre de vrais positifs et celui &agrave; droite, le nombre de correspondances associ&eacute;es &agrave; la paire. La couleur bleue repr&eacute;sente les r&eacute;sultats de la colonne <em>X est le g&eacute;n&eacute;rique de...</em> et la couleur jaune, les r&eacute;sultats de la colonne <em>X est un type de...</em></p> <ul> <li>psychologie.tif:&nbsp;&Eacute;valuation des WS du sous-corpus de psychologie</li> <li>chimie.tif:&nbsp;&Eacute;valuation des WS du sous-corpus de chimie</li> <li>droit.tif:&nbsp;&Eacute;valuation des WS du sous-corpus de droit</li> <li>informatique.tif: &Eacute;valuation des WS du sous-corpus d&rsquo;informatique</li> <li>geographie.tif:&nbsp;&Eacute;valuation des WS du sous-corpus de g&eacute;ographie</li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo48/100

SRBCorp: Corpus of Parliamentary Debates in Serbia

<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the National Assembly of Serbia. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1997-2020 (eight terms) and counts over 300 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): SRBCorp: Corpus of Parliamentary Debates in Serbia (v1.1.1), https://doi.org/10.5281/zenodo.6521648.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - added a new variable for policy category &quot;tag&quot; using ML model trained on known tags of agenda points in the parliaments of Croatia and Bosnia-Herzegovina</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2389); migrated to ZENODO due to limitations concerning the concept DOI</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

CROCorp: Corpus of Parliamentary Debates in Croatia

<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Croatian Parliament (Sabor). The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 2003-2020 (five complete terms) and counts over 500 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): CROCorp: Corpus of Parliamentary Debates in Croatia (v1.1.1), https://doi.org/10.5281/zenodo.6521372.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - improved coding of dummy variable &quot;moderator&quot; (using less error-prone alghoritm for detecting the modertor role)<br> - fixed issue with agenda points which are conncatenated while preserving a unique web link<br> - recoded agenda points tags using better ML model (transformer architecture)</p> <p>v1.0.0<br> - originally posted on GESIS repository (migrated to ZENODO due to limitations concerning the concept DOI)</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina

<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Parliamentary Assembly of Bosnia and Herzegovina. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1998-2018 (six complete terms) and counts over 127 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, Christophe Lesschaeve, and Ensar Muharemović (2022): BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina (v1.1.1),<br> https://doi.org/10.5281/zenodo.6517697.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - fixed a typo in one of the debates&#39; date<br> - fixed minor inconsistencies in the tag column</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2387); migrated to ZENODO due to limitations concerning the concept DOI</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language

<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record