Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
878
datasets available to search
ShareScore release 0.7.1
Dataset results
878 results for “corpus”
S1000 corpus, large-scale tagging results and other supplementary files
<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>
Corpus of critical citations contexts
<p>We present here a corpus of 505 critical citation contexts, i.e. a set of sentences or propositions that contain at least one citation of a study towards which the author(s) has/have a negative opinion. Those contexts come from other existing annotated corpora, from our readings about critical citation and disagreement in science, and from contexts manually annotated by native speakers of English. We have re-annotated all those contexts in order to be sure that they match our definition of critical citations. This corpus can be helpful to train tools dedicated to the automatic retrieval of critical citations. English (2024-02-20)</p>
List of TEI rolename annotations in the ISicily EpiDoc corpus
<p>This CSV file details every instance of a 'roleName' tag in the I.Sicily (sicily.classics.ox.ac.uk) EpiDoc TEI files, reporting the ID number of the file in which it appears, and the value of the @type and @subtype attributes in each case - as such it serves as an index of roleName attestations in the I.Sicily dataset (also recoverable directly from the EpiDoc files). The file will be updated in future.</p>
UNIC Templates for uploading corpus metadata v1.11
<p>The UNIC platform (https://unic.dipintra.it) accepts a JSON file for uploading corpus metadata based on the template here. Alternatively, use the spreadsheet template to input the corpus metadata and convert the resulting .xlsm file to JSON using this application at https://huggingface.co/spaces/nannanliu/UNIC_metadata_conversion. When opening the Excel spreadsheet template, please enable Macros, which will automatically validate your input in the columns. Please do not change the order of the columns because they are embedded with code. To add elements and components not included by the UNIC schema, create new columns after the existing ones.</p>
ZooCor Corpus (fr-it)
<p>ZooCor is a bilingual (fr-it) specialized corpus, consisting of 350 texts concerning marine fauna and conservation biology. The ZooCor corpus covers a time window from 2000 to 2022. The partial metadata related to the subdomain of the threatened species of chelonids (marine turtles) are published. For further information on the entire corpus, please adress to silvia.zollo@uniparthenope.it. </p>
Word Embedding of Amazon Product Review Corpus
<p>A word embedding of the <a href="https://www.cs.uic.edu/~liub/FBS/sentiment-analysis.html#datasets">Amazon Product Review Corpus</a> (<a href="https://www.doi.org/10.1145/1341531.1341560">Jindal and Liu, 2008</a>).</p> <p>Created using <a href="https://code.google.com/archive/p/word2vec/">Word2Vec</a> in CBOW mode, 500 dimensions and window size 5.</p> <p>Words have been lemmatised and particle verbs have been merged into a single token (e.g. <code>calm_down</code>).</p> <ul> </ul> <p> </p> <p><strong>Attribution</strong></p> <p>This dataset was created as part of the following publication:</p> <p>Marc Schulder, Michael Wiegand, Josef Ruppenhofer and Benjamin Roth (2017). <strong>"Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features"</strong>. Proceedings of the 8th International Joint Conference on Natural Language Processing (IJCNLP). Taipei, Taiwan, November 27 - December 3, 2017. <a href="https://doi.org/10.5281/zenodo.3365609">DOI: 10.5281/zenodo.3365609</a>.</p> <p>If you use the data in your research or work, please cite the publication.</p>
BioASQ-QA: A manually curated corpus for Biomedical Question Answering
<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>
Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets
<p>Murreviikko (literally 'Dialect week') is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko ('dialect week') and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, Häme, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni 'my book-my') are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata > pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id's, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union’s Digital Single Market directive (2019/790).</p>
WTC1.1 (WikiTailor corpus v. 1.1)
<p> </p><p><strong>Content:</strong></p><p>List of the 743 domains, their term vocabularies in 10 languages, and the Wikipedia articles associated to each domain extracted by the best model described in:</p><p><i> Cristina España-Bonet, Alberto Barrón-Cedeño and Lluís Màrquez. "Tailoring and Evaluating the Wikipedia for in-Domain Comparable Corpora Extraction." </i>Knowledge and Information Systems, Volume 65, pages 1365-1397. 2023. Springer-Verlag, London Ldt. https://doi.org/10.1007/s10115-022-01767-5</p><p><i> https://github.com/cristinae/WikiTailor</i></p><p> </p><p><strong>Files Description:</strong></p><ul><li>commonCats2015.enesdefrcaareuelrooc.tsv</li></ul><p>Multilingual domains listed one per line, languages are separated by a tab in the order en, es, de, fr, ca, ar, eu, el, ro and oc. For each language we include the pair "ID categoryName" separated by a blank space.</p><ul><li>[LAN].0.tar.bz</li></ul><p>A folder per domain for language [LAN] containing the vocabulary and IDs of the extracted articles by the Wikitailor model 50-WT100.</p><ul><li>extraction[LAN]0.tar.bz</li></ul><p>A folder per domain for language [LAN] containing the text of the extracted articles. The name of the file corresponds to the IDs in [LAN].0.tar.bz.</p><p> </p><p> </p>
Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text
<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>
InVID Fake Video Corpus v1.0
<p>The InVID TV Fake Video Corpus is a small collection of verified fake videos. It was developed in the context of the InVID project with the aim of gaining a perspective of the types of fake video that can be encountered in the real world.</p> <p>Currently the Corpus consists of 59 videos. For each video, information is provided describing the fake, its original source, and the evidence proving it is a fake. As we do not own the videos, the dataset only provides the video URLs and metadata, in the form of a tab-separated value (TSV) file. See the README file for more information.</p>
Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)
<h3>Überblick</h3> <p>Das <strong>Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)</strong> ist einer der größten, frei verfügbaren Datensätze von Plenarprotokollen des Deutschen Bundestages. Er ist eine Zusammenstellung aller Plenarprotokolle von der 1. Wahlperiode bis zur aktuellsten 21. Wahlperiode die im XML-Format auf dem <a href="https://www.bundestag.de/services/opendata">Open Data Portal</a> des Deutschen Bundestages und dem <a href="https://dip.bundestag.de/">Dokumentations- und Informationssystem für parlamentarische Materialien (DIP)</a> bis zum jeweiligen Stichtag veröffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die TXT-Sammlung. Parquet-Dateien sind für Big Data-Anwendungen verfügbar.</p> <p>Der CPP-BT ist der <strong>Zwillings-Korpus</strong> des<strong> </strong><a href="https://doi.org/10.5281/zenodo.4643065"><strong>Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT)</strong></a><strong>. </strong> Durch die Verbindung beider Korpora können Sie Plenarprotokolle und Drucksachen — und damit alle Vorgänge des Bundestages — in einheitlichen Analysen untersuchen.</p> <p> </p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird mehrmals pro Wahlperiode aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p> </p> <h3>NEU in Version 2025-05-24</h3> <ul> <li>Vollständige Aktualisierung der Daten (bis einschließlich aktuellste Wahlperiode)</li> <li>Neukonzeptionierung des Datensatzes als deklarative {targets} Pipeline</li> <li><strong>Wichtige Änderung: </strong>Variable "nummer_original" zu "protokoll_nr" umbenannt</li> <li><strong>Wichtige Änderung:</strong> Variable "datum" zu "sitzung_datum" umbenannt</li> <li>Neues Feature: Alle Einzelreden des Bundestages in tabellarischem Format mit vielen neuen Metadaten verfügbar (ab 18. Wahlperiode)</li> <li>Neues Feature: Datensatz im Parquet-Format verfügbar</li> <li>Neues Feature: Zusätzlicher Bericht zur Qualitätskontrolle</li> <li>Inhaltiche Erweiterung und Verbesserung der TXT-Variante</li> <li>Viele zusätzliche Tests zur Qualitätsprüfung</li> <li>Pipeline ruft automatisch die tagesaktuell neuesten Bundestagsprotokolle ab (API Key notwendig)</li> <li>Pipeline speichert viele Checkpoints und kann jederzeit unterbrochen und fortgesetzt werden</li> <li>Delta Updates möglich</li> <li>Grundlegende Überarbeitung des Codebooks</li> </ul> <p> </p> <h3>Features</h3> <ul> <li>Insgesamt bis zu 35 Variablen in der CSV-Variante</li> <li>Plenarprotokolle von der 1. Wahlperiode bis zur neuesten Wahlperiode am Stichtag</li> <li>Aufteilung in Einzelreden u.a. mit ID, Name, Fraktion und Amt der Redner:in (ab 18. Wahlperiode)</li> <li>Aufteilung in Protokollbestandteile: Inhaltsverzeichnis, Sitzungsverlauf, Anlagen, Rednerliste (ab 18. Wahlperiode)</li> <li>Fortlaufende Aktualisierung (Datensatz kann zusätzlich via Pipeline täglich aktualisiert werden)</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabhängige Formate (PDF, TXT, CSV, XML, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report, um den Erstellungs-Prozess zu erläutern</li> <li>Dutzende Diagramme und Tabellen für alle Zwecke (im ZIP-Archiv 'ANALYSE')</li> <li>Diagramme liegen jeweils in einem für den Druck (PDF) und das Web (PNG) optimierten Format vor</li> <li>Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl für Menschen als auch für Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4542661">Veröffentlichung des Source Codes</a></li> </ul> <h3> </h3> <h3>Eckdaten</h3> <p><em>Stichtag:</em> 24. Mai 2025</p> <p><em>Inhaltlicher Umfang</em>: 4566 Plenarprotokolle / ~362 Millionen Tokens</p> <p><em>Zeitlicher Umfang:</em> 1949 bis 2025</p> <p><em>Wahlperioden:</em> 1. bis 21. Wahlperiode</p> <p><em>Formate:</em><strong> </strong>CSV, TXT, XML und Parquet</p> <h3> </h3> <h3>Source Code und Compilation Report</h3> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Qualitätskontrollen auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enthält den Source Code für die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollständige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>öffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="https://doi.org/10.5281/zenodo.4542661">https://doi.org/10.5281/zenodo.4542661</a></p> <h3> </h3> <h3>Kryptographische Signaturen</h3> <p>Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist.</p> <p>Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner <em>persönlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <h3> </h3> <h3>Kein Urheberrecht: Public Domain</h3> <p>An den Plenarprotokollen besteht gem. § 5 Abs. 2 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollständig urheberrechtsfrei.</p> <p> </p> <h3>Disclaimer</h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zum Deutschen Bundestag oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p> </p> <h3>Alternativen</h3> <ul> <li>Blaette, A., & Leonhardt, C. (2024). GermaParl Corpus of Plenary Protocols (v2.1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12794676">https://doi.org/10.5281/zenodo.12794676</a></li> <li>Richter, F., Koch, P., Franke, O., Kraus, J., Kuruc, F., Thiem, A., Högerl, J., Heine, S., & Schöps, K. (2020). Open Discourse (Version V4) [dataset]. Harvard Dataverse. <a href="https://doi.org/doi:10.7910/DVN/FIKIBO">https://doi.org/doi:10.7910/DVN/FIKIBO</a></li> <li>Rauh, Christian; Schwalbach, 2020, "The ParlSpeech V2 data set: Full-text corpora of 6.3 million parliamentary speeches in the key legislative chambers of nine representative democracies", <a href="https://doi.org/10.7910/DVN/L4OAKN">https://doi.org/10.7910/DVN/L4OAKN</a>, Harvard Dataverse, V1</li> <li>Open Knowledge Foundation, "Offenes Parlament", <a href="https://offenesparlament.de/daten/">https://offenesparlament.de/daten/</a></li> </ul> <p> </p> <h3>Weitere Open Access Veröffentlichungen (Fobbe)</h3> <p>Website<em> </em>—<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data — <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code — <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regulärer Publikationen — <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p> </p> <h3>Kontakt</h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im Issue Tracker auf Codeberg oder kontaktieren Sie mich über <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>
Corpus der Entscheidungen des Bundespatentgerichts (CE-BPatG)
<h3>Überblick</h3> <p>Das <strong>Corpus der Entscheidungen des Bundespatentgerichts (CE-BPatG)</strong> ist der bislang größte, frei verfügbare Datensatz von Entscheidungen des Bundespatentgerichts. Er ist eine Zusammenstellung aller Entscheidungen die in der <a href="https://www.bundespatentgericht.de/">amtlichen Datenbank des Bundespatentgerichts</a> am jeweiligen Stichtag veröffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die PDF-Sammlung.</p> <p>Für Praktiker:innen stelle ich zusätzlich nach Senat sortierte PDF-Sammlungen aller <em>Leitsatzentscheidungen</em> zur Verfügung.</p> <p> </p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird <em>1-2 mal im Jahr</em> aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p> </p> <h3>NEU in Version 2025-07-08</h3> <ul> <li>Vollständige Aktualisierung der Daten</li> <li>NEU: Datensatz im Parquet-Format</li> <li>Expliziter R Package Version Lock für 2024-06-13 (CRAN Date)</li> <li>Überarbeitung des Dockerfiles</li> <li>Vereinfachung der Run-Skripte und stärkere Integration mit Docker Compose</li> <li>Vereinheitlichung der Extraktion von PDF-Dateien, der Berechnung linguistischer Kennzahlen und der Berechnung kryptographischer Hashes</li> <li>Überarbeitung der Dokumentation zu Varianten des Datensatzes</li> <li>Entfernung von exakten Prozentzahlen in den Frequenztabellen</li> <li>Entfernung der Tesseract System Library</li> <li>Entfernung der Nummerierung der Diagramme</li> </ul> <h3> </h3> <h3><strong>Features</strong></h3> <ul> <li>Insgesamt bis zu 32 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabhängige Formate (PDF, TXT, CSV, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erläutern</li> <li>Dutzende Diagramme und Tabellen für alle Zwecke (im ZIP-Archiv 'Analyse')</li> <li>Jedes Diagramm liegt in einem für den Druck (PDF) und das Web (PNG) optimierten Format vor. Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl für Menschen als auch für Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="../doi/10.5281/zenodo.6667305">Veröffentlichung des Source Codes</a></li> </ul> <p> </p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em> 8. Juli 2025</p> <p><em>Inhaltlicher Umfang</em>: 31.203 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 2000 bis 2025</p> <p><em>Formate:</em><strong> </strong>PDF, TXT, CSV und Parquet</p> <p> </p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist ab Version 2022-07-12 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Robustness Checks auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enthält den Source Code für die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollständige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>öffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="../doi/10.5281/zenodo.6667305">https://zenodo.org/doi/10.5281/zenodo.6667305</a></p> <p> </p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist.</p> <p>Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner <em>persönlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p> </p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Entscheidungstexten und amtlichen Leitsätzen besteht gem. § 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollständig urheberrechtsfrei.</p> <p> </p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Behörden, Gerichten oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p> </p> <p><strong>Weitere Open Access Veröffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>—<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data — <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code — <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regulärer Publikationen — <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p> </p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Kommentieren Sie gerne im Issue Tracker auf Codeberg oder kontaktieren Sie mich via <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p> </p>
Corpus der Entscheidungen des Bundesverfassungsgerichts (CE-BVerfG)
<p><strong>Überblick</strong></p> <p>Das <strong>Corpus der Entscheidungen des Bundesverfassungsgerichts (CE-BVerfG)</strong> ist der bislang größte, frei verfügbare Datensatz von Entscheidungen des Bundesverfassungsgerichts. Er ist eine Zusammenstellung aller Entscheidungen die auf der <a href="https://www.bundesverfassungsgericht.de">amtlichen Webseite des Bundesverfassungsgerichts</a> am jeweiligen Stichtag veröffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die PDF-Sammlung.</p> <p>Alle die <em>Corona-Pandemie</em> betreffenden Entscheidungen des Bundesverfassungsgerichts finden Sie zusätzlich separat dokumentiert und analysiert im Datensatz <strong><a href="https://doi.org/10.5281/zenodo.4459405">Corona-Rechtsprechung des Bundesverfassungsgerichts (BVerfG-Corona)</a></strong>.</p> <p>Das <em>CE-BVerfG</em> sollte nicht mit dem <a href="http://doi.org/10.5281/zenodo.3831111"><strong>Corpus der amtlichen Entscheidungssammlung des Bundesverfassungsgerichts (C-BVerfGE)</strong></a> verwechselt werden. Letzterer zielt nur auf eine Abbildung der amtlichen Sammlung ab und ist deutlich kleiner.</p> <p> </p> <p><strong>Aktualisierung</strong></p> <p>Dieser Datensatz wird <em>1-2 mal im Jahr </em>aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p> </p> <p><strong>NEU in Version 2025-09-22</strong></p> <ul> <li>Vollständige Aktualisierung der Daten</li> <li>Neue Variablen: Tenor und Beschlussformel</li> <li>Neues Feature: Datensatz im Parquet-Format verfügbar</li> <li>Variable entfernt: Kurzbeschreibung (wird vom BVerfG nicht mehr angeboten)</li> <li>Amtliche Sammlung bis inklusive Band 169 mit Name, Band und Seite versehen</li> <li>Anpassung der Sammlung von Entscheidungen an neues Datenbankformat der BVerfG-Webseite</li> <li>Überarbeitung der Dokumentation zu den Varianten des Datensatzes</li> <li>Expliziter R Package Version Lock für 2024-06-13 (CRAN Date)</li> <li>Parallelisierung von Quanteda repariert</li> <li>Überarbeitung des Dockerfiles</li> <li>Zusätzliche Tests auf Einzigartigkeit von ECLI und doc_id</li> <li>Definition von Inhalt des Source Code-Archivs via git</li> <li>/tmp in Arbeitsspeicher ausgelagert</li> <li>Fix für Metadaten-Extraktion</li> </ul> <p> </p> <p><strong>Features</strong></p> <ul> <li>Insgesamt bis zu 38 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabhängige Formate (PDF, TXT, CSV, HTML)</li> <li>Entscheidungsnamen und BVerfGE-Fundstelle (nur BVerfGE)</li> <li>Verknüpfung mit PräsidentIn/Vize-PräsidentIn</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erläutern</li> <li>Dutzende Diagramme und Tabellen für alle Zwecke (im ZIP-Archiv 'ANALYSE').</li> <li>Jedes Diagramm liegt in einem für den Druck (PDF) und das Web (PNG) optimierten Format vor. Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl für Menschen als auch für Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4308216">Veröffentlichung des Source Codes</a></li> </ul> <p> </p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em> 22. September 2025</p> <p><em>Inhaltlicher Umfang</em>: 9.302 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 1998 bis 2025, plus vereinzelte Entscheidungen aus anderen Jahren</p> <p><em>Zitationsnetzwerk: </em>ca. 170.000 Zitate zwischen ca. 14.500 Entscheidungen</p> <p><em>Formate:</em><strong> </strong>PDF, TXT, CSV, HTML und GraphML</p> <p> </p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist ab Version 2021-01-08 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Robustness Checks auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enthält die gesamte Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollständige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>öffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="../doi/10.5281/zenodo.3902658">https://doi.org/10.5281/zenodo.4308216</a></p> <p> </p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist.</p> <p>Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner <em>persönlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p> </p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Entscheidungstexten und amtlichen Leitsätzen besteht gem. § 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollständig urheberrechtsfrei.</p> <p> </p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht weder mit dem Bundesverfassungsgericht noch mit den Herausgebern der BVerfGE in Verbindung.</p> <p> </p> <p><strong>Ähnliche Datensätze zum BVerfG<br></strong></p> <p>Wendel, L., & Möllers, C. (2023). Korpus der Entscheidungen des Bundesverfassungsgerichts (2.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10369205">https://doi.org/10.5281/zenodo.10369205</a></p> <p><em>[Senatsentscheidungen 1972-2010] Engst, Benjamin G., Thomas Gschwend, Christoph Hönnige & Caroline E. Wittig. (2020). “The Constitutional Court Database." <a href="https://ccdb.eu/">https://ccdb.eu/</a></em></p> <p><em>[BVerfGE von 1951 bis 2015]</em> Coupette, C. (2019). Juristische Netzwerkforschung: Modellierung, Quantifizierung und Visualisierung relationaler Daten im Recht (Online-Appendix). Zenodo. <a href="https://doi.org/10.1628/978-3-16-157012-4-appendix">https://doi.org/10.1628/978-3-16-157012-4-appendix</a></p> <p> </p> <p><strong>Weitere Open Access Veröffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>—<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data — <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code — <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regulärer Publikationen — <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p> </p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Notieren Sie diese bitte im Issue Tracker auf Codeberg oder schreiben Sie mir via <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>
The Distant Listening Corpus
<p dir="auto">This is a README file for a data repository originating from the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and serves as welcome page for both</p> <ul> <li>the GitHub repo <a href="https://github.com/DCMLab/distant_listening_corpus">https://github.com/DCMLab/distant_listening_corpus</a> and the corresponding</li> <li>documentation page <a href="https://dcmlab.github.io/distant_listening_corpus" rel="nofollow">https://dcmlab.github.io/distant_listening_corpus</a></li> </ul> <p dir="auto">For information on how to obtain and use the dataset, please refer to <a href="https://dcmlab.github.io/distant_listening_corpus/introduction" rel="nofollow">this documentation page</a>.</p> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#the-distant-listening-corpus-a-corpus-of-annotated-scores">The Distant Listening Corpus (A corpus of annotated scores)</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#getting-the-data">Getting the data</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#data-formats">Data Formats</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#opening-scores">Opening Scores</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#opening-tsv-files-in-a-spreadsheet">Opening TSV files in a spreadsheet</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#loading-tsv-files-in-python">Loading TSV files in Python</a></li> </ul> </li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#version-history">Version history</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#questions-suggestions-corrections-bug-reports">Questions, Suggestions, Corrections, Bug Reports</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#cite-as">Cite as</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#license">License</a></li> </ul> </li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#overview">Overview</a> <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#abc">ABC</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bach_en_fr_suites">bach_en_fr_suites</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bach_solo">bach_solo</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#bartok_bagatelles">bartok_bagatelles</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#beethoven_piano_sonatas">beethoven_piano_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#c_schumann_lieder">c_schumann_lieder</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#chopin_mazurkas">chopin_mazurkas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#corelli">corelli</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#couperin_clavecin">couperin_clavecin</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#couperin_concerts">couperin_concerts</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#cpe_bach_keyboard">cpe_bach_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#debussy_suite_bergamasque">debussy_suite_bergamasque</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#dvorak_silhouettes">dvorak_silhouettes</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#frescobaldi_fiori_musicali">frescobaldi_fiori_musicali</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#grieg_lyric_pieces">grieg_lyric_pieces</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#handel_keyboard">handel_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#jc_bach_sonatas">jc_bach_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#kleine_geistliche_konzerte">kleine_geistliche_konzerte</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#kozeluh_sonatas">kozeluh_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#liszt_pelerinage">liszt_pelerinage</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mahler_kindertotenlieder">mahler_kindertotenlieder</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#medtner_tales">medtner_tales</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mendelssohn_quartets">mendelssohn_quartets</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#monteverdi_madrigals">monteverdi_madrigals</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#mozart_piano_sonatas">mozart_piano_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#pergolesi_stabat_mater">pergolesi_stabat_mater</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#peri_euridice">peri_euridice</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#pleyel_quartets">pleyel_quartets</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#poulenc_mouvements_perpetuels">poulenc_mouvements_perpetuels</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#rachmaninoff_piano">rachmaninoff_piano</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#ravel_piano">ravel_piano</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#scarlatti_sonatas">scarlatti_sonatas</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schubert_winterreise">schubert_winterreise</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schulhoff_suite_dansante_en_jazz">schulhoff_suite_dansante_en_jazz</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schumann_kinderszenen">schumann_kinderszenen</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#schumann_liederkreis">schumann_liederkreis</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#sweelinck_keyboard">sweelinck_keyboard</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#tchaikovsky_seasons">tchaikovsky_seasons</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#wagner_overtures">wagner_overtures</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/#wf_bach_sonatas">wf_bach_sonatas</a></li> </ul> </li> </ul> <div dir="auto"> <h1>The Distant Listening Corpus (A corpus of annotated scores)</h1> <a href="https://github.com/DCMLab/distant_listening_corpus/#the-distant-listening-corpus-a-corpus-of-annotated-scores"></a></div> <p dir="auto"><em>A modular infrastructure for the empirical study of (an)notated music</em></p> <p dir="auto">This corpus has been created within the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and employs the <a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p dir="auto">The publication covers the following public corpora (the DOI links always point at the latest version respectively):</p> <ul> <li>J.S. Bach – English and French Suites [<a href="https://doi.org/10.5281/zenodo.14996489" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bach_en_fr_suites">repo</a>][<a href="https://github.com/DCMLab/bach_en_fr_suites/archive/refs/heads/main.zip">ZIP</a>]</li> <li>J.S. Bach – Solo Pieces (A corpus of annotated scores) [<a href="https://doi.org/10.5281/zenodo.14996765" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bach_solo">repo</a>][<a href="https://github.com/DCMLab/bach_solo/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Béla Bartók – 14 Bagatelles, Op. 6 [<a href="https://doi.org/10.5281/zenodo.14996945" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/bartok_bagatelles">repo</a>][<a href="https://github.com/DCMLab/bartok_bagatelles/archive/refs/heads/main.zip">ZIP</a>]</li> <li>François Couperin – L'art de toucher le clavecin [<a href="https://doi.org/10.5281/zenodo.14984598" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/couperin_clavecin">repo</a>][<a href="https://github.com/DCMLab/couperin_clavecin/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Clara Schumann – Lieder [<a href="https://doi.org/10.5281/zenodo.14996952" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/c_schumann_lieder">repo</a>][<a href="https://github.com/DCMLab/c_schumann_lieder/archive/refs/heads/main.zip">ZIP</a>]</li> <li>François Couperin – Concerts Royaux [<a href="https://doi.org/10.5281/zenodo.15027239" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/couperin_concerts">repo</a>][<a href="https://github.com/DCMLab/couperin_concerts/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Carl Philipp Emanuel Bach – Works for Keyboard [<a href="https://doi.org/10.5281/zenodo.14996326" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/cpe_bach_keyboard">repo</a>][<a href="https://github.com/DCMLab/cpe_bach_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Girolamo Frescobaldi (1583-1643) – Fiori Musicali, op. 12 (1635) [<a href="https://doi.org/10.5281/zenodo.14984864" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/frescobaldi_fiori_musicali">repo</a>][<a href="https://github.com/DCMLab/frescobaldi_fiori_musicali/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Georg Friedrich Händel – Grobschmied Variations (The Harmonious Blacksmith), HWV 430 [<a href="https://doi.org/10.5281/zenodo.14996996" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/handel_keyboard">repo</a>][<a href="https://github.com/DCMLab/handel_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>J.C. Bach – Keyboard Sonatas [<a href="https://doi.org/10.5281/zenodo.14996292" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/jc_bach_sonatas">repo</a>][<a href="https://github.com/DCMLab/jc_bach_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Heinrich Schütz – Kleine Geistliche Konzerte [<a href="https://doi.org/10.5281/zenodo.14997003" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/kleine_geistliche_konzerte">repo</a>][<a href="https://github.com/DCMLab/kleine_geistliche_konzerte/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Leopold Koželuch – Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.14997015" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/kozeluh_sonatas">repo</a>][<a href="https://github.com/DCMLab/kozeluh_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Gustav Mahler – Kindertotenlieder [<a href="https://doi.org/10.5281/zenodo.14997022" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mahler_kindertotenlieder">repo</a>][<a href="https://github.com/DCMLab/mahler_kindertotenlieder/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Felix Mendelssohn – String Quartets [<a href="https://doi.org/10.5281/zenodo.14996150" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mendelssohn_quartets">repo</a>][<a href="https://github.com/DCMLab/mendelssohn_quartets/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Claudio Monteverdi – Madrigals [<a href="https://doi.org/10.5281/zenodo.15003026" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/monteverdi_madrigals">repo</a>][<a href="https://github.com/DCMLab/monteverdi_madrigals/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Giovanni Battista Pergolesi – Stabat Mater (1736) [<a href="https://doi.org/10.5281/zenodo.14990099" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/pergolesi_stabat_mater">repo</a>][<a href="https://github.com/DCMLab/pergolesi_stabat_mater/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Jacopo Peri – Euridice (1600) [<a href="https://doi.org/10.5281/zenodo.14996445" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/peri_euridice">repo</a>][<a href="https://github.com/DCMLab/peri_euridice/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Ignaz Pleyel – String Quartets [<a href="https://doi.org/10.5281/zenodo.14997048" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/pleyel_quartets">repo</a>][<a href="https://github.com/DCMLab/pleyel_quartets/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Francis Poulenc – Mouvements Perpetuels [<a href="https://doi.org/10.5281/zenodo.14997053" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/poulenc_mouvements_perpetuels">repo</a>][<a href="https://github.com/DCMLab/poulenc_mouvements_perpetuels/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Sergei Rachmaninoff – Piano Pieces [<a href="https://doi.org/10.5281/zenodo.14984155" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/rachmaninoff_piano">repo</a>][<a href="https://github.com/DCMLab/rachmaninoff_piano/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Maurice Ravel – Piano Pieces [<a href="https://doi.org/10.5281/zenodo.14997064" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/ravel_piano">repo</a>][<a href="https://github.com/DCMLab/ravel_piano/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Domenico Scarlatti – Keyboard Sonatas [<a href="https://doi.org/10.5281/zenodo.14992884" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/scarlatti_sonatas">repo</a>][<a href="https://github.com/DCMLab/scarlatti_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Franz Schubert – Winterreise [<a href="https://doi.org/10.5281/zenodo.14997095" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schubert_winterreise">repo</a>][<a href="https://github.com/DCMLab/schubert_winterreise/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Erwin Schulhoff – Suite dansante en jazz [<a href="https://doi.org/10.5281/zenodo.14997098" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schulhoff_suite_dansante_en_jazz">repo</a>][<a href="https://github.com/DCMLab/schulhoff_suite_dansante_en_jazz/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Robert Schumann – Liederkreis [<a href="https://doi.org/10.5281/zenodo.14997104" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schumann_liederkreis">repo</a>][<a href="https://github.com/DCMLab/schumann_liederkreis/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Jan Sweelinck – Organ Pieces [<a href="https://doi.org/10.5281/zenodo.14997111" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/sweelinck_keyboard">repo</a>][<a href="https://github.com/DCMLab/sweelinck_keyboard/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Richard Wagner – Overtures [<a href="https://doi.org/10.5281/zenodo.14997120" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/wagner_overtures">repo</a>][<a href="https://github.com/DCMLab/wagner_overtures/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Wilhelm Friedemann Bach – Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.14997133" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/wf_bach_sonatas">repo</a>][<a href="https://github.com/DCMLab/wf_bach_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Rammos, Y., Neuwirth, M., Moss, F. C., & Rohrmeier, M. (2024). An annotated corpus of tonal piano music from the long 19th century. Empirical Musicology Review, 18(1), 84–95. <a href="https://doi.org/10.18061/emr.v18i1.8903" rel="nofollow">https://doi.org/10.18061/emr.v18i1.8903</a></em></p> <ul> <li>Ludwig van Beethoven - Piano Sonatas [<a href="https://doi.org/10.5281/zenodo.7473560" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/beethoven_piano_sonatas">repo</a>][<a href="https://github.com/DCMLab/beethoven_piano_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Frédéric Chopin - Mazurkas [<a href="https://doi.org/10.5281/zenodo.7473566" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/chopin_mazurkas">repo</a>][<a href="https://github.com/DCMLab/chopin_mazurkas/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Claude Debussy - Suite Bergamasque [<a href="https://doi.org/10.5281/zenodo.7473568" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/debussy_suite_bergamasque">repo</a>][<a href="https://github.com/DCMLab/debussy_suite_bergamasque/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Antonín Dvořák - Silhouettes [<a href="https://doi.org/10.5281/zenodo.7473576" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/dvorak_silhouettes">repo</a>][<a href="https://github.com/DCMLab/dvorak_silhouettes/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Edvard Grieg - Lyric Pieces [<a href="https://doi.org/10.5281/zenodo.7473578" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/grieg_lyric_pieces">repo</a>][<a href="https://github.com/DCMLab/grieg_lyric_pieces/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Franz Liszt - Années de Pèlerinage [<a href="https://doi.org/10.5281/zenodo.7473580" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/liszt_pelerinage">repo</a>][<a href="https://github.com/DCMLab/liszt_pelerinage/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Nikolai Medtner - Tales [<a href="https://doi.org/10.5281/zenodo.7473528" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/medtner_tales">repo</a>][<a href="https://github.com/DCMLab/medtner_tales/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Robert Schumann - Kinderszenen [<a href="https://doi.org/10.5281/zenodo.7473582" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/schumann_kinderszenen">repo</a>][<a href="https://github.com/DCMLab/schumann_kinderszenen/archive/refs/heads/main.zip">ZIP</a>]</li> <li>Pyotr Tchaikovsky - The Seasons [<a href="https://doi.org/10.5281/zenodo.7473586" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/tchaikovsky_seasons">repo</a>][<a href="https://github.com/DCMLab/tchaikovsky_seasons/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Moss, F. C., Neuwirth, M., & Rohrmeier, M. A. (2021). A semi-automated workflow paradigm for the distributed creation and curation of expert annotations. Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR, 262–269. <a href="https://doi.org/10.5281/ZENODO.5624417" rel="nofollow">https://doi.org/10.5281/ZENODO.5624417</a></em></p> <ul> <li>Arcangelo Corelli – Trio Sonatas [<a href="https://zenodo.org/doi/10.5281/zenodo.7504011" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/corelli">repo</a>][<a href="https://github.com/DCMLab/corelli/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Hentschel, J., Neuwirth, M., & Rohrmeier, M. (2021). The Annotated Mozart Sonatas: Score, harmony, and cadence. Transactions of the International Society for Music Information Retrieval, 4(1), 67–80. <a href="https://doi.org/10.5334/tismir.63" rel="nofollow">https://doi.org/10.5334/tismir.63</a></em></p> <ul> <li>Wolfgang Amadeus Mozart - Piano Sonatas [<a href="https://zenodo.org/doi/10.5281/zenodo.7424962" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/mozart_piano_sonatas">repo</a>][<a href="https://github.com/DCMLab/mozart_piano_sonatas/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <p dir="auto"><em>Neuwirth, M., Harasim, D., Moss, F. C., & Rohrmeier, M. (2018). The Annotated Beethoven Corpus (ABC): A Dataset of Harmonic Analyses of All Beethoven String Quartets. Frontiers in Digital Humanities, 5(July), 1–5. <a href="https://doi.org/10.3389/fdigh.2018.00016" rel="nofollow">https://doi.org/10.3389/fdigh.2018.00016</a></em></p> <ul> <li>Ludwig van Beethoven - String Quartets [<a href="https://zenodo.org/doi/10.5281/zenodo.7441343" rel="nofollow">DOI</a>][<a href="https://github.com/DCMLab/ABC">repo</a>][<a href="https://github.com/DCMLab/ABC/archive/refs/heads/main.zip">ZIP</a>]</li> </ul> <div dir="auto"> <h2>Getting the data</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#getting-the-data"></a></div> <ul> <li>download individual subcorpora as ZIP files using the URLs provided above</li> <li>download a <a href="https://specs.frictionlessdata.io/data-package/" rel="nofollow">Frictionless Datapackage</a> that includes concatenations of the TSV files in the four folders (<code>measures</code>, <code>notes</code>, <code>chords</code>, and <code>harmonies</code>) and a JSON descriptor: <ul> <li><a href="https://github.com/DCMLab/distant_listening_corpus/releases/latest/download/distant_listening_corpus.zip">distant_listening_corpus.zip</a></li> <li><a href="https://github.com/DCMLab/distant_listening_corpus/releases/latest/download/distant_listening_corpus.datapackage.json">distant_listening_corpus.datapackage.json</a></li> </ul> </li> <li>clone the repo (~2.4 GB): <code>git clone --recursive -j12 https://github.com/DCMLab/distant_listening_corpus.git</code></li> </ul> <div dir="auto"> <h2>Data Formats</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#data-formats"></a></div> <p dir="auto">Each piece in this corpus is represented by five files with identical name prefixes, each in its own folder. For example, the <em>Prélude</em> of J.S. Bach’s first English Suite, BWV 806, has the following files:</p> <ul> <li><code>MS3/BWV806_01_Prelude.mscx</code>: Uncompressed MuseScore 3.6.2 file including the music and annotation labels.</li> <li><code>notes/BWV806_01_Prelude.notes.tsv</code>: A table of all note heads contained in the score and their relevant features (not each of them represents an onset, some are tied together)</li> <li><code>measures/BWV806_01_Prelude.measures.tsv</code>: A table with relevant information about the measures in the score.</li> <li><code>chords/BWV806_01_Prelude.chords.tsv</code>: A table containing layer-wise unique onset positions with the musical markup (such as dynamics, articulation, lyrics, figured bass, etc.).</li> <li><code>harmonies/BWV806_01_Prelude.harmonies.tsv</code>: A table of the included harmony labels (including cadences and phrases) with their positions in the score.</li> </ul> <p dir="auto">Each TSV file comes with its own JSON descriptor that describes the meanings and datatypes of the columns ("fields") it contains, follows the <a href="https://specs.frictionlessdata.io/tabular-data-resource/" rel="nofollow">Frictionless specification</a>, and can be used to validate and correctly load the described file.</p> <div dir="auto"> <h3>Opening Scores</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#opening-scores"></a></div> <p dir="auto">After navigating to your local copy, you can open the scores in the folder <code>MS3</code> with the free and open source score editor <a href="https://musescore.org" rel="nofollow">MuseScore</a>. Please note that the scores have been edited, annotated and tested with <a href="https://github.com/musescore/MuseScore/releases/tag/v3.6.2">MuseScore 3.6.2</a>. MuseScore 4 has since been released which renders them correctly but cannot store them back in the same format.</p> <div dir="auto"> <h3>Opening TSV files in a spreadsheet</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#opening-tsv-files-in-a-spreadsheet"></a></div> <p dir="auto">Tab-separated value (TSV) files are like Comma-separated value (CSV) files and can be opened with most modern text editors. However, for correctly displaying the columns, you might want to use a spreadsheet or an addon for your favourite text editor. When you use a spreadsheet such as Excel, it might annoy you by interpreting fractions as dates. This can be circumvented by using <code>Data --> From Text/CSV</code> or the free alternative <a href="https://www.libreoffice.org/download/download/" rel="nofollow">LibreOffice Calc</a>. Other than that, TSV data can be loaded with every modern programming language.</p> <div dir="auto"> <h3>Loading TSV files in Python</h3> <a href="https://github.com/DCMLab/distant_listening_corpus/#loading-tsv-files-in-python"></a></div> <p dir="auto">Since the TSV files contain null values, lists, fractions, and numbers that are to be treated as strings, you may want to use this code to load any TSV files related to this repository (provided you're doing it in Python). After a quick <code>pip install -U ms3</code> (requires Python 3.10 or later) you'll be able to load any TSV like this:</p> <div dir="auto"> <pre><span>import</span> <span>ms3</span> <span>labels</span> <span>=</span> <span>ms3</span>.<span>load_tsv</span>(<span>"harmonies/BWV806_01_Prelude.harmonies.tsv"</span>) <span>notes</span> <span>=</span> <span>ms3</span>.<span>load_tsv</span>(<span>"notes/BWV806_01_Prelude.notes.tsv"</span>)</pre> <div> </div> </div> <div dir="auto"> <h2>Version history</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#version-history"></a></div> <p dir="auto">See the <a href="https://github.com/DCMLab/distant_listening_corpus/releases">GitHub releases</a>.</p> <div dir="auto"> <h2>Questions, Suggestions, Corrections, Bug Reports</h2> <a href="https://github.com/DCMLab/distant_listening_corpus/#questions-suggestions-corrections-bug-reports"></a></div> <p>Please <a href="https://github.com/DCMLab/distant_listening_corpus/issues">create an issue</a> and/or feel free to fork and submit pull requests.</p>
Données supplémentaires: Repérage automatisé de l'hyponymie dans des corpus spécialisés en français à l'aide de Sketch Engine
<p>Ces figures sont des données supplémentaires de l'article suivant :<br> San Martín A., Trekker C., et León-Araúz P. 2022. Repérage automatisé de l’hyponymie dans des corpus spécialisés en français à l’aide de Sketch Engine. <em>Terminology</em>. doi: 10.1075/term.20044.san</p> <p>Les figures suivantes représentent le résultat complet de l’évaluation des WS. La première colonne représente le terme évalué (c’est-à-dire les termes de recherche) et les trois colonnes suivantes, les trois premiers résultats. Enfin, les colonnes suivantes représentent visuellement la précision de chaque paire, le chiffre à gauche étant le nombre de vrais positifs et celui à droite, le nombre de correspondances associées à la paire. La couleur bleue représente les résultats de la colonne <em>X est le générique de...</em> et la couleur jaune, les résultats de la colonne <em>X est un type de...</em></p> <ul> <li>psychologie.tif: Évaluation des WS du sous-corpus de psychologie</li> <li>chimie.tif: Évaluation des WS du sous-corpus de chimie</li> <li>droit.tif: Évaluation des WS du sous-corpus de droit</li> <li>informatique.tif: Évaluation des WS du sous-corpus d’informatique</li> <li>geographie.tif: Évaluation des WS du sous-corpus de géographie</li> </ul>
SRBCorp: Corpus of Parliamentary Debates in Serbia
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the National Assembly of Serbia. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1997-2020 (eight terms) and counts over 300 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): SRBCorp: Corpus of Parliamentary Debates in Serbia (v1.1.1), https://doi.org/10.5281/zenodo.6521648.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - added a new variable for policy category "tag" using ML model trained on known tags of agenda points in the parliaments of Croatia and Bosnia-Herzegovina</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2389); migrated to ZENODO due to limitations concerning the concept DOI</p>
CROCorp: Corpus of Parliamentary Debates in Croatia
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Croatian Parliament (Sabor). The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 2003-2020 (five complete terms) and counts over 500 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): CROCorp: Corpus of Parliamentary Debates in Croatia (v1.1.1), https://doi.org/10.5281/zenodo.6521372.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - improved coding of dummy variable "moderator" (using less error-prone alghoritm for detecting the modertor role)<br> - fixed issue with agenda points which are conncatenated while preserving a unique web link<br> - recoded agenda points tags using better ML model (transformer architecture)</p> <p>v1.0.0<br> - originally posted on GESIS repository (migrated to ZENODO due to limitations concerning the concept DOI)</p>
BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Parliamentary Assembly of Bosnia and Herzegovina. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1998-2018 (six complete terms) and counts over 127 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, Christophe Lesschaeve, and Ensar Muharemović (2022): BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina (v1.1.1),<br> https://doi.org/10.5281/zenodo.6517697.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - fixed a typo in one of the debates' date<br> - fixed minor inconsistencies in the tag column</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2387); migrated to ZENODO due to limitations concerning the concept DOI</p>
Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language
<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.