Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

878

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

878 results for “Corpus”

Learn how ShareScore rates datasets ↗
zenodo44/100

FiloBass: A Dataset and Corpus Based Study of Jazz Basslines

<p>Dataset to accompany the paper "FiloBass: A Dataset and Corpus Based Study of Jazz Basslines" which was published at ISMIR 2023.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

An Annotated Corpus of Tonal Piano Music from the Long 19th Century

<p>This corpus has been created within the&nbsp;<a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a>&nbsp;and employs&nbsp;the&nbsp;<a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p><strong>Version 1</strong>&nbsp;has been released for submitting it as part of the data&nbsp;report&nbsp;<code>Hentschel, J., Rammos, Y., Neuwirth, M., Rohrmeier, M. (forthcoming). An Annotated Corpus of Tonal Piano Music from the Long 19th Century</code>&nbsp;that accompanies nine corpora grouped under the DOI&nbsp;<a href="https://doi.org/10.5281/zenodo.7483349">10.5281/zenodo.7483349</a>.</p> <p><strong>Version 1.1</strong>&nbsp;comes with a complete set of metadata and score headers.&nbsp;Among more accurate composition dates,&nbsp;the&nbsp;metadata now include URIs that identify the compositions in terms of&nbsp;the&nbsp;<a href="https://viaf.org/">Virtual International Authority File (VIAF)</a>,&nbsp;<a href="https://www.wikidata.org/">Wikidata</a>,&nbsp;<a href="https://imslp.org/">IMSLP</a>&nbsp;and&nbsp;<a href="https://musicbrainz.org/">MusicBrainz</a>.&nbsp;The data has been re-extracted from the scores&nbsp;using&nbsp;<a href="https://pypi.org/project/ms3/">ms3 1.1.1</a>.</p> <p>The publication covers the following corpora (the DOI links always point at the latest version respectively):</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7473560">Ludwig van Beethoven - Piano Sonatas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473566">Fr&eacute;d&eacute;ric Chopin - Mazurkas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473568">Claude Debussy - Suite Bergamasque</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473576">Anton&iacute;n Dvoř&aacute;k - Silhouettes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473580">Franz Liszt - Ann&eacute;es de P&egrave;lerinage</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473528">Nikolai Medtner - Tales</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473582">Robert Schumann - Kinderszenen</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473586">Pyotr Tchaikovsky - The Seasons</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473578">Edvard Grieg - Lyric Pieces</a></li> </ul> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2022View details →
zenodo44/100

The Tsez Annotated Corpus Project

<p>Cite the source of the dataset as:</p> <blockquote> <p>Abdulaev, A.K. &amp; I. K. Abdullaev. 2010. Cezyas folklor/Dido (Tsez) folklore/Didojskij (cezskij) fol´klor. Leipzig–Makhachkala: "Lotos".</p> </blockquote>

opencc-by-4.0Sep 2022View details →
zenodo44/100

The grammatically annotated corpus of the pericopes of the Old Lithuanian Postil of Jonas Bretkūnas

<p>This grammatically annoted corpus aims at facilitating linguistic research on Old Lithuanian based on the Postil of Jonas Bretkūnas from the year 1591. In addition to the two subcorpora "pericopes" and "homilies" of the version 1.0, the current version 2.0<br>includes also "passion harmony", "pericope (prophet)" and "prayer". In this new version, the tokens occurring in biblical citations can be specifically searched for. The search results also show the corresponding passage in the BrP pericopes.<br>For full description see the BrP_2.0._documentation.md included in the files.</p> <p>The pericopes were annotated using SIL Toolbox and converted to be used in the search-tool ANNIS using the conversion tool PEPPER.</p> <p>Three formats are provided in this release: 1. the Toolbox files, 2. the transitional Excel files and 3. a zipped folder to be imported into ANNIS.</p> <p>Created in the project B02, <em>Emergence and change of registers: The case of Lithuanian and Latvian</em> of the CRC 1412 "Register" (funded by the Deutsche Forschungsgemeinschaft: DFG, German Research Foundation: 416591334).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo44/100

A Survey of Body Part Construction Metaphors in the Neo-Assyrian Letter Corpus

<p>The dataset consists of approximately 2,400 examples of metaphors in Akkadian of what we term Body Part Constructions (BPC&#39;s) within the letter sub-corpus of the <a href="http://oracc.museum.upenn.edu/saao/">State Archives of Assyria online</a> (SAAo). The dataset was generated by a multi-step process involving the training and application of a spaCy language model to the SAAo letter sub-corpus, converting the resulting annotations to linked open data format amenable to searching for BPC&rsquo;s, and manually adding metalinguistic data to the search results; these files, in CONLLU and TTL formats, as well as the model specific files based on spaCy&#39;s requirements, are also made available in this publication. The BPC dataset is stored as a CSV file, and can serve as an easy starting place for other scholars interested in finding socio-linguistic usage patterns of this construction.</p> <p>The royal archives of the late Neo-Assyrian kings (8th-7th century BCE) constitute an important source for understanding many facets of the Neo-Assyrian empire. Ranging from treaty tablets and legal documents to prophecies, ritual instructions, and even court literature, the approximately five thousand texts in this corpus primarily come from the palatial complex at Nineveh and document the reigns of Sargon II (r. 721-705), Sennacherib (r. 704-681), Esarhaddon (r. 680-669), and Assurbanipal (668-627). Over the past four decades, much of these archives has been published in the State Archives of Assyria (SAA) volumes at the University of Helsinki, and in more recent years has appeared digitally under the <a href="http://www.en.ag.geschichte.uni-muenchen.de/research/mocci/">Munich Open-access Cuneiform Corpus Initiative</a>&nbsp;(LMU Munich) as the SAAo.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Corpus der Entscheidungen des Bundesverwaltungsgerichts (CE-BVerwG)

<p><strong>&Uuml;berblick</strong></p> <p>Das <strong>Corpus der Entscheidungen des Bundesverwaltungsgerichts (CE-BVerwG)</strong> ist der bislang gr&ouml;&szlig;te, frei verf&uuml;gbare Datensatz von Entscheidungen des Bundesverwaltungsgerichts. Er ist eine Zusammenstellung aller Entscheidungen die in der <a href="https://www.bverwg.de/">amtlichen Datenbank des Bundesverwaltungsgerichts</a> am jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die PDF-Sammlung.</p> <p>&nbsp;</p> <p><strong>Aktualisierung</strong></p> <p>Dieser Datensatz wird 1-2 mal im Jahr aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <p><strong>Neu in Version 2024-03-13</strong></p> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten</li> <li>Die Pipeline mit allen Zwischenergebnissen wird nun automatisch in "output/" archiviert</li> <li>Gruppierung von Dateinamen der Diagramme nach Typ</li> <li>Vereinfachung der Repository-Struktur</li> <li>Anpassung von Docker Compose an Debian 11</li> <li>Update GPG Public Key in Repository</li> </ul> <p>&nbsp;</p> <p><strong>Features</strong></p> <ul> <li>31 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV, HTML)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'Analyse').</li> <li>Jedes Diagramm liegt in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor. Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="../doi/10.5281/zenodo.4625134">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <p>&nbsp;</p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em> 13. M&auml;rz 2024</p> <p><em>Inhaltlicher Umfang</em>: 27.200 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 2002 bis 2024, plus vereinzelte Entscheidungen aus anderen Jahren</p> <p><em>Formate:</em><strong> </strong>PDF, TXT und CSV</p> <p>&nbsp;</p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist ab Version 2021-04-15 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Robustness Checks auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Source Code f&uuml;r die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="../doi/10.5281/zenodo.4625134">https://zenodo.org/doi/10.5281/zenodo.4625134</a></p> <p>&nbsp;</p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p>&nbsp;</p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Entscheidungstexten und amtlichen Leits&auml;tzen besteht gem. &sect; 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Beh&ouml;rden, Gerichten oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <p><strong>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im <a href="https://github.com/SeanFobbe/ce-bverwg/issues">Issue Tracker auf GitHub</a> oder schreiben Sie mir eine E-Mail an <a href="mailto:fobbe-data@posteo.de">fobbe-data@posteo.de</a></p> <p>&nbsp;</p>

opencc-zeroJul 2020View details →
zenodo44/100

PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities

<h2><a href="https://marianaossilva.github.io/pportal_ner/" target="_blank" rel="noopener">PPORTAL_ner</a></h2> <h3>An Annotated Dataset of Portuguese Literary Entities</h3> <p>The corpus is tailored to Brazilian and Portuguese literary texts, containing annotations for five entity categories, including PER, LOC, GPE, ORG, and DATE. Within a diverse collection of 25 literary works, it offers a total of 125,059 tokens and 5,266 annotated entities. This dataset contributes to the development of potentially more accurate and context-aware NER models, as well as to encourage further exploration within Portuguese literature.<br><br></p> <h2>Corpus Statistics</h2> <p>Our corpus is sourced from&nbsp;<a href="https://doi.org/10.5281/zenodo.5178063">PPORTAL</a>, an extensive repository of metadata containing over 80,000 public domain literary works in the Portuguese language, predominantly derived from Brazil and Portugal.&nbsp;<a href="https://doi.org/10.5281/zenodo.5178063">PPORTAL</a>&nbsp;aggregates data from three digital libraries:&nbsp;<a href="https://www.dominiopublico.gov.br/">Dom&iacute;nio P&uacute;blico</a>,&nbsp;<a href="https://projectoadamastor.org/">Projecto Adamastor</a>, and&nbsp;<a href="https://www.literaturabrasileira.ufsc.br/">Biblioteca Digital de Literatura dos Pa&iacute;ses Lus&oacute;fonos (BLPL)</a>.</p> <p>To simplify referencing, this new dataset is called PPORTAL_ner. PPORTAL_ner selection process contains a diverse range of 25 individual literary works, spanning different authors and literary styles. All of these texts were published prior to 1953, adhering to the current criteria for public domain status in Brazil, with the majority falling within the timeframe spanning from 1554 to 1938.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

CLDF dataset derived from the DoReCo core corpus

<p>Cite the source of the dataset as:</p> <blockquote> <p>Seifart, Frank, Ludger Paschen &amp; Matthew Stave (eds.). 2022. Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). DOI:10.34847/nkl.7cbfq779</p> </blockquote> <p>Please note that when citing this dataset, it is NOT sufficient to refer to DoReCo as a whole, but the full citation for each individual corpus must be provided, including the name(s) of the creator(s) of each corpus.</p> <blockquote>Ozerov, Pavel. 2022. Anal DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/anal1239. DOI:10.34847/nkl.0dbazp8m</blockquote> <blockquote>Riesberg, Sonja. 2022. Yali (Apahapsili) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/apah1238. DOI:10.34847/nkl.9d91nkq2</blockquote> <blockquote>Cowell, Andrew. 2022. Arapaho DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/arap1274. DOI:10.34847/nkl.36f5r1b6</blockquote> <blockquote>Cobbinah, Alexander Yao. 2022. Baïnounk Gubëeher DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/bain1259. DOI:10.34847/nkl.a332abw8</blockquote> <blockquote>Vanhove, Martine. 2022. Beja DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/beja1238. DOI:10.34847/nkl.edd011t1</blockquote> <blockquote>Seifart, Frank. 2022. Bora DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/bora1263. DOI:10.34847/nkl.6eaf5laq</blockquote> <blockquote>Reiter, Sabine. 2022. Cashinahua DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/cash1254. DOI:10.34847/nkl.a8f9q2f1</blockquote> <blockquote>Däbritz, Chris Lasse and Kudryakova, Nina and Stapert, Eugénie and Arkhipov, Alexandre. 2022. Dolgan DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/dolg1241. DOI:10.34847/nkl.f09eikq3</blockquote> <blockquote>Kazakevich, Olga and Klyachko, Elena. 2022. Evenki DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/even1259. DOI:10.34847/nkl.5e0d27cu</blockquote> <blockquote>Hellwig, Birgit. 2022. Goemai DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/goem1240. DOI:10.34847/nkl.b93664ml</blockquote> <blockquote>Harvey, Andrew. 2022. Gorwaa DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/goro1270. DOI:10.34847/nkl.a4b4ijj2</blockquote> <blockquote>Burenhult, Niclas. 2022. Jahai DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/jeha1242. DOI:10.34847/nkl.6a71xp0p</blockquote> <blockquote>Kim, Soung-U. 2022. Jejuan DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/jeju1234. DOI:10.34847/nkl.06ebrk38</blockquote> <blockquote>Vydrina, Alexandra. 2022. Kakabe DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/kaka1265. DOI:10.34847/nkl.d5aeu9t6</blockquote> <blockquote>Gusev, Valentin and Klooster, Tiina and Wagner-Nagy, Beáta and Arkhipov, Alexandre. 2022. Kamas DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/kama1351. DOI:10.34847/nkl.cdd8177b</blockquote> <blockquote>Hellwig, Birgit and Schneider-Blum, Gertrud and Ismail, Khaleel Bakheet Khaleel. 2022. Tabaq (Karko) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/kark1256. DOI:10.34847/nkl.eea8144j</blockquote> <blockquote>Döhler, Christian. 2022. Komnzo DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/komn1238. DOI:10.34847/nkl.c5e6dudv</blockquote> <blockquote>Bartels, Hauke and Szczepański, Marcin. 2022. Lower Sorbian DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/lowe1385. DOI:10.34847/nkl.6c6e4e9k</blockquote> <blockquote>Haude, Katharina. 2022. Movima DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/movi1243. DOI:10.34847/nkl.da42xf67</blockquote> <blockquote>Ponsonnet, Maïa. 2022. Dalabon DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/ngal1292. DOI:10.34847/nkl.fae299ug</blockquote> <blockquote>Güldemann, Tom and Ernszt, Martina and Siegmund, Sven and Witzlack-Makarevich, Alena. 2022. Nǁng DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/nngg1234. DOI:10.34847/nkl.f6c37fi0</blockquote> <blockquote>Haig, Geoff and Vollmer, Maria and Thiele, Hanna. 2022. Northern Kurdish (Kurmanji) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/nort2641. DOI:10.34847/nkl.ca10ez5t</blockquote> <blockquote>Franjieh, Michael. 2022. Fanbyak DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/orko1234. DOI:10.34847/nkl.02084446</blockquote> <blockquote>Ring, Hiram. 2022. Pnar DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/pnar1238. DOI:10.34847/nkl.5ba1062k</blockquote> <blockquote>Krifka, Manfred. 2022. Daakie DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/port1286. DOI:10.34847/nkl.efeav5l9</blockquote> <blockquote>Seifart, Frank. 2022. Resígaro DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/resi1247. DOI:10.34847/nkl.ffb96lo8</blockquote> <blockquote>Witzlack-Makarevich, Alena and Namyalo, Saudah and Kiriggwajjo, Anatol and Molochieva, Zarina. 2022. Ruuli DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/ruul1235. DOI:10.34847/nkl.fde4pp1u</blockquote> <blockquote>Xu, Xianming and Bai, Bibo. 2022. Sadu DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/sadu1234. DOI:10.34847/nkl.3db4u59d</blockquote> <blockquote>Forker, Diana and Schiborr, Nils Norman. 2022. Sanzhi Dargwa DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/sanz1248. DOI:10.34847/nkl.81934177</blockquote> <blockquote>Wegener, Claudia. 2022. Savosavo DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/savo1255. DOI:10.34847/nkl.b74d1b33</blockquote> <blockquote>Thieberger, Nick. 2022. Nafsan (South Efate) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/sout2856. DOI:10.34847/nkl.ba4f760l</blockquote> <blockquote>Schiborr, Nils Norman. 2022. English (Southern England) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/sout3282. DOI:10.34847/nkl.9c271u5g</blockquote> <blockquote>Avanzi, Mathieu and Béguelin, Marie-José and Corminboeuf, Gilles and Diémoz, Federica and Johnsen, Laure Anne. 2022. French (Swiss) DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/stan1290. DOI:10.34847/nkl.3520l685</blockquote> <blockquote>Teo, Amos. 2022. Sümi DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/sumi1235. DOI:10.34847/nkl.5ad4t01p</blockquote> <blockquote>Gippert, Jost. 2022. Svan DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/svan1243. DOI:10.34847/nkl.9ba054c3</blockquote> <blockquote>Bogomolova, Natalia and Ganenkov, Dmitry and Schiborr, Nils Norman. 2022. Tabasaran DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/taba1259. DOI:10.34847/nkl.ad7f97xr</blockquote> <blockquote>Mosel, Ulrike. 2022. Teop DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/teop1238. DOI:10.34847/nkl.9322sdf2</blockquote> <blockquote>Wichmann, Søren. 2022. Texistepec Popoluca DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/texi1237. DOI:10.34847/nkl.c50ck58f</blockquote> <blockquote>Rose, Françoise. 2022. Mojeño Trinitario DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/trin1278. DOI:10.34847/nkl.cbc3b4xr</blockquote> <blockquote>Griscom, Richard. 2022. Asimjeeg Datooga DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/tsim1256. DOI:10.34847/nkl.f77c7m72</blockquote> <blockquote>Schnell, Stefan. 2022. Vera&#x27;a DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/vera1241. DOI:10.34847/nkl.3e2cu8c4</blockquote> <blockquote>Michaud, Alexis. 2022. Yongning Na DoReCo dataset. In Seifart, Frank, Ludger Paschen and Matthew Stave (eds.). Language Documentation Reference Corpus (DoReCo) 1.2. Berlin &amp; Lyon: Leibniz-Zentrum Allgemeine Sprachwissenschaft &amp; laboratoire Dynamique Du Langage (UMR5596, CNRS &amp; Université Lyon 2). https://doreco.huma-num.fr/languages/yong1270. DOI:10.34847/nkl.abe65p95</blockquote>

opencc-zeroApr 2024View details →
zenodo44/100

AWOFRO : Annotated tweet corpus of mixed Wolof-French for detecting obnoxious messages

<p>These data are tweets of mixed Wolof-French codes annotated by three(3) annotators.&nbsp;<br>They were extracted during the period from 1 January 2021 to 31 May 2023.</p> <p>Content description :</p> <p><strong>Corpora.rar</strong> : The dataset contains 3510 annotated tweets</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Corpus der Entscheidungen des Bundesfinanzhofs (CE-BFH)

<h3><strong>&Uuml;berblick</strong></h3> <p>Das <strong>Corpus der Entscheidungen des Bundesfinanzhofs (CE-BFH)</strong> ist eine m&ouml;glichst vollst&auml;ndige Sammlung der vom Bundesfinanzhof (BFH) ver&ouml;ffentlichten Entscheidungen. Der Datensatz nutzt als seine Datenquelle die <a href="https://www.bundesfinanzhof.de/">amtliche Entscheidungsdatenbank des Bundesfinanzhofs</a> und wertet diese vollst&auml;ndig aus.</p> <p><em>Bitte lesen Sie zuerst das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die PDF-Sammlung.</p> <p>F&uuml;r Praktiker:innen stelle ich zus&auml;tzlich eine Variante mit allen in der amtlichen Sammlung BFHE abgedruckten "V-Entscheidungen" zur Verf&uuml;gung.</p> <p>&nbsp;</p> <h3><strong>Aktualisierung</strong></h3> <p>Dieser Datensatz wird 1-2 mal im Jahr aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <h3><strong>NEU in Version 2025-01-14<br></strong></h3> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten</li> <li>LIZENZ&Auml;NDERUNG: Source Code jetzt unter GNU General Public License Version 3 (GPLv3) oder sp&auml;ter lizenziert</li> <li>NEU: Zitationsnetzwerk des BFH von Aktenzeichen-zu-Aktenzeichen und Aktenzeichen-zu-BFHE als GraphML mit vielen Metadaten</li> <li>NEU: Option f&uuml;r Clean Runs in Konfiguration eingef&uuml;gt (l&ouml;scht alle Daten vor dem eigentlichen Run)</li> <li>NEU: Test auf geringen oder fehlenden Text-Inhalt</li> <li>NEU: Automatische Archivierung der Zwischenergebnisse in der Pipeline als ZIP-Archiv</li> <li>Docker Image auf R 4.4.0 aktualisiert (wegen CVE-2024-27322)</li> <li>Expliziter R Package Version Lock f&uuml;r 2024-06-13 (CRAN Date)</li> <li>&Uuml;berarbeitung des Dockerfiles</li> <li>&Uuml;berarbeitung der Dokumentation zu den Varianten des Datensatzes</li> <li>Vereinheitlichung der Komponenten f&uuml;r PDF-Extraktion, linguistische Statistiken und Berechnung kryptographischer Hashes</li> <li>Vereinfachung der Run-Skripte und st&auml;rkere Integration mit Docker Compose</li> <li>Erweiterung des L&ouml;sch-Skriptes</li> <li>Docker Zeitzone auf Berlin eingestellt</li> <li>Entfernung der Nummerierung von Diagrammen</li> <li>Entfernung der Tesseract Dependencies</li> <li>Entfernung der Python Toolchain</li> <li>Aktualisierung des Public GPG Keys im Repository</li> </ul> <p>&nbsp;</p> <h3><strong>Eckdaten</strong></h3> <p><em>Stichtag:</em> 14. Januar 2025</p> <p><em>Inhaltlicher Umfang:</em> 10.885 Entscheidungen des Bundesfinanzhofs der Bundesrepublik Deutschland</p> <p><em>Zeitlicher Umfang:</em> ab Januar 2010 bis zum Stichtag</p> <p><em>Formate:</em> CSV, GraphML, PDF, TXT und HTML</p> <p>&nbsp;</p> <h3><strong>Features</strong></h3> <ul> <li>Bis zu 34 Variablen in den CSV-Varianten</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Metadaten enthalten u.a. relevante Rechtsnormen, Vorinstanz, Titel der Entscheidung und Leits&auml;tze</li> <li>Zitationsnetzwerk des BFH f&uuml;r Aktenzeichen-zu-Aktenzeichen und Aktenzeichen-zu-BFHE</li> <li>Sowohl f&uuml;r traditionelle Rechtsanwender als auch f&uuml;r Legal Tech-Anwendungen geeignete Formate (CSV, PDF, TXT und HTML)</li> <li>Umfangreicher Compilation Report um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Hochaufl&ouml;sende Diagramme und deskriptive Tabellen f&uuml;r alle Zwecke</li> <li>Diagramme in PDF (Druck) und PNG (Web) verf&uuml;gbar, Tabellen als menschen- und maschinenlesbares CSV</li> <li><a href="Ver&ouml;ffentlichung%20des%20Source%20Codes">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <p>&nbsp;</p> <h3><strong>Source Code und Compilation Report</strong></h3> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Robustness Checks auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Code f&uuml;r die vollst&auml;ndige Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Er ist zusammen mit dem Source Code hinterlegt. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der <em>vollst&auml;ndige Source Code</em> &mdash; sowohl f&uuml;r die Erstellung des Datensatzes, als auch f&uuml;r das Codebook &mdash; ist <em>&ouml;ffentlich einsehbar </em>und<em> dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="https://doi.org/10.5281/zenodo.7691842">https://doi.org/10.5281/zenodo.7691842</a></p> <p>&nbsp;</p> <h3><strong>Kryptographische Signaturen</strong></h3> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p>&nbsp;</p> <h3><strong>Kein Urheberrecht: Public Domain</strong></h3> <p>An den Entscheidungen und Metadaten besteht gem. &sect; 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <h3><strong>Disclaimer</strong></h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Beh&ouml;rden, Gerichten oder anderen &ouml;ffentlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <h3><strong>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</strong></h3> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <h3><strong>Kontakt</strong></h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im <a href="https://github.com/SeanFobbe/ce-bfh/issues">Issue Tracker auf GitHub</a> oder schreiben Sie mir eine E-Mail an <a href="mailto:fobbe-data@posteo.de">fobbe-data@posteo.de</a></p> <p>&nbsp;</p>

opencc-zeroOct 2023View details →
zenodo44/100

FoodSky: A Food-oriented Large Language Model, and FoodEarth: A Foundamental Food Corpus and Instruction Dataset

<p>Food is the cornerstone of both survival and social life. With the increasing complexity of global dietary needs and preferences, there is a growing demand for food intelligence to enable tasks like recipe recommendation and diet-disease correlation discovery. To address this, we introduce the Food-oriented Large Language Model (LLM) FoodSky, which offers fine-grained perception and reasoning of food data. We constructed a food corpus, FoodEarth, from various authoritative sources to enhance FoodSky's knowledge. We also developed the Topic-based Selective State Space Model and Hierarchical Topic Retrieval Augmented Generation algorithms to improve FoodSky's ability to capture fine-grained food semantics and generate context-aware food-relevant text. Extensive experiments show that FoodSky outperforms general-purpose LLMs on the Chinese National Chef Exam and Dietetic Exam, achieving accuracies of 67.2% and 66.4%, respectively. FoodSky not only enhances culinary creativity and promotes healthier eating patterns but also establishes a new standard for domain-specific LLMs tackling real-world food-related issues.</p>

opencc-zeroSep 2024View details →
zenodo44/100

Corpus der Entscheidungen des Bundesgerichtshofs (CE-BGH)

<h3>&Uuml;berblick</h3> <p>Das <strong>Corpus der Entscheidungen des Bundesgerichtshofs (CE-BGH)</strong> ist der bislang gr&ouml;&szlig;te, frei verf&uuml;gbare Datensatz von Entscheidungen des Bundesgerichtshofs. Er ist eine Zusammenstellung aller Entscheidungen ab 2000, die in der <a href="https://www.bundesgerichtshof.de">amtlichen Datenbank des Bundesgerichtshofs</a> am jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die PDF-Sammlung.</p> <p>F&uuml;r Praktiker:innen stelle ich zus&auml;tzlich nach Senat sortierte PDF-Sammlungen aller <em>Leitsatzentscheidungen</em> und aller <em>Entscheidungen mit Namen</em> (z.B. &raquo;Trabrennbahn&laquo;) zur Verf&uuml;gung.</p> <p>Ab Version 2023-03-10 ist auch das Zitationsnetzwerk des BGH (Aktenzeichen, BGHZ und BGHSt) f&uuml;r die einfache Nutzung mit graphischer Software wie&nbsp;<a href="https://gephi.org/">Gephi</a> oder f&uuml;r die maschinelle Weiterverarbeitung als GraphML verf&uuml;gbar. Das Zitationsnetzwerk enth&auml;lt ca. 600.000 Zitate und ca. 100.000 Knoten (Aktenzeichen, BGHZ oder BGHSt).</p> <p>Die strafrechtlichen Entscheidungen des BGH von 1950 bis 1999 finden Sie im Datensatz <a href="https://doi.org/10.5281/zenodo.4540376"><strong>Entscheidungen des Bundesgerichtshofs in Strafsachen aus dem 20. Jahrhundert (BGH-Strafsachen-20Jhd)</strong></a>.</p> <p>&nbsp;</p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird <em>1-2 mal im Jahr</em> aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <h3>NEU in Version 2025-04-07</h3> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten</li> <li>&Uuml;berarbeitung der Dokumentation zu den Varianten des Datensatzes</li> <li>Expliziter R Package Version Lock f&uuml;r 2024-06-13 (CRAN Date)</li> <li>&Uuml;berarbeitung des Dockerfiles</li> <li>Vereinfachung der Run-Skripte und st&auml;rkere Integration mit Docker Compose</li> <li>Vereinheitlichung der Berechnung kryptographischer Hashes</li> <li>/tmp in Arbeitsspeicher ausgelagert</li> <li>Entfernung von exakten Prozentzahlen in den Frequenztabellen</li> <li>Entfernung der Tesseract System Library</li> <li>Entfernung der Nummerierung des Workflow-Diagramms</li> </ul> <p>&nbsp;</p> <h3>Features</h3> <ul> <li>Insgesamt bis zu 36 Variablen in der CSV-Variante</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV)</li> <li>Zitationsnetzwerk zwischen allen Aktenzeichen, BGHZ und BGHSt</li> <li>Verkn&uuml;pfung mit Pr&auml;sidentIn/Vize-Pr&auml;sidentIn</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'ANALYSE')</li> <li>Diagramme liegen jeweils in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor</li> <li>Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4459415">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <p>&nbsp;</p> <h3>Eckdaten</h3> <p><em>Stichtag:</em> 7. April 2025</p> <p><em>Inhaltlicher Umfang</em>: 79.708 Entscheidungen</p> <p><em>Zeitlicher Umfang:</em> 2000 bis 2025</p> <p><em>Formate:</em><strong> </strong>PDF, TXT, CSV und GraphML</p> <h3>&nbsp;</h3> <h3>Source Code und Compilation Report</h3> <p>Der gesamte Erstellungs-Prozess ist ab Version 2021-04-27 vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Robustness Checks auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Source Code f&uuml;r die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="https://doi.org/10.5281/zenodo.4459415">https://doi.org/10.5281/zenodo.4459415</a></p> <p>&nbsp;</p> <h3>Kryptographische Signaturen</h3> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p>&nbsp;</p> <h3>Kein Urheberrecht: Public Domain</h3> <p>An den Entscheidungstexten und amtlichen Leits&auml;tzen besteht gem. &sect; 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <h3>Disclaimer</h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Beh&ouml;rden, Gerichten oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <h3>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</h3> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <h3>Kontakt</h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im Issue Tracker auf Codeberg oder kontaktieren Sie mich &uuml;ber <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>&nbsp;</p>

opencc-zeroJul 2020View details →
zenodo44/100

Kobalt: Extension Corpus and Annotation Guidelines for Verb Classification and Dependency Adjustments

<p>Kobalt (Zinsmeister et al. 2012) is a task-based corpus of essays written by learners and native speakers of German. This repository contains data that was not included in the original corpus and new layers of annotation to the original and the extended corpus, specifically morphological and syntactic classification of verbs and corrections and changes to dependency parses. Please refer to the annotation guidelines included in this repository for further information.<br> &nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Inputlog Copy Task Corpus: Exploring and defining typing skills

<p><strong>Context</strong></p> <p>One of the components that is included in the keystroke logging program Inputlog (<a href="https://www.inputlog.net">https://www.inputlog.net</a>) is the Copy Task component. It consists of a multi-layered set of tasks that measure a person&#39;s typing skill:</p> <table> <tbody> <tr> <td>Tapping task</td> <td>press the &lsquo;d&rsquo; and &lsquo;k&rsquo; key alternatively during 15 s</td> </tr> <tr> <td>Sentence</td> <td>copy a sentence during 30 s</td> </tr> <tr> <td>Word combination 1</td> <td>copy a combination of three words seven times</td> </tr> <tr> <td>Word combination 2</td> <td>copy a combination of three words seven times</td> </tr> <tr> <td>Word combination 3</td> <td>copy a combination of three words seven times</td> </tr> <tr> <td>Word combination 4</td> <td>copy a combination of three words seven times</td> </tr> <tr> <td>Consonant groups</td> <td>copy four blocks of six consonants once</td> </tr> </tbody> </table> <p>The task is currently made available in twelve languages.&nbsp;</p> <p>For more information:&nbsp;<a href="https://doi.org/10.5334/jors.234 ">https://doi.org/10.5334/jors.234&nbsp;</a></p> <p>&nbsp;</p> <p><strong>Interactive Dashboard</strong><br> Visit the webpage with an interactive dashboard to explore, filter, and download the +5K copy task corpus.</p> <p><em><strong>website</strong></em>:&nbsp;<a href="https://www.inputlog.net/copy-task/">https://www.inputlog.net/copy-task/</a><br> <em><strong>dashboard</strong></em>:&nbsp;<a href="https://inputlog-analysis.uantwerpen.be/expert">https://inputlog-analysis.uantwerpen.be/expert</a></p> <p>&nbsp;</p> <p><strong>Corpus</strong></p> <p>We are happy to make a multilingual corpus available (open access) that currently consists of more than 5000 copy tasks.&nbsp;</p> <ul> <li>The + 5K corpus is carefully cleaned and fully anonymized.</li> <li>The Shiny interface allows users to filter the corpus based on about 10 variables.</li> <li>The selection can be downloaded in different formats and levels of aggregation (from raw idfx to synthesized analysis).</li> <li>The selection can be explored using different interactive graph visualizations.</li> <li>Researchers can upload their own corpus (or single copy task file) and compare it to the (selected) corpus.</li> <li>An extra webpage is designed for laypersons wanting to take a copy task to test their typing skills. They get dashboard feedback in a user-friendly and attractive way and can compare their performance with (age-related) participants in the corpus. (Specially designed to further expand the corpus).</li> </ul> <p><strong>Facts and Figures</strong><br> Some facts and figures about the corpus&#39; composition:</p> <p>Languages:</p> <ul> <li>Dutch&nbsp; &nbsp; &nbsp;3130 files</li> <li>English&nbsp; &nbsp; 1163 files</li> <li>German&nbsp; &nbsp; &nbsp;281 files</li> <li>French&nbsp; &nbsp; &nbsp; &nbsp;201 files</li> <li>Other&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;378 file</li> </ul> <p><strong>Gender</strong></p> <ul> <li>Female:&nbsp; &nbsp; &nbsp; &nbsp;3495 files</li> <li>Male:&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;1276 files</li> <li>X or missing&nbsp; 382 files</li> </ul> <p><strong>Age</strong></p> <ul> <li>15-&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;439 files</li> <li>16-20&nbsp; &nbsp;1591 files</li> <li>21-25&nbsp; &nbsp; 2427 files</li> <li>26-35&nbsp; &nbsp; &nbsp; 478 files</li> <li>36-45&nbsp; &nbsp; &nbsp; 126 files</li> <li>46+&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 230 files</li> </ul> <p>A subset of the total corpus has been uploaded here. The subset contains a dataset of about 500 tests (English | 21-25-year-olds).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

The Makerere Gendered Corpus: A Gendered English to Luganda Parallel Corpus

<p>This&nbsp;English-Luganda parallel sentence corpus consists of&nbsp;gendered examples created by a team of researchers from Makerere AI&nbsp; Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under CC-0 and translating these sentences using a crowdsourcing, iterative and opensource approach was done using Pontoon an opensource Translation Management System built by Mozilla. This is a corpus of&nbsp;1,000 parallel sentences.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Corpus of political tweets UK-EU-DEBATE-20-21

<p>&nbsp;</p> <p>The&nbsp;<em>UK-EU-DEBATE-20-21</em>&nbsp;corpus was collected within the framework of the collaborative research project OLiNDiNUM (<em><a href="https://olindinum.huma-num.fr">Observatoire LINguistique du DIscours NUM&eacute;rique</a> /&nbsp;</em>Linguistic Observatory of Online Debate)&nbsp;to be part of a shared research archive of shared corpora and resources.&nbsp;</p> <p>The corpus was selected&nbsp;with a view to examining the UK-EU media debate on the COVID-19 vaccination campaign following a specific transformative moment: the signature of the Brexit withdrawal agreement by the UK and the EU at the end of January 2021.</p> <p>The data were retrieved through the&nbsp;Application Programming Interface&nbsp;of the social networking site Twitter, using the accounts of key political actors in the UK government and EU institutions&nbsp;over a period of 14 months (1 February 2020&ndash;31 March 2021).&nbsp;The composition of the&nbsp;corpus is illustrated in the table.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><em>Political Actor</em></td> <td><em>Role</em></td> <td><em>Account</em></td> <td><em>Tweets</em></td> </tr> <tr> <td>Boris Johnson</td> <td>UK Prime Minister</td> <td>@BorisJohnson</td> <td>1186</td> </tr> <tr> <td>Dominic R. Raab</td> <td>UK Foreign Secretary</td> <td>@DominicRaab</td> <td>1468</td> </tr> <tr> <td>Priti Patel</td> <td>UK Home Secretary</td> <td>@pritipatel</td> <td>941</td> </tr> <tr> <td>Ursula von der Leyen</td> <td>President of the European Commission</td> <td>@vonderleyen</td> <td>1338</td> </tr> <tr> <td>David Sassoli</td> <td>President of the European Parliament</td> <td>@EP_President</td> <td>554</td> </tr> <tr> <td>Charles Michel</td> <td>President of the Council of the European Union</td> <td>@eucopresident</td> <td>675</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>The data are supplied in separate .csv files (tab-delimited format). Each row contains the text of the&nbsp;tweet (<em>data__text</em>) and the tweet identifier (<em>data__id</em>) as a header.&nbsp;The tweet identifier enables swift retrieval of the original tweet by searching&nbsp;https://twitter.com/anyuser/status/<em>data__id.&nbsp;</em></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Oral cancer speech corpus for paper "Detecting and analysing spontaneous oral cancer speech in the wild"

<p>This is the oral cancer speech corpus used in the paper <em>&quot;Detecting and analysing spontaneous oral cancer speech in the wild&quot;.</em></p> <p><strong>Description</strong></p> <p>This dataset contains approximately 3 hours of oral cancer speech data collected from YouTube, including a file with additional metadata. We use this dataset to perform an oral cancer speech detection task in our paper.</p> <p><strong>Funding</strong></p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under Marie Sklodowska-Curie grant agreement No 766287. The Department of Head and Neck Oncology and surgery of the Netherlands Cancer Institute receives a research grant from Atos Medical (Horby, Sweden),<br> which contributes to the existing infrastructure for quality of life research.</p> <p><strong>Citation:</strong></p> <p>If you use this dataset please cite:</p> <pre><code>@misc{halpern2020detecting, title={Detecting and analysing spontaneous oral cancer speech in the wild}, author={Bence Mark Halpern and Rob van Son and Michiel van den Brekel and Odette Scharenborg}, year={2020}, eprint={2007.14205}, archivePrefix={arXiv}, primaryClass={eess.AS} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

SLNET: A Redistributable Corpus of 3rd-party Simulink Models

<p>MSR 2022 Data and Tool Showcase Track (https://conf.researchr.org/track/msr-2022/msr-2022-data-showcase)</p> <p>SLNET is collection of third party Simulink models. It is curated via mining open source repository (GitHub and Matlab Central) using SLNET-Miner (https://github.com/50417/SLNet_Miner)</p> <p>Paper:&nbsp;https://ranger.uta.edu/~csallner/papers/Shrestha22SLNET.pdf</p> <p>The metrics are collected using https://github.com/50417/SLNET_Metrics</p> <p>v2.0: Added database schema&nbsp;</p> <p>Note: SRC_block_type tables houses &quot;block type&quot; and its count per model. For many blocktypes such as MATLAB Function (which are actually built using a Subsystem), Simulink&#39;s get_param API returns block type as &quot;Subsystem&quot; instead of MATLAB Function. Thus,&nbsp;&quot;Subsystem&quot; block type include counts both from user created ones and library blocks (which are implemented using Subsystem). Thus To get a precise number for user created Subsystem, refer to &quot;Agg_subsystem_count&quot; in SRC_metric table.&nbsp;&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

LADDER. Learners' digital communication: a corpus for pragmatic competences in Italian L1/L2

<p>&nbsp;</p> <p><strong>Ladder</strong>. A Corpus of Computer-Mediated Communication for the Analysis of the Acquisition of Pragmalinguistic Competences by German-Speaking Learners of Italian.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Project description:</p> <p>Many recent research projects (Artoni, Benigni, &amp; Nuzzo, 2020; Cort&eacute;s Vel&aacute;squez &amp; Nuzzo, 2017; Nuzzo &amp; Cort&eacute;s Vel&aacute;squez, 2020) have underlined the usefulness of creating and analyzing corpora for teaching pragmatics, which, unlike other linguistic levels such as syntax, cannot be explained by rules but only by reference to tendential values or more or less appropriate choices in a given context. This is even more true for interactions via digital media, such as email and instant-messaging services, which have little place in manuals or L2 courses and for which learners have few reference models (Brocca, 2021; Trubnikova &amp; Garofolin, 2020).</p> <p>Data collection:</p> <p>Data were collected from April 2020 to April 2021 with the help of a discourse completion task (DCT). The data consists of emails and instant messages. The informants are (i) German learners of Italian between A2-C1 level according to the CEFR and most of them are students living in Tyrol (Austria) and (ii) native speakers of Italian most of whom are students from Rome (Italy). The data of the learners were collected by students of the undergraduate seminar &ldquo;Insegnare la pragmatica&rdquo; which is part of the compulsory module 2b for student teachers at the Institute of Didactics of the University of Innsbruck. The data of the native speakers were collected in large part from students in foreign languages at the University RomaTre thanks to the collaboration with Prof. Elena Nuzzo.</p> <p>The DCTs have been conducted with online questionnaires. Along with the texts, metadata were also registered with the help of an online questionnaire giving sociolinguistic information about the informant (age, self-assessed language level, place of residence, native language, etc.). The DCTs aim to elicit linguistic acts of request and refusal in increasing levels of social distance and different media (Taguchi &amp; Roever, 2017, pp. 85, 231; Hinger et al. 2018: 148). The DCTs elicit different speech acts (requests and refusals) with different degrees of formality (study/work or free time), directed at different people (lecturer, friend, boss) and in different media (mail or instant messaging). The scenarios represent authentic circumstances for the students. The following table shows the situations that were studied:</p> <p>&nbsp;</p> <p><strong>Email</strong></p> <p>high level of social distance between sender and recipient</p> <p>Scenario 1: Sender is asking for something that he/she is not entitled to</p> <p>Scenario 2: Sender is asking for something that he/she is entitled to</p> <p><strong><em>WhatsApp</em></strong><strong> messages</strong></p> <p>a) low level of social distance between sender and recipient</p> <p>Scenario 1: Request</p> <p>Scenario 2: Rejecting a request</p> <p>Scenario 3: Short-notice cancellation of an invitation</p> <p>b) medium level of social distance between sender and recipient</p> <p>Scenario 4: Request</p> <p>Scenario 5: Rejecting a request</p> <p>Scenario 6: Short-term rejection of an invitation</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The <em>WhatsApp</em> messages, which are exemplary of the text type instant messaging, were produced directly with the cell phone. The metadata were subsequently associated with the respective messages in an Excel spreadsheet. All personal data were anonymized.</p> <p>The prompts were presented in Italian, as follows:</p> <p>Mail</p> <p><strong>Mail a)</strong> Immagina di star facendo un corso con il Dr. Nicola Brocca. Domani devi fare una presentazione in classe. Non hai avuto tempo per studiare perch&eacute; dovevi prepararti a un esame di inglese e ti accorgi che il materiale da presentare &egrave; pi&ugrave; di quello che avevi previsto. Scrivi una mail al professore: la tua speranza &egrave; spostare la presentazione.</p> <p>Engl: &nbsp;Imagine you are taking a course with Dr. Nicola Brocca. Tomorrow you have to give a presentation in class. You had no time to study because you had to prepare for an English exam, and you realize that there is more material to present than you had imagined. You write an email to the professor: your hope is to reschedule the presentation.</p> <p><strong>Mail b)</strong> Hai fatto un corso con il Dr. Brocca. Hai consegnato il tuo portfolio il 01.02.2020 adesso &egrave; il 01.03.2020 e non hai ancora ricevuto il voto. Ti serve il voto per registrarti per una borsa di studio. Manda una mail al prof.: il tuo obiettivo &egrave; ricevere il voto al pi&ugrave; presto</p> <p>Engl: You have taken a course with Dr. Brocca. You turned in your portfolio on 02/01/2020, it is now 03/01/2020 and you have not received the grade yet. You need the grade to register for a scholarship. Send an email to the professor: your goal is to receive the grade as soon as possible.</p> <p>&nbsp;</p> <p><em>WhatsApp</em> messages</p> <p><strong>1.</strong> Sei in Erasmus in Italia. Avete creato una chat con 10 compagni di corso. Hai perso la tua tessera della biblioteca a vuoi chiedere se qualcuno ti pu&ograve; aiutare perch&eacute; ti serve un libro entro domani...per esempio prestandoti la sua. Cosa scrivi?</p> <p>Engl: You are taking part in the Erasmus program in Italy. You have created a chat with 10 classmates. You lost your library card and want to ask if someone can help you because you need a book by tomorrow.... E.g. by lending you their card. What do you write?</p> <p><strong>2.</strong> Ricevi questo messaggio da un amico/a che fa un seminario con te: &quot;Ciao, sono a corto di tempo. Ho visto che hai preso 30 all&#39;esame. Potresti darmi una mano e restare con me in biblioteca oggi?&quot; Non vuoi aiutare il tuo amico. Come reagisci?</p> <p>Engl: You receive this message from a friend who is attending a seminar with you: &quot;Hello, I&#39;m running out of time. I saw that you got a 30 on the exam. Could you help me and stay with me in the library today?&quot; You don&#39;t want to help the friend. How do you respond?</p> <p><strong>3.</strong> Cinque giorni fa hai promesso ad un/a amico/a che questa sera sareste andati al cinema assieme. Per&ograve; hai cambiato idea. Cosa fai? Cosa scrivi?</p> <p>&nbsp;Engl:&nbsp; Five days ago, you promised a friend that tonight you would go to the movies together. But you changed your mind. What would you do? What do you write?</p> <p>&nbsp;</p> <p><strong>4.</strong> Sei al lavoro e hai smarrito il documento elettronico per entrare nel parcheggio. Sei nuovo in questo gruppo di lavoro e hai solo il numero del tuo diretto superiore. Gli mandi un messaggio per chiedergli se ti pu&ograve; aiutare.</p> <p>Engl: You are at work and have lost your electronic badge to enter the parking lot. You are new to this work group and only have the number of your direct supervisor. You send him/her a message and ask if he/she can help you.</p> <p>&nbsp;</p> <p><strong>5.</strong> Ricevi questo messaggio dal/la tuo/a superiore. &quot;Gentile collega, domani c&#39;&egrave; una scadenza importante. Per caso sarebbe in grado di restare oggi in ufficio oltre l&#39;orario?&quot; Non vuoi restare in ufficio oltre il normale. Come reagisci?</p> <p>Engl: You receive this message from your supervisor. &quot;Dear colleague, tomorrow is an important appointment. Would you be able to stay in the office after hours today?&quot; You don&#39;t want to stay in the office beyond normal working hours. How do you respond?</p> <p>&nbsp;</p> <p><strong>6.</strong> Cinque giorni fa hai promesso al/la tuo/a superiore che oggi saresti andato a una cena di lavoro. Per&ograve; devi disdire. Cosa fai?</p> <p>Engl: Five days ago, you promised your superior that you would go to a business dinner today. However, you have to cancel. What do you do?</p> <p>&nbsp;</p> <p>The corpus, which was first collected in .xlsx format, was exported to XML format and CSV format in cooperation with Joseph Wang-Kathrein (Brenner Archive Research Center). It was ensured that the emoticons and special characters were also transferred unchanged in the conversion process. These formats allow long-term archiving and significantly facilitate data exchange.</p> <p>The size of the corpus (as of May 2021, version Ladder 1.0):</p> <p>The LADDER corpus includes emails and instant-messaging messages amounting to 18,935 tokens and 33,966 tokens respectively. The corpus of <em>WhatsApp</em> messages consists of a total of 1,204 messages from 80 native speakers and 114 learners. The corpus of emails consists of a total of 235 emails from 78 native-speaker informants and 38 learners. The amount of data allows a qualitatively relevant comparison in sub-corpora e.g. language levels.</p> <p>The size of the corpus is necessarily limited quantitatively, as data collection must be done manually through individual DCT management and metadata checking. The major bottleneck is currently the annotation of socio-pragmatic aspects, a process that is difficult to automate and that needs to be conducted through cross-annotation by multiple annotators.</p> <p>Some students&#39; works on the corpus have been collected and are accessible via the following link: https://ladder.hypotheses.org/</p> <p>&nbsp;</p> <p>Bibliography:</p> <p>Artoni, D., Benigni, V., &amp; Nuzzo, E. (2020), &quot;Pragmatic instruction in L2-Russian: a study on requests and advice&quot; in <em>Instructed Second Language Acquisition, 4</em>(1), 62-95. doi:10.1558/isla.39864</p> <p>Brocca, N. (2021), &quot;LADDER: La costruzione e analisi di un corpus di scritture digitali per l&rsquo;insegnamento della pragmatica in L2&quot; in <em>Italiano Lingua Due, 13</em>(1 (2021)).</p> <p>Cort&eacute;s Vel&aacute;squez, D., &amp; Nuzzo, E. (2017), &quot;Disdire un appuntamento: spunti per la didattica dell&#39;italiano L2 a partire da un corpus di parlanti nativi&quot; in <em>Italiano Lingua Due, 1</em>, 17-36.</p> <p>Hinger, B., Stadler, W., Schmiderer, K., Bauer, M., (Hrg.) (2018). Testen und Bewerten fremdsprachlicher Kompetenzen. T&uuml;bingen: Narr Francke Attempto Verlag.</p> <p>Nuzzo, E., &amp; Cort&eacute;s Vel&aacute;squez, D. (2020), &quot;Canceling Last Minute in Italian and Colombian Spanish: A Cross-Cultural Account of Pragmalinguistic Strategies&quot; in <em>Corpus Pragmatics, 4</em>, 1-26. doi:10.1007/s41701-020-00084-y</p> <p>Taguchi, N., &amp; Roever, C. (2017), <em>Second language pragmatics</em>: Oxford: Oxford University Press.</p> <p>Trubnikova, V., &amp; Garofolin, B. (2020), <em>Lingua e interazione. Insegnare la pragmatica a scuola</em>. Pisa: ETS.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

NewsCom-NEG Corpus

<p>The&nbsp;NewsCom&nbsp;corpus consists of 2955 comments posted in response to 18 different news articles&nbsp;obtained from online Spanish newspapers from August 2017 to May 2019. These news articles cover nine topics (two articles per topic): immigration, politics, technology, terrorism, economy, society, religion, refugees, and real estate.&nbsp;The&nbsp;NewsCom&nbsp;corpus contains 2965 negative structures with their corresponding negation marker, scope, and focus.&nbsp;It is a valuable resource that can be used both for the training and evaluation of systems that aim to automatically detect the scope and focus of negation and for the linguistic analysis of negation grounded in real data.</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record