Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
878
datasets available to search
ShareScore release 0.9.0
Dataset results
878 results for “Corpus”
Polifonia Corpus - Books Module Metadata - French Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Dutch Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - German Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Spanish Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Italian Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
BiTe_Corpus
<p>[ENGLISH] The <em>BiTe_Corpus</em> is a text file (txt.) containing the abstracts published in Spanish collected in the <em><a href="https://buske.de/bibliografia-tematica-de-historiografia-linguistica-espanola.html?___store=buske_english&___from_store=buske_english">Bibliografía Temática de Historiografía Lingüística Española: fuentes secundarias</a></em> [Esparza et al. 2008]. It is a document containing 102613 words and 9270 unique words and it is especially conceived for the metahistoriographical study of the history of Hispanic linguistics.</p> <p>This final corpus was the result of two stages. First, the abstracts that were part of the bibliographic records of the Bibliografía Temática de la Historiografía Lingüística Española. Fuentes secundarias (BiTe) (Esparza et al. 2008). At the beginning of this first stage, the corpus consisted of a total of 296 029 words and 31469 vocablos (unique word occurrences) distributed in 2298 abstracts. Secondly, we proceeded to a second edition taking as a starting point the criteria published by Samper Padilla (1998) and maintaining a conservative stance -as opposed to a uniform one- (Fernández Juncal 2013) when making different decisions about the edition. More information about the editing criteria can be requested from the authors.</p> <p>[SPANISH] El <em>BiTe_Corpus</em> es un documento de texto que contiene los resúmenes publicados en español reunidos en la Bibliografía Temática de Historiografía Lingüística Española: fuentes secundarias [Esparza et al. 2008]. Se trata de un documento que contiene 102613 palabras y 9270 vocablos únicos y está especialmente concebido para el estudio metahistoriográfico de la historia de la lingüística hispánica.</p> <p>Este corpus final fue el resultado de dos etapas. En primer lugar, se reunieron los resúmenes que formaban parte de las fichas bibliográficas de la <a href="https://buske.de/bibliografia-tematica-de-historiografia-linguistica-espanola.html?___store=buske_english&___from_store=buske_english">Bibliografía Temática de la Historiografía Lingüística Española. Fuentes secundarias</a> (BiTe) (Esparza et al. 2008). Al inicio de esta primera etapa, el corpus constaba con un total de 296 029 palabras y 31469 vocablos (apariciones únicas de palabras) distribuidos en 2298 resúmenes. En segundo lugar, se procedió a una segunda edición tomando como punto de partida los criterios publicados por Samper Padilla (1998) y manteniendo una postura conservadora –frente a la uniformadora­– (Fernández Juncal 2013) a la hora de tomar diferentes decisiones sobre la edición. Puede solicitar más información acerca de los criterios de edición a las autoras.</p>
Polifonia Mini Textual Corpus
<p>The Polifonia Mini Textual Corpus is made of a balanced selection of sentences representative of the different periods and styles of the corpus. It was created to perform the transformation of sentences into AMR graphs. The reason why the Polifonia Mini Textual Corpus is made of sentences is twofold: (i) the Polifonia Textual Corpus' documents' length prevents them from fitting into GPU memory; (ii) AMR parsers' models are trained to transform no more than a few sentences into AMR graphs; therefore, their performance on large documents parsing is unreliable.</p> <p>Creating a sub-sample of sentences representative of all the modules of the Polifonia Textual Corpus allows for performing iterative validations of the text2AMR pipeline output. For example, it facilitates testing the algorithms developed to minimise the loss of information that segregating texts into sentences inherently brings (such as the loss of co-reference information). It also favours a more manageable Quality Assurance of the AMR graphs produced by the text2AMR models.</p> <p>When the results obtained on processing the Polifonia Mini Textual Corpus are considered satisfactory, we will apply the text2AMR pipeline to the Polifonia Textual Corpus in its entirety.</p> <p>Full description at: <a href="https://github.com/polifonia-project/Polifonia-Knowledge-Extractor">https://github.com/polifonia-project/Polifonia-Knowledge-Extractor</a></p>
OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan
<p>OcWikiDisc is a freely available corpus in Occitan, extracted from the talk pages associated with the Occitan Wikipedia.</p> <p>The corpus contains messages posted by users in direct user-to-user interactions as part of the discussions about the content and the editing policies on Wikipedia. The messages are associated with metadata, such as the username, the date and time of the posting, the discussion title, etc. The corpus has also been annotated with tools for automatic language identification, allowing to filter out content in languages other than Occitan. Using different filtering strategies, four versions of the corpus are published (see documentation for more details). The version with the most restrictive filtering contains 8,000 messages for a total of 618,000 tokens, produced by 520 different users.</p>
InVID Fake Video Corpus v2.0
<p>The InVID TV Fake Video Corpus was developed in the context of the InVID project with the aim of gaining a perspective of the types of fake video that can be encountered in the real world. The dataset does not aspire to serve as an exhaustive list of all forgeries that have circulated the Web in the past, but we intend to maintain and extend it throughout the course of the project as new cases arise. </p> <p>The collection is a collaborative effort between AFP and CERTH-ITI. This is the second version of the dataset, containing 117 fake videos and 110 real videos, alongside annotations and descriptions. As we do not own the rights to the videos, the dataset only contains the video URLs and annotations.</p>
Corpus of Resolutions: UN Security Council (CR-UNSC)
<h2>Overview</h2> <p>The <strong>Corpus of Resolution: UN Security Council (CR-UNSC)</strong> collects and presents for the first time in human and machine-readable form all resolutions, drafts, and meeting records of the UN Security Council, including detailed metadata, as published by the <a href="https://digitallibrary.un.org/">UN Digital Library</a> and revised by the authors.</p> <p>The United Nations Security Council (UNSC) is the most influential of the principal UN organs. Composed of five permanent and ten non-permanent members, its functioning is constrained by the political context in which it operates. During the Cold War, the complex political relationships between the permanent members and their veto powers significantly affected the capacity of the UNSC to address violations of international peace and security, with only 646 resolutions passed from 1946 to 1989. Since the 1990s, the activity of the UN Security Council has increased dramatically and produced 2721 resolutions up to the end of 2023. The length, complexity and thematic breadth of the resolutions has also increased, prompting calls to redefine it as a quasi-legislative body.</p> <p>Under Articles 24 and 25 of the UN Charter, member states have conferred upon the UNSC the "primary responsibility for the maintenance of international peace and security" and have agreed "to accept and carry out" its decisions. The discharge of this function is carried out through the powers bestowed upon it under Chapter VI of the UN Charter, "Pacific Settlement of Disputes", Chapter VII, "Action with Respect to Threats to the Peace, Breaches of the Peace, and Acts of Aggression", Chapter VIII, "Regional Arrangements", and Chapter XII, "International Trusteeship System". </p> <p>Under the peace and security mandate, its areas of activity cover disarmament, pacific settlement of disputes, enforcement, and, until 1994, strategic areas in a trusteeship agreement. Its functions also pertain to the correct working of the United Nations, covering issues of membership, the appointment of the Secretary General, the elections of judges of the International Court of Justice (ICJ), the calling of special and emergency sessions of the General Assembly, the amendment of the Charter and of the ICJ Statute.</p> <p><strong>Please refer to the Codebook for a detailed explanation of the dataset and instructions on how to make use of it.</strong></p> <p> </p> <h2>Updates</h2> <p>The CR-UNSC will be updated at least once per year.</p> <p>In case of serious errors an update will be provided at the earliest opportunity and a highlighted advisory issued on the Zenodo page of the current version. Minor errors will be documented in the GitHub issue tracker and fixed with the next scheduled release.</p> <p>The CR-UNSC is versioned according to the day of the last run of the data pipeline, in the ISO format YYYY-MM-DD. Its initial release version is 2024-05-03.</p> <p>Notifications regarding new and updated data sets will be published on my academic website at www.seanfobbe.com or on the Fediverse at @seanfobbe@fediscience.org</p> <p> </p> <h2>Changelog</h2> <ul> <li>New variant: EN_TXT_BEST containing a write-out of the English resolution texts equivalent to the CSV file text variable</li> <li>New diagrams: bar charts of top M49 regions and sub-regions of countries mentioned in resolution texts</li> <li>Fixed naming mix-up of BIBTEX and GRAPHML zip archives</li> <li>Fixed whitespace character detection in citation extraction (adds ca. 10% more citations)</li> <li>Fixed improper merging of weights in citation network</li> <li>Fixed "cannot xtfrm data frames" warning</li> <li>Improve REGEX detection for certain geographic entities</li> <li>Improve Codebook (headings, citation network docs)</li> </ul> <p> </p> <h2>Key Metrics</h2> <p><em>Version:</em> 2024-05-19</p> <p><em>Scope:</em> UNSC Resolutions from 1 (1946) up to and including 2722 (2024)</p> <p><em>Tokens:</em> 3,704,016 (English resolution texts)</p> <p><em>Languages: </em>English, French, Spanish, Arabic, Chinese, Russian</p> <p> </p> <h2>Features</h2> <ul> <li>82 Variables</li> <li>Resolution texts in all six official UN languages (English, French, Spanish, Arabic, Chinese, Russian)</li> <li>Draft texts of resolutions in English</li> <li>Meeting record texts in English</li> <li>URLs to draft texts in all other languages (French, Spanish, Arabic, Chinese, Russian)</li> <li>URLs to meeting record texts in all other languages (French, Spanish, Arabic, Chinese, Russian)</li> <li>Citation data as GraphML (UNSC-to-UNSC resolutions and UNSC-to-UNGA resolutions)</li> <li>Bibliographic database in BibTeX/OSCOLA format for e.g. Zotero, Endnote and Jabref</li> <li>Extensive Codebook to explain the uses of the dataset</li> <li>Compilation Report and Quality Assurance Report explain construction and validation of the data set</li> <li>Publication quality diagrams for teaching, research and all other purposes (PDF for printing, PNG for web)</li> <li>Open and platform independent file formats (CSV, PDF, TXT, GraphML)</li> <li>Software version controlled with Docker</li> <li>Publication of full data set (Open Data)</li> <li><a href="../doi/10.5281/zenodo.7319783">Publication of full source code (Open Source)</a></li> <li>Data published under Public Domain waiver (CC Zero 1.0)</li> <li>Source Code is Free Software published under the GNU General Public License Version 3 (GNU GPL v3)</li> <li>Secure cryptographic signatures for all files in version of record (SHA2-256 and SHA3-512)</li> </ul> <p> </p> <h2>Recommended Variants</h2> <table> <tbody> <tr> <td><strong>Traditional Scholars</strong></td> <td> <p>ALL_PDF_Resolutions</p> <p>EN_TXT_BEST</p> <p>BIBTEX_OSCOLA</p> </td> </tr> <tr> <td><strong>Quantitative Scholars</strong></td> <td> <p>ALL_CSV_FULL</p> <p>EN_TXT_BEST</p> <p>CITATIONS_GRAPHML</p> </td> </tr> </tbody> </table> <p> </p> <p>Please refer to the Codebook regarding for details on each variant. The ZIP archives include texts in all languages, unless noted in the filename.</p> <p>We strongly recommend using the CSV files for quantitative analysis, but if you find CSV hard to use and want to analyze only the text of resolutions, the EN_TXT_BEST variant is a mix of expert-revised OCR and born digital texts equivalent to the "text" variable in the CSV file.</p> <p> </p> <h2>Compilation Report and Quality Assurance Report</h2> <p>With every compilation of the full data set, an extensive Compilation Report and detailed Quality Assurance Report are created and published in PDF format.</p> <p>The Compilation Report includes the source code for the pipeline architecture, comments and explanations of design decisions, relevant computational results, exact timestamps and a table of contents with clickable internal hyperlinks to each section.</p> <p>The Quality Assurance Report contains a count of all hard tests and expectations, additional visualizations and documented test results for all soft tests that require further interpretation</p> <p>The Compilation Report, Quality Assurance Report and Source Code are published under the following DOI: <a href="../doi/10.5281/zenodo.7319783">https://zenodo.org/doi/10.5281/zenodo.7319783</a></p> <p> </p> <h2>Attribution and Copyright</h2> <p>This data is derived from the United Nations Digital Library at <a href="https://digitallibrary.un.org">https://digitallibrary.un.org</a>. Records were accessed and downloaded on 13 and 26 March 2024, with additional work on revisions and corrections up to and including the date given as the version number.</p> <p>Pursuant to <a href="https://en.wikisource.org/wiki/Administrative_Instruction_ST/AI/189/Add.9/Rev.2">UN Administrative Instruction ST/AI/189/Add.9/Rev.2 of 17 September 1987</a> all official records and United Nations Documents (including resolutions, compilations of resolutions, drafts and meeting records) are in the public domain. We wish to honor the letter and spirit of this UN policy. To ensure the widest possible distribution of official UN documents and to promote the international rule of law we waive any copyright that might have accrued by creating the dataset under a <a href="https://creativecommons.org/public-domain/cc0/">Creative Commons CC0 1.0 Universal (CC0 1.0) Public Domain Dedication</a>. </p> <p> </p> <h2>Disclaimer</h2> <p>This data set is an academic initiative and is not associated with or endorsed by the United Nations or any of its constituent organs and organizations.</p> <p> </p> <h2>Author Websites</h2> <p><a href="https://www.seanfobbe.com">Personal Website of Seán Fobbe</a></p> <p><a href="https://www.santannapisa.it/en/lorenzo-gasbarri">Personal Website of Lorenzo Gasbarri</a></p> <p><a href="https://www.kcl.ac.uk/people/niccolo-ridi">Personal Website of Niccolò Ridi</a></p> <p> </p> <h2>Contact</h2> <p>Did you discover any errors? Do you have suggestions on how to improve the data set? You can either post these to the <a href="https://github.com/SeanFobbe/cr-unsc/issues">Issue Tracker on GitHub</a> or contact Seán Fobbe via <a href="https://seanfobbe.com/contact/">https://seanfobbe.com/contact/</a></p>
BuzzFeed-Webis Fake News Corpus 2016
<p>The corpus comprises the output of 9 publishers in a week close to the US elections. Among the selected publishers are 6 prolific hyperpartisan ones (three left-wing and three right-wing), and three mainstream publishers (see Table 1). All publishers earned Facebook’s blue checkmark, indicating authenticity and an elevated status within the network. For seven weekdays (September 19 to 23 and September 26 and 27), every post and linked news article of the 9 publishers was fact-checked by professional journalists at BuzzFeed. In total, 1,627 articles were checked, 826 mainstream, 256 left-wing and 545 right-wing. The imbalance between categories results from differing publication frequencies.</p>
Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)
<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from 'Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais' de J.-F. Bladé, 'Coundes biarnés, couéilhuts aüs parsàas miéytadès dou péys dé Biarn' de J.-V. Lalanne, 'Contes populaires du Languedoc' de L. Lambert and 'Contes populaires recueillis dans la Grande-Lande' de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>
CrowdTruth Corpus for Open Domain Relation Extraction from Sentences
<p>This repository contains a ground truth corpus for open domain relation extraction from sentences, acquired with crowdsourcing and processed with <strong><a href="http://crowdtruth.org/">CrowdTruth</a></strong> metrics that capture ambiguity in annotations by measuring inter-annotator disagreement.</p> <p>The dataset contains annotations for 4,100 sentences sampled from Angeli et al. (1) and Riedel et al. (2), over 16 relations, with each sentence annotated by 15 workers. The sentences have been pre-processed with Distant Supervision (3) using the Freebase knowledge base, in order to identify the term pairs in each sentence that are likely to express a relation. The crowdsourced data was collected from <a href="http://figure-eight.com/">Figure Eight</a> and <a href="https://www.mturk.com/">Amazon Mechanical Turk</a>.</p> <p>This corpus has been discussed in the following papers:</p> <ul> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1809.00537">Crowdsourcing Semantic Label Propagation in Relation Classification</a></strong>. <a href="http://fever.ai/">FEVER</a> Workshop at <a href="http://emnlp2018.org/">EMNLP 2018</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1711.05186">False Positive and Cross-relation Signals in Distant Supervision Data</a></strong>. <a href="http://www.akbc.ws/">AKBC</a> Workshop at <a href="http://nips.cc/">NIPS 2017</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="http://crowdtruth.org/wp-content/uploads/2017/03/collint17-open-domain.pdf">Disagreement in Crowdsourcing and Active Learning for Better Distant Supervision Quality</a></strong>. <a href="http://collectiveintelligenceconference.org/">Collective Intelligence 2017</a>.</li> </ul> <p>Sentence-level data is available in file: <code>|--data/output/aggregated_sentences.csv</code></p> <p>Worker-level data is available in file: <code>|--data/output/aggregated_workers.csv</code></p> <p>Raw crowdsourcig data is available in folder: <code>|--data/input/</code></p> <p>Results of the relation classification model are available in folder: <code>|--data/model_results/</code></p> <p> </p> <p>References</p> <p>(1) Angeli, Gabor, et al. "Combining distant and partial supervision for relation extraction." Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014.</p> <p>(2) Riedel, Sebastian, et al. "Relation extraction with matrix factorization and universal schemas." Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2013.</p> <p>(3) Mintz, Mike, et al. "Distant supervision for relation extraction without labeled data." Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2-Volume 2. Association for Computational Linguistics, 2009.</p>
The Rodrigo corpus
<p>The Rodrigo<em> </em>corpus was obtained from the digitisation of the book “Historia de España del arçobispo Don Rodrigo”, written in ancient Spanish in 1545. It is a single writer book where most pages consist of a single block of well-separated lines of calligraphical text.</p> <p>This dataset is free available for research purposes. It contains 15,010 images of text lines with their paleographic transcription. It is divided into three partitions: 9000 text lines for training, 1000 for validation and 5010 for testing.</p>
MiRoR11 - P2 - Annotated corpus for semantic similarity of clinical trial outcomes
<p>Outcome similarity corpus</p> <p>This dataset contains annotations of semantic similarity for pairs of primary and reported outcomes.<br> Tab-separated format is used. The files contain the following columns:<br> filename, sentence pair ID, sentence pair text, primary outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_relations_split contains the dataset splits for 10-fold cross-validation.</p>
MiRoR11 - P2 - Annotated corpus for the relation between reported outcomes and their significance levels
<p>Corpus of relations between outcomes and significance levels</p> <p>This dataset contains annotations of the relations between reported outcomes and their significance levels.<br> Tab-separated format is used. The file contains the following comumns:<br> filename, sentence text, outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_sig_rel contains the dataset splits for 10-fold cross-validation.</p>
The Sincere Apology Corpus (SinA-C)
<p>This repository contains the Sincere Apology Corpus (SinA-C). SinA-C is an English speech corpus of acted apologies in various prosodic styles created with the purpose of investigating the attributes of the human voice which convey sincerity.</p> <p>Thirty-two speakers were recorded in a studio at the Columbia Computer Music Center inside a sound-proof recording booth. Audio was recorded with an AKG C414 dynamic microphone. The digital audio workstation Logic Pro 9 was used to collect the audio signals. Recordings were captured at 44.1 kHz and 16 bit in AIFF format and later converted to mono WAV files.</p> <p><strong>Speakers </strong></p> <ul> <li>Gender: 15 male and 17 female</li> <li>Age: 20-60 years old (mean: 29.8 years; std; 9.9 years)</li> <li>Background: 27 American born English native speakers, and 5 from other nationalities (all fluent in spoken English). 24 speakers were professional actors, and the remaining 12 were artists</li> </ul> <p><strong>Recordings</strong></p> <p>Speakers were given a description of the study, a set of 6 sentences (apologies; see Table 1) and a short definition for a set of 4 prosodic styles (see Table 2) to adopt when uttering each sentence (the recordings are not spontaneous, but rather acted). The sentences used were the following:</p> <ul> <li><em>Sorry.</em></li> <li><em>I am sorry for everything I have done to you. </em></li> <li><em>I cannot tell you how sorry I am for everything I did</em>.</li> <li><em>Please allow me to apologise for everything I did to you. I was inappropriate and lacked respect</em>.</li> <li><em>It was never my intention to offend you, for this I am very sorry. </em></li> <li><em>I am sorry but I am going to have to decline your generous offer. Thank you for considering me.</em></li> </ul> <p>The prosodic styles intended to be adopted when uttering each of the sentences were:</p> <ul> <li>monotonic;</li> <li>pitch prominence (labelled as `Stress');</li> <li>fast speaking rate;</li> <li>slow speaking rate.</li> </ul> <p><strong>Annotations</strong></p> <p>The SinA-C audio recordings were labelled in terms of the sincerity perceived by listeners (`<em>How sincere was the apology you just heard?</em>') on a 5-point Likert scale ranging from 0 (Not Sincere) to 4 (Very Sincere) by 22 volunteers (13 male and 9 female; age range: 18-22; μ 19.5 std 1.0). Of the 22 annotators, all reported to have normal hearing, and all were English speakers (6 reported to be bilingual with at least one other language).</p> <p>Raw annotations were standardised to zero mean and unit standard deviation on a per-subject basis in order to eliminate potential individual rating biases. We then computed the mean across all subjects for each utterance. This resulted in a set of ratings ranging from [-1.51, 1.72] (mean -0.002 std 0.60) which are used as the gold-standard for regression experiments. We also converted these ratings to binary labels. Average ratings larger than 0 were labelled as `Sincere` (S), and those smaller of equal to 0 were labelled as `Not Sincere` (NS). This resulted in 478 instances labelled as S and 438 as NS. These labels are the gold-standard for classification tasks.</p> <p><strong>SinA-C Baseline</strong></p> <p>The Baseline for the dataset is described in detail in the INTERSPEECH 2019 publication "Sincerity in Acted Speech: Presenting the Sincere Apology Corpus and Results" [1]. This article presents both classification and regression baseline results. The modelling experiments included both a 3-fold Speaker Independent Nest Cross Validation (SICV) schema as well as Speaker Independent folds (C-SIF) (train, validation, test). C-SIF is provided by the original database baseline from the INTERSPEECH 2016 COMputation PARalinguistics challengE (COMPARE) [2]. For reproducibility, speaker distributions across the two partitioning strategies are provided with the corpus package.</p> <p>The audio descriptors include conventional and state-of-the-art features extracted from the audio files. We used Support Vector Machines (SVM) for classification tests and linear Support Vector Regression (SVR) for the regression ones. In both cases we used linear kernels and both SVM and SVR were implemented using the open-source machine learning toolkit Scikit-Learn. During the development phase, we trained various models (using the training set) with different complexity parameters (C ∈ 10-7, 10-6, 10-5, 10-4, 10-3, 10-2, 10-1, 1), and evaluated their performance on the validation set. After determining the optimal value for C, we concatenated the training and validation sets, re-trained the model with this enlarged training set, and evaluated the performance on the test set. Further detail on the baseline development are given in [1].</p> <p><strong>Comments</strong></p> <p>SinA-C was initially gathered between 2015-2016 at the Columbia University Computer Music Centre (CCMC) in New York City, United States of America. The dataset was also included in the INTERSPEECH 2016 COMPARE challenge [2], and prior to that in 2015 a subset of the dataset was also exhibited as part of a graduate-school art exhibition.</p> <p><strong>Citing this corpus</strong></p> <p>When using the data set for your own research, and within publications, please cite this repository and [1].</p> <p><strong>Bibliography</strong></p> <p>[1] Baird, A., Coutinho, E., Hirschberg, J., & Schuller, B. W. (2019). Sincerity in Acted Speech: Presenting the Sincere Apology Corpus and Results. In <em>Interspeech</em> <em>2019</em>, in press.</p> <p>[2] Schuller, B. W., Steidl, S., Batliner, A., Hirschberg, J., Burgoon, J. K., Baird, A., Elkins, A. C., Zhang, Y., Coutinho, E., & Evanini, K. (2016). The INTERSPEECH 2016 Computational Paralinguistics Challenge: Deception, Sincerity & Native Language. In <em>Interspeech 2016</em>, 2001-2005.</p>
Named Entity Corpus for Occupational Substance Exposure Assessment
<p>This is a corpus consisting of selected sections (i.e., <em>Abstract, Methods</em> and <em>Results</em>) of scientific research articles concerning occupational exposures to two different types of substance, i.e., diesel exhaust (51 articles) and respirable crystalline silica (RCS) (50 articles). The article sections have been annotated by experts in the field with 6 categories of named entities (NEs) relevant to the assessment of occupational substance exposures, particularly in the context of Job Exposure Matrices (JEMs).</p> <p>The corpus is available in two different formats, <a href="https://brat.nlplab.org/standoff.html">brat standoff format</a> and JSON. </p> <p>The corpus and associated NER models are described in more detail in the following atricle, which should be cited if you use the corpus: </p> <p>Thompson, P., Ananiadou, S., Basinas I., Brinchmann, B. C., Cramer, C., Galea, K. S., Ge, C., Georgiadis, P., Kirkeleit, J., Kuijpers, E., Nguyen, N., Nuñez, R., Schlünssen, V., Stokholm, Z. A., Taher, E. A., Tinnerberg, H., Van Tongeren, M. and Xie, Q. (2024).<a href="https://doi.org/10.1371/journal.pone.0307844"> </a><a href="https://doi.org/10.1371/journal.pone.0307844">Supporting the working life exposome: annotating occupational exposure for enhanced literature search</a>. PLoS ONE 19(8): e0307844</p>
Croatian Coronavirus News Comments Corpus News-CommHR
<p>A corpus of readers' news comments posted below news articles on the topic of the covid-19 pandemic, published in major Croatian daily newspapers and news portals in the six-month early pandemic period (March 2020 to September 2020).</p> <div>The corpus is designed to facilitate research on crisis discourses, crisis communication, as well as pandemic-time linguistic innovation. It is available in plain text version and XML with full metadata. The corpus complements a separate corpus of news articles Croatian Coronavirus Corpus NewsHR. Parallel versions from Slovenia and Serbia are also available.</div> <div> </div> <div>The project leading to this publication has received funding from the European Union’s Horizon 2020 research and innovation programme under the <a href="https://cordis.europa.eu/programme/id/H2020-EU.4./en">H2020-EU.4. - SPREADING EXCELLENCE AND WIDENING PARTICIPATION </a>programme Widening fellowships grant agreement No 101038047.</div>
Corpus des Deutschen Bundesrechts (C-DBR)
<p><strong>Überblick</strong></p> <p>Das <strong>Corpus des deutschen Bundesrechts (C-DBR)</strong> ist eine möglichst vollständige Sammlung der konsolidierten Fassungen aller Gesetze und Verordnungen auf Bundesebene. Der Datensatz nutzt als seine Datenquelle das amtliche Internetangebot <a href="http://www.gesetze-im-internet.de">www.gesetze-im-internet.de</a> des Bundesministeriums der Justiz und wertet dieses vollständig aus.</p> <p><em>Bitte lesen Sie zuerst das beiliegende Codebook!</em> Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die PDF-Sammlung.</p> <p>Um das <em>Gesetzgebungsverfahren</em> näher zu beleuchten können Sie zusätzlich auf folgende Datensätze zurückgreifen (jeweils mit Links auf vergleichbare Datensätze anderer Autor:innen):</p> <ul> <li><a href="http://doi.org/10.5281/zenodo.4643065">Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT)</a></li> <li><a href="http://doi.org/10.5281/zenodo.4542661">Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)</a></li> </ul> <p> </p> <p><strong>Aktualisierung</strong></p> <p>Dieser Datensatz wird <em>ca. alle 3 Monate</em> aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p> </p> <p><strong>NEU in Version 2025-10-02</strong></p> <ul> <li>Vollständige Aktualisierung der Daten</li> </ul> <p> </p> <p><strong>Eckdaten</strong></p> <p><em>Stichtag:</em> 2. Oktober 2025</p> <p><em>Umfang:</em> 6838 Bundesgesetze und -verordnungen der Bundesrepublik Deutschland</p> <p><em>Formate:</em> CSV, PDF, EPUB, TXT und XML</p> <p> </p> <p><strong>Features</strong></p> <ul> <li>Einfache Nutzung für statistische Analysen mit CSV-Dateien</li> <li>Bis zu 42 Variablen in den CSV-Varianten</li> <li>Fortlaufende Aktualisierung</li> <li>Urheberrechtsfreiheit</li> <li>Sowohl für traditionelle Rechtsanwender als auch für Legal Tech-Anwendungen geeignete Formate (CSV, PDF, EPUB, TXT und XML)</li> <li>Umfangreicher Compilation Report um den Erstellungs-Prozess zu erläutern</li> <li>Hochauflösende Diagramme und deskriptive Tabellen für alle Zwecke</li> <li>Diagramme in PDF (Druck) und PNG (Web) verfügbar, Tabellen als menschen- und maschinenlesbares CSV</li> <li>Vollständiges tabellarisches Verzeichnis aller Rechtsakte und der vom BMJV gebrauchten Abkürzungen</li> <li>Netzwerk-Strukturen für alle Rechtsakte und Visualisierungen für über 1000 Rechtsakte (experimentell)</li> <li><a href="../doi/10.5281/zenodo.4072934">Veröffentlichung des Source Codes</a></li> </ul> <p> </p> <p><strong>Source Code und Compilation Report</strong></p> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Robustness Checks auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enthält den Code für die vollständige Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Er ist zusammen mit dem Source Code hinterlegt. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der <em>vollständige Source Code</em> — sowohl für die Erstellung des Datensatzes, als auch für das Codebook — ist <em>öffentlich einsehbar </em>und<em> dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="https://doi.org/10.5281/zenodo.8402568">https://zenodo.org/doi/10.5281/zenodo.4072934</a></p> <p> </p> <p><strong>Kryptographische Signaturen</strong></p> <p>Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist.</p> <p>Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner <em>persönlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <p> </p> <p><strong>Kein Urheberrecht: Public Domain</strong></p> <p>An den Normtexten und Metadaten besteht gem. § 5 Abs. 1 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollständig urheberrechtsfrei.</p> <p> </p> <p><strong>Disclaimer</strong></p> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Behörden, Gerichten oder anderen öffentlichen Stellen der Bundesrepublik Deutschland.</p> <p> </p> <p><strong>Alternativen</strong></p> <p><em>[Ab 10.06.2019, nur XML]</em> Beckedorf, Janis/Coupette, Corinna/Hartung, Dirk. 2020. "gesetze-im-internet: A daily archive of https://www.gesetze-im-internet.de". GitHub. <a href="https://github.com/QuantLaw/gesetze-im-internet">https://github.com/QuantLaw/gesetze-im-internet</a></p> <p><em>[Änderungsgesetze]</em> Wehrmeyer, Stefan/Semsrott, Arne/Filter, Johannes. 2021. "OffeneGesetze.de ist eine zivilgesellschaftliche, ehrenamtliche Plattform für amtliche Gesetzesblätter". Open Knowledge Foundation. <a href="https://offenegesetze.de/">https://offenegesetze.de/</a></p> <p><em>[Alte Rechtsakte]</em> Open Knowledge Foundation. 2013. "Bundesgit". GitHub. <a href="https://github.com/bundestag/gesetze">https://github.com/bundestag/gesetze</a></p> <p> </p> <p><strong>Weitere Open Access Veröffentlichungen (Fobbe)</strong></p> <p>Website<em> </em>—<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data — <a href="../communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code — <a href="../communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regulärer Publikationen — <a href="../communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p> </p> <p><strong>Kontakt</strong></p> <p>Fehler gefunden? Anregungen? Kommentieren Sie gerne im <a href="https://codeberg.org/seanfobbe/c-dbr/issues">Issue Tracker</a> oder kontaktieren Sie mich über <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.