Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “multilingual corpus”
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis
<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria—Hausa, Igbo, Nigerian-Pidgin, and Yorùbá—consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>
MaSS - Multilingual corpus of Sentence-aligned Spoken utterances
<p><strong>Abstract</strong></p> <p>The CMU Wilderness Multilingual Speech Dataset is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible), is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 para-lel spoken utterances across 8 languages (56 language pairs).We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for syntactically divergent language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs).</p> <p><a href="https://arxiv.org/pdf/1907.12895.pdf">Paper </a>| <a href="https://github.com/getalp/mass-dataset">GitHub Repository</a> containing the scripts needed to build the data set from scratch (if needed)</p> <p><strong>Project structure</strong></p> <p>This repository contains 8 Numpy files, one for each featured language, pickled with Python 3.6. Each line corresponds to the spectrogram of the file mentioned in the file <em>verses.csv</em>. There is a direct mapping between the ID of the verse and its index in the list (thus verse with ID 5634 is located at index 5634 in the Numpy file). Verses not available for a given language (as stated by the value "Not Available" in the CSV file) are represented by empty lists in the Numpy files, thus ensuring a perfect verse-to-verse alignement between each file.</p> <p>Spectrogram were extracted using Librosa with the following parameters:</p> <pre><code>Pre-emphasis = 0.97 Sample rate = 16000 Window size = 0.025 Window stride = 0.01 Window type = 'hamming' Mel coefficients = 40 Min frequency = 20</code></pre> <p> </p>
Subset of 'MLSUM: The Multilingual Summarization Corpus' for constraints annotation experiment
<p><strong>[EN] Subset of 'MLSUM: The Multilingual Summarization Corpus' for constraints annotation experiment.</strong></p> <ul> <li><strong>Description</strong>: MLSUM is a dataset of newspappers articles aimed at training summaring model. We use it for a constraints annotation experiment on newspapper titles according to their topic classification.</li> <li><strong>Content</strong>: For constraints annotation experiment based on data similarity, this dataset have been subsetted (randomly pick 75 articles in the following 14 most used topics: 'economie', 'politique', 'sport', 'planete' (renamed in 'ecologie'), 'sciences', 'police-justice', 'disparitions', 'emploi', 'sante', 'musiques', 'arts', 'educations', 'climat' (renamed in 'meteo'), 'immobilier') and filtered (keep articles that have an obvious topics regarding their titles, without their bodies). Two reviewers have working on this task in order to limit the subjectivity of the filtering. This subsetted dataset is used (1) to estimate needed time to annotate titles similarity with constraints (MUST-LINK, CANNOT-LINK) and (2) to test interactive clustering methodology (constraints annotation and constrained clustering).</li> <li><strong>Origin</strong>: The dataset is bassed on the original 'MLSUM: The Multilingual Summarization Corpus' dataset (https://doi.org/10.48550/arXiv.2004.14900).</li> </ul> <p><br> <strong>[FR] Echantillon de 'MLSUM: The Multilingual Summarization Corpus' pour une expérience d'annotation de contraintes.</strong></p> <ul> <li><strong>Description </strong>: MLSUM est un ensemble de données d'articles de journaux destinés à l'entraînement d'un modèle de résumé automatique. Nous l'utilisons pour une expérience d'annotation de contraintes sur des titres de journaux en fonction de leur classification thématique.</li> <li><strong>Contenu </strong>: Pour une expérience d'annotation de contraintes basée sur la similarité des données, cet ensemble de données a été échantillonné (sélectionner au hasard de 75 articles dans les 14 sujets les plus utilisés : 'économie', 'politique', 'sport', 'planète' (renommé en « écologie »). ), 'sciences', 'police-justice', 'disparitions', 'emploi', 'sante', 'musiques', 'arts', 'éducations', 'climat' (renommé en 'meteo'), 'immobilier' ) et filtré (conserver les articles qui ont un sujet évident par rapport à leur titre, sans leur corps). Deux relecteurs ont travaillé sur cette tâche afin de limiter la subjectivité du filtrage. Ce sous-ensemble de données est utilisé (1) pour estimer le temps nécessaire pour annoter la similarité des titres avec des contraintes (MUST-LINK, CANNOT-LINK) et (2) pour tester la méthodologie de clustering interactif (annotation de contraintes et clustering contraint).</li> <li><strong>Origine </strong>: L'ensemble de données est basé sur l'ensemble de données original 'MLSUM : The Multilingual Summarization Corpus' (https://doi.org/10.48550/arXiv.2004.1490).</li> </ul>
VivesDebate: A New Annotated Multilingual Corpus of Argumentation in a Debate Tournament
<p>The application of the latest Natural Language Processing breakthroughs in computational argumentation has shown promising results which have raised the interest in this area of research. However, the available corpora with argumentative annotations are often limited to a very specific purpose or are not of adequate size to take advantage of state-of-the-art deep learning techniques (e.g., deep neural networks). In this paper, we present VivesDebate, a large, richly annotated, and versatile professional debate corpus for computational argumentation research. The corpus has been created from 29 transcripts of a debate tournament in Catalan and has been machine-translated into Spanish and English. The annotation contains argumentative propositions, argumentative relations, debate interactions, and professional evaluations of the arguments and argumentation. The presented corpus can be useful for research on a heterogeneous set of computational argumentation underlying tasks such as argument mining, argument analysis, argument evaluation, or argument generation among others. All this makes VivesDebate a valuable resource for computational argumentation research within the context of massive corpora aimed at Natural Language Processing tasks.</p>
BVS Corpus: A Multilingual Parallel Corpus and Translation Experiments of Biomedical Scientific Texts
<p>The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME in agreement with the Pan American Health Organization (OPAS). Abstracts are available in English, Spanish, and Portuguese, with a subset in more than one language, thus being a possible source of parallel corpora. In this article, we present the development of parallel corpora from BVS in three languages: English, Portuguese, and Spanish. Sentences were automatically aligned using the Hunalign algorithm for EN/ES and EN/PT language pairs, and for a subset of trilingual articles also. We demonstrate the capabilities of our corpus by training a Neural Machine Translation (OpenNMT) system for each language pair, which outperformed related works on scientific biomedical articles. Sentence alignment was also manually evaluated, presenting an average 96\% of correctly aligned sentences across all languages. Our parallel corpus is freely available, with complementary information regarding article metadata.</p> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
FONA corpus: Food & Nutrition Abstracts Multilingual corpus
<p>The FONA corpus is a collection of case reports specifically selected to foster the development of Language Technologies, Text Mining and NLP for applications in the domain of food & nutrition.</p> <p> </p> <p>It contains a large collection of documents (titles and abstracts) with metadata information on their MeSH terms. In addition, a subset of the collection contains automatically recognized entities of the following categories:</p> <ul> <li>medical procedures</li> <li>symptoms</li> <li>diseases</li> <li>medications</li> <li>occupational and demographic information</li> <li>species (pathogens)</li> <li>cancer morphology</li> </ul>
The Vuk'uzenzele South African Multilingual Corpus
<pre># The Vuk'uzenzele South African Multilingual Corpus [](https://doi.org/10.5281/zenodo.7598539) Github: <a href="https://github.com/dsfsi/vukuzenzele-nlp">https://github.com/dsfsi/vukuzenzele-nlp</a> ## About dataset The dataset contains editions from the South African government magazine Vuk'uzenzele. Data was scraped from PDFs that have been placed in the [data/raw](data/raw/) folder. The PDFS were obtained from the [Vuk'uzenzele website](https://www.vukuzenzele.gov.za/). The datasets contain government magazine editions in 11 languages, namely: | Language | Code | Language | Code | |------------|-------|------------|-------| | English | (eng) | Sepedi | (sep) | | Afrikaans | (afr) | Setswana | (tsn) | | isiNdebele | (nbl) | Siswati | (ssw) | | isiXhosa | (xho) | Tshivenda | (ven) | | isiZulu | (zul) | Xitstonga | (tso) | | Sesotho | (nso) | ### Number of Aligned Pairs with Cosine Similarity Score >= 0.65 | src_lang | trg_lang | num_aligned_pairs | |----------|----------|-------------------| | ven | zul | 186 | | ssw | xho | 1965 | | sep | xho | 279 | | nbl | zul | 227 | | nso | tsn | 1279 | | nso | tso | 1491 | | tsn | zul | 1346 | | afr | eng | 1369 | | eng | ssw | 1601 | | afr | ssw | 1496 | | nbl | ssw | 264 | | tso | zul | 1758 | | afr | zul | 1384 | | eng | zul | 1888 | | ssw | tsn | 1263 | | sep | tsn | 302 | | nso | xho | 1248 | | sep | tso | 324 | | ssw | tso | 1657 | | tsn | ven | 235 | | eng | nbl | 153 | | nso | sep | 349 | | afr | nbl | 359 | | nbl | ven | 657 | | eng | ven | 243 | | afr | ven | 281 | | tso | ven | 256 | | ven | xho | 215 | | eng | tsn | 1380 | | afr | tsn | 1076 | | nso | ssw | 1132 | | eng | tso | 2016 | | afr | tso | 1139 | | xho | zul | 1895 | | tsn | xho | 1209 | | sep | zul | 223 | | nbl | xho | 204 | | ssw | zul | 2161 | | afr | xho | 1363 | | eng | xho | 1354 | | tso | xho | 1485 | | sep | ssw | 219 | | nbl | tso | 215 | | tsn | tso | 1570 | | nso | zul | 1247 | | nbl | tsn | 140 | | eng | sep | 276 | | afr | sep | 394 | | ssw | ven | 217 | | sep | ven | 1140 | | afr | nso | 962 | | eng | nso | 1721 | | nbl | nso | 151 | | nbl | sep | 843 | | nso | ven | 262 | The dataset is present in several forms on the repo. Generally the dataset is split by edition, eg. `2020-01-ed1` The data directory is broken down as follows ``` ./data ├── external # Data external to this repo ├── interim # I am not really sure - looks like interim in regards to processed. ├── processed # The data from scraping the raw pdfs ├── raw # The raw pdfs of the Vuk'uzenzele magazine ├── sentence_align_output # The output (csv) of the sentence alignment with LASER language encoders └── simple_align_output # The output (csv) of a simple one to one sentence alignment ``` The dataset is split by edition in the [data/processed](data/processed/) folder. Authors ------- - Vukosi Marivate - [@vukosi](https://twitter.com/vukosi) - Andani Madodonga - Daniel Njini - Richard Lastrucci Citation -------- Vukosi Marivate, Andani Madodonga, Daniel Njini, Richard Lastrucci, Isheanesu Dzingirai . **The Vuk'uzenzele South African Multilingual Corpus**, 2023 > @inproceedings{lastrucci-etal-2023-preparing, title = "Preparing the Vuk{'}uzenzele and {ZA}-gov-multilingual {S}outh {A}frican multilingual corpora", author = "Richard Lastrucci and Isheanesu Dzingirai and Jenalea Rajab and Andani Madodonga and Matimba Shingange and Daniel Njini and Vukosi Marivate", booktitle = "Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)", month = may, year = "2023", address = "Dubrovnik, Croatia", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.rail-1.3", pages = "18--25" } > @dataset{marivate_vukosi_2023_7598540, author = {Marivate, Vukosi and Njini, Daniel and Madodonga, Andani and Lastrucci, Richard and Dzingirai, Isheanesu}, title = {The Vuk'uzenzele South African Multilingual Corpus}, month = feb, year = 2023, publisher = {Zenodo}, doi = {10.5281/zenodo.7598539}, url = {https://doi.org/10.5281/zenodo.7598539} } Licences ------- * License for Data - [CC 4.0 BY SA](LICENSE.data.md) * Licence for Code - [MIT License](LICENSE.md)</pre>
NLAS-multi: A Multilingual Corpus of Automatically Generated Natural Language Argumentation Schemes
<p>The multilingual corpus of natural language argumentation schemes (NLAS-multi) consists of 3,810 natural language argumentation schemes of which 1,893 are in English and 1,917 in Spanish. It has a total of 253,516 words distributed in 118,493 words in English and 135,023 words in Spanish. In terms of inferences, our corpus has a total of 7,964 (3,949 in English and 4,015 in Spanish). Furthermore, the NLAS-multi corpus contains a total of 23,781 conflict relations between arguments in the same topic.</p>
MultiCardioNER Corpus: Multilingual Adaptation of Clinical NER Systems to the Cardiology Domain
<h1><strong>MultiCardioNER</strong></h1> <p><strong>MultiCardioNER</strong> is a shared task about the adaptation of clinical NER systems to the cardiology domain. It uses a combination of two existing datasets (DisTEMIST for diseases and the newly-released DrugTEMIST for medications), as well as a new, smaller dataset of cardiology clinical cases annotated using the same guidelines.</p> <p>Participants are provided DisTEMIST and DrugTEMIST as training data to use as they see fit (1,000 documents, with the original partitions splitting them into 750 for training and 250 for testing). The cardiology clinical cases (cardioccc) are meant to be used as a development or validation set (258 documents), although participants are encourage to experiment with the documents and annotations as they see fit. The evaluation is done using a different collection of cardiology clinical cases (250).</p> <p>MultiCardioNER proposes two tracks:</p> <p>- Track 1: Spanish adaptation of disease recognition systems to the cardiology domain.<br>- Track 2: Multilingual (Spanish, English and Italian) adaptation of medication recognition systems to the cardiology domain.</p> <p>Please read the README file attached for more information on folder structure and file format.</p> <p><strong>MultiCardioNER</strong> was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioASQ 2024. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">https://temu.bsc.es/multicardioner</a>. This task is promoted by Spanish and European projects such as DataTools4Heart, AI4HF, BARITONE and AI4ProfHealth.</p> <p><strong>UPDATE MAY 28th 2024: </strong>The test set annotations are now out! We've also included the original background set files, as well as a file with the mappings from the masked filenames used during the evaluation phase to the original filenames. Please check the README for more information.</p> <h2><strong>Resources</strong></h2> <ul> <li><a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">MultiCardioNER website</a></li> <li><a href="http://bioasq.org/" target="_blank" rel="noopener">BioASQ website</a></li> <li><a href="../doi/10.5281/zenodo.6458078" target="_blank" rel="noopener">DisTEMIST Guidelines</a></li> <li><a href="../doi/10.5281/zenodo.11065432" target="_blank" rel="noopener">DrugTEMIST Guidelines</a></li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p> <h2><strong>Additional resources and corpora</strong></h2> <p>If you are interested in MultiCardioNER, you might want to check out these corpora and resources:</p> <ul> <li><a href="../records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT)</li> <li><a href="../records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT)</li> <li><a href="../records/10635215">SympTEMIST</a> (Corpus of clinical findings and normalization to SNOMED CT)</li> <li><a href="../records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization)</li> <li><a href="../records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization)</li> <li><a href="../records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization)</li> <li><a href="../records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI))</li> <li><a href="../records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization)</li> <li><a href="../records/3837305">CodiESP</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version)</li> <li><a href="../records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy)</li> <li><a href="../records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags)</li> <li><a href="../records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries))</li> <li><a href="../records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags)</li> <li><a href="../records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts)</li> </ul>
The South African Gov-ZA multilingual corpus
<pre>The South African Gov-ZA multilingual corpus ============================== Github: https://github.com/dsfsi/gov-za-multilingual Zenodo: About Dataset --------------------- The data set contains cabinet statements from the South African government. Data was scraped from the governments website: https://www.gov.za/cabinet-statements The datasets contain government cabinet statements in 11 languages, namely: | Language | Code | Language | Code | |------------|------|------------|------| | English | (eng) | Sepedi | (nso) | | Afrikaans | (afr) | Setswana | (tsn) | | isiNdebele | (nbl) | Siswati | (ssw) | | isiXhosa | (xho) | Tshivenda | (ven) | | isiZulu | (zul) | Xitstonga | (tso) | | Sesotho | (sot) | The dataset contains the full data in a JSON file (/data/govza-cabinet-statements.json), as well as CSV’s split by each language, eg: “govza-cabinet-statements-en.csv” for english. The dataset does not contain special characters like unicode or ascii. Please see the [data-statement.md](/data_statement.md) for full dataset information. *(TODO)* Number of Aligned Pairs with Cosine Similarity Score >= 0.65 ------------------------------------------------------------ | src_lang | trg_lang | num_aligned_pairs | |----------|----------|-------------------| | afr | eng | 14549 | | afr | nbl | 6621 | | afr | nso | 15388 | | afr | sot | 8834 | | afr | ssw | 15610 | | afr | tsn | 12605 | | afr | tso | 14936 | | afr | ven | 5776 | | afr | xho | 16065 | | afr | zul | 14998 | | nbl | eng | 3616 | | nbl | nso | 6342 | | nbl | sot | 16163 | | nbl | ssw | 4655 | | nbl | tsn | 3369 | | nbl | tso | 4465 | | nbl | ven | 18984 | | nbl | xho | 5213 | | nbl | zul | 3868 | | nso | eng | 15257 | | nso | ssw | 18697 | | nso | tsn | 16179 | | nso | tso | 17617 | | nso | ven | 6367 | | sot | eng | 5212 | | sot | nso | 8077 | | sot | ssw | 5811 | | sot | tsn | 5450 | | sot | tso | 6586 | | sot | ven | 14098 | | ssw | eng | 15721 | | ssw | tso | 17880 | | ssw | ven | 4588 | | tsn | eng | 14544 | | tsn | ssw | 16386 | | tsn | tso | 16681 | | tsn | ven | 3267 | | tso | eng | 16068 | | ven | eng | 3670 | | ven | tso | 4578 | | xho | eng | 16537 | | xho | nso | 18110 | | xho | sot | 7489 | | xho | ssw | 18387 | | xho | tsn | 16571 | | xho | tso | 17954 | | xho | ven | 4559 | | xho | zul | 18145 | | zul | eng | 16149 | | zul | nso | 17630 | | zul | sot | 5975 | | zul | ssw | 18563 | | zul | tsn | 16482 | | zul | tso | 17789 | | zul | ven | 3606 | Authors ------- - Vukosi Marivate - [@vukosi](https://twitter.com/vukosi) - Matimba Shingange - Richard Lastrucci - Isheanesu Joseph Dzingirai - Jenalea Rajab Publications -------</pre> <p>> @inproceedings{lastrucci-etal-2023-preparing,<br> title = "Preparing the Vuk{'}uzenzele and {ZA}-gov-multilingual {S}outh {A}frican multilingual corpora",<br> author = "Richard Lastrucci and Isheanesu Dzingirai and Jenalea Rajab and Andani Madodonga and Matimba Shingange and Daniel Njini and Vukosi Marivate",<br> booktitle = "Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)",<br> month = may,<br> year = "2023",<br> address = "Dubrovnik, Croatia",<br> publisher = "Association for Computational Linguistics",<br> url = "https://aclanthology.org/2023.rail-1.3",<br> pages = "18--25"<br> }</p>
TUNDRA - A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision,
<p>The corpus is described in:</p> <p><strong>A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King,</strong> <em>TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision</em>, In Proc. Interspeech, Lyon, France, August 2013</p> <p> </p> <p> </p> <pre>############################################################### ## ## ## THE SIMPLE4ALL TUNDRA CORPUS ## ## version 1.0 ## ## ## ############################################################### Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech research with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Most audiobooks are from the public domain and allow redistribution. However, some have restricted use, and in those cases the segmented and aligned data cannot be downloaded from our website. --------------------------------------------------------------- LICENCE --------------------------------------------------------------- This work is licensed under a Creative Commons Attribution 3.0 Unported License http://creativecommons.org/licenses/by/3.0/ This licence applies to the selection, segmentation and alignment of the speech and text data. The underlying audio and text are licensed under their specific datasource terms. Please refer to the links below for a full description of them. If you use any part of the corpus in your work, please cite the following paper: A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 --------------------------------------------------------------- SPEECH AND TEXT SOURCES --------------------------------------------------------------- 1) Bulgarian - "Zhetvariat" by Yordan Yovkov audio: http://librivox.org/zhetvariat-by-yordan-yovkov text: http://slovo.bg/showwork.php3?AuID=95&WorkID=9610&Level=1 2)Danish - "Grimms eventyr I udvalg" by Grimm Brothers audio: http://librivox.org/grimms-eventyr-i-udvalg-by-br%C3%B8drene-grimm text: http://www.estrup.org/cms/?mod=text&id=392 3)Dutch- "Anna Karenina" by Leo Tolstoy audio: http://librivox.org/anna-karenina-by-leo-tolstoy text: http://www.gutenberg.org/ebooks/13214 4) English - "Living Alone" by Stella Benson audio: http://librivox.org/living-alone-by-stella-benson text: http://www.gutenberg.org/ebooks/14907 5) Finnish- "Rautatie" by Juhani Aho audio: http://librivox.org/rautatie-by-juhani-aho text: http://www.gutenberg.org/ebooks/10481 6) French - "Candide" by Voltaire audio: http://librivox.org/candide-by-voltaire text: http://www.gutenberg.org/cache/epub/4650/pg4650.txt 7) German - "Das Bildnis des Dorian Gray" by Oscar Wilde audio: http://librivox.org/das-bildnis-des-dorian-gray-by-oscar-wilde text: http://gutenberg.spiegel.de/buch/1836/1 8) Hungarian - "Egri csillagok" by Geza Gardonyi audio: http://gutenberg.spiegel.de/buch/1836/1 text: http://mek.oszk.hu/00600/00656/index.phtml 9) Italian - "Galatea" by Anton Giulio Barrili audio: http://librivox.org/galatea-by-anton-giulio-barrili/ text: http://www.gutenberg.org/ebooks/19427 10) Polish - "Siedem wybranyc opowiadan" by Wladyslaw Orkan audio: http://librivox.org/siedem-wybranych-opowiadan-by-wladyslaw-orkan/ text: http://pl.wikisource.org/wiki/Autor:W%C5%82adys%C5%82aw_Orkan 11) Portuguese - "Senhora" by Jose de Alencar audio: http://librivox.org/senhora-by-jose-de-alencar/ text: http://stat.correioweb.com.br/arquivos/educacao/arquivos/JosdeAlencar-Senhora0.pdf 12) Romanian - "Mara" by Ioan Slavici audio: http://speech.utcluj.ro/corpora/mara.html text: http://ro.wikisource.org/wiki/Mara 13) Russian - "Ucheniye Khrista" by Leo Tolstoy audio: http://librivox.org/teachings-of-christ-rus-by-leo-tolstoy/ text: http://az.lib.ru/t/tolstoj_lew_nikolaewich/text_0520.shtml 14) Spanish - "Don Quijote de la Mancha" by Miguel de Cervantes audio: http://www.quijote.es/IVCentenario_AudioLibro.php text: http://www.gutenberg.org/ebooks/5921 --------------------------------------------------------------- CONTENTS --------------------------------------------------------------- For each audiobook you can download the following information, with the exception of the Spanish audiobook which has a restricted use and the speech data cannot be downloaded from out website: 1) Segmented and aligned data -- http://tundra.simple4all.org/download.html -- an archive containing the results of the lightly supervised segmentation and alignment algorithm; -- folders and files (the following are the same for both training and test data sets): -- wav/ - speech data maintaining the original chapter names, but with additional indexes resulted from the sentence-level segmentation; -- txt/ - raw text files corresponding to each speech file from the wav/ folder; -- txtWithPunctuation/ - text files for each speech segment with punctuation restored from the original book text; -- speech_transcript.txt - a single file for all orthographic transcripts; -- a separate handmade test set data in the handmadeTest/ folder (see below for its description); 2) 1 hour subset of selected data -- http://tundra.simple4all.org/ssw8data.html -- an archive containing approximately 1 hour of selected audio used to train the voices from the Demo section; 3) Synthetic samples -- http://tundra.simple4all.org/ssw8data.html -- an archive containing the synthetic samples obtained with our lightly supervised TTS system for the handmade test set; 4) Chapter-level annotation -- http://tundra.simple4all.org/download.html -- a file with the chapter-level time alignment within the original data and the corresponding text for the confident data. --------------------------------------------------------------- SEGMENTATION AND ALIGNMENT --------------------------------------------------------------- Descriptions of the lightly supervised segmentation and alignment methods can be found in the following papers: 1) A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 2) Adriana STAN, Peter BELL, Simon KING A grapheme-based method for automatic alignment of speech and text data, In Proc. IEEE Workshop on Spoken Language Technology, Miami, Florida, USA, December 2012 3) Yoshitaka Mamiya, Junichi Yamagishi, Oliver Watts, Robert A.J. Clark, Simon King and Adriana STAN Lightly Supervised GMM VAD to use Audiobook for Speech Synthesiser, In Proc. ICASSP, May 2013 The synthetic voice building algorithm is described in detail here: 4) O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from ‘found’ data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 With a similar approach being presented in: 5) O. Watts, A. Stan, A. Suni, M. Burgos, J.M. Montero, The Simple4All entry to the Blizzard Challenge 2013, Blizzard Challenge 2013 --------------------------------------------------------------- TRAIN/TEST DIVISION OF DATA --------------------------------------------------------------- Test material is taken from the ends of books, from enough whole chapters or stories to make up at least 10 min of audio of aligned data (NB more can be harvested from these chapters from the unaligned utterances). The exceptions are the Hungarian and Portuguese audiobooks in which the following chapters have variable recording conditions and are not considered suitable for comparisons: Hungarian: egricsillagok_[19-49] Portuguese: senhora_[14-20] and senhora_[23-41] The following chapters are reserved for testing: Bulgarian: zhetvariat_2{3,4,5}* Danish: eventyr_{08,09,10,11,12}* Dutch: annakarenina_021* German: doriangray_17* English: livingalone_{09,10}* Finnish: rautatie_{7,8}* French: candide_{29,30}* Hungarian: egricsillagok_{17,18}* Italian: galatea_{19,20}* Polish: siedemwybranchopowiadan_7* Portuguese: senhora_{12,13}* Romanian: mara_7{1,2}* Russian: teachingsofchrist_9* Spanish: Parte1_35* For the evaluations published in the following paper: O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from 'found' data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 a hand-segmented test set of about 40 utterances in all languages was prepared from the test chapters, so that various problems with the automatically aligned test utterances (sentence fragments, non-matching transcripts etc.) would not confuse the evaluation results. These hand segmented and aligned utterances are contained in the ./handmadeTest folder. Synthesised samples of this handmade test set are also available for download. --------------------------------------------------------------- CONTRIBUTORS --------------------------------------------------------------- Adriana Stan (Communications Department, Technical University of Cluj-Napoca) Oliver Watts (Centre for Speech Technology Research, University of Edinburgh) Yoshitaka Mamiya (Centre for Speech Technology Research, University of Edinburgh) Junichi Yamagishi (National Institute of Informatics, Tokyo) Mircea Giurgiu (Communications Department, Technical University of Cluj-Napoca) Rob Clark (Centre for Speech Technology Research, University of Edinburgh) Simon King (Centre for Speech Technology Research, University of Edinburgh) --------------------------------------------------------------- CONTACT --------------------------------------------------------------- Please send all you enquires regarding the Tundra Copus to one of the following e-mail addresses: adriana.stan@com.utcluj.ro owatts@inf.ed.ac.uk --------------------------------------------------------------- ACKNOWLEDGEMNTS --------------------------------------------------------------- The research leading to these results has received funding from the European Community's Seventh Framework Programme (FP7/2007-2013) under grant agreement No 287678 (the Simple4All project - http://www.simple4all.org) The research presented here has made use of the resources provided by the Edinburgh Compute and Data Facility (ECDF: http://www.ecdf.ed.ac.uk). The ECDF is partially supported by the eDIKT initiative (http://www.edikt.org.uk). We would like to thank Mihai Nae from Cartea Sonora for releasing the Romanian data, as well as to all the volunteers at Librivox and Gutenberg for dedicating their time to distribute this wide variety of data. </pre>
CT-FAN-22 corpus: A Multilingual dataset for Fake News Detection
<p><strong>Data Access: </strong>The data in the research collection provided may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use it only for research purposes. Due to these restrictions, the collection is not open data. Please download the Agreement at <a href="https://drive.google.com/file/d/1QU-rw4D26r3F04FB63hTvToxOvDKaKdv/view?usp=sharing">Data Sharing Agreement</a> and send the signed form to <a href="mailto:fakenewstask@gmail.com">fakenewstask@gmail.com</a> .</p> <p><strong>Citation</strong></p> <p>Please cite our work as</p> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Subtask 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong> Sub-task A would detect fake news designed as a four-class classification problem. The training data will be released in batches and roughly about 900 articles with the respective label. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <p><strong>Task 3</strong></p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <p><strong>Task 3</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p>Sample file</p> <pre><code>public_id, predicted_domain 1, health 2, crime</code></pre> <p><strong>Additional data for Training</strong></p> <p>To train your model, the participant can use additional data with a similar format; some datasets are available over the web. We don't provide the background truth for those datasets. For testing, we will not use any articles from other datasets. Some of the possible sources:</p> <ul> <li><a href="https://www.kaggle.com/liberoliber/onion-notonion-datasets">Fakenews Classification Datasets</a></li> <li><a href="https://www.kaggle.com/c/fakenewskdd2020/overview">Fake News Detection Challenge KDD 2020</a></li> <li><a href="https://www.kaggle.com/mdepak/fakenewsnet?select=PolitiFact_real_news_content.csv">FakeNewsNet</a></li> </ul> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2021, and the content of fake news is mixed up with several topics like election, COVID-19 etc.</li> </ol> <p><strong>Evaluation Metrics</strong></p> <p>This task is evaluated as a classification task. We will use the F1-macro measure for the ranking of teams. There is a limit of 5 runs (total and not per day), and only one person from a team is allowed to submit runs.</p> <p><strong>Submission Link: </strong><a href="https://codalab.org/">Coming soon</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi, G. K., Struß, J. M., & Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection. <em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeño, A., Míguez, R., Shaar, S., ... & Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In <em>European Conference on Information Retrieval</em> (pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeño, A., Míguez, R., Shaar, S., ... & Kartal, Y. S. (2021, September). Overview of the CLEF–2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In <em>International Conference of the Cross-Language Evaluation Forum for European Languages</em> (pp. 264-291). Springer, Cham.</li> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, “FakeCovid – a multilingualcross-domain fact check news dataset for covid-19,” inWorkshop Proceedings of the 14th International AAAIConference on Web and Social Media, 2020. <a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., & Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter. <em>Online Social Networks and Media</em>, <em>22</em>, 100104. doi: <a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> </ul>
EmoFilm - A multilingual emotional speech corpus
<p><strong>EmoFilm</strong> is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages. The audio clips (with a mean length of 3.5 sec. and std 1.2 sec.) were extracted in wave format (uncompressed, mono, 48 kHz sample rate and 16-bit) from 43 films (original in English and their over-dubbed Italian and Spanish versions). Genres including comedy, drama, horror, and thriller were considered; anger, contempt, happiness, fear, and sadness emotional states were taken into account. EmoFilm has been presented at Interspeech 2018:</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, and Björn Schuller (2018), <em>Categorical vs Dimensional Perception of Italian Emotional Speech</em>, in Proc. of Interspeech, Hyderabad, India, pp. 3638-3642 .</p> <p>We would like to thank Linda Ratz for her contribution in the generation of the transcriptions.</p> <p> </p> <p><strong>How to access EmoFilm</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1pFHfsqk7snF_EVqq0WAC0Dz8FcTD3s9_/view?usp=share_link</p>
Multilingual fine-grained sentiment analysis corpus
<p>A sentiment annotated corpus based on Fallout New Vegas. The corpus has the following sentiments: <em>neutral, anger, disgust, fear, happy, pained, sad, surprised</em> in the following languages: <em>English, German, Italian, Spanish and French</em>.</p> <p>Please cite the following paper: Mika Hämäläinen, Khalid Alnajjar, and Thierry Poibeau. 2022. Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. In <em>FDG’22: Proceedings of the 17th International Conference on the Foundations of Digital Games (FDG ’22)</em></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.