Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “Multilingual data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data for Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake

<p>Data to go with our submission to Climatic Change titled &quot;Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake&quot;.</p> <p>Dataset contains the embeddings (.zip with pickles) as well as the associated document items (idem), the most-closely associated keywords and paragraphs per topic in the final model (.xlsx), the reduced 2d embeddings with all selected paragraphs (.csv utf-8 encoded), as well as an overview with the meta-data per source (.csv utf-8 encoded).</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Multilingual Bottle-Neck Feature Learning from Untranscribed data for track 1 in zerospeech2017 (system 1 -- without VTLN)

<p>We investigate the extraction of bottle-neck features (BNFs) for multiple languages without access to manual transcription. Multilingual BNFs are derived from a multi-task learning deep neural network which is trained with unsupervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models separately trained on untranscribed speech of multiple languages.</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Glosario: A multilingual glossary for computing and data science terms.

<p><code>glosario</code> is an open-source glossary of terms used in data science that is available online and also as a library in both&nbsp;<a href="https://github.com/carpentries/glosario-r/">R</a>&nbsp;and&nbsp;<a href="https://github.com/carpentries/glosario-py/">Python</a>. By adding glossary keys to a lesson&rsquo;s metadata, authors can indicate what the lesson teaches, what learners ought to know before they start, and where they can go to find that knowledge. Authors can also use the library&rsquo;s functions to insert consistent hyperlinks for terms and definitions in their lessons in any of several languages. The master copy of the glossary lives in the <code>glossary.yml</code> file.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data set for "Token-Level Multilingual Epidemic Dataset for Event Extraction"

<p>This is the data for the TPDL 2021 paper &quot;<a href="https://zenodo.org/record/5780020">Token-Level Multilingual Epidemic Dataset for Event Extraction</a>&quot;. If you use this resource, please cite the paper:</p> <pre><code>@inproceedings{mutuvi2021dataset,     title = "Token-level Multilingual Epidemic Dataset for Event Extraction",     author = {Mutuvi, Stephen and Boros, Emanuela and Doucet, Antoine, and Lejeune, Gaël and Jatowt, Adam and Odeo, Moses},     booktitle = "Proceedings of the 25th International Conference on Theory and Practice of Digital Libraries, September 13–17, 2021, TPDL 2021",     year = "2021",     location = "Online" }</code></pre> <p>&nbsp;</p> <p>This work has been supported by the European Union Horizon 2020 research and innovation programme under grants 825153 (Embeddia) and 770299 (NewsEye).</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction

<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models.&nbsp;</p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p>&nbsp; &nbsp; English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br>&nbsp; &nbsp; German: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br>&nbsp; &nbsp; Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br>&nbsp; &nbsp; Greek: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br>&nbsp; &nbsp; Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br>&nbsp; &nbsp; Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p>&nbsp;</p> <p><strong>Dataset Building&nbsp;</strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model&nbsp;</li> <li>Translation into the targeted language</li> <li>Word Alignment&nbsp;</li> <li>Data Augmentation&nbsp;</li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website:&nbsp;</span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Data of Paper "Turning a Multilingual Historical Archive into an Information System through Post-OCR Correction and Content-Based Indexation"

<p>We evaluated our approach on a collection of 946 historical documents belonging to the Biblioteca Nacional de Catalunya (BNC), spanning from 1914 to 1951. Each document is the issue of a magazine, comprising different articles by different authors. This implies that, despite the thematic nature of magazines and specific issues, there is a certain degree of heterogeneity in each document. Magazines were selected based on their relevance w.r.t. art in general and, more specifically, early 20th century avant-garde movements (e.g., Dadaism, Cubism, etc.). For each document, we have the scanning of the original artifact and the plain raw text extracted through ABBYY FineReader OCR tool. To the best of our knowledge, this is the first Catalan-dominated OCR corpus ever released.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Multilingual CoNaLa Datset, train data

<p>Training datasets used in the Multilingual CoNaLa benchmark experimentation.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Multilingual Transformations Data

<p>Data for &quot;Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models&quot; (Findings of ACL 2022).</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

TUNDRA - A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision,

<p>The corpus is described in:</p> <p><strong>A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King,</strong> <em>TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision</em>, In Proc. Interspeech, Lyon, France, August 2013</p> <p>&nbsp;</p> <p>&nbsp;</p> <pre>############################################################### ## ## ## THE SIMPLE4ALL TUNDRA CORPUS ## ## version 1.0 ## ## ## ############################################################### Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech research with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Most audiobooks are from the public domain and allow redistribution. However, some have restricted use, and in those cases the segmented and aligned data cannot be downloaded from our website. --------------------------------------------------------------- LICENCE --------------------------------------------------------------- This work is licensed under a Creative Commons Attribution 3.0 Unported License http://creativecommons.org/licenses/by/3.0/ This licence applies to the selection, segmentation and alignment of the speech and text data. The underlying audio and text are licensed under their specific datasource terms. Please refer to the links below for a full description of them. If you use any part of the corpus in your work, please cite the following paper: A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 --------------------------------------------------------------- SPEECH AND TEXT SOURCES --------------------------------------------------------------- 1) Bulgarian - "Zhetvariat" by Yordan Yovkov audio: http://librivox.org/zhetvariat-by-yordan-yovkov text: http://slovo.bg/showwork.php3?AuID=95&amp;WorkID=9610&amp;Level=1 2)Danish - "Grimms eventyr I udvalg" by Grimm Brothers audio: http://librivox.org/grimms-eventyr-i-udvalg-by-br%C3%B8drene-grimm text: http://www.estrup.org/cms/?mod=text&amp;id=392 3)Dutch- "Anna Karenina" by Leo Tolstoy audio: http://librivox.org/anna-karenina-by-leo-tolstoy text: http://www.gutenberg.org/ebooks/13214 4) English - "Living Alone" by Stella Benson audio: http://librivox.org/living-alone-by-stella-benson text: http://www.gutenberg.org/ebooks/14907 5) Finnish- "Rautatie" by Juhani Aho audio: http://librivox.org/rautatie-by-juhani-aho text: http://www.gutenberg.org/ebooks/10481 6) French - "Candide" by Voltaire audio: http://librivox.org/candide-by-voltaire text: http://www.gutenberg.org/cache/epub/4650/pg4650.txt 7) German - "Das Bildnis des Dorian Gray" by Oscar Wilde audio: http://librivox.org/das-bildnis-des-dorian-gray-by-oscar-wilde text: http://gutenberg.spiegel.de/buch/1836/1 8) Hungarian - "Egri csillagok" by Geza Gardonyi audio: http://gutenberg.spiegel.de/buch/1836/1 text: http://mek.oszk.hu/00600/00656/index.phtml 9) Italian - "Galatea" by Anton Giulio Barrili audio: http://librivox.org/galatea-by-anton-giulio-barrili/ text: http://www.gutenberg.org/ebooks/19427 10) Polish - "Siedem wybranyc opowiadan" by Wladyslaw Orkan audio: http://librivox.org/siedem-wybranych-opowiadan-by-wladyslaw-orkan/ text: http://pl.wikisource.org/wiki/Autor:W%C5%82adys%C5%82aw_Orkan 11) Portuguese - "Senhora" by Jose de Alencar audio: http://librivox.org/senhora-by-jose-de-alencar/ text: http://stat.correioweb.com.br/arquivos/educacao/arquivos/JosdeAlencar-Senhora0.pdf 12) Romanian - "Mara" by Ioan Slavici audio: http://speech.utcluj.ro/corpora/mara.html text: http://ro.wikisource.org/wiki/Mara 13) Russian - "Ucheniye Khrista" by Leo Tolstoy audio: http://librivox.org/teachings-of-christ-rus-by-leo-tolstoy/ text: http://az.lib.ru/t/tolstoj_lew_nikolaewich/text_0520.shtml 14) Spanish - "Don Quijote de la Mancha" by Miguel de Cervantes audio: http://www.quijote.es/IVCentenario_AudioLibro.php text: http://www.gutenberg.org/ebooks/5921 --------------------------------------------------------------- CONTENTS --------------------------------------------------------------- For each audiobook you can download the following information, with the exception of the Spanish audiobook which has a restricted use and the speech data cannot be downloaded from out website: 1) Segmented and aligned data -- http://tundra.simple4all.org/download.html -- an archive containing the results of the lightly supervised segmentation and alignment algorithm; -- folders and files (the following are the same for both training and test data sets): -- wav/ - speech data maintaining the original chapter names, but with additional indexes resulted from the sentence-level segmentation; -- txt/ - raw text files corresponding to each speech file from the wav/ folder; -- txtWithPunctuation/ - text files for each speech segment with punctuation restored from the original book text; -- speech_transcript.txt - a single file for all orthographic transcripts; -- a separate handmade test set data in the handmadeTest/ folder (see below for its description); 2) 1 hour subset of selected data -- http://tundra.simple4all.org/ssw8data.html -- an archive containing approximately 1 hour of selected audio used to train the voices from the Demo section; 3) Synthetic samples -- http://tundra.simple4all.org/ssw8data.html -- an archive containing the synthetic samples obtained with our lightly supervised TTS system for the handmade test set; 4) Chapter-level annotation -- http://tundra.simple4all.org/download.html -- a file with the chapter-level time alignment within the original data and the corresponding text for the confident data. --------------------------------------------------------------- SEGMENTATION AND ALIGNMENT --------------------------------------------------------------- Descriptions of the lightly supervised segmentation and alignment methods can be found in the following papers: 1) A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 2) Adriana STAN, Peter BELL, Simon KING A grapheme-based method for automatic alignment of speech and text data, In Proc. IEEE Workshop on Spoken Language Technology, Miami, Florida, USA, December 2012 3) Yoshitaka Mamiya, Junichi Yamagishi, Oliver Watts, Robert A.J. Clark, Simon King and Adriana STAN Lightly Supervised GMM VAD to use Audiobook for Speech Synthesiser, In Proc. ICASSP, May 2013 The synthetic voice building algorithm is described in detail here: 4) O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from &lsquo;found&rsquo; data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 With a similar approach being presented in: 5) O. Watts, A. Stan, A. Suni, M. Burgos, J.M. Montero, The Simple4All entry to the Blizzard Challenge 2013, Blizzard Challenge 2013 --------------------------------------------------------------- TRAIN/TEST DIVISION OF DATA --------------------------------------------------------------- Test material is taken from the ends of books, from enough whole chapters or stories to make up at least 10 min of audio of aligned data (NB more can be harvested from these chapters from the unaligned utterances). The exceptions are the Hungarian and Portuguese audiobooks in which the following chapters have variable recording conditions and are not considered suitable for comparisons: Hungarian: egricsillagok_[19-49] Portuguese: senhora_[14-20] and senhora_[23-41] The following chapters are reserved for testing: Bulgarian: zhetvariat_2{3,4,5}* Danish: eventyr_{08,09,10,11,12}* Dutch: annakarenina_021* German: doriangray_17* English: livingalone_{09,10}* Finnish: rautatie_{7,8}* French: candide_{29,30}* Hungarian: egricsillagok_{17,18}* Italian: galatea_{19,20}* Polish: siedemwybranchopowiadan_7* Portuguese: senhora_{12,13}* Romanian: mara_7{1,2}* Russian: teachingsofchrist_9* Spanish: Parte1_35* For the evaluations published in the following paper: O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from 'found' data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 a hand-segmented test set of about 40 utterances in all languages was prepared from the test chapters, so that various problems with the automatically aligned test utterances (sentence fragments, non-matching transcripts etc.) would not confuse the evaluation results. These hand segmented and aligned utterances are contained in the ./handmadeTest folder. Synthesised samples of this handmade test set are also available for download. --------------------------------------------------------------- CONTRIBUTORS --------------------------------------------------------------- Adriana Stan (Communications Department, Technical University of Cluj-Napoca) Oliver Watts (Centre for Speech Technology Research, University of Edinburgh) Yoshitaka Mamiya (Centre for Speech Technology Research, University of Edinburgh) Junichi Yamagishi (National Institute of Informatics, Tokyo) Mircea Giurgiu (Communications Department, Technical University of Cluj-Napoca) Rob Clark (Centre for Speech Technology Research, University of Edinburgh) Simon King (Centre for Speech Technology Research, University of Edinburgh) --------------------------------------------------------------- CONTACT --------------------------------------------------------------- Please send all you enquires regarding the Tundra Copus to one of the following e-mail addresses: adriana.stan@com.utcluj.ro owatts@inf.ed.ac.uk --------------------------------------------------------------- ACKNOWLEDGEMNTS --------------------------------------------------------------- The research leading to these results has received funding from the European Community's Seventh Framework Programme (FP7/2007-2013) under grant agreement No 287678 (the Simple4All project - http://www.simple4all.org) The research presented here has made use of the resources provided by the Edinburgh Compute and Data Facility (ECDF: http://www.ecdf.ed.ac.uk). The ECDF is partially supported by the eDIKT initiative (http://www.edikt.org.uk). We would like to thank Mihai Nae from Cartea Sonora for releasing the Romanian data, as well as to all the volunteers at Librivox and Gutenberg for dedicating their time to distribute this wide variety of data. </pre>

opencc-by-4.0Jun 2013View details →
zenodo20/100

Multilingual Stylometry: Data and Code

<p>Data and code accompanying the Multilingual Stylometry Showcase as well as the research paper describing the showcase.&nbsp;</p>

restrictedcc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record