Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

135

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

135 results for “multilingual”

Learn how ShareScore rates datasets ↗
zenodo28/100

Figure 1a from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 1a Engagement metadata of the English Wikipedia article "Vaccine hesitancy". - Some experienced Wikipedia editors use account extensions to turn on the display of article metadata as seen here in the article header. The article is currently graded as "B" class on Wikipedia's quality scale. 1,031 Wikipedia editors have made 3,570 editorial revisions to the article since the article's creation on 15 April 2005. There are 428 registered Wikipedia editors who have put this article in their watchlist, which means that they have alerts about the article's development either whenever they request it or by some push notification. In the 30 days proceeding April 6, 2021 (when this screenshot was taken), this article has received 43,388 Wikipedia pageviews. The image is in the public domain and available via Wikimedia Commons.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 5a from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 5a Sample visualizations of data from within the Wikipedia ecosystem that relate to COVD-19 vaccine hesitancy. - Frequency of terms appearing in a Wikipedia-indexed set of papers on COVID-19 vaccine hesitancy. This screenshot is in the public domain and available via Wikimedia Commons.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 6 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 6 SELF magazine provided this photograph by Heather Hazzan with a free and open copyright license (CC BY 2.0) for use in Wikipedia or any other publication which required vaccine illustrations. It is available via Wikimedia Commons.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 5b from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 5b Sample visualizations of data from within the Wikipedia ecosystem that relate to COVD-19 vaccine hesitancy. - Tools in the Wikipedia ecosystem show the network relationship of terms in academic publications on COVID-19. This screenshot is in the public domain and available via Wikimedia Commons.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 4 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 4 The most accessed language versions of Wikipedia articles for "vaccine hesitancy" in 2020 were English, Italian, German, French, Russian, Spanish, Japanese, Chinese, Polish, Portuguese, and Arabic. This screenshot is available via Wikimedia Commons under the CC0 1.0 Universal Public Domain Dedication.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 3 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 3 Overview of pageviews for articles in the English Wikipedia's category for "vaccine hesitancy" over the course of 2020. This screenshot is available via Wikimedia Commons under the CC0 1.0 Universal Public Domain Dedication.

opencc-by-4.0Jul 2021View details →
zenodo28/100

Figure 2 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042

Figure 2 This is a Pageviews Analysis of multiple English Wikipedia articles related to vaccine hesitancy. Data visualizations such as this give insight into reader demand by reporting user engagement with sets of Wikipedia articles over time. Among other values, the table gives reports for "views", a measure of readership; "class", a grade of the content quality; and "editors", which is a count of the number of people who submitted editorial content. Wikipedia editors use these metrics to prioritize development of popular articles in need of more quality content. This screenshot is available via Wikimedia Commons under the terms of the Expat/ MIT license.

opencc-by-4.0Jul 2021View details →
zenodo28/100

FIGURE 183 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854

FIGURE 183. Morfologia externa básica dos Scarabaeinae.

opennotspecifiedApr 2011View details →
zenodo28/100

FIGURE 186. Diagramme d in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854

FIGURE 186. Diagramme d'un coléoptère coprophage, morphologie externe de base.

opennotspecifiedApr 2011View details →
zenodo28/100

FIGURE 182 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854

FIGURE 182. Basic external morphology of Scarabaeinae.

opennotspecifiedApr 2011View details →
zenodo28/100

FIGURE 185 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854

FIGURE 185. Diagram van een mestkever, fundamentele uitwendige morfologie.

opennotspecifiedApr 2011View details →
zenodo28/100

Amazon cell phone reviews (MT multilingual)

<p>English, Greek and Italian machine translated cell phone reviews&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo28/100

Does Native Multilingualism Lead to Enhanced Executive Functioning in Adulthood? - A Study Examining Inhibitory Control (Stroop Effect) in University Students

<p>The SPSS file lists all the subjects and information on these relevant to the Stroop Effect analysis, whose purpose was to find out whether native multilingualism has some effect on executive control in native multilinguals at their late age.</p>

opencc-by-4.0Sep 2021View details →
ClinicalTrials.gov28/100

A Mixed Method to Study Adherence to Oral Anticancer Medications in a Multilingual and Multicultural Setting

ClinicalTrials.gov study NCT04613765. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Pilot Trial of Multilingual Support Intervention

ClinicalTrials.gov study NCT04935047. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
zenodo24/100

A Comprehensive Dataset of Classified Citations with Identifiers from Multilingual Wikipedia (2024)

<p>This is a collection of translated citation datasets extracted from the Multilingual Wikipedia February 2024 dumps. The same extraction and template harmonization pipeline was used as for English Wikipedia&nbsp;<a title="English Wikipedia citations" href="../records/10782978">https://zenodo.org/records/10782978</a>.&nbsp;</p> <p><strong>Note:&nbsp;</strong> Versions 2 and 3 fix issues with large Italian, French and German datasets that were corrupted (failed to upload in full) in the initial version.</p> <p>In each language, Wikipedia authors can cite sources using language-specific or English templates. Our main effort in compiling these datasets was to assemble lists of citation templates for each language and convert relevant fields into a common English template. We started with known citation templates per each language (typically covering books, journals, web pages and news), and, in some cases, augmented these lists with additional frequently used templates (films, links, webarchives, etc.) which we were able to locate via the XML reference tags vs usage frequency dictionaries. For the list of accepted templates see our source code:&nbsp;<a href="https://github.com/albatros13/wikicite/tree/multilang">https://github.com/albatros13/wikicite/tree/multilang</a> (templates are listed in __init__.py files of the wikiciteparser library).</p> <p>A classification label is assigned to each citation (either 'news', 'book', 'journal' or 'other)' by the deterministic rule-based classifier that analyses available identifiers (see code documentation for details). Please note that these numbers do not represent the overall estimation of the book and journal citation numbers. We count only citations with DOI, PMID, PMC and ISBN identifiers assigned by authors (prior to the lookup process that augments citations with missing identifiers). The number of news citations is dependent on our list of recognised 22.646 <a href="https://github.com/albatros13/wikicite/blob/master/news/domains.txt">news agency domains</a>.&nbsp;&nbsp;</p> <table> <tbody> <tr> <td>Language</td> <td>Acronym</td> <td>Link</td> <td>Dump size</td> <td>Citations</td> <td>Books</td> <td>Journals</td> <td>News</td> </tr> <tr> <td>German&nbsp;</td> <td>de</td> <td><a href="https://dumps.wikimedia.org/dewiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dewiki/20240220/</a></td> <td>6.7GB</td> <td>4.854.945</td> <td>320.179</td> <td>105.542</td> <td>901.091</td> </tr> <tr> <td>French&nbsp;</td> <td>fr</td> <td><a href="https://dumps.wikimedia.org/frwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/frwiki/20240220/</a></td> <td>5.9GB</td> <td>9.552.768</td> <td>798.525</td> <td>264.560</td> <td>1.907.183</td> </tr> <tr> <td>Russian</td> <td>ru</td> <td><a href="https://dumps.wikimedia.org/ruwiki/20240220/">https://dumps.wikimedia.org/ruwiki/20240220/</a></td> <td>5.1GB</td> <td>7.437.100</td> <td>420.828</td> <td>130.470</td> <td>1.370.665</td> </tr> <tr> <td>Spanish</td> <td>es</td> <td><a href="https://dumps.wikimedia.org/eswiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/eswiki/20240220/</a></td> <td>4.2GB</td> <td>6.918.442</td> <td>522.910</td> <td>213.767</td> <td>1.699.396</td> </tr> <tr> <td>Italian</td> <td>it</td> <td><a href="https://dumps.wikimedia.org/itwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/itwiki/20240220/</a></td> <td>3.6GB</td> <td>5.545.082</td> <td>384.816</td> <td>128.366</td> <td>917.517</td> </tr> <tr> <td>Polish</td> <td>pl</td> <td><a href="https://dumps.wikimedia.org/plwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/plwiki/20240220/</a></td> <td>2.4GB</td> <td>4.744.158</td> <td>463.783&nbsp;</td> <td>95.988</td> <td>513.006</td> </tr> <tr> <td>Portuguese</td> <td>pt</td> <td><a href="https://dumps.wikimedia.org/ptwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/ptwiki/20240220/</a></td> <td>2.2GB</td> <td>4.775.025</td> <td>243.593</td> <td>142.216&nbsp;</td> <td>1.176.140</td> </tr> <tr> <td>Dutch</td> <td>nl</td> <td><a href="https://dumps.wikimedia.org/nlwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nlwiki/20240220/</a></td> <td>1.8GB</td> <td>566.549</td> <td>27.074&nbsp;</td> <td>12.706</td> <td>114.110</td> </tr> <tr> <td>Swedish</td> <td>sv</td> <td><a href="https://dumps.wikimedia.org/svwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/svwiki/20240220/</a></td> <td>1.5GB</td> <td>3.802.416</td> <td>112.748</td> <td>155.740&nbsp;</td> <td>869.662&nbsp;</td> </tr> <tr> <td>Catalan</td> <td>ca</td> <td><a href="https://dumps.wikimedia.org/cawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/cawiki/20240220/</a></td> <td>1.2GB</td> <td>2.239.714</td> <td>261.779</td> <td>105.125</td> <td>423.241</td> </tr> <tr> <td>Finnish</td> <td>fi</td> <td><a href="https://dumps.wikimedia.org/fiwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/fiwiki/20240220/</a></td> <td>900.9MB</td> <td>1.697.731</td> <td>209.556</td> <td>12.068</td> <td>286.420</td> </tr> <tr> <td>Turkish</td> <td>tr</td> <td><a href="https://dumps.wikimedia.org/trwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/trwiki/20240220</a></td> <td>883.9MB</td> <td>1.993.177</td> <td>85.079</td> <td>56.202</td> <td>339.122&nbsp;</td> </tr> <tr> <td>Norwegian</td> <td>no</td> <td><a href="https://dumps.wikimedia.org/nowiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nowiki/20240220</a></td> <td>763.7MB</td> <td>796.500</td> <td>43.314</td> <td>12.373</td> <td>151.780</td> </tr> <tr> <td>Danish</td> <td>da</td> <td><a href="https://dumps.wikimedia.org/dawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dawiki/20240220</a></td> <td>413.3MB</td> <td>437.239</td> <td>23.303</td> <td>7.522</td> <td>70.760&nbsp;</td> </tr> </tbody> </table> <p>This datasets can be equipped with identifiers located via the lookup process (no 'acquired_ID_list' field). If there is interest in augmented versions, see the source code for instructions or contact authors for assistance with this task.&nbsp; &nbsp;</p> <p>This research was supported in part by the <a href="https://dsc.uva.nl/">University of Amsterdam Data Science Centre</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo24/100

Multilingual test set for language identification and speech recognition from European Parliament recordings

<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings.&nbsp;</p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes.&nbsp;</p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>

opencc-zeroJul 2024View details →
zenodo24/100

FaceAttDB: A Multilingual Dataset for Facial Attribute Captioning

<p>The FaceCaption dataset is a curated collection specifically created for the purpose of research in the field of facial attribute captioning. It consists of 2,000 portrait images sourced from the CelebA dataset, showcasing a diverse range of facial characteristics such as age, gender, expression, and hair color. The dataset includes five captions per image, providing both English and Google-translated Bangla versions.</p> <p>The dataset is designed to facilitate the exploration of multilingual caption generation on portrait images. Each image in the dataset is accompanied by descriptive and informative captions that accurately describe the visual characteristics present in the image. The captions were generated based on the attribute annotations available in the CelebA dataset, ensuring a close alignment between the captions and the visual attributes.</p> <p>The images in the BanglaFaceCaption dataset are conveniently stored in a single folder, making them easily accessible for training and evaluation purposes. Additionally, an accompanying Excel sheet is provided, linking each image file with its corresponding English and Bangla captions.</p> <p>While the current version of the dataset comprises 2,000 images with five captions each, future work aims to expand the dataset size to enhance the diversity and robustness of models trained on it. The BanglaFaceCaption dataset serves as a valuable resource for researchers and practitioners interested in advancing the field of facial attribute captioning and exploring multilingual caption generation capabilities.</p>

opencc-by-4.0Jul 2023View details →
ClinicalTrials.gov24/100

Evaluating the Impact of CONNECT in a Multilingual Population

ClinicalTrials.gov study NCT07111936. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Clinical Intelligent Management System - Multilingual Exploration

ClinicalTrials.gov study NCT06923410. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record