Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
135
datasets available to search
ShareScore release 0.7.1
Dataset results
135 results for “multilingual”
Figure 1a from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 1a Engagement metadata of the English Wikipedia article "Vaccine hesitancy". - Some experienced Wikipedia editors use account extensions to turn on the display of article metadata as seen here in the article header. The article is currently graded as "B" class on Wikipedia's quality scale. 1,031 Wikipedia editors have made 3,570 editorial revisions to the article since the article's creation on 15 April 2005. There are 428 registered Wikipedia editors who have put this article in their watchlist, which means that they have alerts about the article's development either whenever they request it or by some push notification. In the 30 days proceeding April 6, 2021 (when this screenshot was taken), this article has received 43,388 Wikipedia pageviews. The image is in the public domain and available via Wikimedia Commons.
Figure 5a from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 5a Sample visualizations of data from within the Wikipedia ecosystem that relate to COVD-19 vaccine hesitancy. - Frequency of terms appearing in a Wikipedia-indexed set of papers on COVID-19 vaccine hesitancy. This screenshot is in the public domain and available via Wikimedia Commons.
Figure 6 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 6 SELF magazine provided this photograph by Heather Hazzan with a free and open copyright license (CC BY 2.0) for use in Wikipedia or any other publication which required vaccine illustrations. It is available via Wikimedia Commons.
Figure 5b from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 5b Sample visualizations of data from within the Wikipedia ecosystem that relate to COVD-19 vaccine hesitancy. - Tools in the Wikipedia ecosystem show the network relationship of terms in academic publications on COVID-19. This screenshot is in the public domain and available via Wikimedia Commons.
Figure 4 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 4 The most accessed language versions of Wikipedia articles for "vaccine hesitancy" in 2020 were English, Italian, German, French, Russian, Spanish, Japanese, Chinese, Polish, Portuguese, and Arabic. This screenshot is available via Wikimedia Commons under the CC0 1.0 Universal Public Domain Dedication.
Figure 3 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 3 Overview of pageviews for articles in the English Wikipedia's category for "vaccine hesitancy" over the course of 2020. This screenshot is available via Wikimedia Commons under the CC0 1.0 Universal Public Domain Dedication.
Figure 2 from: Rasberry L, Mietchen D (2021) Wikipedia for multilingual COVID-19 vaccine education at scale. Research Ideas and Outcomes 7: e70042. https://doi.org/10.3897/rio.7.e70042
Figure 2 This is a Pageviews Analysis of multiple English Wikipedia articles related to vaccine hesitancy. Data visualizations such as this give insight into reader demand by reporting user engagement with sets of Wikipedia articles over time. Among other values, the table gives reports for "views", a measure of readership; "class", a grade of the content quality; and "editors", which is a count of the number of people who submitted editorial content. Wikipedia editors use these metrics to prioritize development of popular articles in need of more quality content. This screenshot is available via Wikimedia Commons under the terms of the Expat/ MIT license.
FIGURE 183 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURE 183. Morfologia externa básica dos Scarabaeinae.
FIGURE 186. Diagramme d in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURE 186. Diagramme d'un coléoptère coprophage, morphologie externe de base.
FIGURE 182 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURE 182. Basic external morphology of Scarabaeinae.
FIGURE 185 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURE 185. Diagram van een mestkever, fundamentele uitwendige morfologie.
Amazon cell phone reviews (MT multilingual)
<p>English, Greek and Italian machine translated cell phone reviews </p>
Does Native Multilingualism Lead to Enhanced Executive Functioning in Adulthood? - A Study Examining Inhibitory Control (Stroop Effect) in University Students
<p>The SPSS file lists all the subjects and information on these relevant to the Stroop Effect analysis, whose purpose was to find out whether native multilingualism has some effect on executive control in native multilinguals at their late age.</p>
A Mixed Method to Study Adherence to Oral Anticancer Medications in a Multilingual and Multicultural Setting
ClinicalTrials.gov study NCT04613765. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Pilot Trial of Multilingual Support Intervention
ClinicalTrials.gov study NCT04935047. IPD Sharing: NO. Countries: 1. Publications: 0.
A Comprehensive Dataset of Classified Citations with Identifiers from Multilingual Wikipedia (2024)
<p>This is a collection of translated citation datasets extracted from the Multilingual Wikipedia February 2024 dumps. The same extraction and template harmonization pipeline was used as for English Wikipedia <a title="English Wikipedia citations" href="../records/10782978">https://zenodo.org/records/10782978</a>. </p> <p><strong>Note: </strong> Versions 2 and 3 fix issues with large Italian, French and German datasets that were corrupted (failed to upload in full) in the initial version.</p> <p>In each language, Wikipedia authors can cite sources using language-specific or English templates. Our main effort in compiling these datasets was to assemble lists of citation templates for each language and convert relevant fields into a common English template. We started with known citation templates per each language (typically covering books, journals, web pages and news), and, in some cases, augmented these lists with additional frequently used templates (films, links, webarchives, etc.) which we were able to locate via the XML reference tags vs usage frequency dictionaries. For the list of accepted templates see our source code: <a href="https://github.com/albatros13/wikicite/tree/multilang">https://github.com/albatros13/wikicite/tree/multilang</a> (templates are listed in __init__.py files of the wikiciteparser library).</p> <p>A classification label is assigned to each citation (either 'news', 'book', 'journal' or 'other)' by the deterministic rule-based classifier that analyses available identifiers (see code documentation for details). Please note that these numbers do not represent the overall estimation of the book and journal citation numbers. We count only citations with DOI, PMID, PMC and ISBN identifiers assigned by authors (prior to the lookup process that augments citations with missing identifiers). The number of news citations is dependent on our list of recognised 22.646 <a href="https://github.com/albatros13/wikicite/blob/master/news/domains.txt">news agency domains</a>. </p> <table> <tbody> <tr> <td>Language</td> <td>Acronym</td> <td>Link</td> <td>Dump size</td> <td>Citations</td> <td>Books</td> <td>Journals</td> <td>News</td> </tr> <tr> <td>German </td> <td>de</td> <td><a href="https://dumps.wikimedia.org/dewiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dewiki/20240220/</a></td> <td>6.7GB</td> <td>4.854.945</td> <td>320.179</td> <td>105.542</td> <td>901.091</td> </tr> <tr> <td>French </td> <td>fr</td> <td><a href="https://dumps.wikimedia.org/frwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/frwiki/20240220/</a></td> <td>5.9GB</td> <td>9.552.768</td> <td>798.525</td> <td>264.560</td> <td>1.907.183</td> </tr> <tr> <td>Russian</td> <td>ru</td> <td><a href="https://dumps.wikimedia.org/ruwiki/20240220/">https://dumps.wikimedia.org/ruwiki/20240220/</a></td> <td>5.1GB</td> <td>7.437.100</td> <td>420.828</td> <td>130.470</td> <td>1.370.665</td> </tr> <tr> <td>Spanish</td> <td>es</td> <td><a href="https://dumps.wikimedia.org/eswiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/eswiki/20240220/</a></td> <td>4.2GB</td> <td>6.918.442</td> <td>522.910</td> <td>213.767</td> <td>1.699.396</td> </tr> <tr> <td>Italian</td> <td>it</td> <td><a href="https://dumps.wikimedia.org/itwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/itwiki/20240220/</a></td> <td>3.6GB</td> <td>5.545.082</td> <td>384.816</td> <td>128.366</td> <td>917.517</td> </tr> <tr> <td>Polish</td> <td>pl</td> <td><a href="https://dumps.wikimedia.org/plwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/plwiki/20240220/</a></td> <td>2.4GB</td> <td>4.744.158</td> <td>463.783 </td> <td>95.988</td> <td>513.006</td> </tr> <tr> <td>Portuguese</td> <td>pt</td> <td><a href="https://dumps.wikimedia.org/ptwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/ptwiki/20240220/</a></td> <td>2.2GB</td> <td>4.775.025</td> <td>243.593</td> <td>142.216 </td> <td>1.176.140</td> </tr> <tr> <td>Dutch</td> <td>nl</td> <td><a href="https://dumps.wikimedia.org/nlwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nlwiki/20240220/</a></td> <td>1.8GB</td> <td>566.549</td> <td>27.074 </td> <td>12.706</td> <td>114.110</td> </tr> <tr> <td>Swedish</td> <td>sv</td> <td><a href="https://dumps.wikimedia.org/svwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/svwiki/20240220/</a></td> <td>1.5GB</td> <td>3.802.416</td> <td>112.748</td> <td>155.740 </td> <td>869.662 </td> </tr> <tr> <td>Catalan</td> <td>ca</td> <td><a href="https://dumps.wikimedia.org/cawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/cawiki/20240220/</a></td> <td>1.2GB</td> <td>2.239.714</td> <td>261.779</td> <td>105.125</td> <td>423.241</td> </tr> <tr> <td>Finnish</td> <td>fi</td> <td><a href="https://dumps.wikimedia.org/fiwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/fiwiki/20240220/</a></td> <td>900.9MB</td> <td>1.697.731</td> <td>209.556</td> <td>12.068</td> <td>286.420</td> </tr> <tr> <td>Turkish</td> <td>tr</td> <td><a href="https://dumps.wikimedia.org/trwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/trwiki/20240220</a></td> <td>883.9MB</td> <td>1.993.177</td> <td>85.079</td> <td>56.202</td> <td>339.122 </td> </tr> <tr> <td>Norwegian</td> <td>no</td> <td><a href="https://dumps.wikimedia.org/nowiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nowiki/20240220</a></td> <td>763.7MB</td> <td>796.500</td> <td>43.314</td> <td>12.373</td> <td>151.780</td> </tr> <tr> <td>Danish</td> <td>da</td> <td><a href="https://dumps.wikimedia.org/dawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dawiki/20240220</a></td> <td>413.3MB</td> <td>437.239</td> <td>23.303</td> <td>7.522</td> <td>70.760 </td> </tr> </tbody> </table> <p>This datasets can be equipped with identifiers located via the lookup process (no 'acquired_ID_list' field). If there is interest in augmented versions, see the source code for instructions or contact authors for assistance with this task. </p> <p>This research was supported in part by the <a href="https://dsc.uva.nl/">University of Amsterdam Data Science Centre</a>.</p>
Multilingual test set for language identification and speech recognition from European Parliament recordings
<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings. </p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes. </p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>
FaceAttDB: A Multilingual Dataset for Facial Attribute Captioning
<p>The FaceCaption dataset is a curated collection specifically created for the purpose of research in the field of facial attribute captioning. It consists of 2,000 portrait images sourced from the CelebA dataset, showcasing a diverse range of facial characteristics such as age, gender, expression, and hair color. The dataset includes five captions per image, providing both English and Google-translated Bangla versions.</p> <p>The dataset is designed to facilitate the exploration of multilingual caption generation on portrait images. Each image in the dataset is accompanied by descriptive and informative captions that accurately describe the visual characteristics present in the image. The captions were generated based on the attribute annotations available in the CelebA dataset, ensuring a close alignment between the captions and the visual attributes.</p> <p>The images in the BanglaFaceCaption dataset are conveniently stored in a single folder, making them easily accessible for training and evaluation purposes. Additionally, an accompanying Excel sheet is provided, linking each image file with its corresponding English and Bangla captions.</p> <p>While the current version of the dataset comprises 2,000 images with five captions each, future work aims to expand the dataset size to enhance the diversity and robustness of models trained on it. The BanglaFaceCaption dataset serves as a valuable resource for researchers and practitioners interested in advancing the field of facial attribute captioning and exploring multilingual caption generation capabilities.</p>
Evaluating the Impact of CONNECT in a Multilingual Population
ClinicalTrials.gov study NCT07111936. IPD Sharing: NO. Countries: 1. Publications: 0.
Clinical Intelligent Management System - Multilingual Exploration
ClinicalTrials.gov study NCT06923410. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.