Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “Cross language dataset”

Learn how ShareScore rates datasets ↗
zenodo48/100

CloneCorp: Cross-language clone detection dataset

<p>Data set of mobile apps and code examples to evaluate clone detection algorithms across languages (Kotlin, Swift, and Dart)</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages

<div> <div>SpokeN-100 is a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples.</div> </div>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Cross-Language Vulnerability Dataset with File Changes and Commit Messages

<p>Cross-Language Vulnerability Dataset with File Changes and Commit Messages</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Datasets for "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection"

<p>This repository contains the data and additional resources used for the paper:</p> <p>"Exploring the Potential of Neural Machine Translation for Cross-Language Clinical NLP Resource Generation through Annotation Projection. Rodriguez Miret et al. Information (2024)".</p> <p>There are four different datasets included, namely:</p> <ul> <li>The (1)&nbsp; <strong><a href="https://temu.bsc.es/distemist" target="_blank" rel="noopener">DisTEMIST</a>, </strong>(2)&nbsp;<strong> <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">DrugTEMIST</a> </strong>and<strong>&nbsp;</strong>(3)&nbsp;<strong><a href="https://temu.bsc.es/meddoprof" target="_blank" rel="noopener">MEDDOPROF</a> Spanish corpora and corresponding versions in 10 different languages</strong>, created through Machine Translation and annotation projection techniques. The Catalan annotations, used in the paper's experiments, were validated by bilingual expert annotators, who also provided alternative translations for the annotated terms in case they were wrongly translated. Thus, for Catalan we provide two different versions of the data: (i) the output of the annotation projection process as is, without any further validation, and (ii) the validated version of the data. For the rest of the languages (with the exception of the DrugTEMIST English and Italian data, used for the MultiCardioNER shared task), only an unvalidated version is provided.<strong><br></strong></li> <li>The (4) <strong>Catalan Clinical Case Corpus (CataCCC)</strong>, a collection of <em>200 clinical case reports in originally written in Catalan </em>covering a variety of clinical specialties. This corpus includes manually validated annotations for diseases, medications and professions created by the experts who annotated the corpora mentioned above, using the same guidelines and annottaion criteria. It can this be considered the first clinical Gold Standard corpus for diseases, medications and processions in Catalan.</li> </ul> <p>It is noteworthy that the MEDDOPROF-related data includes annotations for two labels, PROFESION and SITUACION_LABORAL, but only the former was used for training and evaluation in the paper.</p> <p>These are the <strong>10 languages</strong> included in the repository (along with their language codes):</p> <ul> <li><strong>Spanish</strong> (`es-gs`, with `gs` standing for Gold Standard)</li> <li><strong>Catalan</strong> (`cat`)</li> <li><strong>English</strong> (`en`)</li> <li><strong>French</strong> (`fr`)</li> <li><strong>Italian</strong> (`it`)</li> <li><strong>Dutch</strong> (`nl`)</li> <li><strong>Portuguese</strong> (`pt`)</li> <li><strong>Romanian</strong> (`ro`)</li> <li><strong>Swedish</strong> (`sv`)</li> <li><strong>Czech</strong> (`cz`)</li> </ul> <h2><strong>Related Links</strong></h2> <ul> <li><a href="../doi/10.5281/zenodo.13151039" target="_blank" rel="noopener">Validation and Correction Guidelines for the Multilingual Annotation Projection of Gold Standard Corpora</a></li> <li><a href="https://temu.bsc.es/multicardioner/" target="_blank" rel="noopener">MultiCardioNER Shared Task</a>&nbsp;</li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Dataset and Tool for the Paper: Cross-Language Dependencies: An Empirical Study of Kotlin-Java

<ul> <li>Folder <strong>tool-to-extract-kotlin-java-dependencies</strong> contains a runnable jar of our tool to extract Kotlin-Java dependencies, please follow README in this folder to execute this tool.</li> <li>Folder <strong>accuracy-verification&nbsp;</strong>contains the source code of the ground-truth project, along with manual dependency inspections from file to file and from line to line. The extracted dependencies result of this project from our tool is also included in this folder.</li> <li>Folder <strong>rq1-dependencies-graph-for-each-project&nbsp;</strong>contains all dependency graphs of our 23 subjects in Json format.</li> <li>Folder<strong> rq1-circular-layout-for-each-project</strong> contains the circular layout of dependencies among source files of our 23 subjects.</li> <li>Folder <strong>rq2-maintenance-cost-for-file-with-out-kotlin-java-interactions </strong>contains <ul> <li><strong>details</strong> lists the #commits and loc for each type of source files in details <ul> <li>_java_in_cross_language.csv represents Java source files participating in Kotlin-Java interactions</li> <li>_java_not_in_cross_language.csv represents Java source files not participating in Kotlin-Java interactions&nbsp;</li> <li>_kt_in_cross_language.csv represents Kotlin source files participating in Kotlin-Java interactions</li> <li>_kt_not_in_cross_language.csv represents Kotlin source files not participating in Kotlin-Java interactions&nbsp;</li> </ul> </li> </ul> </li> <li>Folder <strong>rq3-common-mistakes&nbsp;</strong>contains 103 cases in the manual inspections.</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo24/100

BuGL - A Cross-Language Dataset for Bug Localization

<p>BuGL - A Cross-Language Dataset for Bug Localization</p>

opencc-by-4.0Feb 2020View details →
zenodo8/100

One Model to Rule Them All: Cross-Language Detection of Malicious Packages for NPM and PyPI - Dataset, Models, Malicious Packages List

<p>This artifact complements the paper &quot;One Model to Rule Them All: Language-Independent Detection of Malicious Packages for PyPI and NPM&quot; submitted at ESEC/FSE 2023.</p> <p>In particular:</p> <ul> <li>&quot;Labelled_Dataset.csv &quot; contains the data used for the features evaluation, models evaluation and models training</li> <li>The .pkl files are the best-performing classifiers obtained after the models evaluation</li> <li>&quot;Malicious_Packages_Discovered.csv&quot; contain the list of malicious packages discovered during our real-world experiment</li> </ul>

restrictedJan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record