Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Cross language dataset”
CloneCorp: Cross-language clone detection dataset
<p>Data set of mobile apps and code examples to evaluate clone detection algorithms across languages (Kotlin, Swift, and Dart)</p>
SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages
<div> <div>SpokeN-100 is a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples.</div> </div>
Cross-Language Vulnerability Dataset with File Changes and Commit Messages
<p>Cross-Language Vulnerability Dataset with File Changes and Commit Messages</p>
Datasets for "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection"
<p>This repository contains the data and additional resources used for the paper:</p> <p>"Exploring the Potential of Neural Machine Translation for Cross-Language Clinical NLP Resource Generation through Annotation Projection. Rodriguez Miret et al. Information (2024)".</p> <p>There are four different datasets included, namely:</p> <ul> <li>The (1) <strong><a href="https://temu.bsc.es/distemist" target="_blank" rel="noopener">DisTEMIST</a>, </strong>(2) <strong> <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">DrugTEMIST</a> </strong>and<strong> </strong>(3) <strong><a href="https://temu.bsc.es/meddoprof" target="_blank" rel="noopener">MEDDOPROF</a> Spanish corpora and corresponding versions in 10 different languages</strong>, created through Machine Translation and annotation projection techniques. The Catalan annotations, used in the paper's experiments, were validated by bilingual expert annotators, who also provided alternative translations for the annotated terms in case they were wrongly translated. Thus, for Catalan we provide two different versions of the data: (i) the output of the annotation projection process as is, without any further validation, and (ii) the validated version of the data. For the rest of the languages (with the exception of the DrugTEMIST English and Italian data, used for the MultiCardioNER shared task), only an unvalidated version is provided.<strong><br></strong></li> <li>The (4) <strong>Catalan Clinical Case Corpus (CataCCC)</strong>, a collection of <em>200 clinical case reports in originally written in Catalan </em>covering a variety of clinical specialties. This corpus includes manually validated annotations for diseases, medications and professions created by the experts who annotated the corpora mentioned above, using the same guidelines and annottaion criteria. It can this be considered the first clinical Gold Standard corpus for diseases, medications and processions in Catalan.</li> </ul> <p>It is noteworthy that the MEDDOPROF-related data includes annotations for two labels, PROFESION and SITUACION_LABORAL, but only the former was used for training and evaluation in the paper.</p> <p>These are the <strong>10 languages</strong> included in the repository (along with their language codes):</p> <ul> <li><strong>Spanish</strong> (`es-gs`, with `gs` standing for Gold Standard)</li> <li><strong>Catalan</strong> (`cat`)</li> <li><strong>English</strong> (`en`)</li> <li><strong>French</strong> (`fr`)</li> <li><strong>Italian</strong> (`it`)</li> <li><strong>Dutch</strong> (`nl`)</li> <li><strong>Portuguese</strong> (`pt`)</li> <li><strong>Romanian</strong> (`ro`)</li> <li><strong>Swedish</strong> (`sv`)</li> <li><strong>Czech</strong> (`cz`)</li> </ul> <h2><strong>Related Links</strong></h2> <ul> <li><a href="../doi/10.5281/zenodo.13151039" target="_blank" rel="noopener">Validation and Correction Guidelines for the Multilingual Annotation Projection of Gold Standard Corpora</a></li> <li><a href="https://temu.bsc.es/multicardioner/" target="_blank" rel="noopener">MultiCardioNER Shared Task</a> </li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
Dataset and Tool for the Paper: Cross-Language Dependencies: An Empirical Study of Kotlin-Java
<ul> <li>Folder <strong>tool-to-extract-kotlin-java-dependencies</strong> contains a runnable jar of our tool to extract Kotlin-Java dependencies, please follow README in this folder to execute this tool.</li> <li>Folder <strong>accuracy-verification </strong>contains the source code of the ground-truth project, along with manual dependency inspections from file to file and from line to line. The extracted dependencies result of this project from our tool is also included in this folder.</li> <li>Folder <strong>rq1-dependencies-graph-for-each-project </strong>contains all dependency graphs of our 23 subjects in Json format.</li> <li>Folder<strong> rq1-circular-layout-for-each-project</strong> contains the circular layout of dependencies among source files of our 23 subjects.</li> <li>Folder <strong>rq2-maintenance-cost-for-file-with-out-kotlin-java-interactions </strong>contains <ul> <li><strong>details</strong> lists the #commits and loc for each type of source files in details <ul> <li>_java_in_cross_language.csv represents Java source files participating in Kotlin-Java interactions</li> <li>_java_not_in_cross_language.csv represents Java source files not participating in Kotlin-Java interactions </li> <li>_kt_in_cross_language.csv represents Kotlin source files participating in Kotlin-Java interactions</li> <li>_kt_not_in_cross_language.csv represents Kotlin source files not participating in Kotlin-Java interactions </li> </ul> </li> </ul> </li> <li>Folder <strong>rq3-common-mistakes </strong>contains 103 cases in the manual inspections.</li> </ul>
BuGL - A Cross-Language Dataset for Bug Localization
<p>BuGL - A Cross-Language Dataset for Bug Localization</p>
One Model to Rule Them All: Cross-Language Detection of Malicious Packages for NPM and PyPI - Dataset, Models, Malicious Packages List
<p>This artifact complements the paper "One Model to Rule Them All: Language-Independent Detection of Malicious Packages for PyPI and NPM" submitted at ESEC/FSE 2023.</p> <p>In particular:</p> <ul> <li>"Labelled_Dataset.csv " contains the data used for the features evaluation, models evaluation and models training</li> <li>The .pkl files are the best-performing classifiers obtained after the models evaluation</li> <li>"Malicious_Packages_Discovered.csv" contain the list of malicious packages discovered during our real-world experiment</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.