Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “machine translation”
Flexico: Sustainable Machine Translation via Self-Adaptation
<p>Data for reproducing the results of paper 'Flexico: Sustainable Machine Translation via Self-Adaptation'</p>
Lesan: Machine Translation for Low Resource Languages
<p>Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.</p>
Parallel text dataset for Neural Machine Translation (French -> Fongbe, French -> Ewe)
<p>This A Multilingual Dataset for Machine Translation from French to Ewe and fongbe, languages respectively from Togo and Benin</p> <p>This is a contribution with the AI4D initiative</p> <p>It contains 70k+ parallel sentences that have been annotated for Neural Machine Translation and sentence classification tasks.</p>
MENYO-20k: A Multi-domain English - Yorùbá Corpus for Machine Translation
<p>MENYO-20k is a multi-domain parallel dataset with texts obtained from news articles, ted talks, movie transcripts, radio transcripts, science and technology texts, and other short articles curated from the web and professional translators. The dataset has 20,100 parallel sentences split into 10,070 training sentences, 3,397 development sentences, and 6,633 test sentences (3,419 multi-domain, 1,714 news domain, and 1,500 ted talks speech transcript domain)</p> <p>The dataset is open but for non-commercial use because some of the data sources like <a href="https://www.ted.com/about/our-organization/our-policies-terms/ted-talks-usage-policy">Ted talks</a> and <a href="https://www.jw.org/en/terms-of-use/#link0">JW news</a> requires permission for commercial use.</p> <p><strong>Acknowledgement</strong>: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4All and Zindi Africa</p>
Datasets for "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection"
<p>This repository contains the data and additional resources used for the paper:</p> <p>"Exploring the Potential of Neural Machine Translation for Cross-Language Clinical NLP Resource Generation through Annotation Projection. Rodriguez Miret et al. Information (2024)".</p> <p>There are four different datasets included, namely:</p> <ul> <li>The (1) <strong><a href="https://temu.bsc.es/distemist" target="_blank" rel="noopener">DisTEMIST</a>, </strong>(2) <strong> <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">DrugTEMIST</a> </strong>and<strong> </strong>(3) <strong><a href="https://temu.bsc.es/meddoprof" target="_blank" rel="noopener">MEDDOPROF</a> Spanish corpora and corresponding versions in 10 different languages</strong>, created through Machine Translation and annotation projection techniques. The Catalan annotations, used in the paper's experiments, were validated by bilingual expert annotators, who also provided alternative translations for the annotated terms in case they were wrongly translated. Thus, for Catalan we provide two different versions of the data: (i) the output of the annotation projection process as is, without any further validation, and (ii) the validated version of the data. For the rest of the languages (with the exception of the DrugTEMIST English and Italian data, used for the MultiCardioNER shared task), only an unvalidated version is provided.<strong><br></strong></li> <li>The (4) <strong>Catalan Clinical Case Corpus (CataCCC)</strong>, a collection of <em>200 clinical case reports in originally written in Catalan </em>covering a variety of clinical specialties. This corpus includes manually validated annotations for diseases, medications and professions created by the experts who annotated the corpora mentioned above, using the same guidelines and annottaion criteria. It can this be considered the first clinical Gold Standard corpus for diseases, medications and processions in Catalan.</li> </ul> <p>It is noteworthy that the MEDDOPROF-related data includes annotations for two labels, PROFESION and SITUACION_LABORAL, but only the former was used for training and evaluation in the paper.</p> <p>These are the <strong>10 languages</strong> included in the repository (along with their language codes):</p> <ul> <li><strong>Spanish</strong> (`es-gs`, with `gs` standing for Gold Standard)</li> <li><strong>Catalan</strong> (`cat`)</li> <li><strong>English</strong> (`en`)</li> <li><strong>French</strong> (`fr`)</li> <li><strong>Italian</strong> (`it`)</li> <li><strong>Dutch</strong> (`nl`)</li> <li><strong>Portuguese</strong> (`pt`)</li> <li><strong>Romanian</strong> (`ro`)</li> <li><strong>Swedish</strong> (`sv`)</li> <li><strong>Czech</strong> (`cz`)</li> </ul> <h2><strong>Related Links</strong></h2> <ul> <li><a href="../doi/10.5281/zenodo.13151039" target="_blank" rel="noopener">Validation and Correction Guidelines for the Multilingual Annotation Projection of Gold Standard Corpora</a></li> <li><a href="https://temu.bsc.es/multicardioner/" target="_blank" rel="noopener">MultiCardioNER Shared Task</a> </li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
Datasets for Machine translations
Open the record for dataset details and reuse information.
An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation
<p>Paper: An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation</p> <p>Authors: Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk</p> <p>Journal: TOSEM 2019 - ACM Transactions on Software Engineering and Methodology </p>
An explorative Investigation into Neural Machine Translation: the Case of Low-Resource Language Pairs in Burkina Faso
<p>We explore the caveats of a promising research field, namely neural machine translation, in the context of preserving indigenous languages in Africa. We face the challenge of dealing with low-resource language pairs. Methodically, we employ some literature approaches and frameworks to learn translation models from Bible data in Moore (a major language in Burkina Faso) and French. Our experiments indeed confirm previous findings in the literature that vanilla neural machine translation models are ineffective for low resource language pairs. Surprisingly, however, we also found that we are not able to even remotely match performance recorded by the state of the art adapted methods for low-resource language pair in the literature. Nevertheless, although<br> word alignment, Byte Pair Encoding and Adam optimization did not successfully bring reasonable performance, we note that there are many remaining insightful approaches for low-resource language pairs.</p>
Error analysis of surname rendering in Finnish-to-English machine translation
<p>This surname list is a part of the dataset that I used in my master’s thesis “Error analysis of surname rendering in Finnish-to-English machine translation”. My master thesis is deposited at the Helsinki University Library.</p> <p>The surname list was extracted from the The Finnish News Agency Archive 2019–2021 (stt-fi-2019-2021-src). Permission to access the corpus can be obtained through the Language Bank of Finland. Corpus cannot be shared with third parties even if permission is granted.</p> <p>I cannot share the full dataset, as it contains sentences from the corpus.</p> <p>I am sharing only the list of surnames that I analyse in my master's thesis. The list contains 4,000 surnames extracted from the corpus. The list also contains the identification numbers of news articles from which the surnames were extracted and other metadata. Anyone with the permission to use the The Finnish News Agency Archive 2019–2021 corpus can use the id number to identify the news articles and recreate the dataset.</p> <p>Reference:</p> <p>STT. (2022). Finnish News Agency Archive 2019-2021, source [text corpus]. Kielipankki. Retrieved March 10, 2023, from http://urn.fi/urn:nbn:fi:lb-2022030202</p>
On Learning Meaningful Code Changes via Neural Machine Translation
<p>Paper: On Learning Meaningful Code Changes via Neural Machine Translation</p> <p>Authors: Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk</p> <p>ICSE 2019 - 41st ACM/IEEE International Conference on Software Engineering, May 25-31, 2019, Montréal, QC, Canada</p>
Machine learning analysis of the orbitofrontal cortex transcriptome of human opioid users identifies Shisa7 as a translational target relevant for heroin-seeking leveraging a male rat model
GEO Series GSE280431. Rattus norvegicus; Homo sapiens. 100 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.