Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

12

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

12 results for “Parallel Corpus”

Learn how ShareScore rates datasets ↗
zenodo44/100

ENGLISH-AKUAPEM TWI PARALLEL CORPUS

<p>This dataset&nbsp;<em><strong>(verified_data.csv)</strong></em>&nbsp;is bilingual machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs.&nbsp;<br> A transformer-based machine translator was used to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native speakers.&nbsp;<br> The main idea of a typical use case for the dataset is for further training of machine translation models in Akuapem Twi.<br> The data can also be used for other downstream NLP tasks such as Named Entity Recognition and POS tagging, with appropriate additional annotations.&nbsp;<br> Another potential application is training unsupervised embeddings for the Akuapem Twi language.<br> In addition a higher quality 697 crowdsourced sentences <em><strong>(crowdsourced_data.csv)&nbsp;</strong></em>are provided for use as an evaluation set for the tasks highlighted above. It is recommended as a testing dataset for machine translation English to Twi and Twi to English models.</p> <p><strong>Acknowledgement</strong>: This project was supported by the&nbsp;<a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a>&nbsp;through K4all and&nbsp;Zindi Africa</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

The Makerere Gendered Corpus: A Gendered English to Luganda Parallel Corpus

<p>This&nbsp;English-Luganda parallel sentence corpus consists of&nbsp;gendered examples created by a team of researchers from Makerere AI&nbsp; Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under CC-0 and translating these sentences using a crowdsourcing, iterative and opensource approach was done using Pontoon an opensource Translation Management System built by Mozilla. This is a corpus of&nbsp;1,000 parallel sentences.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

The Makerere MT Corpus: English to Luganda parallel corpus

<p>This English-Luganda parallel sentence corpus&nbsp;was created by a team of researchers from AI &amp; Data science research Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under CC-0 and translating these sentences using a crowdsourcing, iterative and opensource approach was done using Pontoon an opensource Translation Management System built by Mozilla.</p> <p>Acknowledgment: This project was supported by the&nbsp;<a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a>&nbsp;through K4All and&nbsp;<a href="https://zindi.africa/">Zindi Africa</a>.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Vernacular Parallel Glosses in the Gloss-ViBe Corpus

<p>For this dataset I have collected all those instances in which there is an Old Irish gloss in the Vienna Bede that has a parallel gloss in at least one different language &ndash; Latin or Old Breton/Welsh &ndash; in one of the other three manuscripts recorded in the Gloss-ViBe corpus (https://gams.uni-graz.at/context:glossvibe). This dataset is used in <a href="https://doi.org/10.12688/openreseurope.16006.1">https://doi.org/10.12688/openreseurope.16006.1</a>.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Nupe-English Parallel Corpus

<p>This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP Tasks such as machine translation.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

BVS Corpus: A Multilingual Parallel Corpus and Translation Experiments of Biomedical Scientific Texts

<p>The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME in&nbsp;agreement with the Pan American Health Organization (OPAS). Abstracts are available in English, Spanish, and Portuguese, with a subset in more than one language, thus being a possible source of parallel corpora. In this article, we present the development of parallel corpora from BVS in three languages: English, Portuguese, and Spanish. Sentences were automatically aligned using the Hunalign algorithm for EN/ES and EN/PT language pairs, and for a subset of trilingual articles also. We demonstrate the capabilities of our corpus by training a Neural Machine Translation (OpenNMT) system for each language pair, which outperformed related works on scientific biomedical articles. Sentence alignment was also manually evaluated, presenting an average 96\% of correctly aligned sentences across all languages. Our parallel corpus is freely available, with complementary information regarding article metadata.</p> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Umsuka English - isiZulu Parallel Corpus

<p>We&rsquo;ve developed an open-source, high quality isiZulu parallel corpus that comes from a<br> mixture of domains, taking into account both Southern African context and international<br> English context, by using professional translators. We sourced 5000 English sentences,<br> sampled from News Crawl datasets that were translated into isiZulu. Additionally, we<br> translated 5000 isiZulu sentences, sampled from both the NCHLT monolingual corpus and<br> the open-source documents of the UKZN isiZulu National monolingual corpus, into English.<br> From each set, we separated out 1000 patterns as the evaluation dataset. Since isiZulu is<br> highly morphologically complex, we believe that the English-to-isiZulu evaluation set should<br> be translated at least twice, by different translators which will allow us to calculate<br> human-level BLEU score for the dataset.</p> <p>More details in the provided Data Statement</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Supplementary material for "Using a parallel corpus to study patterns of word order variation: Determiners and quantifiers within the noun phrase in European languages"

<p>- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files.</p> <p>Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

BSL-Hansard: A parallel, multimodal corpus of English and interpreted British Sign Language data from parliamentary proceedings

<p>BSL-Hansard is a novel open source and multimodal resource composed by combining Sign Language video data in BSL and English text from the official transcription of British parliamentary sessions. This paper describes the method followed to compile BSL-Hansard including time alignment of text using the MAUS (Schiel, 2015) segmentation system, gives some statistics about this dataset, and suggests experiments. These primarily include end-to-end Sign Language-to-text translation, but is also relevant for broader machine translation, and speech and language processing tasks.</p> <p>This dataset will be useful for translation between BSL and English, or for studies in BSL or English down to the phonetic level.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

EuroparlExtract - Directional Parallel Corpora Extracted from the European Parliament Proceedings Parallel Corpus

<p>This dataset contains directional parallel corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. EuroparlExtract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>

opencc-by-4.0Nov 2017View details →
zenodo36/100

EuroparlExtract - Comparable Corpora Extracted from the European Parliament Proceedings Parallel Corpus

<p>This dataset contains comparable translational corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. Europarl Extract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>

opencc-by-4.0Nov 2017View details →
zenodo36/100

ClinSpEn Corpus: Parallel English-Spanish COVID-19 Clinical Cases, Terminology and Ontology Concepts

<p><strong>ClinSpEn Parallel Corpus Collection</strong></p> <p>This repository contains the <strong>complete</strong> <strong>ClinSpEn corpus collection</strong>, which was used for the <strong><a href="https://temu.bsc.es/clinspen">ClinSpEn shared task</a></strong> at Biomedical WMT 2022.</p> <p>ClinSpEn is a collection of <strong>Gold Standard EN-ES parallel corpora of different types of clinical data</strong>: case reports, medical controlled vocabularies/ontologies, and clinical terms and entities extracted from medical content. It includes development and test data translated by professional medical translators that can be used <strong>to train and benchmark clinical EN-ES machine translation systems</strong>. Additionally, monolingual background data is provided so that the systems&#39; performance can be analyzed in unseen data.</p> <p>If you use this dataset, please cite:</p> <blockquote> <pre><code>inproceedings{biowmt22, title={Findings of the WMT 2022 Biomedical Translation Shared Task: Monolingual Clinical Case Reports}, author={Neves, Mariana and Yepes, Antonio Jimeno and Siu, Amy and Roller, Roland and Thomas, Philippe and Navarro, Maika Vicente and Yeganova, Lana and Wiemann, Dina and Di Nunzio, Giorgio Maria and Vezzani, Federica and others}, booktitle={WMT22-Seventh Conference on Machine Translation}, pages={694--723}, year={2022} } </code></pre> </blockquote> <p><strong>Data Description</strong></p> <p>ClinSpEn proposes three different sub-tracks, each based on a different type of clinical data:</p> <p><em>1. Clinical Cases:</em></p> <p>Parallel EN-ES COVID-19 clinical case reports. The direction of this sub-track is EN&gt;ES.</p> <p>The dataset&rsquo;s case reports were carefully selected to cover a wide range of aspects related to the disease: different types of patients (children, adults, elderly and pregnant people, babies), different comorbidities (cancer, mental health issues, immunosuppressed patients) and symptomatology (mild and severe presentations, dermatologic, immunologic and psychiatric manifestations, thrombosis, ...). The reports were translated from English to Spanish by a professional medical translator on a first step and revised by a clinical expert on a second step.</p> <p>The sample (dev) set and test set are made up of parallel txt files (50 and 152 documents each, respectively), with the Spanish version having a &ldquo;.es&rdquo; extension and the English files having a &ldquo;.en&rdquo; extension. Each report has been parallelized so that every sentence&rsquo;s line number corresponds to the same sentence&rsquo;s line number in both languages.</p> <p>The background data (9,804 files) is made up of a TSV file with four columns: filename, document number, line number and English line. The clinical cases themselves include COVID-19 case reports as well as diverse content extracted from PubMed.</p> <p>If you need to map the entries in the join test + background document provided in earlier versions, you may use the &quot;clinspen_clinicalcases_test-set_filename_mapping.tsv&quot; file.</p> <p><em>2. Clinical Terminology:</em></p> <p>Parallel EN-ES clinical terms extracted from medical literature and clinical records, with particular focus on diseases, symptoms, findings, procedures and professions and translated and revised by professional medical translators. The direction of this sub-track is ES&gt;EN.</p> <p>The sample (dev) set contains 7,000 terms as a tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms.</p> <p>The test data (12,128 terms) is made up of a TSV file with three columns: term number, English term and Spanish term.</p> <p>The background data (201,890 terms) is made up of a TSV file with two columns: term number and Spanish term.</p> <p>The term number columns can be used to map the entries in the join test + background document provided in earlier versions.</p> <p><em>3. Ontology Concepts:</em></p> <p>Parallel EN-ES concepts extracted from various open biomedical ontologies and taxonomies and then manually translated by a professional medical translator. The direction of this sub-track is EN&gt;ES.</p> <p>The sample (dev) data includes 400 concepts. The terms are presented as tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms. The third column includes the term&rsquo;s origin ontology and its correspondent ID (separated by an underscore), while the fourth one includes a link to the concept in OBO Library.</p> <p>The test data (1,789 concepts) is made up of a TSV file with five columns: term number, English term, Spanish term, ontology id and OBO library URL.</p> <p>The background data (299,408 concepts) is made up of a TSV file with four columns: term number, English term, ontology id and OBO library URL.</p> <p>The term number columns can be used to map the entries in the join test + background document provided in earlier versions.</p> <p>&nbsp;</p> <p><strong>Related Links:</strong></p> <ul> <li> <p>ClinSpEn website with more information: <a href="https://temu.bsc.es/clinspen/">https://temu.bsc.es/clinspen/</a></p> </li> <li> <p>WMT website: <a href="https://www.statmt.org/wmt22/">https://www.statmt.org/wmt22/</a></p> </li> </ul> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any question or suggestion, please contact us at the following addresses:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br> - Darryl Estrada (&lt;darrylestrada97 [at] gmail [dot] com&gt;)<br> - Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record