Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “translation alignment”

Learn how ShareScore rates datasets ↗
zenodo44/100

Aligned translation of Artemidorus Onir. book V

<p>Aligned translation of Artemidorus' Oneirocritica Book 5 in 95 chapters and a prologue divided into four sections. Part of the Open Projects in Digital Classics at the College of Letters and Sciences of the State University of S&atilde;o Paulo in Araraquara, S&atilde;o Paulo, Brazil. That is a second version of the translation. It was aligned on the Ugarit Platform and is visible at <a href="http://ugarit.ialigner.com/userProfile.php?userid=15&amp;tgid=9056">https://ugarit.ialigner.com/userProfile.php?userid=15&amp;tgid=14643</a>. The Greek text source was the digitized Pack's 1963 edition from CTS Perseids: urn:cts:greekLit:tlg0553.tlg001.1st1K-grc1:5. The Portuguese text is the revised translation (urn:cts.greekLit:tlg:0553.tlg001.ferreira2:5) of a previous&nbsp;<a href="https://www.culturaacademica.com.br/catalogo/oneirokritika-de-artemidoro-de-daldis-seculo-ii-d-c/">ebook</a>&nbsp;published by Cultura Acad&ecirc;mica in 2014 (digitized and available at https://furman-editions-in-progress.github.io/UNESP_FU/ as urn:cts:greekLit:tlg0553.tlg001.ferreira1:5).</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Translation Alignment: Ancient Greek to English. Annotation Style Guide and Gold Standard.

<p>This dataset&nbsp;contains guidelines and a gold standard for the alignment of Ancient Greek texts with English translations.</p> <p>The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, and Platonic dialogue, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (<a href="https://scaife.perseus.org/">https://scaife.perseus.org/</a>).</p> <p>The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (<a href="http://ugarit.ialigner.com/">http://ugarit.ialigner.com/</a>).</p> <p>The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.</p> <p>The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.</p> <p>For further information on Ugarit and translation alignment of historical languages, see&nbsp;<a href="http://ugarit.ialigner.com/bib.php">http://ugarit.ialigner.com/bib.php</a>&nbsp;and follow us on Twitter (@ugarit_ty).</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Translation Alignment: Ancient Greek to Latin. Annotation Style Guide and Gold Standard

<p>This dataset contains guidelines and a gold standard for the alignment of Ancient Greek texts with Latin scholarly translations.&nbsp;</p> <p>The gold standard consists of 100 fragments randomly selected from the&nbsp;<em>Digital Fragmenta Historicorum Graecorum&nbsp;</em>(DFHG) (https://www.dfhg-project.org/), which were aligned manually by Chiara Palladino and David J. Wright using Ugarit (https://ugarit.ialigner.com/). The Annotation Style Guide was developed for this project. The resulting Inter-Annotator-Agreement (IAA) is 90.5%.&nbsp;&nbsp;</p> <p>The materials available in this repository can be used to perform and evaluate alignments of various texts in Ancient Greek, to create gold standards, and to train automated translation alignment models.&nbsp;</p> <p>The Guidelines can be further adapted to address similar language pairs including inflected languages, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the scenario of machine translation. Different research questions, such as translation history or pedagogy, may need further tweaking of these guidelines.&nbsp;</p> <p>For further information on Ugarit and translation alignment of historical languages, see http://ugarit.aligner.com/bib.php and follow us on Twitter (@ugarit_ty).&nbsp;<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Sentence-aligned student translations of Crito (Ancient Greek, English, German, Persian)

<p>This dataset is a corpus of five student&#39;s translation of Plato&#39;s Crito aligned at sentence-level with the original Ancient Greek text, one German translation, and two English translations. The Ancient Greek text is the Burnet edition, made available by Perseus Digital Library. For more information, see:&nbsp;<br> https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0169%3Atext%3DCrito%3Asection%3D43a<br> &nbsp;</p> <p>The details of the eight translations (five Persian, two English, and one German translations) are as follows:</p> <ul> <li>German Translation:&nbsp;The German translation of Schleiermmacher available on Project Gutenberg has been aligned at the sentence level.<br> For more information, see: Plato, F. Schleiermacher, Platons Werke, In der Realschulbuchhandlung, 1809<br> Link to the text on Project Gutenberg:<br> https://www.projekt-gutenberg.org/platon/platowr1/kriton.html<br> &nbsp;</li> <li>English Translations: Two different English translations of &quot;Crito&quot;, one by Benjamin Jowett and the other by Harold North Fowler are included in the dataset.<br> For more information on Jowett&#39;s translation, see:<br> Plato, H. N. Fowler, W. Lamb, Plato in Twelve Volumes, Vol. 1 translated by Harold North<br> Fowler; Introduction by W.R.M. Lamb, volume 1, Harvard University Press and Wiliam<br> Heinemann Ltd., Cambridge, MA and London, 1966.<br> Fowler&#39;s translation on Perseus Digital Library:<br> https://www.perseus.tufts.edu/hopper/text?doc=plat.+crito+43a<br> For more information on Jowett&#39;s translation, see:<br> Plato, B. Jowett, Crito, The Internet Classics Archive, Massachusetts Institute of Technology,<br> http://classics.mit.edu/Plato/crito.html.<br> &nbsp;</li> <li>Persian Translations: The dataset consists of five Persian translations by students who have already completed a 30-hour Homeric Greek course. Each translator has translated the text into Persian using treebanks, commentaries, lexicon entries, and English and German translations. The translators themselves aligned the Persian translations to the Greek text at word-level using Ugarit. The alignments are available in their Ugarit profile:<br> Shouresh Assimi: https://ugarit.ialigner.com/userProfile.php?userid=50956<br> Aylar Mahmoudzadeh Sarabi: https://ugarit.ialigner.com/userProfile.php?userid=63464&amp;tgid=9576<br> Nima Mohammadi: https://ugarit.ialigner.com/userProfile.php?userid=52434&amp;tgid=9362<br> Kimia Nikpour: https://ugarit.ialigner.com/userProfile.php?userid=52378<br> Farshid Rahimi: https://ugarit.ialigner.com/userProfile.php?userid=50932&amp;tgid=9727</li> </ul> <p>The group&#39;s initial goal was to produce one finalized translation of Crito to Persian, but due to the intriguing variations in the translations and the text&#39;s intricacy, it was decided to provide three finalized translations rather than one. The finalized translations will be available in Beyond Translation as part of the Perseus Digital Library under a Creative Commons license. For more information on our final versions of Crito, see:&nbsp;http://beyond-translation.perseus.org</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Book One of The Iliad, Persian and Kurdish Translation with Word-level Alignment to the Treebank, Didakta Annotations and Glossary

<p>This dataset includes Persian and Kurdish translation of book one of the Iliad, aligned at word-level with the treebanks. The Persian translation is by Farnoosh Shamsian and the Kurdish translation by Farshid Rahimi. The dataset also includes glossaries in both languages.&nbsp;</p> <p>The treebank data combines both the UD treebank and the Perseus treebank in one spreadsheet. The Perseus treebank is available here:&nbsp;<a href="http://perseusdl.github.io/treebank_data/">http://perseusdl.github.io/treebank_data/</a></p> <p>The UD version is converted by Francesco Mambrini. For more information, see:&nbsp;<a href="https://github.com/francescomambrini/katholou/tree/main/ud_treebanks/agdt/data">https://github.com/francescomambrini/katholou/tree/main/ud_treebanks/agdt/data</a></p> <p>Both translations are aligned word by word to the treebank according to guidelines. The guideline for the Persian alignment is available here:</p> <p>Farnoosh Shamsian. (2023). Alignment Guidelines for Classical Greek-Persian. Zenodo. <a href="https://doi.org/10.5281/zenodo.8039931">https://doi.org/10.5281/zenodo.8039931</a></p> <p>The spreadsheet also contains Didakta annotations. For moe information about Didakta, see:&nbsp;</p> <p>Farnoosh Shamsian. (2023). Didakta Grammar for Annotation (English). Zenodo. <a href="http://https://doi.org/10.5281/zenodo.8318137">https://doi.org/10.5281/zenodo.8318137</a></p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

WMT'16 Biomedical Translation Task - Scielo parallel datasets - GMA alignment files

<p>Aligment files using the GMA tool (https://nlp.cs.nyu.edu/GMA/) for the parallel data from Scielo for the Biomedical Translation Task in the First Conference on Machine Translation (WMT 16) (http://www.statmt.org/wmt16/biomedical-translation-task.html).</p> <p>The parallel data is available here: https://zenodo.org/record/5588265</p>

opencc-by-4.0Jan 2016View details →
geo20/100

Translational activators align mitochondrial mRNAs at the small ribosomal subunit for translation initiation

GEO Series GSE282943. Saccharomyces cerevisiae. 35 samples. Type: Other.

openGEO-OpenJul 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record