Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “Ancient Greek”

Learn how ShareScore rates datasets ↗
zenodo48/100

Benchmark for the Evaluation of Lexical Semantic Change Detection for Ancient Greek

<p>This repository contains a benchmark of Ancient Greek lemmas which underwent semantic change. It is meant as a support for the evaluation of methods detecting lexical semantic change in Ancient Greek. It was created at the University of Groningen, The Netherlands.&nbsp;A publication will follow soon.</p> <p>&nbsp;</p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created by retrieving and selecting from existing scholarship cases of lexemes which underwent semantic change. The evaluation items are 44 Ancient Greek lemmas, accompanied by the following information (see the column headers in the CSV file):</p> <ul> <li><strong>reference:</strong> the literature source of information about the change;</li> <li><strong>which_change: </strong>an explanation of the change in meaning. NB: the older meaning(s) do not necessarily disappear after the change, but it can happen that the new meaning(s) are added to the existing one(s), increasing the polysemy of the lemma;</li> <li><strong>when_changed:&nbsp;</strong>information about the work(s) or time period in which the change was first recorded; this kind of information was not always available or precise;</li> <li><strong>christian_change:</strong> whether the change is triggered by social, religious, or cultural changes related to the spread of Christianity, according to the scholarship.</li> </ul> <p>&nbsp;</p> <p><strong>2. References</strong></p> <p>The literature used to build this benchmark is the following:</p> <p>&nbsp; &nbsp; BUCK, Carl Darling. A dictionary of selected synonyms in the principal Indo-European languages. University of Chicago Press, 1949.</p> <p>&nbsp; &nbsp; FINKELBERG, Aryeh. "On the History of the Greek &Kappa;&Omicron;&Sigma;&Mu;&Omicron;&Sigma;." Harvard Studies in Classical Philology (1998): 103-136.</p> <p>&nbsp; &nbsp; GINGRICH, F. Wilbur. "The Greek New Testament as a landmark in the course of semantic change." <em>Journal of Biblical Literature</em> (1954): 189-196.</p> <p>&nbsp; &nbsp; HORKY, Phillip Sidney. "When did Kosmos become the Kosmos." <em>Cosmos in the Ancient World</em> (2019): 22-41.</p> <p>&nbsp; &nbsp; LURAGHI, Silvia. "The verb ar&eacute;skein in Ancient Greek: Constructions and semantic change." <em>Acta Linguistica Petropolitana. Труды института лингвистических исследований</em> 18-1 (2022): 226-245.</p> <p>&nbsp;</p> <p>These dictionaries of Ancient Greek were also used to double-check the instances of change:</p> <p>&nbsp; &nbsp; LIDDELL, Henry George, and Robert Scott. <em>A Greek-English Lexicon</em>. revised and augmented throughout by. Sir Henry Stuart Jones. with the assistance of. Roderick McKenzie. Oxford. Clarendon Press. 1940.</p> <p>&nbsp; &nbsp; ROCCI, Lorenzo.<em> Vocabolario greco-italiano</em>. Roma. Societ&agrave; editrice Dante Alighieri. 1939.</p> <p>&nbsp; &nbsp; SLUITER, Ineke, and Lucien van Beek, and Ton Kessels, and Albert Rijksbaron. <em>Woordenboek Grieks/Nederlands</em>. 2024. <a href="https://woordenboekgrieks.nl/" target="_blank" rel="noopener">https://woordenboekgrieks.nl/</a></p> <p>&nbsp;</p> <p><strong>3. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br>&nbsp;<br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website&nbsp;<a href="https://www.anchoringinnovation.nl/">www.anchoringinnovation.nl</a>.</div> <div> <p>&nbsp;</p> <p><strong>4. How to cite</strong></p> </div> <div>Until there is no publication about this benchmark, please cite the resource as:</div> <div>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim (2024), <em>Benchmark for the Evaluation of Lexical Semantic Change Detection Measures in Ancient Greek</em>, DOI: 10.5281/zenodo.13364555.</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Aug 2024View details →
zenodo44/100

List of Links to Digital Resources for Latin and Ancient Greek

<p>List of Links to Digital Resources for Latin and Ancient Greek</p> <p>The list was produced as an appendix to the German publication "Wie die Digitalisierung unseren Umgang mit den Alten Sprachen ver&auml;ndert hat" (How Digitization Changed the Way We Deal&nbsp;with Latin and Ancient Greek) in the&nbsp;journal "Forum Classicum", scheduled for release at the end of the year 2020.</p> <p>It contains references to various resources, such as text editions, databases, teaching materials, newspaper articles,&nbsp;tools for natural language processing and more. Most of them are available&nbsp;in English, some only in German. The list is sorted&nbsp;by the appearance of links in the article.</p> <p>Changelog:</p> <p>Version 2.0: Added headings from the paper to indicate topics for each part of the link list. English translations for the German headings are given in brackets.</p> <p>The list:</p> <p>Wie die Digitalisierung unseren Umgang mit den Alten Sprachen ver&auml;ndert hat / Linkliste (How Digitization Changed the Way We Deal with Latin and Ancient Greek / Link List)<br>A. Umgang mit der Literatur und anderen Wissensbest&auml;nden (Dealing with Literature and Other Data Collections)<br>1. Digitale Textsammlungen sind schnell verf&uuml;gbar und unterst&uuml;tzen Lehre und Forschung. (Digital text collections are quickly accessible and support teaching as well as research.)<br>https://www.degruyter.com/view/db/btltll&nbsp;<br>http://stephanus.tlg.uci.edu/&nbsp;<br>https://cil.bbaw.de/&nbsp;<br>https://latin.packhum.org/&nbsp;<br>http://cite-architecture.org/cts/&nbsp;<br>https://referenceworks.brillonline.com/entries/brill-s-new-pauly/ancient-authors-and-titles-of-works-Ancient_Authors_and_Titles_of_Works&nbsp;<br>http://www.perseus.tufts.edu/hopper/collection?collection=Perseus:collection:Greco-Roman&nbsp;<br>https://tesserae.caset.buffalo.edu/<br>2. Digitale Datenbanken erm&ouml;glichen schnelle systematische Suchanfragen in gro&szlig;en Text- oder Informationsbest&auml;nden, auch &uuml;ber disziplin&auml;re Grenzen hinweg. (Digital databases enable quick systematic queries for large collections of texts and other information, even beyond disciplinary boundaries.)<br>https://about.brepolis.net/lannee-philologique-aph/&nbsp;<br>https://www.gbd.digital/metaopac/start.do?View=gnomon&nbsp;<br>https://referenceworks.brillonline.com/browse/brill-s-new-pauly&nbsp;<br>https://www.navigium.de/&nbsp;<br>https://www.navigium.de/latein-unterrichten.html&nbsp;<br>http://lehrerportal.ccbuchner.de/Textanalyse/Default.aspx&nbsp;<br>https://open-educational-resources.de/&nbsp;<br>https://github.com/sommerschield/ancient-text-restoration&nbsp;<br>3. Digitale Datenbest&auml;nde werden vernetzt und f&uuml;r neue Anwendungszwecke kombiniert. (Digital data collections can be interconnected and combined for new use cases.)<br>https://www.w3.org/standards/semanticweb/data&nbsp;<br>https://lila-erc.eu/&nbsp;<br>https://peripleo.pelagios.org/&nbsp;<br>https://medium.com/pelagios/linked-open-data-to-navigate-the-past-using-peripleo-in-class-4286b3089bf3&nbsp;<br>https://topostext.org/&nbsp;<br>4. Die maschinelle sprachliche Vorverarbeitung antiker Texte erleichtert den Zugang f&uuml;r Lernende und Forschende. (Natural language processing of ancient texts facilitates access for both teachers and researchers.)<br>http://www.lemlat3.eu/&nbsp;<br>https://d.iogen.es/&nbsp;<br>https://alpheios.net/</p> <p>B. Umgang mit dem Spracherwerb (Dealing with Language Acquisition)<br>5. Die Digitalisierung f&ouml;rdert einen multimodalen und inklusiven &nbsp;Spracherwerb. (Digitization supports multimodal and inclusive language acquisition.)<br>https://www.hearinglink.org/living/loops-equipment/hearing-loops/what-is-a-hearing-loop/<br>http://www.cross-plus-a.com/balabolka.htm<br>https://propylaeum.de/e-learning/historische-aussprache-des-lateinischen-und-altgriechischen<br>https://www.youtube.com/watch?v=R5vdg_2i_pU<br>https://www.lesediagnostik.de/eye-tracking/<br>https://www.youtube.com/watch?v=8QocWsWd7fc<br>https://www.speechtexter.com/<br>https://etherpad.org/<br>https://moodle.org<br>6. Der Spracherwerb kann flexibel und personalisiert gestaltet werden. (Language acquisition can be designed in a flexible and personalized manner.)</p> <p>C. Umgang mit der &Ouml;ffentlichkeit (Dealing with the Public)<br>https://www.che.de/third-mission/<br>7. Social Media erm&ouml;glichen eine schnelle Interessens- und Wissensvernetzung innerhalb und vor allem au&szlig;erhalb einer definierten Gemeinschaft. (Social Media enable us to quickly connect interests and knowledge inside and especially outside of a specific community.)<br>https://la.wikipedia.org/wiki/Vicipaedia_Latina<br>http://forum.latein24.de/<br>https://twitter.com/RomAthen<br>https://www.projekte.hu-berlin.de/de/callidus/blog-2017-2018<br>https://www.superprof.de/blog/lateinische-begriffe-im-deutschen/<br>https://www.facebook.com/klassphil/?__tn__=%2Cd%2CP-R&amp;eid=ARDXqBAnvPxAePqFMxWrKxnFG2nfqqzKDWdoHdSg1CBNwBmcZbHwF5f8IWuQZXEODH6VKzqzWvUvUzfU<br>https://www.instagram.com/fs_klassphil_tuebingen/<br>https://hu-berlin.academia.edu/MarkusAsper<br>https://www.researchgate.net/profile/Monica_Berti<br>https://www.br.de/alphalernen/faecher/latein/latein-einfach-erklaert-100.html<br>https://www.pinterest.de/pin/5418462037462026/<br>https://www.youtube.com/channel/UChB8TYnAEtSIL1mY7FuBoqA<br>https://learnattack.de/latein/saetze-uebersetzen?utm_campaign=Learnattack_Kanal&amp;utm_source=youtube.com&amp;utm_medium=social&amp;utm_content=saetze-uebersetzen-latein&amp;kanal=youtube#video-wie-du-einen-lateinischen-satz-%C3%BCbersetzt<br>https://vimeo.com/276706092<br>8. Der digitale weltweite Zugang zu und Austausch von Wissen f&ouml;rdert das informelle Lernen und die Open-Science-Bewegung. (The worldwide digital access to and exchange of knowledge supports informal learning and the Open Science movement.)<br>https://www.udemy.com/course/an-introduction-to-classical-latin/<br>https://www.coursera.org/learn/roman-architecture<br>https://www.coursera.org/learn/plato<br>https://www.youtube.com/channel/UCNW1n7ctSkW3cgYFCzKPK3A/videos<br>https://scholar.google.de/<br>https://www.kim.uni-konstanz.de/openscience/onlinekurs-open-science-von-daten-zu-publikationen/<br>https://www.go-fair.org/fair-principles/<br>https://zenodo.org/record/3601182<br>https://zenodo.org/record/3816709<br>https://scm.cms.hu-berlin.de/callidus<br>https://www.ianus-fdz.de/<br>https://opr.degruyter.com/<br>http://ahropenreview.com/<br>https://arxiv.org/help/trackback<br>https://www.propylaeum.de/<br>https://journals.ub.uni-heidelberg.de/index.php/dco/index<br>http://www.pegasus-onlinezeitschrift.de/<br>https://www.schule-bw.de/faecher-und-schularten/sprachen-und-literatur/latein<br>https://www.schule-bw.de/faecher-und-schularten/sprachen-und-literatur/griechisch<br>https://www.bmbf.de/de/citizen-science-wissenschaft-erreicht-die-mitte-der-gesellschaft-225.html<br>https://pleiades.stoa.org/home</p> <p>Fazit (Conclusion)<br>http://pom.bbaw.de/cmg/</p>

opencc-zeroOct 2020View details →
zenodo44/100

Dataset for Case Attraction on Infinitive Clauses of Ancient Greek in Herodotus, Plato and Xenophon

<p>Annotated data used at the MA Dissertation&nbsp;<a href="http://doi.org/10.11606/D.8.2020.tde-12042021-174449">Geraldes, C.A.B. (2020) Case Attraction on Infinitive Clauses of Ancient Greek: a case study on Herodotus, Plato and Xenophon</a>, defended at the Faculdade de Letras e Ci&ecirc;ncias Humanas of Universidade de S&atilde;o Paulo (FFLCH-USP). For further information on the factors annotated at the dataset, please refer to the section on methodology at the dissertation.</p> <p>This studied was financed by The S&atilde;o Paulo Research Foundation, FAPESP, by means of the grant n&ordm; 2017/23334-2, S&atilde;o Paulo Research Foundation (FAPESP); and the international grant n&ordm; 2019/18473-9,&nbsp;S&atilde;o Paulo Research Foundation (FAPESP). This study was also partly financed by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento Pessoal de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Translation Alignment: Ancient Greek to English. Annotation Style Guide and Gold Standard.

<p>This dataset&nbsp;contains guidelines and a gold standard for the alignment of Ancient Greek texts with English translations.</p> <p>The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, and Platonic dialogue, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (<a href="https://scaife.perseus.org/">https://scaife.perseus.org/</a>).</p> <p>The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (<a href="http://ugarit.ialigner.com/">http://ugarit.ialigner.com/</a>).</p> <p>The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.</p> <p>The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.</p> <p>For further information on Ugarit and translation alignment of historical languages, see&nbsp;<a href="http://ugarit.ialigner.com/bib.php">http://ugarit.ialigner.com/bib.php</a>&nbsp;and follow us on Twitter (@ugarit_ty).</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Translation Alignment: Ancient Greek to Latin. Annotation Style Guide and Gold Standard

<p>This dataset contains guidelines and a gold standard for the alignment of Ancient Greek texts with Latin scholarly translations.&nbsp;</p> <p>The gold standard consists of 100 fragments randomly selected from the&nbsp;<em>Digital Fragmenta Historicorum Graecorum&nbsp;</em>(DFHG) (https://www.dfhg-project.org/), which were aligned manually by Chiara Palladino and David J. Wright using Ugarit (https://ugarit.ialigner.com/). The Annotation Style Guide was developed for this project. The resulting Inter-Annotator-Agreement (IAA) is 90.5%.&nbsp;&nbsp;</p> <p>The materials available in this repository can be used to perform and evaluate alignments of various texts in Ancient Greek, to create gold standards, and to train automated translation alignment models.&nbsp;</p> <p>The Guidelines can be further adapted to address similar language pairs including inflected languages, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the scenario of machine translation. Different research questions, such as translation history or pedagogy, may need further tweaking of these guidelines.&nbsp;</p> <p>For further information on Ugarit and translation alignment of historical languages, see http://ugarit.aligner.com/bib.php and follow us on Twitter (@ugarit_ty).&nbsp;<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Cyprus, Σαλαμίνα (Ancient Greek: Σαλαμίς) Salamis.

<p>Cyprus, &Sigma;&alpha;&lambda;&alpha;&mu;ί&nu;&alpha;&nbsp;(Ancient&nbsp;Greek: &Sigma;&alpha;&lambda;&alpha;&mu;ί&sigmaf;)&nbsp;Salamis. As documented in 1972.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Ancient Greek Literature for Advanced Data Processing: A Text Fabric Representation of Open Access Texts in TEI XML

<p>This data set contains a full conversion of Greek texts available in the Perseus Digital Library and the Open Greek and Latin Project to the Text Fabric data format. The main advantage of the Text Fabric datatype over the original TEI XML format is that it utilizes a strict separation of text and annotation in a flat data structure. At the same time, it permits multiple distinct formats of the same text as well as an unlimited depth of (embedded) annotations. Because of its flat data structure, it facilitates easy and clean procedures to analyze, transform, and enrich the available data. Many of these processes are very difficult to conduct while departing from the hierarchically organized XML tree representation.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Pedalion Ancient Greek Dependency Treebank - Euripides: Medea

<p>xml treebank</p> <p>Annotated by Toon Van Hal, with student contributions by Mathieu Cuijpers; Sanderijn Gijbels; Yoran Joosten; Yordi Lenaerts; Eva Uffing; Chiara Van der Hasselt; Lisa Vanhee and Jolien Volders (KU Leuven Bachelor 3, 2018-2019). Based on a preparsed text by Alek Keersmaekers. Controlled by Toon Van Hal, Sanderijn Gijbels and Yoran Joosten.</p>

opencc-by-nc-sa-4.0Dec 2018View details →
zenodo40/100

TranscriboQuest Ancient Greek Team

<p>Dataset produced by the Ancient Greek team during the TranscriboQuest event in Lyon, September 11-13th 2024. See README for further informations.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Ancient Greek Fasttext Word Embeddings

<p>Word embeddings generated with Fasttext and 1 GB of Ancient Greek texts. These embeddings were produced for the study of social networks and social semantics in ancient Greece by the Diogenet project at the University of San Diego, California.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

AGREE: a New Benchmark for the Evaluation of Semantic Models of Ancient Greek

<p>AGREE (Ancient Greek Relatedness Embeddings Evaluation) is a benchmark for the evaluation of semantic models of Ancient Greek created at the University of Groningen (The Netherlands). More information about it can be found in the following publication:</p> <p>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek,&nbsp;<em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392,&nbsp;<a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></p> <p>&nbsp;</p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created from a mix of expert judgements about relatedness between Ancient Greek words and model outputs validated by human experts. The evaluation items are pairs of Ancient Greek lemmas with a high semantic relatedness.</p> <p>The human judgements were collected via two questionnaires, proposing two different tasks to the experts. The evaluation items included in the AGREE benchmark are a selection of the most strictly related pairs of lemmas obtained from the two tasks. Here an overview of the contents of the repository:</p> <ul> <li><strong>1_agree_task1.json</strong>&nbsp;includes all the data collected with the first task. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'frequency': the number of times that the pair was suggested as related by an expert;</li> <li>'POS1': part-of-speech of the first lemma;</li> <li>'POS2': part-of-speech of the second lemma;</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>2_agree_task2.json&nbsp;</strong>includes all the data collected with the second task.&nbsp;The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin':&nbsp; <ul> <li>'common_pair' = one of the two pairs proposed to all participants in the second task;</li> <li>'task1' = pairs proposed by experts in the first task;</li> <li>'models_easy_rel' = output of word2vec models, pair considered as strictly related;</li> <li>'models_task1' = pairs proposed by experts in the first task and also output by word2vec models;</li> <li>'models' = output of word2vec language models;</li> <li>'unrelated' = made up pairs of unrelated lemmas (control pairs);</li> </ul> </li> <li>'respondents': number of experts evaluating a pair;</li> <li>'score': average relatedness score given by the experts on a 0-100 scale;</li> <li>'agreement': inter-annotated agreement between all experts who evaluated the&nbsp;block of pairs to which the current&nbsp;pair belongs&nbsp;(when available, i.e. when the block of pairs was presented to more than one participant);</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>3_agree_final_benchmark.json </strong>includes the&nbsp;final selection of items that constitutes AGREE. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin': <ul> <li>'task1': pair either proposed more than once in the first task or proposed only once, but scored&nbsp;&gt;=&nbsp;70 in the second task;</li> <li>'task2': pair scored by more than one respondent in the second task and with average score &gt;= 70.</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>This updated version of the repository includes the individual answers to the two questionnaires (see files 'answers_Task1_postprocessed.xlsx' and 'raw_answers_Task2.xlsx').</p> <p>&nbsp;</p> <p><strong>2. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br>&nbsp;<br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website <a href="https://www.anchoringinnovation.nl">www.anchoringinnovation.nl</a>.<br>&nbsp;<br>We want to thank the experts of Ancient Greek around the world who shared their knowledge of Ancient Greek semantics and donated some of their precious time. Without them the creation of this benchmark would not have been possible.<br>&nbsp;<br>We also want to thank the many colleagues from the University of Groningen, the National Research School OIKOS, and other Universities abroad who contributed to this work with discussion and advice.</div> <div>&nbsp;</div> <div>&nbsp;<br><strong>3. Citation</strong><br>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek, <em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392, <a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Sentence-aligned student translations of Crito (Ancient Greek, English, German, Persian)

<p>This dataset is a corpus of five student&#39;s translation of Plato&#39;s Crito aligned at sentence-level with the original Ancient Greek text, one German translation, and two English translations. The Ancient Greek text is the Burnet edition, made available by Perseus Digital Library. For more information, see:&nbsp;<br> https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0169%3Atext%3DCrito%3Asection%3D43a<br> &nbsp;</p> <p>The details of the eight translations (five Persian, two English, and one German translations) are as follows:</p> <ul> <li>German Translation:&nbsp;The German translation of Schleiermmacher available on Project Gutenberg has been aligned at the sentence level.<br> For more information, see: Plato, F. Schleiermacher, Platons Werke, In der Realschulbuchhandlung, 1809<br> Link to the text on Project Gutenberg:<br> https://www.projekt-gutenberg.org/platon/platowr1/kriton.html<br> &nbsp;</li> <li>English Translations: Two different English translations of &quot;Crito&quot;, one by Benjamin Jowett and the other by Harold North Fowler are included in the dataset.<br> For more information on Jowett&#39;s translation, see:<br> Plato, H. N. Fowler, W. Lamb, Plato in Twelve Volumes, Vol. 1 translated by Harold North<br> Fowler; Introduction by W.R.M. Lamb, volume 1, Harvard University Press and Wiliam<br> Heinemann Ltd., Cambridge, MA and London, 1966.<br> Fowler&#39;s translation on Perseus Digital Library:<br> https://www.perseus.tufts.edu/hopper/text?doc=plat.+crito+43a<br> For more information on Jowett&#39;s translation, see:<br> Plato, B. Jowett, Crito, The Internet Classics Archive, Massachusetts Institute of Technology,<br> http://classics.mit.edu/Plato/crito.html.<br> &nbsp;</li> <li>Persian Translations: The dataset consists of five Persian translations by students who have already completed a 30-hour Homeric Greek course. Each translator has translated the text into Persian using treebanks, commentaries, lexicon entries, and English and German translations. The translators themselves aligned the Persian translations to the Greek text at word-level using Ugarit. The alignments are available in their Ugarit profile:<br> Shouresh Assimi: https://ugarit.ialigner.com/userProfile.php?userid=50956<br> Aylar Mahmoudzadeh Sarabi: https://ugarit.ialigner.com/userProfile.php?userid=63464&amp;tgid=9576<br> Nima Mohammadi: https://ugarit.ialigner.com/userProfile.php?userid=52434&amp;tgid=9362<br> Kimia Nikpour: https://ugarit.ialigner.com/userProfile.php?userid=52378<br> Farshid Rahimi: https://ugarit.ialigner.com/userProfile.php?userid=50932&amp;tgid=9727</li> </ul> <p>The group&#39;s initial goal was to produce one finalized translation of Crito to Persian, but due to the intriguing variations in the translations and the text&#39;s intricacy, it was decided to provide three finalized translations rather than one. The finalized translations will be available in Beyond Translation as part of the Perseus Digital Library under a Creative Commons license. For more information on our final versions of Crito, see:&nbsp;http://beyond-translation.perseus.org</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Ancient Greek language models

<p>In this repository, we release a series of vector space models of Ancient Greek, trained following different architectures and with different hyperparameter values.&nbsp;</p> <p>Below is a breakdown of all the models released, with an indication of the training method and hyperparameters. The models are split into &lsquo;<strong>Diachronica&rsquo; </strong>and &lsquo;<strong>ALP&rsquo; </strong>models, according to the published paper they are associated with.</p> <blockquote> <p>[<strong>Diachronica</strong>:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray &amp; Malvina Nissim. Forthcoming. Natural Language Processing for Ancient Greek: Design, Advantages, and Challenges of Language Models, <em>Diachronica</em>.</p> <p>[<strong>ALP</strong>:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray &amp; Malvina Nissim. 2023. Evaluation of Distributional Semantic Models of Ancient Greek: Preliminary Results and a Road Map for Future Work. <em>Proceedings of the Ancient Language Processing Workshop associated with the 14th International Conference on Recent Advances in Natural Language Processing (RANLP 2023)</em>. 49-58. Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-087-8.2023_006</p> </blockquote> <h1><em>Diachronica</em> models</h1> <h2>Training data</h2> <p>Diorisis corpus (Vatri &amp; McGillivray 2018). Separate models were trained for:</p> <ol> <li>Classical subcorpus</li> <li>Hellenistic subcorpus</li> <li>Whole corpus</li> </ol> <p>Models are named according to the (sub)corpus they are trained on (i.e. <code>hel_</code> or <code>hellenestic</code> is appended to the name of the models trained on the Hellenestic subcorpus, <code>clas_</code> or <code>classical</code> for the Classical subcorpus, <code>full_</code> for the whole corpus).</p> <h2>Models</h2> <h3><strong><em>Count-based</em></strong></h3> <blockquote> <p>Software used: LSCDetection (Kaiser et al. 2021;&nbsp;<a href="https://github.com/Garrafao/LSCDetection">https://github.com/Garrafao/LSCDetection</a>)</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; With Positive Pointwise Mutual Information applied (folder PPMI spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (<code>_stopfilt</code> is appended to the model names). Hyperparameter values: <code>window=5</code>, <code>k=1</code>, <code>alpha=0.75</code>.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; With both Positive Pointwise Mutual Information <em>and</em> dimensionality reduction with Singular Value Decomposition applied (folder PPMI+SVD spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (<code>_stopfilt</code> is appended to the model names). Hyperparameter values: <code>window=5</code>, <code>dimensions=300</code>, <code>gamma=0.0</code>.</p> <h3><strong><em>Word2Vec</em></strong></h3> <blockquote> <p>Software used: CADE (Bianchi et al. 2020;&nbsp;<a href="https://github.com/vinid/cade">https://github.com/vinid/cade</a>).</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; Continuous-bag-of-words (CBOW). Hyperparameter values: <code>size=30</code>, <code>siter=5</code>, <code>diter=5</code>, <code>workers=4</code>, <code>sg=0</code>, <code>ns=20</code>.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; Skipgram with Negative Sampling (SGNS). Hyperparameter values: <code>size=30</code>, <code>siter=5</code>, <code>diter=5</code>, <code>workers=4</code>, <code>sg=1</code>, <code>ns=20</code>.</p> <h3><strong><em>Syntactic word embeddings</em></strong></h3> <p>Syntactic word embeddings were also trained on the Ancient Greek subcorpus of the PROIEL treebank (Haug &amp; J&oslash;hndal 2008), the Gorman treebank (Gorman 2020), the PapyGreek treebank (Vierros &amp; Henriksson 2021), the Pedalion treebank (Keersmaekers et al. 2019), and the Ancient Greek Dependency Treebank (Bamman &amp; Crane 2011) largely following the SuperGraph method described in Al-Ghezi &amp; Kurimo (2020) and the Node2Vec architecture (Grover &amp; Leskovec 2016) (see <a href="https://github.com/npedrazzini/ancientgreek-syntactic-embeddings#graph-based-syntactic-word-embeddings">https://github.com/npedrazzini/ancientgreek-syntactic-embeddings</a> for more details). Hyperparameter values: window=1, min_count=1.</p> <h1><em>ALP</em> models</h1> <h2>Training data</h2> <p>Archaic, Classical, and Hellenistic portions of the Diorisis corpus (Vatri &amp; McGillivray 2018) merged, stopwords removed according to the list&nbsp;made by Alessandro Vatri, available at https://figshare.com/articles/dataset/Ancient_Greek_stop_words/9724613.</p> <h2>Models</h2> <h3><strong><em>Count-based</em></strong></h3> <blockquote> <p>Software used: LSCDetection (Kaiser et al. 2021;&nbsp;<a href="https://github.com/Garrafao/LSCDetection">https://github.com/Garrafao/LSCDetection</a>)&nbsp;</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; With Positive Pointwise Mutual Information applied (folder ppmi_alp).&nbsp; Hyperparameter values: <code>window=5</code>, <code>k=1</code>, <code>alpha=0.75</code>. Stopwords were removed from the training set.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; With both Positive Pointwise Mutual Information <em>and</em> dimensionality reduction with Singular Value Decomposition applied (folder ppmi_svd_alp). Hyperparameter values: <code>window=5</code>, <code>dimensions=300</code>, <code>gamma=0.0</code>. Stopwords were removed from the training set.</p> <h3><strong><em>Word2Vec</em></strong></h3> <blockquote> <p>Software used: Gensim library (Řehůřek and Sojka, 2010)</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; Continuous-bag-of-words (CBOW). Hyperparameter values: <code>size=30</code>, <code>window=5</code>, <code>min_count=5</code>, <code>negative=20</code>, <code>sg=0</code>. Stopwords were removed from the training set.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; Skipgram with Negative Sampling (SGNS). Hyperparameter values: <code>size=30</code>, <code>window=5</code>, <code>min_count=5</code>, <code>negative=20</code>, <code>sg=1</code>. Stopwords were removed from the training set.</p> <h1><strong>References</strong></h1> <p>Al-Ghezi, Ragheb &amp; Mikko Kurimo. 2020. Graph-based syntactic word embeddings. In Ustalov, Dmitry, Swapna Somasundaran, Alexander Panchenko, Fragkiskos D. Malliaros, Ioana Hulpuș, Peter Jansen &amp; Abhik Jana (eds.), <em>Proceedings of the Graph-based Methods for Natural Language Processing (TextGraphs)</em>, 72-78.</p> <p>Bamman, D. &amp; Gregory Crane. 2011. The Ancient Greek and Latin dependency treebanks. In Sporleder, Caroline, Antal van den Bosch &amp; Kalliopi Zervanou (eds.), <em>Language Technology for Cultural Heritage. Selected Papers from the LaTeCH [Language Technology for Cultural Heritage] Workshop Series. Theory and Applications of Natural Language Processing</em>, 79-98. Berlin, Heidelberg: Springer.</p> <p>Gorman, Vanessa B. 2020. Dependency treebanks of Ancient Greek prose. <em>Journal of Open Humanities Data</em> 6(1).</p> <p>Grover, Aditya &amp; Jure Leskovec. 2016. Node2vec: scalable feature learning for networks. In <em>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD &lsquo;16)</em>, 855-864.</p> <p>Haug, Dag T. T. &amp; Marius L. J&oslash;hndal. 2008. Creating a parallel treebank of the Old Indo-European Bible translations. In <em>Proceedings of the Second Workshop on Language Technology for Cultural Heritage Data (LaTeCH)</em>, 27&ndash;34.</p> <p>Keersmaekers, Alek, Wouter Mercelis, Colin Swaelens &amp; Toon Van Hal. 2019. Creating, enriching and valorizing treebanks of Ancient Greek. In Candito, Marie, Kilian Evang, Stephan Oepen &amp; Djam&eacute; Seddah (eds.), <em>Proceedings of the 18th International Workshop on Treebanks and Linguistic Theories</em> <em>(TLT, SyntaxFest 2019)</em>, 109-117.</p> <p>Kaiser, Jens, Sinan Kurtyigit, Serge Kotchourko &amp; Dominik Schlechtweg. 2021. Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change Detection. In <em>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics</em>.</p> <p>Schlechtweg, Dominik, Anna H&auml;tty, Marco del Tredici &amp; Sabine Schulte im Walde. 2019. A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In <em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</em>, 732-746, Florence, Italy. ACL.</p> <p>Vatri, Alessandro &amp; Barbara McGillivray. 2018. The Diorisis Ancient Greek Corpus: Linguistics and Literature.&nbsp;<em>Research Data Journal for the Humanities and Social Sciences</em>&nbsp;3, 1, 55-65, Available From: Brill&nbsp;<a href="https://doi.org/10.1163/24523666-01000013" target="_blank" rel="noopener">https://doi.org/10.1163/24523666-01000013</a></p> <p>Vierros, Marja &amp; Erik Henriksson. 2021. PapyGreek treebanks: a dataset of linguistically annotated Greek documentary papyri. <em>Journal of Open Humanities Data</em> 7.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Dataset for the paper "Ideology and identity in grammar: A diachronic-quantitative approach to language standardisation processes in Ancient Greek"

<p>This is the dataset for the paper &quot;Ideology and identity in grammar: A diachronic-quantitative approach to language standardisation processes in Ancient Greek&quot;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Laginos (greek ancient vase)

Laginos is a Hellenistic type of an ancient wine jug, also known as ''oinochoe'' which is a key form of ancient Greek pottery. It has a flat base, a sharp or more rarely curved shoulder and a circular mouth. Source: Objaverse 1.0 / Sketchfab

opencc-byJan 2022View details →
zenodo36/100

Ancient Greek Amphora

Amphora with scratches, dirt, dust and damaged parts. Typical orange black styled drawing. Textures made in SubstancePainter. Source: Objaverse 1.0 / Sketchfab

opencc-byMay 2022View details →
zenodo36/100

Stylized ancient greek catapult

Stylized simple with ancient greek look model. Sampi as the symbol. Source: Objaverse 1.0 / Sketchfab

opencc-byAug 2019View details →
zenodo36/100

Ancient greek theater Acropolis

Odeon of Herodes Atticus, 161AD, Athens Greece. http://odysseus.culture.gr/h/2/eh251.jsp?obj_id=6622 Let's digitize the world and share it :) Source: Objaverse 1.0 / Sketchfab

opencc-byDec 2018View details →
zenodo36/100

Ancient Greek Aulos

Hello, all. I am a beginner creator with no real training. My name is Hannah and I am an aspiring musicologist with a passion for ancient music and instruments. With that interest, I also got an interest in creating some models for ancient/historical woodwind instruments, and I have started with the fascinating double reed, doubled piped instrument from ancient Greece, the aulos. I created this model based off of [this picture](http://https://www.researchgate.net/figure/The-Louvre-aulos-E10962-photograph-by-Stefan-Hagel-courtesy-Louvre-Museum_fig1_316277175) of the Louvre Aulos. This was modeled in Blender and Quixel Mixer was used to create the textures. It is considered a work in progress. For one, reeds have not been created for it yet. However, I've been sitting on it for a year now and figured I would share. Source: Objaverse 1.0 / Sketchfab

opencc-bySep 2021View details →
zenodo36/100

Ancient Greek Soccer Players

This is a scan of a sculpture from Friederichienabend Museum in Vienna, Austria. Dated 2nd century A.D., marble. A group of young men are presumably representing early stages of football game. Made with Momento Beta. Unfortunately I was unable to preserve textures. Source: Objaverse 1.0 / Sketchfab

opencc-byMay 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record