Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “Lexical Semantics”
Benchmark for the Evaluation of Lexical Semantic Change Detection for Ancient Greek
<p>This repository contains a benchmark of Ancient Greek lemmas which underwent semantic change. It is meant as a support for the evaluation of methods detecting lexical semantic change in Ancient Greek. It was created at the University of Groningen, The Netherlands. A publication will follow soon.</p> <p> </p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created by retrieving and selecting from existing scholarship cases of lexemes which underwent semantic change. The evaluation items are 44 Ancient Greek lemmas, accompanied by the following information (see the column headers in the CSV file):</p> <ul> <li><strong>reference:</strong> the literature source of information about the change;</li> <li><strong>which_change: </strong>an explanation of the change in meaning. NB: the older meaning(s) do not necessarily disappear after the change, but it can happen that the new meaning(s) are added to the existing one(s), increasing the polysemy of the lemma;</li> <li><strong>when_changed: </strong>information about the work(s) or time period in which the change was first recorded; this kind of information was not always available or precise;</li> <li><strong>christian_change:</strong> whether the change is triggered by social, religious, or cultural changes related to the spread of Christianity, according to the scholarship.</li> </ul> <p> </p> <p><strong>2. References</strong></p> <p>The literature used to build this benchmark is the following:</p> <p> BUCK, Carl Darling. A dictionary of selected synonyms in the principal Indo-European languages. University of Chicago Press, 1949.</p> <p> FINKELBERG, Aryeh. "On the History of the Greek ΚΟΣΜΟΣ." Harvard Studies in Classical Philology (1998): 103-136.</p> <p> GINGRICH, F. Wilbur. "The Greek New Testament as a landmark in the course of semantic change." <em>Journal of Biblical Literature</em> (1954): 189-196.</p> <p> HORKY, Phillip Sidney. "When did Kosmos become the Kosmos." <em>Cosmos in the Ancient World</em> (2019): 22-41.</p> <p> LURAGHI, Silvia. "The verb aréskein in Ancient Greek: Constructions and semantic change." <em>Acta Linguistica Petropolitana. Труды института лингвистических исследований</em> 18-1 (2022): 226-245.</p> <p> </p> <p>These dictionaries of Ancient Greek were also used to double-check the instances of change:</p> <p> LIDDELL, Henry George, and Robert Scott. <em>A Greek-English Lexicon</em>. revised and augmented throughout by. Sir Henry Stuart Jones. with the assistance of. Roderick McKenzie. Oxford. Clarendon Press. 1940.</p> <p> ROCCI, Lorenzo.<em> Vocabolario greco-italiano</em>. Roma. Società editrice Dante Alighieri. 1939.</p> <p> SLUITER, Ineke, and Lucien van Beek, and Ton Kessels, and Albert Rijksbaron. <em>Woordenboek Grieks/Nederlands</em>. 2024. <a href="https://woordenboekgrieks.nl/" target="_blank" rel="noopener">https://woordenboekgrieks.nl/</a></p> <p> </p> <p><strong>3. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br> <br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website <a href="https://www.anchoringinnovation.nl/">www.anchoringinnovation.nl</a>.</div> <div> <p> </p> <p><strong>4. How to cite</strong></p> </div> <div>Until there is no publication about this benchmark, please cite the resource as:</div> <div>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim (2024), <em>Benchmark for the Evaluation of Lexical Semantic Change Detection Measures in Ancient Greek</em>, DOI: 10.5281/zenodo.13364555.</div> <div> </div> <div> </div>
Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Språkbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. "Korp-the corpus infrastructure of Språkbanken." <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dannélls, and Nina Tahmasebi. "Exploring the Quality of the Digital Historical Newspaper Archive KubHist." <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p> </p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>
JeSemE models for lexical semantic change
<p>Models for diachronic lexical semantics used by the <a href="http://jeseme.org">Jena Semantic Explorer (JeSemE)</a> web site described in our <a href="http://aclweb.org/anthology/C18-2003">COLING 2018 paper "JeSemE: A Website for Exploring Diachronic Changes in Word Meaning and Emotion"</a>.</p> <p>Also described and applied in Johannes Hellrich's Ph.D. thesis "Word Embeddings: Reliability & Semantic Change" who was funded by the Deutsche Forschungsgemeinschaft (DFG) within the graduate school "The Romantic Model" (GRK 2041/1).</p> <p>One ZIP file per corpus, each containing several CSV files:</p> <ul> <li>CHI.csv with χ<sup>2 </sup>word association values (structure: word-id, word-id, time, value)</li> <li>EMBEDDING.csv with SVD-PPMI word embeddings (aligned; structure: word-id, time, values)</li> <li>EMOTION.csv with VAD word emotion values (structure: word-id, time, values)</li> <li>FREQUENCY.csv with relative word frequency values (structure: word-id, time, value)</li> <li>PPMI.csv with PPMI<sup> </sup>word association values (structure: word-id, word-id, time, value)</li> <li>SIMILARITY.csv with word embedding derived word similarity values (structure: word-id, word-id, time, value)</li> <li>WORDIDS.csv mapping words to their corpus specific IDs</li> </ul> <p>Corpora are:</p> <ul> <li> <p>coha: Corpus of Historical American English</p> </li> <li> <p>dta: Deutsches Textarchiv 'German Text Archive'</p> </li> <li> <p>google_fiction: Google Books N-Gram corpus, English fiction subcorpus</p> </li> <li> <p>google_german: Google Books N-Gram corpus, German subcorpus</p> </li> <li> <p>rsc: Royal Society Corpus </p> </li> </ul>
Lexical Semantic Change Cause-Type-Definitions Benchmark
<p>The Lexical Semantic Change Cause-Type-Definitions (LSC-CTD) Benchmark is a digitised dataset that builds on and extends the Blank's seminal 1997 taxonomy of semantic change. This collection categorises 657 instances of linguistic evolution across the vocabulary of the Romance languages, with additional entries of German and English instances. Each entry is accompanied by a new pair (Old and New Meaning) of english definitions, manually curated by a historical linguist.</p> <p>The dataset includes a detailed classification of causes of change such as semantic wear, lexical gap, orphaned word, lexical complexity, atypical actant structure, frame, socio-cultural change, abstract concept, atypical part of speech, new concept, taboo, expressivity and prototype. It also includes types of semantic shift as classified by Blank, i.e. specialisation, generalisation, co-hyponymous transfer, auto-antonym, metaphor, antiphrasis, metonymy, auto-converse, ellipsis, folk etymology, analogy, meaning dilution, meaning reinforcement and doubtful cases. </p> <p><br><strong>Reference</strong></p> <p>The accompanying paper where this resource is described in detail will be published at ACL 2024.<br><br><span>Pierluigi Cassotti, Stefano De Pascale, and Nina Tahmasebi. 2024. <a href="https://aclanthology.org/2024.acl-long.249">Using Synchronic Definitions and Semantic Relations to Classify Semantic Change Types</a>. In <em>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, pages 4539–4553, Bangkok, Thailand. Association for Computational Linguistics.</span></p>
The evolution of lexical semantics dynamics, directionality, and drift: S4
<p>Supplementary Material (S4) for the study "The evolution of lexical semantics dynamics, directionality, and drift", for Frontiers in Communication, Special Issue "<a href="https://www.frontiersin.org/research-topics/38650/the-evolution-of-meaning-challenges-in-quantitative-lexical-typology?fbclid=IwAR3AeXx_11P-8CZG0UavlOgqvbVo6MlMxX8AejtPiREAKmeLyy3LIB3Ux24">The Evolution of Meaning: Challenges in Quantitative Lexical Typology</a>", ed. Gerd Carling & Annemarie Verkerk</p>
Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes
<p>The dataset was made for the purposes of the author's master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">"Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo’s Core Vocabulary"</a>. The dataset includes two worksheets. The first is titled "Main table", and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled "Extended table", includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively). <br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">at this link</a>. </p>
SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p><strong>Authors</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi</p> <p><strong>Description</strong></p> <p>This data collection contains the <strong>post-evaluation</strong> data for <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:</p> <ul> <li>the starting kit to download data, and examples for competing in the CodaLab challenge including baselines</li> <li>the true binary change scores of the targets for Subtask 1, and their true graded change scores for Subtask 2 (<code>test_data_truth/</code>),</li> <li>the scoring program used to score submissions against the true test data in the evaluation and post-evaluation phase (<code>scoring_program/</code>),</li> <li>the results of the evaluation phase including <ul> <li>the final rankings of the participating teams by their best submission (<code>results/rankings_teams.csv</code>),</li> <li>the submitted files of each team (<code>results/submissions/</code>),</li> <li>an overview of the results for each submission ordered by team (<code>results/submissions_results.csv</code>),</li> <li>analysis plots (<code>plots/</code>) displaying the results: <ul> <li>under <code>per_target/</code> we provide the gold change scores and the normalized prediction error of target words plotted against their frequency and polysemy statistics,</li> <li>under <code>per_team/</code> we provide the model predictions from the best submission per team (per subtask) plotted against frequency/polysemy statistics and performance on gold data (gray lines give the correlation with the respective variable in the gold data); we also provide plots of visualizing the teams' prediction similarities.</li> </ul> </li> </ul> </li> </ul> <p>Some remarks:</p> <ul> <li>the paper referenced below remains the only source for the rankings between teams,</li> <li>some teams were disqualified, and are thus removed from the analyses and the rankings present in the paper,</li> <li>some teams have changed names, resulting in a discrepancy between team names under <code>results/</code> and team names in the paper. The paper contains a key to match old names with new names.</li> </ul> <p><strong>Test Data </strong>for SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection can be found using the links below:</p> <ul> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-eng/">English</a></li> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-ger/">German</a></li> <li><a href="https://zenodo.org/record/3734089">Latin</a></li> <li><a href="https://zenodo.org/record/3730550">Swedish</a></li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi. 2020. <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. SemEval@COLING2020.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <pre><code>@inproceedings{schlechtweg2020semeval, title = "{S}em{E}val-2020 {T}ask 1: {U}nsupervised {L}exical {S}emantic {C}hange {D}etection", author = "Schlechtweg, Dominik and McGillivray, Barbara and Hengchen, Simon and Dubossarsky, Haim and Tahmasebi, Nina", booktitle = "To appear in Proceedings of the 14th International Workshop on Semantic Evaluation", year = "2020", address = "Barcelona, Spain", publisher = "Association for Computational Linguistics"}</code></pre> <p> </p>
GlossReader at LSCDiscovery: Train to Select a Proper Gloss in English -- Discover Lexical Semantic Change in Spanish
<pre>Precomputed vectors for the GlossReader system. LSCDiscovery Competition: https://codalab.lisn.upsaclay.fr/competitions/2243. </pre>
Sampled sentence pairs from SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p>Each dataset consists of samples, containing two sentences with positions of one of the given target words. In every sample, first sentence is taken from corpus1 and second from corpus2. Initial sentences were taken from https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd/. </p>
LEXICAL SEMANTIC FEATURES OF SPECIFIC GENDER AFFILIATION IN THE LANGUAGE COMMUNITY
Open the record for dataset details and reuse information.
Rehabilitation of post-stroke aphasia by a single protocol targeting phonological, lexical, and semantic deficits with speech output tasks
<p>This study assessed the effectiveness of a novel rehabilitation protocol (PHOLEXSEM), focused on PHonological, SEmantic, and LExical deficits, aiming at improving lexical retrieval, and, generally, spoken output. The study is published here: https://doi.org/10.23736/S1973-9087.24.08576-9</p> <p><span>These data cannot be made publicly available because they include sensitive information, but can be made available to interested researchers upon reasonable request. Please, forward your request to Prof. Nadia Bolognini (n.bolognini@auxologico.it).</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.