Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

13

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

13 results for “Lexical Data”

Learn how ShareScore rates datasets ↗
zenodo44/100

Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Spr&aring;kbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. &quot;Korp-the corpus infrastructure of Spr&aring;kbanken.&quot; <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dann&eacute;lls, and Nina Tahmasebi. &quot;Exploring the Quality of the Digital Historical Newspaper Archive KubHist.&quot; <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on:&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>&nbsp;</p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>

opencc-by-2.0Feb 2020View details →
zenodo44/100

SiLeNe data -- a Sinitic Lexical Network

<p>This dataset is the semi-raw data used to build the Sinitic Lexical Network, browsable at https://silene.magistry.fr/</p> <p>It is a compilation of various sources of lexical description of different languages for which sinograms are a traditional script. These sources were released as Open Data by&nbsp; different third parties and are combined here for convenience in cross-lingual linguistic studies.</p> <p>This work is a graph-based extention of a work started with Fabienne Marc on Mandarin Chinese in the early 2000&#39;s. It has since then shifted to a multilingual perspective on the Chinese script, with a larger team of young researchers.</p> <p>For the sake of simplicity, the format is a single CSV file where each row describes the reading of a sinogram in a specific word in a specific language. Its design was dictated by our need for the work on the book &laquo;&Agrave; L&#39;&Eacute;coute des Sinogrammes&raquo; (ALES), but it can be converted to build the SiLeNe graph.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Wordlist files of lexical data from Papua New Guinea and western Solomons Oceanic languages collated for Ross's 1986 PhD thesis and 1988 publication thereof

<p>It occurs to me that the files containing&nbsp;Western Oceanic lexical data&nbsp;that&nbsp;I collected in the late 70s/early 80s for my PhD (Ross 1988) might be useful to someone. They are also used in the volumes of <em>The lexicon of Proto&nbsp;Oceanic </em>(Ross, Pawley &amp; Osmond 1998, 2003, 2011, 2016, 2023). In any case, it is right that they be made publicly available, something that wasn&#39;t so easy back then. Most of the material is from wordlists that I collected during fieldwork in Papua New Guinea from around 1978 to 1982. The file cor06 is omitted because it contains SE Solomonic data (outside Western Oceanic) drawn from Tryon &amp; Hackman 1983.</p> <p>I keyed the data into text files in a format such that each line was the entry for a single word, and each field within an entry was marked by a backslash code (I adapted this format from SIL&#39;s conventions at the time), then arranged them in cognate sets, each set separated from the next by an empty line. This work was done between 1983 and 1985, when text files were the best way to store data. They were entered on a terminal connected to a mainframe computer at the ANU. I have converted the ASCII symbols used in the original files into UTF-8 here in the interests of readability. The conversion was largely automatic, and I have not done a full check of each file, so there may be glitches.</p> <p>Each file contains languages from a region, as listed below (and the regions sometimes cut across subgroups determined by the comparative method). Three-letter abbreviations are used for language names, and two key files are also provided, one (COR-abbrevs) ordered by regions (determined by the numerals that start each line), the other by alphabetical order of&nbsp;language name (COR-abbrevs-alph). Some three-letter codes are followed by a hyphen and an extra letter. These are dialects. For example, MUM stands for Mumeng&nbsp;and MUM-P for the Patep dialect of Mumeng.</p> <p>Data files are labelled with COR (for &#39;correspondence sets&#39;) plus a numeral. The numerals are: 1-3 New Ireland; 4 Willaumez Peninsula (New Britain) area; 5 NW Solomonic; 7+8 Papuan Tip; 9 Vitiaz Strait area and NG north coast; 10 Huon Gulf and Markham Valley; 11 South and west New Britain. 7+8 are partial only. When I keyed the files,&nbsp;I had to rely on a mainframe&#39;s nightly back-up onto tape spools. One night the system failed, and so did the restore, and I lost some data.</p> <p>The backslash codes in the data files are: \l language; \p protolanguage; \w word; \g gloss; \n note; \s source. The formatting of these files is a little odd, since they served as input to routines I wrote to pull out sound correspondences. Anything after &#39;%&#39; is the elicited form: what immediately precedes &#39;%&#39; has had something &#39;undone&#39;, e.g. metathesis.</p> <p>The orthography of the files is phonemic and largely obvious. The conventions are set out in the introductions to the volumes of&nbsp;<em>The lexicon of Proto&nbsp;Oceanic.</em></p> <p>Finally, the files also contain reconstructions at various interstages at the top of a cognate set. These were inserted for heuristic reasons during my research. Many of them did not survive into my PhD thesis, and they should preferably be ignored. The reader who is interested in current Oceanic reconstructions should turn to the volumes of <em>The lexicon of Proto&nbsp;Oceanic.</em></p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Ume Saami lexical data from Álgu database with modernized spelling

<p>The files in this archive contain several processed versions of Ume Saami dictionary data originating from (Schlachter 1958). The data has been retrieved from the &Aacute;lgu database &lt;<a href="http://kaino.kotus.fi/algu/">http://kaino.kotus.fi/algu/</a>&gt; and checked against the original dictionary for headword variants. The final versions have the headword mechanically converted to (approximate) the current Ume Saami orthography (as in Barruk 2018). In addition to an alphabetized list, a reverse-alphabetized (a tergo) file is provided.</p> <p>For convenience, .xlsx versions of the final files are&nbsp;provided in addition to the plain-text (.tsv) versions (in zip file).<br> The orthography conversion program can be found at &lt;<a href="https://doi.org/10.5281/zenodo.4162535">https://doi.org/10.5281/zenodo.4162535</a>&gt;.</p> <p><strong>NOTE: The mechanically converted lemmas do not always fully correspond to the modern orthography and should be checked before use in further applications!</strong></p> <p>Authors: Juha Kuokkala (data processing and programming), Wolfgang Schlachter (original dictionary data).</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Sidwell Austroasiatic lexical data set for phylogenetic analyses 2015 version

<p>This is an Excel spreadsheet presenting lexical data for 122 Austroasiatic (AA) doculects, based on a 200 item list of semantic values. Each lexical item is scored for cognate judgement within each numbered semantic value. The purpose is to permit and test phylogenetic and lexicostatistical analyses.&nbsp;It is anticipated that the data set contains unrecognised errors, and is generally deserving of improvement, and scholars are heartily invited to report any and all shortcomings, or to make such offers of improvement as they may feel appropriate; such with be received with humility and acted upon with enthusiasm.</p>

opencc-by-nc-sa-4.0Nov 2015View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p><strong>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Supporting data for: "Steady state visual evoked potentials in reading aloud: Effects of lexicality, frequency and orthographic familiarity"

<p>Supporting data&nbsp;for the article&nbsp;&quot;Steady state visual evoked potentials in reading aloud: Effects of lexicality, frequency and orthographic familiarity&quot;.</p> <p>The dataset consists of the 16 original .bdf files.</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p>Original source of the data:</p> <blockquote> <p>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</p> </blockquote>

opencc-by-4.0Nov 2019View details →
zenodo32/100

A machine readable collection of lexical data on the Burmish languages

<p>This dataset includes lexical lists, mostly semantically normalized to WordNet, that brings together with near comprehensiveness the published data on Burmish languages. It was compiled by the ERC project &#39;Asia: Beyond Boundaries&#39; in collaboration with the Center for Research in Computational Linguistics.</p> <p>The Bibtex references can be found at this Zenodo deposit--</p> <p>Hill, Nathan, List, Johann-Mattis, &amp; Gong, Xun. (2020). Asia.bib: A bibtex bibliography for Asian Historical Linguistics [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3759114</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

Indo-European lexical cognate data

<p>This lexical cognate data was exported from the <a href="http://ielex.mpi.nl/">Indo-European Lexical Database</a> (IELex) supporting the publication:</p> <blockquote> <p>Verkerk, Annemarie. (2018). Detecting non-tree-like signal using multiple tree topologies. Journal of Historical Linguistics.</p> </blockquote> <p>If you are looking to use the IELex data in your own research, we recommend you use <a href="https://zenodo.org/record/5556801">the latest version</a>, which is available along with documentation, phylogenetic tree samples and example BEAST control scripts.</p>

opencc-by-4.0Jun 2018View details →
dryad28/100

Data from: Impact of lexical and sentiment factors on the popularity of scientific papers

We investigate how textual properties of scientific papers relate to the number of citations they receive. Our main finding is that correlations are nonlinear and affect differently the most cited and typical papers. For instance, we find that, in most journals, short titles correlate positively with citations only for the most cited papers, whereas for typical papers, the correlation is usually negative. Our analysis of six different factors, calculated both at the title and abstract level of 4.3 million papers in over 1500 journals, reveals the number of authors, and the length and complexity of the abstract, as having the strongest (positive) influence on the number of citations.

opencc-zeroDec 2015View details →
zenodo28/100

French liaison is allomorphy, not allophony: Evidence from lexical statistics (Data and code)

<p>This repository contains the supplementary materials to my paper entitled "French liaison is allomorphy, not allophony: Evidence from lexical statistics". Please refer to the README file for information about how to use the files in this repository to reproduce the analyses in the paper. Note that you should adapt the file paths to your own computer to run the scripts (the folder organization that I used is described in the README file).</p> <p>The three datafiles that were used for the final analyses in the paper are:&nbsp;</p> <ul> <li>study1-data.csv</li> <li>study2-data.csv</li> <li>study3-data.csv</li> </ul> <p>The three R scripts that were used to produce the visualizations and to run statistical analyses are:&nbsp;</p> <ul> <li>study1-script.R</li> <li>study2-script.R</li> <li>study3-script.R</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
dryad28/100

Data from: Impact of lexical and sentiment factors on the popularity of scientific papers

Open the record for dataset details and reuse information.

publicMay 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record