Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “Lexical Dataset”
CLDF dataset accompanying Daniel et al.'s "Lingua Franca as Lexical Donors: Evidence from Daghestan" from 2021
<p>Cite the source of the dataset as:</p> <blockquote> <p>Daniel et al. (2021) Lingua Franca as Lexical Donors: Evidence from Daghestan. Language.</p> </blockquote>
CLDF dataset accompanying Chechuro et al.'s "Small-scale multilingualism through the prism of lexical borrowing" from 2021
<p>Cite the source of the dataset as:</p> <blockquote> <p>Chechuro et al. (2021) Small-scale multilingualism through the prism of lexical borrowing. International Journal of Bilingualism.</p> </blockquote>
ANEE Lexical Networks v. 2.0 - the dataset
<p>The Dataset of the <a href="https://urn.fi/urn:nbn:fi:lb-2024121101" target="_blank" rel="noopener">ANEE Lexical Networks v.2.0</a> created in the Semantic Domains in Akkadian Texts project at the University of Helsinki. This version 1.1 also includes the networks as GEXF files.</p>
CLDF dataset derived from Carroll et al. "Yamfinder: The Southern New Guinea Lexical Database"
<p>Cite the source of the dataset as:</p> <blockquote> <p>Carroll, Matthew J., Barth, Wolfgang, Nicholas Evans, I Wayan Arka, Christian Döhler, Eri Kashima, Volker Gast, Tina Gregor, Kate L. Lindsey, Julia Miller, Emil Mittag, Bruno Olsson, Dineke Schokkin, Jeff Siegel, Charlotte van Tongeren, Kyla Quinn. [DATE ACCESSED]. Yamfinder: Southern New Guinea Lexical Database. Available online at: http://www.yamfinder.com</p> </blockquote>
CLDF dataset derived from Ugarte et al.'s "NorthPeruLex - A Lexical Dataset of Small Language Families and Isolates from Northern Peru (forthcoming).
<p>Cite the source of the dataset as:</p> <blockquote> <p>Ugarte, Carlos and Blum, Frederic and Ingunza, Adriano and Gonzales, Rosa and Peña, Jaime. Forthcoming. NorthPeruLex - A Lexical Dataset of Small Language Families and Isolates from Northern Peru.</p> </blockquote>
CLDF dataset derived from Deepadung et al.'s "Lexical Comparison of Palaung Dialects" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Deepadung, Sujaritlak; Buakaw, Supakit; and Rattanapitak, Ampica (2015): A lexical comparison of the Palaung dialects spoken in China, Myanmar, and Thailand. Mon-Khmer Studies 44. 19-38.</p> </blockquote>
CLDF dataset derived from Bodt's "Lexical Cognates in Western Kho-Bwa" from 2019
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bodt, Timotheus Adrianus and List, Johann-Mattis (2019): Testing the predictive strength of the comparative method: An ongoing experiment on unattested words in Western Kho-Bwa languages. Papers in Historical Phonology 4.1: 22-44.</p> </blockquote>
CLDF dataset derived from Sagart et al.'s "Sino-Tibetan Database of Lexical Cognates" from 2019
<p>Cite the source of the dataset as:</p> <blockquote> <p>Laurent Sagart, Jacques, Guillaume, Yunfan Lai, and Johann-Mattis List (2019): Sino-Tibetan Database of Lexical Cognates. Jena: Max Planck Institute for the Science of Human History.</p> </blockquote>
CLDF dataset derived from Galucio et al.'s "Lexical Distances within the Tupian Linguistic family" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Galucio, Ana Vilacy and Meira, Sérgio and Birchall, Joshua and Moore, Denny and Gabas Júnior, Nilson and Drude, Sebastian and Storto, Luciana and Picanço, Gessiane and Rodrigues, Carmen Reis. (2015). Genealogical relations and lexical distances within the Tupian linguistic family. Boletim do Museu Paraense Emílio Goeldi. Ciências Humanas, 10(2), 229-274. https://dx.doi.org/10.1590/1981-81222015000200004</p> </blockquote>
CLDF dataset derived from MaJeLeD (Macro-Je Lexical Database)
<p>Cite the source of the dataset as:</p> <blockquote> <p>Gerardi, F. F., Tresoldi, T. (2024)</p> </blockquote>
Common20LS: A Lexical Simplification Dataset with Demographic Information
<p>Common20LS is a dataset for the task of Lexical Simplification that contains demographic information about the annotators. It consists on 20 Lexical Simplification problems annotated by 262 people. Each annotated instance is composed of a sentence, a target complex word or phrase, and a set of simplifications suggested by humans ranked by simplicity.</p>
Dataset of Middle Dutch lexical stress patterns and syllabifications
<p>This dataset consists of <strong>48.219 Middle Dutch words</strong> taken from in total 205 rhymed texts of the <em>Cd-rom Middelnederlands </em>(1998). All of these words have been <strong>assigned a syllabification and lexical stress pattern</strong>.</p> <p>E.g.: <em>proevede</em> is syllabified as <em>proe-ve-de</em> and has a stress index set at -3, which means that – counting from the rightmost syllable – the third syllable receives stress.</p> <p>This upload contains the following files:</p> <ul> <li>The <strong>JSON-file</strong> (compressed), which was used as input data for a machine learning algorithm trained for the automatic syllabification and stress assignment of Middle Dutch polysyllabic words (for the code of this experiment, see <a href="https://github.com/WHaverals/stresser">GitHub</a>)</li> <li>An <strong>Excel-file</strong>, containing the same data as the JSON (for more convenient reference)</li> <li>A <strong>split file </strong>(compressed), used in the training proces of the above-mentioned experiment</li> <li>A pdf-file with some <strong>insightful illustrations</strong> about the contents of the dataset</li> </ul> <p>This dataset is part of the research of <a href="https://www.uantwerpen.be/en/staff/wouter-haverals/research/">Wouter Haverals</a> (FWO, University of Antwerp), carried out under the supervision of prof. Mike Kestemont and em. prof. Frank Willaert.</p>
Frequency dataset for "Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch."
<p>The dataset is structured according to:</p> <ul> <li>the lexical field (CLOTHING, TRAFFIC, IT, and EMOTION);</li> <li>the part of speech (noun or adjective);</li> <li>the corpus;</li> <li>the concept;</li> <li>the term.</li> </ul> <p>It first gives the absolute frequency as found in the corpus and also after it was disambiguated. The concept frequency and relative frequency is calculated based on the "disambiguated" absolute frequency.</p> <p>More information can be found in this dissertation:</p> <p>Daems, Jocelyne. 2022. <em>Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch. </em>KU Leuven.</p>
cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification
<p><strong>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. & LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</strong></p>
BenchLS: A Reliable Dataset for Lexical Simplification
<p>To create our dataset we combined two resources: the LexMTurk (Horn et al., 2014) and LSeval (De Belder and Moens, 2012) datasets. The instances in both datasets, 929 in total, contain a sentence, a target complex word, and several candidate substitutions ranked according to their simplicity. The candidates in both datasets were suggested and ranked by English speakers from the U.S. To increase its reliability, we applied the following corrections over each instance of our dataset:</p> <ol> <li> <p>Spelling Filtering: We discard any misspelled can- didates using Norvig’s algorithm. We trained our spelling model over the News Crawl corpus.</p> </li> <li> <p>Inflection Correction: We inflected all candidates to the tense of the target word using the Text Adorning module of LEXenstein (Paetzold and Specia, 2015; Burns, 2013).</p> </li> </ol> <p>The resulting dataset – BenchLS – contains 929 instances, with an average of 7.37 candidate substitutions per complex word. </p>
cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification
<p>Original source of the data:</p> <blockquote> <p>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. & LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</p> </blockquote>
Dataset of Latin lexical reciprocals
<p>Dataset for the paper:</p> <p>Inglese, Guglielmo. Forthcoming. <em><span>Nisi paria non pugnant</span></em><span>: </span><span>argument structure alternations with lexical reciprocal verbs in Latin. <em>Journal of Latin Linguistics. </em></span></p>
Weak lexical equivalents in Russian → Belarusian translation: approaches to identification (dataset)
<p>Data accompanying the talk:<br> А. А. Ваўчок, У. В. Парыцкі. <a href="https://www.academia.edu/109056239">Слабыя лексічныя адпаведнікі ў руска-беларускім перакладзе: падыходы да ідэнтыфікацыі</a> // XI Міжнародны Кангрэс даследчыкаў Беларусі, Гданьск, 23.09.2023 [Oksana Volchek, Vladislav Poritski. Weak lexical equivalents in Russian → Belarusian translation: approaches to identification // Presented at 11th International Congress of Belarusian Studies, Gdańsk, 23.09.2023]</p> <p>The file <code>data.csv</code> is a test battery of 100 contexts in Russian, expected to cause difficulties when attempting to translate them into Belarusian. This is because each context contains a pair of words distinct in Russian but lacking clearly distinct Belarusian lexical equivalents. The contexts were taken, often with some simplifications, from various sources available on the web, i.e. they are representative of real usage, not artificially constructed.</p> <p>The file has four columns:</p> <ul> <li><code>w1</code>, <code>w2</code> – two words belonging to the same part of speech (typically adjective, noun, or verb), <code>w1</code> lexicographically preceding <code>w2</code>;</li> <li><code>context</code> – a snippet of text with both words (typically a single sentence, sometimes longer);</li> <li><code>inspired_by</code> – URL of a web page or file containing the original version of the context, which may have been slightly abridged or reworded for illustrative purposes. Archive copies of most URLs, with the exception of Google Books links, are available in the Wayback Machine (https://web.archive.org) or in https://archive.today.</li> </ul>
Datasets from `Discovering and analysing lexical variation in social media text'
<p>This repository contains the datasets that were used in the following three papers, which are also included within P. Shoemark's PhD dissertation `Discovering and analysing lexical variation in social media text':</p> <ul> <li>P. Shoemark, D. Sur, L. Shrimpton, I. Murray, and S. Goldwater. <a href="https://www.aclweb.org/anthology/E17-1116/"><em>Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media.</em></a> 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 2017. </li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W17-4908/"><em>Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data.</em> </a>Workshop on Stylistic Variation at EMNLP. 2017.</li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W18-6101/"><em>Inducing a lexicon of sociolinguistic variables from code-mixed text.</em> </a>Workshop on Noisy User Generated Text at EMNLP. 2018. </li> </ul> <p>Datasets consist of tab-separated-values files, in which rows correspond to tweets, with columns for user ID, tweet ID, and timestamp. </p> <p>The text of the tweets (and associated metadata) can be re-downloaded (in batches of 100 per request) using Twitter's <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint <em>(NB: Tweets which have been deleted or made private since the original datasets were collected can <strong>not </strong>be re-downloaded, so it may not be possible to reconstruct the original datasets in their entirety). </em></p> <p> </p> <p>Most of these datasets were originally drawn from <a href="https://developer.twitter.com/en/docs/tweets/sample-realtime/api-reference/get-statuses-sample">the Sample endpoint of Twitter’s Streaming API (</a>a.k.a. the ‘Spritzer’), which provides a random 1% sample of all public tweets in near real-time:</p> <p> </p> <p><strong>GU Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-UK.zip?versionId=a9cc222e-1d07-4026-9b55-cf2189b58191">Geotagged-UK.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which are geotagged to locations within the UK.</em></p> <p>The file <strong>GU_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within the UK.<strong> </strong><em><strong>• Tweets: </strong>1,768,334<strong> • Unique Users: </strong>455,075 <strong>• </strong></em></p> <p>The file <strong>GU.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GU dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>1,654,204<strong> • Unique Users: </strong>446,510 <strong>• </strong><sub>(the number of users in the GU dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>GS Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-Scotland.zip?versionId=146a0589-ddc0-4c42-86f6-7da79686ae56">Geotagged-Scotland.zip </a></strong></p> <p><em>The subset of Tweets in the GU dataset which are geo-tagged to locations within Scotland, specifically.</em></p> <p>The file <strong>GS_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within Scotland. <em><strong>• Tweets: </strong>178,401<strong> • Unique Users: </strong>41,685 <strong>• </strong></em></p> <p>The file <strong>GS.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GS dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>166,992<strong> • Unique Users: </strong>40,837 <strong>• </strong><sub>(the number of users in the GS dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>IT Dataset & Controls: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Indyref-Tweets.zip?versionId=41042a3e-f247-4493-9abe-29fb9ae81ee1">Indyref-Tweets.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which contain hashtags relating to the 2014 Scottish Independence Referendum (plus 'control' tweets which are by the same users but do not contain referendum-related hashtags)</em></p> <p>The file <strong>IT_</strong><strong>pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and contain at least one of 47 hashtags we identified as relating to the 2014 Scottish Independence Referendum (see paper for hashtag list). <em><strong>• Tweets: </strong>77,708<strong> • Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IT dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers, and tweets which do not contain hashtags that we judged to <em>unambiguously</em> relate to the referendum. <em><strong>• Tweets: </strong>59,664 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p>The file <strong>IT_controls_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and do <em><strong>not</strong></em> contain any of the hashtags we identified as relating to the 2014 Scottish Independence Referendum, but are authored by a user who has <em><strong>also</strong> </em>authored a tweet in <strong>IT_</strong><strong>pre-filtering.tsv</strong>. <em><strong>• Tweets: </strong>1,354,701 <strong>• Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT_controls</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> set of Control tweets that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, i.e. tweets which do not contain referendum-related hashtags but are authored by users who also have also authored tweets in <strong>IT</strong><strong>.tsv</strong>. <em><strong>• Tweets: </strong>881,679 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p> </p> <p><strong>SG-Users’ and IH-Users' Autumn 2014 Timeline Datasets: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Autumn-2014_Timelines.zip?versionId=c456b530-361e-4b88-9006-5bebd4a43c92">Autumn-2014_Timelines.zip</a></strong></p> <p><em>Complete tweet histories from Aug-Oct 2014 for users from the GS and IT datasets.</em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the GS dataset, i.e. users we know to have used Scottish geotags. This dataset is not restricted to tweets which appear in the ‘Spritzer’ sample; instead it consists of complete User Timelines for the months concerned, retrieved using the <a href="https://developer.twitter.com/en/docs/tweets/timelines/api-reference/get-statuses-user_timeline">statuses/user timeline</a> endpoint of Twitter’s REST API in March 2017. Because there are limits on the number of tweets that can be retrieved using this endpoint, we were not able to retrieve complete Autumn 2014 tweet histories for <em>all</em> of the users in the GS dataset. <em><strong>• Tweets: </strong>3,014,029 </em> <em><strong>• Unique Users: </strong>18,274 <strong>• </strong></em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> SG-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. This dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. <em><strong>• Tweets: </strong>1,112,931</em> <em><strong>• Unique Users: </strong>10,103 <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the IT dataset, i.e. users we know to have used Indyref-related hashtags. This dataset was collected in the same manner as SG-Users_Autumn_2014_timelines_pre-filtering.tsv; however, due to an error in this process, <strong>the IDs of most of the tweets in this dataset were not recorded</strong>. For such tweets the tweet ID column instead contains a placeholder tweet ID of the form _<user_ID>_<month>_<integer>, where the integer denotes the tweet's position in the reverse-chronological list of tweets that were retrieved for that user from that month (e.g. _147527441_09_286 is the placeholder tweet ID we assigned to the 286th September tweet we retrieved from the user whose Account ID is 147527441). Unfortunately, therefore, the tweets in this file whose 'IDs' begin with an underscore cannot be straightforwardly re-downloaded using Twitter's free <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint; but since their user IDs and timestamps are intact, it would still be possible to retrieve them using the (paid-for) <a href="http:// https://developer.twitter.com/en/docs/tutorials/choosing-historical-api">Historical APIs</a>. <em><strong>• Tweets: </strong>6,997,858 <strong>• Tweets whose IDs were recorded: </strong>288,394</em><em><strong> </strong> <strong>• Unique Users: </strong>14,645</em><em> <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IH-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. Like the SG-Users dataset, this dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. As with the pre-filtered version, most of the tweet IDs are unfortunately missing in this dataset. <em><strong>• Tweets: </strong>2,165,320 <strong>• Tweets whose IDs were recorded: </strong>115,366</em><em><strong> </strong> <strong>• Unique Users: </strong>10,784 <strong>• </strong></em></p> <p> </p> <p> </p> <p><strong>US Geotags: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-USA.zip">Geotagged-USA.zip</a></strong></p> <p><em>Tweets from June 2013 - July 2016 which are geotagged to locations within the USA.</em></p> <p>The file <strong>GUSA.tsv</strong> contains all tweets from the ‘Spritzer’ sample which were posted between June 30th 2013 to July 1st 2016, are classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets, do not contain urls or embedded media, are not by users with more than 1000 friends or followers, and are geotagged to locations within the USA. This dataset (along with the GU Dataset) was used in our <a href="http://www.aclweb.org/anthology/W18-6101/">WNUT 2018</a> paper. <em><strong>• Tweets: </strong></em> <em>8,375,573 </em> <em><strong>• Unique Users: </strong>1</em>,<em>826,260</em><em> <strong>• </strong></em></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.